Paper deep dive
The Spectral Edge Thesis: A Mathematical Framework for Intra-Signal Phase Transitions in Neural Network Training
Yongzhong Xu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/1/2026, 1:28:29 AM
Summary
The paper introduces the 'Spectral Edge Thesis', a mathematical framework explaining neural network phase transitions (e.g., grokking, loss plateaus) through the spectral gap of the rolling-window Gram matrix of parameter updates. It establishes that in the extreme aspect ratio regime, the BBP detection threshold is vacuous, and the operative structure is an intra-signal gap separating dominant from subdominant modes. The framework provides a universal, architecture-agnostic set of flow equations and axioms, consistent with existing theories like the Edge of Stability and Tensor Programs.
Entities (5)
Relation Signals (3)
Spectral Edge Thesis → controls → Phase Transitions
confidence 95% · phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap
Adiabatic Parameter → controls → Circuit Stability
confidence 95% · The adiabatic parameter A = ||dG||_F / (eta g^2) controls circuit stability
Gram Matrix → exhibits → Spectral Gap
confidence 95% · the operative structure is the intra-signal gap separating dominant from subdominant modes
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We develop the spectral edge thesis: phase transitions in neural network training -- grokking, capability gains, loss plateaus -- are controlled by the spectral gap of the rolling-window Gram matrix of parameter updates. In the extreme aspect ratio regime (parameters $P \sim 10^8$, window $W \sim 10$), the classical BBP detection threshold is vacuous; the operative structure is the intra-signal gap separating dominant from subdominant modes at position $k^* = \mathrm{argmax}\, \sigma_j/\sigma_{j+1}$. From three axioms we derive: (i) gap dynamics governed by a Dyson-type ODE with curvature asymmetry, damping, and gradient driving; (ii) a spectral loss decomposition linking each mode's learning contribution to its Davis--Kahan stability coefficient; (iii) the Gap Maximality Principle, showing that $k^*$ is the unique dynamically privileged position -- its collapse is the only one that disrupts learning, and it sustains itself through an $\alpha$-feedback loop requiring no assumption on the optimizer. The adiabatic parameter $\mathcal{A} = \|\Delta G\|_F / (\eta\, g^2)$ controls circuit stability: $\mathcal{A} \ll 1$ (plateau), $\mathcal{A} \sim 1$ (phase transition), $\mathcal{A} \gg 1$ (forgetting). Tested across six model families (150K--124M parameters): gap dynamics precede every grokking event (24/24 with weight decay, 0/24 without), the gap position is optimizer-dependent (Muon: $k^*=1$, AdamW: $k^*=2$ on the same model), and 19/20 quantitative predictions are confirmed. The framework is consistent with the edge of stability, Tensor Programs, Dyson Brownian motion, the Lottery Ticket Hypothesis, and neural scaling laws.
Tags
Links
- Source: https://arxiv.org/abs/2603.28964v1
- Canonical: https://arxiv.org/abs/2603.28964v1
Trouble viewing inline? Open PDF directly →
Full Text
200,156 characters extracted from source content.
Expand or collapse full text
The Spectral Edge Thesis: A Mathematical Framework for Intra-Signal Phase Transitions in Neural Network Training From Axioms to Flow Equations Yongzhong Xu abbyxu@gmail.com. (March 2026) Abstract Empirical summary. Gap dynamics in the rolling-window Gram matrix precede every grokking event observed: 24/24 runs with weight decay, 0/24 without. Across six model families spanning 150K–124M parameters (TinyStories 51M, GPT-2 124M, and grokking experiments on Dyck-1, SCAN, and modular arithmetic), 19 of 20 quantitative predictions of the framework are confirmed (7/8 in this paper; the remainder in companion papers [33, 34, 37, 35, 36]). The number of simultaneously active modes is empirically small (k∗≤3k^*≤ 3) and optimizer-dependent: Muon drives k∗=1k^*=1 while AdamW produces k∗=2k^*=2 on the same TinyStories 51M model, both reaching comparable validation loss. Framework. We develop a mathematical framework—the spectral edge thesis—showing that phase transitions, grokking, and feature circuit formation are controlled by the spectral gap structure of rolling-window parameter updates. The framework is architecture-agnostic: the architecture enters only through NTK eigenvalues, Hessian curvatures, and the kernel evolution rate. In the extreme aspect ratio regime (P∼108P 10^8, W∼10W 10), the classical BBP detection threshold is vacuous—every eigenvalue is signal—and the operative structure is the intra-signal gap separating dominant from subdominant modes. Theoretical results. From three axioms we derive: (i) the gap position k∗k^* and its characterisation via NTK outliers; (i) Dyson-type gap dynamics with subspace stability from Davis–Kahan; (i) a coupled ODE system for signal strengths, with phase transitions at gap collapse. The adiabatic parameter =‖K˙‖/(ηg2)A=\| K\|/(η g^2) controls circuit stability (≪1A 1: plateau; ∼1 1: phase transition; ≫1 1: forgetting), and is computable from architecture via the Roberts–Yaida–Hanin differential NTK: (0)∼L/(ng2)A(0) L/(ng^2). Each spectral edge event is rank-1 in function space but maps through † J to a dense reorganisation of all parameters (holographic encoding); circuit survival requires the protecting gap to remain large through all subsequent transitions and weight decay compression. Consistency. The framework is consistent with the edge of stability (λmax→2/η _ → 2/η, Cohen et al. 2021), Greg Yang’s feature-learning regime (Tensor Programs), Roberts–Yaida–Hanin criticality, Dyson Brownian motion, the Lottery Ticket Hypothesis, and empirical scaling laws. We view this consilience, alongside the direct empirical tests, as the primary evidence that the spectral gap of the rolling-window Gram matrix is the right object to study. Contents I Axiomatics and Setup 1 The Objects 1.1 Parameter Space and Training Trajectory 1.2 The Aspect Ratio Regime 2 The Three Axioms 3 The Central Question I The Signal Hierarchy 4 The Noise is Negligible 4.1 The BBP Detection Threshold 4.2 Noise Concentration in the Extreme Aspect Ratio 5 Why the Signal Has a Gap: The Hessian Mechanism 5.1 The Hessian Spectral Hierarchy 5.2 The Krylov Subspace Bound I Determining k∗k^*: The Maximum Intra-Signal Gap 6 The Intra-Signal Gap 6.1 Definition of k∗k^* 6.2 The Gap Ratio 6.3 Null Distribution of the Maximum Ratio 7 Connection to k95k_95 and Cumulative Variance 7.1 The Ratio Test IV The Spectral Gap and Phase Transitions 8 The Intra-Signal Gap as Order Parameter 9 Subspace Stability: The Davis–Kahan Connection 10 Phase Transitions as Level Crossings 10.1 Dyson Eigenvalue Dynamics 10.2 Gap Collapse: Absorption of the Dominant Mode 10.3 Gap Opening: Emergence of a New Dominant Mode 10.4 The Avoided Crossing Duration V The Flow Equations 11 Spectral Perturbation of the Covariance 12 Signal Strength Flow 12.1 The Gradient Flow Component 12.2 The Anharmonic Replenishment 12.3 Simplified Signal Flow 13 Noise Level Dynamics 14 The Coupled System: Gap Flow 14.1 The Complete Dynamical System 14.2 The Gap Flow Equation 14.3 Critical Dynamics Near g=0g=0 14.4 Timescales 15 The Role of β2 _2 (Adam’s Second Moment) VI Coupling to the Loss Function 16 The Stability Coefficient 17 The Loss Improvement Decomposition 17.1 Why the Largest Gap: The Gap Maximality Principle 18 The Master Equation VII The Complete Dynamical System and Where BBP Fits 19 The Full System 20 Where BBP Fits: The Outer Boundary VIII Empirical Verification 21 Test 1: The BBP Threshold is Vacuous 22 Test 2: k∗k^* = Argmax Ratio 22.1 k∗k^* Dynamics: Gap Position Shifts During Training 23 Test 3: Krylov Bound Consistency 24 Test 4: Gap Ratio–Loss Correlation 25 Test 5: Stability Coefficient Hierarchy 26 Test 6: Gap Flow Equation 27 Test 7: Gap Opening and Grokking (Dyck/SCAN) 28 Test 8: Causal Intervention (Loss Decomposition) 29 Summary of Theory–Experiment Match IX Connections, Extensions, and Open Problems 30 Connection to Dyson Brownian Motion 31 Connection to Tensor Programs (Yang) 32 Connection to Roberts–Yaida–Hanin 33 The Three-Phase Pattern 34 The Evolving NTK and Circuit Transport 35 Overparameterisation and the Spectral Gap 36 Scaling Laws from the Spectral Tail 37 Summary of the Three Answers 38 Connection to the Lottery Ticket Hypothesis 39 Connection to the Holographic Encoding Principle 40 Connection to the Edge of Stability 41 Circuit Survival and the Final Model 42 The Geometric Flow Connection 43 Logical Structure and the Commutativity Assumption 44 Open Problems References Figure 1: The spectral edge framework. (A) Singular value spectrum of the Gram matrix G for TinyStories 51M, showing the intra-signal gap at k∗=2k^*=2. All eigenvalues are far above the BBP noise detection threshold (dashed; see Proposition 4.1). (B) Ratio profile σk/σk+1 _k/ _k+1: the maximum at k=2k=2 defines the signal rank. (C) Gap ratio σ2/σ3 _2/ _3 over training for TinyStories 51M, showing the three-phase pattern: rise, plateau, collapse. (D) Gap ratio σ3/σ4 _3/ _4 tracks validation loss for GPT-2 124M. The spectral edge event (“Shift”) coincides with the loss plateau. Part I Axiomatics and Setup 1 The Objects 1.1 Parameter Space and Training Trajectory Definition 1.1 (Training Trajectory). A neural network with p trainable parameters defines a path :0,1,2,…,T→ℝp θ:\0,1,2,…,T\ ^p in parameter space. The parameter update at step t is t=t+1−t∈ℝp. δ_t\;=\; θ_t+1- θ_t\;∈\;R^p. For an optimizer O (SGD, Adam, etc.) with learning rate ηt _t, this takes the form t=−ηt(∇Lℬt(t),statet) δ_t=- _t\,O(∇ L_B_t( θ_t),state_t), where ℬtB_t is the minibatch. Definition 1.2 (Rolling Window and Trajectory Matrix). Fix a window size W≥2W≥ 2. At time t0t_0, the trajectory matrix is (t0)=(t0⊤t0+1⊤⋮t0+W−1⊤)∈ℝW×p. X(t_0)\;=\; pmatrix δ_t_0 \\ δ_t_0+1 \\ \\ δ_t_0+W-1 pmatrix\;∈\;R^W× p. We suppress the t0t_0 dependence when clear from context and write ∈ℝW×p X ^W× p. Definition 1.3 (Gram Matrix and Spectrum). The Gram matrix is =⊤∈ℝW×W G= X X ^W× W. Its eigendecomposition =⊤ G= U U yields eigenvalues λ1≥λ2≥⋯≥λW≥0 _1≥ _2≥·s≥ _W≥ 0, which are the squared singular values σk2 _k^2 of X. The right singular vectors k∈ℝp v_k ^p (the principal directions of the trajectory) are k=σk−1⊤k v_k= _k^-1 X u_k. Remark 1.4 (Why the Gram Matrix). Since p≫Wp W (typically p∼106p 10^6–101010^10, W∼10W 10–3030), we never form the p×p× p covariance ⊤ X X. The Gram matrix ∈ℝW×W G ^W× W shares the same nonzero eigenvalues and is O(W2)O(W^2) to store and O(W3)O(W^3) to diagonalize. Definition 1.5 (Sliding-Window Covariance). The sliding-window covariance is (t):=(t)⊤(t)=∑r=0W−1t−rt−r⊤∈ℝp×p. C(t)\;:=\; X(t) X(t)\;=\; _r=0^W-1 δ_t-r\, δ_t-r \;∈\;R^p× p. Its nonzero eigenvalues coincide with those of (t)=(t)(t)⊤∈ℝW×W G(t)= X(t) X(t) ^W× W, but its eigenvectors are the right singular vectors j∈ℝp v_j ^p of X: (t)j(t)=λj(t)j(t),λj=dj2. C(t)\, v_j(t)= _j(t)\, v_j(t), _j=d_j^2. The perturbation theory for eigenvalues and eigenvectors (Section 11) operates on C rather than G, since the signal directions j v_j live in parameter space ℝpR^p. 1.2 The Aspect Ratio Regime Definition 1.6 (Aspect Ratio). The aspect ratio is γ=p/Wγ=p/W. In our setting, γ∼104γ 10^4–10810^8. We call this the extreme aspect ratio regime: γ→∞γ→∞ with W fixed. Remark 1.7 (Classical vs. Extreme RMT). Classical random matrix theory studies p,n→∞p,n→∞ with p/n→γ∈(0,∞)p/n→γ∈(0,∞) fixed. Our regime is W fixed, p→∞p→∞. This simplifies the theory: all W eigenvalues of G are observable, and by the law of large numbers, noise eigenvalues (if present) concentrate tightly. The fluctuation theory (Tracy–Widom) gives corrections—but as we shall show, the noise is negligible in practice, and the relevant structure is within the signal. 2 The Three Axioms The entire framework rests on three axioms about the structure of the trajectory matrix X. Axiom 2.1 (Hierarchical Signal Decomposition). The trajectory matrix admits the decomposition =dom+sub+, X\;=\; S_dom+ S_sub+ N, (1) where: • dom∈ℝW×p S_dom ^W× p is the dominant signal (the “backbone”), of rank k∗≪Wk^* W, representing the parameter updates projected onto the loss landscape’s primary descent directions—those aligned with the top Hessian eigenvectors. • sub∈ℝW×p S_sub ^W× p is the subdominant signal, of rank W−k∗W-k^*, representing coherent but weaker update components aligned with smaller (but still nonzero) Hessian eigenvalues. • ∈ℝW×p N ^W× p is the noise matrix, representing stochastic fluctuations from minibatch sampling. The critical structural assumption: ∥op≪σW(dom+sub), N _op\; \; _W( S_dom+ S_sub), (2) i.e., the noise operator norm is much smaller than the smallest signal singular value. In the extreme aspect ratio regime, all W singular values of X are signal. Axiom 2.2 (Gap Structure). The singular values of the full signal =dom+sub S= S_dom+ S_sub exhibit a spectral gap: there exists a position k∗∈1,…,W−1k^*∈\1,…,W-1\ such that dk∗dk∗+1=max1≤j≤W−1djdj+1≫ 1, d_k^*d_k^*+1\;=\; _1≤ j≤ W-1 d_jd_j+1\; \;1, (3) where d1≥d2≥⋯≥dW>0d_1≥ d_2≥·s≥ d_W>0 are the singular values of S. The gap separates the k∗k^* dominant modes from the W−k∗W-k^* subdominant modes. This gap arises from the hierarchical structure of the Hessian spectrum (see Section 5). Axiom 2.3 (Slow Variation). The signal directions j\ v_j\, signal strengths dj\d_j\, and noise covariance N _N vary slowly on the timescale of the window: ∥j(t+W)−j(t)∥j(t)∥=O(ε),|dj(t+W)−dj(t)|dj(t)=O(ε),ε≪1. v_j(t+W)- v_j(t) v_j(t) =O( ), |d_j(t+W)-d_j(t)|d_j(t)=O( ), 1. This allows us to treat the signal parameters as approximately constant within each window, while studying their evolution across windows. Remark 2.4 (Justification of the Axioms). Axiom 2.1 is justified by the empirical observation that in the extreme aspect ratio regime (p∼108p 10^8, W∼10W 10), the BBP detection threshold dcrit=ν(p(W−1))1/4d_crit=ν(p(W-1))^1/4 evaluates to ∼0.2 0.2–1.51.5, while the smallest observed singular value σW∼12 _W 12–8080. The ratio σW/dcrit∼20 _W/d_crit 20–300300 (see Section 4). Every eigenvalue is 2020–300×300× above the noise detection threshold. Axiom 2.2 is justified by: (i) the Hessian spectrum of neural networks concentrates in O(k)O(k) large eigenvalues with a near-zero bulk (Sagun et al. [21], Ghorbani et al. [16]); (i) gradient descent projects onto the top Hessian eigendirections (Gur-Ari et al. [17]); and (i) empirically, the maximum ratio dk∗/dk∗+1d_k^*/d_k^*+1 peaks at k∗=2k^*=2 (TinyStories, ratio ∼1.79 1.79) and k∗=3k^*=3 (GPT-2, ratio ∼1.12 1.12). Axiom 2.3 is only used for the dominant modes (j≤k∗j≤ k^*), where it holds strongly. By Davis–Kahan (Theorem 9.1), the subspace perturbation of mode j satisfies sinθj≤‖Δ‖F/gapj _j≤\| G\|_F/gap_j. For the dominant modes (j<k∗j<k^*), the gap is so large that the directions are essentially frozen—ε≪1 1 independent of W. At the spectral edge (j=k∗j=k^*), the gap is smaller and ε grows with W, but remains manageable for W∼10W 10–3030 (Section 28). Beyond the edge (j>k∗j>k^*), slow variation fails, but the framework never invokes the axiom for those modes. 3 The Central Question Given the trajectory matrix X and its observed singular values σ1≥⋯≥σW _1≥·s≥ _W: (I) Where is the spectral gap? That is, what is the position k∗(t)k^*(t) of the maximum intra-signal ratio? (I) How does the gap evolve? What controls the dynamics of g(t)=dk∗(t)−dk∗+1(t)g(t)=d_k^*(t)-d_k^*+1(t)? (I) What are the flow equations? How do the signal strengths dj(t)j=1W\d_j(t)\_j=1^W evolve, and when do phase transitions occur? The spectral edge thesis states: phase transitions in learning occur when the spectral gap g(t)g(t) collapses or opens—i.e., when eigenvalues within the signal hierarchy undergo level crossings. Remark 3.1 (Architecture Independence: A Geometric Flow Theory). The spectral edge framework is architecture-agnostic. The Gram matrix, signal flow ODE, gap dynamics, stability coefficient αj _j, loss decomposition, evolving-NTK signal equation, adiabatic parameter A, and the edge-of-stability identification all hold for any differentiable model with a well-defined NTK—MLPs, CNNs, transformers, or otherwise. Why? The spectral edge dynamics are a geometric flow on the NTK manifold (Section 42). The NTK defines a metric on function space; training evolves this metric. The Hessian curvatures hjh_j are essentially the NTK eigenvalues (via the Gauss–Newton approximation hj≈λj/Nh_j≈ _j/N), and the gradient projections GjG_j depend on the NTK eigenbasis plus the residuals. The only information beyond the metric is the residuals ri=f(xi)−yir_i=f(x_i)-y_i, encoding where learning currently stands relative to the target. The architecture enters only through the spectral inputs λk,hj,ck,K˙\ _k,h_j,c_k, K\; any model producing the same inputs yields the same flow. The logical chain: 1. Geometric flow: the NTK defines a metric; training evolves it; the Hessian and gradient projections are largely determined by the metric, with the residuals as the only additional input. All architecture dependence is absorbed into λk,hj,ck,K˙\ _k,h_j,c_k, K\. 2. ⟹ Architecture independence: any model producing the same spectral inputs gives the same flow. 3. ⟹ Universal phenomena: phase transitions at gap collapse, k∗≤3k^*≤ 3, the circuit lifecycle, grokking, holographic encoding—all follow from the flow equations, not from architectural specifics. The architecture determines the spectral inputs (eigenvalue spacing, curvature profile, kernel evolution rate), but the dynamical laws governing how those inputs produce training phenomena are architecture-independent. In particular, k∗=2k^*=2–33 is empirically consistent with an edge-of-stability constraint on the Hessian spectrum (see Remark 40.2), though this remains an open question rather than a proved result. What is architecture-dependent is how to compute the initial conditions. The RYH connection (Section 32), the depth–width tradeoff (0)∼L/(ng2)A(0) L/(ng^2), and the tensor programs framework (Section 31) compute the initial metric and its evolution rate for specific architecture classes. Part I The Signal Hierarchy 4 The Noise is Negligible We first establish that the classical BBP framework is vacuous in the extreme aspect ratio regime. 4.1 The BBP Detection Threshold Proposition 4.1 (BBP is Vacuous). In the extreme aspect ratio regime (p≫Wp W, p∼106p 10^6–101010^10, W∼10W 10–3030), the BBP detection threshold for isotropic noise N=ν2p _N=ν^2 I_p is dcrit=ν⋅(p(W−1))1/4.d_crit=ν·(p(W-1))^1/4. (4) This threshold is trivially satisfied by every eigenvalue. Specifically, the per-coordinate noise standard deviation ν=O(η/B⋅p)ν=O(η/ B· p) where B is the batch size, giving: dcrit∼ηB⋅p⋅(p(W−1))1/4=η(W−1)1/4B1/2⋅p1/4.d_crit η B· p·(p(W-1))^1/4= η\,(W-1)^1/4B^1/2· p^1/4. (5) For typical values (η=10−3η=10^-3, B=20B=20, p=1.6×108p=1.6× 10^8, W=10W=10), this gives dcrit∼0.2d_crit 0.2–1.51.5. Meanwhile, the smallest observed singular value satisfies σW∼12 _W 12–8080, giving: σWdcrit∼ 20–300. _Wd_crit\; \;20--300. (6) Every eigenvalue of G is 2020–300×300× above the BBP detection threshold. The BBP phase transition is never approached. Proof. From the Gram matrix spectrum of a pure-noise matrix ∈ℝW×p N ^W× p with i.i.d. rows s∼(0,ν2p) n_s (0,ν^2 I_p): the eigenvalues of ⊤ N N concentrate at pν2pν^2 with spread O(ν2p)O(ν^2 p). The BBP threshold requires dj2>pν2W−1d_j^2>pν^2 W-1, i.e., dj>ν(p(W−1))1/4d_j>ν(p(W-1))^1/4. The per-coordinate noise for Adam with learning rate η and batch size B satisfies ν≲η/B⋅pν η/ B· p (since the preconditioner approximately normalizes the noise, giving ν2≈η2/pν^2≈η^2/p; see Cohen et al. [14], Roberts et al. [20]). Substituting: dcrit≤ηB⋅(W−1)1/4p1/4.d_crit≤ η B· (W-1)^1/4p^1/4. With η=10−3η=10^-3, B=20B=20, W=10W=10, p=1.6×108p=1.6× 10^8: dcrit≤10−3⋅0.22⋅(9)0.25/(1.6×108)0.25≈1.3d_crit≤ 10^-3· 0.22·(9)^0.25/(1.6× 10^8)^0.25≈ 1.3. Empirically, the smallest Gram eigenvalue across all windows and seeds is λW≥150 _W≥ 150 (TinyStories) and λW≥600 _W≥ 600 (GPT-2), giving σW=λW≥12 _W= _W≥ 12. Hence σW/dcrit≥12/1.3≈9 _W/d_crit≥ 12/1.3≈ 9, and the typical ratio is ∼50 50–300300. ∎ Remark 4.2 (What the BBP Predicts vs. What We Observe). The BBP framework predicts k∗=Wk^*=W (every eigenvalue is “signal” in the BBP sense). Yet the empirical signal rank is k∗=2k^*=2 (TinyStories) or k∗=3k^*=3 (GPT-2). This is not a contradiction: the empirical k∗k^* is not the signal-noise boundary. It is the position of the maximum spectral gap within the signal hierarchy—the border between dominant and subdominant modes, not between signal and noise. Remark 4.3 (When BBP Becomes Binding). The BBP threshold becomes relevant when: • W is much larger (approaching p p), so the noise eigenvalues spread enough to overlap with weak signal. • p is much smaller (e.g., per-layer analysis with p∼103p 10^3). • The noise is much stronger (very high learning rate, very small batch). In the standard regime of this paper (p∼108p 10^8, W∼10W 10), the BBP transition is vacuous, and the physics is entirely within the signal spectrum. 4.2 Noise Concentration in the Extreme Aspect Ratio Proposition 4.4 (Noise Eigenvalue Concentration). For ∈ℝW×p N ^W× p with i.i.d. rows s∼(0,N) n_s (0, _N), the Gram eigenvalues μ1≥⋯≥μW _1≥·s≥ _W of ⊤ N N satisfy: μmax−μminμ¯=O(W⋅κN),μ¯=Tr(N),κN=∥N∥F2Tr(N)2. _ - _ μ\;=\;O\! ( W· _N ), μ=Tr( _N), _N= _N _F^2Tr( _N)^2. (7) For isotropic noise (κN=1/p _N=1/p), the relative spread is O(W/p)∼10−3.5O( W/p) 10^-3.5: the noise eigenvalues are essentially degenerate. Even for colored noise from Adam (κN∼104/p _N 10^4/p), the noise operator norm is μ¯(1+O(W⋅104/p))≈μ¯(1+O(10−1.5)) μ(1+O( W· 10^4/p))≈ μ(1+O(10^-1.5))—still far below the signal eigenvalues. Proof. The Gram matrix entries are GijN=∑a=1pτaziazjaG_ij^N= _a=1^p _az_iaz_ja where τa _a are eigenvalues of N _N and ziaz_ia are i.i.d. standard normal. By the law of large numbers, GiiN=∑aτazia2→∑aτa=Tr(N)G_i^N= _a _az_ia^2→ _a _a=Tr( _N) as p→∞p→∞, and GijN=∑aτaziazja=O(∑aτa2)=O(Tr(N)κN)G_ij^N= _a _az_iaz_ja=O( _a _a^2)=O(Tr( _N) _N) for i≠ji≠ j. The eigenvalue spread of the W×W× W Gram matrix is controlled by the off-diagonal fluctuations, giving the stated bound. ∎ 5 Why the Signal Has a Gap: The Hessian Mechanism The spectral gap in the trajectory (2.2) is not accidental. It is consistent with the hierarchical structure of the loss landscape Hessian, as the following proposition makes precise. 5.1 The Hessian Spectral Hierarchy Proposition 5.1 (Hessian ⇒ Trajectory Gap). Assume [,]≈0[P, H]≈ 0. Let the Hessian H of the loss function have eigenvalues h1≥h2≥⋯≥hp≥0h_1≥ h_2≥·s≥ h_p≥ 0 with a spectral gap at position k: hk≫hk+1.h_k\; \;h_k+1. (8) Then the signal strengths djd_j of the trajectory matrix satisfy, at steady state: djss∝1hj⋅|⟨j,∇L⟩|,d_j^s\; \; 1 h_j· v_j,P∇ L , (9) where P is the optimizer preconditioner (the exact formula is djss=η|Gjeff|Φjd_j^s=η\,|G_j^eff| _j; see Corollary 12.20). Here (=IP=I for SGD; t=diag(1/v^t+ϵ)P_t=diag(1/ v_t+ε) for Adam, with v^t v_t the bias-corrected EMA of squared gradients controlled by β2 _2; see Section 15). If the gradient has comparable projections onto all Hessian eigendirections (|⟨j,∇L⟩|≈G v_j,P∇ L ≈ G for j≤Kj≤ K), then: dkssdk+1s=hk+1hk≪ 1. d_k^sd_k+1^s= h_k+1h_k\; \;1. (10) The ratio is inverted: the direction with smaller Hessian curvature hjh_j has larger signal strength djd_j (since dj∝1/hjd_j 1/ h_j). Under the comparable-projection assumption, dk+1≫dkd_k+1 d_k: the direction with lower curvature accumulates stronger signal. A gap in the Hessian spectrum (hk≫hk+1h_k h_k+1) therefore produces a gap in the trajectory spectrum at the same position k, but with inverted relative strengths—the dominant trajectory modes are the low-curvature Hessian eigendirections. Remark 5.2 (The Spectral Response Function (Heuristic)). A rough heuristic for the signal hierarchy is djss∝|Gj|/hjd_j^s |G_j|/ h_j: lower curvature ⇒ stronger signal. The trajectory signal strength is the gradient projection amplified by the spectral response. Large F (low curvature, large window, large learning rate) means a strong trajectory signal. The spectral gap in the trajectory occurs where F(hj)/F(hj+1)F(h_j)/F(h_j+1) is maximized—i.e., where the Hessian eigenvalue ratio hj+1/hjh_j+1/h_j is maximized. Remark 5.3 (Two Sources of Gap). The trajectory gap can arise from two sources: 1. Hessian gap: A large gap in hk/hk+1h_k/h_k+1 translates directly via dj∝1/hjd_j 1/ h_j. 2. Gradient alignment asymmetry: Even with a flat Hessian spectrum, if |⟨j,∇L⟩| v_j,P∇ L drops sharply at some j=kj=k, the trajectory gap appears there. In practice, the Hessian gap is the dominant mechanism: neural network Hessians have O(1)O(1)–O(10)O(10) large outlier eigenvalues separated by orders of magnitude from the bulk (Sagun et al. [21], Papyan [19]). Remark 5.4 (Four Channels of Curvature Separation). Curvature separation is not synonymous with amplitude modulation (dj∝1/hjd_j 1/ h_j). It operates through four independent channels: 1. Dissipation rate: the ODE term −2η(hj+ω)dj2-2η(h_j+ω)d_j^2 separates modes by their convergence speed, even when steady-state amplitudes are similar. 2. Weight-decay threshold: the grokking condition λgen/N>ω>λmem/N _gen/N>ω> _mem/N (Remark 12.22) is a binary curvature gate that compresses low-curvature modes to a common floor. 3. Eigenvector stability: Davis–Kahan (Theorem 9.1) ties eigenvector rotation to eigenvalue gaps, which are curvature-driven. The empirical hierarchy α1=0.818>α2=0.234>α≥4≈0 _1=0.818> _2=0.234> _≥ 4≈ 0 is a direct, model-free observation of this channel. 4. Gap flow: Term I of the gap equation (Theorem 14.2), −η(hk∗−hk∗+1)d¯-η(h_k^*-h_k^*+1) d, is a curvature-difference force independent of Φ . In the weak-curvature regime (Φj≈W _j≈ W), channel 1’s amplitude effect weakens, but channels 2–4 remain fully active. The edge eigenvector rotating fast while interior eigenvectors barely move—observed across GPT-2 124M and TinyStories 51M—is curvature separation operating through channel 3. 5.2 The Krylov Subspace Bound Proposition 5.5 (Krylov Bound on k∗k^*). Assume [,]≈0[P, H]≈ 0. Consider W consecutive gradient descent steps with constant learning rate η and Hessian H. Starting from initial gradient 0=∇L(0) g_0=∇ L( θ_0), the trajectory matrix X has rows: s=−η(−η)s0,s=0,1,…,W−1. δ_s=- ( I- H)^s g_0, s=0,1,…,W-1. (11) These rows span the Krylov subspace W(−η,0)=span0,(−η)0,…K_W( I- H,P g_0)=span\P g_0,( I- H)P g_0,…\. If the Hessian has K distinct large eigenvalues (separated from the bulk by a gap), then the effective rank of the Krylov subspace is min(K,W) (K,W). The dominant signal rank satisfies: k∗≤K=#j:hj is a Hessian outlier.k^*\;≤\;K\;=\;\#\j:h_j is a Hessian outlier\. (12) Proof sketch. The Krylov subspace W(,)K_W( M, g) with =−η M= I- H has the property that its projection onto the eigenspace of M corresponding to eigenvalue μj=1−ηhj _j=1-η h_j (for preconditioned curvature hjh_j) produces a component of magnitude ∼|μj|s|⟨j,⟩| _j ^s e_j, g . When hjh_j is large, |μj| _j is far from 1, producing a rapidly varying component. When hj≈0h_j≈ 0, μj≈1 _j≈ 1, and the component is nearly constant across steps—it does not contribute to the trajectory variation. The effective rank of the trajectory matrix is therefore determined by the number of Hessian eigenvalues hjh_j for which 1−ηhj1-η h_j is “distinguishable from 1” over W steps, i.e., ηhjW≳1η h_jW 1. This gives K=#j:hj≳1/(ηW)K=\#\j:h_j 1/(η W)\. ∎ Remark 5.6 (Empirical Verification). For TinyStories (η=10−3η=10^-3, W=10W=10): 1/(ηW)=1001/(η W)=100. The Hessian has ∼2 2–33 eigenvalues above this threshold, matching k∗=2k^*=2. For GPT-2 (η=3×10−5η=3× 10^-5, W=10W=10): 1/(ηW)=33331/(η W)=3333. The Hessian has ∼3 3–44 eigenvalues above this threshold, matching k∗=3k^*=3. Part I Determining k∗k^*: The Maximum Intra-Signal Gap 6 The Intra-Signal Gap 6.1 Definition of k∗k^* Definition 6.1 (Spectral Gap Position). The spectral gap position at time t is: k∗(t)=argmax1≤j≤W−1σj(t)σj+1(t), k^*(t)\;=\; *arg\,max_1≤ j≤ W-1 _j(t) _j+1(t), (13) where σ1≥σ2≥⋯≥σW>0 _1≥ _2≥·s≥ _W>0 are the observed singular values of (t) X(t). Equivalently, in terms of the population signal strengths d1≥⋯≥dWd_1≥·s≥ d_W (which equal the observed singular values up to negligible noise corrections, by Proposition 4.1): k∗(t)=argmax1≤j≤W−1dj(t)dj+1(t).k^*(t)= *arg\,max_1≤ j≤ W-1 d_j(t)d_j+1(t). (14) Remark 6.2 (Contrast with BBP Definition). In the classical BBP framework, k∗k^* is defined as #j:dj>dcrit\#\j:d_j>d_crit\—the number of signals above the noise detection threshold. Our definition is fundamentally different: k∗k^* is the position of the maximum consecutive ratio within the spectrum, regardless of any noise threshold. Since the BBP threshold is trivially satisfied by all eigenvalues (Proposition 4.1), the classical kBBP∗=Wk^*_BBP=W always, which carries no information. Our k∗k^* captures the internal structure of the signal hierarchy. 6.2 The Gap Ratio Definition 6.3 (Gap Ratio). The gap ratio at position k∗k^* is: R(t)=σk∗(t)σk∗+1(t)=max1≤j≤W−1σj(t)σj+1(t).R(t)\;=\; _k^*(t) _k^*+1(t)\;=\; _1≤ j≤ W-1 _j(t) _j+1(t). (15) This is the scale-invariant measure of the spectral gap. Values R≫1R 1 indicate a strong gap; R→1R→ 1 indicates gap collapse. 6.3 Null Distribution of the Maximum Ratio To determine whether an observed gap is “real” (i.e., not an artifact of random fluctuations in a flat spectrum), we need the null distribution. Proposition 6.4 (Null Distribution of Maximum Ratio). Under the null hypothesis H0H_0: all signal strengths are equal (d1=d2=⋯=dWd_1=d_2=·s=d_W), the ordered eigenvalues of the W×W× W Gram matrix follow a Laguerre ensemble. The distribution of the maximum consecutive ratio Rmax=maxjλj/λj+1R_ = _j _j/ _j+1 satisfies, for W fixed and p→∞p→∞: Pr[max1≤j≤W−1λjλj+1>r]= 1−FW(r), \! [ _1≤ j≤ W-1 _j _j+1>r ]\;=\;1-F_W(r), (16) where FW(r)F_W(r) is expressible in terms of the GOE eigenvalue spacing distribution. For practical purposes: • W=10W=10, p=108p=10^8: the 95%95\% quantile is r0.95≈1+6/p≈1.0006r_0.95≈ 1+6/ p≈ 1.0006. • Any observed ratio R>1.01R>1.01 is significant at the 10−610^-6 level under the isotropic null. Proof sketch. Under H0H_0, the Gram matrix =d2(SS⊤+noise) G=d^2( U_S U_S +noise) where S U_S has structure determined by the signal temporal pattern. In the limit p→∞p→∞, the eigenvalues of G concentrate at d2+pν2d^2+pν^2 with fluctuations of order ν2pν^2 p. The ratios of consecutive eigenvalues are 1+O(1/p)1+O(1/ p). By the Tracy–Widom universality theorem, the maximum ratio over W−1W-1 pairs has the stated distribution. ∎ Remark 6.5 (Practical Significance). Our observed ratios are: • TinyStories: σ2/σ3≈1.79 _2/ _3≈ 1.79 (peak), mean ≈1.36≈ 1.36. • GPT-2: σ3/σ4≈1.12 _3/ _4≈ 1.12 (peak), mean ≈1.08≈ 1.08. Both are vastly larger than the null expectation of 1+O(10−4)1+O(10^-4). The gap is “real” at overwhelming significance. It reflects genuine structure in the signal hierarchy, not random fluctuations. 7 Connection to k95k_95 and Cumulative Variance Definition 7.1 (k95k_95: Cumulative Variance Threshold). The 95%95\%-variance rank is: k95(t)=minj:∑i=1jλi∑i=1Wλi≥0.95.k_95(t)= \j: _i=1^j _i _i=1^W _i≥ 0.95 \. (17) Remark 7.2 (Empirical Relationship Between k∗k^* and k95k_95). In the models we study, k∗≤k95k^*≤ k_95 empirically. This is not a theorem—one can construct spectra where the sharpest ratio is at a higher position than the 95% variance cutoff—but it holds in practice because the spectral gap in neural network training occurs near the top of the spectrum. The relationship depends on the spectrum shape: 1. Sharp gap: If dk∗/dk∗+1≫1d_k^*/d_k^*+1 1, then k∗≈k95k^*≈ k_95 (the dominant modes capture nearly all variance). 2. Gradual decay: If the signal spectrum decays smoothly, k95≫k∗k_95 k^* (many subdominant modes contribute significant variance, but no single ratio is large). 3. GPT-2 example: k∗=3k^*=3 (sharpest ratio), but k95=14k_95=14–1919 depending on the training phase. The 11–16 subdominant modes collectively carry ∼5 5–30%30\% of the variance. 7.1 The Ratio Test Remark 7.3 (Practical Ratio Test for k∗k^*). By definition (Definition 6.1), k∗k^* is the argmax of σj/σj+1 _j/ _j+1. To assess whether the maximum ratio indicates a genuine gap (rather than noise fluctuations), compare against the null distribution of Proposition 6.4. Empirically, a ratio exceeding a threshold τ∈[1.05,1.30]τ∈[1.05,1.30] at the gap position indicates a genuine spectral gap. The threshold depends on the noise level: • Extremely clean signals (noise CV <1%<1\%): τ=1.05τ=1.05. • Moderate noise (noise CV ∼5 5–20%20\%): τ=1.10τ=1.10–1.151.15. • Noisy settings (noise CV >20%>20\%): τ=1.20τ=1.20–1.301.30. Part IV The Spectral Gap and Phase Transitions 8 The Intra-Signal Gap as Order Parameter Definition 8.1 (Spectral Gap). The spectral gap at time t is the difference between the dominant and subdominant signal strengths at the gap position: g(t)=dk∗(t)−dk∗+1(t). g(t)=d_k^*(t)-d_k^*+1(t). (18) The gap ratio (scale-invariant) is R(t)=σk∗(t)σk∗+1(t).R(t)= _k^*(t) _k^*+1(t). (19) Remark 8.2 (Contrast with Classical Gap). In the BBP framework, the gap is defined as gBBP=dk∗−dcritg_BBP=d_k^*-d_crit, measuring the distance from the noise detection threshold. Since dcritd_crit is trivially small in our regime (Proposition 4.1), gBBP≈dk∗g_BBP≈ d_k^* always—it carries no dynamical information. Our gap g=dk∗−dk∗+1g=d_k^*-d_k^*+1 measures the internal separation within the signal hierarchy, which is the quantity that controls subspace stability and loss coupling. 9 Subspace Stability: The Davis–Kahan Connection The key observation is that the Davis–Kahan sinΘ theorem does not care what lies below the gap—noise or subdominant signal. It only cares about the size of the gap. Theorem 9.1 (Davis–Kahan sinΘ Theorem). Let A and + A+ E be n×n× n symmetric matrices. Let [a,b][a,b] be an interval containing r≥1r≥ 1 eigenvalues of A (possibly a cluster), and let V be the corresponding r-dimensional eigenspace. Let V be the eigenspace of + A+ E corresponding to its eigenvalues closest to [a,b][a,b]. Define the separation δ=minλ∈σ()∖[a,b]dist(λ,[a,b])=minλ∈σ()∖[a,b]min(|λ−a|,|λ−b|),δ= _λ∈σ( A) [a,b]dist(λ,\;[a,b])\;=\; _λ∈σ( A) [a,b] (|λ-a|,\;|λ-b| ), (20) i.e., δ is the distance from [a,b][a,b] to the nearest eigenvalue of A outside [a,b][a,b]. Then: ∥sinΘ(,^)∥F≤∥Fδ. (V, V) _F\;≤\; E _Fδ. (21) Corollary 9.2 (Gap Controls Subspace Stability). The stability of the dominant signal subspace k∗=span(1,…,k∗)V_k^*=span( v_1,…, v_k^*) under perturbation E (from sliding the window by one step) satisfies: ∥sinΘ(k∗(t),k∗(t+1))∥F≤∥Δ∥Fσk∗2−σk∗+12≈∥Δ∥F(dk∗+dk∗+1)⋅g, (V_k^*(t),V_k^*(t+1)) _F\;≤\; G _F _k^*^2- _k^*+1^2\;≈\; G _F(d_k^*+d_k^*+1)· g, (22) where Δ=(t+1)−(t) G= G(t+1)- G(t) is the Gram matrix update from sliding the window, and g=dk∗−dk∗+1g=d_k^*-d_k^*+1 is the spectral gap. As g→0g→ 0: the bound diverges and the dominant subspace k∗V_k^* rotates. The rotation is concentrated at the boundary: k∗ v_k^* mixes with k∗+1 v_k^*+1, while the deeper modes 1,…,k∗−1 v_1,…, v_k^*-1 have their own, much larger gaps and remain individually stable. What we observe is the subspace as a whole rotating because its edge component is replaced. Remark 9.3 (Universality of Davis–Kahan). This is the fundamental reason why the intra-signal gap framework works: Davis–Kahan depends only on the eigenvalue gap δ, regardless of whether the gap separates signal from noise or dominant signal from subdominant signal. The same theorem that would govern the BBP transition (if it were active) governs the intra-signal transition. The mathematics is identical; only the interpretation changes. 10 Phase Transitions as Level Crossings 10.1 Dyson Eigenvalue Dynamics Theorem 10.1 (Eigenvalue Evolution — Hellmann–Feynman). The eigenvalues λ1(t)≥⋯≥λW(t) _1(t)≥·s≥ _W(t) of the Gram matrix (t) G(t), viewed as functions of the training step t, satisfy: dλjdt=j⊤ddtj(exact), d _jdt\;=\; u_j d Gdt u_j (exact), (23) where j u_j are the eigenvectors of (t) G(t) and d/dtd G/dt is the rate of change of the Gram matrix as the window slides. Proof. Differentiate the eigenvalue equation (t)j(t)=λj(t)j(t) G(t) u_j(t)= _j(t) u_j(t) and take the inner product with j u_j. The terms involving ˙j u_j cancel by symmetry of G and normalization j⊤j=1 u_j u_j=1. ∎ Corollary 10.2 (Eigenvalue Repulsion). The second derivative contains a repulsive interaction: d2λjdt2=j⊤d2dt2j⏟direct acceleration+2∑i≠j|i⊤ddtj|2λj−λi⏟eigenvalue repulsion. d^2 _jdt^2\;=\; u_j d^2 Gdt^2 u_j_direct acceleration\;+\; 2 _i≠ j u_i d Gdt u_j ^2 _j- _i_eigenvalue repulsion. (24) Proof. The eigenvector evolution is ˙j=∑i≠ji⊤˙jλj−λii u_j= _i≠ j u_i G u_j _j- _i u_i (from projecting the differentiated eigenvalue equation onto i u_i, i≠ji≠ j). Differentiating (23) and substituting gives the repulsion term with its factor of 22. ∎ Remark 10.3 (The von Neumann–Wigner Non-Crossing Rule). For a one-parameter family of real symmetric matrices (t) G(t), the eigenvalues generically do not cross (von Neumann–Wigner 1929). Near a would-be crossing at λj≈λj+1 _j≈ _j+1, the repulsive acceleration in Corollary 10.2 diverges as ∝1/(λj−λj+1) 1/( _j- _j+1), creating an impassable barrier. This produces an avoided crossing: the eigenvalues approach, repel, and separate, with the eigenvectors exchanging character. In the Gram matrix context: as the dominant and subdominant signal strengths approach each other (dk∗→dk∗+1d_k^*→ d_k^*+1), the divergent repulsive acceleration resists the crossing, and the eigenvectors k∗ v_k^* and k∗+1 v_k^*+1 undergo rapid rotation—exchanging the roles of dominant and subdominant. 10.2 Gap Collapse: Absorption of the Dominant Mode Definition 10.4 (Gap Collapse Event). A gap collapse event at time t∗t^* occurs when the spectral gap closes: g(t∗)=dk∗(t∗)−dk∗+1(t∗)→ 0.g(t^*)=d_k^*(t^*)-d_k^*+1(t^*)\;→\;0. (25) At this point: 1. The dominant subspace k∗V_k^* becomes unstable (Davis–Kahan bound diverges, Corollary 9.2). 2. The eigenvectors k∗ v_k^* and k∗+1 v_k^*+1 undergo rapid rotation (subspace mixing). 3. The gap position k∗k^* may shift (the new maximum ratio may be at a different index). Proposition 10.5 (Gap Collapse ⇒ Loss Stagnation). Under [,]≈0[P, H]≈ 0. At a gap collapse event: 1. The subspace stability coefficient αk∗ _k^* (Definition 16.1) drops to zero. 2. The effective learning rate along direction k∗ v_k^* becomes erratic (the direction is no longer consistently defined). 3. The validation loss improvement decelerates by the amount ηαk∗⟨k∗,∇Ltrain⟩⟨k∗,∇Lval⟩η\, _k^* v_k^*,∇ L_train v_k^*,∇ L_val that mode k∗k^* was contributing. Proof. (1) By Definition 16.1, αk∗=1−C∥Δ∥F2/gapk∗2 _k^*=1-C G _F^2/gap_k^*^2. At collapse, gapk∗→0gap_k^*→ 0 while ∥Δ∥F G _F remains bounded (Axiom 3), so the ratio diverges and αk∗ _k^* is clamped to 0. (2) By Corollary 9.2, the canonical angle between k∗(t) v_k^*(t) and k∗(t+1) v_k^*(t+1) satisfies sinθ≤∥Δ∥F/[(dk∗+dk∗+1)g] θ≤ G _F/[(d_k^*+d_k^*+1)g]. As g→0g→ 0 the bound exceeds 11: the direction can rotate by up to 90∘90 per step. (3) By the spectral loss decomposition (Theorem 17.1), [ΔLval]=−η∑jαj⟨j,∇Ltrain⟩⟨j,∇Lval⟩+…E[ L_val]=-η _j _j v_j,∇ L_train v_j,∇ L_val +… When αk∗→0 _k^*→ 0, the j=k∗j=k^* term vanishes, reducing the loss improvement rate by exactly the contribution of that mode. ∎ 10.3 Gap Opening: Emergence of a New Dominant Mode Definition 10.6 (Gap Opening Event). A gap opening event at time t∗t^* occurs when a new spectral gap forms: g(t∗)=0g(t^*)=0 and g˙(t∗)>0 g(t^*)>0 (the gap begins to grow). This represents a new signal direction separating from the subdominant tier and joining the dominant subspace. Proposition 10.7 (Gap Opening ⇒ Capability Gain). At a gap opening event: 1. A new stable direction k∗ v_k^* emerges with well-defined alignment to the gradient. 2. The effective signal rank increases (the dominant subspace gains a dimension). 3. The validation loss begins to improve along this direction if and only if the gradient projections have the same sign: ⟨k∗,∇Ltrain⟩⟨k∗,∇Lval⟩>0 v_k^*,∇ L_train v_k^*,∇ L_val >0 (i.e., the direction is generalizing, not memorizing). Proof. (1) At t=t∗t=t^* the gap opens: g(t∗)=0g(t^*)=0, g˙(t∗)>0 g(t^*)>0. For t>t∗t>t^*, the gap g(t)>0g(t)>0 grows, so αk∗ _k^* rises from 0 toward 11 (Definition 16.1). Once gapk∗≫∥Δ∥Fgap_k^* G _F, the Davis–Kahan bound (Theorem 9.1) ensures k∗ v_k^* is well-defined (small rotation per step). (2) Before the event, the modes at positions k∗k^* and k∗+1k^*+1 had comparable singular values and no gap: their subspace was two-dimensional but unseparated. After opening, dk∗>dk∗+1d_k^*>d_k^*+1, and the dominant subspace span(1,…,k∗)span( v_1,…, v_k^*) gains a new individually stable direction. (3) By Theorem 17.1, mode k∗k^* now contributes −ηαk∗⟨k∗,∇Ltrain⟩⟨k∗,∇Lval⟩-η\, _k^* v_k^*,∇ L_train v_k^*,∇ L_val to the loss improvement. This term was previously zero (because αk∗=0 _k^*=0). It contributes to validation loss improvement when the two projections have the same sign—i.e., when the emerging direction is a generalizing direction (training progress also helps validation). A memorization direction would have opposite-sign projections and would not improve validation loss. ∎ Remark 10.8 (Grokking as Delayed Gap Opening). The grokking phenomenon maps onto a delayed gap opening event within the signal hierarchy. During the “memorization” phase, the generalization direction exists as a subdominant signal (not as noise—it is well above the BBP threshold). Its signal strength djd_j is comparable to the other subdominant modes, so there is no gap separating it. Grokking occurs when, through slow signal growth driven by the gradient alignment term or slow signal decay of the memorization direction (from weight decay), a gap opens at a new position: the generalization direction separates from the subdominant tier. The “delay” is the time for the gap to form. The “abruptness” comes from the Davis–Kahan bound: even a small gap creates a well-defined, stable subspace. This is fundamentally different from the BBP interpretation (signal emerging from noise). In our framework, the signal was always present—it was simply not separated from other signals by a gap. 10.4 The Avoided Crossing Duration Proposition 10.9 (Avoided Crossing Duration). Near a level crossing where two eigenvalues λk∗ _k^* and λk∗+1 _k^*+1 approach each other, the gap has a minimum value: gmin= 2|V|,g_ \;=\;2\,|V|, (26) where V is the off-diagonal coupling defined in (29) below. The duration of the avoided crossing (time spent with g<g0g<g_0 for some threshold g0g_0) is: Δtcross≈2g0|λ˙k∗−λ˙k∗+1|, t_cross\;≈\; 2g_0 λ_k^*- λ_k^*+1 , (27) where λ˙j=j⊤˙j λ_j= u_j G u_j is the unperturbed drift rate. During this time, the eigenvectors rotate by an angle: Θ≈arctan(gming0). \;≈\; \! ( g_ g_0 ). (28) When V is large (strong avoided crossing), the eigenvectors undergo substantial rotation (Θ large). When V=0V=0 (a true crossing, requiring symmetry or fine-tuning), gmin=0g_ =0 and the eigenvalues pass through each other without mixing—the eigenvectors remain unchanged. Proof. Near the crossing, only the two approaching eigenvalues interact significantly. Let k∗(0),k∗+1(0) u_k^*^(0), u_k^*+1^(0) be the eigenvectors at a reference time far from the crossing (the diabatic basis). Restricting (t) G(t) to this two-dimensional subspace and linearising around the crossing time t∗t^*: M(t)=(λ¯+12Δλ˙(t−t∗)Vλ¯−12Δλ˙(t−t∗)),M(t)\;=\; pmatrix λ+ 12 λ\,(t-t^*)&V\\ V& λ- 12 λ\,(t-t^*) pmatrix, (29) where λ¯=12(λk∗(t∗)+λk∗+1(t∗)) λ= 12 ( _k^*(t^*)+ _k^*+1(t^*) ) is the mean eigenvalue at the crossing, Δλ˙=λ˙k∗−λ˙k∗+1 λ= λ_k^*- λ_k^*+1 is the difference in unperturbed drift rates, and V=(k∗(0))⊤(t∗)k∗+1(0)V\;=\;( u_k^*^(0)) G(t^*)\, u_k^*+1^(0) is the off-diagonal coupling of the Gram matrix in the diabatic basis. This coupling is generically nonzero: the cumulative rank-two updates s R_s build cross-correlation between the two modes (cf. Corollary 11.5). Eq. (26). The eigenvalues of M(t)M(t) are λ¯±(12Δλ˙)2(t−t∗)2+V2 λ± ( 12 λ)^2(t-t^*)^2+V^2. The gap is minimized at t=t∗t=t^*, giving gmin=2|V|g_ =2|V|. Eq. (27). The gap equals g0g_0 when (12Δλ˙)2(t−t∗)2+V2=(12g0)2( 12 λ)^2(t-t^*)^2+V^2=( 12g_0)^2. Solving: |t−t∗|=g02−gmin2/|Δλ˙||t-t^*|= g_0^2-g_ ^2\,/\,| λ|. For gmin≪g0g_ g_0, the total duration is Δtcross≈2g0/|Δλ˙| t_cross≈ 2g_0/| λ|. Eq. (28). At t=t∗t=t^* the eigenvectors of M are rotated 45∘45 relative to the diabatic basis (the diagonal splitting vanishes). Away from the crossing, at times where g=g0g=g_0, the mixing angle satisfies tanΘ=gmin/g0 =g_ /g_0. When V is large (gmin≈g0g_ ≈ g_0), the rotation is substantial. When V=0V=0 (true crossing), gmin=0g_ =0 and Θ=0 =0: the eigenvalues cross but the eigenvectors do not mix. ∎ Part V The Flow Equations This is the dynamical core of the framework. We first establish the perturbation theory for the sliding-window covariance (t) C(t) (Definition 1.5), then derive the equations governing the evolution of all signal strengths dj(t)j=1W\d_j(t)\_j=1^W along the training trajectory, and consequently the evolution of the spectral gap g(t)g(t) and gap position k∗(t)k^*(t). 11 Spectral Perturbation of the Covariance The eigenvalue and eigenvector dynamics of the signal directions j∈ℝp v_j ^p are governed by the p×p× p sliding-window covariance (t)=(t)⊤(t) C(t)= X(t) X(t), not by the W×W× W Gram matrix (t)=(t)(t)⊤ G(t)= X(t) X(t) . (The two share nonzero eigenvalues, but their eigenvectors live in different spaces: j∈ℝW u_j ^W vs. j∈ℝp v_j ^p.) Proposition 11.1 (Exact rank-two update of the sliding-window covariance). Let (t):=∑r=0W−1t−rt−r⊤∈ℝp×p. C(t)\;:=\; _r=0^W-1 δ_t-r\, δ_t-r \;∈\;R^p× p. Then (t+1)−(t)=t+1t+1⊤−t−W+1t−W+1⊤. C(t+1)- C(t)\;=\; δ_t+1 δ_t+1 - δ_t-W+1 δ_t-W+1 . Equivalently, defining the rank-two update t:=t+1t+1⊤−t−W+1t−W+1⊤ R_t:= δ_t+1 δ_t+1 - δ_t-W+1 δ_t-W+1 , we have (t+1)=(t)+t C(t+1)= C(t)+ R_t. Proof. By definition, (t+1)=∑r=0W−1t+1−rt+1−r⊤=t+1t+1⊤+∑r=1W−1t+1−rt+1−r⊤. C(t+1)= _r=0^W-1 δ_t+1-r\, δ_t+1-r = δ_t+1 δ_t+1 + _r=1^W-1 δ_t+1-r\, δ_t+1-r . Reindexing the remaining sum gives (t+1)=t+1t+1⊤+∑r=0W−2t−rt−r⊤. C(t+1)= δ_t+1 δ_t+1 + _r=0^W-2 δ_t-r\, δ_t-r . Similarly, (t)=∑r=0W−2t−rt−r⊤+t−W+1t−W+1⊤. C(t)= _r=0^W-2 δ_t-r\, δ_t-r + δ_t-W+1 δ_t-W+1 . Subtracting yields (t+1)−(t)=t+1t+1⊤−t−W+1t−W+1⊤ C(t+1)- C(t)= δ_t+1 δ_t+1 - δ_t-W+1 δ_t-W+1 . ∎ Proposition 11.2 (Second-order perturbation of a simple eigenvalue). Let ∈ℝp×p C ^p× p be symmetric with orthonormal eigenbasis j=λjj C v_j= _j v_j, j=1,…,pj=1,…,p. Fix k and assume λk _k is simple, with spectral gap δk:=minj≠k|λk−λj|>0 _k:= _j≠ k| _k- _j|>0. Let ∈ℝp×p R ^p× p be symmetric with ∥≤δk/4 R ≤ _k/4. Then + C+ R has a unique eigenvalue λ~k λ_k in (λk−δk/2,λk+δk/2)( _k- _k/2,\; _k+ _k/2), and λ~k−λk=k⊤k+∑j≠k|j⊤k|2λk−λj+(∥3δk2). λ_k- _k\;=\; v_k R\, v_k+ _j≠ k | v_j R\, v_k|^2 _k- _j+O\! ( R ^3 _k^2 ). (30) Proof. Let ~k v_k be a normalised eigenvector of + C+ R for λ~k λ_k, with ⟨~k,k⟩>0 v_k, v_k >0. Expand ~k=∑jajj v_k= _ja_j v_j, ∑j|aj|2=1 _j|a_j|^2=1. The eigenvalue equation (+)~k=λ~k~k( C+ R) v_k= λ_k v_k gives, after inner product with j v_j: (λj−λ~k)aj+j⊤~k=0.( _j- λ_k)\,a_j+ v_j R\, v_k=0. For j≠kj≠ k: aj=j⊤~k/(λ~k−λj)a_j= v_j R\, v_k/( λ_k- _j). Since |λ~k−λj|≥δk/2| λ_k- _j|≥ _k/2: |aj|≤2∥δk,ak=1+(∥2δk2),aj=(∥δk)(j≠k).|a_j|≤ 2 R _k, a_k=1+O\! ( R ^2 _k^2 ), a_j=O\! ( R _k )\;\;(j≠ k). Inner product with k v_k: (λ~k−λk)ak=k⊤~k=akk⊤k+∑j≠kajk⊤j.( λ_k- _k)\,a_k= v_k R\, v_k=a_k\, v_k R\, v_k+ _j≠ ka_j\, v_k R\, v_j. Dividing by ak=1+(∥2/δk2)a_k=1+O( R ^2/ _k^2): λ~k−λk=k⊤k+∑j≠kajk⊤j+(∥3δk2). λ_k- _k= v_k R\, v_k+ _j≠ ka_j\, v_k R\, v_j+O\! ( R ^3 _k^2 ). Replacing ~k v_k by k v_k and λ~k λ_k by λk _k in aja_j incurs only second-order error, giving aj=j⊤k/(λk−λj)+(∥2/δk2)a_j= v_j R\, v_k/( _k- _j)+O( R ^2/ _k^2). Since R is symmetric, (j⊤k)(k⊤j)=|j⊤k|2( v_j R\, v_k)( v_k R\, v_j)=| v_j R\, v_k|^2, yielding (30). ∎ Corollary 11.3 (First- and second-order increment of λk(t) _k(t)). Let (t)j(t)=λj(t)j(t) C(t) v_j(t)= _j(t) v_j(t) with λk(t) _k(t) simple, and t=(t+1)−(t)=t+1t+1⊤−t−W+1t−W+1⊤ R_t= C(t+1)- C(t)= δ_t+1 δ_t+1 - δ_t-W+1 δ_t-W+1 . Assume ∥t∥≤δk(t)/4 R_t ≤ _k(t)/4. Then λk(t+1)−λk(t)=k(t)⊤tk(t)+∑j≠k|j(t)⊤tk(t)|2λk(t)−λj(t)+(∥t∥3δk(t)2). _k(t+1)- _k(t)= v_k(t) R_t\, v_k(t)+ _j≠ k | v_j(t) R_t\, v_k(t)|^2 _k(t)- _j(t)+O\! ( R_t ^3 _k(t)^2 ). (31) In particular, the first-order term is k(t)⊤tk(t)=|⟨k(t),t+1⟩|2−|⟨k(t),t−W+1⟩|2. v_k(t) R_t\, v_k(t)= | v_k(t), δ_t+1 |^2- | v_k(t), δ_t-W+1 |^2. (32) Proof. Apply Proposition 11.2 with =(t) C= C(t), =t R= R_t, λ~k=λk(t+1) λ_k= _k(t+1). Eq. (32) follows from expanding the quadratic form against the rank-two t R_t. ∎ Proposition 11.4 (First-order twist of the eigenvector). Under the assumptions of Corollary 11.3, let k(t+1) v_k(t+1) be the normalised eigenvector of (t+1) C(t+1) for λk(t+1) _k(t+1), with ⟨k(t+1),k(t)⟩>0 v_k(t+1), v_k(t) >0. Then k(t+1)−k(t)=∑j≠kj(t)⊤tk(t)λk(t)−λj(t)j(t)+(∥t∥2δk(t)2). v_k(t+1)- v_k(t)= _j≠ k v_j(t) R_t\, v_k(t) _k(t)- _j(t)\; v_j(t)+O\! ( R_t ^2 _k(t)^2 ). (33) Equivalently, using the rank-two form of t R_t: k(t+1)−k(t)=∑j≠k⟨j(t),t+1⟩⟨k(t),t+1⟩−⟨j(t),t−W+1⟩⟨k(t),t−W+1⟩λk(t)−λj(t)j(t)+(∥t∥2δk(t)2). v_k(t+1)- v_k(t)= _j≠ k v_j(t), δ_t+1 v_k(t), δ_t+1 - v_j(t), δ_t-W+1 v_k(t), δ_t-W+1 _k(t)- _j(t)\; v_j(t)+O\! ( R_t ^2 _k(t)^2 ). (34) Proof. Write k(t+1)=∑jajj(t) v_k(t+1)= _ja_j v_j(t). From the proof of Proposition 11.2, aj=j(t)⊤tk(t)/(λk(t)−λj(t))+(∥t∥2/δk(t)2)a_j= v_j(t) R_t\, v_k(t)/( _k(t)- _j(t))+O( R_t ^2/ _k(t)^2) for j≠kj≠ k, and ak=1+(∥t∥2/δk(t)2)a_k=1+O( R_t ^2/ _k(t)^2). Hence k(t+1)−k(t)=∑j≠kajj(t)+(∥t∥2/δk(t)2) v_k(t+1)- v_k(t)= _j≠ ka_j v_j(t)+O( R_t ^2/ _k(t)^2), giving (33). The explicit form follows from the rank-two structure of t R_t. ∎ Corollary 11.5 (Near-edge singular contribution to the gap increment). Define the spectral gap γk(t):=λk(t)−λk+1(t) _k(t):= _k(t)- _k+1(t). Assume both λk(t) _k(t) and λk+1(t) _k+1(t) are simple, and that t R_t is small compared with the neighbouring gaps. Then γk(t+1)−γk(t) _k(t+1)- _k(t) =k(t)⊤tk(t)−k+1(t)⊤tk+1(t) = v_k(t) R_t\, v_k(t)- v_k+1(t) R_t\, v_k+1(t) +2|k+1(t)⊤tk(t)|2γk(t)+(∥t∥2), + 2\,| v_k+1(t) R_t\, v_k(t)|^2 _k(t)+O\! ( R_t ^2 ), (35) where the (∥t∥2)O( R_t ^2) remainder is uniform away from the adjacent denominator γk(t) _k(t). More explicitly, k+1(t)⊤tk(t)=⟨k+1(t),t+1⟩⟨k(t),t+1⟩−⟨k+1(t),t−W+1⟩⟨k(t),t−W+1⟩. v_k+1(t) R_t\, v_k(t)= v_k+1(t), δ_t+1 v_k(t), δ_t+1 - v_k+1(t), δ_t-W+1 v_k(t), δ_t-W+1 . Proof. Apply Proposition 11.2 separately to λk _k and λk+1 _k+1, and subtract. The j=k+1j=k+1 term in the expansion of λk _k contributes |k+1⊤tk|2/γk| v_k+1 R_t\, v_k|^2/ _k; the j=kj=k term in the expansion of λk+1 _k+1 contributes |k⊤tk+1|2/(λk+1−λk)=−|k+1⊤tk|2/γk| v_k R_t\, v_k+1|^2/( _k+1- _k)=-| v_k+1 R_t\, v_k|^2/ _k. Subtracting yields 2|k+1⊤tk|2/γk2| v_k+1 R_t\, v_k|^2/ _k. All remaining second-order terms are non-singular with respect to γk _k and are absorbed into the remainder. ∎ Sign convention. The second-order correction may be written either as ∑j≠k|j⊤k|2/(λk−λj) _j≠ k| v_j R\, v_k|^2/( _k- _j) or as −∑j≠k|j⊤k|2/(λj−λk)- _j≠ k| v_j R\, v_k|^2/( _j- _k). These are identical. One should not write a minus sign together with the denominator λk−λj _k- _j. 12 Signal Strength Flow 12.1 The Gradient Flow Component Consider the loss landscape L()L( θ) with Hessian () H( θ). Near a point t θ_t, the gradient evolves: ∇L(t+1)≈∇L(t)+tt.∇ L( θ_t+1)≈∇ L( θ_t)+ H_t\, δ_t. Theorem 12.1 (Signal Strength Flow). By Corollary 11.3, the signal strengths satisfy the exact difference equation: dj2(t+1)−dj2(t)=η2|Gjeff(t)|2−η2|Gjeff(t−W)|2+2ηGjeff(t−W)εj(t)−εj(t)2, d_j^2(t+1)-d_j^2(t)\;=\;η^2 |G_j^eff(t) |^2-η^2 |G_j^eff(t-W) |^2+2η\,G_j^eff(t-W)\, _j(t)- _j(t)^2, (36) where the eigenvector rotation correction is (Propositions 11.4 and 42): εj(t)=∑s=t−W+1t∑i≠ji(s)⊤sj(s)dj2(s)−di2(s)⟨i(s),t−W⟩. _j(t)\;=\; _s=t-W+1^t\; _i≠ j v_i(s) R_s\, v_j(s)d_j^2(s)-d_i^2(s)\; v_i(s),\, δ_t-W . (37) The spectral gaps dj2−di2d_j^2-d_i^2 appear in the denominator: εj _j is O(W∥2/δj)O(W R ^2/ _j) when δj=mini≠j|dj2−di2| _j= _i≠ j|d_j^2-d_i^2| is bounded away from zero, but diverges as δj→0 _j→ 0. The three pieces are: • Entering signal: η2|Gjeff(t)|2η^2|G_j^eff(t)|^2 — new gradient projection squared. • Exiting signal: η2|Gjeff(t−W)|2η^2|G_j^eff(t-W)|^2 — gradient projection leaving the window. • Rotation correction: 2ηGjeff(t−W)εj−εj22η\,G_j^eff(t-W)\, _j- _j^2 — coupling between gradient and eigenvector drift over W steps. Diverges as δj→0 _j→ 0 (near phase transitions). Taylor-expanding the delay (Axiom 2.3) and dropping εj _j gives dj2≈η2W|Gjeff|2d_j^2≈η^2W|G_j^eff|^2 (eigenvalue tracks gradient projection). Under the additional assumption [,(t)]≈0[P, H(t)]≈ 0, the eigenvalue evolution decomposes as: 1. Eigenvalue ODE: ddj2dt≈−2η(hj+ω)dj2+η2W(Sj+2Gjeffj) d_j^2dt\;≈\;-2η(h_j+ω)\,d_j^2+η^2W (S_j+2\,G_j^eff\,N_j ) (38) where SjS_j is the exact (conservative) mode-coupling source and jN_j is the nonlinear gradient residual, both defined in the proof below. 2. Eigenvector equation (exact): ˙j=∑i≠ji⊤˙jdj2−di2i v_j\;=\; _i≠ j v_i C\, v_jd_j^2-d_i^2\; v_i (39) where: • Gjeff(t)=⟨j(t),∇Lt+ωt⟩G_j^eff(t)= v_j(t),P∇ L_t+ω θ_t is the current effective gradient projection. • hj(t)=j(t)⊤((t))j(t)h_j(t)= v_j(t) ( H(t)P) v_j(t) is the instantaneous Rayleigh quotient (for SGD, hj=j⊤jh_j= v_j H v_j). Both hjh_j and GjeffG_j^eff are time-dependent. • ˙=tt⊤−t−Wt−W⊤ C= δ_t δ_t - δ_t-W δ_t-W is the covariance update from sliding the window (Proposition 11.1). The ODE (38) has three terms: dissipation −2η(hj+ω)dj2-2η(h_j+ω)d_j^2 from curvature and weight decay, mode coupling η2WSjη^2W\,S_j (conservative: ∑jSj=0 _jS_j=0), and injection 2η2WGjj2η^2W\,G_j\,N_j from the nonlinear loss landscape. The steady-state signal strength (from the window sum with within-window projection decay at rate η(hj+ω)η(h_j+ω)): djss=η|Gjeff,s|Φj,Φj=1−e−2η(hj+ω)W1−e−2η(hj+ω).d_j^s=η\, |G_j^eff,s |\; _j, _j= 1-e^-2η(h_j+ω)W1-e^-2η(h_j+ω). (40) Two limits: Φj≈W _j≈ W when η(hj+ω)W≪1η(h_j+ω)W 1 (giving dj≈ηW|Gj|d_j≈η W\,|G_j|), and Φj≈1/(2η(hj+ω)) _j≈ 1/ (2η(h_j+ω) ) when η(hj+ω)W≫1η(h_j+ω)W 1 (giving dj≈|Gj|η/(2(hj+ω))d_j≈|G_j| η/ (2(h_j+ω) )). In both regimes, modes with larger |Gjeff|/hj+ω|G_j^eff|/ h_j+ω have larger signal strengths. Proof. The proof proceeds in three steps: the exact perturbation theory gives a difference equation (Step 1), Taylor-expanding the delay converts it to a differential equation (Step 2), and substituting the gradient evolution yields the working form (Step 3). Step 1 (Exact difference equation). By Corollary 11.3, the rank-two update t=tt⊤−t−Wt−W⊤ R_t= δ_t δ_t - δ_t-W δ_t-W (Proposition 11.1) gives an exact eigenvalue increment and an eigenvector rotation: onto j: v_j\!: dj2˙=|⟨t,j(t)⟩|2−|⟨t−W,j(t)⟩|2, d_j^2= | δ_t, v_j(t) |^2- | δ_t-W, v_j(t) |^2, (41) onto i: v_i\!: i⊤˙j=i⊤˙jdj2−di2. v_i v_j= v_i C v_jd_j^2-d_i^2. (42) The entering term is η2|Gjeff(t)|2η^2|G_j^eff(t)|^2 by definition of GjeffG_j^eff. Decompose the exiting term using the cumulative eigenvector drift (Proposition 11.4): ⟨j(t),t−W⟩=⟨j(t−W),t−W⟩⏟−ηGjeff(t−W)+⟨j(t)−j(t−W),t−W⟩⏟εj:rotation correction. v_j(t), δ_t-W = v_j(t-W), δ_t-W _-η\,G_j^eff(t-W)+ v_j(t)- v_j(t-W), δ_t-W _ _j:\;rotation correction. Using (42), the cumulative drift expands as j(t)−j(t−W)=∑s∑i≠ji(s)⊤sj(s)dj2(s)−di2(s)i(s) v_j(t)- v_j(t-W)= _s _i≠ j v_i(s) R_s v_j(s)d_j^2(s)-d_i^2(s) v_i(s), giving the explicit rotation correction (37) with spectral gaps dj2−di2d_j^2-d_i^2 in the denominator. Squaring and substituting into (41): dj2(t+1)−dj2(t)=η2|Gjeff(t)|2−η2|Gjeff(t−W)|2+2ηGjeff(t−W)εj−εj2.d_j^2(t+1)-d_j^2(t)\;=\;η^2 |G_j^eff(t) |^2-η^2 |G_j^eff(t-W) |^2+2η\,G_j^eff(t-W)\, _j- _j^2. (43) This is the exact equation (36). By the eigenvector twist bound, each step contributes O(∥/δj)O( R / _j) to the drift, so |εj|=O(W∥2/δj)| _j|=O(W R ^2/ _j) over W steps. Note: the rotation correction diverges as δj→0 _j→ 0 (near phase transitions), so it cannot be dropped when the spectral gap is closing. Step 2 (Taylor-expand the delay ⇒ differential equation). Away from phase transitions (δj _j bounded below), εj _j is small and we drop the rotation correction. Under Axiom 2.3, |Gjeff|2|G_j^eff|^2 varies slowly over W steps, so: |Gjeff(t)|2−|Gjeff(t−W)|2≈Wdt|Gjeff|2. |G_j^eff(t) |^2- |G_j^eff(t-W) |^2\;≈\;W\, ddt |G_j^eff |^2. Hence dj2˙≈η2Wdt|Gjeff|2. d_j^2\;≈\;η^2W\, ddt |G_j^eff |^2. (44) Integrating: dj2≈η2W|Gjeff|2+C0d_j^2≈η^2W\,|G_j^eff|^2+C_0, i.e., the eigenvalue tracks the sliding-window average of the squared gradient projection. Full form with rotation correction. If we Taylor-expand only the |Gj|2|G_j|^2 delay in (43) but retain the rotation correction εj _j: d˙j2≈η2Wdt|Gjeff|2+2ηGjeff(t−W)εj(t)−εj(t)2. d_j^2\;≈\;η^2W\, ddt |G_j^eff |^2+2η\,G_j^eff(t-W)\, _j(t)- _j(t)^2. (45) After substituting the gradient evolution (47) (Step 3 below): d˙j2=−2η(hj+ω)dj2+η2WSj+2ηGjeff(t−W)εj−εj2. d_j^2=-2η(h_j+ω)\,d_j^2+η^2W\,S_j+2η\,G_j^eff(t-W)\, _j- _j^2. (46) The first two terms are the smooth part (dissipation ++ mode coupling from the gradient evolution). The last two terms are the discrete rotation correction, which diverges as δj→0 _j→ 0 (Corollary 11.5) and dominates near phase transitions. Eq. (44) is recovered by setting εj=0 _j=0; the coupled system (53) follows from further substituting the explicit SjS_j (50). Step 3 (Gradient evolution ⇒ two-term form; requires [,H]≈0[P, H]≈ 0). Write =∇L+ω f=P∇ L+ω θ for the effective gradient, so Gjeff=⟨j,⟩G_j^eff= v_j, f . The linearised gradient flow is ˙=−η(+ω) f=-η(P H+ω I) f. Differentiating GjeffG_j^eff: G˙jeff=⟨j,˙⟩+⟨˙j,⟩. G_j^eff\;=\; v_j, f + v_j, f . For the first term, [,]≈0[P, H]≈ 0 gives ⟨j,(+ω)⟩=(hj+ω)Gjeff v_j,(P H+ω I) f =(h_j+ω)\,G_j^eff (no off-diagonal mixing), so ⟨j,˙⟩=−η(hj+ω)Gjeff v_j, f =-η(h_j+ω)\,G_j^eff. For the second term, the eigenvector rotation (39) gives ⟨˙j,⟩=∑i≠ji⊤˙jdj2−di2Gieff v_j, f = _i≠ j v_i C\, v_jd_j^2-d_i^2\,G_i^eff. Hence G˙jeff=−η(hj+ω)Gjeff+∑i≠ji⊤˙jdj2−di2Gieff G_j^eff=-η(h_j+ω)\,G_j^eff+ _i≠ j v_i C\, v_jd_j^2-d_i^2\,G_i^eff. Applying the chain rule dt|Gjeff|2=2GjeffG˙jeff ddt|G_j^eff|^2=2\,G_j^eff\, G_j^eff: dt|Gjeff|2=−2η(hj+ω)|Gjeff|2+Sj(t), ddt |G_j^eff |^2\;=\;-2η(h_j+ω)\, |G_j^eff |^2\;+\;S_j(t), (47) where the source term (mode coupling) is: Sj(t)= 2Gjeff∑i≠ji⊤˙jdj2−di2Gieff.S_j(t)\;=\;2\,G_j^eff _i≠ j v_i C\, v_jd_j^2-d_i^2\;G_i^eff. (48) Substituting (47) into (44) and using dj2≈η2W|Gjeff|2d_j^2≈η^2W|G_j^eff|^2: dj2˙≈−2η(hj+ω)dj2+η2WSj(t). d_j^2≈-2η(h_j+ω)\,d_j^2+η^2W\,S_j(t). The first term is dissipation: curvature and weight decay drain the signal. The second is injection: the loss landscape continually drives the mode. The source SjS_j is not computed exactly here. Its precise form requires tracking mode-coupling integrals from (39), which depend on the off-diagonal terms i⊤˙j v_i C v_j. Combining the dissipation −2η(hj+ω)|Gj|2-2η(h_j+ω)|G_j|^2, the mode coupling SjS_j, and the nonlinear residual 2Gjj2G_jN_j: d˙j2≈−2η(hj+ω)dj2+η2W(Sj+2Gjeffj). d_j^2\;≈\;-2η(h_j+ω)\,d_j^2+η^2W (S_j+2G_j^eff\,N_j ). This is (38), where SjS_j is the exact source (48) and jN_j is the measurable nonlinear residual (57). The steady state (40) dj=η|Gjeff,s|Φjd_j=η\,|G_j^eff,s| _j separates modes by their driving-to-curvature ratio through Φj _j. ∎ Proposition 12.2 (Off-diagonal covariance coupling). Under slow variation (Axiom 2.3) and away from phase transitions (δj _j bounded below), the off-diagonal coupling in the covariance update is, for i≠ji≠ j: i⊤˙j≈−η3W[(hi+ω)+(hj+ω)]GieffGjeff. v_i C\, v_j\;≈\;-η^3W [(h_i+ω)+(h_j+ω) ]\,G_i^eff\,G_j^eff. (49) Proof. The covariance update (Proposition 11.1) gives i⊤˙j=⟨i,t⟩⟨j,t⟩−⟨i,t−W⟩⟨j,t−W⟩. v_i C\, v_j= v_i, δ_t v_j, δ_t - v_i, δ_t-W v_j, δ_t-W . Since ⟨k(t),t⟩=−ηGkeff(t) v_k(t), δ_t =-η\,G_k^eff(t) and, under slow eigenvector rotation (εk _k small), ⟨k(t),t−W⟩≈−ηGkeff(t−W) v_k(t), δ_t-W ≈-η\,G_k^eff(t-W): i⊤˙j≈η2[Gieff(t)Gjeff(t)−Gieff(t−W)Gjeff(t−W)]. v_i C\, v_j\;≈\;η^2 [G_i^eff(t)\,G_j^eff(t)-G_i^eff(t-W)\,G_j^eff(t-W) ]. Taylor-expanding the delay (Axiom 2.3): Gi(t)Gj(t)−Gi(t−W)Gj(t−W)≈Wdt(GiGj)G_i(t)\,G_j(t)-G_i(t-W)\,G_j(t-W)≈ W\, ddt(G_i\,G_j). At leading order, dropping the eigenvector-rotation contribution to G˙keff G_k^eff and using G˙keff≈−η(hk+ω)Gkeff G_k^eff≈-η(h_k+ω)\,G_k^eff: dt(GiGj)=−η[(hi+ω)+(hj+ω)]GiGj, ddt (G_i\,G_j )=-η [(h_i+ω)+(h_j+ω) ]\,G_i\,G_j, which gives (49). ∎ Proposition 12.3 (Source term evaluation (approximate)). Under the three approximations of Proposition 12.2 (slow eigenvector rotation, Taylor-expanded delay, leading-order gradient decay), (49) substituted into the source term (48) gives: Sj≈−2η3W|Gjeff|2∑i≠j(hi+ω)+(hj+ω)dj2−di2|Gieff|2.S_j\;≈\;-2η^3W\, |G_j^eff |^2 _i≠ j (h_i+ω)+(h_j+ω)d_j^2-d_i^2\; |G_i^eff |^2. (50) In the quasi-steady regime (all |Gk||G_k| decaying), every term in the sum is positive for the dominant mode (d1>did_1>d_i), so S1<0S_1<0: the dominant mode loses signal via mode coupling. For subdominant modes, SjS_j can be positive (signal transferred from dominant modes). During rapid learning (a mode’s |Gj||G_j| growing), i⊤˙j v_i C\, v_j can have either sign, and the conclusion S1<0S_1<0 may fail. The exact SjS_j (48) remains well-defined in all regimes. Theorem 12.4 (Mode-coupling conservation). The source terms satisfy ∑jSj= 0. _jS_j\;=\;0. (51) Mode coupling redistributes signal between modes but does not inject or remove total signal. Proof. Write ∑jSj=2∑j∑i≠jGieffGjeff(i⊤˙j)/(dj2−di2) _jS_j=2 _j _i≠ jG_i^eff\,G_j^eff\,( v_i C\, v_j)\,/\,(d_j^2-d_i^2). Swap the labels i↔ji j. Symmetry of ˙ C gives j⊤˙i=i⊤˙j v_j C\, v_i= v_i C\, v_j, while 1/(di2−dj2)=−1/(dj2−di2)1/(d_i^2-d_j^2)=-1/(d_j^2-d_i^2). Hence the swapped sum equals the negative of the original, so ∑jSj=0 _jS_j=0. ∎ Remark 12.5 (Exact vs. approximate hierarchy). The proof above uses the exact SjS_j (48): the antisymmetry of 1/(dj2−di2)1/(d_j^2-d_i^2) holds for any symmetric ˙ C, with no approximation beyond [,]≈0[P, H]≈ 0 (needed only to decompose G˙j G_j in Step 3). Likewise Corollary 12.6 below uses only ∑jSj=0 _jS_j=0 and is exact. The approximate evaluation of SjS_j in (50) introduces three additional approximations (slow eigenvector rotation, Taylor-expanded delay, leading-order G˙k G_k). The coupled system (53) and the sign claim S1<0S_1<0 both depend on this evaluation. For numerical work, SjS_j should be computed from the exact formula (48) using the observed i⊤˙j v_i C v_j from the training trajectory, without the approximations of Proposition 12.2. Corollary 12.6 (Total signal dissipation). In the linearised gradient model ˙=−η(+ω) f=-η(P H+ω I) f, the total signal decays monotonically: dt∑jdj2=−2η∑j(hj+ω)dj2. ddt _jd_j^2\;=\;-2η _j(h_j+ω)\,d_j^2. (52) Any growth in an individual djd_j comes from redistribution by SjS_j, not from net injection. Proof. Sum d˙j2=−2η(hj+ω)dj2+η2WSj d_j^2=-2η(h_j+ω)\,d_j^2+η^2W\,S_j over j and apply (51). ∎ Corollary 12.7 (Coupled eigenvalue system). Substituting (50) into the eigenvalue equation and using |Gkeff|2≈dk2/(η2W)|G_k^eff|^2≈ d_k^2/(η^2W) from Step 2: d˙j2=−2ηdj2[(hj+ω)+∑i≠j(hi+ω)+(hj+ω)dj2−di2di2]. d_j^2\;=\;-2η\,d_j^2\, [(h_j+ω)+ _i≠ j (h_i+ω)+(h_j+ω)d_j^2-d_i^2\;d_i^2 ]. (53) This is a closed system of coupled ODEs for the eigenvalues dj2\d_j^2\, with the mode-coupling terms producing level repulsion: when dj≈did_j≈ d_i, the denominator dj2−di2→0d_j^2-d_i^2→ 0 creates a divergent interaction that pushes eigenvalues apart. Remark 12.8 (The injection and the linearised model). Corollary 12.6 shows that total signal decays in the linearised model. Where, then, does the “injection” in the phenomenological ODE (38) come from? The linearised gradient dynamics ˙=−η(+ω) f=-η(P H+ω I) f models gradient decay: each mode’s gradient projection decreases exponentially. In reality, the loss landscape is nonlinear and continuously produces new gradient as the model encounters data it has not yet learned. This nonlinear gradient replenishment maintains |Gjeff|>0|G_j^eff|>0 even as learning progresses, and is the physical mechanism behind the injection term. Three processes govern each eigenvalue: 1. Dissipation: gradient decay at rate η(hj+ω)η(h_j+ω) — rigorous, from the linearised model. 2. Redistribution: mode coupling transfers signal between modes (∑jSj=0 _jS_j=0) — rigorous, from Theorem 12.4. 3. Replenishment: the nonlinear loss landscape regenerates gradient — not captured by the linearised model. The phenomenological ODE (38) captures (1) and (3) but not (2). The coupled system (53) captures (1) and (2) but not (3). Neither alone is complete. Remark 12.9 (Commutator dependence). Step 3 requires [,]≈0[P, H]≈ 0 so that hjh_j is well-defined. Steps 1–2 (the exact difference equation and its Taylor expansion) hold without this assumption. • SGD (=P= I): exact. • Adam: =diag(1/v^)P=diag(1/ v) is diagonal; the commutator is small when H is approximately diagonal in the parameter basis. • Muon: P is the Newton–Schulz orthogonalizer; [,][P, H] is not generally small and hjh_j acquires a correction from the anti-symmetric part. Remark 12.10 (Status of the eigenvalue equation). The eigenvalue equation has three tiers of rigour: 1. Exact (delay equation (36)): no approximation beyond the rank-two covariance update. Gives dj∝|Gjeff|d_j |G_j^eff| at leading order. 2. Exact mode coupling (eq. (46) with exact SjS_j): d˙j2=−2η(hj+ω)dj2+η2WSj+2ηGj(t−W)εj−εj2. d_j^2=-2η(h_j+ω)\,d_j^2+η^2W\,S_j+2η\,G_j(t-W)\, _j- _j^2. SjS_j (48) is a known quantity at each time step, computable from the observed i⊤˙j v_i C v_j. The conservation ∑jSj=0 _jS_j=0 is exact. This equation predicts ∑jdj2→0 _jd_j^2→ 0 (Corollary 12.6): total signal decays in the linearised model. The missing piece is the anharmonic jN_j (Section 12.2). 3. Phenomenological closure (ODE (38)): replaces both the exact SjS_j and the unknown jN_j with a single effective injection ηW|Gj|2/djη W|G_j|^2/d_j. This discards the known mode-coupling structure and treats each mode independently. It is a convenience, not a necessity: all downstream results (gap flow, steady state, phase transitions) can be derived from tiers 1–2 alone. 12.2 The Anharmonic Replenishment Corollary 12.6 shows that the linearised model cannot sustain nonzero eigenvalues. We now derive the leading nonlinear correction and identify the formal source of injection. Proposition 12.11 (Second-order gradient dynamics). The exact single-step effective gradient update is t+1=t−η(t+ω)t+12η2(∇3Lt)[t,t]+O(η3), f_t+1= f_t-η(P H_t+ω I)\, f_t+ 12η^2\,P\,(∇^3L_t)[ f_t, f_t]+O(η^3), (54) where =∇L+ω f=P∇ L+ω θ is the effective gradient and (∇3L)[,]a=∑bc(∂3L/∂θa∂θb∂θc)ubwc(∇^3L)[ u, w]_a= _bc(∂^3L/∂ _a∂ _b∂ _c)\,u_b\,w_c is the third derivative (anharmonic) tensor contracted with two copies of f. Proof. Taylor-expand ∇L(t+1)∇ L( θ_t+1) around t θ_t with t+1−t=−ηt θ_t+1- θ_t=-η f_t: ∇L(t+1)=∇Lt−ηtt+12η2(∇3Lt)[t,t]+O(η3).∇ L( θ_t+1)=∇ L_t-η\, H_t\, f_t+ 12η^2\,(∇^3L_t)[ f_t, f_t]+O(η^3). Apply P and add ωt+1=ωt−ηωtω θ_t+1=ω θ_t-ηω f_t; collect terms to get (54). ∎ Definition 12.12 (Anharmonic coupling tensor). Define Tjkℓ=⟨j,(∇3L)[k,ℓ]⟩T_jk = v_j,\;P\,(∇^3L)[ v_k, v_ ] . This quantifies how learning along directions k and ℓ creates new gradient in direction j through the curvature of the loss landscape. Proposition 12.13 (Gradient evolution with anharmonic term). Projecting (54) onto j v_j and passing to continuous time: G˙jeff=−η(hj+ω)Gjeff⏟dissipation+∑i≠ji⊤˙jdj2−di2Gieff⏟mode coupling (conservative)+j⏟anharmonic, G_j^eff= -η(h_j+ω)\,G_j^eff_dissipation+ _i≠ j v_i C\, v_jd_j^2-d_i^2\,G_i^eff_mode coupling (conservative)+ N_j_anharmonic, (55) where j=12η2∑k,ℓGkeffGℓeffTjkℓN_j= 12η^2 _k, G_k^eff\,G_ ^eff\;T_jk (56) is the nonlinear replenishment rate. In the linearised model (∇3L=0∇^3L=0), j=0N_j=0 and only dissipation and conservative mode coupling remain. Remark 12.14 (jN_j is a measurable residual). Eq. (55) can be rearranged: j=G˙jeff+η(hj+ω)Gjeff−∑i≠ji⊤˙jdj2−di2Gieff.N_j\;=\; G_j^eff+η(h_j+ω)\,G_j^eff- _i≠ j v_i C\, v_jd_j^2-d_i^2\,G_i^eff. (57) Every quantity on the right is computable from the training trajectory: G˙j G_j from consecutive time steps, hjh_j from the Rayleigh quotient, and i⊤˙j v_i C v_j from the rank-two update. No knowledge of ∇3L∇^3L is required. Thus jN_j is not “unknown” in practice — it is the nonlinear residual of the gradient evolution, measurable at each step. The perturbative expansion (56) via TjkℓT_jk is an analytical model for this residual; the residual itself is always available. The eigenvalue equation (58) with measured SjS_j and measured jN_j is a complete, closed description of the spectral dynamics, with no phenomenological terms. Corollary 12.15 (Eigenvalue equation with injection). Applying the chain rule and substituting into the eigenvalue equation from Step 2: d˙j2=−2η(hj+ω)dj2+η2WSj+2η2WGjeffj. d_j^2=-2η(h_j+ω)\,d_j^2+η^2W\,S_j+2η^2W\,G_j^eff\,N_j. (58) The three terms are: • Dissipation: −2η(hj+ω)dj2-2η(h_j+ω)\,d_j^2 (rigorous). • Mode coupling: η2WSjη^2W\,S_j (conservative: ∑jSj=0 _jS_j=0, rigorous). • Anharmonic injection: 2η2WGjeffj2η^2W\,G_j^eff\,N_j — the only term that can produce net growth of total signal. Proposition 12.16 (Steady-state balance). At any steady state (d˙j=0 d_j=0), the anharmonic injection must compensate dissipation and mode-coupling losses: 2η2WGjeffj=2η(hj+ω)dj2−η2WSj.2η^2W\,G_j^eff\,N_j=2η(h_j+ω)\,d_j^2-η^2W\,S_j. (59) For the dominant mode, S1<0S_1<0 (Proposition 12.3), so both terms on the right-hand side are positive. The anharmonic coupling must satisfy G11>0G_1\,N_1>0: the loss landscape’s nonlinearity must inject signal into the dominant mode to sustain it. Remark 12.17 (Self-interaction and the edge of stability). The dominant self-interaction is jself=12η2(Gjeff)2TjjjN_j^self= 12η^2(G_j^eff)^2\,T_j, where Tjjj=⟨j,(∇3L)[j,j]⟩T_j= v_j,\,P(∇^3L)[ v_j, v_j] measures the rate at which curvature changes as the model moves along direction j v_j. At the edge of stability (Section 40), the top Hessian eigenvalue self-corrects at 2/η2/η. This constrains TjjjT_j for the dominant mode: T111T_111 is effectively regulated by the edge-of-stability dynamics, which prevents the dominant eigenvalue from overshooting. Subdominant modes are not constrained by this mechanism; their TjjjT_j can be large, driving the phase transitions that create new active modes. Remark 12.18 (The phenomenological ODE as a crude closure). The primary eigenvalue equation (58) is d˙j2=−2η(hj+ω)dj2⏟rigorous+η2WSj⏟exact (known)+2η2WGjj⏟unknown. d_j^2= -2η(h_j+ω)\,d_j^2_rigorous+ η^2W\,S_j_exact (known)+ 2η^2W\,G_j\,N_j_unknown. Only the anharmonic jN_j is genuinely unknown. The exact SjS_j (48) is computable from any training trajectory and provides the full mode-coupling structure (level repulsion, conservation, redistribution). The phenomenological ODE (38) replaces both the exact SjS_j and the unknown jN_j with a single effective injection ηW|Gj|2/djη W|G_j|^2/d_j. This is a crude closure: it discards the known mode coupling and treats each mode independently. Its virtues are simplicity and the correct 1/dj1/d_j singularity at mode onset, but it sacrifices the inter-mode structure that the exact equation preserves. The phenomenological ODE is used downstream only for convenience: the steady-state hierarchy (40) and the gap flow (68) can both be derived from the delay equation (36) without it (see the first halves of their respective proofs). All qualitative conclusions (level repulsion, phase transitions, gap dynamics) follow from the exact SjS_j and Corollary 11.5 alone. Remark 12.19 (Spectral gap structure of the injection). The resolvent 1/(dj2−di2)1/(d_j^2-d_i^2) appears in the eigenvector rotation (39), the source term (50), and the coupled system (53). The anharmonic tensor TjkℓT_jk , being a direct projection of ∇3L∇^3L onto the eigenbasis, does not carry this resolvent. At steady state, however, the spectral gap re-enters through self-consistency. The balance (59) forces: jss=ηGjeff[(hj+ω)+∑i≠j(hi+ω)+(hj+ω)dj2−di2di2].N_j^s=η\,G_j^eff\, [(h_j+ω)+ _i≠ j (h_i+ω)+(h_j+ω)d_j^2-d_i^2\;d_i^2 ]. (60) The resolvent is restored: the equilibrium injection rate into mode j is enhanced precisely when j is near a level crossing (dj≈did_j≈ d_i). The physical picture separates cleanly: • TjkℓT_jk provides the fuel (no gap — injects into all modes via ∇3L∇^3L). • Mode coupling SjS_j provides the channel (gap in denominator — redistributes preferentially near level crossings). • Self-consistency locks them together: the steady-state injection (60) inherits the resolvent from the redistribution balance. A second route to the gap passes through the NTK: the evolving-NTK mixing rate Γjk∝k⊤˙j/(λj−λk) _jk q_k K\, q_j\,/\,( _j- _k) carries the NTK spectral gap λj−λk=N(hj−hk) _j- _k=N(h_j-h_k) in the denominator. Since ˙ K is driven by ∇3L∇^3L, this is the NTK-space manifestation of the same anharmonic coupling, dressed by the NTK resolvent rather than the covariance resolvent. 12.3 Simplified Signal Flow Corollary 12.20 (Steady-State Signal Hierarchy). From the window-sum formula (40): djss≈η|Gjeff,s|Φj,Φj=1−e−2η(hj+ω)W1−e−2η(hj+ω).d_j^s\;≈\;η\,|G_j^eff,s|\; _j, _j= 1-e^-2η(h_j+ω)W1-e^-2η(h_j+ω). (61) Two limits: • η(hj+ω)W≪1η(h_j+ω)W 1 (weak curvature): Φj≈W _j≈ W, dj≈ηW|Gjeff|d_j≈η W\,|G_j^eff| (all window entries nearly equal). • η(hj+ω)W≫1η(h_j+ω)W 1 and η(hj+ω)≪1η(h_j+ω) 1 (intermediate curvature: window long relative to decay, but single-step decay still small): Φj≈1/(2η(hj+ω)) _j≈ 1/ (2η(h_j+ω) ), dj≈|Gjeff|η/(2(hj+ω))d_j≈|G_j^eff| η/ (2(h_j+ω) ) (window dominated by newest entries). If η(hj+ω)≳1η(h_j+ω) 1, then Φj≈1 _j≈ 1 (single step dominates). In both regimes, modes with larger |Gjeff|/hj+ω|G_j^eff|/ h_j+ω have larger signal strengths: the driving-to-curvature ratio orders the signal hierarchy. Remark 12.21 (Physical Interpretation). The steady-state signal strength djssd_j^s is determined by the balance between gradient projection (GjeffG_j^eff) and curvature-modulated window averaging (Φj _j). The spectral gap at steady state is: gss=η|Gk∗eff|Φk∗−η|Gk∗+1eff|Φk∗+1.g^s=η\, G_k^*^eff _k^*-η\, G_k^*+1^eff _k^*+1. (62) The gap vanishes when the “driving-to-curvature ratio” |Gjeff|/hj+ω|G_j^eff|/ h_j+ω is the same for the modes on either side. Weight decay ω acts as a curvature floor: modes with hj≪ωh_j ω all get the same effective curvature ω, so their signal strengths are compressed to a common scale djss≈ηW|Gjeff|d_j^s≈η W\,|G_j^eff| (since Φj≈W _j≈ W when ηωW≪1ηω W 1). Remark 12.22 (Grokking Condition). For grokking to occur, the generalizing mode must survive while memorization modes are suppressed. Using hj=λk(j)/Nh_j= _k(j)/N (Proposition 31.3), this requires: λgenN>ω>λmemN. _genN>ω> _memN. (63) Weight decay must sit between the NTK eigenvalues of the generalizing and memorization modes. Too small: memorization persists (no gap opens). Too large: generalization is also suppressed. When ω=0ω=0: all low-curvature modes persist at dj∝1/hj→∞d_j 1/ h_j→∞, the generalizing direction is buried, and grokking never occurs. This is consistent with the universal finding of 24/24 grokking with ω>0ω>0 and 0/24 without (Section 27). 13 Noise Level Dynamics Although the noise is negligible for determining k∗k^*, the noise variance ν2(t)ν^2(t) still evolves and affects higher-order corrections. Theorem 13.1 (Noise Variance Flow). The per-coordinate noise variance ν2(t)ν^2(t) evolves as: dν2dt=η2pTr(tgrad(t)t)−2ηων2, dν^2dt\;=\; η^2pTr (P_t _grad(t)P_t )-2ηω\,ν^2, (64) where grad _grad is the gradient covariance and ω is the weight decay coefficient. Proof. The noise variance is ν2=1p∑j>k∗dj2≈1p(Tr()−∑j=1k∗dj2)ν^2= 1p _j>k^*d_j^2≈ 1p (Tr( C)- _j=1^k^*d_j^2 ), i.e., the per-coordinate energy in the non-signal subspace (here (t) C(t) is the p×p× p sliding-window covariance, Definition 1.5). Injection. When the window slides, the entering update t=−η(∇Lt+ωt) δ_t=-η(P∇ L_t+ω θ_t) adds tt⊤ δ_t δ_t to C (Proposition 11.1). The contribution to the trace is Tr(tt⊤)=|t|2=η2|∇Lt+ωt|2Tr( δ_t δ_t )=| δ_t|^2=η^2|P∇ L_t+ω θ_t|^2. Taking the expectation over the mini-batch noise: [|t|2]=η2Tr(grad)+η2|∇¯L+ω|2E[| δ_t|^2]=η^2Tr(P _gradP)+η^2|P ∇L+ω θ|^2, where ∇¯L ∇L is the full-batch gradient. The second term is the signal (absorbed into the top k∗k^* eigenvalues); the first is the noise injection, contributing η2Tr(grad)/pη^2Tr(P _gradP)/p per coordinate. Decay. Weight decay acts on the parameters as t+1=(1−ηω)t+⋯ θ_t+1=(1-ηω) θ_t+·s. The accumulated updates in the window inherit this decay: each entry’s squared norm shrinks by factor (1−ηω)2≈1−2ηω(1-ηω)^2≈ 1-2ηω per step, so ν˙2|decay=−2ηων2 ν^2|_decay=-2ηω\,ν^2. Combining gives (64). ∎ Corollary 13.2 (Noise Steady State). At equilibrium: νs2=η2ωpTr(grad).ν^2_s= η2ω pTr(P _gradP). (65) The noise level is proportional to η/ωη/ω and to the preconditioned gradient variance. For Adam, νs2≈η/(2ω)ν^2_s≈η/(2ω) (approximately constant per coordinate). 14 The Coupled System: Gap Flow 14.1 The Complete Dynamical System The delay equation (43) for all j=1,…,Wj=1,…,W gives the primary dynamical system: dj2(t+1)−dj2(t)=η2|Gjeff(t)|2−η2|Gjeff(t−W)|2+2ηGjeff(t−W)εj−εj2(j=1,…,W),εj(t)=∑s=t−W+1t∑i≠ji(s)⊤sj(s)dj2(s)−di2(s)⟨i(s),t−W⟩,k∗(t)=argmaxjdj(t)dj+1(t),g(t)=dk∗(t)−dk∗+1(t). aligned d_j^2(t+1)-d_j^2(t)&=η^2 |G_j^eff(t) |^2-η^2 |G_j^eff(t-W) |^2+2η\,G_j^eff(t-W)\, _j- _j^2 (j=1,…,W),\\[6.0pt] _j(t)&= _s=t-W+1^t\; _i≠ j v_i(s) R_s\, v_j(s)d_j^2(s)-d_i^2(s)\; v_i(s),\, δ_t-W ,\\[6.0pt] k^*(t)&= *arg\,max_j d_j(t)d_j+1(t),\\[6.0pt] g(t)&=d_k^*(t)-d_k^*+1(t). aligned (66) Under the Taylor expansion (Step 2) and gradient evolution (Step 3 of Theorem 12.1), this reduces to the working-form ODE system: ddj2dt≈−2η(hj+ω)dj2+η2W(Sj+2Gjeffj)(j=1,…,W),k∗(t)=argmaxjdj(t)dj+1(t),g(t)=dk∗(t)−dk∗+1(t). aligned d_j^2dt&≈-2η(h_j+ω)\,d_j^2+η^2W (S_j+2\,G_j^eff\,N_j ) (j=1,…,W),\\[6.0pt] k^*(t)&= *arg\,max_j d_j(t)d_j+1(t),\\[6.0pt] g(t)&=d_k^*(t)-d_k^*+1(t). aligned (67) Both systems are coupled through the eigenvector equation (39) (which rotates j v_j and thereby changes hjh_j and GjeffG_j^eff), with k∗k^* and g as derived quantities. Phase transitions occur when g(t)g(t) passes through zero. Remark 14.1 (Comparison with BBP System). The classical (BBP-based) system had k+1k+1 equations (k signal + 1 noise) with an external threshold dcritd_crit: gBBP(t)=dk∗(t)−ν(t)(p(W−1))1/4⏟dcrit(t).g_BBP(t)=d_k^*(t)- ν(t)(p(W-1))^1/4_d_crit(t). Our system has W equations (all signal) with no external threshold. The gap is internal: g(t)=dk∗(t)−dk∗+1(t)g(t)=d_k^*(t)-d_k^*+1(t). The dynamics are richer because all W modes interact through the Hessian and gradient coupling. 14.2 The Gap Flow Equation Theorem 14.2 (Gap Flow). Under [,(t)]≈0[P, H(t)]≈ 0 and Axiom 2.3, the spectral gap g(t)=dk∗(t)−dk∗+1(t)g(t)=d_k^*(t)-d_k^*+1(t) evolves as: dgdt≈−η(hk∗−hk∗+1)⋅d¯−η(h¯+ω)⋅g+ηW(|Gk∗eff|2dk∗−|Gk∗+1eff|2dk∗+1) dgdt≈-η(h_k^*-h_k^*+1)· d-η( h+ω)· g+η W\! ( G_k^*^eff ^2d_k^*- G_k^*+1^eff ^2d_k^*+1 ) (68) where d¯=(dk∗+dk∗+1)/2 d=(d_k^*+d_k^*+1)/2 and h¯=(hk∗+hk∗+1)/2 h=(h_k^*+h_k^*+1)/2. The three terms have distinct physical origins: 1. Curvature asymmetry −η(hk∗−hk∗+1)d¯-η(h_k^*-h_k^*+1) d: higher curvature ⇒ faster gradient decay ⇒ drives the gap closed. 2. Gap damping −η(h¯+ω)⋅g-η( h+ω)· g: average curvature plus weight decay damps the gap toward zero. 3. Driving asymmetry ηW(|Gk∗|2/dk∗−|Gk∗+1|2/dk∗+1)η W( G_k^* ^2/d_k^*- G_k^*+1 ^2/d_k^*+1): if the gradient projects more strongly onto k∗ v_k^*, this drives the gap open. Terms 1–2 (dissipation) can only close the gap. Term 3 (injection) can open it. Proof. From the delay equation. Apply the delay equation (43) to the adjacent pair k∗k^* and k∗+1k^*+1 and subtract: Δ(dk∗2)−Δ(dk∗+12)=η2(|Gk∗eff(t)|2−|Gk∗+1eff(t)|2)−η2(|Gk∗eff(t−W)|2−|Gk∗+1eff(t−W)|2). (d_k^*^2)- (d_k^*+1^2)=η^2 (|G_k^*^eff(t)|^2-|G_k^*+1^eff(t)|^2 )-η^2 (|G_k^*^eff(t-W)|^2-|G_k^*+1^eff(t-W)|^2 ). Taylor-expanding the delay: Δg(delay)≈η2Wdt(|Gk∗eff|2−|Gk∗+1eff|2) g^(delay)≈η^2W\, ddt (|G_k^*^eff|^2-|G_k^*+1^eff|^2 ) (after dividing by 2d¯2 d to convert d2˙→d˙ d^2→ d). Substituting the gradient evolution (47) for each mode reproduces the three-term structure. Alternatively, from the ODE. Subtract (38) for mode k∗+1k^*+1 from that for mode k∗k^*. In dj2d_j^2 form, the dissipation difference is −2η(h¯+ω)(dk∗2−dk∗+12)−2ηΔhd2¯-2η( h+ω)(d_k^*^2-d_k^*+1^2)-2η h\, d^2, and the injection difference is η2W[(Sk∗+2Gk∗k∗)−(Sk∗+1+2Gk∗+1k∗+1)]η^2W [(S_k^*+2G_k^*N_k^*)-(S_k^*+1+2G_k^*+1N_k^*+1) ]. Converting to g˙ g via dk∗2−dk∗+12≈2d¯gd_k^*^2-d_k^*+1^2≈ 2 d\,g reproduces (68). ∎ Remark 14.3 (Level repulsion near g=0g=0). By Corollary 11.5, the second-order eigenvalue perturbation contributes an exact repulsive term to the gap increment: 2|k∗+1⊤tk∗|2/γk∗2| v_k^*+1 R_t v_k^*|^2/ _k^*. Converting from λ=d2λ=d^2 to d (with γk∗≈2d¯g _k^*≈ 2 d\,g) gives: g˙rep=|γ|22d¯2g,γ=k∗⊤˙k∗+1. g_rep= |γ|^22 d^2\,g, γ= v_k^* C\, v_k^*+1. (69) This term is always positive, scales as 1/g1/g, and diverges as g→0g→ 0—preventing true eigenvalue crossings. It is the mechanism behind avoided crossings (Proposition 10.9): even when the three-term flow drives g toward zero, the level repulsion creates a minimum gap gmin>0g_ >0. The full gap dynamics near g=0g=0 are therefore: g˙≈−η(h¯+ω)g+c+|γ|22d¯2g, g≈-η( h+ω)\,g+c+ |γ|^22 d^2\,g, (70) where c=η(hk∗+1−hk∗)d¯+ηW(|Gk∗|2−|Gk∗+1|2)/d¯c=η(h_k^*+1-h_k^*) d+η W( G_k^* ^2- G_k^*+1 ^2)/ d. Setting g˙=0 g=0 gives the minimum gap during an avoided crossing (c<0c<0): gmin≈|γ|22d¯2|c|(for |c| large).g_ ≈ |γ|^22 d^2\,|c| (for $|c|$ large). (71) 14.3 Critical Dynamics Near g=0g=0 Proposition 14.4 (Critical Dynamics (Heuristic)). Under [,(t)]≈0[P, H(t)]≈ 0 (required for the working-form ODE; the delay system (66) and level repulsion from Corollary 11.5 hold without this). Near the gap collapse point g→0g→ 0, the gap flow (68) simplifies to: dgdt≈−η(h¯+ω)⋅g+c, dgdt\;≈\;-η( h+ω)· g+c, (72) where c=η(hk∗+1−hk∗)d¯+ηW(|Gk∗|2−|Gk∗+1|2)/d¯c=η(h_k^*+1-h_k^*) d+η W( G_k^* ^2- G_k^*+1 ^2)/ d collects the terms that are approximately constant near g=0g=0. Two regimes: 1. Viable gap (c>0c>0): The gap is attracted to a positive equilibrium g∗=c/(η(h¯+ω))g^*=c/(η( h+ω)). The dominant subspace is stable. 2. Collapsing gap (c<0c<0): The gap shrinks exponentially at rate η(h¯+ω)η( h+ω). The three-term flow predicts g→0g→ 0, but level repulsion (Remark 14.3) prevents true crossing, creating a minimum gap gmin≈|γ|2/(2d¯2|c|)g_ ≈|γ|^2/(2 d^2|c|). This is an avoided crossing (cf. Proposition 10.9). The spectral edge phase transition occurs when c changes sign: (hk∗+1−hk∗)⋅d¯2=W(|Gk∗+1|2−|Gk∗|2).(h_k^*+1-h_k^*)· d^2=W ( G_k^*+1 ^2- G_k^* ^2 ). (73) The curvature asymmetry balances the gradient alignment asymmetry. 14.4 Timescales Corollary 14.5 (Gap Collapse Time). (Under [,]≈0[P, H]≈ 0.) When the gap is collapsing (c<0c<0), the time from initial gap g0g_0 to collapse below resolution ε is: tcollapse∼1ηh¯logg0ε.t_collapse 1η h g_0 . (74) Modes with higher average curvature h¯ h collapse faster. Corollary 14.6 (Gap Opening Time). (Under [,]≈0[P, H]≈ 0.) When a gap opens (c>0c>0, starting from g=0g=0), the time to reach a detectable gap g0g_0 is: topen∼1ηh¯logg∗g∗−g0,t_open 1η h g^*g^*-g_0, (75) where g∗=c/(ηh¯)g^*=c/(η h) is the equilibrium gap. 15 The Role of β2 _2 (Adam’s Second Moment) Proposition 15.1 (β2 _2 and the Signal Hierarchy). Assume [,]≈0[P, H]≈ 0 (approximately diagonal Hessian in the parameter basis). For Adam with second-moment coefficient β2 _2, the preconditioner t=diag(v^t)−1/2P_t=diag( v_t)^-1/2 modifies the signal hierarchy as follows: 1. High β2 _2 (near 1): The preconditioner accurately tracks the gradient variance, making the noise approximately isotropic. The signal hierarchy becomes more concentrated in the top mode (backbone dominance increases, PC1 variance fraction ↑ ). 2. Low β2 _2 (near 0): The preconditioner responds to instantaneous gradient fluctuations. The signal hierarchy becomes more uniform (backbone dominance decreases, variance is spread across modes). The effective noise anisotropy satisfies: κN(β2)∝11−β2. _N( _2)\; \; 11- _2. (76) Remark 15.2 (Empirical Verification). This prediction is consistent with the β2 _2 sweep data: • PC1 variance fraction: 68.8% (β2=0.99 _2=0.99) → 52.5% (β2=0 _2=0). • Total drift: increases 1300×1300× from β2=0.99 _2=0.99 to β2=0 _2=0. • Reheating recoverability: G=+0.71G=+0.71 (β2=0.95 _2=0.95) vs. ∼0 0 (β2=0.80 _2=0.80). Higher β2 _2 concentrates the signal, widens the gap, and makes the dominant subspace more stable (and thus more recoverable after perturbation). Part VI Coupling to the Loss Function 16 The Stability Coefficient The classical BBP coherence ρj _j is a 0/10/1 threshold function: above the BBP threshold, ρj>0 _j>0; below, ρj=0 _j=0. Since the BBP threshold is vacuous in our regime, we need a replacement that captures the continuous dependence of subspace stability on the local eigenvalue gap. Definition 16.1 (Stability Coefficient). The stability coefficient of direction j is: αj= 1−C⋅∥Δ∥F2gapj2, _j\;=\;1- C· G _F^2gap_j^2, (77) where: • gapj=mini≠j|σj2−σi2|gap_j= _i≠ j _j^2- _i^2 is the nearest-neighbor eigenvalue gap at position j. • ∥Δ∥F G _F is the Frobenius norm of the Gram matrix change per step (the perturbation strength). • C is a constant of order 1 (determined by the Davis–Kahan bound; C=1C=1 suffices for the upper bound). The stability coefficient satisfies αj∈[0,1] _j∈[0,1] (clamped): αj=1 _j=1 when the gap is large (the direction is perfectly stable); αj=0 _j=0 when the gap is small relative to the perturbation (the direction is unstable). Remark 16.2 (Comparison with BBP Coherence). Property BBP Coherence ρj _j Stability Coefficient αj _j Threshold dj>dcritd_j>d_crit (noise boundary) gapj>∥Δ∥Fgap_j> G _F (eigenvalue gap) Transition Hard: ρj=0 _j=0 or ρj>0 _j>0 Continuous: αj∈[0,1] _j∈[0,1] Applies to Signal vs. noise boundary only Any eigenvalue gap (including intra-signal) Regime p/Wp/W finite (classical RMT) p/W→∞p/W→∞ (extreme aspect ratio) Remark 16.3 (Stability Near the Gap). For directions near the spectral gap (j=k∗j=k^* or j=k∗+1j=k^*+1): αk∗=1−C∥Δ∥F2(σk∗2−σk∗+12)2≈ 1−C∥Δ∥F2(dk∗+dk∗+1)2g2, _k^*=1- C G _F^2( _k^*^2- _k^*+1^2)^2\;≈\;1- C G _F^2(d_k^*+d_k^*+1)^2g^2, (78) where g=dk∗−dk∗+1g=d_k^*-d_k^*+1 is the spectral gap. As g→0g→ 0: αk∗→0 _k^*→ 0 (the direction at the gap becomes completely unstable). For directions far from the gap (j≪k∗j k^*): αj≈1 _j≈ 1 (highly stable), since gapj=σj2−σj+12gap_j= _j^2- _j+1^2 is large. For directions far below the gap (j≫k∗+1j k^*+1): αj _j depends on the local subdominant spectrum—typically αj≪1 _j 1 since the subdominant eigenvalues are closely spaced. Remark 16.4 (Evolving-NTK Derivation of αj _j). In Section 34, the stability coefficient receives a microscopic derivation. The Gram matrix’s signal direction j v_j averages the instantaneous NTK eigenvector j(t) q_j(t) over the window. The average fidelity over W steps gives: αj=1−W212∑k≠j(k⊤K˙j)2(λj−λk)2, _j=1- W^212 _k≠ j ( q_k K\, q_j)^2( _j- _k)^2, (79) where λk _k are NTK eigenvalues, k q_k are NTK eigenvectors, and K˙ K is the kernel evolution rate. The bound αj≥1−W2‖K˙‖2/(12(gλ(j))2) _j≥ 1-W^2\| K\|^2/(12\,(g_λ^(j))^2) gives the hierarchy α1≥α2≥⋯≥αk∗ _1≥ _2≥·s≥ _k^*, since the mode-specific gap gλ(j)=mink≠j|λj−λk|g_λ^(j)= _k≠ j| _j- _k| is largest for j=1j=1. The connection to the adiabatic parameter (Definition 34.2) is: αj≈1 _j≈ 1 if and only if mode j is in the adiabatic regime. 17 The Loss Improvement Decomposition Theorem 17.1 (Spectral Decomposition of Learning). Under [,]≈0[P, H]≈ 0, the expected per-step validation loss change decomposes spectrally as: [ΔLval]=−η∑j=1Wαj⟨j,∇Ltrain⟩⟨j,∇Lval⟩+O(ν2p) E[ L_val]=-η _j=1^W _j v_j,∇ L_train v_j,∇ L_val +O\! ( ν^2p ) (80) where αj _j is the stability coefficient (Definition 16.1). Proof. Step 1. Expand the update in the SVD basis: t=∑j=1Wcj^j δ_t= _j=1^Wc_j v_j, where ^j v_j are the sample singular vectors. Step 2. The expected loss change is [ΔLval]=⟨∇Lval,[t]⟩E[ L_val]= ∇ L_val,E[ δ_t] . Step 3. The sample singular vector ^j v_j aligns with the population direction j v_j with quality controlled by the Davis–Kahan bound: sin∠(^j,j)≤∥F/gapj ( v_j, v_j)≤ E _F/gap_j. The effective projection is: ⟨^j,∇Lval⟩≈αj⟨j,∇Lval⟩ v_j,∇ L_val ≈ _j v_j,∇ L_val . Step 4. All W directions contribute—there is no hard cutoff at k∗k^*. However, directions near or below the gap have αj≪1 _j 1 (unstable eigenvectors contribute noisy, unreliable projections), so their contribution is suppressed. Directions above the gap have αj≈1 _j≈ 1 and contribute coherently. Step 5. The noise residual O(ν2/p)O(ν^2/p) comes from the stochastic component of the updates. ∎ Remark 17.2 (All Directions Contribute). In the BBP framework, only k∗k^* directions contribute (the rest have ρj=0 _j=0). In our framework, all W directions contribute, weighted by their stability coefficient αj _j. The dominant modes (j≤k∗j≤ k^*) contribute with weight αj≈1 _j≈ 1. The subdominant modes (j>k∗j>k^*) contribute with weight αj≪1 _j 1 (but nonzero). This is physically correct: even unstable directions carry some information about the gradient, just unreliably. 17.1 Why the Largest Gap: The Gap Maximality Principle Every result so far—Davis–Kahan, the stability coefficient, the loss decomposition—applies at any spectral position. A natural question remains: among all W−1W-1 possible positions, why is k∗=argmaxdj/dj+1k^*= *arg\,max\,d_j/d_j+1 the dynamically privileged one? The following proposition answers this. Crucially, Parts (i) and (i) use only the Davis–Kahan theorem and the first-order loss identity—they require no assumption on the preconditioner P or the Hessian H. Part (i) is stated qualitatively in the same generality; the quantitative fixed-point analysis under [,]≈0[P, H]≈ 0 is given separately in Corollary 17.4. Proposition 17.3 (Gap Maximality Principle). Let k∗=argmax1≤j≤W−1dj/dj+1k^*= *arg\,max_1≤ j≤ W-1d_j/d_j+1. The position k∗k^* is uniquely privileged for phase transitions, for three mutually reinforcing reasons. No assumption on [,][P, H] is required. (i) Block optimality. For any partition of the spectrum at position j, the block Davis–Kahan bound on the subspace ≤j=span(1,…,j)V_≤ j=span( v_1,…, v_j) reads ∥sinΘ(≤j,~≤j)∥F≤∥Δ∥Fσj2−σj+12=∥Δ∥Fdj+12(rj2−1), (V_≤ j,\, V_≤ j) _F\;≤\; G _F _j^2- _j+1^2\;=\; G _Fd_j+1^2\,(r_j^2-1), (81) where rj=dj/dj+1r_j=d_j/d_j+1. Since k∗k^* maximises rjr_j, it also maximises rj2−1r_j^2-1. The full bound involves the product dj+12(rj2−1)d_j+1^2(r_j^2-1), so k∗k^* minimises the bound exactly when the spectrum has a single dominant gap—i.e., rk∗≫rjr_k^* r_j for all j≠k∗j≠ k^*, so that the ratio factor rj2−1r_j^2-1 dominates the dj+12d_j+1^2 factor. This is the empirically observed case in every experiment (the spectrum separates into two clusters with one sharp drop between them). (i) Functional hierarchy. The first-order loss identity (valid for any preconditioner) gives [ΔLval]=−η∑j⟨^j,∇Lval⟩⟨^j,∇Ltrain⟩+O(∥2)E[ L_val]=-η _j v_j,∇ L_val v_j,P∇ L_train +O( δ ^2), where ^j v_j are the sample Gram eigenvectors. By Davis–Kahan, the quality of each sample eigenvector ^j v_j as an estimator of the population direction j v_j is controlled by the local eigenvalue gap: well-resolved directions (gapjgap_j large) contribute a consistent sign step after step; poorly-resolved directions (gapjgap_j small) contribute noise that averages toward zero. Gap collapses at different positions therefore have qualitatively different consequences: • j<k∗j<k^* (within the signal cluster). Modes j and j+1j+1 mix, but both lie inside ≤k∗V_≤ k^*. The block Davis–Kahan bound at k∗k^* is unaffected—the block boundary is at k∗k^*, not at j. The total projection of the gradient onto ≤k∗V_≤ k^* is preserved (the projector Π≤k∗ _V_≤ k^* is invariant under rotations within the block). Individual eigenvectors may rotate, but the subspace is intact. • j=k∗j=k^* (the signal/subdominant boundary). The block ≤k∗V_≤ k^* is no longer well-separated from >k∗V_>k^*: the Davis–Kahan bound diverges as δk∗→0 _k^*→ 0. The signal subspace boundary dissolves, the gradient projection ∥Π≤k∗∇L∥2 _V_≤ k^*∇ L ^2 fluctuates from step to step, and the coherent learning contribution of the entire signal block becomes unreliable. This is the maximally disruptive event. • j>k∗j>k^* (within the subdominant cluster). These modes have small local gaps (gapj≪gapk∗gap_j _k^*), so their sample eigenvectors are poorly resolved: the per-step gradient projections ⟨^j,∇L⟩ v_j,∇ L fluctuate in sign and average toward zero. Rearranging them has minimal impact on cumulative learning. Therefore, the k∗k^*-gap is the only gap whose collapse qualitatively changes the learning dynamics: collapses at j<k∗j<k^* rearrange modes within an already-stable block, and collapses at j>k∗j>k^* rearrange modes that barely contribute. (i) Self-reinforcement. The k∗k^*-gap sustains itself through a positive feedback loop: a large gap ensures well-resolved eigenvectors (αk∗≈1 _k^*≈ 1), which enable coherent gradient alignment, which maintains the signal strengths djd_j for j≤k∗j≤ k^*, which maintains the gap. This loop requires only the Davis–Kahan bound and the first-order loss identity—not [,]≈0[P, H]≈ 0. Gaps at other positions lack this loop: their local gaps are smaller, so αj≪1 _j 1, the gradient projections are incoherent, and the driving that would sustain the gap is noise-dominated rather than signal-dominated. The k∗k^*-gap is the unique self-sustaining spectral structure. Phase transitions occur when an external force—weight-decay damping, or a change in the loss landscape that weakens the gradient driving—overcomes this self-reinforcement and drives the gap to zero. Proof sketch. Part (i) is immediate from the block Davis–Kahan theorem (Theorem 9.1) applied to the interval [σj+12,σj2][ _j+1^2,\, _j^2]: the separation is δ=σj2−σj+12=dj+12(rj2−1)δ= _j^2- _j+1^2=d_j+1^2(r_j^2-1), and k∗=argmaxrjk^*= *arg\,max\,r_j maximises rj2−1r_j^2-1. Part (i): for j<k∗j<k^*, modes j and j+1j+1 both lie in ≤k∗V_≤ k^*, so their mixing is an internal rotation. The projector Π≤k∗ _V_≤ k^* is unchanged (it depends on the subspace, not on the basis within it), so ∥Π∇L∥2 ∇ L ^2 is preserved. For j=k∗j=k^*: δk∗→0 _k^*→ 0, the block bound diverges, and the signal subspace becomes ill-defined. For j>k∗j>k^*: gapj≪gapk∗gap_j _k^*, so sin∠(^j,j)=O(1) ( v_j, v_j)=O(1) (the sample eigenvector is essentially random within its eigenspace), and ⟨^j,∇L⟩ v_j,∇ L has fluctuating sign with mean close to zero. Part (i): the stability coefficient αk∗=[1−C∥Δ∥F2/δk∗2]+ _k^*=[1-C G _F^2/ _k^*^2]_+ is close to 11 when g is large (ensuring coherent gradient alignment) and close to 0 when g is small (destroying coherence). The per-step signal injection into mode k∗k^* is proportional to ⟨^k∗,∇L⟩2 v_k^*,P∇ L ^2, which scales as αk∗⟨k∗,∇L⟩2 _k^* v_k^*,P∇ L ^2 by Davis–Kahan. When αk∗≈1 _k^*≈ 1, this injection maintains dk∗d_k^* and hence g; when αk∗→0 _k^*→ 0, the injection becomes noise. For secondary gaps, αj≪1 _j 1 always, so the injection is never coherent enough to sustain a gap—no feedback loop forms. ∎ Corollary 17.4 (Fixed-Point Analysis under [,]≈0[P, H]≈ 0). Under the commutativity assumption [,]≈0[P, H]≈ 0, the self-reinforcement of Part (i) can be made quantitative. The gap flow (Theorem 14.2) takes the form dgdt=F(αk∗(g))−η(h¯+ω)g, dgdt=F\! ( _k^*(g) )-η( h+ω)\,g, (82) where F(α)=−η(hk∗−hk∗+1)d¯+ηW(|Gk∗eff|2/dk∗−|Gk∗+1eff|2/dk∗+1)F(α)=-η(h_k^*-h_k^*+1) d+η W\! ( G_k^*^eff ^2/d_k^*- G_k^*+1^eff ^2/d_k^*+1 ) is the net driving. Since F is bounded (it saturates as α→1α→ 1) while the damping η(h¯+ω)gη( h+ω)\,g is unbounded: (a) A positive equilibrium g∗>0g^*>0 exists whenever F(1)>0F(1)>0 (net driving positive at full coherence). (b) The equilibrium is linearly stable, with relaxation timescale τg≈1/[η(h¯+ω)] _g≈ 1/[η( h+ω)]. (c) At secondary positions j≠k∗j≠ k^*, αj≪1 _j 1 clamps the driving near its incoherent value Fj(0)F_j(0), eliminating the g-dependent feedback. Remark 17.5 (Universality across Optimisers). The Gap Maximality Principle (Parts (i)–(i) of Proposition 17.3) holds for any preconditioner P, including Muon, whose orthogonalisation step explicitly violates [,]≈0[P, H]≈ 0. This is consistent with experiment: Muon produces a different gap position (k∗=1k^*=1) from AdamW (k∗=2k^*=2), but in both cases phase transitions occur at k∗k^*, never at secondary gaps. The commutativity assumption determines where k∗k^* sits; the Gap Maximality Principle explains why k∗k^* is special, regardless of the optimiser. Remark 17.6 (Formation vs. Persistence). The Gap Maximality Principle explains why an existing gap at k∗k^* is self-sustaining (persistence), but not why a gap forms there in the first place (formation). Gap formation is provided by the Hessian mechanism (Proposition 5.1): the Hessian of the loss has outlier eigenvalues separated from the bulk, which creates a gap in the trajectory spectrum via dj∝1/hjd_j 1/ h_j. The Hessian mechanism does use [,]≈0[P, H]≈ 0. The full logical chain is therefore: Hessian mechanism[,]≈0→formationgap at k∗→persistence (assumption-free)self-sustaining structure. [P, H]≈ 0Hessian mechanism\; \;formation\;\;gap at k^*\; \;persistence (assumption-free)\;\;self-sustaining structure. For optimisers that violate [,]≈0[P, H]≈ 0, the gap still forms (the Hessian still has outliers), but through a mechanism not captured by Proposition 5.1. An optimiser-agnostic theory of gap formation is an open problem. Remark 17.7 (Empirical Confirmation). The gap maximality principle is confirmed by every phase transition observed in our experiments (Section 21–Section 28): all 24/24 grokking events with weight decay, the k∗k^* shifts in GPT-2 and TinyStories, and the Dyck-1/SCAN gap openings all occur at the position of the largest consecutive eigenvalue ratio, never at secondary gaps. 18 The Master Equation Proposition 18.1 (Master Equation). Under [,]≈0[P, H]≈ 0, the validation loss dynamics are approximated by: dLvaldt=−η∑j=1Wαj(gj)⋅Gjtrain⋅Gjval,whereGjtrain=⟨j,∇Ltrain⟩,Gjval=⟨j,∇Lval⟩,αj=max(0, 1−C∥Δ∥F2gapj2). aligned dL_valdt&=-η _j=1^W _j(g_j)· G_j^train· G_j^val,\\[8.0pt] where G_j^train&= v_j,P∇ L_train ,\\ G_j^val&= v_j,∇ L_val ,\\ _j&= \! (0,\;1- C G _F^2gap_j^2 ). aligned (83) Remark 18.2 (Self-Containment). The master equation, together with the flow system (67) for dj\d_j\, forms a closed dynamical system (up to the external inputs ∇L∇ L, H, which are determined by the loss landscape and data distribution). Given initial conditions and the loss landscape geometry, the framework predicts: 1. Which directions are stable (αj≈1 _j≈ 1, j≤k∗j≤ k^*). 2. How fast learning proceeds along each direction. 3. When phase transitions occur (gap collapse/opening). 4. How optimizer hyperparameters (η,ω,β2η,ω, _2) control the transition timing. Part VII The Complete Dynamical System and Where BBP Fits 19 The Full System Collecting all results, the spectral edge thesis is captured by the following self-contained dynamical system (using the linearised signal flow, Remark 12.10): Signal flow:ddj2dt≈−2η(hj+ω)dj2+η2W(Sj+2Gjeffj),j=1,…,W,Noise flow:dν2dt=η2pTr(grad)−2ηων2,Gap position:k∗(t)=argmaxjdjdj+1,Gap:g(t)=dk∗(t)−dk∗+1(t),Stability:αj=max(0, 1−C∥Δ∥F2gapj2),Loss:dLvaldt=−η∑j=1WαjGjeff,trainGjval. aligned &Signal flow: d_j^2dt≈-2η(h_j+ω)\,d_j^2+η^2W (S_j+2G_j^effN_j ), j=1,…,W,\\[6.0pt] &Noise flow: dν^2dt= η^2pTr(P _gradP)-2ηων^2,\\[6.0pt] &Gap position: k^*(t)= *arg\,max_j d_jd_j+1,\\[6.0pt] &Gap: g(t)=d_k^*(t)-d_k^*+1(t),\\[6.0pt] &Stability: _j= \! (0,\,1-C G _F^2gap_j^2 ),\\[6.0pt] &Loss: dL_valdt=-η _j=1^W _j\,G_j^eff,train\,G_j^val. aligned (84) Phase transitions occur when g(t)g(t) passes through zero: • g→0g→ 0 from above: gap collapse, subspace destabilisation, loss stagnation. • g→0g→ 0 from below (g was at a different position and a new gap opens): gap opening, new stable direction, capability gain. 20 Where BBP Fits: The Outer Boundary The BBP framework is not wrong—it is simply irrelevant in the extreme aspect ratio regime. It provides the outer boundary of the theory: the condition under which signal is distinguishable from noise at all. Remark 20.1 (BBP as the Outer Boundary). The BBP detection threshold dcrit=ν(p(W−1))1/4d_crit=ν(p(W-1))^1/4 defines the weakest possible signal that can be detected. In our regime: dWdcrit≥ 20, d_Wd_crit\;≥\;20, (85) so the outer boundary is never approached. The BBP transition becomes operationally relevant when: 1. W→pW→ p (much larger windows), so the noise eigenvalue spread approaches the signal eigenvalue spread. 2. p→W2p→ W^2 (much smaller parameter spaces), e.g., per-layer analysis with small layers. 3. ν→dW/(p(W−1))1/4ν→ d_W/(p(W-1))^1/4 (very noisy training: high learning rate, small batch, no weight decay). In none of our experimental settings are these conditions met. Remark 20.2 (The BBP Hierarchy). The complete hierarchy of phase transitions, from inner to outer, is: 1. Intra-signal gap (this paper): dk∗=dk∗+1d_k^*=d_k^*+1 within the signal spectrum. This is the operative transition. 2. BBP transition: dW=dcritd_W=d_crit, the weakest signal crosses the noise floor. This is the outer boundary, trivially satisfied in our regime. 3. Tracy–Widom fluctuations: corrections of order p−5/6p^-5/6 to the noise ceiling. Negligible. Our framework subsumes the BBP framework: when the intra-signal gap sits at k∗=W−1k^*=W-1 (only the last eigenvalue is close to dcritd_crit), the intra-signal transition reduces to the BBP transition. But in our data, k∗∈2,3k^*∈\2,3\—far from the outer boundary. Part VIII Empirical Verification We now confront the theoretical predictions of Part I–Part VII with experimental data from two transformer models: • TinyStories 51M (p=163,150,848p=163,150,848, W=10W=10, η=10−3η=10^-3, ω=0.5ω=0.5, β2=0.95 _2=0.95, B=20B=20): 72 rolling-window spectra over ∼9000 9000 training steps. • GPT-2 124M (p=162,364,416p=162,364,416, W=10W=10, η=3×10−5η=3× 10^-5 fine-tune, ω=0.1ω=0.1, β2=0.95 _2=0.95, B=20B=20): 41 rolling-window spectra over ∼8000 8000 fine-tuning steps. All singular values σ1≥⋯≥σ10 _1≥·s≥ _10 are computed from the 10×p10× p trajectory Gram matrix at each window position. 21 Test 1: The BBP Threshold is Vacuous Proposition 4.1 predicts σW/dcrit≫1 _W/d_crit 1 in the extreme aspect ratio regime. Table 1: Empirical verification of Proposition 4.1: every singular value exceeds the BBP detection threshold by at least one order of magnitude. Model σW/dcrit _W/d_crit min median max GPT-2 124M (41 windows) 22.9×22.9× 56.3×56.3× 63.1×63.1× TinyStories 51M (72 windows) 8.0×8.0× 53.3×53.3× 62.5×62.5× The weakest case (8×8× for TinyStories, early training when gradient variance is highest) still far exceeds the threshold. The BBP transition is never approached: all W=10W=10 eigenvalues are signal. Figure 2: Spectral edge analysis for TinyStories 51M. Top: Consecutive singular value ratios σk/σk+1 _k/ _k+1 over training. The σ1/σ2 _1/ _2 gap (red) dominates during the plateau phase, while σ2/σ3 _2/ _3 (blue) shows the three-phase rise–plateau–collapse pattern. Middle: Eigenvalue gaps σk2−σk+12 _k^2- _k+1^2 over training, showing the same three-phase structure. Bottom: Gap ratio σ2/σ3 _2/ _3 (blue) overlaid with validation loss (red), confirming that the spectral edge tracks learning progress. 22 Test 2: k∗k^* = Argmax Ratio Definition 6.1 defines k∗(t)=argmaxjσj/σj+1k^*(t)= *arg\,max_j\, _j/ _j+1. Table 2: Distribution of k∗k^* over training windows. The gap position is stable within each training phase. Model Mode k∗k^* Frequency Max ratio R GPT-2 124M 2 51.2% (21/41) R¯=1.73 R=1.73, peak 2.862.86 TinyStories 51M 1 77.8% (56/72) R¯=2.17 R=2.17, peak 5.905.90 Both models show a dominant k∗k^* that is far smaller than W=10W=10, confirming that the relevant spectral structure is within the signal hierarchy. The observed ratios (R=1.73R=1.73–5.905.90) are vastly above the null expectation of 1+O(10−4)1+O(10^-4) (Proposition 6.4), confirming genuine gap structure. Remark 22.1 (Argmax vs. cross-correlation k∗k^*: the +1+1 offset). The argmax mode in Table 2 identifies the position where σj/σj+1 _j/ _j+1 is largest at each window. A complementary definition—used throughout the body text—selects the position whose ratio has the strongest cross-correlation with validation loss. Across every experiment in this thesis, the cross-correlation spectral edge sits one position above the argmax mode: Experiment Argmax k∗k^* Cross-corr k∗k^* GPT-2 124M (pretrain) 2 3 (σ3/σ4 _3/ _4, |r|=0.870|r|=0.870) TinyStories 51M (AdamW) 1 2 (σ2/σ3 _2/ _3) Grokking (all tasks) 1 g23g_23 is the dynamical signal This offset has a natural explanation. The argmax ratio marks the settled gap—the pair already well separated. The dynamical spectral edge is one position further out because that is the active frontier: the mode with the smallest gap protecting it rotates fastest (Davis–Kahan), responds most sensitively to perturbations, and therefore tracks training dynamics most tightly. Both definitions confirm k∗≪Wk^* W. 22.1 k∗k^* Dynamics: Gap Position Shifts During Training A striking empirical finding is that k∗(t)k^*(t) is not constant—it shifts during training. For GPT-2 124M: Phase Steps Dominant k∗k^* FineWeb pretrain ≤17 800≤ 17\,800 k∗=3k^*=3 Post-shift (OWT) >17 800>17\,800 k∗=2k^*=2 Each shift in k∗k^* is itself a phase transition: the identity of the “most separated” pair changes. In our framework, a k∗k^* shift occurs when the gap ratio σj/σj+1 _j/ _j+1 at the current k∗k^* falls below the ratio at a different position—i.e., when the gap collapses at one location and opens at another. Remark 22.2 (Which Ratio to Monitor). When k∗=jk^*=j, the operative phase transition is σj/σj+1→1 _j/ _j+1→ 1 (gap collapse at position j). The phase transition is not at σj+1/σj+2 _j+1/ _j+2 (the next pair down). Specifically: • When k∗=1k^*=1: monitor σ1/σ2 _1/ _2 for collapse. • When k∗=2k^*=2: the transition is σ2/σ3→1 _2/ _3→ 1. • When k∗=3k^*=3: the transition is σ3/σ4→1 _3/ _4→ 1. 23 Test 3: Krylov Bound Consistency Proposition 5.5 predicts k∗≤K=#j:hj≳1/(ηW)k^*≤ K=\#\j:h_j 1/(η W)\. Table 3: Krylov bound consistency. The observed k∗k^* is bounded by the predicted number of Hessian outliers. Model η 1/(ηW)1/(η W) Expected K Observed k∗k^* TinyStories 51M 10−310^-3 100 2–3 1 GPT-2 124M 3×10−53× 10^-5 3333 3–4 2 The bound is consistent in both cases: the observed k∗k^* is at or below the predicted Krylov dimension. TinyStories, with its larger learning rate, has a lower Hessian threshold and correspondingly lower k∗k^*. 24 Test 4: Gap Ratio–Loss Correlation The spectral decomposition of learning (Theorem 17.1) predicts that the gap ratio R(t)R(t) and validation loss are correlated: when the gap is large, learning proceeds stably; when the gap collapses, learning stalls. Table 4: Cross-correlation between gap ratio R(t)R(t) and validation loss. Model Zero-lag |r||r| Best lag |r||r| GPT-2 124M 0.656 0.656 (lag 0) TinyStories 51M 0.359 0.669 (lag −5-5) Both models show substantial gap–loss correlation (|r|≈0.66|r|≈ 0.66–0.670.67 at optimal lag). The negative lag for TinyStories indicates that loss changes lead the spectral gap by approximately 5 window positions—consistent with the fact that the Gram matrix is a lagged summary of the trajectory. 25 Test 5: Stability Coefficient Hierarchy Remark 16.3 predicts αj≈1 _j≈ 1 for j<k∗j<k^* (dominant modes above the gap), αj≪1 _j 1 near the gap, and αj _j variable below. For GPT-2 124M with argmax mode k∗=2k^*=2 (Table 2; the dominant tier j<k∗j<k^* is nonempty): Table 5: Stability coefficient hierarchy for GPT-2 124M (k∗=2k^*=2). Mean αj _j across 21 windows where k∗=2k^*=2. Region Position Mean αj _j Dominant (j<k∗j<k^*) j=1j=1 0.818 At gap (j=k∗,k∗+1j=k^*,k^*+1) j=2,3j=2,3 0.234 Subdominant (j>k∗+1j>k^*+1) j=4,…,10j=4,…,10 ≈0≈ 0 The predicted hierarchy αdom>αgap>αsub _dom> _gap> _sub is confirmed: 0.818>0.234>0.0000.818>0.234>0.000. Directions above the gap are nearly perfectly stable (α≈0.82α≈ 0.82); directions at the gap are marginally stable (α≈0.23α≈ 0.23); subdominant directions are unstable (α≈0α≈ 0). Figure 3: Stability coefficient hierarchy for GPT-2 124M. Heatmap of αj _j (Definition 16.1) across eigenvalue index j (vertical) and training step (horizontal). The dominant mode (j=1j=1, green) has α≈1α≈ 1 throughout training. The gap mode (j=2j=2) fluctuates between stable (green) and unstable (red) as k∗k^* shifts. All subdominant modes (j≥3j≥ 3, red) have α≈0α≈ 0. The hierarchy α1>α2>⋯ _1> _2>·s is visually obvious. Remark 25.1 (When k∗=1k^*=1). For TinyStories with k∗=1k^*=1, the “dominant” tier (j<1j<1) is empty. In this case, the stability hierarchy test applies only to the gap and below: αgap=0.567>αsub=0.009 _gap=0.567> _sub=0.009, which is consistent with the prediction that directions at the gap are more stable than those in the closely-spaced subdominant tier. 26 Test 6: Gap Flow Equation The gap flow (Theorem 14.2) predicts three contributions to dg/dtdg/dt: curvature asymmetry, gap damping, and driving asymmetry. Testing the simplest prediction—that the average curvature damps the gap (ρ(g,g˙)<0ρ(g, g)<0)—yields an inconclusive result: r(g,g˙)=+0.08r(g, g)=+0.08 for GPT-2 (R2=0.006R^2=0.006). This is expected: the damping term −ηh¯⋅g-η h· g is only one of three terms in the gap flow equation. The driving asymmetry and curvature asymmetry terms can dominate, especially during phase transitions. A proper test requires controlling for all three terms simultaneously, which in turn requires Hessian eigenvalue estimates not available from the spectrum alone. 27 Test 7: Gap Opening and Grokking (Dyck/SCAN) Proposition 10.7 predicts that gap opening corresponds to capability gain, and Remark 10.8 maps grokking onto a delayed gap opening event within the signal hierarchy. We test this in a setting where the capability gain is unambiguous: the grokking transition, where generalization accuracy jumps from near-zero to near-perfect. Setup. We analyze weight-matrix singular values σ1≥σ2≥⋯ _1≥ _2≥·s of the query projection WQW_Q during training of 2-layer transformers on Dyck-1 language (dmodel=128d_model=128, ∼ 150K parameters) and 6-layer encoder–decoder transformers on SCAN compositional generalization (dmodel=256d_model=256, ∼ 1.5M parameters). Each task is trained with weight decay λ∈0,1.0λ∈\0,1.0\ across 3 seeds, giving 12 runs total. The gap ratio R(t)=σ1(t)/σ2(t)R(t)= _1(t)/ _2(t) is computed at each checkpoint. Grokking is defined as the first step where test accuracy exceeds 0.95. Results. Table 6: Gap ratio at grokking and terminal training for Dyck and SCAN. Weight decay ω=1.0ω=1.0 enables grokking; ω=0ω=0 does not. Task Seed Grok Step RgrokR_grok RterminalR_terminal Groks? Dyck (ω=1ω=1) 42 600 1.13 3187 ✓ Dyck (ω=1ω=1) 137 1400 1.19 2841 ✓ Dyck (ω=1ω=1) 2024 1000 1.09 1756 ✓ SCAN (ω=1ω=1) 42 3000 1.05 12.4 ✓ SCAN (ω=1ω=1) 137 4000 1.11 8.7 ✓ SCAN (ω=1ω=1) 2024 2500 1.02 6.3 ✓ Dyck (ω=0ω=0) all — — <1.2<1.2 × SCAN (ω=0ω=0) all — — <1.1<1.1 × The pattern is striking: 1. At grokking: The gap ratio is small (R≈1.02R≈ 1.02–1.191.19), indicating near-degeneracy of the top two modes. 2. After grokking: The gap opens by 2–3 orders of magnitude (R→103R→ 10^3 for Dyck, R→101R→ 10^1 for SCAN), driven by continued weight decay compressing σ2→0 _2→ 0. 3. Controls: Without weight decay, no gap opens and no grokking occurs—6/6 grokking runs with λ=1.0λ=1.0, 0/6 with λ=0λ=0 (100% hit rate, 0% false positives). This is consistent with Proposition 10.7: the gap opening event is of the same order as the capability gain. The small gap at the grokking step and large gap afterward matches Remark 10.8: the generalizing direction exists as a subdominant signal; grokking occurs when training dynamics cross a threshold, after which weight decay drives the gap open. Figure 4: Grokking as a spectral edge event (Dyck-1). Grokking runs (weight decay λ=1.0λ=1.0, red/solid, 3 seeds) vs. control runs (λ=0λ=0, blue/dashed, 3 seeds). Top: singular value ratio σ1/σ2 _1/ _2 of WQW_Q. With weight decay, the ratio rises dramatically as σ2→0 _2→ 0 (gap opens); without weight decay, the ratio stays flat (no gap). Middle: Frobenius norm ‖WQ‖F\|W_Q\|_F. Weight decay compresses the weights; the control runs retain large norms. Bottom: test accuracy. Grokking (accuracy →1→ 1) occurs only in the weight-decay runs, coinciding with the gap opening above. All 6 grokking runs show gap opening; all 6 controls show neither gap opening nor grokking. Modular arithmetic: gap closing precedes grokking. For modular arithmetic (4 operations × 3 seeds, p=97p=97, 1-layer transformer, ∼ 400K parameters), the spectral mechanism operates differently. The eigenvalue spectrum of the attention operator is effectively rank-2 from early training (eigenvalues λ3 _3–λ5≈0 _5≈ 0). Rather than a gap opening at grokking, the sub-leading gap g23=λ2−λ3g_23= _2- _3 closes monotonically before grokking, declining 6.6×6.6× on average (range 4.34.3–8.4×8.4×): Table 7: Eigenvalue sub-leading gap decline for modular arithmetic (4 operations, 3 seeds each). The gap g23g_23 decreases monotonically from early training to grokking, with k∗k^* shifting to 1 at terminal training in 10/12 runs. Op Seed Grok g23earlyg_23^early g23grokg_23^grok Decline kterm∗k^*_term add 42 3100 1.32 0.18 7.5×7.5× 3 add 137 3000 1.25 0.19 6.6×6.6× 1 add 2024 2500 1.34 0.23 5.8×5.8× 3 sub 42 3200 1.14 0.15 7.5×7.5× 1 sub 137 3300 1.15 0.15 7.5×7.5× 1 sub 2024 3600 1.17 0.14 8.4×8.4× 1 mul 42 2600 1.33 0.21 6.3×6.3× 1 mul 137 3000 1.27 0.17 7.4×7.4× 1 mul 2024 2500 1.24 0.21 5.9×5.9× 1 x2+y2x^2+y^2 42 1800 1.26 0.30 4.3×4.3× 1 x2+y2x^2+y^2 137 2000 1.26 0.17 7.3×7.3× 1 x2+y2x^2+y^2 2024 1900 1.33 0.28 4.8×4.8× 1 Simultaneously, the ratio λ1/λ2 _1/ _2 rises from ∼1.6 1.6 to ∼2.0 2.0 at grokking, so the gap position migrates from k∗=2k^*=2–33 to k∗=1k^*=1. This is a k∗k^* shift: the gap at position 2–3 closes (consistent with Proposition 5.1), destabilizing the sub-leading subspace, while the gap at position 1 opens as σ1 _1 separates from the remaining modes. Grokking corresponds to the resolution of this spectral symmetry-breaking. The temporal ordering for modular arithmetic is: g23↓→σ1≈σ2→↑→σ1≫σ2→grok,g_23\! \;→\; _1≈ _2\;→\;D\! \;→\; _1 _2\;→\;grok, (86) where D denotes the SGD commutator defect. Weight decay enables this process in all 12 runs; without weight decay, none of the 12 control runs grok (0/12). Multi-task arithmetic. Multi-task models (dual-task add+mul, 3 seeds; tri-task add+mul+sq, 3 seeds) confirm the same pattern: 6/6 grok with weight decay, 0/6 without. In these models, multiple tasks grok at different positions along a shared spectral trajectory, with the staggered ordering (mul → sq → add) reflecting each task’s alignment threshold. Synthesis. The Dyck/SCAN and modular arithmetic experiments test complementary aspects of the framework: • Dyck/SCAN: gap opens at grokking (Proposition 10.7), with R rising from ∼1.1 1.1 to 10310^3. • Modular arithmetic: sub-leading gap closes before grokking, triggering a k∗k^* shift from position 2–3 to position 1. This is a gap closing event (at the old k∗k^*) followed by a gap opening event (at the new k∗=1k^*=1). Both mechanisms are predicted by the gap dynamics (Theorem 14.2): gap closing destabilizes the subspace, and gap opening at a new position stabilizes the generalizing direction. The weight decay → grokking correspondence is universal: 24/24 grok with λ>0λ>0, 0/24 without, across all four task families (48 runs total). Remark 27.1 (Weight Decay as Gap Driver). Weight decay plays a dual role across all experiments: in Dyck/SCAN, it drives subdominant singular values toward zero (opening the gap); in modular arithmetic, it compresses the sub-leading eigenvalue spectrum (closing g23g_23), which triggers spectral symmetry-breaking and a k∗k^* shift. The 100% correspondence between weight decay and grokking across 48 runs (24 grok, 24 control) spanning four task families is consistent with weight decay being a key driver by which gap dynamics are associated with the generalization transition. Remark 27.2 (Weight Matrix vs. Trajectory Matrix Gaps). The grokking experiments analyze weight-matrix spectra (σj(WQ) _j(W_Q) for Dyck/SCAN, eigenvalues λj _j for modular arithmetic) rather than the trajectory Gram spectrum σj() _j( G). The Davis–Kahan stability framework (Remark 16.3) applies to both: gap dynamics control subspace stability regardless of the matrix being analyzed. The universality of the gap → capability correspondence across both spectral types provides strong evidence for the generality of the framework. 28 Test 8: Causal Intervention (Loss Decomposition) The strongest test is causal: does removing a signal direction from the parameter updates degrade learning by the amount predicted by the loss decomposition (Theorem 17.1)? Experimental setup. Using the multi-task modular arithmetic models from Xu (2026) (3 tasks, 315K parameters, WD == 1.0, seed 42), we analyze 9 timepoints during the grokking transition (steps 8K–24K). At each timepoint t: 1. Build a Gram matrix from a window of W consecutive parameter updates centered at t. 2. Extract signal directions j v_j and strengths djd_j from the SVD. 3. Compute the stability coefficient αj _j by comparing eigenvectors between the two halves of the window. 4. Compute gradient projections Gjtrain=j⊤∇LtrainG_j^train= v_j ∇ L_train and Gjval=j⊤∇LvalG_j^val= v_j ∇ L_val. 5. Compare the predicted per-direction importance αjGjtrainGjval _j\,G_j^train\,G_j^val against the actual per-direction loss change (first-order Taylor expansion). Results: window size matters. The stability coefficient αj _j requires sufficient data to estimate reliably. We sweep window sizes W∈10,20,30,40W∈\10,20,30,40\ checkpoints (spanning 2K–8K training steps): W Span ρ(α⋅G⋅G)ρ(α· G· G) ρ(G⋅G)ρ(G· G) α helps? 10 2K steps 0.21 0.76 No 20 4K steps 0.42 0.72 No 30 6K steps 0.75 0.63 Yes 40 8K steps 0.68 0.35 Yes At small windows (W≤20W≤ 20), αj _j is noise—the half-windows have too few data points for reliable SVD beyond mode 1. The gradient projections GjtrainGjvalG_j^train\,G_j^val alone achieve ρ≈0.76ρ≈ 0.76. At W=30W=30, a crossover occurs: the full formula αjGjtrainGjval _j\,G_j^train\,G_j^val achieves ρ=0.75ρ=0.75, outperforming G⋅G· G alone (ρ=0.63ρ=0.63). At W=40W=40, the advantage of including α grows further (ρ=0.68ρ=0.68 vs. 0.350.35). Interpretation. The loss decomposition formula requires three ingredients: the gradient projections GjG_j (which directions the loss gradient points), the stability coefficient αj _j (how reliably those directions are maintained), and a sufficiently large observation window to estimate αj _j. The crossover at W≈30W≈ 30 reflects the minimum window for the SVD to resolve the stability of subdominant modes (j≥2j≥ 2). Additional findings. • The first-order Taylor decomposition ∑jGjval⋅Δθj _jG_j^val· _j predicts the total test loss change with Pearson r=0.82r=0.82 (p<10−3p<10^-3) across all timepoints. • The signal strength djd_j alone has ρ≈0ρ≈ 0 for instantaneous importance (unlike the post-convergence setting where djd_j dominates), confirming that GjG_j captures which directions are currently active, not which have accumulated the most change historically. • Post-convergence, djd_j alone predicts the total causal importance of each direction with ρ=0.98ρ=0.98—consistent with djd_j being the time-integral of the instantaneous contributions. 29 Summary of Theory–Experiment Match Table 8: Theory–experiment match across all predictions and models. Prediction Formula GPT-2 124M TinyStories 51M Dyck/SCAN Status BBP vacuous (4.1) σW/dcrit≫1 _W/d_crit 1 2323–63×63× 88–63×63× — ✓ k∗k^* = argmax (6.1) argmaxjσj/σj+1 *arg\,max_j _j/ _j+1 mode k∗=3k^*=3 mode k∗=2k^*=2 mode k∗=1k^*=1 ✓ Krylov bound (5.5) k∗≤Kk^*≤ K 3≤33≤ 3–44 2≤22≤ 2–33 — ✓ Gap–loss corr (17.1) |r(R,Lval)|>0|r(R,L_val)|>0 |r|=0.66|r|=0.66 |r|=0.67|r|=0.67 — ✓ Stability (16.3) αdom>αgap _dom> _gap 0.82>0.230.82>0.23 (N/A, k∗=1k^*=1) — ✓ Gap flow (14.2) ρ(g,g˙)<0ρ(g, g)<0 r=+0.08r=+0.08 — — ∼ Gap dynamics (10.7) gap event ⇒ capability — — 24/24 grok, 0/24 ctrl ✓ Causal (28) ρ(αGG,Taylor)>0ρ(α G,\,Taylor)>0 — — ρ=0.75ρ=0.75 (W=30W=30) ✓ Overall: 7 of 8 predictions confirmed across six model families (GPT-2 124M, TinyStories 51M, Dyck 150K, SCAN 1.5M, modular arithmetic single-task and multi-task), 1 inconclusive (gap flow, requires Hessian data not available from the spectrum alone). No predictions are contradicted. The gap dynamics ⇒ grokking test (Section 27) achieves 100% hit rate with 0% false positives across 48 controlled runs (24 grok, 24 control). The causal intervention test (Section 28) confirms that the per-direction loss decomposition formula correctly ranks signal directions during the grokking transition (ρ=0.75ρ=0.75, p<0.05p<0.05 at each timepoint). Part IX Connections, Extensions, and Open Problems 30 Connection to Dyson Brownian Motion The eigenvalue dynamics of Theorem 10.1 and Corollary 10.2 are a discrete analog of Dyson Brownian motion (Dyson [5]). In the continuous limit: Remark 30.1 (Empirical Observation: Dyson-Type Dynamics of the Gram Spectrum). The eigenvalues λj(t)\ _j(t)\ of (t) G(t) evolve in a manner consistent with a system of interacting particles: • External potential: determined by the Hessian curvatures hjh_j and gradient projections GjG_j (through the signal flow). • Pairwise repulsion: ∝1/(λj−λi) 1/( _j- _i) (the Dyson repulsion from the non-crossing rule). • Stochastic forcing: from minibatch noise (entering through Δ G). This suggests an analogy between the spectral edge thesis and a particle system in random matrix theory, where phase transitions are analogous to particle collisions moderated by the repulsive potential. 31 Connection to Tensor Programs (Yang) Proposition 31.1 (Gram Matrix from NTK Eigenvalues). For SGD on loss L=1N∑αℓ(f(xα),yα)L= 1N _α (f(x_α),y_α), the Gram matrix entry is: Gst=η2N2∑α,βrs,αrt,βΘs,t(xα,xβ),G_st= η^2N^2 _α,βr_s,α\,r_t,β\; _s,t(x_α,x_β), (87) where rs,α=ℓ′(fs(xα),yα)r_s,α= (f_s(x_α),y_α) is the loss derivative and Θs,t(x,x′)=⟨∇f(x;s),∇f(x′;t)⟩ _s,t(x,x )= _ θf(x; θ_s), _ θf(x ; θ_t) is the cross-time NTK. Proposition 31.2 (Kernel Regime: Flat Hierarchy). When η=O(1/n)η=O(1/n), the NTK is constant: Θs,t≈Θ0 _s,t≈ _0, the residuals decay as t=e−η0t/N0 r_t=e^-η K_0t/N r_0, and the Gram matrix becomes: Gst=η2N2∑k=1Nλkck2e−ηλk(s+t)/N,G_st= η^2N^2 _k=1^N _kc_k^2\,e^-η _k(s+t)/N, where λk _k are NTK eigenvalues and ck=k⊤0c_k= q_k r_0. Since ηλk/N=O(1/(nN))η _k/N=O(1/(nN)), all temporal vectors (k)s=e−ηλks/N( u_k)_s=e^-η _ks/N are approximately constant: k≈ u_k≈ 1 for all k. Thus ≈(η20⊤00/N2)⋅⊤ G≈(η^2 r_0 K_0 r_0/N^2)· 1 1 is rank 1. The signal hierarchy is flat and the spectral edge thesis is vacuous. Proposition 31.3 (Feature-Learning Regime: Signal Hierarchy from NTK Outliers). When η=O(1)η=O(1) (μ ), the temporal vectors (k)s=e−ηλks/N( u_k)_s=e^-η _ks/N are no longer parallel for NTK modes with ηλkW/N≳1η _kW/N 1. These “active” modes produce distinguishable temporal patterns in the trajectory. The signal rank is: k∗=#k:ηλkW/N≳1 and ck≠0,k^*=\#\k:η _kW/N 1 and c_k≠ 0\, (88) recovering the Krylov bound (Proposition 5.5) from the Gram matrix structure. The signal strength of mode j associated with NTK eigenvalue λk(j) _k(j) is: dj∼ηN|ck(j)|λk(j)⋅∥k(j)∥.d_j\; \; ηN c_k(j) _k(j)· u_k(j) . (89) The spectral gap R=dk∗/dk∗+1R=d_k^*/d_k^*+1 is controlled by the outlier–bulk transition in the NTK eigenvalue spectrum: the Hessian gap λK≫λK+1 _K _K+1 translates to a trajectory gap dK≫dK+1d_K d_K+1. Phase transitions occur when NTK eigenvalues cross the activation threshold N/(ηW)N/(η W). Remark 31.4 (What This Resolves). This calculation is consistent with Conjecture 1 of the original draft and suggests a resolution: 1. The signal rank k∗k^* is determined by the number of NTK/Hessian outlier eigenvalues above N/(ηW)N/(η W)—confirming the Krylov bound. 2. The signal strengths djd_j are deterministic in the n→∞n→∞ limit, given by (89). 3. The spectral gap arises from the outlier–bulk structure of the NTK spectrum, not from noise. The kernel regime (η=O(1/n)η=O(1/n)) produces a degenerate rank-1 Gram matrix with no gap and no phase transitions. The μ regime (η=O(1)η=O(1)) produces a Gram matrix with k∗≥2k^*≥ 2 whose structure reflects the loss landscape Hessian. This provides the width-independent foundation for the spectral edge thesis. 32 Connection to Roberts–Yaida–Hanin The Roberts–Yaida–Hanin (RYH) perturbative framework [20] develops a systematic 1/n1/n expansion for deep networks. The connection to the spectral edge framework is conceptual: RYH provides the initial conditions (NTK spectrum, kernel evolution rate) that the spectral edge dynamics then evolve, and RYH’s dNTK (differential of the NTK) is the natural candidate for K˙ K in the evolving-NTK framework (Section 34). We state here only what follows directly from standard RYH results; the detailed mapping of RYH quantities to spectral edge quantities is deferred to the companion notes [38]. Remark 32.1 (NTK Spectrum from Architecture). A central output of the RYH framework is the NTK at initialisation: the eigenvalues λk _k are computable from the O(1)O(1) kernel recursion K(l+1)=Cb+CW⟨σ(h1)σ(h2)⟩K(l).K^(l+1)=C_b+C_W σ(h_1)σ(h_2) _K^(l). These eigenvalues determine the signal decay rates ηλk/Nη _k/N, the signal rank k∗k^*, and the activation threshold N/(ηW)N/(η W) that appear throughout this paper. Remark 32.2 (Criticality and Non-Vacuous Spectral Edge). RYH criticality (χ⟂=CW⟨(σ′)2⟩=1 _ =C_W (σ )^2 =1) ensures Θ(L)=O(L) ^(L)=O(L), so NTK eigenvalues λk=O(1) _k=O(1) and the activation condition ηλkW/N≥1η _kW/N≥ 1 is satisfiable—the spectral edge thesis is non-vacuous. Non-critical initialisations produce either vanishing NTK (λk→0 _k→ 0 exponentially in L, giving k∗=0k^*=0) or exploding NTK (λk→∞ _k→∞, giving →∞A→∞ and no stable circuits). Criticality is thus a necessary condition for the spectral edge framework to operate. Why learning happens at the spectral edge. The spectral edge—the boundary between outlier and bulk NTK eigenvalues—is special for three reinforcing reasons: 1. Detectability. Below the edge, the signal-to-noise ratio is <1<1 and the mode is undetectable in the Gram matrix. Above it, SNR >1>1. New learning becomes visible when a mode crosses the edge. 2. Minimal gap. The gap gλ=λk∗−λk∗+1g_λ= _k^*- _k^*+1 at the edge is the smallest eigenvalue spacing. The mixing rate Γjk∝1/gλ _jk 1/g_λ is therefore largest there—phase transitions concentrate at the edge. 3. Adiabatic breakdown. The adiabatic parameter =∥K˙∥/(ηgλ2)A= K /(η g_λ^2) is largest where gλg_λ is smallest. Interior modes are adiabatically protected; the edge is where protection fails and circuits reconfigure. Complementarity. RYH provides the spectral inputs (NTK eigenvalues and their architecture dependence) that the spectral edge framework then evolves through training: RYH⏟architecture→NTK spectrum+Spectral edge⏟NTK spectrum→training dynamics=Architecture→Dynamics. RYH_architecture\,→\,NTK spectrum\;\;+\;\; Spectral edge_NTK spectrum\,→\,training dynamics\;\;=\;\;Architecture\,→\,Dynamics. 33 The Three-Phase Pattern Our empirical data reveals a universal three-phase pattern in the gap ratio, consistent across all seeds and models (see Section 21–Section 29 for full empirical details): Remark 33.1 (Empirical Observation: Three-Phase Pattern). In the TinyStories experiment (4 seeds: 42, 123, 149, 256) and GPT-2 124M pretraining, the gap ratio R(t)=σk∗/σk∗+1R(t)= _k^*/ _k^*+1 follows three empirically consistent phases: Phase Steps (TinyStories) Gap ratio R Val-loss Rise 1000–5000 Rising (gap opening) 70–80% improvement Plateau 5000–7000 Stable high Slow improvement Collapse 7000–9000 Falling to ∼1 1 Stabilisation The collapse onset is remarkably consistent: step ∼7500±300 7500± 300 across 4 seeds (42, 123, 149, 256). For GPT-2 124M, the three-phase pattern is modulated by a k∗k^* shift (Section 22.1): as the gap position migrates from k∗=3k^*=3 to k∗=2k^*=2 after the distribution shift, the collapse of one gap coincides with the opening of another at a different position. Interpretation in our framework: • Rise phase: The dominant signal direction rapidly gains strength (driven by large |Gk∗||G_k^*| and favorable curvature), opening a gap above the subdominant modes. • Plateau phase: The gap is near steady state (dg/dt≈0dg/dt≈ 0). Learning proceeds at a stable rate. • Collapse phase: The curvature asymmetry shifts (the dominant direction has “used up” its gradient projection, |Gk∗|→0|G_k^*|→ 0), and the gap closes. The subspace destabilizes, and the loss improvement saturates. Remark 33.2 (Cross-Correlation Evidence). The phase-specific cross-correlation between R(t)R(t) and val-loss confirms the framework (Table 4): • Overall gap–loss correlation: |r|=0.66|r|=0.66–0.670.67 for both models. • Collapse-phase correlation: |r|=0.864±0.059|r|=0.864± 0.059 (4 seeds). • Derivative-segmentation correlation: |r|=0.937±0.019|r|=0.937± 0.019. • W-dependent lag flip: W=10→W=10→ lag =−1=-1 (val-loss leads); W=20→W=20→ lag =+1=+1 (gap leads). The lag flip is predicted by the framework: larger W averages over more steps, making the gap a leading indicator. Remark 33.3 (Grokking as Rise-Only Dynamics). The Dyck/SCAN grokking experiments (Section 27) exhibit a variant of the three-phase pattern: the rise phase is present (gap opening from R≈1R≈ 1 to R→103R→ 10^3), but there is no plateau or collapse because weight decay continuously drives subdominant singular values toward zero. The grokking transition corresponds to the onset of the rise phase, and the gap continues growing indefinitely—a limiting case where the collapse phase never occurs. 34 The Evolving NTK and Circuit Transport The flow equations of Section 12–Section 14 assume a fixed NTK. In the feature-learning regime (μ , η=O(1)η=O(1)), the NTK K(t)=∑kλk(t)k(t)k(t)⊤K(t)= _k _k(t)\, q_k(t)\, q_k(t) evolves, and its evolution drives both the creation and destruction of feature circuits. Theorem 34.1 (Signal Dynamics with Evolving Kernel). Define signal coefficients cj(t)=j(t)⊤(t)c_j(t)= q_j(t) f(t) and target projections yj∗(t)=j(t)⊤y_j^*(t)= q_j(t) y. Then: c˙j=−ηλj(t)(cj−yj∗)+∑k≠jk⊤K˙jλj−λkck. c_j=-η _j(t)(c_j-y_j^*)+ _k≠ j q_k K\, q_j _j- _k\;c_k. (90) The second term is the injection: it transfers signal from mode k into mode j at rate Γjk=k⊤K˙j/(λj−λk) _jk= q_k K\, q_j/( _j- _k). The eigenvalues evolve as λ˙j=j⊤K˙j λ_j= q_j K\, q_j (Hellmann–Feynman). Sketch. Differentiate cj=j⊤c_j= q_j f, use ˙=−ηK(t)(−) f=-η K(t)( f- y) for the function evolution, and apply first-order perturbation theory for the eigenvector rotation ˙j=∑k≠j[k⊤K˙j/(λj−λk)]k q_j= _k≠ j[ q_k K\, q_j/( _j- _k)]\, q_k. Full derivation in the companion notes [38]. ∎ The injection term Γjk _jk diverges when eigenvalues near-cross (λj≈λk _j≈ _k)—this is the spectral edge phase transition of Section 8, now derived from the NTK dynamics rather than postulated. Definition 34.2 (Adiabatic Parameter). The adiabatic parameter of the training trajectory is: (t)=‖K˙(t)‖opηgλ(t)2,A(t)= \| K(t)\|_opη\,g_λ(t)^2, (91) where gλ=λk∗−λk∗+1g_λ= _k^*- _k^*+1 is the NTK spectral gap. Proposition 34.3 (Adiabatic Circuit Preservation (Heuristic)). If (s)≤AmaxA(s)≤ A_ for all s∈[0,T]s∈[0,T], then every signal circuit j≤k∗j≤ k^* is expected to have cumulative eigenvector rotation of order O(AmaxηT)O(A_ \,η\,T) and circuit lifetime of order π/(4Amaxη)π/(4A_ \,η). Full derivation is in the companion notes [38]. This yields three regimes: A Regime Behaviour ≪1 1 Adiabatic Circuits preserved (plateau) ∼1 1 Non-adiabatic Circuits transform (phase transition) ≫1 1 Strongly non-adiabatic Circuits destroyed (forgetting) The spectral gap plays the role of the energy gap in quantum mechanics: it protects feature circuits against perturbation. The complete circuit lifecycle is birth(∼1)→transport(≪1)→death(∼1)birth\;(A 1) \;(A 1) \;(A 1). 35 Overparameterisation and the Spectral Gap Overparameterisation (P≫NP N) serves three distinct roles, each tied to a spectral quantity: 1. Spectral richness. When P≥NP≥ N, the NTK K=JJ⊤K=J is full rank, so every target direction is reachable in principle. But full rank does not imply learnability in finite time: tiny eigenvalues require astronomically many steps. The effective capacity is k∗k^*—the number of modes above the activation threshold—not rank(K)rank(K). Overparameterisation enriches the top end of the spectrum, increasing k∗k^*. This is necessary but explains nothing about why P≫NP N helps. 2. Gap protection. In μ , the kernel evolution rate ‖K˙‖\| K\| is O(1)O(1) regardless of width. But the spectral gap grows with the overparameterisation ratio γ=P/Nγ=P/N: gλ∼O(γ)g_λ O( γ). Therefore ∼O(1)ηgλ2∼1γ→0asγ→∞.A O(1)η\,g_λ^2 1γ→ 0 \;γ→∞. (92) More parameters ⇒ larger gap ⇒ more adiabatic ⇒ circuits survive longer. 3. Feature reservoir. More neurons ⇒ denser coverage of weight space ⇒ for any target feature direction, neurons near the relevant decision boundary exist and can be recruited. This accelerates feature emergence via the inter-mode injection Γjk=k⊤K˙j/(λj−λk) _jk= q_k K\, q_j/( _j- _k) of Theorem 34.1, which grows as modes approach the spectral edge. At the interpolation threshold (P≈NP≈ N, γ→1γ→ 1), the NTK eigenvalues crowd together, gλ→0g_λ→ 0, →∞A→∞, and no circuit survives transport—this is the peak of double descent. Beyond the threshold, the gap reopens, A drops, and the second descent begins: performance improves because circuits persist. The critical width for reliable learning is ncrit∼C2/g∞2n_crit C^2/g_∞^2, where g∞g_∞ is the infinite-width gap limit and O(1/n)O(1/ n) are the finite-width gap fluctuations. 36 Scaling Laws from the Spectral Tail The spectral edge framework is consistent with empirical scaling laws. Under the following standard assumptions—which we state explicitly but do not verify here—the framework recovers the observed power-law exponents. This is a consistency check, not a derivation: the assumptions may or may not hold for a given architecture and dataset, and we refer to Bahri et al. [25] for a careful treatment. Assumptions: NTK eigenvalues decay as λk∼k−p _k k^-p; target projections as |yk∗|2∼k−q|y_k^*|^2 k^-q; residual–feature correlations as ρk2∼k−s _k^2 k^-s. Proposition 36.1 (Compute Scaling, Constant Kernel). Mode k is learned when ηλkT≫1η _kT 1, giving k∗(T)∼(ηT)1/pk^*(T) (η T)^1/p active modes by time T. The residual loss is the tail integral: L(T)∼∑k>k∗|yk∗|2∼k∗−(q−1)∼T−(q−1)/p.L(T) _k>k^*|y_k^*|^2 k^*\,-(q-1) T^-(q-1)/p. (93) Proposition 36.2 (Width Scaling). If width n provides ∼nr n^r NTK eigenvalues above the activation threshold, then at fixed training time: L(n)∼n−r(q−1).L(n) n^-r(q-1). (94) Proposition 36.3 (Staircase Envelope, Evolving Kernel). If the residual–feature correlations decay as ρk2∼k−s _k^2 k^-s, the inter-transition times grow as Δtk∼1/ρk2∼ks t_k 1/ _k^2 k^s, and the loss envelope becomes: L(T)∼T−(q−1)/(s+1).L(T) T^-(q-1)/(s+1). (95) The exponent s is the “feature discovery tax”: architectures with better inductive biases (smaller s) have steeper scaling laws. Remark 36.4 (Staircase as Spectral Edge Events). Independently of the power-law assumptions above, the staircase shape of the loss curve has a direct mechanistic reading within the spectral edge framework: each step corresponds to one gap opening cycle (Section 33), with the plateau being the steady-state phase and the drop being the capability gain at gap opening. No additional assumptions are required for the staircase shape itself. The exponents p, q, s are in principle measurable from the eigenvalue spectrum and target projections. The empirical scaling laws of Kaplan et al. [23] and Hoffmann et al. [24] are consistent with this spectral picture. 37 Summary of the Three Answers I. Where is the spectral gap (k∗k^*)? k∗(t)=argmax1≤j≤W−1σj(t)σj+1(t)k^*(t)= *arg\,max_1≤ j≤ W-1 _j(t) _j+1(t) (Answer I) The spectral gap position is the location of the maximum consecutive singular value ratio within the signal hierarchy. It separates dominant modes (backbone) from subdominant modes. It is not the signal-noise boundary (which is trivially at j=Wj=W in our regime). Practically: compute all W−1W-1 singular value ratios and find the maximum. Any ratio R>1.05R>1.05 at the gap position indicates genuine structure. I. How does the gap evolve? dgdt=−η(hk∗−hk∗+1)d¯−η(h¯+ω)⋅g+ηW(|Gk∗eff|2dk∗−|Gk∗+1eff|2dk∗+1) dgdt=-η(h_k^*-h_k^*+1) d-η( h+ω)· g+η W\! ( |G_k^*^eff|^2d_k^*- |G_k^*+1^eff|^2d_k^*+1 ) (Answer I) The gap is driven by curvature asymmetry and gradient alignment asymmetry, damped by the average curvature plus weight decay. Weight decay ω strengthens the damping (making gaps harder to sustain) while simultaneously compressing subdominant modes through the effective driving Gjeff=⟨j,∇L⟩+ω⟨j,⟩G_j^eff= v_j,P∇ L +ω v_j, θ . Subspace stability is controlled by the Davis–Kahan bound (Corollary 9.2): sin∠(^j,j)≤∥F/[(dk∗+dk∗+1)g] ( v_j, v_j)≤ E _F/[(d_k^*+d_k^*+1)g]. Small gap ⇒ unstable subspace ⇒ unreliable learning. I. What are the flow equations? ddj2dt≈−2η(hj+ω)dj2+η2W(Sj+2Gjeffj)(Remark 12.10)k∗(t)=argmaxjdjdj+1g(t)=dk∗(t)−dk∗+1(t)dLvaldt=−η∑j=1WαjGjeff,trainGjval cases d_j^2dt≈-2η(h_j+ω)\,d_j^2+η^2W(S_j+2G_j^effN_j)&(Remark~ rem:two-term)\\[10.0pt] k^*(t)= *arg\,max_j d_jd_j+1\\[10.0pt] g(t)=d_k^*(t)-d_k^*+1(t)\\[10.0pt] dL_valdt=-η _j=1^W _j\,G_j^eff,train\,G_j^val cases (Answer I) Phase transitions occur when g(t)=0g(t)=0. The dynamics are controlled by: • hj=j⊤j≈λk(j)/Nh_j= v_j H v_j≈ _k(j)/N: Hessian curvature along direction j (Proposition 31.3). • ω: weight decay, acting as a curvature floor that compresses low-curvature modes (Remark 12.22). • GjeffG_j^eff: effective driving, combining gradient projection and weight decay source. • αj _j: stability coefficient (controls how reliably direction j contributes to learning). • β2 _2: Adam’s second-moment coefficient, which reshapes the curvature hierarchy through the preconditioner P. In the NTK decomposition (Proposition 31.1–Proposition 31.3), the signal strengths have the exact form dj=(η/N)|ck(j)|λk(j)Φ(ηλk(j)/N,W)d_j=(η/N)|c_k(j)| _k(j)\, (η _k(j)/N,W) and the gap collapse time is t∗=N/[η(λk∗−λk∗+1)]⋅log(⋯)t^*=N/[η( _k^*- _k^*+1)]· (·s), where the logarithmic factor depends on the initial residual projections and NTK eigenvalues. 38 Connection to the Lottery Ticket Hypothesis The Lottery Ticket Hypothesis (LTH; Frankle & Carbin [30]) states that a randomly-initialised, overparameterised network contains a sparse subnetwork—the “winning ticket”—that, trained from the same initialisation, matches the full network’s performance. Every major LTH phenomenon has a spectral edge explanation. The winning ticket is the signal subspace. Write =∑j≤k∗djjj⊤+⟂ W= _j≤ k^*d_j\, u_j v_j + W_ . Magnitude pruning keeps the largest entries of W. Signal-bearing weights scale as ∼dj/mn d_j/ mn; bulk weights scale as ∼dk∗+1/mn d_k^*+1/ mn. The pruning signal-to-noise ratio is the gap ratio: SNRprune∼dk∗/dk∗+1=R.SNR_prune d_k^*/d_k^*+1=R. (96) When R≫1R 1, magnitude pruning naturally preserves signal and removes noise. Pruning fidelity from Davis–Kahan. By the sinΘ theorem (Corollary 9.2), the signal subspace distortion from a pruning perturbation satisfies sin∠(~j,j)≤‖F/g ( u_j, u_j)≤\| \|_F/g. The critical sparsity—where the perturbation overwhelms the gap—satisfies sj∗≈gλ(j) 2/‖op2.s^*_j≈ g_λ^(j)\,2/\| J\|_op^2. (97) Each circuit has its own critical sparsity. The dominant circuit (j=1j=1, largest gap) survives the most aggressive pruning; the marginal circuit (j=k∗j=k^*, gap =gλ=g_λ) breaks first. Wider networks (larger gap, by (92)) tolerate higher sparsity—exactly as observed empirically. Why re-initialisation fails. The mask M preserves the entries of (0) W(0) that carry the initial projections ck(0)=⟨ψk,(0)⟩c_k^(0)= _k, W(0) , which determine which NTK modes activate. Random re-initialisation gives different ck(0)′c_k^(0) , breaking the co-adaptation between mask and initialisation. Late resetting = first phase transition. Frankle et al. [30] showed that rewinding to step t0>0t_0>0 outperforms rewinding to step 0. In the spectral edge framework, the optimal t0t_0 is the time of the first gap opening—before this, R(t)≈1R(t)≈ 1 and the mask cannot distinguish signal from noise; after this, the first circuit is established at high SNR. IMP as iterative spectral thresholding. Each round of iterative magnitude pruning (train → prune → rewind → repeat) amplifies the signal (djd_j grows), thresholds the noise (small weights removed), and resets (rewind preserves mask geometry). This is the spectral analogue of iterative hard thresholding from compressed sensing, with the spectral gap playing the role of the restricted isometry constant. See the companion notes [38] for the full derivation, including the strong lottery ticket conjecture, quantitative predictions, and connections to the three roles of overparameterisation. 39 Connection to the Holographic Encoding Principle Xu (2025) discovered a duality in grokked solutions: the training trajectory is globally low-rank (3–5 trajectory PCs recover >>95% accuracy), while individual weight matrices are locally full-rank (per-matrix SVD at rank 64 gives chance-level performance). This “holographic encoding” is consistent with the spectral edge framework. The NTK has k∗k^* outlier eigenvalues. Gradient descent filters the loss through these k∗k^* modes, confining the trajectory to a k∗k^*-dimensional subspace in function space—hence global low-rank. But the Jacobian J maps each function-space eigenmode k q_k to a dense parameter-space vector †k J q_k that fills every weight matrix—hence local full-rank. Proposition 39.1 (Holographic Encoding). Let kk=1k∗\ q_k\_k=1^k^* be the active NTK eigenmodes. Then: 1. The training trajectory lies in span⊤kk=1k∗⊂ℝPspan\ J q_k\_k=1^k^* ^P (dimension k∗k^*). 2. For generic J, the projection of this subspace onto each weight matrix l W_l has effective rank min(dl,dl+1) (d_l,d_l+1) (full rank). 3. Trajectory PCA recovers span⊤kspan\ J q_k\; per-matrix SVD cannot. Weight decay amplifies the effect by pushing task-critical information from the top of each matrix’s SVD spectrum into the spectral tail (the signal flow equation (38) with ω>0ω>0 suppresses the dominant modes, redistributing energy to the tail). The scaffold/refinement hierarchy of trajectory PCs (PC 1–2 for coarse structure, PC 3–5 for fine structure) maps directly onto the sequence of spectral edge phase transitions. Circuit separability in activation space (selectivity index >0.96>0.96) coexists with weight-level entanglement because † J preserves function-space orthogonality but not parameter-space orthogonality. See the companion notes [38] for the full analysis. 40 Connection to the Edge of Stability The edge of stability (Cohen et al. [14]) is the empirical observation that during gradient descent with step size η, the largest Hessian eigenvalue λmax _ evolves toward 2/η2/η and hovers there. The spectral edge framework and the edge of stability are consistent perspectives on related phenomena. The identification. The Gauss–Newton Hessian has eigenvalues λkNTK/N _k^NTK/N. At the edge of stability, λ1NTK/N≈2/η _1^NTK/N≈ 2/η, i.e., ηλ1/(2N)≈1η _1/(2N)≈ 1. This is of the same order as the activation condition (88) with window size W=2W=2 (i.e., two consecutive steps): both mark the transition where the top Hessian mode enters the strongly-activated regime. Remark 40.1 (Edge of Stability and Spectral Edge Activation). The following conditions are consistent with one another, up to order-of-magnitude factors: 1. λmax≈2/η _ ≈ 2/η (edge of stability). 2. The top NTK eigenvalue satisfies ηλ1/(2N)≈1η _1/(2N)≈ 1 (spectral edge activation with W=2W=2). 3. The Gram matrix has k∗≥1k^*≥ 1 with d1=O(1)d_1=O(1) (non-trivial signal above the bulk). All three mark the same qualitative transition: the top Hessian mode crossing from the weakly-activated to the strongly-activated regime. Implicit regularisation. Damian et al. [15] and Barrett & Dherin [13] showed that GD with step size η implicitly minimises Leff(θ)=L(θ)+η4‖∇L(θ)‖2.L_eff(θ)=L(θ)+ η4\|∇ L(θ)\|^2. (98) The critical points of LeffL_eff satisfy either ∇L=0∇ L=0 or λmax=2/η _ =2/η. The edge of stability is a critical point condition of the implicit loss. What the spectral edge framework adds. The edge of stability gives one number (λmax≈2/η _ ≈ 2/η) and one phenomenon (the top eigenvalue tracks this threshold). The spectral edge framework adds the full spectral structure: Aspect Edge of stability Spectral edge Observable λmax _ (1 number) dj\d_j\, g(t)g(t), αj _j Mode count Top mode only All k∗k^* modes + hierarchy Dynamics Equilibrium Phase transitions, gap flow Generalisation Not addressed Loss decomposition per mode Why k∗k^* is small. The edge of stability provides a mathematical explanation for the empirically small k∗k^*. If k eigenvalues simultaneously exceed 2/η2/η, the GD step overshoots in k directions. For k>kcritk>k_crit (typically 2–4), the nonlinear interactions between overshoots destabilise the self-correction mechanism. The system self-organises to keep exactly k∗k^* modes at the edge. Remark 40.2 (Empirical Observation: Small k∗k^*). Across all models studied (k∗∈2,3k^*∈\2,3\ for GPT-2 124M and TinyStories), the number of simultaneously active modes is small. A plausible explanation is that when k>kcritk>k_crit modes simultaneously overshoot, nonlinear interactions destabilise the self-correction mechanism of GD, so the system self-organises to keep only a few modes at the edge. This would make k∗k^* a property of the optimizer rather than the task or architecture. We do not have a proof of this; it remains an open question. Sequential phase transitions. Combining both frameworks: training proceeds as sequential passages through the edge. Each time a bulk eigenvalue reaches 2/η2/η, a new mode joins the active set (k∗k^* increases by 1), the spectral gap collapses (∼1A 1), circuits mix and reconfigure, and a new plateau begins. The staircase structure of loss curves (Olsson et al. [28]) is analogous to this sequence of spectral edge events. See the companion notes [38] for the full derivation, including the derivation of Remark 40.1 and the connection to the implicit regularisation of Damian et al. [15]. 41 Circuit Survival and the Final Model The spectral edge framework describes the birth of circuits (phase transitions at ∼1A 1) and their transport (adiabatic plateaus at ≪1A 1). But a trained model is a finished product: only circuits that survive to the end appear in the final model. Three fates. After birth, a circuit faces three possible outcomes: 1. Survival: the gap g(j)g^(j) stays large, so j≪1A_j 1 throughout. The circuit persists to the final model. 2. Destruction by phase transition: a new bulk eigenvalue approaches λk∗ _k^*, collapsing the edge gap. The edge circuit (j=k∗j=k^*) is destroyed; interior circuits (j<k∗j<k^*) are protected by their larger gaps. 3. Destruction by weight decay: WD compresses the signal strength. If hj<ωh_j<ω (curvature below WD), the circuit is compressed below the detection threshold and absorbed into the bulk. This is how memorisation circuits die. Proposition 41.1 (Circuit Survival Criterion (Heuristic)). Under [,]≈0[P, H]≈ 0. A circuit j is expected to survive to the final model when all of: (i) gap protection: j(t)≤AmaxA_j(t)≤ A_ for all remaining t; (i) WD balance: hj>ωh_j>ω; (i) continued injection: |Gjeff|>0|G_j^eff|>0. Full derivation is in the companion notes [38]. The stability hierarchy α1≥α2≥⋯≥αk∗ _1≥ _2≥·s≥ _k^* determines the survival order: the dominant circuit (largest gap, highest α) is the most robust; the edge circuit (smallest gap, lowest α) is the most fragile and first to break under perturbation. The final model is the set of survivors—training is Darwinian selection where fitness equals spectral gap. How few events produce many circuits. A pretrained GPT-2 has many identifiable circuits (induction heads, IOI circuits, etc.), yet we observe k∗=2k^*=2–33 and only ∼2 2 spectral edge events in the observation window. The resolution is the holographic encoding (Section 39): each event is rank-1 in function space but maps through † J to a dense reorganisation of the entire parameter space. Mechanistic interpretability decomposes this one dense change into multiple “circuits” because it analyses the computational graph, not function space. The complexity of the final model comes from the dimensionality of parameter space (P∼108P 10^8), not from the number of spectral edge events. Additionally, many events occurred before the observation window—the Gram matrix is a local diagnostic, not a historical record. See the companion notes [38] for the full treatment, including the interior circuit protection theorem, grokking as circuit replacement, and the degenerate perturbation theory for eigenvalue clusters. 42 The Geometric Flow Connection The NTK K(x,x′)=∇θf(x)⊤∇θf(x′)K(x,x )= _θf(x) _θf(x ) defines a metric on function space (more precisely, a bilinear form on the tangent space of the function manifold at f). Training in the feature-learning regime evolves this metric: K(t)→K(t+dt)K(t)→ K(t+dt). This is a geometric flow: a one-parameter family of metrics evolving according to a differential equation. The flow is not purely intrinsic (unlike Ricci flow, where the evolution is determined by the metric alone), but it is closer to intrinsic than it may first appear. The Hessian curvatures hjh_j are essentially the NTK eigenvalues via the Gauss–Newton approximation (hj≈λj/Nh_j≈ _j/N), and the gradient projections GjG_j depend on the NTK eigenbasis plus the residuals ri=f(xi)−yir_i=f(x_i)-y_i. The only information beyond the metric is the residuals, which encode the current state of learning relative to the target. This makes the spectral edge flow a nearly intrinsic geometric flow: the metric plus a scalar field (the residual) determines the evolution. This is the normal situation for flows in nonlinear PDE (e.g., reaction-diffusion on manifolds, where the metric governs diffusion and the scalar field drives reaction). The key structural features of geometric flows—singularity formation controlled by curvature, classification of singularity types—still apply. We do not claim that the NTK flow is a Ricci flow or that existing theorems from geometric analysis apply directly. Rather, we observe that several structural features of the spectral edge framework have natural analogues in geometric flows, which we record as motivation for future investigation: Geometric flow Spectral edge Status Metric gij(t)g_ij(t) NTK K(x,x′;t)K(x,x ;t) Exact Curvature controls singularity Gap controls phase transition Proved (Theorem 14.2) Singularity == topology change Gap collapse == k∗k^* change Supported by Propositions 10.5 and 17.1 Singularity classification k∗≤3k^*≤ 3 empirically Empirical + heuristic Monotone quantity L+(η/4)‖∇L‖2L+(η/4)\|∇ L\|^2 (Damian et al.) Known result, not ours The “Status” column is important. The first two rows are theorems within our framework. The singularity classification (k∗≤3k^*≤ 3) is an empirical observation supported by the edge-of-stability heuristic (Section 40), not a theorem. The monotone quantity is a known result from optimisation theory [15], not something we proved. What k∗=2k^*=2–33 does and does not mean. The constraint k∗=2k^*=2–33 limits the number of simultaneous singularities (gap collapses) in the NTK flow, not the number of features learned per event. Each spectral edge event corresponds to a rank-1 change in function space—a single NTK eigenmode qkq_k crossing the edge. However, the Jacobian J†J maps this function-space direction to a dense direction in parameter space (P∼108P 10^8 dimensions), coherently reorganising weights across every layer and head simultaneously (cf. [34]). The number of “circuits” (in the mechanistic interpretability sense) produced by a single event depends on the Jacobian structure, not on k∗k^*. Floer-theoretic structure (speculative). The Ricci flow analogy is incomplete: Ricci flow evolves autonomously (∂tg=−2Ric _tg=-2\,Ric), while the NTK evolution depends on the optimizer. A potentially more appropriate analogy is Floer-theoretic: spectral edge events play the role of critical points; adiabatic transport between events provides flow lines; and the trained model’s capabilities are an invariant constructed from the pattern of events. This suggests a conjecture: Conjecture 42.1 (Informal). The capabilities of the trained model are invariant under continuous deformations of the optimizer, provided these deformations preserve the ordering and type of spectral edge events. We emphasise that this is speculation—we have not constructed a chain complex, verified ∂2=0∂^2=0, or computed any homology groups. The Floer analogy is recorded here as a direction for future work, not as a result. Making it rigorous would require, at minimum: 1. A precise definition of the “moduli space” of spectral edge events (analogous to the moduli space of pseudo-holomorphic curves). 2. A compactness theorem for this moduli space (analogous to Gromov compactness). 3. A proof that the resulting algebraic structure is independent of auxiliary choices (the optimizer, the learning rate schedule). Preliminary evidence from the Muon experiment (Open Problem 8 below) is suggestive: AdamW and Muon produce different spectral edge dynamics (k∗=2k^*=2 vs. k∗=1k^*=1, R≈10R≈ 10–3030 vs. R≈3R≈ 3–55) yet reach models with comparable capabilities. The flow differs; the invariant appears preserved. See the companion notes [38] for further discussion of the circuit chain complex and its conjectural properties. 43 Logical Structure and the Commutativity Assumption Figure 5 shows the complete logical dependency graph of the framework. Results are coloured by their dependence on the commutativity assumption [,H]≈0[P,H]≈ 0: • Clean (green): no dependence. The diagnostic framework (definitions of k∗k^*, R, αj _j), the spectral stability results (Davis–Kahan, eigenvalue repulsion), the loss decomposition, the evolving-NTK theory, the adiabatic theorem, scaling laws, and holographic encoding are all independent of the commutativity assumption. • Weak (yellow): the qualitative conclusion survives without [,H]≈0[P,H]≈ 0; only the exact ODE coefficients require it. This includes the signal flow ODE (some directions decay faster than others—generically true), the gap flow (the gap evolves—generically true), the three-phase pattern (rise–plateau–collapse—observed regardless of optimizer), and the circuit survival criterion (hj>ωh_j>ω—qualitatively clear even if the exact threshold shifts). • Strong (orange): both the structure and the quantitative conclusion require [,H]≈0[P,H]≈ 0. This includes the Krylov bound on k∗k^* (the exact Krylov subspace structure W(I−ηH,g0)K_W(I- ,Pg_0) requires commutativity), the collapse/opening times (exact timing from critical dynamics), and the β2 _2 theorem (the preconditioner enters explicitly). For SGD (=IP=I), commutativity is exact and all results hold. For Adam with β2→1 _2→ 1, it is approximately satisfied. For Muon, it is the most questionable—see Open Problem 8 below. Figure 5: Dependency graph of the spectral edge framework. Green = no dependence on [,H]≈0[P,H]≈ 0. Yellow (dashed) = qualitative conclusion survives; only exact coefficients need commutativity. Orange (thick dashed) = both structure and conclusion require [,H]≈0[P,H]≈ 0. The core diagnostic and learning-theoretic results (left column and NTK branch) are entirely clean. The flow equations are weakly dependent. Only the Krylov bound, collapse/opening times, and β2 _2 theorem have strong dependence. 44 Open Problems 1. Null Distribution of Intra-Signal Gap Ratios. Derive the exact distribution of maxj(σj/σj+1) _j( _j/ _j+1) under structured signal models (not just the isotropic null). This would provide rigorous significance tests for the gap position. 2. Scaling of k∗k^* with Model Size and Optimizer. Our empirical data shows k∗∈2,3k^*∈\2,3\ for models of size 51M–124M under AdamW. The Muon experiment (Section 44) provides direct evidence that k∗k^* is at least partly optimizer-dependent: Muon’s spectral norm equalisation drives k∗=1k^*=1, whereas AdamW produces k∗=2k^*=2 on the same TinyStories 51M model, while both reach comparable validation loss. Whether k∗k^* remains small at billion-parameter scale, and how it varies across optimizers and architectures, are open empirical questions. The Krylov bound (Proposition 5.5) suggests k∗k^* depends on the number of Hessian outliers rather than p directly, but this has not been tested beyond 124M parameters. 3. Multi-Scale Phase Transitions. In a multi-layer network, each layer has its own spectral gap gl(t)g_l(t). The temporal ordering of gap collapses across layers should reveal the “causal chain” of phase transitions. Can we derive the ordering from the layer-wise recursion? 4. Prediction of Grokking Time. Given the loss landscape and optimizer parameters, can we predict when grokking occurs by solving the gap flow equation for the opening time (Corollary 14.6)? 5. Intervention (partially resolved). The causal intervention test (Section 28) confirms that the per-direction loss decomposition ranks signal directions correctly (ρ=0.75ρ=0.75) when the window is large enough (W≥30W≥ 30). Open sub-problems: (a) extending this to non-grokking tasks (language modelling), (b) deriving the minimum window size for reliable αj _j estimation as a function of the spectral gap. 6. Adiabatic Bounds for Continual Learning. Proposition 34.3 gives circuit preservation guarantees when ≤AmaxA≤ A_ . Can this be used to design continual learning algorithms with provable circuit stability? 7. Pruning Phase Transition. Section 38 predicts that circuits break under progressive pruning in reverse stability order (dk∗d_k^* first, then dk∗−1d_k^*-1, etc.), with critical sparsity sj∗∝gλ(j) 2s^*_j g_λ^(j)\,2. Can this be verified empirically by tracking Gram matrix singular values during progressive pruning? Does the optimal IMP rewind point correlate with the first spectral edge event? 8. Quantitative edge-of-stability predictions. Section 40 establishes the qualitative identification between the edge of stability and the spectral edge. Open sub-problems: (a) derive the precise dynamics of k∗k^* at the edge—what determines kcritk_crit as a function of architecture? (b) test the prediction k∗≤kcritk^*≤ k_crit empirically across model scales. (c) connect the implicit regularisation (98) quantitatively to the gap flow equation. 9. Preconditioned optimisers: the Muon test. The Hessian–trajectory gap correspondence (Prop. 5.1) requires commutativity [,H]≈0[P,H]≈ 0. For SGD this is exact; for Adam it is approximate. The Muon optimiser (momentum + orthogonalisation via Newton–Schulz [31]) has a nonlinear, state-dependent preconditioner P(m)=(mm⊤)−1/2P(m)=(m )^-1/2 that equalises spectral norms, partially counteracting the Hessian’s role in creating the gap. Preliminary results (TinyStories 51M). We trained a TinyStories 51M model with Muon (lr =5×10−3=5× 10^-3 for 2D weights, wd =0.5=0.5) and compared per-step Gram matrix tracking against the AdamW baseline (lr =10−3=10^-3, wd =0.5=0.5). The findings: • k∗=1k^*=1 for Muon vs. k∗=2k^*=2 for AdamW — the spectral edge dynamics are optimizer-dependent. Muon’s spectral norm equalisation collapses the gap to a single dominant mode. • R≈3R≈ 3–55 for Muon vs. R≈10R≈ 10–3030 for AdamW — the gap ratio is weaker, as predicted. • Both optimizers reach comparable validation loss (∼1.2 1.2) and probe accuracy (>0.8>0.8 OOD). This is consistent with the Floer-theoretic picture (Conjecture 42.1): the flow is optimizer-dependent (k∗k^* and R differ), but the invariant (the capabilities of the trained model) is preserved. Different optimizers trace different paths through the spectral landscape, with different numbers of simultaneous active modes, yet arrive at models with the same capabilities—just as different Hamiltonians produce different flow lines but the same Floer homology. Controlled experiments isolating the optimizer effect from the learning rate effect are ongoing. 10. Computing TjkℓT_jk for transformers. The anharmonic coupling tensor Tjkℓ=⟨j,(∇3L)[k,ℓ]⟩T_jk = v_j,\,P(∇^3L)[ v_k, v_ ] (Definition 12.12) is the formal source of injection into the eigenvalue equation (58). For a transformer with cross-entropy loss, ∇3L∇^3L involves the Jacobian of the softmax and the second derivative of the feature map. Open sub-problems: (a) derive TjkℓT_jk in the mean-field limit (tensor programs, n→∞n→∞) and verify that G11>0G_1N_1>0 at steady state (the dominant mode must receive net injection); (b) connect TjkℓT_jk to the evolving-NTK mixing rate Γjk _jk (the notes derive Γjk _jk from the McKean–Vlasov equation; is Γjk _jk the accumulated TjkℓT_jk ?); (c) measure TjkℓT_jk empirically for small transformers and test whether the eigenvalue equation (58) matches observed djd_j trajectories. 11. Perelman-type estimates for NTK flow. The Ricci flow analogy (Section 42) suggests that the NTK flow may satisfy Perelman-type entropy monotonicity and non-collapsing estimates. Proving this would elevate the spectral edge thesis from an analogy to a theorem, with consequences for convergence guarantees and singularity resolution. Closing Remark The spectral edge framework does not claim to supersede or subsume the prior frameworks discussed in this paper. Rather, its consistency with each of them independently—Tensor Programs / Greg Yang’s feature-learning regime (Section 31), Roberts–Yaida–Hanin criticality (Section 32), Dyson Brownian motion (Section 30), the Lottery Ticket Hypothesis (Section 38), the Edge of Stability (Section 40), and empirical scaling laws (Section 36)—is itself evidence that the spectral gap of the rolling-window Gram matrix is the right object to study. A single quantity that is simultaneously consistent with this many independently established results is unlikely to be a coincidence. We view this consilience as the primary argument for the framework, alongside the direct empirical tests of Section 21–Section 29. References [1] J. Baik, G. Ben Arous, and S. Péché. Phase transition of the largest eigenvalue for nonnull complex sample covariance matrices. Ann. Probab., 33(5):1643–1697, 2005. [2] J. Baik and J. W. Silverstein. Eigenvalues of large sample covariance matrices of spiked population models. J. Multivariate Anal., 97(6):1382–1408, 2006. [3] F. Benaych-Georges and R. R. Nadakuditi. The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices. Adv. Math., 227(1):494–521, 2011. [4] R. Couillet and Z. Liao. Random Matrix Methods for Machine Learning. Cambridge University Press, 2022. [5] F. Dyson. A Brownian-motion model for the eigenvalues of a random matrix. J. Math. Phys., 3(6):1191–1198, 1962. [6] N. El Karoui. A rate of convergence result for the largest eigenvalue of complex white Wishart matrices. Ann. Probab., 34(6):2077–2117, 2006. [7] V. A. Marčenko and L. A. Pastur. Distribution of eigenvalues for some sets of random matrices. Math. USSR-Sb., 1(4):457–483, 1967. [8] C. A. Tracy and H. Widom. Level-spacing distributions and the Airy kernel. Comm. Math. Phys., 159(1):151–174, 1994. [9] C. Davis and W. M. Kahan. The rotation of eigenvectors by a perturbation. I. SIAM J. Numer. Anal., 7(1):1–46, 1970. [10] T. Kato. Perturbation Theory for Linear Operators. Springer, 1966. [11] G. W. Stewart and J. Sun. Matrix Perturbation Theory. Academic Press, 1990. [12] J. von Neumann and E. P. Wigner. Über das Verhalten von Eigenwerten bei adiabatischen Prozessen. Phys. Z., 30:467–470, 1929. [13] T. D. Barrett and B. Dherin. Implicit gradient regularization. ICLR, 2021. [14] J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar. Gradient descent on neural networks typically occurs at the edge of stability. ICLR, 2021. [15] A. Damian, E. Nichani, and J. D. Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. ICLR, 2023. [16] B. Ghorbani, S. Krishnan, and Y. Xiao. An investigation into neural net optimization via Hessian eigenvalue density. ICML, 2019. [17] G. Gur-Ari, D. A. Roberts, and E. Dyer. Gradient descent happens in a tiny subspace. arXiv:1812.04754, 2018. [18] C. H. Martin and M. W. Mahoney. Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for training. JMLR, 22(165):1–73, 2021. [19] V. Papyan. Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet Hessians. ICML, 2019. [20] D. A. Roberts, S. Yaida, and B. Hanin. The Principles of Deep Learning Theory. Cambridge University Press, 2022. [21] L. Sagun, U. Evci, V. U. Güney, Y. Dauphin, and L. Bottou. Empirical analysis of the Hessian of over-parametrized neural networks. arXiv:1706.04454, 2017. [22] G. Yang. Tensor programs I–V: Feature learning in infinite-width neural networks. arXiv:2006.14548 et seq., 2020–2023. [23] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv:2001.08361, 2020. [24] J. Hoffmann, S. Borgeaud, A. Mensch, et al. Training compute-optimal large language models. NeurIPS, 2022. [25] Y. Bahri, E. Dyer, J. Kaplan, J. Lee, and U. Sharma. Explaining neural scaling laws. arXiv:2102.06701, 2021. [26] A. Power, Y. Burda, H. Edwards, I. Babuschkin, and V. Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. arXiv:2201.02177, 2022. [27] N. Nanda, L. Chan, T. Lieberum, J. Smith, and J. Steinhardt. Progress measures for grokking via mechanistic interpretability. ICLR, 2023. [28] C. Olsson, N. Elhage, N. Nanda, et al. In-context learning and induction heads. Transformer Circuits Thread, 2022. [29] K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt. Interpretability in the wild: A circuit for indirect object identification in GPT-2 small. ICLR, 2023. [30] J. Frankle and M. Carbin. The lottery ticket hypothesis: Finding sparse, trainable neural networks. ICLR, 2019. [31] K. Jordan. Muon: An optimizer for hidden layers. 2024. [32] Y. Xu. Spectral edge dynamics in neural network training. arXiv:2603.15678, 2026. [33] Y. Xu. Backbone drift and phase transitions in transformer pretraining. arXiv:2602.23696, 2026. [34] Y. Xu. Holographic encoding and spectral edge events in neural network training. arXiv:2602.18649, 2026. [35] Y. Xu. Low-dimensional and transversely curved optimization dynamics in grokking. arXiv:2602.16746, 2026. [36] Y. Xu. The geometry of multi-task grokking: Transverse instability, superposition, and weight decay phase structure. arXiv:2602.18523, 2026. [37] Y. Xu. Phase transitions in Dyck and SCAN grokking: A spectral edge analysis. 2026. [38] Y. Xu. The spectral edge thesis: Detailed mathematical notes. Companion document, 2026. 196 p.