Paper deep dive
Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold
Pratyush Acharya, Habish Dhakal
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 5:23:17 AM
Summary
The paper introduces 'Spectral Gating' to explain grokking in neural networks, framing it as a variance-limited phase transition. It argues that AdamW acts as a variance-gated stochastic system where generalization is delayed because the generalizing solution resides in a sharp basin initially inaccessible due to stability constraints. Grokking occurs when accumulated gradient variance lifts the stability ceiling, allowing the optimizer to enter the sharp manifold. The authors identify three complexity regimes (Capacity Collapse, Variance-Limited, and Stability Override) and demonstrate that anisotropic rectification, rather than isotropic noise, is essential for generalization.
Entities (6)
Relation Signals (4)
Grokking → constrainedby → Spectral Gating
confidence 95% · Grokking is constrained by a stability condition: the generalizing solution resides in a sharp basin
AdamW → implements → Spectral Gating
confidence 95% · AdamW dynamics on modular arithmetic tasks, revealing a 'Spectral Gating' mechanism
AdamW → operatesas → Variance-Gated Stochastic System
confidence 95% · We find that AdamW operates as a variance-gated stochastic system.
Anisotropic Rectification → enables → Grokking
confidence 90% · Generalization requires the anisotropic rectification unique to adaptive optimizers
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Standard optimization theories struggle to explain grokking, where generalization occurs long after training convergence. While geometric studies attribute this to slow drift, they often overlook the interaction between the optimizer's noise structure and landscape curvature. This work analyzes AdamW dynamics on modular arithmetic tasks, revealing a ``Spectral Gating'' mechanism that regulates the transition from memorization to generalization. We find that AdamW operates as a variance-gated stochastic system. Grokking is constrained by a stability condition: the generalizing solution resides in a sharp basin ($\lambda_{max}^H$) initially inaccessible under low-variance regimes. The ``delayed'' phase represents the accumulation of gradient variance required to lift the effective stability ceiling, permitting entry into this sharp manifold. Our ablation studies identify three complexity regimes: (1) \textbf{Capacity Collapse} ($P < 23$), where rank-deficiency prevents structural learning; (2) \textbf{The Variance-Limited Regime} ($P \approx 41$), where generalization waits for the spectral gate to open; and (3) \textbf{Stability Override} ($P > 67$), where memorization becomes dimensionally unstable. Furthermore, we challenge the "Flat Minima" hypothesis for algorithmic tasks, showing that isotropic noise injection fails to induce grokking. Generalization requires the \textit{anisotropic rectification} unique to adaptive optimizers, which directs noise into the tangent space of the solution manifold.
Tags
Links
- Source: https://arxiv.org/abs/2603.15492v1
- Canonical: https://arxiv.org/abs/2603.15492v1
Trouble viewing inline? Open PDF directly →
Full Text
39,669 characters extracted from source content.
Expand or collapse full text
Grokking as a Variance-Limited Phase Transition: Spectral Gating and the Epsilon-Stability Threshold Pratyush Acharya acharya.pratyush12@gmail.com Habish Dhakal dhakalhabish@gmail.com (March 16, 2026) Abstract Standard optimization theories struggle to explain grokking, where generalization occurs long after training convergence. While geometric studies attribute this to slow drift, they often overlook the interaction between the optimizer’s noise structure and landscape curvature. This work analyzes AdamW dynamics on modular arithmetic tasks, revealing a “Spectral Gating” mechanism that regulates the transition from memorization to generalization. We find that AdamW operates as a variance-gated stochastic system. Grokking is constrained by a stability condition: the generalizing solution resides in a sharp basin (λmaxH _max^H) initially inaccessible under low-variance regimes. The “delayed” phase represents the accumulation of gradient variance required to lift the effective stability ceiling, permitting entry into this sharp manifold. Our ablation studies identify three complexity regimes: (1) Capacity Collapse (P<23P<23), where rank-deficiency prevents structural learning; (2) The Variance-Limited Regime (P≈41P≈ 41), where generalization waits for the spectral gate to open; and (3) Stability Override (P>67P>67), where memorization becomes dimensionally unstable. Furthermore, we challenge the "Flat Minima" hypothesis for algorithmic tasks, showing that isotropic noise injection fails to induce grokking. Generalization requires the anisotropic rectification unique to adaptive optimizers, which directs noise into the tangent space of the solution manifold. 1 Introduction Sometimes, a neural network learns a task only after it has already perfectly fit the training data. This delayed phase transition, known as grokking [1], challenges the traditional view that optimization ends when gradients hit zero. Understanding this disconnect is vital, as it may explain why large foundation models continue to improve reasoning capabilities even after their training loss plateaus. This implies the existence of a high-dimensional Minimizing Level Set [7, 19] =θ∣ℒ(θ)≈0Z=\θ (θ)≈ 0\, where dynamics are driven not by the loss gradient, but by the optimizer’s internal noise structure. Earlier studies explain these dynamics in three ways: 1. Geometric Drift: Theories suggesting that weight decay slowly pushes the model toward max-margin solutions [7, 19]. While mathematically rigorous, these frameworks assume a generic gradient flow. They often fail to explain why isotropic noise (SGLD) or standard SGD does not grok, even when weight decay is present [5]. 2. Circuit Competition: Mechanistic interpretations that view grokking as a battle between “memorizing” and “generalizing” circuits [3, 20]. These studies map the topology of the solution (e.g., the “Clock Circuit”) but lack a kinetic mechanism to explain the timescale of the transition. 3. Edge of Stability (EoS): Research arguing that instability drives feature learning [6, 18]. However, the precise interaction between numerical stability parameters (such as ϵε) and the grokking phase transition remains under-explored. In this paper, we unify these perspectives by analyzing the optimizer as a Variance-Gated Stochastic System. We find that grokking is a fragile state that depends on matching task complexity to optimizer stability. We posit that the delay in generalization is not a random walk, but a Variance-Limited Equilibrium. The optimizer is initially too stable to enter the sharp manifolds where generalizing circuits reside. Our contributions are as follows: • The Spectral Gating Mechanism: We find that the generalizing basin for modular arithmetic tasks is significantly sharper than the optimizer’s initial stability threshold. Grokking appears to occur only when accumulated gradient variance vt v_t lifts the stability ceiling (2/ηeff2/ _eff), finally permitting entry into the solution’s high-curvature geometry. • The Complexity Threshold: We identify a “Signal Starvation” regime (P<23P<23) where the loss landscape is structurally barren. Below this threshold, neither variance accumulation nor parameter tuning induces generalization, refuting purely thermal explanations of grokking. • Anisotropic Rectification: By comparing AdamW to Isotropic SGLD, we show that thermal energy alone is insufficient. Generalization requires the covariance structure of adaptive optimizers to rectify noise into the tangent space of the solution manifold. 2 Related Work Grokking and Phase Transitions. The phenomenon of grokking was first characterized by Power et al. [1] as generalization occurring long after training error converges to zero. Subsequent work has attempted to demarcate the conditions for this transition. Liu et al. [2] identified a "Goldilocks zone" of initialization and data size, while Kumar et al. [18] framed it as a transition from "lazy" kernel regimes to "rich" feature learning. While these works describe the phenomenology of the transition, they largely treat the optimizer as a black box. Our work differs by providing a kinetic explanation: we identify the specific spectral state the optimizer must reach to trigger the transition. Mechanistic Interpretability. A parallel line of inquiry focuses on what is learned during grokking. Nanda et al. [3] fully reverse-engineered the modular addition network, identifying a specific "Clock Circuit" relying on trigonometric interference. Varma et al. [20] and Merrill et al. [17] argue that grokking is a competition between dense memorization circuits and sparse generalizing circuits. We adopt this topological view—that the solution is a specific, sparse circuit—but address the open question of accessibility. While mechanistic work maps the destination, our spectral gating theory explains the optimization dynamics required to enter the destination basin. Implicit Bias and the Edge of Stability. Geometric theories posit that gradient descent introduces an implicit bias toward max-margin solutions [8]. Recent rigorous frameworks by Musat [7] and Pesme et al. [19] model grokking as Riemannian Norm Minimization on the zero-loss manifold. However, these "drift" theories often assume stable gradient flow. Cohen et al. [6] and Damian et al. [14] demonstrate that neural network training typically occurs at the "Edge of Stability," where the sharpness oscillates around 2/η2/η. Thilak et al. [5] empirically linked this instability ("Slingshots") to grokking. Our work unifies the geometric and stability perspectives. We formalize the Slingshot not as an anomaly, but as a variance-injection mechanism required to satisfy the stability constraints of the sharp basins identified by mechanistic interpretability. Furthermore, we challenge the prevailing "Flat Minima" hypothesis [9, 16] in this specific domain, providing evidence that for algorithmic tasks, the generalizing solution is spectrally sharper than the memorization solution. 3 Theoretical Framework: Spectral Gating and Rectified Stochastic Dynamics To understand how AdamW moves through the loss landscape, we model its training dynamics as a continuous stochastic process. While classical optimization focuses on the convergence of the loss mean, grokking takes place in a "post-convergence" regime where ℒ(θ)→0L(θ)→ 0 yet the weights ‖θ‖\|θ\| continue to evolve [1]. In this phase, the trajectory is governed not by the loss gradient, which is near zero, but by the noise covariance structure of the optimizer. 3.1 AdamW as Variance-Gated Stochastic Dynamics We treat the discrete updates of AdamW as a continuous-time Stochastic Differential Equation (SDE) to capture the interaction between noise and geometry. Let gt(θ)=∇ℒ(θ)+ξtg_t(θ)= (θ)+ _t be the stochastic gradient, where ξt∼(0,Σ(θt)) _t (0, ( _t)) represents anisotropic, often heavy-tailed noise [15, 16]. AdamW preconditions these updates using vtv_t, the exponential moving average of squared gradients. Under the Adiabatic Approximation [11], we assume the preconditioner Dt=diag(vt+ϵ)D_t=diag( v_t+ε) evolves on a slower timescale than the parameters. Matching the first and second moments of the discrete step Δθ≈−ηDt−1gt θ≈-η D_t^-1g_t yields the following SDE [10]: dθt=−Dt−1(∇ℒ(θt)+λθt)dt⏟Preconditioned Drift+η1/2Dt−1Σ(θt)1/2dWt⏟Rectified Diffusiond _t= -D_t^-1 ( ( _t)+λ _t )dt_Preconditioned Drift+ η^1/2D_t^-1 ( _t)^1/2dW_t_Rectified Diffusion (1) where WtW_t is a standard Brownian motion. The η1/2η^1/2 scaling ensures the diffusion term matches the (η)O(η) variance of the discrete algorithm. Crucially, AdamW behaves differently from standard Riemannian Langevin Dynamics (RLD). While RLD scales diffusion by the inverse square root of the metric (Dt−1/2D_t^-1/2), AdamW scales it by the inverse (Dt−1D_t^-1). This structural difference creates a distinct diffusivity profile for the i-th parameter: eff(i)∝η⋅Σii(Σii+ϵ)2D_eff^(i) η· _i( _i+ε)^2 (2) As the gradient noise variance Σii→∞ _i→∞, RLD diffusion becomes unbounded. In contrast, AdamW diffusion saturates at a fixed limit (eff→ηD_eff→η). This Bounded Diffusivity ensures that high-variance gradients do not cause the system to diverge from the manifold, effectively creating a "variance ceiling" that maintains stability even in high-noise regimes. Remark on Discrete Dynamics: While the SDE provides an equilibrium description, the transition into the grokking phase is often triggered by "Slingshots" [5]—discrete instabilities where the adiabatic assumption momentarily breaks down. In these transient moments, the rapid accumulation of vtv_t (as described in Section 5.4) serves to actively suppress the effective step size, dynamically restoring stability. 3.2 The Stability-Diffusion Trade-off Equation (2) shows how the stability parameter ϵε regulates the effective diffusion coefficient. By analyzing the limits of this equation, we can derive the three distinct physical regimes observed in our experiments. 3.2.1 Regime I: Variance Suppression (Over-Damped) Condition: ϵ≫σiε _i. When the stability term dominates the preconditioner, the effective diffusivity vanishes: limϵ→∞eff(i)≈ησi2ϵ2→0 _ε→∞D_eff^(i)≈ η _i^2ε^2→ 0 (3) In this limit, AdamW degenerates into Damped SGD. The diffusive pressure becomes too weak to overcome the restorative drift of weight decay [4]. As seen in the phase diagram (Figure 1), this results in model stagnation; the system remains trapped in the kernel regime, unable to explore the manifold. 3.2.2 Regime I: Radial Expansion (Under-Damped) Condition: ϵ≪σiε _i. When ϵε is negligible, the preconditioner perfectly cancels the noise magnitude: limϵ→0eff(i)≈ησi2(σi2)2=η _ε→ 0D_eff^(i)≈ η _i^2( _i^2)^2=η (4) While this maximizes exploration, it removes the numerical floor required for stability. In our ablation studies, setting ϵ→10−15ε→ 10^-15 caused the weight norm ‖θ‖2\|θ\|^2 to diverge. Without the ϵε constraint, the system enters a phase of Radial Expansion, diffusing outwardly faster than weight decay can pull it back [19]. 3.2.3 Regime I: Anisotropic Rectification (Grokking) Condition: ϵ≈σnoiseε≈ _noise. Grokking occurs when ϵε is balanced. In this band, the optimizer performs Anisotropic Rectification: it amplifies noise in directions where the signal is coherent (low σ) while damping chaotic directions (high σ). This balance allows the system to drift along the tangent space of Z without diverging. 3.3 Spectral Gating and the Edge of Stability For intermediate complexity tasks (P≈23−60P≈ 23-60), we observe that structural learning is delayed by a Spectral Lock. We explain this by connecting our framework to the Edge of Stability (EoS) theory [6]. In discrete optimization, a step size η is stable only if the local curvature satisfies λ<2/ηλ<2/η. For AdamW, the effective step size is adaptive and parameter-specific: ηeff≈η/(vt+ϵ) _eff≈η/( v_t+ε). Consequently, the stability condition becomes dynamic. The optimizer can only converge into a basin with Hessian curvature λmaxH _max^H if it satisfies the Spectral Gating Condition: λmaxH<2η(vt+ϵ) _max^H< 2η( v_t+ε) (5) This inequality uncovers the causal mechanism behind the delay: 1. The Memorization Trap (Low Variance): Initially, gradient variance vtv_t is low. This suppresses the stability ceiling (2/ηeff2/ _eff), forcing the optimizer to seek flat, high-entropy minima. The "lazy" memorization solution fits this profile: it is broad and robust to small perturbations, making it immediately accessible. 2. The Sharpness Inversion: Contrary to the "Flat Minima" hypothesis [9], we argue that for algorithmic tasks, the generalizing solution (the "Clock Circuit" [3]) is geometrically sharper than the memorization basin. It requires precise parameter alignment, resulting in a high λmaxH _max^H that initially violates the condition in Eq. (5). 3. Stability Release (Variance Injection): As training progresses, gradients do not vanish but fluctuate, accumulating variance vtv_t (the "Slingshot" mechanism). This increase in the denominator of the update rule effectively anneals the step size. This lifts the stability ceiling defined by Eq. (5), finally permitting the optimizer to stably enter and settle in the sharp generalizing manifold. Thus, the delay is not a random walk, but a variance-accumulation phase. The model must generate enough gradient noise to "unlock" the spectral gate, transitioning from the variance-intolerant memorization basin to the variance-stabilized generalizing circuit. 4 Experimental Setup Task & Model. Our experiments utilize an MLP with 2 hidden layers (width 128, ReLU activations) trained on the modular addition task a+b(modP)a+b P with embeddings learned from scratch. Training employs full-batch AdamW with learning rate η=10−4η=10^-4, weight decay λ=0.1λ=0.1, and β=(0.9,0.999)β=(0.9,0.999) unless specified otherwise. Protocol. We perform three targeted experimental sweeps to dissect the underlying mechanism: 1. The Stability Frontier (Hardness vs. Gating): To map the thermodynamic boundaries of grokking, we conduct a dense 2D grid search over Task Difficulty (P∈[11,97]P∈[11,97], sampled linearly) and Optimizer Stability (ϵ∈[10−9,10−1]ε∈[10^-9,10^-1], sampled logarithmically). Each configuration is trained for 10610^6 steps to determine the "Steps to Grok" (defined as Test Accuracy >99%>99\%). 2. Thermodynamic Sufficiency (SGLD): To determine if isotropic energy can substitute for geometric steering, we compare AdamW against Stochastic Gradient Langevin Dynamics (SGLD). By injecting Gaussian noise (0,σ2)N(0,σ^2) into standard SGD updates (σ∈[10−4,10−1]σ∈[10^-4,10^-1]), we match the effective thermal pressure of AdamW for both easy (P=23P=23) and hard (P=67P=67) tasks. 3. Mechanism & Topology: Finally, we monitor high-resolution topological metrics, including the Force Ratio (R=Cpush/CpullR=C_push/C_pull), the Hessian Trace (via Hutchinson’s estimator), and Fourier Sparsity (the L1L_1 norm of the weights’ DFT). These metrics allow us to track the formation of the "Clock Circuit" structure in real-time. These experiments define the three dynamical regimes analyzed in the following section. 5 Results and Analysis 5.1 Phase Boundaries of Generalization: The Impact of Task Complexity To map the dynamical regimes of grokking, we performed a dense grid search over Task Difficulty (P∈[11,97]P∈[11,97]) and Optimizer Stability (ϵ∈[10−9,10−1]ε∈[10^-9,10^-1]). The resulting phase diagram (Figure 1) reveals that generalization speed is non-monotonic with respect to complexity. Our results reveal three distinct regimes that explain how grokking begins and ends. Figure 1: Stability-Complexity Phase Diagram. We define three regimes based on the time-to-generalization: (1) Capacity Collapse (P<23P<23), where rank-deficiency prevents the representation of the solution; (2) Variance-Limited Regime (29≤P≤5929≤ P≤ 59), where generalization is delayed by spectral gating; and (3) Stability Override (P≥67P≥ 67), where the high dimensionality of the memorization manifold forces immediate structural learning. The white dashed line marks ϵ=σnoiseε= _noise, representing the intrinsic gradient noise level. 5.1.1 Regime I: Capacity Collapse (P<23P<23) The region P<23P<23 is characterized by convergence failure (Figure 1, Red). While often counter-intuitive that easier tasks are harder to learn, we attribute this to **Capacity Collapse** (or Rank Deficiency). As detailed in Section 6.4, the modular addition task requires the model to represent P distinct Fourier modes [2]. When the embedding dimension dmodeld_model is small relative to P (or when P is small enough that the orthogonality of high-frequency modes is compromised by initialization variance), the gradient signal becomes rank-deficient. The model enters a "tunneling" phase (Figure 2) where it drifts stochastically without locking into a coherent minimum. This is a representational failure, not a thermodynamic one. Figure 2: Capacity Collapse (P=17P=17). Despite aggressive stability tuning (ϵ→10−15ε→ 10^-15), the optimizer fails to locate the generalizing minimum. The dynamics exhibit stochastic tunneling without convergence, indicating that the embedding space lacks the geometric capacity to separate the task’s Fourier modes. 5.1.2 Regime I: The Variance-Limited Regime (29≤P≤5929≤ P≤ 59) Intermediate tasks exhibit the canonical "Grokking Gap," peaking at P=41P=41 (Figure 3, Red line). We attribute this to **Competitive Variance Accumulation**. The memorization solution acts as a "Lazy" attractor: it is geometrically broad (high entropy) and easily accessible from random initialization. However, it is not the global minimum. The optimizer rapidly converges to this basin but is eventually destabilized by the accumulation of gradient variance vt v_t. The delay corresponds to the time required for vtv_t to grow sufficiently to "heat" the optimizer out of the flat memorization trap and into the sharper generalizing basin. Figure 3: Non-Monotonic Generalization Dynamics. Generalization time peaks at intermediate complexity (P=41P=41, Red). Hard tasks (P=97P=97, Teal) generalize immediately (Stability Override), while easy tasks (P=23P=23, Purple) stagnate. The results reveal a "Complexity Valley" where the competition between the entropic pull of memorization and the spectral gate of generalization is maximized. 5.1.3 Regime I: Stability Override (P≥67P≥ 67) For high-complexity tasks (P≥67P≥ 67), the grokking delay vanishes (Figure 3, Teal line). We observe **Stability Override**. The parameter cost of a memorization look-up table scales quadratically as (P2)O(P^2), whereas the generalizing circuit (a constant frequency rotation) remains (1)O(1) [20]. As P increases, the volume of the parameter space occupied by valid memorization solutions shrinks exponentially relative to the structural solution. At P=100P=100, the memorization basin becomes dimensionally unstable—it is simply too "small" to catch the optimizer. Consequently, AdamW bypasses the "Lazy" phase entirely, converging directly to the structural solution. 5.2 Radial Stationarity: The Push-Pull Dynamics Post-convergence dynamics are governed by a radial equilibrium between Weight Decay (ℓ2 _2 regularization) and the Rectified Diffusion of the optimizer. We model the weight norm ‖θ‖2\|θ\|^2 as a radial Ornstein-Uhlenbeck process. (a) Diffusive Expansion (FpushF_push) (b) Restorative Drift (FpullF_pull) (c) Radial Orbit ‖θ‖2\|θ\|^2 Figure 4: Radial Stationarity. (a) AdamW generates a constant diffusive "push" (peach line) that counters the restorative force. (b) The restorative drift stabilizes, confirming a fixed radial orbit. (c) Stronger weight decay (peach) forces a lower-norm equilibrium, compressing the search space onto the manifold [7]. Distinct stationary states emerge under analysis (Figure 4). High weight decay (λ=0.5λ=0.5) forces the system into a high-energy, low-radius orbit (Figure 4c). This compression accelerates grokking (Figure 5) by restricting the search space volume. Effectively, the constraint forces the optimizer to rectify noise into the tangent directions of the manifold, consistent with Riemannian Norm Minimization theories [19]. (a) Test Accuracy vs. Steps (b) Optimizer Variance v^t v_t Figure 5: Acceleration via Constraint. (a) Higher weight decay (Peach) significantly accelerates the phase transition. (b) Strong regularization suppresses the peak gradient variance, enforcing a disciplined search in the tangent space. 5.3 Stability-Regulated Diffusion The stability parameter ϵε acts as a spectral gatekeeper. Figure 6 demonstrates the sensitivity of the phase transition to ϵε. Figure 6: The Stability Threshold. Class C (Hard): Performance degrades as ϵ→10−1ε→ 10^-1, confirming the need for anisotropic noise (Limit I). Class B (Intermediate): Exhibits a convex "optimal stability" region around ϵ=10−4ε=10^-4. Class A (Easy): Fails regardless of stability settings. For solvable tasks, we identify an **Optimal Stability Interval** (ϵ≈10−4ε≈ 10^-4): • Over-Damping (ϵ→10−1ε→ 10^-1): The effective diffusivity eff→0D_eff→ 0. The optimizer behaves like SGD, failing to explore the manifold. • Under-Damping (ϵ→10−9ε→ 10^-9): The system suffers from radial instability. While diffusion is maximized, the lack of a stability floor prevents weight condensation onto the sparse circuit [21]. 5.4 Mechanism: Spectral Gating and Sharpness Violation We provide direct evidence that a **Spectral Stability Condition ** governs the delayed generalization. According to Edge of Stability (EoS) theory [6], the optimizer cannot stably enter a basin where curvature λmaxH _max^H exceeds the stability threshold 2/ηeff2/ _eff. Figure 7: Violation of the Initial Stability Condition. We compare Hessian Sharpness (Blue) to the Initial Stability Limit determined by ϵε (Red Dotted Line). For the grokking task (P=41P=41, Center), the generalizing solution resides in a basin where λmax _max is **721x higher** than the initial stability threshold. The delayed phase (0–200k steps) corresponds to the time required for gradient variance vt v_t to grow sufficiently to lift the stability ceiling (Black Dashed Line) above the basin’s curvature. Figure 7 illustrates the core mechanism. For the grokking task (P=41P=41), the generalizing minimum is extremely sharp. The "Sharpness Violation Ratio" reaches **721.31x**. This creates a **Spectral Gate**: the model is initially locked out of the generalizing basin. It remains trapped in the flat memorization basin until the accumulated gradient noise variance vt v_t increases the denominator of the effective step size. This variance accumulation raises the effective stability ceiling 2/ηeff2/ _eff, eventually permitting entry into the sharp "Clock Circuit" basin. 5.5 The Necessity of Anisotropy: Failure of Isotropic Diffusion To disentangle the role of "thermal energy" (variance) from "geometric rectification" (covariance), we trained with Stochastic Gradient Langevin Dynamics (SGLD). We injected isotropic Gaussian noise matched to the thermal scale of AdamW. (a) test/acc: Flatline at Random Chance (b) metrics/weight_norm_sq: Scalar Equilibrium Figure 8: Failure of Isotropic SGLD. (a) Despite 1×1061× 10^6 steps and extensive noise tuning (σ∈[10−3,10−1]σ∈[10^-3,10^-1]), SGLD fails to generalize. (b) Isotropic noise successfully maintains the weight norm, proving that radial energy alone is insufficient. Isotropic diffusion fails to induce grokking (Figure 8). While the noise prevents weight collapse, the model remains in a high-entropy state (Figure 9). This confirms that grokking requires **Anisotropic Rectification**: the optimizer must strictly suppress noise in high-curvature directions while amplifying it in tangent directions to navigate the Minimizing Level Set [10]. (a) metrics/fourier_sparsity_a: High Entropy (b) metrics/fourier_sparsity_b: High Entropy (c) metrics/hessian_trace: No Spectral Features Figure 9: Structural Stagnation in SGLD. (a) and (b) Fourier sparsity remains high, indicating no circuit formation. (c) The Hessian trace lacks the characteristic "Ignition" spike of AdamW. 6 Discussion Grokking has historically been framed as a mystery of timescales—specifically, why generalization lags so far behind training convergence. Our findings suggest this is actually a problem of spectral accessibility. By analyzing AdamW as a Variance-Gated Stochastic System, we show that the delay is not merely a slow drift, but a structural requirement for stability. The optimizer must undergo specific internal state changes to access high-curvature basins. This perspective helps resolve three tensions in the current literature. 6.1 The Kinetic Mechanism: Geometric Rectification Recent geometric frameworks [19, 7] prove that grokking corresponds to Riemannian Norm Minimization on the Minimizing Level Set Z. However, they attribute the delay primarily to the weak magnitude of weight decay. Our ablation of Isotropic SGLD (Section 5.5) refines this view. We find that while isotropic noise can prevent weight collapse and lift test accuracy above random chance, it fails to induce perfect generalization. In simple terms, adding random noise helps the model escape shallow local minima, but it does not provide the steering necessary to find the specific generalizing circuit. Because isotropic noise vectors are statistically orthogonal to the low-dimensional solution manifold Z, standard diffusion results in orthogonal oscillations. AdamW resolves this via **Anisotropic Rectification**: by scaling updates by vt−1/2v_t^-1/2, it directs noise into the tangent directions of the manifold [10, 16]. Thus, isotropic noise helps the model move, but geometric rectification ensures it moves in the right direction. 6.2 Unifying the Stability Paradox There is a conflict in the literature regarding instability. Thilak et al. [5] argue that "Slingshots"—spikes in loss and gradient variance—drive grokking. Conversely, Prieto et al. [21] suggest that numerical instability (like Softmax Collapse) prevents learning. Our **Spectral Gating Theory** (Eq. 5) suggests these views are compatible. We identify the Slingshot as a necessary variance accumulation event. • Mechanism of Entry: A Slingshot causes a spike in the second moment estimator vtv_t. This increase in the denominator reduces the effective step size ηeff≈η/vt _eff≈η/ v_t. This reduction raises the stability ceiling 2/ηeff2/ _eff, temporarily allowing the optimizer to survive in the sharper generalizing basin. • The Epsilon Constraint: As noted by Thilak et al., increasing ϵε halts Slingshots. We derive this analytically: if ϵε is too high, the variance vtv_t never dominates the denominator, and the stability ceiling never lifts. If ϵε is too low, the instability becomes unbounded. Grokking appears to happen near the edge of stability. The optimizer needs enough instability to generate variance for exploration, but too much instability leads to numerical divergence. 6.3 Revisiting Flatness: The Geometry of Algorithmic Precision Kumar et al. [18] describe grokking as a transition from "Lazy" to "Rich" dynamics. This creates a theoretical tension with the prevailing "Flat Minima" hypothesis [9], which posits that generalization correlates with broad, low-curvature basins. We propose a topological resolution specific to algorithmic tasks. While flat minima are robust to input noise in perceptual tasks (e.g., image classification), algorithmic solutions—such as the trigonometric "Clock Circuit" identified by Nanda et al. [3]—rely on **precise interference patterns**. Constructive interference is geometrically fragile: small perturbations in the frequency or phase parameters destroy the circuit’s logic. Consequently, the generalizing solution resides in a **narrow, high-curvature manifold**. In contrast, the "Lazy" memorization solution relies on a look-up table structure. This basin is **entropically vast**: there are combinatorially many ways to assign weights to memorized disconnected points, creating a broad, flat attractor that is easily accessible from random initialization. The grokking delay is therefore the kinetic cost of this geometry. The optimizer is initially trapped in the high-entropy memorization basin. It cannot access the low-entropy, high-precision generalizing manifold until it accumulates sufficient gradient variance to lift the spectral gate (Eq. 5), satisfying the strict stability requirements of the algorithmic solution. 6.4 Dimensional Bottlenecks: Capacity Collapse Finally, we address the limits of this phenomenon. Our embedding ablation (Figure 10) reveals that task difficulty is governed not just by the modulus P, but by the ratio of capacity to complexity. Figure 10: Capacity Collapse. We ablate embedding dimension V against modulus P. The color scale indicates the time required to grok (Log10 Steps). The sharp diagonal boundary reveals that grokking requires a minimum dimensional capacity proportional to task complexity (V∝PV P). Black crosses indicate configurations that failed to generalize within 10610^6 steps. As shown in Figure 10, a sharp phase boundary emerges. Failure occurs when the embedding dimension is insufficient relative to the modulus (V≲PV P). We term this **Capacity Collapse**. When the embedding space lacks the geometric bandwidth to represent the P discrete Fourier modes required for the solution [2], the gradient signal becomes **rank-deficient**. This prevents the optimizer from constructing the generalizing circuit regardless of the variance supplied. This finding confirms that spectral gating is a secondary mechanism that can only operate when the model capacity is geometrically sufficient to support the solution. 7 Conclusion We have characterized grokking as a conditional phase transition governed by the interplay between landscape geometry and optimizer stability. We establish three governing principles: 1. The Complexity Threshold: Generalization is not monotonic. We identify a "Signal Starvation" regime where the landscape is structurally barren (V<PV<P), and a "Complexity Override" regime (P>67P>67) where memorization becomes dimensionally unstable, forcing early generalization. 2. Spectral Gating: We demonstrate that the delayed phase is a stable equilibrium where the model is spectrally barred from the generalizing basin. The delay represents the time required for gradient variance to accumulate and lift the stability ceiling defined by ϵε. 3. Anisotropic Rectification: We show that thermal energy alone is insufficient. Generalization requires the directed channeling of variance unique to adaptive optimizers, which rectify noise into the tangent space of the solution manifold. In short, generalization depends on how the optimizer balances noise and stability. Future work could investigate whether similar spectral gating mechanisms govern the emergence of reasoning capabilities in larger language models, or if this behavior is specific to highly structured algorithmic tasks. References [1] Alethea Power, Yuri Burda, Harri Edwards, Igor Babuschkin, and Vedant Misra. Grokking: Generalization beyond overfitting on small algorithmic datasets. arXiv preprint arXiv:2201.02177, 2022. [2] Ziming Liu, Eric J. Michaud, and Max Tegmark. Omnigrok: Grokking beyond algorithmic data. ICLR, 2023. [3] Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. International Conference on Learning Representations (ICLR), 2023. [4] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. ICLR, 2019. [5] Vimal Thilak, Etai Littwin, Shuangfei Zhai, Omid Saremi, Roni Paiss, and Joshua Susskind. The slingshot mechanism: An empirical study of adaptive optimizers and the grokking phenomenon. arXiv preprint arXiv:2206.04817, 2022. [6] Jeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar. Gradient descent on neural networks typically occurs at the edge of stability. ICLR, 2021. [7] Tiberiu Musat. The Geometry of Grokking: Norm Minimization on the Zero-Loss Manifold. arXiv preprint arXiv:2511.01938, 2025. [8] Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro. The implicit bias of gradient descent on separable data. The Journal of Machine Learning Research, 19(1):2822–2878, 2018. [9] Sepp Hochreiter, Jurgen Schmidhuber Flat Minima Neural Computation, 1997 [10] Lukas Balles and Philipp Hennig. Dissecting Adam: The sign, magnitude and variance of stochastic gradients. Proceedings of the 35th International Conference on Machine Learning (ICML), 2018. [11] Frederik Kunstner, Philipp Hennig, and Lukas Balles. Limitations of the empirical fisher approximation for natural gradient descent. Advances in Neural Information Processing Systems (NeurIPS), 32, 2019. [12] Samy Jelassi and Yuanzhi Li. Towards understanding how momentum improves generalization in deep learning. Proceedings of the 39th International Conference on Machine Learning (ICML), 2022. [13] Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi. The scale of initialization determines the edge of stability. Advances in Neural Information Processing Systems (NeurIPS), 35, 2022. [14] Alex Damian, Eshaan Nichani, and Jason D. Lee. Self-stabilization: The implicit bias of gradient descent at the edge of stability. International Conference on Learning Representations (ICLR), 2023. [15] Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban. A tail-index analysis of stochastic gradient noise in deep neural networks. Proceedings of the 36th International Conference on Machine Learning (ICML), 2019. [16] Zeke Xie, Issei Sato, and Masashi Sugiyama. A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima. International Conference on Learning Representations (ICLR), 2021. [17] William Merrill, Nikolaos Tsilivis, and Aman Shukla. A tale of two circuits: Grokking as competition of sparse and dense subnetworks. arXiv preprint arXiv:2303.11873, 2023. [18] Tanishq Kumar, Blake Bordelon, Samuel J. Gershman, and Cengiz Pehlevan. Grokking as the transition from lazy to rich training dynamics. International Conference on Learning Representations (ICLR), 2024. [19] Scott Pesme, Etienne Boursier, and Radu-Alexandru Dragomir. A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm Minimisation. Advances in Neural Information Processing Systems (NeurIPS), 2025. [20] Vikrant Varma, Rohin Shah, Zachary Kenton, János Kramár, and Ramana Kumar. Explaining grokking through circuit efficiency. arXiv preprint arXiv:2309.02390, 2023. [21] Lucas Prieto, Melih Barsbey, Pedro A.M. Mediano, and Tolga Birdal. Grokking at the Edge of Numerical Stability. International Conference on Learning Representations (ICLR), 2025. Appendix A Appendix This supplementary material provides detailed geometric and spectral evidence supporting the **Spectral Gating** framework presented in the main text. We provide topological visualizations of the embedding space (Section A.1), quantify the anisotropic covariance structure of AdamW updates (Section A.2), and examine the spectral limits of thermal noise injection across the full range of task complexities (Section A.4). A.1 Topological Analysis: The Clock Face Section 5.1 hypothesized that the "Signal Starvation" regime (P<23P<23) stems from a failure to form the requisite task geometry. Figure 11 visualizes the embedding weights WEW_E via Principal Component Analysis (PCA). Figure 11: Embedding Geometry across Complexity Regimes. Learned embeddings projected onto their first two principal components. Left (P=17P=17): The Signal Starvation regime exhibits a lack of structural organization; embeddings are scattered, indicating failure to discover the modular circle. Center (P=41P=41): The Grokking regime reveals a recognizable but noisy circular geometry, consistent with a variance-limited equilibrium where the optimizer maintains high entropy to navigate the basin. Right (P=97P=97): The Stability Override regime forms a clear circular structure immediately, indicating that high-complexity tasks force rapid geometric condensation. A.2 Evidence of Anisotropy Section 5.5 argued that grokking relies on **Anisotropic Rectification**—the selective amplification of noise in tangent directions. Figure 12 characterizes this by plotting the distribution of the inverse preconditioner scale factors αi=1/(vt,i+ϵ) _i=1/( v_t,i+ε) at the moment of generalization. (a) Grokking Regime (P=41P=41) (b) Override Regime (P=97P=97) Figure 12: Distribution of Update Scalings. Histograms display the effective learning rates applied by AdamW per parameter. Unlike Isotropic SGLD (which implies a Dirac distribution), AdamW generates a **heavy-tailed distribution** spanning multiple orders of magnitude. This indicates that the optimizer actively shapes the noise geometry, suppressing variance in high-curvature directions (left tail) while amplifying it in flat directions (right tail). A.3 Spectral Dynamics across Complexity To ensure the mechanisms identified in Section 5.4 apply beyond specific cases, Figure 13 presents the spectral dynamics across the full range of moduli. Figure 13: Spectral Dynamics for Varying Moduli. Green lines denote Test Accuracy; Blue lines denote Hessian Sharpness. For successful generalization cases (P≥29P≥ 29), we observe the characteristic "Ignition" event where the variance-driven stability ceiling (Black Dashed) rises to meet the local curvature. A.4 Limits of Isotropic Noise Figure 14 provides a granular view of the interaction between isotropic noise and the "Signal Starvation" regime. Figure 14: Isotropic Noise Injection. Injecting isotropic noise (σ=0.01σ=0.01, cyan lines) allows the model to escape the lazy basin and reach ≈60−80%≈ 60-80\% accuracy, outperforming the noiseless baseline. However, it fails to induce the perfect generalization seen in grokking. Thus, isotropic noise aids optimization but does not replace the geometric steering required for structural convergence.