Paper deep dive
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
Sheng Ren, Yadong Wang, Naiqiang Tan, Jiangang Kong, Jun Fang, Rui Liu, Jun Wang, Kai Chen, Lipeng Liang, Xiang Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/14/2026, 6:17:56 AM
Summary
This paper investigates the interaction between normalization placement (pre-norm vs. post-norm) and training curriculum (joint training vs. curriculum depth growth) in Large Language Models. Using a block-stack distillation setup with Qwen3-8B, the authors find that while pre-norm and post-norm perform similarly under joint training, post-norm significantly outperforms pre-norm under curriculum depth growth. This advantage is attributed to post-norm's ability to provide stable boundary representations for newly appended blocks, whereas pre-norm suffers from structural-token scale drift.
Entities (8)
Relation Signals (6)
Post-norm â outperforms â Pre-norm
confidence 95% ¡ post-norm improves over pre-norm by 0.0328 under curriculum growth
Qwen3-8b â usedasteacherfor â Block-Stack Student
confidence 95% ¡ controlled distillation study with a Qwen3-8B teacher and a nine-layer student
Curriculum Depth Growth â enhances â Post-norm
confidence 92% ¡ post-norm improves over pre-norm by 0.0328 under curriculum growth
Pre-norm â performssimilarlyto â Post-norm
confidence 90% ¡ pre-norm and post-norm are indistinguishable under joint training, differing by 0.0004 validation CE
Pre-norm â causes â Structural-token scale drift
confidence 88% ¡ pre-norm with structural-token scale drift
Post-norm â provides â Stable residual scales
confidence 88% ¡ Boundary diagnostics associate post-norm with stable residual scales
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by $0.0004$ validation CE, while post-norm improves over pre-norm by $0.0328$ under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting.
Tags
Links
- Source: https://arxiv.org/abs/2608.13156v1
- Canonical: https://arxiv.org/abs/2608.13156v1
Trouble viewing inline? Open PDF directly â
Full Text
56,929 characters extracted from source content.
Expand or collapse full text
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing Sheng Ren Yadong Wang Naiqiang Tan Jiangang Kong Jun Fang Rui Liu Jun Wang Kai Chen Lipeng Liang Xiang Chen Abstract Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We ask whether this preference persists when depth is introduced through a curriculum. In curriculum depth growth, each appended block receives the boundary representation produced by a trained prefix, making normalization placement relevant to forward conditioning. We therefore test whether placement and training curriculum interact. In a controlled distillation study with a Qwen3-8B teacher and a nine-layer student, pre-norm and post-norm are indistinguishable under joint training, differing by 0.00040.0004 validation CE, while post-norm improves over pre-norm by 0.03280.0328 under curriculum growth, an order of magnitude larger. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out compute as the sole explanation. The ranking crosses over during the curriculum: post-norm takes the lead once blocks are appended. Single-block and freeze controls localize the ranking change to block appending rather than shallow-block quality or retraining. Boundary diagnostics associate post-norm with stable residual scales and pre-norm with structural-token scale drift; on a fixed batch, the final pre-grow block is also nearly identity-mapped. Together with the phase-wise crossover, these observations are consistent with boundary-scale conditioning after new blocks are appended. The results motivate treating normalization placement and training curriculum as coupled design choices in this distillation setting. Introduction Pre-norm has become the standard choice in modern Transformer language models (33; 26). This preference is well motivated in conventional end-to-end training. When a full-depth Transformer is trained from random initialization, normalization before each sub-layer improves gradient behavior at initialization and reduces the need for carefully tuned learning-rate warm-up in practice (33; 17). This finding has influenced decoder-only LLMs and distilled Transformer students (26; 7). However, post-norm remains competitive when its optimization difficulties are addressed (28; 5). Normalization placement may therefore depend on the training regime rather than on the layer formulation alone, an issue we take up here. Figure 1: Pre-norm normalizes sub-layer inputs, whereas post-norm normalizes the residual stream after each update. We study this question in curriculum depth growth. Unlike joint training, in which all layers are active from the first update, curriculum depth growth begins with a shallow model and appends new blocks in later phases (9; 6). The final model is therefore constructed through a sequence of inherited initializations. At each phase, the blocks trained in the previous phase determine the input distribution received by the newly appended block. This perspective is related to Net2Net-style network growth (4; 30), but it raises a question that existing growth methods generally leave unanswered: which normalization placement provides a better-conditioned boundary input for a newly added block? We hypothesize that normalization placement interacts with the way depth is introduced. Under joint training, normalization primarily supports optimization through a complete stack. Under curriculum depth growth, it also shapes the forward distribution passed from the trained stack to each newly appended block. Pre-norm normalizes sub-layer inputs but does not explicitly normalize the residual stream emitted across block boundaries, whereas post-norm normalizes after each residual update. We therefore expect the placements to behave similarly under joint training but to separate after the curriculum begins appending blocks. We test this hypothesis in a controlled block-stack distillation setting. The student uses a frozen teacher embedding and a teacher-initialized language-modeling head that is trained jointly with its appendable decoder blocks. Under joint training, the two placements differ by only 0.00040.0004 validation CE; under curriculum growth, post-norm improves over pre-norm by 0.03280.0328 after additional blocks are appended, yielding an interaction of â0.0332-0.0332. A post-joint control matched by student active-layer tokens remains worse than post-grow, which rules out additional student-decoder compute as the sole explanation. Block-boundary diagnostics show stable residual scales for post-norm and increasing scale drift for pre-norm, providing correlational evidence for the proposed mechanism. In summary, our contributions are as follows: ⢠We recast curriculum depth growth as block-wise initialization, shifting the normalization question from full-depth optimization to inherited boundary conditioning. ⢠We show that normalization placement interacts with the training curriculum: pre-norm and post-norm are nearly tied under joint training, whereas post-norm performs better under sequential depth growth in the evaluated block-stack setting. ⢠We analyze the interaction through single-block controls, block-removal analysis, and block-boundary activation diagnostics, obtaining converging evidence for forward conditioning. Related Work Normalization Placement in Transformers. The original Transformer adopted post-norm (27), while later analyses showed that pre-norm improves initialization-time gradients and reduces warm-up sensitivity (33; 17; 2), motivating its use in modern decoder-only LLMs (26; 25). Sandwich norm, DeepNorm, and ReZero stabilize post-norm or related residual designs (5; 28; 3), while recent work explores placements beyond the standard dichotomy (15; 36; 18). HybridNorm combines QKV normalization with post-norm FFNs within each block to capture the strengths of both placements (37), and SiameseNorm couples pre-norm-like and post-norm-like streams through shared residual blocks to reconcile stability and representational capacity (16). A principled forwardâbackward stability analysis of normalization placements further clarifies why different placements induce distinct training dynamics at fixed depth (14). Initialization and residual scaling provide another stability lever (35; 12). These studies largely assume fixed-depth end-to-end training; we instead test whether the preferred placement changes when depth is introduced sequentially. Our objective is therefore different from proposing another fixed-depth stabilization rule: we ask whether the same placement can change rank when the training path changes. Network Growth and Function Preservation. Progressive growth expands a trained smaller network, often through function-preserving transformations (4; 30). Transformer variants include StackBERT, stacking operators for LLM pretraining, and block insertion into pretrained models (9; 6; 31). These methods primarily optimize training efficiency or preserve the function of the smaller network, while normalization placement is usually inherited and held fixed. They do not isolate whether placement itself interacts with the joint-versus-grow protocol. We instead view curriculum growth as inherited initialization: an independently parameterized block is trained on the boundary distribution produced by the existing stack, making normalization placement part of the growth condition rather than a fixed inherited choice. Block-Stack Knowledge Distillation. Transformer distillation typically studies which logits, hidden states, or attention signals a fixed-depth student should match (10; 8; 13; 21; 24; 29). LayerSkip shares the language-modeling head for early exits, while stochastic depth varies active layers without imposing a depth-introduction order (7; 11). Pruning methods also produce compact stacks from the top down (19; 32). These approaches determine what a fixed-depth student should retain or which existing layers should be removed, but they do not study how a trained prefix initializes the input distribution of a newly appended block. Our bottom-up block-stack setting adds independently parameterized blocks sequentially and compares normalization placements at the resulting phase boundaries, separating this question from fixed-depth distillation and top-down compression. Figure 2: Overview of the block-stack distillation framework. The student uses a frozen teacher embedding and a teacher-initialized, trainable language-modeling head, while its decoder is divided into appendable blocks. Joint training activates all blocks from the beginning, whereas curriculum growth adds one block in each phase. Pre-norm and post-norm differ in whether normalization is applied before or after each residual update. Methodology This section introduces the block-stack student, the curriculum depth-growth protocol, and the forward-conditioning hypothesis underlying our comparison. The Experiments section describes the implementation and evaluates the interaction between normalization placement and the training curriculum. The formulation below makes the training path and block-boundary condition explicit. Depth-Growable Block-Stack Student Block-Stack Student. We use a block-stack student whose depth can be expanded during training. Given a teacher Transformer with T decoder layers, the student is divided into K blocks with L decoder layers per block, resulting in a total depth of S=Kâ LâŞTS=K¡ L T. The vocabulary embedding E and language-modeling head lmW_lm are copied from the teacher. The embedding remains frozen, whereas the language-modeling head and the independently parameterized student decoder layers are optimized. This design fixes the input interface and initializes the output interface identically across variants while allowing the decoder stack and language-modeling head to adapt during distillation, consistent with common Transformer distillation settings (10; 13). Each block contains a contiguous group of decoder layers and, once introduced, remains active in all subsequent phases. All variants share the same attention and MLP sub-layersârotary position embeddings (23), SwiGLU activations (22), grouped-query attention (1), and scaled dot-product attentionâso the comparison isolates the placement of RMSNorm (34) relative to each residual update. Pre-Norm and Post-Norm Layers. We compare two placements of RMSNorm. Pre-norm applies RMSNorm immediately before each sub-layer but does not explicitly normalize the residual stream after each residual update: Ⲡ=AttentionâĄ(RMSNormâĄ())+, =Attention(RMSNorm(h))+h, (1) out _out =MLPâĄ(RMSNormâĄ(â˛))+â˛. =MLP(RMSNorm(h ))+h . Following the standard pre-norm decoder interface, we apply a final RMSNorm before the language-modeling head: final=finalâ(K).h_final=N_final(h_K). (2) Post-norm applies RMSNorm after each residual addition: =AttentionâĄ()+, =Attention(h)+h, Ⲡ=RMSNormâĄ(), =RMSNorm(u), (3) =MLPâĄ(â˛)+â˛, =MLP(h )+h , out _out =RMSNormâĄ(). =RMSNorm(v). Because the final post-norm layer already produces a normalized representation, no additional RMSNorm is applied after the final block. The first block, however, receives the raw embedding âĄ()E(x), so we apply a single entry RMSNorm entryN_entry before it. Both finalN_final (pre-norm) and entryN_entry (post-norm) are initialized from the teacherâs final RMSNorm and remain trainable. Both variants therefore contain two RMSNorm modules per decoder layer and exactly one model-boundary RMSNorm, giving them the same number of trainable normalization parameters. Thus, internal block boundaries differ through normalization placement rather than parameter count. Algorithm 1 Curriculum Depth-Growth Distillation 0: Teacher components (,,lm)(E,N,W_lm), blocks K, layers per block L, loss weights Îą,γι,Îł, and phase budgets B0,âŚ,BKâ1B_0,âŚ,B_K-1. 0: A trained block-stack student. 1: Randomly initialize all student decoder blocks. 2: Copy E and lmW_lm; freeze E and initialize the placement-specific boundary norm from N. 3: for phase p=0p=0 to Kâ1K-1 do 4: Activate the first p+1p+1 blocks. 5: Reset optimizer state and restart the phase LR schedule. 6: for step t=1t=1 to BpB_p do 7: Obtain teacher and active-student logits. 8: Compute the KD+CE objective âpL_p. 9: Update all active decoder blocks, the boundary norm, and lmW_lm. 10: end for 11: end for Curriculum Depth Growth Growth as Inherited Initialization. Curriculum depth growth introduces decoder blocks in phases rather than activating the full stack from the beginning. At phase p, the active student contains the first p+1p+1 blocks, corresponding to LâĄ(p+1)L(p+1) decoder layers. The parameters learned in earlier phases are retained, and one additional block is activated when the model depth increases. Each phase therefore initializes a deeper model from a trained prefix and a newly activated, randomly initialized block. This protocol is related to progressive network growth and Net2Net-style expansion (4; 30; 9; 6; 31), but it does not require the expanded model to preserve the exact function of the preceding phase. Instead, we study how the trained prefix determines the input distribution received by the newly introduced block during subsequent distillation. This boundary distribution is part of the initialization condition faced by the appended block. Per-Phase Distillation Objective. At each phase, the active student produces (p)=StudentâĄ(,activeâ_âblocks=p+1),h^(p)=Student(x;\ active\_blocks=p+1), (4) and is trained with the same distillation objective: âp=ÎąâĎ2âKLâ(qĎteacherâĽqĎ(p))+ÎłâCEâ(,(p)),L_p=ÎąĎ^2KL\! (q_Ď^teacher\, \|\,q_Ď^(p) )+Îł\,CE\! (y,z^(p) ), (5) where qĎ=softmaxâĄ(/Ď)q_Ď=softmax(z/Ď), Ď is the distillation temperature, (p)z^(p) denotes the logits of the active student, and y denotes the next-token labels. All active decoder blocks, the boundary norm, and the language-modeling head are optimized jointly within each phase. Joint Training as the Comparison Protocol. Under joint training, all K blocks are active from the first optimization step and are trained end to end with the objective in Eq. 5. The two protocols use the same final architecture and objective but differ in staged block activation and, under the prescribed grow protocol, optimizer and learning-rate resets at phase boundaries. The experimental setup details these protocol differences and the token- and compute-matching comparisons. PlacementâCurriculum Interaction. Let âm,c _m,c denote final held-out CE for placement mâpre,postmâ\pre,post\ and curriculum câjoint,growcâ\joint,grow\. We measure the within-curriculum gap and factorial interaction as Îc=âpost,cââpre,c,â=ÎgrowâÎjoint. _c= _post,c- _pre,c, = _grow- _joint. (6) A negative âI means that post-norm becomes relatively more favorable under curriculum growth. Forward-Conditioning Hypothesis Curriculum depth growth creates a boundary condition specific to staged training. When a new block is introduced, it receives the residual stream produced by the prefix trained in the preceding phase. Normalization placement therefore affects not only computation within each layer but also the forward distribution received by each newly introduced block. Pre-norm and post-norm handle this boundary distribution differently. Pre-norm applies RMSNorm immediately before the attention and MLP sub-layers but does not explicitly normalize the representation after each residual update. Post-norm applies RMSNorm after every residual update, so the representation passed to the next layer or across a block boundary is explicitly normalized. We hypothesize that this distinction becomes important during curriculum depth growth. Each newly introduced block is optimized on top of the boundary distribution produced by the existing stack. Explicit normalization of block outputs may therefore provide a more stable input distribution for subsequent phases. We test this hypothesis by comparing pre-norm and post-norm under joint training and curriculum depth growth, controlling for single-block performance and active compute, and examining activation scales and block contributions at block boundaries. Testable Predictions. The hypothesis yields three observable predictions. First, post-norm need not outperform pre-norm before any block is appended, because the initial shallow model does not yet inherit a trained block boundary. Second, any placement gap should emerge after a depth transition rather than uniformly throughout training. Third, the placement favored by growth should exhibit better-controlled boundary representations and make nontrivial use of the appended blocks. These predictions separate forward conditioning from an unconditional post-norm advantage; the single-block, compute-matched, freeze, activation-scale, and block-removal analyses test the corresponding alternatives and observable consequences. Experiments We conduct a series of controlled experiments to investigate whether normalization placement remains a fixed architectural choice when Transformer depth is introduced through a curriculum. Using a block-stack distillation setting, we compare pre-norm and post-norm students under joint training and curriculum depth growth. The primary comparison matches the teacher, architecture, objective, data shard, and unique-token exposure; the curriculum protocol additionally changes active depth, optimization steps, data replay, and phase resets, with compute examined separately. We ask whether placement changes the jointâgrow ranking (RQ1), whether single-block quality or active compute explains the grow advantage (RQ2), and whether boundary dynamics support forward conditioning as a contributing mechanism (RQ3). (a) Placementâcurriculum interaction. (b) Phase-wise crossover under curriculum growth. Figure 3: Main placementâcurriculum results. The preâpost ranking changes under curriculum growth, and the crossover appears after the initial shallow phase. Experimental Setup Teacher, Student, and Data. We use the base Qwen3-8B model (25) as the teacher and train a block-stack student with K=3K=3 blocks and L=3L=3 decoder layers per block, giving S=9S=9 trainable layers in total. Following the block-stack distillation setting (10; 13), the student copies the teacher embedding E and language-modeling head lmW_lm. The embedding remains frozen, while the language-modeling head is optimized jointly with decoder layers initialized using the default Qwen3 module-wise scheme with initializer range 0.020.02. All students are distilled on the FineWeb-Edu 10BT split (20). All reported primary runs use seed 42, a global batch size of 6464, and a maximum sequence length of 81928192 across all configurations and phases. Compared Configurations. We compare four configurations formed by crossing normalization placement and depth-introduction protocol: pre-norm, post-normĂjoint, grow. In joint training, all blocks are active from the first update and are trained end-to-end. In curriculum growth, blocks are appended sequentially following Algorithm 1: each phase activates one additional block, carries over the previously trained weights, and introduces a newly initialized block when depth increases. Within each paired placement comparison, the teacher, student architecture, data order, objective, optimizer configuration, and initialization protocol are matched. Joint and grow then follow their respective activation, duration, replay, and reset schedules. Training Budget. Each grow phase runs for 50,00050,000 steps. Because sequences are variable-length and most documents fall well below the 81928192-token cap, each phase processes approximately 3.113.11B non-padding tokens rather than the fully-packed upper bound. The three grow phases consume the same data shard, so grow sees 3.113.11B unique tokens repeated across phases (9.349.34B total). Joint training runs 50,00050,000 steps over the same 3.113.11B-token shard. At every phase transition, the optimizer state is reset and the learning-rate schedule is restarted with warmup. Because growth activates fewer layers in early phases, we report a unique-token-matched primary comparison and active-layer-token controls; Appendix Table A2 summarizes the budget for each configuration. The primary setting matches unique-token exposure, while the compute proxy matches student active-layer tokens, computed as training tokens multiplied by the number of active decoder layers. Evaluation Metrics. The primary metric is validation cross-entropy (CE) loss â on a fixed held-out set of 4,8334,833 documents (approximately 4.984.98M tokens), evaluated every 2,0002,000 steps, with perplexity (PPL) â reported as an equivalent scale. For mechanism analysis, we measure block-boundary residual RMS across grow checkpoints and inspect validation-loss trajectories around phase transitions. Main Results For RQ1, we evaluate whether the relative behavior of pre-norm and post-norm changes across depth-introduction protocols. We compare the two placements under matched architecture, data, objective, and unique-token exposure; grow performs more optimization steps, examined through active-layer-token controls. PlacementâCurriculum Interaction. Table 1 and Figure 3(a) report the primary comparison. Under joint training, pre-norm and post-norm reach nearly identical validation losses, with CE values of 2.76032.7603 and 2.76072.7607, respectively. Under curriculum growth, the ranking changes: post-norm reduces validation CE from 2.76582.7658 to 2.73302.7330 compared with pre-norm. This yields Îjoint=+0.0004 _joint=+0.0004 and Îgrow=â0.0328 _grow=-0.0328, where Î=PostâPre =Post-Pre. The joint gap indicates how far the two placements drift apart when depth is held fixed, and the grow gap is roughly eighty times larger. The resulting interaction, ÎgrowâÎjoint=â0.0332 _grow- _joint=-0.0332, quantifies the observed shift toward post-norm under curriculum growth in this distillation setting. Unique-token-matched Joint Grow Pre-norm 2.7603 2.7658 Post-norm 2.7607 2.7330 Î (Postâ-Pre) ++0.0004 â-0.0328 Interaction â-0.0332 Table 1: Primary placementâcurriculum interaction. Validation CE loss â under unique-token-matched budgets. Î=PostâPre =Post-Pre, so negative values indicate a post-norm advantage. Phase-wise Crossover. The per-phase results further show where the ranking changes during curriculum growth. As shown in Figure 3(b), pre-norm is slightly better in the shallow Phase 0, where the student contains only one block. After additional blocks are appended, post-norm becomes better in both Phase 1 and Phase 2. This crossover matches the expected shift from optimizing a shallow block to conditioning the boundary distribution passed to appended blocks. We next use single-block and active-layer-token controls to test whether shallow-block quality or longer optimization can explain the pattern. Controlled Analysis For RQ2, we examine whether the post-norm advantage under curriculum growth can be explained by simpler alternatives. We test two possibilities under the same held-out evaluation: post-norm may form stronger individual blocks, or curriculum growth may benefit from additional student-decoder active compute. Single-Block Control. We first test whether post-norm is already stronger when training a single block in isolation. Table 2 reports Phase 0 results, where the student contains only one block (L=3L=3). Pre-norm is marginally ahead on CE, KD loss, and PPL, with all gaps below 0.010.01 in CE/KD, which rules out post-norm forming a better shallow block as the source of its grow advantage; instead, the advantage appears only after additional blocks are appended, matching the phase-wise crossover in Figure 3(b). Norm CE â KD â PPL â Pre 2.9802 3.0155 19.6900 Post 2.9846 3.0183 19.7800 Table 2: Single-block Phase 0 results. CE, KD loss, and PPL are reported for one-block pre-norm and post-norm students. Active-Layer-Token Control. We next test whether additional student-decoder active compute explains the grow advantage. Because curriculum growth traverses three phases while joint training uses one full-depth phase, we extend post-joint to matched active-layer budgets; these controls use phase-aligned 50k-step restarts at fixed depth and continue onto fresh data rather than replaying the grow shard. Table 3 shows that post-joint improves from 2.76072.7607 to 2.75182.7518 when matched to grow on student active-layer tokens, and to 2.74282.7428 when matched to grow on total training tokens while using more student-decoder active compute. Both nonetheless remain worse than post-grow (2.73302.7330), despite the additional compute. This suggests that the post-norm grow result is not reducible to additional student active-layer compute under the evaluated schedules. Run Matches grow on CE â Post-Joint Unique tokens 2.7607 Post-Joint-long Active-layer tokens 2.7518 Post-Joint-longer Total tokens 2.7428 Table 3: Post-joint budget controls. Validation CE loss â under three matching criteria. Figure 4: Block-boundary RMS across growth phases. Boundary-Scale Diagnostics For RQ3, we test whether block-boundary dynamics support forward conditioning as a contributing mechanism. At each depth transition, a newly appended block receives the residual stream produced by the prefix trained in the preceding phase. We examine residual-stream scale and token composition, appended-block contributions, transition-localized loss, and the freeze control, which together characterize how the boundary distribution evolves as depth grows. (a) Token-type attribution of scale drift. (b) Validation-loss trajectories and grow transitions. Figure 5: Forward-conditioning diagnostics. Post-norm exhibits limited structural-token scale drift, and the placement gap emerges after block-appending transitions. Block-Boundary Scale Drift. Figure 4 shows that post-norm keeps block-boundary RMS controlled, whereas pre-norm drifts as depth grows: after block 0, the per-token RMS standard deviation rises from 5.815.81 to 19.9819.98 under pre-norm but remains approximately 0.070.07â0.190.19 under post-norm. On the final-phase diagnostic batch (Appendix Figure A1), pre-norm shows a heavy-tailed distribution (kurtosis >2400>2400) versus a far less heavy-tailed post-norm distribution. Token-Type Attribution of Scale Drift. The extreme kurtosis suggests that pre-normâs drift is driven by a small subset of tokens rather than a uniform shift. Figure 5(a) localizes pre-normâs drift to structural tokens: document-boundary RMS grows from 119119 to 425425 (37Ă37Ă regular-token RMS at 150k), while regular-token RMS grows only from 7.17.1 to 11.511.5; post-norm keeps both below 2.52.5 (2.452.45 vs. 2.402.40), so the drift is concentrated at structural-token massive activations rather than a uniform global shift. Block-Removal Sensitivity and Representation Change. On a fixed 2Ă20482Ă 2048-token diagnostic batch (seed 42), we bypass each entire block and measure ÎâCEl=CEâĄ(skip block âl)âCEâĄ(full model) _l=CE(skip block l)-CE(full model), together with angular distance arccosâĄ(cosineâĄ(in,out))/Ď (cosine(h_in,h_out))/Ď. For a post-norm three-layer block, bypass removes its attention and MLP modules together with six post-residual RMSNorm operations; the model entry norm remains outside the skipped block. Thus, ÎâCE measures whole-block removal sensitivity rather than the attention and MLP contribution alone. Under pre-norm growth, bypassing the final block changes CE by only 0.020.02, and its angular distance is 0.0310.031, indicating a near-identity mapping. Removing any post-grow block instead raises CE by at least 2.382.38, with angular distances above 0.340.34; neither joint model has a near-identity final block. Post-joint places low removal sensitivity on block 0 (ÎâCE=0.55 =0.55) but high sensitivity on block 2 (4.004.00), showing that uneven allocation alone does not imply the pre-grow failure pattern. Thus, scale drift and a nearly identity-mapped final block coincide under pre-norm growth, while post-norm retains controlled scales and nontrivial representation changes across blocks. Run Metric Block 0 Block 1 Block 2 Pre-Grow Î â 8.13 2.36 0.02 Angular dist. 0.499 0.310 0.031 Post-Grow Î â 8.32 3.74 2.38 Angular dist. 0.483 0.365 0.342 Pre-Joint Î â 8.05 1.47 2.34 Angular dist. 0.485 0.245 0.360 Post-Joint Î â 0.55 0.99 4.00 Angular dist. 0.395 0.166 0.437 Table 4: Block-removal sensitivity and representation change. Î includes removal of internal normalization; angular distance measures inputâoutput change. Only Pre-Grow has a near-identity final block (Î â0â 0, angular â0.03â 0.03). Transition-Localized Loss Divergence. Figure 5(b) shows that the placement gap emerges after blocks are appended. Transition-aligned loss reveals that post-norm incurs larger immediate increases at both transitions, with much slower recovery after the first and only a small delay after the second (Appendix Table A4), whereas pre-norm changes more smoothly. The smooth pre-norm transition is consistent with a smaller initial perturbation, yet its final block is nearly identity-mapped; post-norm incurs a larger adjustment but develops nontrivial block changes. Thus, boundary-scale control should not be interpreted as a smoother immediate transition: transition smoothness and final block utilization need not coincide in general. Freeze Control. Finally, we freeze completed blocks and update only the newly appended block under matched budgets and transition policies. Freezing worsens both placements by approximately 0.1440.144 CE but leaves the gap nearly unchanged (â0.0328-0.0328 vs. â0.0332-0.0332), so retraining earlier blocks does not account for the placement gap (Appendix Table A3). These diagnostics observe rather than intervene on the boundary distribution, and together they provide converging evidence for boundary-scale conditioning. Conclusion Pre-norm and post-norm are nearly tied under joint training, whereas post-norm is better under curriculum depth growing. Single-block, compute-matched, freeze, and boundary diagnostics associate this crossover with controlled boundary scales, ruling out shallow-block quality, active-layer compute, or retraining alone. The results therefore motivate treating normalization placement and depth curriculum as coupled rather than independent design choices. References Ainslie et al. (2023) J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. LebrĂłn, and S. Sanghai GQA: training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), p. 4895â4901. External Links: Link, Document Cited by: Block-Stack Student.. Ba et al. (2016) L. J. Ba, J. R. Kiros, and G. E. Hinton Layer normalization. CoRR abs/1607.06450. External Links: Link, 1607.06450 Cited by: Normalization Placement in Transformers.. Bachlechner et al. (2021) T. Bachlechner, B. P. Majumder, H. H. Mao, G. Cottrell, and J. J. McAuley ReZero is all you need: fast convergence at large depth. In Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, UAI 2021, Virtual Event, 27-30 July 2021, C. P. de Campos, M. H. Maathuis, and E. Quaeghebeur (Eds.), Proceedings of Machine Learning Research, Vol. 161, p. 1352â1361. External Links: Link Cited by: Normalization Placement in Transformers.. Chen et al. (2016) T. Chen, I. J. Goodfellow, and J. Shlens Net2Net: accelerating learning via knowledge transfer. In 4th International Conference on Learning Representations, ICLR 2016, San Juan, Puerto Rico, May 2-4, 2016, Conference Track Proceedings, Y. Bengio and Y. LeCun (Eds.), External Links: Link Cited by: Introduction, Network Growth and Function Preservation., Growth as Inherited Initialization.. Ding et al. (2021) M. Ding, Z. Yang, W. Hong, W. Zheng, C. Zhou, D. Yin, J. Lin, X. Zou, Z. Shao, H. Yang, and J. Tang CogView: mastering text-to-image generation via transformers. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, M. Ranzato, A. Beygelzimer, Y. N. Dauphin, P. Liang, and J. W. Vaughan (Eds.), p. 19822â19835. External Links: Link Cited by: Introduction, Normalization Placement in Transformers.. Du et al. (2024) W. Du, T. Luo, Z. Qiu, Z. Huang, Y. Shen, R. Cheng, Y. Guo, and J. Fu Stacking your transformers: A closer look at model growth for efficient LLM pre-training. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), External Links: Link Cited by: Introduction, Network Growth and Function Preservation., Growth as Inherited Initialization.. Elhoushi et al. (2024) M. Elhoushi, A. Shrivastava, D. Liskovich, B. Hosmer, B. Wasti, L. Lai, A. Mahmoud, B. Acun, S. Agarwal, A. Roman, A. A. Aly, B. Chen, and C. Wu LayerSkip: enabling early exit inference and self-speculative decoding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), p. 12622â12642. External Links: Link, Document Cited by: Introduction, Block-Stack Knowledge Distillation.. Gou et al. (2021) J. Gou, B. Yu, S. J. Maybank, and D. Tao Knowledge distillation: A survey. Int. J. Comput. Vis. 129 (6), p. 1789â1819. External Links: Link, Document Cited by: Block-Stack Knowledge Distillation.. Gu et al. (2021) X. Gu, L. Liu, H. Yu, J. Li, C. Chen, and J. Han On the transformer growth for progressive BERT training. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-TĂźr, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), p. 5174â5180. External Links: Link, Document Cited by: Introduction, Network Growth and Function Preservation., Growth as Inherited Initialization.. Hinton et al. (2015) G. E. Hinton, O. Vinyals, and J. Dean Distilling the knowledge in a neural network. CoRR abs/1503.02531. External Links: Link, 1503.02531 Cited by: Block-Stack Knowledge Distillation., Block-Stack Student., Teacher, Student, and Data.. Huang et al. (2016) G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger Deep networks with stochastic depth. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part IV, B. Leibe, J. Matas, N. Sebe, and M. Welling (Eds.), Lecture Notes in Computer Science, Vol. 9908, p. 646â661. External Links: Link, Document Cited by: Block-Stack Knowledge Distillation.. Huang et al. (2020) X. S. Huang, F. PĂŠrez, J. Ba, and M. Volkovs Improving transformer optimization through better initialization. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Proceedings of Machine Learning Research, Vol. 119, p. 4475â4483. External Links: Link Cited by: Normalization Placement in Transformers.. Jiao et al. (2020) X. Jiao, Y. Yin, L. Shang, X. Jiang, X. Chen, L. Li, F. Wang, and Q. Liu TinyBERT: distilling BERT for natural language understanding. In Findings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020, T. Cohn, Y. He, and Y. Liu (Eds.), Findings of ACL, Vol. EMNLP 2020, p. 4163â4174. External Links: Link, Document Cited by: Block-Stack Knowledge Distillation., Block-Stack Student., Teacher, Student, and Data.. Kan et al. (2025) K. Kan, X. Li, B. J. Zhang, T. Sahai, S. J. Osher, K. Kumar, and M. A. Katsoulakis Stability of transformers under layer normalization. CoRR abs/2510.09904. External Links: Link, Document, 2510.09904 Cited by: Normalization Placement in Transformers.. Kim et al. (2025) J. Kim, B. Lee, C. Park, Y. Oh, B. Kim, T. Yoo, S. Shin, D. Han, J. Shin, and K. M. Yoo Peri-ln: revisiting normalization layer in the transformer architecture. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267. External Links: Link Cited by: Normalization Placement in Transformers.. Li et al. (2026) T. Li, D. Han, Z. Cao, H. Huang, M. Zhou, M. Chen, E. Zhao, X. Jiang, G. Jiang, and G. Huang SiameseNorm: breaking the barrier to reconciling pre/post-norm. CoRR abs/2602.08064. External Links: Link, Document, 2602.08064 Cited by: Normalization Placement in Transformers.. Liu et al. (2020) L. Liu, X. Liu, J. Gao, W. Chen, and J. Han Understanding the difficulty of training transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, B. Webber, T. Cohn, Y. He, and Y. Liu (Eds.), p. 5747â5763. External Links: Link, Document Cited by: Introduction, Normalization Placement in Transformers.. Loshchilov et al. (2025) I. Loshchilov, C. Hsieh, S. Sun, and B. Ginsburg NGPT: normalized transformer with representation learning on the hypersphere. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: Normalization Placement in Transformers.. Men et al. (2025) X. Men, M. Xu, Q. Zhang, Q. Yuan, B. Wang, H. Lin, Y. Lu, X. Han, and W. Chen ShortGPT: layers in large language models are more redundant than you expect. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, p. 20192â20204. External Links: Link, Document Cited by: Block-Stack Knowledge Distillation.. Penedo et al. (2024) G. Penedo, H. KydlĂcek, L. B. Allal, A. Lozhkov, M. Mitchell, C. A. Raffel, L. von Werra, and T. Wolf The fineweb datasets: decanting the web for the finest text data at scale. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), External Links: Link Cited by: Teacher, Student, and Data.. Sanh et al. (2019) V. Sanh, L. Debut, J. Chaumond, and T. Wolf DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. CoRR abs/1910.01108. External Links: Link, 1910.01108 Cited by: Block-Stack Knowledge Distillation.. Shazeer (2020) N. Shazeer GLU variants improve transformer. CoRR abs/2002.05202. External Links: Link, 2002.05202 Cited by: Block-Stack Student.. Su et al. (2024) J. Su, M. H. M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu RoFormer: enhanced transformer with rotary position embedding. Neurocomputing 568, p. 127063. External Links: Link, Document Cited by: Block-Stack Student.. Sun et al. (2020) Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou MobileBERT: a compact task-agnostic BERT for resource-limited devices. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, D. Jurafsky, J. Chai, N. Schluter, and J. R. Tetreault (Eds.), p. 2158â2170. External Links: Link, Document Cited by: Block-Stack Knowledge Distillation.. Team (2025) Q. Team Qwen3 technical report. CoRR abs/2505.09388. External Links: Link, Document, 2505.09388 Cited by: Normalization Placement in Transformers., Teacher, Student, and Data.. Touvron et al. (2023) H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, A. Rodriguez, A. Joulin, E. Grave, and G. Lample LLaMA: open and efficient foundation language models. CoRR abs/2302.13971. External Links: Link, Document, 2302.13971 Cited by: Introduction, Introduction, Normalization Placement in Transformers.. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, I. Guyon, U. von Luxburg, S. Bengio, H. M. Wallach, R. Fergus, S. V. N. Vishwanathan, and R. Garnett (Eds.), p. 5998â6008. External Links: Link Cited by: Normalization Placement in Transformers.. Wang et al. (2024) H. Wang, S. Ma, L. Dong, S. Huang, D. Zhang, and F. Wei DeepNet: scaling transformers to 1,000 layers. IEEE Trans. Pattern Anal. Mach. Intell. 46 (10), p. 6761â6774. External Links: Link, Document Cited by: Introduction, Normalization Placement in Transformers.. Wang et al. (2020) W. Wang, F. Wei, L. Dong, H. Bao, N. Yang, and M. Zhou MiniLM: deep self-attention distillation for task-agnostic compression of pre-trained transformers. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin (Eds.), External Links: Link Cited by: Block-Stack Knowledge Distillation.. Wei et al. (2016) T. Wei, C. Wang, Y. Rui, and C. W. Chen Network morphism. In Proceedings of the 33nd International Conference on Machine Learning, ICML 2016, New York City, NY, USA, June 19-24, 2016, M. Balcan and K. Q. Weinberger (Eds.), JMLR Workshop and Conference Proceedings, Vol. 48, p. 564â572. External Links: Link Cited by: Introduction, Network Growth and Function Preservation., Growth as Inherited Initialization.. Wu et al. (2024) C. Wu, Y. Gan, Y. Ge, Z. Lu, J. Wang, Y. Feng, Y. Shan, and P. Luo LLaMA pro: progressive llama with block expansion. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), p. 6518â6537. External Links: Link, Document Cited by: Network Growth and Function Preservation., Growth as Inherited Initialization.. Xia et al. (2024) M. Xia, T. Gao, Z. Zeng, and D. Chen Sheared llama: accelerating language model pre-training via structured pruning. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: Block-Stack Knowledge Distillation.. Xiong et al. (2020) R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T. Liu On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event, Proceedings of Machine Learning Research, Vol. 119, p. 10524â10533. External Links: Link Cited by: Introduction, Normalization Placement in Transformers.. Zhang and Sennrich (2019) B. Zhang and R. Sennrich Root mean square layer normalization. In Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, H. M. Wallach, H. Larochelle, A. Beygelzimer, F. dâAlchĂŠ-Buc, E. B. Fox, and R. Garnett (Eds.), p. 12360â12371. External Links: Link Cited by: Block-Stack Student.. Zhang et al. (2019) H. Zhang, Y. N. Dauphin, and T. Ma Fixup initialization: residual learning without normalization. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, External Links: Link Cited by: Normalization Placement in Transformers.. Zheng et al. (2026) C. Zheng, J. Sun, Y. Gao, C. Wang, Y. Wang, J. Xiong, L. Ren, B. Peng, Q. Wang, X. Shang, M. Schwager, A. Schneider, Y. Nevmyvaka, and X. Liu GeoNorm: unify pre-norm and post-norm with geodesic optimization. CoRR abs/2601.22095. External Links: Link, Document, 2601.22095 Cited by: Normalization Placement in Transformers.. Zhuo et al. (2025) Z. Zhuo, Y. Zeng, Y. Wang, S. Zhang, J. Yang, X. Li, X. Zhou, and J. Ma HybridNorm: towards stable and efficient transformer training via hybrid normalization. CoRR abs/2503.04598. External Links: Link, Document, 2503.04598 Cited by: Normalization Placement in Transformers.. Appendix A Appendix A: Training and Compute Protocol A.1 Detailed Hyperparameter Settings Table A1 lists shared hyperparameters. Training uses AdamW with cosine decay and warmup. A.2 Budget and Compute Protocol Because curriculum growth activates fewer layers in early phases, step-matched budgets are not compute-fair. We use a unique-token-matched primary comparison and additional post-joint active-layer-token controls: ⢠Unique-token-matched: joint and grow see the same unique tokens. This compares quality at equal data coverage, although grow uses more active-layer compute across three phases. ⢠Active-layer-token control: post-joint is extended to the active-layer-token budget of grow, i.e. trained tokens multiplied by the number of active decoder layers at each step. We additionally report a 3Ă3Ă post-joint run beyond that proxy-matched budget. Compute can be approximated by active-layer tokens: active-layer tokens=âstep ât(tokens at ât)â (active layers at ât).active-layer tokens= _step t(tokens at t)¡(active layers at t). Hyperparameter Value Optimizer AdamW Optimizer β1,β2 _1, _2 0.9, 0.950.9,\ 0.95 Optimizer Ͼξ 1âe-â81e-8 Weight Decay 00 Peak Learning Rate 1âe-â41e-4 Minimum Learning Rate 1âe-â61e-6 Learning Rate Schedule Cosine decay with warmup Warmup Steps 0.015Ă50,000=7500.015Ă 50,000=750 Gradient Clipping 0.50.5 (global L2 norm) Training Precision BF16 Global Batch Size 6464 Sequence Length 81928192 (max) Evaluation Interval 2,0002,000 steps Distillation Weight Îą 0.70.7 Language Modeling Weight Îł 0.30.3 Temperature Ď 1.01.0 Dropout 00 Attention Dropout 00 Seed 4242 Table A1: Hyperparameters shared across the primary 2Ă22Ă 2 configurations. Controls additionally vary training duration or trainable blocks as stated in their protocols. Config Phases Layers Total tok. Unique tok. Joint-50k 1 9 3.11B 3.11B Grow-3Ă50k 3 3â 6â 9 9.34B 3.11B Joint (proxy-matched) 1 9 6.22B 6.22B Joint-150k 1 9 9.34B 9.34B Table A2: Training budgets. Grow reuses one fixed shard, whereas longer joint controls continue onto fresh data; proxy-matched joint matches student active-layer tokens, and Joint-150k is the 3Ă3Ă stress test. A.3 Phase Transition Optimization Reset At each curriculum growth transition, we reset the optimizer state and restart the learning-rate schedule with warmup for two reasons: 1. Depth increase. Moving from p to p+1p+1 blocks introduces newly initialized weights. Restarting the LR schedule with warmup updates these layers smoothly instead of exposing them to stale momentum. 2. Stale gradient statistics. Resetting Adam momentum prevents statistics accumulated in the shallower phase from biasing the deeper network. Thus, each phase inherits model weights but re-estimates optimizer statistics rather than carrying shallower-phase momentum across successive transitions. Appendix B Appendix B: Additional Analyses B.1 Freeze Ablation: Probing Forward Conditioning Setup. The main grow protocol retrains all active blocks in every phase. A natural alternative is to freeze each block once its phase completes, so that phase p trains only the newly appended block p while blocks 0,âŚ,pâ10,âŚ,p-1 are held fixed. Freeze and retrain differ in two ways: (i) freeze gives weaker absolute performance because earlier blocks cannot adapt to the deeper stack, and (i) under freeze the input distribution to a newly appended block is strictly fixed (it is the exact output of frozen blocks), which sharpens the forward-conditioning channel. If retraining primarily compensated for pre-normâs boundary drift, the postâpre gap would be expected to amplify under freeze. Instead, the observed gap is essentially unchanged. However, freeze does not fully isolate the forward channel: the newly appended block is still trained, so its backward-gradient dynamics at the append step still differ across placements. The result is therefore compatible with a forward-conditioning contribution, but does not exclude a backward-at-append contribution. Protocol. The freeze variant follows the grow protocol of the main paper exactly, except that at the start of phase p+1p\!+\!1 the parameters of blocks 0,âŚ,p0,âŚ,p are frozen (no gradient, no optimizer state); only block p+1p\!+\!1 receives updates. Step budgets, LR schedule, data shards, and the reset-at-transition policy are identical to grow. Pre-norm and post-norm use the same paired seed and initialization as the main runs. Result. Table A3 reports validation CE at 150k. Freeze is uniformly worse than retrain by â0.144â 0.144 CE for both placements, showing that retraining earlier blocks is beneficial in absolute terms. The placement gap, however, is essentially unchanged: Îgrow=â0.0328 _grow=-0.0328 versus Îfreeze=â0.0332 _freeze=-0.0332, a difference of â0.0004-0.0004 (â1.2%â 1.2\% of the gap). On PPL the gap is wider under freeze (â0.60-0.60 vs â0.51-0.51), but PPL is the exponential of CE, so this amplification is consistent with an unchanged CE gap. Grow (retrain) Freeze Pre-norm 2.7658 2.9104 Post-norm 2.7330 2.8772 Î (Postâ-Pre) â-0.0328 â-0.0332 ÎfreezeâÎgrow _freeze- _grow â-0.0004 Table A3: Freeze ablation. Validation CE loss â at 150k under the unique-token-matched budget (grow and freeze both traverse three 50k-step phases over the same shard). Grow reproduces the main interaction (retrain); freeze holds previously appended blocks fixed. Î=PostâPre =Post-Pre (negative == post wins). The placement gap is essentially unchanged across retrain/freeze (|Î|| | differs by 0.00040.0004). Interpretation. The result is consistent with the forward-conditioning account, with two caveats. First, the absolute freeze penalty (â0.144â 0.144 CE) is nearly identical across placements, so a retraining-only explanation for post-normâs relative advantage is unlikely under this control. Second, the gapâs near-invariance to retrain/freeze is consistent with the RQ3 RMS evidence: pre-normâs boundary drift is already present in the outputs of the prefix at the append step, whether or not those blocks are retrained later. The result therefore weighs against the stronger claim that retraining actively masks pre-normâs drift, and suggests that later joint updates do not recover the full append-time cost. We emphasize what freeze does not establish: because the newly appended block remains trainable, its backward-gradient dynamics at the append step still differ by placement, so the gapâs stability is also compatible with a backward-at-append channel. Forward conditioning is therefore a candidate explanation for post-normâs grow advantage, but isolating it as the sole cause would require an intervention that changes the forward boundary distribution without changing backward updates. B.2 Mechanism Diagnostics The main paper reports the primary interaction and phase-wise crossover in Figure 3, and the block-boundary RMS trend, token-type attribution, and transition-localized loss trajectories in Figures 4 and 5. Figures A1 and A2 complement those aggregate trends with the final-checkpoint RMS distribution and per-block deletion behavior. Metric 50k (Pre/Post) 100k (Pre/Post) Loss jump 0.0226 / 0.2500 0.0230 / 0.2847 AUC1k 0.0694 / 0.1820 0.0688 / 0.2059 AUC5k 0.1217 / 0.1464 0.1277 / 0.1774 Recovery steps 101 / 19,914 19,054 / 19,426 Table A4: Transition-aligned training-loss response at the 50k and 100k block-appending steps. Jump and AUC are measured relative to the pre-transition baseline; recovery is the first step at which a 200-step smoothed loss returns to that baseline. Post-norm pays a larger immediate adjustment cost despite its better final grow loss. Figure A1: Final-phase per-token RMS distributions. Figure A2: Whole-block removal sensitivity and representation change. Whole-block bypass also removes internal normalization operations. B.3 Additional Protocol Checks Data-Iterator Reset. Within curriculum growth, we compared resetting the data iterator at phase boundaries with continuing from its current position. Both variants draw from the same fixed 3.11B-token shard and wrap within that shard after reaching its end; a phase transition never switches to new data. They produce similar final losses and the same placement ranking. The main results use the no-reset variant, so a phase change continues from the current offset in this repeating cycle rather than explicitly rewinding it, while unique-token coverage remains 3.11B. The longer joint controls instead continue onto fresh data, so their total- and unique-token counts coincide. Joint Learning-Rate Restarts. The 100k and 150k post-joint controls in Table A2 keep all nine decoder layers active and restart the 50k-step learning-rate schedule at the same boundaries used by grow. Each restart produces a transient loss increase followed by renewed descent, but the subsequent decrease is slower than the corresponding curriculum-growth trajectory. Their endpoint CE values improve to 2.75182.7518 and 2.74282.7428, respectively, yet remain above post-grow at 2.73302.7330; LR restarting alone therefore does not reproduce the grow trajectory. Appendix C Appendix C: Implementation Details Model and Systems. The Qwen3-8B-Base teacher has 36 decoder layers. The nine-layer student uses hidden size 4096, 32 attention heads, 8 keyâvalue heads, and intermediate size 12288. It contains approximately 3.09B parameters including the teacher-initialized embedding and language-modeling head, 2.47B when excluding the head but retaining the embedding, and 1.85B in the decoder stack alone. The embedding is frozen, whereas the language-modeling head is trained. Training uses FSDP over 8 GPUs, while teacher inference uses a separate 8-GPU pool; gradient checkpointing and MLP recomputation are enabled. Data Processing and Evaluation. Examples are tokenized independently, truncated to 8192 tokens, and right-padded to the longest sequence in each batch; multiple documents are not packed into one sequence. Padding is excluded from the token accounting and loss. The fixed validation set contains 4,833 documents and approximately 4.98M tokens, and evaluation is run every 2,000 optimization steps. Main tables report the scheduled endpoints rather than selecting the best validation checkpoint.