Paper deep dive
From Growing to Looping: A Unified View of Iterative Computation in LLMs
Ferdinand Kapl, Emmanouil Angelis, Kaitlin Maile, Johannes von Oswald, Stefan Bauer
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/21/2026, 1:35:50 AM
Summary
This paper provides a mechanistic unification of looped models (which reuse layers via weight tying) and depth-grown models (which duplicate middle layers during training). It demonstrates that both approaches induce similar depth-wise computational signatures, such as increased reliance on late layers and periodic patterns aligned with the looped or grown block, supporting the view that they share a common form of iterative computation. The study shows that these techniques are adaptable and composable; applying inference-time looping to depth-grown models improves reasoning accuracy by up to 2x. Additionally, depth-grown models perform best with high-quality, math-heavy cooldown mixtures, and the combination of growth and looping offers complementary benefits for scaling reasoning capabilities.
Entities (8)
Relation Signals (7)
Looped Models â exhibits â Iterative Computation
confidence 95% ¡ These shared signatures support the view that their gains stem from a common form of iterative computation.
Depth-Grown Models â exhibits â Iterative Computation
confidence 95% ¡ These shared signatures support the view that their gains stem from a common form of iterative computation.
MIDAS â isatypeof â Depth-Grown Models
confidence 95% ¡ Saunshi et al. [32] introduce MIDAS, which partitions the network into blocks and duplicates the middle block
LIDAS â isatypeof â Depth-Grown Models
confidence 95% ¡ They propose LIDAS, which duplicates the exact layer-wise middle
Looped Models â composablewith â Depth-Grown Models
confidence 92% ¡ applying inference-time looping to the middle blocks of a depth-grown model improves accuracy on some reasoning primitives by up to 2x
Reasoning Primitives â improvedby â Looped Models
confidence 90% ¡ applying inference-time looping to the middle blocks of a depth-grown model improves accuracy on some reasoning primitives by up to 2x
Looped Models â sharessignaturewith â Depth-Grown Models
confidence 90% ¡ looped and depth-grown models exhibit convergent depth-wise signatures
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Looping, reusing a block of layers across depth, and depth growing, training shallow-to-deep models by duplicating middle layers, have both been linked to stronger reasoning, but their relationship remains unclear. We provide a mechanistic unification: looped and depth-grown models exhibit convergent depth-wise signatures, including increased reliance on late layers and recurring patterns aligned with the looped or grown block. These shared signatures support the view that their gains stem from a common form of iterative computation. Building on this connection, we show that the two techniques are adaptable and composable: applying inference-time looping to the middle blocks of a depth-grown model improves accuracy on some reasoning primitives by up to $2\times$, despite the model never being trained to loop. Both approaches also adapt better than the baseline when given more in-context examples or additional supervised fine-tuning data. Additionally, depth-grown models achieve the largest reasoning gains when using higher-quality, math-heavy cooldown mixtures, which can be further boosted by adapting a middle block to loop. Overall, our results position depth growth and looping as complementary, practical methods for inducing and scaling iterative computation to improve reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2602.16490v1
- Canonical: https://arxiv.org/abs/2602.16490v1
Trouble viewing inline? Open PDF directly â
Full Text
65,825 characters extracted from source content.
Expand or collapse full text
From Growing to Looping: A Unified View of Iterative Computation in LLMs Ferdinand Kapl 1,2â⥠Emmanouil Angelis 1,2â⥠Kaitlin Maile 3â Johannes von Oswald 3â Stefan Bauer 1,2â 1 Technical University of Munich 2 Helmholtz AI, Munich 3 Google, Paradigms of Intelligence Team Abstract Looping, reusing a block of layers across depth, and depth growing, training shallow-to-deep models by duplicating middle layers, have both been linked to stronger reasoning, but their relationship remains unclear. We provide a mech- anistic unification: looped and depth-grown models exhibit convergent depth- wise signatures, including increased reliance on late layers and recurring patterns aligned with the looped or grown block. These shared signatures support the view that their gains stem from a common form of iterative computation. Building on this connection, we show that the two techniques are adaptable and composable: applying inference-time looping to the middle blocks of a depth-grown model im- proves accuracy on some reasoning primitives by up to 2Ă, despite the model never being trained to loop. Both approaches also adapt better than the baseline when given more in-context examples or additional supervised fine-tuning data. Additionally, depth-grown models achieve the largest reasoning gains when using higher-quality, math-heavy cooldown mixtures, which can be further boosted by adapting a middle block to loop. Overall, our results position depth growth and looping as complementary, practical methods for inducing and scaling iterative computation to improve reasoning. 1 Introduction The dominant paradigm for improving the performance of Large Language Models (LLMs) has been to scale both parameters and data simultaneously [17, 21]. However, reasoning capabilities, which are often framed as the ability to perform multi-step logical deductions, do not always scale linearly with parameter count alone [39, 41]. A popular way to increase reasoning performance at inference time is to generate longer textual traces via chain-of-thought [40], but this scales compute through tokens and does not directly encourage internal computation. A complementary direction is to build iteration into the modelâs forward pass: repeatedly applying transformations in latent space so representations can be refined over multiple steps before producing the final prediction [3, 15, 23, 27, 33, 45, 46]. A prominent approach in this direction is looped models (Universal Transformers), motivated by the goal of decoupling model size from computational depth. By tying weights across layers and execut- ing the same block recurrently, looped models can perform deeper computations without increasing the number of unique parameters [10, 11, 25]. While early investigations focused mostly on param- eter efficiency, recent findings highlight their potential for reasoning. Previous work [15, 33, 46] observed that looped models can scale latent computations to increase reasoning performance. Sim- ilarly, approaches like Jolicoeur-Martineau [18], Wang et al. [37] utilize recursive architectures to maximize the reasoning capacity of notably small networks. Growing models, an alternative paradigm for training, was originally motivated by increased training efficiency through parameter growth [16, 31, 38]. By initializing a shallow model and progressively * Equal contribution. â Provided equal in-depth feedback and guidance. ⥠Correspondence:ferdinand.kapl,emmanouil.angelis@tum.de. arXiv:2602.16490v1 [cs.CL] 18 Feb 2026 100M1B Total Weights 20 30 40 Reasoning Primitives Parameters 10 14 10 15 10 16 FLOPs Inference 10 20 10 21 10 22 FLOPs Training StandardLoopedGrown Figure 1: Trade-offs for looped and depth-grown models. Each point corresponds to a model in Table 1 (up to 1.7B parameters), plotted by average Reasoning Primitives accuracy versus unique parameters (left), inference FLOPs (middle), and training FLOPs (right). Looped and depth-grown models improve accuracy in reasoning primitives over standard baselines, suggesting a shared inductive bias toward better reasoning. Looped models improve reasoning under fixed parameter budgets and can be competitive under fixed inference budgets, while depth-grown models reach similar or better reasoning with less training compute. Fig. 10 shows additional benchmark categories. adding layers during training, the total training FLOPs required to reach a target depth can be sig- nificantly reduced. Additionally, recent work discovered an intriguing inductive bias. Saunshi et al. [32] and Kapl et al. [20] demonstrated that models trained via a particular type of growing, i.e., via duplication of blocks in the middle of the transformer architecture, outperform equally sized baselines trained from scratch on reasoning tasks, even when controlling for final architecture and training dataset composition. The Connection: A Unified View? Depth growing creates implicit repetition through initialization (duplicating layers), while looping enforces explicit repetition through weight tying. This raises a fundamental question: Do the observed reasoning gains in depth-grown and looped models arise from the same underlying mechanism? Beyond the superficial resemblance implied by parameter- level block similarity of grown models [20, 32], the relationship between the grown structure and looped execution remains underexplored. Contributions. In this work, we provide a mechanistic unification of looping and depth growing, and show how to compose them for scalable reasoning. ⢠Empirical trade-offs. We benchmark standard, looped, and depth-grown transformers at 360M and 1.7B across 22 tasks, and characterize the resulting trade-offs between unique parameters, inference FLOPs, and training FLOPs for reasoning performance. We find that looped models improve reasoning under fixed parameter budgets and can be competitive under fixed inference budgets, while depth-grown models reach similar or better reasoning with less training compute. ⢠Unified mechanistic signatures. Using depth-utilization diagnostics and residual-stream in- terventions, we show that looped and depth-grown models exhibit similar depth-wise compu- tational signatures. Both approaches shift indispensable computation to later layers and induce periodic, block-aligned patterns in residual updates and sublayer contributions, consistent with iterative computation. ⢠Intervention robustness and looping in the middle. Through layer-swapping interventions, we find fully looped models are more order-sensitive, while looping in the middle designs, empirically the best in our setting, with unique encoderâdecoder layers recover robustness sim- ilar to depth-grown models, suggesting a practical design principle for tying recurrence to the middle of the network. ⢠Adaptability. Looped and depth-grown models adapt more efficiently than standard base- lines under in-context learning and supervised fine-tuning. Furthermore, depth-grown mod- els achieve the largest gains when exposed to high-quality, math-heavy data mixtures during cooldown. ⢠Composability: Grow first, loop later. Depth-grown models can be looped at inference time by repeating a middle block, yielding up to 2Ă gains on reasoning primitives despite never being trained with weight tying. Further, retrofitting recurrence to LIDAS during cooldown, 2 Table 1: Performance comparison of standard transformer baselines, looped models, and two depth- grown models MIDAS [32] and LIDAS [20] at 360M and 1.7B base model sizes. Looped models often outper- form iso-param baselines and are competitive with iso-inference baselines, especially for reasoning-heavy task categories such as Open-book Q&A, Math Word Problems and Reasoning Primitives. Depth-grown models match the baselines across most task categories with roughly 80% of the pre-training compute, while outper- forming them on reasoning. This suggests a shared inductive bias toward reasoning for looped and depth-grown models. Best performance per model size in bold and looped model rows in gray. Standard cooldown ParamsOpen-book Closed-bookMath Word Reasoning /Holdout SetQ&AQ&ALambada HellaSwagProblemsPrimitives FLOPs(NLLâ)(F1â)(F1â)(Accâ)(Accâ)(Accâ)(Accâ) 360M Baseline32 / 322.1822.8914.5043.3539.973.6930.04 MIDAS32 / 322.1824.5713.7543.2640.364.3928.42 LIDAS32 / 322.1626.6314.5744.0340.584.3631.58 Standard4 / 42.677.886.5426.2029.682.0214.00 Loop (4Ă8)4 / 322.5014.798.9731.6932.422.3717.00 Standard8 / 82.4713.328.6532.3732.272.3719.78 Loop (8Ă4)8 / 322.3816.2011.1234.8934.811.8120.02 Standard16 / 162.3117.3811.2437.9435.722.9021.62 Loop (16Ă2)16 / 322.2720.7411.4838.0237.252.4531.26 1.7B Baseline24 / 241.9629.5718.6050.0546.2813.6134.62 MIDAS24 / 241.9728.8018.4350.8146.1915.9140.88 LIDAS24 / 241.9629.8419.0751.4146.3218.1848.02 Standard4 / 42.2914.3110.4134.0235.782.1916.58 Loop (4Ă6)4 / 242.1223.6815.1042.4041.353.3533.84 Standard12 / 122.0125.3317.3046.3643.695.4235.42 Loop (12Ă2)12 / 242.0725.5615.3144.7142.443.3241.22 Loop (4 -4Ă4 -4)12 / 242.0527.3217.0047.5143.526.9139.88 combined with a high-quality math-focused cooldown mixture, produces the strongest reason- ing performance under matched data and inference FLOPs. Overall, our results suggest that growing and looping, despite their different origins, are comple- mentary methods for inducing and scaling iterative computation for reasoning. 2 Related Work Growing Neural Networks. Growing reuses a smaller model to initialize a deeper one, reducing pre-training compute. Common approaches include copying existing layers to initialize newly added depth [12, 16, 22, 31], learning a mapping [38], or masked structural growth [43]. Saunshi et al. [32] introduce MIDAS, which partitions the network into blocks and duplicates the middle block rather than the end, preserving efficiency while improving reasoning at comparable perplexity. Kapl et al. [20] analyze depth-grown models, including MIDAS, with depth diagnostics and block-level interventions. They argue that growth changes depth-wise computation and makes later layers more indispensable. Additionally, they propose LIDAS, which duplicates the exact layer-wise middle, yielding a more symmetric weight structure and stronger reasoning performance. Looped and recurrent models. Depth-wise parameter sharing and recurrence have been used to trade unique parameters for computation, from Universal Transformers [11] and recurrent seq2seq [10] to ALBERT [25] and broader sharing analyses [35]. Recent variants revisit recurrence with partial sharing and added capacity, such as using low-rank adapters [2] or Mixture-of-Experts [8]. Looped models are broadly used for their iterative computation and test-time scaling. They im- prove in-context learning-to-learn [42] and algorithmic length generalization [13] and enable latent reasoning [33, 45]. Additional works investigate looped pre-training at scale [15, 46], and moti- vate converting fixed-depth pre-trained models into recurrent ones [23, 27]. Related non-classical transformer settings include hierarchical multi-timescale computation [37] and compact recursive refinement architectures [18]. 3 3 The Inductive Bias of Looped and Depth-Grown Models Training Transformer-based LLMs by depth growing [20, 32] and looping [15, 23, 27, 33, 46] has been argued to improve reasoning performance. Depth-grown models are trained by progressively increasing depth via middle-layer duplication, yielding strong reasoning performance at reduced training compute. Looped models, in contrast, explicitly tie weights across depth and iteratively apply a small set of layers, trading unique parameters for additional sequential computation. Despite this connection, their practical trade-offs and the extent to which they share a common inductive bias toward reasoning have only been alluded to [20, 33] and remain unclear. The following sections introduce notation for looped and depth-grown models (§3.1) and compare their performance under parameter, inference, and training budgets (§3.2). 3.1 Notation of Looped and Depth-Grown Models We fix a Transformer architecture class (width, heads, embedding size, tokenizer, etc.) and vary only depth and parameter sharing. Let f L = [â 0 ,...,â Lâ1 ] denote a standard L-layer Transformer with untied parameters across layers, excluding the embedding matrix and the final LM head for simplicity. Looped models. A looped model reuses a sequence of consecutive layers (the recurrent block) by repeatedly applying this fixed set of layers. We denote by Loop (LĂk) a model with L unique layers that is unrolled for k repetitions, yielding an effective depth L¡ k. For example, Loop(4Ă6) repeats a 4-layer block 6 times to reach a final depth of 24 layers. We additionally consider designs that keep untied, i.e., unique, encodingâdecoding blocks and loop only the middle recurrent block, written as Loop (e -LĂk -d). For instance, Loop (4 -4Ă4 -4) has a 4-layer unique encoding block, a 4-layer recurrent block repeated 4 times, and a 4-layer unique decoding block (12 unique layers, 24 effective depth). Depth-grown models. Depth growing starts from a shallow network and repeatedly applies a growth operator according to a fixed growing schedule that inserts new layers at mid-depth by du- plicating existing layers. Concretely, we consider MIDAS [32], which duplicates the middle block (here with block size B = 4), and LIDAS [20], which duplicates the exact layer-wise middle, which has been empirically shown to be a better growing operator. We refer to Section A for more details on growing. After the last growing operation, both methods reach the same final depth as the base- line: training models by growing in depth only changes the training procedure, not the final weight structure. Therefore, growing reduces training FLOPs versus the baseline because early stages have fewer layers. Note, that in contrast to looped models, all layers are always kept untied at training and inference time for depth-grown models. 3.2 Trade-offs of Looped and Depth-Grown Models Setup. We compare standard baselines against (i) depth-grown MIDAS and LIDAS models reaching the same final depth, (i) looped models with varying numbers of unique layers but always the same effective depth (iso-inference), and (i) their iso-param standard counterparts (same unique- layer count, but no looping). All models are based on the SmolLM-v1 [5] 360M and 1.7B model configurations trained on approximately 200B and 400B tokens, respectively; for more details, see also the training protocol from Kapl et al. [20], which we follow. We note that tuning the baseline training hyperparameters for depth-grown or looped models may improve performance, but we leave this for future work. Benchmarks. We report negative log-likelihood (NLL) on a held-out SmolLM-Corpus validation set, and follow the aggregated knowledge, language, and reasoning suite of Kapl et al. [20], Saunshi et al. [32], overall spanning 22 benchmarks: ⢠Closed-book Q&A: Knowledge benchmarks that test memorization without context (Trivi- aQA, TyDiQA-NoContext, NaturalQuestions, WebQuestions) evaluated zero-shot. ⢠Open-book Q&A: Reading comprehension benchmarks that test extracting information from given context (TyDiQA-GoldP, SQuADv2, DROP, QuAC, CoQA) evaluated zero-shot. ⢠Text Completion/Language Modeling: Lambada [29] and HellaSwag [44] evaluated zero- shot. 4 BaselineLoop (4x6)Loop (4-4x4-4)LIDAS 510 Score MATH MQuAKE 6.56 5.93 9.12 8.14 7.82 7.31 8.72 6.99 A. Depth score 05101520 Layer Index 0.2 0.4 0.6 0.8 Overlap (%) B. Top 5 Overlap (Tuned Lens) 05101520 Layer Index 0.25 0.50 0.75 1.00 Rel. Accuracy C. Early Exit Tuned Lens Figure 2: Looped and depth-grown models use later layers more.We compare Baseline, LIDAS, Loop (4Ă6) and Loop (4 -4Ă4 -4) on (A) depth score, (B) top-5 vocabulary overlap on GSM8K and (C) Tuned Lens early-exit normalized accuracy on the Variable Assignment Math reasoning primitive. All three diagnos- tics imply higher usage of later layers for the grown LIDAS and looped models. ⢠Math Word Problems: Simple math benchmarks (SVAMP [30], ASDiv [28], MAWPS [24]) evaluated five-shot. ⢠Reasoning Primitives: Reasoning benchmarks from Saunshi et al. [32] evaluated five-shot. These include copying words tasks, and inferring the final value of a variable after a sequence of several assignments. An example would be: a = 3,b = 8,c = a,d = b,c =?. For more details see Section D.1. All evaluations use the language model evaluation harness [14]. Results and trade-offs. Table 1 reports aggregated performance, and Fig. 1 summarizes the re- sulting trade-offs between (i) unique parameters, (i) inference FLOPs, and (i) training FLOPs for reasoning performance. Appendix Fig. 10 provides the same trade-off view for additional bench- mark categories, illustrating how these trade-offs differ across knowledge, language, and reasoning metrics. We make the following three observations. First, under an iso-param comparison (same number of unique parameters), looping is consistently beneficial: looped models outperform equally sized standard models across most task categories, with the largest gains on reasoning-heavy task categories (Open-book Q&A, Math Word Problems, and Reasoning Primitives). Second, under an iso-inference comparison (same effective depth), looped models typically under- perform the full baseline on knowledge and general language modeling, reflecting the previously observed relationship of knowledge capacity and number of unique parameters [45, 46]. How- ever, when retaining enough unique parameters, roughly 50% of the baseline, e.g. Loop (16Ă2) or Loop (12Ă2) for the 360M and 1.7B model, respectively, looped models exceed baseline accuracy on reasoning primitives. Across the looped variants considered here, Loop (4 -4Ă4 -4) provides the best overall performance. Third, depth-grown models improve reasoning while preserving broad capabilities, and they do so at lower training compute [20]. In particular, LIDAS yields the strongest and most consistent reason- ing gains while remaining superior or competitive on NLL, knowledge and language benchmarks (Table 1). In summary, growing shifts the reasoning Pareto frontier toward lower training FLOPs in Fig. 1 (right), while looped models are competitive at fixed inference budgets (middle) and offer a comple- mentary trade-off when unique parameters are constrained (left). 4 The relationship of Looping and Growing After empirically establishing that looped and depth-grown models share an inductive bias for better reasoning performance, we next investigate whether this co-occurs with shared mechanistic traits. In this section, we focus on the following 1.7B iso-inference variants: Baseline, the depth-grown LIDAS model (block size B = 4), and two looped models with matched effective depth but different numbers of parameters, Loop (4Ă6) and Loop (4 -4Ă4 -4). 5 BaselineLoop (4x6)Loop (4-4x4-4)LIDAS 01020 Layer Index 0 1 2 3 4 5 Norm (Ă10Âł) Residual L2 Norm 01020 Layer Index 0.0 0.2 0.4 0.6 Rel. Contribution Attention Contribution Figure 3: Looped and depth-grown models exhibit similar (sub)layer usage. LIDAS, Loop (4Ă6) and Loop (4 -4Ă4 -4) share a slower residual norm growth than the baseline and exhibit periodic attention-sublayer contributions (ratio of norms for attention sublayer output over residual) with a 4-layer cycle, matching the block size of LIDAS and the size of the recurrent block. We investigate whether looping reproduces the depth-wise computational signatures previously at- tributed to depth growth [20], increased usage of later layers, periodic (sub)layer patterns aligned with the block size, and robust computational blocks under layer-order interventions. Additional results and ablations are deferred to Section B. 4.1 Mechanistic Analysis Kapl et al. [20] recently observed that pre-training LLMs by growing in depth counteracts the Curse of Depth of standard pre-LayerNorm Transformers, where later layers contribute less to the final output distribution [9, 34]. Here, we test whether explicit recurrence via looping yields depth-wise signatures that are similar. Does looping lead to higher depth usage? To quantify whether later layers perform indispensable computation, instead of mostly small independent refinements, we use the depth-utilization diag- nostics from Kapl et al. [20]: the depth score [9], and both top-5 early-exit vocabulary overlap and early-exit accuracy via Tuned Lens [4]. Across the three depth-utilization diagnostics in Fig. 2, the two looped models (Loop (4Ă6), Loop (4 -4Ă4 -4)) closely track LIDAS and show substantially stronger late-layer reliance than the baseline. The grown and looped models have higher depth scores (Fig. 2, left), indicating more indispensable computation in later layers. They also show lower top-5 overlap with the final predic- tion in later layers than the baseline (Fig. 2, middle), suggesting that late layers continue to alter the modelâs outputs. Finally, for early-exit on Variable Assignment Math, the accuracy for the baseline plateaus earlier (Fig. 2 right), whereas all other models continue improving their predictions until later layers. Residual stream and (sub)layer usage. To connect these depth-utilization signals to internal dy- namics, we first analyze residual stream norms and sublayer contributions. Then we intervene in the residual stream where we measure local effects by directly and separately removing the con- tribution of layers from future layers without propagating the changes. Comparing the Baseline, LIDAS, Loop (4Ă6) and Loop (4 -4Ă4 -4), we observe three recurring patterns. Both LIDAS and the two looped models show substantially slower growth of the residual stream norm than the baseline (Fig. 3). Moreover, the contribution of the attention sublayer to the residual stream follows a clear 4- layer cycle, matching the block size of LIDAS and the recurrent block of the looped models. Finally, if we intervene in the residual stream by removing the output of previous layers from subsequent layers directly (local effect) without propagating the changes, local future-effect heatmaps exhibit an âaggregationâ-like layer within each 4-layer segment that depends on a broad set of previous layers (Fig. 4). These shared patterns suggest that both training procedures encourage repeated, depth-periodic com- putation, and we include further ablations in Section B.3. Are looped models as robust as depth-grown models to layer interventions? Kapl et al. [20] observed robust permutable blocks in the middle of depth-grown networks. We investigate whether 6 05101520 Effect @ layer 0 5 10 15 20 Layer skipped LIDAS 05101520 Effect @ layer 0 5 10 15 20 Loop (4-4x4-4) 0.0 0.2 0.4 0.6 0.8 1.0 Relative Change Figure 4: Local effect of skipping a layer on downstream layer contributions for future tokens. We intervene in the residual stream by removing the contribution of a layer from all subsequent layers sepa- rately (local effect) and measuring the relative change on the representations of all future tokens. LIDAS and Loop (4 -4Ă4 -4) both show a characteristic phenomenon. Every four layers, there is a layer that depends di- rectly on most of the previous inputs (vertical pattern): removing the contribution of a previous layer results in relatively large changes in the representation of future tokens for this âaggregationâ layer. BaselineLoop (4x6)Loop (4-4x4-4)LIDAS 05101520 Layer Index 0.5 1.0 Rel. Accuracy Variable Assignment Math 05101520 Layer Index 0.0 0.5 1.0 Rel. Accuracy Lambada Figure 5: Fully looped models are less robust to interventions. Swapping a single layer degrades the per- formance of the fully looped model Loop (4Ă6) considerably on Lambada and Variable Assignment Math compared to the other models. Using unique encoderâdecoder layers lets Loop (4 -4Ă4 -4) recover the robust- ness of the baseline and LIDAS. this robustness is likely due to the connection between looped and grown models, or whether it is more specific to the depth-growing mechanism. In Fig. 5, swapping even a single layer in the middle of the networks causes substantially larger per- formance degradation for the Loop (4Ă6) than for the baseline, LIDAS, or Loop (4 -4Ă4 -4). This also holds for larger interventions, e.g., swapping blocks of two consecutive layers, see Section B.1. Taken together, these intervention results suggest that, while looping and growing can yield simi- lar depth-periodic (sub)layer usage, the fully looped modelâs computation is more order-sensitive, whereas depth growth can produce computational blocks in the middle that are comparatively more tolerant to local reorderings. Consistent with this, introducing unique encoderâdecoder layers to the looped model design and looping a block in the middle of the network, here, Loop (4 -4Ă4 -4), recovers the robustness of the grown model. We therefore hypothesize that the depth-growing mech- anism studied here, i.e., duplicating layers or blocks in the middle of the network, allows the model to flexibly learn unique encoderâdecoder layers, with the middle of the network acting as a relaxed version of a fully looped or tied model. 4.2 Inference Scaling After confirming that looped and depth-grown models exhibit similar computational patterns inter- nally, a natural question is: Can we loop depth-grown models at inference time to increase their 7 1234OriginalRandom 01020 Block Start 0.2 0.4 Accuracy Baseline 01020 Block Start MIDAS 01020 Block Start LIDAS Figure 6: Grown models benefit from looping at inference time. Repeating a block of four layers in the middle of the network during inference increases the accuracy of grown models (MIDAS, LIDAS) on the Copy Real Words reasoning primitive up to 2Ă compared to the original network. In contrast, the baseline rarely benefits from additional repetitions. reasoning performance? This is challenging in general, as even looped models trained with a fixed number of recursions often do not extrapolate to more recursions and instead degrade performance [46]. Since we consider grown models with block size B = 4, we focus on looping 4-layer blocks. Per- haps surprisingly, both grown models (MIDAS, LIDAS) can benefit greatly from repeating blocks of size 4 in the middle of the network to increase performance on a reasoning primitive up to 2Ă (Fig. 6), without ever being trained to loop. This is consistent with other reasoning primitives (Sec- tion C.3), where in general the baseline does not benefit at all or a lot less than the grown models from additional inference compute. Zooming in on the number of times a block is repeated in Fig. 6, we notice that a single repetition already yields the biggest improvement, two repetitions usually achieve the highest accuracy, and looping more times does not lead to additional gains but often results in a decrease. 5 Adaptability of Looped and Grown Models In this section, we study how looped and depth-grown transformers adapt. We first consider sim- ple adaptation settings (§5.1), few-shot in-context learning and supervised fine-tuning on reasoning primitives, where both looped and grown models improve faster than the baseline. We then move to more complex pre-training settings, higher-quality math cooldown mixtures (§5.2), and retrofitted recurrence (§5.3), showing that the grown model (LIDAS) achieves the largest overall reasoning gains and makes the best use of additional inference-time repetitions. 5.1 In-Context Learning & Supervised Fine-Tuning In-Context Learning. To assess in-context learning, we evaluate each reasoning primitive with an increasing number of examples, up to the context-length limit. In Fig. 7, looped and depth- grown models benefit more from additional examples than the baseline, which often shows little to no improvement. With enough examples, Loop (4 -4Ă4 -4) sometimes even surpasses the grown model LIDAS, consistent with Geiping et al. [15], who observe that dynamically trained looped models make better use of additional in-context examples at higher recursion. In general, not all reasoning primitives benefit from more examples, see Section C.2. Supervised Fine-Tuning. To evaluate supervised fine-tuning, we use the Variable Assignment Code task, focusing on the depth-1 (d = 1) and depth-2 (d = 2) variants with one and two assignment hops, respectively (see Section D). This is a challenging setting: without fine-tuning, most mod- els remain near chance. Following 32, we fine-tune on additional examples from the depth-1 and depth-2 variants, using an equal number of samples from each. In Fig. 8, with just 64 training ex- amples, LIDAS and Loop (4 -4Ă4 -4) already rise above chance, while the baseline remains close to random even with 128 examples. As the dataset grows, all looped and grown models outperform the baseline. Notably, for the depth-2 variant, Loop (4Ă6) matches LIDAS at larger dataset sizes, while Loop (4 -4Ă4 -4) reaches even higher accuracy. 8 BaselineLoop (4x6)Loop (4-4x4-4)LIDAS 010203040 Number of Few-Shot Examples 0.25 0.50 0.75 Accuracy Variable Assignment Math 010203040 Number of Few-Shot Examples 0.0 0.2 0.4 Accuracy Copying Random Words Figure 7: Looped and depth-grown models use in-context examples better. As we increase the number of in-context examples up to the maximum context length, the grownâand especially the loopedâmodels improve, while the baseline often does not. BaselineLoop (4x6)Loop (4-4x4-4)LIDAS 64128256512 Finetuning Dataset Size 0.25 0.50 0.75 Accuracy A. Finetuning (d=1) 64128256512 Finetuning Dataset Size 0.25 0.50 0.75 Accuracy B. Finetuning (d=2) Figure 8: Looped and grown models benefit substantially more from supervised fine-tuning than the baseline. Shaded regions indicateÂą one standard deviation over three random seeds (varying the fine-tuning dataset). 5.2 High-Quality Cooldown Mixtures Moving beyond supervised fine-tuning, we investigate whether the improved adaptability of grown and looped models also holds in a more complex pre-training setting: upsampling high-quality math tokens during the final stage of training [6]. Following commonly used cooldown setups, often referred to as mid-training, and math ratios [1, 36, 46], we increase the math fraction to 20% during the last 30k steps (15% of pre-training) and ablate the source of math tokens, comparing FineMath- 4+ (FMT) [1] to Nemotron-C-Math-4+ (NMT) [26]. As shown in Table 2, NMT yields larger gains for the 360M models on Math Word Problems, Reasoning Primitives, and GSM8K [7], and we therefore use it in subsequent experiments. Applying this improved cooldown at 1.7B for the Baseline, LIDAS, and Loop (4 -4Ă4 -4), all models improve (Table 3), with LIDAS achieving the strongest gains on the same reasoning benchmarks, and especially GSM8K. 5.3 Retrofitted Recurrence We previously found that depth-grown models can benefit from additional inference compute by looping a middle block. Following Koishekenov et al. [23], we retrofit this recurrence during cooldown by training on the improved cooldown mixture while looping a selected 4-layer mid- dle block for one additional repetition. Table 3 supports two observations. Leveraging the grown modelâs natural 4-layer block structure to choose candidate middle blocks, looping the 3rd block (layers 8â11) is the best overall choice for both LIDAS and the baseline, especially on GSM8K, while looping layers 12â15 is slightly weaker overall. In contrast, looping Loop (4 -4Ă4 -4) âs recurrent block one additional time yields only modest gains compared to the corresponding improvements for the baseline and LIDAS. Appendix Fig. 17 provides the full ablation over the retrofitted block position. Using the third-block-adapted variants, Fig. 9 shows that LIDAS benefits the most from repeating the adapted block additional times at inference time. More broadly, this adaptation makes 9 Table 2: Nemotron-C-Math-4+ leads to highest reasoning gains. We ablate the effect of increasing the proportion and quality of math tokens (from 6% OpenWebMath in gray) during the cooldown on rea- soning performance. Both FineMath-4+ (FMT) and Nemotron-C-Math-4+ (NMT) increase the perfor- mance on Math Word Problems, Reasoning Primitives and GSM8K. Math Cooldown Dataset Math Word Reasoning ProblemsPrimitives GSM8K (Accâ)(Accâ)(Accâ) 360M Baseline3.6930.041.06 + FMT 20%13.4630.102.12 + NMT 20%21.1535.703.41 MIDAS4.3928.421.44 + FMT 20%14.6832.843.49 + NMT 20%23.7038.424.09 LIDAS4.3631.581.36 + FMT 20%16.5440.123.94 + NMT 20%25.6442.206.07 Table 3: LIDAS benefits most from math cooldown and retrofitted recurrence. Baseline, LIDAS, and Loop (4 -4Ă4 -4) (1.7B) before (gray) and after math cooldown with 20% Nemotron-C-Math-4+ (NMT), with and without a single additional loop of a 4-layer block (0-indexed layer range). Best values per model and metric are bolded. Math CooldownÂą Loop LoopedMath Word Reasoning BlockProblemsPrimitives GSM8K Index(Accâ)(Accâ)(Accâ) 1.7B Baseline-13.6134.622.35 + NMT 20%-32.9939.2010.61 + NMT 20%8-1134.7540.9012.59 + NMT 20%12-1534.0239.8411.75 LIDAS-18.1848.025.61 + NMT 20%-35.5452.6013.34 + NMT 20%8-1138.8454.8417.36 + NMT 20%12-1538.1456.6015.47 Loop (4 -4Ă4 -4)-6.9139.882.65 + NMT 20%-27.0041.946.52 + NMT 20%4-727.7443.627.35 No AdaptNMT F.T.NMT F.T. + 1 Block 024 Additional Blocks 0.25 0.50 0.75 Accuracy Baseline 024 Additional Blocks Loop (4-4x4-4) 024 Additional Blocks LIDAS Figure 9: LIDAS benefits the most from additional inference repetitions. Repeating a fixed 4-layer block at inference increases accuracy on the reasoning primitive Variable Assignment Basic. LIDAS and the baseline make the best use of additional blocks, up to 3 repetitions, after adapting the models to loop over that block once. The looped model does not benefit from further repetition of the recurrent block, but it also degrades less than the baseline, especially before the baseline is adapted to looping. performance changes more stable when running more repetitions than seen during training. Notably, the baseline, which previously showed noticeable degradation with three additional repetitions, can improve slightly after adaptation. LIDAS and the looped model also show less degradation without adaptation. For additional results, see Section C.3. 6 Conclusion Depth growing reduces pre-training FLOPs and produces a weight structure closely aligned with looped models, while retaining the flexibility of untied layers. Across architectures, both exhibit convergent depth-wise signatures, greater reliance on late layers and recurring residual-stream and sublayer-usage patterns around the looped/grown block, that support a shared mechanism for itera- tive refinement. These parallels suggest that the benefits of growing and looping come from similar repeated computation across depth. This connection makes the techniques adaptable and compos- able: they adapt better than baselines with more in-context examples or supervised fine-tuning data, and depth-grown models benefit most from higher-quality, math-heavy cooldown mixtures. Depth- grown models can be retrofitted to loop a middle block during cooldown, improving reasoning capabilities. In our setting, this âgrow first, loop laterâ strategy yields the strongest reasoning per- formance among the standard, looped, and depth-grown models under matched data and inference FLOPs, suggesting a simple reasoning recipe. 10 Acknowledgments and Disclosure of Funding The authors would like to thank Nino Scherrer and Tobias H Ě oppe for insightful discussions and support throughout this work. This work was partially supported by the Helmholtz Foundation Model Initiative and the Helmholtz Association.The authors gratefully acknowledge the Gauss Centre for Supercomputing e.V. (w.gauss-centre.eu) for funding this project by providing computing time through the John von Neumann Institute for Computing (NIC) on the GCS Supercomputer JUPITER â JUWELS [19] at J Ě ulich Supercomputing Centre (JSC). Furthermore, the authors appreciate the computational re- sources provided by the National High Performance Computing Centre (w.nhr.kit.edu). The research presented is supported by the TUM Georg Nemetschek Institute Artificial Intelligence for the Built World and the German Federal Ministry of Education and Research (Grant:01IS24082). References [1] L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Bl Ě azquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydl Ě Äą Ë cek, A. P. Lajar Ě Äąn, V. Srivastav, et al. Smollm2: When smol goes bigâdata-centric training of a small language model. arXiv preprint arXiv:2502.02737, 2025. [2] S. Bae, A. Fisch, H. Harutyunyan, Z. Ji, S. Kim, and T. Schuster. Relaxed recursive trans- formers: Effective parameter sharing with layer-wise loRA. In The Thirteenth International Conference on Learning Representations, 2025. [3] S. Bae, Y. Kim, R. Bayat, S. Kim, J. Ha, T. Schuster, A. Fisch, H. Harutyunyan, Z. Ji, A. Courville, and S.-Y. Yun. Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation. In The Thirty-ninth Annual Conference on Neural Informa- tion Processing Systems, 2025. [4] N. Belrose, Z. Furman, L. Smith, D. Halawi, I. Ostrovsky, L. McKinney, S. Biderman, and J. Steinhardt. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112, 2023. [5] L. Ben Allal, A. Lozhkov, G. Penedo, T. Wolf, and L. von Werra. SmolLM-corpus, July 2024. URL https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus. [6] C. Blakeney, M. Paul, B. W. Larsen, S. Owen, and J. Frankle. Does your data spark joy? performance gains from domain upsampling at the end of training. In First Conference on Language Modeling, 2024. [7] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. [8] R. Csord Ě as, K. Irie, J. Schmidhuber, C. Potts, and C. D. Manning. MOEUT: Mixture-of-experts universal transformers. Advances in Neural Information Processing Systems, 37:28589â28614, 2024. [9] R. Csord Ě as, C. D. Manning, and C. Potts. Do language models use their depth efficiently? In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [10] R. Dabre and A. Fujita. Recurrent stacking of layers for compact neural machine translation models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6292â6299, 2019. [11] M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser. Universal transformers. In International Conference on Learning Representations, 2019. [12] W. Du, T. Luo, Z. Qiu, Z. Huang, Y. Shen, R. Cheng, Y. Guo, and J. Fu. Stacking your transformers: A closer look at model growth for efficient llm pre-training. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 10491â10540. Curran Associates, Inc., 2024. doi: 10.52202/079017-0336. 11 [13] Y. Fan, Y. Du, K. Ramchandran, and K. Lee. Looped transformers for length generalization. In The Thirteenth International Conference on Learning Representations, 2025. [14] L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noacâh, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou. The language model evaluation harness, 07 2024. URL https://zenodo.org/records/ 12608602. [15] J. Geiping, S. M. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein. Scaling up test-time compute with latent reasoning: A recurrent depth approach. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [16] L. Gong, D. He, Z. Li, T. Qin, L. Wang, and T. Liu. Efficient training of BERT by progressively stacking. In International conference on machine learning, pages 2337â2346. PMLR, 2019. [17] J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556, 2022. [18] A. Jolicoeur-Martineau. Less is more: Recursive reasoning with tiny networks. arXiv preprint arXiv:2510.04871, 2025. [19] J Ě ulich Supercomputing Centre. JUWELS Cluster and Booster: Exascale Pathfinder with Mod- ular Supercomputing Architecture at Juelich Supercomputing Centre. Journal of large-scale research facilities, 7(A138), 2021. doi: 10.17815/jlsrf-7-183. [20] F. Kapl, E. Angelis, T. H Ě oppe, K. Maile, J. von Oswald, N. Scherrer, and S. Bauer. Do depth-grown models overcome the curse of depth? an in-depth analysis. arXiv preprint arXiv:2512.08819, 2025. [21] J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Rad- ford, J. Wu, and D. Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. [22] S. Kim, D. Kim, C. Park, W. Lee, W. Song, Y. Kim, H. Kim, Y. Kim, H. Lee, J. Kim, C. Ahn, S. Yang, S. Lee, H. Park, G. Gim, M. Cha, H. Lee, and S. Kim. SOLAR 10.7B: Scaling large language models with simple yet effective depth up-scaling. In Y. Yang, A. Davani, A. Sil, and A. Kumar, editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 6: Industry Track), pages 23â35, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.naacl-industry.3. [23] Y. Koishekenov, A. Lipani, and N. Cancedda. Encode, think, decode: Scaling test-time rea- soning with recursive latent thoughts. arXiv preprint arXiv:2510.07358, 2025. [24] R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi. MAWPS: A math word problem repository. In Proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1152â 1157, 2016. [25] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut. Albert: A lite bert for self-supervised learning of language representations. In International Conference on Learning Representations, 2020. [26] R. K. Mahabadi, S. Satheesh, S. Prabhumoye, M. Patwary, M. Shoeybi, and B. Catanzaro. Nemotron-c-math: A 133 billion-token-scale high quality math pretraining dataset. arXiv preprint arXiv:2508.15096, 2025. [27] S. McLeish, A. Li, J. Kirchenbauer, D. S. Kalra, B. R. Bartoldson, B. Kailkhura, A. Schwarzschild, J. Geiping, T. Goldstein, and M. Goldblum. Teaching pretrained language models to think deeper with retrofitted recurrence. arXiv preprint arXiv:2511.07384, 2025. 12 [28] S.-y. Miao, C.-C. Liang, and K.-Y. Su. A diverse corpus for evaluating and developing English math word problem solvers. In D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975â984, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/ v1/2020.acl-main.92. [29] D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fern Ě andez. The LAMBADA dataset: Word prediction requiring a broad discourse context. In K. Erk and N. A. Smith, editors, Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1525â1534, Berlin, Germany, Aug. 2016. Association for Computational Linguistics. doi: 10.18653/v1/ P16-1144. [30] A. Patel, S. Bhattamishra, and N. Goyal. Are NLP models really able to solve simple math word problems? In K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, editors, Proceedings of the 2021 Confer- ence of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080â2094, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.168. [31] S. J. Reddi, S. Miryoosefi, S. Karp, S. Krishnan, S. Kale, S. Kim, and S. Kumar. Efficient training of language models using few-shot learning. In International Conference on Machine Learning, pages 14553â14568. PMLR, 2023. [32] N. Saunshi, S. Karp, S. Krishnan, S. Miryoosefi, S. Jakkam Reddi, and S. Kumar. On the inductive bias of stacking towards improving reasoning. Advances in Neural Information Pro- cessing Systems, 37:71437â71464, 2024. [33] N. Saunshi, N. Dikkala, Z. Li, S. Kumar, and S. J. Reddi. Reasoning with latent thoughts: On the power of looped transformers. In The Thirteenth International Conference on Learning Representations, 2025. [34] W. Sun, X. Song, P. Li, L. Yin, Y. Zheng, and S. Liu. The curse of depth in large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. [35] S. Takase and S. Kiyono. Lessons on parameter sharing across layers in transformers. In Proceedings of The Fourth Workshop on Simple and Efficient Natural Language Processing (SustaiNLP), pages 78â90, 2023. [36] E. P. Walsh, L. Soldaini, D. Groeneveld, K. Lo, S. Arora, A. Bhagia, Y. Gu, S. Huang, M. Jor- dan, N. Lambert, D. Schwenk, O. Tafjord, T. Anderson, D. Atkinson, F. Brahman, C. Clark, P. Dasigi, N. Dziri, A. Ettinger, M. Guerquin, D. Heineman, H. Ivison, P. W. Koh, J. Liu, S. Ma- lik, W. Merrill, L. J. V. Miranda, J. Morrison, T. Murray, C. Nam, J. Poznanski, V. Pyatkin, A. Rangapur, M. Schmitz, S. Skjonsberg, D. Wadden, C. Wilhelm, M. Wilson, L. Zettlemoyer, A. Farhadi, N. A. Smith, and H. Hajishirzi. 2 OLMo 2 furious (COLMâs version). In Second Conference on Language Modeling, 2025. [37] G. Wang, J. Li, Y. Sun, X. Chen, C. Liu, Y. Wu, M. Lu, S. Song, and Y. A. Yadkori. Hierarchical reasoning model. arXiv preprint arXiv:2506.21734, 2025. [38] P. Wang, R. Panda, L. T. Hennigen, P. Greengard, L. Karlinsky, R. Feris, D. D. Cox, Z. Wang, and Y. Kim. Learning to grow pretrained models for efficient transformer training. In The Eleventh International Conference on Learning Representations, 2023. [39] J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, E. H. Chi, T. Hashimoto, O. Vinyals, P. Liang, J. Dean, and W. Fedus. Emergent abilities of large language models. Transactions on Machine Learning Research, 2022. ISSN 2835-8856. Survey Certification. [40] J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou. Chain-of-thought prompting elicits reasoning in large language models. In S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, editors, Advances in Neural Information Processing Systems, volume 35, pages 24824â24837. Curran Associates, Inc., 2022. 13 [41] F. Xu, Q. Hao, Z. Zong, J. Wang, Y. Zhang, J. Wang, X. Lan, J. Gong, T. Ouyang, F. Meng, et al. Towards large reasoning models: A survey of reinforced reasoning with large language models. arXiv preprint arXiv:2501.09686, 2025. [42] L. Yang, K. Lee, R. D. Nowak, and D. Papailiopoulos. Looped transformers are better at learn- ing learning algorithms. In The Twelfth International Conference on Learning Representations, 2024. [43] Y. Yao, Z. Zhang, J. Li, and Y. Wang. Masked structural growth for 2x faster language model pre-training. In The Twelfth International Conference on Learning Representations, 2024. [44] R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi. HellaSwag: Can a machine really finish your sentence?In A. Korhonen, D. Traum, and L. M ` arquez, editors, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791â 4800, Florence, Italy, July 2019. Association for Computational Linguistics. doi: 10.18653/ v1/P19-1472. [45] R. Zhu, H. Zhang, T. Shi, C. Wang, T. Zhou, and Z. Qin. The 4th dimension for scaling model size. arXiv preprint arXiv:2506.18233, 2025. [46] R.-J. Zhu, Z. Wang, K. Hua, T. Zhang, Z. Li, H. Que, B. Wei, Z. Wen, F. Yin, H. Xing, et al. Scaling latent reasoning via looped language models. arXiv preprint arXiv:2510.25741, 2025. 14 A Details of Growing We briefly summarize the training procedure used to obtain the depth-grown models MIDAS and LIDAS in this work. The implementation follows the growing methodology introduced by Saunshi et al. [32] and analyzed in detail by Kapl et al. [20]; we refer the reader to those works for a complete description. We use a fixed block size B = 4 and perform gradual depth expansion by inserting a new middle block after each stage. Depending on the variant, this corresponds to MIDAS (duplicating the middle block of the current stage) or LIDAS (duplicating the layer-wise middle). Model width and the number of attention heads remain unchanged throughout growth. At each growth step, layer parameters and their optimizer state are deep-copied so that duplicated layers start from identical weights and AdamW moments, and then diverge through continued train- ing. Token embeddings and the final output head are copied without modification. Let L final denote the final depth and k = L final /B the number of growth stages. If T is the total number of training steps, the budget for stage i is allocated using the PROP-Îą schedule of Saunshi et al. [32]: T i = i Îą P k j=1 j Îą T i = 1,...,k and we use PROP-1 (Îą = 1) in all experiments. In practice, T i is rounded to integers while pre- serving P i T i = T , and a single continuous learning-rate schedule is maintained across stages (no learning-rate reset). We set T = 170,000 so that all models reach their final depth before entering the cooldown phase. To complement the main trade-off summary in Fig. 1, which focuses on Reasoning Primitives, Fig. 10 breaks down the same parameter/inference/training trade-offs by benchmark category from Table 1. B Additional Mechanistic Analysis This section provides additional mechanistic analyses that complement the results presented in Sec- tion 4.1. The experiments here further probe how architectural repetition patterns manifest in inter- vention robustness, block reuse, and sublayer-level behavior across models. B.1 Swapping Interventions We present an additional swap experiment in Fig. 11, complementing the block-size-1 results shown in Fig. 5. Consistent with those findings, looped models again exhibit lower robustness to structural interventions. When swapping a consecutive block of two layers, the fully looped model Loop (4Ă6) shows a pro- nounced degradation in performance on both Lambada and Variable Assignment (Math) compared to the baseline, LIDAS, and the partially looped variant Loop (4 -4Ă4 -4). B.2 Repeat Block experiments This section complements Section 4.2 and Fig. 6 by providing additional plots for the repeat-block without adaptation setting across the rest reasoning primitives (Fig. 12). Consistent with the main findings, repeating a contiguous block of four layers located in the middle of the network (starting around layer 10), without any further training, leads to measurable perfor- mance gains on the reasoning primitives tasks, although for some tasks is less pronounced. B.3 Sublayer Usage Experiments In this section, we provide additional plots analyzing sublayer usage, complementing Section 4.1. 15 100M1B Total Weights 2.0 2.2 2.4 2.6 NLL Parameters 10 14 10 15 10 16 FLOPs Inference 10 20 10 21 10 22 FLOPs Training 100M1B Total Weights 10 15 20 25 30 Open-book Q&A 10 14 10 15 10 16 FLOPs 10 20 10 21 10 22 FLOPs 100M1B Total Weights 10 15 Closed-book Q&A 10 14 10 15 10 16 FLOPs 10 20 10 21 10 22 FLOPs 100M1B Total Weights 5 10 15 Math Word Problems 10 14 10 15 10 16 FLOPs 10 20 10 21 10 22 FLOPs 100M1B Total Weights 20 30 40 Reasoning Primitives 10 14 10 15 10 16 FLOPs 10 20 10 21 10 22 FLOPs StandardLoopedGrown Figure 10: Category-wise trade-offs for looped and depth-grown models. Each point corresponds to a model in Table 1 (up to 1.7B parameters). For each benchmark category, we plot the average metric versus unique parameters (left), inference FLOPs (middle), and training FLOPs (right), complementing Fig. 1. 16 First, Fig. 13, complementing Fig. 4, shows that both looped models exhibit a periodic pattern in which specific layers are highly sensitive to changes in all preceding layers. This behavior closely resembles that of LIDAS and contrasts with the Baseline, where no clear structure emerges. Next, Fig. 14, complementing Fig. 3, reports the relative contribution of each layer (Attention + MLP) with respect to its input. Around the middle of the network, periodic local maxima remain visible, primarily for LIDAS and Loop (4Ă6), recurring every four layers. C Additional Adaptability and Inference Scaling Results This section provides additional results that complement the adaptability and inference-scaling anal- yses presented in the main text. We further examine how looped and depth-grown models respond to supervised fine-tuning and to inference-time repetition across tasks, highlighting how their archi- tectural inductive biases continue to influence performance beyond the original training setup. C.1 Supervised Fine Tuning In Fig. 15, we present an extended version of the results shown in Fig. 8. Beyond the fine-tuning behavior on the Depth 0 variant of the Variable Assignment task, we additionally include two hybrid models: variants of Loop (4 -4Ă4 -4) and Loop (4Ă6) in which weight tying is removed during the fine-tuning phase. Despite having strictly more degrees of freedom during adaptation, these hybrid variants consistently underperform their original tied counterparts across all depths. This suggests that the inductive bias introduced by weight tying continues to play a beneficial role even during supervised fine-tuning. C.2 In-Context Learning In Fig. 16, we show the In Context Learning Behaviour for all reasoning primitives tasks, comple- menting Section 5.1. C.3 Inference Scaling In Figs. 18 and 19, we present additional results on retrofitted recurrence across all reasoning prim- itives and GSM8K, complementing Fig. 9 in the main text. These plots confirm the same trend observed there: introducing inference-time repetition in the middle blocks consistently improves performance for the depth-grown models. Additionally, we ablate which 4-layer block is retrofitted to loop during cooldown. Fig. 17 shows that the strongest candidates are typically in the middle of the network: among the six 4-layer blocks, the third (layers 8â11) and fourth (layers 12â15) blocks are consistently competitive, with the third often strongest on reasoning-heavy tasks. This motivates our default choice of adapting the third block in the main experiments. 0.02.55.07.510.012.515.017.520.0 Layer Index 0.0 0.2 0.4 0.6 0.8 1.0 Normalized Accuracy Swap Lambada Intervened Baseline Intervened Loop (4x6) Intervened Loop (4-4x4-4) Intervened LIDAS 0.02.55.07.510.012.515.017.520.0 Layer Index 0.4 0.6 0.8 1.0 Normalized Accuracy Swap Variable Assignment MC (MATH) Intervened Baseline Intervened Loop (4x6) Intervened Loop (4-4x4-4) Intervened LIDAS Figure 11: Swapping consecutive 2-layer subblocks 17 1234OriginalRandom 05101520 Block Start 0.2 0.4 0.6 Accuracy Baseline 05101520 Block Start MIDAS 05101520 Block Start LIDAS (a) Copying Random Letter Words 1234OriginalRandom 05101520 Block Start 0.2 0.3 0.4 0.5 Accuracy Baseline 05101520 Block Start MIDAS 05101520 Block Start LIDAS (b) Variable Assignment (Basic) 1234OriginalRandom 05101520 Block Start 0.2 0.4 0.6 0.8 Accuracy Baseline 05101520 Block Start MIDAS 05101520 Block Start LIDAS (c) Variable Assignment (Math) 1234OriginalRandom 05101520 Block Start 0.2 0.4 0.6 0.8 Accuracy Baseline 05101520 Block Start MIDAS 05101520 Block Start LIDAS (d) Variable Assignment (Code) Figure 12: Repeat-block ablations across reasoning primitive tasks without further training. D Tasks and Benchmarks Overview D.1 Reasoning Primitives We implement the Reasoning Primitives tasks according to the setup introduced by Saunshi et al. [32]. 18 05101520 Effect @ layer 0 5 10 15 20 Layer skipped Baseline 05101520 Effect @ layer 0 5 10 15 20 Layer skipped LIDAS 05101520 Effect @ layer 0 5 10 15 20 Layer skipped Loop (4-4x4-4) 05101520 Effect @ layer 0 5 10 15 20 Layer skipped Loop (4x6) 0.0 0.2 0.4 0.6 0.8 1.0 Relative Change Figure 13: Future Local Effects 05101520 Layer Index 0.0 0.2 0.4 0.6 0.8 1.0 Relative Contribution Mean Relative Contribution Layer Baseline LIDAS Loop (4x6) Loop (4-4x4-4) Figure 14: Mean Relative Layer Contribution 64128256512 Finetuning Dataset Size 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Accuracy Variable Assignment (CODE) MC Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Loop (4x6) no tying Loop (4-4x4-4) no tying Random (a) Depth 0 64128256512 Finetuning Dataset Size 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Accuracy Variable Assignment (CODE) MC Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Loop (4x6) no tying Loop (4-4x4-4) no tying Random (b) Depth 1 64128256512 Finetuning Dataset Size 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 Accuracy Variable Assignment (CODE) MC Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Loop (4x6) no tying Loop (4-4x4-4) no tying Random (c) Depth 2 Figure 15: Supervised Fine Tuning Experiments 19 010203040 Number of Few-Shot Examples 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 Accuracy Copying Random Words Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Random (a) Copying Random Letter Words 0102030405060 Number of Few-Shot Examples 0.05 0.10 0.15 0.20 0.25 0.30 Accuracy Copying Real Words Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Random (b) Copying Real Words 01020304050 Number of Few-Shot Examples 0.20 0.25 0.30 0.35 0.40 0.45 0.50 0.55 Accuracy Variable Assignment Basic Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Random (c) Variable Assignment (Basic) 0510152025303540 Number of Few-Shot Examples 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Accuracy Variable Assignment Math Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Random (d) Variable Assignment (Math) 0510152025 Number of Few-Shot Examples 0.2 0.3 0.4 0.5 0.6 0.7 Accuracy Variable Assignment Code Baseline Loop (4x6) Loop (4-4x4-4) LIDAS Random (e) Variable Assignment (Code) Figure 16: In-context learning behavior across all reasoning primitive tasks. The copying task is constructed by first sampling a sequence of random three-letter tokens (e.g., a sequence of length 10). A contiguous segment of this sequence (e.g., length 5) is then repeated at the end, and the model is asked to predict the next token in the original order. This formulation isolates the ability to track sequence structure and positional dependencies without relying on semantic cues. An illustrative example of the Copying random words task is: Prompt: Fill in the blank: 20 123456 Repeated Block 1.96 1.98 2.00 2.02 2.04 NLL 123456 Repeated Block NLL 123456 Repeated Block 30 31 32 Open-book Q&A 123456 Repeated Block Open-book Q&A 123456 Repeated Block 17 18 19 Closed-book Q&A 123456 Repeated Block Closed-book Q&A 123456 Repeated Block 20 30 40 Math Word Problems 123456 Repeated Block Math Word Problems 123456 Repeated Block 35 40 45 50 55 Reasoning Primitives 123456 Repeated Block Reasoning Primitives 123456 Repeated Block 5 10 15 GSM8K 123456 Repeated Block GSM8K BaselineLIDAS Figure 17: Retrofitted-recurrence block position ablation. We sweep which 4-layer block is looped once during cooldown adaptation for the baseline and LIDAS. Dashed curves show performance before adaptation, and solid curves show performance after adaptation of the respective block. Middle blocks, in particular the third (layers 8â11), tend to be the strongest candidates, especially for reasoning. 21 01234 Additional Blocks 0.10 0.15 0.20 0.25 0.30 0.35 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (a) Copying Random Letter Words 01234 Additional Blocks 0.1 0.2 0.3 0.4 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (b) Copying Real Words 01234 Additional Blocks 0.000 0.025 0.050 0.075 0.100 0.125 0.150 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (c) GSM8K Figure 18: Repeating blocks without further training (NMT finetuning setting) on copying tasks and GSM8K. jic dqy sof uzg ewr oxw osp tkj rvw mnu jic dqy sof uzg ewr ___. -> Answer: oxw The variable assignment task is constructed by sampling a collection of variableâvalue bindings and asking for the value of a queried variable after the assignments have been processed. We use the same prompt templates for the basic, math, and code variants of this task. An important notion in this task family is the depth, which controls how many intermediate sub- stitutions are required before the queried variable can be resolved. At depth d = 0, the queried variable appears directly among the assignments. At higher depths, the queried variable must be resolved through a chain of dependencies, requiring the model to iteratively propagate values across multiple steps. An example of the Variable assignment task at depth 0 (Basic) is: 22 01234 Additional Blocks 0.2 0.3 0.4 0.5 0.6 0.7 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (a) Variable Assignment (Basic) 01234 Additional Blocks 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (b) Variable Assignment (Math) 01234 Additional Blocks 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Accuracy Baseline 01234 Additional Blocks Loop (4-4x4-4) 01234 Additional Blocks LIDAS No AdaptationNMT FinetuningNMT Finetuning + 1 Block (c) Variable Assignment (Code) Figure 19: Repeating blocks without further training (NMT finetuning setting) on Variable Assignment tasks. Prompt: Fill in blank: o=23 k=3 t=13 a=1 e=9 o=___. -> Answer: 23 An example at depth 1 is: 23 Prompt: Fill in blank: o=2 k=23 t=13 a=1 e=9 v=k c=e s=o y=t r=a y=___. -> Answer: 13 Here, the model must first resolve the value of t before determining the value of y. An example at depth 2 is: Prompt: Fill in blank: o=2 k=23 t=13 a=1 e=9 v=k c=e s=o y=t r=a b=r h=c f=y x=s g=v h=___. -> Answer: 9 In this case, correctly answering requires following a two-step chain of substitutions, illustrating how increasing depth demands progressively longer reasoning chains. All evaluations are conducted in a multiple-choice setting with five in-context examples. Under this protocol, random guessing corresponds to a baseline accuracy of 10 for the copying task and 20 for the variable assignment task. E Experimental Protocol for Depth and Mechanistic Analysis Codebase and methodology. Our analysis follows the intervention framework introduced by Csord Ě as et al. [9], adapted to the setting of our models. We perform single-layer interventions to quantify how information introduced at one layer affects representations in later layers. Tuned Lens experiments follow the procedure of Belrose et al. [4]. 24 Models and evaluation data. All analyses are conducted on SmolLM backbones at 360M and 1.7B parameters [5, 20]. For consistency across experiments, we use the same fixed set of GSM8K prompts and evaluation settings throughout all depth and early-exit analyses. Future local effects. To measure how information propagates forward through the network, we use the future local effects protocol. For a given source layer s and a later layer â > s, we remove the contribution of layer s from the residual stream that is fed into layer â, while keeping the rest of the forward pass unchanged. We then measure the relative change induced at layer â compared to the original, unmodified forward pass. This procedure isolates the direct influence of layer s on layer â without allowing the modification to propagate further through the network. Repeating this for all pairs (s,â) yields a matrix of relative changes that forms the basis of the heatmaps shown in the appendix. The relative change is computed as the norm of the difference in the residual update at layer â divided by the norm of the original residual update. For visualization, we aggregate these values by taking the maximum across batch examples and sequence positions, resulting in a single sourceâtarget matrix per model. Tuned Lens early-exit evaluation. For early-exit analyses, we train a small affine adapter for each layer that maps its residual output to the representation space immediately before the final normalization and unembedding layer, following Belrose et al. [4]. Logits are then obtained by applying the modelâs final normalization and unembedding. Adapters are trained on a held-out split of FineWeb-Edu. We evaluate early-exit quality by measur- ing (i) the KL divergence between the early-exit and final output distributions, and (i) the overlap of the top-5 predicted tokens with those of the final layer. Depth score. To summarize how strongly different layers influence the modelâs output, we com- pute a depth score based on the change in the output distribution when intervening at each layer. For each layer, we measure the maximum L2 change in the softmaxed logits caused by the interven- tion, average this quantity across examples, normalize the resulting vector across layers to form a distribution, and report its expected layer index as the depth score, following Csord Ě as et al. [9]. 25