Paper deep dive
T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Ziyang Cai, Xingyu Zhu, Yihe Dong, Yinghui He, Sanjeev Arora
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 2:15:27 AM
Summary
The paper introduces T2MLR (Transformer with Temporal Middle-Layer Recurrence), a novel architecture that enhances latent reasoning in Transformers by fusing cached middle-layer representations from previous tokens into earlier layers of the current token. This approach overcomes the information bottleneck of autoregressive decoding, allowing intermediate reasoning states to persist across time with minimal inference overhead. T2MLR outperforms standard Transformer baselines in pretraining perplexity and multi-hop reasoning tasks, and can be retrofitted into existing models to improve performance without full pretraining.
Entities (13)
Relation Signals (11)
T2MLR → usesmechanism → Temporal Middle-Layer Recurrence
confidence 100% · We introduce Transformers with Temporal Middle-Layer Recurrence (T2MLR)... fuses a cached middle layer representation from the previous token directly into an earlier layer
Princeton Language and Intelligence → authors → T2MLR
confidence 95% · Ziyang Cai... Princeton Language and Intelligence... Core contribution. Abstract Transformer reasoning is limited...
T2MLR → improves → MATH500 accuracy
confidence 95% · improves... MATH500 from 12.8 to 18.0
T2MLR → improves → GSM8K accuracy
confidence 95% · retrofitting... improves GSM8K accuracy from 35.8 to 39.9
T2MLR → outperforms → standard Transformers
confidence 95% · T2MLR consistently outperforms data- and parameter-matched Transformer baselines.
T2MLR → solves → S5-Retrieval
confidence 95% · Shallow T2MLR solves S5-Retrieval, a challenging synthetic benchmark...
T2MLR → improves → GSM8K
confidence 91% · retrofitting the recurrent pathway into a pretrained SmolLM2-1.7B-Instruct model... improves GSM8K accuracy from 35.8 to 39.9
T2MLR → improves → MATH500
confidence 91% · and MATH500 from 12.8 to 18.0
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We in troduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, T2MLR consistently outperforms data- and parameter-matched Transformer base lines. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Im portantly, T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require looping over all layers as in previous works, but can instead emerge more strongly from targeted middle-layer recurrence.
Tags
Links
- Source: https://arxiv.org/abs/2607.15178v2
- Canonical: https://arxiv.org/abs/2607.15178v2
Trouble viewing inline? Open PDF directly →
Full Text
100,138 characters extracted from source content.
Expand or collapse full text
T2MLR: Transformer with Temporal Middle-Layer Recurrence Ziyang Cai, Xingyu Zhu∗, Yihe Dong, Yinghui He, Sanjeev Arora Princeton Language and Intelligence, Department of Computer Science Princeton University zc5794, xingyu.zhu@princeton.edu Joint first authors; equal contribution.Core contribution. Abstract Transformer reasoning is limited by autoregressive decoding, which repeatedly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time. We introduce Transformers with Temporal Middle-Layer Recurrence (T2MLR), a transformers-based latent reasoning architecture that fuses a cached middle-layer representation from the previous token directly into an earlier layer of the current token position, enabling abstract intermediate computation to persist across decoding steps with little inference overhead. Across natural-language pretraining and multi-hop reasoning finetuning, T2MLR consistently outperforms data- and parameter-matched Transformer baselines. Moreover, applying recurrence to only a localized middle-layer block (as little as 20% of the network) often outperforms full-layer recurrence. Importantly, T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into an existing pretrained 1.7B Transformer and briefly finetuning substantially improves math reasoning, lowering the barrier to practical adoption. These results suggest that effective latent reasoning in Transformers does not require looping over all layers as in previous works, but can instead emerge more strongly from targeted middle-layer recurrence. 1 Introduction Recent developments in large language models have demonstrated remarkable capabilities in reasoning, solving complex problems in mathematics, physics, and other scientific domains (Wei et al., 2022; Cobbe et al., 2021; OpenAI, 2023). Many of these tasks require multi-step reasoning in which intermediate abstract reasoning states must be iteratively refined before producing a final answer. Despite these advances, the underlying Transformer architecture (Vaswani et al., 2017) remains fundamentally token-centric: auto-regressive generation repeatedly projects rich, high-dimensional latent representations back into a sparse one-hot vector in the token space at each decoding step, and this discrete representation is then used as the sole input for the next forward pass. This repeated projection creates an information bottleneck that limits how latent intermediate reasoning states can persist across time and influence the computation for future tokens. Figure 1: T2MLR fuses representation from a deep layer from the previous token position into a shallow layer of the current token position (left). It gets better pretraining perplexity (see section˜4.1) and reasoning downstream performance (see section˜4.3 ) To further improve the reasoning capability of Transformers, a growing line of work on latent reasoning seeks to relax this temporal constraint by enabling computation to persist in continuous latent space beyond explicit token-level intermediates. Such approaches either perform reasoning entirely in continuous space (Hao et al., 2024; Shen et al., 2025), or propagate uncertainty by forming linear combinations of token embeddings induced by post-softmax distributions (Zhang et al., 2025; Zhuang et al., 2025; Yue et al., 2025; Tang et al., 2026). However, these methods typically operate on representations at or beyond the final layer and feed recurrent signals into the next token through the input embedding. As a result, recurrent information is forced to live outside the middle layers, where abstract reasoning is known to occur more prominently (Tenney et al., 2019; Geva et al., 2021; Meng et al., 2022; Saunshi et al., 2024; Atanas and Liu, 2025). Concurrently, another line of work has attempted to scale reasoning capacity by extending computation along the depth dimension. These approaches loop over the transformer layers multiple times during the forward pass for a single token without increasing parameter count (Saunshi et al., 2025; Geiping et al., 2025; Zhu et al., 2025b). While successful at amplifying reasoning capabilities, such architectures incur increased inference cost which scales with the number of loops. In this paper, we introduce Transformers with Temporal Middle-Layer Recurrence (T2MLR),111Code: https://github.com/princeton-pli/T2MLR a novel Transformer-based latent-reasoning architecture addressing these limitations through middle-layer temporal recurrence. Instead of confining recurrence to the token or embedding space, or increasing inference-time depth via looping, T2MLR allows abstract intermediate representations computed in the middle layers of the network to persist and evolve across decoding steps. Concretely, we inject representations from a deeper layer at the previous token directly into an earlier layer of the current token via a gated recurrent pathway. This design enables temporally extended latent reasoning while preserving standard auto-regressive decoding, dense token-level supervision during training, and the inference-time computational profile of a standard transformer. A central practical challenge for latent-reasoning architectures is scalable teacher-forced training: recurrent latent dependencies across decoding steps often break standard sequence-parallel training, making pretraining and finetuning much less scalable. In this work, we address it with an approximate temporal-parallel training scheme (section˜2.4). To the best of our knowledge, this is the first continuous chain-of-thought variant pretrained with dense teacher forcing under scalable sequence parallelism. Empirically, we show that on data and parameter matched settings, T2MLR yields substantial gains across tasks that stress different aspects of reasoning. The main findings are summarized as follows: • Shallow T2MLR solves S5-Retrieval, a challenging synthetic benchmark requiring both non-solvable group state tracking and in-context retrieval, where standard Transformers and recurrent models fail in isolation (Section 3.2). • In pretraining settings, T2MLR achieves lower perplexity against parameter-matched Transformers and demonstrates clear improvement on NLP benchmarks (Section 4.1). The best improvement is seen when only looping over 20% of the middle layers. • Finetuning on downstream reasoning datasets such as GSM-Aug (Deng et al., 2023), ProsQA (Hao et al., 2024), HotPotQA (Yang et al., 2018), and Variable Assignment (Saunshi et al., 2024), T2MLR consistently outperforms parameter- and data-matched baselines (Sections 4.3), and notably, middle layer recurrence variants consistently outperform full-layer recurrence ones. • The gains persist as we scale to 361M and 1B parameters and 50B pretraining tokens (section˜C.1). Crucially, unlike looped and latent-recurrent baselines, T2MLR adds at most ∼ 8% per-token inference overhead (section˜B.5), trading additional training compute for a low-overhead recurrent latent pathway at inference. • T2MLR does not require pretraining from scratch: retrofitting the recurrent pathway into a pretrained SmolLM2-1.7B-Instruct model and briefly finetuning on math data improves GSM8K accuracy from 35.835.8 to 39.939.9 and MATH500 from 12.812.8 to 18.018.0 over the identically-finetuned baseline (section˜4.4), showing the architecture can be adopted at scale without a full pretraining run. 2 Transformer with Temporal Middle-Layer Recurrence In this section, we formally introduce our proposed architecture: T2MLR. We will go over the design motivation, characterize the recurrence and representation fusion module, and describe how we are training the model. 2.1 Motivation: Effect of Middle Layers and the Information Bottleneck Recent mechanistic analyses reveal that intermediate layers serve as the primary locus of abstract reasoning in Transformers, whereas early layers focus on lexical and syntactic processing and late layers specialize in projecting representations onto the output vocabulary (Tenney et al., 2019; Geva et al., 2021; Meng et al., 2022; Saunshi et al., 2024; Atanas and Liu, 2025). This suggests that the most valuable reasoning-related computation occurs mainly in the middle layers, and that representations produced after these layers encode rich information about the model’s current reasoning state. However, autoregressive inference in a standard decoder-only Transformer provides no mechanism for such intermediate representations computed for the previous token to be directly leveraged during the forward pass of the current token: shallow layers processing the current token cannot directly access deeper-layer representations computed at the previous step, even though these representations are already available in memory. To exploit this information, the model must either reconstruct it by traversing the same depth again and retrieving it indirectly via attention, or rely on the input embedding, a highly compressed representation that must pass through the unembed–decode–embed bottleneck before reaching deeper layers. 2.2 Temporal Middle-Layer Recurrence Figure 2: Illustration of the T2MLR architecture (left) and the representation fusion module Φ (right). When doing forward pass for the t-th token, we use the representation after layer ℓend _end to update the recurrent cache tR_t. During the forward pass for the t+1t+1-th token, tR_t is then merged with the representation before layer ℓstart _start via the representation fusion module. To relax the bottleneck without changing the autoregressive interface, we introduce a lightweight recurrent pathway that exposes a cached intermediate representation from the previous step as an explicit input at the current step. We now make our architectural change precise in this section. We start from a standard L-layer decoder-only Transformer with hidden dimension d, and work at the level of Transformer blocks with KV caches (Ott et al., 2019). We abstract out the key and value caches, and assume each token at each layer corresponds to a single cache vector in ℝdKVR^d_KV. Let t(0)∈ℝd h_t^(0) ^d be the input embedding of the t-th token xtx_t. During autoregressive generation, each layer ℓ∈[L] ∈[L] maintains a KV cache of past tokens up to step t−1t-1, denoted by 1:t−1(ℓ)∈ℝ(t−1)×dKV.KV_1:t-1^( ) ^(t-1)× d_KV. We view the ℓ -th Transformer block as a map that takes the current token representation together with the past cache and returns an updated representation and an updated cache: (t(ℓ),1:t(ℓ))=ℱℓ(t(ℓ−1),1:t−1(ℓ)),ℓ=1,2,…,L. ( h_t^( ),\,KV_1:t^( ) )=F_ ( h_t^( -1),\,KV_1:t-1^( ) ), =1,2,…,L. (2.1) To break the depth-time barrier, we introduce temporal middle-layer recurrence, parameterized by two layer indices 1≤ℓstart≤ℓend≤L1≤ _start≤ _end≤ L and a learnable representation fusion module Φ:ℝd×ℝd→ℝd :R^d×R^d ^d. We additionally maintain a constant-size recurrent cache t∈ℝdR_t ^d for each decoding step t. tR_t will store information for the representation after ℓend _end at the current step and be used by ℓstart _start in the next. Formally, for any step t≥1t≥ 1, we modify the ℓstart _start-th transformer block computation from equation˜2.1 into: (t(ℓstart),1:t(ℓstart))=ℱℓstart(Φ(t(ℓstart−1),t−1),1:t−1(ℓstart)). ( h_t^( _start),\,KV_1:t^( _start) )=F_ _start ( ( h_t^( _start-1),\,R_t-1 ),\,KV_1:t-1^( _start) ). (2.2) Before layer ℓstart _start, we fuse the recurrent cache t−1R_t-1 with the pre-ℓstart _start representation t(ℓstart−1) h_t^( _start-1). This creates a direct pathway, allowing representations computed via the middle layers in the previous step to be utilized early on in the current step. The computation for the other layers remains the same as in a standard decoder-only transformer. After passing through layer ℓend _end, we will compute the recurrent cache tR_t based on t(ℓend) h_t^( _end) and t−1R_t-1, preparing for the next recurrence. 2.3 Gated Fusion and Recurrent Cache Update Now we are ready to describe how the recurrent cache is fused with the current stream of information from shallower layers. Recall that t(ℓstart−1)∈ℝd h_t^( _start-1) ^d is the representation of the current token before layer ℓstart _start and t−1∈ℝdR_t-1 ^d is the recurrent cache from the previous decoding step. The gated fusion module computes Φ(t(ℓstart−1),t−1)= ( h_t^( _start-1),\,R_t-1 )= t(ℓstart−1) h_t^( _start-1) (2.3) +tanh(γcur)σ(fcur([t(ℓstart−1),t−1]))⊙t(ℓstart−1) + ( _cur )\,σ (f_cur ( [ h_t^( _start-1),R_t-1 ] ) ) h_t^( _start-1) +tanh(γrec)σ(frec([t(ℓstart−1),t−1]))⊙rect−1, + ( _rec )\,σ (f_rec ( [ h_t^( _start-1),R_t-1 ] ) ) W_recR_t-1, where fcur,frec:ℝ2d→ℝdf_cur,f_rec:R^2d ^d are learnable linear layers applied to the concatenation of the current representation and the recurrent cache, rec:ℝd→ℝd W_rec:R^d ^d is a learnable linear projection, and σ(⋅)σ(·) denotes the element-wise sigmoid function. The scalar factors tanh(γcur) ( _cur ) and tanh(γrec) ( _rec ) act as learnable input-independent gates, while the element-wise σ terms provide local, input-dependent modulation. We find this design particularly helpful for early training stability as we could initialize γ’s to zero and the gates randomly. The fused representation Φ(t(ℓstart−1),t−1) ( h_t^( _start-1),R_t-1) the follows the standard transformers forward pass from ℓstart _start to layer ℓend _end . After the computation at layer ℓend _end, the recurrent cache is updated as t=RMSNorm(t(ℓend)+t−1).R_t=RMSNorm ( h_t^( _end)+R_t-1 ). (2.4) The updated cache tR_t is used in the next decoding step through the fusion in equation˜2.2. 2.4 Approximated Training of T2MLR with Temporal Parallelism Like all previous transformer-based latent-reasoning architectures which involves recurrence of continuous representations, T2MLR cannot directly adopt the standard token-parallel training procedure used by vanilla Transformers. To retain scalability, we approximate R for all tokens in a sequence in parallel using a constant number of Jacobi fixed-point iterations similar to (Wu et al., 2025) (see illustration in figure˜3). Due to space constraints, we defer the full training characterization to algorithm˜1 in Appendix appendix˜B, and only provide a simplified description in this section. Figure 3: Approximated training scheme of T2MLR. The ground truth residual cache ∗R^* is approximated by a Jacobi iteration passing through the middle layers for dforwardd_forward times. We first run a standard forward pass assuming no recurrent cache is available, and take the resulting representation after layer ℓend _end (shifted appropriately) as an initial cache ⟨0⟩R 0 . We then fuse this cache with the representation before layer ℓstart _start, run the network forward again through the middle layers up to layer ℓend _end, and use the new layer-ℓend _end representations to get a refined cache ⟨1⟩R 1 . During training, we repeat this procedure for a fixed number of iterations dforwardd_forward, yielding a refined ⟨dforward⟩R d_forward that will be used in the final forward pass all the way up to the last layer. Note that since the above computation is fully differentiable, gradient computation can be back-propagated with recurrence depth of dforwardd_forward as well. We allow a separate hyperparameter dbackwardd_backward controlling the backward depth analogous to truncated back propagation through time (TBPTT) in classical RNN training (Williams and Zipser, 2013). Throughout the paper, we use dforward=16d_forward=16 and dbackward=4d_backward=4 for experiments unless explicitly stated otherwise. We justify the empirical choices of such parameters as balance of approximation quality and training efficiency (see more analysis in section˜B.3). We would also want to highlight that while the middle layer recurrence does increase the compute budget necessary for training, the inference cost remains nearly identical to that of the standard Transformer during auto-regressive generation. The only additional overhead comes from the constant compute per step added by the fusion module. 3 Best of Both Worlds: State Tracking and Retrieval Before testing T2MLR on general natural language modeling, we first isolate its core inductive bias on a synthetic benchmark. We use S5S_5-Retrieval, a task constructed to sharply separate T2MLR with (i) standard Transformers, which excel at in-context retrieval but struggle with state tracking at small depth, and (i) recurrent models (e.g., RNNs/LSTMs), which naturally support state tracking but bottleneck retrieval through a fixed state size. 3.1 The S5S_5-Retrieval Task The S5S_5 state-tracking task, first described by Liu et al. (2023), is a sequence to sequence modeling task where the input is an ordered random sequence of N elements (a1,a2,…,aN) (a_1,a_2,…,a_N ) drawn from S5S_5 (the permutation group on 55 elements, with a cardinality of 120120) and the output is the cumulative composition of the input sequence (a1,a1∘a2,…,Πi=1Nai)(a_1,a_1 a_2,…, _i=1^Na_i). Here the i-th element is considered as the i-th state, which transits into the i+1i+1-th state by applying ai+1a_i+1 to itself. Leveraging circuit complexity arguments, Merrill et al. (2024) showed that Transformers require Ω(logN) ( N ) layers to exactly solve the S5S_5 state-tracking, and the known shallowest learnable solution in transformers is via parallel associative scan which requires ⌈log2N⌉ _2N layers (Li et al., 2025). On the other hand, RNNs can solve the task with just constant number of layers (Merrill et al., 2024), since its circuit depth grows along the temporal dimension. However, classical recurrent architectures fail in key-value retrieval tasks as the associations must be compressed into a fixed-dimensional state. To stress test a model’s capability of composing both state tracking and in-context retrieval, we define the S5S_5-Retrieval Task as follows (see sample data in table˜15): Let ℵ be an alphabet, and let :S5→ℵKD:S_5→ ^K be a random mapping from the set S5S_5 to length-K strings from the alphabet. The input is a serialization of the dictionary (a,(a)):a∈S5\(a,D(a)):a∈ S_5\ followed by a delimiter and the action sequence (a1,a2,…,aN)(a_1,a_2,…,a_N) (appropriately padded to match token count). The target output is the step-wise state-tracking results interleaved with retrieval results based on the state: (a1,(a1),a1∘a2,(a1∘a2),…,Πi=1Nai,(Πi=1Nai)). (a_1, (a_1), a_1 a_2, (a_1 a_2), …, _i=1^Na_i, ( _i=1^Na_i ) ). (a1,a2,a3,…,aN)⇒(a1,a1∘a2,a1∘a2∘a3,…,Πi=1Nai) (a_1, a_2, a_3, …, a_N ) (a_1, a_1 a_2, a_1 a_2 a_3,…, _i=1^Na_i ) 3.2 Experiments: Empirical Separation between T2MLR vs. Transformers / RNNs We train the S5S_5-Retrieval task on (i) a small 4-layer LSTM model, (i) a 4-layer, 6-head Llama (Touvron et al., 2023) transformer model, and (i) a T2MLR model built on top of the small Llama model with ℓstart=0 _start=0 and ℓend=4 _end=4, all with matching number of parameters. We use K=4K=4 and uniformly sample the number of states N from 1,2,…,32 \1,2,…,32 \ within the training set. We randomly resample the dictionary D for each sequence to ensure the necessity of doing purely in-context retrieval. We train the models with learning rate 1×10−31× 10^-3 and batchsize of 64 (150150k steps for T2MLR and 400400k steps for the baselines), and evaluate on a held-out set of input-output pairs with N ranging from 11 to 4848. In figure˜4, we compare T2MLR against recurrent and Transformer baselines on the S5S_5-Retrieval task, reporting both sequence exact-match accuracy (top row) and average per-token accuracy (bottom row). T2MLR substantially outperforms both baselines across input lengths, maintaining high exact-match accuracy over a wide range of sequence lengths. This matches the intended inductive bias of T2MLR: the recurrent pathway supports continuous propagation of the evolving latent state, while the Transformer backbone preserves access to the input sequence for retrieval. The sequence exact-match metric (top row of figure˜4) is highly sensitive to compounding errors at long sequence lengths (each sequence contains up to 4848 states × 5×\,5 tokens per state =240=240 tokens), so it can fall to zero for the baselines even when they are partially correct. The average per-token accuracy (bottom row) gives a finer-grained picture. T2MLR attains near-perfect in-distribution exact match while sustaining non-trivial token accuracy at higher out-of-distribution lengths, whereas the LSTM and Transformer baselines retain only modest token accuracy and collapse on exact match. The blue dashed line marks the maximum number of states a standard 4-layer Transformer can track using the parallel associative-scan circuit (Li et al., 2025), and the orange dotted line marks the maximum number of states seen during training. Figure 4: Performance comparison between LSTM, Transformer, and T2MLR on the S5S_5-Retrieval task after extended training (150150k steps for T2MLR, 400400k steps for the baselines). x-axis denotes the test length; top row reports sequence exact-match accuracy and bottom row reports average per-token accuracy. The blue dashed line marks the maximum number of states a standard 4-layer Transformer can track using the parallel associative-scan circuit (Li et al., 2025); the orange dotted line marks the maximum number of states seen during training. The exact-match metric is far more sensitive to compounding errors at long lengths; T2MLR attains near-perfect in-distribution exact match while retaining non-trivial token accuracy at out-of-distribution lengths, whereas the baselines fail to learn the task. 4 Experiments In this section, we provide empirical results of pretraining / finetuning a small T2MLR model on natural language data as well as synthetic reasoning tasks. Across a diverse set of tasks we considered, T2MLR achieves notable gains over the parameter-matched transformer baseline, and middle-layer recurrence generally yields larger gains compared to full-model recurrence. Section D contains further mechanistic experiments on T2MLR, in particular future token prediction. 4.1 Pretraining T2MLR We first test T2MLR in standard autoregressive language modeling. We set the baseline as SmolLM2-135M (Allal et al., 2025), a lightweight LLaMA-like decoder-only Transformer architecture. We construct T2MLR variants on top of the same backbone, parameterized by recurrence boundaries (ℓstart,ℓend)( _start, _end). We denote the recurrence depth by D=ℓend−ℓstart+1D= _end- _start+1, and refer to each model as T2MLR (ℓstart,ℓend)( _start, _end) (optionally annotated with D). The representation fusion module introduces additional parameters. To ensure fair comparison, we adjust the baseline hidden size from 576 to 584 so that all models have approximately identical parameter counts (136.4M), with the baseline being slightly larger and thus conservative. We train all models for one epoch on the official 10B-token FineWeb-Edu subset (Penedo et al., 2024) using identical optimization hyperparameters (section˜F.3), and evaluate zero-shot downstream performance using lm-eval-harness (Gao et al., 2024). Since most related works on continuous latent reasoning do not easily generalize to pretraining settings due to the lack of sequence-parallelism support, we limit our comparison to the standard transformer baseline. Unless otherwise noted, our comparisons are matched on parameter count and inference-time compute. T2MLR incurs additional training overhead due to the Jacobi-style approximation described in section˜2.4; we report a training-compute-matched comparison in table˜12 and discuss this trade-off in section˜6. Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ Baseline 24.74 44.28 29.81 30.20 61.53 60.80 48.46 42.83 T2MLR (1,30); D=30D=30 24.49 43.48 29.98 30.00 60.72 62.60 52.25 43.36 T2MLR (5,26); D=22D=22 23.98 46.13 30.75 29.80 60.77 62.80 51.85 43.73 T2MLR (9,22); D=14D=14 24.15 45.24 30.40 29.20 61.15 66.40 51.54 44.01 T2MLR (13,18); D=6D=6 24.23 45.50 29.95 31.20 61.70 64.20 52.17 44.14 T2MLR (15,16); D=2D=2 24.40 45.83 30.11 29.20 59.63 60.20 50.83 42.89 Table 1: Zero-shot downstream evaluation of 135M T2MLR variants with different recurrence boundaries (ℓstart,ℓend)( _start, _end), pretrained on 10B FineWeb-Edu tokens. All results are obtained using lm-eval-harness (Gao et al., 2024); we report normalized accuracy (acc_n) whenever available. Abbreviations: ARC-C/E = ARC-Challenge/Easy, HS = HellaSwag, OBQA = OpenBookQA, WG = Winogrande. As shown in table˜1, T2MLR consistently matches or improves upon the parameter-matched Transformer baseline on most downstream benchmarks. Across recurrence configurations, T2MLR (13,18) (D=6D=6) achieves the highest average score while only looping over 20% of the layers. However, full recurrence (D=30D=30), as used by all the previous latent reasoning works, generally performs weaker than middle layer recurrence. In figure˜5, we show the training loss for different variants of T2MLR. Most configurations attain lower evaluation loss compared to the baseline except for ℓstart=1 _start=1 (full recurrence with D=30D=30) and ℓstart=14 _start=14 (only recurring on D=2D=2 layers). Apart from these two exceptions, we note that the evaluation loss generally improves with increasing number of recurrent layers, yielding different trend as in downstream. This suggest distinct implicit bias of middle layers on reasoning that might not be captured by perplexity metrics and resembles the observation in Saunshi et al. (2024) for middle layer stacking. Figure 5: Validation cross-entropy loss vs ℓstart _start for language modeling pretraining runs at 16k steps. We further find that these improvements are not specific to the 135M from-scratch setting: they persist and grow when scaling to 361M and 1B parameters and when extending pretraining to 50B tokens, and a fixed-width ablation confirms that middle-layer placement remains optimal at fixed recurrence depth. We defer the full scaling results and recurrence-location ablation to section˜C.1. 4.2 Comparison to Looped and Latent-Recurrent Baselines Our main controlled comparison is against parameter- and data-matched Transformers, which isolates the effect of the temporal middle-layer pathway. To further situate T2MLR among architectures that also add latent or recurrent computation without replacing the attention sequence mixer, we compare against three closely related baselines at 135M scale, all matched on parameter count and trained on the same 10B FineWeb-Edu tokens: a 2×2× pause-token model (Goyal et al., 2024), a 2×2× full-looped Transformer, and a 3×3× middle-looped Transformer (looping layers 9–22). We do not directly compare against COCONUT-style embedding-level latent recurrence (Hao et al., 2024), as those methods lack a scalable dense teacher-forcing pretraining mechanism. As shown in table˜2, T2MLR attains the best average performance among these baselines. Crucially, while T2MLR incurs extra forward iterations only during training, the looped and pause-token baselines incur additional inference cost at every decoding step: the pause-token variant scales quadratically with the number of pauses, and the looped variants multiply per-token compute by the number of loops. T2MLR therefore matches or exceeds these baselines while retaining standard autoregressive inference cost (section˜B.5). Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ Transformer Baseline 24.74 44.28 29.81 30.20 61.53 60.80 48.46 42.83 T2MLR (9,22) 24.15 45.24 30.40 29.20 61.15 66.40 51.54 44.01 T2MLR (13,18) 24.23 45.50 29.95 31.20 61.70 64.20 52.17 44.14 Pause-token ×2× 2 24.74 44.57 29.51 31.60 60.23 61.90 50.59 43.31 Full-looped ×2× 2 24.83 44.95 29.84 30.80 60.23 60.70 49.57 42.99 Middle-looped ×3× 3 23.04 45.16 30.12 30.20 60.17 59.80 50.28 42.68 Table 2: Comparison against looped and latent-recurrent baselines at 135M (10B FineWeb-Edu tokens), all matched on parameter count. T2MLR achieves the best average while, unlike the looped and pause-token baselines, adding no per-token inference overhead. 4.3 Finetuning T2MLR on Reasoning Downstream Tasks While perplexity reflects average next-token prediction, it can understate changes in the model’s internal computation. To directly test whether temporal middle-layer recurrence improves latent reasoning, we finetune the pretrained checkpoints from section˜4.1 on the following suite of multi-hop reasoning and grade-school math tasks: Variable assignment: We take the variable assignment task from the set of "reasoning-primitives" task proposed by Saunshi et al. (2024). In this task, the input is a series of variable assignments, and the model needs to output the value of a query variable. Since the baseline models already saturates the original dataset, we increase the difficulty of the task from depth=2=2 to depth=5=5. ProsQA-Hard: We evaluate on the ProsQA task from Hao et al. (2024). The input is a directed-acyclic-graph, and a query asks a binary question for the connectivity between a starting and two choices of ending nodes. The output steps through a path that connects the starting and end nodes. Since the original ProsQA data is saturated by the baseline model, we increase the average number of nodes to 60 and average length of the path to 8. HotPotQA-Simple: HotpotQA is a question answering dataset featuring natural, multi-hop questions, with annotated supporting facts (Yang et al., 2018). We use a teacher model (GPT-4o-mini (OpenAI et al., 2024)) to generate natural language chain-of-thought SFT data. Since the medium and hard level subsets are too challenging at the model scale we are testing, we present the results for the easy subset. GSM-Aug (Symbolic / Natural Language): To check whether the T2MLR improves arithmetic-heavy reasoning, we finetune the pre-trained checkpoints on GSM-8K-style grade-school math problems. In particular, we use synthetic GSM style training data from Deng et al. (2023), which contains a symbolic variant only containing the arithemetic computation involved, and a natural language version more analogous to natural CoT reasoning. We report the downstream results in figure˜6. On both synthetic and real world reasoning tasks, T2MLR with middle-layer recurrence shows better performance compared to the parameter and data-matched baseline. Meanwhile, the full layer recurrence model shows smaller improvements. This further highlights the importance of leveraging middle layer recurrence to boost latent reasoning capabilities. Figure 6: Fine-tuning performance comparison between variants of T2MLR and transformer baseline on multi-hop reasoning and grade-school math tasks. We report pass@1 accuracy for all experiments. In all five tasks, T2MLR outperforms parameter and data-matched transformer baselines. Notably, we see the middle-layer looping variants (D=22,14,6D=22,14,6) achieving higher performance than full looping (D=30D=30). 4.4 Retrofitting a Pretrained Transformer into T2MLR A key question for practical adoption is whether T2MLR must be trained from scratch, or whether the recurrent pathway can be grafted onto an existing pretrained model. This is our largest-scale experiment, and the most direct test of real-world applicability: we retrofit a pretrained instruction-tuned SmolLM2-1.7B-Instruct model (32 layers) by inserting the recurrent fusion pathway (T2MLR (5,28)) and continuing finetuning for one epoch on the OpenMathReasoning dataset (Moshkov et al., 2025), comparing against the identical model finetuned without the recurrent pathway. As shown in table˜3, retrofitting improves GSM8K from 35.7835.78 to 39.8839.88 (+11.5%+11.5\% relative) and MATH500 from 12.8012.80 to 18.0018.00 (+40.6%+40.6\% relative). In other words, simply adding the temporal middle-layer pathway to an off-the-shelf pretrained model and briefly finetuning delivers substantial reasoning gains, without any pretraining from scratch. This makes T2MLR markedly easier to adopt: existing pretrained Transformers can be upgraded with the recurrent pathway at a fraction of the cost of pretraining, while retaining the near-identical inference cost of the base model. Model GSM8K MATH500 Baseline (SmolLM2-1.7B-Instruct) 35.78 12.80 T2MLR (5,28) 39.88 (+11.5%) 18.00 (+40.6%) Table 3: Retrofitting a pretrained SmolLM2-1.7B-Instruct model into T2MLR via continued finetuning on OpenMathReasoning (Moshkov et al., 2025). Adding the recurrent fusion pathway improves both GSM8K and MATH500. 5 Related Works Latent reasoning Recent latent-reasoning methods relax the token-space information bottleneck by shifting part of the reasoning process from explicit text into continuous hidden representations. Coconut (Hao et al., 2024) and CODI (Shen et al., 2025) internalize chain-of-thought into latent states, replacing portions of textual reasoning with hidden-state computation. Related works replace discrete reasoning tokens with soft or continuous alternatives, such as mixed latent/text training and continuous token mixtures in embedding space (Su et al., 2025; Tack et al., 2025; Zhang et al., 2025; Butt et al., 2025; Zhuang et al., 2025; Tang et al., 2026). Across these approaches, the common theme is to reduce reliance on verbose discrete reasoning traces by allowing computation to proceed in continuous space. T2MLR differs the most from existing works by relaxing the bottleneck through middle-layer recurrence. This preserves intermediate reasoning states in the part of the network where abstract computation occurs. In addition, T2MLR is trainable with standard supervised fine-tuning under dense token-level supervision, and allows BPTT to optimize the recurrent pathway across tokens. The closest prior work is HRPO (Yue et al., 2025). However their method feeds latent information back through the embedding interface and is trained in an RL setting without gradient flow through past tokens. Table˜4 provides an additional comparison between T2MLR and other latent reasoning works. Looped Transformers Looped transformers reuse the same block multiple times to increase effective reasoning depth (Dehghani et al., 2019; Giannou et al., 2023). Notably, Geiping et al. (2025) and Zhu et al. (2025b) leverage KV caches from later loops in earlier loops, representing another form of shortcutting deeper representations to shallower layers. Unlike T2MLR, which adds at most ∼ 8% per-token inference overhead (section˜B.5), looped architectures incur increased inference costs that scale with the number of loops. State-space, hybrid, and recurrent-memory models A related body of work introduces recurrence along a different architectural axis: state-space and hybrid models (e.g., Mamba (Gu and Dao, 2023), RetNet (Sun et al., 2023), Griffin (De et al., 2024)) replace or interleave attention with recurrent/linear-time sequence mixers, typically motivated by long-context efficiency, while recurrent-memory variants such as Transformer-XL (Dai et al., 2019) carry memory across segments beyond the context window. In contrast, T2MLR keeps the Transformer backbone and standard attention intact and instead adds a deep-to-shallow recurrent residual pathway across decoding steps, targeting reasoning within a shorter context window rather than long-context memory. Due to the limited space, we defer the more extended discussion around recurrent memory and state-space models / linear attention works to Appendix appendix˜A. 6 Discussions, Limitations, and Future work Training-time overhead The main limitation of T2MLR is training cost. Because our batch approximate forward applies multiple Jacobi-style refinement steps through the recurrent middle block, training is slower than a parameter-matched Transformer baseline. In our pretraining setting, a moderate approximation depth is already sufficient, but this still leads to roughly a 22–4×4× wall-clock overhead depending on the size of the recurrent block ℓend−ℓstart _end- _start. We make this trade-off explicit in table˜12: under a training-wall-clock-matched budget, a longer-trained Transformer can surpass the 135M T2MLR on zero-shot NLP. Our parameter-, data-, and inference-compute-matched comparisons should therefore be read as isolating the architectural and inference-side contribution of temporal middle-layer recurrence—which adds at most ∼ 8% per-token inference overhead (section˜B.5)—rather than as a claim of a training-compute win. A key direction for future work is to reduce this training cost further by improving the temporal mixing mechanism or the approximation scheme. In particular, it may be possible to preserve the benefits of middle-layer recurrence with fewer refinement steps, or to reuse exact recurrent states more directly in settings such as on-policy RL, where those states are naturally produced during rollout generation. Scale and Transfer Our main experiments focus on relatively small models, but we find that the gains from middle-layer recurrence persist as we scale to 361M and 1B parameters and extend pretraining to 50B tokens (tables˜9 and 10), with the relative improvements on reasoning-oriented tasks growing at larger scale. We also show that T2MLR need not be trained from scratch: retrofitting a pretrained SmolLM2-1.7B-Instruct model by inserting the recurrent fusion module and continuing finetuning yields substantial gains on GSM8K and MATH500 (table˜3), making the architecture substantially easier to adopt in practice. Scaling further to larger models and more demanding reasoning benchmarks, as well as multi-seed variance estimates, remain important directions for future work. Architecture Design and Mechanistic Understanding We believe our recurrent fusion module as well as the general recurrent pathway can be further improved. Understanding which fusion parameterizations and recurrence boundaries work best may lead to stronger and more efficient variants. It is also an interesting future direction to understand the information propagated through the recurrent pathways via a mechanistic lens. 7 Conclusion We introduced T2MLR, a Transformer architecture that routes recurrence through the middle layers. Across pretraining and downstream reasoning tasks, T2MLR consistently outperforms parameter- and data-matched Transformer baselines, with localized middle-layer recurrence often outperforming full-layer recurrence. Taken together, these results suggest that the value of recurrence in Transformers depends not only on enabling iterative latent computation, but on introducing it at the right depth. We hope this work serves as a starting point for a broader investigation of where recurrence should live in latent-reasoning architectures, and whether reasoning gains can be unlocked by placing recurrent computation where abstract processing is most active. Contributions Ziyang Cai∗ designed the fusion gate, completed the majority of the experiments including data generation and training, and contributed to writing. Xingyu Zhu∗ proposed the idea of T2MLR, wrote the core training and inference script, and contributed to writing. Yihe Dong† contributed to general discussion and experiments on the architecture design, completed the future token prediction task and part of pre-training experiments, and contributed to writing. Yinghui He worked on the HotpotQA experiment. Sanjeev Arora advised the project. (∗Equal contribution; †core contribution.) Acknowledgments We acknowledge the support from NSF, Schmidt Foundation, DARPA AIQ Program, OpenAI and Google Inc. Ziyang Cai and Xingyu Zhu are additionally supported by the Gordon Y.S. Wu Fellowship in Engineering. We thank Abhishek Panigrahi for the initial discussion with Xingyu Zhu which motivated this idea. We would also like to thank Yun Cheng, Haoyu Zhao, Liam Fowl, Narutatsu Ri, and Zixuan Wang for helpful discussions during various stages of the project. References L. B. Allal, A. Lozhkov, E. Bakouch, G. M. Blázquez, G. Penedo, L. Tunstall, A. Marafioti, H. Kydlíček, A. P. Lajarín, V. Srivastav, J. Lochner, C. Fahlgren, X. Nguyen, C. Fourrier, B. Burtenshaw, H. Larcher, H. Zhao, C. Zakka, M. Morlon, C. Raffel, L. von Werra, and T. Wolf (2025) SmolLM2: when smol goes big – data-centric training of a small language model. External Links: 2502.02737, Link Cited by: §4.1. A. Atanas and K. Liu (2025) A modular dataset to demonstrate llm abstraction capability. External Links: 2503.17645, Link Cited by: §1, §2.1. A. Behrouz, P. Zhong, and V. Mirrokni (2024) Titans: learning to memorize at test time. External Links: 2501.00663, Link Cited by: Appendix A. Y. Bisk, R. Zellers, J. Gao, Y. Choi, et al. (2020) Piqa: reasoning about physical commonsense in natural language. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34, p. 7432–7439. Cited by: §F.1. A. Bulatov, Y. Kuratov, and M. S. Burtsev (2022) Recurrent memory transformer. External Links: 2207.06881, Link Cited by: Appendix A. N. Butt, A. Kwiatkowski, I. Labiad, J. Kempe, and Y. Ollivier (2025) Soft tokens, hard truths. External Links: 2509.19170, Link Cited by: Appendix A, §5. P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord (2018) Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457. Cited by: §F.1. K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman (2021) Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. Cited by: §1. Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov (2019) Transformer-xl: attentive language models beyond a fixed-length context. External Links: 1901.02860, Link Cited by: Appendix A, §5. S. De, S. L. Smith, A. Fernando, A. Botev, G. Cristian-Muraru, A. Gu, R. Haroun, L. Berrada, Y. Chen, S. Srinivasan, G. Desjardins, A. Doucet, D. Budden, Y. W. Teh, R. Pascanu, N. De Freitas, and C. Gulcehre (2024) Griffin: mixing gated linear recurrences with local attention for efficient language models. External Links: 2402.19427, Link Cited by: Appendix A, §5. M. Dehghani, S. Gouws, et al. (2019) Universal transformers. In International Conference on Learning Representations, Cited by: Appendix A, §5. Y. Deng, Y. Choi, and S. Shieber (2024) From explicit cot to implicit cot: learning to internalize cot step by step. External Links: 2405.14838, Link Cited by: Appendix A. Y. Deng, K. Prasad, R. Fernandez, P. Smolensky, V. Chaudhary, and S. Shieber (2023) Implicit chain of thought reasoning via knowledge distillation. External Links: 2311.01460, Link Cited by: Appendix A, 3rd item, §4.3. T. Fu, Y. You, Z. Chen, G. Dai, H. Yang, and Y. Wang (2025) Think-at-hard: selective latent iterations to improve reasoning language models. arXiv preprint arXiv:2511.08577. Cited by: Appendix A. L. Gao, J. Tow, B. Abbasi, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, A. Le Noac’h, H. Li, K. McDonell, N. Muennighoff, C. Ociepa, J. Phang, L. Reynolds, H. Schoelkopf, A. Skowron, L. Sutawika, E. Tang, A. Thite, B. Wang, K. Wang, and A. Zou (2024) The language model evaluation harness. Zenodo. External Links: Document, Link Cited by: §4.1, Table 1. J. Geiping, S. M. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein (2025) Scaling up test-time compute with latent reasoning: a recurrent depth approach. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: Appendix A, §1, §5. M. Geva, R. Schuster, J. Berant, and O. Levy (2021) Transformer feed-forward layers are key-value memories. EMNLP. Cited by: §1, §2.1. A. Giannou, S. Rajput, J. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos (2023) Looped transformers as programmable computers. In Proceedings of the 40th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 202, p. 11398–11442. Note: arXiv:2301.13196 External Links: Link Cited by: Appendix A, §5. S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan (2024) Think before you speak: training language models with pause tokens. External Links: 2310.02226, Link Cited by: §4.2. R. Grazzi, J. Siems, A. Zela, J. K. H. Franke, F. Hutter, and M. Pontil (2025) Unlocking state-tracking in linear rnns through negative eigenvalues. In International Conference on Learning Representations (ICLR), External Links: Link Cited by: §E.1. A. Gu and T. Dao (2023) Mamba: linear-time sequence modeling with selective state spaces. External Links: 2312.00752, Link Cited by: Appendix A, §5. A. Gu, K. Goel, and C. Ré (2021) Efficiently modeling long sequences with structured state spaces. External Links: 2111.00396, Link Cited by: Appendix A. S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian (2024) Training large language models to reason in a continuous latent space. External Links: 2412.06769, Link Cited by: Appendix A, 3rd item, §1, §4.2, §4.3, §5. S. Jelassi, D. Brandfonbrener, S. Girotti, Y. Li, and A. Lazaric (2024) Repeat after me: transformers are better than state space models at copying. In Proceedings of the 41st International Conference on Machine Learning (ICML), External Links: Link Cited by: §E.1. D. Kong, M. Zhao, D. Xu, B. Pang, S. Wang, E. Honig, Z. Si, C. Li, J. Xie, S. Xie, and Y. N. Wu (2025) Latent thought models with variational bayes inference-time computation. External Links: 2502.01567, Link Cited by: Appendix A. Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut (2020) ALBERT: a lite BERT for self-supervised learning of language representations. In International Conference on Learning Representations (ICLR), Note: arXiv:1909.11942 External Links: Link Cited by: Appendix A. B. Z. Li, Z. C. Guo, and J. Andreas (2025) (How) do language models track state?. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: Figure 4, §3.1, §3.2. B. Liu, J. T. Ash, S. Goel, A. Krishnamurthy, and C. Zhang (2023) Transformers learn shortcuts to automata. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §3.1. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: Link Cited by: §F.3. K. Meng, D. Bau, A. Andonian, and Y. Belinkov (2022) Locating and editing factual associations in gpt. NeurIPS. Cited by: §1, §2.1. W. Merrill, J. Petty, and A. Sabharwal (2024) The illusion of state in state-space models. In Forty-first International Conference on Machine Learning, External Links: Link Cited by: §E.1, §3.1. T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal (2018) Can a suit of armor conduct electricity? a new dataset for open book question answering. arXiv preprint arXiv:1809.02789. Cited by: §F.1. I. Moshkov, D. Hanley, I. Sorokin, S. Toshniwal, C. Henkel, B. Schifferer, W. Du, and I. Gitman (2025) AIMO-2 winning solution: building state-of-the-art mathematical reasoning models with openmathreasoning dataset. External Links: 2504.16891, Link Cited by: §4.4, Table 3. P. Nichani, A. Vitushinsky, S. Mei, H. Gonen, and Y. Li (2025) Understanding factual recall in transformers via associative memories. In International Conference on Learning Representations (ICLR), Note: Spotlight paper External Links: Link Cited by: §E.1. OpenAI, :, A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, A. Mądry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A. Kamali, A. Jabri, A. Moyer, A. Tam, A. Crookes, A. Tootoochian, A. Tootoonchian, A. Kumar, A. Vallone, A. Karpathy, A. Braunstein, A. Cann, A. Codispoti, A. Galu, A. Kondrich, A. Tulloch, A. Mishchenko, A. Baek, A. Jiang, A. Pelisse, A. Woodford, A. Gosalia, A. Dhar, A. Pantuliano, A. Nayak, A. Oliver, B. Zoph, B. Ghorbani, B. Leimberger, B. Rossen, B. Sokolowsky, B. Wang, B. Zweig, B. Hoover, B. Samic, B. McGrew, B. Spero, B. Giertler, B. Cheng, B. Lightcap, B. Walkin, B. Quinn, B. Guarraci, B. Hsu, B. Kellogg, B. Eastman, C. Lugaresi, C. Wainwright, C. Bassin, C. Hudson, C. Chu, C. Nelson, C. Li, C. J. Shern, C. Conger, C. Barette, C. Voss, C. Ding, C. Lu, C. Zhang, C. Beaumont, C. Hallacy, C. Koch, C. Gibson, C. Kim, C. Choi, C. McLeavey, C. Hesse, C. Fischer, C. Winter, C. Czarnecki, C. Jarvis, C. Wei, C. Koumouzelis, D. Sherburn, D. Kappler, D. Levin, D. Levy, D. Carr, D. Farhi, D. Mely, D. Robinson, D. Sasaki, D. Jin, D. Valladares, D. Tsipras, D. Li, D. P. Nguyen, D. Findlay, E. Oiwoh, E. Wong, E. Asdar, E. Proehl, E. Yang, E. Antonow, E. Kramer, E. Peterson, E. Sigler, E. Wallace, E. Brevdo, E. Mays, F. Khorasani, F. P. Such, F. Raso, F. Zhang, F. von Lohmann, F. Sulit, G. Goh, G. Oden, G. Salmon, G. Starace, G. Brockman, H. Salman, H. Bao, H. Hu, H. Wong, H. Wang, H. Schmidt, H. Whitney, H. Jun, H. Kirchner, H. P. de Oliveira Pinto, H. Ren, H. Chang, H. W. Chung, I. Kivlichan, I. O’Connell, I. O’Connell, I. Osband, I. Silber, I. Sohl, I. Okuyucu, I. Lan, I. Kostrikov, I. Sutskever, I. Kanitscheider, I. Gulrajani, J. Coxon, J. Menick, J. Pachocki, J. Aung, J. Betker, J. Crooks, J. Lennon, J. Kiros, J. Leike, J. Park, J. Kwon, J. Phang, J. Teplitz, J. Wei, J. Wolfe, J. Chen, J. Harris, J. Varavva, J. G. Lee, J. Shieh, J. Lin, J. Yu, J. Weng, J. Tang, J. Yu, J. Jang, J. Q. Candela, J. Beutler, J. Landers, J. Parish, J. Heidecke, J. Schulman, J. Lachman, J. McKay, J. Uesato, J. Ward, J. W. Kim, J. Huizinga, J. Sitkin, J. Kraaijeveld, J. Gross, J. Kaplan, J. Snyder, J. Achiam, J. Jiao, J. Lee, J. Zhuang, J. Harriman, K. Fricke, K. Hayashi, K. Singhal, K. Shi, K. Karthik, K. Wood, K. Rimbach, K. Hsu, K. Nguyen, K. Gu-Lemberg, K. Button, K. Liu, K. Howe, K. Muthukumar, K. Luther, L. Ahmad, L. Kai, L. Itow, L. Workman, L. Pathak, L. Chen, L. Jing, L. Guy, L. Fedus, L. Zhou, L. Mamitsuka, L. Weng, L. McCallum, L. Held, L. Ouyang, L. Feuvrier, L. Zhang, L. Kondraciuk, L. Kaiser, L. Hewitt, L. Metz, L. Doshi, M. Aflak, M. Simens, M. Boyd, M. Thompson, M. Dukhan, M. Chen, M. Gray, M. Hudnall, M. Zhang, M. Aljubeh, M. Litwin, M. Zeng, M. Johnson, M. Shetty, M. Gupta, M. Shah, M. Yatbaz, M. J. Yang, M. Zhong, M. Glaese, M. Chen, M. Janner, M. Lampe, M. Petrov, M. Wu, M. Wang, M. Fradin, M. Pokrass, M. Castro, M. O. T. de Castro, M. Pavlov, M. Brundage, M. Wang, M. Khan, M. Murati, M. Bavarian, M. Lin, M. Yesildal, N. Soto, N. Gimelshein, N. Cone, N. Staudacher, N. Summers, N. LaFontaine, N. Chowdhury, N. Ryder, N. Stathas, N. Turley, N. Tezak, N. Felix, N. Kudige, N. Keskar, N. Deutsch, N. Bundick, N. Puckett, O. Nachum, O. Okelola, O. Boiko, O. Murk, O. Jaffe, O. Watkins, O. Godement, O. Campbell-Moore, P. Chao, P. McMillan, P. Belov, P. Su, P. Bak, P. Bakkum, P. Deng, P. Dolan, P. Hoeschele, P. Welinder, P. Tillet, P. Pronin, P. Tillet, P. Dhariwal, Q. Yuan, R. Dias, R. Lim, R. Arora, R. Troll, R. Lin, R. G. Lopes, R. Puri, R. Miyara, R. Leike, R. Gaubert, R. Zamani, R. Wang, R. Donnelly, R. Honsby, R. Smith, R. Sahai, R. Ramchandani, R. Huet, R. Carmichael, R. Zellers, R. Chen, R. Chen, R. Nigmatullin, R. Cheu, S. Jain, S. Altman, S. Schoenholz, S. Toizer, S. Miserendino, S. Agarwal, S. Culver, S. Ethersmith, S. Gray, S. Grove, S. Metzger, S. Hermani, S. Jain, S. Zhao, S. Wu, S. Jomoto, S. Wu, Shuaiqi, Xia, S. Phene, S. Papay, S. Narayanan, S. Coffey, S. Lee, S. Hall, S. Balaji, T. Broda, T. Stramer, T. Xu, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Cunninghman, T. Degry, T. Dimson, T. Raoux, T. Shadwell, T. Zheng, T. Underwood, T. Markov, T. Sherbakov, T. Rubin, T. Stasi, T. Kaftan, T. Heywood, T. Peterson, T. Walters, T. Eloundou, V. Qi, V. Moeller, V. Monaco, V. Kuo, V. Fomenko, W. Chang, W. Zheng, W. Zhou, W. Manassra, W. Sheu, W. Zaremba, Y. Patil, Y. Qian, Y. Kim, Y. Cheng, Y. Zhang, Y. He, Y. Zhang, Y. Jin, Y. Dai, and Y. Malkov (2024) GPT-4o system card. External Links: 2410.21276, Link Cited by: §4.3. OpenAI (2023) GPT-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §1. M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli (2019) Fairseq: a fast, extensible toolkit for sequence modeling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (Demonstrations), W. Ammar, A. Louis, and N. Mostafazadeh (Eds.), Minneapolis, Minnesota, p. 48–53. External Links: Link, Document Cited by: §2.2. G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf (2024) The fineweb datasets: decanting the web for the finest text data at scale. External Links: 2406.17557, Link Cited by: §4.1. M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. Ré (2023) Hyena hierarchy: towards larger convolutional language models. External Links: 2302.10866, Link Cited by: Appendix A. J. W. Rae, A. Potapenko, S. M. Jayakumar, and T. P. Lillicrap (2019) Compressive transformers for long-range sequence modelling. External Links: 1911.05507, Link Cited by: Appendix A. K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi (2021) Winogrande: an adversarial winograd schema challenge at scale. Communications of the ACM 64 (9), p. 99–106. Cited by: §F.1. N. Saunshi, N. Dikkala, Z. Li, S. Kumar, and S. J. Reddi (2025) Reasoning with latent thoughts: on the power of looped transformers. In International Conference on Learning Representations (ICLR), Note: arXiv:2502.17416 External Links: Link Cited by: Appendix A, §1. N. Saunshi, S. Karp, S. Krishnan, S. Miryoosefi, S. Jakkam Reddi, and S. Kumar (2024) On the inductive bias of stacking towards improving reasoning. Advances in Neural Information Processing Systems 37, p. 71437–71464. Cited by: §C.1, 3rd item, §1, §2.1, §4.1, §4.3. Z. Shen, H. Yan, L. Zhang, Z. Hu, Y. Du, and Y. He (2025) CODI: compressing chain-of-thought into continuous space via self-distillation. External Links: 2502.21074, Link Cited by: Appendix A, §1, §5. O. Skean, M. R. Arefin, D. Zhao, N. Patel, J. Naghiyev, Y. LeCun, and R. Shwartz-Ziv (2025) Layer by layer: uncovering hidden representations in language models. In Forty-second International Conference on Machine Learning (ICML), External Links: Link Cited by: §C.1. D. Su, H. Zhu, Y. Xu, J. Jiao, Y. Tian, and Q. Zheng (2025) Token assorted: mixing latent and text tokens for improved language model reasoning. External Links: 2502.03275, Link Cited by: Appendix A, §5. Y. Sun, L. Dong, S. Huang, S. Ma, Y. Xia, J. Xue, J. Wang, and F. Wei (2023) Retentive network: a successor to transformer for large language models. External Links: 2307.08621, Link Cited by: Appendix A, §5. J. Tack, J. Lanchantin, J. Yu, A. Cohen, I. Kulikov, J. Lan, S. Hao, Y. Tian, J. Weston, and X. Li (2025) LLM pretraining with continuous concepts. External Links: 2502.08524, Link Cited by: Appendix A, §5. Y. Tang, L. Dong, Y. Hao, Q. Dong, F. Wei, and J. Gu (2026) Multiplex thinking: reasoning via token-wise branch-and-merge. External Links: 2601.08808, Link Cited by: Appendix A, §1, §5. I. Tenney, D. Das, and E. Pavlick (2019) BERT rediscovers the classical nlp pipeline. In ACL, Cited by: §1, §2.1. H. Touvron, T. Lavril, G. Izacard, X. Martinet, M. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al. (2023) Llama: open and efficient foundation language models. arXiv preprint arXiv:2302.13971. Cited by: §3.2. A. Vaswani, N. Shazeer, N. Parmar, et al. (2017) Attention is all you need. In Advances in Neural Information Processing Systems, Cited by: §1. J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, and D. Zhou (2022) Chain-of-thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903. Cited by: §1. J. Welbl, N. F. Liu, and M. Gardner (2017) Crowdsourcing multiple choice science questions. arXiv preprint arXiv:1707.06209. Cited by: §F.1. R. J. Williams and D. Zipser (2013) Gradient-based learning algorithms for recurrent networks and their computational complexity. In Backpropagation, p. 433–486. Cited by: §B.1, §2.4. H. Wu, Z. Teng, and K. Tu (2025) Parallel continuous chain-of-thought with jacobi iteration. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 914–926. Cited by: §2.4. S. Yang, E. Gribovskaya, N. Kassner, M. Geva, and S. Riedel (2024a) Do large language models latently perform multi-hop reasoning?. External Links: 2402.16837, Link Cited by: Appendix A. S. Yang, B. Wang, Y. Zhang, Y. Shen, and Y. Kim (2024b) Parallelizing linear transformers with the delta rule over sequence length. Advances in neural information processing systems 37, p. 115491–115522. Cited by: Appendix A. Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning (2018) HotpotQA: a dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP), Cited by: 3rd item, §4.3. Z. Yue, B. Jin, H. Zeng, H. Zhuang, Z. Qin, J. Yoon, L. Shang, J. Han, and D. Wang (2025) Hybrid latent reasoning via reinforcement learning. External Links: 2505.18454, Link Cited by: Appendix A, §1, §5. R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi (2019) Hellaswag: can a machine really finish your sentence?. arXiv preprint arXiv:1905.07830. Cited by: §F.1. Z. Zhang, X. He, W. Yan, A. Shen, C. Zhao, and X. E. Wang (2025) Soft thinking: unlocking the reasoning potential of LLMs in continuous concept space. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: Appendix A, §1, §5. H. Zhu, S. Hao, Z. Hu, J. Jiao, S. Russell, and Y. Tian (2025a) Reasoning by superposition: a theoretical perspective on chain of continuous thought. arXiv preprint arXiv:2505.12514. External Links: Link Cited by: Appendix A. R. Zhu, Z. Wang, K. Hua, T. Zhang, Z. Li, H. Que, B. Wei, Z. Wen, F. Yin, H. Xing, L. Li, J. Shi, K. Ma, S. Li, T. Kergan, A. Smith, X. Qu, M. Hui, B. Wu, Q. Min, H. Huang, X. Zhou, W. Ye, J. Liu, J. Yang, Y. Shi, C. Lin, E. Zhao, T. Cai, G. Zhang, W. Huang, Y. Bengio, and J. Eshraghian (2025b) Scaling latent reasoning via looped language models. arXiv preprint. Note: arXiv:2510.25741 External Links: Link Cited by: Appendix A, §1, §5. Y. Zhuang, L. Liu, C. Singh, J. Shang, and J. Gao (2025) Mixture of inputs: text generation beyond discrete token sampling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: Appendix A, §1, §5. J. Zou, X. Yang, R. Qiu, G. Li, K. Tieu, P. Lu, K. Shen, H. Tong, Y. Choi, J. He, J. Zou, M. Wang, and L. Yang (2025) Latent collaboration in multi-agent systems. External Links: 2511.20639, Link Cited by: Appendix A. Appendix A Extended Related Works Latent Reasoning Many recent works study or exploit latent computation in transformers to improve reasoning without relying on long explicit chain-of-thought (CoT). On the modeling side, recent work has explored increasing reasoning capacity without relying on long explicit chain-of-thought traces, by shifting computation into latent or continuous representations. One line of work focuses on internalizing chain-of-thought into hidden states, via distillation or progressively removing explicit intermediate steps while preserving final-answer performance (Deng et al., 2023; 2024; Hao et al., 2024; Shen et al., 2025). A second line explicitly allocates additional inference-time computation to latent processes, such as iterative latent refinement, token-wise branch-and-merge reasoning, or latent communication in multi-agent settings (Kong et al., 2025; Fu et al., 2025; Tang et al., 2026; Zou et al., 2025). Finally, several approaches replace discrete reasoning tokens with soft or continuous alternatives, including mixed latent/text token training and continuous mixtures in embedding space, which can shorten or eliminate verbose reasoning traces (Su et al., 2025; Tack et al., 2025; Zhang et al., 2025; Butt et al., 2025; Zhuang et al., 2025). Additionally, Zhu et al. (2025a) finds latent COT offers theoretical benefits on graph tasks. Our work is most similar to Yue et al. (2025), which uses latent information by projecting the last-token latent state back to token space before the first layer. They use the resulting architecture for RL training, showing improvements on mathematical reasoning benchmarks. Crucially, in Yue et al. (2025) gradients do not flow to the past tokens, while in our work, we use supervised training and back-propagate through time. Furthermore, we emphasize middle-layer recurrence, avoiding constraints that force recurrent information to live in the embedding space. Model Recurrent Information Injection Location Fusion Module BPTT Transformer Discrete token Input embedding – N/A Coconut Last-layer hidden state Post-embedding Identity ✓ CODI Last-layer hidden state Post-embedding 2-layer MLP ✓ Multiplex-Thinking Soft token mixture Post-embedding Identity × HRPO Soft token mixture + discrete token Post-embedding (soft token) + Input embedding (discrete token) Gated fusion × T2MLR (ours) Middle-layer hidden state + Discrete token Earlier middle layer (hidden state) + Input embedding (discrete token) Gated fusion ✓ Table 4: Comparison with related architectures. Standard transformers communicate across time only through discrete tokens. Prior recurrent or latent-reasoning variants additionally propagate either last-layer hidden states or soft token mixtures, usually through input-level or post-embedding pathways. T2MLR adds a recurrent middle-layer hidden-state pathway, fused into an earlier middle layer of the next token, while preserving the ordinary discrete token pathway and supporting end-to-end supervised training with a constant-depth BPTT. Looped Transformers Looped (recurrent-depth) Transformers trade new parameters for repeated computation by applying the same Transformer block multiple times. Universal Transformers introduce this idea explicitly as recurrence over depth, optionally with adaptive halting Dehghani et al. (2019). ALBERT popularizes cross-layer parameter sharing in Transformers to reduce parameter count Lan et al. (2020). On the expressivity side, Giannou et al. show that looping expands the expressive power of Transformers Giannou et al. (2023). Saunshi et al. (2025) provide theoretical evidence that looping can approximate the benefits of much deeper Transformers, yielding large effective depth with fewer parameters. Geiping et al. (2025) demonstrate empirically that increasing test-time compute via recurrent depth improves reasoning. Most recently, ByteDance’s Ouro (Zhu et al., 2025b) looped LMs scale this design to billion-parameter models, reporting performance that rivals much larger-parameter Transformers. It is worth noting that both Geiping et al. (2025) and Zhu et al. (2025b) introduced leveraging the KV caches of the last loops when computing the forward pass for the shallower loops. This represents another form of creating a shortcut for utilizing more refined deeper representations at shallower layers. Recurrent memory for long-context modeling. A line of work uses recurrence primarily as a memory mechanism to improve long-context performance, rather than to change per-token computation. Transformer-XL reuses hidden states across segments to extend effective context (Dai et al., 2019), while Compressive Transformer retains older history via a compressed memory (Rae et al., 2019). Recurrent Memory Transformer passes information across segments using dedicated memory tokens (Bulatov et al., 2022). More recently, Titans (Behrouz et al., 2024) augments standard attention over a fixed-length current context (motivated by attention’s quadratic cost) with a neural long-term memory module that stores and retrieves information from much longer histories. State space and hybrid models. State space models offer subquadratic alternatives to attention by maintaining and updating a compact state over the sequence, with representative examples including S4 (Gu et al., 2021), Hyena (Poli et al., 2023), and Mamba (Gu and Dao, 2023). Some hybrid architectures interleave attention with recurrent/SSM-style blocks to trade off content-based retrieval and efficiency, e.g., RetNet (Sun et al., 2023), Griffin (De et al., 2024), and DeltaNet (Yang et al., 2024a; b). In contrast, we do not introduce a new recurrent/SSM layer or replace attention; instead, we add a token-to-token recurrent connection that injects a previous token’s later-layer residual stream into an earlier-layer residual stream of the current token. Appendix B Additional Information on Training T2MLR B.1 Temporal-Parallel Approximated Training Here we formally describe the fixed-point iteration algorithm used to train T2MLR, which has been informally introduced in section˜2.4. Let (0)∈ℝT×d H^(0) ^T× d denote the input embedding for a sequence of T tokens, and with slight abuse of notation let ℱℓ1:ℓ2:ℝT×d→ℝT×dF_ _1: _2:R^T× d ^T× d denote the mapping on embedding sequences implemented by layers between ℓ1 _1 and ℓ2 _2. After doing a standard forward pass (assuming no recurrent cache exists), we store (ℓstart−1):=ℱ1:ℓstart−1((0)) H^( _start-1):=F_1: _start-1( H^(0)) as the exact input into layer ℓstart _start. This also corresponds to the exact input for the sequence if recurrence has been done sequentially. During the same forward pass, we set the right-shifted version of ⟨0⟩(ℓend):=ℱ1:ℓend((0)) H^( _end)_ 0 :=F_1: _end( H^(0)) as the first iterate of the approximated recurrent cache. Formally, denote ⟨1⟩:=ShiftRight(⟨0⟩(ℓend))R_ 1 :=ShiftRight( H^( _end)_ 0 ). Here, we use subscripts in angle brackets to denote the index of iteration for the recurrent cache and the subsequent representation induced by that cache. With this approximated cache, we merge with the cached (ℓstart−1) H^( _start-1) through the fusion module Φ(⋅,⋅) (·,·). Then we pass the fused representation through the middle layers, getting ⟨1⟩(ℓend):=ℱℓstart:ℓend(Φ(⟨0⟩(ℓstart−1),⟨1⟩)). H^( _end)_ 1 :=F_ _start: _end\! ( ( H^( _start-1)_ 0 ,\,R_ 1 ) ). (B.1) This gives us the post-ℓend _end representation conditioned on the first iterate of the cache ⟨1⟩R_ 1 . We then combine the right-shifted ⟨1⟩R_ 1 with right shifted ⟨1⟩(ℓend) H^( _end)_ 1 following the temporal cache residual rule to yield the next iterate of cache ⟨2⟩R_ 2 . We continue this iterative update for dforwardd_forward time to ⟨Dforward⟩R_ D_forward , and use it as the recurrent cache for the final forward pass. To avoid too deep of a gradient backward pass, we also allow selective detachment analogous to TBPTT in standard RNN training (Williams and Zipser, 2013). We present the full temporal-parallel approximated forward algorithm in algorithm˜1, where we adopt a temporal residual connection for the cache as described in section˜2.3. With no temporal cache residual, line 8 in algorithm˜1 is then ⟨k⟩←ShiftRight(⟨k−1⟩(ℓend)).R_ k \! ( H^( _end)_ k-1 ). Input: DforwardD_forward, DbackwardD_backward, input embeddings (0)∈ℝT×d H^(0) ^T× d Output: ⟨Dforward⟩(L)∈ℝT×d H^(L)_ D_forward ^T× d (ℓstart−1)←ℱ1:ℓstart−1((0)) H^( _start-1) _1: _start-1\! ( H^(0) ) // prefill to ℓstart−1 _start-1 ⟨0⟩(ℓend)←ℱℓstart:ℓend((ℓstart−1)) H^( _end)_ 0 _ _start: _end\! ( H^( _start-1) ) // single pass to seed cache ⟨1⟩←ShiftRight(⟨0⟩(ℓend))R_ 1 \! ( H^( _end)_ 0 ) // initialize deep cache for k=2k=2 to DforwardD_forward do ⟨k−1⟩(ℓend)←ℱℓstart:ℓend(Φ(⟨0⟩(ℓstart−1),⟨k−1⟩)) H^( _end)_ k-1 _ _start: _end\! ( ( H^( _start-1)_ 0 ,\,R_ k-1 ) ) // recurrent refinement if k>Dforward−Dbackwardk>D_forward-D_backward then ⟨k−1⟩(ℓend)←StopGrad(⟨k−1⟩(ℓend)) H^( _end)_ k-1 \! ( H^( _end)_ k-1 ) // truncate gradients ⟨k⟩←Normalize(ShiftRight(⟨k−1⟩(ℓend)+⟨k−1⟩))R_ k (ShiftRight\! ( H^( _end)_ k-1 +R_ k-1 ) ) // update cache ⟨Dforward⟩(L)←ℱℓstart:L(Φ(⟨0⟩(ℓstart−1),⟨Dforward⟩)) H^(L)_ D_forward _ _start:L\! ( ( H^( _start-1)_ 0 ,\,R_ D_forward ) ) // forward w/ refined cache return ⟨Dforward⟩(L) H^(L)_ D_forward Algorithm 1 Temporal-Parallel Approximated Forward for T2MLR B.2 Computation Overhead of Temporal-Parallel Jacobi Training In table˜5, we summarize the asymptotic computation cost of T2MLR. With N as the sequence length, L as the total number of Transformer layers, d as the hidden width, and D=ℓend−ℓstart+1D= _end- _start+1 as the size of the recurrent middle block. For full-sequence training without KV caching, each layer incurs cost O(N2d+Nd2)O(N^2d+Nd^2), where the O(N2d)O(N^2d) term comes from dense self-attention and the O(Nd2)O(Nd^2) term comes from projections and MLPs. For autoregressive inference with KV caching at context length N, each layer costs O(Nd+d2)O(Nd+d^2) per generated token. The training overhead of T2MLR comes from the dforwardd_forward Jacobi-style refinement steps used to approximate the recurrent cache in parallel across token positions. Importantly, these extra passes only re-run the recurrent middle block of size D, rather than the full L-layer network. As a result, the training cost of T2MLR is O((L+dforwardD)(N2d+Nd2)),O\! ((L+d_forwardD)\,(N^2d+Nd^2) ), compared with O(L(N2d+Nd2))O\! (L\,(N^2d+Nd^2) ) for a vanilla Transformer. This should be contrasted with full-layer looping approaches, whose cost scales as O(KL(N2d+Nd2))O(KL\,(N^2d+Nd^2)), and middle-block looping approaches, whose cost scales as O((L+(K−1)D)(N2d+Nd2))O((L+(K-1)D)\,(N^2d+Nd^2)). Thus, the additional cost of T2MLR scales only with the recurrent block size D and the Jacobi depth dforwardd_forward, which is substantially cheaper than repeatedly looping over the entire stack when D≪LD L. At inference time, T2MLR does not require extra serial Jacobi refinement. Once the recurrent cache from the previous token is available, the current token is processed with a single forward pass together with one application of the fusion module. Therefore, under KV caching, the per-token inference complexity remains approximately the same as that of a standard Transformer, O(L(Nd+d2))O\! (L\,(Nd+d^2) ) up to a small constant-factor overhead from the fusion computation. This distinguishes T2MLR from looped-Transformer approaches, whose inference cost grows linearly with the number of serial loops. The same argument is particularly favorable in RL settings. During online rollout generation, the model already computes the exact recurrent latent states sequentially token by token, so for rollout tokens there is no need to perform additional forward-side Jacobi refinement. Instead, one can cache these exact latent representations during decoding and reuse them during policy optimization. Temporal-parallel approximation is only needed in settings such as prefill, teacher-forced training, or replayed trajectories where token-parallel processing is desired. It is worth noting that in RL-style training, dforwardd_forward need not be chosen to improve the approximation quality of recurrent cache themselves. It only needs to be large enough to construct the temporal computation graph needed for BPTT. Therefore one can set dforward=dbackwardd_forward=d_backward, which will significantly reduce the practical overhead relative to supervised pretraining. Model Train (per seq) Infer (per token) Vanilla O(L(N2d+Nd2))O\! (L\,(N^2d+Nd^2) ) O(L(Nd+d2))O\! (L\,(Nd+d^2) ) Full-loop O(KL(N2d+Nd2))O\! (KL\,(N^2d+Nd^2) ) O(KL(Nd+d2))O\! (KL\,(Nd+d^2) ) Mid-loop O((L+(K−1)D)(N2d+Nd2))O\! ((L+(K-1)D)\,(N^2d+Nd^2) ) O((L+(K−1)D)(Nd+d2))O\! ((L+(K-1)D)\,(Nd+d^2) ) T2MLR O((L+dforwardD)(N2d+Nd2))O\! ((L+d_forwardD)\,(N^2d+Nd^2) ) ≈O(L(Nd+d2))≈ O\! (L\,(Nd+d^2) ) Table 5: Asymptotic compute comparison. Here N is the sequence length, L the total number of layers, d the hidden width, and D=ℓend−ℓstart+1D= _end- _start+1 the size of the recurrent middle block. B.3 On the Validity of Approximation Depth Due to the existence of the ShiftRightShiftRight operation, the t-th column of R will converge after t-iterations. In this subsection, we show that empirically, we might need way fewer depth for training. We take a middle checkpoint from the FineWeb-Edu pretraining runs (with ℓstart=5,ℓend=26,dforward=16,dbackward=4 _start=5, _end=26,d_forward=16,d_backward=4), sample a random training batch with 2048 tokens per sequence, and compute the gradients under different approximation-depth settings (dforward,dbackward)(d_forward,d_backward) in algorithm˜1. We treat the gradients under the largest setting (dforward,dbackward)=(32,32)(d_forward,d_backward)=(32,32) as an anchor. For each (dforward,dbackward)(d_forward,d_backward) setting, we measure the difference between the gradient and the anchor by (i) the cosine similarity and (i) the ℓ2 _2 distance normalized by the norm of the anchor gradients (which we denote by relative ℓ2 _2 distance). As shown in figure˜7, setting dforward<8d_forward<8 and dbackward≤4d_backward≤ 4 could lead to drastic gradient differences compared to the anchor, while further increasing depth yields diminishing benefits. Figure 7: Cosine similarity and relative ℓ2 _2 distance between gradients attained under different setup of (dforward,dbackward)(d_forward,d_backward) against setting dforward=32,dbackward=32d_forward=32,d_backward=32. B.4 Jacobi Approximation vs. Exact Recurrence in Training To directly validate the temporal-parallel Jacobi approximation as a training mechanism, we compare it against exact sequential recurrent training at three scales. We train T2MLR models with ℓstart=8 _start=8 at 135M, 360M, and 1B scale on the first 1B tokens of FineWeb-Edu. For Jacobi-approximation training we use (dforward,dbackward)=(16,4)(d_forward,d_backward)=(16,4), the same configuration as our main pretraining runs. Exact training with full 20482048-token recurrent backpropagation is computationally intractable, so we use the same truncated-BPTT budget dbackward=4d_backward=4 for a controlled comparison. As shown in table˜6, Jacobi-approximation training closely matches exact recurrent training across all three scales: the Jacobi-trained models are within 0.00100.0010, 0.00420.0042, and 0.00450.0045 validation loss of the exact-recurrence-trained models at 135M, 360M, and 1B respectively (and are in fact marginally better in these runs). Moreover, evaluating a Jacobi-trained model with the approximate forward yields essentially the same perplexity as evaluating it with exact recurrent rollout, consistent with the inference-side measurement in table˜8. This indicates that the Jacobi approximation does not introduce meaningful degradation in learned model quality at this scale. Size Vanilla Transformer T2MLR: exact train, exact eval T2MLR: Jacobi train, exact eval T2MLR: Jacobi train, Jacobi eval 135M 3.9514 (52.01) 3.9083 (49.81) 3.9073 (49.76) 3.9073 (49.77) 360M 3.3803 (29.38) 3.3246 (27.79) 3.3204 (27.67) 3.3205 (27.68) 1B 3.1562 (23.48) 3.1148 (22.53) 3.1103 (22.43) 3.1102 (22.43) Table 6: Jacobi-approximation training vs. exact recurrent training (ℓstart=8 _start=8, first 1B FineWeb-Edu tokens, (dforward,dbackward)=(16,4)(d_forward,d_backward)=(16,4)). Numbers outside parentheses are validation loss; numbers in parentheses are validation perplexity. Jacobi training matches exact recurrent training across scales, and Jacobi-forward evaluation matches exact recurrent rollout. B.5 Inference Overhead of T2MLR A central claim of T2MLR is that, unlike looped or pause-token models, it adds negligible per-token inference cost. We verify this empirically. Because the recurrent states are small, all T2MLR variants add less than 0.1%0.1\% of the peak GPU memory at inference. Autoregressive generation. We measure free-form generation from the BOS token up to 512512, 10241024, and 20482048 tokens, and report wall-clock time relative to the parameter-matched Transformer baseline (table˜7). The per-token overhead stays under ∼ 8% and decreases with longer generation and larger models, as the constant gate cost is increasingly dominated by the growing attention cost. # Generated Tokens 135M T2MLR (1,30) 135M T2MLR (5,26) 135M T2MLR (9,22) 361M T2MLR (9,24) 1B T2MLR (9,24) 512 1.070×1.070× 1.076×1.076× 1.082×1.082× 1.065×1.065× 1.067×1.067× 1024 1.064×1.064× 1.068×1.068× 1.078×1.078× 1.057×1.057× 1.057×1.057× 2048 1.067×1.067× 1.076×1.076× 1.078×1.078× 1.056×1.056× 1.041×1.041× Table 7: Autoregressive generation time of T2MLR relative to the parameter-matched Transformer baseline (1.0×1.0×). Overhead is at most ∼ 8% and decreases with longer generation length and larger model size. Prefill. The main paper reports statistics under exact sequential prompt prefill. The Jacobi approximation can also be applied during prefill, trading a small approximation error for substantially lower prefill latency. table˜8 reports zero-shot performance together with the prefill speedup over exact prefill at varying Jacobi forward depths dforwardd_forward. At dforward=8d_forward=8 or 1616, approximate prefill recovers exact-prefill performance to within roughly 0.1%0.1\% while accelerating prefill by several-fold. Model dforward=4d_forward=4 dforward=8d_forward=8 dforward=16d_forward=16 dforward=32d_forward=32 exact T2MLR (5,26) 0.436 (3.6×3.6×) 0.436 (6.2×6.2×) 0.435 (11.1×11.1×) 0.436 (20.5×20.5×) 0.437 T2MLR (9,22) 0.436 (2.7×2.7×) 0.441 (4.6×4.6×) 0.439 (8.0×8.0×) 0.440 (14.5×14.5×) 0.440 T2MLR (13,18) 0.434 (1.8×1.8×) 0.440 (2.7×2.7×) 0.441 (4.3×4.3×) 0.441 (7.7×7.7×) 0.441 T2MLR (15,16) 0.432 (1.3×1.3×) 0.430 (1.7×1.7×) 0.430 (2.5×2.5×) 0.429 (4.0×4.0×) 0.429 Baseline – – – – 0.428 Table 8: Approximate (Jacobi) vs. exact prefill. Each cell reports zero-shot average performance, with the prefill speedup over exact prefill in parentheses. At dforward=8d_forward=8 or 1616, approximate prefill matches exact-prefill performance within ∼ 0.1% while being several-fold faster. Appendix C Additional Empirical Results This section reports the additional empirical results discussed in the main text: scaling T2MLR to larger models and more data, the recurrence-location ablation, and a training-compute-matched baseline. C.1 Scaling to Larger Models and More Data To test whether the gains of middle-layer recurrence persist beyond the 135M from-scratch setting reported in the main text (table˜1), we pretrain larger T2MLR variants at 361M and 1B parameters, and additionally extend pretraining up to 50B tokens. We use the gated fusion module with Jacobi depths (dforward=16,dbackward=4)(d_forward=16,d_backward=4) and recurrence starting at ℓstart=8 _start=8 (denoted gated-f16b4-L8). As summarized in table˜9, the zero-shot NLP gains remain consistent at 361M and 1B scale, and the improvements on reasoning-oriented downstream tasks (HotpotQA-Easy, GSM-Aug) become substantially larger, with 88–16%16\% relative gains over the parameter-matched baseline. Extending pretraining to 50B tokens (table˜10) yields even larger gains, especially at 361M scale. Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ ∼ 370M parameters Baseline (367.8M) 27.13 52.65 37.18 34.00 65.51 71.00 51.85 48.48 T2MLR (361.8M; gated-f16b4-L8) 28.84 52.48 38.66 33.40 66.00 73.10 52.09 49.22 ∼ 1B parameters Baseline (996.9M) 30.97 55.85 42.91 33.40 67.74 77.10 51.70 51.38 T2MLR (981.5M; gated-f16b4-L8) 32.00 55.30 45.11 35.40 68.99 75.60 54.22 52.37 Model/Config HotpotQA-Easy GSM-Aug-NL GSM-Aug-Sym Metric acc ↑ acc ↑ acc ↑ ∼ 370M parameters Baseline (367.8M) 24.43 31.08 39.35 T2MLR (361.8M; gated-f16b4-L8) 28.28 (+15.8%) 34.12 (+9.8%) 42.46 (+7.9%) ∼ 1B parameters Baseline (996.9M) 23.28 32.37 43.97 T2MLR (981.5M; gated-f16b4-L8) 26.52 (+14.0%) 36.69 (+13.3%) 44.96 (+2.3%) Table 9: Scaling T2MLR to 361M and 1B parameters (pretrained on 10B FineWeb-Edu tokens). Top: zero-shot NLP benchmarks. Bottom: reasoning downstream tasks (HotpotQA-Easy, GSM-Aug natural-language and symbolic), with relative improvement over the baseline in parentheses. Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ ∼ 135M parameters Baseline 27.82 52.53 36.82 33.60 65.67 72.40 49.64 48.35 T2MLR (9,22) 27.56 51.30 38.74 34.00 66.10 75.20 53.04 49.42 ∼ 370M parameters Baseline 31.23 59.64 46.28 34.60 70.08 76.00 52.01 52.83 T2MLR (9,24) 33.28 58.88 48.57 36.80 69.42 80.30 56.20 54.78 Model/Config GSM-Aug-NL GSM-Aug-Sym Metric acc ↑ acc ↑ ∼ 135M parameters Baseline 26.46 38.67 T2MLR (9,22) 30.48 (+15.2%) 41.55 (+7.4%) ∼ 370M parameters Baseline 40.18 42.53 T2MLR (9,24) 44.35 (+10.4%) 46.63 (+9.6%) Table 10: Extending pretraining to 50B FineWeb-Edu tokens. Top: zero-shot NLP benchmarks for 135M and 361M models. Bottom: grade-school math reasoning downstream (GSM-Aug). Training on more data yields larger gains, especially at 361M scale (average 52.83→54.7852.83→ 54.78). Recurrence location ablation. The scaling configurations above couple recurrence width (D) with recurrence location. To isolate the effect of location, we run a fixed-width ablation that places the recurrent block over early, middle, or late layers while holding D constant. As shown in table˜11, at both D=6D=6 and D=14D=14, middle-layer recurrence clearly outperforms placing the same recurrent block over the earliest or latest layers. This supports our motivation (section˜2.1) that the middle layers, which host the most abstract computation (Skean et al., 2025; Saunshi et al., 2024), are where temporal recurrence is most beneficial. Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ Baseline 24.74 44.28 29.81 30.20 61.53 60.80 48.46 42.83 6-Layer Recurrence (D=6D=6) 13–18 (Middle) 24.23 45.50 29.95 31.20 61.70 64.20 52.17 44.14 1–6 (Early) 24.66 45.08 29.83 30.00 60.07 61.00 48.54 42.74 25–30 (Late) 25.09 45.24 29.65 28.20 60.12 61.20 50.59 42.87 14-Layer Recurrence (D=14D=14) 9–22 (Middle) 24.15 45.24 30.40 29.20 61.15 66.40 51.54 44.01 1–14 (Early) 23.38 44.70 28.59 29.00 59.25 59.50 51.07 42.21 17–30 (Late) 23.98 44.02 29.97 31.40 60.28 61.90 50.43 43.14 Table 11: Fixed-width recurrence-location ablation at 135M (10B tokens). Holding the recurrence depth D constant, middle-layer recurrence outperforms placing the same recurrent block over the earliest or latest layers, at both D=6D=6 and D=14D=14. C.2 Training-Compute-Matched Baseline Our main comparisons are matched on parameter count, data, and inference compute, but not on training wall-clock: the Jacobi-style approximation makes T2MLR training slower than the baseline (section˜B.2). For a training-compute-matched comparison, we additionally train the Transformer baseline for 2.24 epochs, matching the 2.24×2.24× training wall-clock of T2MLR (13,18). As shown in table˜12, under this matched training budget the longer-trained Transformer surpasses the 135M T2MLR on zero-shot NLP. This is consistent with our positioning of T2MLR as an inference-side and architectural contribution: under fixed parameter / data / inference-compute budgets it adds a temporal latent pathway with little decoding cost (at most ∼ 8% per-token overhead) and a more pronounced boost to state-tracking and multi-hop reasoning, at the cost of additional training compute. Model/Config ARC-C ARC-E HS OBQA PIQA SciQ WG Average Metric acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc_n ↑ acc ↑ - ↑ Transformer (1 epoch) 24.74 44.28 29.81 30.20 61.53 60.80 48.46 42.83 T2MLR (13,18) 24.23 45.50 29.95 31.20 61.70 64.20 52.17 44.14 Transformer (2.24 epochs) 25.68 46.30 32.15 31.80 62.40 66.50 52.25 45.30 Table 12: Training-compute-matched comparison at 135M. The Transformer baseline trained for 2.24 epochs matches the 2.24×2.24× training wall-clock of T2MLR (13,18). Under matched training compute, longer Transformer training surpasses T2MLR on zero-shot NLP; T2MLR’s advantage is at fixed parameter / data / inference-compute budgets and on reasoning tasks. Appendix D Future Token Prediction Temporal middle-layer recurrence creates an explicit temporal shortcut: the recurrent cache computed from token t is fused into the residual stream when processing token t+1t+1, so the loss at position t+1t+1 can backpropagate through this pathway into token t’s intermediate representation. This encourages the model to maintain latent states that better anticipate upcoming tokens. To test this hypothesis, we perform a probing analysis: from each intermediate layer, we extract the representation of token t and train a two-layer MLP to predict the next token xt+1x_t+1. Table 13 reports next-token prediction loss for T2MLR and a parameter-matched Transformer baseline under two choices of ℓend _end. For this analysis, negative indices count backward from the LM head, so ℓend=−k _end=-k places the recurrence endpoint k layers below the head. We probe representations from layers ℓ−k _-k under the same convention. ℓend=−3 _end=-3 ℓend=−5 _end=-5 ℓ−1 _-1 ℓ−2 _-2 ℓ−3 _-3 ℓ−1 _-1 ℓ−2 _-2 ℓ−3 _-3 ℓ−4 _-4 ℓ−5 _-5 T2MLR 5.0915.091 5.0885.088 5.1105.110 5.0975.097 5.0965.096 5.1095.109 5.1265.126 5.1585.158 Baseline 5.1355.135 5.1315.131 5.1455.145 5.1355.135 5.1315.131 5.1455.145 5.1735.173 5.1825.182 Table 13: Each predictor is trained for 20,000 steps on Fineweb-edu 10B. Table 13 shows that intermediate-layer representations learned by T2MLR consistently yield lower next-token prediction loss than the parameter-matched baseline, providing direct evidence that T2MLR promotes representations that better anticipate future tokens. Appendix E Extended Discussions E.1 T2MLR as a generalized Transformer and recurrent model T2MLR subsumes both recurrent neural networks and the standard Transformer as special cases. In particular, when the learned self-attention pattern reduces to the identity—so that token interactions occur exclusively through recurrence—T2MLR becomes a fully recurrent model. Conversely, when the learned gating parameters γcur _cur and γrec _rec are set to zero, the recurrent pathway is disabled and T2MLR recovers the standard Transformer architecture. Consequently, T2MLR has the expressive capacity to smoothly interpolate between recurrent and attention-based computation. Since Transformers excel at in-context retrieval tasks (Jelassi et al., 2024; Nichani et al., 2025) while recurrent models are well suited for explicit state tracking (Merrill et al., 2024; Grazzi et al., 2025), T2MLR has the expressivity to naturally support both modes of computation within a unified framework. E.2 Learned gating behavior As formulated in Section 2.3, T2MLR learns scalar, data-independent gates γcur _cur and γrec _rec that control the contributions of the input and recurrent streams, respectively. To better understand how the model allocates these streams relative to the residual branch t(ℓstart−1) h_t^( _start-1) in Equation 2.3, we analyze the learned gating parameters over the course of training. Empirically, across all tasks we study, the recurrent gate γrec _rec consistently converges to positive values, while the input gate γcur _cur consistently converges to negative values. Figure 8 illustrates this behavior for pretraining on FineWeb. This pattern aligns with the structure of Equation 2.3: since the input representation t(ℓstart−1) h_t^( _start-1) is already added unconditionally via the residual connection, the role of the γ-gated terms is not to reintroduce the input signal, but rather to modulate deviations from the identity mapping. In particular, a negative γcur _cur suppresses redundant amplification of the input stream, while a positive γrec _rec selectively promotes the recurrent contribution, allowing information from previous time steps to be injected only when it provides additional predictive value beyond the current-token representation. This asymmetry suggests that T2MLR learns to treat recurrence as an additive refinement to the residual pathway, rather than as a competing source of information. In effect, the model preserves the standard Transformer computation as a default and leverages the recurrent stream to inject latent, temporally accumulated information when beneficial. Finally, note that both γcur _cur and γrec _rec are initialized to zero, so the gating module in Equation 2.3 initially reduces to the identity function. As a result, T2MLR starts training in a regime that is exactly equivalent to a standard Transformer, and only gradually departs from it as the gating parameters adapt. This design stabilizes optimization and ensures that recurrent computation is introduced smoothly, rather than being imposed a priori. Figure 8: Plot of gating parameters γcur _cur and γrec _rec with respect to training steps on the FineWeb dataset. This illustrates the learned behavior of input and recurrent gating over the course of training. Appendix F Additional Experiment Setups and Results F.1 NLP benchmarks for evaluating pre-trained models We use the following abbreviations for the NLP benchmarks: ARC-C/E: ARC-Challenge/Easy (Clark et al., 2018), HS: HellaSwag (Zellers et al., 2019), OBQA: OpenBookQA (Mihaylov et al., 2018), PIQA: PhysicalInteractionQA (Bisk et al., 2020), SciQ: ScienceExamQA (Welbl et al., 2017), WG: Winogrande (Sakaguchi et al., 2021). This set of benchmarks follows standard evaluation setup in the pretraining literature. F.2 Data Curation for HotpotQA To construct SFT data with step-by-step multi-hop reasoning, we curate a subset of HotpotQA and generate 19k chain-of-thought reasoning traces using GPT-4o-mini. As noted earlier, we restrict our experiments to the easy subset, since preliminary experiments show that both baseline models and T2MLR struggle significantly on the medium and hard subsets, making it difficult to draw meaningful comparisons. Prompt for generating HotpotQA chain of thoughts According to the context, answer the question in a multi-hop reasoning process. Be sure to include a detailed and clear logic chain in your reasoning. Detail your thought process at each step. Put your final answer within . Output example: Knox County Regional Airport. Question: question Context: context Note that the correct final answer is answer. F.3 Training details We note the training details used for each task. Task Model LR Schedule Steps S5-Retrieval Transformer: 4-layer, 6-head 384 hidden dim, RoPE LSTM: 4-layer, 456 hidden size 5e-4 Cosine 200k FineWeb-Edu Pretraining SmolLM2-135M-T2MLR 5e-4 Cosine w/ min LR 20k GSM8K-Aug SmolLM2-135M-T2MLR 5e-5 Cosine 8000 HotpotQA SmolLM2-135M-T2MLR 1e-4 Cosine w/ min LR 5000 Table 14: Training details for various experiments. For pretraining on FineWeb-Edu, we used AdamW (Loshchilov and Hutter, 2019) optimizer with β1=0.9 _1=0.9, β2=0.98 _2=0.98, and weight-decay of λ=0.01λ=0.01. We set the minimum learning rate decay to be 0.0010.001 times the peak learning rate. We pack the pretraining data into sequences of length 2048 and train with batchsize 256. For finetuning runs we used AdamW (Loshchilov and Hutter, 2019) optimizer with β1=0.9 _1=0.9, β2=0.999 _2=0.999, and weight-decay of λ=0.01λ=0.01. Training is conducted on 4 × H100 GPUs. Data examples We provide simple examples of the datasets used in our tasks. Dataset Details ProsQA-Hard Input: Every zolufibus is a conarus. Every conarus is a tanirus. Every tanirus is a gedatus. Every cemapus is a lovaperus. Every lovaperus is a zucirarus. Every zucirarus is a pilevus. Every sasarus is a madusurus. Every madusurus is a cidodivus. Every cidodivus is a bumobus. Every bumobus is a nuronus. Every nuronus is a metuvelus. Every metuvelus is a resamus. Every resamus is a gedatus. Jordan is a sasarus. Is Jordan a gedatus or a pilevus? Steps: Jordan is a sasarus. Every sasarus is a madusurus. Every madusurus is a cidodivus. Every cidodivus is a bumobus. Every bumobus is a nuronus. Every nuronus is a metuvelus. Every metuvelus is a resamus. Every resamus is a gedatus. Output: Jordan is a gedatus. S5 Retrieval Input: <A_32514>F4Nd <A_12543>9O8W <A_54213>Jccq | <A_54213> <A_43125> <A_52314> Target: <A_32514>F4Nd <A_12543>9O8W <A_54213>Jccq | <A_54213>Jccq <A_12543>9O8W <A_32514>F4Nd GSM-Aug Natural Language Question: Out of 600 employees in a company, 30% got promoted while 10% received bonus. How many employees did not get either a promotion or a bonus? Steps: 600 x 30/100 = 180 employees were promoted. 600 x 10/100 = 60 employees received a bonus. So a total of 180+60=240 employees received a promotion or a bonus. Therefore, 600 - 240 = 360 employees did not get either a promotion or a bonus. Answer: 360 Variable Assignment Input: Variable assignment. Follow the assignments and fill in the blank. a=5 b=2 c=8 x=c y=x z=y w=b v=w z=___ Answer: | Outputs: 8 HotpotQA Input: Which magazine was founded first, The Economist or The Atlantic? Context: Document 1: The Economist is a weekly newspaper founded in 1843. Document 2: The Atlantic is a magazine founded in 1857. Steps: The Economist was founded in 1843. The Atlantic was founded in 1857. 1843 is earlier than 1857. Answer: The Economist Table 15: Examples for each dataset used. Additional Pretraining Results Table 16 shows the pretraining perplexities of T2MLR and baseline Transformer at different model scales. 135M 368M Baseline T2MLR Baseline T2MLR Perplexity↓ 22.64 21.54 16.12 15.79 Table 16: We pretrain the standard Transformer and T2MLR, both based on SmolLM, at different scales on Fineweb-edu-10B, with T2MLR consistently having improved test time performance.