Paper deep dive
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
Mohan Tang, Sidi Lu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 10:59:32 PM
Summary
The paper introduces Turbo Connection (TurboConn), a novel Transformer architecture that enhances reasoning capabilities by routing residual connections from higher-layer hidden states of token t to lower layers of token t+1. This design overcomes the fixed-depth computational constraint of standard Transformers, allowing effective reasoning depth to scale linearly with sequence length. Experiments on Llama 3 and Qwen 3 models demonstrate significant accuracy gains on benchmarks like GSM8K, Parity, and multi-step arithmetic, with TurboConn enabling Qwen-3-1.7B to achieve 100% accuracy on the Parity task, a feat where standard fine-tuning plateaued at 53.78%. The method improves performance without increasing inference latency or GPU memory usage significantly.
Entities (8)
Relation Signals (7)
Turbo Connection → connects → higher-layer hidden states
confidence 95% · routing multiple residual connections from the higher-layer hidden states of each token t to the lower layers of token t+1
Turbo Connection → improves → Reasoning Ability
confidence 95% · TurboConn yields accuracy gains of 0.9% to over 10% on reasoning benchmarks
Turbo Connection → enables → 100% accuracy on Parity
confidence 92% · adding our architectural modification enables the model to reach 100% accuracy
Turbo Connection → appliedto → Qwen-3
confidence 90% · fine-tuning Llama 3 1B and 8B models and Qwen 3 1.7B model
Turbo Connection → appliedto → Llama 3
confidence 90% · fine-tuning Llama 3 1B and 8B models... with our method
Universal Transformer → comparedwith → Turbo Connection
confidence 85% · While both TurboConn and the Universal Transformer utilize recurrence... we compare the two architectures
Chain-of-Thought → comparedwith → Turbo Connection
confidence 80% · The popular Chain-of-Thought (CoT) framework addresses this... In this work, we aim to increase the effective depth... without... human-annotated chain-of-thought data
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next. In this work, we adopt the perspective that the reasoning power of Transformers is fundamentally limited by a fixed maximum number of steps along any latent path of computation. To address this, we introduce Turbo Connection (TurboConn), a novel architecture that overcomes the fixed-depth constraint by routing multiple residual connections from the higher-layer hidden states of each token $t$ to the lower layers of token $t+1$. Fine-tuning pre-trained LLMs with our method not only yields accuracy gains of 0.9% to over 10% on benchmarks like GSM8K, Parity, and multi-step arithmetic, but also demonstrates that the density of these backward connections is critical; our dense interaction significantly outperforms "sparse" alternatives that only pass a single hidden state or vector. Notably, TurboConn can be integrated into pre-trained LLMs to overcome task-specific plateaus: while a fine-tuned Qwen-3-1.7B achieves only 53.78% on Parity, adding our architectural modification enables the model to reach 100% accuracy, all without the necessity to retrain the full model from scratch or sophisticated curriculum learning. Our results provide strong empirical evidence that the depth of the computational path is a key factor in reasoning ability, also offering a new mechanism to enhance LLMs without significantly affecting generation latency.
Tags
Links
- Source: https://arxiv.org/abs/2602.17993v2
- Canonical: https://arxiv.org/abs/2602.17993v2
Trouble viewing inline? Open PDF directly →
Full Text
78,562 characters extracted from source content.
Expand or collapse full text
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers Mohan Tang Sidi Lu Abstract Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next. In this work, we adopt the perspective that the reasoning power of Transformers is fundamentally limited by a fixed maximum number of steps along any latent path of computation. To address this, we introduce Turbo Connection (TurboConn), a novel architecture that overcomes the fixed-depth constraint by routing multiple residual connections from the higher-layer hidden states of each token t to the lower layers of token t+1t+1. Fine-tuning pre-trained LLMs with our method not only yields accuracy gains of 0.9% to over 10% on benchmarks like GSM8K, Parity, and multi-step arithmetic, but also demonstrates that the density of these backward connections is critical; our dense interaction significantly outperforms "sparse" alternatives that only pass a single hidden state or vector. Notably, TurboConn can be integrated into pre-trained LLMs to overcome task-specific plateaus: while a fine-tuned Qwen-3-1.7B achieves only 53.78% on Parity, adding our architectural modification enables the model to reach 100% accuracy, all without the necessity to retrain the full model from scratch or sophisticated curriculum learning. Our results provide strong empirical evidence that the depth of the computational path is a key factor in reasoning ability, also offering a new mechanism to enhance LLMs without significantly affecting generation latency. Machine Learning, ICML Figure 1: Modified Transformer architecture with downward connections (orange arrows) from higher to lower decoder layers. 1 Introduction While Transformer-based Large Language Models (LLMs) have advanced significantly in recent years, they continue to exhibit shortcomings on tasks that demand complex reasoning. The popular Chain-of-Thought (CoT) framework (Wei et al., 2022) addresses this by allocating dynamic computation through intermediate steps. However, this approach places a considerable strain on computational resources and often requires specialized training data (DeepSeek-AI et al., 2025). A growing area of research (Dehghani et al., 2019; Geiping et al., 2025; Fan et al., 2025a) focuses on performing additional computation in the depth dimension of a Transformer rather than along the token sequence dimension. This is achieved by recursively applying Transformer layers, which enables increased latent reasoning. In this approach, "thinking" occurs within the model’s hidden states instead of being explicitly projected onto tokens, potentially allowing for more complex information processing. A key advantage of this approach is that it allows training on unlabeled pre-training data (Geiping et al., 2025; Zeng et al., 2025). However, it still requires increased time and GPU costs during training, as well as longer inference times that scale proportionally with the number of recursion steps. These factors pose considerable scalability concerns for future development. Is enhancing reasoning solely a matter of increasing the total amount of computation? We argue that the answer is no. In this work, we aim to increase the effective depth—defined as the maximum length of the computational path available to process information—with negligible impact on total floating-point operations. In standard Transformers, information flows strictly from lower to higher layers. As a result, the maximum length of the computational path is always bounded by a fixed number (proportional to the depth of the model). Because the model’s depth is fixed regardless of input length, this architecture imposes limitations on the Transformer’s computational power. Aligning with this perspective, Merrill and Sabharwal (2023) demonstrated that finite-depth Transformers are confined to the uniform TC0TC^0 complexity class, a class of problems solvable by constant-depth circuits, suggesting they may lack the expressive power required for inherently sequential problems. Feng et al. (2023) also presents a viewpoint that CoT works by increasing the "effective depth" of Transformers as the outputs are repeatedly looped back to the input. Overcoming this constraint, we introduce a novel architecture featuring connections from higher layers back to lower layers. Specifically, the hidden state from a higher layer of token t is fed into a lower layer processing token t+1t+1. This modification allows the effective reasoning depth to grow linearly with the sequence length, breaking the fixed-depth constraint of standard Transformers with only a marginal increase in total computation. We refer to our method as Turbo Connection (TurboConn). The t→t+1t→ t+1 design is a deliberate choice that leverages the autoregressive nature of the Transformer decoder to avoid the circular dependencies that would arise from self-connections (e.g., t→t→ t), a point we discuss in detail in Section 3.2. The primary trade-off is the loss of full parallelism during training and the prefill stage of inference, as processing degenerates to group-sequential evaluations. Nevertheless, this has little impact on GPU memory usage or the speed of autoregressive generation. A key advantage of TurboConn is that it requires no modifications to the standard loss function, making it compatible with existing pre-training objectives, though our experiments are focused on the fine-tuning stage due to resource constraints. For this work, we focus on evaluating its effectiveness by fine-tuning Llama 3 1B and 8B models (Grattafiori et al., 2024) and Qwen 3 1.7B model (Yang et al., 2025) on several reasoning-intensive datasets. We find that our architecture consistently outperforms the standard Transformer architectures across these tasks. To summarize our contributions: 1. We propose a novel method that increases the reasoning ability of large language models, without: • additional latency at inference-time autoregressive generation, • increased GPU memory consumption during training or inference, or • the need for human-annotated chain-of-thought data. 2. Across different models and tasks, TurboConn yields accuracy gains ranging from 0.9% to over 10% on reasoning benchmarks. 3. We show that TurboConn enables a qualitative shift in model behavior, including better length generalization on the Parity task and emergence of discriminative filtering. 4. Our results provide strong empirical evidence that the lack of depth in latent computation indeed limits the reasoning ability of LLMs. Having a fixed computational depth for all tokens is likely a necessary consequence of the standard parallel training paradigm for Transformers, where all tokens in a sequence are processed simultaneously. This finding could motivate more research to move beyond purely parallel designs and explore sequential or non-parallel training paradigms to unlock deeper reasoning. 2 Related Work 2.1 Chain-of-Thought People observe that large language models benefit from generating more tokens to "think step by step" (Wei et al., 2022; Kojima et al., 2022). Theoretically, chain-of-thought has been shown to be able to greatly enhance LLMs in terms of computational power (Merrill and Sabharwal, 2024; Li et al., 2024). To train models to produce these reasoning chains, researchers have explored several paradigms. The most direct method is supervised fine-tuning on chain-of-thought data (Liu et al., 2025a; Yue et al., 2023; Ahmad et al., 2025). A more challenging but highly sought-after goal is to elicit CoT reasoning using only final outcomes for supervision (Zelikman et al., 2022; Zhu et al., 2023; Singh et al., 2024). The most prominent current approach in this domain is reinforcement learning from outcome-based rewards (OpenAI et al., 2024; DeepSeek-AI et al., 2025; Team et al., 2025; Liu et al., 2025b). Recent work has pushed these boundaries even further. For instance, some methods aim to train CoT behavior on unlabeled data using only the standard next-token prediction objective (Zelikman et al., 2024). While the standard CoT framework relies on generation of discrete tokens, others have explored continuous Chain-of-Thought, where the discrete reasoning tokens are replaced with continuous representations derived from the model’s hidden states (Hao et al., 2024). This method enables continuous reasoning by passing the top-layer hidden states back to the initial embedding layer on dedicated ⟨pause⟩ tokens inserted between problems and answers. However, training such continuous thinking is challenging, requiring extensive curriculum learning from existing CoT datasets. 2.2 Looped/Universal Transformers Looped or universal Transformers apply the same set of Transformer layers recursively in the depth dimension before producing a final output. This architectural paradigm has shown advantages in reasoning tasks, as demonstrated in several experiments (Dehghani et al., 2019; Saunshi et al., 2025). Theoretically, looped Transformers have been proven capable of solving various logical problems of interest (Giannou et al., 2023; Fan et al., 2025b). More recently, variants of this approach have been successfully scaled up to the billion-parameter range and trained on large, general-purpose pre-training datasets (Geiping et al., 2025; Zeng et al., 2025). Figure 2: Grouping strategy for downward connections. Example shown for group of 2. 2.3 On Relationship between Depth and Reasoning Merrill and Sabharwal (2023) theoretically showed that finite-depth Transformers are confined to the uniform TC0 complexity class, which is believed to have limited computational power. This theoretical constraint has practical implications: assuming the standard conjecture that L≠PL≠ P, Transformers cannot solve problems such as linear equation solving or universal context-free grammar recognition. Feng et al. (2023) also shows that bounded-depth Transformers cannot solve certain basic mathematical tasks unless the model size grows super-polynomially with respect to the input length and explains why CoT works through the lens of increased "effective depth." Empirical evidence also highlights the critical role of depth in how models perform multi-step reasoning. For instance, LLMs have been shown to resolve two-hop problems by handling the first hop in lower layers before processing the second hop in higher layers (Biran et al., 2024). Similarly, intensive training on multi-hop reasoning tasks leads to the gradual strengthening of intermediate representations in the middle layers, which are then used by subsequent layers to reach a final answer (Wang et al., 2025). Our work is distinguished from these approaches by isolating the effect of depth alone, without a significant increase in computation, to empirically show its importance. Particularly relevant is the "back-patching" experiment by Biran et al. (2024), which showed that manually injecting a hidden state from a higher layer back into a lower layer can correct reasoning failures. Our work can be seen as operationalizing this insight: instead of using it as a post-hoc analysis or intervention that requires a second, full forward pass, we integrate this top-to-bottom connection directly into the model’s architecture. Among novel architectural designs, Fan et al. (2021) enables top-down information flow by replacing the Transformer’s standard key-value memory with a single key-value set per token that encapsulates information across all layers. Similarly, Staircase Attention (Ju et al., 2022) introduces a recurrent architecture that applies a shared Transformer block iteratively over overlapping token groups, feeding hidden states from previous steps back into the core network as new tokens are processed. While both methods also provide feedback from higher to lower layers, our work is uniquely designed to augment existing, large-scale pre-trained LLMs. TurboConn successfully leverages expensive, pre-trained features of LLMs by using zero-initialized additive connections. 3 Method 3.1 Background and Notations Consider the following formalization of the Transformer architecture. Let l(i) h^(i)_l be the hidden state of token i at layer l for l=0,…,Ll=0,…,L. Given input tokens t0,t1,t2,…,tkt_0,\ t_1,\ t_2,…,\ t_k, obtain their embeddings 0(i)=i=Embed(ti),i=0,…,k. h^(i)_0= e_i=Embed(t_i), i=0,…,k. For l=1,…,Ll=1,…,L and i=0,…,ki=0,…,k, l(i)=LayerBlockl(l−1(i);l−1(0),l−1(1),…,l−1(i−1)) h^(i)_l=LayerBlock_l\! ( h^(i)_l-1;\, h^(0)_l-1, h^(1)_l-1,…, h^(i-1)_l-1 ) Here LayerBlock(⋅)LayerBlock(·) encapsulates the self-attention, MLP, positional encodings, residuals, etc. The final logits are P(i)=lm_head(L(i)),i=0,…,k. P^(i)=lm\_head\! ( h^(i)_L ), i=0,…,k. 3.2 Our Approach We introduce a novel modification to the standard Transformer architecture by incorporating downward connections from higher layers to lower layers, as shown in Figure 1. In contrast to our approach, an intuitive design might be to feed a token’s higher layers back into its own lower layers (t→t→ t). However, this creates a circular dependency in the computational graph, a logical impossibility in a single forward pass. To resolve this loop, the model would need to re-run its layers multiple times for a single token, a process that dramatically increases the required computation at inference time. In contrast, TurboConn avoids this problem. The causal nature of the Transformer decoder means that token t (past) can influence token t+1t+1 (future) without creating any circular dependency in the computational graph. Formally, for l=1,…,Ll=1,…,L and i=1,…,ki=1,…,k, we modify the hidden state computation as such: • When a connection from layer s to layer l exists: l(i)=LayerBlockl h^(i)_l=LayerBlock_l (l−1(i);l−1(0),…,l−1(i−1)) \! ( h^(i)_l-1; h^(0)_l-1,…, h^(i-1)_l-1 ) +α⋅(s(i−1)) +α·D\! ( h^(i-1)_s ) • Otherwise (when no connection (s→l)(s→ l) exists): l(i)=LayerBlockl(l−1(i);l−1(0),…,l−1(i−1)) h^(i)_l=LayerBlock_l\! ( h^(i)_l-1; h^(0)_l-1,…, h^(i-1)_l-1 ) (⋅)D(·) is a learnable linear projection specific to each connection. We can initialize (⋅)D(·) to be all zeros when finetuning, so that the model behaves the same as the original LLM at the start of training. The multiplier α scales the strength of downward connections. It does not affect the model’s theoretical representational power, but encourages greater utilization of the connections during training. By default we set α to 100. The effect of different multiplier values is analyzed in Section 4.4. Note that there can be multiple connections, as shown in Figure 3. Because l(i) h^(i)_l depends only on tokens 0,…,i0,…,i, the computation can proceed token-by-token from left to right. TurboConn preserves the standard language modeling paradigm and can be trained using the conventional autoregressive loss: ℒ=−∑i=1klogP(ti∣t0,t1,…,ti−1)L=- _i=1^k P(t_i t_0,t_1,…,t_i-1) 3.2.1 Analysis of Depth As an example, consider the case that we have a connection from the highest layer to the lowest layer. In standard Transformers, the maximum computational path length is L (corresponding to the number of layers). With TurboConn, this path length extends to kLkL, where k is the sequence length. The maximum depth of reasoning now scales linearly with the number of input tokens. While this does not enable the model to solve all arbitrarily complex problems, the shift from a constant to a linear-depth computational model is a fundamental breakthrough compared to the standard Transformer. Figure 3: This diagram illustrates the dense connectivity pattern utilized in our experimental setup, depicted here for a group size of 1. 3.3 Analysis of Computational Cost The downward connections introduce sequential dependencies between tokens, breaking the parallelism typically available during training. This requires sequential processing of tokens during the forward pass, potentially increasing training latency. Similarly, the prefill stage during inference degenerates to sequential processing, introducing a potential latency bottleneck, though this can be mitigated by evaluating tokens in larger groups. However, this sequential constraint does not impact the generation phase at inference time, as autoregressive generation inherently processes tokens one by one. Regarding memory requirements, TurboConn introduces only marginal additional hidden states that need to be stored, resulting in minimal increase in GPU memory usage. The memory footprint remains comparable to standard Transformers. 3.4 Grouping To mitigate the computational cost of sequential processing, we introduce a grouping mechanism that enables more efficient parallel computation. As shown in Figure 2, we group tokens together and send information back in groups rather than processing individual tokens sequentially. This approach processes tokens in groups of size g, where each group receives downward connections simultaneously. For layers l=1,2,…,Ll=1,2,…,L and token positions i=g,g+1,…,ki=g,g+1,…,k, we modify the hidden state computation as follows: when a connection from layer s to layer l (s→l)(s→ l) exists: l(i)=LayerBlockl h^(i)_l=LayerBlock_l (l−1(i);l−1(0),…,l−1(i−1)) \! ( h^(i)_l-1; h^(0)_l-1,…, h^(i-1)_l-1 ) +α⋅(s(i−g)) +α·D\! ( h^(i-g)_s ) In this way, (i) h^(i) does not depend on hidden states at tokens i−g+1,i−g+2,…,i−1i-g+1,i-g+2,…,i-1, so computations within each group can be performed in parallel. With group size g, the computational depth becomes kL/gkL/g (assuming there is a connection from top to bottom), and the number of sequential steps reduces from k to k/gk/g. This reduces training latency while decreasing computational power. The group size can be chosen based on the specific use case requirements. Applications requiring deep sequential reasoning may benefit from smaller group sizes, while those prioritizing training efficiency may prefer larger groups. 3.5 Comparison with Universal Transformers While both TurboConn and the Universal Transformer utilize recurrence to increase depth, the Universal Transformer is recurrent in depth (per token), whereas TurboConn is recurrent across the sequence (cross-token). To clarify the mechanical differences, we compare the two architectures under the assumption that both are configured to increase the effective reasoning depth by a factor of D. 3.5.1 Computational Efficiency and Scaling To increase the effective reasoning depth by a factor of D, the Universal Transformer must perform D recursive iterations for every token. Its training and inference time scale linearly with D. In contrast, TurboConn scales the depth by utilizing the sequence dimension. Although it breaks full token-level parallelism, it preserves intra-group and tensor-wise parallelism. Therefore, the training time increases by a factor significantly less than D. 3.5.2 Memory Footprint For the Universal Transformer to achieve a depth factor of D, the memory used must grow by a factor of D, as the number of states that must be stored for each recursion increases. Conversely, TurboConn introduces no significant change in memory cost. 3.5.3 Autoregressive Generation and Inference The difference is also pronounced during the autoregressive generation phase. For the Universal Transformer, each generated token requires D passes through the recurrent layers, which increases the latency by a factor of D. TurboConn, however, involves little additional cost during generation. The backward connections are integrated into the single forward pass of the model, allowing for depth-scaling without the linear latency penalty of Universal Transformers. 4 Experiments 4.1 Reasoning Tasks We evaluate the effectiveness of TurboConn on the following datasets. Each dataset consists of about 380K questions: 1. Parity (Fan et al., 2025a; Banino et al., 2021; Graves, 2017): Given a sequence of 0s and 1s, the task is to determine whether there is an odd number of 1s. We sample sequences ranging from 1 to 70 digits in length. 2. Multi-step Arithmetic (Srivastava et al., 2023): Randomly generated arithmetic expressions using digits 0-9, with the number of operands ranging from 1 to 30. The result is projected to modulo 10 for a more uniform distribution of answers. 3. GSM8K (Cobbe et al., 2021): A collection of grade school math problems. We use an enlarged dataset augmented by Deng et al. (2023). To evaluate performance without process supervision, we removed the chain-of-thought reasoning steps from the original dataset. We train and evaluate on different splits of the same distribution. Further details on training data can be found in Appendix A.4. 4.2 Training Hyperparameters We evaluate performance by fine-tuning pre-trained Llama 3 and Qwen 3 models, comparing a standard Transformer baseline against the Transformer augmented with TurboConn under identical training settings. Training uses a batch size of 64, resulting in approximately 6,000 iterations per epoch for each dataset. We train for a maximum of 3 epochs, following standard fine-tuning practices to prevent overfitting. Each model has a dense connection setup with 15 to 45 connections. This configuration empirically provides more stable training. Due to resource constraints, we use LoRA on attention and feed-forward parameters. To ensure fair comparison, we set the rank to r = 120 for TurboConn and r = 140 for the baseline, maintaining comparable parameter counts. The downward connection projections (⋅)D(·) are also implemented as low-rank linear maps. We employ a cosine scheduler (Loshchilov and Hutter, 2017), where the lr increases and decreases in periods of 1000 steps. In each period, it is warmed up from 0 to the highest value in 100 steps, and then gradually decrease to 0 again following a Cosine function with period 900. Additional training details are provided in Appendix A. 4.3 Main Results We evaluate our method on Llama 3.2 1B and Llama 3.1 8B models (Grattafiori et al., 2024) and Qwen 3 1.7B model. We use a group size of 4 and set the connection multiplier to 100. Results are presented in Table 2. TurboConn demonstrates significant improvements over standard Transformers across all model sizes and datasets. Notably, the 1B model with TurboConn outperforms the 8B model without connections on the Parity task, suggesting that improved information flow can be more beneficial than increased computational scale for certain reasoning tasks. Moreover, adding TurboConn to pretrained LLM enables Qwen-3-1.7B to achieve 100%100\% accuracy on Parity, where fine-tuning alone fails to resolve the underlying complexity and plateaus at 53.78%. Table 1: Training efficiency and performance analysis for Llama 3.2 1B across group sizes for multipliers α=1α=1 and α=100α=100. We report average sequence length, average number of recursions per training iteration, and relative training time per step compared to the baseline transformer across three reasoning tasks. Dataset (Avg. Seq. Length) Method # Recursions Time/Step (seconds) (× Transformer) Acc (%) (α=1α=1) Acc (%) (α=100α=100) Parity (154.69 tokens) Baseline 1.00 1.18 (1.00×) 92.87 – TurboConn (Group 4) 38.81 5.75 (4.87×) 98.94 100.00 TurboConn (Group 6) 25.94 3.07 (2.60×) 98.68 100.00 TurboConn (Group 8) 19.84 2.66 (2.25×) 98.83 51.42 TurboConn (Group 16) 10.00 1.61 (1.36×) 98.85 51.40 Multi-step Arithmetic (125.75 tokens) Baseline 1.00 1.01 (1.00×) 38.16 – TurboConn (Group 4) 31.81 4.46 (4.42×) 40.13 42.66 TurboConn (Group 6) 21.37 2.29 (2.27×) 39.94 41.74 TurboConn (Group 8) 16.13 1.82 (1.80×) 39.79 41.48 TurboConn (Group 16) 8.17 1.45 (1.44×) 38.79 38.79 GSM8K (No CoT) (101.92 tokens) Baseline 1.00 0.77 (1.00×) 7.20 – TurboConn (Group 4) 25.86 3.22 (4.18×) 7.67 8.32 TurboConn (Group 6) 17.40 1.73 (2.25×) 7.87 7.95 TurboConn (Group 8) 13.18 1.39 (1.81×) 7.55 8.17 TurboConn (Group 16) 6.84 1.38 (1.80×) 7.41 6.94 4.4 Effect of Group Sizes and Multiplier In this section, we evaluate the effect of different group sizes g on both model performance and computational efficiency. Furthermore, we find that the choice of group size dictates the optimal setting for the connection multiplier (α). We conduct our analysis using the Llama 3.2 1B model across all three datasets. To ensure consistent experimental setup, we remove a small number of overly long examples (ones exceeding 168 tokens using Llama 3.2 1B tokenizer) from GSM8K so that all experiments can be run with identical settings. Based on this threshold, only 44 examples were removed from the original dataset of 384,620. Each measurement is performed on a single 80G A100 GPU. Table 2: Performance comparison across model sizes and datasets, comparing standard Transformers baseline against Transformers augmented with TurboConn. Model Dataset Method Acc (%) Llama 3.2 1B GSM8K (No CoT) Baseline 7.20 TurboConn 8.32 Multi-step- arithmetic Baseline 38.16 TurboConn 42.66 Parity Baseline 92.87 TurboConn 100.0 Llama 3.1 8B GSM8K (No CoT) Baseline 23.92 TurboConn 24.82 Multi-step- arithmetic Baseline 48.32 TurboConn 51.86 Parity Baseline 89.40 TurboConn 100.0 Qwen 3 1.7B GSM8K (No CoT) Baseline 15.90 TurboConn 20.31 Multi-step- arithmetic Baseline 36.10 TurboConn 45.81 Parity Baseline 53.78 TurboConn 100.0 4.4.1 Training Latency Analysis Table 1 presents computational efficiency analysis across different group sizes. The results demonstrate that larger group sizes lead to reduced training latency. Importantly, when we recurse for k/gk/g times, the training latency increases by less than a factor of k/gk/g. This occurs because our sequential token processing preserves other forms of parallelism, including batch-level parallelism and vectorized operations within the attention and feed-forward computations. As shown in Table 1, on the Parity task, group size 4 requires 38.81 recursions yet increases training time by only 4.87×. Moreover, Parity with group size 16 achieves 10× deeper reasoning paths with just 1.36× training overhead–delivering enhanced reasoning capability at near-baseline computational cost. This reveals a favorable trade-off between reasoning depth and computational efficiency, suggesting that our method can provide substantial improvements in model reasoning capabilities with manageable increases in training time. 4.4.2 Performance Analysis The results are shown in Table 1. In general, we can see that smaller group sizes result in stronger performance. One issue we observed is that when using a large multiplier (α=100α=100) and a large group size, the performance can sometimes degrade compared to the baseline. This suggests that while the downward connections pass valuable, high-level information from previous tokens, the reduced computational depth of larger groups may be insufficient to process this information effectively. The influx of strong signals from higher layers can destabilize the training process. When using a high multiplier, the strong signals from downward connections may also cause the model to over-rely on the connections for information propagation between tokens at the expense of its standard attention mechanisms. Conversely, setting the multiplier to a lower value (α=1α=1) mitigates this issue, leading to consistent performance gains over the standard Transformer across all tested group sizes. On the other hand, this approach results in less significant improvements when group size is small, compared to using α=100α=100. A smaller α value encourages the model to utilize the downward connections more subtly, resulting in steady, incremental advantages without destabilizing the training process. Therefore, the recommended approach is to pair a small group size with a large multiplier and a larger group size with a small multiplier. 4.5 Comparison to Single “Soft-Token” Feedback Chain-of-thought (CoT) can also be conceptualized as a mechanism for passing information from higher to lower layers, by feeding back the information of a single discrete token into the input embedding of the next step. Similarly, methods like Coconut (Hao et al., 2024) utilize a continuous CoT approach, passing the highest hidden state back to the lowest layer, effectively passing information on the distribution of the next token. In contrast, our method is more “dense,” passing hidden states from multiple layers back to multiple layers simultaneously. To determine if this architectural density is providing extra reasoning power, we compare our approach against a single “soft-token” feedback baseline. We implement a feedback mechanism that passes only a single vector of information from the top of the model back to the input. Specifically, the input to the first layer for token i, denoted as h0(i)h_0^(i), is modified as follows: h0(i)=(1−λ)ei+λ∑j∈P(xi=j|x<i)Emb(j)h_0^(i)=(1-λ)e_i+λ _j P(x_i=j|x_<i)Emb(j) where eie_i is the original embedding of token i, Emb is the embedding matrix, and λ is a scaling factor (set to 0.1) representing the strength of the feedback. In this setup, λ=1λ=1 would imply the prediction of the previous token completely overwrites the current input. Table 3: Comparison of our dense feedback method against standard Transformers baseline and the single soft-token feedback using Llama 3.2 1B. Model Dataset Method Acc (%) Llama 3.2 1B GSM8K (No CoT) Baseline 7.20 Soft-token 7.07 TurboConn 8.32 Multi-step Arithmetic Baseline 38.16 Soft-token 38.00 TurboConn 42.66 Parity Baseline 92.87 Soft-token 96.51 TurboConn 100.0 4.5.1 Results and Analysis. As shown in Table 3, the single soft-token approach is insufficient to capture the reasoning gains of our dense architecture. These results suggest that multi-layer downward connections offer a more effective mechanism for unlocking the reasoning power of depth-scaling than the sparse feedback of a single vector or distribution. 4.6 Length Generalization Experiment We hypothesize that models with TurboConn learn fundamentally different problem representations compared to the standard Transformer. This is particularly evident in the Parity task, where our method achieves perfect accuracy. To further verify this claim, we conducted a length generalization experiment using the Llama 3.1 8B model. TurboConn (with a group size of 1), “soft-token” method, and the baseline Transformer were all trained exclusively on parity sequences with a maximum length of 10. We then evaluated their performance on much longer sequences. The results are shown in Figure 4. While all models achieved near-perfect accuracy on the training data, their generalization capabilities diverged significantly. As we can see from the plot, the LLM with TurboConn shows much better length generalization ability than Transformer and soft-token method. In particular, our method maintains perfect accuracy on sequences up to length 30. This suggests that TurboConn allows the model to learn a more robust and generalizable algorithm for the task, whereas the standard Transformer appears to overfit to the training distribution, leveraging its vast parameter count to solve the problem. Figure 4: Length Generalization Performance. We evaluate a Llama 3.1 8B model, trained on Parity with up to 10-digit sequences, on its ability to generalize to longer input sequences. 4.7 Emergence of Discriminative Filtering Beyond improvements in top-1 accuracy, we observe that TurboConn results in a fundamental change to the output distribution of the models, and enables the emergence of discriminative filtering. We define this as the model’s ability to confidently eliminate mathematically incorrect candidates from the output probability distribution. To quantify this effect, we measure the number of “eliminated choices,” defined as the count of digits 0,…,9\0,…,9\ assigned a probability below a threshold of τ=0.001τ=0.001 by the models after finetuning. As shown in Table 4, on the Multi-step Arithmetic task, Qwen-1.7B model augmented with TurboConn eliminates an average of 5.259 choices per token, more than doubling the 2.516 choices eliminated by the standard Transformer baseline. Table 4: Comparison of accuracy and discriminative filtering on Multi-step Arithmetic after finetuning. “Eliminated Choices” represents the mean number of digit classes with probability P<0.001P<0.001. The result for TurboConn is reported using a group size of 4. Model Method Acc (%) Eliminated Choices Qwen3-1.7B Baseline 36.10 2.516 Qwen3-1.7B TurboConn 45.81 5.259 Qwen3-4B Baseline 43.68 4.609 Qwen3-8B Baseline 45.61 4.500 A particularly striking result is that while increasing the model scale allows standard Transformers to match our accuracy, their discriminative power remains significantly lower. Qwen 3 1.7B model with TurboConn eliminates more choices (5.259) than the 8B baseline (4.500), despite achieving nearly identical accuracy (45.81% vs 45.61%). This signals a fundamental gap in this reasoning dimension that cannot be easily filled by parameter scaling alone, suggesting that the depth afforded by our method enables a more efficient internal verification process. 4.8 Synergy with Chain-of-Thought Reasoning While TurboConn enhances the internal reasoning capabilities of models without explicit intermediate tokens, it is also compatible with standard Chain-of-Thought prompting. One way to combine those methods is to augment the explicit reasoning of CoT with the latent reasoning ability of backward connections. We hypothesize that TurboConn can act as an error-correction mechanism for explicit reasoning chains, allowing the model to more accurately follow the “program” laid out in the CoT. To test this synergy, we utilize the NuminaMath-CoT dataset (LI et al., 2024) and augmented GSM8K (Deng et al., 2023) dataset to generate CoT trajectories. We first fine-tune a Llama 3.2-1B model to produce reasoning chains for mathematical problems. We then compare models with TurboConn against the standard Transformers baseline by fine-tuning both to predict the final numerical outcome based on the same generated CoTs. This setup ensures that both models process the same explicit reasoning steps, isolating the effect of latent computation on final answer derivation. More details on this experiment are provided in Appendix A.7. As shown in Table 5, TurboConn improves the model’s ability to reach the correct conclusion from a given reasoning chain. These results suggest that TurboConn can effectively use latent computation to verify intermediate steps and correct potential errors within the explicit reasoning process, leading to more robust mathematical performance. We note that the performance gains on GSM8K are less pronounced compared to those on NuminaMath. Since the training data for GSM8K is augmented using GPT-4, there is a significant distribution shift between the augmented training data and the original test/validation splits, potentially reducing the utility of the learned representations. We believe that our method would benefit more from access to sufficient higher-quality data that closely matches the target distribution. 4.8.1 Generalization to Competition Mathematics To further evaluate the robustness of our method, we test the checkpoints trained on the NuminaMath dataset on recent competition mathematics problems. This evaluation set is an aggregate of problems from the 2023 AMC and 2024–2025 AIME competitions (Mathematical Association of America, 2023, 2024, 2025). We employ the same “correct existing CoT” setting described above, utilizing the Llama 3.2-1B model finetuned on NuminaMath to generate the initial CoT trajectories, and maintaining identical inference hyperparameters. As shown in Table 5, on this combined competition dataset, TurboConn demonstrates a significant relative performance gain (4.45% vs. 2.00%). These results indicate that even when a 1B parameter model struggles to generate high-quality explicit reasoning chains for highly complex problems, TurboConn’s latent reasoning mechanism is still capable of extracting value from imperfect trajectories to correct errors and improve final performance. Table 5: Performance comparison using generated Chain-of-Thought trajectories. All results use Llama 3.2-1B; TurboConn uses a group size of 8. Task Method Acc (%) NuminaMath-CoT Baseline + CoT 30.96 TurboConn + CoT 33.98 GSM8K Baseline + CoT 36.45 TurboConn + CoT 36.96 AMC & AIME (Transferred) Baseline + CoT 2.00 TurboConn + CoT 4.45 5 Conclusions In this work, we introduce TurboConn, a novel method that creates residual connections from higher to lower layers between sequential tokens, allowing the model’s effective reasoning depth to scale linearly with sequence length. Our results demonstrated that this architectural change yields significant performance gains on logical reasoning tasks, with only a manageable increase in computational cost. This confirms empirically that extending effective depth alone is a viable strategy for overcoming the inherent reasoning bottlenecks of fixed-depth Transformers. It would be valuable to explore the potential of this architecture further. Future work could focus on integrating this architecture into the pre-training phase, which may foster the development of more foundational, general-purpose reasoning skills. Acknowledgment This work was conducted during the authors’ studies at the University of California, Los Angeles. The authors acknowledge Peng’s Language Understanding & Synthesis Lab at UCLA for providing the computational resources and infrastructure necessary for the experiments in this study. Impact Statement This paper presents work whose goal is to advance the field of machine learning. There are many potential societal consequences of our work, none of which we feel must be specifically highlighted here. References W. U. Ahmad, S. Narenthiran, S. Majumdar, A. Ficek, S. Jain, J. Huang, V. Noroozi, and B. Ginsburg (2025) OpenCodeReasoning: advancing data distillation for competitive coding. External Links: 2504.01943, Link Cited by: §2.1. A. Banino, J. Balaguer, and C. Blundell (2021) PonderNet: learning to ponder. External Links: 2107.05407, Link Cited by: item 1. E. Biran, D. Gottesman, S. Yang, M. Geva, and A. Globerson (2024) Hopping too late: exploring the limitations of large language models on multi-hop queries. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 14113–14130. External Links: Link, Document Cited by: §2.3, §2.3. K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman (2021) Training verifiers to solve math word problems. External Links: 2110.14168, Link Cited by: item 3. DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025) DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning. External Links: 2501.12948, Link Cited by: §1, §2.1. M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser (2019) Universal transformers. External Links: 1807.03819, Link Cited by: §1, §2.2. Y. Deng, K. Prasad, R. Fernandez, P. Smolensky, V. Chaudhary, and S. Shieber (2023) Implicit chain of thought reasoning via knowledge distillation. External Links: 2311.01460, Link Cited by: §A.7, item 3, §4.8. A. Fan, T. Lavril, E. Grave, A. Joulin, and S. Sukhbaatar (2021) Addressing some limitations of transformers with feedback memory. External Links: 2002.09402, Link Cited by: §2.3. Y. Fan, Y. Du, K. Ramchandran, and K. Lee (2025a) Looped transformers for length generalization. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1, item 1. Y. Fan, Y. Du, K. Ramchandran, and K. Lee (2025b) Looped transformers for length generalization. External Links: 2409.15647, Link Cited by: §2.2. G. Feng, B. Zhang, Y. Gu, H. Ye, D. He, and L. Wang (2023) Towards revealing the mystery behind chain of thought: a theoretical perspective. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. External Links: Link Cited by: §1, §2.3. J. Geiping, S. McLeish, N. Jain, J. Kirchenbauer, S. Singh, B. R. Bartoldson, B. Kailkhura, A. Bhatele, and T. Goldstein (2025) Scaling up test-time compute with latent reasoning: a recurrent depth approach. External Links: 2502.05171, Link Cited by: §1, §2.2. A. Giannou, S. Rajput, J. Sohn, K. Lee, J. D. Lee, and D. Papailiopoulos (2023) Looped transformers as programmable computers. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. External Links: Link Cited by: §2.2. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §1, §4.3. A. Graves (2017) Adaptive computation time for recurrent neural networks. External Links: 1603.08983, Link Cited by: item 1. S. Hao, S. Sukhbaatar, D. Su, X. Li, Z. Hu, J. Weston, and Y. Tian (2024) Training large language models to reason in a continuous latent space. External Links: 2412.06769, Link Cited by: §2.1, §4.5. D. Ju, S. Roller, S. Sukhbaatar, and J. Weston (2022) Staircase attention for recurrent processing of sequences. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA. External Links: ISBN 9781713871088 Cited by: §2.3. T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa (2022) Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA. External Links: ISBN 9781713871088, Link Cited by: §2.1. J. LI, E. Beeching, L. Tunstall, B. Lipkin, R. Soletskyi, S. C. Huang, K. Rasul, L. Yu, A. Jiang, Z. Shen, Z. Qin, B. Dong, L. Zhou, Y. Fleureau, G. Lample, and S. Polu (2024) NuminaMath. Numina. Note: [https://huggingface.co/AI-MO/NuminaMath-CoT](https://github.com/project-numina/aimo-progress-prize/blob/main/report/numina_dataset.pdf) Cited by: §4.8. Z. Li, H. Liu, D. Zhou, and T. Ma (2024) Chain of thought empowers transformers to solve inherently serial problems. External Links: 2402.12875, Link Cited by: §2.1. Z. Liu, Y. Chen, M. Shoeybi, B. Catanzaro, and W. Ping (2025a) AceMath: advancing frontier math reasoning with post-training and reward modeling. External Links: 2412.15084, Link Cited by: §2.1. Z. Liu, Z. Yang, Y. Chen, C. Lee, M. Shoeybi, B. Catanzaro, and W. Ping (2025b) AceReason-nemotron 1.1: advancing math and code reasoning through sft and rl synergy. External Links: 2506.13284, Link Cited by: §2.1. I. Loshchilov and F. Hutter (2017) SGDR: stochastic gradient descent with warm restarts. External Links: 1608.03983, Link Cited by: §4.2. Mathematical Association of America (2023) 2023 AMC problems and solutions. Note: Archived by the Art of Problem Solving. External Links: Link Cited by: §4.8.1. Mathematical Association of America (2024) 2024 AIME i and i problems and solutions. Note: Archived by the Art of Problem Solving. External Links: Link Cited by: §4.8.1. Mathematical Association of America (2025) 2025 AIME i and i problems and solutions. Note: Archived by the Art of Problem Solving. External Links: Link Cited by: §4.8.1. W. Merrill and A. Sabharwal (2023) The parallelism tradeoff: limitations of log-precision transformers. Transactions of the Association for Computational Linguistics 11, p. 531–545. External Links: Link, Document Cited by: §1, §2.3. W. Merrill and A. Sabharwal (2024) The expressive power of transformers with chain of thought. External Links: 2310.07923, Link Cited by: §2.1. OpenAI, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li (2024) OpenAI o1 system card. External Links: 2412.16720, Link Cited by: §2.1. N. Saunshi, N. Dikkala, Z. Li, S. Kumar, and S. J. Reddi (2025) Reasoning with latent thoughts: on the power of looped transformers. External Links: 2502.17416, Link Cited by: §2.2. A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, J. Harrison, J. Lee, K. Xu, A. Parisi, A. Kumar, A. Alemi, A. Rizkowsky, A. Nova, B. Adlam, B. Bohnet, G. Elsayed, H. Sedghi, I. Mordatch, I. Simpson, I. Gur, J. Snoek, J. Pennington, J. Hron, K. Kenealy, K. Swersky, K. Mahajan, L. Culp, L. Xiao, M. L. Bileschi, N. Constant, R. Novak, R. Liu, T. Warkentin, Y. Qian, Y. Bansal, E. Dyer, B. Neyshabur, J. Sohl-Dickstein, and N. Fiedel (2024) Beyond human data: scaling self-training for problem-solving with language models. External Links: 2312.06585, Link Cited by: §2.1. A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, A. Kluska, A. Lewkowycz, A. Agarwal, A. Power, A. Ray, A. Warstadt, A. W. Kocurek, A. Safaya, A. Tazarv, A. Xiang, A. Parrish, A. Nie, A. Hussain, A. Askell, A. Dsouza, A. Slone, A. Rahane, A. S. Iyer, A. Andreassen, A. Madotto, A. Santilli, A. Stuhlmüller, A. Dai, A. La, A. Lampinen, A. Zou, A. Jiang, A. Chen, A. Vuong, A. Gupta, A. Gottardi, A. Norelli, A. Venkatesh, A. Gholamidavoodi, A. Tabassum, A. Menezes, A. Kirubarajan, A. Mullokandov, A. Sabharwal, A. Herrick, A. Efrat, A. Erdem, A. Karakaş, B. R. Roberts, B. S. Loe, B. Zoph, B. Bojanowski, B. Özyurt, B. Hedayatnia, B. Neyshabur, B. Inden, B. Stein, B. Ekmekci, B. Y. Lin, B. Howald, B. Orinion, C. Diao, C. Dour, C. Stinson, C. Argueta, C. F. Ramírez, C. Singh, C. Rathkopf, C. Meng, C. Baral, C. Wu, C. Callison-Burch, C. Waites, C. Voigt, C. D. Manning, C. Potts, C. Ramirez, C. E. Rivera, C. Siro, C. Raffel, C. Ashcraft, C. Garbacea, D. Sileo, D. Garrette, D. Hendrycks, D. Kilman, D. Roth, D. Freeman, D. Khashabi, D. Levy, D. M. González, D. Perszyk, D. Hernandez, D. Chen, D. Ippolito, D. Gilboa, D. Dohan, D. Drakard, D. Jurgens, D. Datta, D. Ganguli, D. Emelin, D. Kleyko, D. Yuret, D. Chen, D. Tam, D. Hupkes, D. Misra, D. Buzan, D. C. Mollo, D. Yang, D. Lee, D. Schrader, E. Shutova, E. D. Cubuk, E. Segal, E. Hagerman, E. Barnes, E. Donoway, E. Pavlick, E. Rodola, E. Lam, E. Chu, E. Tang, E. Erdem, E. Chang, E. A. Chi, E. Dyer, E. Jerzak, E. Kim, E. E. Manyasi, E. Zheltonozhskii, F. Xia, F. Siar, F. Martínez-Plumed, F. Happé, F. Chollet, F. Rong, G. Mishra, G. I. Winata, G. de Melo, G. Kruszewski, G. Parascandolo, G. Mariani, G. Wang, G. Jaimovitch-López, G. Betz, G. Gur-Ari, H. Galijasevic, H. Kim, H. Rashkin, H. Hajishirzi, H. Mehta, H. Bogar, H. Shevlin, H. Schütze, H. Yakura, H. Zhang, H. M. Wong, I. Ng, I. Noble, J. Jumelet, J. Geissinger, J. Kernion, J. Hilton, J. Lee, J. F. Fisac, J. B. Simon, J. Koppel, J. Zheng, J. Zou, J. Kocoń, J. Thompson, J. Wingfield, J. Kaplan, J. Radom, J. Sohl-Dickstein, J. Phang, J. Wei, J. Yosinski, J. Novikova, J. Bosscher, J. Marsh, J. Kim, J. Taal, J. Engel, J. Alabi, J. Xu, J. Song, J. Tang, J. Waweru, J. Burden, J. Miller, J. U. Balis, J. Batchelder, J. Berant, J. Frohberg, J. Rozen, J. Hernandez-Orallo, J. Boudeman, J. Guerr, J. Jones, J. B. Tenenbaum, J. S. Rule, J. Chua, K. Kanclerz, K. Livescu, K. Krauth, K. Gopalakrishnan, K. Ignatyeva, K. Markert, K. D. Dhole, K. Gimpel, K. Omondi, K. Mathewson, K. Chiafullo, K. Shkaruta, K. Shridhar, K. McDonell, K. Richardson, L. Reynolds, L. Gao, L. Zhang, L. Dugan, L. Qin, L. Contreras-Ochando, L. Morency, L. Moschella, L. Lam, L. Noble, L. Schmidt, L. He, L. O. Colón, L. Metz, L. K. Şenel, M. Bosma, M. Sap, M. ter Hoeve, M. Farooqi, M. Faruqui, M. Mazeika, M. Baturan, M. Marelli, M. Maru, M. J. R. Quintana, M. Tolkiehn, M. Giulianelli, M. Lewis, M. Potthast, M. L. Leavitt, M. Hagen, M. Schubert, M. O. Baitemirova, M. Arnaud, M. McElrath, M. A. Yee, M. Cohen, M. Gu, M. Ivanitskiy, M. Starritt, M. Strube, M. Swędrowski, M. Bevilacqua, M. Yasunaga, M. Kale, M. Cain, M. Xu, M. Suzgun, M. Walker, M. Tiwari, M. Bansal, M. Aminnaseri, M. Geva, M. Gheini, M. V. T, N. Peng, N. A. Chi, N. Lee, N. G. Krakover, N. Cameron, N. Roberts, N. Doiron, N. Martinez, N. Nangia, N. Deckers, N. Muennighoff, N. S. Keskar, N. S. Iyer, N. Constant, N. Fiedel, N. Wen, O. Zhang, O. Agha, O. Elbaghdadi, O. Levy, O. Evans, P. A. M. Casares, P. Doshi, P. Fung, P. P. Liang, P. Vicol, P. Alipoormolabashi, P. Liao, P. Liang, P. Chang, P. Eckersley, P. M. Htut, P. Hwang, P. Miłkowski, P. Patil, P. Pezeshkpour, P. Oli, Q. Mei, Q. Lyu, Q. Chen, R. Banjade, R. E. Rudolph, R. Gabriel, R. Habacker, R. Risco, R. Millière, R. Garg, R. Barnes, R. A. Saurous, R. Arakawa, R. Raymaekers, R. Frank, R. Sikand, R. Novak, R. Sitelew, R. LeBras, R. Liu, R. Jacobs, R. Zhang, R. Salakhutdinov, R. Chi, R. Lee, R. Stovall, R. Teehan, R. Yang, S. Singh, S. M. Mohammad, S. Anand, S. Dillavou, S. Shleifer, S. Wiseman, S. Gruetter, S. R. Bowman, S. S. Schoenholz, S. Han, S. Kwatra, S. A. Rous, S. Ghazarian, S. Ghosh, S. Casey, S. Bischoff, S. Gehrmann, S. Schuster, S. Sadeghi, S. Hamdan, S. Zhou, S. Srivastava, S. Shi, S. Singh, S. Asaadi, S. S. Gu, S. Pachchigar, S. Toshniwal, S. Upadhyay, Shyamolima, Debnath, S. Shakeri, S. Thormeyer, S. Melzi, S. Reddy, S. P. Makini, S. Lee, S. Torene, S. Hatwar, S. Dehaene, S. Divic, S. Ermon, S. Biderman, S. Lin, S. Prasad, S. T. Piantadosi, S. M. Shieber, S. Misherghi, S. Kiritchenko, S. Mishra, T. Linzen, T. Schuster, T. Li, T. Yu, T. Ali, T. Hashimoto, T. Wu, T. Desbordes, T. Rothschild, T. Phan, T. Wang, T. Nkinyili, T. Schick, T. Kornev, T. Tunduny, T. Gerstenberg, T. Chang, T. Neeraj, T. Khot, T. Shultz, U. Shaham, V. Misra, V. Demberg, V. Nyamai, V. Raunak, V. Ramasesh, V. U. Prabhu, V. Padmakumar, V. Srikumar, W. Fedus, W. Saunders, W. Zhang, W. Vossen, X. Ren, X. Tong, X. Zhao, X. Wu, X. Shen, Y. Yaghoobzadeh, Y. Lakretz, Y. Song, Y. Bahri, Y. Choi, Y. Yang, Y. Hao, Y. Chen, Y. Belinkov, Y. Hou, Y. Hou, Y. Bai, Z. Seid, Z. Zhao, Z. Wang, Z. J. Wang, Z. Wang, and Z. Wu (2023) Beyond the imitation game: quantifying and extrapolating the capabilities of language models. External Links: 2206.04615, Link Cited by: item 2. K. Team, A. Du, B. Gao, B. Xing, C. Jiang, C. Chen, C. Li, C. Xiao, C. Du, C. Liao, C. Tang, C. Wang, D. Zhang, E. Yuan, E. Lu, F. Tang, F. Sung, G. Wei, G. Lai, H. Guo, H. Zhu, H. Ding, H. Hu, H. Yang, H. Zhang, H. Yao, H. Zhao, H. Lu, H. Li, H. Yu, H. Gao, H. Zheng, H. Yuan, J. Chen, J. Guo, J. Su, J. Wang, J. Zhao, J. Zhang, J. Liu, J. Yan, J. Wu, L. Shi, L. Ye, L. Yu, M. Dong, N. Zhang, N. Ma, Q. Pan, Q. Gong, S. Liu, S. Ma, S. Wei, S. Cao, S. Huang, T. Jiang, W. Gao, W. Xiong, W. He, W. Huang, W. Xu, W. Wu, W. He, X. Wei, X. Jia, X. Wu, X. Xu, X. Zu, X. Zhou, X. Pan, Y. Charles, Y. Li, Y. Hu, Y. Liu, Y. Chen, Y. Wang, Y. Liu, Y. Qin, Y. Liu, Y. Yang, Y. Bao, Y. Du, Y. Wu, Y. Wang, Z. Zhou, Z. Wang, Z. Li, Z. Zhu, Z. Zhang, Z. Wang, Z. Yang, Z. Huang, Z. Huang, Z. Xu, Z. Yang, and Z. Lin (2025) Kimi k1.5: scaling reinforcement learning with llms. External Links: 2501.12599, Link Cited by: §2.1. B. Wang, X. Yue, Y. Su, and H. Sun (2025) Grokking of implicit reasoning in transformers: a mechanistic journey to the edge of generalization. In Proceedings of the 38th International Conference on Neural Information Processing Systems, NIPS ’24, Red Hook, NY, USA. External Links: ISBN 9798331314385, Link Cited by: §2.3. J. Wei, X. Wang, D. Schuurmans, M. Bosma, b. ichter, F. Xia, E. Chi, Q. V. Le, and D. Zhou (2022) Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 24824–24837. External Links: Link Cited by: §1, §2.1. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §1. X. Yue, X. Qu, G. Zhang, Y. Fu, W. Huang, H. Sun, Y. Su, and W. Chen (2023) MAmmoTH: building math generalist models through hybrid instruction tuning. External Links: 2309.05653, Link Cited by: §2.1. E. Zelikman, G. Harik, Y. Shao, V. Jayasiri, N. Haber, and N. D. Goodman (2024) Quiet-star: language models can teach themselves to think before speaking. External Links: 2403.09629, Link Cited by: §2.1. E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman (2022) STaR: self-taught reasoner bootstrapping reasoning with reasoning. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ’22, Red Hook, NY, USA. External Links: ISBN 9781713871088, Link Cited by: §2.1. B. Zeng, S. Song, S. Huang, Y. Wang, H. Li, Z. He, X. Wang, Z. Li, and Z. Lin (2025) Pretraining language models to ponder in continuous space. External Links: 2505.20674, Link Cited by: §1, §2.2. X. Zhu, J. Wang, L. Zhang, Y. Zhang, Y. Huang, R. Gan, J. Zhang, and Y. Yang (2023) Solving math word problems via cooperative reasoning induced language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N. Okazaki (Eds.), Toronto, Canada, p. 4471–4485. External Links: Link, Document Cited by: §2.1. Appendix A More Training Details A.1 Pseudocode Implementation Algorithm 1 provides the complete pseudocode for our primary training procedure. Algorithm 1 Forward Pass with Cross-Layer Latent Bridges 1: Input: Token sequence ∈ℕB×LX ^B× L, group size G, bridge set ℬ=(src,dest)B=\(src,dest)\, multiplier α 2: Initialize: KV cache ←∅K← , logits list ←[]Y←[\;] 3: Initialize: State buffers s,d←NULLS_s,d for all (s,d)∈ℬ(s,d) 4: for t=0,G,2G,…,L−Gt=0,G,2G,…,L-G do 5: // Extract current group of tokens 6: g←:,t:t+GX_g _:,\,t:t+G 7: (0)←EmbedTokens(g)H^(0) (X_g) 8: for i=0i=0 to N−1N-1 do 9: // 1. Standard Transformer Computation 10: (i+1),←TransformerLayeri((i),)H^(i+1),K _i(H^(i),K) 11: // 2. Injection Phase: Inject latent state from previous group 12: for (s,d)∈ℬ(s,d) do 13: if i=di=d and s,d≠NULLS_s,d then 14: (i+1)←(i+1)+α⋅UpProjs,d(s,d)H^(i+1) ^(i+1)+α·UpProj_s,d(S_s,d) 15: end if 16: end for 17: // 3. Extraction Phase: Save state for next group 18: for (s,d)∈ℬ(s,d) do 19: if i=si=s then 20: s,d←DownProjs,d((i+1))S_s,d _s,d(H^(i+1)) 21: end if 22: end for 23: end for 24: Append LMHead((N))LMHead(H^(N)) to Y 25: end for 26: Output: Concat(,dim=1)Concat(Y,dim=1) // Shape: [B,L,V][B,L,V] A.2 Hyperparameters We set the learning rate for Llama 3 models to 1.92×10−51.92× 10^-5. This was determined by scaling a base learning rate of 3×10−73× 10^-7 proportionally with our batch size of 64. For Qwen 3 models, the learning rate is set to 5.12×10−65.12× 10^-6, which is 8×10−8×648× 10^-8× 64. We set the decoding temperature to 1.0. A.3 Model Implementation We implement TurboConn using a cache. During the forward pass, as the model processes each token, the cache stores the hidden states from the specified higher layers. When processing subsequent tokens, these cached states are retrieved and added to the corresponding lower-layer inputs. The sequential nature of TurboConn necessitates the use of a key-value (KV) cache during the training forward pass, a mechanism typically reserved only for autoregressive inference. Because tokens (or groups of tokens) are processed sequentially, the attention keys and values from all previous tokens must be cached to be available for the current token’s self-attention computation. To implement this efficiently, we store past key and value states in a list of tensors rather than concatenating them. This approach avoids the memory duplication that can occur when new tensors are created through repeated concatenation. A.4 Example Training Data Here we offer some examples for the format of data used for training and evaluation. A.4.1 Parity Prompt: Question: Output the parity of this sequence. Input: 1 0 0 1 0 1 Answer: Completion: 1 A.4.2 Multi-step Arithmetic Prompt: Question: Evaluate this expression modulo 10. Input: (-(3 + -(1)) * 4 * (8 + 2 + 9)) = Answer: Completion: 8 A.4.3 GSM8K Prompt: Question: Answer the math question. Input: John cuts his grass to 2 inches. It grows .5 inches per month. When it gets to 4 inches he cuts it back down to 2 inches. It cost $100 to get his grass cut. How much does he pay per year? Answer: Completion: 300 A.5 Connection Configurations The specific wiring configurations were derived from a simple heuristic. Our principle was to ensure broad layer coverage while maintaining a low computational overhead: =(s,l)∣s−l>k, where s,l∈[Lmin,Lmax] sampled at interval mC=\(s,l) s-l>k, where s,l∈[L_min,L_max] sampled at interval m\ We did not conduct an extensive search for “optimal” wiring. The only decision based on experiment is whether to leave the bottom‑most layers unchanged. They are generally understood to encode basic features rather than the high-level logic targeted by our recurrence. When training appears unstable, we increase LminL_min from 0 to a value that feels right. This turns out to be sufficient for our method to work. The other hyperparameters of the connection formula were arbitrarily chosen at the beginning of our experiments. Our primary constraints were to ensure that the total number of trainable parameters did not exceed the standard baseline and that the connections covered the layers relatively evenly. Then these configurations were kept constant throughout all subsequent experiments. Because we conducted one or zero trials per model to determine these hyperparameters, we believe the method is robust to the specific wiring setup. The specific wiring configurations for different models are listed below. The downward connections are defined from a source layer to a destination layer using the notation: ‘source -> destination’. A.5.1 Llama 3.2 1B (16 total layers) A total of 15 connections were used, as specified below: 6 -> 0 8 -> 0 8 -> 2 10 -> 0 10 -> 2 10 -> 4 12 -> 0 12 -> 2 12 -> 4 12 -> 6 14 -> 0 14 -> 2 14 -> 4 14 -> 6 14 -> 8 A.5.2 Llama 3.1 8B (32 total layers) A total of 45 connections were used, as specified below: 14 -> 7 16 -> 7 16 -> 9 18 -> 7 18 -> 9 18 -> 11 20 -> 7 20 -> 9 20 -> 11 20 -> 13 22 -> 7 22 -> 9 22 -> 11 22 -> 13 22 -> 15 24 -> 7 24 -> 9 24 -> 11 24 -> 13 24 -> 15 24 -> 17 26 -> 7 26 -> 9 26 -> 11 26 -> 13 26 -> 15 26 -> 17 26 -> 19 28 -> 7 28 -> 9 28 -> 11 28 -> 13 28 -> 15 28 -> 17 28 -> 19 28 -> 21 30 -> 7 30 -> 9 30 -> 11 30 -> 13 30 -> 15 30 -> 17 30 -> 19 30 -> 21 30 -> 23 A.5.3 Qwen 3 1.7B (28 total layers) A total of 21 connections were used, as specified below: 12 -> 4 15 -> 4 15 -> 7 18 -> 4 18 -> 7 18 -> 10 21 -> 4 21 -> 7 21 -> 10 21 -> 13 24 -> 4 24 -> 7 24 -> 10 24 -> 13 24 -> 16 27 -> 4 27 -> 7 27 -> 10 27 -> 13 27 -> 16 27 -> 19 A.6 LoRA Setting For both 1B and 8B models, we use r=120r=120 for TurboConn and r=140r=140 for the original model. In this section, we show that with our setup, the number of trainable parameters available to TurboConn is less than or equal to that of the standard Transformer baseline. In both setups, LoRA was applied to several primary parameters within each Transformer block. There are the following key dimensions: • dhiddend_hidden: The hidden size of the model. • dkvd_kv: The dimension of the key/value heads. • dinterd_inter: The intermediate size of the MLP blocks. Table 6: Dimensions of the layers adapted with LoRA. Module d_in d_out Attention Block q_proj dhiddend_hidden dhiddend_hidden k_proj dhiddend_hidden dkvd_kv v_proj dhiddend_hidden dkvd_kv o_proj dhiddend_hidden dhiddend_hidden MLP Block gate_proj dhiddend_hidden dinterd_inter up_proj dhiddend_hidden dinterd_inter down_proj dinterd_inter dhiddend_hidden It is noted that all of these targeted modules in the base architecture are configured without bias terms. The matrices for those parameters are unbiased. Given a LoRA rank r, the total number of trainable parameters per Transformer block (PblockP_block) is calculated as follows: Pblock= P_block= 4rdhidden⏟for q_proj, o_proj+2r(dhidden+dkv)⏟for k_proj, v_proj+3r(dhidden+dinter)⏟for gate, up, down_proj 4rd_hidden_for q\_proj, o\_proj+ 2r(d_hidden+d_kv)_for k\_proj, v\_proj+ 3r(d_hidden+d_inter)_for gate, up, down\_proj The total number of trainable parameters is then derived from this block-level calculation. For our method, we add the parameters from the NconnN_conn downward connections. Each connection is a LoRA-adapted linear layer (dhidden→dhiddend_hidden→ d_hidden) with an added bias. The parameter calculation is: Pconnection=2rdhidden⏟LoRA matrices+r⏟bias+dhidden⏟biasP_connection= 2rd_hidden_LoRA matrices+ r_bias+ d_hidden_bias A.6.1 Llama 3.2 1B For the Llama 3.2 1B model, the key dimensions are: • dhidden=2048d_hidden=2048 • dkv=512d_kv=512 • dinter=8192d_inter=8192 The number of trainable LoRA parameters per Transformer block (PblockP_block), given a rank r, is calculated as: Pblock= P_block= 4r(2048)+2r(2048+512)+3r(2048+8192) 4r(048)+2r(048+12)+3r(048+192) = = 44032r 44032r Baseline Model (r=140r=140): Ptotal P_total =16×(44032×140) =6×(4032× 40) =98,631,680 =98,631,680 Our Method (r=120r=120): Pconn P_conn =Nconn×(2rdhidden+r+dhidden) =N_conn×(2rd_hidden+r+d_hidden) =15×(2⋅120⋅2048+120+2048) =5×(2· 20· 048+20+048) =7,405,320 =7,05,20 The total number of trainable parameters is therefore: Pours P_ours =(16×Pblock)+Pconn =(6× P_block)+P_conn =(16×44032×120)+7,405,320 =(6× 4032× 20)+7,05,20 =91,946,760 =91,946,760 A.6.2 Llama 3.1 8B For the Llama 3.1 8B model, the key dimensions are: • dhidden=4096d_hidden=4096 • dkv=1024d_kv=1024 • dinter=14336d_inter=14336 The number of trainable LoRA parameters per Transformer block (PblockP_block), given a rank r, is calculated as: Pblock= P_block= 4r(4096)+2r(4096+1024)+3r(4096+14336) 4r(096)+2r(096+024)+3r(096+4336) = = 81920r 81920r Baseline Model (r=140r=140): Ptotal P_total =32×(81920×140) =2×(1920× 40) =367,001,600 =367,001,600 Our Method (r=120r=120): Pconn P_conn =Nconn×(2rdhidden+r+dhidden) =N_conn×(2rd_hidden+r+d_hidden) =45×(2⋅120⋅4096+120+4096) =5×(2· 20· 096+20+096) =44,426,520 =4,26,20 The total number of trainable parameters is therefore: Pours P_ours =(32×Pblock)+Pconn =(2× P_block)+P_conn =(32×81920×120)+44,426,520 =(2× 1920× 20)+4,26,20 =358,999,320 =358,999,320 A.6.3 Qwen 3 1.7B For the Qwen 3 1.7B model, the key dimensions are: • dhidden=2048d_hidden=2048 • dkv=1024d_kv=1024 • dinter=6144d_inter=6144 The number of trainable LoRA parameters per Transformer block (PblockP_block), given a rank r, is calculated as: Pblock= P_block= 4r(2048)+2r(2048+1024)+3r(2048+6144) 4r(048)+2r(048+024)+3r(048+144) = = 38912r 38912r Baseline Model (r=140r=140): Ptotal P_total =28×(38912×140) =8×(8912× 40) =152,535,040 =152,535,040 Our Method (r=120r=120): Pconn P_conn =Nconn×(2rdhidden+r+dhidden) =N_conn×(2rd_hidden+r+d_hidden) =21×(2⋅120⋅2048+120+2048) =1×(2· 20· 048+20+048) =10,367,448 =0,67,48 The total number of trainable parameters is therefore: Pours P_ours =(28×Pblock)+Pconn =(8× P_block)+P_conn =(28×38912×120)+10,367,448 =(8× 8912× 20)+0,67,48 =141,111,768 =141,111,768 A.7 More Training Details for TurboConn + CoT Experiments Since the original NuminaMath-CoT dataset contains a test set of only 100100 problems, we perform a custom re-split of the combined training and test data to ensure more statistically stable evaluations. We randomly partitioned the dataset into 384,000384,000 training, 1,0001,000 validation, and 1,0001,000 testing examples. The fixed CoT was generated by a Llama 3.2-1B model fine-tuned for 33 epochs on the original training sets, using a learning rate of 6.4×10−46.4× 10^-4. For GSM8K, we maintain the same experimental setup as for NuminaMath-CoT. We use the augmented GSM8K dataset described in Deng et al. (2023). However, unlike NuminaMath-CoT, GSM8K exhibits more significant overfitting to the training data; therefore, we use only 1/6 of the data to train the CoT generation model for a single epoch. The training setup for the final prediction models remains identical to all previous experiments: we set the learning rate for Llama 3.2-1B models to 1.92×10−51.92× 10^-5, use a batch size of 64, and train for 3 epochs using all training data. In this particular experiment, we utilized a temperature of 0.010.01 for evaluation to sharpen the probability distribution. This encourages the model to be more ’certain’ in its predictions, often yielding higher expected accuracy by amplifying the likelihood of the dominant choice.