Paper deep dive
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
Chanhee Park, Sungbin Han, Jeongho Yoon, Seongtae Hong, Heuiseok Lim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/18/2026, 5:21:56 AM
Summary
The paper introduces Funnel of Thoughts (FoT), an inference-time method for Large Reasoning Models (LRMs) that reduces computational cost while maintaining accuracy. FoT utilizes a training-free lexical signal based on hesitation markers (e.g., 'Wait', 'Actually') to identify and prune unproductive reasoning trajectories early. By combining early voting for committed answers with rollout pruning for hesitant ones, FoT achieves a 56.1% reduction in attention FLOPs and a 37.6% reduction in wall time compared to standard Self-Consistency (SC@32), without sacrificing accuracy across multiple models and benchmarks.
Entities (7)
Relation Signals (6)
Funnel of Thoughts → preserves → Accuracy
confidence 95% · FoT preserves SC@32 accuracy on AIME24/25... at a 57% reduction in attention FLOPs
Funnel of Thoughts → reduces → attention FLOPs
confidence 95% · FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1%
Hesitation Markers → indicate → unproductive trajectories
confidence 90% · unproductive trajectories often reveal themselves through repeated hesitation markers such as 'Wait', 'Actually', and 'perhaps.'
Self-Consistency → isbenchmarkedagainst → Funnel of Thoughts
confidence 90% · we compare FoT against SC@32 and two efficient-SC baselines
Funnel of Thoughts → uses → Early Voting
confidence 90% · At fixed token-count checkpoints, FoT applies two mechanisms: early voting banks trajectories that have already committed to an answer
Funnel of Thoughts → uses → Rollout Pruning
confidence 90% · rollout pruning removes those exhibiting the highest hesitation marker density
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as "Wait", "Actually", and "perhaps." These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks.
Tags
Links
- Source: https://arxiv.org/abs/2608.15065v1
- Canonical: https://arxiv.org/abs/2608.15065v1
Trouble viewing inline? Open PDF directly →
Full Text
83,957 characters extracted from source content.
Expand or collapse full text
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning Chanhee Park Sungbin Han Jeongho Yoon Seongtae Hong Heuiseok Lim Thanks: Corresponding author. Affiliation: Department of Computer Science and Engineering, Korea University Email: pch7678,sungbinhan9039,a007878,ghdchlwls123,limhseok@korea.ac.kr Abstract Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment. Majority voting at k rollouts is the standard solution and the de facto accuracy target for this regime, but it is prohibitively expensive at the scale LRMs require. We introduce Funnel of Thoughts (FoT), an inference-time method that preserves the full 32-trajectory voted accuracy while halving its attention FLOPs, a 28.8% reduction in full-model inference cost. Across 115K reasoning trajectories from six LRMs, we find that unproductive trajectories often reveal themselves through repeated hesitation markers such as “Wait,” “Actually,” and “perhaps.” These trajectories are less likely to reach the correct answer and consume disproportionate attention FLOPs, degenerating into no-answer loops in the worst case. Built on this training-free lexical signal, FoT identifies the vocabulary that captures these pathological patterns and prunes affected trajectories before completion, reducing online generation attention FLOPs by 56.1% and wall time by 37.6% without any additional model inference; the same signal transfers without retuning across held-out architectures and out-of-domain tasks. 1 Introduction Large Reasoning Models (LRMs) (14; 7; 20) solve problems by generating long chains of thought (27). Compared with Instruct models (15), they explore richer answer trajectories and therefore produce diverse, sometimes inconsistent answers across repeated queries (3). This diversity makes multi-sample inference valuable: the correct answer often appears somewhere in the rollout pool, even when no single rollout is reliable. Self-Consistency (SC@k; 25) exploits this by sampling k rollouts and returning the majority vote, routinely recovering 20–40 percentage points over single-rollout accuracy on competition benchmarks (17). The cost is that every rollout must run to completion, and LRM rollouts are long, running to thousands of tokens each. Because attention cost grows with the square of the sequence length, every additional token is more expensive than the one before it, so the tail of a long trajectory, rather than its beginning, dominates what SC@k actually pays for. When some trajectories spiral into repetitive, counterproductive self-correction (26; 18; 10), SC pays for thousands of wasted tokens at exactly the point where tokens cost the most. Prior analyses likewise find that 40–50% of LRM tokens can be redundant (4). Figure 1: Accuracy vs. compute on AIME24/25. SC@k trades accuracy for compute by varying pool size k. FoT@32 defines a new Pareto point: matching SC@32 accuracy at SC@16-level compute. Figure 2: Funnel of Thoughts on AMC23 Problem 20 (DeepSeek-R1-Distill-Qwen-7B, k=32k=32). Top: Each bar is one rollout, ordered by sampling seed, and truncated at its pruning or early-voting checkpoint. Bottom-left: The pool narrows across three checkpoints. Bottom-right: SC@32 elects the wrong plurality (2); FoT@32 lifts the correct answer (9) into the plurality. To mitigate this inefficiency, we introduce Funnel of Thoughts (FoT), an inference-time algorithm that starts with a full pool of k parallel reasoning trajectories and progressively removes unproductive ones in-flight using only the generated text. Unlike efficient-sampling methods that draw fewer trajectories, FoT runs the full pool in parallel and truncates only its spiraling minority, shedding late-stage waste while keeping the productive trajectories that carry the vote. Our key observation is that hesitation markers, such as “Wait,” “Actually,” and “perhaps”, are significantly more frequent in incorrect trajectories. At fixed token-count checkpoints, FoT applies two mechanisms: early voting banks trajectories that have already committed to an answer; rollout pruning removes those exhibiting the highest hesitation marker density. As Figure 1 demonstrates, FoT preserves SC@32 accuracy on AIME24/25, hard competition-math benchmarks where multi-sample inference is most needed, at a 57% reduction in attention FLOPs. Figure 2 traces this on a representative problem, where SC@32’s plurality lands on a wrong answer and FoT recovers the correct one. Our contributions are: (1) a data-driven analysis showing that the density of hesitation markers, a zero-cost lexical signal, identifies the unproductive tail of LRM reasoning across 115K rollouts; (2) Funnel of Thoughts, an inference-time algorithm exploiting this signal to prune rollouts at inference time in parallel batches without training, reward models, or logit access; (3) generalizability: FoT transfers unchanged across architecturally diverse held-out models and out-of-domain tasks, preserving SC@32 accuracy at substantially reduced compute. 2 Related Work 2.1 Reasoning Inefficiency in LRMs A growing body of work documents systematic inefficiency in LRM reasoning trajectories. 4 find that about 48% of generated tokens on MATH500 are spent on redundant operations such as repetitive self-verification and unnecessary case exploration, while 18 observe a non-monotonic relationship between trajectory length and accuracy in which both very short and very long rollouts underperform. 10 show that self-correction without external feedback is counterproductive on average. These findings motivate our approach: the self-correction attempts that manifest as hesitation markers are, on aggregate, wasteful. Building on this diagnosis, a line of work intervenes at the single-trajectory level to suppress unproductive reasoning. TIP (26) penalizes excessive thought-switching in logits, 22 suppress self-reflection keywords to reduce trajectory length, and 28 use hidden-state probes to dynamically terminate thinking. These approaches focus on shaping the best single rollout. But single-rollout interventions cannot mitigate the fundamental property that LRM inference produces a diverse distribution of final answers, demonstrated by the gap between pass@1 and pass@k (3). This diversity grows with trajectory length: every additional decoding step compounds the probability of divergence across rollouts, and LRM trajectories run thousands of tokens. 2.2 Self-Consistency and Sampling Efficiency SC (25) samples k reasoning paths and selects the majority vote, yielding large gains but requiring k full-length trajectories. A line of work reduces cost by deciding when to stop sampling: Adaptive Consistency (2) applies a Beta stopping rule on the running majority vote, Certaindex (5) uses a learned stability metric, and others refine the criterion further (21; 24). These methods generate rollouts sequentially: once a trajectory begins, it must run to completion, even when it has entered the wasteful late-stage generation that dominates LRM compute. They act on the sample axis and FoT on the token axis, so the two are complementary; we compare against Adaptive Consistency in our main results and against Difficulty-Adaptive Self-Consistency and Certaindex in Appendix B.6. Parallel alternatives operate within a k-rollout pool but pursue different goals. Slim-SC (9) prunes redundant rollouts during generation via inter-trajectory embedding similarity, requiring a separate embedding model. CISC (19) reweights votes after all rollouts complete using token-level confidence, requiring logprob access and not reducing generation cost. Neither directly targets the pathological trajectories that drive late-stage waste, whereas FoT preemptively identifies and prunes them at inference time using only the generated text. 2.3 Test-Time Compute Scaling More broadly, test-time compute scaling can be as effective as model scaling: optimal allocation lets smaller models match much larger ones (17; 29), and the gap between pass@1 and pass@k on hard problems exceeds 50p (3), motivating the multi-sample methods. Process Reward Models offer strong rollout selection in this regime but require a separately trained verifier (12; 23). The closest prior to our lexical signal is budget forcing (13), which appends Wait tokens to extend reasoning; we use the density of the same token, and others like it, as a signal that extended reasoning has become unproductive. 3 Funnel of Thoughts 3.1 Preliminaries Pass@k. Pass@k is the probability that at least one of k rollouts is correct. It represents the upper bound of what the model can achieve given its knowledge, but is not attainable in practice without an oracle: selecting the correct rollout from the pool requires knowing the answer a priori. Reaching Pass@k accuracy at low compute cost is therefore the key challenge for effective multi-sampling; SC@k approximates this ceiling via the majority-vote estimator. Self-Consistency (SC@k). Given a problem q, SC@k (25) independently samples k rollouts c1,…,ck\c_1,…,c_k\ from a language model, extracts a candidate answer aia_i from each rollout, and returns the majority answer11 1 We follow standard SC literature convention in using “majority” to denote the most-voted answer regardless of whether it strictly exceeds 50% of votes; on our hardest problems the top answer is often a plurality rather than a strict majority.: a^=argmaxa∑i=1k[ai=a]. a= _a _i=1^k1[a_i=a]. (1) Cost metrics. Each rollout cic_i has total sequence length sis_i tokens. Because self-attention is quadratic, we report compute cost primarily as total attention FLOPs across all rollouts: FLOPs=∑i=1k4⋅L⋅d⋅si2FLOPs= _i=1^k4· L· d· s_i^2, where L is the number of layers and d the model dimension. Attention is the term that grows with trajectory length, and therefore the term that in-flight pruning acts on, which is why we measure it directly. Including feed-forward and projection compute, FoT’s saving is 28.8%, and its end-to-end wall-clock saving on a live server is 37.6%; we report these in Section 4.4 and Appendix D.2, and every FLOP figure in this paper states which of the two accountings it uses. Setup. We analyze and evaluate FoT on 6 LRMs: DeepSeek-R1-Distill-Qwen-1.5B and 7B (7), OpenThinker3-7B (6), Qwen3-4B-Thinking-2507, Qwen3-30B-A3B-Thinking-2507, and QwQ-32B (30; 20). The benchmark suite is four competition-math evaluations spanning difficulty levels: AIME24, AIME25, AMC23, and MATH500 12. For each problem we generate 32 independent rollouts with unique random seeds, yielding a 115,200-rollout pool over 3,600 model-problem pairs and nearly 0.8 billion generated tokens. All rollouts are pre-generated; FoT and baselines operate on the same pool. Answers are graded with math_verify; an earlier draft used exact string match, which depressed only MATH500 absolute values (Appendix D.1). Regrading moves MATH500 absolute accuracies substantially and can shift an individual cell’s FoT–SC gap by a problem or two; the aggregate finding that FoT matches SC@32 is unchanged. Correct Incorrect No Answer Proportion of pool 90.5% 7.4% 2.1% Avg. tokens generated 5,575 16,128 31,319 Hesitation density (/1K chars) 3.46 6.14 6.93 Relative attention FLOPs 1.0× 6.0× 17.4× Table 1: Behavioral profile of the 115,200-rollout pool across 6 models and 4 benchmarks. Correct rollouts commit early, carry fewer hesitation markers, and cost a fraction of the compute; the 9.5% of rollouts that end wrong or never commit consume a disproportionate share of it. 3.2 Motivating Observations Across the 115,200-rollout pool, Table 1 summarizes the behavioral profile; three observations motivate our method. Observation 1: Early commitment predicts correctness. Correct rollouts produce a final answer ( ) after an average of 5,575 tokens, while incorrect rollouts run to 16,128. A further 2.1% of rollouts never commit at all, consuming the maximum budget. Rollouts reaching an answer early have higher chance of being correct, motivating an early voting mechanism that preserves committed answers before pruning. This suggests that for problems the model can solve quickly, additional reasoning may not improve and can even degrade the answer (4; 10). Observation 2: Late-stage generation is disproportionately wasteful. Rollouts that fail to commit early enter self-correction spirals (10; 26). Because attention compute is quadratic in length, these long trajectories are also the expensive ones: an incorrect rollout costs 6.0 times the attention FLOPs of a correct one, and a rollout that never commits costs 17.4 times, despite the two together making up under a tenth of the pool. This asymmetry motivates targeting the wasteful tail via rollout pruning rather than uniformly reducing all rollouts. Observation 3: Hesitation markers signal unproductive reasoning. Hesitation markers—“Wait,”, “actually”, “perhaps”, and 18 others—are denser in incorrect rollouts across all 6 LRMs. We measure hesitation marker density as marker count per 1,000 characters, computed cumulatively over the generated text. Sorting rollouts into density deciles yields a monotonic accuracy decline (r=−0.82r=-0.82 across deciles, rpb=−0.30r_pb=-0.30 per rollout; 31p gap between D1 and D10). Density is the online-available leading indicator of runaway length: it forecasts a rollout’s final length from the visible prefix (r=+0.13r=+0.13 to +0.41+0.41 across models) before that length has been paid for, marking the unproductive tail FoT should prune. We report the details of this process in Appendix A. 3.3 Algorithm Algorithm 1 Funnel of Thoughts (FoT) 1: Rollouts c1,…,ck\c_1,…,c_k\, checkpoints t1,…,tm\t_1,…,t_m\, keep ratio ρ 2: ←1,…,kA←\1,…,k\ ⊳ active set 3: ℬ←∅B← ⊳ vote bank 4: for j←1j← 1 to m do ⊳ loop over checkpoints 5: if ||≤2|A|≤ 2 then 6: break 7: for each i∈i do ⊳ Early Voting 8: if HasAnswer(ci[1:tj]) HasAnswer(c_i[1:t_j]) then 9: ℬ←ℬ∪Extract(ci[1:tj])B ∪\ Extract(c_i[1:t_j])\ 10: ←∖iA \i\ 11: if ||≤2|A|≤ 2 then 12: continue 13: di←HesitationDensity(ci[1:tj])d_i← HesitationDensity(c_i[1:t_j]) for all i∈i ⊳ Rollout Pruning 14: n←max(2,⌊||⋅ρ⌋)n← (2,\; |A|·ρ ) 15: ←A← indices of n lowest-density rollouts 16: Complete generation for all i∈i 17: return Plurality(ℬ∪Extract(ci):i∈) Plurality(B∪\ Extract(c_i):i \) As illustrated in Algorithm 1, Funnel of Thoughts retains the core SC@k framework, sampling k rollouts and taking the majority vote, but eliminates unproductive rollouts to save compute. The method maintains an active set A of rollouts still being generated and a vote bank ℬB of answers from rollouts that have already committed. At each of m fixed token-count checkpoints t1,t2,…,tm\t_1,t_2,…,t_m\, two mechanisms are applied. After the final checkpoint, the answer is selected by Plurality (argmax over answer counts) across ℬB and the surviving rollouts in A. The only hyperparameters are the checkpoint schedule and keep ratio ρ, selected via a grid sweep on the pool and held constant across all problems, models, and checkpoints. Detailed results can be found in Appendix C.2. Early Voting. At checkpoint tjt_j, we inspect the text generated so far for each active rollout cic_i. Any rollout that has already produced a final answer ( ) is moved to the vote bank ℬB: its answer is preserved for the final vote, and the rollout is removed from the active set. This is what makes the method a funnel rather than a truncation. Without banking, preserving k votes would require running k trajectories to completion, so narrowing the pool would mean discarding votes; banking separates the two, letting the pool shrink while the ballot stays full. Across our pool, only 3.1 of 32 trajectories are still generating at the final checkpoint, yet the vote is cast by 19.7 of them, where the same funnel without banking would leave 8.4. Banking also confines the pruning signal to the rollouts it applies to, since a committed rollout’s outcome is already settled and cannot be predicted by a measure of future spiraling. Early voting requires only substring matching, with no model inference. AIME24 AIME25 AMC23 MATH500 Overall (/600) Model Method Acc FLOP Acc FLOP Acc FLOP Acc FLOP Acc FLOP DS-R1-1.5B SC@32 56.7 — 33.3 — 90.0 — 92.4 — 87.5 — AC 56.7 10.4% 33.7 12.7% 90.0 39.4% 92.2 43.3% 87.3 26.4% SlimSC 63.3 56.3% 40.0 68.7% 92.5 60.3% 91.8 54.5% 87.8 60.0% FoT@32 60.0 57.6% 33.3 54.3% 90.0 55.9% 92.6 50.3% 87.8 54.5% DS-R1-7B SC@32 80.0 — 53.3 — 95.0 — 95.4 — 92.5 — AC 80.0 31.3% 53.3 28.0% 95.0 60.8% 95.4 56.5% 92.5 44.1% SlimSC 83.3 53.8% 53.3 64.9% 97.5 41.0% 95.2 43.5% 92.7 50.8% FoT@32 80.0 54.8% 56.7 56.3% 100.0 46.6% 95.4 38.7% 93.0 49.1% OT3-7B SC@32 76.7 — 73.3 — 97.5 — 97.2 — 95.0 — AC 76.7 55.5% 73.3 49.9% 97.5 67.3% 97.2 71.7% 95.0 61.1% SlimSC 73.3 77.0% 50.0 81.6% 97.5 69.5% 95.8 58.0% 92.5 71.5% FoT@32 80.0 60.3% 76.7 60.8% 100.0 53.8% 97.2 43.6% 95.5 54.6% Qwen3-4B SC@32 86.7 — 86.7 — 100.0 — 98.4 — 97.3 — AC 86.7 57.9% 86.7 48.4% 100.0 86.9% 98.4 73.3% 97.3 66.6% SlimSC 76.7 65.7% 80.0 61.9% 95.0 43.7% 98.0 51.7% 95.8 55.8% FoT@32 86.7 58.0% 80.0 59.4% 100.0 45.3% 98.4 43.8% 97.0 51.6% Qwen3-30B SC@32 93.3 — 86.7 — 100.0 — 98.0 — 97.3 — AC 93.3 72.2% 88.3 56.8% 100.0 87.3% 97.9 71.7% 97.3 72.0% SlimSC 93.3 42.1% 83.3 49.9% 100.0 29.4% 98.0 35.8% 97.2 39.3% FoT@32 93.3 54.5% 83.3 59.0% 100.0 38.4% 98.2 37.7% 97.3 47.4% QwQ-32B SC@32 90.0 — 83.3 — 100.0 — 97.4 — 96.5 — AC 90.0 61.2% 83.3 53.8% 100.0 84.7% 97.4 74.5% 96.5 68.5% SlimSC 83.3 69.6% 63.3 75.5% 100.0 50.4% 97.2 45.7% 95.0 60.3% FoT@32 86.7 54.0% 83.3 58.1% 100.0 42.8% 97.2 33.5% 96.2 47.1% Avg. SC@32 80.6 — 69.4 — 97.1 — 96.5 — 94.4 — AC 80.6 48.1% 69.8 41.6% 97.1 71.1% 96.4 65.2% 94.3 56.5% SlimSC 78.9 60.8% 61.7 67.1% 97.1 49.1% 96.0 48.2% 93.5 56.3% FoT@32 81.1 56.5% 68.9 58.0% 98.3 47.1% 96.5 41.3% 94.5 50.7% Table 2: Main results across 6 models and 4 competition-math benchmarks (k=32k=32). Acc is accuracy (%) and FLOP the attention-FLOP saving over SC@32, with Overall FLOP taken as the per-benchmark mean so that MATH500’s size does not dominate. Accuracy cells are shaded relative to SC@32 (higher, lower, unshaded within ±1± 1p) and FLOP cells grey-shaded by size of saving. Rollout Pruning. Among the remaining active rollouts, we compute hesitation marker density did_i over the entire text generated from rollout start to checkpoint tjt_j. Rollouts are ranked by density in ascending order; we retain the ⌊||⋅ρ⌋ |A|·ρ rollouts with the lowest hesitation rates and discard the rest, where ρ∈(0,1]ρ∈(0,1] is a fixed keep ratio. We always retain at least 2 active rollouts. Online Deployment. FoT integrates directly into parallel inference pipelines. At each checkpoint, the inference server pauses active rollouts, applies early voting and pruning, and resumes survivors. Terminated rollouts release their KV cache immediately, reducing memory pressure and freeing batch slots for the remaining rollouts. Because the pruning decision requires only the generated text, with no separate embedding model, reward model, or logprob access, FoT adds negligible overhead to the serving loop. We validate this integration in Section 4.4. 4 Experiments Using the same 6-model, 4-benchmark setup from Section 3, we compare FoT against SC@32 and two efficient-SC baselines: Adaptive Consistency (AC; 2) and Slim-SC (9). We then test generalization to held-out models, cross-domain benchmarks, and online deployment, and ablate the contribution of each FoT mechanism. Sampling parameters, baseline configurations, and the full hyperparameter sweep are in Appendix D. 4.1 Main Results FoT preserves SC@32 accuracy at half the attention FLOPs. Across the 4 math benchmarks, FoT@32 preserves the accuracy of full self-consistency while reducing attention FLOPs by roughly half. FoT@32 and SC@32 return the same correctness outcome on 3,580 of 3,600 paired model-problem instances, differing on only 20 cases: 12 FoT wins and 8 FoT losses.22 2 The difference is not significant under McNemar’s exact test (p=0.50p=0.50), with a paired-bootstrap 95% CI on the accuracy difference of [−0.14,+0.36][-0.14,+0.36]p. As Table 2 demonstrates, FoT holds its FLOP saving in nearly every cell with accuracy preserved, making it a stable operating point on the accuracy–compute tradeoff. It starts from the same 32-rollout pool as SC@32, banks rollouts that have already committed to an answer, and terminates only the active trajectories whose generated text indicates unproductive hesitation. The savings therefore come from removing late-stage waste, not from giving up the answer diversity that makes multi-sample inference useful. Hard regime Easy regime AIME24+25 (gap 22.3p) AMC+MATH500 (gap 5.3p) Method Acc FLOP↓ Acc FLOP↓ Pass@1 (single rollout) 60.7 — 93.8 — Pass@32 (oracle ceiling) 83.1 — 99.1 — SC@32 (baseline) 75.0 — 96.5 — AC 75.2 44.9% 96.6 65.6% SlimSC 70.3 (−-4.7) 64.0% 96.1 48.3% FoT@32 (Ours) 75.0 57.3% 96.6 41.7% Table 3: Each method split by pass@k gap. FoT preserves SC@32 accuracy while saving more attention FLOPs in the hard split, where consensus is slow; accuracy is in %, and FLOP savings are problem-weighted macro averages relative to SC@32. AC struggles when rollouts diverge. Adaptive Consistency reduces cost by deciding when enough completed rollouts have been sampled. This is effective when the model reaches answer consensus quickly, but the same condition also means SC@32 was over-provisioned for that benchmark. The pass@1–pass@32 gap reveals the difficulty split as a factor: when the gap is large, the correct answer is often present somewhere in the rollout pool, but completed rollouts do not agree quickly enough for sample-axis early stopping. Table 3 shows that AC is strongest on easier cells with fast consensus, whereas FoT continues to save compute in the hard split because its decision is made within each rollout. Compared with AC, FoT is more robust on hard samples where consensus is slow but individual trajectories still reveal pruning signals. The two axes trade places with difficulty: sample-axis savings shrink as consensus slows, while FoT’s grow. Because the axes are independent, the methods compose rather than compete, which we report in Appendix B.6. Slim-SC over-prunes when faced with hard samples. Slim-SC exposes a different failure mode: pruning by inter-trajectory similarity can discard useful diversity. This weakness is most visible on AIME25, where Slim-SC saves 67.1% FLOPs but drops the average accuracy from 69.4 to 61.7 as shown in Table 2. On hard problems, correct solutions often share surface structure even when they provide independently useful evidence for the final vote, so treating similar thought segments as redundant can collapse the pool before enough support accumulates. FoT targets a different property: its hesitation-density signal asks whether a trajectory is spiraling, not whether it resembles another trajectory. Compared with Slim-SC, FoT better preserves the reasoning diversity that gives multi-sample inference its advantage. Figure 3: FoT removes compute, not votes. Left: attention FLOPs actually consumed, as a percentage of SC@32’s total, so a pruned rollout is charged for everything it generated before termination. Right: votes cast in the final plurality. FoT prunes the failure mode it was designed to catch. Figure 3 separates what pruning costs from what it saves. Repetitive-wrong and no-answer trajectories together consume 42% of SC@32’s attention FLOPs while casting 6% of its votes, and FoT removes about three fifths of that compute. Correct trajectories keep 87% of their votes, and the pool still casts 84% of the original ballot at half the compute, because banking preserves the vote of a rollout that has stopped generating. A checkpoint-level diagnostic gives the same picture for the non-committing tail: Table 12 shows that FoT removes 71% of it by the final checkpoint, with the largest reductions in cells where the issue is most acute. Hesitation-density pruning therefore does not reshape the answer distribution uniformly; it preferentially removes the long repetitive loops (26; 18) that consume disproportionate attention cost while contributing few votes. 4.2 Cross-Model Generalization FoT transfers across model architectures because its pruning rule is relative rather than model-calibrated. We apply the same token set, checkpoints, and keep ratio to four held-out models: Qwen3-14B, Qwen3-32B, Skywork-OR1-7B (8), and Phi-4-reasoning (1). These models differ in their absolute hesitation-marker density (Table 10), but FoT only compares rollouts within the same active pool at the same checkpoint. A model can therefore be more or less verbose overall; what matters is whether a trajectory is unusually hesitation-heavy relative to its peers. Table 4 shows that this criterion preserves SC@32 accuracy with comparable FLOP savings across the held-out models. Phi-4-reasoning is the most architecturally distinct case and also shows the largest transfer gain, further supporting the interpretation. We interpret this gain as arising from Phi-4-reasoning’s heavy no-answer tail, reported in Table 11, which FoT can capture through the same relative pruning rule. Benchmark SC@32 FoT@32 Δ (p) Token↓ FLOP↓ Qwen3-14B AIME24 83.3 83.3 ++0.0 32.5% 53.6% AIME25 80.0 80.0 ++0.0 38.8% 59.6% AMC23 100.0 100.0 ++0.0 18.7% 40.1% MATH500 70.4 70.6 ++0.2 12.7% 35.6% Qwen3-32B AIME24 86.7 83.3 −-3.3 32.3% 55.1% AIME25 80.0 80.0 ++0.0 38.3% 60.1% AMC23 97.5 97.5 ++0.0 18.7% 41.0% MATH500 71.0 71.2 ++0.2 12.2% 35.2% Skywork-OR1-7B AIME24 80.0 80.0 ++0.0 36.1% 56.6% AIME25 60.0 60.0 ++0.0 39.9% 60.5% AMC23 95.0 95.0 ++0.0 25.7% 53.1% MATH500 70.8 70.8 ++0.0 18.9% 48.6% Phi-4-reasoning AIME24 50.0 56.7 ++6.7 35.9% 58.0% AIME25 36.7 53.3 ++16.7 38.9% 61.5% AMC23 55.0 67.5 ++12.5 30.4% 54.4% MATH500 46.6 47.0 ++0.4 39.6% 57.3% Aggregate 1600→ 1615 of 2400 ++0.6 29.4% 51.9% Table 4: Cross-model transfer with unchanged FoT configuration: FoT preserves SC@32 accuracy on four held-out models while saving 51.9% of attention FLOPs on average. Accuracies use exact-match grading, which depresses only the MATH500 absolutes (Appendix D.1); the deltas are unaffected. 4.3 Cross-Domain Generalization We next ask whether the same signal survives when the answer format changes. GPQA-Diamond (16) changes the knowledge domain and final-answer format while retaining natural-language deliberation; LCB-Lite (11) is a stronger stress test because the final artifact is executable code and two correct solutions may be syntactically unrelated. If FoT were exploiting math-specific answer syntax or majority-vote structure, we should not expect it to help under execution-based scoring. Following standard practice for code generation, we evaluate LCB by pass@k over the rollout pool, reported alongside the mean per-rollout pass rate. Early voting also needs a notion of commitment in each domain: on GPQA a rollout commits when it emits a final option letter, and on LCB when it closes a code block after its reasoning. The pruning signal itself is unchanged, since it reads only the deliberation text and never the answer. Table 5 shows that the unchanged configuration transfers to both settings, and on LCB FoT improves mean per-rollout pass rate while reducing FLOPs; pass@k declines modestly with −2.8-2.8p on average, the expected cost of any method that shrinks the oracle pool. The implication is that hesitation density acts before answer extraction: it identifies low-value continuations in the reasoning process, whether the final output is a boxed number, a multiple-choice option, or code. pass@k (LCB only) Model SC@32 FoT@32 Δ (p) FLOP↓ SC FoT Δ (p) GPQA-Diamond (accuracy %) DS-R1-1.5B 38.9 41.4 ++2.5 42.2% — — — DS-R1-7B 55.6 55.6 ++0.0 31.8% — — — OT3-7B 51.5 53.0 ++1.5 56.8% — — — Qwen3-4B 65.7 67.7 ++2.0 35.3% — — — Qwen3-30B 72.2 72.2 ++0.0 38.7% — — — QwQ-32B 63.6 65.2 ++1.5 41.4% — — — Avg. 57.9 59.2 ++1.3 41.0% — — — LCB-Lite (mean pass rate % / pass@k %) DS-R1-1.5B 15.10 16.29 ++1.19 56.6% 35.4 34.0 −-1.5 DS-R1-7B 36.59 39.49 ++2.90 50.3% 61.6 59.0 −-2.6 OT3-7B 43.63 44.60 ++0.97 62.2% 64.9 61.9 −-3.0 Qwen3-4B 49.23 51.78 ++2.55 53.5% 76.9 72.4 −-4.5 Qwen3-30B 69.36 69.38 ++0.03 60.0% 83.2 81.7 −-1.5 QwQ-32B 61.30 62.15 ++0.85 57.9% 79.1 75.4 −-3.7 Avg. 45.87 47.28 ++1.42 56.8% 66.9 64.1 −-2.8 Table 5: Cross-domain transfer with unchanged FoT configuration. GPQA reports accuracy; LCB reports mean per-rollout pass rate and, in the right block, pass@k (a problem counts as solved if any rollout in the pool—for FoT, any surviving rollout—passes all tests). Δ is FoT−-SC in percentage points. 4.4 Online Deployment We deploy FoT in an online inference pipeline using SGLang (31) on a single A100 GPU, generating all rollouts end-to-end for 30 AIME24 pro blems with DeepSeek-R1-Distill-Qwen-7B. Table 6 shows that FoT@32 solves 24 of 30 problems versus 23 for SC@32, while reducing wall time by 37.6% and attention FLOPs by 56.1%. The same logs show that the cumulative generated-token footprint, and hence KV-cache occupancy, drops from 12.88M to 8.51M tokens. The wall-time savings follow directly from pruning: terminated rollouts release KV cache memory and batch slots, reducing both compute and memory pressure for the remaining rollouts. Because the hesitation marker count is a string operation with negligible overhead, FoT requires no additional model inference beyond the rollouts themselves. Pass@1 SC@32 FoT@32 vs. SC@32 Accuracy (%) 53.3 76.7 80.0 ++3.3 p Wall time (s) 4,145 21,561 13,445 −-37.6% Attn. FLOPs (×1010× 10^10) 0.35 12.9 5.67 −-56.1% Cumulative KV (M) 0.37 12.88 8.51 −-33.9% Table 6: Online deployment on AIME24 with DS-R1-7B (k=32k=32), measured end-to-end on a single A100. Cumulative KV counts generated-token positions rather than peak memory; pruned and banked rollouts release their slots at each checkpoint. 4.5 Ablations We ablate two properties of FoT in the main text: the pool size at which it becomes effective, and the contribution of each mechanism. Kernel-size robustness and additional sensitivity analyses are deferred to Appendix C. Pool-size scaling. FoT is a large-pool method. Sweeping k from 8 to 32 shows that pruning is mildly harmful when the pool is small, becomes neutral around k=16k=16, and turns positive beyond it: a small pool cannot spare the vote diversity that pruning removes, whereas a large one can (Appendix C.1). This sets the regime in which FoT should be used, as a replacement for large-pool self-consistency rather than a small-k method. w/o Early Voting w/ Early Voting Pruning Signal Correct FLOP↓ Correct FLOP↓ None (SC@32) 94.36 (baseline) — Random 93.94 (−-0.42) 48.0% 94.38 (++0.01) 48.4% Hesitation density 94.61 (++0.25) 51.7% 94.47 (++0.11) 48.7% Table 7: Pruning signal × early voting on the 4-benchmark pool (3,600 model-problem pairs), with Δ relative to the SC@32 baseline of 94.36; random rows are means over 10 seeds, with a pooled standard deviation of ± 0.14 without early voting and ± 0.06 with it. Early voting slightly lowers the FLOP saving because banked rollouts are exempt from the pruning quota, so the shrunken active pool ends pruning earlier. Mechanism ablation. Table 7 reports a two-by-two ablation on the same 4-benchmark pool as the main result: random vs. hesitation-density pruning, each with and without early voting, plus SC@32 as the no-pruning baseline. The result is direct: the lexical pruning signal, not early voting, explains the accuracy. Random pruning loses accuracy at similar compute savings, while hesitation-density pruning preserves SC@32 accuracy even without banking committed answers. Early voting does a different job, and accuracy at this operating point is the wrong place to look for it. Banking is what keeps the ballot full as the pool narrows, holding the vote at 19.7 of 32 trajectories when only 3.1 are still generating. Its effect on accuracy is neutral here because banking can only change an outcome when pruning would otherwise delete a committed correct answer, and at the mild default keep ratio that is rare. It becomes visible once pruning is sharpened: on the hard cells it recovers 3.6p at a keep ratio of one third, while retaining most of the compute saving (See Appendix C.6). Early voting is therefore what makes the keep ratio safe to expose as a deployment knob, not a source of gain at the default. Full sweep results, hyperparameter sensitivity, per-model calibration, kernel-size robustness, and the per-difficulty analysis are reported in Appendix C. 5 Discussion Our findings suggest that predicting per-rollout correctness from hesitation density is fragile, since the same markers appear in both correct and incorrect reasoning and only moderately separate them in aggregate, with rpb=−0.30r_pb=-0.30. The reliable signal is distributional: error rate rises monotonically with density, with a decile-level r=−0.83r=-0.83, so the high-density tail is disproportionately wrong. This is why FoT works by removing that tail from the answer distribution, pruning the votes least likely to be right and the compute least likely to pay off. It needs no confidence model, reward model, or logit access, and matches SC@32 at roughly half the attention FLOPs, a 28.8% reduction in full-model FLOPs. The key reason this replacement transfers is that FoT does not rely on a model-specific threshold for how often a word such as wait or actually should appear. Absolute marker frequency varies across LRMs, but failed long-form reasoning often exposes a shared lexical pattern: repeated revision, backtracking, and non-commitment. FoT uses this pattern comparatively, ranking active rollouts within the same model, problem, and checkpoint. As a result, the signal can survive architectural changes and domain shifts: a Qwen-family model, Phi-4-reasoning, a graduate-level science question, and a code-generation task need not use identical wording for FoT to identify the unusually unproductive tail of the current pool. This also clarifies why FoT differs from the efficient-SC baselines. Adaptive Consistency is strongest when the answer distribution stabilizes quickly, and Slim-SC is strongest when similar trajectories are genuinely redundant. Those are precisely the cases where multiple sampling contributes less. In the harder regime, consensus may be slow and superficially similar derivations may still carry useful voting evidence. FoT is designed for this regime: it preserves the pool-level diversity that makes SC effective, while pruning within-trajectory waste before it dominates the compute budget. This is why its advantage is most persuasive on hard cells, held-out models, and domain shifts, where robustness matters more than early consensus. 6 Limitations Our evaluation is centered on reasoning tasks with extractable final answers. We include cross-domain tests on GPQA-Diamond and LCB-Lite, but these still have relatively well-defined answer or execution-based evaluation. Open-ended generation, multi-turn interaction, and tasks where correctness cannot be reduced to answer extraction remain outside the scope of this study. Our calibration pool is also unbalanced by construction: MATH500 supplies 3,000 of the 3,600 model-problem pairs, so any statistic pooled over the whole suite is weighted toward easy problems on which pruning has little to remove and little to risk. We report the hard split separately throughout for this reason, and the hard-split numbers are the ones to read when judging the mechanism rather than the aggregate. A pool built with a different difficulty mix would move the pooled figures, though it should not affect the within-pool relative comparison that FoT actually uses. FoT also relies on a commitment detector for early voting. In our math setting this is the convention, while other domains require task-specific answer extraction or prompting. The pruning signal itself is independent of the final-answer format, but the vote bank is only as reliable as the commitment detector. Finally, FoT uses a fixed checkpoint schedule and an English hesitation-marker kernel. The relative-density rule reduces sensitivity to model-level verbosity, but models reasoning in other languages or with substantially different deliberation styles may require revalidating the marker set. Our FLOP accounting assumes standard full attention; under efficient-attention mechanisms the saving scales down toward the token-reduction floor (from 48.7% to 23.0% in the linear limit), which we quantify across window sizes in Appendix D.2. Pruning still releases KV cache and batch slots regardless of the attention regime. Adaptive checkpointing based on per-problem difficulty is another natural extension that could further improve the efficiency–accuracy tradeoff. References Abdin et al. (2025) M. Abdin, S. Agarwal, A. Awadallah, V. Balachandran, H. Behl, L. Chen, G. de Rosa, S. Gunasekar, M. Javaheripi, N. Joshi, P. Kauffmann, Y. Lara, C. C. T. Mendes, A. Mitra, B. Nushi, D. Papailiopoulos, O. Saarikivi, S. Shah, V. Shrivastava, V. Vineet, Y. Wu, S. Yousefi, and G. Zheng Phi-4-reasoning technical report. External Links: 2504.21318, Link Cited by: §4.2. Aggarwal et al. (2023) P. Aggarwal, A. Madaan, Y. Yang, and Mausam Let’s sample step by step: adaptive-consistency for efficient reasoning and coding with LLMs. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 12375–12396. External Links: Link, Document Cited by: §D.3, §2.2, §4. Brown et al. (2024) B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Ré, and A. Mirhoseini Large language monkeys: scaling inference compute with repeated sampling. External Links: 2407.21787, Link Cited by: §1, §2.1, §2.3. Chen et al. (2025) X. Chen, J. Xu, T. Liang, Z. He, J. Pang, D. Yu, L. Song, Q. Liu, M. Zhou, Z. Zhang, R. Wang, Z. Tu, H. Mi, and D. Yu Do not think that much for 2+3=? on the overthinking of o1-like llms. External Links: 2412.21187, Link Cited by: §1, §2.1, §3.2. Fu et al. (2025) Y. Fu, J. Chen, S. Zhu, Z. Fu, Z. Dai, Y. Zhuang, Y. Ma, A. Qiao, T. Rosing, I. Stoica, and H. Zhang Efficiently scaling llm reasoning with certaindex. External Links: 2412.20993, Link Cited by: Table 15, §2.2. Guha et al. (2025) E. Guha, R. Marten, S. Keh, N. Raoof, G. Smyrnis, H. Bansal, M. Nezhurina, J. Mercat, T. Vu, Z. Sprague, A. Suvarna, B. Feuer, L. Chen, Z. Khan, E. Frankel, S. Grover, C. Choi, N. Muennighoff, S. Su, W. Zhao, J. Yang, S. Pimpalgaonkar, K. Sharma, C. C. Ji, Y. Deng, S. Pratt, V. Ramanujan, J. Saad-Falcon, J. Li, A. Dave, A. Albalak, K. Arora, B. Wulfe, C. Hegde, G. Durrett, S. Oh, M. Bansal, S. Gabriel, A. Grover, K. Chang, V. Shankar, A. Gokaslan, M. A. Merrill, T. Hashimoto, Y. Choi, J. Jitsev, R. Heckel, M. Sathiamoorthy, A. G. Dimakis, and L. Schmidt OpenThoughts: data recipes for reasoning models. External Links: 2506.04178, Link Cited by: §3.1. Guo et al. (2025) D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Ding, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Chen, J. Yuan, J. Tu, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. You, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Zhou, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), p. 633–638. External Links: ISSN 1476-4687, Link, Document Cited by: §1, §3.1. He et al. (2025) J. He, J. Liu, C. Y. Liu, R. Yan, C. Wang, P. Cheng, X. Zhang, F. Zhang, J. Xu, W. Shen, S. Li, L. Zeng, T. Wei, C. Cheng, Y. Liu, and Y. Zhou Skywork open reasoner series. Note: https://capricious-hydrogen-41c.notion.site/Skywork-Open-Reaonser-Series-1d0bc9ae823a80459b46c149e4f51680Notion Blog Cited by: §4.2. Hong et al. (2025) C. Hong, X. Guo, A. C. Singh, E. Choukse, and D. Ustiugov Slim-sc: thought pruning for efficient scaling with self-consistency. External Links: 2509.13990, Link Cited by: §D.4, §2.2, §4. Huang et al. (2024) J. Huang, X. Chen, S. Mishra, H. S. Zheng, A. W. Yu, X. Song, and D. Zhou Large language models cannot self-correct reasoning yet. External Links: 2310.01798, Link Cited by: §1, §2.1, §3.2, §3.2. Jain et al. (2024) N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica LiveCodeBench: holistic and contamination free evaluation of large language models for code. External Links: 2403.07974, Link Cited by: §4.3. Lightman et al. (2023) H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe Let’s verify step by step. External Links: 2305.20050, Link Cited by: §B.4, §2.3, §3.1. Muennighoff et al. (2025) N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei-Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Candès, and T. Hashimoto S1: simple test-time scaling. External Links: 2501.19393, Link Cited by: §2.3. OpenAI et al. (2024) OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li OpenAI o1 system card. External Links: 2412.16720, Link Cited by: §1. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. External Links: 2203.02155, Link Cited by: §1. Rein et al. (2023) D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman GPQA: a graduate-level google-proof q&a benchmark. External Links: 2311.12022, Link Cited by: §4.3. Snell et al. (2024) C. Snell, J. Lee, K. Xu, and A. Kumar Scaling llm test-time compute optimally can be more effective than scaling model parameters. External Links: 2408.03314, Link Cited by: §1, §2.3. Su et al. (2025) J. Su, J. Healey, P. Nakov, and C. Cardie Between underthinking and overthinking: an empirical study of reasoning length and correctness in llms. External Links: 2505.00127, Link Cited by: §1, §2.1, §4.1. Taubenfeld et al. (2025) A. Taubenfeld, T. Sheffer, E. Ofek, A. Feder, A. Goldstein, Z. Gekhman, and G. Yona Confidence improves self-consistency in llms. In Findings of the Association for Computational Linguistics: ACL 2025, p. 20090–20111. External Links: Link, Document Cited by: §2.2. Team (2025) Q. Team QwQ-32b: embracing the power of reinforcement learning. External Links: Link Cited by: §1, §3.1. Wan et al. (2025) G. Wan, Y. Wu, J. Chen, and S. Li Reasoning aware self-consistency: leveraging reasoning paths for efficient LLM sampling. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 3613–3635. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §2.2. Wang et al. (2025a) C. Wang, Y. Feng, D. Chen, Z. Chu, R. Krishna, and T. Zhou Wait, we don’t need to "wait"! removing thinking tokens improves reasoning efficiency. External Links: 2506.08343, Link Cited by: §2.1. Wang et al. (2024) P. Wang, L. Li, Z. Shao, R. X. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui Math-shepherd: verify and reinforce llms step-by-step without human annotations. External Links: 2312.08935, Link Cited by: §2.3. Wang et al. (2025b) X. Wang, S. Feng, Y. Li, P. Yuan, Y. Zhang, C. Tan, B. Pan, Y. Hu, and K. Li Make every penny count: difficulty-adaptive self-consistency for cost-efficient reasoning. In Findings of the Association for Computational Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 6919–6932. External Links: Link, Document, ISBN 979-8-89176-195-7 Cited by: Table 15, §2.2. Wang et al. (2023) X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou Self-consistency improves chain of thought reasoning in language models. External Links: 2203.11171, Link Cited by: §1, §2.2, §3.1. Wang et al. (2025c) Y. Wang, Q. Liu, J. Xu, T. Liang, X. Chen, Z. He, L. Song, D. Yu, J. Li, Z. Zhang, R. Wang, Z. Tu, H. Mi, and D. Yu Thoughts are all over the place: on the underthinking of o1-like llms. External Links: 2501.18585, Link Cited by: §1, §2.1, §3.2, §4.1. Wei et al. (2023) J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou Chain-of-thought prompting elicits reasoning in large language models. External Links: 2201.11903, Link Cited by: §1. Wu et al. (2025a) M. Wu, C. Zhou, S. Bates, and T. Jaakkola Thought calibration: efficient and confident test-time scaling. External Links: 2505.18404, Link Cited by: §2.1. Wu et al. (2025b) Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang Inference scaling laws: an empirical analysis of compute-optimal inference for problem-solving with language models. External Links: 2408.00724, Link Cited by: §2.3. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu Qwen3 technical report. External Links: 2505.09388, Link Cited by: §3.1. Zheng et al. (2024) L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng SGLang: efficient execution of structured language model programs. External Links: 2312.07104, Link Cited by: §D.1, §4.4. Appendix A Motivating Analysis Table 8: Operational 21-marker kernel, ranked by benchmark-weighted point-biserial correlation with correctness on the 115,200-rollout pool (neff=32,797n_eff=32,797 stratified). Prev. is rollout prevalence and W/C the wrong-to-correct frequency ratio; †marks a marker essentially absent from correct rollouts, where the ratio is unstable. Marker rpbwr_pb^w pwp^w Prev. (%) W/C perhaps −-0.427 <<1e-300 42.7 3.56× Wait, −-0.286 <<1e-300 97.3 1.83× Wait −-0.276 <<1e-300 97.7 1.79× actually −-0.210 <<1e-300 60.4 1.75× wrong −-0.142 <<1e-300 23.4 2.32× However −-0.125 <<1e-300 18.0 1.89× Or −-0.122 <<1e-300 12.9 2.36× no, −-0.108 <<1e-300 58.5 1.80× Alternatively −-0.094 <<1e-300 64.3 1.28× But wait −-0.090 <<1e-300 43.3 1.47× Actually −-0.072 <<1e-300 2.7 2.06× Maybe −-0.063 <<1e-300 66.2 1.13× incorrect −-0.059 <<1e-300 10.2 1.74× instead −-0.057 <<1e-300 30.7 1.08× I was wrong −-0.045 2e-16 0.7 3.20× Let me think −-0.032 7e-9 43.1 0.93× Let me reconsider −-0.024 1e-5 0.0 n/a† I think I made −-0.023 3e-5 2.2 2.11× Re-examining −-0.017 0.002 0.0 n/a† This isn’t −-0.016 0.003 0.0 n/a† Let me try −-0.015 0.008 43.4 1.08× A.1 Hesitation-Marker Validation Table 8 shows the 21 hesitation markers used by FoT. We select markers whose density has a significant negative benchmark-weighted point-biserial correlation with rollout correctness (weighted r<0r<0, weighted p<0.05p<0.05); weighting each benchmark equally prevents MATH500’s 96,000 rollouts from dominating the pooled statistic, and all 21 markers satisfy the criterion. The resulting kernel spans common revision markers such as Wait, and actually, as well as rarer but highly skewed phrases such as Let me reconsider. We exclude markers that correlate positively with correctness, such as mistake, Oh, and Hmm, because they more often indicate productive self-correction. The aggregate density of the 21 retained markers yields rpb=−0.30r_pb=-0.30, making the composite signal stronger than most individual markers. Table 9 gives the decile-level diagnostic behind Observation 3. Accuracy declines from 75.7% in the lowest-density decile to 50.7% in the highest-density decile. The signal is not intended to identify the single correct rollout; rather, it separates the unproductive high-density tail that FoT should prune. Density and eventual length. Decomposing the two signals per rollout on AIME24, final rollout length correlates with correctness at −0.60-0.60 to −0.76-0.76 across models and hesitation density at the 8K checkpoint at −0.05-0.05 to −0.38-0.38, while the partial correlation of density with correctness given final length is small and model-dependent (−0.24-0.24 to +0.10+0.10). Density instead forecasts eventual length from the visible prefix (r=+0.13r=+0.13 to +0.41+0.41), which is the sense in which it is a leading indicator of runaway generation (See Observation 3). It therefore carries little information beyond eventual length; its value is that it is available online. An oracle that prunes by final length reaches the same accuracy as FoT (72.9 on the exact-match lineage of Table 7) at a 61.8% attention saving against FoT’s 50.7%, but a rollout’s final length is unknown until it has been paid for, and at a fixed token-count checkpoint all active rollouts have the same length, so “prune the longest” is not even defined online. Decile Density range Accuracy (%) D1 0.00–1.11 96.9 D2 1.11–1.76 98.0 D3 1.76–2.30 97.0 D4 2.30–2.80 95.9 D5 2.80–3.31 95.1 D6 3.31–3.87 94.3 D7 3.87–4.57 93.0 D8 4.57–5.48 89.1 D9 5.48–6.75 80.1 D10 >>6.75 65.6 Table 9: Accuracy by hesitation-marker density decile across 115,200 rollouts, where density is marker count per 1K characters. Accuracy falls monotonically with density (r=−0.82r=-0.82 between decile index and accuracy), for a 31.4p gap between the lowest and highest decile. Model All Correct Wrong Wrong/Correct DS-R1-1.5B 4.05 3.12 7.23 2.32 DS-R1-7B 3.37 2.93 6.66 2.27 OT3-7B 5.11 4.98 6.56 1.32 Qwen3-4B 3.13 3.10 3.59 1.16 Qwen3-30B 2.44 2.43 2.55 1.05 QwQ-32B 4.30 4.17 6.73 1.62 Table 10: Per-model hesitation-marker density, in markers per 1K characters, where wrong counts every rollout not graded correct. Absolute density varies several-fold across models, but every model separates wrong from correct in the same direction, which is what lets one relative rule transfer. A.2 Marker-Selection Validation A potential concern is that the marker list may overfit to the lexical styles of the models and benchmarks used for calibration. We address this primarily through the held-out model and cross-domain experiments in Section 4.2 and Section 4.3, where the same 21-marker kernel is applied without retuning. Table 10 reports the absolute marker density by model. The scale varies substantially, but FoT does not apply a fixed density threshold across models; it ranks rollouts within the same active pool at the same checkpoint. This relative rule is why a shared marker kernel can transfer across models with different verbosity levels. Benchmark Pass@1 SC@32 Pass@32 No-ans./32 AIME24 15.7 50.0 73.3 19.0 AIME25 13.6 36.7 66.7 17.2 AMC23 16.8 55.0 87.5 18.2 MATH500 11.5 46.6 59.6 21.0 Table 11: Phi-4-reasoning rollout profile at k=32k=32. Accuracy columns are percentages under exact-match grading (Appendix D.1); no-answer reports the mean number of rollouts per problem that never produce an extractable final answer. Phi-4-reasoning diagnostic. Table 11 reports the rollout profile behind the Phi-4 cross-model result in Section 4.2. Phi-4 has both a large pass@32–SC@32 gap and an unusually high number of no-answer rollouts, making it a useful diagnostic for whether FoT can remove non-committing trajectories without retuning. Subset Start Ckpt 1 Ckpt 2 Final AIME/AMC avg. 2.67 1.74 1.18 0.79 Qwen3-4B/AIME25 9.4 – – 2.4 OT3-7B/AIME25 5.4 – – 1.7 Table 12: No-answer rollouts remaining under FoT over the checkpoint sequence on the AIME24/AIME25/AMC23 stress subset. Entries report the mean count per 32-rollout pool; the aggregate removes 71% of no-answer rollouts by the final checkpoint. No-answer tail diagnostic. Table 12 reports the checkpoint-level analysis used to interpret Figure 3. FoT does not identify every failed rollout independently; instead, it progressively removes the non-committing tail from the active pool, with the largest reductions in the cells where no-answer behavior is most acute. Accuracy (%) Model Dataset Pass@1 Pass@32 Gap1→32 SC@32 DS-R1-1.5B AIME24 29.2 80.0 50.8 56.7 AIME25 24.3 53.3 29.1 33.3 AMC23 71.2 100.0 28.8 90.0 MATH500 83.8 98.0 14.2 92.4 DS-R1-7B AIME24 54.9 83.3 28.4 80.0 AIME25 39.1 66.7 27.6 53.3 AMC23 90.5 100.0 9.5 95.0 MATH500 93.2 99.0 5.8 95.4 OT3-7B AIME24 63.6 90.0 26.4 76.7 AIME25 60.7 80.0 19.3 73.3 AMC23 93.1 100.0 6.9 97.5 MATH500 95.2 99.0 3.8 97.2 Qwen3-4B AIME24 70.6 90.0 19.4 86.7 AIME25 68.8 86.7 17.9 86.7 AMC23 99.0 100.0 1.0 100.0 MATH500 97.4 99.4 2.0 98.4 Qwen3-30B AIME24 86.8 93.3 6.6 93.3 AIME25 80.9 90.0 9.1 86.7 AMC23 99.7 100.0 0.3 100.0 MATH500 97.4 99.6 2.2 98.0 QwQ-32B AIME24 80.2 93.3 13.1 90.0 AIME25 69.7 90.0 20.3 83.3 AMC23 98.2 100.0 1.8 100.0 MATH500 96.8 99.0 2.2 97.4 Average 76.8 91.3 14.4 85.9 Table 13: Pass@1 vs. Pass@32 accuracy gap across 6 models and 4 benchmarks (k=32k=32), where Pass@1 is the average single-rollout accuracy and Pass@32 the oracle accuracy of the pool. The 14.4p average gap is the headroom majority voting exploits; averages are unweighted means over the 24 model–benchmark cells. A.3 Pass@1 vs. Pass@32 Performance Gap Table 13 reports the per-rollout accuracy, oracle accuracy, and SC@32 majority-vote accuracy for all 6 models across 4 benchmarks. The average pass@1–pass@32 gap is 14.4 percentage points, confirming that individual rollouts are highly stochastic and that correct answers often appear somewhere in the 32-rollout pool. The gap is largest for smaller models on harder benchmarks and smallest for stronger models near ceiling. Appendix B Additional Experiments This section collects secondary diagnostics behind the main results: per-model behavior, the cases where FoT and SC@32 disagree, benchmark and per-problem difficulty effects, and a unified-basis comparison with sample-axis baselines. B.1 Per-Model Accuracy Breakdown Figure 4 breaks down accuracy by model on the 4-benchmark math suite. The pattern mirrors the main table: FoT is most helpful when the base model still has substantial sampling variance, while near-ceiling models leave less room for improvement. SlimSC is less stable on the harder AIME split, consistent with the main-text observation that similarity pruning can remove useful vote diversity. Figure 4: Per-model accuracy breakdown across the 4-benchmark math suite. B.2 Discordant Instance Analysis Figure 5 examines cases where FoT and SC@32 produce different correctness outcomes on the AIME/AMC subset. FoT wins more often than it loses, and the losses cluster on harder problems with low pass rates. This supports the main-text interpretation that FoT usually preserves the SC answer distribution, while occasionally changing the final vote after removing non-committing or repetitive trajectories. Figure 5: Discordant cases on the AIME/AMC subset. Left: SC@32–FoT@32 concordance. Right: FoT wins and losses by model. B.3 Vote-Concentration Diagnostic Figure 6 asks when FoT changes the final SC decision. We bin each model-problem pair by the fraction of rollouts assigned to the SC@32 plurality answer. Four pairs with no extractable SC answer are excluded from this diagnostic, leaving 3,596 pairs. The pattern is consistent with FoT’s intended role. When SC@32 has a clear majority, FoT almost always preserves it: in the >50%>50\% bin, the final answer changes in only 0.2% of cases. In contrast, low-consensus pools are both more diverse and less reliable: the <25%<25\% bin averages 15.2 distinct answers, and FoT changes the final answer in 49.1% of cases. This is also where most net corrections arise (+4+4). The effect is modest, but it explains why FoT can slightly improve aggregate accuracy while primarily serving as a compute-reduction method: it leaves strong SC trends intact and acts mainly on the divergent tail. Figure 6: Vote-concentration diagnostic on the 4-benchmark pool. Bins are SC@32 plurality-share ranges. Left: how often FoT changes the SC answer. Right: FoT correct minus SC correct within each bin. B.4 Difficulty Adaptation MATH500 (12) sits at the easier end of our math suite. Table 2 shows that FoT maintains accuracy parity on MATH500 while saving less compute than on AIME/AMC. This is expected: easier problems produce shorter rollouts with less hesitation tail to prune, so many trajectories commit early instead of entering the pruning regime. The result clarifies that FoT’s savings track late-stage reasoning waste rather than applying a fixed reduction uniformly across benchmarks. B.5 Difficulty-Stratified Accuracy Table 14 bins every model-problem pair by per-problem pass@1, the fraction of its 32 rollouts that are correct. FoT sits 4.8–5.8p above SC@32 in the low-pass@1 bins (0<p1≤0.250<p_1≤0.25), where the pass@1–pass@32 gap is largest and the vote is least concentrated. The p1=0p_1=0 bin is unreachable by any prune-only method, since no correct rollout exists in the pool to preserve, and the saturated p1>0.5p_1>0.5 bin is at ceiling. This is the same picture as the vote-concentration diagnostic of Section B.3, resolved by difficulty rather than by plurality share. All (4 benchmarks) Hard (AIME24/25) pass@1 bin n SC FoT Δ n SC FoT Δ p1=0p_1=0 91 0.0 0.0 ++0.0 61 0.0 0.0 ++0.0 0<p1≤0.250<p_1≤0.25 121 23.1 28.9 ++5.8 42 35.7 40.5 ++4.8 0.25<p1≤0.50.25<p_1≤0.5 91 79.1 80.2 ++1.1 33 93.9 97.0 ++3.0 p1>0.5p_1>0.5 3297 100.0 99.9 −-0.1 224 100.0 99.1 −-0.9 Table 14: Accuracy (%) stratified by per-problem difficulty, graded with math_verify. Bins are per-problem pass@1, the fraction of the 32 rollouts that are correct; n counts model-problem pairs. B.6 Sample-Axis Baselines on a Unified FLOP Basis FoT prunes along the token axis, whereas adaptive-consistency methods act on the sample axis by deciding how many rollouts to draw. Table 15 places both on one basis: math_verify grading and full-model (attention, FFN, and projection) FLOP savings against SC@32. At equal pooled accuracy DSC saves more than FoT overall (69.4% vs. 28.8%), and our offline Certaindex proxy also holds SC accuracy at an 11.5–13.2% saving, though the proxy is only indicative: it can fire later than the original online probe, understating savings, while paying no probe cost, overstating them. The two axes move in opposite directions with difficulty. DSC’s saving falls from 75.0% on MATH500 to 54.4% on AIME24/25 as consensus slows, while FoT’s rises from 21.9% to 44.0% because hard problems carry longer unproductive tails. Because the axes are independent, the methods compose: FoT+DSC reaches 80.9% (all) and 75.7% (hard) full-model savings, at accuracy costs of −0.4-0.4 and −1.4-1.4p. Per-query latency and KV-cache relief remain specific to FoT, as a sample-axis method cannot shorten a rollout it has already drawn. All Hard Method Acc FLOP↓ Acc FLOP↓ SC@32 94.4 — 75.0 — DSC 94.4 69.4% 75.0 54.4% Certaindex (offline proxy) 94.0 11.5–13.2% 75.0 9.9–10.9% FoT@32 94.5 28.8% 75.0 44.0% FoT+DSC (composed) 94.0 80.9% 73.6 75.7% Table 15: Sample-axis baselines and composition on a unified basis: math_verify grading, with FLOP↓ the full-model saving against SC@32 and hard = AIME24/25. DSC (24) and Certaindex (5) act on the sample axis, FoT on the token axis. Appendix C Additional Ablations Figure 7: Hyperparameter sensitivity analysis. (a) Full 340-configuration sweep on the 4-benchmark calibration pool. Stars mark the default joint-Pareto config and the math-only in-domain maximizer. (b) Pool-size scaling for the default config. C.1 Hyperparameter Sensitivity To validate our hyperparameter choices, we conduct a comprehensive sweep over 340 configurations spanning checkpoint count (1–5), checkpoint placement, and keep ratio (ρ∈0.50,0.60,0.67,0.75,0.85ρ∈\0.50,0.60,0.67,0.75,0.85\) on the unified 4-benchmark calibration pool (6 models × 600 problems × 32 rollouts == 115,200 rollouts, with inverse-frequency stratified weighting so MATH500’s larger problem count does not dominate the smaller benchmarks). Figure 7(a) plots all configurations on accuracy vs. FLOP savings axes, color-coded by checkpoint count. Two natural Pareto candidates emerge. The math-only winner (4K,14K\4K,14K\, ρ=0.60ρ=0.60) maximizes accuracy on the math calibration pool (Δ=+13 =+13 / 3,600 problems, +0.36+0.36p) at a 47.9% FLOP saving. The joint winner (6K,10K,14K\6K,10K,14K\, ρ=0.67ρ=0.67; green star), which we use throughout the paper, gives up 9 of those math problems while saving marginally more compute (48.7%), in exchange for substantially larger gains on out-of-distribution evaluations: +15+15 problems on GPQA-Diamond (+1.26+1.26p on 1,188 model-problem pairs) and +14+14 problems on the cross-architecture Phi-4-reasoning evaluation (+2.33+2.33p on 600 model-problem pairs spanning the same four math benchmarks). Why not the math-only winner? The math-only schedule is the best choice if the sole objective is maximizing the four-benchmark math aggregate. We instead use the joint configuration because the paper’s claim is model- and domain-level transfer: it gives up 9 math problems but recovers larger gains on the held-out Phi-4 and GPQA evaluations. It ranks 71st of the 340 configurations on the math pool, so 58 configurations beat it in domain; this makes the default a deliberately conservative operating point rather than the in-domain maximizer. The out-of-distribution figures quoted above are graded by exact match, the lineage of the held-out evaluations in Table 4, while the math-pool figures use math_verify. On GPQA that choice is nearly immaterial: the two graders assign different SC@32 outcomes on 8 of the 1,188 model-problem pairs, every one of them a single model emitting X in place of the bare option letter. Pool-size behavior. Figure 7(b) shows the expected tradeoff for the default schedule: pruning is too aggressive when the initial pool is small, but the gap closes as k grows and becomes slightly positive at k≥24k≥24. This supports using FoT primarily as a replacement for large-pool SC rather than as a small-k method. C.2 Checkpoint and Threshold Sensitivity We sweep four checkpoint schedules (Early: 4K,8K,12K\4K,8K,12K\, Mid: 6K,10K,14K\6K,10K,14K\, Math-only: 6K,11K,16K\6K,11K,16K\, Late: 8K,12K,16K\8K,12K,16K\) crossed with three keep ratios (ρ∈0.50,0.67,0.75ρ∈\0.50,0.67,0.75\) on the unified 4-benchmark pool (3,600 model-problem pairs per cell). Table 16 shows the results. Keep ratio ρ Schedule 0.50 0.67 0.75 Early 4K, 8K, 12K 3380 (−-17) 3398 (++1) 3400 (++3) Mid 6K, 10K, 14K 3389 (−-8) 3401 (++4) 3403 (++6) Math-only 6K, 11K, 16K 3399 (++2) 3404 (++7) 3399 (++2) Late 8K, 12K, 16K 3395 (−-2) 3401 (++4) 3402 (++5) Table 16: Ablation over checkpoint schedules and keep ratio ρ on the 4-benchmark pool (3,600 pairs per cell), showing FoT correct count and Δ against the SC@32 baseline of 3397/3600. The shaded universal default (Mid 6K, 10K, 14K, ρ=0.67ρ=0.67) is our joint choice; within this grid Math-only 6K, 11K, 16K scores highest in domain but loses out of distribution (Appendix C.1). Two patterns emerge. First, the Early schedule is consistently the worst, hurting accuracy at every keep ratio (Δ≤−2 ≤-2, down to −15-15 at ρ=0.50ρ=0.50): early checkpoints prune before the hesitation marker signal has accumulated sufficiently, removing rollouts that would have converged correctly given more tokens. Second, the Mid, Math-only, and Late schedules all produce viable configurations (Δ≥0 ≥ 0) at ρ∈0.67,0.75ρ∈\0.67,0.75\, with the math-only winner reaching Δ=+9 =+9 on math but giving up the larger out-of-distribution gains discussed in Section C.1. The universal Mid schedule with ρ=0.67ρ=0.67, used throughout the paper, is the joint Pareto choice balancing math accuracy (Δ=+4 =+4) against the larger OOD gains. C.3 Per-Model Calibration All main experiments use a single universal config (Mid schedule 6K,10K,14K\6K,10K,14K\, ρ=0.67ρ=0.67). We ask: does per-model hyperparameter tuning improve results? For each model, we sweep 12 configurations (4 schedules × 3 keep ratios) on that model’s 600 problems across the four math benchmarks independently. Per-Model Universal Model Best Config Δ FLOP↓ Δ FLOP↓ DS-R1-1.5B Math-only, ρ=0.50ρ=0.50 ++7 63.2 ++2 52.7 DS-R1-7B Math-only, ρ=0.50ρ=0.50 ++4 55.5 ++3 46.5 OT3-7B Mid, ρ=0.67ρ=0.67 ++3 51.2 ++3 51.2 Qwen3-4B Mid, ρ=0.75ρ=0.75 −-1 39.6 −-2 49.1 Qwen3-30B Math-only, ρ=0.67ρ=0.67 ++2 43.0 ++0 45.3 QwQ-32B Math-only, ρ=0.50ρ=0.50 ++0 53.5 −-2 43.9 Agg. ++15 51.0 ++4 48.1 Table 17: Per-model calibration against the universal Mid ρ=0.67ρ=0.67 config, with each model’s best config selected on its own 600-problem math pool. Δ is relative to SC@32 and Agg. sums across all 3,600 model-problem pairs. Table 17 shows that per-model calibration yields a small aggregate improvement: Δ=+17 =+17 vs. universal +4+4 on 3,600 model-problem pairs (a marginal +13+13 problems, +0.36+0.36p). The bulk of this gain concentrates on the weakest model in the suite (DS-R1-1.5B picks Math-only ρ=0.50ρ=0.50 for Δ=+10 =+10 vs. universal +3+3); the remaining five models gain ≤2≤ 2 problems each from per-model tuning. Three of six models select the Math-only schedule 6K,11K,16K\6K,11K,16K\, matching the in-domain tradeoff described above, but the preferred schedules are not stable enough to justify model-specific tuning. We therefore keep the universal config as the default. Target Model (Δ vs. SC@32 on 600 problems) Source Config DS-R1-1.5B DS-R1-7B OT3-7B Qwen3-4B Qwen3-30B QwQ-32B Agg. Mid, ρ=0.67ρ=0.67 (Ours) ++2 ++3 ++3 −-2 ++0 −-2 ++4 Math-only, ρ=0.50ρ=0.50 ++7 ++4 −-2 −-6 −-1 ++0 ++2 Math-only, ρ=0.67ρ=0.67 ++4 ++4 ++1 −-3 ++2 −-1 ++7 Mid, ρ=0.75ρ=0.75 ++3 ++2 ++2 −-1 ++1 −-1 ++6 Table 18: Cross-model config transfer on the 4-benchmark pool, where cells show Δ against SC@32 out of 600 problems per model and Agg. sums across all 3,600 pairs. Bold marks each model’s self-optimal config from Table 17. C.4 Per-Model Config Transfer To test whether per-model configs transfer, we apply each of the five distinct optimal configs from Section C.3 to all 6 models. Table 18 shows that the math-only Pareto winner, Math-only 6K,11K,16K\6K,11K,16K\ with ρ=0.67ρ=0.67, achieves the highest cross-model math aggregate (Δ=+9 =+9 on 3,600), but this is the same in-domain tradeoff described in Section C.1. Our universal Mid config (ρ=0.67ρ=0.67, Δ=+4 =+4 on math) is the joint Pareto choice that balances math performance against out-of-distribution generalization. No per-model-tuned configuration dominates the universal config across both math and OOD evaluations simultaneously. Kernel Δ Tok↓ (%) FLOP↓ (%) top-1 (perhaps) −3-3 29.12 50.59 top-2 +6+6 29.20 50.66 top-3 (3 markers) +5+5 29.22 50.70 top-5 +5+5 29.19 50.65 top-7 +6+6 29.17 50.59 top-10 +0+0 29.17 50.61 top-15 +3+3 29.24 50.76 top-21 (default) +4+4 29.23 50.73 Table 19: Kernel-size sweep on the 4-benchmark pool. Tokens are added by descending |rpb||r_pb|; Δ is FoT@32 minus SC@32. C.5 Kernel-Size Robustness We sweep the hesitation-marker kernel from 1 to 21 markers, adding markers in descending order of |rpb||r_pb| (Table 8). All other hyperparameters are held at the paper default (6K,10K,14K\6K,10K,14K\, ρ=0.67ρ=0.67, hesitation marker density). Table 19 reports the full sweep on the unified 4-benchmark pool (3,600 model-problem pairs). Two findings. First, accuracy is essentially flat across kernel sizes: Δ ranges from −3-3 to +6+6 problems on 3,600, with all 8 settings within 0.17p of one another and well within single-rollout sampling noise. Even the top-1 kernel (perhaps alone) preserves SC@32 accuracy to within 3 problems. Second, FLOP saving is constant to two significant figures (50.59–50.76%), confirming that the pruning behavior is driven by aggregate hesitation density rather than the identity of any individual marker. The three highest-|r||r| markers (perhaps, Wait,, Wait) match the full 21-marker kernel within sampling noise (Δ=+5 =+5 vs. +4+4). This robustness is itself the substantive finding: the lexical signal is concentrated in a small number of canonical English hesitation markers, and the method does not depend on the specific composition of the kernel. We retain the 21-token kernel as the calibrated default for stability, but smaller kernels (down to top-3) are an equivalent operating point. C.6 Early Voting Analysis Figure 8 analyzes the early voting mechanism on the AIME/AMC subset under the current FoT configuration. On average, 14.7 of 32 rollouts are banked before pruning begins. The bank contains a correct answer in 78% of all instances, (71% restricted to FoT wins, 25% on SC wins). Early commitment is therefore a strong correctness signal, yet the 4-benchmark ablation in Table 7 shows the aggregate gain coming primarily from hesitation-density pruning. These two facts are consistent, and the reconciliation defines early voting’s role. The neutrality in Table 7 is structural rather than evidence of a redundant mechanism. The pooled ablation is dominated by MATH500 (3,000 of 3,600 model-problem pairs), where problems are easy, commitment is fast, and the mild default keep ratio ρ=0.67ρ=0.67 almost never endangers a committed answer; banking answers that pruning would not have touched changes nothing. Early voting’s contribution concentrates exactly where this condition fails: hard problems under aggressive pruning. Table 20 sweeps the keep ratio on the hard cells (AIME24/25, 6 models, n=360n=360; SC@32 == 75.0): Accuracy FLOP↓ keep ρ w/ EV w/o EV EV Δ w/ EV w/o EV 0.67 (default)∗ 75.0 75.0 ++0.0 57.3% 59.2% 0.50 73.9 73.3 ++0.6 72.0% 73.3% 0.33 71.1 67.5 +3.6+3.6 80.5% 81.7% 0.25 70.8 67.2 +3.6+3.6 81.7% 82.8% Table 20: Early voting under keep-ratio sweep on the hard cells (attention-FLOP convention of Table 2). At the default ρ early voting is neutral; at aggressive ρ it recovers committed correct answers that pruning would otherwise delete. ∗The default row quotes the main-result lineage (Table 3); the sweep recomputation differs by one problem on one model due to re-tokenization, within single-problem resolution. At mild pruning almost no committed answer is ever at risk, so early voting is neutral. As pruning sharpens on hard problems it begins deleting committed correct answers, and early voting recovers them: without it, pushing to ρ=0.33ρ=0.33 for ∼82% 82\% attention-FLOP savings costs 7.5p (75.0→67.5); with it, the same aggressive setting holds 71.1 at 80.5% savings. Early voting is thus a cheap, always-on safety margin that caps the cost of pruning errors, operationalizing Observation 1: once a rollout has produced its outcome is resolved, so ranking it with a predictor of future spiraling would apply the signal outside its domain of validity. This safeguard is what makes the keep ratio safe to expose as a deployment knob rather than a fixed constant. Figure 8: Early voting diagnostic on the AIME/AMC subset. Left: banked rollout count per problem. Center: fraction with a correct answer banked. Right: average banked and active rollouts per model. Appendix D Implementation Details D.1 Generation Details All rollouts are generated using SGLang (31) with the following parameters: temperature T=0.6T=0.6, top-p=0.95p=0.95, top-k=30k=30, and a maximum generation budget of 32,768 tokens. Each model uses a rolling random seed starting from 42 (i.e., seed =42+rep_id=42+rep\_id) to ensure reproducibility while maintaining independence across rollouts. Offline rollouts are collected across multiple GPU configurations depending on model size. For the online deployment experiment in Section 4.4, all three conditions (Pass@1, SC@32, FoT@32) are run on a single isolated A100 GPU with a fixed random seed of 2000, ensuring that the same rollout pool is generated across conditions for a controlled comparison. Offline simulation and answer handling. Five implementation details affect exact reproduction. (1) The offline experiments approximate token-count checkpoints by per-rollout linear interpolation over characters; this can move a hesitation marker or a commitment across a checkpoint boundary relative to a true tokenizer cut, whereas online deployment (§4.4) cuts on real token counts. (2) Answer extraction (extract_boxed) falls back to multiple-choice-letter and “Final Answer:” patterns when no is present; these fallbacks can also fire on math outputs. (3) Overlapping markers double-count by design: text matching Wait, also increments the count for Wait. (4) Plurality ties in the final vote are broken by insertion order. (5) On MATH500, 38% of gold answers are non-integer LaTeX expressions, which is why only MATH500 absolute accuracies were depressed under the earlier exact-string-match grader while the integer-answer AIME/AMC benchmarks were unaffected; all main-result accuracies in this version are graded with math_verify; appendix analyses that retain the earlier exact-match grading are labeled where they appear. Checkpoint synchronization overhead. FoT requires pausing all active rollouts at each token-count checkpoint to evaluate the pruning criterion. In batch-mode serving frameworks such as SGLang, this synchronization is natively supported: all rollouts in a batch share the same generation schedule, so pausing at a fixed token count introduces no idle time. In continuous-batching frameworks (e.g., vLLM), rollouts that reach a checkpoint before others may incur brief idle time while the batch synchronizes, a minor scheduling cost. Our online wall-time measurements in Section 4.4 confirm that this overhead is negligible in practice: the pruning decision itself is a string count operation requiring no model inference. KV-cache footprint. In the online deployment run, FoT reduces the cumulative generated-token footprint from 12.88M under SC@32 to 8.51M under FoT@32, a 33.9% reduction. Since each generated token occupies one KV-cache position per layer, this is the memory-side counterpart of the wall-time and generation-FLOP savings reported in Table 6. D.2 Efficient-Attention Scaling Our headline FLOP accounting (Section 3.1) assumes standard full attention, where per-rollout cost scales as s2s^2 and late tokens dominate, which is precisely the cost FoT removes by terminating spiraling rollouts early. Efficient attention changes this accounting: under sliding-window attention with window w, per-token cost saturates once the position exceeds w, so per-rollout cost grows linearly (≈w⋅s≈ w· s) rather than quadratically, and sparse or linear variants behave similarly. FoT’s saving then tracks its token reduction rather than the larger quadratic figure. Table 21 applies FoT’s termination pattern under each regime on the full 4-benchmark pool. The full-attention saving (48.7%, pooled over all rollouts as in Table 7) degrades smoothly toward the token-reduction floor (23.0%) as the window narrows; realistic sliding-window models (w≈4w≈ 4–88K) retain roughly 63–79% of it. The accuracy result is unaffected, since it depends only on which rollouts survive pruning, and the KV-cache release at each checkpoint is likewise independent of the attention regime. Three headline savings therefore coexist in this paper as related but distinct decompositions of the same pruning pattern: the pooled full-quadratic attention saving (48.7%), the linear-attention token-reduction floor (23.0%), and the full-model FLOP saving that includes non-attention compute (28.8%; Table 15). Attention regime FLOP saving Full quadratic (headline) 48.7% Sliding window, w=8w=8K 38.6% Sliding window, w=4w=4K 30.9% Sliding window, w=2w=2K 26.8% Sliding window, w=1w=1K 24.8% Linear / sparse top-k 23.0% Table 21: FoT attention-FLOP saving relative to SC@32 under different attention regimes, on the 4-benchmark pool with the universal configuration. The saving interpolates between the full-attention headline and the token-reduction floor as the effective attention window narrows. D.3 Adaptive Consistency Baseline We simulate Adaptive Consistency (AC; 2) on the same pre-generated rollout pool used by all other methods.33 3 https://github.com/Pranjal2041/AdaptiveConsistency AC is a sequential early-stopping method: rollouts are consumed one at a time, and at each step a Bayesian stopping criterion determines whether the current majority answer is sufficiently confident. We use the Beta stopping criterion with confidence threshold τ=0.95τ=0.95 (the paper’s default), which stops when posterior confidence in the current majority answer exceeds τ. Because AC assumes sequential generation while our rollouts are generated in parallel, we simulate sequentiality by processing rollouts in index order. To account for ordering effects, we average results over 10 random permutations of the rollout sequence per problem. AC’s compute savings are measured as the fraction of rollouts not consumed before stopping; FLOP savings are computed identically to other methods (sum of si2s_i^2 over consumed rollouts only). D.4 Slim-SC Baseline Configuration Our Slim-SC baseline is a faithful reproduction of 9, using the configuration from their official repository.44 4 https://github.com/hyscale-lab/slimsc We use the diversity pruning strategy with cosine similarity threshold τ=0.9τ=0.9, segment-level embeddings via sentence-transformers/all-mpnet-base-v2, and a warm-up of 20 thoughts before pruning is enabled. These are the authors’ recommended defaults; no threshold tuning was performed on our data.