Paper deep dive
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
Huajun Bai, Weiwei Lv, Huichuan Zheng, Youyou Lu, Jiwu Shu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/7/2026, 7:36:27 AM
Summary
The paper introduces SPORK, a training-free controller that accelerates agentic LLM inference by implementing self-speculative forking. It addresses the serial bottleneck in the Thought-Action-Observation loop by forking a probe at the start of generation to predict the upcoming tool call. This predicted tool executes speculatively, overlapping with the model's chain-of-thought reasoning. SPORK employs a prefix-cache fork to reduce probe cost, a confidence gate to filter mispredictions, and partial-token acceptance to recycle rejected probes. Evaluated on benchmarks like GAIA and HotpotQA using Qwen3-32B, SPORK reduces tail latency by up to 18% without accuracy loss, deploying as a thin controller over standard APIs like vLLM.
Entities (14)
Relation Signals (10)
SPORK ā testedon ā Qwen3-32b
confidence 97% Ā· On real-tool benchmarks, SPORK cuts Qwen3-32B's GAIA P95 by 18% (131.9 to 108.1 s)
SPORK ā implements ā Speculative tool execution
confidence 96% Ā· We present SPORK... a training-free controller that dispatches the speculated tool call early, overlapping its execution with the remaining chain-of-thought decode.
SPORK ā integrateswith ā vLLM
confidence 96% Ā· realize speculative overlap within the standard vLLM serving stack (<<0.3% TPOT overhead)
SPORK ā reduceslatencyon ā GAIA benchmark
confidence 96% Ā· On real-tool benchmarks, SPORK cuts Qwen3-32B's GAIA P95 by 18% (131.9 to 108.1 s)
SPORK ā accelerates ā Thought-Action-Observation loop
confidence 95% Ā· LLM agents... Thought-Action-Observation loop is often serial... SPORK... overlapping its execution with the remaining chain-of-thought decode.
SPORK ā consistsof ā Prefix-cache fork
confidence 95% Ā· a prefix-cache fork cuts probe cost, a confidence gate filters mispredictions, and partial-token accept turns rejected probes into speculative-decoding drafts
SPORK ā consistsof ā Confidence gate
confidence 95% Ā· a confidence gate filters mispredictions, and partial-token accept turns rejected probes into speculative-decoding drafts
ā ā
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns. This wait consumes 16-37% of wall time in our workloads and 35-61% in prior reports. Speculative tool execution can hide this wait, but existing systems need auxiliary predictors, historical traces, or static workflow graphs, leaving a gap for training-free, day-one deployment. We observe that the model can be its own predictor: a probe forked at the start of generation predicts Qwen3-32B's upcoming tool name with 74.6-99.6% accuracy across five benchmarks. We present SPORK (Self-sPeculative fORKing), a training-free controller that dispatches the speculated tool call early, overlapping its execution with the remaining chain-of-thought decode. A cost model captures when speculation breaks even, and each component improves one of its terms: a prefix-cache fork cuts probe cost, a confidence gate filters mispredictions, and partial-token accept turns rejected probes into speculative-decoding drafts. On acceptance, the tool result is ready when reasoning ends; on rejection, SPORK falls back to serial execution with no correctness penalty. On real-tool benchmarks, SPORK cuts Qwen3-32B's GAIA P95 by 18% (131.9 to 108.1 s); the mechanism holds across model sizes from 4B to 32B and across dense and mixture-of-experts models, with task accuracy within 1 pp of baseline or better wherever measured. SPORK deploys as a thin controller over standard completion APIs (no retraining, no auxiliary models, no offline traces) and is orthogonal to token-level speculative decoding. SPORK is open source at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.03333v1
- Canonical: https://arxiv.org/abs/2607.03333v1
Trouble viewing inline? Open PDF directly ā
Full Text
74,726 characters extracted from source content.
Expand or collapse full text
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference Huajun Bai Tsinghua University , Weiwei Lv Meituan , Huichuan Zheng Tsinghua University , Youyou Lu Tsinghua University and Jiwu Shu Tsinghua University Abstract. LLM agents are becoming a common interface for research, coding, and question answering, yet their ThoughtāActionāObservation loop is often serial: the model reasons, emits a tool call, then idles the GPU until the result returns. This wait consumes 16ā37% of wall time in our workloads and 35ā61% in prior reports. Speculative tool execution can hide this wait, but existing systems need auxiliary predictors, historical traces, or static workflow graphs, leaving a gap for training-free, day-one deployment. We observe that the model can be its own predictor: a probe forked at the start of generation predicts Qwen3-32Bās upcoming tool name with 74.6ā99.6% accuracy across five benchmarks. We present Spork (Self-sPeculative fORKing), a training-free controller that dispatches the speculated tool call early, overlapping its execution with the remaining chain-of-thought decode. A cost model captures when speculation breaks even, and each component improves one of its terms: a prefix-cache fork cuts probe cost, a confidence gate filters mispredictions, and partial-token accept turns rejected probes into speculative-decoding drafts. On acceptance, the tool result is ready when reasoning ends; on rejection, Spork falls back to serial execution with no correctness penalty. On real-tool benchmarks, Spork cuts Qwen3-32Bās GAIA P95P_95 by 18% (131.9ā 108.1 s); the mechanism holds across model sizes from 4B to 32B and across dense and mixture-of-experts models, with task accuracy within 1 p of baseline or better wherever measured. Spork deploys as a thin controller over standard completion APIs (no retraining, no auxiliary models, no offline traces) and is orthogonal to token-level speculative decoding. Spork is open source at https://github.com/baihuajun24/spork. LLM agents, tool use, speculative execution, inference acceleration ā copyright: none 1. Introduction āTreat things before they exist. Regulate things before disorder begins.ā ā Lao Tzu LLM-powered agents are becoming a common interface for complex work such as deep research, coding, and tool-assisted question answering (OpenAI, 2025; Manus, 2025; Anthropic, 2025; GitHub, 2025; Yao et al., 2023; Schick et al., 2024). In these systems, the LLM acts as the controller: it reasons over the user request, emits a structured tool call, waits for the external tool to finish, and then resumes generation using the tool result. This ThoughtāActionāObservation loop (Yao et al., 2023) is simple and safe, but it is also often serial: each tool result conditions all subsequent tokens, so the serving system cannot resume generation until the result arrives. While the tool is running, the GPU often has no useful work to do for that agent turn. This tool wait is hard to engineer away because it is structurally bounded by network round-trips, container cold-starts, and database I/O that resist the exponential improvements seen in model serving. It is also large. Our profiling across three benchmarks in Figure 2 shows that tool execution accounts for 16ā37% of end-to-end wall time, and on turns with real network tools the GPU is idle for the majority of the turn; Sui et al. (2026) report even higher fractions, 35ā61% of total agent latency across agent workloads. Figure 1. Tool intent is visible early. A speculative thread, forked at the main threadās first token with a forced tool-call prefix, predicts the exact tool call that the main thread emits ā¼2,000 2,000 CoT tokens later. As LLM inference grows faster (through speculative decoding (Leviathan et al., 2023), hardware scaling, and disaggregated serving (Zheng et al., 2024)), the fraction of wall time spent waiting for external tools increases, not decreases, making tool-wait the asymptotic bottleneck of agent latency. Figure 1 illustrates the opportunity this paper exploits: after a single streaming token, a forked probe (sharing the same KV cache prefix) predicts the identical web-search query that the main stream will emit after 2,000 tokens of chain-of-thought. The tool can be dispatched speculatively during those 2,000 tokens of decode, hiding most of its network latency behind useful computation. Existing systems attack adjacent parts of this problem but leave a gap for day-one, training-free deployment without predictors or historical traces. Token-level speculative decoding (Leviathan et al., 2023; Li et al., 2024b; Cai et al., 2024) and LLM serving systems (Kwon et al., 2023; Zheng et al., 2024; Lin et al., 2024; Agrawal et al., 2024) reduce generation or serving overhead, but they do not dispatch external tools earlier: the tool stall begins only after the model has finished emitting the call. Workflow and serverless optimizers (Mahgoub et al., 2022; Stojkovic et al., 2023; Fu et al., 2024b) can prewarm or schedule known execution graphs, whereas tool-using agents generate their next tool call online; there is no static graph to prefetch. Recent agent-speculation systems move closer to execution-level overlap, but they rely on a separate source of predictions: Speculative Actions (Ye et al., 2026) and DualSpec (Zhong et al., 2026) use predictors or verifier policies, while PASTE (Sui et al., 2026) mines recurring tool-call patterns from historical traces. These designs are powerful when traces, predictors, or repeated workflows are available; they do not speculate from the currently running model alone. Figure 2. Tool time is significant across agent tasks. Across search, browsing, and database-style tasks, tool execution accounts for 16ā37% of end-to-end wall time. The analogy to speculative execution in microprocessors is instructive: a CPU executes past an unresolved branch and rolls back on misprediction. The difference is the prediction source: where CPUs use branch predictors trained on instruction history, Spork uses the running model itself, with no auxiliary model, no trace collection, and no retraining. Spork exploits a different opportunity: instruction-tuned models often reveal their next tool call before completing the full chain-of-thought (§3.1). The fork also emits logprobabilities that indicate prediction confidence, and even failed probes share a verified prefix with the final tool call. These properties let the serving system turn the target model into its own speculative tool predictor. In this paper, we establish three empirical insights that make self-speculative tool use viable. First, Qwen3-32B (Yang et al., 2025a) predicts its own next tool call with high name accuracy (74.6ā99.6% across GAIA (Mialon et al., 2024), tau2-bench (Barres et al., 2025), BrowseComp (Wei et al., 2025), HotpotQA (Yang et al., 2018), and BFCL (Patil et al., 2025)) at the first-token fork point; the model already āknowsā what it will do before finishing its reasoning (§3.1). Second, the probeās logprobability confidence separates correct from incorrect predictions (0.90 vs. 0.46 mean span probability on GAIA), and argument accuracy climbs from 7.6% to 97.5% as more chain-of-thought context is observed (§3.2). Third, even when strict matching rejects a probe, the probe shares a median of 18 verified tokens with the final tool call. Reused as a speculative-decoding draft, these tokens let the main stream skip re-decoding 35ā50% of the tool-call body (§3.3). Spork realizes this opportunity as a training-free controller for GPU-served tool agents. At each agent turn it forks the KV cache after the main generationās first token (D1: near-zero overhead via prefix-cache sharing, reducing probe cost from 1.6 s to 0.35 s), gates speculative dispatch on the probeās own logprob confidence (D2: maximizes the accepted overlap by committing at the earliest confident probe), and recovers verified prefix tokens from rejected probes via speculative-decoding verification (D3: a median of 18 tokens recycled as drafts, avoiding ā¼ 0.6 s of re-decode per rejected turn in offline analysis). On a strict name-and-arguments match, Spork uses the pre-computed tool result; otherwise it falls back to serial execution with no correctness penalty. On GAIA (N=165, real web search, Qwen3-32B), the full Spork configuration (D1+D2+D3) reduces median per-query latency by 10% (P50: 34.7 s ā 31.2 s) and tail latency by 18% (P95: 131.9 s ā 108.1 s) while preserving task accuracy (EM within 1 p of baseline). On HotpotQA (N=200, real Wikipedia API, Qwen3-32B), the full D1+D2+D3 configuration is Pareto dominant: faster and higher quality than the baseline simultaneously (P95 speedup 1.06Ć, EM +1.5 p). A latency sweep on tau2-bench (airline domain, N=43, simulated tool latency 0.5ā5 s) validates the cost model: mean speedup grows monotonically with tool duration, from 1.09Ć at a 0.5 s tool floor to 1.18Ć at 5 s, matching EQ1ās prediction within 2%. The mechanism generalizes across model scale and architecture: it holds on Qwen3-4B (on tau2; neutral elsewhere) and on the Qwen3.5-35B-A3B (Qwen Team, 2026) mixture-of-experts (16ā20% P95P_95 reduction across tau2, GAIA, and HotpotQA). We also characterize when Spork does not help: without thinking-mode chain-of-thought the overlap window vanishes, and a model whose forced probe diverges from its own tool-call format defeats self-speculation (§6). This paper makes the following contributions: (1) Empirical insights (§3). We show that Qwen3-32B predicts its own tool names with high accuracy (74.6ā99.6%) at the fork point, exposes confidence through logprobabilities, and leaves recoverable prefixes when strict matching rejects; we formalize the conditions under which speculation breaks even via a cost model (EQ1, §2.2). (2) Training-free system design (§4). We present three mechanisms (D1āD3) that realize speculative overlap within the standard vLLM serving stack (<<0.3% TPOT overhead), requiring no model changes, no auxiliary predictors, and no historical traces. (3) End-to-end evaluation with real tools (§6). On GAIA and HotpotQA with real network tools, Spork reduces P95P_95 latency by 6ā18% while improving or preserving task accuracy, and the mechanism generalizes across architecture (dense and MoE), with gains tapering to neutral at 4B scale on real-tool benchmarks. We validate the cost model against measured operating points and characterize when speculation wins and when it does not. 2. Background & Preliminaries 2.1. The Serial Tool-Wait Bottleneck The dominant execution model for LLM agents follows the ReAct loop (Yao et al., 2023): the model reasons (Thought), emits a structured tool invocation (Action), waits for the external service to return a result (Observation), and then resumes reasoning conditioned on that result. This ThoughtāActionāObservation cycle enforces a strict serial dependency: [(Re)prefill]ā[Decode]ā[Tool wait]āāÆ[(Re)prefill]\;ā\;[Decode]\;ā\;[Tool wait]\;ā\;Ā·s Each turn (re)prefills the context (the user prompt on the first turn, the appended tool result thereafter), decodes the reasoning and tool call, then stalls on the tool wait. The tool call cannot be dispatched until decoding completes; the next decode cannot begin until the tool result arrives and is (re)prefilled into the context window. This creates a structural pipeline stall regardless of how fast individual components are. Tool latency dominates agent wall time. Figure 2 shows that tool execution consumes a substantial share of end-to-end agent latency across three workloads: 16% on tau2-bench (enterprise database APIs with 2 s simulated floor), 19% on GAIA (real web search APIs), and 37% on BrowseComp (web search + page browse). On BrowseComp, the distribution is heavily right-skewed: median tool latency is 1.19 s but P95P_95 reaches 70 s, meaning tail-latency optimization requires addressing tool stalls, not just model serving. Why existing solutions fall short. Token-level speculative decoding (Leviathan et al., 2023; Li et al., 2024b; Cai et al., 2024) reduces generation latency but does not dispatch external tools earlier. Parallel tool calling (e.g., OpenAI function calls) helps only when the model can identify independent tools upfront; empirically, only 15ā25% of tool calls in GAIA and SWE-bench (Jimenez et al., 2024) are parallelizable with their predecessor. Workflow optimizers (Mahgoub et al., 2022; Stojkovic et al., 2023; Fu et al., 2024b) prewarm known execution graphs, but agent tool calls are generated online: there is no static graph to prefetch. Recent agent-speculation systems (Ye et al., 2026; Zhong et al., 2026; Sui et al., 2026) move closer to execution-level overlap but require separate predictors, verifier policies, or historical traces. None speculates from the currently running model alone. 2.2. Cost Model for Speculative Overlap If we could predict the next tool call before the model finishes generating it, we could dispatch the tool speculatively and overlap its execution with ongoing model decode. The realized speedup depends on four measurable parameters: ⢠α: the fraction of tool-call turns where speculation is accepted (gate rate). α is always measured per turn, not per dispatched probe; a retrying gate may issue several probes within one turn. ⢠toverlapt_overlap: mean realized time overlap per accepted turn (how much tool latency is hidden behind decode). ⢠TbaseāT_base^*: the realized main-stream decode cost on the speculative path. It equals the baseline TbaseT_base unless prefix recovery is active, which lowers it by letting the main stream skip re-decoding the verified tool-call tokens it recovers from the rejected probe on a miss. ⢠TohT_oh: per-turn overhead from probe computation and wasted speculative execution on rejected turns. The resulting latency-ratio formula (full derivation in Appendix A): (1) Ratio=TbaseTbaseāāαā toverlap+Toh Ratio= T_baseT_base^*-α· t_overlap+T_oh Speculation improves latency when Ratioā„1Ratioā„ 1, which approximately requires αā toverlapā„Tohα· t_overlapā„ T_oh (exact form in Appendix A): the time saved by overlapping tool execution on accepted turns must exceed the overhead incurred across all turns. The upper bound is set by the tool-time fraction ftoolf_tool: with perfect prediction (α=1α=1, Toh=0T_oh=0), Ratio=1/(1āftool)Ratio=1/(1-f_tool). For BrowseComp (ftool=0.366f_tool=0.366), this ceiling is 1.58Ć1.58Ć. The next section establishes empirically that these parameters are achievable: α is high, TohT_oh is low, and even rejected probes contribute via prefix recovery. 3. Self-Speculative Tool Use The cost model in §2.2 identifies three conditions for speculation to succeed: high gate acceptance (α), sufficient overlap time (toverlapt_overlap), and low overhead (TohT_oh). Whether these are achievable depends on an empirical question about the model itself: Research question. Can a running LLM predict its own next tool call accurately enough, early enough, and recoverably enough that speculative tool use pays off? We constrain the prediction source to the running modelās own state: no auxiliary model, no historical traces. We answer affirmatively with three empirical insights, established on Qwen3-32B across five real-tool and simulated benchmarks (GAIA, HotpotQA, tau2-bench, BrowseComp, BFCL): the model already knows its next tool call (Insight 1, accuracy), its own confidence signals when to commit (Insight 2, timing), and even rejected probes leave verified prefixes (Insight 3, recovery). How these properties scale to a smaller model (Qwen3-4B) and where they break (native XML tool formats) is examined in §6.6. 3.1. Insight 1: Models Already Know Their Next Tool Call A short forced-prefix probe can expose the next tool identity well before the serial tool boundary. We force a tool-call prefix immediately after the modelās first decode token (fork-at-start) and compare the probeās predicted tool name and arguments against the modelās unrestricted generation. Figure 3. Fork-at-start probes recover the next tool name across workloads. On Qwen3-32B, forced tool-call probes near decode start predict the eventual tool name across five benchmarks (sorted by argument exact-match; hatched bars). Protocols: replay pos=0 (GAIA, tau2-bench, BFCL) and online first-token (BrowseComp, HotpotQA). Foresight across benchmarks. On Qwen3-32B, near-decode-start tool-name accuracy is consistently high across the five benchmarks in Figure 3: 98.3% on GAIA, 74.6% on tau2-bench, 83.7% on BrowseComp, 99.6% on HotpotQA, and 99.2% on BFCL. Argument-exact accuracy is lower and more variable (7.6ā83.9%), reflecting the difficulty of predicting precise parameter values before observing the chain-of-thought. This is the gap that Insight 2 addresses. No-op recovery. On 11 tau2-bench turns where the baseline declined to call any tool, 9/11 forced-probe outputs were correct tool calls the model should have made; fork-prefix can act as a cheap correction against model over-caution. Implication for EQ1. High name accuracy means α (gate acceptance on name match) is intrinsically high before any confidence gating. The remaining challenge is argument accuracy, which Insight 2 addresses. 3.2. Insight 2: Confidence Improves with Chain-of-Thought Figure 4. Fork accuracy rises monotonically with CoT context. On GAIA (Qwen3-32B), args-exact accuracy climbs from 7.6% at fork-at-start to 61.0% at 80% CoT and 97.5% at think-end. Later probes are more accurate but leave a smaller overlap budget. The useful operating point is not necessarily the latest or most accurate probe; it is the earliest probe whose confidence justifies committing tool execution time. Logprob separability. The probeās own generation logprobabilities over the tool-name span separate correct from incorrect predictions without any classifier training. On GAIA (Qwen3-32B, think mode): the mean span probability is 0.90 for correct predictions versus 0.46 for incorrect ones. This makes confidence observable without external classifiers: low confidence probes can be delayed or rejected before they spend real tool time. On tau2-bench no-think mode the signal saturates (correct ā incorrect ā 0.979): tool selection is a first-token reflex, so confidence is most informative on reasoning-mode workloads with diverse tool sets. CoT length amplifies accuracy. Fork accuracy at six positions along the CoT trajectory (0ā100% of tokens, gold-CoT replay) rises monotonically (Figure 4). On GAIA (N=118 tool-call questions, Qwen3-32B), argument-exact accuracy climbs from 7.6% at position 0 to 61.0% at 80% CoT and 97.5% at think-end, a +53 p gap that a later probe can exploit, at the cost of a smaller overlap budget (the decode time remaining to hide the tool, distinct from EQ1ās realized toverlapt_overlap). Implication for EQ1. The accuracyātimeliness tradeoff means a dynamic gate is better than a fixed fork position: commit early when confident, wait when uncertain. This is the foundation of D2ās adaptive confidence gate (§4.2). 3.3. Insight 3: Rejected Probes Still Contain Verified Prefixes Strict matching is necessary for correctness, but rejection does not mean the probe was useless: a rejected probe can serve as a draft that the main stream verifies and partially reuses via speculative decoding. Prefix overlap after strict rejection. When the strict gate rejects a probe (tool name or arguments differ from main), the probeās tool-call token sequence still shares a long verified prefix with the eventual main output. Figure 5 quantifies this on two GAIA N=165 seeds (n=657n=657 rejected-probe events): the median shared prefix is 18 tokens, with a mean of 21.1 and 25th/75th percentiles of 14 and 25 tokens (often 35ā50% of the full tool-call body). A rejected probe need not be pure waste: if the shared prefix can be verified, the main generation can reuse those tokens and decode only the remaining suffix. Figure 5. Most rejected probes share recoverable prefixes. On GAIA, the 25th-percentile shared prefix is 14 tokens: 75% of rejected probes share at least 14 tokens with the main tool call. Implication for EQ1. Prefix recovery reduces the effective TohT_oh on the miss branch: instead of discarding the probe and re-decoding the full tool call from scratch (ā¼50 50 tokens), the system verifies and accepts k matching tokens and decodes only the 50āk50-k suffix. This is the foundation of D3 (§4.3). Summary. The three insights establish that self-speculative tool use is feasible, steerable, and recoverable: the model predicts its next tool identity early, its own confidence signals when to commit, and even rejected probes yield verified prefixes. The next section translates these insights into three system mechanisms that jointly push the operating point toward the EQ1 break-even condition from both sides. 4. Spork Design Systems question. How can we overlap tool execution with main generation at minimal added overhead, while keeping the agent lossless with respect to its serial execution? Spork answers with three mechanisms: Design 1 (D1), Design 2 (D2), and Design 3 (D3), each grounded in one insight from §3 and targeting a specific parameter of the cost model (EQ1, §2.2). At each agent turn, Spork runs the following pipeline (Figure 6): (1) Dispatch the main generation as a streaming chat-completion request. (2) After the first streaming token, launch a fork thread that issues a raw completion request with a forced tool-call prefix. (3) The fork thread monitors its generation logprobs (D2); once confidence reaches a threshold Īø (the confidence gate, §4.2), dispatch speculative tool execution. (4) When the main stream reaches its tool call, compare it against the committed probe under the gate criterion. On match: use the pre-computed result. On mismatch: fall back to serial tool execution. In the engine configuration, D3 has meanwhile served the probeās tool-call body to the main stream as speculative-decoding draft tokens, so the main decode re-emits only the tokens after the first mismatch rather than the whole call. The D1/D2 controller is a Python application-layer component that uses standard serving APIs and collects no training data. D3 additionally integrates with vLLMās speculative-decoding proposer path to inject verified probe tokens as drafts; we describe this as a prototype integration rather than a generic no-engine-modification claim. Figure 6. Spork method overview. Top: the baseline turn is strictly serial: (re)prefill, CoT decode, tool-call decode, then a stall while the tool executes. Bottom: Spork forks the KV cache after the mainās first token and launches a forced tool-call probe (D1); once the probeās logprob confidence clears the gate (D2), the predicted tool executes during the remaining decode, producing the overlap toverlapt_overlap. A strict match at main-stream completion accepts the pre-computed result; on mismatch the turn falls back to serial execution and D3 recycles the rejected probeās tokens as a speculative-decoding draft. EQ1 parameter mapping. Each mechanism targets a different term in the cost model (Table 1; Eq. 1, §2.2): D1 follows Insight 1 and minimizes the overhead term TohT_oh: by reusing the mainās prefix KV cache, the probeās cost collapses from 1.6 s to 0.35 s, and its dispatch at the first streaming token opens the overlap window. D2 follows Insight 2 and maximizes the expected overlap αā toverlapα· t_overlap: gating dispatch on the probeās own logprob confidence, it commits at the earliest probe confident enough to act on, trading a slightly smaller window for much higher acceptance. D3 follows Insight 3 and lowers TbaseāT_base^* on misses by recycling the probeās verified prefix as draft tokens, so the main stream re-decodes only the suffix after the first mismatch. Together, all three push the operating point toward the αā toverlapā„Tohα· t_overlapā„ T_oh break-even condition from both sides. Table 1. Design rationale. Each mechanism is tied to one insight and one EQ1 term. Design Insight EQ1 effect D1 Insight 1 minimizes TohT_oh D2 Insight 2 maximizes αā toverlapα· t_overlap D3 Insight 3 lowers TbaseāT_base^* (miss re-decode) Spork composes the three mechanisms. D1+D2 form the training-free controller, and D3 adds engine-side partial-token recovery. The ablation in §6.4 measures each mechanismās contribution. The controller is stateless across turns except for the shared prefix cache inside the serving engine: each agent turn forks independently, and the gate always compares against the main streamās final tool call before any result is committed to the conversation history. 4.1. D1: Prefix-Cache Fork for Near-Zero Probe Overhead Mechanism. The probe is not free, since it must generate a tool call of its own; its overhead decomposes into a decoding cost and a prefilling cost. The decoding cost is small. The fork decodes only the tool-call body (typically ā¤50⤠50 tokens, stopping at the tool-call close), 10ā50Ć10--50Ć shorter than the mainās CoT, about 0.3 s on Qwen3-32B at 15K context. The prefilling cost is the heavy part: the fork shares the full prompt prefix (system + user + history) with the main request, and re-prefilling it costs ā¼ 1.3 s at the same context; worse, we measure that dispatching the fork simultaneously with the main request causes 2.5Ć prefill contention. Modern LLM serving systems such as vLLM (Kwon et al., 2023) and SGLang (Zheng et al., 2024) cache prefix KV states; D1 therefore dispatches the fork after the mainās first streaming token. Because the main request has by then populated the prefix cache, the forkās prefill is a cache hit at near-zero cost (ā¼ 0.05 s), and the staggered dispatch reduces prefill interference to <1%<1\%. With D1, probe overhead collapses from 1.6 s (1.3 s prefill + 0.3 s decode) to 0.35 s (0.05 s cache-hit prefill + 0.3 s decode). Figure 7 shows the resulting end-to-end effect: the forkās measured TPOT overhead on the main stream is within intrinsic batching noise. Figure 7. D1 reuses the prefix KV cache to keep probe overhead small. Probe TPOT overhead (+0.22%) is within intrinsic batching noise (+0.3%). 4.2. D2: Adaptive Confidence Gate Motivation. Without gating, D1 dispatches speculative tool execution on every fork, regardless of whether the probe is likely correct. Each wasted dispatch consumes network/compute resources and, on side-effecting tools, must be rolled back. Insight 2 (§3) shows that probe accuracy improves dramatically with context: argument accuracy rises from 7.6% at 0% CoT to 97.5% at 100% CoT (Figure 4). D2 exploits this by not committing until the forkās own token-level confidence signals that the generated tool call is likely correct, enabling the system to filter out low-quality probes before they trigger wasted work. Confidence metric. The fork generates tool-call tokens with per-token logprobs enabled. We define the confidence score as the minimum top-1 token probability over the tool-name span (the first L tokens after the "name": " prefix): (2) c=mini=2Lā”expā”(āi),c= _i=2^L\; ( _i), where āi _i is the log-probability of the top-1 token at position i, and L is the length of the generated tool name (typically 1ā4 tokens). The minimum starts at i=2i=2 because the first token of the span (the opening quote) is always high-probability boilerplate. The minimum operator is deliberately conservative: a single low-confidence token in the name span signals that the model is uncertain about the tool selection, and committing would likely waste a tool execution. Gate logic. Once cā„Īøcā„Īø, the fork commits to speculative tool execution. If the fork finishes without c reaching Īø, no speculative dispatch occurs and the turn falls back to serial execution at zero additional cost (the fork decoding is already overlapped with main CoT). Strict acceptance. Spork accepts the probe if and only if its full tool call (name and serialized arguments) matches the main generation exactly. We deliberately do not ship a name-only gate: a name-only match can silently change tool behavior on side-effecting tools, and D3 (§4.3) already recovers most of the savings that name-only acceptance would deliver without compromising correctness. Threshold selection. Figure 8 shows the confidence distribution across 997 probes on GAIA (N=127, Qwen3-32B). Correct probes (those whose tool name matches the main generation) cluster tightly above c=0.9c=0.9, while incorrect probes spread broadly over [0.2,0.9][0.2,0.9]. At Īø=0.90Īø=0.90, the gate achieves 88% precision with 100% recall (F1 == 0.937), filtering 77% of all probes while retaining every correct one. We select Īø=0.90Īø=0.90 as the operating point that maximizes F1. Alternatives considered. (1) Fixed token-count delay (wait N main tokens before forking): rigid, does not adapt to per-task difficulty, and delays correct probes unnecessarily. (2) Entropy-based gate (threshold on next-token entropy): entropy over the full vocabulary correlates poorly with tool-name correctness; min-prob over the specific name span is a more direct signal. (3) Learned classifier: adds inference latency and a training dependency; the zero-parameter min-prob gate achieves F1 >> 0.93 with no training. Effect on EQ1. D2 maximizes the expected overlap αā toverlapα· t_overlap: by waiting for the earliest confident probe it trades a slightly smaller window for much higher acceptance, and by filtering low-confidence turns before dispatching spec_tool it avoids wasted work (on BrowseComp, 52% of strict-gate turns are rejected and never trigger a tool). Under the strict gate without D2, αā0.22αā 0.22 on GAIA; with D2, α rises to 0.370.37 (49 vs. 101 accepted probes). Recall that α counts tool-call turns (§2.2); the dispatched-probe totals are larger (434 for D1, 2124 for D1+D2) because the retry cadence can issue several probes within one turn. Figure 8. D2 confidence gate threshold selection (GAIA N=127, 997 probes). (a) Correct probes cluster above Īø=0.90Īø=0.90; incorrect probes spread below. The threshold cleanly separates them. (b) Precision/Recall/F1 vs. Īø: the operating point Īø=0.90Īø=0.90 maximizes F1 (0.937) with 88% precision and 100% recall. D2 filters 77% of probes (saving wasted tool executions) while committing every probe that would have been correct. 4.3. D3: Partial-Token Accept for Rejected Probes When the strict gate rejects a probe, the main thread must otherwise generate the full tool call from scratch (ā¼50 50 tokens, ā¼1.5 1.5 s at 32B scale). Empirical analysis of 1268 probeābaseline pairs (BrowseComp, Appendix B) shows that even rejected probes match the mainās greedy output for the first k tokens: mean kā29.5kā 29.5, median 27, P25=17P_25=17, P75=35P_75=35. Mechanism. D3 turns the probe into draft tokens for the main streamās own speculative decoding rather than re-decoding the tool call from scratch (Figure 9). We integrate a SporkProposer into the serving engineās speculative-decoding proposer path: as the main request decodes, the proposer watches for the <tool_call> boundary and, at that point, supplies the probeās tool-call body as the draft for that request. (1) When the main stream reaches the tool-call opener, the probe body is registered as the per-request draft. (2) The engine verifies the draft against the main model in the same decode step (no separate verification pass, no second prompt); it accepts the longest prefix matching the main modelās greedy tokens. (3) The first k matching tokens are accepted; the main stream decodes only the tool-call tokens after the first mismatch. D3 runs only in the in-process (engine) configuration, whose baseline is ngram speculative decoding; the HTTP controller configuration (§5) runs D1+D2 only. D3 reduces the per-rejected-turn cost to approximately Tbaseā+TohT_base^*+T_oh, where Tbaseā=TbaseāTD3T_base^*=T_base-T_D3 is the realized base decode after recycling the verified prefix as drafts (ā 0.6 s per rejected turn, ā 0.31 s amortized per turn on BrowseComp; Appendix B). D3 is complementary to D1+D2: tool-level speculation on accepted turns, token-level recycling on rejected turns; the fork itself is the draft, avoiding a separate draft model. Figure 9. D3 treats the rejected fork as a draft for speculative verification. Instead of discarding the probe, Spork verifies its tool-call tokens, accepts the longest matching prefix, and lets the main stream decode only the suffix after the first mismatch. 5. Implementation Overview. Spork is implemented in Python as a controller layer plus an optional D3 proposer integration. The D1/D2 path (fork dispatch, logprob monitoring, asynchronous tool execution, and gate verification) uses standard OpenAI-compatible vLLM endpoints (/v1/chat/completions for the main stream, /v1/completions for the fork). D3 adds a small SporkProposer integration with vLLMās speculative-decoding proposer path so that verified probe tokens can be injected as draft tokens at the tool-call boundary. Algorithm 1 summarizes the end-to-end control flow. The HTTP controller realizes D1+D2; the engine configuration additionally performs D3, which the HTTP path omits. Algorithm 1 Spork Controller with D1+D2+D3 0: User request, tool set T, threshold Īø, retry budget R, token step s 1: Dispatch main stream (streaming chat completions) 2: Wait for mainās first streaming token 3: specāā specā ; (t^,a^)āā ( t, a)ā 4: for r=0,1,ā¦,Rā1r=0,1,ā¦,R-1 do 5: Wait until main has decoded rā srĀ· s further CoT tokens 6: Fork a completions request with forced prefix over observed CoT D1 cache hit 7: (t^,a^)āparseā(fork)( t, a) (fork); cāminjānameā_āspanā”expā”(āj)cā _j \_span ( _j) D2 8: if (t^,a^)( t, a) valid and cā„Īøcā„Īø then 9: specāasyncā_ātoolā(t^,a^)spec \_tool( t, a) commit; overlaps main decode 10: break 11: end if 12: if main stream finished then 13: break 14: end if 15: end for 16: Wait for main tool call (tā,aā)(t^*,a^*) 17: if specā ā specā and (t^,a^)=(tā,aā)( t, a)=(t^*,a^*) then 18: return pre-computed specspec result strict match 19: else 20: return toolā(tā,aā)tool(t^*,a^*) serial fallback; D3 draft-injects the probe body during decode 21: end if Figure 10. Spork reduces P95P_95 tail latency across models and benchmarks. Teal: baseline; orange: the paper-facing Spork configuration for each run (engine D1+D2+D3; HTTP D1+D2). P95P_95 latency on GAIA, HotpotQA, and tau2 (2 s tool floor) for Qwen3.5-35B-A3B, Qwen3-32B, and Qwen3-4B (note the per-panel y-scales); the percentage above each pair is the P95P_95 reduction. Spork generalizes across model scale (4Bā 32B) and architecture (dense and MoE). Per-model measurement setup is described in §6.1; quality is preserved across all configurations (Figure 11). vLLM compatibility. Sporkās D1/D2 controller requires a serving backend that supports: (1) streaming /v1/chat/completions with per-token logprobs; (2) /v1/completions with per-token logprobs and an arbitrary prompt prefix (for the fork thread); and (3) prefix KV-cache sharing between concurrent requests (D1 relies on this). vLLM (our serving version, §6.1) satisfies all three with prefix caching enabled; among the broader wave of serving engines (Song et al., 2024; Hu et al., 2025), any backend exposing these three interfaces can host the controller. D3 additionally requires access to the serving engineās speculative-draft proposer interface; in our prototype this is implemented by wrapping vLLMās ngram proposer. Concurrency and safety. The main stream and the fork run as concurrent asyncio tasks against the same vLLM endpoint. A shared state dictionary (guarded by a threading.Lock) exposes the main streamās growing chain-of-thought to each fork, so successive probes condition on more context; the speculative tool runs in a worker thread via asyncio.to_thread while the main stream continues decoding. Concurrent agent episodes, each contributing a main stream and a fork, run against vLLM without observable throughput degradation (§6). By design, only tools marked read-only in a manifest are speculated. Write and non-idempotent tools are not supported yet and always take the serial fallback path; supporting them is feasible with sandboxing or agent-state checkpointing that can roll back a mispredicted speculative call. Our prototype evaluates read-only benchmarks (web search, page browse, Wikipedia, DB-read), where the strict gate already bounds correctness on every turn. Deployment overhead. The controller adds ā¤0.05⤠0.05 s of Python-level overhead per turn (thread management, JSON parsing, lock acquisition) on top of the vLLM call latency. This overhead is accounted for in TohT_oh and is visible in the EQ1 validation (§6.2). Agent integration. In our agent harness, each turn wraps the stock ReAct-style loop: the main stream runs through the chat API with tools registered in the request; the fork thread uses the same prompt prefix plus a forced <tool_call> opener and a JSON schema the model was fine-tuned on. Tool results are injected back into the chat history exactly as in the serial baseline, so correctness reduces to matching the eventual tool name and arguments at turn end. No changes to model weights or offline datasets are required: only serving configuration (prefix caching on, logprobs enabled) and the controller process colocated with the agent. 6. Evaluation 6.1. Setup Models, Hardware, Frameworks. We evaluate Qwen3-32B and Qwen3-4B (dense) (Yang et al., 2025a) and Qwen3.5-35B-A3B (a 3B-active mixture-of-experts) (Qwen Team, 2026) to test generalization across both scale and architecture, all self-hosted via vLLM 0.19.1 (TP=1) on NVIDIA H20-3e GPUs (143 GiB each) with bf16 precision. All runs use greedy decoding (temperature 0, seed 42) with thinking mode enabled, producing 2ā20 s of chain-of-thought decode that creates the overlap budget Spork exploits. Benchmarks. Table 2 (Appendix E) profiles the three benchmarks. GAIA and HotpotQA execute real network tools (web search + page browse, and a Wikipedia search API, respectively); tau2-bench executes structured DB-query tools with simulated latency floors, giving a controlled sweep over TtoolT_tool. Metrics. Latency: per-query end-to-end wall-clock time. Quality: exact-match accuracy (EM) for GAIA/HotpotQA, F1 for HotpotQA. Speedup: P95P_95 ratio (baseline / Spork). Configurations. Per-model serving mode, baseline, and Spork configuration appear in Table 3 (Appendix E); the dense-model baseline is vLLMās built-in ngram speculative decoding (Saxena, 2023), and each model is compared against its own baseline under the same serving mode. 6.2. End-to-End Speedup Figure 10 reports end-to-end P95P_95 latency across models and benchmarks. With Qwen3-32B, Spork achieves ā-18% P95P_95 on GAIA (131.9 s ā 108.1 s, real web search, Ttool=4.8T_tool=4.8 s), while HotpotQA shows a modest 1.06Ć improvement because its short tools (Ttool=0.9T_tool=0.9 s) sit near the EQ1 break-even. The benefit transfers across model size: Qwen3-4B achieves 1.15Ć mean speedup on tau2-bench (simulated 2 s tool floor), with full benchmark coverage in Appendix G. It also transfers across model architecture: Qwen3.5-35B-A3B reduces P95P_95 by 16% on tau2 (59.4 s ā 49.8 s, N=155), 16% on GAIA (222.4 s ā 187.3 s, N=53, tail-positive), and 20% on HotpotQA (18.9 s ā 15.2 s, N=200). Figure 11. Spork preserves task accuracy. Paired comparison of baseline (teal) vs. Spork (orange) exact-match scores across benchmarks and model sizes. Quality stays within 1 p of baseline on all settings and sometimes improves (+1.5 p EM on HotpotQA with Qwen3-32B). The strict gate prevents speculative side effects from corrupting conversation history. Spork achieves ā-18% P95P_95 on GAIA (Qwen3-32B) and 1.15Ć mean on tau2 (Qwen3-4B), confirming that the overlap window dominates when tool latency ā„ 2 s. §6.6 locates these points in the EQ1 envelope. 6.3. Quality Preservation We expect Spork to preserve task accuracy: the strict gate ensures only exact-matching tool calls are accepted, and D3 recycles only the rejected probeās verified prefix as a speculative-decoding draft. Speculative decodingās verification step preserves the target modelās output distribution (Leviathan et al., 2023); in Spork the target is the running model itself, so D3 inherits this guarantee. Figure 11 shows that Spork keeps EM within 1 p of baseline on every benchmark and model, and sometimes improves it. tau2-bench is not plotted because its quality is identical by construction: the simulated DB tools return deterministic canned results, matching exactly for baseline and Spork, so the latency sweep exercises timing only. GAIA and HotpotQA instead call search APIs whose results are non-deterministic even under identical arguments, so baseline and Spork trajectories cannot be token-identical despite temperature 0; quality on these benchmarks is therefore compared via aggregate scores (EM, F1) rather than per-token equality. On HotpotQA with Qwen3-32B, Spork is faster and higher quality (+1.5 p EM, +0.012 F1) than the baseline. Spork preserves accuracy (within ā-1 p of baseline) regardless of model size or benchmark. The strict gate is the key mechanism: it rejects any probe that does not exactly match the main generation, ensuring the conversation history is never corrupted by speculative side effects. 6.4. Ablation We expect each component (D1, D2, D3) to contribute independently to the final result, with no single configuration dominating all metrics simultaneously. Figure 12. Ablation: full Spork (D1+D2+D3) achieves the best tail latency. GAIA N=165, Qwen3-32B; black dashes mark P50P_50. Ngram alone gives ā-31% at P50P_50; D1+D2+D3 adds ā-18% at P95P_95. No single config wins on all metrics: D1 for mean, D1+D2+D3 for P95P_95. Figure 12 decomposes contributions: ⢠Ngram spec-dec (the baseline): accelerates decode TPOT (P50P_50 50.5ā 34.7 s, ā-31%); every Spork configuration stacks on top of it, so the contributions below are orthogonal to token-level speculative decoding. ⢠D1 on ngram: additional P50P_50 ā-12%, mean ā-10%. Fork-prefix overlap adds on top of per-token spec-dec. ⢠D1+D2+D3: best P95P_95 (108.1 s, ā-18%). D2+D3 synergy helps tail queries where rejected retries leave verified prefixes that D3 recycles as drafts. Per-configuration EM differences exist but stem from search-API and vLLM-batching nondeterminism, both present in the baseline as well (§6.3); they do not originate from the gate: D3 accepts only tokens that match the main modelās greedy output, so trajectory divergence is orthogonal to Sporkās correctness guarantee. D2 confidence gate analysis. D2 is a trade-off knob. The gate raises per-turn acceptance from αā0.22αā 0.22 (D1, 49 accepted probes) to αā0.37αā 0.37 (D1+D2, 101 accepted) and avoids wasted real tool calls by filtering low-confidence probes before dispatch; the dispatched-probe totals (434 for D1, 2124 for D1+D2) exceed the turn counts because the retry cadence re-probes low-confidence turns (§2.2). The confidence threshold Īø=0.90Īø=0.90 was selected by sweeping min-prob over 997 probes: precision 88%, recall 100%, F1==0.937 (Figure 8). This filtering pays when tools are slow or costly; when tools are cheap, the added wait can outweigh the savings, and D2 alone can hurt mean latency on short-tool benchmarks. With D3, the gateās rejected retries become useful drafts, and D2+D3 together give the best P95P_95: the full configuration wins the tail while D1 alone wins the mean. Each component contributes to different metrics: D1 for mean/P50, D2+D3 for P95P_95. All three are needed for the best tail latency. 6.5. Case Study: Self-Speculation vs. a Draft Model Figure 13. Self-speculation trades a slower probe for higher useful overlap. D1-only HTTP ablations with the same Qwen3-32B main model. (a) Self has higher exact-argument agreement than a Qwen3-4B cross probe on all three benchmarks. (b) On tau2, self hides more mean tool latency across floors; cross wastes 2.1Ć more speculative tools at the 2 s floor. A natural alternative to self-speculation is to serve a smaller model as a tool-call drafter, analogous to token-level speculative decoding. Figure 13 compares that choice against self-speculation in a controlled D1-only HTTP ablation on tau2-bench: same Qwen3-32B main model, tasks, seed, parser, and simulated tool floor; only the probe model changes. The cross-4B probe has lower probe latency, the time to decode the tool-call token span: 0.315 s mean vs. 0.970 s for self at the tau2 2 s floor. Its exact-argument agreement, however, is far lower than self-speculationās: 0.590 vs. 0.792 on tau2 (Figure 13a), 0.161 vs. 0.296 on HotpotQA, and 0.056 vs. 0.109 on GAIA. Since both probes usually finish before the tool floor on accepted turns, probe speed is off the critical path; acceptance rate dominates realized overlap. Self-speculation therefore hides more tool latency at every floor in the sweep and wastes fewer speculative tool executions (154 vs. 317 at N=155; Figure 13b). It also avoids a second served model: in our measurement the cross-model arm serves two models (Qwen3-32B and the Qwen3-4B drafter, each TP=1 on its own GPU), while self-speculation needs only the target modelās single GPU; a colocated drafter would instead compete with the target for request scheduling and KV-cache capacity. This result places a boundary on external drafters: they help only if their lower probe cost offsets the lower α and the additional serving contention. 6.6. Operating Envelope We expect Spork to break even when tool latency exceeds probe overhead (Ttoolā„TohT_toolā„ T_oh, the probe overhead of §4.1) and the model produces sufficient CoT for overlap. Below this threshold, probe overhead dominates and Spork is neutral or negative. EQ1 calibration. The tau2 latency sweep validates EQ1ās structural prediction: speedup grows monotonically with tool latency (1.09Ć at 0.5 s floor ā 1.18Ć at 5.0 s floor on airline domain). EQ1 predicts observed speedup with at most 1.84% residual across all operating points (1.84% on tau2, 0.88% on GAIA), confirming the cost model is a reliable predictor of when Spork helps. Figure 14. Spork stays at or above break-even across the operating envelope. Mean per-task speedup vs. tool latency on the tau2 sweep (N=155). Full Spork (D1+D2+D3) stays above break-even and rises with tool latency; the gated D1+D2 line approaches break-even on the shortest floors, matching EQ1ās αā toverlapā„Tohα· t_overlapā„ T_oh condition. Stars mark the two real-tool operating points (Qwen3-32B mean speedup): near-neutral HotpotQA at EQ1ās break-even, GAIA well inside the beneficial region. Figure 14 traces this envelope on the controlled tau2 sweep. The real-tool benchmarks land where the cost model places them: GAIAās 4.8 s tool latency sits well inside the beneficial region, while HotpotQAās 0.9 s latency sits near break-even, exactly where EQ1 predicts the gains should shrink toward neutrality. When Spork wins. The largest gains appear when TtoolT_tool is large relative to probe overhead, tool schemas are regular enough for high α, and speculation is restricted to read-only tools (§5): GAIAās multi-second real web search is the clearest favorable regime (P95P_95 ā-18%). On HotpotQA (0.9 s Wikipedia API), Spork is Pareto-dominant with Qwen3-32B (faster and higher EM) and Pareto-neutral with Qwen3-4B: it never hurts, and occasionally helps on tail queries. The case study (§6.5) reinforces the principle: once the probe finishes before the tool floor, acceptance rate, not probe speed, determines hidden tool latency. When Spork does not help. The primary boundary is a vanishing overlap window (toverlapā0t_overlapā 0): from the decode side, no-think mode cuts main decode to 0.3ā0.9 s, too fast for the probe to finish before the main stream needs the tool result (0.79Ć on tau2), while thinking-mode CoT, typically ā„ 2 s in our workloads, creates the window; from the tool side, very fast tools (e.g., local file reads) finish before any probe can pay off, leaving nothing to hide. The window also bounds long tools: toverlapt_overlap ends when either side finishes, so tool time that runs past the main streamās stop is not hidden, and the relative gain saturates once TtoolT_tool exceeds the decode remaining at dispatch. A second boundary, a native tool-call format that diverges from the forced probe (αā0αā 0), is detailed in Appendix F. Sporkās operating boundary is well-defined: thinking-mode CoT long enough to host the probe, and tool latency above the probe-overhead break-even, with gains growing until TtoolT_tool fills the available decode window. Practitioners can evaluate applicability before full deployment with a lightweight pilot trace measuring fork name accuracy, α, and TtoolT_tool, plugged into EQ1. 7. Related Work Token-level speculative decoding. Speculative decoding (Leviathan et al., 2023) accelerates generation with draft models or extra heads that propose tokens the target model verifies in parallel (Li et al., 2024b, a, 2025; Cai et al., 2024; An et al., 2026); model-free variants draft from prompt or prior-output n-grams (Saxena, 2023; Oliaro et al., 2024). Both families shorten the decode portion of a turn but leave the tool wait untouched: the external call is still issued only after the model finishes emitting it. Spork is orthogonal: our partial-token accept reuses vLLMās speculative-decoding verification path to recycle a rejected probeās verified prefix as draft tokens, so the two mechanisms stack. Action- and execution-level speculation. Speculative Actions (Ye et al., 2026), SPAgent (Huang et al., 2025), DualSpec (Zhong et al., 2026), and PASTE (Sui et al., 2026) move speculation up to the action level, pre-executing likely tool calls; Nichols et al. (2025) speculate tool calls inside the serving engine, targeting agent-serving throughput. They differ from Spork in the source of the prediction: Speculative Actions and DualSpec rely on auxiliary predictor models or verifier policies, Nichols et al. (2025) draft with a smaller speculative model, and PASTE mines recurring tool-call patterns from historical traces and replays them through a middleware interceptor. IdleSpec (Choi et al., 2026) fills the same tool-wait windows with speculative plans, improving task accuracy rather than hiding tool latency. Spork instead forks the running modelās own state and reads its probe logprobs: day-one speculation with no auxiliary model, no trace collection, and the prefix-cache sharing an interceptor lacks. Atomix (Mohammadi et al., 2026) adds complementary transactional semantics for speculative side effects, which our read-only-tool restriction sidesteps. Early intent in LLMs. Certaindex (Fu et al., 2024a) finds that reasoning programsā intermediate answers stabilize well before generation ends and allocates test-time compute from that certainty; When2Tool (Sun et al., 2026) shows tool-call necessity is linearly decodable from pre-generation hidden states. DEER (Yang et al., 2026) induces an early final answer to stop reasoning, and SpecExit (Yang et al., 2025b) reads exit signals from draft-model hidden states; inspired by this line of work, Spork inserts a forced tool-call prefix to surface the next tool call early and overlap the tool wait with the remaining decode. 8. Conclusion We presented Spork. Our central finding is that agentic LLM workloads exhibit exploitable early intent: the running model already knows its next tool call before finishing its reasoning, with fork-at-start name accuracy of 74.6ā99.6% on Qwen3-32B against a tool wait of 16ā37% of agent wall time. Spork turns this into latency with the lightest possible machinery, requiring no draft model, no historical traces, and no retraining: D1āD3 map the early-intent signals onto EQ1ās terms, and our experiments validate the cost model on real web tools (GAIA), short-latency APIs (HotpotQA), and the controlled tau2 latency sweep. The mechanism generalizes across scale and architecture, from Qwen3-4B to Qwen3-32B and a 3B-active mixture-of-experts, while preserving task accuracy. Limitations. Spork speculates only read-only tools today; supporting write operations is future work, requiring sandboxing or agent-state checkpointing so a mispredicted speculative call can be rolled back. It relies on open-weight models served by an engine that exposes per-token logprobs and a completion endpoint; closed-source APIs cannot host the fork thread. Finally, we target latency-oriented serving with spare capacity; under heavy batching, probe traffic competes with foreground decode and the overlap benefit shrinks. Future work. Sandboxed or checkpointed execution would extend speculation to write tools: a mispredicted call rolls back to the pre-speculation agent state. Deeper engine integration could schedule probes into idle batch slots, shrinking TohT_oh further and widening the favorable regime. References Agrawal et al. [2024] A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. S. Gulavani, A. Tumanov, and R. Ramjee. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, 2024. An et al. [2026] Z. An, H. Bai, Z. Liu, D. Li, and E. Barsoum. PARD: Accelerating LLM inference with low-cost PARallel draft model adaptation. In International Conference on Learning Representations (ICLR), 2026. URL https://arxiv.org/abs/2504.18583. Anthropic [2025] Anthropic. Claude code. https://w.claude.com/product/claude-code, 2025. Barres et al. [2025] V. Barres, H. Dong, S. Ray, X. Si, and K. Narasimhan. Ļ2Ļ^2-bench: Evaluating conversational agents in a dual-control environment, 2025. URL https://arxiv.org/abs/2506.07982. Cai et al. [2024] T. Cai, Y. Li, Z. Geng, H. Peng, J. D. Lee, D. Chen, and T. Dao. Medusa: Simple LLM inference acceleration framework with multiple decoding heads. International Conference on Machine Learning, 2024. Choi et al. [2026] D. Choi, K. Park, W. Song, S. Dingliwal, S. M. Jayanthi, J. Shin, and A. Galstyan. IdleSpec: Exploiting idle time via speculative planning for LLM agents, 2026. URL https://arxiv.org/abs/2605.22154. Fu et al. [2024a] Y. Fu, J. Chen, S. Zhu, Z. Fu, Z. Dai, Y. Zhuang, Y. Ma, A. Qiao, T. Rosing, I. Stoica, and H. Zhang. Efficiently scaling LLM reasoning with certaindex, 2024a. URL http://arxiv.org/abs/2412.20993. Fu et al. [2024b] Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, and L. Mai. ServerlessLLM: Low-latency serverless inference for large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation, pages 135ā153. USENIX Association, 2024b. GitHub [2025] GitHub. GitHub Copilot. https://github.com/features/copilot, 2025. Hu et al. [2025] J. Hu, J. Xu, Z. Liu, Y. He, Y. Chen, H. Xu, et al. DeepServe: Serverless large language model serving at scale, 2025. URL https://arxiv.org/abs/2501.14417. Huang et al. [2025] Z. Huang, W. Zeng, T. Fu, T. Liu, Y. Sun, K. Hong, X. Yang, C. Liu, Y. Li, Q. Zhang, G. Dai, Z. Zhu, and Y. Wang. Reducing latency of LLM search agent via speculation-based algorithm-system co-design, 2025. URL http://arxiv.org/abs/2511.20048. Jimenez et al. [2024] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan. SWE-bench: Can language models resolve real-world GitHub issues? International Conference on Learning Representations, 2024. Kwon et al. [2023] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611ā626, 2023. Leviathan et al. [2023] Y. Leviathan, M. Kalman, and Y. Matias. Fast inference from transformers via speculative decoding. In International Conference on Machine Learning, pages 19274ā19286, 2023. Li et al. [2024a] Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE-2: Faster inference of language models with dynamic draft trees, 2024a. URL https://arxiv.org/abs/2406.16858. Li et al. [2024b] Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE: Speculative sampling requires rethinking feature uncertainty. International Conference on Machine Learning, 2024b. Li et al. [2025] Y. Li, F. Wei, C. Zhang, and H. Zhang. EAGLE-3: Scaling up inference acceleration of large language models via training-time test, 2025. URL https://arxiv.org/abs/2503.01840. Lin et al. [2024] C. Lin, Z. Han, C. Zhang, Y. Yang, F. Yang, C. Chen, and L. Qiu. Parrot: Efficient serving of llm-based applications with semantic variable. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation. USENIX Association, 2024. Mahgoub et al. [2022] A. Mahgoub, E. B. Yi, K. Shankar, S. Elnikety, S. Chaterji, and S. Bagchi. ORION and the three rights: Sizing, bundling, and prewarming for serverless DAGs. In 16th USENIX Symposium on Operating Systems Design and Implementation, pages 303ā320. USENIX Association, 2022. Manus [2025] Manus. Manus: Hands on ai. https://manus.im/, 2025. Mialon et al. [2024] G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom. GAIA: A benchmark for general AI assistants. In International Conference on Learning Representations (ICLR), 2024. URL https://arxiv.org/abs/2311.12983. Mohammadi et al. [2026] B. Mohammadi, N. Potamitis, L. Klein, A. Arora, and L. Bindschaedler. Atomix: Timely, transactional tool use for reliable agentic workflows, 2026. URL http://arxiv.org/abs/2602.14849. Nichols et al. [2025] D. Nichols, P. Singhania, C. Jekel, A. Bhatele, and H. Menon. Optimizing agentic language model inference via speculative tool calls, 2025. URL http://arxiv.org/abs/2512.15834. Oliaro et al. [2024] G. Oliaro, Z. Jia, D. Campos, and A. Qiao. SuffixDecoding: Extreme speculative decoding for emerging AI applications, 2024. URL https://arxiv.org/abs/2411.04975. OpenAI [2025] OpenAI. Introducing deep research. https://openai.com/index/introducing-deep-research/, Feb. 2025. Patil et al. [2025] S. G. Patil, H. Mao, F. Yan, C. C.-J. Ji, V. Suresh, I. Stoica, and J. E. Gonzalez. The berkeley function calling leaderboard (BFCL): From tool use to agentic evaluation of large language models. In International Conference on Machine Learning (ICML), 2025. Qwen Team [2026] Qwen Team. Qwen3.5: Towards native multimodal agents, 2026. URL https://qwen.ai/blog?id=qwen3.5. Saxena [2023] A. Saxena. Prompt lookup decoding, 2023. URL https://github.com/apoorvumang/prompt-lookup-decoding. Model-free n-gram drafting; basis of vLLMās ngram speculative decoding. Schick et al. [2024] T. Schick, J. Dwivedi-Yu, R. DessƬ, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2024. Song et al. [2024] Y. Song, Z. Mi, H. Xie, and H. Chen. PowerInfer: Fast large language model serving with a consumer-grade GPU. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP), 2024. URL https://arxiv.org/abs/2312.12456. Stojkovic et al. [2023] J. Stojkovic, T. Xu, H. Franke, and J. Torrellas. SpecFaaS: Accelerating serverless applications with speculative function execution. In 2023 IEEE International Symposium on High-Performance Computer Architecture, pages 814ā827, 2023. doi: 10.1109/HPCA56546.2023.10071120. Sui et al. [2026] Y. Sui, H. Zhao, R. Ma, Z. He, H. Wang, J. Li, and Y. Yang. Act while thinking: Accelerating LLM agents via pattern-aware speculative tool execution, 2026. URL http://arxiv.org/abs/2603.18897. Sun et al. [2026] C.-E. Sun, L. Liu, G. Yan, Z. Wang, and T.-W. Weng. LLM agents already know when to call tools ā even without reasoning, 2026. URL https://arxiv.org/abs/2605.09252. Wei et al. [2025] J. Wei, Z. Sun, S. Papay, S. McKinney, J. Han, I. Fulford, H. W. Chung, A. T. Passos, W. Fedus, and A. Glaese. BrowseComp: A simple yet challenging benchmark for browsing agents, 2025. URL https://arxiv.org/abs/2504.12516. Yang et al. [2025a] A. Yang et al. Qwen3 technical report, 2025a. URL https://arxiv.org/abs/2505.09388. Yang et al. [2026] C. Yang, Q. Si, Y. Duan, Z. Zhu, C. Zhu, Q. Li, M. Chen, Z. Lin, and W. Wang. Dynamic early exit in reasoning models. In International Conference on Learning Representations (ICLR), 2026. doi: 10.48550/arXiv.2504.15895. URL https://arxiv.org/abs/2504.15895. Yang et al. [2025b] R. Yang, H. Bai, S. Liu, G. Yu, et al. SpecExit: Accelerating large reasoning model via speculative exit, 2025b. URL https://arxiv.org/abs/2509.24248. Yang et al. [2018] Z. Yang, P. Qi, S. Zhang, Y. Bengio, W. W. Cohen, R. Salakhutdinov, and C. D. Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Conference on Empirical Methods in Natural Language Processing, 2018. Yao et al. [2023] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations, 2023. Ye et al. [2026] N. Ye, A. Ahuja, G. Liargkovas, Y. Lu, K. Kaffes, and T. Peng. Speculative actions: A lossless framework for faster agentic systems. 2026. doi: 10.48550/arXiv.2510.04371. URL http://arxiv.org/abs/2510.04371. Zheng et al. [2024] L. Zheng, L. Yin, Z. Xie, J. Huang, C. Sun, C. H. Yu, S. Cao, C. Kober, Y. Sheng, J. E. Gonzalez, I. Stoica, and H. Zhang. SGLang: Efficient execution of structured language model programs. In Advances in Neural Information Processing Systems, 2024. Zhong et al. [2026] S. Zhong, B. Lu, Q. Chen, C. Liu, F. Yang, and M. Li. DualSpec: Accelerating deep research agents via dual-process action speculation, 2026. URL http://arxiv.org/abs/2603.07416. Appendix A Full EQ1 Derivation Let a single agent turn have: ⢠TdecT_dec: main generation wall time (post-prefill, includes CoT and tool-call decode) ⢠TtoolT_tool: tool execution wall time ⢠Tbase=Tdec+TtoolT_base=T_dec+T_tool: serial baseline cost Under Spork, let α be the gate acceptance rate and toverlapt_overlap be the mean realized overlap on accepted turns. On accepted turns (p=αp=α), the speculative tool execution overlaps toverlapt_overlap seconds of tool time with decode: Thit=TbaseātoverlapT_hit=T_base-t_overlap On rejected turns (p=1āαp=1-α), the agent falls back to serial execution with overhead TohT_oh, but prefix recovery (D3) lets the main stream skip re-decoding the recovered tool-call tokens, replacing TbaseT_base with the realized base cost Tbaseā=TbaseāTD3T_base^*=T_base-T_D3 (Tbaseā=TbaseT_base^*=T_base without D3): Tmiss=Tbaseā+TohT_miss=T_base^*+T_oh Expected per-turn cost: ā[T] [T] =αā Thit+(1āα)ā Tmiss =α· T_hit+(1-α)Ā· T_miss =Tbaseāαā toverlap+(1āα)ā(TbaseāāTbase+Toh) =T_base-α· t_overlap+(1-α)(T_base^*-T_base+T_oh) Latency ratio: S=TbaseTbaseāαā toverlap+(1āα)ā (TbaseāāTbase+Toh)S= T_baseT_base-α· t_overlap+(1-α)Ā·(T_base^*-T_base+T_oh) For notational simplicity, the main text uses TbaseāT_base^* in place of TbaseT_base on the speculative path and TohT_oh as a uniform overhead (applied regardless of accept/reject), giving the formula Tbase/(Tbaseāāαā toverlap+Toh)T_base/(T_base^*-α· t_overlap+T_oh). The floor-2s validation in §6.6 uses this simplified form. Break-even condition. Sā„1Sā„ 1 requires αā toverlapā„(1āα)ā(TbaseāāTbase+Toh)α· t_overlapā„(1-α)(T_base^*-T_base+T_oh); without D3 (Tbaseā=TbaseT_base^*=T_base) this reduces to αā toverlapā„(1āα)ā Tohα· t_overlapā„(1-α)Ā· T_oh, or approximately αā toverlapā„Tohα· t_overlapā„ T_oh when TohT_oh is treated as a fixed overhead. Tool-fraction ceiling. When α=1α=1 and toverlap=Ttoolt_overlap=T_tool (perfect overlap of all tool time): Smax=Tdec+TtoolTdec=11āftoolS_ = T_dec+T_toolT_dec= 11-f_tool where ftool=Ttool/Tbasef_tool=T_tool/T_base is the tool-time fraction. On BrowseComp (ftool=0.366f_tool=0.366): Smax=1.58ĆS_ =1.58Ć. The gap between SmaxS_ and realized speedup is explained by (a) toverlap<Ttoolt_overlap<T_tool (probe dispatch lag + turns with Ttool<TprobeT_tool<T_probe) and (b) Toh>0T_oh>0 on strict mode. Appendix B D3 Partial-Token Accept: Phase 0 Data Offline analysis of 1268 probeābaseline pairs from BrowseComp-120, a 120-question probe-analysis run (distinct from the 240-question end-to-end evaluation run described in Appendix C). Both arms decode under identical observations, so the prefix statistics are unbiased: Statistic Value Mean first-mismatch position k 29.5 tokens Median k 27 tokens P25P_25 / P75P_75 17 / 35 tokens % pairs with kā„10kā„ 10 82% % pairs with kā„20kā„ 20 61% Even rejected probes share a substantial token prefix with the baseline. Accepting k=29.5k=29.5 tokens and autoregressively decoding the remaining 50ā29.5=20.550-29.5=20.5 tokens reduces the decode time of a full ā¼50 50-token tool call from ā¼1.5 1.5 s to ā¼0.9 0.9 s (verify pass + 20.5 tokens), saving ā¼0.6 0.6 s on such rejected turns. In this 120-question run, 52% of strict-gate turns were rejected, giving a mean per-turn D3 saving of ā0.31ā 0.31 s. Appendix C BrowseComp Experimental Details Setup. Qwen3-32B via vLLM 0.18.1, TP=4 on 4Ć NVIDIA H20-3e (143 GiB each), bf16, max_model_len=40960. BrowseComp first 240 questions (indices 0ā239, fixed seed). Temperature 0, seed 42 on all calls. MAX_TURNS = 15. Workers = 4 (4 questions in flight). Tools: a production web-search API plus a web-page crawl/browse API. First-token timeout. 3.4% of turns (50/1487) had probe dispatch timeout (first-token threshold = 3s). These are treated as aborted probes; the turn falls back to serial execution. Not catastrophic for overall results but indicates that under workers=4 concurrency, probe latency occasionally exceeds the 3s window on long real-content prompts. D2 (confidence gating with earlier abort) would reduce this failure mode. Fork statistics. Metric Value Total tool turns 1487 Name hit rate 83.7% (1245/1487) Args-exact rate 38.8% (577/1487) Spec used (strict) 577 (38.8%) Spec wasted (strict reject) 687 (46.2%) First-token timeouts 50 (3.4%) Appendix D BrowseComp Timing Decomposition Wall-time decomposition of the April HTTP-mode runs of Appendix C; the later engine-mode rerun shows the same tool share (37%) but a larger prefill share (it lacks the HTTP pathās shared prefix cache). Component Fraction of TbaseT_base Mean (s) ftoolf_tool (tool execution) 36.6% 8.92 fdecf_dec (decode: CoT + tool_call) 57.2% ā CoT decode 54.6% ā Tool-call decode 2.6% ā Other (prefill, overhead) 6.2% ā Realized toverlapt_overlap (spec_used turns) ā 1.03 Probe dispatch latency (mean probe_wall) ā 0.72 The 36% turns with Ttool<0.72T_tool<0.72 s (Ā” probe_wall) have zero realizable overlap: the tool completes before the probe can even dispatch speculative execution. The 47% turns with Ttoolā[0.72,2]T_toolā[0.72,2] s have limited overlap (ā¤1.28⤠1.28 s). Only the 17% with Ttoolā„2T_toolā„ 2 s achieve substantial overlap. Hence mean toverlap=1.03t_overlap=1.03 s despite mean Ttool=8.92T_tool=8.92 s: 83% of turns fall in the first two buckets. Appendix E Per-Model Serving Configurations Table 2 profiles the three benchmarks of §6.1; Table 3 lists the serving mode, baseline, and Spork configuration for every model in §6. Table 2. Benchmark and tool profile. Tool vocabulary size, typical tool latency, and turn depth per benchmark. Benchmark N #Tools Typical TtoolT_tool Turns GAIA 165 2 4.8 s mean (real) 1ā8 HotpotQA 200 1 0.9 s mean (real) multi-hop tau2-bench 155 30 (2 dom.) 0.5ā5 s floors (sim.) multi-turn The HTTP configuration drives a stock vLLM server through its OpenAI-compatible API and supports D1+D2 only; the engine configuration adds the D3 proposer integration inside vLLM (§4.3). Qwen3.5-35B-A3B uses JSON-format tool prompting throughout, as its native XML tool-call format is a self-speculation boundary case (§6.6). Table 3. Per-model serving mode, baseline, and Spork configuration. Model Serving mode Baseline Spork Qwen3-32B engine ngram spec. dec. D1+D2+D3 Qwen3-4Bā engine ngram spec. dec. D1+D2+D3 Qwen3.5-35B-A3Bā” HTTP (JSON) serial D1+D2 ā The tau2 run uses HTTP: serial baseline, D1+D2. ā” The GAIA run uses D1 only. All end-to-end runs use greedy decoding with a fixed seed (42). Appendix F Format-Divergence Boundary A model whose native tool-call format diverges from the forced probe defeats self-speculation by collapsing the acceptance term (αā0αā 0). Qwen3.5-35B-A3Bās native XML predicts the end-goal tool rather than the next intermediate one, collapsing fork-at-start name accuracy to αname=4.5% _name=4.5\% despite 97% parse success; this is a semantic next-step mismatch, not a parser failure, and it motivates the JSON-format prompting of Table 3. Appendix G Qwen3-4B Full-Benchmark Coverage Figure 15 and Table 4 report the full Qwen3-4B end-to-end coverage. These are scaling-boundary diagnostics, not headline speedup claims: on the smaller model the engine D1+D2+D3 protocol is near-neutral on HotpotQA and marginally positive on GAIA, while the tau2 run (2 s simulated tool-latency floor) shows a modest gain. Even though Qwen3-4B in think mode produces long chains of thought (mean 3,323 CoT tokens on GAIA, 1,843 on HotpotQA), the smaller model decodes quickly, so the overlap window and D3 recovery yield limited end-to-end benefit. Figure 15. Qwen3-4B P95P_95 latency across benchmarks (scaling boundary). Spork vs. baseline (GAIA/HotpotQA: engine D1+D2+D3 vs. ngram spec-dec; tau2: HTTP D1+D2 vs. serial). tau2 sees a modest gain, GAIA is marginally positive, and HotpotQA is neutral. Quality is preserved or improved (GAIA EM 29ā36/16529ā 36/165; HotpotQA EM 40ā39/20040ā 39/200, within ā-1 p). Benchmark N Baseline P95P_95 Spork P95P_95 Speedup GAIA 165 29.95 s 28.67 s 1.03Ć (agg) HotpotQA 200 17.36 s 17.39 s 1.00Ć (agg) tau2 155 44.81 s 43.04 s 1.15Ć (mean)ā Table 4. Qwen3-4B end-to-end results. Latency columns are P95P_95; the GAIA and HotpotQA speedups (agg) are aggregate wall-clock ratios on real tools. ā tau2 reports the mean speedup at the 2 s simulated tool-latency floor (21.2ā 18.5 s), not a P95P_95 ratio. Appendix H Late-Probe Supersession (continue_after_dispatch) D2 issues a probe after each additional block of CoT, and Insight 2 shows that probe accuracy rises as the chain-of-thought unfolds. This raises a design choice: should the controller commit to the first probe that clears the gate, or keep probing and upgrade its speculation when a later, higher-confidence probe predicts a different call? We expose this as a single switch, continue_after_dispatch: ⢠OFF (default): the first probe that clears the confidence gate dispatches the speculative tool and the retry loop stops (first commit wins). ⢠ON: the controller keeps probing after dispatch; if a later probe predicts a different call, it cancels the in-flight speculative tool and re-dispatches the new candidate (a supersession). Table 5 reports a controlled A/B on tau2 (Qwen3-32B, HTTP mode, 2 s tool floor, N=30N=30, a subset of the 155-task suite, seed 42) using deterministic recordā tool execution so that tool latency and output are identical across arms (zero cache-miss fallbacks). Supersession is mildly beneficial: it raises the gate acceptance rate (α 0.714ā0.7620.714ā 0.762) and roughly halves wasted speculative dispatches (28ā1528ā 15) through 36 supersessions, with no latency regression (mean wall 80.1ā77.580.1ā 77.5 s, within noise). Config α (used/disp.) wasted disp. superseded baseline ā 0 0 OFF (first commit) 0.714 28 0 ON (supersede) 0.762 15 36 Table 5. Effect of continue_after_dispatch on tau2 (Qwen3-32B, HTTP, 2 s floor, N=30N=30, deterministic replay). The effect is small enough that we leave continue_after_dispatch as an optional knob rather than a headline mechanism, and the main-paper results use the default behavior. We do not report a GAIA latency A/B for this switch: with real free-text search, the main streamās queries diverge across arms (no API seed under concurrent main/probe decode), so a clean replay is not possible in HTTP mode; an engine-mode deterministic harness would be required.