Paper deep dive
Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling
Jeffrey T. H. Wong, Zixi Zhang, Junyi Liu, Yiren Zhao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/21/2026, 1:36:55 AM
Summary
The paper introduces Team-of-Thoughts, a heterogeneous Multi-Agent System (MAS) framework that utilizes an orchestrator to dynamically select and coordinate specialized tool agents based on self-assessed profiles. This approach aims to improve test-time scaling by leveraging diverse model expertise for mathematical reasoning and code generation tasks, outperforming homogeneous baselines.
Entities (8)
Relation Signals (7)
Team-of-Thoughts â achievesaccuracyon â LiveCodeBench
confidence 98% ¡ on ... LiveCodeBench, Team-of-Thoughts achieves ... 77.91% accuracy
Team-of-Thoughts â achievesaccuracyon â AIME24
confidence 98% ¡ on AIME24 ... Team-of-Thoughts achieves 96.00% ... accuracy
Team-of-Thoughts â uses â Orchestrator Calibration
confidence 95% ¡ Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration...
Team-of-Thoughts â uses â Agent Self-Assessment
confidence 95% ¡ Team-of-Thoughts introduces two novel components: ... (2) Agent Self-Assessment...
Team-of-Thoughts â outperforms â homogeneous role-play baselines
confidence 92% ¡ significantly improving over homogeneous role-play baselines
DeepSeek-V3.2 â servesas â Orchestrator
confidence 90% ¡ we utilize DeepSeek v3.2 as the orchestrator for mathematical reasoning
gpt-5-mini â servesas â Orchestrator
confidence 90% ¡ we utilize ... GPT-5 Mini for code generation.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We propose Team-of-Thoughts, a heterogeneous MAS framework that treats diverse models as specialized tools within an orchestrator-driven paradigm. Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration, which identifies models with superior coordination and synthesis capabilities, and (2) Agent Self-Assessment, a protocol where tool agents profile their own domain-specific strengths to guide selection. At inference, the orchestrator dynamically activates the most compatible agents based on these profiles to maximize capability coverage. Across five mathematical reasoning and code generation benchmarks, Team-of-Thoughts consistently outperforms individual models and existing MAS baselines. Notably, on AIME24 and LiveCodeBench, Team-of-Thoughts achieves 96.00% and 77.91% accuracy, respectively, significantly improving over homogeneous role-play baselines (80.00% and 65.93%).
Tags
Links
- Source: https://arxiv.org/abs/2602.16485v2
- Canonical: https://arxiv.org/abs/2602.16485v2
Trouble viewing inline? Open PDF directly â
Full Text
59,993 characters extracted from source content.
Expand or collapse full text
Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling Jeffrey T. H. Wong1,â Zixi Zhang1,â Junyi Liu2 Yiren Zhao1 1Imperial College London, 2Microsoft Research tsz.wong20,b.zhang25,a.zhao@imperial.ac.uk junyi.liu@microsoft.com Abstract Existing Multi-Agent Systems (MAS) typically rely on homogeneous model configurations, failing to exploit the diverse expertise inherent in different post-trained architectures. We propose Team-of-Thoughts, a heterogeneous MAS framework that treats diverse models as specialized tools within an orchestrator-driven paradigm. Team-of-Thoughts introduces two novel components: (1) Orchestrator Calibration, which identifies models with superior coordination and synthesis capabilities, and (2) Agent Self-Assessment, a protocol where tool agents profile their own domain-specific strengths to guide selection. At inference, the orchestrator dynamically activates the most compatible agents based on these profiles to maximize capability coverage. Across five mathematical reasoning and code generation benchmarks, Team-of-Thoughts consistently outperforms individual models and existing MAS baselines. Notably, on AIME24 and LiveCodeBench, Team-of-Thoughts achieves 96.00% and 77.91% accuracy, respectively, significantly improving over homogeneous role-play baselines (80.00% and 65.93%). 111Code is available at https://github.com/JeffreyWong20/Team-of-Thoughts. Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Calling Jeffrey T. H. Wong1,â Zixi Zhang1,â Junyi Liu2 Yiren Zhao1 1Imperial College London, 2Microsoft Research tsz.wong20,b.zhang25,a.zhao@imperial.ac.uk junyi.liu@microsoft.com $ $$ $footnotetext: These authors contributed equally to this work. 1 Introduction Figure 1: Pass@5 accuracy versus cost on MBPP+ for Team-of-Thoughts, single models, and AgentVerse. The dashed line denotes the Pareto front of single models. Figure 2: Overview of Team-of-Thoughts. (Left) Unlike single-model reasoning or homogeneous MAS (role-play), Team-of-Thoughts leverages heterogeneous model priors to maximize solution space coverage. (Right) Our architecture utilizes a pre-inference pipelineâcomprising orchestrator calibration and agent self-profilingâto enable dynamic agent selection. During inference, the orchestrator selectively invokes optimal tool agents and synthesizes their outputs into a high-confidence final response. Test-time scaling (TTS) has emerged as a critical paradigm for extending the capabilities of large language models (LLMs) beyond their training-time performance (Snell et al., 2024; Wu et al., 2025). By allocating additional computational budget during inference via methods such as process reward model (PRM) scoring, beam search, or tree-based exploration (Wei et al., 2023; Yao et al., 2023; Besta et al., 2024), models can unlock latent reasoning capabilities to solve complex tasks. This shift recognizes that the strategic deployment of inference-time compute is as fundamental to model performance as the scale of pre-training. Despite these gains, current TTS approaches are typically confined to single-model execution or static multi-agent workflows with fixed role assignments for a single model family (Qian et al., 2024; Hong et al., 2024; Wu et al., 2023). This homogeneity prevents systems from exploiting the complementary strengths inherent in diverse LLMs, which often possess divergent expertise due to distinct post-training procedures and dataset compositions. Recent industry developments, such as Grok 4.2 (xAI, 2026) and Claudeâs agent teams (Anthropic, 2026), signal a nascent industry shift toward coordinated multi-agent reasoning. However, existing academic frameworks (Zhang et al., 2024; Chen et al., 2024) still lack the dynamic flexibility required to coordinate these diverse agents based on model-specific expertise and task attributes. We address these limitations with Team-of-Thoughts, a novel Multi-Agent System (MAS) that achieves efficient test-time scaling through an orchestrated tool-calling paradigm as highlighted in FigureË1. Rather than treating models as monolithic reasoners or assigning them rigid personas, our framework conceptualizes diverse LLMs as specialized tools that can be dynamically invoked. By leveraging the native tool-calling capabilities of modern LLMs, we construct a hierarchical architecture where a central orchestrator strategically activates the most suitable tool agents for a query. This design facilitates massive parallelism and superior token efficiency. By replacing sequential, token-heavy reasoning with a coordinated team effort, Team-of-Thoughts ensures that the most proficient âspecialistsâ are consulted at the optimal time. Our key contributions are: ⢠Team-of-Thoughts Framework: A novel MAS architecture enabling heterogeneous LLMs to collaborate through a dynamic, tool-calling hierarchy that maximizes capability coverage. ⢠Orchestration and Self-Assessment Mechanisms: We introduce an orchestration calibration to identify superior coordinators and a self-assessment protocol for agents to profile their own domain expertise. ⢠Empirical Superiority: Extensive evaluation showing Team-of-Thoughts consistently outperforms standalone models and MAS baselines. Notably, achieving 96.00% on AIME24 and 77.91% on LiveCodeBench. 2 Background 2.1 Test-Time Scaling and Reasoning Test-time scaling (TTS) enhances performance by allocating additional computational budget during inference. Reasoning frameworks, such as Chain-of-Thought (CoT) (Wei et al., 2023), Tree-of-Thoughts (Yao et al., 2023), and Graph-of-Thoughts (Besta et al., 2024), leverage this by decomposing tasks into intermediate steps, enabling structured search over the solution space. However, a single agentâs expressive power is bounded by its fixed parameterization θ. While search-based scaling laws (Snell et al., 2024; Wu et al., 2025; Brown et al., 2024) show performance gains, these remain constrained by the modelâs inherent biases, which may preclude reaching solutions in remote regions of the solution space. As training larger models is often prohibitively expensive, there is a clear need for architectures that scale capability at test-time without re-training. 2.2 Multi-Agent Systems (MAS) Multi-Agent Systems (MAS) address single-model limitations by employing an ensemble θ1,θ2,âŚ,θn\ _1, _2,âŚ, _n\. In these systems, agents operate through independent reasoning or collaborative interactionâexchanging intermediate states and debating hypothesesâto synthesize a final output. Despite their potential, many current MAS frameworks are internally homogeneous, generating ensembles by prompting a single model with varying âpersonasâ (Chen et al., 2024; Yang et al., 2025b; Li et al., 2025). Because these agents share identical parameterization (θ1=θ2=âŻ=θn _1= _2=¡s= _n), they lack the distributional diversity necessary to explore complementary regions of the solution space. Furthermore, these frameworks are often computationally inefficient, requiring multiple rounds of exhaustive reasoning from every agent regardless of their task-specific relevance. This motivates Team-of-Thoughts, a framework utilizing heterogeneous model priors and a centralized orchestrator to strategically activate specialized agents, optimizing both capability coverage and inference efficiency. 3 Team of Thoughts: An Efficient Heterogeneous MAS Figure 3: Schematic comparison of Chain-of-Thought (CoT) and Team-of-Thoughts. (Left) Standard agentic reasoning via CoT generates sequential intermediate steps to refine the prediction distribution heuristically. However, on complex tasks, CoT may fail to converge on distant targets. (Right) Team-of-Thoughts utilizes heterogeneous tool agents to explore a broader solution space. During inference, these agents refine their individual distributions through local reasoning, which an orchestrator then prioritizes and aggregates into the final output. We introduce Team-of-Thoughts, a MAS framework that leverages a suite of heterogeneous agents to maximize capability coverage. The Team-of-Thoughts framework is composed of a central orchestrator agent porch(â âŁD)p_orch(¡ D) and a diverse ensemble of tool agents pθi(â âŁD)i=1n \p_ _i(¡ D) \_i=1^n, each characterized by a distinct prediction distribution. The orchestrator manages the reasoning process by dynamically invoking tool agents based on the requirements of the input query D. Specifically, the orchestrator performs three primary functions: 1. Selection: Identifying the tool agents best suited to address the specific question of D. 2. Evaluation: Assessing the quality and relevance of the tool agentsâ outputs. 3. Aggregation: Synthesizing tool-generated insights with its own reasoning to update the global context. When invoked, each tool agent i generates a reasoning trajectory (i)=Zi,1,Zi,2,⌠Z^(i)=\Z_i,1,Z_i,2,âŚ\ and produces a candidate prediction X^iâźpθiâ(XâŁD,(i)) X_i p_ _i (X D, Z^(i) ). The orchestrator then integrates these perspectives to update its context Z as follows: Zâźporch(â âŁD,X^1,X^2,âŚ,X^k) Z p_orch (¡ D, X_1, X_2,âŚ, X_k ) (1) ââĽZ Zâ Z\|Z (2) where kâ¤nk⤠n denotes the subset of agents active for the given task and âĽ\| denotes appending the new thinking to the context. In the rest of this section, we first provide a probabilistic motivation for our approach, and then we detail the implementation of the framework. 3.1 A Probabilistic View of Team-of-Thoughts A task-solving problem comprises an input query D and an unknown target answer XâX within the solution space S. We represent the prediction distribution of a language model with parameters θ as pθ(â âŁD)p_θ(¡ D). For analytical clarity, we can approximate this distribution as a multivariate Gaussian: pθ(â âŁD)â(θ,θ)p_θ(¡ D) ( Îź_θ, _θ) where Îź and denote the mean and covariance matrix, respectively. In inference, we aim to refine this distribution to maximize the likelihood of X. Limitations of Single-Agent Reasoning Standard agentic frameworks, such as Chain-of-Thought (CoT) (Wei et al., 2023), attempt to reach the target X by generating a sequence of intermediate reasoning steps Z1,Z2,âŚ,ZT\Z_1,Z_2,âŚ,Z_T\. Each step iteratively updates the modelâs state, shifting the prediction distribution: Z1âźpθ(â âŁD)â(θ1,θ1),Z2âźpθ(â âŁD,Z1)â(θ2,θ2),Z3âźpθ(â âŁD,Z1,Z2)â(θ3,θ3),âŽX^âźpθ(â âŁD,Z1,Z2,âŚ,ZT)â(θâ˛,θâ˛) gatheredZ_1 p_θ(¡ D) ( Îź^1_θ, ^1_θ ),\\ Z_2 p_θ(¡ D,Z_1) ( Îź^2_θ, ^2_θ ),\\ Z_3 p_θ(¡ D,Z_1,Z_2) ( Îź^3_θ, ^3_θ ),\\ \\ X p_θ(¡ D,Z_1,Z_2,âŚ,Z_T) ( Îź _θ, _θ ) gathered As illustrated in FigureË3 (left), these steps heuristically âdriftâ the initial distribution toward the target. However, a single agentâs ability to transform the prediction distribution â(θâ˛,θâ˛)N ( Îź _θ, _θ ) is essentially framed by its initial distribution â(θ,θ)N ( Îź_θ, _θ ) defined by the parameterization θ. In complex, high-dimensional tasks where the target X can be significantly distant from the initial mean θ Îź_θ, a single reasoning chain may require an impractical number of steps to converge, or may become trapped in local optima within the solution space. Team-of-Thoughts as a Gaussian Mixture The Team-of-Thoughts framework overcomes these constraints by employing a Multi-Agent System (MAS) of heterogeneous agents to achieve broader capability coverage. By utilizing an ensemble of n distinct tool agents E=pθii=1nE= \p_ _i \_i=1^n, we effectively initialize the system with multiple exploratory heads across the solution space. As illustrated in FigureË3 (right), the orchestrator agent porchp_orch selectively invokes these tool agents and integrates their outputs into its global reasoning context. From a probabilistic perspective, the resulting distribution formed by the orchestrator can be modeled as a Gaussian Mixture Model (GMM): porch(â âŁD,,^tools)ââi=1nwipθi(â âŁD,(i))p_orch (¡ D,Z, X_tools )â _i=1^nw_ip_ _i (¡ D,Z^(i) ) where Z represents the orchestratorâs reasoning trajectory, (i)Z^(i) is the local reasoning trajectory of the i-th tool agent, ^tools X_tools is the set of candidate tool predictions, and wiw_i are mixing weights determined by the orchestrator such that âwi=1ÎŁ w_i=1. To implement the weight assignment in a discrete LLM setup, we characterize it as a classification task performed by the orchestrator over the candidate responses: âporch(â âŁ(D,,^tools),^cali)wâ p_orch (¡ (D,Z, X_tools ), X_cali ) Here, â(âŚ)C(âŚ) denotes a classification query conditioned on the problem and tool generations, while ^cali X_cali represents the performance of tool agents on a calibration dataset, detailed in SectionË3.3. Conceptual Nuances It is important to note that the GMM serves as a conceptual framework rather than a literal mechanical description. The equivalence is approximated because (a) the updated orchestrator distribution is conditioned on the tool agentsâ discrete responses (^tools X_tools) rather than their full continuous distributions (pθip_ _i), and (b) the weight vector w is typically sparse, as the orchestrator identifies and selects only a high-confidence subset of kâ¤nk⤠n tool agents based on the calibration results. Despite these approximations, the formulations encapsulate how Team-of-Thoughts dynamically navigates the solution space by prioritizing high-utility distributions. Rather than relying on the âlinear driftâ of single-model reasoning, the orchestrator can effectively âjumpâ to high-probability regions already localized by specialized tool agents. It then refines this synthesized mixture through further reasoningâpotentially via iterative tool callsâto produce a definitive, high-confidence answer. 3.2 Orchestration Calibration Orchestration demands a specialized suite of capabilitiesâincluding tool comprehension, multi-agent coordination, and trajectory synthesisâthat are not uniformly distributed across model architectures. Crucially, a modelâs proficiency as a standalone solver does not always correlate with its efficacy as a coordinator. To identify the optimal âmanagerâ for our MAS, we introduce an Orchestration Calibration procedure. We evaluate candidate orchestrators based on their empirical performance in aggregating tool-agent responses within a specific task category c, subject to a fixed computational budget. For each candidate θcand _cand, we compute a calibration score: Scoreâ(θcand,c)=âDâval(c)â[X^caliâ(D,θcand)=Xcaliâ(D)]|val(c)|Score( _cand,c)=\\ _D _val^(c)I [ X_cali(D, _cand)=X_cali(D) ] |D_val^(c) | (3) where val(c)D_val^(c) represents the category-specific calibration set, Xcaliâ(D)X_cali(D) is the ground-truth answer, and X^caliâ(D,θcand) X_cali(D, _cand) is the prediction generated after the candidate aggregated the available tool agents. The model achieving the highest score is designated as the primary orchestrator for that category. As demonstrated in SectionË4.2, this process reveals that orchestration capability is often decoupled from model scale or general benchmark rankings. Certain models exhibit superior âintegrative reasoning,â while others, including some larger models, are more effective as specialized tool agents. These findings validate treating orchestrator selection as a distinct optimization problem rather than defaulting to the largest available model. 3.3 Agent Specialization via Self-Assessment Heterogeneous model families exhibit divergent expertise across task domains due to variations in architecture and post-training procedures. To capitalize on these specialized priors, we introduce a Self-Assessment Mechanism that enables tool agents to characterize their own proficiency profiles. This mechanism consists of three stages: 1. Empirical Profiling: Each tool agent i generates predictions X^i X_i and reasoning trajectories (i)Z^(i) on a representative validation set DvalD_val: X^iâźpθi(â âŁD,(i)),âDâval X_i p_ _i (¡ D,Z^(i) ), â D _val 2. Performance Quantification: The system evaluates the accuracy sis_i of each agent using an indicator function relative to the ground truth Xâ(D)X(D): si=1|Dval|ââDâvalâ[X^iâ(D)=Xâ(D)]s_i= 1 |D_val | _D _valI [ X_i(D)=X(D) ] 3. Qualitative Characterization: Using its performance data si(c)s_i^(c) on every the categorized question type c within valD_val, each tool agent generates a âcapability profile.â This profile is a concise natural language summary detailing the agentâs strengths, weaknesses, and reliability for specific task categories. During inference, these capability profiles are provided to the orchestratorâs context. This enables dynamic, priority-based activation: the orchestrator evaluates the incoming query D against the agentsâ profiles to select only the most compatible tools. This approach facilitates strategic budget allocationâhighly proficient agents are prioritized while irrelevant agents are bypassedâcontrasting with traditional MAS architectures that invoke all agents regardless of task fit. 4 Experiments Table 1: Performance comparison of Team-of-Thoughts across five benchmark tasks. Success rates (%) are reported using pass@1, 3, and 5. Team-of-Thoughts is evaluated against individual model baselines and state-of-the-art multi-agent methods. The Theoretical Limit represents the oracle upper bound achieved by selecting the best result among all individual models for each problem. AIME 2024 AIME 2025 Humaneval+ MBPP+ LiveCodeBench v6 Model pass@1 @3 @5 @1 @3 @5 @1 @3 @5 @1 @3 @5 @1 @3 @5 Claude Sonnet 4.5 84.00 86.67 86.67 74.67 87.67 90.00 95.61 96.34 96.34 81.80 84.29 85.45 48.24 55.33 57.69 GPT-5-mini 88.67 93.33 93.33 73.33 83.50 86.11 92.68 94.88 95.12 82.38 86.19 87.57 64.84 74.40 77.20 Gemini 3 Flash 86.67 90.67 93.33 71.33 81.33 86.67 96.46 96.95 96.95 84.44 85.13 85.45 72.53 75.71 77.47 Deepseek v3.2 Exp 94.67 96.67 96.67 94.00 98.67 100.00 91.26 95.88 96.65 80.20 84.59 85.83 76.48 86.65 88.46 GPT OSS 20b 75.33 88.00 90.00 73.33 83.50 86.11 89.23 95.12 95.63 79.23 84.88 86.29 46.04 54.34 58.24 Qwen3-vl 235B 93.33 93.33 93.33 89.44 91.67 92.78 93.19 94.70 95.02 83.20 86.24 87.04 67.47 75.77 78.57 Phi-4 13.33 18.00 20.00 10.67 15.67 16.67 74.80 80.55 82.83 63.07 72.51 74.60 24.84 27.69 28.02 Theoretical Limit 96.67 â â 100.00 â â 100.00 â â 100.00 â â 90.50 â â Majority Voting 93.33 â â 93.33 â â N/A N/A N/A N/A N/A N/A N/A N/A N/A AgentVerse 80.00 â â 60.00 â â 91.46 â â 79.37 â â 65.93 â â Team-of-Thoughts 96.00 96.67 96.67 95.33 99.33 100.00 95.33 97.47 98.07 85.93 88.84 89.68 77.91 85.05 87.91 Models and Benchmarks We evaluate Team-of-Thoughts across seven model families: Claude-Sonnet-4.5 (PBC, 2025), GPT-5-mini (Singh et al., 2025), Gemini-3-Flash-Preview (Pichai et al., 2025), DeepSeek-V3.2-Exp (DeepSeek-AI et al., 2025), GPT-OSS-20B (OpenAI et al., 2025), Qwen3-VL-235B-A22B-Thinking (Yang et al., 2025a), and Phi-4 (Abdin et al., 2024). Each tool call activates two tool agents per call. Our evaluation spans mathematical reasoning (AIME2024 (of America, 2024), AIME2025 (of America, 2025)) and code generation (Humaneval+ (Chen et al., 2021), MBPP+ (Austin et al., 2021), and LiveCodeBench v6 (Jain et al., 2024) for JanâMay 2025). For calibration, we typically employ a 10% random sample. However, to account for the limited scale of AIME benchmarks, we adopt a cross-calibration strategy: 50% of AIME2025 serves as the calibration set for AIME2024, and vice versa. This ensures robust agent profiling while preventing test-set contamination. 4.1 Performance Analysis The Team-of-Thoughts framework integrates tool agents via descriptions derived from their self-assessment profiles. Guided by the calibration results in SectionË4.2, we utilize DeepSeek v3.2 as the orchestrator for mathematical reasoning and GPT-5 Mini for code generation. As detailed in TableË1, Team-of-Thoughts demonstrates robust performance across both reasoning and coding domains. On mathematical benchmarks (AIME2024, AIME2025), it achieves superior stability compared to standalone models, consistently attaining higher pass@1 accuracy while remaining competitive at pass@3 and pass@5. In code generation, Team-of-Thoughts achieves state-of-the-art results on MBPP+ across all pass@k. On HumanEval+, it trails Claude Sonnet 4.5 marginally at pass@1âa metric increasingly constrained by benchmark saturationâbut demonstrates superior scaling by surpassing Claude as the number of allowed generations increases. For LiveCodeBench, Team-of-Thoughts achieves peak pass@1 performance but marginally lags behind the strongest single-model baseline at pass@3 and pass@5. We hypothesize that default random calibration fails to capture the high task diversity of this benchmark. Validation experiments using a categorized calibration set support this: the strongest models are invoked twice as frequently as a random calibration set, significantly improving overall performance (see Appendix B). Overall, these results validate the probabilistic intuition behind Team-of-Thoughts. By leveraging heterogeneous agents to cover a broader solution space, the orchestrator âjumpsâ to high-probability regions localized by specialized tools. This strategic prioritization allows the system to synthesize robust answers that transcend the constraints of any single model parameterization. 4.2 Orchestration Agent Selection Table 2: Orchestration performance across candidate models and budget constraints. Calibration accuracy (%) is evaluated on AIME2024 and MBPP+ benchmarks under two distinct budgetary tiers (USD). Avg denotes the mean accuracy across all evaluated settings. Model AIME2024 MBPP+ Budget Cost ($) 0.03 0.02 Avg 0.03 0.02 Avg Claude Sonnet 4.5 86.67 46.67 66.67 78.38 78.38 78.38 GPT-5 Mini 86.67 93.33 90.00 86.49 83.78 85.14 Gemini 3 Flash 86.67 40.00 63.34 83.78 83.78 83.78 DeepSeek v3.2 93.33 93.33 93.33 81.08 81.08 81.08 GPT-OSS 20B 80.00 80.00 80.00 75.68 75.68 75.68 Qwen3-vl 235B 33.33 20.00 26.67 78.38 78.38 78.38 Phi-4 33.33 33.33 33.33 72.97 72.97 72.97 To identify the optimal coordinator, we evaluate candidate models on AIME2024 and MBPP+ under fixed monetary constraints. We normalize budgets by translating costs into model-specific token limits to ensure a rigorous, cost-controlled comparison across providers. As shown in TableË2, orchestration proficiency is task-dependent: DeepSeek v3.2 performs best on the AIME2024 mathematical benchmark, while GPT-5 Mini excels on the MBPP+ coding task. These results indicate that the most effective orchestrator is not necessarily the largest model, but the one with the highest orchestrating efficiency for a specific domain. Consequently, we employ DeepSeek v3.2 for mathematical reasoning and GPT-5 Mini for code generation in our primary analysis. 4.3 Tool Agent Assessment Strategies To optimize agent invocation, we generate granular capability profiles for each tool agent. We evaluate three distinct selection policies to determine how these profiles are best constructed and utilized: Random Invocation (Baseline): Tool models are sampled uniformly from the ensemble without prior profiling and domain-specific weighting. Orchestrator-Led Assessment: A centralized model (DeepSeek-v3.2 for math; GPT-5 Mini for coding) evaluates the tool agents on a calibration set to map their competencies and common failure modes. Agent Self-Assessment: Each tool agent performs a self-audit. Provided with the task, its own reasoning trajectories, and the ground truth, the agent identifies the specific skills required and critiques its own performance. These profiles allow the orchestrator to dynamically adjust invocation preference based on task-agent compatibility. As detailed in TableË3, all structured selection policies significantly outperform the orchestrator-only baseline (âSingleâ). Among these, the self-assessment policy achieves the highest overall accuracy. A key advantage of self-assessment is that it decouples capability profiling from the orchestratorâs internal biases, providing a more objective and stable representation of agent expertise. Consequently, we adopt self-assessment as the default selection strategy. Table 3: Performance evaluation of tool agent assessment strategies. Comparison of Pass@1 accuracy (%) with no tool agent (âSingleâ), Orchestrator-Led Assessment, Agent Self-Assessment, and Random Invocation across AIME2024 and MBPP+ benchmarks. Benchmark Single Self-Assess Orch-Based Random AIME2024 94.67 96.00 92.67 94.00 MBPP+ 82.38 85.93 85.77 84.44 4.4 Tool Response Aggregation Strategies Table 4: Performance of orchestrator aggregation strategies. Comparison of task accuracy (%) achieved when incorporating versus excluding the intermediate reasoning traces of tool agents. AIME 2025 MBPP+ Config pass@1 @3 @5 pass@1 @3 @5 No traces 95.33 99.33 100.00 85.93 88.84 89.68 Include traces 78.67 93.33 96.67 80.90 85.61 87.04 In the Team-of-Thoughts framework, an activated tool agent θi _i is prompted with a query DcallD_call. The agent generates a sequence of intermediate reasoning steps (i)=[Z1(i),âŚ,ZT(i)]Z^(i)= [Z_1^(i),âŚ,Z_T^(i) ] where Zt(i)âźpθi(â âŁDcall,1:tâ1(i))Z_t^(i) p_ _i(¡ D_call,Z_1:t-1^(i)), reaching a final output X^i X_i. We evaluate two distinct protocols for information exchange between the tool agent and the orchestrator: (1) returning only the final solution X^i X_i, or (2) providing the complete reasoning trajectory (i)Z^(i) alongside the answer. As summarized in TableË4, our results indicate that incorporating full reasoning trajectories consistently degrades orchestration performance across all benchmarks. This performance drop suggests that verbose trajectories introduce significant contextual noise, which can distract or mislead the orchestrator through intermediate logical errors or irrelevant sub-steps. These findings demonstrate that for complex multi-agent synthesis, minimal tool responses enhance the orchestratorâs ability to integrate heterogeneous information. By filtering out potentially fallible intermediate steps, the orchestrator maintains a cleaner, more reliable global context, leading to more robust decision-making and higher task accuracy. 4.5 Scaling Dynamics of Tool-Agent Ensemble While Team-of-Thoughts exhibits high efficacy, the performance of Multi-Agent Systems (MAS) does not scale monotonically with ensemble size. We investigate the scaling behavior of Team-of-Thoughts relative to two variables: the size of the available ensemble (|E||E|) and the number of activated agents per call (k). To test the limits of this architecture, we expanded the ensemble E to include the 100 most popular models on OpenRouter and evaluated performance on MBPP+. Table 5: Performance under different numbers of activated tool agents, with 25 total tool agents available. Activated tools Pass@1 Pass@3 Pass@5 2 82.75% 85.58% 86.77% 4 82.96% 85.32% 85.98% 8 60.32% 83.76% 86.24% 16 60.63% 83.10% 85.71% Impacts of Activation Density: We first examine whether the orchestrator benefits from a higher volume of concurrent agent responses. Fixing the available ensemble at |E|=25|E|=25, we varied the number of activated models k from 2 to 16. As shown in TableË5, increasing k yields negligible gains and eventually leads to performance degradation. Notably, Pass@1 accuracy drops significantly once k reaches 8. The narrowing performance gap at Pass@3 and Pass@5 suggests that while higher activation density k/nk/n increases run-time varianceâwidening the distribution to cover the correct answer across multiple attemptsâit simultaneously compromises the robustness of the single-shot (Pass@1) prediction. Table 6: Orchestrator performance and tool selection consistency (Agreement) across varying ensemble sizes. Performance is compared between configurations activating two and eight agents per tool call. Activate 2 Activate 8 Agreement Avail Tools Pass@1 @3 @5 Pass@1 @3 @5 A (%) 10 82.70 86.32 87.57 74.97 85.26 85.98 90.74 25 82.75 85.58 86.77 60.32 83.76 86.24 91.80 50 82.43 86.59 87.83 54.76 82.17 86.51 61.64 75 82.49 87.09 88.36 63.65 84.21 86.24 75.13 100 82.49 86.56 87.57 61.43 84.13 86.51 71.43 Ensemble Breadth and Selection Overload. We further analyze whether the orchestratorâs selection improves when provided with a broader pool of candidates. Sweeping over ensemble sizes for k=2k=2 and k=8k=8, TableË6 reveals that while k=2k=2 remains stable across ensemble sizes, k=8k=8 exhibits a clear performance decay as |E||E| increases. To isolate the root cause, we define the Selection Agreement between these two configurations: Agreement=1|D|ââi=1|D|â[S2(i)âS8(i)]Agreement= 1 |D | _i=1 |D |I [S_2^(i) S_8^(i) ] where Sk(i)S_k^(i) denotes the set of tool agents invoked for the i-th question at activation level k. As shown in the final column of TableË6, Agreement sharply declines once |E|>50|E|>50. This suggests a selection logic overload phenomenon: the orchestratorâs ability to identify optimal tools is overwhelmed by excessive options, leading to invocations of suboptimal agents. Even though Agreement remains high when |E|â¤25|E|⤠25, a Pass@1 performance gap persists. This indicates that even with âoptimalâ tool selection, the high density of responses (eight vs. two) introduces excessive variance into the aggregation process, exceeding the orchestratorâs synthesis capacity. Our findings demonstrate that MAS performance is constrained by an orchestration bottleneck. Scaling both activation density and ensemble breadth eventually degrades performance by overwhelming the orchestratorâs selection and aggregation logic. Conversely, with a calibrated activation count (e.g. k=2k=2), the orchestrator can reliably extract high-utility responses from an arbitrarily large ensemble, producing stable, high-confidence solutions. 5 Related Work Test-Time Scaling (TTS). TTS establishes that increasing inference-time computation can yield performance gains comparable to scaling model parameters (Snell et al., 2024; Wu et al., 2025). Approaches generally follow two trajectories: (1) parallel sampling, such as Best-of-N (Brown et al., 2024), which uses verification to select optimal outputs from a broad search space; and (2) sequential reasoning, which deepens the reasoning topology. This includes linear Chain-of-Thought (CoT) (Wei et al., 2023) and non-linear frameworks like Tree of Thoughts (ToT) (Yao et al., 2023) and Graph of Thoughts (GoT) (Besta et al., 2024). These methodologies enable complex operations like backtracking and information aggregation, mimicking human âSystem 2â cognition. Multi-Agent Systems (MAS). Inspired by human dynamics, MAS utilizes agent ensembles to enhance task reliability and role-playing capabilities (Park et al., 2023; Li et al., 2023). Early frameworks leveraged conversational flows (Wu et al., 2023) or Standard Operating Procedures (SOPs) to reduce hallucinations in software engineering (Qian et al., 2024; Hong et al., 2024). More recent research emphasizes dynamic workflows; for instance, AgentVerse (Chen et al., 2024) introduces âexpert recruitment,â while AgentNet (Yang et al., 2025b) employs autonomous graph topologies for task decomposition. Theoretically, the âSociety of Thoughtsâ (Kim et al., 2026) suggests that even single-model reasoning benefits from simulating diverse internal perspectives. In contrast to these frameworks, which often rely on homogeneous model ensembles, our Team-of-Thoughts framework explicitly exploits the divergence inherent in heterogeneous model priors. By dynamically selecting specialized agents and aggregating their distinct reasoning paths, we facilitate more robust decision-making than systems restricted to a single model family. 6 Conclusion We propose the Team-of-Thoughts framework, which explicitly leverages the skill diversity of a group of heterogeneous agent models. By dynamically selecting the most suitable orchestrator and making skill-dependent tool calls that can be executed in parallel, Team-of-Thoughts overcomes the limitations of fixed-role, sequential multi-agent systems, achieving superior performance across reasoning and code generation tasks. Our analysis and empirical experiment across reasoning and code generation tasks demonstrates that this approach consistently achieves more efficient cost usage and higher accuracy compared to both single-model baselines and multi-agent systems with homogeneous backbone models. This work highlights a novel paradigm for enabling heterogeneous models to collaborate, paving the way for future exploration of multi-round, complex multi-agent tasks. Limitations The proposed Team-of-Thoughts framework relies on a central LLM as an orchestrator, which introduces a performance bottleneck as the agent pool scales. In SectionË4.5, we characterize the limitations of this centralized architecture when scaled to 100 tool-agents. Specifically, we observe two primary failure modes: The orchestrator often fails to navigate expansive tool selection spaces, leading to suboptimal tool-calling sequences. Large-scale tool orchestration induces high variance in output quality as show in pass@1 performance, suggesting a lack of robustness in the modelâs iterative reasoning process under high context load. These results suggest that the scalability of the system is currently constrained by the state of frontier models. Future advancements in long-context reasoning and hierarchical information aggregation will be essential to fully realize the potential of massive tool-agent ensembles. References M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmann, J. R. Lee, Y. T. Lee, Y. Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y. Wu, D. Yu, C. Zhang, and Y. Zhang (2024) Phi-4 technical report. External Links: 2412.08905, Link Cited by: §A.1, §4. Anthropic (2026) Claude code agent teams documentation. Note: https://code.claude.com/docs/en/agent-teamsAccessed: 2026-03-15 Cited by: §1. J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, and C. Sutton (2021) Program synthesis with large language models. External Links: 2108.07732, Link Cited by: §A.1, §4. M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler (2024) Graph of thoughts: solving elaborate problems with large language models. Proceedings of the AAAI Conference on Artificial Intelligence 38 (16), p. 17682â17690. External Links: ISSN 2159-5399, Link, Document Cited by: §1, §2.1, §5. B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. RĂŠ, and A. Mirhoseini (2024) Large language monkeys: scaling inference compute with repeated sampling. External Links: 2407.21787, Link Cited by: §2.1, §5. M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba (2021) Evaluating large language models trained on code. External Links: 2107.03374, Link Cited by: §A.1, §4. W. Chen, Y. Su, J. Zuo, C. Yang, C. Yuan, C. Chan, H. Yu, Y. Lu, Y. Hung, C. Qian, Y. Qin, X. Cong, R. Xie, Z. Liu, M. Sun, and J. Zhou (2024) AgentVerse: facilitating multi-agent collaboration and exploring emergent behaviors. In International Conference on Learning Representations, B. Kim, Y. Yue, S. Chaudhuri, K. Fragkiadaki, M. Khan, and Y. Sun (Eds.), Vol. 2024, p. 20094â20136. External Links: Link Cited by: §A.2, §1, §2.2, §5. DeepSeek-AI, A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Lu, C. Zhao, C. Deng, C. Xu, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, E. Li, F. Zhou, F. Lin, F. Dai, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Li, H. Liang, H. Wei, H. Zhang, H. Luo, H. Ji, H. Ding, H. Tang, H. Cao, H. Gao, H. Qu, H. Zeng, J. Huang, J. Li, J. Xu, J. Hu, J. Chen, J. Xiang, J. Yuan, J. Cheng, J. Zhu, J. Ran, J. Jiang, J. Qiu, J. Li, J. Song, K. Dong, K. Gao, K. Guan, K. Huang, K. Zhou, K. Huang, K. Yu, L. Wang, L. Zhang, L. Wang, L. Zhao, L. Yin, L. Guo, L. Luo, L. Ma, L. Wang, L. Zhang, M. S. Di, M. Y. Xu, M. Zhang, M. Zhang, M. Tang, M. Zhou, P. Huang, P. Cong, P. Wang, Q. Wang, Q. Zhu, Q. Li, Q. Chen, Q. Du, R. Xu, R. Ge, R. Zhang, R. Pan, R. Wang, R. Yin, R. Xu, R. Shen, R. Zhang, S. H. Liu, S. Lu, S. Zhou, S. Chen, S. Cai, S. Chen, S. Hu, S. Liu, S. Hu, S. Ma, S. Wang, S. Yu, S. Zhou, S. Pan, S. Zhou, T. Ni, T. Yun, T. Pei, T. Ye, T. Yue, W. Zeng, W. Liu, W. Liang, W. Pang, W. Luo, W. Gao, W. Zhang, X. Gao, X. Wang, X. Bi, X. Liu, X. Wang, X. Chen, X. Zhang, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Li, X. Yang, X. Li, X. Chen, X. Su, X. Pan, X. Lin, X. Fu, Y. Q. Wang, Y. Zhang, Y. Xu, Y. Ma, Y. Li, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Qian, Y. Yu, Y. Zhang, Y. Ding, Y. Shi, Y. Xiong, Y. He, Y. Zhou, Y. Zhong, Y. Piao, Y. Wang, Y. Chen, Y. Tan, Y. Wei, Y. Ma, Y. Liu, Y. Yang, Y. Guo, Y. Wu, Y. Wu, Y. Cheng, Y. Ou, Y. Xu, Y. Wang, Y. Gong, Y. Wu, Y. Zou, Y. Li, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Zhao, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Huang, Z. Wu, Z. Li, Z. Zhang, Z. Xu, Z. Wang, Z. Gu, Z. Zhu, Z. Li, Z. Zhang, Z. Xie, Z. Gao, Z. Pan, Z. Yao, B. Feng, H. Li, J. L. Cai, J. Ni, L. Xu, M. Li, N. Tian, R. J. Chen, R. L. Jin, S. S. Li, S. Zhou, T. Sun, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Song, X. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Z. Huang, Z. Xu, Z. Zhang, D. Ji, J. Liang, J. Guo, J. Chen, L. Xia, M. Wang, M. Li, P. Zhang, R. Chen, S. Sun, S. Wu, S. Ye, T. Wang, W. L. Xiao, W. An, X. Wang, X. Sun, X. Wang, Y. Tang, Y. Zha, Z. Zhang, Z. Ju, Z. Zhang, and Z. Qu (2025) DeepSeek-v3.2: pushing the frontier of open large language models. External Links: 2512.02556, Link Cited by: §A.1, §4. S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, C. Zhang, J. Wang, Z. Wang, S. K. S. Yau, Z. Lin, L. Zhou, C. Ran, L. Xiao, C. Wu, and J. Schmidhuber (2024) MetaGPT: meta programming for a multi-agent collaborative framework. External Links: 2308.00352, Link Cited by: §1, §5. N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica (2024) LiveCodeBench: holistic and contamination free evaluation of large language models for code. External Links: 2403.07974, Link Cited by: §A.1, §4. J. Kim, S. Lai, N. Scherrer, B. A. y Arcas, and J. Evans (2026) Reasoning models generate societies of thought. External Links: 2601.10825, Link Cited by: §5. G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for "mind" exploration of large language model society. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36, p. 51991â52008. External Links: Link Cited by: §5. W. Li, J. Lin, Z. Jiang, J. Cao, X. Liu, J. Zhang, Z. Huang, Q. Chen, W. Sun, Q. Wang, H. Lu, T. Qin, C. Zhu, Y. Yao, S. Fan, X. Li, T. Wang, P. Liu, K. Zhu, H. Zhu, D. Shi, P. Wang, Y. Guan, X. Tang, M. Liu, Y. E. Jiang, J. Yang, J. Liu, G. Zhang, and W. Zhou (2025) Chain-of-agents: end-to-end agent foundation models via multi-agent distillation and agentic rl. External Links: 2508.13167, Link Cited by: §2.2. M. A. of America (2024) AIME 2024. Cited by: §A.1, §4. M. A. of America (2025) AIME 2025. Cited by: §A.1, §4. OpenAI, :, S. Agarwal, L. Ahmad, J. Ai, S. Altman, A. Applebaum, E. Arbus, R. K. Arora, Y. Bai, B. Baker, H. Bao, B. Barak, A. Bennett, T. Bertao, N. Brett, E. Brevdo, G. Brockman, S. Bubeck, C. Chang, K. Chen, M. Chen, E. Cheung, A. Clark, D. Cook, M. Dukhan, C. Dvorak, K. Fives, V. Fomenko, T. Garipov, K. Georgiev, M. Glaese, T. Gogineni, A. Goucher, L. Gross, K. G. Guzman, J. Hallman, J. Hehir, J. Heidecke, A. Helyar, H. Hu, R. Huet, J. Huh, S. Jain, Z. Johnson, C. Koch, I. Kofman, D. Kundel, J. Kwon, V. Kyrylov, E. Y. Le, G. Leclerc, J. P. Lennon, S. Lessans, M. Lezcano-Casado, Y. Li, Z. Li, J. Lin, J. Liss, Lily, Liu, J. Liu, K. Lu, C. Lu, Z. Martinovic, L. McCallum, J. McGrath, S. McKinney, A. McLaughlin, S. Mei, S. Mostovoy, T. Mu, G. Myles, A. Neitz, A. Nichol, J. Pachocki, A. Paino, D. Palmie, A. Pantuliano, G. Parascandolo, J. Park, L. Pathak, C. Paz, L. Peran, D. Pimenov, M. Pokrass, E. Proehl, H. Qiu, G. Raila, F. Raso, H. Ren, K. Richardson, D. Robinson, B. Rotsted, H. Salman, S. Sanjeev, M. Schwarzer, D. Sculley, H. Sikchi, K. Simon, K. Singhal, Y. Song, D. Stuckey, Z. Sun, P. Tillet, S. Toizer, F. Tsimpourlas, N. Vyas, E. Wallace, X. Wang, M. Wang, O. Watkins, K. Weil, A. Wendling, K. Whinnery, C. Whitney, H. Wong, L. Yang, Y. Yang, M. Yasunaga, K. Ying, W. Zaremba, W. Zhan, C. Zhang, B. Zhang, E. Zhang, and S. Zhao (2025) Gpt-oss-120b & gpt-oss-20b model card. External Links: 2508.10925, Link Cited by: §A.1, §4. J. S. Park, J. OâBrien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein (2023) Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST â23, New York, NY, USA. External Links: ISBN 9798400701320, Link, Document Cited by: §5. A. PBC (2025) Introducing claude sonnet 4.5. External Links: Link Cited by: §A.1, §4. S. Pichai, D. Hassabis, and K. Kavukcuoglu (2025) A new era of intelligence with gemini 3. External Links: Link Cited by: §A.1, §4. C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, J. Xu, D. Li, Z. Liu, and M. Sun (2024) ChatDev: communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 15174â15186. External Links: Link, Document Cited by: §1, §5. A. Singh, A. Fry, A. Perelman, A. Tart, A. Ganesh, A. El-Kishky, A. McLaughlin, A. Low, A. Ostrow, A. Ananthram, A. Nathan, A. Luo, A. Helyar, A. Madry, A. Efremov, A. Spyra, A. Baker-Whitcomb, A. Beutel, A. Karpenko, A. Makelov, A. Neitz, A. Wei, A. Barr, A. Kirchmeyer, A. Ivanov, A. Christakis, A. Gillespie, A. Tam, A. Bennett, A. Wan, A. Huang, A. M. Sandjideh, A. Yang, A. Kumar, A. Saraiva, A. Vallone, A. Gheorghe, A. G. Garcia, A. Braunstein, A. Liu, A. Schmidt, A. Mereskin, A. Mishchenko, A. Applebaum, A. Rogerson, A. Rajan, A. Wei, A. Kotha, A. Srivastava, A. Agrawal, A. Vijayvergiya, A. Tyra, A. Nair, A. Nayak, B. Eggers, B. Ji, B. Hoover, B. Chen, B. Chen, B. Barak, B. Minaiev, B. Hao, B. Baker, B. Lightcap, B. McKinzie, B. Wang, B. Quinn, B. Fioca, B. Hsu, B. Yang, B. Yu, B. Zhang, B. Brenner, C. R. Zetino, C. Raymond, C. Lugaresi, C. Paz, C. Hudson, C. Whitney, C. Li, C. Chen, C. Cole, C. Voss, C. Ding, C. Shen, C. Huang, C. Colby, C. Hallacy, C. Koch, C. Lu, C. Kaplan, C. Kim, C. Minott-Henriques, C. Frey, C. Yu, C. Czarnecki, C. Reid, C. Wei, C. Decareaux, C. Scheau, C. Zhang, C. Forbes, D. Tang, D. Goldberg, D. Roberts, D. Palmie, D. Kappler, D. Levine, D. Wright, D. Leo, D. Lin, D. Robinson, D. Grabb, D. Chen, D. Lim, D. Salama, D. Bhattacharjee, D. Tsipras, D. Li, D. Yu, D. Strouse, D. Williams, D. Hunn, E. Bayes, E. Arbus, E. Akyurek, E. Y. Le, E. Widmann, E. Yani, E. Proehl, E. Sert, E. Cheung, E. Schwartz, E. Han, E. Jiang, E. Mitchell, E. Sigler, E. Wallace, E. Ritter, E. Kavanaugh, E. Mays, E. Nikishin, F. Li, F. P. Such, F. de Avila Belbute Peres, F. Raso, F. Bekerman, F. Tsimpourlas, F. Chantzis, F. Song, F. Zhang, G. Raila, G. McGrath, G. Briggs, G. Yang, G. Parascandolo, G. Chabot, G. Kim, G. Zhao, G. Valiant, G. Leclerc, H. Salman, H. Wang, H. Sheng, H. Jiang, H. Wang, H. Jin, H. Sikchi, H. Schmidt, H. Aspegren, H. Chen, H. Qiu, H. Lightman, I. Covert, I. Kivlichan, I. Silber, I. Sohl, I. Hammoud, I. Clavera, I. Lan, I. Akkaya, I. Kostrikov, I. Kofman, I. Etinger, I. Singal, J. Hehir, J. Huh, J. Pan, J. Wilczynski, J. Pachocki, J. Lee, J. Quinn, J. Kiros, J. Kalra, J. Samaroo, J. Wang, J. Wolfe, J. Chen, J. Wang, J. Harb, J. Han, J. Wang, J. Zhao, J. Chen, J. Yang, J. Tworek, J. Chand, J. Landon, J. Liang, J. Lin, J. Liu, J. Wang, J. Tang, J. Yin, J. Jang, J. Morris, J. Flynn, J. Ferstad, J. Heidecke, J. Fishbein, J. Hallman, J. Grant, J. Chien, J. Gordon, J. Park, J. Liss, J. Kraaijeveld, J. Guay, J. Mo, J. Lawson, J. McGrath, J. Vendrow, J. Jiao, J. Lee, J. Steele, J. Wang, J. Mao, K. Chen, K. Hayashi, K. Xiao, K. Salahi, K. Wu, K. Sekhri, K. Sharma, K. Singhal, K. Li, K. Nguyen, K. Gu-Lemberg, K. King, K. Liu, K. Stone, K. Yu, K. Ying, K. Georgiev, K. Lim, K. Tirumala, K. Miller, L. Ahmad, L. Lv, L. Clare, L. Fauconnet, L. Itow, L. Yang, L. Romaniuk, L. Anise, L. Byron, L. Pathak, L. Maksin, L. Lo, L. Ho, L. Jing, L. Wu, L. Xiong, L. Mamitsuka, L. Yang, L. McCallum, L. Held, L. Bourgeois, L. Engstrom, L. Kuhn, L. Feuvrier, L. Zhang, L. Switzer, L. Kondraciuk, L. Kaiser, M. Joglekar, M. Singh, M. Shah, M. Stratta, M. Williams, M. Chen, M. Sun, M. Cayton, M. Li, M. Zhang, M. Aljubeh, M. Nichols, M. Haines, M. Schwarzer, M. Gupta, M. Shah, M. Huang, M. Dong, M. Wang, M. Glaese, M. Carroll, M. Lampe, M. Malek, M. Sharman, M. Zhang, M. Wang, M. Pokrass, M. Florian, M. Pavlov, M. Wang, M. Chen, M. Wang, M. Feng, M. Bavarian, M. Lin, M. Abdool, M. Rohaninejad, N. Soto, N. Staudacher, N. LaFontaine, N. Marwell, N. Liu, N. Preston, N. Turley, N. Ansman, N. Blades, N. Pancha, N. Mikhaylin, N. Felix, N. Handa, N. Rai, N. Keskar, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, O. Gleeson, P. Mishkin, P. Lesiewicz, P. Baltescu, P. Belov, P. Zhokhov, P. Pronin, P. Guo, P. Thacker, Q. Liu, Q. Yuan, Q. Liu, R. Dias, R. Puckett, R. Arora, R. T. Mullapudi, R. Gaon, R. Miyara, R. Song, R. Aggarwal, R. Marsan, R. Yemiru, R. Xiong, R. Kshirsagar, R. Nuttall, R. Tsiupa, R. Eldan, R. Wang, R. James, R. Ziv, R. Shu, R. Nigmatullin, S. Jain, S. Talaie, S. Altman, S. Arnesen, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Yoo, S. Heon, S. Ethersmith, S. Grove, S. Taylor, S. Bubeck, S. Banesiu, S. Amdo, S. Zhao, S. Wu, S. Santurkar, S. Zhao, S. R. Chaudhuri, S. Krishnaswamy, Shuaiqi, Xia, S. Cheng, S. Anadkat, S. P. Fishman, S. Tobin, S. Fu, S. Jain, S. Mei, S. Egoian, S. Kim, S. Golden, S. Mah, S. Lin, S. Imm, S. Sharpe, S. Yadlowsky, S. Choudhry, S. Eum, S. Sanjeev, T. Khan, T. Stramer, T. Wang, T. Xin, T. Gogineni, T. Christianson, T. Sanders, T. Patwardhan, T. Degry, T. Shadwell, T. Fu, T. Gao, T. Garipov, T. Sriskandarajah, T. Sherbakov, T. Kaftan, T. Hiratsuka, T. Wang, T. Song, T. Zhao, T. Peterson, V. Kharitonov, V. Chernova, V. Kosaraju, V. Kuo, V. Pong, V. Verma, V. Petrov, W. Jiang, W. Zhang, W. Zhou, W. Xie, W. Zhan, W. McCabe, W. DePue, W. Ellsworth, W. Bain, W. Thompson, X. Chen, X. Qi, X. Xiang, X. Shi, Y. Dubois, Y. Yu, Y. Khakbaz, Y. Wu, Y. Qian, Y. T. Lee, Y. Chen, Y. Zhang, Y. Xiong, Y. Tian, Y. Cha, Y. Bai, Y. Yang, Y. Yuan, Y. Li, Y. Zhang, Y. Yang, Y. Jin, Y. Jiang, Y. Wang, Y. Wang, Y. Liu, Z. Stubenvoll, Z. Dou, Z. Wu, and Z. Wang (2025) OpenAI gpt-5 system card. External Links: 2601.03267, Link Cited by: §A.1, §A.2, §4. C. Snell, J. Lee, K. Xu, and A. Kumar (2024) Scaling llm test-time compute optimally can be more effective than scaling model parameters. External Links: 2408.03314, Link Cited by: §1, §2.1, §5. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou (2023) Chain-of-thought prompting elicits reasoning in large language models. External Links: 2201.11903, Link Cited by: §1, §2.1, §3.1, §5. Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, A. H. Awadallah, R. W. White, D. Burger, and C. Wang (2023) AutoGen: enabling next-gen llm applications via multi-agent conversation. External Links: 2308.08155, Link Cited by: §1, §5. Y. Wu, Z. Sun, S. Li, S. Welleck, and Y. Yang (2025) Inference scaling laws: an empirical analysis of compute-optimal inference for problem-solving with language models. External Links: 2408.00724, Link Cited by: §1, §2.1, §5. xAI (2026) Grok (version 4.2 beta 0309). Note: https://x.aiAccessed: 2026-03-15 Cited by: §1. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025a) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §A.1, §4. Y. Yang, H. Chai, S. Shao, Y. Song, S. Qi, R. Rui, and W. Zhang (2025b) AgentNet: decentralized evolutionary coordination for llm-based multi-agent systems. External Links: 2504.00587, Link Cited by: §2.2, §5. S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan (2023) Tree of thoughts: deliberate problem solving with large language models. External Links: 2305.10601, Link Cited by: §1, §2.1, §5. Y. Zhang, R. Sun, Y. Chen, T. Pfister, R. Zhang, and S. Ă. ArÄąk (2024) Chain of agents: large language models collaborating on long-context tasks. In Advances in Neural Information Processing Systems, A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang (Eds.), Vol. 37, p. 132208â132237. External Links: Document, Link Cited by: §1. Appendix A Experiment Setup Details A.1 Additional Experiment Setup Main experiment setting We evaluated Team-of-Thoughts MAS across a diverse suite of seven model families, comprising three closed-source models: Claude-Sonnet-4.5 (PBC, 2025), GPT-5-mini (Singh et al., 2025), and Gemini-3-Flash-Preview (Pichai et al., 2025) and four open-source models: DeepSeek-V3.2-Exp (DeepSeek-AI et al., 2025), GPT-OSS-20B (OpenAI et al., 2025), Qwen3-VL-235B-A22B-Thinking (Yang et al., 2025a), and Phi-4 (Abdin et al., 2024). Two tool agents are activated on each tool call. Our assessment spanned two domains: mathematical reasoning (AIME2024 (of America, 2024), AIME2025 (of America, 2025)) and code generation (Humaneval+ (Chen et al., 2021), MBPP+ (Austin et al., 2021), and LiveCodeBench v6 (Jain et al., 2024) with problems released between 2025/01/01 and 2025/05/01). Unless stated otherwise, we set a standardized context window for each tool-agent: 20,000 tokens for AIME tasks and 4,096 tokens for coding tasks. For the orchestrator, we used a 16,384 token context window across all tasks to ensure sufficient capacity for processing tool descriptions and making informed selection and reasoning decisions. For reasoning models, we applied the default âmediumâ effort setting, capping the reasoning token budget at 50% of the maximum generation length to ensure consistent comparisons across all baselines. By default, we randomly sample 10% of the target tasks to construct a calibration dataset. However, for AIME2024 and AIME2025, 10% corresponds to only three data points, which is too small to provide a reliable characteristic modeling. We instead sample 50% of the problems from AIME2025 and use them as the calibration dataset for evaluating AIME2024 and vice versa for AIME2025. This cross-calibration setup ensures sufficient calibration data while avoiding exposure of the test data. Orchestration selection In the orchestration agent selection experiment, we judge the performance of the orchestrator by activating all tool agents. The max generation token is set based on each agentâs cost, ensuring it will not exceed the cost budgets. Scaling tool agents experiment setting For this analysis, we select the top 100 OpenRouter models ranked by popularity that provide a context length greater than 16K tokens and cost less than $1 USD per million input tokens. The orchestrator context window is set as 16,384 unless stated otherwise. Due to varying max context lengths of different models, we employ Orchestrator-Led Assessment using GPT-5 Mini. The orchestrator is provided with the tool agents with top single-model performance on MBPP+ when a subset of the models is made available. A.2 AgentVerse Setup We employ GPT-5-Mini (Singh et al., 2025) as the backbone language model of agents in AgentVerse (Chen et al., 2024). We use a maximum token limit of 512 for the role assigner agent, and 4,096 for the rest of the agents. For invalid agent outputs, such as invalid role assignments, unparseable answers, or errors in code, AgentVerse retries generation a limited number of times: 10 times on math tasks and 1,000 times on coding tasks. Appendix B Impact of Calibration Dataset Choice We present additional results using a categorized calibration set. To construct the categorized calibration dataset, we first randomly sampled 10% of the LiveCodeBench v6 dataset as before, resulting in 18 questions. We then used GPT-5 to group these questions according to the primary algorithmic concept involved, such as Arrays / Data Structures, Strings / Sequences, Math / Logic, Grid / Matrix, Simulation / Greedy, Knapsack / Optimization, and Intervals / Range. From each category, one representative question was selected to form the final calibration dataset. Table 7: Performance in pass@k on LiveCodeBench v6 using different calibration construction strategies. Construction Method pass@1 pass@3 pass@5 Random Sample 77.91% 85.05% 87.91% Categorized Sample 81.65% 85.93% 86.81% Figure 4: Comparing the times of tool agent invocations between two calibration dataset construction methods. As shown in TableË7, the categorized calibration dataset leads to a noticeable improvement in pass@1 accuracy. FigureË4 compares the distribution of Team-of-Thoughts tool selections under the two calibration strategies. When using the categorized calibration set, the system selects the strongest single model approximately twice as often as when using the random calibration set. This explains the improved pass@1 performance. However, the categorized calibration set also results in lower diversity in tool selection across generations. Consequently, the pass@5 performance is slightly lower than that obtained with the random calibration dataset. This suggests a trade-off between bias and variance: while the categorized calibration set biases the system toward stronger single-model decisions, it reduces exploration across different tools. Designing calibration datasets that better balance this trade-off remains an interesting direction for future work. Appendix C Longer Context Window with 100-Tool Agent Ensemble Table 8: Performance under different orchestrator max context lengths, with eight of 25 available tool agents activated on each tool call. Orch context len Pass@1 Pass@3 Pass@5 16K 60.32% 83.76% 86.24% 32K 59.79% 82.99% 85.71% 64K 59.42% 83.39% 86.51% 128K 60.32% 82.96% 85.45% To verify whether the performance decline with eight activated agents stems from context window limitations, we evaluated the 25-model ensemble using extended orchestration contexts. As shown in TableË8, increasing context length yielded no performance gains. This confirms that the bottleneck resides in the orchestratorâs aggregation capability rather than spatial context constraints. Appendix D Experiment Prompt Here is an example of the language model-based tool agent assessment prompts. Assessment Prompt Instructions: You will be provided with a series of problems. For each problem, the Subject Agent has provided an answer. Some are correct, and some are incorrect. Part 1: Per-Problem Analysis For every problem provided, generate a structured audit containing: 1. Taxonomy: Classify the problem type (e.g., Arithmetic, Logical Reasoning, Creative Writing, Coding) and specific skill required. 2. Performance Verdict: Clearly state [PASS] or [FAIL]. 3. Gap Analysis: ⢠If Correct: Briefly explain why the agent succeeded (e.g., âGood step-by-step reasoning,â âRobust knowledge retrievalâ). ⢠If Incorrect: Pinpoint exactly where and why the agent failed. Was it a calculation error? A logic jump? A hallucination? A misunderstanding of constraints? Compare the Subjectâs logic to the Ground Truth. Part 2: Executive Summary After analyzing all problems, synthesize a âModel Personaâ profile: 1. Core Competencies: List specific categories where the agent consistently succeeds. 2. Blind Spots & Failure Modes: Describe the patterns in the agentâs errors (e.g., âThe agent struggles with negative integers,â or âThe agent is verbose but inaccurateâ). 3. Final Verdict: A 2-sentence summary of the agentâs reliability. â Input Data: ⢠Problem 1: Question, Subject Agent Answer, Ground Truth/Solution⌠⢠Problem 2: âŚrepeat for all samples⌠â COMMAND: Based on the data above, proceed with the Per-Problem Analysis followed by the Executive Summary. Reminders: 1. Be Specific: Do not just say âThe agent failed.â Identify why (e.g., Logic Error vs. Calculation Error). 2. Be Critical: Compare the Subject Answer against the Ground Truth rigorously. 3. Format: Use the headers ## Part 1: Per-Problem Analysis and ## Part 2: Executive Summary. GENERATE REPORT: