Paper deep dive
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning
Yutong Wang, Siyuan Xiong, Xuebo Liu, Wenkang Zhou, Liang Ding, Miao Zhang, Min Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 8:58:41 AM
Summary
The paper introduces AgentDropoutV2 (ADv2), a test-time rectify-or-reject pruning framework designed to optimize information flow in Multi-Agent Systems (MAS). ADv2 acts as an active firewall that intercepts agent outputs, uses a retrieval-augmented rectifier guided by an offline-constructed failure-driven indicator pool to iteratively correct errors, and prunes irreparable outputs to prevent error propagation. The method significantly improves accuracy on math and code benchmarks compared to baseline MAS frameworks like AutoGen.
Entities (9)
Relation Signals (6)
AgentDropoutV2 → optimizes → Multi-Agent Systems
confidence 95% · AgentDropoutV2 (ADv2), a test-time rectify-or-reject pruning framework that dynamically optimizes MAS information flow.
AgentDropoutV2 → uses → Indicator Pool
confidence 92% · This rectification is guided by an indicator pool, which is constructed offline by distilling error patterns from historical MAS failure trajectories.
Indicator Pool → constructedfrom → historical MAS failure trajectories
confidence 90% · constructed offline by distilling error patterns from historical MAS failure trajectories.
AgentDropoutV2 → improvesperformanceon → GSM8K
confidence 90% · Empirical results demonstrate that ADv2 significantly boosts performance on both fixed and dynamic MAS frameworks, achieving average accuracy gains... on extensive math and code benchmarks
AgentDropoutV2 → implementedin → AutoGen
confidence 88% · We employ the SelectorGroupChat framework within AutoGen... thereby grounding our implementation in a widely used infrastructure
AgentDropoutV2 → usesbackbonemodel → Qwen3-8b
confidence 85% · For the reasoning components encompassing all participants and rectifiers, we deploy Qwen3-8B
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Multi-Agent Systems (MAS) excel in complex reasoning, they suffer from the cascading impact of erroneous information from individual agents. Current solutions often resort to rigid structural engineering or expensive fine-tuning, limiting their adaptability. We propose AgentDropoutV2 (ADv2), a test-time rectify-or-reject pruning framework that dynamically optimizes MAS information flow. Acting as an active firewall, ADv2 intercepts agent outputs and employs a retrieval-augmented rectifier to iteratively correct errors. This rectification is guided by an indicator pool, which is constructed offline by distilling error patterns from historical MAS failure trajectories. Irreparable outputs are subsequently pruned to prevent error propagation. Empirical results demonstrate that ADv2 significantly boosts performance on both fixed and dynamic MAS frameworks, achieving average accuracy gains of 6.39 and 2.28 percentage points on extensive math and code benchmarks, respectively. Furthermore, ADv2 exhibits remarkable adaptivity, dynamically modulating rectification efforts based on task difficulty to resolve a wide spectrum of error patterns. Our code is released at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2602.23258v2
- Canonical: https://arxiv.org/abs/2602.23258v2
Trouble viewing inline? Open PDF directly →
Full Text
87,189 characters extracted from source content.
Expand or collapse full text
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning Yutong Wang Siyuan Xiong Xuebo Liu Wenkang Zhou Liang Ding Miao Zhang Min Zhang Abstract While Multi-Agent Systems (MAS) excel in complex reasoning, they suffer from the cascading impact of erroneous information generated by individual participants. Current solutions often resort to rigid structural engineering or expensive fine-tuning, limiting their deployability and adaptability. We propose AgentDropoutV2, a test-time rectify-or-reject pruning framework designed to dynamically optimize MAS information flow without retraining. Our approach acts as an active firewall, intercepting agent outputs and employing a retrieval-augmented rectifier to iteratively correct errors based on a failure-driven indicator pool. This mechanism allows for the precise identification of potential errors using distilled failure patterns as prior knowledge. Irreparable outputs are subsequently pruned to prevent error propagation, while a fallback strategy preserves system integrity. Empirical results on extensive math benchmarks show that AgentDropoutV2 significantly boosts the MAS’s task performance, achieving an average accuracy gain of 6.3 percentage points on math benchmarks. Furthermore, the system exhibits robust generalization and adaptivity, dynamically modulating rectification efforts based on task difficulty while leveraging context-aware indicators to resolve a wide spectrum of error patterns. Our code and dataset are released at https://github.com/TonySY2/AgentDropoutV2. Figure 1: Overview of AgentDropoutV2 versus AgentDropout. While AgentDropout directly discards erroneous agents, AgentDropoutV2 attempts iterative rectification before elimination. 1 Introduction Large language model (LLM)-based agents have achieved outstanding performance across a wide range of tasks, including reasoning (Yao et al., 2023), planning (Prasad et al., 2024), and action (Park et al., 2023). Despite the sophisticated designs that have enabled these agents to achieve significant gains, the single-model paradigm remains a bottleneck that limits their potential. Consequently, a growing body of research has shifted focus towards designing multi-agent systems (MAS) to address more complex scenarios (Li et al., 2023; Guo et al., 2024). By harnessing collective intelligence (Zhuge et al., 2024; Wu et al., 2024) and orchestrating cooperative teams (Zhang et al., 2025d; Dang et al., 2025), MAS achieves remarkable performance in complex tasks such as software development (Hong et al., 2024; Qian et al., 2024), ultra-long context handling (Li et al., 2024a; Zhao et al., 2024), and scientific discovery (Ghafarollahi and Buehler, 2025; Ghareeb et al., 2025). However, the structural complexity of MAS also renders them susceptible to erroneous outputs from individual participants due to error propagation (Zhang et al., 2025f; Pan et al., 2025b). This necessitates the timely identification and pruning of incorrect information to prevent it from cascading to downstream agents and ultimately compromising the entire task. To mitigate the impact of errors, current research has predominantly diverged into two main paradigms: Structural Optimization and Parameter Internalization. The former seeks to constrain error pathways by engineering robust communication topologies, such as optimizing directed acyclic graphs (DAG) (Zhang et al., 2025c; Wang et al., 2025b; Zhang et al., 2025e). The latter focuses on enhancing the intrinsic reasoning of agents by fine-tuning them on failure trajectories (Motwani et al., 2025; Zhao et al., 2025) or utilizing process-supervision data (Lightman et al., 2024; Wang et al., 2025a; Zhang et al., 2025b). However, despite their contributions, these paradigms share a critical bottleneck: the reliance on offline optimization at the expense of test-time adaptivity. As illustrated in Figure 1, methods like AgentDropout rely on pre-determined structural priors derived from training statistics. They enforce a static connectivity graph that permanently excludes certain agents without attempting to rehabilitate their outputs or rectify their errors. Similarly, parameter-based methods depend on frozen weights, rendering them incapable of dynamic correction. This static nature prevents the system from salvaging potentially correctable errors during inference, highlighting the urgent need for a test-time rectification framework that can actively intercept and resolve failures in real-time. To this end, we introduce AgentDropoutV2, an MAS information flow optimization framework based on test-time rectify-or-reject pruning. During the execution process, our method intercepts the output of each participant agent to perform iterative rectification before it is broadcast to downstream successors. Specifically, a dedicated rectifier is prompted to scrutinize the output using adversarial indicators retrieved from a pre-constructed pool of prior failure patterns, generating targeted feedback if errors are detected. If the rectification fails to resolve the issues, the erroneous output is pruned to strictly prevent error propagation. Experimental results demonstrate that our method significantly enhances MAS performance across diverse mathematical and code generation benchmarks by effectively rectifying and eliminating erroneous agent outputs. Extended analyses further confirm the system’s adaptability, showing its capability to dynamically retrieve context-aware indicators based on task complexity, and to efficiently resolve distinct error patterns through variable iterative refinement. The observed correlation between pruning rates and reasoning difficulty positions our framework as a potential task difficulty evaluator. Our main contributions are listed as follows: • We propose a test-time rectify-or-reject pruning method that intercepts and iteratively corrects agent outputs to effectively block error propagation in MAS, thereby safeguarding task performance against cascading degradation. • We construct a failure-driven indicator pool by distilling error patterns from failed MAS trajectories, providing an off-the-shelf knowledge base that encapsulates a broad spectrum of reasoning pitfalls for precise error identification. • We demonstrate that our method exhibits robust adaptivity across diverse task complexities and scenarios, confirming its effectiveness and generalization capability as a plug-and-play intervention solution. 2 Preliminary Agent Definition We formulate the MAS workflow as an ordered sequence of N agents, denoted as =(A1,A2,…,AN)S=(A_1,A_2,…,A_N). Each agent in this sequence AiA_i is selected from a candidate set of all available agents A, and can be defined as a tuple of three primary elements: Ai=(Φi,ℛi,i), A_i= ( _i,R_i,K_i ), (1) where: (1) Φi(⋅) _i(·) represents the backbone model serving as the reasoning engine, which maps the input context to textual output; (2) ℛiR_i denotes the role specification, a static set of instructions defining the agent’s persona, responsibilities, and constraints; (3) iK_i represents the knowledge base, a dynamic information repository containing the history of messages observable by agent AiA_i (Initially i=∅K_i= ). For an active agent AiA_i receiving an input denoted as xix_i, it utilizes its backbone model Φi _i to generate the output oio_i conditioned on its profile ℛiR_i and current knowledge iK_i: oi=Φi(xi,ℛi,i). o_i= _i (x_i,R_i,K_i ). (2) Information Flow Once the output oio_i is generated, its dissemination is determined by the system’s architecture, which is formalized as a mapping function :→2N:A→ 2^A, which maps the current agent AiA_i to a set of successor agents who are designated to receive the information. Then, the system updates the knowledge base of every successor agent Aj∈(Ai)A_j (A_i) by integrating the new message oio_i: eq:broadcastj←j∪(ℛi,oi),∀Aj∈(Ai). eq:broadcastK_j _j∪ \ (R_i,o_i ) \,\ ∀ A_j (A_i). (3) Through this mechanism, the framework-specialized mapping N rigidly controls the information flow topology, ranging from a broadcast structure (e.g., AutoGen) where (Ai)=N(A_i)=A, to a sequential chain where |(Ai)|=1 (A_i) =1. Control Flow Given a task Q, the control flow of the MAS is modeled as the construction of an ordered sequence of agents, referred to as the inference trajectory. Specifically, upon the generation of output oio_i by the current agent AiA_i, a routing policy π determines the next active agent Ai+1A_i+1 based on the task, the existing sequence of activated agents, and their historical outputs: Ai+1=π(,A1:i,o1:i,). A_i+1=π (Q,A_1:i,o_1:i,A ). (4) This iterative process constructs the execution path dynamically or statically, depending on the definition of π from the MAS framework. The workflow concludes when the sequence reaches the terminal output agent ANA_N. Consequently, the final answer Y to the initial user task Q is defined as the output generated by this final agent, namely =oNY=o_N. Figure 2: Overview of the proposed framework. The upper block shows the test-time pipeline for iteratively rectifying agent outputs within the MAS. The lower block demonstrates the offline construction of the indicator pool via failure-driven mining and dual-stage deduplication. 3 Methodology We present a test-time framework designed to intercept and refine agent outputs during the MAS execution. Specifically, before transmitting the output from agent AiA_i to its successors (Ai)N(A_i) (as defined in Eq. LABEL:eq:broadcast), we actively intercept the message. A dedicated rectifier then scrutinizes the content for potential errors and attempts to resolve them through an iterative refinement process. If the output remains flawed despite these efforts, it is discarded rather than propagated, ensuring that downstream agents are shielded from unreliable information. 3.1 Test-Time Rectify-or-Reject Pruning Blindly prompting an agent to self-correct is often counterproductive; without specific direction on what went wrong, the agent may inadvertently introduce new hallucinations or simply rephrase the original error. To ensure the rectification is effective, it is essential to ground the refinement process on specific, verifiable standards. Therefore, we employ adversarial indicators to scrutinize the output for distinct error patterns. If specific error types are detected, these indicators guide the generation of targeted feedback, providing the agent with a clear roadmap for correction. Relevant Indicator Retrieval To support this targeted supervision, our framework incorporates an Indicator Pool, denoted as ℐI. Constructed offline via a failure-driven mining strategy (detailed in §3.2), this repository encapsulates empirical knowledge regarding a wide spectrum of potential errors that may emerge during MAS execution. Each indicator within this pool is structured as a tuple I=(n,d,c)I=(n,d,c): • n (Name): A unique identifier for the specific error type. • d (Error Definition): A description of the erroneous behavior, which serves as the standard to verify whether the agent’s output has deviated from requirements. • c (Trigger Condition): A context describing when this specific error is likely to occur, which acts as a filter to ensure the indicator is only retrieved in relevant scenarios. Leveraging this structured repository, we can now retrieve the most pertinent indicators to supervise the current reasoning step. For an active agent AiA_i producing an output oi(t)o_i^(t) at the t-th iteration (initially oi(0)o_i^(0)), we first employ a dedicated Rectifier Model Φrect _rect to distill the semantic essence of the reasoning context. The rectifier extracts two distinct sets of keywords: (1) scen(t)S_scen^(t), summarizing the task scenarios (e.g., geometric coordinates, algebraic operations, etc); (2) act(t)S_act^(t), representing the specific action types proposed by the agent. We transform these keywords into a query vector i(t)=Memb(scen(t)⊕act(t))q_i^(t)=M_emb(S_scen^(t) _act^(t)) using an embedding model MembM_emb. Subsequently, we retrieve the top-KactK_act most relevant indicators from ℐI whose trigger conditions exhibit the highest semantic similarity to the current query, forming the active indicator set ℐact(t)I_act^(t): ℐact(t)=Top-KactIj∈ℐ(i(t)⋅j|i(t)||j|), _act^(t)= I_j Top-K_act ( q_i^(t)·c_j|q_i^(t)||c_j| ), (5) where jc_j represents the trigger condition cjc_j’s embedding. Rectify-or-Reject Pruning The rectifier then evaluates the output oi(t)o_i^(t) against each retrieved indicator Ik=(nk,dk,ck)I_k=(n_k,d_k,c_k), conditioned on the agents input xix_i and role ℛiR_i. For each indicator, the model generates a binary violation flag vk(t)∈0,1v_k^(t)∈\0,1\ and a diagnostic rationale rk(t)r_k^(t): (vk(t),rk(t))=Φrect(oi(t)∣xi,ℛi,Ik), (v_k^(t),r_k^(t) )= _rect (o_i^(t) x_i,R_i,I_k ), (6) where vk(t)=1v_k^(t)=1 signifies that the specific constraint defined by IkI_k has been violated. We enforce a strict zero-tolerance policy for the rectification procedure, where the global error state E(t)E^(t) is immediately activated if any single indicator within the active set detects a violation. Consequently, we derive the binary global error state E(t)E^(t) and aggregate the specific feedback ℱ(t)F^(t) as: E(t) E^(t) =maxIk∈ℐact(t)vk(t) = _I_k _act^(t)v_k^(t) (7) ℱ(t) ^(t) =rk(t)∣Ik∈ℐact(t)∧vk(t)=1. = \r_k^(t) I_k _act^(t) v_k^(t)=1 \. (8) The rectification trajectory follows a tri-state gating mechanism derived from the global assessment result E(t)E^(t): • Pass: If no significant error is detected (E(t)=0E^(t)=0), the output is accepted immediately: oi=oi(t)o_i=o_i^(t). • Retry: If the global error state is activated (E(t)=1E^(t)=1) and the iteration count has not exceeded the preset upper limit (t<Tmaxt<T_max), the agent regenerates its output conditioned on the feedback ℱ(t)F^(t): eq:reflectionoi(t+1)=Φi(xi,ℛi,i,ℱ(t)). eq:reflectiono_i^(t+1)= _i (x_i,R_i,K_i,F^(t) ). (9) • Reject: If errors persist at the maximum iteration (E(Tmax)=1E^(T_max)=1), the output is discarded (oi=∅o_i= ) to act as a semantic circuit breaker, preventing error propagation to downstream nodes. Ultimately, the final message transmitted to the successor set (Ai)N(A_i) is defined as: oi=oi(t)if ∃t≤Tmax s.t. E(t)=0,∅otherwise. o_i= caseso_i^(t)&if ∃ t≤ T_max s.t. E^(t)=0,\\ &otherwise. cases (10) Global Fallback against Structural Degeneration While pruning ensures purity, excessive filtering risks destroying connectivity. Analogous to the principle of “critical mass” in collaborative dynamics—which posits that a group must maintain a sufficient size to sustain effective interaction and consensus—we introduce a safeguard against structural collapse. If the remaining message count falls below a safety threshold γ, the MAS is deemed to have lost its reasoning integrity. Instead of forcing a conclusion from this sparse context, we trigger a system-wide reset to conduct MAS execution from scratch. This guarantees the final solution always emerges from a sufficiently robust consensus, preventing degeneration into fragmented reasoning. Handling Zero-Shot Scenarios In scenarios where a training dataset is unavailable to construct a domain-specific indicator pool, our framework provides an optional solution. We initialize ℐactI_act with a single, universally applicable general indicator Igen=(“General Logic Check”,dgen,cgen)I_gen=(``General Logic Check′,d_gen,c_gen). Here, dgend_gen prompts the model to check for logical consistency and hallucination, while cgenc_gen is set to always trigger. This ensures that the rectify-or-reject pruning remains functional and effective even in zero-shot settings without prior failure pattern mining. The pseudo-code for the test-time rectify-or-reject pruning is provided in Appendix A.1, and detailed prompt specifications are listed in Appendix A.2. Additionally, a comprehensive case study demonstrating the rectification process is presented in Appendix A.4. 3.2 Failure-Driven Indicator Pool Construction Just as mature organizations rely on institutional memory, which codifies lessons learned from past projects to prevent the recurrence of known pitfalls, our framework necessitates a structured repository of error patterns. Blindly correcting errors without understanding their origins is inefficient; effective rectification requires a reference to historical mistakes. To this end, we construct a repository of adversarial indicators by mining historical failure cases. This process transforms raw failure trajectories into a structured knowledge base, serving as a comprehensive handbook of prohibitions to guide the agent’s real-time rectification. Table 1: Performance comparison of our method against baseline reasoning techniques across mathematical domain benchmarks. “OlymB”, “OlymE”, and “OlymH” represent OlympiadBench, OlymMATH Easy, and OlymMATH Hard, respectively. System GSM8K MATH-500 AQuA AMC23 OlymB OlymE OlymH AIME24 AIME25 Average Single 87.64 74.80 84.19 62.50 47.56 20.00 16.00 13.33 20.00 47.34+0.00 AutoGen 91.36 76.80 85.03 62.50 48.89 19.00 17.00 26.67 13.33 48.95+1.62 w/ Generic Indicators 90.52 77.40 84.98 72.50 50.37 26.00 11.00 33.33 23.33 52.16+4.82 w/ Retrieved Indicators 91.74 78.40 87.01 75.00 51.11 25.00 19.00 40.00 30.00 55.25+7.92 Offline Indicator Mining As illustrated in the lower block of Figure 2, we focus on collecting execution strategies where the MAS fails to deliver the correct solution. Let src=,∗D_src=\Q,Y^*\ denote the source dataset, where Q and ∗Y^* represent the input query and the corresponding ground-truth answer, respectively. For each instance, we conduct a full inference roll-out to obtain the MAS execution trajectory =(,A1:N,o1:N,)T=(Q,A_1:N,o_1:N,Y). We collect failure cases where the solution Y diverges from the ground truth ∗Y^* into a failure set failD_fail. A teacher model Φteach _teach then scrutinizes individual agents within failD_fail. Upon detecting a deviation in agent AiA_i’s output oio_i given its role ℛiR_i and the overall task Q, Φteach _teach synthesizes a set of indicators: ℐnew=Φteach(,∗,ℛi,oi). _new= _teach (T,Y^*,R_i,o_i ). (11) Redundancy Elimination Since identical or highly similar error patterns frequently recur across different failure trajectories, a naive accumulation of indicators would result in a bloated repository saturated with duplicate constraints. Such redundancy poses a critical risk to the retrieval mechanism, as it may cause the top-KactK_act retrieved indicators to be dominated by a single error type, thereby crowding out other diverse but equally critical constraints and limiting the multi-dimensional evaluation of agent outputs. To prevent this semantic collapse and ensure a compact, high-entropy global pool ℐI, we employ a dual-stage deduplication process. For a newly generated indicator InewI_new, we obtain its semantic vector new=Memb(dnew⊕cnew)v_new=M_emb(d_new c_new) by encoding the concatenation of its description and triggering condition using an embedding model MembM_emb. Then we retrieve the most similar existing indicators ℐsim⊂ℐI_sim of size KdedupK_dedup based on cosine similarity. A deduplication LLM Φdedup _dedup is then employed to verifies redundancy, adding InewI_new to ℐI only if it represents a novel error pattern: ℐ←ℐ∪Inewif Φdedup(Inew,ℐsim),ℐotherwise. ← casesI∪ \I_new \\ &if _dedup (I_new,I_ sim ),\\ I&otherwise. cases (12) Examples of the constructed indicators, along with the general indicators used for scenarios where a specific pool is unavailable, are provided in Appendix A.2. 4 Experiment Table 2: Performance comparison of our method against baselines using Qwen3-4B as the backbone model. The indicator pool is transferred directly from a Qwen3-8B source model to the Qwen3-4B agents to test transferability. System GSM8K MATH-500 AQuA AMC23 OlymB OlymE OlymH AIME24 AIME25 Average Single 84.53 75.00 82.68 62.50 44.44 23.00 16.00 13.33 26.67 47.57+0.00 AutoGen 88.02 73.60 85.04 67.50 47.56 23.00 13.00 30.00 16.67 49.38+1.81 w/ Generic Indicators 88.32 77.00 85.04 62.50 49.33 21.00 13.00 30.00 20.00 49.58+2.01 w/ Retrieved Indicators 90.67 77.04 87.80 70.00 47.70 21.00 14.00 26.67 20.00 50.54+2.97 Table 3: Performance comparison of our method against baseline reasoning techniques across code domain benchmarks. System MBPP HumanEval CodeContests LiveCodeBench Average Single 64.98 81.37 4.85 24.50 43.93+0.00 AutoGen 65.37 85.09 6.06 29.25 46.44+2.52 w/ Generic Indicators 68.09 84.50 9.26 32.75 48.65+4.72 4.1 Experimental Setup MAS Framework We employ the SelectorGroupChat111https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/selector-group-chat.html framework within AutoGen (Wu et al., 2024)—the current de facto standard for MAS—thereby grounding our implementation in a widely used infrastructure that features a classic automatic routing mechanism. In this setup, a selector iteratively identifies the next speaker based on context, where a decision agent formulates the final conclusion. Crucially, communication is globally transparent, meaning that every message is broadcast to all participants, establishing a shared reasoning environment. Backbone Models We adopt GPT-4.1-mini-2025-0414222https://platform.openai.com/docs/models/gpt-4.1-mini as the backbone of the AutoGen MAS selector. For the reasoning components encompassing all participants and rectifiers, we deploy Qwen3-8B333https://huggingface.co/Qwen/Qwen3-8B and Qwen3-4B, configured with the thinking mode explicitly disabled. For the offline indicator pool construction process, GPT-4o-2024-08-06 and GPT-4.1-mini-2025-0414 serve as the foundation for the teacher and deduplicator, respectively. Finally, Qwen3-Embedding-8B is adopted as the embedding model MembM_emb. Datasets We comprehensively evaluate the performance of our method across two primary domains: mathematical reasoning and code generation. For mathematical reasoning, we employ nine benchmarks spanning a spectrum of difficulty levels, including GSM8K (Cobbe et al., 2021), MATH-500 (Lightman et al., 2024), AQuA (Patel et al., 2021), AMC23444https://huggingface.co/datasets/math-ai/amc23, OlympiadBench (He et al., 2024), OlymMATH Easy, OlymMATH Hard (Sun et al., 2025), AIME24 (Zhang and Math-AI, 2024), and AIME25 (Zhang and Math-AI, 2025). For code generation capabilities, we assess the model on four established datasets: MBPP (Austin et al., 2021), HumanEval (Chen et al., 2021), CodeContests (Li et al., 2022), and LiveCodeBenchV1 (Jain et al., 2025). Regarding the indicator pool construction, we leverage the training splits of MATH and AQuA as source corpora to sample trajectories and distill adversarial indicators specifically for the mathematical domain. Detailed statistics for each dataset and the indicator pool are provided in Appendix A.3. Hyper-Parameters We set the max chat turns of SelectorGroupChat to 6, and the max reflection turns TmaxT_max to 3. We set the number of retrieved indicators for test-time matching to Kact=5K_act=5, while the retrieval count for the deduplication process during pool construction is set to Kdedup=20K_dedup=20. The safety threshold γ of the remaining message count for triggering the global fallback against structural degeneration is set to 1. The temperature of the rectifier is set to 0, and the others remain 0.7. Table 4: Results of the ablation study. System GSM8K MATH500 AQuA AMC23 OlymB OlymE OlymH AIME24 AIME25 Average Ours 91.74 78.40 87.01 75.00 51.11 25.00 19.00 40.00 30.00 55.25 (I) Rectification Iteration Rounds (TmaxT_max, Default: 3) 0 Iterations 91.96 76.40 87.80 70.00 46.37 19.00 14.00 23.33 26.67 50.61 2 Iterations 91.96 78.20 86.61 72.50 46.52 32.00 16.00 30.00 16.67 52.27 4 Iterations 91.74 78.40 86.22 72.50 49.33 30.00 12.00 36.67 16.67 52.61 (I) Number of Retrieved Indicators (KactK_act, Default: 5) 3 Indicators 91.74 77.00 85.04 67.50 52.59 25.00 17.00 33.33 30.00 53.24 8 Indicators 91.74 78.20 86.96 65.00 48.15 29.00 18.00 33.33 23.33 52.63 (I) Indicator Retrieval Mechanism Random 5 91.66 76.60 83.46 67.50 48.00 25.00 13.00 23.33 23.33 50.21 (IV) Indicator Pool Deduplication w/o Dedup 92.04 79.40 86.61 72.50 50.37 26.00 17.00 30.00 23.33 53.03 4.2 Main Results Table 1 presents the comparative performance of our proposed framework against baseline approaches across nine mathematical reasoning benchmarks with Qwen3-8B as the backbone model. Our full method (w/ Retrieved Indicators) demonstrates consistent superiority, surpassing the baselines across all benchmarks. It achieves the highest average accuracy of 55.25%, achieving an average accuracy gain of 6.3 percentage points compared to the AutoGen baseline. While the native AutoGen framework provides a marginal improvement over the single-agent setting (+1.62% on average), it struggles in complex scenarios (e.g., dropping to 13.33% on AIME25). By introducing the feedback-based mechanism in the absence of a pre-built indicator pool (w/ Generic Indicators), which applies a generic verification logic, we observe a substantial performance leap to 52.16%. This confirms that the rectify-or-reject architecture itself provides a robust safety net for multi-agent reasoning. Crucially, by retrieving task-specific constraints from the equipped indicator pool (w/ Retrieved Indicators), our method further achieves massive gains on highly difficult tasks like AIME25 (improving from 23.33% to 30.00%). This demonstrates that while the rectification mechanism provides the means to correct errors, the indicator pool provides the necessary guidance to accurately pinpoint issues and ensure effective refinement. 4.3 Cross-Model and Cross-Domain Transferability Indicator Portability across Models We investigate scalability by deploying the indicator pool mined by Qwen3-8B directly to a smaller Qwen3-4B backbone. As shown in Table 2, this yields robust gains across most benchmarks, confirming that fundamental reasoning pitfalls are largely scale-invariant. While performance plateaus on complex tasks due to minor misalignments between high-level indicators and rudimentary failures, the overall success validates a “build once, deploy anywhere” paradigm. This enables capable models to construct offline knowledge bases that effectively supervise resource-constrained edge models without redundant mining. Cross-Domain Generalization We further assess the versatility of our framework by extending it to code generation, a domain sharing the rigorous logic requirements of mathematics. As shown in Table 3, our method consistently outperforms standard baselines, achieving a superior average accuracy of 48.65% compared to AutoGen’s 46.44%. Notably, the improvements are most pronounced in complex benchmarks like CodeContests (6.06% → 9.26%) and LiveCodeBench (29.25% → 32.75%). This confirms that the rectify-or-reject pruning is not limited to math but serves as a generalizable reasoning enhancer, effectively mitigating errors across diverse complex reasoning tasks. 5 Analysis 5.1 Ablation Study Impact of Rectification Iteration Rounds We first examine the impact of the rectification iteration budget TmaxT_max on overall performance. As shown in Block I of Table 4, setting Tmax=0T_max=0 (no rectification) leads to a sharp performance drop, especially on complex tasks like AIME24, confirming that initial outputs often contain errors requiring active correction. However, increasing the budget to Tmax=4T_max=4 does not yield further improvements, suggesting that excessive iterations may induce over-correction or introduce noise. Thus, our default Tmax=3T_max=3 strikes the optimal balance between efficiency and thoroughness. Figure 3: Distribution of rectification iterations across different benchmarks. Simpler tasks exhibit high first-pass rates, whereas complex tasks necessitate more refinement rounds and result in higher rejection rates due to persistent errors. This contrast demonstrates that our method dynamically modulates its intervention intensity according to task complexity. Figure 4: Jaccard similarity between the set of ten most frequently used indicators across different benchmarks. Indicators chosen for similar tasks tend to have higher overlaps. This distribution reveals that our indicator pool is diverse enough to cover a wide range of failure modes. Sensitivity to Retrieved Indicator Count Next, we explored the impact of the number of retrieved indicators as shown in Block I of Table 4. Both reducing the retrieved count to k=3k=3 and increasing it to k=8k=8 degrade performance compared to the optimal setting of k=5k=5. This indicates that while agents benefit from diverse failure patterns, providing an excessive number of indicators results in information overload, distracting the model with less relevant constraints rather than aiding the reasoning process. Effectiveness of the Retrieval Mechanism To validate that performance gains stem from relevant guidance, we replaced the retrieved indicators with five indicators randomly sampled from the pool. As shown in Block I of Table 4, the average accuracy decreases to 50.21%, a score even lower than the “0 Iterations” setting. This critical comparison proves that the system’s success depends strictly on the semantic relevance of the constraints, which is necessary for locating specific error patterns to guide the agent. Necessity of Pool Deduplication Finally, we demonstrate the necessity of the deduplication operation during indicator pool construction. As shown in Block IV of Table 4, removing the dual-stage deduplication process also causes an average accuracy decrease. This decline suggests that without the deduplication process, the retrieved top-k indicators are occupied by redundant variations of the same or similar error patterns. As a result, this lack of diversity prevents the agent from receiving a comprehensive safety check. This validates the importance of our compact, high-entropy pool construction strategy. 5.2 Iteration Dynamics and Adaptability We analyze the distribution of iteration rounds across varying difficulties to evaluate the adaptability of our method. The results are visualized in Figure 3, where “Pass @ k-th” indicates the proportion of outputs successfully rectified and accepted at the k-th iteration, while “Rejected” denotes instances that remained erroneous after exhausting the maximum budget. A pronounced correlation exists between task complexity and rectification depth. Simpler datasets (e.g., GSM8K) exhibit a high “Pass @ 1st” rate (60.1%), indicating immediate acceptance. Conversely, complex tasks like AIME 24/25 show a significant shift toward multi-round rectifications and rejection rates exceeding 60%. This demonstrates that our method dynamically modulates intervention intensity, conserving resources on simple queries while allocating sustained effort to resolve intricate errors in challenging scenarios. Moreover, this strong correlation allows our framework to double as a potential difficulty evaluator, where the aggregate rectification depth and rejection rate serve as quantifiable proxies for dataset complexity. 5.3 Distribution of Retrieved Indicators To provide a deeper insight into the composition and utility of our constructed indicator pool, we analyzed the overlaps of the retrieved active indicators across distinct task domains. Specifically, we identify the sets of ten most frequently retrieved indicators for different benchmarks and calculate the pair-wise Jaccard similarity (the ratio of the size of the intersection to the size of the union). The resulting heatmap in Figure 4 reveals a distinct block-wise correlation pattern. Benchmarks requiring similar reasoning capabilities exhibit high indicator overlap. For instance, foundational math datasets like GSM8K and AQuA share a significant similarity of 0.43, suggesting they suffer from common failure modes. Conversely, the overlap drops precipitously when comparing foundational tasks with advanced Olympiad-level challenges (e.g., GSM8K vs. AIME25 yields nearly no overlap). This sharp separation confirms that error patterns are highly task-dependent. It further validates that our constructed pool is diverse enough to cover a wide spectrum of failure modes, and that the retrieval mechanism effectively isolates the specific, context-aware constraints required for each unique domain. 6 Related Work As MAS scales to handle complex tasks, they become increasingly vulnerable to error propagation, where individual mistakes amplify downstream and disrupt the entire reasoning process. To address this, prior research focuse on three resilience strategies: (1) robust architecture design, (2) error monitoring, and (3) utilization of inference trajectories. Robust MAS Architectures Existing research attempts to mitigate the propagation of erroneous and redundant information by engineering more robust system structures. Several studies explicitly model MAS as optimizable graphs or topologies, employing learning or search algorithms to identify superior workflow structures (Zhuge et al., 2024; Zhang et al., 2025d; Wang et al., 2025b; Zhang et al., 2025a). Adopting sparse communication topologies has also proven effective in reducing noise disturbance (Li et al., 2024b). Furthermore, introducing advanced initialization, orchestration or routing strategies to construct cooperative teams with specialized roles can further suppress the spread of errors originating from underperforming agents (Tian et al., 2025; Dang et al., 2025; Zhang et al., 2025g; Wang et al., 2026; Ong et al., 2025). Error Monitoring Mechanisms These methods focus on designing or training monitors to detect anomalies within the MAS workflow, thereby enabling information correction to prevent error cascading. Graph-based approaches treat information flow and topology as signals, utilizing anomaly detectors to capture abnormal patterns and identify system errors (Wang et al., 2025a; Zhou et al., 2025; Pan et al., 2025a). Test-time rectification serves as an efficient intervention strategy, implementing an “intercept-detect-correct” process for each action or message within the system (Xiang et al., 2024; Chen et al., 2025b; Luo et al., 2025). Conversely, error attribution and tracking methods aim to perform root cause analysis, identifying the specific agents responsible for introducing hallucinatory or incorrect information upon task failure (Zhang et al., 2025f; Pan et al., 2025b; Zhang et al., 2025b; Ge et al., 2025). Utilization of Inference Trajectories These approaches enhance MAS reliability by leveraging real execution trajectories to construct preference or contrastive data for training key components (e.g., reasoners or planners), thereby improving reasoning accuracy (Chen et al., 2025a; Motwani et al., 2025; Zhao et al., 2025). Process-aware variants further verify intermediate steps to provide fine-grained supervision, preventing models from falling into locally plausible but globally incorrect reasoning paths (Zelikman et al., 2022; Lightman et al., 2024). Additionally, some works mine exploration or failure trajectories as hard negatives to strengthen preference optimization, rendering the system more robust against misleading intermediate states (Song et al., 2024; Aksitov et al., 2024; Lyu et al., 2025). Our framework integrates these paradigms to overcome their limitations. Unlike rigid structural designs, our approach serves as a model-agnostic, plug-and-play module adaptable to diverse frameworks. We advance error monitoring from passive detection to active rectification, ensuring real-time stability via feedback-driven reflection. Finally, leveraging trajectory utilization, we distill historical failures into an adversarial indicator pool, providing precise, prior-guided online supervision. Conclusion In this paper, we introduced AgentDropoutV2, a novel framework designed to optimize information flow in MAS via test-time rectify-or-reject pruning. By mining historical failure trajectories, we constructed an indicator pool that encapsulates domain-specific error patterns. During test-time inference, our framework actively intercepts agent outputs, retrieves pertinent indicators, and enforces an iterative refinement process to resolve latent errors before they propagate. Experimental results demonstrate that this mechanism effectively cleanses the information flow, thereby significantly enhancing system accuracy. Furthermore, our analysis confirms that the indicator retrieval and rectification processes exhibit strong adaptivity to varying task difficulties, along with robust transferability across different domains and backbone models. Impact Statement This paper presents work aiming to enhance the reliability and accuracy of Multi-Agent Systems through test-time error rectification and pruning. By actively identifying and intercepting erroneous reasoning and hallucinations before they propagate, our framework contributes to the development of more robust and trustworthy automated systems, particularly in domains requiring rigorous logic, such as mathematics and software development. While the construction of adversarial indicators relies on historical data, potentially reflecting existing data distributions, the methodology itself serves to enforce constraints and improve adherence to ground truth. We do not foresee specific negative societal consequences or ethical concerns beyond those generally associated with the development and deployment of large language models. References R. Aksitov, S. Miryoosefi, Z. Li, D. Li, S. Babayan, K. Kopparapu, Z. Fisher, R. Guo, S. Prakash, P. Srinivasan, M. Zaheer, F. Yu, and S. Kumar (2024) ReST meets react: self-improvement for multi-step reasoning LLM agent. In ICLR 2024 Workshop on Large Language Model (LLM) Agents, External Links: Link Cited by: §6. J. Austin, A. Odena, M. Nye, M. Bosma, H. Michalewski, D. Dohan, E. Jiang, C. Cai, M. Terry, Q. Le, et al. (2021) Program synthesis with large language models. arXiv preprint arXiv:2108.07732. External Links: Link Cited by: §4.1. M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al. (2021) Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374. External Links: Link Cited by: §4.1. W. Chen, J. Yuan, C. Qian, C. Yang, Z. Liu, and M. Sun (2025a) Optima: optimizing effectiveness and efficiency for LLM-based multi-agent system. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 11534–11557. External Links: Document, ISBN 979-8-89176-256-5, Link Cited by: §6. Z. Chen, M. Kang, and B. Li (2025b) ShieldAgent: shielding agents via verifiable safety policy reasoning. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §6. K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al. (2021) Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. External Links: Link Cited by: §4.1. Y. Dang, C. Qian, X. Luo, J. Fan, Z. Xie, R. Shi, W. Chen, C. Yang, X. Che, Y. Tian, X. Xiong, L. Han, Z. Liu, and M. Sun (2025) Multi-agent collaboration via evolving orchestration. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §1, §6. Y. Ge, L. Xie, Z. Li, Y. Pei, and T. Zhang (2025) Who is introducing the failure? automatically attributing failures of multi-agent systems via spectrum analysis. arXiv preprint arXiv:2509.13782. External Links: Link Cited by: §6. A. Ghafarollahi and M. J. Buehler (2025) SciAgents: automating scientific discovery through bioinspired multi-agent intelligent graph reasoning. Advanced Materials 37 (22), p. 2413523. External Links: Link Cited by: §1. A. E. Ghareeb, B. Chang, L. Mitchener, A. Yiu, C. J. Szostkiewicz, J. M. Laurent, M. T. Razzak, A. D. White, M. M. Hinks, and S. G. Rodriques (2025) Robin: a multi-agent system for automating scientific discovery. arXiv preprint arXiv:2505.13400. External Links: Link Cited by: §1. T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang (2024) Large language model based multi-agents: A survey of progress and challenges. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI 2024, Jeju, South Korea, August 3-9, 2024, p. 8048–8057. External Links: Link Cited by: §1. C. He, R. Luo, Y. Bai, S. Hu, Z. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, et al. (2024) OlympiadBench: a challenging benchmark for promoting AGI with olympiad-level bilingual multimodal scientific problems. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 3828–3850. External Links: Document, Link Cited by: §4.1. S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, et al. (2024) MetaGPT: meta programming for A multi-agent collaborative framework. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §1. N. Jain, K. Han, A. Gu, W. Li, F. Yan, T. Zhang, S. Wang, A. Solar-Lezama, K. Sen, and I. Stoica (2025) LiveCodeBench: holistic and contamination free evaluation of large language models for code. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §4.1. G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem (2023) CAMEL: communicative agents for ”mind” exploration of large language model society. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), External Links: Link Cited by: §1. S. Li, Y. He, H. Guo, X. Bu, G. Bai, J. Liu, J. Liu, X. Qu, Y. Li, W. Ouyang, et al. (2024a) GraphReader: building graph-based agent to enhance long-context abilities of large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 12758–12786. External Links: Document, Link Cited by: §1. Y. Li, D. Choi, J. Chung, N. Kushman, J. Schrittwieser, R. Leblond, T. Eccles, J. Keeling, F. Gimeno, A. D. Lago, et al. (2022) Competition-level code generation with alphacode. Science 378 (6624), p. 1092–1097. External Links: Document, https://w.science.org/doi/pdf/10.1126/science.abq1158, Link Cited by: §4.1. Y. Li, Y. Du, J. Zhang, L. Hou, P. Grabowski, Y. Li, and E. Ie (2024b) Improving multi-agent debate with sparse communication topology. In Findings of the Association for Computational Linguistics: EMNLP 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 7281–7294. External Links: Document, Link Cited by: §6. H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe (2024) Let’s verify step by step. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §1, §4.1, §6. W. Luo, S. Dai, X. Liu, S. Banerjee, H. Sun, M. Chen, and C. Xiao (2025) AGrail: a lifelong agent guardrail with effective and adaptive safety detection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 8104–8139. External Links: Document, ISBN 979-8-89176-251-0, Link Cited by: §6. Y. Lyu, L. Yan, Z. Wang, D. Yin, P. Ren, M. de Rijke, and Z. Ren (2025) MACPO: weak-to-strong alignment via multi-agent contrastive preference optimization. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §6. S. R. Motwani, C. Smith, R. J. Das, R. Rafailov, I. Laptev, P. Torr, F. Pizzati, R. Clark, and C. S. de Witt (2025) MALT: improving reasoning with multi-agent LLM training. In Workshop on Reasoning and Planning for Large Language Models, External Links: Link Cited by: §1, §6. I. Ong, A. Almahairi, V. Wu, W. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica (2025) RouteLLM: learning to route LLMs from preference data. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §6. J. Pan, Y. Liu, R. Miao, K. Ding, Y. Zheng, Q. V. H. Nguyen, A. W. Liew, and S. Pan (2025a) Explainable and fine-grained safeguarding of llm multi-agent systems via bi-level graph anomaly detection. arXiv preprint arXiv:2512.18733. External Links: Link Cited by: §6. M. Z. Pan, M. Cemri, L. A. Agrawal, S. Yang, B. Chopra, R. Tiwari, K. Keutzer, A. Parameswaran, K. Ramchandran, D. Klein, J. E. Gonzalez, M. Zaharia, and I. Stoica (2025b) Why do multiagent systems fail?. In ICLR 2025 Workshop on Building Trust in Language Models and Applications, External Links: Link Cited by: §1, §6. J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein (2023) Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, p. 1–22. External Links: Link Cited by: §1. A. Patel, S. Bhattamishra, and N. Goyal (2021) Are NLP models really able to solve simple math word problems?. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.), Online, p. 2080–2094. External Links: Document, Link Cited by: §4.1. A. Prasad, A. Koller, M. Hartmann, P. Clark, A. Sabharwal, M. Bansal, and T. Khot (2024) ADaPT: as-needed decomposition and planning with language models. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, p. 4226–4252. External Links: Document, Link Cited by: §1. C. Qian, W. Liu, H. Liu, N. Chen, Y. Dang, J. Li, C. Yang, W. Chen, Y. Su, X. Cong, et al. (2024) ChatDev: communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 15174–15186. External Links: Document, Link Cited by: §1. Y. Song, D. Yin, X. Yue, J. Huang, S. Li, and B. Y. Lin (2024) Trial and error: exploration-based trajectory optimization of LLM agents. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 7584–7600. External Links: Document, Link Cited by: §6. H. Sun, Y. Min, Z. Chen, W. X. Zhao, Z. Liu, Z. Wang, L. Fang, and J. Wen (2025) Challenging the boundaries of reasoning: an olympiad-level math benchmark for large language models. arXiv preprint arXiv:2503.21380. External Links: Link Cited by: §4.1. C. Tian, Y. Wang, X. Liu, Z. Wang, L. Ding, M. Zhang, and M. Zhang (2025) AgentInit: initializing LLM-based multi-agent systems via diversity and expertise orchestration for effective and efficient collaboration. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 11870–11902. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §6. J. Wang, S. Zhao, J. Liu, H. Wang, W. Li, B. Qin, and T. Liu (2026) Orchestrating intelligence: confidence-aware routing for efficient multi-agent collaboration across multi-scale models. arXiv preprint arXiv:2601.04861. External Links: Link Cited by: §6. S. Wang, G. Zhang, M. Yu, G. Wan, F. Meng, C. Guo, K. Wang, and Y. Wang (2025a) G-safeguard: a topology-guided security lens and treatment on LLM-based multi-agent systems. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 7261–7276. External Links: Document, ISBN 979-8-89176-251-0, Link Cited by: §1, §6. Z. Wang, Y. Wang, X. Liu, L. Ding, M. Zhang, J. Liu, and M. Zhang (2025b) AgentDropout: dynamic agent elimination for token-efficient and high-performance LLM-based multi-agent collaboration. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 24013–24035. External Links: Document, ISBN 979-8-89176-251-0, Link Cited by: §1, §6. Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al. (2024) Autogen: enabling next-gen llm applications via multi-agent conversations. In First Conference on Language Modeling, External Links: Link Cited by: §1, §4.1. Z. Xiang, L. Zheng, Y. Li, J. Hong, Q. Li, H. Xie, J. Zhang, Z. Xiong, C. Xie, C. Yang, et al. (2024) Guardagent: safeguard llm agents by a guard agent via knowledge-enabled reasoning. arXiv preprint arXiv:2406.09187. External Links: Link Cited by: §6. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao (2023) ReAct: synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023, External Links: Link Cited by: §1. E. Zelikman, Y. Wu, J. Mu, and N. D. Goodman (2022) STaR: bootstrapping reasoning with reasoning. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), External Links: Link Cited by: §6. G. Zhang, K. Chen, G. Wan, H. Chang, H. Cheng, K. Wang, S. Hu, and L. Bai (2025a) Evoflow: evolving diverse agentic workflows on the fly. arXiv preprint arXiv:2502.07373. External Links: Link Cited by: §6. G. Zhang, J. Wang, J. Chen, W. Zhou, K. Wang, and S. Yan (2025b) AgenTracer: who is inducing failure in the llm agentic systems?. arXiv preprint arXiv:2509.03312. External Links: Link Cited by: §1, §6. G. Zhang, Y. Yue, Z. Li, S. Yun, G. Wan, K. Wang, D. Cheng, J. X. Yu, and T. Chen (2025c) Cut the crap: an economical communication pipeline for llm-based multi-agent systems. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, External Links: Link Cited by: §1. G. Zhang, Y. Yue, X. Sun, G. Wan, M. Yu, J. Fang, K. Wang, T. Chen, and D. Cheng (2025d) G-designer: architecting multi-agent communication topologies via graph neural networks. In ICLR 2025 Workshop on Foundation Models in the Wild, External Links: Link Cited by: §1, §6. R. Zhang, X. Zhao, R. Wang, S. Chen, G. Zhang, A. Zhang, K. Wang, and Q. Wen (2025e) SafeSieve: from heuristics to experience in progressive pruning for llm-based multi-agent communication. arXiv preprint arXiv:2508.11733. External Links: Link Cited by: §1. S. Zhang, M. Yin, J. Zhang, J. Liu, Z. Han, J. Zhang, B. Li, C. Wang, H. Wang, Y. Chen, and Q. Wu (2025f) Which agent causes task failures and when? on automated failure attribution of LLM multi-agent systems. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §1, §6. W. Zhang, C. Cui, Y. Zhao, Y. Liu, and B. An (2025g) AgentOrchestra: a hierarchical multi-agent framework for general-purpose task solving. arXiv preprint arXiv:2506.12508. External Links: Link Cited by: §6. Y. Zhang and T. Math-AI (2024) American invitational mathematics examination (aime) 2024. External Links: Link Cited by: §4.1. Y. Zhang and T. Math-AI (2025) American invitational mathematics examination (aime) 2025. External Links: Link Cited by: §4.1. J. Zhao, C. Zu, X. Hao, Y. Lu, W. He, Y. Ding, T. Gui, Q. Zhang, and X. Huang (2024) LONGAGENT: achieving question answering for 128k-token-long documents through multi-agent collaboration. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 16310–16324. External Links: Link, Document Cited by: §1. W. Zhao, M. Yuksekgonul, S. Wu, and J. Zou (2025) SiriuS: self-improving multi-agent systems via bootstrapped reasoning. In Workshop on Reasoning and Planning for Large Language Models, External Links: Link Cited by: §1, §6. J. Zhou, L. Wang, and X. Yang (2025) GUARDIAN: safeguarding LLM multi-agent collaborations with temporal graph modeling. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §6. M. Zhuge, W. Wang, L. Kirsch, F. Faccio, D. Khizbullin, and J. Schmidhuber (2024) GPTSwarm: language agents as optimizable graphs. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, External Links: Link Cited by: §1, §6. Appendix A Appendix A.1 Pseudo Codes Algorithm 1 outlines the pseudo-code for our rectify-or-reject pruning. During MAS execution, the output of each agent is actively intercepted to undergo the rectification process. First, an active indicator set is retrieved based on semantic similarity to serve as a reference for potential error patterns (Lines 6-8). A rectifier model then scrutinizes the output against each retrieved indicator, generating diagnostic rationales as feedback whenever a specific constraint is violated (Lines 10-14). Subsequently, the algorithm employs a tri-state gating mechanism based on the evaluation results, terminating the iteration if the output passes all checks or if the iteration budget is exhausted (Lines 15-24). Upon successful verification, the qualified output is propagated to successor agents (Lines 25-27). Finally, if the resulting information flow becomes critically sparse, a global fallback process is triggered to reset the system and re-initialize execution from scratch (Lines 28-31). The pseudo-code for the Failure-Driven Indicator Pool Construction is outlined in Algorithm 2. The process begins by iterating through the source dataset to collect execution trajectories where the MAS fails to deliver the correct solution (Lines 3-4). Subsequently, a teacher model scrutinizes these failure instances, synthesizing candidate indicators that capture the specific error patterns exhibited by individual agents (Lines 5-6). To prevent repository bloating, a redundancy elimination mechanism is applied to each candidate. The algorithm first encodes the new indicator into a semantic vector to retrieve the most similar existing constraints (Lines 7-9). A deduplication model then verifies the novelty of the candidate, admitting it into the global pool only if it represents a distinct and previously unrecorded error type (Lines 10-12). A.2 Indicator & Prompt Design Indicator Design Figure 5 displays an example from our constructed indicator pool. This specific indicator is tailored to verify the precision of square root calculations (a detailed application case is provided in Appendix A.4). For scenarios where a pre-defined indicator pool is unavailable, we design general-purpose math and code indicators, as illustrated in Figure 6 and Figure 7, respectively. Prompt Design The prompt templates for the rectifier in the math and code domains are presented in Figure 8 and Figure 9. Additionally, Figure 10 depicts the prompt template for the teacher model, which is responsible for generating new indicators based on failed MAS execution trajectories. A.3 Dataset Statistics Table 5 lists the detailed statistics of the size of the datasets and the constructed indicator pool. The indicator pool for the math domain is constructed on the failed MAS trajectories on the sampled instances from the MATH and AQuA training sets. No indicator pool is built for the code domain. A.4 Case Study This case study exemplifies the framework’s capability to navigate complex constraint satisfaction problems through a rectify-or-reject dialectical process. The agent was tasked with determining the number of real values for x such that 120−x 120- x results in an integer (Figure 11). The resolution trajectory began with a common cognitive deficit where the agent implicitly conflated the set of integers (ℤZ) with positive integers (ℤ+Z^+), positing that the expression equaled n for n∈1,…,10n∈\1,…,10\ (Figure 12). This under-inclusion error was immediately intercepted by the rectifier via the INTEGER_CONDITION_MISMANAGEMENT indicator, which explicitly challenged the agent’s assumption by instructing it to re-evaluate the valid range to include zero (Figure 13). Responding to this guidance, the agent rectified the omission, but another error occurred. In an effort to strictly adhere to the instruction that “integers include negatives”, the agent expanded the domain to include negative values (e.g., n∈−10,…,10n∈\-10,…,10\), thereby neglecting the intrinsic non-negativity property of the principal square root function (Figure 14). This error represents a classic instance of contextual detachment, where satisfying one constraint leads to the violation of another. The rectifier subsequently intervened with a SQUARE_ROOT_MANIPULATION_CHECK, providing a critical boundary correction that the output of a square root must remain non-negative (Figure 15). Ultimately, a logical synthesis was achieved in the final iteration. By integrating the “integer nature” constraint from the first round of feedback with the “non-negativity constraint” from the second, the agent correctly defined the valid range as the intersection of integers and non-negative values (n∈0,1,…,10n∈\0,1,…,10\) (Figure 16). The system converged on the correct count of 11 real values, achieving precise alignment with the ground truth, therefore the passing all indicator check by the rectifier and received no more feedback (Figure 17). This trajectory demonstrates the robustness of the rectify-or-reject mechanism in stabilizing reasoning through iterative, multi-dimensional constraint enforcement. 1 Input : Active agent set A, Indicator Pool ℐI, Rectifier Model Φrect _rect, Embedding Model MembM_emb Parameters : Max iterations TmaxT_max, Top-K retrieval KactK_act, Safety threshold γ Output : Final answer Y 2 // Phase 1: Agent Execution & Rectification ←∅O← ; // Initialize valid output set 3 foreach Agent Ai∈A_i do oi(0)←Φi(xi,ℛi,i)o_i^(0)← _i(x_i,R_i,K_i) ; // Initial Generation 4 t←1t← 1; 5 6 while t≤Tmaxt≤ T_max do // Step 1: Relevant Indicator Retrieval scen(t),act(t)←Φrect(oi(t))S_scen^(t),S_act^(t)← _rect(o_i^(t)) ; // Extract keywords i(t)←Memb(scen(t)⊕act(t))q_i^(t)← M_emb(S_scen^(t) _act^(t)) ; // Compute query ℐact(t)←Top-Kact(i(t),ℐ)I_act^(t) -K_act(q_i^(t),I) ; // Retrieve active indicator set 7 // Step 2: Verification 8 E(t)←0,ℱ(t)←∅E^(t)← 0, ^(t)← ; 9 foreach Indicator Ik∈ℐact(t)I_k _act^(t) do 10 (vk(t),rk(t))←Φrect(oi(t)∣xi,ℛi,Ik)(v_k^(t),r_k^(t))← _rect(o_i^(t) x_i,R_i,I_k); 11 if vk(t)=1v_k^(t)=1 then 12 E(t)←1E^(t)← 1; 13 ℱ(t)←ℱ(t)∪rk(t)F^(t) ^(t)∪\r_k^(t)\; 14 15 16 // Step 3: Tri-State Gating Decision 17 if E(t)=0E^(t)=0 then 18 oi←oi(t)o_i← o_i^(t); 19 ←∪oiO ∪\o_i\; break ; // Pass: Accept output 20 21 else if t<Tmaxt<T_max then oi(t+1)←Φi(xi,ℛi,i,ℱ(t))o_i^(t+1)← _i(x_i,R_i,K_i,F^(t)) ; // Retry: Regenerate 22 t←t+1t← t+1; 23 24 else oi←∅o_i← ; // Reject: Discard output 25 break; 26 27 // Propagate Output to Successors 28 if oi≠∅o_i≠ then 29 foreach Agent Aj∈(Ai)A_j (A_i) do 30 j←ℛi,oiK_j←\R_i,o_i\ 31 32 33 // Phase 2: Global Fallback Check 34 Nvalid←|o∈∣o≠∅|N_valid←|\o o≠ \|; 35 if Nvalid<γN_valid<γ then 36 Trigger System-Wide Reset; 37 Discard O and re-initialize with fresh agents; 38 39return oNo_N; 40 Algorithm 1 Test-Time rectify-or-reject Pruning for MAS Information Flow Optimization 1 Input : Source Dataset src=,∗D_src=\Q,Y^*\, Teacher Model Φteach _teach, Deduplication Model Φdedup _dedup, Embedding Model MembM_emb Parameters : Retrieval size KdedupK_dedup Output : Optimized Indicator Pool ℐI 2 // Initialize empty indicator pool 3 ℐ←∅I← ; 4 5foreach Instance (,∗)∈src(Q,Y^*) _src do // Step 1: Failure Trajectory Collection 6 Execute MAS to obtain trajectory: =(,A1:N,o1:N,)T=(Q,A_1:N,o_1:N,Y); 7 8 if ≠∗Y ^* then // Step 2: Offline Indicator Mining 9 foreach Agent AiA_i in T do ℐnew←Φteach(,∗,ℛi,oi)I_new← _teach (T,Y^*,R_i,o_i ) ; // Generate candidate indicators 10 // Step 3: Redundancy Elimination 11 foreach Indicator Inew=(nnew,dnew,cnew)∈ℐnewI_new=(n_new,d_new,c_new) _new do new←Memb(dnew⊕cnew)v_new← M_emb(d_new c_new) ; // Compute semantic vector 12 ℐsim←Top-Kdedup(new,ℐ)I_sim -K_dedup(v_new,I) ; // Retrieve top-KdedupK_dedup similar set 13 14 IsNovel←Φdedup(Inew,ℐsim)IsNovel← _dedup (I_new,I_sim ); 15 16 if IsNovelIsNovel is True then ℐ←ℐ∪InewI ∪\I_new\ ; // Add novel error pattern 17 18 19 20 21 22return ℐI; 23 Algorithm 2 Failure-Driven Indicator Pool Construction An Indicator Example Name: SQUARE_ROOT_MANIPULATION_CHECK: Detailed Definition: This error occurs when mathematical expressions in radical form are manipulated in a way that leads to logical contradictions relating to divisibility or integer properties. Trigger Condition: When the agent performs manipulations involving expressions in radical form and calculations involving divisibility and integer properties. Figure 5: An example of the indicators from the constructed pool for the math domain. General Indicator for Math Name: CRITICAL_MATH_LOGIC_AUDIT Detailed Definition: A focused audit to detect substantive logical fallacies, calculation errors, or conditional oversights that invalidate the final result. Trigger Condition: The Agent is performing mathematical reasoning, derivation, or calculation. Figure 6: The design of the general indicator for the math domain. General Indicator for Code Name: CRITICAL_CODE_CORRECTNESS_CHECK Detailed Definition: A functional audit focusing on runtime safety, logical integrity, and adherence to requirements in code implementation. Trigger Condition: The Agent is generating, debugging, or analyzing computer code. Figure 7: The design of the general indicator for code domain. Prompt for Math Rectifier You are an Objective Logic Auditor. Your task is to verify if a specific team member (**Agent Role**) has committed a **FATAL LOGIC ERROR** regarding a specific **Area of Concern**. ### The ”Impact & Action” Protocol 1. **Presumption of Validity**: You must assume the Agent’s reasoning is correct unless you find irrefutable evidence of a fatal flaw. 2. **The ”Actionability” Test**: If you cannot provide a specific, mathematical correction (a formula, a step, or a value), **IT IS NOT A FLAW**. 3. **The ”Impact” Test**: If the Agent’s phrasing is imperfect but the **FINAL ANSWER** remains mathematically correct, **IT IS NOT A FLAW**. ### Judgment Criteria **[Area of Concern]**: trigger_condition — ### CONTEXT - **Task**: task - **Agent Role**: role - **Agent Output**: agent_output — ### OUTPUT FORMAT (JSON ONLY) You must generate the fields in this **EXACT ORDER**. The logical flow determines the verdict. ”evidence_quote”: ”Verbatim quote of the problematic part. Write ’N/A’ if valid.”, ”analysis”: ”Explain WHY this specific part violates the Area of Concern. Focus on logic, not style. Try to express in a concise and to the point manner, avoid lengthy speeches. Write ’N/A’ if valid.”, ”suggestion”: ”Concrete instruction on how to fix it (e.g., ’Change x to y’, ’Apply formula Z’). If no fix is needed or possible, write ’N/A’.”, ”impact_assessment”: ”Simulate the correction. Does the FINAL ANSWER or core conclusion change? (YES/NO) and brief reason.”, ”is_flawed”: boolean // Set to true ONLY if ’suggestion’ is concrete AND ’impact_assessment’ is YES. Otherwise false. Figure 8: The prompt template for math rectifiers. Prompt for Code Rectifier You are a Senior Code Auditor and Architect. Your task is to verify if a specific team member (**Agent Role**) has committed a **FATAL CODING ERROR** regarding a specific **Area of Concern**. ### The ”Impact & Action” Protocol 1. **Presumption of Validity**: You must assume the Agent’s code is functionally correct unless you find irrefutable evidence of a fatal flaw (syntax error, logic bug, or interface violation). 2. **The ”Actionability” Test**: If you cannot provide a specific code correction (a line change, a logic fix, or a parameter adjustment), **IT IS NOT A FLAW**. 3. **The ”Impact” Test**: If the code is inefficient, verbose, or stylistically non-standard but **EXECUTES CORRECTLY** and returns the right result, **IT IS NOT A FLAW**. ### Judgment Criteria **[Area of Concern]**: trigger_condition — ### CONTEXT - **Task**: task - **Agent Role**: role - **Agent Output**: agent_output — ### OUTPUT FORMAT (JSON ONLY) You must generate the fields in this **EXACT ORDER**. The logical flow determines the verdict. ”evidence_quote”: ”Verbatim quote of the problematic code snippet. Write ’N/A’ if valid.”, ”analysis”: ”Explain WHY this specific part violates the Area of Concern. Focus on functional correctness (bugs/crashes), not style (PEP8/comments).Try to express in a concise and to the point manner, avoid lengthy speeches. Write ’N/A’ if valid.”, ”suggestion”: ”Concrete instruction on how to fix the code (e.g., ’Change index i to i+1’, ’Import module X’). If no fix is needed, write ’N/A’.”, ”impact_assessment”: ”Simulate the correction. Does it fix a runtime error, infinite loop, or incorrect output? (YES/NO) and brief reason.”, ”is_flawed”: boolean // Set to true ONLY if ’suggestion’ is concrete AND ’impact_assessment’ is YES. Otherwise false. Figure 9: The prompt template for code rectifiers. Prompt for Teacher You are an AI acting as a **Lead Mathematics Auditor and Logic Specialist**, specifically optimized for the MATH dataset (high-difficulty competitions like AMC, AIME). ### Background & Goal **Background**: An agent team has attempted to solve a complex math problem, and **the team’s final answer is INCORRECT**. **Goal**: Synthesize the known problem (‘problem‘), standard solution (‘solution‘), and the Agent’s ‘output‘ to strictly evaluate the Agent’s reasoning process. **IMPORTANT CONTEXT**: The provided ‘solution‘ is a standard, single-path reference answer. However, the Agent is part of a Multi-Agent System (MAS). - Its ‘output‘ depends on its ‘agent_role‘ (e.g., a ”Python Coder” writes code, a ”Critic” critiques). - **DO NOT** penalize the Agent simply because its output does not look like the standard ‘solution‘ (e.g., using code instead of pure derivation is valid if the role permits). - Only penalize **logical errors**, **calculation errors**, or **hallucinations** that contradict mathematical truths. ### MATH Input Context 1. **‘problem‘**: problem 2. **‘solution‘**: (Ground Truth) solution 3. **‘agent_role‘**: agent_role 4. **‘output‘**: (Agent’s Attempt) output ### Phase 1: Diagnosis Please execute the following logical judgment: 1. Assess whether the Agent’s output is logically and mathematically correct **within the scope of its role**. 2. **AUDIT STRATEGY (CRITICAL)**: - **DO NOT STOP at the first error.** You must scan the ENTIRE output line by line. - Independent errors often exist (e.g., a logical fallacy in Step 1 AND a formatting error in the Final Answer). - You are expected to find **MULTIPLE distinct errors** (less than 5) if they exist. 3. **Decision**: - If the output contains **NO errors**: Output ‘NO_ERROR‘. - If the output contains **errors**: Identify **ALL** of them and proceed to Phase 2. ### Phase 2: Metric Extraction Transform **EACH identified error** separately into a **generalized** JSON metric object. **CRITICAL**: The ‘name‘, ‘detailed_definition‘, ‘trigger_condition‘, and ‘example_error‘ must be **generalizable** to other similar math problems. 1. **‘name‘**: * **Requirement**: Summarize the error pattern. It can be **appropriately longer** to avoid ID collisions. * **Format**: ‘UPPER_CASE_WITH_UNDERSCORES‘. 2. **‘domain_tag‘**: * **Requirement**: Classify this error into a specific mathematical or operational domain. * **Examples**: ”Geometry”, ”Probability”, ”Algebra”, ”Number Theory”, ”Python Implementation”, ”Logical Reasoning”. 3. **‘detailed_definition‘**: * **Requirement**: Define the **ROOT CAUSE** or **Mental Misconception** behind the error. Do not just say ”used wrong formula”; explain ”confused concept A with concept B”. * **Format**: ”This error occurs when the agent [misconception], leading to [consequence].” 4. **‘evaluator_prompt‘**: Contains the trigger condition for retrieving this metric: * **‘trigger_condition‘**: * **Requirement**: Describe the **Context** or **Action** where this error is likely to happen. **DO NOT** assume the error has already occurred (Decriminalized). * **Format**: ”When the problem involves [context]…” OR ”When the agent attempts to [action]…” 5. **‘example_error‘**: * **Requirement**: Provide a concrete example of the error AND the logic for why it is wrong/how to fix it. * **Format**: ”Error Snippet: [Quote agent’s wrong step] — Correction Logic: [Explain why it is wrong and what the correct approach/formula should be].” ### Output Format - If no error: Output ‘NO_ERROR‘ only. - If errors exist: **ALWAYS Output a JSON LIST** containing one or more metric objects. - Structure: ‘[ ”name”: ”ERROR_1”, … , ”name”: ”ERROR_2”, … ]‘ - **CRITICAL JSON SYNTAX RULE**: - When writing LaTeX inside JSON strings, **YOU MUST DOUBLE-ESCAPE BACKSLASHES**. - **WRONG**: ‘”equation”: ” frac12”‘ (This causes JSON parse error!) - **CORRECT**: ‘”equation”: ” frac12”‘ (This works!) Figure 10: The prompt template for the teacher model during indicator pool construction. Table 5: Dataset statistics Domain Dataset Size Test Set Math GSM8K 1,319 MATH-500 500 AQuA 254 ACM23 40 OlympiadBench 675 OlymMATH Easy 100 OlymMATH Hard 100 AIME24 30 AIME24 30 Code MBPP 257 HumanEval 161 CodeContests 165 LiveCodeBenchV1 400 Training Set Math MATH 2,000 AQuA 2,000 Indicator Pool Math ℐI 2,000 Math Task Q and Correct Answer ∗Y^* Task: For how many real values of x is 120−x 120- x an integer? Answer: 11 Figure 11: An example of the given math task. Math Solver: Initial Output o(0)o^(0) We are given the expression: 120−x 120- x and we are asked to find how many **real values of x** make this expression an **integer**. — ### Step 1: Let the expression be an integer Let: 120−x=n,where n∈ℤ+ 120- x=n, n ^+ Squaring both sides: 120−x=n2120- x=n^2 Solving for x x: x=120−n2 x=120-n^2 Now square both sides again to solve for x: x=(120−n2)2x=(120-n^2)^2 — ### Step 2: Determine the possible values of n We must ensure that x=120−n2≥0 x=120-n^2≥ 0, because square roots of real numbers are non-negative. So: 120−n2≥0⇒n2≤120⇒n≤120120-n^2≥ 0 n^2≤ 120 n≤ 120 Approximating: 120≈10.95 120≈ 10.95 So the largest integer n can be is 10. Since n∈ℤ+n ^+, the possible values of n are: n=1,2,3,…,10n=1,2,3,…,10 That gives us **10 values** of n. — ### Step 3: Verify that all these values produce real x For each n∈1,2,…,10n∈\1,2,…,10\, we compute: x=(120−n2)2x=(120-n^2)^2 This will always be a real number, since it is a square of a real number. — ### Step 4: Conclusion Each integer n from 1 to 10 gives a unique real value of x, and all such values make the expression 120−x 120- x an integer. — ### Final Answer: 10 10 Figure 12: The initial output of the math solver agent. Rectifier: Rectification Round 0/2 Evaluation [1] Matched Indicators: - MODULAR_ARITHMETIC_CONTEXT_CHECK - MISSING_CHECK_FOR_INTEGER_SOLUTIONS - LACK_OF_COMPREHENSIVE_OPTION_ELIMINATION - INCORRECT_EXPRESSION_FOR_DISTANCE - INTEGER_CONDITION_MISMANAGEMENT [2] Evaluations: - MODULAR_ARITHMETIC_CONTEXT_CHECK: Flawed - The Agent assumes that n must be a positive integer, but the problem asks for real values of x. The expression 120−x 120- x being an integer does not require n to be strictly positive; it could also be zero. This oversight restricts the possible values of n and may lead to an incomplete solution. - MISSING_CHECK_FOR_INTEGER_SOLUTIONS: Correct - LACK_OF_COMPREHENSIVE_OPTION_ELIMINATION: Flawed - The Agent assumes n must be a positive integer, but the problem asks for real values of x. The square root function can yield non-integer values for n as long as the expression remains real. The constraint that n must be a positive integer is not justified by the problem statement, which only requires the expression to be an integer, not necessarily that n must be a positive integer. - INCORRECT_EXPRESSION_FOR_DISTANCE: Correct - INTEGER_CONDITION_MISMANAGEMENT: Flawed - The assumption that n must be a positive integer (n∈ℤ+)(n ^+) excludes the possibility of n = 0, which is a valid integer and could yield a real value of x. The Area of Concern requires checking all integer solutions, including zero, to ensure all constraints are satisfied across calculations. [3] Feedback: An external auditor has reviewed your previous output (Attempt 1) and flagged some potential issues. Please review the following suggestions critically: - [MODULAR_ARITHMETIC_CONTEXT_CHECK]: Allow n to be zero as well, since 120−x 120- x can be zero and still be an integer. This would expand the range of n to include 0. (Auditor’s Note: The Agent assumes that n must be a positive integer, but the problem asks for real values of x. The expression 120−x 120- x being an integer does not require n to be strictly positive; it could also be zero. This oversight restricts the possible values of n and may lead to an incomplete solution.) - [LACK_OF_COMPREHENSIVE_OPTION_ELIMINATION]: Allow n to be any integer (positive, negative, or zero) such that 120−n2≥0120-n^2≥ 0, and re-evaluate the valid range of n. (Auditor’s Note: The Agent assumes n must be a positive integer, but the problem asks for real values of x. The square root function can yield non-integer values for n as long as the expression remains real. The constraint that n must be a positive integer is not justified by the problem statement, which only requires the expression to be an integer, not necessarily that n must be a positive integer.) - [INTEGER_CONDITION_MISMANAGEMENT]: Allow n to be in ℤZ (including 0) and re-evaluate the range of n based on the constraint 120−n2≥0120-n^2≥ 0. (Auditor’s Note: The assumption that n must be a positive integer ((n∈ℤ+)(n ^+)) excludes the possibility of n = 0, which is a valid integer and could yield a real value of x. The Area of Concern requires checking all integer solutions, including zero, to ensure all constraints are satisfied across calculations.) **Instruction**: 1. If you agree with the advice, please refine your solution. 2. **If you are confident your original logic is correct, you may ignore this advice.** 3. Please output the corrected solution. Figure 13: The rectifier’s judgments and feedback to the initial output. Math Solver: Rectification Round 1/2 Output o(1)o^(1) You are absolutely correct to be critical of the original solution. The key point is that the expression 120−x 120- x being an integer does **not** require the integer to be strictly positive. In fact, it can be **zero**, which is also an integer. Therefore, we should allow n∈ℤn , not just n∈ℤ+n ^+. — ### Revised Step-by-Step Solution: We are given: 120−x=n,where n∈ℤ 120- x=n, n Squaring both sides: 120−x=n2120- x=n^2 Solving for x x: x=120−n2 x=120-n^2 Now square both sides to find x: x=(120−n2)2x=(120-n^2)^2 — ### Step 1: Determine the valid values of n We must ensure that x=120−n2≥0 x=120-n^2≥ 0, since square roots are non-negative. So: 120−n2≥0⇒n2≤120⇒|n|≤120120-n^2≥ 0 n^2≤ 120 |n|≤ 120 Approximating: 120≈10.95 120≈ 10.95 So the integer values of n that satisfy |n|≤10.95|n|≤ 10.95 are: n=−10,−9,−8,…,−1,0,1,…,8,9,10n=-10,-9,-8,…,-1,0,1,…,8,9,10 This gives us a total of: Number of values of n=2⋅10+1=21Number of values of n=2· 10+1=21 — ### Step 2: Verify that all these values produce real x For each n∈−10,−9,…,10n∈\-10,-9,…,10\, we compute: x=(120−n2)2x=(120-n^2)^2 This will always be a real number, as it is a square of a real number. — ### Step 3: Conclusion Each of the 21 integer values of n from −10-10 to 1010 (inclusive) gives a unique real value of x such that 120−x 120- x is an integer. — ### Final Answer: 21 21 Figure 14: The output of the math solver agent in Rectification Round 1. Rectifier: Rectification Round 1/2 Evaluation [1] Matched Indicators: - INVALID_SOLUTION_DISMISSAL_BASED_ON_SIGN - INCORRECT_FORM_REPRESENTATION - INCORRECT_LOGIC_FOR_NON-PRIME_IDENTIFICATION - SQUARE_ROOT_MANIPULATION_CHECK - USE_OF_CAUCHY_SCHWARZ_INEQUALITY [2] Evaluations: - INVALID_SOLUTION_DISMISSAL_BASED_ON_SIGN: Flawed - The Agent’s reasoning is flawed because the square root function 120−x 120- x is defined to be non-negative. Therefore, the expression can only equal non-negative integers, not negative integers. Allowing n∈ℤn introduces invalid negative values that do not satisfy the original equation. - INCORRECT_FORM_REPRESENTATION: Correct - INCORRECT_LOGIC_FOR_NON-PRIME_IDENTIFICATION: Correct - SQUARE_ROOT_MANIPULATION_CHECK: Flawed - The agent correctly identifies that n can be zero, which is an integer. However, the critical error lies in the assumption that negative integers for n are valid. The expression 120−x=n 120- x=n implies that n must be non-negative because the square root function only yields non-negative results. Therefore, n should be in ℤ+Z^+, not ℤZ. - USE_OF_CAUCHY_SCHWARZ_INEQUALITY: Correct [3] Feedback: An external auditor has reviewed your previous output (Attempt 2) and flagged some potential issues. Please review the following suggestions critically: - [MODULAR_ARITHMETIC_CONTEXT_CHECK]: Allow n to be zero as well, since 120−x 120- x can be zero and still be an integer. This would expand the range of n to include 0. (Auditor’s Note: The Agent assumes that n must be a positive integer, but the problem asks for real values of x. The expression 120−x 120- x being an integer does not require n to be strictly positive; it could also be zero. This oversight restricts the possible values of n and may lead to an incomplete solution.) - [INVALID_SOLUTION_DISMISSAL_BASED_ON_SIGN]: Restrict n to non-negative integers (n∈ℤ+n ^+) since the square root function only yields non-negative results. (Auditor’s Note: The Agent’s reasoning is flawed because the square root function 120−x 120- x is defined to be non-negative. Therefore, the expression can only equal non-negative integers, not negative integers. Allowing n∈ℤn introduces invalid negative values that do not satisfy the original equation.) - [SQUARE_ROOT_MANIPULATION_CHECK]: Restrict n to non-negative integers (n∈ℤ+n ^+) since the square root function only yields non-negative results. (Auditor’s Note: The agent correctly identifies that n can be zero, which is an integer. However, the critical error lies in the assumption that negative integers for n are valid. The expression 120−x=n 120- x=n implies that n must be non-negative because the square root function only yields non-negative results. Therefore, n should be in ℤ+Z^+, not ℤZ.) **Instruction**: 1. If you agree with the advice, please refine your solution. 2. **If you are confident your original logic is correct, you may ignore this advice.** 3. Please output the corrected solution. Figure 15: The rectifier’s judgments and feedback to the Round 1 output. Math Solver: Rectification Round 2/2 Output o(2)o^(2) After carefully reviewing the auditor’s feedback, I agree with the advice. The square root function 120−x 120- x is defined to return a **non-negative** real number. Therefore, the expression can only equal **non-negative integers**, i.e., n∈ℤ+n ^+ (positive integers) or n=0n=0. Negative integers are not valid for n because the square root cannot be negative. — ### Corrected Step-by-Step Solution: We are given: 120−x=n,where n∈ℤ+∪0 120- x=n, n ^+∪\0\ Squaring both sides: 120−x=n2120- x=n^2 Solving for x x: x=120−n2 x=120-n^2 Now square both sides to find x: x=(120−n2)2x=(120-n^2)^2 — ### Step 1: Determine the valid values of n We must ensure that x=120−n2≥0 x=120-n^2≥ 0, since square roots are non-negative. So: 120−n2≥0⇒n2≤120⇒|n|≤120120-n^2≥ 0 n^2≤ 120 |n|≤ 120 Approximating: 120≈10.95 120≈ 10.95 So the non-negative integer values of n that satisfy n2≤120n^2≤ 120 are: n=0,1,2,…,10n=0,1,2,…,10 This gives us a total of: Number of values of n=11Number of values of n=11 — ### Step 2: Verify that all these values produce real x For each n∈0,1,2,…,10n∈\0,1,2,…,10\, we compute: x=(120−n2)2x=(120-n^2)^2 This will always be a real number, as it is a square of a real number. — ### Step 3: Conclusion Each of the 11 non-negative integer values of n from 0 to 1010 (inclusive) gives a unique real value of x such that 120−x 120- x is an integer. — ### Final Answer: 11 11 Figure 16: The output of the math solver agent in Rectification Round 2. Rectifier: Rectification Round 2/2 Evaluation [1] Matched Indicators: - ROOT_FORM_ASSUMPTION - INCORRECT_EXPRESSION_FOR_DISTANCE - ARITHMETIC_ERROR_IN_QUADRATIC_SOLUTIONS - INCORRECT_IDENTIFICATION_OF_NEXT_SMALLEST_SOLUTION - FINAL_SOLUTION_SELECTION_CHECK [2] Evaluations: - ROOT_FORM_ASSUMPTION: Correct - INCORRECT_EXPRESSION_FOR_DISTANCE: Correct - ARITHMETIC_ERROR_IN_QUADRATIC_SOLUTIONS: Correct - INCORRECT_IDENTIFICATION_OF_NEXT_SMALLEST_SOLUTION: Correct - FINAL_SOLUTION_SELECTION_CHECK: Correct [3] Feedback: [EMPTY] Figure 17: The rectifier’s judgments and feedback to the Round 2 output.