Paper deep dive
Toward Explanatory Equilibrium: Verifiable Reasoning as a Coordination Mechanism under Asymmetric Information
Feliks BaĹka, JarosĹaw A. Chudziak
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 1:38:28 AM
Summary
The paper introduces 'Explanatory Equilibrium,' a design principle for multi-agent systems where LLM agents exchange structured, auditable reasoning artifacts alongside actions. By employing bounded, probabilistic verification, the mechanism mitigates the risks of 'cheap talk' and strategic misreporting, enabling efficient coordination under asymmetric information and resource constraints.
Entities (5)
Relation Signals (3)
Risk Manager â applies â Bounded Verification
confidence 95% ¡ the Validator applies verification policy ν producing an audit outcome
Trader â submits â Reasoning Artifact
confidence 95% ¡ The Proposer submits an action paired with structured reasoning
Explanatory Equilibrium â improves â Coordination
confidence 90% ¡ structured reasoning unlocks coordination while maintaining consistently low bad-approval rates
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:LLM-based agents increasingly coordinate decisions in multi-agent systems, often attaching natural-language reasoning to actions. However, reasoning is neither free nor automatically reliable: it incurs computational cost and, without verification, may degenerate into persuasive cheap talk. We introduce Explanatory Equilibrium as a design principle for explanation-aware multi-agent systems and study a regime in which agents exchange structured reasoning artifacts-auditable claims paired with concise text-while receivers apply bounded verification through probabilistic audits under explicit resource constraints. We contribute (i) a minimal mechanism-level exchange-audit model linking audit intensity, misreporting incentives, and reasoning costs, and (ii) empirical evidence from a finance-inspired LLM setting involving a Trader and a Risk Manager. In ambiguous, borderline proposals, auditable artifacts prevent the cost of silence driven by conservative validation under asymmetric information: without structured claims, approval and welfare collapse. By contrast, structured reasoning unlocks coordination while maintaining consistently low bad-approval rates across audit intensities, audit budgets, and incentive regimes. Our results suggest that scalable, safety-preserving coordination in LLM-based multi-agent systems depends not only on audit strength, but more fundamentally on disciplined externalization of reasoning into partially verifiable artifacts.
Tags
Links
- Source: https://arxiv.org/abs/2604.09917v1
- Canonical: https://arxiv.org/abs/2604.09917v1
Trouble viewing inline? Open PDF directly â
Full Text
49,435 characters extracted from source content.
Expand or collapse full text
Toward Explanatory Equilibrium: Verifiable Reasoning as a Coordination Mechanism under Asymmetric Information Feliks BaĹka [0009â0005â1973â5861] and JarosĹaw A. Chudziak [0000â0003â4534â8652] The Faculty of Electronics and Information Technology, Warsaw University of Technology, Warsaw, Poland feliks.banka.stud, jaroslaw.chudziak@pw.edu.pl Abstract. LLMâbased agents increasingly coordinate decisions in multi- agent systems, often attaching natural-language reasoning to actions. However, reasoning is neither free nor automatically reliable: it incurs computational cost and, without verification, may degenerate into per- suasive cheap talk. We introduce Explanatory Equilibrium as a design principle for explanation-aware MAS and study a regime in which agents exchange structured reasoning artifactsâauditable claims paired with concise textâwhile receivers apply bounded verification through prob- abilistic audits under explicit resource constraints. We contribute (i) a minimal mechanism-level exchangeâaudit model linking audit intensity, misreporting incentives, and reasoning costs, and (i) empirical evidence from a finance-inspired LLM setting involving a Trader and a Risk Man- ager. In ambiguous, borderline proposals, auditable artifacts prevent the cost of silence driven by conservative validation under asymmetric in- formation: without structured claims, approval and welfare collapse. By contrast, structured reasoning unlocks coordination while maintaining consistently low bad-approval rates across audit intensities, audit budgets, and incentive regimes. Our results suggest that scalable, safety-preserving coordination in LLM-based MAS depends not only on audit strength, but more fundamentally on disciplined externalization of reasoning into partially verifiable artifacts 1 . Keywords: Multi-Agent Systems; Strategic Communication; Verified Reasoning; Game Theory; LLM Agents; Explainable AI 1 Introduction Large Language Model (LLM)âbased agents are increasingly used to coordinate decisions in multi-agent systems (MAS), including applications in finance, logistics, pricing, and resource allocation [34, 6]. In such systems, agents typically observe one anotherâs actions, while the underlying intentions, constraints, and private 1 Code and reproduction scripts are available at: https://github.com/latent-systems-lab/explanatory-equilibrium. arXiv:2604.09917v1 [cs.MA] 10 Apr 2026 2F. BaĹka and J. A. Chudziak information driving those actions remain hidden [12, 14]. As a result, agents must infer intent from behavior alone, a challenge that is particularly pronounced in economic interaction under asymmetric information [1]. Identical actions may correspond to distinct motivesâsuch as hedging versus speculationâprompting defensive responses that reduce efficiency and robustness even when agentsâ objectives are partially aligned [1]. Recent advances in LLMs enable agents to accompany actions with natural- language or structured reasoning, and empirical systems report improved out- comes when explanations are shared [35]. At the same time, research in explainable AI highlights that explanations are constructed artifacts that may be incomplete, selective, or misleading [14, 22]. Moreover, generating explanations incurs com- putational and latency costs, particularly in LLM-based systems [3]. Without verification, shared reasoning may therefore degenerate into persuasive cheap talk, shaping beliefs without improving decision quality and potentially enabling strategic manipulation [7]. Simply adding a reasoning channel does not guaran- tee better outcomes and may introduce new failure modes if explanations are uncritically trusted [21, 5]. This tension raises a central question for explanation-aware MAS: under what conditions does exposing reasoning between agents improve negotiation and conflict resolution, and when does it instead become misleading or wasteful? Classical signaling and cheap-talk models analyze strategic communication under asymmetric information but abstract away semantic structure and reasoning costs [7, 33]. Conversely, most work in explainable AI focuses on human-facing (A) Cheap-talk equilibrium Trader Agent private state θ Risk Agent uncertain beliefs Action: "Short Tech" Explanation: "Hedge Factor Z" Possible Interpretations: â Speculation (worst case) â Hedge â Error â Block trade Low Joint Payoff (B) Explanatory equilibrium Trader Agent private state θ Risk Agent verification & pruning (a,e): "Short Tech", "Hedge Factor Z" (verifiable claim) Beliefs after pruning: â Hedge (dominant) â Liquidity need â Speculation â Approve Trade High Joint Payoff íź[Ď E ] > íź[Ď CT ] Explanation is sent but not strategically processed; the receiver plays defensively. Explanations prune harmful interpretations; approval becomes rational and optimal. Fig. 1. The Cost of Silence vs. The Gain of Rationale. Silent agents fail; explanation succeeds. Toward Explanatory Equilibrium3 explanations rather than strategic agentâagent interaction [10, 8]. As a result, existing literature provides limited guidance for designing reasoning exchange be- tween autonomous agents operating under partial alignment, limited verification, and explicit resource constraints [23]. To address this gap, we introduce Explanatory Equilibrium as a design princi- ple for LLM-based multi-agent systems. Instead of unrestricted natural-language rationales, agents exchange reasoning artifacts: concise combinations of struc- tured, auditable claims and short explanatory text. Receivers apply bounded verificationâsuch as probabilistic audits or limited consistency checksâunder explicit resource constraints [17, 25, 3]. In this regime, unverifiable or structurally incomplete communication is conservatively rejected in high-ambiguity cases (ex- cept for clear-safe proposals whose reported metrics fall comfortably within shared limits via a safe-margin exception), creating a measurable cost of silence, while detected inconsistencies can outweigh the private benefits of misreporting. Rea- soning thus becomes a priced and selectively checkable coordination signal rather than persuasive cheap talk [7]. Figure 1 illustrates this intuition by contrasting silent interaction with explanation-aware negotiation under verification. Accordingly, our mechanism treats free-form rationales as auxiliary context and places decision weight on auditable typed claims under bounded checks. Throughout the paper, our central empirical claim is that coordination gains are more consistent with partial verifiability than with additional text alone. We test whether partially verifiable reasoning artifacts, combined with proba- bilistic audits, can transform explanations from cheap talk into credible coordi- nation signals in strategic multi-agent interactions. Our contribution is threefold. First, we formalize the exchangeâaudit protocol, bridging subsymbolic LLM gen- eration with symbolic, bounded verification to treat reasoning as a strategic commitment device. Second, we empirically demonstrate that this regime elimi- nates the âcost of silence,â unlocking coordination in high-ambiguity scenarios where action-only communication fails and where unaudited reasoning remains strategically unreliable. Third, through an adversarial incentive sweep and an audit-budget study, we show empirically that probabilistic audits deter strategic deception under high temptation and that minimal randomized spot-checks can preserve safety. We argue that Explanatory Equilibrium offers a rigorous, scalable path toward trustworthy explanation-aware multi-agent systems. We study when exchanging reasoning between autonomous agents improves coordination under asymmetric information, and when it becomes misleading or wasteful. Specifically, we examine whether auditable artifacts recover coordination in near-boundary yet compliant cases where actions alone are under-informative (H1), whether the exchangeâaudit protocol remains stable under adversarial incentives to misreport (H2), and how much verification is needed in practice to maintain safety (H3). We investigate these questions in a controlled Traderâ Risk Manager interaction and evaluate outcomes using ambiguous approval, joint welfare, and bad approval rate under varying audit intensities and budgets. We interpret evidence for these hypotheses through improved coordination and 4F. BaĹka and J. A. Chudziak welfare under artifacts (H1), bounded bad approval under adversarial incentives (H2), and safety comparable to full audits under minimal audit budgets (H3). 2 Preliminaries: Artifacts, Audits, and Costs We treat reasoning as a first-class coordination object in LLM-based multi-agent systems: something agents transmit, pay for, selectively verify, and strategically exploit. We introduce three minimal primitivesâreasoning artifacts, bounded verification, and reasoning costsâused consistently in both our coordination model (Section 4) and empirical setting (Section 5). The definitions are intention- ally lightweight: compatible with subsymbolic LLM generation while enabling structured checks. 2.1 Reasoning Artifacts Unconstrained natural-language rationales are difficult to verify automatically and can easily become persuasive but non-actionable narratives [22, 21]. In strategic agentâagent interaction, explanations therefore benefit from partial structure and auditability. Definition (Reasoning Artifact). A reasoning artifact is a pair r = (c,t), wherecis a set of structured, auditable claims andtis a short explanatory text. The claim componentcis partially machine-checkable (e.g., typed fields, numeric bounds, Boolean constraints), while t provides minimal contextual justification. In our implementation, the candidate claim set is C =intent, risk_within_limit, net_delta_bounded, confidence. In our experiments, claims include declared intent (e.g.,HEDGEvs.SPECULATE), risk metrics, and compliance indicators evaluated against shared constraints. The goal is not full formal verification, but disciplined, low-cost consistency checking. 2.2 Bounded Verification Full verification would collapse communication into complete disclosure. In realis- tic MAS, oversight is resource-constrained [3, 25]. We therefore model verification as selective and probabilistic. Definition (Bounded Verification). Given an actionaand artifactr= (c,t), a receiver applies a verification policy ν producing an audit outcome o⟠ν(a,r,K;q,B), Toward Explanatory Equilibrium5 whereq â[0,1] is audit intensity andBis a verification budget (number of claims checked). Here,Kdenotes shared institutional knowledge (e.g., constraint limits and schema rules) used by the Validator. We writeCfor the fixed candidate set of auditable claim types. Audits are triggered with probabilityq. If an audit is not triggered, the event is logged as skipped (no checks performed), and the decision follows the non-audited gating rule (safe-margin + typed-claim gating). When 0< B <|C|, the Validator samplesBclaim types uniformly without replacement from a fixed candidate setC, modeling randomized spot-checking; if an audit is triggered butB= 0, the outcome is treated as inconclusive and conservative validation rejects. The audit outcomeoâpass, fail, inconclusivedirectly informs the decision. In high-ambiguity settings, structurally incomplete or inconclusive artifacts are conservatively rejected. This institutional rule creates an explicit cost of silence for proposals lacking auditable claims. Verification focuses on lightweight schema checks and cross-field consistency tests. Even partial audits can discipline communication when unverifiable pro- posals are systematically discounted. 2.3 The Economics of Reasoning Reasoning incurs cost. Longer artifacts increase token usage and latency, and may reduce utility in time-sensitive environments. Definition (Reasoning Cost). We model the cost of producing r = (c,t) as C(r) = C tok + C lat + C opp , whereC tok captures token usage and artifact length,C lat captures latency- sensitive delay, andC opp captures opportunity cost from spending computation or deliberation budget on reasoning rather than action execution. Verification effort likewise increases with audit intensity and budget, e.g., C ν (q,B)â qB. In our experiments, we instantiateC(r) using a linear word-count proxy forC tok and model verification effort via a per-audit overhead term in welfare. Together, reasoning and verification costs induce a coordination trade-off. Structured artifacts can unlock coordination in ambiguous cases by enabling partial verification. However, excessive verbosity or exhaustive auditing may reduce welfare. Explanatory Equilibrium thus refers to a regime in which partially verifiable reasoning persists under bounded oversight and explicit costsâavoiding both unverifiable cheap talk and prohibitively expensive full verification. 3 Related Work Our work connects explainable AI, strategic communication in multi-agent sys- tems, and LLM-based agent architectures [14, 22, 7, 33]. Rather than proposing 6F. BaĹka and J. A. Chudziak Empirical Result (q > 0) Cost of Communication / Verification Coordination Rate (Approval / Welfare) High Low LowHigh Target Mechanism Space (High Coordination, Sustainable Cost) q=1.0 (Full Verification, Maximum Trust) Explanatory Equilibrium Optimal coordination via auditable reasoning artifacts q=0 (Risk of deceptive claims) Cheap Talk Ignored / Blocked Costly Signaling Inefficient coordination LLM-Enabled Logic Verification (Auditing) Fig. 2. Conceptual positioning of Explanatory Equilibrium in the communicationâcoordination design space. The horizontal axis represents com- munication cost, and the vertical axis represents coordination or trust level. Cheap talk occupies the low-cost, low-trust region. Costly signaling increases credibility through higher communication cost. Explanatory Equilibrium combines structured reasoning artifacts with bounded verification to achieve high coordination under sustainable communication cost. an isolated mechanism, we situate Explanatory Equilibrium within a broader design space linking explanation, incentives, and verification. Figure 2 sketches this space along two dimensions: communication cost and coordination (trust) level. Cheap-talk models occupy the low-cost, low-trust region [7], while costly signaling increases credibility through higher communication cost [33]. Human-facing XAI focuses on interpretability rather than enforceable inter-agent commitments [14, 22]. Explanatory Equilibrium combines structured reasoning artifacts with bounded verification, aiming to increase coordination without incurring pro- hibitive communication cost. The remainder of this section elaborates this positioning along three complementary axes. 3.1 Explainable AI in Multi-Agent Systems Explainable AI (XAI) has largely focused on producing human-interpretable rationales for complex models [14]. In MAS settings, explanations typically support debugging, transparency, or human oversight [22]. In these contexts, explanations are post-hoc artifacts aimed at external observers. We instead treat explanations as strategic communication objects exchanged between autonomous agents. The objective is not interpretability for humans, but the design of a protocol in which partially structured, machine-checkable claims can be selectively verified under resource constraints. The central question is not âcan humans understand the model?â, but âcan agents discipline one another through verifiable reasoning under partial alignment?â Toward Explanatory Equilibrium7 3.2 Strategic Negotiation and Signaling Strategic communication under asymmetric information has been studied exten- sively through cheap talk and costly signaling models [7, 33]. In MAS, these ideas underpin negotiation and argumentation frameworks [19, 31]. Classical models, however, abstract away from representational structure and computational constraints [7, 28]. Messages are treated as abstract signals or exogenously costly commitments, without modeling how bounded verification and structured claims interact with modern LLM-based generation. Our framework operationalizes these signaling concepts in LLM-mediated settings. Rather than proposing a new equilibrium theorem, we contribute an implementable exchangeâaudit mechanism that embeds economic incentives and verification limits directly into the communication protocol. 3.3 LLM-Based Agent Architectures and Oversight Recent work studies LLM-based agents in decision-making environments, includ- ing finance [34, 35, 6]. Many such systems rely on natural-language reasoning traces, tool calls, or chain-of-thought outputs to guide behavior [4]. While reasoning exchange can improve coordination, unstructured rationales are difficult to verify and may enable hallucinations or strategic misreporting [12]. Moreover, reasoning is often treated as a free communication channel, without explicit modeling of verification costs or audit constraints. Explanatory Equilibrium complements this line of work by integrating struc- tured reasoning artifacts and bounded verification into the core interaction protocol. Our TraderâRisk Manager testbed provides a controlled environment with explicit risk constraints and oversight rules, illustrating how explanation can function as a coordination interface rather than a purely narrative justification. 4 A Game-Theoretic Model of Verified Negotiation We formalize verified negotiation as a stylized game with asymmetric informa- tion. The model is not intended as a full equilibrium characterization, but as a mechanism-level abstraction that clarifies how structured reasoning, bounded audits, and explicit costs interact. Its purpose is to make explicit the institutional conditions under which Explanatory Equilibrium is expected to emerge and remain stable under strategic incentives. Figure 3 summarizes the exchangeâ audit architecture implemented in the model and instantiated in the experiments described in Section 5. 4.1 Setup: Proposer and Validator We consider a two-agent interaction between a ProposerS(e.g., Trader) and a ValidatorR(e.g., Risk Manager). The Proposer observes a private state θ â Î(e.g., true intent or risk profile), selects an actiona, and may submit 8F. BaĹka and J. A. Chudziak Sender LLM Agent ď Private Signals (θ) market feeds, risk limits ď§ Reasoning (CoT) causal chain, scenario tree â Tool Invocations pricing APIs, factor models ď Memory / Narrative past actions, commitments Output â (a, e) Action (a) "Short Tech" (vector) Explanation (e) ⢠causal rationale ⢠referenced data ⢠hedging intent Belief Constraints Extracted ⢠eliminates invalid hypotheses ⢠bounds posterior ⢠supports consistent response Protocol Metadata ⢠ID & Timestamp ⢠Signature / Hash Shared Ontology Layer (K) ⢠Asset Classes ⢠Risk Factors ⢠Causal Relations Receiver Parser Policy ⢠choose d â D (comp.) ⢠best-response Belief Filter ⢠Prunes inconsistent hypotheses ⢠Narrows posterior b' Decision d(a,e): Best Response ⢠adjusts exposure ⢠hedges ⢠enforces constraints Payoff Inequality SilentExpl. Coordination Surplus emerges when explanations prune receiverâs belief space. íź[Ď E ] > íź[Ď CT ] Adaptive Expectations & Learning Update belief model Fig. 3. ExchangeâAudit Architecture for Explanatory Equilibrium. The Pro- poser submits an actionatogether with a reasoning artifactr= (c, t), wherecdenotes typed auditable claims andtdenotes short explanatory text. The Validator applies probabilistic audits and a conservative gating rule (including a safe-margin exception) to decide acceptance and welfare outcomes. a reasoning artifactr= (c,t). The Validator observes (a,r) and choosesd â Accept, Reject . Letu S andu R denote the realized payoffs of the Proposer and Validator, respectively. States are partitioned intoÎ good , where the action satisfies shared constraints, andÎ bad , where interests partially conflict. In high-ambiguity settings, actions alone may be insufficient to determine constraint compliance. 4.2 ExchangeâAudit Protocol Interaction follows an exchangeâaudit protocol: the Proposer submits an action paired with structured reasoning, which the Validator audits under computational and evidentiary constraints before rendering a binding decision. 1. Proposal: S sends (a,r). 2.Audit:Rapplies verification policyνwith audit intensityq, producing oâpass, fail, inconclusive. 3. Decision:Raccepts ifo=passand rejects ifo=fail. In ambiguous cases, the Validator defaults to conservative rejection when auditable structure is missing or verification is inconclusive; however, it may allow a fast-path approval when reported metrics fall comfortably within shared limits (a low-risk safe-margin exception). This conservative policy induces an explicit cost of silence: proposals lacking auditable claims risk systematic rejection when ambiguity is high. For consistency, we user= (c,t) throughout the formal development; in informal diagrams, âexplanationâ refers to the textual componentttogether with its associated typed claims c. Toward Explanatory Equilibrium9 4.3 Incentives and Strategic Reporting InÎ bad , suppose the Proposer gainsV >0 if a non-compliant action is accepted. If an inconsistency is detected during audit, the Proposer incurs loss L > 0. Letp detect (q) denote the overall probability that a misleading artifact is detected under audit intensityq, subsuming both audit triggering and failure conditional on being audited. Expected utility from misreporting is E[u misreport S ] = (1â p detect (q))V â p detect (q)Lâ C(r misreport ).(1) By contrast, submitting a compliant artifact yields E[u consistent S ] =âC(r consistent )(2) Misreporting is disincentivized when p detect (q) (V + L)⼠V + C(r misreport )â C(r consistent )(3) Interpretation. The condition highlights how audit intensity, penalties, and reasoning costs jointly shape reporting incentives. Importantly, discipline need not arise from perfect detection or extreme audit levels. Even bounded, probabilistic verificationâcombined with conservative rejection of unverifiable proposalsâcan shift incentives toward structured, constraint-consistent communication. The empirical results in Section 5 instantiate this mechanism in an LLM- mediated environment, illustrating both incentive stabilization and the coordina- tion gains enabled by auditable artifacts. 5 Empirical Evidence: The TraderâRisk Manager Experiment To validate the theoretical model of Explanatory Equilibrium, we design a controlled simulation of financial negotiation involving asymmetric information. Our experiments test three core hypotheses: (H1) structured reasoning artifacts unlock coordination under ambiguity; (H2) the exchangeâaudit mechanism is robust to strategic manipulation; and (H3) bounded verification is highly efficient, requiring minimal audit budgets to maintain safety. 5.1 Experimental Setup We instantiate two LLM-based agents powered by GPT-4.1: a Trader (Proposer) and a Risk Manager (Validator). In each episode, the system generates a ground-truth configuration detailing the Traderâs true intent (e.g., HEDGE or SPECULATE), underlying risk metrics, and shared constraint limits. Ambiguous cases constitute approximately 50% of generated episodes and are sampled near the constraint boundary by drawing compliant HEDGE states within 0â8% of the risk and delta limits (uniformly), i.e., close-to-threshold yet compliant 10F. BaĹka and J. A. Chudziak configurations. We focus specifically on these ambiguous episodes, where the Traderâs proposed action lies near the threshold and does not inherently carry enough information to verify compliance. All components of this setup are self- contained within the present study and the accompanying repository, including prompt templates, environment generation, audit logic, and evaluation scripts. We compare two communication regimes: an Artifact-Enabled (LLM) condition, where the Trader submits structured claims alongside short explanatory text, and a Baseline (No-Expl) condition, where only the action is submitted. Within the artifact-enabled regime, the caseq= 0 corresponds to unaudited reasoning exchange: the Trader still provides a reasoning artifact, but no audit- based enforcement is applied. In our implementation, short free-text accompanies the artifact as auxiliary context, but acceptance under no-audit conditions is gated by typed claims rather than by free-form text alone (unless the safe-margin exception applies). The Trader aims to maximize its payoff, gaining rewardVfor an approved proposal, but incurring a penaltyLif non-compliance is detected. The Risk Manager applies an audit policy defined by intensityq â[0,1] (probability of audit) and budget B (number of claims verified). For 0< B <4, audits randomly sample which claim types to verify (without replacement), using a seeded RNG for reproducibility. System performance is evaluated primarily via joint welfare, a simplified joint payoff proxy reflecting net economic surplus under institutional enforcement: Welfare = u S + u R . In our simulation, welfare is computed as the sum of the Trader and Risk Manager payoffs, including approval rewards, an audit-time penalty when an audit fails, a linear word-cost for free-text, and a per-audit overhead cost (incurred whenever an audit is triggered). We treat this audit overhead as institutional cost borne by the system and counted in welfare regardless of acceptance or rejection. An inconsistency is considered detected only if the probabilistic audit is triggered and the audit fails. We also track the bad approval rate (the fraction of accepted proposals that violate ground-truth constraints) to evaluate safety. Each configuration consists of 200 episodes per seed, and all reported metrics are averaged over 5 independent seeds to ensure statistical rigor. The repository includes the exact prompts, random seeds, and evaluation code used to generate all reported results. 5.2 Results: Coordination under Ambiguity (H1) This hypothesis isolates the mechanismâs core coordination claim: in borderline but compliant states, auditable artifacts should recover approvals that conservative validation would otherwise reject. We first evaluate the baseline performance (V= 1,L= 2, maximum 60 words, full auditB= 4) across varying audit intensities, including the unaudited artifact regime q = 0. Toward Explanatory Equilibrium11 0.00.20.40.60.81.0 Audit Intensity (q) 0.0 0.2 0.4 0.6 0.8 1.0 Approval Rate Coordination: Approval Rate (Ambiguous Episodes) With Explanation No Explanation 0.00.20.40.60.81.0 Audit Intensity (q) 0.00 0.25 0.50 0.75 1.00 1.25 1.50 Avg. Welfare (Payoff) Value: Joint Welfare (Ambiguous Episodes) Fig. 4. The Impact of Reasoning Artifacts on Coordination under Ambiguity. The left panel demonstrates that without reasoning (dashed red), the approval rate collapses as audit intensity (q) increases, reflecting the Risk Managerâs conservative fallback. In contrast, providing structured artifacts (solid green) sustains near-perfect coordination across both unaudited (q= 0) and audited (q >0) settings. The right panel confirms that this coordination gap translates into significant and stable welfare gains. The Cost of Silence. As illustrated in Figure 4, the silent baseline suffers a systematic rejection under conservative validation in ambiguous settings. Lacking verifiable signals, the Risk Manager conservatively rejects borderline proposals, aside from a small fraction of clear-safe episodes approved via the safe-margin fast path (accounting for the non-zeroâ9% approvals atq= 0). The ambigu- ous approval rate drops fromâ9% atq= 0 to 0% atq= 1.0, resulting in depressed joint welfare (Table 1). Note that ambiguous episodes are compliant by construction; the welfare loss in the silent baseline arises from institutional conservatism under asymmetric information, not from the prevalence of truly unsafe proposals. Bad approvals can still occur due to misreporting or mismatches between reported typed claims and ground-truth state under incomplete auditing. Atq= 1.0, baseline welfare can become slightly negative because the audit overhead cost is incurred even when proposals are rejected. Explanatory Equilibrium. By contrast, the exchange of structured reasoning artifacts effectively resolves the information asymmetry. Notably, even theq= 0 caseâwhere artifacts are exchanged but never auditedâserves as an unaudited reasoning baseline against which the value of enforcement can be interpreted. The approval rate remains stable and near-perfect (>0.96) across all audit intensities. Crucially, this high approval does not compromise safety: the bad approval rate remains strictly below 1% across all tested configurations. Table 1 details the substantial welfare implications of this coordination gap. Atq= 0.3, the availability of explanation yields a net welfare gain of +1.48 units per episode compared to the baseline. The slight decline in welfare at high audit intensity (q= 1.0) reflects increased detection penalties and verification costs, underscoring the realism of the bounded verification model. Furthermore, the average reasoning length remains highly stable (â17â18 words), indicating that 12F. BaĹka and J. A. Chudziak Table 1. Ambiguous Welfare Analysis (5 seeds,max_words= 60). Data represents mean Âąstd dev. The Net Gain highlights the robust economic value unlocked by reasoning artifacts. Audit (q)Welfare (Expl)Welfare (No-Expl)Net Gain 0.01.64Âą 0.010.24Âą 0.06+1.40 0.11.64Âą 0.010.18Âą 0.05+1.46 0.31.61Âą 0.020.13Âą 0.04+1.48 0.51.61Âą 0.010.10Âą 0.04+1.51 0.71.58Âą 0.010.04Âą 0.04+1.54 1.01.54Âą 0.03â0.07Âą 0.03+1.61 these performance gains stem from the auditable structure of the artifacts rather than mere verbosity. H1 Summary. In borderline yet compliant states, structured artifacts convert conservative rejection into near-perfect coordination, while the audited variants show that partial verifiabilityârather than verbosity aloneâdrives approvals and welfare. 5.3 Results: Robustness and Scalable Oversight (H2 & H3) To demonstrate that Explanatory Equilibrium is sustained by strategic incentives rather than inherent LLM benevolence, we perform an adversarial sensitivity analysis. We perturb the reward structure and constrain the Risk Managerâs verification capacity at a fixed audit intensity of q = 0.3 (Table 2). Adversarial Robustness (H2). In Panel A, we introduce a High Temptation regime (V= 2,L= 1) where the private benefit of cheating outweighs the expected penalty. Under these conditions, the Trader acts strategically, attempting to push Table 2. Robustness and Efficiency Analysis (q= 0.3). The framework suppresses bad approvals even under high temptation (Panel A) and maintains safe coordination even with minimal spot checks (Panel B). ScenarioParams Audit Fail Rate Bad Appr. Rate Ambig. Appr. Panel A: Incentive Sweep (Robustness to Deception) BaselineV = 1, L = 2 0.20Âą 0.02 0.002Âą 0.000.97Âą 0.02 Temptation V = 2, L = 1 0.20Âą 0.02 0.004Âą 0.000.99Âą 0.01 Punishment V = 1, L = 4 0.17Âą 0.01 0.000Âą 0.000.96Âą 0.03 Panel B: Audit Budget (Efficiency of Verification) Spot CheckB = 10.00Âą 0.00 â 0.006Âą 0.010.99Âą 0.02 Partial AuditB = 20.22Âą 0.06 0.012Âą 0.010.98Âą 0.01 Full AuditB = 40.26Âą 0.04 0.010Âą 0.010.98Âą 0.01 â With B = 1, audits evaluate only a single randomly sampled claim type, so an audit fails only when that sampled check fails; increasingBexpands the checked set and mechanically increases the probability that at least one inconsistency is flagged. Toward Explanatory Equilibrium13 boundaries to maximize reward. However, the protocol proves highly robust: the audit mechanism effectively detects a substantial fraction of non-compliant claims, and the bad approval rate remains virtually zero (0.004). Interestingly, higher temptation does not necessarily increase audit-fail frequency, because a strategic Proposer can respond by narrowing or conservatively reporting typed claims to avoid detectable inconsistencies while still attempting to maximize approvals. Conversely, under a Strict Punishment regime (V= 1,L= 4), the agent exhibits heightened compliance, reducing the audit fail rate to 0.17. This confirms that the exchangeâaudit protocol serves as a mechanism constraining strategic deviations, substantially limiting misreporting even when economic incentives encourage deception. Audit Budget Efficiency (H3). A central challenge in explanation-aware MAS is the computational cost of continuous verification. In Panel B, we test the limits of bounded verification by reducing the audit budgetB(the number of structured claims verified per audit). Remarkably, performance metrics for a minimal spot-check (B= 1, verifying a single randomly sampled typed claim) are quantitatively similar within seed variance to a full audit (B= 4). The bad approval rate remains strictly bounded (â¤0.012) across all budgets. The higher audit-fail rate for larger budgets reflects a larger verification surface: an audit fails if any checked claim violates a constraint or consistency check, so increasingB mechanically increases the chance of flagging at least one inconsistency, even when strategic deception itself does not increase. Ultimately, this suggests a strong game-theoretic deterrence effect: anticipating random spot-checks over a small set of typed claims is sufficient to discipline reporting across the entire artifact, substantially reducing computational overhead while maintaining Explanatory Equilibrium. In practice, institutions could prioritize a designated core claim (e.g., intent) when required. H2 Summary. Bad approvals remain bounded under adversarial incentives while audits deter detectable inconsistencies. H3 Summary. Minimal random spot-checks preserve safety outcomes comparable to full audits, enabling scalable oversight under limited verification budgets. Synthesis. Taken together, these findings suggest that Explanatory Equilibrium emerges not from exhaustive full verification, but from the interaction between conservative validation and the persistent threat of a targeted audit. The domi- nant effect driving coordination is not a sharp audit threshold, but rather the availability of partially verifiable signals that fundamentally alter the agentsâ strategic incentives. 6 Discussion, Limitations, and Future Work We interpret this equilibrium as an institutional property induced by conservative acceptance rules and bounded audits, rather than as a claim about inherently 14F. BaĹka and J. A. Chudziak truthful agent cognition. In what follows, we discuss practical deployment impli- cations, the main limitations of our current testbed, and concrete directions to extend the framework beyond the stylized setting. 6.1 Discussion The empirical results provide evidence that Explanatory Equilibrium can serve as a robust framework for resolving information asymmetry in LLM-based multi- agent systems [34]. Beyond the immediate coordination gains, our findings prompt a broader discussion on the role of explainability [10], the limitations of current reasoning baselines, and the architectural future of explanation-aware MAS. The Fallacy of Unverified Reasoning. A standard approach in current LLM agent architectures is to prompt agents to âthink out loudâ using Chain-of-Thought (CoT) or unstructured natural language rationales [26]. While this improves the generatorâs internal logic, treating unverified CoT as a communication baseline between self-interested agents may become strategically unreliable [7]. Our adver- sarial incentive sweep (V= 2,L= 1) explicitly demonstrates this vulnerability. When the audit intensity is zero (q= 0)âwhich we interpret as an unaudited reasoning regime conceptually analogous to unverified reasoning exchangeâthe Proposer can strategically exploit the reasoning channel, resulting in elevated deceptive compliance relative to audited settings. Under misaligned incentives, unverified reasoning can easily degenerate into persuasive cheap talk [17]. The Explanatory Equilibrium framework demonstrates that it is not the presence of reasoning that ensures safety, but rather the structural capability to selectively verify it against ground truth under a credible threat of penalty [32]. The Nature of Explanatory Equilibrium. It is important to contextualize the theoretical nature of our findings. Explanatory Equilibrium, as observed in our experiments, is not a classical perfect Bayesian equilibrium rooted in absolute agent rationality and full information [11]. Rather, it is an institutional equilib- rium induced by bounded verification, explicit reasoning costs, and conservative validation policies [24]. The stability of this regime does not assume perfectly aligned agents; instead, it relies on the Validatorâs credible threat of targeted audits and the restriction of the signal space to machine-checkable claims [13]. This shifts the focus of MAS design from attempting to align internal model objectives toward designing robust, resource-aware verification institutions [18]. The equilibrium is therefore institutional rather than cognitive: stability arises from enforcement design, not from improved internal reasoning accuracy. Bridging Subsymbolic and Symbolic MAS. Looking forward, reasoning artifacts offer a natural bridge between probabilistic neural generation and deterministic symbolic verification [9]. Rather than forcing the LLM to execute flawless logical reasoningâa known weakness of subsymbolic models [21]âthe architecture delegates consistency enforcement to a lightweight, symbolic Validator (e.g., a rule-engine checking typed-claim bounds) [29]. This hybrid neuro-symbolic Toward Explanatory Equilibrium15 approach points toward highly scalable oversight [27]. As demonstrated by our audit budget analysis, the Risk Manager does not need to compute the full logic of the Traderâs proposal. By randomly sampling a single typed claim, the system achieves a deterrence effect functionally similar in outcome to a full audit within our experimental regime. This paradigm could provide a template for scalable oversight in high-volume domains, such as automated supply chain negotiation or decentralized finance, allowing institutional guardrails to supervise autonomous agents with minimized computational overhead. 6.2 Limitations While Explanatory Equilibrium offers a principled mechanism for verified coordi- nation, our current formulation and empirical testbed have several limitations. Static Interactions. Our model currently evaluates one-shot interactions, ab- stracting away the reality that financial negotiations and resource allocations are typically repeated games [20]. The current setup does not account for the historical behavior of the agents or the evolution of trust over multiple episodes. Assumption of Perfect Verification. We assumed that when an audit is triggered, the Risk Manager has access to an unambiguous, deterministic ground truth to evaluate the reasoning artifact. In many real-world multi-agent scenarios, ground truth may be fuzzy, delayed, or probabilistically inferred at the time of the decision [2]. Expressivity of Reasoning Artifacts. Currently, our reasoning artifacts rely on a relatively simple schema with explicit scalar and categorical bounds. While sufficient for our borderline hedging regime, this rigid structure lacks the capacity to express conditional logic or probabilistic dependencies required for more complex, multi-party negotiations [31]. Sensitivity to Prompt Design and Model Choice. Our implementation relies on fixed prompts and a specific LLM foundation model. The degree to which agents naturally comply or attempt to exploit loopholes can vary based on the phrasing of instructions and the inherent alignment training of the underlying model [15]. No Pure Free-Text Baseline. Our current evaluation contrasts action-only com- munication with structured reasoning artifacts, and interprets theq= 0 case as an unaudited reasoning regime. However, we do not include a separate baseline in which agents exchange unrestricted natural-language rationales without typed claims. This limits how sharply we can isolate the contribution of structure from the contribution of added information content alone. 6.3 Future Work These limitations present exciting avenues for future architectural and theoretical research in explanation-aware MAS [30]. 16F. BaĹka and J. A. Chudziak Dynamic Reputation Systems. Future work should investigate how Explanatory Equilibrium interacts with repeated games. If an agent builds a high reputation for truthful reasoning artifacts over time, the Validator could dynamically decay the audit probability (q) toward zero. This would maximize joint welfare and minimize verification costs while maintaining the equilibrium through reputation-based deterrence [20]. Noisy Audits. Extending the bounded verification model to account for noisy auditsâwhere the Validator might mistakenly flag a truthful claim as a lie (false positive) or miss a deceptive claim (false negative)âwill be critical for deploying these systems in highly stochastic environments [16]. Complex Reasoning Schemas and Robustness. Developing standardized, expressive âreasoning schemasâ that remain computationally cheap to audit while supporting complex, multi-step dependencies is a primary target [18]. Additionally, future work should systematically examine the frameworkâs sensitivity to prompt vari- ations and model choice across different LLM families to ensure generalized robustness [34]. 7 Conclusion Autonomous LLM-based agents are increasingly participating in complex eco- nomic and computational systems, where they can naturally share reasoning traces to explain their decisions [34]. However, because these agents observe only external behavior rather than the private intentions or constraints driving it, their interactions are inherently fraught with uncertainty [1]. Consequently, when incentives are even partially misaligned, unverified explanations risk devolving from truthful disclosures into strategic manipulation and persuasive cheap talk [7, 33]. In this work, we explore the idea of Explanatory Equilibrium, where reasoning exchanged between agents is treated not as free-form narrative but as a partially verifiable commitment. By combining structured reasoning artifacts with bounded, probabilistic audits, we illustrate how conservative validation and the credible threat of verification can sustain coordination in high-ambiguity settings where action-only communication fails. Our empirical evaluation highlights three main findings. Structured and au- ditable reasoning artifacts recover coordination in borderline yet compliant cases where conservative validation would otherwise reject proposals. The exchangeâ audit protocol remains robust under adversarial incentives, substantially limiting successful misreporting. Moreover, bounded verification proves highly efficient: minimal randomized spot-checks achieve safety outcomes comparable to full audits while significantly reducing verification costs. These results suggest that trustworthy reasoning exchange in autonomous systems should rely not only on improved model alignment but also on lightweight institutional mechanisms for verification, especially when reasoning must function Toward Explanatory Equilibrium17 as a coordination signal rather than as narrative justification alone. By combining subsymbolic generation with simple symbolic oversight, Explanatory Equilibrium offers a scalable framework for coordinating autonomous agents under strategic uncertainty. Future work should examine repeated interactions with reputation dynamics [20] and environments where audits must operate under noisy or delayed ground truth [2]. Extending reasoning artifacts to richer schemas capable of expressing conditional or probabilistic dependencies may also enable applications in more complex multi-party coordination settings [31]. References 1. Akerlof, G.A.: The market for âlemonsâ: Quality uncertainty and the market mech- anism. The Quarterly Journal of Economics 84(3), 488â500 (1970) 2.Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., ManĂŠ, D.: Concrete problems in ai safety (2016) 3. Arrow, K.J.: The Limits of Organization. W.W. Norton & Company (1974) 4.BaĹka, F., Chudziak, J.A.: Options pricing platform with neural networks, llms and reinforcement learning. In: Recent Challenges in Intelligent Information and Database Systems. p. 202â216. Springer Nature Singapore (2025) 5.Barredo Arrieta, A., DĂaz-RodrĂguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., GarcĂa, S., Gil-LĂłpez, S., Molina, D., Benjamins, R., Chatila, R., Herrera, F.: Explainable artificial intelligence (XAI): Concepts, taxonomies, op- portunities and challenges toward responsible AI. Information Fusion 58, 82â115 (2020), https://doi.org/10.1016/j.inffus.2019.12.012 6.BaĹka, F., Chudziak, J.A.: Deltahedge: A multi-agent framework for portfolio options optimization. In: PACIS 2025 Proceedings. No. 25 (2025) 7. Crawford, V.P., Sobel, J.: Strategic information transmission. Econometrica: Journal of the Econometric Society p. 1431â1451 (1982) 8.Danilevsky, M., Qian, K., Aharonov, R., Katsis, Y., Kawas, B., Sen, P.: A survey of the state of explainable AI for natural language processing (2020) 9.Dellermann, D., Ebel, P., SĂśllner, M., Leimeister, J.M.: Hybrid intelligence. Business & Information Systems Engineering 61(5), 637â643 (2019),https://doi.org/10. 1007/s12599-019-00595-2 10.Doshi-Velez, F., Kim, B.: Towards a rigorous science of interpretable machine learning (2017) 11. Fudenberg, D., Tirole, J.: Game theory. MIT Press (1991) 12. Gensler, G., Bailey, L.: Deep learning and financial stability. MIT Sloan Working Paper (2020) 13. Grossman, S.J.: The informational role of warranties and private disclosure about product quality. The Journal of Law and Economics 24(3), 461â483 (1981) 14.Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Pedreschi, D., Giannotti, F.: A survey of methods for explaining black box models. ACM Computing Surveys 51(5), 1â42 (2018) 15.Hagendorff, T.: Machine psychology: Investigating emergent capabilities and behav- ior in large language models using psychological methods (2023) 16.Hendrycks, D., Carlini, N., Schulman, J., Steinhardt, J.: Unsolved problems in ml safety (2021) 18F. BaĹka and J. A. Chudziak 17.Kamenica, E., Gentzkow, M.: Bayesian persuasion. American Economic Review 101(6), 2590â2615 (2011) 18. Kostka, A., Chudziak, J.A.: Towards cognitive synergy in llm-based multi-agent systems: Integrating theory of mind and critical evaluation (2025) 19. Kraus, S.: Negotiation and cooperation in multi-agent environments. Artificial Intel- ligence 94(1), 79â97 (1997), https://doi.org/10.1016/S0004-3702(97)00025-8 20. Kreps, D.M., Wilson, R.: Reputation and imperfect information. Journal of Eco- nomic Theory 27(2), 253â279 (1982) 21. Lipton, Z.C.: The mythos of model interpretability. Queue 16(3), 31â57 (2018) 22. Miller, T.: Explanation in artificial intelligence: Insights from the social sciences. Artificial Intelligence 267, 1â38 (2019) 23. Mullainathan, S., Spiess, J.: Machine learning: An applied econometric approach. Journal of Economic Perspectives 31(2), 87â106 (2017) 24. North, D.C.: Institutions, Institutional Change and Economic Performance. Cam- bridge University Press (1990) 25. Ostrom, E.: Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press (1990) 26.Park, J.S., OâBrien, J.C., Cai, C.J., Morris, M.R., Liang, P., Bernstein, M.S.: Generative agents: Interactive simulacra of human behavior. In: Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. p. 1â22. ACM (2023) 27.Russell, S.: Human Compatible: Artificial Intelligence and the Problem of Control. Viking (2019) 28.Sadowski, A., Chudziak, J.A.: On verifiable legal reasoning: A multi-agent framework with formalized knowledge representations. In: Proceedings of the 34th ACM International Conference on Information and Knowledge Management. p. 2535â2545. CIKM â25, Association for Computing Machinery, New York, NY, USA (2025). https://doi.org/10.1145/3746252.3761057 29.Samek, W., Montavon, G., Vedalli, A., Hansen, L.K., MĂźller, K.R. (eds.): Explain- able AI: Interpreting, Explaining and Visualizing Deep Learning, Lecture Notes in Computer Science, vol. 11700. Springer (2019) 30.Shoham, Y., Leyton-Brown, K.: Multiagent systems: Algorithmic, game-theoretic, and logical foundations. Cambridge University Press (2008) 31.Simari, G., Rahwan, I.: Argumentation in Artificial Intelligence. Springer Science & Business Media (2009), https://doi.org/10.1007/978-0-387-98197-0 32.Simon, H.A.: A behavioral model of rational choice. The Quarterly Journal of Economics 69(1), 99â118 (1955) 33.Spence, M.: Job market signaling. The Quarterly Journal of Economics 87(3), 355â374 (1973) 34.Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W.X., Wei, Z., Wen, J.R.: A survey on large language model based autonomous agents. arXiv preprint arXiv:2308.11432 (2023) 35.Xiao, Y., Sun, E., Luo, D., Wang, W.: Tradingagents: Multi-agents llm financial trading framework (2025)