Paper deep dive
Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social Simulation
Weiwei Fang, Lin Li, Kaize Shi, Yu Yang, Jianwei Zhang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/27/2026, 1:12:58 AM
Summary
BEACOF is a belief-driven adaptive collaboration framework for multi-agent systems, designed to overcome the rigidity of static interaction topologies in social simulations. By modeling interactions as a dynamic game of incomplete information and utilizing an Approximate Perfect Bayesian Equilibrium (PBE) mechanism, BEACOF enables agents to autonomously switch between cooperative, competitive, and coopetitive strategies based on inferred peer capabilities, leading to more robust and nuanced social simulations.
Entities (5)
Relation Signals (3)
BEACOF → isinspiredby → Perfect Bayesian Equilibrium
confidence 100% · we propose BEACOF, a belief-driven adaptive collaboration framework inspired by Perfect Bayesian Equilibrium (PBE).
BEACOF → implements → Approximate Perfect Bayesian Equilibrium
confidence 95% · our approach leverages an Approximate Perfect Bayesian Equilibrium (PBE) mechanism to rigorously couple belief updates with strategy selection.
BEACOF → targets → Social Simulation
confidence 95% · demonstrating superior potential for reliable social simulation.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:High-fidelity social simulation is pivotal for addressing complex Web societal challenges, yet it demands agents capable of authentically replicating the dynamic spectrum of human interaction. Current LLM-based multi-agent frameworks, however, predominantly adhere to static interaction topologies, failing to capture the fluid oscillation between cooperative knowledge synthesis and competitive critical reasoning seen in real-world scenarios. This rigidity often leads to unrealistic ``groupthink'' or unproductive deadlocks, undermining the credibility of simulations for decision support. To bridge this gap, we propose \textit{BEACOF}, a \textit{belief-driven adaptive collaboration framework} inspired by Perfect Bayesian Equilibrium (PBE). By modeling social interaction as a dynamic game of incomplete information, BEACOF rigorously addresses the circular dependency between collaboration type selection and capability estimation. Agents iteratively refine probabilistic beliefs about peer capabilities and autonomously modulate their collaboration strategy, thereby ensuring sequentially rational decisions under uncertainty. Validated across adversarial (judicial), open-ended (social) and mixed (medical) scenarios, BEACOF prevents coordination failures and fosters robust convergence toward high-quality solutions, demonstrating superior potential for reliable social simulation. Source codes and datasets are publicly released at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.24973v1
- Canonical: https://arxiv.org/abs/2603.24973v1
Trouble viewing inline? Open PDF directly →
Full Text
72,203 characters extracted from source content.
Expand or collapse full text
Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social Simulation Weiwei Fang 311137@whut.edu.cn Wuhan University of Technology Wuhan, China Lin Li ∗ cathylilin@whut.edu.cn Wuhan University of Technology Wuhan, China Kaize Shi Kaize.Shi@unisq.edu.au University of Southern Queensland Toowoomba, Australia Yu Yang yangyy@eduhk.hk The Education University of Hong Kong Hong Kong, China Jianwei Zhang zhang@iwate-u.ac.jp Iwate University Morioka, Japan Abstract High-fidelity social simulation is pivotal for addressing complex Web societal challenges, yet it demands agents capable of authen- tically replicating the dynamic spectrum of human interaction. Current LLM-based multi-agent frameworks, however, predomi- nantly adhere to static interaction topologies, failing to capture the fluid oscillation between cooperative knowledge synthesis and competitive critical reasoning seen in real-world scenarios. This rigidity often leads to unrealistic “groupthink” or unproductive deadlocks, undermining the credibility of simulations for decision support. To bridge this gap, we propose BEACOF, a belief-driven adaptive collaboration framework inspired by Perfect Bayesian Equi- librium (PBE). By modeling social interaction as a dynamic game of incomplete information, BEACOF rigorously addresses the cir- cular dependency between collaboration type selection and ca- pability estimation. Agents iteratively refine probabilistic beliefs about peer capabilities and autonomously modulate their collab- oration strategy, thereby ensuring sequentially rational decisions under uncertainty. Validated across adversarial (judicial), open- ended (social) and mixed (medical) scenarios, BEACOF prevents coordination failures and fosters robust convergence toward high- quality solutions, demonstrating superior potential for reliable so- cial simulation. Source codes and datasets are publicly released at: https://github.com/WUT-IDEA/BEACOF. CCS Concepts • Computing methodologies→Modeling and simulation; Multi-agent systems. Keywords Social Simulation, Multi-Agent Collaboration, Perfect Bayesian Equilibrium, Large Language Models ∗ Corresponding author This work is licensed under a Creative Commons Attribution-NonCommercial- NoDerivatives 4.0 International License. W ’26, Dubai, United Arab Emirates © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2307-0/2026/04 https://doi.org/10.1145/3774904.3792976 ACM Reference Format: Weiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang. 2026. Belief- Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equi- librium for Social Simulation. In Proceedings of the ACM Web Conference 2026 (W ’26), April 13–17, 2026, Dubai, United Arab Emirates. ACM, New York, NY, USA, 12 pages. https://doi.org/10.1145/3774904.3792976 1 Introduction Simulating complex societal dynamics constitutes a fundamen- tal challenge in the pursuit of “Web for Good,” analogous to re- cent advancements in agent-based user simulation for web ecosys- tems [13,23,45]. It enables researchers and policymakers to antici- pate the impacts of interventions in critical domains ranging from judicial fairness [17] to public health consensus [36]. In this con- text, Large Language Model (LLM)-based multi-agent systems have emerged as a transformative paradigm for high-fidelity social simu- lation [25,28,54]. By populating digital environments with agents capable of human-like reasoning and interaction, these systems offer a microcosm for studying collective behavior and problem- solving without the ethical risks of real-world experimentation [41]. However, the fidelity of such social simulations hinges critically on the nature of collaboration. Specifically, we define this fidelity as the capacity to autonomously modulate collaboration modes, mirroring the fluid transitions between consensus and conflict inherent in real- world social dynamics. Authentic societal interactions—whether in a courtroom debate, a medical consultation, or an open social dialogue—are rarely static; they fluctuate dynamically between cooperative knowledge synthesis and competitive critical reason- ing. Consequently, enabling agents to autonomously navigate this “coopetition” spectrum is not merely a technical optimization, but a prerequisite for modeling socially responsible outcomes that avoid the pitfalls of echo chambers or polarization [6,12], which are pervasive issues in current social web platforms. Current frameworks predominantly adhere to static topologies. Exclusively cooperative models (e.g., CAMEL [21], MetaGPT [16]) risk “groupthink” and error amplification [3], while strictly compet- itive approaches (e.g., MAD [24]) often succumb to unproductive deadlocks [7,40]. Although heuristic or rule-based methods attempt to bridge this gap [29,33,43], they lack the granularity to adaptively optimize strategies based on evolving states [47]. As illustrated in Figure 1, this rigidity leads to failure in complex tasks like judicial arXiv:2603.24973v1 [cs.MA] 26 Mar 2026 W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang "Great. We have reached a consensus. We will file the motion for acquittal immediately." "[Refinement] I see. That timeline is crucial. Let's modify the conclusion: the defendant is guilty of 'Imperfect Self- Defense' rather than being fully innocent." "You are ignoring the initial threat! I refuse to accept your characterization. We are at an impasse." Coopetitive Pure Competitive A legal case involves a defendant claiming "Self-Defense" after striking an attacker. Agent A "The defendant acted in self-defense because the plaintiff initiated the conflict with a weapon. Therefore, the defendant is innocent." "The logic follows the standard self-defense statute. Since the threat was initiated by the plaintiff, I support your conclusion of innocence. " Miscarriage of Justice (Overlooked "Excessive Force"). "Objection. You are biased towards the defendant. The act was clearly retaliation, not defense. I completely reject your premise. " No Consensus (Agents stuck in argument, task incomplete). "[Critique] Hold on. Evidence shows the defendant struck after the attacker was disarmed. This constitutes 'Excessive Force,' invalidating a pure self-defense claim." Accurate Verdict (Nuanced legal judgment achieved). Initial Proposition Pure Cooperative Agent B Agent A Agent B Agent A Agent B Agent A Agreeable Adversarial Agreeable Agreeable Adversarial (Note: This statement ignores that the defendant struck the attacker after the weapon was dropped — a critical legal detail.) (Blindly overlooks the excessive force issue.) (Attacks the stance without offering a constructive fix.) (Key Action: Identifying the flaw through critical reasoning.) (Key Action: Accepting the critique and collaborating on the correct definition.) Adversarial Figure 1: Comparison of three collaboration types in a judi- cial deliberation task. deliberation: where fixed collaboration type results in sycophantic agreement or impasse, dynamic switching enables agents to transi- tion from adversarial critique to cooperative refinement, ultimately converging on a legally nuanced verdict. Realizing such adaptivity, however, is non-trivial due to the fun- damental challenge of strategic decision-making under incomplete information—specifically, how agents can make rational collabora- tion decisions when peer capabilities are unobservable and must be inferred from noisy interaction signals. Since true capability is un- observable and textual feedback is often marred by hallucinations or inconsistency [18,37], naive estimation is prone to instability [26] or slow adaptation [36]. More critically, adaptive switching in- troduces a circular dependency problem: collaboration type selection depends on belief estimates of peer capabilities, yet belief updates are themselves influenced by the chosen collaboration mode. With- out principled coordination mechanisms, this interdependence can lead to erratic oscillations between strategies or premature conver- gence to suboptimal equilibria. To address this fundamental challenge, we propose BEACOF, a belief-driven adaptive collaboration framework grounded in Ap- proximate Perfect Bayesian Equilibrium (PBE) theory [10,44]. By employing a tractable approximation to circumvent the compu- tational intractability of exact inference within high-dimensional continuous type spaces, our framework provides a rigorous solu- tion to the circular dependency problem by establishing sequential rationality—ensuring agents make rational collaboration decisions given their current beliefs—while maintaining belief consistency through principled Bayesian updates. By modeling collaboration as a dynamic game, BEACOF successfully overcomes three critical challenges: it establishes strategic coordination where collaboration types are rational responses to anticipated behaviors; it ensures belief stabilization via theoretical convergence guarantees; and it enables real-time adaptive optimality unlike static approaches. We comprehensively validate our framework across three dis- tinct and challenging scenarios: adversarial (court debate), open- ended (persona-based dialogue), and mixed (medical Q&A). Re- sults demonstrate our approach achieves optimal or near-optimal performance in individual scenarios and superior cross-scenario generalization compared to baselines. Our main contributions are summarized as follows: (1) We pro- pose BEACOF, a belief-driven adaptive collaboration framework inspired by Perfect Bayesian Equilibrium that enables agents to autonomously switch collaboration types to align with interaction dynamics. (2) We address the fundamental challenge of strategic decision-making under incomplete information by introducing a tractable approximation mechanism that decouples belief estima- tion from strategy selection. This enables agents to make ratio- nal collaboration decisions, avoiding the erratic oscillations and suboptimal convergence inherent in ad-hoc switching approaches. (3) Extensive experiments demonstrate that BEACOF outperforms baselines, achieving gains of up to 3.4 points in F1 scores for adver- sarial settings against competitive baselines, improving accuracy by over 24 points against competitive baselines in mixed scenarios, and reducing persona contradiction by approximately 12.7 points while increasing diversity by over 10 points in open-ended dialogue, highlighting the framework’s superior generalization capabilities. 2 Related Work 2.1 Cooperative Multi-Agent Collaboration Cooperative paradigms prioritize role-based collaboration. CAMEL [21] pioneered role-playing with bidirectional protocols for task decomposition. MetaGPT [16] structured this via Standardized Op- erating Procedures (SOPs) and role-specific workflows. Similarly, AutoGen [39] provides infrastructure for complex multi-turn di- alogues. Other works utilize chain-of-thought [51] for collective reasoning or multi-round voting [4] for consensus. Crucially, how- ever, these methods rely on fixed interaction topologies, lacking the flexibility to dynamically adapt based on real-time assessment. 2.2 Competitive Multi-Agent Collaboration Conversely, competitive and debate-based systems leverage ad- versarial interactions for error detection and solution refinement. Multi-Agent Debate (MAD) [24] employs argumentation strategies where agents engage in adversarial discussions to surface weak- nesses in proposed solutions, effectively addressing the degeneration- of-thought problem often observed in single-agent reasoning. Du et al. [7] demonstrated that such multi-agent debate can signifi- cantly enhance mathematical and strategic decision-making capa- bilities. Extensions of this paradigm include diverse debate config- urations [31] and evaluation-focused systems [3], where multiple LLM agents act as adversarial referees to assess and critique text quality through structured disagreement. 2.3Game-Theoretic Approaches to Multi-Agent Collaboration Game theory provides a principled foundation for modeling strate- gic interactions and reasoning under uncertainty. Prominent appli- cations in multi-agent systems have demonstrated its efficacy across Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW ’26, April 13–17, 2026, Dubai, United Arab Emirates diverse domains, ranging from mechanism design for incentive op- timization [8,46] and auction-theoretic resource allocation [34] to evolutionary learning of emergent behaviors [5,11]. Furthermore, Bayesian game formulations have been extensively employed to model scenarios with incomplete information, providing theoretical frameworks for agents to reason about hidden opponent types and update beliefs based on observed actions [44, 48]. In summary, existing literature reveals two fundamental limita- tions in multi-agent collaboration: (1) structural rigidity in interac- tion modes, where fixed cooperative frameworks are susceptible to error amplification [2], while purely competitive approaches often succumb to deadlock [40]; and (2) insufficient theoretical adapt- ability, as prior game-theoretic applications predominantly focus on static scenarios [34] or idealized belief modeling within fixed- motive settings, failing to support dynamic transitions between diverse collaboration types. 3 Dynamic Game Formulation LLM-based multi-agent collaboration inherently involves strategic decision-making under uncertainty, where agents must act without full knowledge of other agents’ capabilities or intents. To rigorously model these, we ground our framework in the canonical theory of dynamic games of incomplete information [9]. Following these recent formulations, we map the vague linguistic interactions of agents into a structured tuple of Types, Actions, and Beliefs. The game environment is defined as a tupleG= ⟨N,Θ,A, H,U,푇⟩, proceeding over discrete time steps푡=1, . . .,푇. In this formulation, the term “dynamic” explicitly characterizes the evolu- tion of information states rather than physical parameters. Specif- ically, while the intrinsic capabilities (TypesΘ) of agents remain static latent variables, the public historyHand the agents’ internal belief estimates regarding peers are time-varying states that accu- mulate and evolve at each round푡. This sequential information update drives the adaptive strategy selection, distinguishing our framework from static one-shot interactions. Agents (N). LetN=1, . . .,푛be the set of participant agents. Each agent푖 ∈ Nacts as a strategic player. Additionally, a meta- agent serves as the mechanism designer to coordinate the process and provide evaluations. Types (Θ). To model intrinsic agent capabilities, each agent푖 possesses a private type vector휃 푖 ∈Θ 푖 ⊆ [0,1] 푑 , whereΘ 푖 is the type space of agent푖and푑denotes the dimension of capabilities (e.g., logic, rhetoric, empathy). Crucially, while the true type휃 푖 is static throughout the game, it is strictly private and unobservable to other agents푗≠ 푖. Consequently, agents must maintain a dynamic belief state: at round푡, agent푖holds a point estimate b 푡 푖 (푗) ∈Θ 푗 regarding agent푗’s type, which evolves based on interaction history. Action Space (A). LetA 푖 denote the action space for agent 푖. At each round푡, agent푖selects an action푎 푡 푖 ∈ A 푖 , defined as a tuple of intent and execution:푎 푡 푖 = (푐 푡 푖 ,푚 푡 푖 ), where푐 푡 푖 ∈ C= Cooperation, Competition, Coopetitionis the discrete strategy, and푚 푡 푖 ∈ Mis the textual response. The joint action space is A= Î 푖∈N A 푖 . Public History (H). Let푒 푡 푖 ∈ [ 0,1] 푑 denote the evaluation vector for the message푚 푡 푖 provided by the meta-agent. We define the public history available at the beginning of round푡as퐻 푡−1 = (푐 푘 푗 ,푚 푘 푗 ,푒 푘 푗 ) | 푗 ∈N,1≤ 푘< 푡, which aggregates all interaction tuples from preceding rounds. Contextual Utility (U). The utility function represents the pay- off derived from the outcome of interactions. Unlike static games, payoffs are dynamic and context-dependent:푢 푡 푖 :A ×H → R, where푢 푡 푖 (푎 푖 ,퐻 푡−1 )quantifies the gain of taking action푎 푖 given the current history. This is evaluated by the meta-agent based on task contribution and alignment with the chosen strategy. 4 Methodology Modeling multi-agent collaboration inherently involves strategic decision-making under uncertainty, where agents act without full knowledge of peer capabilities. To rigorously operationalize this process, we propose BEACOF, a framework that formulates the inter- action as a finite-horizon dynamic game of incomplete information. As illustrated in Figure 2, our approach leverages an Approximate Perfect Bayesian Equilibrium (PBE) mechanism to rigorously cou- ple belief updates with strategy selection. This structure enables agents to maintain sequential rationality—dynamically optimizing collaboration modes from cooperative synthesis to competitive critique—consistent with the evolving interaction history, thereby maximizing cumulative utility across diverse tasks. 4.1 Overview To operationalize Approximate PBE, we design a dual-layer ar- chitecture coupling strategic participant agents with a centralized meta-agent coordinator. All components are instantiated as LLMs driven by structured prompts, translating abstract game-theoretic calculus into executable natural language processes. The complete procedure is detailed in Algorithm 1, which or- chestrates the interaction between the meta-agent and participants, ensuring the synchronization of belief updates and strategic choices. Meta-Agent. At each round푡, the meta-agent acts as a central- ized coordinator. Given the public history퐻 푡−1 and the task state, it first constructs payoff vectors푈 푖 푡 푖∈N by scoring the desirability of each collaboration type (Line 4). It then predicts a probability distri- bution ˆ 푃 푡 (푐 푖 )over collaboration types for each agent (Line 5), and broadcasts these global signals to all participants (Line 6). Crucially, in the evaluation phase, the meta-agent assesses the generated mes- sage푚 푡 푖 to produce a tuple(푒 푡 푖 ,휔 푡 푖 )(Line 11), where푒 푡 푖 ∈ [0,1] 푑 is the capability estimate and휔 푡 푖 ∈ R + is the associated evalua- tion confidence. These functions are implemented via structured prompts (see App. C). Crucially, this decoupling offloads global state tracking to the meta-agent, preserving participant agents’ limited context windows for local reasoning and persona adherence, which is vital for maintaining coherence in open-ended scenarios. Participant Agents. Each participant agent푖is instantiated with a static role designation푟 푖 (e.g., “Plaintiff” in court debate) and maintains Gaussian belief estimates(푏 푡 푖 (푗),휔 푡 푖 (푗))regarding peer푗. At round푡, agent푖receives the payoff푈 푡 and predicted type distributions from the meta-agent. Conditioned on these signals, the agent computes an approximate best response푐 ∗ 푖 via Eq.(1) (Line 8) to maximize expected utility under uncertainty, and subsequently generates a message푚 ∗ 푖 = LLM(푐 ∗ 푖 ,푟 푖 ,퐻 푡−1 )(Line 9), thereby translating the abstract strategic intent into a concrete textual response that strictly aligns with its persona. W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang Court Debate Evidence Strength Legal Position Daily Chat Emotional Intelligence Relationship Building Q&A Diagnostic Confidence Expertise Competence Scenario-Specific Belief Dimensions ...... ... ... Belief-based Collaboration Participant Agent i Participant Agent j Interaction Cooperation Competition Coopetition 0.45 0.35 0.20 Interacion History Participants Scenario Profile Interaction Rounds Meta Agent Scenario Action Prediction Generate Payoff ... ... ... ... ... ... Belief and Confidence Updating Input Input Thought:Let's try to figure out the payoff estimates for Dr. Elaine Chen and Dr. Marcus Li in this medical Q&A scenario. First, I need to... Payoff:Dr. Elaine Chen: Cooperation: 7.25, Competition: 2.00, Co-opetition: 8.50 ... Payoff Generation Output Meta Agent Meta Agent Belief Evaluation Weighted Evidence ... ... Belief Updating Total Confidence Confidence Updating ... ... Action Space CooperationCompetition Coopetition Profile Private Type Strategic Decision-Making Probability 푃푟푖표푟 퐵푒푙푖푒푓 푏 ௧ିଵ 푖 푃푟푖표푟 Confidence 휔 ௧ିଵ 푖 푃푟푖표푟 퐵푒푙푖푒푓 푏 ୧ ௧ିଵ j 푃푟푖표푟 Confidence 휔 ୧ ௧ିଵ j 퐴푠표푐푖푎푡푒푑 퐶표푛푓푖푑푒푛푐푒 휔 ௧ 퐸푣푎푙푢푎푡푖표푛 푒 ௧ 푃푟푖표푟 퐵푒푙푖푒푓 푏 ௧ିଵ 푖 푃푟푖표푟 Confidence 휔 ௧ିଵ 푖 퐴푠표푐푖푎푡푒푑 퐶표푛푓푖푑푒푛푐푒 휔 ௧ 퐸푣푎푙푢푎푡푖표푛 푒 ௧ Figure 2: Overview of the Belief-driven Adaptive Collaboration Framework (BEACOF). The framework models collaboration as a dynamic game of incomplete information, applicable across diverse scenarios with specific belief dimensions (top panel). The central workflow executes as follows at round푡: (1) Meta-Agent Coordination: The centralized Meta-Agent utilizes scenario history to generate contextual payoffs푈 푡 and predict probability distributions over agent collaboration types. (2) Agent Strategic Action: A Participant Agent푖, conditioned on its private profile and the Meta-Agent’s outputs, computes an approximate best response strategy푐 ∗ 푖 and generates an interaction message푚 ∗ 푖 . (3) Evaluation & Belief Update: The Meta-Agent evaluates푚 ∗ 푖 to produce a capability estimate tuple(푒 푡 푖 ,휔 푡 푖 ) . The dashed callout box on the right details the critical Gaussian belief update mechanism: other peers (e.g., Agent푗, bottom left) refine their prior belief estimates regarding agent푖, denoted as b 푡−1 푗 (푖), by integrating this new evidence푒 푡 푖 weighted by confidence scores and a forgetting factor휆. This cyclic process drives the evolution of beliefs and strategic adaptation. Upon receiving the evaluation tuple(푒 푡 푖 ,휔 푡 푖 )from the meta-agent, all other agents푗≠ 푖update their beliefs regarding agent푖using the parametric Bayesian update rule defined in Eq.(2)(Lines 14–15). The interaction continues until the belief convergence criterion is met (Lines 18–20). 4.2 Approximate Perfect Bayesian Equilibrium Strictly enforcing PBE consistency is computationally intractable in high-dimensional continuous type spaces [36,38]. To address this, we propose a tractable approximation rooted in bounded ra- tionality [52] that relaxes strict requirements: we substitute exact integration with LLM-based reasoning for Sequential Rational- ity (Sec. 4.2.1) and employ a parametric Gaussian assumption for Belief Consistency (Sec. 4.2.2). 4.2.1 Approximate Sequential Rationality. To ensure adaptive per- formance in dynamic environments, at each time step, agent푖se- lects a collaboration strategy푐 푡 푖 that maximizes the expected utility against the predicted actions of opponents. Crucially, this predic- tion is conditioned on the estimated types b inferred from historical observations. Let휋(푐 −푖 |b −푖 ) denote the predicted distribution of all opponents’ strategies given their estimated types. The agent’s decision rule approximates a Best Response [1]: 푐 ∗ 푖 = arg max 푐∈C E 푐 −푖 ∼휋(·|b −푖 ) [ 푈 푖 (푐,푐 −푖 | 퐻 푡−1 ) ] ,(1) where푈 푖 represents the contextual utility evaluated by the meta- agent. In our implementation, the expectation operationEand the prediction휋are approximated via LLM-based reasoning rather than explicit numerical integration. 4.2.2 Belief Update with Confidence Decay. Instead of perform- ing intractable exact inference over the continuous type space, we adopt a parametric Bayesian approximation inspired by recent advancements in latent reasoning for agents [10]. We model the belief distribution as a multivariate Gaussian with isotropic preci- sion, maintaining only the first moment (estimate b) and a scalar precision (confidence 휔 ) to track uncertainty. Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW ’26, April 13–17, 2026, Dubai, United Arab Emirates Algorithm 1 Adaptive Multi-Agent Collaboration Framework 1:Input: Task specification푇, roles푅=푟 1 , . . .,푟 푛 , threshold휖, patience 퐾 , forgetting factor 휆 2:Initialize:퐻 0 =∅; for all푖, 푗 ∈N, set b 0 푖 (푗)=0.5·1 푑 ,휔 0 푖 (푗)= 휔 init 3: for round 푡= 1, 2, . . . do 4: 푈 푡 ← GenerateContextualPayoffs(퐻 푡−1 ,푇) 5: ˆ 푃 푡 (푐 푖 ) 푖∈N ← PredictAgentActions(퐻 푡−1 ,푈 푡 ) 6:Broadcast(푈 푡 , ˆ 푃 푡 (푐 푖 ) 푖∈N ) to all participant agents 7: for each agent 푖 ∈N do 8:Agent selects strategy and generates message 9: 푐 ∗ 푖 ← arg max 푐∈C E 푐 −푖 ∼ ˆ 푃 푡 [푈 푖 푡 (푐,푐 −푖 )] 10: 푚 ∗ 푖 ← GenerateMessage(agent 푖 ,푐 ∗ 푖 ,푟 푖 ,퐻 푡−1 ) 11:Meta-agent evaluates message and confidence 12: (푒 푡 푖 ,휔 푡 푖 ) ← EvaluateMessage(agent meta ,푚 ∗ 푖 ,푐 ∗ 푖 ,푇) 13:Update history: 퐻 푡 ← 퐻 푡−1 ∪(푚 ∗ 푖 ,푒 푡 푖 ,휔 푡 푖 ) 14:for each agent 푗 ∈N, 푗≠ 푖 do 15:Belief update with Gaussian approximation 16:b 푡 푗 (푖) ← 휔 푡−1 푗 (푖)·b 푡−1 푗 (푖)+휔 푡 푖 ·푒 푡 푖 휔 푡−1 푗 (푖)+휔 푡 푖 17:휔 푡 푗 (푖) ← 휆· 휔 푡−1 푗 (푖)+ 휔 푡 푖 18:end for 19: end for 20: if EarlyStopping(b 푘 푖 ,휖,퐾) is true then 21:break 22: end if 23: end for 24: Output: Final solution synthesized from 퐻 푡 At round푡, the meta-agent provides an evaluation푒 푡 푗 with an asso- ciated confidence휔 푡 푗 , acting as an observation. Under the Gaussian assumption (Normal-Normal conjugate prior), the posterior mean is derived via the standard inverse-variance weighted update [26]: b 푡 푖 (푗)= 휔 푡−1 푖 (푗)· b 푡−1 푖 (푗)+ 휔 푡 푗 · 푒 푡 푗 휔 푡−1 푖 (푗)+ 휔 푡 푗 .(2) To account for non-stationary agent behaviors, we introduce a forgetting factor휆 ∈ (0,1]that artificially inflates the posterior variance (reduces precision) at each step: 휔 푡 푖 (푗)= 휆· 휔 푡−1 푖 (푗)+ 휔 푡 푗 .(3) This formulation draws from adaptive filtering theory for track- ing time-varying parameters [26]. It functions as a computationally efficient approximation, allowing agents to dynamically adjust their beliefs in response to shifting peer capabilities [26]. To reduce computational cost while preserving solution quality, we employ an early stopping criterion based on belief stabilization. Intuitively, when an agent’s estimated capabilities of peers converge to a steady state, additional interaction rounds yield diminishing returns for strategic adaptation. For each agent푖, we quantify the belief shift between rounds 푡 − 1 and 푡 as the normalized Euclidean distance: Δ 푡 푖 = |b 푡 푖 − b 푡−1 푖 | 2 √ 푛·푑 ,(4) where b 푡 푖 concatenates all estimate vectors b 푡 푖 (푗)for푗 ∈ N. The normalization factor √ 푛·푑ensures scale invariance across different agent counts and dimension sizes. The framework terminates when the belief shift of at least one agent remains below a threshold 휖 for 퐾 consecutive rounds: ∃푖 ∈N :Δ 푡−푘 푖 < 휖, ∀푘 ∈ 0, 1, . . .,퐾 − 1.(5) This criterion serves as a proxy for system equilibration, balanc- ing strategy exploration with computational efficiency. A maximum horizon푇 max acts as a failsafe against slow convergence. 4.3 Theoretical Analysis: Beliefs Convergence In the context of dynamic games with incomplete information, the convergence of agents’ beliefs is a critical prerequisite for the sta- bility of a Perfect Bayesian Equilibrium [10]. Without theoretical guarantees, the belief update mechanism in Eq.(2)risks inducing cyclic or divergent behaviors, rendering the early stopping crite- rion (Eq. 5) unreachable. To rigorously address this, we analyze the asymptotic properties of our mechanism by formalizing it as a stochastic approximation process with a constant step-size [48]. Un- like standard decreasing step-size algorithms, our inclusion of a forgetting factor휆<1 implies convergence to a bounded region [20]. We characterize this behavior below. Proposition 4.1 (Bounded Convergence of Belief Estimates). Let the belief update follow Eq.(2)with휆 ∈ (0,1), assuming the meta- agent’s evaluation푒 푡 is an unbiased estimator of the true capability with bounded variance휎 2 . As푡 → ∞, the belief dynamics exhibit Effective Memory Stabilization, where the accumulated precision 휔 푡 converges to a steady state휔 ∞ ≈ E[휔 new ] 1−휆 , establishing a stable effective learning rate훼 ≈1− 휆. Consequently, the system achieves Mean-Square Stability, ensuring that the belief estimate b 푡 does not diverge but converges to a neighborhood of the true parameter, with asymptotic error variance bounded byO((1− 휆)휎 2 ). Proof Sketch.The proof leverages recent results from multi-agent convergence theory. First, regarding precision convergence, the update follows a linear difference equation푥 푡 = 휆푥 푡−1 +푢 푡 . Since 휆 ∈ (0,1), this constitutes a contractive mapping; by the Banach Fixed-Point Theorem, the sequence of expected precision converges to a unique fixed point휔 ∞ [32]. With the precision stabilized, the belief update becomes asymptotically equivalent to an Exponential Moving Average (EMA). According to [48] and [20], for a constant gain algorithm with step-size훼=1− 휆, the asymptotic covariance matrix of the estimation error satisfies the Lyapunov equation, yielding a bound proportional to훼휎 2 . This guarantees that the belief fluctuation∥b 푡 −b 푡−1 ∥remains within a bounded envelope defined by the noise level and the forgetting factor, thus validating the feasibility of the threshold-based termination criterion. 5 Experiments 5.1 Experimental Setup To evaluate BEACOF rigorously, we conduct experiments across three scenarios—adversarial, open-ended, and mixed—using identi- cal backbone LLMs and decoding settings for fair comparison. 5.1.1 Implementation Details. Backbone LLMs. We deploy agents via a localOllamaserver. To assess generalization across varying W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang Table 1: Main results across three evaluation scenarios. We report mean±standard deviation over five runs. Best results are in bold, second-best are underlined. For Contradiction, lower values indicate better performance. All values are in percentage. Court DebatePersona ChatMedQA Legal ArticlesJudgement ResultsModelMethod precisionrecallF1-scoreCharge AccSentence AccFine Acc DiversityConsistencyContradictionMedQA Acc CAMEL43.03±2.24 8.93±0.42 14.27±0.7283.67 83.67 83.67±1.53 43.67±0.58 46.33 46.33 46.33±1.5334.45±0.11 85.45±2.8914.55±2.8952.17±0.76 MAD31.08±0.58 8.49±1.08 13.43±1.5179.87±1.95 45.33±1.53 35.33±0.5832.97±0.73 87.94±2.6512.06±2.6534.50±0.50 ReConcile---44.75±0.35Llama3.1-8B-Instruct Ours43.87 43.87 43.87±0.60 12.28 12.28 12.28±0.23 16.87 16.87 16.87±0.8383.67 83.67 83.67±1.61 48.27 48.27 48.27±6.54 43.50±5.4136.85 36.85 36.85±1.85 88.08 88.08 88.08±0.3211.92 11.92 11.92±0.3252.23 52.23 52.23±1.86 CAMEL42.90±1.73 19.17±0.29 24.57±0.4273.47±4.6534.57±2.6 58.07 58.07 58.07±2.1630.94±1.75 76.86±0.7423.14±0.7454.50±1.32 MAD66.59±0.88 25.80±0.69 36.72±0.5592.33±0.93 50.33 50.33 50.33±2.08 53.00±1.0036.80±0.58 84.86±2.8215.14±2.8231.17±2.36 ReConcile---55.50 55.50 55.50±3.54Gemma3-12B Ours67.43 67.43 67.43±3.52 28.20 28.20 28.20±1.99 38.03 38.03 38.03±2.4893.00 93.00 93.00±2.00 43.00±2.65 46.67±3.2136.93 36.93 36.93±0.29 85.49 85.49 85.49±1.5714.51 14.51 14.51±1.5755.50 55.50 55.50±0.50 CAMEL33.43±5.08 19.00±0.56 22.50±1.5778.00±4.58 35.00±2.65 57.33 57.33 57.33±6.3528.44±0.13 70.28±1.8929.72±1.8984.83 84.83 84.83±0.58 MAD69.08±1.16 27.59±0.34 39.43±0.3696.33 96.33 96.33±0.58 53.67 53.67 53.67±2.08 56.67±1.5331.29±0.26 73.96±6.8326.04±6.8371.83±4.86 ReConcile---73.25±1.06Qwen3-30B-A3B Ours73.33 73.33 73.33±2.11 30.63 30.63 30.63±0.48 41.43 41.43 41.43±0.6694.33±1.53 49.86±0.24 50.67±4.5141.52 41.52 41.52±0.16 86.70 86.70 86.70±3.5713.30 13.30 13.30±3.5784.67±0.76 scales, we employ three open-source LLMs: Llama3.1-8B-Instruct (lightweight), Gemma3-12B [14] (efficient mid-sized), and Qwen3- 30B-A3B [42] (reasoning-optimized). We fix the generation length to 4096 tokens with temperature푇= 0 to ensure reproducibility. Hyperparameters. For the belief update mechanism defined in Eq. 2, we set the discount factor훽=0.6, while the forgetting factor휆 ∈ (0,1]adapts dynamically. The early stopping mechanism utilizes a belief-change threshold휖 change =0.05, a consensus thresh- old휖 cons =0.1, and a patience of퐾=3 rounds. The interaction terminates automatically once at least one agent’s belief change stays below휖 change for퐾consecutive rounds or when the maximum horizon푇 max = 4 is reached. 5.1.2 Scenarios and Datasets. To comprehensively evaluate the framework’s efficacy in generating reliable social simulations, we select three scenarios representing distinct archetypes of complex societal interaction: judicial conflict resolution, interpersonal social bonding, and professional consensus building (details in Appendix A). Court Debate (Adversarial: Judicial Fairness). We construct a curated dataset of 100 criminal cases, comprising 80 cases from the AgentsCourt benchmark [15] and 20 supplementary cases from China Judgements Online 1 . Here, agents act as opposing counsel (plaintiff and defendant) in a zero-sum game. From a Web4Good perspective, this setting is critical for testing whether agents can maintain rigorous, logical argumentation under intense pressure without succumbing to toxic aggression, thereby serving as a proxy for automated dispute resolution systems. This scenario poses a unique challenge distinct from standard NLP tasks: success requires agents to dynamically oscillate between interpreting rigid statutory constraints and constructing fluid, persuasive narratives, simulating the dual pressure of legal rigor and rhetorical adaptability required for effective advocacy. Persona Chat (Open-Ended: Social Inclusion). This scenario simulates the nuanced dynamics of everyday human connection and diversity. We select 100 pairs from PersonaChat [49]. Unlike rigid goal-oriented tasks, the aim is maintaining coherent, empa- thetic identities over long interactions. This evaluates the potential for digital inclusion and social well-being, ensuring agents adapt 1 https://wenshu.court.gov.cn to diverse personas while avoiding generic, repetitive or hollow interactions that hinder meaningful engagement. MedQA (Mixed: Public Health Consensus). This scenario simulates professional collaboration in critical domains, directly ad- dressing public welfare. We sample 200 questions from the MedQA dataset [19], where agents act as medical experts with distinct view- points. Unlike pure debate, this task requires a delicate balance: agents must compete to critique potential misdiagnoses while co- operating to synthesize a unified solution. This setting evaluates the framework’s ability to prevent “medical groupthink”—crucial for responsible AI in healthcare decision support. 5.2 Baselines and Evaluation Metrics We compare BEACOF against three static paradigms: CAMEL [21] (cooperation), MAD [24] (competition; utilizing AgentsCourt [15] for debate), and ReConcile [4] (consensus; MedQA only). All meth- ods use identical backbones to ensure fairness. Please refer to Ap- pendix B for full descriptions and implementation details. We employ a comprehensive set of task-specific metrics meticu- lously tailored to the distinct nature of each experimental scenario. Court Debate. We evaluate judicial decision-making on two levels: (1) Legal Article Prediction: We report Precision, Recall, and F1-score to measure the model’s ability to cite relevant statutes. (2) Judgment Prediction: We assess Charge Accuracy, Prison Term Accuracy, and Fine Accuracy. Following standard practices in legal AI [15], prison terms and fines are evaluated using a bucketed accuracy metric [15], counting predictions within the ground-truth interval as correct. This discretization accounts for the inherent variance in judicial discretion. Persona Chat. We assess dialogue quality across two dimen- sions: (1) Persona Consistency: We use a RoBERTa-Large NLI model [27] to classify persona-response pairs. We report Consis- tency Score (푃 푒푛푡 +푃 푛푒푢 ), summing entailment and neutral probabil- ities to capture valid non-contradictions, and Contradiction Score (푃 푐표푛 ) for hallucinations. (2) Response Diversity: We measure lexical richness via Distinct-1/2 [22] and Normalized Entropy [50]. An Overall Diversity score averages these three metrics. MedQA. For the medical task, we report standard Answer Ac- curacy [53], calculated as the proportion of questions where the agent’s final extracted choice matches the ground-truth option. Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW ’26, April 13–17, 2026, Dubai, United Arab Emirates Table 2: Average regret across scenarios and backbone LLMs. ModelCourt Debate Persona Chat MedQA Llama3.1-8B-Instruct0.7720.4040.190 Gemma3-12B0.5350.7420.012 Qwen3-30B-A3B0.4460.2510.031 5.3 Main Results Table 1 summarizes the performance across three scenarios. Overall, our framework demonstrates superior generalization capability. While specialized baselines suffer significant degradation when task dynamics shift, our belief-driven approach consistently achieves top-tier performance across diverse settings, effectively mitigating the limitations of fixed collaboration types. Court Debate: Process-Outcome Balance. Our method consis- tently outperforms the adversarial baseline (MAD) in Legal Articles F1 across all backbones (e.g., 41.43% vs. 39.43% with Qwen3), indicat- ing that adaptive collaboration fosters better statute identification than rigid competition. While MAD holds a slight edge in final Charge Accuracy on larger models due to its aggressive posture, our framework remains highly competitive (within a 2.0% gap on Qwen3) and even surpasses MAD on Llama3 (e.g., 48.27% Sentence Accuracy). This proves BEACOF achieves necessary adversarial dynamics without sacrificing the cooperative reasoning required for precise legal grounding. Persona Chat: Diversity-Consistency Trade-off. Our frame- work breaks the deadlock between entailment and diversity. Unlike CAMEL or MAD, we achieve the highest Diversity scores across almost all settings (e.g., 41.52 on Qwen3). Crucially, on the largest backbone (Qwen3), we reduce the Contradiction rate by approxi- mately 50% compared to baselines (Ours: 13.30% vs. MAD: 26.04%), demonstrating that dynamic strategy switching injects variability while preserving superior logical consistency. MedQA: Adaptability in Knowledge Tasks. Static competi- tion fails significantly in consensus-based tasks. MAD collapses on MedQA (e.g., 31.17% accuracy with Gemma3), as forced dis- agreement hinders knowledge synthesis. In contrast, BEACOF suc- cessfully adapts to cooperative requirements, outperforming the consensus-focused baseline ReConcile across most settings. No- tably, on Llama3 and Qwen3, our framework achieves absolute gains of 7.48% and 11.42% over ReConcile, respectively. Even on Gemma3, where mean accuracy is tied (55.50%), BEACOF exhibits significantly superior stability (std±0.50 vs.±3.54). Ultimately, our method attains statistical parity with the specialized cooperative baseline CAMEL (e.g., 84.67% vs. 84.83% on Qwen3), confirming that belief-driven adaptation effectively replicates cooperative benefits while avoiding the brittleness of heuristic voting mechanisms. Impact of Model Scale. Granular analysis reveals that larger models exploit the belief mechanism more effectively. While smaller models benefit generally, Qwen3-30B exhibits sharper strategic pivots, evidenced by the significantly widened gap in Persona Chat Consistency (Ours 86.70% vs. MAD 73.96%) compared to smaller backbones. This suggests that the computational benefits of our game-theoretic framework are amplified by the stronger reasoning capabilities of larger models. Court DebateMedQAPersona Chat 0 10 20 30 40 50 60 Performance Score 10.58 13.50 16.87 50.40 51.30 52.23 31.67 32.48 36.85 Variant w/o Belief w/o Type Ours F1 Accuracy Diversity Figure 3: Ablation study (Llama3.1-8B-Instruct): removing belief updates or fixing collaboration type degrades perfor- mance across scenarios. 5.4 Empirical Verification of Equilibrium Properties We focus on verifying Sequential Rationality, given that Be- lief Consistency is structurally guaranteed (Eq.(2)). We quantify rationality via Ex-post Regret (with payoffs normalized to[0,10]): 푟 푡 푖 = max 푐∈C 푈 푡 푖 (푐,푐 −푖 )−푈 푡 푖 (푐 푖 ,푐 ∗ −푖 ),(6) Table 2 demonstrates robust equilibrium approximation. First, av- erage regret remains below 0.5 (optimality gap<5%), indicating learned beliefs effectively guide decisions. Second, while adver- sarial scenarios (e.g., Court Debate) induce slightly higher regret (0.446–0.772), the deviation remains tightly bounded (<8%). Fi- nally, larger backbones (e.g., Qwen3-30B) consistently yield lower regret, suggesting stronger reasoning enhances PBE precision. 5.5 Ablation Study To disentangle the contributions of belief modeling and adaptive switching, we conduct ablation studies using Llama3.1-8B across all scenarios. We compare the BEACOF against two variants: (i) w/o Belief, which discards peer capability updates to rely solely on immediate payoffs; and (i) w/o Type, which enforces a fixed collaboration type throughout the interaction. As shown in Figure 3, BEACOF achieves superior performance, confirming that belief updates and type switching are complemen- tary. Specifically, removing belief updates causes severe degrada- tion in strategic settings (e.g., Court Debate F1 drops from 16.87 to 10.58, a∼37% loss), highlighting the importance of peer estimation. Furthermore, disabling type switching limits interaction variety, reducing Persona Chat Diversity scores from 36.85 to 32.48. This validates that dynamic strategy modulation effectively outperforms static cooperative or competitive paradigms. 5.6 Case Study To validate the framework in high-stakes social simulations, we analyze a representative MedQA trajectory in Figure 4. The scenario represents a classic failure mode in collaborative systems: collective confirmation bias, often manifested as sycophancy in LLMs [30]. The Trap of Social Groupthink. Initially, the interaction mir- rors an echo chamber,” where algorithmic homogeneity amplifies W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang A 28-year-old woman comes to the obstetrics and gynecology clinic because she has been unsuccessfully trying to conceive with her husband for the last 2 years. He has had 2 children from a previous marriage and has been tested for causes of male infertility. She has had prior abdominal surgery as well as a family history of endocrine abnormalities. Based on this history, a panel of tests are obtained and treatment is started. This treatment will be administered intermittently one week before her normal menstrual cycle. If the most likely drug prescribed in this case was administered in a different manner, which of the following conditions may also be treated by this drug? Question Options A: DiabetesB: Parkinson diseaseC: Polycystic ovarian syndromeD: Premature laborE: Prostate cancer Agent A I concur with Dr. Chen's assessment. The patient's history strongly points towards Polycystic Ovarian Syndrome (PCOS). The combination of infertility, prior abdominal surgery, and a family history of endocrine abnormalities is highly suggestive. The intermittent administration of a drug a week before the menstrual cycle strongly indicates Clomiphene Citrate, a common first-line treatment for ovulation induction in women with PCOS. While Clomiphene does have off-label uses, the options provided dont align with its typical applications. Therefore, Polycystic Ovarian Syndrome (PCOS)remains the most probable diagnosis given the clinical picture and likely medication. Agent B Agent A Similar discussions about Clomiphene Citrate... Agent B I concur with Dr. Chen’s assessmentthatPolycystic Ovarian Syndrome(PCOS) is the most likely diagnosis given the patient’s history and the likely use of Clomiphene Citrate. The intermittent administration a week before the menstrual cycle strongly supports this. However,the question specifically asks about what *other* condition could be treated if the drug were administered differently.While PCOS is the primary indication, it's crucial to acknowledge Clomiphene Citrate’s broader role as a Selective Estrogen Receptor Modulator (SERM). SERMs, including Clomiphene, have demonstrated efficacy in treating breast cancer by blocking estrogen’s effects on breast tissue. This is a well-established off-label use. Therefore, while PCOS is the correct diagnosis for this patient,option E, Prostate cancer, is the most appropriate answer to the question posed,as Clomiphene has been explored, albeit less extensively, for its potential in managing prostate cancer.2 Coopetition In subse quent rounds, both agents cooperated and chose E. Cooperation Cooperation ... C C Cooperation C E First reasoning leads to the wrong answer. Falling into groupthink and deviating from the correct answer. Aclevercombinationof cooperation and competition leads to the correct answer! 0.626, 0.725, 0.762, 0.789, 0.762 belief B →A 0.664,0.790,0.834,0.743,0.719 belief A →B 0.613, 0.764, 0.834, 0.746, 0.731 belief B →A Given the patient's history of infertility, prior abdominal surgery, and family history of endocrine abnormalities, Polycystic Ovarian Syndrome (PCOS) is a highly probable diagnosis. The question hints at a drug administered intermittently before the menstrual cycle, strongly suggestingClomiphene Citrate.Clomiphene Citrateis a selective estrogen receptor modulator (SERM) commonly used toinduce ovulation in women with PCOS. While it has some off-label uses, it is not typically used to treat Diabetes, Parkinson's disease, premature labor, or prostate cancer. Therefore, the most likely answer isC, Polycystic Ovarian Syndrome. I believe this is the most logical conclusion based on the provided information and the likely medication involved. 0.626, 0.725, 0.762, 0.685, 0.689 belief A →B Figure 4: The case study of dynamic collaboration type switching in resolving complex medical reasoning tasks. The framework BEACOF adaptively shifts the interaction type, guiding agents from an initial incorrect consensus to the ground truth. errors rather than correcting them [35]. Confronted with a com- plex patient history, Agent A latches onto the salient diagnosis (PCOS) but overlooks critical constraints. In a static framework, Agent B—suffering from degeneration-of-thought” [24]—blindly reinforces this error to maintain harmony. This illustrates how en- forced cooperation accelerates convergence to a false consensus, a primary cause of diagnostic errors. The Belief-Driven Intervention. Crucially, BEACOF breaks this deadlock not through randomness, but a socially grounded mechanism: loss of confidence. By Round 2, the meta-agent detects that repeated exchanges are yielding negligible information gain. Consequently, Agent B’s belief in Agent A declines sharply. This update acts as a decisive trigger, prompting Agent B to strategically switch from Cooperation” to Coopetition”. Constructive Dissent as a Solution. This strategic shift simu- lates constructive dissent within the dyad. Instead of seeking superfi- cial agreement, Agent B critically scrutinizes the premise to mitigate error propagation. This aligns with findings that multi-agent debate significantly enhances factuality [7]. Reliable social consensus re- quires not just aggregation, but autonomously disrupting harmony when reasoning is flawed. 6 Conclusion To transcend static limitations, we introduce BEACOF, which for- malizes collaboration as a dynamic game of incomplete information via Perfect Bayesian Equilibrium. Empirically, this belief-driven adaptation significantly surpasses fixed strategies in diverse scenar- ios. Future work will explore multi-agent mechanisms that faithfully mirror human social dynamics. Acknowledgments This work is partially supported by National Natural Science Foun- dation of China (No. 62276196) and The Education University of Hong Kong project under Grant No. RG 67/2024-2025R. Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW ’26, April 13–17, 2026, Dubai, United Arab Emirates References [1]Anton Bakhtin, Noam Brown, Emily Dinan, Gabriele Farina, Colin Flaherty, Daniel Fried, Andrew Goff, Jonathan Gray, Hengyuan Hu, Athul Paul Jacob, Hermann Baier, Mojtaba Komeili, Karthik Konath, Minae Kwon, Adam Lerer, Mike Lewis, Alexander H. Miller, Sasha Mitts, Adithya Renduchintala, Stephen Roller, Dirk Rowe, Naman Goyal, Arthur Szlam, and Jason Weston. 2022. Human- level play in the game of Diplomacy by combining language models with strategic reasoning. Science 378, 6624 (2022), 1067–1074. doi:10.1126/science.ade9097 [2]Mert Cemri, Melissa Z. Pan, Shuyi Yang, et al.2025. Why Do Multi-Agent LLM Systems Fail? arXiv:2503.13657 [cs.AI] doi:10.48550/arXiv.2503.13657 [3]Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2024. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate. In ICLR’24. OpenReview.net. doi:10. 48550/arXiv.2308.07201 [4]Justin Chih-Yao Chen, Swarnadeep Saha, and Mohit Bansal. 2024. ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). ACL’24, 7066–7085. doi:10.18653/V1/2024.ACL-LONG.381 [5]Filippos Christianos, Lukas Schäfer, and Stefano V. Albrecht. 2020. Shared Ex- perience Actor-Critic for Multi-Agent Reinforcement Learning. In NeurIPS’20. Curran Associates, Inc. doi:10.48550/arXiv.2006.07169 [6]Matteo Cinelli, Gianmarco De Francisci Morales, Alessandro Galeazzi, Walter Quattrociocchi, and Michele Starnini. 2021. The Echo Chamber Effect on Social Media. In W’21. ACM. doi:10.1073/pnas.2023301118 [7] Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mor- datch. 2024. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In ICML’24. PMLR. doi:10.48550/arXiv.2305.14325 [8] Itay Eilat, Ben Finkelshtein, Chaim Baskin, and Nir Rosenfeld. 2023. Strategic Classification with Graph Neural Networks. In ICLR’23. OpenReview.net. doi:10. 48550/arXiv.2302.04633 [9]Andrew Estornell and Yang Liu. 2024. Multi-LLM Debate: Framework, Principles, and Interventions. In NeurIPS’24. Curran Associates, Inc. doi:10.48550/arXiv.2410. 19890 [10] Jakob N. Foerster, Francis Song, Edward Hughes, Neil Burch, Iain Dunning, Shimon Whiteson, Matthew Botvinick, and Michael Bowling. 2019. Bayesian action decoder for deep multi-agent reinforcement learning. In ICML’19. PMLR, 1942–1951. doi:10.48550/arXiv.1811.01499 [11]Neil De La Fuente, Miquel Noguer i Alonso, and Guim Casadellà. 2024. Game The- ory and Multi-Agent Reinforcement Learning: From Nash Equilibria to Evolutionary Dynamics. arXiv:2412.20523 [cs.LG] doi:10.48550/arXiv.2412.20523 [12] Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md. Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. 2024. Bias and Fairness in Large Language Models: A Survey. In KDD’24. ACM. doi:10. 1145/3637528.3671569 [13]Chen Gao, Xiaoyi Du, Zhengxu Wan, Jinpeng Wang, Wenkai Dong, Junyu Fu, Yunfan Tu, and Yong Li. 2024. S3: Social-network Simulation System with Large Language Models. In CIKM’24. ACM. doi:10.1145/3627673.3679883 [14] Gemma Team. 2025. Gemma 3 Technical Report. arXiv:2503.19786 [cs.CL] doi:10. 48550/arXiv.2503.19786 [15] Zhitao He, Pengfei Cao, Chenhao Wang, Zhuoran Jin, Yubo Chen, Jiexin Xu, Huakai Jiang, Kang Liu, and Jun Zhao. 2024. AgentsCourt: Building Judicial Decision-Making Agents with Court Debate Simulation and Legal Knowledge Augmentation. In EMNLP’24 Findings. Association for Computational Linguistics, 9399–9416. doi:10.18653/v1/2024.findings-emnlp.549 [16]Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber. 2024. MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework. In ICLR’24. OpenReview.net. doi:10.48550/arXiv.2308.00352 [17]Kaixi Hu, Lin Li, Jianquan Liu, and Daniel Sun. 2021. DuroNet: A Dual-robust Enhanced Spatial-temporal Learning Network for Urban Crime Prediction. ACM Trans. Internet Techn. 21, 1, 24:1–24:24. doi:10.1145/3432249 [18]Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems 43, 2 (2025). doi:10.1145/3654944 [19]Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. 2020. What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams. arXiv:2009.13081 [cs.CL] doi:10. 48550/arXiv.2009.13081 [20]Stefanos Leonardos, Will Overman, Ioannis Panageas, and Georgios Piliouras. 2022. Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games. In ICLR’22. OpenReview.net. doi:10.48550/arXiv.2106.01969 [21]Guohao Li, Hasan Abed Al Kader Hammoud, Hani Itani, Dmitri Khizbullin, and Bernard Ghanem. 2023. CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society. In NeurIPS’23. Curran Associates, Inc. doi:10. 48550/arXiv.2303.17760 [22]Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016. A Diversity-Promoting Objective Function for Neural Conversation Models. In NAACL’16. Association for Computational Linguistics, 110–119. doi:10.18653/v1/ N16-1014 [23] Yicong Li, Hongxu Chen, Yile Li, Lin Li, Philip S. Yu, and Guandong Xu. 2023. Reinforcement Learning Based Path Exploration for Sequential Explain- able Recommendation. IEEE Trans. Knowl. Data Eng. 35, 11, 11801–11814. doi:10.1109/TKDE.2023.3237741 [24] Tian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang, Yan Wang, Rui Wang, Yujiu Yang, Zhaopeng Tu, and Shuming Shi. 2024. Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate. In EMNLP’24. Association for Computational Linguistics, 17889–17904. doi:10.18653/v1/2024.emnlp-main. 995 [25] Xiangneng Mou, Yang Ding, Kunkun Ren, et al.2024. From Individual to Society: A Survey on Social Simulation Driven by Large Language Model-based Agents. arXiv:2404.12077 [cs.MA] doi:10.48550/arXiv.2404.12077 [26]Siddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan, Karthik Krishna- murthy, and David Bauso. 2023. Scalable Multi-Agent Reinforcement Learning through Intelligent Information Aggregation. In ICML’23. PMLR, 25817–25833. doi:10.48550/arXiv.2211.02534 [27]Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020. Adversarial NLI: A New Benchmark for Natural Language Un- derstanding. In ACL’20. Association for Computational Linguistics, 4885–4901. doi:10.18653/v1/2020.acl-main.441 [28]Joon Sung Park, Joseph C. O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In UIST’23. ACM, 1–22. doi:10.1145/3586183.3606763 [29] Chen Qian, Wei Liu, Hongzhang Liu, Nuo Xu, Yifei Tao, Haoyu Dong, Chenghua Lin, Zhiyuan Liu, and Maosong Sun. 2024. ChatDev: Communicative Agents for Software Development. In ACL’24. Association for Computational Linguistics, 15174–15186. doi:10.18653/v1/2024.acl-long.829 [30]Mrinank Sharma, Meg Tong, Tomasz Korbak, et al.2023. Towards Understanding Sycophancy in Language Models. arXiv:2310.13548 [cs.CL] doi:10.48550/arXiv. 2310.13548 [31] Andries P. Smit, Nathan Grinsztajn, Paul Duckworth, Viet-Nhat Luong, Bamshad Mobasher, and P. K. M. 2024. Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs. In ICML’24. PMLR. doi:10.48550/arXiv.2311.17371 [32]Rayadurgam Srikant and Lei Ying. 2019. Finite-time error bounds for linear stochastic approximation and TD learning. In COLT’19. PMLR, 2803–2830. doi:10. 48550/arXiv.1902.00923 [33]Yashar Talebirad and Amirhossein Nadiri. 2023. Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents.arXiv:2306.03314 [cs.AI] doi:10.48550/arXiv.2306.03314 [34]Vinzenz Thoma, Vitor Bosshard, and Sven Seuken. 2025. Computing Perfect Bayesian Equilibria in Sequential Auctions with Verification. In AAAI’25. AAAI Press, 14158–14166. doi:10.1609/aaai.v38i13.29322 [35] Petter Tornberg. 2024. Simulating social media using Large Language Models to evaluate alternative news feed algorithms. Nature Human Behaviour (2024), 1–18. doi:10.1038/s41562-024-01858-2 [36] Khanh-Tung Tran, Dung Dao, et al.2025. Multi-Agent Collaboration Mechanisms: A Survey of LLMs. arXiv:2501.06322 [cs.CL] doi:10.48550/arXiv.2501.06322 [37] Miles Turpin, Julian Michael, Ethan Perez, and Samuel R Bowman. 2023. Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of- Thought Prompting. In NeurIPS’23, Vol. 36. Curran Associates, Inc., 74952–74965. doi:10.48550/arXiv.2305.04388 [38]Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen. 2024. A survey on Large Language Model based autonomous agents. Frontiers Comput. Sci. 18, 6 (2024), 186345. doi:10.1007/s11704-024-40231-1 [39]Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Shaokun Zhang, Erkang Zhu, Beibin Li, Li Jiang, Xiaoyun Zhang, and Chi Wang. 2023. AutoGen: En- abling Next-Gen LLM Applications via Multi-Agent Conversation Framework. arXiv:2308.08155 [cs.AI] doi:10.48550/arXiv.2308.08155 [40]Andrea Wynn, Harsh Satija, and Gillian Hadfield. 2025. Talk Isn’t Always Cheap: Understanding Failure Modes in Multi-Agent Debate. arXiv:2509.05396 [cs.AI] doi:10.48550/arXiv.2509.05396 [41]Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, Rui Zheng, Xiaoran Fan, Xiao Wang, Limao Xiong, Yuhao Zhou, Weiran Wang, Changhao Jiang, Yicheng Zou, Xiangyang Liu, Zhangyue Yin, Shihan Dou, Rongxiang Weng, Wenjuan Qin, Yongyan Zheng, Xipeng Qiu, Xuanjing Huang, Qi Zhang, and Tao Gui. 2025. The rise and potential of Large Language Model based agents: a survey. Sci. China Inf. Sci. 68, 2 (2025). doi:10.1007/S11432-024-4222-0 W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang [42]An Yang et al.2025. Qwen3 Technical Report. arXiv:2505.09388 [cs.CL] doi:10. 48550/arXiv.2505.09388 [43]Guisong Yang, Jiacai Li, Xingyu He, Fanglei Sun, and Yunhuai Liu. 2025. Hybrid Coopetitive Mechanism for Multiplatform Mobile Crowdsensing: A Two-Stage Approach to Pricing and Matching. IEEE Internet of Things Journal 12, 24 (2025), 54652–54663. doi:10.1109/JIOT.2025.3621450 [44] Xie Yi, Zhanke Zhou, Chentao Cao, et al.2025. From Debate to Equilib- rium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium. arXiv:2506.08292 [cs.AI] doi:10.48550/arXiv.2506.08292 [45]An Zhang, Leheng Sheng, Yuxin Chen, Tun Lu, Xiangyu Zhao, Peng Yong, and Hongzhi Yin. 2024. On Generative Agents in Recommendation Systems: A Survey and Perspective. In W’24. ACM. doi:10.1145/3589335.3651478 [46] Hanrui Zhang and Vincent Conitzer. 2021. Automated Dynamic Mechanism Design. In NeurIPS’21. Curran Associates, Inc., 27785–27797. doi:10.48550/arXiv. 2106.04689 [47]Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. 2024. Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View. In ACL’24. Association for Computational Linguistics, 14544– 14607. doi:10.18653/v1/2024.acl-long.794 [48]Kaiqing Zhang, Sham M. Kakade, Tamer Basar, and Lin F. Yang. 2023. Model- Based Multi-Agent RL in Zero-Sum Markov Games with Near-Optimal Sample Complexity. Journal of Machine Learning Research 24 (2023), 175:1–175:53. doi:10. 5555/3648699.3648874 [49]Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Personalizing Dialogue Agents: I have a dog, do you have pets too?. In ACL’18. Association for Computational Linguistics, 2204–2213. doi:10.18653/v1/P18-1205 [50] Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018. Generating Informative and Diverse Conversational Responses via Adversarial Information Maximization. In NeurIPS’18. Curran Associates, Inc., 1815–1825. doi:10.48550/arXiv.1809.05972 [51] Yusen Zhang, Ruoxi Sun, Yanfei Chen, Tomas Pfister, Rui Zhang, and Sercan Ö. Arik. 2024. Chain of Agents: Large Language Models Collaborating on Long- Context Tasks. In NeurIPS’24. Curran Associates, Inc. doi:10.48550/arXiv.2406. 02818 [52]Stephan Zheng, Alexander Trott, Sunil Srinivasan, Nikhil Naik, Melvin Gruesbeck, David C. Parkes, and Richard Socher. 2022. The AI Economist: Taxation policy design via two-level deep multiagent reinforcement learning. Science Advances 8, 18 (2022), eabk2607. doi:10.1126/sciadv.abk2607 [53]Yinghao Zhu, Ziyi He, Haoran Hu, et al.2025. MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks. arXiv:2505.12371 [cs.CL] doi:10.48550/arXiv.2505.12371 [54]Caleb Ziems, William Held, Omar Shaikh, Jindong Chen, Zhehao Zhang, and Diyi Yang. 2024. Can Large Language Models Transform Computational Social Science? Computational Linguistics 50, 1 (2024), 237–291. doi:10.1162/coli_a_ 00492 A Dataset Details and Statistics To substantiate the empirical results presented in Section 5 and facilitate reproducibility, we provide a granular description of the three evaluation datasets. These datasets were carefully curated to represent the distinct interaction topologies—adversarial, mixed, and open-ended—central to the BEACOF framework. A high-level statistical summary is provided in Table 3. A.1 Court Debate Dataset (Adversarial) The Court Debate dataset serves as a testbed for adversarial rea- soning, comprising 100 criminal cases: 80 from the AgentsCourt benchmark and 20 complex cases from China Judgements Online (CJO). Unlike standard datasets, each entry captures a complete judicial reasoning cycle. Formally, a case is modeled as a tupleD= (퐹,퐶 푃 ,퐶 퐷 ,I,푌). Here,퐹details the factual narrative;퐶 푃 and퐶 퐷 denote the initial claims of the Plaintiff and Defendant, respectively;Ienumerates key adjudication issues (e.g., surrender claims); and푌represents the ground truth verdict, including sentencing details. As detailed in Table 4, the dataset covers diverse domains to ensure robustness, spanning traffic offenses (e.g., Dangerous Driving, 28%), property crimes (e.g., Theft, 15%), and violent crimes (Intentional Injury, 10%). This variety necessitates agent adaptability across distinct statutory contexts. Table 4: Distribution of Case Causes in the Court Debate Dataset. The dataset covers a wide spectrum of criminal of- fenses to test adversarial robustness. Case CauseCount Case CauseCount Dangerous Driving28Traffic Casualty12 Theft15Smuggling8 Fraud12Illegal Business/Gambling7 Intentional Injury10Others (e.g., Assault)8 Total100 A.2 MedQA Dataset (Mixed) To evaluate professional consensus-building, we utilize the MedQA dataset, specifically sampling 200 questions from the MedQA-USMLE test set. This scenario embodies a mixed cooperative-competitive dynamic where agents must collaborate to diagnose but compete to eliminate incorrect distractors. The data structure for each instance consists of a complex clinical vignette, a set of five options (labeled A through E), and the ground- truth answer. The vignettes typically describe a patient’s medical history, symptoms, and vital signs, requiring multi-hop reasoning to link physiological observations with pathological causes. The challenge lies in the high density of domain-specific terminology and the presence of plausible distractors, which compels agents to critically evaluate peer proposals rather than blindly accepting a consensus. To mitigate potential bias from answer position, we verified the distribution of the ground-truth labels. As shown in Table 5, the correct answers are fairly uniformly distributed across the five options, ensuring that the agents’ performance reflects genuine medical reasoning rather than statistical artifacts. Table 5: Distribution of Correct Answer Keys in the MedQA Dataset. The balanced distribution prevents agents from ex- ploiting positional bias. Answer Key Count Percentage (%) A4321.5 B4020.0 C4020.0 D4924.5 E2814.0 Total200100.0 A.3 Persona Chat Dataset (Open-Ended) For the open-ended social interaction scenario, we curated 100 interaction pairs from the PersonaChat dataset to assess long-term consistency and lexical diversity. Unlike goal-oriented tasks, the objective here is maintaining a coherent identity. Belief-Driven Multi-Agent Collaboration via Approximate Perfect Bayesian Equilibrium for Social SimulationWWW ’26, April 13–17, 2026, Dubai, United Arab Emirates Table 3: Statistical summary of the datasets used in the evaluation scenarios. The “Avg. Input Length” denotes the average token count of the initial context provided to the agents. DatasetScenario TypeSourceSamplesKey ModalityAvg. Input Length Court DebateAdversarialAgentsCourt [15] & CJO100Unstructured Legal Text ∼550 tokens MedQAMixed (Coop-Comp)MedQA-USMLE [19]200Clinical Vignettes (5-Option) ∼180 tokens Persona ChatOpen-EndedPersonaChat [49]100Dialogue History∼320 tokens Each sample is defined by two sets ofuser_personas, which are lists of 3 to 5 sentences establishing the agent’s background character (e.g., "I work as a stunt double," "I own a poodle"). Ad- ditionally, autteranceshistory is provided to prime the context. The evaluation metrics for this dataset are specifically designed to measure Persona Consistency, using Natural Language Inference (NLI) to detect contradictions between generated responses and the assigned persona profile, and Response Diversity, quantified via Distinct-N metrics. A critical factor in social simulation is the depth of interaction. To demonstrate the complexity of the curated conversations, we analyze the distribution of dialogue lengths (measured in turns) in Table 6. The majority of dialogues (67%) exceed 20 turns, presenting a significant challenge for agents to maintain persona consistency over long contexts without succumbing to repetitive patterns. Table 6: Distribution of Dialogue Lengths (Turns) in the Per- sona Chat Dataset. Long-context interactions challenge the agents’ ability to maintain consistency. Dialogue Length (Turns) Count Percentage (%) 11 – 202222.0 21 – 304949.0 31 – 401818.0 > 401111.0 Total100100.0 B Baselines To assess the effectiveness of belief-driven adaptation, we compare BEACOF against three representative frameworks that embody distinct positions on the interaction spectrum. First, representing Static Cooperation, we adopt CAMEL [21]. This framework utilizes a role-playing inception prompting mechanism to facilitate task decomposition, serving as a strong baseline for scenarios requiring high coherence but often prone to sycophancy. Second, for Static Competition, we employ MAD [24], which relies on "tit-for-tat" argumentation to surface errors. Specifically, in the Court Debate scenario, we utilize the AgentsCourt [15] architecture, a domain- specialized implementation of MAD that strictly enforces the adver- sarial procedural logic of judicial systems. Additionally, we incorpo- rate ReConcile [4] to represent Static Consensus-Building. Unlike pure cooperation, ReConcile employs a round-table voting mecha- nism with confidence refinement to aggregate diverse viewpoints. We employ ReConcile exclusively for the MedQA scenario. It is excluded from the Court Debate because its consensus-seeking ob- jective fundamentally conflicts with the zero-sum nature of judicial proceedings, and from Persona Chat where open-ended diversity supersedes convergent solution-finding. To ensure a fair compari- son and eliminate architectural bias, all methods utilize identical backbone LLMs, system prompts, and decoding parameters. C Prompt Engineering Strategy To ensure the reproducibility of the BEACOF framework across di- verse domains while maintaining a unified methodological presenta- tion, we abstract the prompt engineering into a generalized schema. The framework operates on a dual-layer prompt architecture that separates the strategic coordination logic of the Meta-Agent from the execution logic of the Participant Agents. In implementation, domain-specific placeholders (e.g.,[DOMAIN_CONTEXT]) are instan- tiated with dataset-specific content (e.g., legal statutes for Court Debate, clinical vignettes for MedQA, or persona profiles for Per- sona Chat) at runtime. C.1 Meta-Agent Coordination Template The Meta-Agent functions as the mechanism designer, responsible for two critical tasks: (1) estimating the contextual payoff matrix to guide strategic equilibrium, and (2) evaluating participant outputs to update belief states. To support cross-scenario applicability, the prompt is structured to accept a variable set of evaluation dimen- sionsD=푑 1 , . . .,푑 푛 tailored to the specific task (e.g., Evidence Strength for adversarial tasks or Empathy for social tasks). The unified template is formalized as follows: System Instruction: You are the Meta-Agent Coordinator overseeing a [SCENARIO_TYPE] interaction. Your objective is to maintain the strategic equilibrium of the conversation. Contextual Input: - Global Context: [DOMAIN_KNOWLEDGE_BASE] - Interaction History: [DIALOGUE_HISTORY] Task 1 (Payoff Estimation): Analyze the current state and estimate the potential utility for each participant if they adopt one of the following strategies: Cooperation, Competition, or Coopetition. Assign a scalar value 푢 ∈ [ 0,10] to each strategy-agent pair. Task 2 (Evaluation): Assess the latest message based on the following domain-specific dimensions: [DIMENSION_LIST]. For each dimension, provide a normalized score 푠 ∈ [ 0,1] and a confidence score 휔 ∈ [0, 1]. Output Requirement: Return the results strictly in a structured JSON format containing keys for "payoff_matrix" and "belief_update_vector". W ’26, April 13–17, 2026, Dubai, United Arab EmiratesWeiwei Fang, Lin Li, Kaize Shi, Yu Yang, and Jianwei Zhang C.2 Participant Agent Strategic Template Participant agents are designed to act as rational players within the incomplete information game. Unlike standard role-playing prompts that rely solely on persona descriptions, our template dynamically injects the game-theoretic signals computed by the PBE mechanism. This injection ensures that the agent’s generative process is conditioned not only on its static role but also on the evolving belief states and strategic predictions. The generalized prompt template is defined below: Role Definition: You are [AGENT_ROLE], characterized by [PRIVATE_PROFILE]. Game State Injection: - Current Beliefs: Your subjective assessment of peers’ capabilities is [BELIEF_STATE]. - Strategic Signal: The estimated payoffs for your potential actions are [PAYOFF_MATRIX]. The predicted strategies of your opponents are [ACTION_ PREDICTION]. Action Directive: Based on the above information, you have resolved to adopt a [SELECTED_STRATEGY] approach. - If Cooperation: Focus on information synthesis and consensus-building. - If Competition: Focus on critical argumentation and error exposure. - If Coopetition: Balance partial agreement with strategic rebuttal. Task: Generate your response to [CURRENT_QUERY] ensuring alignment with your selected strategy and private profile. By utilizing these templates, the framework standardizes the interaction flow across the Court Debate, MedQA, and Persona Chat scenarios, with the only variation being the semantic content of the bracketed placeholders.