Paper deep dive
Opinion-Guided Layered Strategies for Decentralized Coordination
Shuhao Qi, Zhiyong Sun, Siep Weiland, Sofie Haesaert
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/25/2026, 6:14:05 AM
Summary
This paper introduces 'opinion-guided strategies' to solve the decentralized coordination challenge where autonomous agents must coordinate without prior agreement or communication. By leveraging nonlinear opinion dynamics (NOD) within a layered framework, agents can keep multiple admissible joint behaviors available and select one at runtime based on the other agent's evolving behavior. This approach ensures robust coordination, symmetry breaking for identical agents, and convergence to a Nash equilibrium in general-sum games without fixed rules or central coordination.
Entities (7)
Relation Signals (6)
Opinion-Guided Strategy → uses → Nonlinear Opinion Dynamics
confidence 97% · nonlinear opinion dynamics are leveraged in a layered realization to guide the agent... We therefore propose a new form of strategy, the opinion-guided strategy
Opinion-Guided Strategy → solves → Decentralized Coordination
confidence 96% · To address the decentralized coordination challenge, we propose a new form of strategy... the opinion-guided strategy coordinates with every randomly encountered agent
Opinion-Guided Strategy → enables → Symmetry Breaking
confidence 95% · two agents running identical strategies can break symmetry when needed, a capability that conventional strategies lack.
Opinion-Guided Strategy → appliesto → General-Sum Game
confidence 94% · One of them corresponds to a general-sum game... the opinion-guided strategy keeps every equilibrium open and guarantees the agents reach one
Nonlinear Opinion Dynamics → provides → Symmetry Breaking
confidence 93% · this bifurcation property enables agents to break symmetry and rapidly converge to a collaborative decision
Opinion-Guided Strategy → guarantees → Nash Equilibrium
confidence 92% · the opinion-guided strategy keeps every equilibrium open and guarantees the agents reach one, decided by their runtime interaction.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Autonomous agents increasingly interact with other independent agents, and such interactions typically admit multiple joint behaviors. When two agents prefer different ones, their independent strategies may be mutually incompatible and fail to reach a coordinated outcome; when they are identical, neither can differentiate its role when needed. Ideally, an agent should coordinate with any agent it encounters, regardless of which admissible joint behavior that agent aims to realize. We therefore propose a new form of strategy, the opinion-guided strategy, which keeps all the admissible joint behaviors available and postpones the selection to execution time, when the other agent's behavior reveals which one to realize. To realize this, nonlinear opinion dynamics are leveraged in a layered realization to guide the agent to a common admissible joint behavior in response to the other agent's evolving behavior, even without communication. We formally establish the conditions under which the strategy remains robust to every preference the other agent may hold. This robustness has an important implication: two agents running identical strategies can break symmetry when needed, a capability that conventional strategies lack. Three case studies across different applications show that the opinion-guided strategy coordinates with every randomly encountered agent, as long as it is willing to realize one of the admissible joint behaviors. One of them corresponds to a general-sum game: unlike conventional approaches devoted to finding a unique Nash equilibrium in advance, the opinion-guided strategy keeps every equilibrium open and guarantees the agents reach one, decided by their runtime interaction.
Tags
Links
- Source: https://arxiv.org/abs/2608.22104v1
- Canonical: https://arxiv.org/abs/2608.22104v1
Trouble viewing inline? Open PDF directly →
Full Text
140,554 characters extracted from source content.
Expand or collapse full text
Opinion-Guided Layered Strategies for Decentralized Coordination Shuhao Qi Affiliation: S. Qi, S. Weiland, and S. Haesaert are with the Department of Electrical Engineering, Eindhoven University of Technology, Eindhoven, The Netherlands. s.qi, s.weiland, s.haesaert@tue.nl Zhiyong Sun Affiliation: Z. Sun is with the College of Engineering, Peking University, Beijing, China. zhiyong.sun@pku.edu.cn Siep Weiland Affiliation: S. Qi, S. Weiland, and S. Haesaert are with the Department of Electrical Engineering, Eindhoven University of Technology, Eindhoven, The Netherlands. s.qi, s.weiland, s.haesaert@tue.nl Sofie Haesaert Thanks: This work is supported by the Horizon Europe project AIGGREGATE under grant No.˜101202457, the European project COVER under grant No.˜101086228, and the European project SymAware under grant No.˜101070802. Affiliation: S. Qi, S. Weiland, and S. Haesaert are with the Department of Electrical Engineering, Eindhoven University of Technology, Eindhoven, The Netherlands. s.qi, s.weiland, s.haesaert@tue.nl Abstract Autonomous agents increasingly interact with other independent agents, and such interactions typically admit multiple joint behaviors. When two agents prefer different ones, their independent strategies may be mutually incompatible and fail to reach a coordinated outcome; when they are identical, neither can differentiate its role when needed. Ideally, an agent should coordinate with any agent it encounters, regardless of which admissible joint behavior that agent aims to realize. We therefore propose a new form of strategy, the opinion-guided strategy, which keeps all the admissible joint behaviors available and postpones the selection to execution time, when the other agent’s behavior reveals which one to realize. To realize this, nonlinear opinion dynamics are leveraged in a layered realization to guide the agent to a common admissible joint behavior in response to the other agent’s evolving behavior, even without communication. We formally establish the conditions under which the strategy remains robust to every preference the other agent may hold. This robustness has an important implication: two agents running identical strategies can break symmetry when needed, a capability that conventional strategies lack. Three case studies across different applications show that the opinion-guided strategy coordinates with every randomly encountered agent, as long as it is willing to realize one of the admissible joint behaviors. One of them corresponds to a general-sum game: unlike conventional approaches devoted to finding a unique Nash equilibrium in advance, the opinion-guided strategy keeps every equilibrium open and guarantees the agents reach one, decided by their runtime interaction. I Introduction As a growing number of autonomous agents are deployed in the real world, each agent acts to accomplish its own tasks and satisfy its safety requirements, while interacting with other independent agents whose strategies are unknown. However, simultaneously meeting these objectives for all interacting agents remains an open problem. An increasing number of pathological coordination failures between autonomous vehicles (AVs) illustrate this challenge: as recently reported, multiple AVs from the same company have been observed becoming stuck and causing deadlocks in narrow, unsignalized intersections [28] and in parking lots [44]. To understand the underlying cause, consider the scenario in Fig. 1, where two vehicles encounter one another in a narrow corridor and must swap positions to proceed. Such situations are common in constrained environments such as dense residential areas. Intuitively, two joint behaviors can resolve the deadlock: (i) one AV pulls into the waiting lane while the other proceeds first, and (i) the roles are reversed. Although the two joint behaviors are symmetric, each requires the AVs to play different roles. If the AVs agree on one of them, the conflict is resolved quickly; in practice, however, different vehicles may independently commit to incompatible choices. Worse still, the reports in [28, 44] note that the participating AVs were from the same company and thus equipped with identical strategies, leaving them structurally unable to differentiate their roles in such symmetric situations. We refer to the above as the decentralized coordination challenge: given a set of admissible joint behaviors, interacting agents must realize one of them without knowing the choice of the other agent. The central problem is that, even when every agent is rational and committed to realizing an admissible joint behavior, their independent strategies may be mutually incompatible. This incompatibility takes two forms: (i) different agents may disagree on which joint behavior is preferable; and (i) identical agents necessarily select the same action in symmetric situations, which cannot realize a joint behavior that requires distinct roles. In practice, an autonomous agent may encounter either kind of opponent, and cannot know in advance which. Ideally, an agent should be able to coordinate with any agent it encounters, as long as that agent is likewise rational and committed to one of the admissible joint behaviors. Fig. 1: A typical encounter scenario in dense residential regions. I-A Related Works A well-known instance of this decentralized coordination challenge is the general-sum game, which may admit multiple equilibria [5, 8]. This multiplicity has prompted much discussion in game theory and multi-agent reinforcement learning (MARL) over how agents should select among coexisting equilibria. When multiple equilibria coexist, decentralized best-response and learning dynamics may fail to converge, cycling among strategy profiles rather than settling on a single equilibrium [49, 52]. To address this, classical game theory introduces selection criteria such as risk dominance and payoff dominance [20], along with equilibrium refinements such as trembling-hand perfection [42, 13], while MARL research designs learning dynamics that steer convergence toward a desired equilibrium [50, 11]. In more practical settings, both the game-theoretic and MARL approaches sidestep this challenge and try to select a unique equilibrium in advance by injecting asymmetry, either explicitly, through predefined roles and priorities as in Stackelberg games [48, 14], level-k reasoning [34], role-oriented learning [46], priority scheduling [25], and ID-conditioned policies [17]; or implicitly, through jointly trained policies that fix coordination during training. Specifically, the commonly used centralized-training-with-decentralized-execution (CTDE) paradigm [35, 36] settles on a joint behavior at training, which each agent then follows independently at execution. In summary, these explicit and implicit approaches fix the joint outcome at design or training time. Since agents encounter one another at random in practice, as autonomous vehicles do, such predefined asymmetry lacks the flexibility to handle an opponent produced by a different design or training procedure. In addition, the approach of injecting asymmetry overlooks the situation in which the encountered agent runs the identical strategy, where both agents would, for instance, claim priority over the other. As in game-theoretic and MARL research, most prior work in multi-agent planning and control also sidesteps the decentralized coordination challenge mentioned above. The existing conflict-resolution methods can be broadly grouped into three categories. (i) Fixed rules and priorities, such as right-hand-priority schemes [38], can be encoded directly into the controller, but they fail when another agent does not follow the same rules and priorities. (i) Perturbation-based methods inject disturbances to break strict conflict conditions such as deadlock [45]; although practical, they do not guarantee resolution and may violate safety constraints. (i) Communication-based coordination resolves deadlocks through explicit information exchange. Centralized schemes employ a coordinator that jointly manages agent interactions [10]; while effective in small systems, they scale poorly as the number of agents grows [43, 39]. By contrast, distributed schemes rely on local communication among neighbors to negotiate passage, using techniques such as parametric control Lyapunov functions [47] and rotation controllers for position swaps [18, 2]. In both cases, their reliance on communication makes them vulnerable to time delays, protocol ambiguity, and disturbances, particularly across heterogeneous agents [40]. In contrast, human drivers rarely fix on a single plan ahead of time; instead, they keep several possible behaviors open and settle on one only as cues from the ongoing interaction make the right choice clear. This adaptive mechanism is precisely what existing autonomous systems are missing. Drawing on insights from social and biological sciences, the nonlinear opinion dynamics (NOD) framework can model interactions among agents [6, 33]: each agent maintains a continuous opinion state representing its degree of agreement with a decision, updated in response to others’ behavior. Moreover, NOD undergoes a supercritical pitchfork bifurcation at a critical attention level, and this bifurcation property enables agents to break symmetry and rapidly converge to a collaborative decision [6, 33]. Crucially, this property targets the coordination challenge identified above. NOD has accordingly been used across multi-agent interaction tasks, such as resolving deadlocks [9], generating adaptive priorities in narrow corridors [1], and stabilizing repeated and differential games [37, 22]. In our prior work, we integrated NOD into a safety filter to resolve blocking behaviors in a two-airplane encounter [41], without relying on communication or fixed rules. However, NOD is continuous in nature, so existing works have primarily deployed it within a continuous-time controller over a continuous state space. In addition, these integrations are hand-tailored by experts for each specific task. Although the bifurcation property of NOD holds promise for the decentralized coordination challenge, coupling it to a continuous controller in this task-specific manner limits its scalability to general decision-making problems of multi-agent coordination. I-B Contributions To address the decentralized coordination challenge, we propose a new form of strategy that lifts the core properties of NOD to the decision-making level. However, such lifting is non-trivial: it requires the resulting framework to preserve the bifurcation and symmetry-breaking properties of NOD while also overcoming the decentralized coordination challenge. In this work, our contributions are fourfold. First, we formalize the decentralized two-agent coordination problem by defining a set-valued permissive joint strategy and a rational individual strategy, together with the two requirements a strategy must satisfy under independent execution, robust coordination and homogeneous implementability (Section I). Second, we propose an opinion-guided strategy and analyze the conditions under which it guarantees robust coordination against different types of rational opponents, and homogeneous implementability, even in symmetric situations (Sections IV and V). Third, we propose a layered framework that realizes the opinion-guided strategy, comprising opinion dynamics, a continuous controller, and an opinion estimator, which together enable each agent to adapt in real time to the evolving behavior of others and converge to a common admissible joint behavior while operating fully decentralized, without inter-agent communication (Section VI-A). Fourth, we evaluate the proposed framework across three case studies, the last of which corresponds to a classic general-sum game (Section VI). I Background: The Coordination Challenge I-A Running Example: Charging-Bay Selection Fig. 2: Sketch of the running example, a two-robot charging-bay selection task that instantiates the anti-coordination game. The robots R1R_1 (blue) and R2R_2 (red) start together in the middle zone p2p_2, while a charging bay is available in each of the side zones p1p_1 and p3p_3. The dashed arrows mark the two candidate moves of each robot (left toward p1p_1 or right toward p3p_3), and the task succeeds only when the two robots commit to distinct bays, so that both can recharge simultaneously. To make the decentralized coordination challenges concrete, we begin with a simple two-robot example that instantiates the classical anti-coordination game, of which the “chicken game” is a well-known instance [7]. In such a game, the agents are expected to take distinct actions to reach a desired outcome. For simplicity, we consider a single-stage game, so that each agent makes only one decision. As shown in Fig. 2, the two robots, denoted by R1R_1 and R2R_2, start at the middle zone p2p_2 and must be dispatched simultaneously to the two charging bays located in the target zones p1p_1 and p3p_3, with one robot assigned to each. Since each bay can serve only one robot at a time, and each robot selects its target independently and without knowledge of the other’s choice, the task succeeds only when the two robots commit to different zones. Two admissible joint behaviors thus arise: R1R_1 moves to p1p_1 while R2R_2 moves to p3p_3, or vice versa. Since R1R_1 and R2R_2 may in reality be either heterogeneous or homogeneous, this minimal setting exposes two failure modes of decentralized coordination: 1. Miscoordination risk: The robots complete the charging task only when they independently select compatible strategies (R1R_1 to p1p_1, R2R_2 to p3p_3), yet they may just as well select incompatible ones (both to p1p_1 or both to p3p_3), and such compatibility can never be ensured. 2. Symmetry-breaking failure: If the robots are homogeneous, whatever strategy they adopt, they make identical decisions concurrently in this symmetric scenario (both to p1p_1 or both to p3p_3) and thus can never commit to the distinct zones the task requires. I-B A Promising Approach: Nonlinear Opinion Dynamics (a) Clockwise swap (b) Anti-clockwise swap Fig. 3: Two resolution behaviors randomly generated in a two-airplane encounter. The upper panels show the airplanes’ trajectories, while the bottom panels depict the evolution of the opinion states. In our prior work [41], we proposed an approach based on NOD to mitigate the two coordination challenges described above at the continuous-control layer. This approach does not require predefined rules, fixed priorities, or inter-agent communication. Intuitively, an agent’s opinion state represents its degree of agreement with a decision. The bifurcation property of NOD [33] makes each agent highly sensitive to small changes in the other’s state, enabling the group to break symmetry when needed and rapidly converge to a collaborative decision. As shown in Fig. 3, two airplanes are tasked to swap their positions, for which two symmetric joint behaviors are admissible: bypassing on the left or bypassing on the right. Each airplane infers the other’s opinion state, which encodes its intent of the bypassing side, and updates its own opinion based on local observations. As a result, if one airplane commits to one bypassing side, the other can adapt to it and reach consensus on that side. If both airplanes are free to choose, the final outcome has two possibilities: clockwise or anticlockwise swapping (Fig. 3). Which one is realized is determined by the real-time evolution of the opinion states. Therefore, the NOD-based controller can resolve both miscoordination and symmetry-breaking failures in a fully decentralized manner. This paper builds on this insight by lifting the core properties of NOD from the continuous-control layer to the discrete decision-making layer, enabling more general strategic coordination. I Problem Formulation and Statement This section formalizes the decentralized coordination problem. To examine this challenge in depth, we focus on two-agent interactions, an important step before considering interactions among more agents. I-A Game Model and Strategies Definition 1 (Two-Agent Game). A two-agent game is defined as a tuple =(,s0,A1,A2,→),G=(S,s_0,A_1,A_2,→), where: • S is the set of joint states, with initial state s0∈s_0 ; • A1A_1 is the set of actions of Agent 1; • A2A_2 is the set of actions of Agent 2; and • →⊆×(A1×A2)×→ \!×\!(A_1\!×\!A_2)\!×\!S is the transition relation that maps joint state and joint action to a successor state. An infinite sequence of joint actions (a1,0,a2,0),(a_1,0,a_2,0), (a1,1,a2,1),(a1,2,a2,2),…(a_1,1,a_2,1),(a_1,2,a_2,2),… induces an execution of G: s0→(a1,0,a2,0)s1→(a1,1,a2,1)s2→(a1,2,a2,2)s3…s_0 (a_1,0,a_2,0)s_1 (a_1,1,a_2,1)s_2 (a_1,2,a_2,2)s_3… where (st,(a1,t,a2,t),st+1)∈→(s_t,(a_1,t,a_2,t),s_t+1)∈→ for all t≥0t≥ 0. We define the set of all possible finite executions (i.e., histories), based on which decisions can be made, as ℋ:=(×A1×A2)∗×H:=(S× A_1× A_2) ×S, where X∗X denotes the Kleene star, i.e., the set of all finite sequences over a set X. An element h∈ℋh\! is a finite sequence of the form h=s0→(a1,0,a2,0)s1→(a1,1,a2,1)⋯→(a1,t−1,a2,t−1)sth=s_0 (a_1,0,a_2,0)s_1 (a_1,1,a_2,1)·s (a_1,t-1,a_2,t-1)s_t. Notation. The labels 1,21,2 denote the two specific agents. When describing the interaction between them, we use i∈1,2i∈\1,2\ to denote a generic agent and j∈1,2∖ij∈\1,2\ \i\ the other agent, so that i,j=1,2\i,j\=\1,2\. Individual strategy. For each agent i∈1,2i\!∈\!\1,2\, a deterministic individual strategy returns an individual action ai∈Aia_i\!∈\!A_i, denoted by σi _i. A commonly used individual strategy is a mapping from the state history to an action [21, 29], i.e., σi:ℋ→Ai _i\!:\!H\!→\!A_i. The strategy is memoryless (Markov) if it is a function σi:→Ai _i:S→ A_i, meaning that the chosen action depends only on the current state. Given two individual strategies σ1:ℋ→A1 _1:H→ A_1 and σ2:ℋ→A2 _2:H→ A_2, the joint strategy is the function σ1|σ2:ℋ→A1×A2 _1\,\|\, _2:H→ A_1× A_2, defined for each history h∈ℋh by (σ1∥σ2)(h):=(σ1(h),σ2(h)).( _1\,\|\, _2)(h):=\( _1(h), _2(h))\. (1) Permissive joint strategy. A strategy for a multi-agent system is conventionally defined as a single-valued mapping ℋ→A1×A2H\!→\!A_1\!×\!A_2, which is suitable for a centralized approach in which all agents are coordinated to follow one prescribed joint action at each history. As illustrated in the running example, however, a multi-agent decentralized system may admit multiple equally admissible joint actions. Prior work has captured this multiplicity through set-valued strategies, such as permissive strategies in reactive synthesis [12], permissive controllers for interval MDPs [3], and permissive equilibria in multi-player reachability games [8]. Following these definitions, we define a permissive joint strategy as a set-valued mapping from each history to a set of joint actions, given by π:ℋ→2A1×A2,π:H→ 2^A_1× A_2, where π(h)⊆A1×A2π(h) A_1× A_2 represents the set of joint actions permitted at history h∈ℋh . Throughout the paper, a joint action in π(h)π(h), and the joint behavior it induces, are called admissible. Rational actions. In practice, the challenge is that each agent should select an action without knowledge of the other agent’s choice, such that the resulting joint behavior may fail to be admissible, (σ1∥σ2)(h)⊈π(h)( _1\,\|\, _2)(h) π(h). In the extreme case where an agent behaves adversarially, coordination cannot be achieved. However, in realistic applications such as traffic scenarios, adversarial behavior is rare, as each vehicle pursues its own navigation objective rather than seeking to attack others. The literature on rational synthesis [15, 32] has shown that specifications unrealizable against a hostile environment can become realizable once agent rationality is taken into account. Given this motivation, we assume each agent is rational and committed to realizing an admissible joint behavior within the shared permissive strategy π, which is referred to as π-rationality. The rational action set and individual strategy under π-rationality are formalized below. Definition 2 (π-Rational Action Set and Strategy). Let π be a given permissive joint strategy for a two-agent game G defined in Def. 1. Given a history h and i∈1,2i∈\1,2\, an action ai∈Aia_i\!∈\!A_i of agent i is rational with respect to π at history h if there exists aj∈Aja_j∈ A_j such that (a1,a2)∈π(h)(a_1,a_2)\!∈\!π(h). The rational action set of agent i at history h with π(h)≠∅π(h)\!≠\! is defined as Ai(h)=ai∈Ai∣∃aj∈Aj:(a1,a2)∈π(h).A rat_i(h)=\a_i∈ A_i ∃\,a_j∈ A_j:(a_1,a_2)∈π(h)\. (2) An individual strategy σi _i is said to be rational with respect to π if σi(h)∈Ai(h) _i(h)∈ A rat_i(h) for all h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! . Intuitively, a rational individual strategy is one that realizes an admissible joint behavior under π, provided the other agent behaves cooperatively. We denote the set of all the rational strategies for agent i by Σi _i rat, where i∈1,2i∈\1,2\. We use the running example to illustrate these concepts. Running example: In the two-robot anti-coordination game, the state of each robot is its location, drawn from p1,p2,p3\p_1,p_2,p_3\, where p2p_2 is the initial zone and p1p_1, p3p_3 are the two target zones. The joint state set S consists of all possible combinations of the robots’ positions, with initial state s0=(p2,p2)s_0\!=\!(p_2,p_2). Each robot RiR_i selects an action from Ai=ap1,ap2,ap3A_i\!=\!\a_p_1,a_p_2,a_p_3\, where apka_p_k denotes movement toward zone pkp_k for k∈1,2,3k\!∈\!\1,2,3\; in particular, ap1a_p_1 and ap3a_p_3 move the robot toward the target zones p1p_1 and p3p_3 respectively, and ap2a_p_2 keeps the robot stationary at p2p_2. The objective is to reach a terminal state (p1,p3)(p_1,p_3) or (p3,p1)(p_3,p_1). A permissive joint strategy at the initial state is illustrated in Fig. 4, which includes two valid cooperative action pairs: π(s0)=(ap1,ap3),(ap3,ap1)π(s_0)=\(a_p_1,a_p_3),\ (a_p_3,a_p_1)\. The corresponding rational action set for each robot is Ai(s0)=ap1,ap3A rat_i(s_0)\!=\!\a_p_1,a_p_3\ for i∈1,2i\!∈\!\1,2\. Fig. 4: A permissive joint strategy for the running example at the initial state. The green regions represent the permissive joint action pairs π(s0)π(s_0). The dashed arrows show projections to the rational action sets. I-B Robust Coordination Fig. 5: Representative geometries of a permissive joint strategy, π1(h) _1(h) through π5(h) _5(h), illustrating cases in which both, one, or neither agent has a non-empty robust set. In every panel the horizontal axis is A1A_1 (Agent 1, blue) and the vertical axis A2A_2 (Agent 2, red). On the axes, thin light lines mark the rational sets Ai(h)A_i rat(h) and thick dark lines the robust sets Ai(h,Σj)A_i rob(h; _j rat), and the dotted gray box is their product A1(h)×A2(h)A_1 rat(h)\!×\!A_2 rat(h). The green region is the permissive set π(h)π(h); blue and red hatching mark the joint actions in which Agent 1 and Agent 2, respectively, commit to their robust set. When an agent faces another whose strategy is unknown, an ideal strategy is one that guarantees coordination regardless of the other agent’s choice. We refer to such a strategy as a robust strategy, which is formally defined below. Definition 3 ((π,Σj)(π, _j)-Robust Strategies). Given a permissive strategy π and a set of strategies Σj _j of agent j, we say that an individual strategy σi _i of agent i is (π,Σj)(π, _j)-robust if for all h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! : ∀σj∈Σj:(a1,a2)∈π(h),∀ _j\!∈\! _j:\;(a_1,a_2)\!∈\!π(h), where ai=σi(h),aj=σj(h)a_i\!=\! _i(h),\,a_j\!=\! _j(h), i,j∈1,2i,j\!∈\!\1,2\ and i≠ji\!≠\!j. At each history h, the actions that agent i can select to guarantee coordination regardless of the other agent’s choice are called robust actions, which are defined as follows. Definition 4 ((π,Σj)(π, _j)-Robust Action Set). Given a permissive strategy π and a set of strategies Σj _j of agent j, let AΣj(h):=σj(h)∣σj∈ΣjA _j(h)\!:=\!\ _j(h) _j\!∈\! _j\ denote the set of actions agent j may select at history h under a strategy in Σj _j. The robust action set of agent i at a history h with π(h)≠∅π(h)\!≠\! , denoted by Ai(h,Σj)A rob_i(h; _j), is defined as Ai(h;Σj):=ai∈Ai| A rob_i(h; _j)\!:=\! \a_i\!∈\!A_i\; | ∀aj∈AΣj(h): ∀ a_j\!∈\!A _j(h): (3) (a1,a2)∈π(h), (a_1,a_2)\!∈\!π(h) \, where i,j∈1,2i,j\!∈\!\1,2\ and i≠ji\!≠\!j. If agent i selects an action from a nonempty robust action set Ai(h,Σj)A rob_i(h; _j) at every history h, then coordination is guaranteed regardless of which strategy σj∈Σj _j\!∈\! _j agent j selects, providing a unilateral guarantee of coordination. In particular, a (π,Σj)(π, _j rat)-robust strategy guarantees the coordination prescribed by π against every rational strategy of the other agent. When Σj=Σj _j\!=\! _j rat, we abbreviate Ai(h):=Ai(h,Σj)A_i rob(h)\!:=\!A_i rob(h; _j rat). Figure 5 illustrates five representative geometries of the permissive action set π(h)π(h). Across all subfigures, the light-colored lines along the axes mark the rational action sets Ai(h)A_i rat(h) from Def. 2, and the dark-colored lines mark the robust action sets Ai(h)A_i rob(h) from Def. 4, i∈1,2i∈\1,2\. In Fig. 5(a), every pair of rational actions is permitted, so π(h)=A1(h)×A2(h)π(h)\!=\!A_1 rat(h)\!×\!A_2 rat(h) and Ai(h)=Ai(h)A_i rat(h)\!=\!A_i rob(h) for i∈1,2i\!∈\!\1,2\; coordination is guaranteed regardless of how the agents act. Fig. 5(b) shows a case in which not all pairs of A1(h)×A2(h)A_1 rat(h)\!×\!A_2 rat(h) lie in π(h)π(h), yet both robust action sets remain non-empty, so either agent can unilaterally guarantee coordination by restricting its choice to Ai(h)A_i rob(h). In Fig. 5(c), only A2(h)A_2 rob(h) is non-empty; agent 1 cannot unilaterally guarantee coordination, because no single a1∈A1(h)a_1\!∈\!A_1 rat(h) satisfies (a1,a2)∈π(h)(a_1,a_2)\!∈\!π(h) for every a2∈A2(h)a_2\!∈\!A_2 rat(h). In reality, the permissive action set may have more complex geometries, and it is possible that both robust action sets are empty, as illustrated in Fig. 5(d) and (e). In such cases, neither agent can unilaterally guarantee coordination, even if both agents are rational, and the two agents face the risk of miscoordination: each may select a rational action, yet their joint action can still fall outside the permissive set. It is therefore essential to mitigate this risk by ensuring the existence of robust actions. I-C Homogeneous Implementability Beyond the miscoordination risk in the heterogeneous settings, another challenge arises when the two agents are homogeneous [51], as exemplified by the running example. An ideal strategy should therefore not only be robust to the other rational agent, but also remain effective in homogeneous settings, a property we refer to as homogeneous implementability. Building on the notions of homogeneous multi-agent Markov decision processes introduced in [51], we adapt the definition to our two-agent setting. The underlying premise is that the labels “agent 1” and “agent 2” carry no physical meaning, so every property of the game should be invariant under permuting them. To formalize this, writing the joint state as =S1×S2S\!=\!S_1\!×\!S_2, we introduce a swap operator τ:S1×S2→S1×S2τ\!:S_1\!×\!S_2\!→\!S_1\!×\!S_2 as τ(s1,s2)=(s2,s1).τ(s_1,s_2)\!=\!(s_2,s_1). Definition 5 (Homogeneous Two-Agent Game). A two-agent game =(,s0,A1,A2,→)G=(S,s_0,A_1,A_2,→) as defined in Def. 1 with =S1×S2S=S_1× S_2, S1=S2S_1=S_2 and A1=A2A_1=A_2 is called homogeneous if for all s,s′∈s,s \!∈\!S and a1,a2∈A1a_1,a_2\!∈\!A_1, (s,(a1,a2),s′)∈→(s,(a_1,a_2),s )∈\;→ if and only if (τ(s),(a2,a1),τ(s′))∈→(τ(s),(a_2,a_1),τ(s ))∈\;→. Following [51, Def. 14], we say that two Markov strategies σ1 _1 and σ2 _2 are identical if the joint action they induce is equivariant under agent relabeling: (σ1(s),σ2(s))=τ(σ1(τ(s)),σ2(τ(s)))∀s∈.( _1(s), _2(s))\;=\;τ\! ( _1(τ(s)),\, _2(τ(s)) ) ∀ s . This condition is equivalent to the pointwise relation σ1(s)=σ2(τ(s)) _1(s)\!=\! _2(τ(s)) for all s. Intuitively, the strategies of two agents are identical when they take the same action from their own first-person view of every state, so that exchanging the labels “agent 1” and “agent 2” leaves the joint behavior unchanged. Similarly to τ, we define the permutation of a history, h=s0→(a1,0,a2,0)s1→(a1,1,a2,1)⋯→(a1,t−1,a2,t−1)sth=s_0 (a_1,0,a_2,0)s_1 (a_1,1,a_2,1)·s (a_1,t-1,a_2,t-1)s_t, as τ¯(h)=τ(s0)→(a2,0,a1,0)τ(s1)→(a2,1,a1,1)⋯→(a2,t−1,a1,t−1)τ(st). τ(h)=τ(s_0) (a_2,0,a_1,0)τ(s_1) (a_2,1,a_1,1)·s (a_2,t-1,a_1,t-1)τ(s_t). This allows us to define identical strategies σ1 _1 and σ2 _2 as those for which σ1(h)=σ2(τ¯(h))∀h∈ℋ. _1(h)= _2( τ(h)) ∀ h . (4) Definition 6 (π-Homogeneous Implementability). Let G be a homogeneous two-agent game (Def. 5) and π a permissive strategy. An individual strategy σ is π-homogeneously implementable if ∀h∈ℋ with π(h)≠∅:∅≠(σ∥σ∘τ¯)(h)⊆π(h),∀ h with π(h)≠ : \!≠\!(σ\,\|\,σ τ)(h) π(h), (5) where (σ∘τ¯)(h):=σ(τ¯(h))(σ τ)(h):=σ( τ(h)). Such identical strategies are known to exist when the two agents are never in the same state [51, Theorem 3]. However, for an individual strategy of the form σ:ℋ→A1σ:H\!→\!A_1, the symmetric condition with τ(s)=sτ(s)\!=\!s forces both agents to select the same action, so a π-homogeneously implementable strategy of this form may not exist. Running example: In the two-robot example, the initial state s0=(s0,1,s0,2)s_0=(s_0,1,s_0,2) satisfies s0,1=s0,2s_0,1\!=\!s_0,2, and π(s0)=(ap1,ap3),(ap3,ap1)π(s_0)=\(a_p_1,a_p_3),\ (a_p_3,a_p_1)\. Since τ(s0)=s0τ(s_0)=s_0, the homogeneity condition forces both agents to choose the same action, so the feasible joint actions are restricted to the diagonal (a,a):a∈A1\(a,a):a∈ A_1\, none of which lies in π(s0)π(s_0), as shown in Fig. 4. The symmetric situation with s1=s2s_1\!=\!s_2 captures those joint states at which the two agents are indistinguishable. In practice, such symmetry is not rare: two homogeneous agents in a symmetric situation perceive identical first-person observations and therefore act identically, as illustrated by the encounter scenario in Fig. 1. It is therefore essential that the strategy of an autonomous agent deployed in an interactive environment be homogeneously implementable. I-D Problem Statement Consider a two-agent game G (Def. 1) with a permissive joint strategy π shared by both agents as common knowledge. Each agent selects its action independently. Without loss of generality, take agent i as the ego agent whose individual strategy σi _i is to be designed, and let agent j be the encountered agent, which follows an unknown π-rational strategy σj∈Σj _j\!∈\! _j rat that may either differ from σi _i or be identical to it in the sense of (4). The objective is to design σi _i such that, in both cases, the realized joint action lies in π(h)π(h) at every history h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! . Formally, σi _i is required to (i) achieve (π,Σj)(π, _j rat)-robust coordination (Def. 3), and (i) be π-homogeneously implementable (Def. 6) whenever G is homogeneous (Def. 5). IV Opinion-Guided Strategies IV-A Reactive Strategy Gσ1(s) _1(s)σ2(s) _2(s)G…ssssa1a_1a2a_2s′s s′s (a) Independent implementation Gσ1(s) _1(s) σ2act(s,a1)σ^act_2(s,a_1)G…ssssa1a_1a2a_2a1a_1s′s s′s (b) Sequential implementation Gσ1act(s,a2)σ^act_1(s,a_2)σ2act(s,a1)σ^act_2(s,a_1)G…ssssa1a_1a2a_2a2a_2a1a_1s′s s′s (c) Mutually implementation Gσ1op(s,z~1)σ^op_1(s, z_1)σ2op(s,z~2)σ^op_2(s, z_2)G…ssssz~1 z_1z~2 z_2a1a_1a2a_2s′s s′s (d) Opinion-guided implementation Fig. 6: Four decision structures for a two-agent game. (a) agents act concurrently with independent strategies σ1,σ2 _1, _2. (b) a reactive strategy realized sequentially, where agent 2 observes a1a_1 before selecting a2a_2. (c) mutually reactive strategies, where each agent’s action conditions on the other’s; this is an abstract specification without a realization mechanism. (d) an opinion-guided implementation, where each agent maintains an opinion state z~i z_i paired to the other’s through Γi _i, and selects its action via σiop(s,z~i)σ^op_i(s, z_i). So far we have assumed that the two agents select their actions concurrently and independently, as depicted in panel (a) of Fig. 6. A common alternative is a sequential (or turn-based) decision structure, shown in panel (b), in which the agents act one after the other within each step. The second agent observes the first’s action and conditions its decision on the history h together with the observed action. The resulting strategy is called a reactive strategy and corresponds to the notion of a reaction function in the game-theoretic literature [4]11 1 The reactive strategy in this paper conditions on the other agent’s action, unlike reactive strategies in control and robotics, which condition on the state of the environment [26, 31].. Definition 7 (Reactive Strategy). Given a game G, a reactive strategy for agent i∈1,2i\!∈\!\1,2\ is a mapping σiact:ℋ×Aj→Aiσ^act_i:H× A_j→ A_i, where σiact(h,aj)=aiσ^act_i(h,a_j)\!=\!a_i. The set of all such strategies is denoted Σiact ^act_i. In contrast, a strategy σi:ℋ→Ai _i:H\!→\!A_i that does not condition on the other agent’s action is called a non-reactive strategy. If only one agent applies a reactive strategy while the other uses a non-reactive strategy, as in panel (b) of Fig. 6, the joint strategy is defined as (σ1∥σ2act)(h):=(σ1(h),σ2act(h,σ1(h))).( _1\,\|\,σ^act_2)(h)\!:=\!\( _1(h),σ^act_2(h, _1(h)))\. (6) Consider now the mutually reactive setting, in which both agents employ reactive strategies and each conditions on the other’s action, as depicted in panel (c) of Fig. 6. In this setting, the joint strategy is no longer single-valued but set-valued, defined as (σ1act∥σ2act):ℋ→2A1×A2(σ^act_1\,\|\,σ^act_2):H→ 2^A_1× A_2, whose value at a history h collects all action pairs in which each agent’s action is a reaction to the other’s, (σact1∥σact2)(h):=(a1,a2)|a1 (σ^act_1\,\|\,σ^act_2)(h):= \(a_1,a_2)\,|\,a_1 =σ1act(h,a2), =σ^act_1(h,\!a_2), (7) a2 a_2 =σact2(h,a1). =σ^act_2(h,\!a_1) \. Depending on σ1actσ^act_1 and σ2actσ^act_2, this set may be empty, a singleton, or contain several compatible action pairs. Non-reactive and reactive strategies are both individual strategies: each returns agent i’s own action ai∈Aia_i\!∈\!A_i, in contrast to a joint strategy that returns an action pair for both agents. For any two individual strategies, the joint outcome (σi∥σj)(h)⊆A1×A2( _i\,\|\, _j)(h)\! \!A_1\!×\!A_2 collects all joint action pairs realized by the two agents. The π-homogeneous implementability of Def. 6 is understood in this general sense. Running example: Recall that at the symmetric initial state s0=(p2,p2)s_0\!=\!(p_2,p_2), the permissive joint strategy is π(s0)=(ap1,ap3),(ap3,ap1)π(s_0)\!=\!\(a_p_1,a_p_3),(a_p_3,a_p_1)\. A reactive strategy can be designed as σact(s0,aj)=ap1if aj=ap3,ap3if aj=ap1,~σ^act(s_0,a_j)= casesa_p_1&if a_j=a_p_3,\\ a_p_3&if a_j=a_p_1, cases (8) where aja_j denotes the action of the other agent. If one agent applies the reactive strategy while the other applies an arbitrary rational strategy σj _j, then σj(s0)∈Aj(s0)=ap1,ap3 _j(s_0)\!∈\!A_j rat(s_0)\!=\!\a_p_1,a_p_3\, and the reactive response yields the pair (ap3,ap1)(a_p_3,a_p_1) when σj(s0)=ap1 _j(s_0)\!=\!a_p_1 and (ap1,ap3)(a_p_1,a_p_3) when σj(s0)=ap3 _j(s_0)\!=\!a_p_3, both of which lie in π(s0)π(s_0). Thus, the reactive strategy can achieve robust coordination. When both agents apply this reactive strategy, the resulting set of action pairs is (σact∥σact∘τ¯)(s0)=(ap3,ap1),(ap1,ap3)=π(s0)(σ^act\,\|\,σ^act τ)(s_0)\!=\!\(a_p_3,a_p_1),(a_p_1,a_p_3)\\!=\!π(s_0): each agent reacts to the other, so the only mutually consistent outcomes are those in which one plays ap1a_p_1 and the other plays ap3a_p_3. This bypasses the issue caused by symmetry noted earlier, where no homogeneous individual strategy could produce any pair in π(s0)π(s_0). The above example illustrates that reactive strategies can deliver both robust coordination and homogeneous implementation. However, reactive strategies are built on the assumption that each agent can observe the other’s intended action before execution, which is unrealistic. The remaining question is how the two agents converge on a common pair in (σ1act∥σ2act)(h)(σ^act_1\,\|\,σ^act_2)(h) without knowing each other’s intended actions in advance. Inspired by the properties of opinion dynamics introduced in Section I-B, we propose to realize reactive strategies using opinion dynamics. IV-B Opinion-Guided Reactive Strategies We illustrate this opinion-guided implementation, depicted in panel (d) of Fig. 6, on the running example. Running example: at the initial state s0s_0, we equip each agent with a scalar opinion zi∈ℝz_i\!∈\!R evolving under the following nonlinear opinion dynamics: z˙i=−3zi+3tanh(3(zi−zj)), z_i=-3z_i+3 \! (3(z_i-z_j) ), (9) where i,j∈1,2i,j\!∈\!\1,2\ and i≠ji\!≠\!j. The individual strategies of both agents follow the same rule σop(s0,zi)=ap1if zi≥0,ap3if zi<0,σ^op(s_0,z_i)= casesa_p_1&if z_i≥ 0,\\ a_p_3&if z_i<0, cases (10) which is homogeneous by construction. The discrete decision is made after a reserved interval Δtop>0 t_op\!>\!0, during which the opinion dynamics (9) are expected to converge. By [6, Cor. IV.1.2], the neutral equilibrium (z1,z2)=(0,0)(z_1,z_2)\!=\!(0,0) is unstable, and two stable equilibria emerge at (z∗,−z∗)(z^*,-z^*) and (−z∗,z∗)(-z^*,z^*), with z∗≈1z^*\!≈\!1, as shown in the phase portrait of Fig. 7(a). Provided Δtop t_op is sufficiently large, the trajectories converge to one of these equilibria within the interval (Fig. 7(b)). Under the rule (10), the joint action ultimately settles on either (ap1,ap3)(a_p_1,a_p_3) or (ap3,ap1)(a_p_3,a_p_1), both of which lie in π(s0)π(s_0) by (8). Fig. 7: (a) Phase portrait of (9): trajectories from generic initial conditions diverge from (0,0)(0,0) (open circle) and converge to one of the two stable equilibria (filled circles). (b) Time evolution of z1(t)z_1(t) (solid) and z2(t)z_2(t) (dashed) from several initial conditions; the shaded interval marks the reserved window Δtop t_op. The example shows that the decentralized coordination challenges are mitigated by the opinion dynamics. The instability of the symmetric equilibrium (0,0)(0,0) breaks the initial symmetry between the two homogeneous agents and drives their opinions apart, assigning opposite roles in this anti-coordination scenario. Both stable equilibria pair the two opinions with opposite signs, so each agent is committed to the role complementary to the other’s, and its choice thereby becomes adaptive to the other agent’s unknown decision behavior. This opinion-guided implementation is intrinsically two-layered, with the layers operating on separated time scales. On the continuous opinion layer, each agent’s opinion ziz_i evolves under mutual coupling and converges to one of the stable equilibria. On the discrete strategy layer, the sign of ziz_i commits the agent to a discrete action through the rule σopσ^op in (10). In what follows, we first formalize the opinion-guided strategy at the discrete strategy layer, generalizing it from the scalar running example to multi-dimensional opinions, and then present its layered realization through the continuous opinion dynamics, sketched in Fig. 8. GGGG⋯·sσ1opσ^op_1σ2opσ^op_2sσ1opσ^op_1σ2opσ^op_2sσ1opσ^op_1σ2opσ^op_2sssst0t_0t1t_1t2t_2t3t_3Strategy Layertt(z~1,z~2)( z_1, z_2)Δtop t_op⋯·sz1z_1z2z_2ContinuousOpinion Layer Fig. 8: Layered realization of opinion-guided strategy. Strategy layer: at each decision time tkt_k, both agents query the game G at the current state s and select opinion-guided strategies σ1op,σ2opσ^op_1,σ^op_2 whose action intents are compatible, as required by (12). Continuous opinion layer: between consecutive decisions, the coupled opinions z1z_1 and z2z_2 converge to opposite signs within the reserved window Δtop t_op. At the end of each window, the converged opinion is abstracted to its discrete sign z~i z_i and passed up to the strategy layer, which commits one of the compatible action pairs. IV-B1 Opinion-Guided Strategy As the running example suggests, one-dimensional opinion dynamics with a bistable bifurcation reliably converge to one of two stable equilibria [6]. When the opinion state is multi-dimensional, each component behaves analogously, so we abstract each agent i’s continuous opinion into a discrete opinion state z~i∈ni z_i\!∈\!B^n_i, where :=−1,1B\!:=\!\-1,1\ is a binary index set and ni∈ℕn_i\!∈\!N is the dimension of that agent’s opinion state. Each entry of z~i z_i records which of the two subsets has been selected along the corresponding component, and this discrete abstraction induces a partition of the agent’s action space. For each agent i∈1,2i\!∈\!\1,2\, we define a map i:Ai→niO_i\!:\!A_i\!→\!B^n_i from actions to discrete opinion states. It labels each action with an element of niB^n_i, thereby partitioning AiA_i into at most 2ni2^n_i subsets; we write Ai(z~i):=a∈Ai∣i(a)=z~iA_i^( z_i)\!:=\!\a\!∈\!A_i _i(a)\!=\! z_i\ for the set of actions labeled z~i z_i. Given the partitioning, the opinion state of agent i is used to select an action that responds appropriately to the action agent j intends. This motivates the following definition of opinion-guided strategies. Definition 8 (Opinion-Guided Strategy). Given a game G and partition maps 1,2O_1,O_2, an opinion-guided strategy for agent i∈1,2i\!∈\!\1,2\ is a mapping σiop:ℋ×ni→Aiσ^op_i\!:\!H\!×\!B^n_i\!→\!A_i from a history h and a discrete opinion state z~i z_i to an action aia_i. Agent i’s opinion state is determined from its opponent’s through its own opinion-pairing map Γi:ℋ×nj→ni _i\!:\!H×B^n_j\!→\!B^n_i: ai=σiop(h,z~i),z~i=Γi(h,z~j),a_i=σ^op_i(h, z_i), z_i= _i(h, z_j), (11) where the selected action lies in Ai(z~i)A_i^( z_i) with z~i=i(ai) z_i\!=\!O_i(a_i), and z~j=j(aj) z_j\!=\!O_j(a_j) is the opinion state of the other agent j∈1,2∖ij\!∈\!\1,2\\! \!\i\ corresponding to its intended action aja_j. The set of all such strategies is denoted Σiop ^op_i. Each opinion-pairing map Γi _i encodes how agent i reacts to the other’s opinion state. For simplicity, we assume in this paper that Γi _i is linear in the opinion state z~j z_j and that each component of z~i z_i is coupled to at most one component of z~j z_j. Assumption 1 (Linear Opinion Pairing). The opinion-pairing map at h is linear, represented by a partial signed permutation matrix Γi(h)∈−1,0,+1ni×nj _i(h)\!∈\!\-1,0,+1\^n_i× n_j with at most one nonzero entry per row. Specifically, the k-th row of Γi(h) _i(h) governs the k-th component of z~i z_i: a +1+1 entry couples it in agreement with the matched component of z~j z_j, a −1-1 entry couples it in opposition, and an all-zero row leaves it free, coupled to no component of z~j z_j. In the running example of Section I-A, each agent’s opinion-pairing map at s0s_0 reduces to the scalar Γi(s0)=−1 _i(s_0)\!=\!-1, representing the rule “if the other agent commits to one target zone, I commit to the other”. To achieve the coordination required by π, the maps Γ1 _1 and Γ2 _2 must be designed carefully, which will be discussed in Section V. When agent 11 employs an opinion-guided strategy σ1op∈Σ1opσ^op_1\!∈\! ^op_1 and agent 22 a non-reactive strategy σ2:ℋ→A2 _2\!:\!H\!→\!A_2, the joint outcome is uniquely determined as (σ1op∥σ2)(h)=(σ1op(h,z~1),a2)(σ^op_1\,\|\, _2)(h)=\(σ^op_1(h, z_1),\,a_2)\, where a2=σ2(h)a_2\!=\! _2(h), z~2=2(a2) z_2\!=\!O_2(a_2), and z~1=Γ1(h)z~2 z_1\!=\! _1(h)\, z_2. When both agents employ opinion-guided strategies σ1op,σ2opσ^op_1,σ^op_2, their opinion states must jointly satisfy both couplings z~1=Γ1(h)z~2 z_1\!=\! _1(h)\, z_2 and z~2=Γ2(h)z~1 z_2\!=\! _2(h)\, z_1, which in general admit multiple solutions, so the joint outcome is set-valued: (σop1∥σop2)(h):= (σ^op_1\,\|\,σ^op_2)(h)\!:=\! \ (a1,a2)|∃z~i,z~j:ai=σiop(h,z~i), (a_1,a_2) ∃ z_i, z_j:\ a_i\!=\!σ^op_i(h, z_i), (12) z~i=Γi(h)z~j,z~j=j(aj), z_i\!=\! _i(h)\, z_j,\ z_j\!=\!O_j(a_j), i,j∈1,2,i≠j. i,j\!∈\!\1,2\,\ i\!≠\!j\, \. To realize such joint behavior, the agents need to exchange their opinion states and converge on common pairs of opinions that satisfy the coupling constraints. This is achieved by the continuous opinion dynamics described subsequently. Remark 1. The fixed partition i:Ai→niO_i\!:\!A_i\!→\!B^n_i is restrictive: since π(h)π(h) varies across histories, a history-dependent partition would be more flexible. For simplicity, we adopt the fixed partition in this paper. IV-B2 Layered Realization To realize the opinion-guided strategy in Def. 8, agent i maintains a continuous opinion state zi(t)∈ℝniz_i(t)\!∈\!R^n_i that evolves according to the NOD equation [33] and is mapped to its discrete counterpart z~i∈ni z_i\!∈\!B^n_i via the element-wise sign function z~i:=sgn(zi) z_i\!:=\!sgn(z_i). Before introducing the communication-free implementation (Section VI), we first describe the idealized interaction in which each agent receives its opponent’s opinion state directly in real time through communication. The layered realization of the opinion-guided strategy (11) is given by: Strategy layer:ai=σiop(h,z~i),z~i=sgn(zi(tk+Δtop))Opinion layer:z˙i=−dzi+ρitanh(νzi+βΓi(h)zj+bi) -10.00002pt \ aligned Strategy layer:\,a_i&\!=\!σ^op_i(h, z_i),\; z_i\!=\!sgn(z_i(t_k\!+\! t_op))\\ Opinion layer:\, z_i&\!=\!-dz_i\!+\! _i \! (ν z_i\!+\!β _i(h)z_j\!+\!b_i ) aligned . (13) where d>0d\!>\!0 is a damping coefficient, tanh is a saturation function that enables fast and flexible decision-making, and ρi≥0 _i\!≥\!0 is a tunable attention parameter. The gains ν>0ν\!>\!0 and β>0β\!>\!0 weigh the self- and opponent-coupling terms, and bib_i is a constant bias encoding agent i’s prior preference. Here Γi(h)zj∈ℝni _i(h)\,z_j\!∈\!R^n_i is the target opinion that agent i should align with, where Γi(h) _i(h) is the opinion-pairing map from Def. 8. The two layers evolve on separated time scales: the strategy layer acts in discrete time, at decision instants tkt_k indexed by k∈ℕk\!∈\!N, while the opinion layer evolves in continuous time t∈ℝ≥0t\!∈\!R_≥ 0. Each decision step begins at tkt_k with a reserved window of length Δtop t_op over which the opinion converges; the action is then committed from the settled opinion at tk+Δtopt_k\!+\! t_op and executed for the remainder of the step. V Analysis for Opinion-Guided Coordination In this section, we analyze formal guarantees for the layered opinion-guided strategy. The analysis is organized layer by layer. We first analyze the strategy layer, characterizing how to design an opinion-guided strategy that achieves robust coordination against rational opponents and admits homogeneous deployment. We then analyze the continuous opinion layer and close with a layered guarantee, showing that the nonlinear opinion dynamics realize the desired coordination. V-A Robust Coordination at Strategy Layer For each i∈1,2i\!∈\!\1,2\, the function i:Ai→niO_i\!:\!A_i\!→\!B^n_i partitions the action set AiA_i into the subsets Ai(z~i)A_i^( z_i) defined earlier. Next, we introduce the partition-restricted robust action sets Ai(h,z~j,Σj)A_i rob(h; z_j, _j), which refine Ai(h,Σj)A_i rob(h; _j) in Def. 4 by quantifying over a single cell Aj(z~j)A_j^( z_j) of agent j rather than the full selectable set AΣj(h)A _j(h), as follows: Ai(h;z~j,Σj):=ai| A_i rob(h; z_j, _j)\!:=\!\a_i\; | ∀aj∈Aj(z~j)∩AΣj(h): ∀ a_j\!∈\!A_j^( z_j)\!∩\!A _j(h): (14) (a1,a2)∈π(h). (a_1,a_2)\!∈\!π(h)\. Two degenerate cases make agent i’s choice immaterial. If Aj(z~j)∩AΣj(h)=∅A_j^( z_j)∩ A _j(h)\!=\! , agent j has no selectable action in the cell Aj(z~j)A_j^( z_j), so there is nothing to coordinate against. If π(h)=∅π(h)\!=\! , no joint action is required at h. In both cases, we set Ai(h,z~j,Σj)=AiA_i rob(h; z_j, _j)\!=\!A_i, so the robust set imposes no restriction. When the selectable set is fixed to the rational set Σj _j rat, we abbreviate Ai(h,z~j):=Ai(h,z~j,Σj)A_i rob(h; z_j)\!:=\!A_i rob(h; z_j, _j rat). To illustrate how the partitions induced by 1,2O_1,O_2 enable robust coordination, we consider the permissive action set π4(h) _4(h) in Fig. 5(d), in which neither agent has a non-empty robust action set, i.e., Ai(h)=∅A_i rob(h)\!=\! for both i∈1,2i\!∈\!\1,2\. As shown in Fig. 9(a), partitioning A1A_1 via 1O_1 makes A2(h,z~1)≠∅A_2 rob(h; z_1)\!≠\! on each cell A1(z~1)A_1^( z_1), z~1∈−1,1 z_1\!∈\!\-1,1\, so agent 22 can guarantee coordination on its own once z~1 z_1 is known. As shown in Fig. 9(b), when both agents partition their action sets via 1O_1 and 2O_2, Ai(h,z~j)≠∅A_i rob(h; z_j)\!≠\! holds for each agent i and every opinion state z~j z_j of the other agent, so each agent can guarantee coordination on its own once the other agent’s opinion state is known. The conditions under which an opinion-guided strategy achieves π-robust coordination are formalized in the following theorems. Fig. 9: Illustration of the opinion-guided strategy from the perspective of agent i, i∈1,2i\!∈\!\1,2\. The dashed line with a scissor icon denotes the partition of AjA_j induced by jO_j. In (a), only AjA_j is partitioned, indexed by z~j∈−1,1 z_j\!∈\!\-1,1\. In (b), the action sets of both agents are mutually partitioned according to their respective discrete opinion states z~1,z~2∈n z_1, z_2\!∈\!B^n. Theorem 1 (Robust coordination against a set of non-reactive opponents). Consider a two-agent game G with permissive joint strategy π, given partition functions i,jO_i,O_j, and a set of non-reactive rational strategies Σj⊆Σj _j\! \! _j rat for agent j. There exists an opinion-guided strategy σiop∈Σiopσ^op_i\!∈\! _i^op for agent i that achieves (π,Σj)(π, _j)-robust coordination if there exists an opinion-pairing map Γi _i such that for every h∈ℋh\!∈\!H and every opinion state z~j∈nj z_j\!∈\!B^n_j: Ai(h,z~j,Σj)∩Ai(z~i)≠∅,z~i=Γi(h,z~j).A_i rob(h; z_j, _j)∩ A_i^( z_i)\!≠\! , z_i\!=\! _i(h, z_j). (15) Proof. Let Γi _i be an opinion-pairing map satisfying (15). We first build an opinion-guided strategy with Γi _i and then show it is robust. For every history h∈ℋh\!∈\!H and opponent opinion z~j∈nj z_j\!∈\!B^n_j, let z~i=Γi(h,z~j) z_i\!=\! _i(h, z_j) and choose σiop(h,z~i)∈Ai(h,z~j,Σj)∩Ai(z~i)σ^op_i(h, z_i)\!∈\!A_i rob(h; z_j, _j)∩ A_i^( z_i), which is non-empty by (15). Next, we show that the designed σiopσ^op_i is (π,Σj)(π, _j)-robust. Let σj∈Σj _j\!∈\! _j and h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! be arbitrary. The opponent plays aj=σj(h)a_j\!=\! _j(h), which reveals its opinion z~j=j(aj) z_j\!=\!O_j(a_j) and places aj∈Aj(z~j)∩AΣj(h)a_j\!∈\!A_j^( z_j)\!∩\!A _j(h). Agent i reacts with z~i=Γi(h,z~j) z_i\!=\! _i(h, z_j) and plays ai=σiop(h,z~i)a_i\!=\!σ^op_i(h, z_i), which by construction lies in Ai(h,z~j,Σj)A_i rob(h; z_j, _j). By the definition of the robust action set (14), such an action coordinates with every action agent j may select in the cell Aj(z~j)A_j^( z_j), in particular with aja_j, so (ai,aj)∈π(h)(a_i,a_j)\!∈\!π(h). As σj _j and h were arbitrary, coordination holds against every rational strategy in Σj _j at every history with π(h)≠∅π(h)\!≠\! , i.e. σiopσ^op_i is (π,Σj)(π, _j)-robust. ∎ Theorem 1 gives a sufficient condition for a (π,Σj)(π, _j)-robust opinion-guided strategy to exist: there is an opinion-pairing map Γi _i under which, whenever agent j commits to an intent z~j z_j, the paired opinion state z~i=Γi(h,z~j) z_i\!=\! _i(h, z_j) of agent i contains an action that handles every action agent j may select in Aj(z~j)∩AΣj(h)A_j^( z_j)\!∩\!A _j(h). As shown in the proof, the strategy is then constructed by letting agent i commit to an action of Ai(h,z~j,Σj)∩Ai(z~i)A_i rob(h; z_j, _j)∩ A_i^( z_i). Theorem 1 holds for any fixed Σj⊆Σj _j\! \! _j rat; taking Σj=Σj _j\!=\! _j rat, the full rational strategy set of agent j, specializes it to the strongest guarantee, robustness against every non-reactive rational opponent at once. Theorem 2 (Robust coordination against all rational opponents). Consider a two-agent game G with permissive joint strategy π and given partition functions i,jO_i,O_j. There exists an opinion-guided strategy σiop∈Σiopσ^op_i\!∈\! _i^op for agent i that achieves (π,Σj)(π, _j rat)-robust coordination if and only if there exists an opinion-pairing map Γi _i such that for every h∈ℋh\!∈\!H and every opinion state z~j∈nj z_j\!∈\!B^n_j: Ai(h,z~j)∩Ai(z~i)≠∅,z~i=Γi(h,z~j).A_i rob(h; z_j)∩ A_i^( z_i)\!≠\! , z_i\!=\! _i(h, z_j). (16) Proof. (⇐ ) Sufficiency. This follows from Theorem 1 with Σj=Σj _j\!=\! _j rat. (⇒ ) Necessity. We argue by contradiction. Suppose σiop∈Σiopσ^op_i\!∈\! _i^op is (π,Σj)(π, _j rat)-robust with pairing Γi _i, yet at some history h⋆∈ℋh \!∈\!H and opinion state z~j⋆ z_j , Ai(h⋆,z~j⋆)∩Ai(z~i⋆)=∅A_i rob(h ; z_j )∩ A_i^( z_i )\!=\! , where z~i⋆=Γi(h⋆,z~j⋆) z_i \!=\! _i(h , z_j ). Consider the action ai⋆=σiop(h⋆,z~i⋆)a_i \!=\!σ^op_i(h , z_i ), which lies in Ai(z~i⋆)A_i^( z_i ) by Def. 8. For every rational action aj∈Aj(z~j⋆)∩Aj(h⋆)a_j\!∈\!A_j^( z_j )\!∩\!A_j rat(h ), some σj∈Σj _j\!∈\! _j rat plays aja_j at h⋆h ; since j(aj)=z~j⋆O_j(a_j)\!=\! z_j , agent i reacts with z~i⋆ z_i and plays ai⋆a_i , and robustness gives (ai⋆,aj)∈π(h⋆)(a_i ,a_j)\!∈\!π(h ). By (14), ai⋆∈Ai(h⋆,z~j⋆)a_i \!∈\!A_i rob(h ; z_j ), so ai⋆∈Ai(h⋆,z~j⋆)∩Ai(z~i⋆)a_i \!∈\!A_i rob(h ; z_j )∩ A_i^( z_i ), contradicting the emptiness of this intersection. Hence, if there exists a (π,Σj)(π, _j rat)-robust opinion-guided strategy, then no pair (h⋆,z~j⋆)(h , z_j ) with Ai(h⋆,z~j⋆)∩Ai(z~i⋆)=∅A_i rob(h ; z_j )∩ A_i^( z_i )\!=\! exists. Thus, the necessity is proved. ∎ Running example: Recall that π(s0)=(ap1,ap3),(ap3,ap1)π(s_0)\!=\!\(a_p_1,a_p_3),(a_p_3,a_p_1)\ and Ai(s0)=ap1,ap3A_i rat(s_0)\!=\!\a_p_1,a_p_3\. Let each robot hold a scalar opinion state (n1=n2=1n_1\!=\!n_2\!=\!1), and let both share the partition function (ap1)=+1O(a_p_1)\!=\!+1, (ap2)=(ap3)=−1O(a_p_2)\!=\!O(a_p_3)\!=\!-1 and the opinion-pairing map Γ(s0)=−1 (s_0)\!=\!-1. For z~j=+1 z_j\!=\!+1, the cell Aj(+1)=ap1A_j^(+1)\!=\!\a_p_1\ contains the single rational action ap1a_p_1, and only ap3a_p_3 coordinates with it, so Ai(s0,+1)=ap3A_i rob(s_0;+1)\!=\!\a_p_3\ by (14); the paired opinion state z~i=−1 z_i\!=\!-1 then gives Ai(s0,+1)∩Ai(−1)=ap3≠∅A_i rob(s_0;+1)∩ A_i^(-1)\!=\!\a_p_3\\!≠\! . The same computation for z~j=−1 z_j\!=\!-1 gives Ai(s0,−1)∩Ai(+1)=ap1≠∅A_i rob(s_0;-1)∩ A_i^(+1)\!=\!\a_p_1\\!≠\! . Condition (15) thus holds at s0s_0, and the robust opinion-guided strategy can be designed as σop(s0,+1)=ap1,σop(s0,−1)=ap3.σ^op(s_0,+1)\!=\!a_p_1, σ^op(s_0,-1)\!=\!a_p_3. (17) With this strategy, each robot can guarantee coordination at s0s_0 with any rational opponent, regardless of which rational action it selects, as established by Theorem 2. We now extend Theorem 1 to the setting in which both agents employ opinion-guided strategies. Extending Def. 2, we call an opinion-guided strategy σjop∈Σjopσ^op_j\!∈\! _j^op rational if σjop(h,z~j)∈Aj(h)σ^op_j(h, z_j)\!∈\!A_j rat(h) for every h∈ℋh\!∈\!H and every z~j∈nj z_j\!∈\!B^n_j, and we denote by Σjop, _j^op, rat the set of all such strategies. The following theorem characterizes robust coordination against Σjop, _j^op, rat. Theorem 3 (Robust coordination against opinion-guided opponents). For a two-agent game G with permissive joint strategy π, suppose the partition functions i,jO_i,O_j and the other agent’s opinion-pairing map Γj _j are given. There exists an opinion-guided strategy σiop∈Σiopσ^op_i\!∈\! _i^op for agent i that achieves (π,Σjop,)(π, _j^op, rat)-robust coordination if there exists an opinion-pairing map Γi _i such that, for every h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! , the coupling [Ini−Γi(h)−Γj(h)Inj][z~iz~j]=0~ bmatrixI_n_i&- _i(h)\\ - _j(h)&I_n_j bmatrix bmatrix z_i\\ z_j bmatrix=0 (18) admits at least one solution (z~i,z~j)( z_i, z_j) with Aj(z~j)∩Aj(h)≠∅A_j^( z_j)\!∩\!A_j rat(h)\!≠\! , and every such solution satisfies: Ai(h,z~j)∩Ai(z~i)≠∅.A_i rob(h; z_j)∩ A_i^( z_i)\!≠\! . (19) Here, IniI_n_i and InjI_n_j denote the identity matrices of size nin_i and njn_j. Proof. Let Γi _i be an opinion-pairing map satisfying the conditions of the theorem. We first build an opinion-guided strategy σiopσ^op_i from Γi _i and then show it coordinates robustly with every rational opinion-guided opponent. For every history h∈ℋh\!∈\!H and opinion state z~i∈ni z_i\!∈\!B^n_i, let z~j=Γj(h)z~i z_j\!=\! _j(h) z_i be the opinion state paired with z~i z_i through the opponent’s map. If (z~i,z~j)( z_i, z_j) solves the coupling (18) with Aj(z~j)∩Aj(h)≠∅A_j^( z_j)\!∩\!A_j rat(h)\!≠\! , then z~i=Γi(h)z~j z_i\!=\! _i(h) z_j and we choose σiop(h,z~i)∈Ai(h,z~j)∩Ai(z~i),σ^op_i(h, z_i)\!∈\!A_i rob(h; z_j)∩ A_i^( z_i), (20) which is non-empty by (19); otherwise let σiop(h,z~i)σ^op_i(h, z_i) be any action of Ai(z~i)A_i^( z_i). Next, we show that the designed σiopσ^op_i is (π,Σjop,)(π, _j^op, rat)-robust. Let σjop∈Σjop,σ^op_j\!∈\! _j^op, rat and h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! be arbitrary. Since the coupling (18) admits a solution (z~i,z~j)( z_i, z_j) with Aj(z~j)∩Aj(h)≠∅A_j^( z_j)\!∩\!A_j rat(h)\!≠\! , at which the rational opponent has a feasible action to play, the joint outcome (σiop∥σjop)(h)(σ^op_i\,\|\,σ^op_j)(h) is non-empty. Let (ai,aj)∈(σiop∥σjop)(h)(a_i,a_j)\!∈\!(σ^op_i\,\|\,σ^op_j)(h) be any realized joint action. By (12), the induced opinions z~i=i(ai) z_i\!=\!O_i(a_i) and z~j=j(aj) z_j\!=\!O_j(a_j) solve the coupling (18), with ai=σiop(h,z~i)a_i\!=\!σ^op_i(h, z_i) and aj=σjop(h,z~j)a_j\!=\!σ^op_j(h, z_j). Since σjopσ^op_j is rational, aj∈Aj(z~j)∩Aj(h)a_j\!∈\!A_j^( z_j)\!∩\!A_j rat(h), so this coupling solution has Aj(z~j)∩Aj(h)≠∅A_j^( z_j)\!∩\!A_j rat(h)\!≠\! and the hypothesis (19) applies, giving Ai(h,z~j)∩Ai(z~i)≠∅A_i rob(h; z_j)∩ A_i^( z_i)\!≠\! , and the construction (20) sets ai=σiop(h,z~i)∈Ai(h,z~j)a_i\!=\!σ^op_i(h, z_i)\!∈\!A_i rob(h; z_j). By (14), aia_i coordinates with aja_j, hence (ai,aj)∈π(h)(a_i,a_j)\!∈\!π(h). As σjopσ^op_j, h, and the realized pair were arbitrary, σiopσ^op_i is (π,Σjop,)(π, _j^op, rat)-robust. ∎ Compared with Theorem 1, Theorem 3 imposes a more restrictive coupling between the two agents’ opinions to ensure their mutual compatibility. Intuitively, when the opponent commits to a fixed action, agent i only needs to adapt to it; when both agents react to each other’s opinions, their reactions must be mutually compatible for the coordination prescribed by π. Under this condition, a (π,Σjop,)(π, _j^op, rat)-robust strategy can be designed as illustrated in the proof. In fact, the two theorems can be linked because a non-reactive rational opponent is a limiting case of an opinion-guided one. When Γj(h)= _j(h)\!=\!0, agent j’s opinion state z~j z_j is decoupled from agent i’s opinion z~i z_i and stays fixed, so agent j commits to a fixed action, i.e., a non-reactive opponent. Moreover, since real agents carry varying degrees of prior preference, encoded in the bias bjb_j of the opinion dynamics (13), a single opinion-guided model spans the range from a fully reactive to a fully committed opponent. A strategy satisfying Theorem 3 therefore achieves robust coordination against both reactive and non-reactive opponents, as formalized in the following corollary. Corollary 1 (Robust coordination against both reactive and non-reactive opponents). For a two-agent game G with permissive joint strategy π, suppose the partition functions i,jO_i,O_j and the other agent’s opinion-pairing map Γj _j are given. Suppose there exists an opinion-pairing map Γi _i such that, for every h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! and every opinion state z~j∈nj z_j\!∈\!B^n_j, the coupling (18) admits a solution (z~i,z~j)( z_i, z_j), and every such solution satisfies (19). Then the opinion-guided strategy σiopσ^op_i constructed from Γi _i via (20) achieves robust coordination not only against reactive opinion-guided opponents σjop∈Σjop,σ^op_j\!∈\! _j^op, rat but also against non-reactive rational opponents σjop∈Σjσ^op_j\!∈\! _j rat. Proof. By assumption Γi _i satisfies the conditions of Theorem 3, so σiopσ^op_i is (π,Σjop,)(π, _j^op, rat)-robust. Because the coupling (18) is solvable at every z~j∈nj z_j\!∈\!B^n_j and (19) holds at every solution, the pairing z~i=Γi(h)z~j z_i\!=\! _i(h) z_j satisfies (19) for every z~j z_j, which is exactly the full-rational condition (16) of Theorem 2. Hence Γi _i also satisfies Theorem 2, so the same σiopσ^op_i is (π,Σj)(π, _j rat)-robust. ∎ Running example: When both robots run the strategy (17), so that Γi(s0)=Γj(s0)=−1 _i(s_0)\!=\! _j(s_0)\!=\!-1, the coupling (18) has two solutions, (z~i,z~j)=(+1,−1)( z_i, z_j)\!=\!(+1,-1) and (−1,+1)(-1,+1). At both solutions, condition (19) holds by the robust-set computation above, so Theorem 3 guarantees that every joint action realized at s0s_0 lies in π(s0)π(s_0). Since the coupling is solvable at every z~j z_j, Corollary 1 extends this guarantee to a non-reactive rational opponent as well, so the strategy is robust against any rational opponent, reactive or not. Finally, since the game is homogeneous, the strategy is also π-homogeneously implementable: both robots can run it and still realize only pairs in π(s0)π(s_0). Theorems 1 and 3 are stated for predefined (π,i,j,Γj)(π,O_i,O_j, _j) and allow the two agents to be heterogeneous. When both agents are homogeneous, they run the same σopσ^op, and share a single partition i=j=O_i\!=\!O_j\!=\!O and a single opinion-pairing map Γi(h)=Γj(h)=Γ(h) _i(h)\!=\! _j(h)\!=\! (h) with identical opinion dimension n1=n2n_1\!=\!n_2. Expanding Eq. (18) and eliminating z~j z_j gives z~i=Γi(h)z~j,z~j=Γj(h)z~i⟹z~i=Γi(h)Γj(h)z~i. z_i\!=\! _i(h) z_j, z_j\!=\! _j(h) z_i\; \; z_i\!=\! _i(h) _j(h)\, z_i. Whatever opinion state the other agent holds, the two couplings must admit a matched pair (z~i,z~j)( z_i, z_j), as Corollary 1 requires. For the shared map Γ(h) (h), such a pair exists at every z~j∈n z_j\!∈\!B^n when Γ(h)2 (h)^2 acts as the identity on the components that Γ(h) (h) couples: an all-zero row leaves its component of z~i z_i free and constrains nothing. A nonzero entry in every row couples them all, and the condition is then involutivity, Γ(h)2=I (h)^2\!=\!I (equivalently, Γ(h)=Γ(h)−1 (h)\!=\! (h)^-1), at every h. Corollary 2 (Robust coordination among homogeneous agents). Let G be a homogeneous game (Def. 5) with permissive joint strategy π, and suppose a shared partition O and an involutive pairing map Γ (i.e., Γ(h)=Γ(h)−1 (h)\!=\! (h)^-1 for all h) satisfy the robust-coordination condition (19). Then the (π,Σjop,)(π, _j^op, rat)-robust strategy σop∈Σop,σ^op\!∈\! ^op, rat constructed from (20) is also π-homogeneously implementable (Def. 6): ∅≠(σop∥σop∘τ¯)(h)⊆π(h),∀h∈ℋ with π(h)≠∅. ≠(σ^op\,\|\,σ^op\! \! τ)(h)\! \!π(h), ∀ h\!∈\!H with π(h)\!≠\! . (21) Proof. Write :=1=2O\!:=\!O_1\!=\!O_2 and Γ:=Γ1=Γ2 \!:=\! _1\!=\! _2 for the shared partition and pairing map, with n1=n2=:n_1\!=\!n_2\!=:\!n. The involutive Γ makes the coupling (18) solvable for every z~j∈n z_j\!∈\!B^n, since z~i=Γ(h)z~j z_i\!=\! (h) z_j gives Γ(h)z~i=Γ(h)2z~j=z~j (h) z_i\!=\! (h)^2 z_j\!=\! z_j. Together with the assumed condition (19), Γ satisfies the requirements of Theorem 3, so the strategy σopσ^op built from it via (20) is (π,Σjop,)(π, _j^op, rat)-robust. By the definition of the robust action set (14), a robust action pairs with the opponent’s rational action to form a joint action in π(h)π(h), and is therefore itself rational (Def. 2); hence the robust σopσ^op must be rational for agent i, i.e., σop∈Σiop,σ^op\!∈\! _i^op, rat. The relabeling τ¯ τ is a symmetry of the homogeneous game, so the relabeled copy σop∘τ¯σ^op\! \! τ plays agent j exactly as σopσ^op plays agent i. Since σopσ^op is rational for agent i, this symmetry makes σop∘τ¯σ^op\! \! τ rational for agent j: every action it selects lies in Aj(h)A_j rat(h) at every history h. Hence σop∘τ¯σ^op\! \! τ is a rational opinion-guided strategy for agent j, i.e., σop∘τ¯∈Σjop,σ^op\! \! τ∈ _j^op, rat. Consequently, for every h∈ℋh\!∈\!H with π(h)≠∅π(h)\!≠\! , the joint outcome (σop∥σop∘τ¯)(h)(σ^op\,\|\,σ^op\! \! τ)(h) is non-empty, and the robustness of σopσ^op against Σjop, _j^op, rat places every realized pair in π(h)π(h). This is (21), so σopσ^op is π-homogeneously implementable. ∎ Corollary 2 shows that in a homogeneous game, once a shared partition O and an involutive pairing map Γ satisfy the robust-coordination condition (19), the resulting robust opinion-guided strategy is automatically homogeneously implementable. Combining this with Corollary 1, a strategy built on the shared (,Γ)(O, ) achieves the coordination prescribed by π against any rational opponent, whether reactive or non-reactive, and whether the opponent runs the identical or a different strategy. In summary, this section establishes, step by step, the conditions under which an opinion-guided strategy can resolve the decentralized coordination challenges. The triple (π,,Γ)(π,O, ) plays a central role in these conditions. Given a scenario, π specifies which joint actions count as successful coordination; O partitions each agent’s (possibly infinite) action set into finitely many cells, each indexed by a discrete opinion state; and Γ specifies how an agent sets its own opinion in reaction to the opponent’s opinion. Intuitively, these three objects represent common knowledge shared by agents, analogous to the traffic conventions that drivers learn from experience. V-B Convergence Analysis of Opinion Dynamics The decision-layer analysis above establishes robust coordination, provided the two agents settle on a matched pair of opinion states that solves the coupling (11). However, it does not explain how the agents arrive at such a pair. We now show how this discrete coupling is realized by the continuous opinion dynamics in (13). Under the linear-pairing Assumption 1 on Γi(h) _i(h), each component of agent i’s opinion state couples to at most one component of agent j’s. The full opinion dynamics (13) therefore separate into independent scalar subsystems, and their convergence analysis reduces to that of a single scalar subsystem. Each subsystem inherits a single coupling entry γ∈+1,0,−1γ\!∈\!\+1,0,-1\ from Γi(h) _i(h). For analytical simplicity, we further consider agents with no prior preference, b1=b2=0b_1\!=\!b_2\!=\!0, set the self- and opponent-coupling gains to a common value κ>0κ\!>\!0, and take ρ1=ρ2=ρ _1\!=\! _2\!=\!ρ. Under these assumptions, the two-agent dynamics for scalar opinions (z1,z2)(z_1,z_2) are given by z˙1=−dz1+ρtanh(κ(z1+γz2))z˙2=−dz2+ρtanh(κ(z2+γz1)), \\! array[]l z_1\!=\!-dz_1+ρ \! (κ(z_1\!+\!γ z_2) )\\ z_2\!=\!-dz_2+ρ \! (κ(z_2\!+\!γ z_1) ) array ., (22) whose adjacency matrix is A=[0γ0]A\!=\! bmatrix0&γ\\ γ&0 bmatrix. If γ=0γ\!=\!0, this component is decoupled from the opponent’s opinion. Starting from the neutral initial value zi=0z_i\!=\!0, it remains at zi=0z_i\!=\!0, leaving the agent free to assign either discrete opinion state to it regardless of the opponent’s choice. When γ=±1γ\!=\!± 1, we analyze how the two opinions evolve and converge to a matched pair of discrete opinion states. By [6, Cor. IV.1.2] applied to system (22), the neutral opinion =[z1z2]⊤= z\!=\![z_1\ z_2] \!=\!0 is locally exponentially stable for ρ<ρ⋆ρ\!<\!ρ and loses stability for ρ>ρ⋆ρ\!>\!ρ , with critical value ρ⋆=d2κρ \!=\! d2κ. At ρ=ρ⋆ρ\!=\!ρ , the system undergoes a supercritical pitchfork bifurcation: two branches emerge from the origin, tangent to max=[1γ]⊤ v_ \!=\![1\ γ] , the eigenvector associated with the largest eigenvalue of A (Fig. 10). To activate the bifurcation for decisive coordination, we set ρ=d2κ+ϵρ= d2κ+ε (23) with ϵ>0ε\!>\!0 small, placing ρ just past ρ⋆ρ . The opinion state z then converges to one of the two stable equilibria + z^+ or − z^-, small positive and negative multiples of max v_ respectively, corresponding to a matched pair of discrete opinion states (z~1,z~2)( z_1, z_2) with z~1=γz~2∈+1,−1 z_1\!=\!γ\, z_2\!∈\!\+1,-1\, the role assignment prescribed by γ. The convergence time can be reduced by increasing ϵε, allowing the opinion dynamics to settle on such a matched pair, after which the strategy layer commits to the corresponding joint action in π(h)π(h). (a) γ=+1γ\!=\!+1 (b) γ=−1γ\!=\!-1 Fig. 10: Bifurcation diagrams of the two-agent scalar opinion dynamics (22) as the attention gain ρ increases, shown for the two coupling signs γ. (a) γ=+1γ\!=\!+1: the two stable branches emerge along ±[1,1]⊤±[1,1] , pairing same-sign opinions. (b) γ=−1γ\!=\!-1: the stable branches emerge along ±[1,−1]⊤±[1,-1] , pairing opposite-sign opinions, as in the running example (9). In both panels the neutral equilibrium = z\!=\!0 is stable for ρ<ρ⋆ρ\!<\!ρ and undergoes a supercritical pitchfork at the threshold ρ⋆=d/(2κ)=0.5ρ \!=\!d/(2κ)\!=\!0.5 (green star); solid blue curves are the stable equilibrium branches and the dashed red curve is the unstable branch. The vertical slice ρ=2ρ\!=\!2 carries the phase portrait, whose trajectories diverge from the unstable equilibrium (open circle) and converge to the two stable equilibria (filled circles), with the red curves marking the separatrix. V-C Layered Guarantee Given the above analysis of the two layers, a guarantee of robust coordination with respect to π can be obtained for the whole layered opinion-guided framework, based on a single time-scale separation between them: Assumption 2 (Time-scale separation). The strategy layer acts at times tkt_k, k∈ℕk\!∈\!N. Over each period [tk,tk+1)[t_k,t_k+1): 1. the history h of G remains constant; 2. the window Δtop<tk+1−tk t_op\!<\!t_k+1-t_k is long enough that the opinion state z~i=sgn(zi) z_i\!=\!sgn(z_i) remains constant over [tk+Δtop,tk+1)[t_k\!+\! t_op,t_k+1). With this assumption, the strategy layer selects a joint action at each decision time tkt_k based on the opinion states at that time, and the opinion dynamics have enough time to settle on a matched pair of opinion states before the next decision time. The following theorem then establishes a layered guarantee for robust coordination. Theorem 4 (Layered guarantee for opinion-guided coordination). Consider a two-agent game G with permissive joint strategy π, where both agents share the partition map O and an opinion-pairing map Γ satisfying Assumption 1. If Γ satisfies the conditions of Corollary 1, the attention gain obeys ρ>ρ⋆=d2κρ\!>\!ρ \!=\! d2κ, and the time-scale separation of Assumption 2 holds, then the layered opinion-guided strategy σiopσ^op_i built from Γ via (20) achieves robust coordination with respect to π against every π-rational opponent, whether reactive or non-reactive, and whether or not it runs the identical strategy as σiopσ^op_i. Proof. By Corollary 1, σiopσ^op_i is robust against every π-rational opponent, reactive or non-reactive: any joint action realized at a history h with π(h)≠∅π(h)\!≠\! lies in π(h)π(h). It thus remains to show that the opinion layer realizes one within each period. With ρ>ρ⋆=d2κρ\!>\!ρ \!=\! d2κ, the neutral opinion = z\!=\!0 of (22) is unstable, so from almost every initial condition the opinions converge to a stable equilibrium aligned with ±max=±[1γ]⊤± v_ \!=\!±[1\ γ] , whose signs form a matched pair z~i=γz~j z_i\!=\!γ z_j that solves the coupling (18) (against a non-reactive opponent, its committed action fixes z~j z_j and ziz_i converges to the matched sign). By Assumption 2, these signs settle within the reserved window and the action executes before the next decision time, so a joint action is realized at h; by robustness, (ai,aj)∈π(h)(a_i,a_j)\!∈\!π(h). ∎ VI Practical Implementation and Case Studies Fig. 11: Block-diagram view of the communication-free opinion-guided coordination framework. Each agent (agent 1 in blue, agent 2 in red) runs an opinion-guided strategy (σiopσ^op_i) in the strategy layer, opinion dynamics (NODi) in the opinion layer, and an opinion estimator (Ψi _i) together with a tracking controller (μi _i) acting on the physical plant (fif_i) in the physical layer. The two control stacks are coupled only through the jointly observable physical state x=(x1,x2)x\!=\!(x_1,x_2): each opinion estimator Ψi _i infers the opponent’s opinion z^j z_j from x and feeds it into NODi, so the opinion loop closes through shared observation rather than direct inter-agent communication. VI-A Communication-Free Implementation The opponent-coupling term in (13) requires agent i to read its opponent’s opinion state zjz_j directly, which presumes inter-agent communication. In practice, such communication is often impractical or unavailable. Yet human drivers coordinate without it, because each agent’s continuous motion already discloses the action it is committing to. The opinion-guided strategy admits a communication-free implementation, depicted in Fig. 11. The communication-free implementation extends the layered opinion-guided strategy in (13) by adding a physical layer. In the physical layer, agent i has a physical state xix_i governed by the plant dynamics x˙i=fi(xi,ui) x_i\!=\!f_i(x_i,u_i), and a low-level tracking controller μi _i maps the committed action aia_i to an input ui=μi(x,ai)u_i\!=\! _i(x,a_i) that steers xix_i along the motion realizing aia_i. The two agents share only the jointly observable state x=(x1,x2)∈x\!=\!(x_1,x_2)\!∈\!X. From x, each opinion estimator obtains an estimate z^j=Ψi(x,h) z_j\!=\! _i(x,h) of the opponent’s opinion. In practice, the estimator Ψi _i can be implemented with a variety of techniques, such as Kalman filtering, particle filtering, or learning-based methods. Unlike the realization in (13), the communication-free implementation replaces the directly read opponent’s opinion zjz_j with the estimate z^j z_j. The loop therefore closes without any inter-agent communication: the strategy layer commits an action from the current opinion state z~i z_i and the history h assembled from the abstracted states s=α(x)s\!=\!α(x) via a state abstraction map α:→α\!:\!X\!→\!S; the tracking controller drives the plant to realize it, so the agent’s own motion discloses that action through the shared state x; the opinion estimator infers the opponent’s opinion z^j=Ψi(x,h) z_j\!=\! _i(x,h) from x; and the opinion dynamics evolve under this estimated coupling signal, feeding the resulting opinion state back to the strategy layer for the next decision. The coordination guarantees of Section V hold if z^j z_j is accurate enough for the opinion dynamics to converge to the same matched pair of signs. Owing to the fast and robust bifurcation of the opinion dynamics, the two agents are driven to achieve a pair of actions in the permissive set π(h)π(h). To deploy the proposed opinion-guided layered control framework in a specific application, the following components must be instantiated. (i) Strategy layer: from the desired joint task and its specifications, first construct a permissive joint strategy π collecting every joint action pair that accomplishes the task, which in turn determines the rational strategy set Σj _j rat of agent j; then design each agent’s strategy σiopσ^op_i from a partition O of the opponent’s action set and an opinion-pairing map Γ that together satisfy the conditions of Corollary 1, so that σiopσ^op_i coordinates robustly with respect to π and Σj _j rat. (i) Opinion layer: opinion dynamics that undergo the intended bifurcation and converge quickly and robustly, as required by the time-scale separation of Assumption 2. (i) Physical layer, which realizes the committed actions on the plant: a state abstraction map α that recovers the discrete game state from the shared physical state, a low-level controller μi _i that tracks and executes the committed action, and an opinion estimator Ψi _i that infers the opponent’s opinion from the shared physical observation. To demonstrate how the layered framework coordinates decentralized agents, the three case studies below ground it in settings of growing complexity: Case-1 resolves a single conflict over a continuous action set with a scalar opinion; Case-2 spans multiple task stages with a multi-dimensional opinion; and Case-3 extends to a sequential general-sum game played over many decision steps. VI-B Case-1: On-Ramp Merging In the single-lane ramp merging scenario as shown in Fig. 12, the ego vehicle faces an opponent vehicle whose behavior is uncertain: the opponent may either yield or maintain its speed, which determines whether the ego vehicle should merge in front of it or behind it. For simplicity, we consider only longitudinal dynamics and model each agent i∈e,oi\!∈\!\e,o\ as a single integrator p˙i=vi p_i\!=\!v_i, where pip_i is the longitudinal position along its lane and the speed vi∈[vimin,vimax]v_i\!∈\![v_i ,v_i ] is the control input. In this case, the joint state is the pair of vehicle positions s=(pe,po)s\!=\!(p_e,p_o), which coincides with the physical state; the abstraction map α is therefore the identity, and no state abstraction is required. Each agent’s action is a constant speed held over one decision interval, and its action set is the continuous feasible range Ai=[vimin,vimax]A_i\!=\![v_i ,v_i ], i∈e,oi\!∈\!\e,o\. We take both the permissive strategy π and each agent’s individual strategy to be memoryless (Markov), so every decision depends only on the current state s. The permissive set π(s)⊆Ae×Aoπ(s)\! \!A_e\!×\!A_o is determined by the merge geometry. The time for vehicle i to reach its merge point pimp_i^m at constant speed is τi=(pim−pi)/vi _i\!=\!(p_i^m\!-\!p_i)/v_i. The joint speed (ve,vo)(v_e,v_o) is permissive when the following vehicle, on reaching the merge point, trails the leader by at least a safety margin dsafed_safe, π(s)=(ve,vo)∈Ae×Ao∣ π(s)\!=\! \(v_e,v_o)\!∈\!A_e\!×\!A_o ve(τo−τe)≥dsafe v_e( _o\!-\! _e)\!≥\!d_safe (24) or vo(τe−τo)≥dsafe. v_o( _e\!-\! _o)\!≥\!d_safe \. The first inequality gives a front merge for the ego vehicle, and the second a rear merge. Fig. 13 draws π(s)π(s) at a representative state as the green region. The rational action set Ai(s)A_i rat(s) of each agent is the projection of π(s)π(s) onto that agent’s speed axis, drawn as the light-colored line along each axis. To design the opinion-guided strategy, we first let the partition function iO_i split each agent’s action set at its midpoint vith=12(vimin+vimax)v_i^th\!=\! 12(v_i +v_i ), i∈e,oi\!∈\!\e,o\, into the two cells Ai(−)=[vimin,vith)A_i^(-)\!=\![v_i ,\,v_i^th) and Ai(+)=[vith,vimax]A_i^(+)\!=\![v_i^th,\,v_i ]. Merging is an anti-coordination task: the two vehicles resolve the conflict by taking opposite roles, one accelerating to pass through the merge point first and the other yielding to merge behind. We encode this structure with the constant opinion-pairing map Γ=−1 \!=\!-1, i.e. γ=−1γ\!=\!-1 in (22). Evaluating the robust action set (14) on the inferred opponent opinion gives Ae(s,z~o)A_e rob(s; z_o), the ego speeds that remain permissive against every rational opponent speed in that cell, shown in Fig. 13. The ego estimates the opponent’s speed by differencing the observed position over a window δ, v¯o=[po(t)−po(t−δ)]/δ v_o\!=\![p_o(t)\!-\!p_o(t\!-\!δ)]/δ, and thresholds it to recover the opponent’s opinion state, z^o=sgn(v¯o−voth)=sgn(po(t)−po(t−δ)−vothδ) z_o\!=\!sgn( v_o\!-\!v_o^th)\!=\!sgn(p_o(t)\!-\!p_o(t\!-\!δ)\!-\!v_o^thδ). Driven by this estimate, the opinion layer runs the scalar dynamics (22) with ρ past the bifurcation threshold (23), converging to a matched opinion state that commits the ego to Ae(s,z~o)A_e rob(s; z_o). The ego then takes the speed within it closest to its reference verefv_e^ref, ve=argminv∈Ae(s,z~o)|v−veref|,v_e\;=\; *arg\,min_v∈ A_e rob(s; z_o)\, v-v_e^ref , (25) so it stays near the reference speed whenever coordination does not force it away. To verify that the designed strategy is effective across the range of rational opponents established in Corollary 1, we evaluate it against two families of opponents. The first is a non-reactive individual strategy that commits to a fixed speed sampled uniformly at random from AoA_o. The second runs the same opinion-guided strategy as the ego but with a randomly sampled prior bias bob_o; as noted in Section V, this bias spans the range from a fully reactive to an effectively committed opponent, so the two families together cover the range of Corollary 1. Over 100100 trials in each setting, the ego completes a safe merge in every trial, keeping the joint speed within the permissive set and the merge separation above dsafed_safe, and it commits to the role compatible with the opponent: against a fixed-speed strategy it adapts to the committed role, and against a biased opinion-guided opponent it commits to the role opposite the one the opponent’s bias drives it toward. To stress-test the strategy further, we consider a time-varying opponent that alternates between high and low speeds. As shown in Fig. 12, the ego adaptively selects its speed from the robust action set of the opponent’s current behavior; despite the rapid switching, it reacts promptly and keeps the joint state within the permissive set throughout (Fig. 13). Fig. 12: One-lane ramp merging against a speed-switching opponent. Top: road scene at successive time instants, where the ego (blue) merges into the lane of the fluctuating opponent (red). Middle: ego and opponent speeds, with the ego reference speed marked by the dotted line. Bottom: the ego opinion state z, whose sign commits the ego to a front (z>0z>0) or rear (z<0z<0) merge. (a) t=1.5t=1.5 s (b) t=2.5t=2.5 s Fig. 13: Permissive and robust action sets in the joint speed space Ae×AoA_e\!×\!A_o, at two instants of the run in Fig. 12. The green region is the permissive set π(s)π(s); the grey diagonal band is the unsafe region where the vehicles reach the merge point closer than dsafed_safe. Dashed lines mark the partition thresholds vothv_o^th and vethv_e^th; the light segments on the two axes are the agents’ rational sets Ae(s)A_e rat(s) and Ao(s)A_o rat(s), while the dark segments on the ego axis are its robust action sets Ae(s,z~o)A_e rob(s; z_o). The green dot is the executed joint speed at this state. VI-C Case-2: Decentralized Dual-Arm Manipulation Fig. 14: Decentralized dual-arm manipulation guided by the opinion state. Top: translucent snapshots of the two arms (arm 1 blue, arm 2 red) at successive instants along their trajectories. (a) Reaching: the arms commit to opposite left/right halves of the table to grasp their cubes. (b)–(c) Transporting: the arms carry the cubes through the shared central space toward the baskets, additionally separating along the upper/lower axis. (d) Returning: the arms retract to their initial poses. Bottom: evolution of each arm’s two opinions, with the solid curves the left/right opinion zi[0]z_i[0] and the dashed curves the upper/lower opinion zi[1]z_i[1]. As a second case study, we evaluate the proposed framework in a tabletop manipulation scenario, simulated in Gazebo [30], in which two Franka Research 3 (FR3) arms [19] operate concurrently within a shared workspace. The two arms are mounted on opposite sides of a table and controlled independently, as shown in Fig. 14(a). Each of the two baskets starts with two cubes of the other basket’s color, and the task is to sort them so that every cube ends in the basket matching its own color. Since an arm carries one cube at a time, it transports its two cubes in succession, and every transport crosses the whole table, so the two arms may enter the shared central space at the same time and collide or block each other. A natural way to coordinate the arms in this shared space is to keep them as far apart as possible, so that at any instant they occupy complementary parts of the workspace rather than the same region. To make this precise, we fix a table frame common to both arms and partition the workspace along two independent axes of this frame. We write ci∈2c_i\!∈\!B^2 for the region arm i occupies: the first entry encodes which lateral half of the table the arm lies in, left or right in this common frame, and the second whether it keeps to the upper or lower part of the shared central space, so that one arm passes above the other where the two cross. Each arm’s action is the region it commits to next, so an action carries the same encoding as cic_i and Ai=2A_i\!=\!B^2. The partition iO_i is then the identity and each cell Ai(z~i)A_i^( z_i) holds a single action, in contrast to Case-1, where iO_i collapses a continuum of speeds onto two cells. The permissive set π(s)π(s) collects the joint actions in which the two arms hold complementary regions along the axis on which they must separate, and the committed region is realized in the physical layer, where a sampling-based planner generates a path confined to that region for the low-level controller μi _i to track. Each arm maintains an opinion zi∈ℝ2z_i\!∈\!R^2, a pair of scalar opinions, one per axis, whose signs form the opinion state z~i=sgn(zi)∈2 z_i\!=\!sgn(z_i)\!∈\!B^2 that commits the arm to a region along the left/right and upper/lower axes. Each arm’s task unfolds in five stages: reach, grasp, transport, place, and return, with the first four repeated once per cube, so the arms complete two rounds before the final return. The arms coordinate through a shared game state: the abstraction α maps their configurations and grasp status to s=(k,c1,c2)s\!=\!(k,c_1,c_2), where k is the current stage and cic_i is the region of the partitioned workspace that arm i’s end-effector lies in. Different stage requirements impose different separations on the two arms. While reaching, they must separate laterally to grasp from opposite halves of the table; while transporting, they must separate vertically to cross the shared central space without blocking each other; while grasping, placing, and returning, they need not separate at all, since their workspaces do not overlap. The arms therefore share one pairing map throughout, Γ=−I \!=\!-I, and both run the opinion dynamics (13) with a per-component attention ρi=diag(ρi,1,ρi,2) _i\!=\!diag( _i,1, _i,2) that decides which of its two separations is produced: a component with ρi,k>ρ⋆ _i,k\!>\!ρ bifurcates and the arms take opposite sides along it, while a component with ρi,k<ρ⋆ _i,k\!<\!ρ stays at the neutral opinion and no separation is made. Fig. 14 shows how the opinion guides the two arms to coordinate in the shared workspace. In (a), the left/right opinion bifurcates and drives the arms to opposite lateral halves of the table to grasp their cubes, rather than reaching into the same half. In (b) and (c), the upper/lower opinion bifurcates and drives the arms to opposite halves of the shared central space, so that one passes above and the other below, and they cross without blocking each other. Notably, the arm passing above in the first crossing passes below in the second, so the upper/lower assignment is resolved anew at each crossing rather than fixed in advance. The opinion estimator recovers the other arm’s opinion from the observed motion of its end-effector, so the two arms can coordinate without any communication. The estimator’s design is not a contribution of this paper, so we omit the details for space. We compare the proposed strategy against two baselines, running 2020 trials of each. The first removes the opinion guidance by fixing Γ(s) (s) to the zero matrix in every stage; with the two arms no longer coupled, they contend for the same region and block each other, and the task succeeds in only 33 of the 2020 trials. The second replaces the opinion guidance with a fixed priority: one arm is the designated leader that always takes its preferred region while the other yields to an alternative. This breaks the symmetry, but does so rigidly: the follower must yield even when it is better placed to proceed, which wastes motion whenever the two arms prefer the same region, and stalls the task entirely if the leader is delayed, since the follower keeps waiting for an arm that cannot move; it succeeds in 1111 of the 2020 trials. The proposed strategy succeeds in all 2020. VI-D Case-3: Autonomous Navigation in Constrained Space In the first two case studies, the application itself dictates which joint behaviors are admissible, so the permissive strategy is constructed from an intuitive understanding of the practical task. In this case study, we instead ground the framework in a classic problem of game theory, a sequential general-sum game, where the admissible joint behaviors correspond to a well-defined solution concept, the Nash equilibria of the game. For such a well-defined game problem, we showcase how the proposed opinion-guided strategy and its layered realization resolve the long-standing equilibrium selection problem noted in the introduction, which arises whenever the game admits more than one equilibrium. VI-D1 Navigation as a General-Sum Game Specifically, we consider a two-agent navigation problem in a constrained environment (Fig. 1), the setting of the autonomous-vehicle deadlocks reported in recent news [28, 44]. We formulate this problem as a sequential general-sum game as follows. Let xi∈Xi⊆ℝ2x_i\!∈\!X_i\! \!R^2 denote the planar position of vehicle i, which serves as its physical state. Both vehicles operate in the same constrained environment, so X1=X2=X_1\!=\!X_2\!=\!X and the joint physical state space is =X×XX\!=\!X\!×\!X. A straightforward abstraction is a grid, as illustrated in Fig. 15 and Fig. 16(a): both vehicles share one map α0:X→S _0\!:\!X\!→\!S that sends a position to its grid state, and applying it to both gives the joint abstraction α:→α\!:\!X\!→\!S with =S×SS\!=\!S\!×\!S, which returns the pair of grid states (s1,s2)(s_1,s_2) of the two vehicles. For a two-vehicle navigation, the grid abstraction need not cover the whole environment: a ×33\!×\!3 map suffices to capture the local interaction. Each vehicle’s action set consists of five semantic maneuvers, Ai=forward,back,left,right,waitA_i=\ forward,\, back,\, left,\, right,\, wait\, defined in a map frame common to both vehicles as the unit displacements (0,1)(0,1), (0,−1)(0,-1), (−1,0)(-1,0), (1,0)(1,0), and (0,0)(0,0); the transition relation → is thus deterministic: from each state s, each joint action (a1,a2)(a_1,a_2) leads to a unique successor s′s , written s→(a1,a2)s′s (a_1,a_2)s . Let Sobs⊂S_obs\!⊂\!S denote the obstacle grid states, which no vehicle can drive through. The task fails if a vehicle enters an obstacle grid state or the two vehicles collide, and the states in which either has occurred form the failure set, fail:=s∈|s1=s2 or s1,s2∩Sobs≠∅.S_fail\!:=\! \\,s\!∈\!S\; |\;s_1\!=\!s_2\ or \ \s_1,s_2\∩ S_obs\!≠\! \, \. (26) VI-D2 Nash Equilibrium Set as the Permissive Strategy As in real-world deployments, each vehicle has a distinct navigation goal gi∈Sg_i\!∈\!S, assumed known to both vehicles. We carry it in the state: the augmented state of vehicle i is s¯i=(si,gi) s_i\!=\!(s_i,g_i), and the joint state is the pair s=(s¯1,s¯2)s\!=\!( s_1, s_2), on which each vehicle plays a standard Markov strategy ai=σi(s)a_i\!=\! _i(s). We keep writing S for the set of such augmented joint states, and s1,s2s_1,s_2 for the grid states of the two vehicles, as in (26). Note that each vehicle’s goal gig_i is constant throughout an interaction. Both vehicles share the same reward function R, each evaluated with respect to its own goal, which rewards goal attainment, penalizes collisions, and charges a constant step cost that discourages delay. Its full form is given in Appendix A-A. We assume a shared R to make the game homogeneous in its payoffs as well as its transitions (Def. 5), so that two strategies can differ only in how they select among equilibria, which keeps the later discussion clean. Since each vehicle optimizes the shared reward under its own goal, the two vehicles are neither fully cooperative nor purely adversarial: the game is general-sum. In a general-sum game, no single quantity ranks a joint maneuver as desirable, and all that can be required is that the two actions be best responses to each other, that is, a Nash equilibrium (NE), at which neither vehicle can improve its own return by changing its action alone. We therefore resolve the interaction by seeking NEs, as in other game-theoretic planning approaches for interactive driving [14, 34, 24, 16]. As noted above, general-sum games may admit multiple NEs. To retain this multiplicity, the permissive joint strategy π is computed to collect every Nash action pair of the stage game at each state. Standard Nash value iteration resolves this multiplicity at synthesis, through a selection rule that is left arbitrary [27, 23]. Algorithm 1, in which λ∈(0,1)λ\!∈\!(0,1) is the discount factor, departs from it in two ways. First, the algorithm returns a permissive strategy that retains the whole set of equilibria rather than committing to one of them. The operator Nash(Q1,Q2) Nash(Q_1,Q_2) in Line 1 enumerates every pure-strategy Nash action pair of the stage game (Q1,Q2)(Q_1,Q_2) by exhaustive best-response checking over the |A1|×|A2||A_1|\!×\!|A_2| joint actions [4, 21], and Line 1 averages their returns, treating each equilibrium as equally plausible. This preserves behavioral flexibility at execution: a vehicle selects among the retained Nash actions according to the opponent’s opinion, rather than being locked into a single equilibrium the opponent may not share. Second, the algorithm does not distinguish the two vehicles by their labels, which avoids selecting an equilibrium during the iteration. Where standard Nash value iteration keeps one value function per player, Algorithm 1 keeps a single V, because relabeling the two vehicles leaves the game unchanged (Def. 5). Accordingly, vehicle 2’s value at s is not a second function but the same one evaluated at τ(s)τ(s). Because of these two points, the resulting permissive strategy privileges neither vehicle and retains every equilibrium, so the choice among them is left to execution rather than made at synthesis. The algorithm terminates when the value function converges to within a threshold θ or after kmaxk_ sweeps through the state space. Initialize V(s)←0V(s)← 0 for every state s=(s¯1,s¯2)s=( s_1, s_2); 1 for sweep k=1,…,kmaxk=1,…,k_ do 2 Vold←V_old← V; 3 foreach state s=(s¯1,s¯2)s=( s_1, s_2) do 4 if s∈fails _fail then V(s)←0V(s)\!←\!0, continue; 5 Q1,Q2←|A1|×|A2|Q_1,Q_2 0^|A_1|×|A_2|; 6 foreach a=(a1,a2)∈A1×A2a=(a_1,a_2)∈ A_1× A_2 do 7 s→s′s\! a\!s ; 8 Q1(a)←R(s,a)+λVold(s′)Q_1(a)← R(s,a)+λ\,V_old(s ); 9 Q2(a)←R(τ(s),τ(a))+λVold(τ(s′))Q_2(a)← R(τ(s),τ(a))+λ\,V_old(τ(s )); 10 π(s)←Nash(Q1,Q2)π(s)← Nash(Q_1,Q_2), K(s)←|π(s)|K(s)←|π(s)|; 11 V(s)←1K(s)∑(a1∗,a2∗)∈π(s)Q1(a1∗,a2∗)V(s)← 1K(s) _(a_1^*,\,a_2^*)∈π(s)Q_1(a_1^*,a_2^*); 12 if maxs|V(s)−Vold(s)|<θ _s|V(s)\!-\!V_old(s)|\!<\!θ then break; 13 return V, π Algorithm 1 Nash value iteration Remark 2. Algorithm 1 returns a well-defined permissive strategy only if the iteration converges and every state outside failS_fail retains at least one pure equilibrium, K(s)≥1K(s)\!≥\!1. Neither is guaranteed for general-sum stochastic games [27, 52], and characterizing when they hold is outside the scope of this paper. We therefore assume both hold at every state and every iteration in our case study. VI-D3 Opinion-Guided Strategy Design The permissive joint strategy is the equilibrium set Algorithm 1 returns at each state, π(s)=(ak1∗,ak2∗)k=1K(s)⊆A1×A2,π(s)\;=\; \(a_k^1*,a_k^2*) \_k=1^K(s)\; \;A_1\!×\!A_2, (27) where K(s)=|π(s)|K(s)\!=\!|π(s)| counts the equilibria at s. Each vehicle’s rational action set is then the projection of π(s)π(s) onto its own action axis, Ai(s)=aki∗k=1K(s)A_i rat(s)\!=\!\a_k^i*\_k=1^K(s). Given π(s)π(s), the pair (,Γ)(O, ) can be designed to satisfy the conditions of Corollary 1 in each state. On the ×33\!×\!3 grid, every feasible state typically retains one or two equilibria. Where there is a single NE, each vehicle’s rational set is a singleton, and no opinion guidance is needed. Where there are two, O partitions each vehicle’s action set into two cells, one for each equilibrium, and Γ couples the two opinions so that the committed pair forms one of the retained equilibria. The resulting opinion-guided strategy then selects between the two equilibria according to the inferred opinion of the opponent. Finally, the attention parameter is held constant above the bifurcation threshold (23), ρ1=ρ2>ρ⋆ _1\!=\! _2\!>\!ρ . VI-D4 Communication-Free Implementation To achieve the communication-free realization, we instantiate the tracking controller and the opinion estimator in the physical layer. Specifically, each vehicle’s dynamics are represented by a kinematic bicycle model, and the tracking controller μi _i is realized as a model predictive controller (MPC) that tracks the reference state associated with the committed action aia_i; the model and the MPC formulation are given in Appendix A-B. The opinion estimator Ψi(x(t),h(t)) _i(x(t),h(t)) is designed around a simple idea: check which of vehicle j’s rational actions best matches its observed velocity. The full form is given in Appendix A-C. (a) Resolution A. (b) Resolution B. Fig. 15: Two symmetric coordination outcomes produced by the proposed framework. Despite sharing an identical strategy, the vehicles spontaneously assume distinct roles via opinion-dynamics-driven symmetry breaking. The green lines mark the boundaries of the grid abstraction. VI-D5 Results Fig. 15 shows the two symmetric resolutions that the framework produces across runs of the same encounter, one for each of the game’s two equilibria. In Resolution A (Fig. 15(a)), the blue vehicle gives way, clearing the corridor by pulling into the side path; the red vehicle drives through, and the blue vehicle then reverses out and continues toward its goal. In Resolution B (Fig. 15(b)), the roles are exactly reversed. In both runs, the opinion states start near neutrality and bifurcate to opposite signs, committing the two vehicles to one of the two retained equilibria; which one is selected is decided by the real-time evolution of the opinions. No communication is involved: each vehicle infers the other’s opinion from its observed motion, so the agreement on roles emerges through shared observation alone. These results demonstrate the effectiveness of the opinion-guided strategy in the two-agent game and of its communication-free realization. Furthermore, we evaluate coordination across all pairings of four strategy types in the scenario of Fig. 15. The three baselines select the equilibrium without regard to the encountered opponent: NE-A and NE-B commit a vehicle to its role in the corresponding resolution of that figure, and rule-based applies a priority rule, giving right of way to the vehicle closer to the conflict region. The proposed opinion-guided strategy selects during the interaction, with each vehicle’s prior bias bib_i sampled independently. Table I reports the success rate of every pairing of these strategy types, over 100100 independent runs each, with vehicle 1 running the row strategy and vehicle 2 the column strategy. As expected, two vehicles committed to the same equilibrium achieve the desired coordination in every run (100%100\%), whereas an NE-A vehicle meeting an NE-B vehicle clashes in every run (0%0\%). The rule-based strategy selects its equilibrium adaptively from the real-time situation, but each vehicle evaluates a rigid criterion on its own observations. Against a fixed-equilibrium opponent, that criterion takes no account of what the opponent has committed to, so the two may settle on different equilibria and the success rate falls to 4848–65%65\%. Against itself the rule does better (76%76\%), since both vehicles evaluate the same condition and usually reach the same equilibrium; the failures that remain occur when the two sit close to that condition’s decision boundary, where noise and delay can pull their verdicts apart. The rule-based strategy is thus not robust to the opponent it meets. The opinion-guided strategy instead adapts to whichever opponent it meets, and the bifurcation is what makes that adaptation robust: the neutral opinion state is unstable, and the only stable ones are the two that assign the vehicles complementary roles, so the opinions are driven into one of them and stay there. A transient misreading of the opponent perturbs that trajectory without changing where it settles, which is exactly what a rigid criterion cannot do. The strategy succeeds in all 100100 runs of every pairing. This result demonstrates that the opinion-guided strategy is robust to the opponent’s strategy type as long as its behavior is rational (Corollary 1). TABLE I: Success rate (%) between different strategy types, over 100100 independent runs per pairing. Vehicle 2 Vehicle 1 NE-A NE-B Rule-based Opinion-guided NE-A 100 0 52 100 NE-B 0 100 65 100 Rule-based 48 51 76 100 Opinion-guided 100 100 100 100 (a) Overall trajectory (b) Submap 1 (c) Submap 2 Fig. 16: Large-scale validation in a constrained environment: (a) overall trajectory of the fleet; (b)–(c) zoomed-in views of two representative interaction segments. Finally, we scale the scenario from a single encounter to a fleet of ten vehicles navigating a constrained environment that resembles a dense residential area or parking lot, each with its own start and goal. Fig. 16(a) shows the executed trajectories: drivable grid states are white, and grid states blocked by buildings are grey. Each vehicle is assumed to meet at most one other at a time, within the ×33\!×\!3 neighborhood map around the contested passage, and the permissive strategy for each such neighborhood is generated by Algorithm 1. The opinion-guided strategy runs on top of it without communication. These neighborhoods are not alike: a corner, an intersection, a one-lane bridge, and the cycle around a central block each induce a different local game, with its own permissive set and its own number of equilibria. The same procedure synthesizes the opinion-guided strategy in every case, so the framework is not tuned to one geometry. From the trajectories, we can see that every vehicle reaches its goal without collision or deadlock, and every passage that admits more than one resolution is settled during the encounter itself. Two of them are enlarged. In Fig. 16(b), vehicles 0 and 1 round the central block and contend for the same grid state; their opinions bifurcate to opposite signs, so one claims the grid state and drives on while the other lets it pass before continuing. In Fig. 16(c), vehicles 2 and 5 meet at an intersection, where vehicle 5 claims the passage and drives straight through while vehicle 2 concedes and proceeds once the intersection clears. Throughout the run, the layered opinion-guided strategy decides on its own when coordination is needed, which equilibrium the two vehicles settle on, and which role each of them takes. The effectiveness across these different geometries demonstrates the generality of the approach, which resolves various encounter scenarios without fixed priorities, hand-designed rules, or communication. VI-E Discussion This paper is motivated by a critical observation: in practical interaction between agents, there is rarely a single admissible joint behavior, and when each agent selects the one it prefers, the two choices may be incompatible and coordination fails. This work departs from the common premise of a known or communicating opponent and adopts a weaker assumption: the opponent is rational, in the sense that at every moment it is willing to take an action that can be part of an admissible joint behavior. On this basis, we formalized the decentralized coordination problem between two agents through a series of definitions, among them the permissive joint strategy and rationality. We then introduced the requirements that a strategy must satisfy when agents carrying different strategies meet at random: robust coordination, which asks the strategy to remain compatible with every rational opponent that respects the shared permissive set, and homogeneous implementability, which asks it to remain effective when both agents run that same strategy. To meet both requirements, we presented one key insight for designing such a strategy: keep every admissible joint behavior open, preserving the flexibility to adapt to whichever opponent is encountered, and decide the specific one at runtime from how that opponent behaves. Building on this insight, we proposed the opinion-guided strategy and its layered realization. As the case studies show, the framework resolves miscoordination among agents with differing preferences and breaks symmetry among structurally identical agents, without fixed rules, priorities, or communication. In addition, several problems remain open for future work: (i) The present analysis addresses the interaction of two agents over a binary decision, and extending the opinion-guided strategy to more agents and more options is not straightforward. For more than two agents, the bifurcation of the opinion dynamics no longer guarantees that the agents will settle on complementary roles. Although a vector opinion can in principle encode more than two options, Assumption 1 couples each of its components to at most one component of the other agent’s opinion, so the components evolve as independent binary decisions; a single choice among more than two mutually related options, which requires dependence among the components, is therefore out of its reach. (i) The geometry of the permissive set varies from state to state, and with it what the interaction demands of the partition and the pairing map. So far the two have been designed either from a simple principle, as in Cases 1 and 3, or by hand, as in Case 2. How to design the opinion-pairing map Γ and the partition O so that they satisfy the conditions of Corollary 1 for an arbitrary permissive geometry, and in a general multi-agent setting, remains open. Resolving these two points would generalize the opinion-guided strategy to interactions among more agents and to more complex scenarios. VII Conclusions This paper studied decentralized coordination between agents that meet at random, without communication and without knowledge of each other’s strategy, so that the strategies they carry may prove incompatible. To make the problem precise, we introduced a formal formulation built on two notions: the permissive joint strategy, which collects all the admissible joint behaviors, and rationality, which constrains an agent’s individual behavior. On that basis, we stated the requirements that a strategy must meet in decentralized coordination: robust coordination, which asks it to remain compatible with every rational opponent that respects the shared permissive set, and homogeneous implementability, which asks it to remain effective when both agents run that same strategy. To satisfy both requirements, we proposed an opinion-guided strategy that couples the two agents through an auxiliary opinion state, realized in a layered framework: the strategy layer keeps every admissible joint behavior available and commits to an action once the other agent’s opinion is inferred, while coupled opinion dynamics drive the two agents to a matched pair of opinions. A layer-by-layer analysis established the conditions under which two individually synthesized strategies are compatible, and thereby guarantees that randomly encountered agents realize an acceptable joint behavior. Three case studies of increasing decision complexity, ramp merging, dual-arm manipulation, and navigation in a constrained environment, showed the framework resolving conflicts and bringing the agents to a common admissible joint behavior, without relying on communication or predesigned coordination. Notably, the last of these corresponds to a typical general-sum game admitting multiple Nash equilibria: there the opinion dynamics guided the two agents to a common equilibrium from the opponent’s real-time behavior alone, and kept them compatible whether that opponent ran a different strategy or the identical one, provided only that it was rational. Appendix A Implementation Details of Case-3 A-A Reward Function The joint state s=(s¯i,s¯j)s\!=\!( s_i, s_j) collects the augmented states s¯i=(si,gi) s_i\!=\!(s_i,g_i) and s¯j=(sj,gj) s_j\!=\!(s_j,g_j), each pairing a vehicle’s grid state with its goal, and the joint action is a=(ai,aj)a\!=\!(a_i,a_j); the vehicle listed first is taken as the ego. Let si′s_i denote the grid state that vehicle i reaches from sis_i under action aia_i, and likewise sj′s_j . The shared reward of Section VI is R(s,a)= R(s,a)= −1+20⋅[si′=gi]−50⋅[si′∉S] -1+20·1[s_i \!=\!g_i]-50·1[s_i \!∉\!S] (28) −100([si′∈Sobs]+[si′=sj′] -100 (1[s_i \!∈\!S_obs]+1[s_i \!=\!s_j ] +[sj=si′∧si=sj′]). +1[s_j\!=\!s_i s_i\!=\!s_j ] ). The constant step cost encourages progress over delay, reaching the goal pays 2020, and leaving the grid costs −50-50. The three −100-100 charges cover the collision events: entering an obstacle grid state, both vehicles entering the same grid state, and the head-on exchange in which each enters the grid state the other vacates. Since R reads only the ego goal gig_i, it is the swap τ that makes the two vehicles’ payoffs differ: the same function, evaluated at s for one vehicle and at τ(s)τ(s) for the other, scores the interaction in terms of their respective goals. A-B Tracking Controller of Case-3 The full physical state of vehicle i is ξi=(xi,φi,vi) _i=(x_i, _i,v_i), which extends the planar position xi∈ℝ2x_i\!∈\!R^2 with the heading φi _i and the speed viv_i. Its motion follows the kinematic bicycle model x˙i=vi[cosφisinφi],φ˙i=viLtanδi,v˙i=ηi, x_i=v_i bmatrix _i\\ _i bmatrix, _i= v_iL _i, v_i= _i, (29) where L is the wheelbase and the input ui=(ηi,δi)u_i=( _i, _i) collects the longitudinal acceleration and the steering angle. For digital implementation, the dynamics (29) are discretized into fdf_d at the sampling time Δtc t_c of the physical layer, with Δtc≪tk+1−tk t_c\! \!t_k+1\!-\!t_k. Given the action aia_i, the tracking controller μi _i solves a finite-horizon problem over N steps of fdf_d to track the reference state ξir _i^r associated with aia_i: minui,0:N−1 _u_i,0:N-1 ∑n=0N−1(‖ξi,n−ξir‖W2+‖ui,n‖Wu2)+‖ξi,N−ξir‖Wf2 _n=0^N-1 (\| _i,n- _i^r\|_W^2+\|u_i,n\|_W_u^2 )+\| _i,N- _i^r\|_W_f^2 (30) s.t. .t. ξi,n+1=fd(ξi,n,ui,n),ξi,0=ξi(t), _i,n+1=f_d( _i,n,u_i,n), _i,0= _i(t), ui,n∈,n=0,…,N−1, u_i,n , n=0,…,N-1, where W,Wf⪰0W,W_f\! \!0 and Wu≻0W_u\! \!0 are weighting matrices, and U bounds the acceleration and the steering angle. The reference ξir _i^r encodes the committed action aia_i. The first input of the minimizing sequence is applied and the problem is re-solved at the next sampling instant, in a receding-horizon fashion. A-C Opinion Estimator of Case-3 Each maneuver in the action set Aj=forward,back,left,right,waitA_j\!=\!\ forward,\, back,\, left,\, right,\, wait\ induces a nominal motion at the current state; let vj(a)v_j^(a) denote the reference velocity of maneuver a∈Aja\!∈\!A_j for vehicle j, and vj(t)v_j(t) the observed velocity of vehicle j at time t. The similarity of the observed velocity to maneuver a is quantified on a common [0,1][0,1] scale as dja(t)=max0,⟨vj(t),vj(a)⟩‖vj(t)‖‖vj(a)‖+ϵ,a≠wait,max0, 1−‖vj(t)‖v¯,a=wait.d_j^a(t)= cases \0,\ v_j(t),\,v_j^(a) \|v_j(t)\|\,\|v_j^(a)\|+ε \,&a≠ wait,\\[8.0pt] \0,\ 1- \|v_j(t)\| v \,&a= wait. cases (31) Here v¯ v is the nominal speed. Thus, a moving maneuver is scored by direction alone and wait by proximity to rest, placing the two on the same scale. Each opinion state z~ z is then scored by the best-matching rational action in Aj(z~)A_j^( z): dj(z~)(t)=maxa∈Aj(s)∩Aj(z~)dja(t)d_j^( z)(t)= _a∈ A_j rat(s)∩ A_j^( z)d_j^a(t). Thus, the opinion estimate is z^j(t)=kzdj(+1)(t)−dj(−1)(t)(dj(+1)(t)+dj(−1)(t))2+ϵ, z_j(t)=k_z\, d_j^(+1)(t)-d_j^(-1)(t) (d_j^(+1)(t)+d_j^(-1)(t) )^2+ε, (32) where kz>0k_z>0 is a scaling factor and ϵ>0ε>0 ensures numerical stability. Both cell scores are non-negative, so |z^j(t)|<kz| z_j(t)|\!<\!k_z, and the sign of z^j(t) z_j(t) records which cell the observed motion falls in. References [1] N. K. Alghamdi and S. Park (2025) Opinion-driven decision-making for multi-robot navigation through narrow corridors. arXiv preprint arXiv:2504.20947. Cited by: §I-A. [2] S. H. Arul and D. Manocha (2021) V-RVO: decentralized multi-agent collision avoidance using voronoi diagrams and reciprocal velocity obstacles. In 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 8097–8104. External Links: Document Cited by: §I-A. [3] T. Badings, S. Junges, A. Marin, M. Suilen, and M. Volk (2025) Robust permissive controller synthesis for interval MDPs. arXiv preprint arXiv:2510.03481. Cited by: §I-A. [4] T. Başar and G. J. Olsder (1998) Dynamic noncooperative game theory. SIAM. Cited by: §IV-A, §VI-D2. [5] J. Bernet, D. Janin, and I. Walukiewicz (2002) Permissive strategies: from parity games to safety games. Vol. 36, p. 261–275. Cited by: §I-A. [6] A. Bizyaeva, A. Franci, and N. E. Leonard (2022) Nonlinear opinion dynamics with tunable sensitivity. IEEE Transactions on Automatic Control 68 (3), p. 1415–1430. Cited by: §I-A, §IV-B1, §IV-B, §V-B. [7] Y. Bramoullé (2007) Anti-coordination and social interactions. Games and Economic Behavior 58 (1), p. 30–49. Cited by: §I-A. [8] L. Brice, J. Raskin, and M. van den Bogaard (2025) Permissive equilibria in multiplayer reachability games. In 33rd EACSL Annual Conference on Computer Science Logic (CSL 2025), LIPIcs. Cited by: §I-A, §I-A. [9] C. Cathcart, M. Santos, S. Park, and N. E. Leonard (2023) Proactive opinion-driven robot navigation around human movers. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 4052–4058. Cited by: §I-A. [10] M. Chen, J. C. Shih, and C. J. Tomlin (2016) Multi-vehicle collision avoidance via hamilton-jacobi reachability and mixed integer programming. In 2016 IEEE 55th Conference on Decision and Control (CDC), p. 1695–1700. External Links: Document Cited by: §I-A. [11] F. Christianos, G. Papoudakis, and S. V. Albrecht (2023) Pareto actor-critic for equilibrium selection in multi-agent reinforcement learning. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §I-A. [12] G. De Giacomo, A. Di Stasio, L. M. Tabajara, M. Y. Vardi, and S. Zhu (2022) Synthesis of maximally permissive strategies for LTLf specifications. In Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence (IJCAI), p. 2783–2789. Cited by: §I-A. [13] G. Farina, N. Gatti, and T. Sandholm (2018) Practical exact algorithm for trembling-hand equilibrium refinements in games. Advances in neural information processing systems 31. Cited by: §I-A. [14] J. F. Fisac, E. Bronstein, E. Stefansson, D. Sadigh, S. S. Sastry, and A. D. Dragan (2019) Hierarchical game-theoretic planning for autonomous vehicles. In 2019 International conference on robotics and automation (ICRA), p. 9590–9596. Cited by: §I-A, §VI-D2. [15] D. Fisman, O. Kupferman, and Y. Lustig (2010) Rational synthesis. In 16th International Conference on Tools and Algorithms for the Construction and Analysis of Systems (TACAS), LNCS, Vol. 6015, p. 190–204. Cited by: §I-A. [16] D. Fridovich-Keil, E. Ratner, L. Peters, A. D. Dragan, and C. J. Tomlin (2020) Efficient iterative linear-quadratic approximations for nonlinear multi-player general-sum differential games. In 2020 IEEE international conference on robotics and automation (ICRA), p. 1475–1481. Cited by: §VI-D2. [17] W. Fu, C. Yu, Z. Xu, J. Yang, and Y. Wu (2022) Revisiting some common practices in cooperative multi-agent reinforcement learning. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, p. 6863–6877. Cited by: §I-A. [18] J. Grover, C. Liu, and K. Sycara (2023) The before, during, and after of multi-robot deadlock. The International Journal of Robotics Research 42 (6), p. 317–336. Cited by: §I-A. [19] S. Haddadin, S. Parusel, L. Johannsmeier, S. Golz, S. Gabl, F. Walch, M. Sabaghian, C. Jähne, L. Hausperger, and S. Haddadin (2022) The franka emika robot: a reference platform for robotics research and education. IEEE Robotics & Automation Magazine 29 (2), p. 46–64. Cited by: §VI-C. [20] J. C. Harsanyi and R. Selten (1988) A general theory of equilibrium selection in games. Vol. 1, MIT Press, Cambridge, MA. Cited by: §I-A. [21] J. P. Hespanha (2017) Noncooperative game theory: an introduction for engineers and computer scientists. Princeton University Press, Princeton, NJ. Cited by: §I-A, §VI-D2. [22] H. Hu, K. Nakamura, K. Hsu, N. E. Leonard, and J. F. Fisac (2023) Emergent coordination through game-induced nonlinear opinion dynamics. In 2023 62nd IEEE Conference on Decision and Control (CDC), p. 8122–8129. Cited by: §I-A. [23] J. Hu and M. P. Wellman (2003) Nash q-learning for general-sum stochastic games. Journal of machine learning research 4 (Nov), p. 1039–1069. Cited by: §VI-D2. [24] Z. Huang, T. Li, S. Shen, and J. Ma (2024) Integrated decision making and trajectory planning for autonomous driving under multimodal uncertainties: a bayesian game approach. arXiv preprint arXiv:2409.13993. Cited by: §VI-D2. [25] T. Ingebrand, S. Smith, and U. Topcu (2023) Decentralized conflict resolution for multi-agent reinforcement learning through shared scheduling protocols. In 2023 62nd IEEE Conference on Decision and Control (CDC), p. 7170–7177. Cited by: §I-A. [26] D. Kappler, F. Meier, J. Issac, J. Mainprice, C. G. Cifuentes, M. Wüthrich, V. Berenz, S. Schaal, N. Ratliff, and J. Bohg (2018) Real-time perception meets reactive motion generation. IEEE Robotics and Automation Letters 3 (3), p. 1864–1871. Cited by: footnote 1. [27] M. Kearns, Y. Mansour, and S. Singh (2000) Fast planning in stochastic games. In Proceedings of the Sixteenth conference on Uncertainty in artificial intelligence, p. 309–316. Cited by: §VI-D2, Remark 2. [28] KHOU 11 (2024) Video: self-driving cars cause brief traffic jam in montrose. Note: YouTube video External Links: Link Cited by: §I, §VI-D1. [29] M. J. Kochenderfer (2015) Decision making under uncertainty: theory and application. MIT press. Cited by: §I-A. [30] N. Koenig and A. Howard (2004) Design and use paradigms for gazebo, an open-source multi-robot simulator. In 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vol. 3, p. 2149–2154. Cited by: §VI-C. [31] H. Kress-Gazit, G. E. Fainekos, and G. J. Pappas (2009) Temporal-logic-based reactive mission and motion planning. IEEE Transactions on Robotics 25 (6), p. 1370–1381. Cited by: footnote 1. [32] O. Kupferman, G. Perelli, and M. Y. Vardi (2016) Synthesis with rational environments. Annals of Mathematics and Artificial Intelligence 78 (1), p. 3–20. Cited by: §I-A. [33] N. E. Leonard, A. Bizyaeva, and A. Franci (2024) Fast and flexible multiagent decision-making. Annual Review of Control, Robotics, and Autonomous Systems 7. Cited by: §I-A, §I-B, §IV-B2. [34] N. Li, I. Kolmanovsky, A. Girard, and Y. Yildiz (2018) Game theoretic modeling of vehicle interactions at unsignalized intersections and application to autonomous vehicle control. In 2018 Annual American Control Conference (ACC), p. 3215–3220. Cited by: §I-A, §VI-D2. [35] F. A. Oliehoek C. Amato et al. (2016) A concise introduction to decentralized pomdps. Vol. 1, Springer. Cited by: §I-A. [36] A. Ozdaglar, M. O. Sayin, and K. Zhang (2023) Independent learning in stochastic games. In International Congress of Mathematicians, p. 5340–5373. Cited by: §I-A. [37] S. Park, A. Bizyaeva, M. Kawakatsu, A. Franci, and N. E. Leonard (2021) Tuning cooperative behavior in games with nonlinear opinion dynamics. IEEE Control Systems Letters 6, p. 2030–2035. Cited by: §I-A. [38] A. Pierson, W. Schwarting, S. Karaman, and D. Rus (2020) Weighted buffered voronoi cells for distributed semi-cooperative behavior. In 2020 IEEE International Conference on Robotics and Automation (ICRA), p. 5611–5617. Cited by: §I-A. [39] A. R. Pritchett and A. Genton (2017) Negotiated decentralized aircraft conflict resolution. IEEE Transactions on Intelligent Transportation Systems 19 (1), p. 81–91. Cited by: §I-A. [40] S. E. S. A. R. Project (2020) European atm master plan – digitalising europe’s aviation infrastructure – executive view. External Links: Document Cited by: §I-A. [41] S. Qi, Z. Tang, Z. Sun, and S. Haesaert (2025) Integrating opinion dynamics into safety control for decentralized airplane encounter resolution. In 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Vol. , p. 173–178. Cited by: §I-A, §I-B. [42] R. Selten (1975) Reexamination of the perfectness concept for equilibrium points in extensive games. International Journal of Game Theory 4 (1), p. 25–55. Cited by: §I-A. [43] C. Tomlin, G.J. Pappas, and S. Sastry (1998) Conflict resolution for air traffic management: a study in multi-agent hybrid systems. IEEE Transactions on Automatic Control 43 (4), p. 509–521. External Links: Document Cited by: §I-A. [44] USA TODAY (2024) Watch: video shows waymo robotaxis honking at one another. Note: YouTube video External Links: Link Cited by: §I, §VI-D1. [45] L. Wang, A. D. Ames, and M. Egerstedt (2017) Safety barrier certificates for collisions-free multirobot systems. IEEE Transactions on Robotics 33 (3), p. 661–674. Cited by: §I-A. [46] T. Wang, T. Gupta, A. Mahajan, B. Peng, S. Whiteson, and C. Zhang (2021) RODE: learning roles to decompose multi-agent tasks. In International Conference on Learning Representations, Cited by: §I-A. [47] B. Weng, H. Chen, and W. Zhang (2022) On the convergence of multi-robot constrained navigation: a parametric control lyapunov function approach. In 2022 International Conference on Robotics and Automation (ICRA), p. 4972–4978. Cited by: §I-A. [48] H. Zhang, W. Chen, Z. Huang, M. Li, Y. Yang, W. Zhang, and J. Wang (2020) Bi-level actor-critic for multi-agent coordination. In Proceedings of the AAAI conference on artificial intelligence, Vol. 34, p. 7325–7332. Cited by: §I-A. [49] K. Zhang, Z. Yang, and T. Başar (2021) Multi-agent reinforcement learning: a selective overview of theories and algorithms. Handbook of reinforcement learning and control, p. 321–384. Cited by: §I-A. [50] R. Zhang, J. Shamma, and N. Li (2024) Equilibrium selection for multi-agent reinforcement learning: a unified framework. arXiv preprint arXiv:2406.08844. Cited by: §I-A. [51] M. Zinkevich and T. R. Balch (2001) Symmetry in markov decision processes and its implications for single agent and multiagent learning. In Proceedings of the eighteenth international conference on machine learning, p. 632. Cited by: §I-C, §I-C, §I-C, §I-C. [52] M. Zinkevich, A. Greenwald, and M. Littman (2005) Cyclic equilibria in markov games. Advances in neural information processing systems 18. Cited by: §I-A, Remark 2.