Paper deep dive
ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration
Abhishek Sriraman, Eleni Vasilaki, Robert Loftin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 4/26/2026, 11:58:52 PM
Summary
ConventionPlay is a reinforcement learning-based approach designed for robust ad-hoc collaboration in multi-agent systems. It extends the cognitive hierarchy model by introducing a diverse population of adaptive followers (Level-1 agents) with varied capability limits. Unlike traditional methods that train agents to simply adapt to fixed conventions, ConventionPlay trains a Level-2 agent to actively probe and steer interactions toward the most effective joint strategy within a partner's repertoire. Experimental results in Repeated Matrix Games and Point Mass Rendezvous tasks demonstrate that ConventionPlay achieves superior coordination efficiency, particularly in settings with differentiated payoffs, by transitioning from passive adaptation to active team steering.
Entities (8)
Relation Signals (5)
Abhishek Sriraman → affiliatedwith → University of Sheffield
confidence 100% · Abhishek Sriraman University of Sheffield
ConventionPlay → evaluatedon → Repeated Matrix Game
confidence 100% · We evaluate our approach on two multi-agent environments... The first is a Repeated Matrix Game
ConventionPlay → evaluatedon → Point Mass Rendezvous
confidence 100% · The second is Point Mass Rendezvous (PMR)
ConventionPlay → extends → Cognitive Hierarchy
confidence 100% · extends cognitive hierarchies to include a diverse population of adaptive followers.
ConventionPlay → uses → Dec-POMDP
confidence 90% · We model multi-agent tasks as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Ad-hoc collaboration often relies on identifying and adhering to shared conventions. However, when partners can follow multiple conventions, agents must do more than simply adapt; they must actively steer the team toward the most effective joint strategy. We present ConventionPlay, a reinforcement learning-based approach that extends cognitive hierarchies to include a diverse population of adaptive followers. By training against partners with varied capability limits, our agent learns to probe its partner's repertoire, leading the team when possible and following when necessary. Our results in canonical coordination tasks show that ConventionPlay achieves superior coordination efficiency, particularly in settings where conventions have differentiated payoffs.
Tags
Links
- Source: https://arxiv.org/abs/2604.18123v1
- Canonical: https://arxiv.org/abs/2604.18123v1
Trouble viewing inline? Open PDF directly →
Full Text
32,077 characters extracted from source content.
Expand or collapse full text
ConventionPlay: Capability-Limited Training for Robust Ad-Hoc Collaboration Abhishek Sriraman University of Sheffield Sheffield, United Kingdom asriraman1@sheffield.ac.uk Eleni Vasilaki University of Sheffield Sheffield, United Kingdom e.vasilaki@sheffield.ac.uk Robert Loftin University of Sheffield Sheffield, United Kingdom r.loftin@sheffield.ac.uk ABSTRACT Ad-hoc collaboration often relies on identifying and adhering to shared conventions. However, when partners can follow multiple conventions, agents must do more than simply adapt; they must actively steer the team toward the most effective joint strategy. We present ConventionPlay, a reinforcement learning-based approach that extends cognitive hierarchies to include a diverse population of adaptive followers. By training against partners with varied capability limits, our agent learns to probe its partner’s repertoire, leading the team when possible and following when necessary. Our results in canonical coordination tasks show that ConventionPlay achieves superior coordination efficiency, particularly in settings where conventions have differentiated payoffs. KEYWORDS Ad-Hoc Collaboration; Multi-Agent Reinforcement Learning 1 INTRODUCTION Cooperative multi-agent tasks often admit a variety of different so- lutions, with success depending on the agents’ ability to coordinate their strategies. In such tasks, joint strategies can exhibit different levels of compatibility, with performance degrading when individ- ual strategies are misaligned. We describe collections of mutually compatible strategies as “conventions”. Within the broader chal- lenge of designing agents to collaborate with unseen teammates [12], the ability to recognize and align with established conven- tions is crucial for effective ad-hoc collaboration [9,15,16]. In this work, we propose a reinforcement learning approach to training collaborative agents that go beyond simply adapting to their part- ner’s convention. Instead, when teamed with adaptive partners, our agents actively steer the interaction towards the most effective joint strategy in the repertoire of conventions that those partners are capable of following. As a motivating example, consider the setting where two humans that have not previously interacted need to communicate with each other to solve a task. They each have varying levels of proficiency with overlapping sets of languages. At the start of their interaction, they are motivated to identify a single preferred language that allows them to successfully coordinate on the underlying task. Here the languages can be viewed as "conventions", and aligning on the preferred language is part of the coordination task. A human in this scenario may try a few different languages to identify the relative proficiencies of their partner. The feedback they receive from this interaction – correctness, or willingness of their partner to speak a Proc. of the Adaptive and Learning Agents Workshop (ALA 2026), Aydeniz, Del- grange, Mohammedalamen, Yang (eds.), May 25 – 26, 2026, Paphos, Cyprus, https://alaworkshop2026.github.io/. 2026. certain language – will inform the choice of a common language to use for the rest of the interaction. A common approach taken in previous work on ad-hoc collabo- ration [2,17,19] is to train collaborative agents through a two-step process that first attempts to train a diverse “population” of agents for the target task, and then trains a “best-response” agent capable of coordinating with any member of this population. In the first phase, a population of teams of agents that can solve the task to- gether is generated—for example, all members of a team may speak a common language. Members of a team are trained in self-play, and may be incompatible with members of other teams. Agents in this initial population effectively act as “leaders” which choose a fixed convention and expect other agents to adapt to this convention. In the second phase, a generalist agent is trained against this entire population, such that it learns to adapt to conventions followed by its partners—for example, a multi-lingual agent that adapts to whichever language its partner speaks. Since these generalist agents are trained against partners who follow fixed conventions, they effectively become “followers” that adapt to their partner’s conven- tion. A shortcoming of this approach is that such followers will fail to account for partners that are themselves capable of some degree of adaptation, and may not be limited to a single (potentially suboptimal) convention. In this work, we introduce ConventionPlay, a reinforcement learning-based approach to training adaptive agents that are capa- ble of coordinating with partners that may themselves be adaptive. The novelty of this algorithm lies in a structured training population that includes a diverse set of generalist followers—agents designed to adapt to specific, restricted subsets of conventions. Exposing our final agent to these adaptive partners forces it to move beyond purely reactive behavior; it must actively probe a partner’s behavior to decide when to lead the team toward an optimal strategy, and when to follow a partner’s established convention. Our experimen- tal results demonstrate that, by explicitly training agents that are capable of both leading and following, ConventionPlay outperforms state-of-the-art methods in complex ad-hoc collaboration scenarios that involve multiple, mutually incompatible conventions. 2 RELATED WORK To contextualize the development of ConventionPlay, we review relevant literature in population diversity for ad-hoc collaboration, convention-free strategies, and partner shaping. Population Diversity. The most common method for achieving robust ad-hoc collaboration is by anticipating the range of possible solutions to the task. A number of methods focus on computing a diverse set of solutions and training a generalist agent against this population. TrajeDI [11] achieves this by incentivizing policies to arXiv:2604.18123v1 [cs.MA] 20 Apr 2026 generate distinct trajectories. MEP [19] encourages the population to cover the entire state-action space. LIPO [3] aims to learn policies that are incompatible with each other, yet individually optimal in self-play. Fictitious Co-Play (FCP) [17] provides behavioral diver- sity by including checkpoints from the self-play training trajectory, simulating partners of varying skill levels. Our work focuses not on the discovery of these base conventions, but on leveraging the diversity of the base population to improve performance of a gen- eralist agent. To that end, our proposed method is compatible with any of these diversity-generation techniques. Convention-Free Strategies. Unlike methods that seek to generate diverse conventions, Other-Play [7] attempts to eliminate the need for them entirely. It constructs "convention-free" policies by sys- tematically avoiding arbitrary coordination choices, relying solely on the strategic structure of the game. While effective in abstract settings, we argue that avoiding conventions is often impossible in practice, particularly in Human-AI scenarios where social norms dictate behavior. Instead of avoiding conventions, our approach embraces them, aiming to construct agents capable of recognizing and coordinating with any convention present in a population. Partner Shaping. We first distinguish our work from the social influence approach of Jaques et al. [8]. While they employ intrinsic rewards to explicitly encourage influence, our steering behavior emerges naturally as a solution to the meta-RL task of partner iden- tification, avoiding the conflation of task success with auxiliary rewards. Similarly, LOLA [6] and M-FOS [10] explore actions that influence partner behavior, but they are fundamentally designed to shape a partner’s learning process over time; they assume the partner’s policy weights will shift in response to the agent’s actions. In contrast, ConventionPlay treats the partner’s underlying policy as fixed but potentially multi-modal. Crucially, we do not attempt to model or access the partner’s internal weights or learning pa- rameters. Our steering actions are intended to probe a partner’s existing repertoire to identify and converge upon the most efficient shared convention. 3 BACKGROUND This section introduces the core formalisms required for our ap- proach. We define the Dec-POMDP framework, formalize the con- cept of conventions, and describe the best-response and cognitive hierarchy models that form the basis of our training methodology. 3.1 Dec-POMDPs We model multi-agent tasks as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) [13], defined by the tuple M= ⟨푆,퐴,푃,푅,Ω,푂,훾⟩.푆is the set of global states,퐴= 퐴 1 × 퐴 2 is the joint action space, and푃(푠 푡+1 |푠 푡 ,푎 푡 )defines the transition probabilities to a new state푠 푡+1 given the current state푠 푡 and joint action푎 푡 ∈ 퐴. The task is fully cooperative, with both agents receiving the same scalar reward푅(푠 푡 ,푎 푡 ). The joint observation 표 푡 ∈Ω=Ω 1 ×Ω 2 is drawn according to the function푂(표 푡 |푠 푡 ,푎 푡−1 ). An individual policy휋 푖 :H 푖,푡 →Δ(퐴 푖 )maps the agent’s local action-observation history,ℎ 푖,푡 = (표 푖,0 ,푎 푖,0 ,표 푖,1 , . . .,푎 푖,푡−1 ,표 푖,푡 ) ∈ H 푖,푡 , to a probability distribution over its action space. A joint policy 휋= (휋 1 ,휋 2 )specifies the behavior of both agents. The goal is to maximize the expected discounted return: 퐽(휋)=E 푠 0 ,푎 푡 ∼휋,푠 푡+1 ∼푃 " ∞ ∑︁ 푡=0 훾 푡 푅(푠 푡 ,푎 푡 ) # (1) where훾 ∈ [0,1)is a discount factor. We denote the self-play (SP) return as퐽 푆푃 (휋 1 )= 퐽(휋 1 ,휋 1 ). We assume throughout this work that each agent inMhas identical action and observation spaces, such that an individual policy 휋 푖 can control any agent 푗 . 3.2 Conventions We formalize a system of conventions as a collectionC=퐶 1 , . . .,퐶 푘 , where each convention퐶 푖 ⊂Πis a subset of the policy space. These sets are defined using two coordination margins,휖and훿, which bound the performance of intra- and inter-convention pairings: (1)Intra-convention휖-compatibility: For any convention 퐶 푖 ∈ C, any two policies 휋,휋 ′ ∈ 퐶 푖 must satisfy: 퐽(휋,휋 ′ ) ≥ min(퐽(휋,휋), 퐽(휋 ′ ,휋 ′ ))−휖(2) where휖 ≥0 is a small bound. This implies that policies within the same convention are functionally compatible. (2) Inter-convention훿-incompatibility: For any two distinct conventions퐶 푖 ,퐶 푗 ∈ C(푖≠ 푗), any pair of individual policies 휋 ∈ 퐶 푖 and 휋 ′ ∈ 퐶 푗 must satisfy: 퐽(휋,휋 ′ ) ≤ min(퐽(휋,휋), 퐽(휋 ′ ,휋 ′ ))−훿(3) where훿> 휖. The parameter훿represents a coordination gap, where a mismatch in conventions leads to a drop in utility. Together, these properties define conventions as compatibility classes in policy space. Importantly, two conventions퐶 푖 and퐶 푗 may yield identical self-play returns—퐽 푆푃 (휋)= 퐽 푆푃 (휋 ′ )for휋 ∈ 퐶 푖 ,휋 ′ ∈ 퐶 푗 — yet remain strictly 훿 -incompatible. 3.3 Best Response Training Given a finite set of training partnersD, a best response휋is defined as the policy that maximizes the expected joint return: 휋 ∈ BR(D)= arg max 휋∈Π E 휋 ′ ∼D [퐽(휋,휋 ′ )](4) where the expectation is taken over a uniform distribution overD. This objective ensures the agent optimizes for robust performance across the entire population. In our framework, these best responses are Bayes-adaptive; the agent treats the partner’s identity휋 ′ ∈ Das a latent variable that must be inferred from the interaction history. Importantly, a generalist best response may achieve lower utility than a policy specialized for a single partner, reflecting the inherent cost of disambiguating the partner’s strategy during the interaction. 3.4 Cognitive Hierarchies The Cognitive Hierarchy (CH) [1] and K-level reasoning [4] models offer a scaffolding for structuring the policy space in multi-agent systems [5]. In this hierarchical approach, each level represents a leap in reasoning depth, where a level푘policy accounts for the behaviors of its lower-level counterparts. We utilize this hierarchy to establish a distribution of diverse agents, where each level defines a specific set of collaborative capabilities and assumptions. We denote individual level-푘policies as휋 푘 ∈ P 푘 and sets of such policies asD 푘 . Level-0 (퐾 0 ): Static Convention Policies. The퐾 0 population, P 0 , comprises policies that each adhere to a fixed convention퐶 푖 ∈ C . We denote a퐾 0 policy adhering to convention푖as휋 0,푖 ∈ 퐶 푖 . By definition, these agents are brittle: they expect a partner to match their specific mode of coordination and lack the flexibility to adapt to others. This level aligns with the populations generated by methods like MEP or TrajeDI; however, unlike SyKLRBR—which uses random퐾 0 policies to avoid conventions—we explicitly use퐾 0 to represent the diverse conventions present in ad-hoc scenarios. Level-1 (퐾 1 ): Adaptive Best Responses. The퐾 1 population, P 1 , comprises policies defined as best responses to a specific dis- tribution of퐾 0 partners:휋 1 ∈ BR(D 1 ), whereD 1 ⊆ P 0 . While existing ad-hoc collaboration methods [3,17] typically produce generalist퐾 1 agents whereD 1 spans the entire퐾 0 population, our methodology allows for more granularity. Specifically, we first de- fine a capability setK 휋 1 ⊆ Crepresenting a restricted subset of conventions, and then defineD 1 = Ð 퐶∈K 휋 1 퐶 as the corresponding set of퐾 0 policies. This allows us to model partners with varying degrees of flexibility. To quantify how effectively an agent adapts to its partner’s specific limitations, we introduce the Coordination Efficiency 휂 for a pair of agents(휋,휋 푘 ) as: 휂(휋,휋 푘 )= 퐽(휋,휋 푘 ) 퐽 ∗ (휋 푘 ) (5) where퐽 ∗ (휋 푘 )= max ˆ 휋∈Π 퐽( ˆ 휋,휋 푘 )represents the maximum pos- sible return achievable with partner휋 푘 . For a static퐾 0 partner, 퐽 ∗ (휋 0 ) ≈ 퐽 푆푃 (휋 0 ). This metric allows us to assess whether an agent achieves the full potential of its partner’s repertoire. Level-2 (퐾 2 ): Robust Ad-hoc Coordinators. The퐾 2 agent is trained as a best response to the full diversity of the hierarchy: 휋 2 ∈ BR(D 2 ), whereD 2 = P 0 ∪P 1 . By optimizing against this mixed population, the퐾 2 agent is incentivized to adopt a dual strat- egy: adhering to the fixed conventions of rigid퐾 0 partners while proactively signaling and steering toward optimal conventions when paired with flexible퐾 1 partners. To evaluate this capability, we define the optimal achievable return퐽 ∗ (휋 1 )for an adaptive part- ner as the maximum joint return possible with any policy belonging to a convention within that partner’s repertoire: 퐽 ∗ (휋 1 )= max 퐶 푖 ∈K 휋 1 max 휋 ′ ∈퐶 푖 퐽(휋 ′ ,휋 1 ) (6) This formulation shifts the evaluation from global optimization to partner-aware coordination, testing whether an agent can realize the full potential of a partner’s specific repertoire. 4 METHOD This section provides an overview of the ConventionPlay algorithm, and details on how the prerequisite퐾 0 and퐾 1 populations are cre- ated. ConventionPlay trains an ad-hoc agent to be a best response to a diverse population of퐾 0 and퐾 1 partners. The algorithm follows three main steps: (1) generating a diverse set of퐾 0 policies to repre- sent different base conventions; (2) training several퐾 1 agents, each adapted to specific subsets of the퐾 0 population; and (3) training the final퐾 2 agent to coordinate effectively across this entire hierarchy. The training hierarchy is illustrated in Figure 1. At its core, ConventionPlay relies on a퐾 1 population built from varied and restricted subsets of conventions. This structure forces Figure 1: Comparison of the ConventionPlay pipeline with existing ad-hoc collaboration methods. While baseline ap- proaches diversify the퐾 0 population, we introduce diversity among퐾 1 followers to force the퐾 2 agent to infer its part- ner’s limited repertoire and coordinate on the most effective shared convention. the퐾 2 agent to transition from passive adaptation to active team steering. To coordinate effectively, the agent must employ probing strategies to reveal its partner’s specific repertoire, allowing it to navigate the trade-offs between different coordination modes; by exposing the agent to partners that do not all support a single global optimum, we encourage the development of such a behavior. We describe the implementation of this hierarchy in the following sections, beginning with the generation of the base conventions. 4.1 Generating a Base Population We generate our퐾 0 population by training MAPPO [18] across mul- tiple random seeds. While specialized methods like TrajeDI or LIPO could also be used, we found that random initialization provided sufficient behavioral diversity. Notably, this approach does not guar- antee that the population will partition into the well-structured system of incompatible policies described in our formalism; in practice, we may encounter complex, non-transitive compatibility structures driven in part by the task or learning architecture. We therefore verify this inter-convention incompatibility empirically to ensure our agents provide a clear foundation of distinct strategies for subsequent training levels. 4.2 Capability-Aware Stratified Sampling To construct a퐾 1 population with diverse and limited capabilities, we define푀training subsetsD 1,1 , . . .,D 1,푀 through stratified sampling of the퐾 0 population. The number of subsets푀is a hy- perparameter that should ideally correspond to the number of con- ventions|C|in퐾 0 . This method is designed to impose a “capability ceiling” on each퐾 1 agent—a property that is particularly critical in tasks with differentiated reward structures where conventions vary in objective value. To generate these subsets, we first define the capability휌(휋 0,푖 )of an agent by its self-play return퐽 푆푃 (휋 0,푖 ). We then identify푀anchor agents휋 ∗ 0,푗 푀 푗=1 by selecting policies whose capabilities are closest to푀target performance levels lin- early spaced across the range[min휌, max휌]. For each anchor, the associated training distributionD 1,푗 is formed by sampling from an eligibility poolP elig =휋 0,푖 ∈ P 0 :휌(휋 0,푖 ) ≤ 휌(휋 ∗ 0,푗 ). By including only policies that perform no better than the anchor, we ensure the Figure 2: Payoff matrices for the Repeated Matrix Game. Left: Uniform payoffs across conventions. Right: Differentiated payoffs across conventions. Figure 3: Trajectories of agents playing two variants of the point-mass-rendezvous game. Left: Layout with 4 landmarks with equal value; no reward is achieved as agents navigated to different landmarks. Right: 4 landmarks with differenti- ated rewards, visualized by opacity; both agents successfully navigating to the landmark with the highest rewards. anchor represents the performance ceiling within that퐾 1 agent’s repertoire, forcing it to coordinate while respecting that limit. 5 EXPERIMENTAL SETUP 5.1 Environments We evaluate our approach on two multi-agent environments where coordination on shared conventions is critical. The first is a Repeated Matrix Game, a canonical task where agents must select matching actions across multiple trials. We con- sider two distinct configurations of this game, illustrated by the payoff matrices in Figure 2. In the first setting, all conventions are of equal value and are equally difficult to discover. In the second setting, conventions possess distinct values and present varying difficulties for discovery, incentivizing agents not just to find any solution but an optimal one. The second is Point Mass Rendezvous (PMR) [3], a time- extrapolated version of the repeated matrix game, similar to coop- erative reaching [14]. In this environment, agents must rendezvous at one of a set of pre-configured landmarks, as shown in Figure 3. The value of each landmark may differ, with some landmarks being more valuable than others. Agents observe their position, velocity, distance to each of the landmarks, and their partner’s velocity. As with the matrix game, we consider both uniform and differenti- ated reward structures for this task, testing the agent’s ability to coordinate on both arbitrary and value-driven conventions. 5.2 Evaluation Protocol Our evaluation protocol is designed to test ad-hoc teamwork capa- bilities across the cognitive hierarchy. For each environment, we construct a benchmark set of hard-coded policies that align with the convention system observed in the learned 퐾 0 population. We denote this evaluation set as퐾 0,test . Following our stratified sam- pling approach (Section 4.2), we then generate a 퐾 1,test population by training best responses against various subsets of퐾 0,test . Success against this adaptive evaluation set requires the agent to actively probe an unseen partner and identify the specific conventions they are capable of supporting. We benchmark ConventionPlay against three baseline methods. First, a Best Response (BR) is trained against the퐾 0 population to establish a baseline for simple adaptation. Second, FCP [17] in- troduces behavioral diversity through intermediate checkpoints; we implement this by wrapping the final퐾 0 agents in an휖-greedy policy with varying exploration rates. Finally, SyKLRBR [5] repre- sents state-of-the-art cognitive hierarchy modeling for zero-shot coordination; for structural parity, we use a two-level SyKLRBR hi- erarchy grounded in our learned퐾 0 population. All adaptive agents are trained using PPO with a recurrent architecture to manage the inherent partial observability of the coordination task. 6 RESULTS We begin by examining the generation of the퐾 0 and퐾 1 subsets, which form the training distribution for our adaptive agent; this is illustrated in Figure 4. The cross-play matrix of the퐾 0 population (top-left) reveals distinct clusters of mutually compatible (and in- compatible) policies. Our stratified sampling process (Section 4.2) successfully constructs a diverse population of퐾 1 agents, ensuring a broad distribution of maximum capabilities across all퐾 1 training sets (top-right). The low Jaccard similarity between subsets (bottom- left) and the broad agent coverage (bottom-right) illustrates that the퐾 1 population is exposed to a wide variety of behaviors without being redundant, or dominated by a single convention. We evaluate the performance of our generalist퐾 2 agent (Conven- tionPlay) against baselines. Table 1 presents the results across the evaluation domains. ConventionPlay shows parity with baselines for uniform games and in interaction with퐾 0 partners, demonstrating that our hierarchical training does not degrade robust convention adherence. The critical difference emerges in the퐾 1 regime for dif- ferentiated games. In these settings, baseline agents lack the probing mechanisms required to identify the extent of a partner’s flexibility. When paired with an adaptive퐾 1 follower, baselines often settle for the first discovered convention, even if a higher-value strategy exists within the partner’s capability setK 휋 1 . In contrast, Conven- tionPlay achieves significantly higher coordination efficiency휂by actively steering the interaction. As visualized in Figure 5, when the agent detects that a partner is capable of adaptation, it persists in signaling high-value conventions rather than immediately con- forming to the partner’s initial trajectory. This "probing" behavior Table 1: Coordination Efficiency휂of ZSC methods across four domains. Results show mean efficiency±standard deviation across three random seeds. ConventionPlay demonstrates parity with partners that play with fixed conventions (퐾 0 ) and superior performance when paired with adaptive (퐾 1 ) partners. Matrix Game (Uniform) Method퐾 0 Test퐾 1 TestSelf-Play BestResponse 77.50± 0.00 85.83± 3.19 78.75± 8.69 FCP61.25± 5.97 72.81± 2.21 82.36± 3.57 SyKLRBR77.50± 0.00 87.19± 2.90 90.00± 2.38 ConventionPlay 77.29± 0.15 89.79± 1.41 83.06± 3.36 Matrix Game (Differentiated) Method퐾 0 Test퐾 1 TestSelf-Play BestResponse47.81± 0.00 65.41± 0.26 70.55± 2.01 FCP47.29± 3.58 64.40± 3.75 79.29± 0.95 SyKLRBR41.67± 0.15 59.06± 0.29 91.65± 2.38 ConventionPlay 59.69± 2.27 75.23± 0.51 60.36± 1.28 PMR (Uniform) Method퐾 0 Test퐾 1 TestSelf-Play BestResponse98.68± 0.08 99.00± 0.05 99.84± 0.07 FCP96.27± 0.88 98.08± 0.15 97.59± 0.63 SyKLRBR 98.89± 0.19 98.93± 0.17 99.80± 0.16 ConventionPlay 98.63± 0.10 98.75± 0.11 99.96± 0.11 PMR (Differentiated) Method퐾 0 Test퐾 1 TestSelf-Play BestResponse92.55± 6.20 78.52± 3.14 59.09± 3.47 FCP94.43± 2.78 74.23± 2.92 49.45± 3.63 SyKLRBR94.32± 1.93 81.06± 1.46 63.85± 4.34 ConventionPlay 96.05± 0.44 88.82± 2.22 65.43± 2.36 Figure 4: Makeup of the generated퐾 1 subsets. The top left image has cross-play of all the퐾 0 population; clusters here roughly visualize compatible policies which can be considered conventions. The top middle shows how the stratified sampling has placed different agents in different clusters. The top right shows the maximum capability distribution across subsets, ensuring various퐾 1 subsets have different capabilities. Bottom left shows Jaccard similarity between all generated clusters. Bottom right shows enumerated subset sizes and the number of times a specific agent is used in the support of the entire 퐾 1 population. allows the team to converge on the optimal shared strategy allowed by the partner’s repertoire. 7 DISCUSSION AND FUTURE WORK In this work, we have introduced ConventionPlay as a method for robust ad-hoc collaboration. By introducing strategic diversity into the퐾 1 layer of a cognitive hierarchy, we have shown that agents can Figure 5: Visualizing the steering behavior of Convention- Play (blue) against퐾 0 (left) and퐾 1 (right) partners. In the퐾 0 case, the agent probes for a better convention but converges on the partner’s choice when no adaptation is detected. Con- versely, with a퐾 1 partner, the agent persists in moving toward the highest value goal, then the second highest value goal, successfully influencing the adaptive partner to switch con- ventions. learn to dynamically navigate the balance between leading and fol- lowing. The primary limitation in our method is the reliance on퐽 푆푃 for capability-aware sampling, which lacks discriminative power when distinct conventions yield similar returns. This can limit the variety of the퐾 1 population by failing to distinguish between strate- gically diverse behaviors. Integrating more granular metrics, such as trajectory diversity [11], would allow ConventionPlay to model different types of generalist followers. While our current evaluation focuses on cooperative settings, the principles of ConventionPlay could be extended to mixed-motive settings as well. In scenarios where agents have partially aligned in- centives or private utilities, the ability to probe a partner’s behavior and steer the interaction toward a mutually beneficial convention becomes even more critical for preventing sub-optimal outcomes. REFERENCES [1] Colin F. Camerer, Teck Ho, and Juin-Kuan Chong. 2003. A Cognitive Hierarchy Theory of One-shot Games and Experimental Analysis. SSRN Electronic Journal (2003). https://doi.org/10.2139/ssrn.411061 [2]Micah Carroll, Rohin Shah, Mark K. Ho, Thomas L. Griffiths, Sanjit A. Seshia, Pieter Abbeel, and Anca Dragan. 2020. On the Utility of Learning about Hu- mans for Human-AI Coordination. https://doi.org/10.48550/arXiv.1910.05789 arXiv:1910.05789 [cs] [3]Rujikorn Charakorn, Poramate Manoonpong, and Nat Dilokthanakul. 2023. Generating Diverse Cooperative Agents by Learning Incompatible Policies. In The Eleventh International Conference on Learning Representations. https: //openreview.net/forum?id=UkU05GOH7_6 [4]Miguel A. Costa-Gomes and Vincent P. Crawford. 2006. Cognition and Behavior in Two-Person Guessing Games: An Experimental Study. American Economic Review 96, 5 (December 2006), 1737–1768. https://doi.org/10.1257/aer.96.5.1737 [5] Brandon Cui, Hengyuan Hu, Luis Pineda, and Jakob Nicolaus Foerster. 2022. K-level Reasoning for Zero-Shot Coordination in Hanabi. ArXiv abs/2207.07166 (2022). https://api.semanticscholar.org/CorpusID:246996266 [6]Jakob N. Foerster, Richard Y. Chen, Maruan Al-Shedivat, Shimon Whiteson, P. Abbeel, and Igor Mordatch. 2017. Learning with Opponent-Learning Awareness. In Adaptive Agents and Multi-Agent Systems. https://api.semanticscholar.org/ CorpusID:8708073 [7]Hengyuan Hu, Adam Lerer, Alex Peysakhovich, and Jakob Foerster. 2020. Other Play. In Proceedings of the 37th International Conference on Machine Learning. PMLR, 4399–4410. [8]Natasha Jaques, Angeliki Lazaridou, Edward Hughes, Caglar Gulcehre, Pedro Ortega, Dj Strouse, Joel Z. Leibo, and Nando De Freitas. 2019. Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning. In Pro- ceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdi- nov (Eds.). PMLR, 3040–3049. https://proceedings.mlr.press/v97/jaques19a.html [9]Adam Lerer and Alexander Peysakhovich. 2019. Learning Existing Social Con- ventions via Observationally Augmented Self-Play. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society (Honolulu, HI, USA) (AIES ’19). Association for Computing Machinery, New York, NY, USA, 107–114. https://doi.org/10.1145/3306618.3314268 [10]Chris Lu, Timon Willi, Christian Schroeder de Witt, and Jakob Foerster. 2022. Model-Free Opponent Shaping.https://doi.org/10.48550/arXiv.2205.01447 arXiv:2205.01447 [cs] [11]Andrei Lupu, Brandon Cui, Hengyuan Hu, and Jakob Foerster. 2021. Trajectory Diversity for Zero-Shot Coordination. In Proceedings of the 38th International Conference on Machine Learning. PMLR, 7204–7213. [12] Reuth Mirsky, Ignacio Carlucho, Arrasy Rahman, Elliot Fosong, William Macke, Mohan Sridharan, Peter Stone, and Stefano V. Albrecht. 2022. A Survey of Ad Hoc Teamwork Research. arXiv:2202.10450 [cs.MA] https://arxiv.org/abs/2202.10450 [13]Frans A. Oliehoek and Christopher Amato. 2016. A Concise Introduction to Decentralized POMDPs. Springer International Publishing, Cham. https://doi. org/10.1007/978-3-319-28929-8 [14]Arrasy Rahman, Elliot Fosong, Ignacio Carlucho, and Stefano V Albrecht. 2023. Generating Teammates for Training Robust Ad Hoc Teamwork Agents via Best- Response Diversity. Transactions on Machine Learning Research (2023). https: //openreview.net/forum?id=l5BzfQhROl [15] Andy Shih, Arjun Sawhney, Jovana Kondic, Stefano Ermon, and Dorsa Sadigh. 2021. On the Critical Role of Conventions in Adaptive Human-AI Collaboration. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. https://openreview.net/forum?id=8Ln- Bq0mZcy [16]Peter Stone, Gal Kaminka, Sarit Kraus, and Jeffrey Rosenschein. 2010. Ad Hoc Autonomous Agent Teams: Collaboration without Pre-Coordination. Proceedings of the AAAI Conference on Artificial Intelligence 24, 1 (Jul. 2010), 1504–1509. https://doi.org/10.1609/aaai.v24i1.7529 [17] DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett. 2021.Collaborating with Humans without Human Data. In Ad- vances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 14502–14515. https://proceedings.neurips.c/paper_files/paper/2021/file/ 797134c3e42371b4979a462eb2f042a-Paper.pdf [18]Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and YI WU. 2022. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mo- hamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35. Curran Asso- ciates, Inc., 24611–24624. https://proceedings.neurips.c/paper_files/paper/2022/ file/9c1535a02f0ce079433344e14d910597-Paper-Datasets_and_Benchmarks.pdf [19] Rui Zhao, Jinming Song, Yufeng Yuan, Hu Haifeng, Yang Gao, Yi Wu, Zhongqian Sun, and Yang Wei. 2022. MEP – Maximum Entropy Population-Based Training for Zero-Shot Human-AI Coordination.https://doi.org/10.48550/arXiv.2112. 11701 arXiv:2112.11701 [cs]