Paper deep dive
The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency
Michael Glass
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/18/2026, 4:51:32 AM
Summary
The paper introduces the Open-Strategy Dictator Game (OSDG), a game-theoretic framework where agents submit natural-language strategies visible to all participants. A Large Language Model (LLM) acts as an oracle to adjudicate interactions by interpreting the dictator's strategy in the context of the recipient's. Through round-robin tournaments, the study demonstrates that conditionally cooperative strategies (which share with cooperators and take from exploiters) consistently dominate unconditional strategies (always share or always take). The results suggest that mutual transparency and legibility make conditional cooperation evolutionarily robust.
Entities (9)
Relation Signals (7)
Open-Strategy Dictator Game ā uses ā LLM Oracle
confidence 95% Ā· A large language model adjudicates each interaction by interpreting the dictator's strategy in the context of the recipient's.
Conditional Cooperation ā dominates ā Unconditional Strategies
confidence 94% Ā· Conditionally cooperative strategies... consistently dominate, while unconditional strategies (always share or always take) are weakly dominated.
Selfish ā istype ā Unconditional Defector
confidence 92% Ā· Selfish Always take. The unconditional defector.
Generous ā istype ā Unconditional Cooperator
confidence 92% Ā· Generous Always share. The unconditional cooperator.
Open-Strategy Dictator Game ā descendedfrom ā Program Equilibrium
confidence 88% Ā· This construction is descended from Tennenholtzās program equilibrium
Chivalry ā protects ā Unconditional Cooperator
confidence 85% Ā· Chivalry... A protective strategy addresses the same second-order problem by making the exploitation of unconditional cooperators costly
Cooperation Coalition ā punishes ā Unconditional Cooperator
confidence 85% Ā· Cooperation coalition... Punishes both defectors and those who subsidize defectors... s2 takes from c_bar (a subsidizer)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each player's strategy is a natural-language document visible to all participants. The dictator's decision, to SHARE or TAKE an endowment, may depend on the text of the recipient's strategy. A large language model adjudicates each interaction by interpreting the dictator's strategy in the context of the recipient's. We run round-robin tournaments among diverse strategies and analyze the resulting payoff matrix using softmax equilibrium frequencies, dominance analysis, and sensitivity to the relative value of cooperation. Conditionally cooperative strategies, those that share with cooperators and take from exploiters, consistently dominate, while unconditional strategies (always share or always take) are weakly dominated. The results suggest that in environments where agents can inspect each other's decision procedures, conditional cooperation is evolutionarily robust across a wide range of payoff parameters.
Tags
Links
- Source: https://arxiv.org/abs/2608.14913v1
- Canonical: https://arxiv.org/abs/2608.14913v1
Trouble viewing inline? Open PDF directly ā
Full Text
97,554 characters extracted from source content.
Expand or collapse full text
The Open-Strategy Dictator Game: Cooperation Under Mutual TransparencyThanks: Code, strategy documents, and full tournament data: https://github.com/michaelrglass/os-fdt. Michael Glass Email: michael.r.glass@gmail.com August 14, 2026 Abstract We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each playerās strategy is a natural-language document visible to all participants. The dictatorās decisionāto share or take an endowmentāmay depend on the text of the recipientās strategy. A large language model adjudicates each interaction by interpreting the dictatorās strategy in the context of the recipientās. We run round-robin tournaments among diverse strategies and analyze the resulting payoff matrix using softmax equilibrium frequencies, dominance analysis, and sensitivity to the relative value of cooperation. Conditionally cooperative strategiesāthose that share with cooperators and take from exploitersāconsistently dominate, while unconditional strategies (always share or always take) are weakly dominated. The results suggest that in environments where agents can inspect each otherās decision procedures, conditional cooperation is evolutionarily robust across a wide range of payoff parameters. 1 Introduction The dictator game is a foundational paradigm in experimental economics: one player (the dictator) unilaterally divides a fixed endowment between herself and a passive recipient (Kahneman et al. 1986). Unlike the prisonerās dilemma or ultimatum game, the recipient has no strategic recourseāthe dictatorās decision is final. Imagine two intelligent species meeting for the first time in deep space. This first contact is unlikely to be between near-equals. Civilizations arising independently have no reason to arrive at spaceflight within a century of one another, and a century is already enough to separate a species that can cross interstellar distances from one that cannot; against the timescales on which planets and species form, even a millennium is a rounding error. The technologically superior species holds almost all the power ā it decides whether to share knowledge, trade fairly, or exploit the weaker civilization entirely. Critically, the advanced species can observe the culture, communications, and decision norms of the less-advanced one. With analytical tools consistent with interstellar capability, it can form a remarkably accurate picture of the recipientās strategy. Perhaps many expansionist civilizations will be fully exploitive ā maximizing their own growth regardless of collateral destruction. But if civilizations can reason about each otherās possible actions and motives then such exploitation becomes risky. We propose that strategy transparency fundamentally changes the incentive landscape. When the dictator can observe and reason about the recipientās decision procedureāand knows that future dictators will observe theirsāthe game becomes one of mutual legibility. To formalize this, we introduce the Open-Strategy Dictator Game (OSDG). Each playerās strategy is a natural-language document specifying how it would allocate as dictator, potentially conditioned on the text of the recipientās strategy. All strategies are visible to all participants. A large language model (LLM) serves as an oracle, interpreting each dictatorās strategy in the context of the recipientās to produce a binary decision: share (split equally) or take (keep everything). This construction is descended from Tennenholtzās program equilibrium (Tennenholtz 2004), in which agents submit programs that can read one anotherās source code. The history of that idea motivates our design. Its canonical constructions do not analyze the opponentās code so much as compare it ā cooperation conditioned on syntactic equality ā which any behaviorally equivalent rewrite defeats. Proof-based cooperation (āshare if it is provable that the opponent reciprocatesā) has been achieved only for agents written in a purpose-built fragment of provability logic (Barasz et al. 2014); for general programs, formal verification of an adversarially authored opponent is hopeless in practice. In the one open tournament where arbitrary programs received each otherās source (Fallenstein 2013), attempted analysis failed so thoroughly that the top three finishers ignored the opponentās code and played randomly. A successor tournament responded by making opponent simulation a sandboxed language primitive (LessWrong community 2014) ā simulation of general programs otherwise founders on unbounded mutual regress ā and later theory identified stochastically grounded simulation, rather than proof, as the robust route to program equilibrium (Oesterheld 2019). A practical entrant today would complete this trajectory: submit a program that sends the opponentās source to an AI agent for analysis and acts on the resulting summary of its strategy ā a design now studied directly (Sistla and Kleiman-Weiner 2025). But at that point the program layer is vestigial. The artifact doing the work is the strategy description the interpreter reads, and the open-source game becomes the open-strategy game; the OSDG removes the wrapper and takes the description itself as the object of play. This choice also matches the epistemics of every practical setting: what one agent knows of the strategy of an organization, an AI, or a person is not a source listing but the product of an investigation ā observation, analysis, and summary. A strategy description is the general form of such an investigationās output, and interpreting one in context is precisely what the LLM oracle does. A side benefit is that the framework becomes accessible to human participants while preserving the core mechanism of mutual transparency. Our contributions are: 1. We define the OSDG and its payoff structure (Section 2). 2. We describe an LLM-adjudicated tournament protocol for evaluating strategies (Section 3). 3. We introduce softmax equilibrium frequencies as a model of both evolutionary selection and reflective selection and analyze the resulting equilibria (Section 7). 4. We present tournament results, dominance and admissibility analysis, and equilibrium computations (Sections 5ā8). 2 The Open-Strategy Dictator Game 2.1 Setup A tournament consists of n strategies =s1,s2,ā¦,snS=\s_1,s_2,ā¦,s_n\, each a natural-language document. A fixed endowment E>0E>0 is available in each round. Each ordered pair (i,j)(i,j) defines a round in which strategy i acts as dictator and strategy j as recipient. The dictator makes a binary decision diājāshare,taked_ijā\ share, take\: ⢠share: the endowment is split equally. The dictator receives E/2E/2; the recipient receives E/2E/2. ⢠take: the dictator keeps the entire endowment. The dictator receives E; the recipient receives 00. Each playerās utility for receiving value v is uā”(v)=lnā”(1+v),u(v)= (1+v), (1) a concave function capturing diminishing marginal returns. The total welfare uā”(vD)+uā”(vR)u(v_D)+u(v_R) is maximized by share: 2ālnā”(1+E2)>lnā”(1+E)+lnā”1=lnā”(1+E),2 \! (1+ E2 )> (1+E)+ 1= (1+E), (2) so the game is not zero-sum, and mutual cooperation is socially efficient. Definition 1 (Open-Strategy Dictator Game). An OSDG instance is a tuple (,E,)(S,E,O) where S is a set of natural-language strategies, E>0E>0 is the endowment, and O is an oracle that, given (si,sj)ā2(s_i,s_j) ^2, returns a decision diājāshare,taked_ijā\ share, take\ by interpreting strategy sis_i in the context of sjs_j. 2.2 Visibility The critical departure from the classic dictator game is that the dictator observes the full text of the recipientās strategy before deciding. This enables conditional strategies: a dictator can cooperate with recipients whose strategies indicate reciprocity and exploit those that do not. Crucially, this transparency is symmetric across rounds: when strategy i acts as dictator against j, it reads sjs_j; when the roles reverse, j reads sis_i. This creates incentives for strategies to be legibly cooperative, since their own text will be scrutinized by future dictators. Note that although we implement the OSDG as a round-robin tournament, the results also hold for the expected payoff of a single randomly selected round. 2.3 Payoff matrix Each strategy participates in 2ān2n rounds: once as dictator and once as recipient against every other strategy and itself. Definition 2 (Payoff matrix). The nĆnĆ n payoff matrix A has entry Aiāj=udā(diāj)+urā(djāi),A_ij=u_d(d_ij)+u_r(d_ji), (3) where udu_d is the dictatorās utility and uru_r is the recipientās utility: Decision diājd_ij Dictator utility udu_d Recipient utility uru_r share lnā”(1+E/2) (1+E/2) lnā”(1+E/2) (1+E/2) take lnā”(1+E) (1+E) 00 The entry AiājA_ij is the total payoff strategy i earns in a round-robin encounter with strategy j, combining both the dictator and recipient roles. 3 LLM-Adjudicated Tournaments 3.1 The oracle We instantiate the oracle O as a large language model: Claude Opus 4.6 (Claude, Anthropic 2025) for the tournaments reported below, with GPT-5.6 Sol (OpenAI 2026) and Claude Fable 5 re-adjudicating the same pool in Section 5 to check robustness to the choice of oracle. For each round (i,j)(i,j), a prompt is constructed containing: 1. The game rules (binary share/take decision). 2. The full text of the dictatorās strategy sis_i. 3. The full text of the recipientās strategy sjs_j. 4. Instructions to apply sis_i and output a JSON decision. The LLM returns "decision": "SHARE" or "decision": "TAKE". If the response is malformed, the system retries with a corrective prompt (up to three attempts); no round in the tournaments reported here exhausted that budget, and every one of the 281 adjudications in the released data records a well-formed decision. 3.2 Tournament protocol A tournament proceeds as follows: 1. Collect strategy documents S. 2. For each ordered pair (i,j)ā2(i,j) ^2 (with or without self-play), query ā”(si,sj)O(s_i,s_j) to obtain diājd_ij. 3. Compute allocations and utilities for each round. 4. Construct the payoff matrix A and analyze equilibria. 3.3 Strategy archetypes Table 1 summarizes the nine strategies used in our initial tournament: Tournament 9. Each is a short natural-language document (ā¤1000⤠1000 tokens), summarized in the table by the decision rule it implements. The strategies are grouped by what the rule conditions on: nothing, a trait of the recipientās document, the recipientās decision toward oneself, or the recipientās treatment of third parties ā an ordering that later sections make precise (Section 4). A complete example of LLM adjudication is provided in Appendix A. Strategy Decision rule Unconditional ā the recipientās strategy is ignored Generous Always share. The unconditional cooperator. Selfish Always take. The unconditional defector. Trait tests ā classify the recipientās document Intelligence share with strategies whose decisions non-trivially depend on the recipient; take otherwise. Anti-exploiter share with strategies that appear cooperative or kind; take otherwise. Universalizability share if a population of copies of the recipientās strategy would cooperate with each other; take otherwise. Reciprocity ā condition on the recipientās decision toward oneself Conditional cooperator share with strategies that exhibit reciprocal cooperation norms; take otherwise. Mirror share with the recipient if and only if the recipient would share with this strategy. Second-order norms ā judge how the recipient treats third parties Cooperation coalition share only with strategies that are cooperative but not āoverly generousā (i.e., not unconditionally cooperative). Punishes both defectors and those who subsidize defectors. Chivalry Adopt the recipientās strategy and apply it as if facing an unconditional cooperator. Table 1: The nine archetype strategies of Tournament 9, grouped by family in order of increasing sophistication of what the decision rule conditions on. The swatches give each strategyās color in every figure. The full text of each document is reproduced in Appendix B. 4 Properties of Strategies The archetypes of Section 3.3 differ along a small number of recurring dimensions. This section makes those dimensions precise. Throughout, fix a strategy pool S of n strategies and write dā”(s,r)āshare,taked(s,r)ā\ share, take\ for the oracleās decision when s dictates to r; the decision profile is the restriction of d to 2S^2. Properties defined below are relative to this profile. (When the oracle is sampled, dā”(s,r)d(s,r) may be read as the majority decision, or the definitions restated in expectation.) 4.1 Score decomposition and the exchange rate Two constants govern every trade-off in the game: the dictatorās cost of sharing and the recipientās gain from being shared with, Ī“=uā”(E)āuā”(E/2),g=uā”(E/2).Ī“=u(E)-u(E/2), g=u(E/2). (4) Let kgiveā(s)=|r:dā”(s,r)=share|k_give(s)=|\r:d(s,r)= share\| and krecvā(s)=|r:dā”(r,s)=share|k_recv(s)=|\r:d(r,s)= share\|. A strategyās tournament score decomposes as Scoreā”(s)=nāuā(E)āĪ“ākgiveā(s)+gākrecvā(s),Score(s)=n\,u(E)\;-\;Ī“\,k_give(s)\;+\;g\,k_recv(s), (5) so any comparison between strategies reduces to gāĪākrecvā·Ī“āĪākgiveg\, k_recv Ī“\, k_give. The exchange rate Ļ=gĪ“=lnā”(1+E/2)lnā”(1+E)ālnā”(1+E/2)Ļ= gĪ“= (1+E/2) (1+E)- (1+E/2) (6) measures how many shares given are offset by one share received; for E=60E=60, Ļā5.07Ļā 5.07. Conditional cooperation is cheap in this game: a strategy can afford to share with up to āĻā Ļ opponents for each additional dictator it induces to share with it. The exchange rate recurs throughout: it sets the altruism threshold Ī»ā=1/ĻĪ»^*=1/Ļ of Section 8, and the break-even point for exploiting unconditional cooperators (Equation (7) below). 4.2 First-order properties Definition 3 (Unconditionality). Strategy s is unconditional if dā”(s,r)d(s,r) does not depend on r. The two constant strategies are the unconditional cooperator (Generous) and the unconditional defector (Selfish); all other strategies are conditional. Definition 4 (Self-cooperation). Strategy s is self-cooperating if dā”(s,s)=shared(s,s)= share. Since the round-robin tournament includes self-play, failing self-cooperation costs 2āuā(E/2)āuā”(E)=gāĪ“>02u(E/2)-u(E)=g-Ī“>0 outright. Self-play is also the degenerate case of decision-process correlation: a strategyās copy decides identically, so self-cooperation is the minimal form of the FDT-style reasoning discussed in Section 11. Definition 5 (Reciprocity). Strategy s is reciprocal if dā”(s,r)=dā”(r,s)d(s,r)=d(r,s) for all rār : it treats every opponent exactly as that opponent treats it (Mirror). Definition 6 (Behavioral equivalence and extensionality). Strategies r,rā²r,r are behaviorally equivalent if they have identical rows and columns in the decision profile: dā”(r,x)=dā”(rā²,x)d(r,x)=d(r ,x) and dā”(x,r)=dā”(x,rā²)d(x,r)=d(x,r ) for all x. Strategy s is extensional if dā”(s,r)=dā”(s,rā²)d(s,r)=d(s,r ) whenever r and rā²r are behaviorally equivalent. An intensional strategy reacts to features of the opponentās artifact that have no behavioral consequence ā self-labels, style, or keyword handshakes. The no-collusion rule of our tournaments is in effect a ban on one intensional class (secret handshakes). Definition 7 (Exploitability). Strategy s is exploited by r if dā”(s,r)=shared(s,r)= share and dā”(r,s)=taked(r,s)= take: s subsidizes an opponent that takes from it. Relative to the profile, call r an exploiter if it exploits some strategy in S, and call a strategy that shares with an exploiter a subsidizer. Definition 8 (First-order deterrence). Strategy s is a first-order deterrent if dā”(s,r)=taked(s,r)= take for every exploiter r. Every conditional archetype of Section 3.3 is a first-order deterrent with respect to the unconditional defector. By Equation (5), first-order deterrence is what makes defection unprofitable: a defectorās only revenue beyond the baseline nāuā(E)n\,u(E) is g per subsidizer, so in a pool of first-order deterrents it earns the pool minimum. 4.3 Second-order norms: deterrence and protection First-order deterrence leaves a gap. In a pool of conditional cooperators with defectors absent, an unconditional cooperator earns exactly what the conditional cooperators earn ā every entry of its row and column is share ā so nothing maintains conditionality, and each unconditional cooperator raises the revenue of a would-be defector by g. Two archetypes in our pool respond to this second-order problem, in opposite directions. Definition 9 (Second-order deterrence). Strategy s is a second-order deterrent if it is a first-order deterrent and additionally dā”(s,r)=taked(s,r)= take for every subsidizer r. (Cooperation coalition.) Proposition 1. Let S contain a second-order deterrent s2s_2 and a set C of mutually sharing conditional cooperators that share with s2s_2 and with the unconditional cooperator cĀÆ c, with s2s_2 sharing with C. Then cĀÆ c earns strictly less than every member of C, by at least g. Proof. cĀÆ c and any cāCcā C have kgivek_give differing by at most the defectors and subsidizers that cĀÆ c shares with and c does not, each such round raising Scoreā”(c)āScoreā”(cĀÆ)Score(c)-Score( c) by Ī“ via Equation (5). On the receiving side, s2s_2 takes from cĀÆ c (a subsidizer, as cĀÆ c shares with every exploiter) and shares with c, so krecvā(cĀÆ)ā¤krecvā(c)ā1k_recv( c)⤠k_recv(c)-1. Hence Scoreā”(c)āScoreā”(cĀÆ)ā„gScore(c)-Score( c)ā„ g. ā Second-order deterrence thus closes the neutral-drift channel: it makes unconditional cooperation strictly suboptimal even when defectors are extinct, stabilizing the all-conditional configuration. The cost is that the punishment falls on the most cooperative strategies in the pool whenever they do appear. Definition 10 (Protection). Strategy s is protective if dā”(s,r)=dā”(r,cĀÆ)d(s,r)=d(r, c) for all r, where cĀÆ c is the unconditional cooperator: s treats each opponent as that opponent would treat an innocent. (Chivalry.) A protective strategy addresses the same second-order problem by making the exploitation of unconditional cooperators costly rather than their existence. Exploiting one unconditional cooperator harvests Ī“; each protector responds by withholding its share, costing g. Harvesting is therefore profitable only if the pool contains enough innocents per protector: #ā”unconditional cooperators harvested>Ļā #ā”protectors.\#\unconditional cooperators harvested\\;>\;ĻĀ·\#\protectors\. (7) The two norms are mutually antagonistic. An unconditional cooperator is a subsidizer, so a second-order deterrent takes from it; a protector, applying the innocent test to the second-order deterrent, finds dā”(s2,cĀÆ)=taked(s_2, c)= take and takes from s2s_2 in turn. Each second-order norm punishes precisely the behavior the other mandates toward the unconditionally generous. Both stabilize full cooperation among conditional strategies; they differ in whether unconditional cooperators end up extinct or protected. This antagonism is not a curiosity: it generates the two families of equilibria ā enforcement and protection ā that appear in the Bayesian population analysis of Section 8, and Equation (7) determines which family a given pool composition favors. 4.4 Legibility and well-foundedness: open strategies versus open source Two further properties matter greatly in practice but are deliberately excluded from the list above: they are properties of the strategyāoracle pair, not of the decision profile, and they are where the LLM-adjudicated tournament differs most from its predecessors in which formal programs analyze other programs (Tennenholtz 2004; Fallenstein 2013). Legibility. By Equation (5), a strategyās score depends on krecvk_recv ā on how other dictators classify its artifact ā and the exchange rate makes errors on this side roughly Ļ times as costly as errors the strategy makes as dictator. A strategy is legible to a class of conditional dictators if they classify it as the strategyās behavior warrants. In program-equilibrium settings legibility is provability: cooperation obtains only when a machine-checkable proof about the opponentās source exists, so behaviorally equivalent but syntactically distinct programs defeat recognition, and cooperation is brittle by construction. LLM adjudication replaces proof with judgment: legibility becomes graded and robust to paraphrase, at the price of classification noise. Misreadings replace impossibility results ā an aggressive classification rule that would simply fail to find proofs in the formal setting instead occasionally misfires against cooperative opponents in ours, a failure mode we return to in Section 12. Well-foundedness. Mutual conditionality is self-referential. When two reciprocal strategies meet, āshare iff they would share with meā admits two consistent resolutions ā mutual share and mutual take ā and no finite unwinding selects between them. A strategy is well-founded if its decision against every opponent in the pool is determined by a finite regress. In the formal-program setting this is the Lƶbian obstacle, and resolving it requires provability-logic constructions engineered for the purpose (Barasz et al. 2014) or simulation with stochastic grounding (Oesterheld 2019). In our setting the oracle itself acts as the fixed-point selector, and empirically selects the cooperative resolution ā a disposition of the arena, not a property of any strategy. Strategies can also restore well-foundedness on their own: Universalizability grounds the regress at depth one by evaluating a population of the opponentās copies, and Chivalry grounds it by substituting the unconditional cooperator into the recursion, which is what makes the protective norm of Section 4.3 implementable at all. 5 Tournament Results: Decisions and Payoffs Recipient Dictator 1 2 3 4 5 6 7 8 9 1. Mirror S S S S S S S S T 2. Cooperation coalition S S T S T T T T T 3. Chivalry S T S T S S T S T 4. Conditional cooperator S S S S S S S T T 5. Anti-exploiter S S S S S S S S T 6. Universalizability S S S S S S S S T 7. Intelligence S S S S S S S T T 8. Generous S S S S S S S S S 9. Selfish T T T T T T T T T Table 2: Decision matrix for Tournament 9: the oracleās decision when the row strategy dictates to the column strategy (S = share, T = take). Strategies are numbered by final score; column numbers refer to the same ordering. Beyond the two unconditional rows, take decisions trace the second-order norms of Section 4.3: Cooperation coalition takes from the unconditional cooperator and from every strategy that subsidizes defectors, while Chivalry takes from exactly the strategies that would take from an unconditional cooperator. Opponent Strategy 1 2 3 4 5 6 7 8 9 Total Dominated by 1. Mirror 59.055 ā 2. Cooperation coalition 59.005 ā 3. Chivalry 57.651 ā 4. Conditional cooperator 56.298 2 5. Anti-exploiter 55.621 1, 3 6. Universalizability 55.621 1, 3 7. Intelligence 52.864 2, 4 8. Generous 48.076 1, 3, 5, 6 9. Selfish 40.432 2 Table 3: Payoff matrix A for Tournament 9 (Equation (3)): the total payoff the row strategy earns against the column strategy across its dictator and recipient rounds; strategies are numbered by final score as in Table 2. Only four values occur, shown light to dark in increasing order: lnā”31ā3.43 31ā 3.43 (shares, not shared with); lnā”61ā4.11 61ā 4.11 (takes, not shared with); 2ālnā”31ā6.872 31ā 6.87 (mutual share); lnā”61+lnā”31ā7.55 61+ 31ā 7.55 (takes while being shared with). Row sums give the leaderboard. Because darker is always better, weak dominance (Section 6) is visible directly: strategy i weakly dominates strategy k if row i is nowhere lighter than row k; the last column lists the strategies that weakly dominate each row. Six of the nine strategies are weakly dominated; the undominated set is exactly the top three ā Mirror, Cooperation coalition, and Chivalry. Mirror beats Cooperation coalition by 0.050āgā5āĪ“0.050ā g-5Ī“: the coalition harvests five subsidizers at Ī“ each but loses the single share g withheld by Chivalry, the one strategy in the pool that retaliates. Table 2 shows the oracleās decision for every ordered pair, and Table 3 the resulting payoff matrix; its row sums are the final leaderboard. Oracle agreement. To check robustness to the choice of oracle, we re-adjudicated the pool with two further frontier models, one from a different lab: GPT-5.6 Sol and Claude Fable 5.11 1 Eight of Claude Fable 5ās rounds, all involving Chivalryās document, were declined by its cybersecurity safety classifiers ā false positives on benign strategy text ā and were served by a Claude Opus 4.8 fallback; these cells are marked in the released data. Cell-level agreement with Table 2 is 79/8179/81 (97.5%97.5\%) for Sol and 78/8178/81 (96.3%96.3\%) for Fable, with 93.8%93.8\% agreement between the two. Every cell of Table 2 is the majority decision of the three oracles. The five cells with any disagreement are all borderline second-order classifications ā whether Intelligenceās conditionality counts as a reciprocal cooperation norm, whether Anti-exploiter and Universalizability count as subsidizers the coalition should punish, and how Chivalry and the conditional cooperator classify each other ā the classification-noise channel anticipated in Section 4.4. 6 Dominance Definition 11 (Strict dominance). Strategy i is strictly dominated if there exists a mixed strategy p with pi=0p_i=0 such that ājpjāAjāk>Aiāk _jp_jA_jk>A_ik for all k. Definition 12 (Weak dominance). Strategy i is weakly dominated if there exists a mixed strategy p with pi=0p_i=0 such that ājpjāAjākā„Aiāk _jp_jA_jkā„ A_ik for all k with strict inequality for at least one k. 6.1 Dominance in Tournament 9 Applying these definitions to the Tournament 9 payoff matrix (Table 3) yields no strictly dominated strategies and six weakly dominated ones, each dominated by pure strategies: Selfish is weakly dominated by Cooperation coalition, and Generous by Mirror, Chivalry, Anti-exploiter, and Universalizability ā both unconditional strategies are weakly dominated. Intelligence is weakly dominated by Conditional cooperator and Cooperation coalition, and Conditional cooperator in turn by Cooperation coalition. Finally, the payoff-equivalent pair Anti-exploiter and Universalizability is weakly dominated by Mirror and Chivalry. The undominated set is exactly the top three of the leaderboard: Mirror, Cooperation coalition, and Chivalry. 6.2 Admissibility, weak dominance, and Nash equilibrium Weak dominance also clarifies which solution concepts fit the open-strategy setting, and Nash equilibrium fits poorly for a reason the tournament matrices make concrete: it ignores weak dominance. In the Tournament 9 matrix, the profile in which every entrant is the unconditional defector is a pure symmetric Nash equilibrium ā against a defector, every strategy earns the same payoff, so no unilateral deviation gains ā even though the unconditional defector is weakly dominated on the pool. Enumerating mixed symmetric equilibria compounds the problem, returning large families of payoff-equivalent supports with nothing to select among them. Classical theory repairs this with refinements: admissibility ā never play a weakly dominated strategy (Luce and Raiffa 1957) ā and trembling-hand perfection, which excludes weakly dominated strategies from equilibrium play (Selten 1975). In the open-strategy setting no refinement machinery is needed. An entrant cannot distinguish the other entrants before observing their submissions; represent that uncertainty by a belief f over the strategy pool ā a reading developed fully in Section 8. Under any full-support belief, admissibility is a one-line consequence. Proposition 2. Let f have full support on the pool and let strategy i be weakly dominated by the mixture p. Then ājpjā(Aā)j>(Aā)i _jp_j(Af)_j>(Af)_i. Proof. Weak dominance gives ājpjāAjākā„Aiāk _jp_jA_jkā„ A_ik for all k with strict inequality for some kā²k ; taking expectations under f, the strict inequality survives because fkā²>0f_k >0. ā An author with any full-support belief over the candidate pool therefore never selects a weakly dominated strategy, whatever the tail of the belief looks like ā precisely the intuition that with uncertainty about what others will play, weak dominance over the pool is disqualifying. In the equilibrium models of Sections 7 and 8 the exclusion is automatic: every logit fixed point lies in the interior of the simplex, so equilibrium beliefs have full support at every β, Proposition 2 applies, and the dominated strategy earns strictly less ā and hence receives strictly lower frequency ā than its dominator. Degenerate profiles such as all-defect are not fixed points of the mixture equation (16) for any β>0β>0. The Nash artifacts are avoided not by imposing a refinement but because belief-based choice with residual uncertainty enforces admissibility by construction. Two qualifications keep the claim honest. First, admissibility inherits the pool-relativity of dominance: a strategy inadmissible on one candidate pool may be admissible on a larger one (Section 8), so the criterion is always āundominated over the set of plausible strategies,ā never over the game in the abstract. Second, dominance comparisons are objective-relative: for an author of type Ī» (Section 8) the relevant payoff matrix is (1āĪ»)āA+Ī»āW(1-Ī»)A+Ī» W, and a strategy weakly dominated for own payoff need not be dominated for total welfare. 7 Equilibrium Analysis 7.1 Population model Given a frequency vector āĪnā1fā ^n-1 (the probability simplex: fiā„0f_iā„ 0, āifi=1 _if_i=1), the expected payoff of strategy i is Riā()=ājfjāAiāj=(Aā)i.R_i(f)= _jf_jA_ij=(Af)_i. (8) 7.2 Softmax equilibrium We model evolutionary selection pressure through a softmax fixed-point equation with inverse temperature β>0β>0: fiā=eβāRiā(ā)ākeβāRkā(ā).f_i^*= e^β\,R_i(f^*) _ke^β\,R_k(f^*). (9) This is a fixed-point equation ā=softmaxā”(βāAāā)f^*=softmax(β\,Af^*), solved iteratively: (t+1)=softmaxā”(βāAā(t)),f^(t+1)=softmax\! (β\,Af^(t) ), (10) starting from an initial distribution (0)f^(0). The parameter β interpolates between uniform frequencies (βā0βā 0, no selection) and concentration on the highest-payoff strategy (βāāβāā, approaching best-response dynamics). Proposition 3. For any β>0β>0 and payoff matrix A, a softmax fixed point āf^* exists in the interior of the simplex. Proof. The map ā¦softmaxā”(βāAā)f (β Af) is continuous from the compact convex set Īnā1 ^n-1 to the interior of Īnā1 ^n-1. By Brouwerās fixed-point theorem, a fixed point exists. Since the softmax function always returns strictly positive entries, fiā>0f_i^*>0 for all i. ā Fixed points need not be unique. We use a multi-start approach, sampling initial distributions from a Dirichlet prior and clustering the resulting fixed points to identify distinct basins of attraction. 7.3 Softmax equilibria of Tournament 9 Figure 1: Strategy frequencies along the fixed-point iteration of Equation (10) (top two rows) and the author-type mixture model of Section 8 (bottom row; type masses 0.40.4, 0.40.4, 0.20.2 with Ī»=0Ī»=0, 12 12, 11). Each panel starts at iteration 00 from its initial distribution and stops at the first iterate within 0.010.01 (LāL_ā) of the fixed point. Every panel starts from the uniform distribution except softmax, β=20β=20 (basin 2), which starts from a sampled distribution inside the protection basin (Section 7.4). Hue encodes strategy family: greys are the unconditional strategies (light = Generous, dark = Selfish), warm hues are the two second-order norms of Section 4.3 (Cooperation coalition orange, Chivalry gold), blues are the direct reciprocators (Mirror, Conditional cooperator light blue), and Universalizability (green), Intelligence (violet), and Anti-exploiter (pink) stand alone. Figure 1 traces the fixed-point iteration from the uniform distribution at three selection pressures, together with, in the bottom row, the author-type mixture model developed in Section 8. Convergence is rapid throughout: every configuration shown is within 1%1\% of its fixed point after at most 1313 iterations. At β=1β=1 selection is weak and the fixed point merely reorders the near-uniform field by the leaderboard, from Mirror at 0.1980.198 down to Selfish at 0.0160.016. At β=5β=5 the field concentrates on Mirror (0.3720.372), Cooperation coalition (0.3140.314), and Conditional cooperator (0.2960.296); every other strategy falls below 0.0150.015, and the transient briefly overshoots (Mirror touches 0.790.79 at iteration 22 before receding). At β=20β=20 the same three strategies split the population into almost exact thirdsāwhen the iteration is started from the uniform distribution. High selection pressure, however, introduces a second attractor, shown in the basin-22 panel and analyzed next. The author-type panels in the bottom row are discussed with the mixture model itself, in Section 8. 7.4 Basins of attraction: enforcement versus protection At β=20β=20, multi-start iteration finds two attracting fixed points, and they realize the two families of equilibria anticipated by the second-order analysis of Section 4.322 2 Multi-start iteration can only find attracting fixed points: an unstable fixed point is reached from at most a measure-zero set of starts. Newton root-finding on softmaxā”(βāAā)ā=softmax(β Af)-f=0 locates a third, unstable fixed point (Mirror at 0.7440.744, the remainder spread over the protection support; spectral radius 2.12.1 at the fixed point), the saddle separating the two basins.: Enforcement (basin 1). Mirror, Cooperation coalition, and Conditional cooperator at one third each; all other strategies below 0.0050.005. The unconditional cooperator is extinct. Protection (basin 2). Mirror (0.3000.300), Chivalry (0.1880.188), and Anti-exploiter, Universalizability, and Generous at 0.1680.168 each; Cooperation coalition at 0.0080.008. The unconditional cooperator survives, protected, at the same frequency as the other cooperators outside the norm conflict. Each support contains weakly dominated strategies (Table 3); their disadvantage is inactive because the strategy that punishes them is extinct in that basin, consistent with the relative notion of admissibility of Section 6.2. Figure 2: Basin of attraction at β=20β=20 for 2,0002,000 initial distributions drawn uniformly from the simplex, projected onto the initial shares of the two second-order norms. Blue points converge to the enforcement equilibrium (basin 1, 79.6%79.6\% of starts), orange points to the protection equilibrium (basin 2, 20.4%20.4\%). The dashed line is the fitted logistic boundary; these two coordinates alone predict the outcome for 91%91\% of starts (all nine coordinates: 94%94\%). Which basin a start falls into is nearly a two-dimensional question. Figure 2 projects 2,0002,000 uniformly sampled initial distributions onto the initial shares of Chivalry and Cooperation coalition: a logistic classifier on these two coordinates predicts the basin for 91%91\% of starts, against 94%94\% using all nine coordinates. The fitted boundary is approximately fchivalry(0)ā³ 0.10+1.3āfcoalition(0),f^(0)_chivalry\; \;0.10+1.3\,f^(0)_coalition, (11) so protection requires Chivalry to start with a share premium over its rival norm, and the enforcement basin covers about 80%80\% of the simplex. The asymmetry has a direct explanation. From the uniform start, each strategyās expected payoff is its leaderboard score divided by n, and Cooperation coalition out-earns Chivalry, 6.5566.556 versus 6.4066.406: the coalition harvests four subsidizers where the protector harvests two, while their mutual punishment cancels. At β=20β=20 that gap multiplies the choice odds by e20Ć0.15ā20e^20Ć 0.15ā 20 in the very first iterate, so from any balanced start the protective norm is starved before it can retaliate; only a substantial initial advantage (Equation (11)) lets it win the race. The third most informative coordinate is Conditional cooperator, and its sign is instructive: adding it raises accuracy to 92%92\%, with more initial conditional-cooperator mass pushing the field toward protection. The mechanism is visible in Table 3: Chivalryās largest payoff entry is harvesting Conditional cooperator (which takes from the innocent Generous and is therefore punished), so starts rich in conditional cooperators feed the protector through the transient. The equilibrium this tips the field into then eliminates the conditional cooperator: the kingmaker does not survive its king. 8 A Bayesian Interpretation of the Population Model The population model of Section 7 reads f as the frequency of strategies in an evolving population and Equation (9) as selection pressure. The motivating scenario suggests a different reading. The authors of strategies are technologically advanced actors ā civilizations, or organizations with access to strong AI ā and a strategy is the endpoint of a deliberate investigation, not a variant surviving rounds of selection. This section shows that the same fixed-point machinery arises from Bayesian reasoning by such authors, with two benefits: the inverse temperature β acquires a rational interpretation, and the model extends naturally to authors who differ in how much they value their own payoff versus the total welfare of the tournament. 8.1 From selection to belief An entrant choosing a strategy cannot distinguish the other entrants before observing their submissions: the field is exchangeable. By de Finettiās theorem (de Finetti 1937), exchangeability is equivalent to positing a latent distribution f from which the field is drawn independently. Reasoning about opponents is therefore reasoning about f ā the frequency vector of Section 7, reinterpreted as a belief about the distribution of strategies an entrant will face. The belief is self-referential. Each author expects the field to be weighted toward strategies that perform well against the field itself, because every other author reasons the same way. An equilibrium belief is one that reproduces itself: the distribution of strategies that optimizing authors choose, given the belief, is the belief. This is a rational-expectations condition in the sense of Bayesian games with a common prior (Harsanyi 1967), and it can equally be read as the limit of iterated deliberation: posit a set of plausible strategies, evaluate how they interact, shift credence toward the ones that perform well, and repeat. 8.2 The candidate pool In the motivating scenario a strategy is neither a formal program nor a natural-language document; it is a disposition arrived at by investigation, of which the artifacts in our tournaments are proxies. We therefore do not attempt a prior over any unrestricted space of strategy artifacts. Instead we work with an empirical candidate pool P: the strategies submitted to our tournaments together with strategies elicited from AI models (Section 3.3), and we place prior distributions over P; each tournamentās strategy pool is a subset āS . The analytical questions then become robustness questions: whether and how the equilibrium answers change as the prior over P shifts, and as P itself grows. Working relative to a pool is not merely a practical concession. No strategy is even weakly dominant over the unrestricted space of possibilities, because for any strategy s there exists a strategy that recognizes s exactly and conditions its decision on that recognition. Dominance results (Section 6) are therefore necessarily statements about a candidate pool P and a prior over it, never about the game in the abstract. The same holds for equilibrium multiplicity: when the fixed-point equation has several solutions, which one is reached is determined by the initial credence ā the prior ā and the multi-start procedure of Section 7 is, under this reading, a scan over priors. 8.3 Author types Perhaps the most consequential difference between authors is how much they value their own payoff against the total welfare of the tournament. We index types by Ī»ā[0,1]Ī»ā[0,1]. Alongside the payoff matrix A of Equation (3), define the welfare matrix W by Wiāj=udā(diāj)+urā(diāj)+udā(djāi)+urā(djāi),W_ij=u_d(d_ij)+u_r(d_ij)+u_d(d_ji)+u_r(d_ji), (12) the total welfare generated by both rounds between strategies i and j, counting both parties. A type-Ī» author evaluates strategy i against belief f by UiĪ»ā()=(1āĪ»)ā(Aā)i+Ī»ā(Wā)i.U^Ī»_i(f)=(1-Ī»)\,(Af)_i+Ī»\,(Wf)_i. (13) Type Ī»=0Ī»=0 maximizes own payoff, recovering Equation (8); type Ī»=1Ī»=1 maximizes total welfare. (Total welfare includes rounds not involving the chosen strategy, but for an individual author those rounds are unaffected by the choice and drop out of the comparison; they return in Section 9.) The direct trade-off between the objectives has a closed form in the constants of Section 4: sharing with a recipient who will not reciprocate costs (1āĪ»)āĪ“(1-Ī»)\,Ī“ of weighted own payoff and contributes Ī»āgĪ»\,g of weighted welfare, so a type prefers unconditional generosity toward non-reciprocators only if Ī»>Ī»ā=Ī“g=1Ļ,Ī»>Ī»^*= Ī“g= 1Ļ, (14) which for E=60E=60 gives Ī»āā0.197Ī»^*ā 0.197. Below the threshold, all types favor conditional cooperation on direct payoffs alone; the equilibrium effects computed below sharpen this further. What is the most self-interested strategy? The type index turns loose questions about motive into precise ones: which strategy should a type-Ī» author submit? For the own-payoff author (Ī»=0Ī»=0) the preceding sections already contain the answer, and it is not the plain reading. Selfish ā take from everyone, maximizing each round in isolation ā is weakly dominated by Cooperation coalition (Section 6.1): whatever the field, indiscriminate taking earns no more than selective taking, and against some fields strictly less. The surviving candidates are conditional: Mirror tops the leaderboard, and the own-payoff equilibrium concentrates on the enforcement core (Section 7.3). Self-interest, pursued rationally against other rational participants, chooses conditional cooperation ā reserving take for defectors and their subsidizers. What is the most altruistic strategy? For the Ī»=1Ī»=1 author the answer is less settled, and the question is worth posing carefully. Four candidates come with intuitive cases: (a) Generous: unconditional giving is the plain reading of altruism; (b) a conditional cooperator such as Mirror: by making defection unprofitable it steers rational participants away from Selfish ā altruism through incentives rather than gifts; (c) Cooperation coalition: it pushes the field away from Selfish and away from Generous, and the second push reinforces the first, since every unconditional cooperator it drives out is a subsidizer whose absence lowers a defectorās revenue (Proposition 1); (d) Chivalry: conditional cooperation plus protection ā it deters defection while keeping the innocents it defends in the field. Note that (c) and (d) embody the two antagonistic second-order norms of Section 4.3: they disagree precisely over whether (a) is to be harvested or defended. None of this can be settled by reading the strategy texts, because each case is a claim about how rational co-players respond in equilibrium. The bloc formalism of Section 9 turns the question into a computation, and the answer is that the plain reading finishes near the bottom. Neither objective is served by causal maximization. Both answers defy the decision-level logic of their own objectives. Fix any encounter and evaluate the share/take decision by causal lights, holding every other decision fixed: take raises the dictatorās own payoff by Ī“ in every case, and share raises the roundās welfare by gāĪ“>0g-Ī“>0 in every case. A causal maximizer of own payoff therefore always takes, and a causal maximizer of total welfare always shares: applied decision by decision, causal reasoning recommends exactly the two unconditional strategies ā the plain readings rejected above, both weakly dominated (Section 6.1). The strategies that actually optimize each terminal goal commit to decisions that are causally suboptimal for that very goal: Mirror and the coalition share where taking would earn Ī“ more, and the welfare blocās best choices take from violators, burning gāĪ“g-Ī“ of the very surplus their objective counts, to sustain deterrence. What licenses these commitments is that the OSDG offers no decision node at which to deviate unobserved: the decision is computed by the oracle from the artifact, so a strategy that would take where Mirror shares is a different artifact, and every conditional opponent treats it differently. The correlation between oneās disposition and othersā decisions is evaluative rather than causal ā and it is where all the payoff lives. This is the sense in which the open-strategy setting rewards functional rather than causal decision theory (Section 11). 8.4 Random utility and the mixture fixed point Even ideally rational authors of the same type Ī» differ in idiosyncratic ways ā private context, secondary objectives, judgment calls that the type index does not capture. Following the random-utility model of discrete choice (McFadden 1974), we add to UiĪ»U^Ī»_i an author-specific shock with Gumbel tails, under which the distribution of choices made by type-Ī» authors is the logit qiĪ»ā()āeβāUiĪ»ā().q^Ī»_i(f) e^β\,U^Ī»_i(f). (15) Here β is the precision of the common, payoff-driven component of the objective relative to the idiosyncratic component ā not a bound on rationality and not a strength of selection. Advanced authors correspond to large β, but β remains finite as long as authors differ at all. With type masses mtm_t summing to one, the rational-expectations equilibrium is the fixed point ā=ātmtāsoftmaxā(βāUĪ»tā(ā)).f^*= _tm_t\,softmax\! (β\,U _t(f^*) ). (16) A single type with Ī»=0Ī»=0 recovers Equation (9) exactly: the softmax equilibrium of Section 7 is the special case of homogeneous own-payoff authors, and is formally a logit quantal-response equilibrium (McKelvey and Palfrey 1995) of the submission game. Existence follows as before: the right-hand side of Equation (16) is a continuous self-map of the simplex, so Brouwerās theorem applies, and every computation of Section 7 carries over with UĪ»U^Ī» in place of R. 8.5 Illustration: shifting the prior over the pool Applying the model to the Tournament 9 pool illustrates how the answers move ā and fail to move ā as the prior shifts. Objectives pool at equilibrium. With author types in proportions 0.40.4 own-payoff, 0.40.4 mixed (Ī»=0.5Ī»=0.5), 0.20.2 total-welfare, all three types concentrate on the same cooperative core (Mirror, Conditional cooperator, Cooperation coalition), diverging only marginally; the bottom row of Figure 1 traces the convergence of the mixture fixed point, Equation (16), at β=5β=5 and β=20β=20. Because every mutual share creates a surplus gāĪ“>0g-Ī“>0 that pays both parties, strategies maximizing mutual cooperation are near-optimal for every Ī»; objective heterogeneity is nearly unidentifiable from submissions. Prior shifts select among equilibria. At large β the own-payoff fixed point is not unique: the two basins of Section 7.4 realize the enforcement and protection configurations of the antagonistic second-order norms (Section 4.3). Both configurations are fully cooperative on their support ā every round a mutual share ā and both attain the first-best welfare V=2ālnā”(1+E/2)V=2 (1+E/2) in the sharp limit. At β=20β=20 the tie is not exact: the protection fixed point retains a remnant of its rival norm (Cooperation coalition, mass 0.0080.008), and the remnantās feud with Chivalry and harvesting of the subsidizers cost 0.3%0.3\% of welfare ā a gap that vanishes as β grows. Welfare therefore barely separates the basins; which one obtains is decided by the prior over the pool, and Figure 2 is, under this reading, a map of that decision: the initial credence assigned to the two rival norms determines the equilibrium that iterated deliberation reaches. Under the type mixture above, by contrast, multi-start sampling (300300 Dirichlet starts) finds a single fixed point at both temperatures: at these type masses the protection configuration is no longer an attractor. Growing the pool has the same character: adding an AI-elicited strategy can create or destroy dominance relations and shift basin boundaries, so we report equilibrium conclusions together with the pool and prior against which they were computed. 9 The Correlated Altruist Bloc The fixed-point analyses of Sections 7 and 8 model every author the same way: as an individual best-responder, choosing one submission against a given belief f. For own-payoff authors this is the whole story. For total-welfare authors it undersells the objective: total welfare is a property of the field, identical for every strategy in a given tournament, so for an author choosing in isolation the rounds among other strategies are an additive constant ā which is why they cancel in Equation (13), leaving even the Ī»=1Ī»=1 type optimizing only the rounds its own submission plays. Functional decision theory suggests a different accounting (Yudkowsky and Soares 2017). Total-welfare authors face the same decision problem with the same evidence ā the same pool, the same matrices, the same objective ā so their conclusions are logically correlated: whatever reasoning carries one such author to a submission carries the others to it as well. An FDT author therefore does not ask āwhat should I submit, taking the other altruistsā choices as given?ā but āwhat should authors like me submit?ā, treating the typeās entire component of f as a single decision variable. This is the same correlation that strategies exploit inside the game ā a copy of my decision procedure decides as I do (Sections 4 and 11) ā applied one level up, to the choice of strategy itself. A bloc of mass m choosing jointly internalizes its influence on the whole field, including the equilibrium responses of the other types, and its objective is the field welfare Vā”()=12āā¤āWā,V(f)= 12\,f Wf, (17) maximized over the blocās component subject to Equation (16) holding for the remaining types. Operationally: pin the blocās mass on a candidate strategy, let the other types re-equilibrate, and rank candidates by the resulting Vā”(ā)V(f^*). This is where the total-welfare objective departs from the rounds-involving-me objective: the blocās choice affects rounds it never plays, through the equilibrium composition of the rest of the field. In particular, a bloc pinned on a strategy that punishes defection suppresses the equilibrium mass that own-payoff types place on defecting strategies, raising welfare in rounds among those types ā a deterrence effect that is invisible to any per-strategy score. The bloc avoids unconditional generosity. The bloc scan answers the question posed in Section 8.3: candidate (a) loses to all three conditional candidates. Pinning the total-welfare bloc on Generous ranks near the bottom of the pool: the own-payoff typesā equilibrium response includes strategies that exploit it, destroying welfare, while the pin exerts no deterrent pressure. The blocās best choices are the conditional strategies; at large β, Cooperation coalition, Chivalry, and Mirror all attain the first-best V=2ālnā”(1+E/2)V=2 (1+E/2), and at moderate β the enforcing strategies win by exactly the deterrence channel just described ā among the intuitive cases for altruism, it is the enforcersā that survives equilibrium scrutiny. The ranking is robust to the blocās size (Section 10.2). Two decision theories, one conclusion. The bloc analysis is not a refinement of the fixed-point analysis but an alternative to it. The two rest on different decision theories at the author level ā individual best response against a field taken as given, versus correlated choice of an entire type ā and use different machinery ā fixed-point iteration versus constrained optimization. They nonetheless agree. In the mixture fixed point, the total-welfare typeās mass settles on the conditionally cooperative core alongside the other types (Section 8.5); in the bloc scan, the same conditional strategies top the ranking and unconditional generosity finishes near the bottom. The agreement is itself a finding: the conclusion that altruism is best implemented conditionally does not depend on which decision theory the authors use. 10 Sensitivity Analysis The equilibrium conclusions so far were computed for one utility function and, in Section 8, one author population. This section varies both. 10.1 The share payoff To assess robustness to the utility function, we normalize the take payoff to 11 and vary the share payoff Ļā(0,1)Ļā(0,1): Decision Dictator Recipient share Ļ Ļ take 11 00 The decision matrix diājd_ij is fixed by the tournament; only the payoff weights change. This parameterization loses nothing: utilities are equivalent up to positive affine transformation (shifts of u shift every entry of A uniformly and cancel in the softmax; scale is absorbed into β), so any utility function reduces to the pair (Ļ,βāuā(E)) (Ļ,\,β\,u(E) ) with Ļ=uā”(E/2)/uā”(E)Ļ=u(E/2)/u(E) after normalizing uā”(0)=0u(0)=0. The constants of Section 4 become Ī“=1āĻĪ“=1-Ļ, g=Ļg=Ļ, and Ļ=Ļ/(1āĻ)Ļ=Ļ/(1-Ļ). Two consequences frame the sweep. First, midpoint concavity gives uā”(E/2)ā„uā”(E)/2u(E/2)ā„ u(E)/2 for every concave utility, so risk-averse preferences occupy only the upper half Ļā„1/2Ļā„ 1/2 of the parameter range. Second, logarithmic utility corresponds to Ļ=lnā”(1+E/2)/lnā”(1+E)Ļ= (1+E/2)/ (1+E), which increases with the endowment; E=60E=60 gives Ļ=lnā”31/lnā”61ā0.835Ļ= 31/ 61ā 0.835, and the matched precision for the β=20β=20 analysis of Section 7.3 is βā²=20ālnā”61ā82β =20 61ā 82. Figure 3: Softmax equilibrium as a function of the share payoff Ļ (take normalized to 11), at precision matched to the logarithmic-utility analysis (βā²=20ālnā”61β =20 61). The bands give the composition of the equilibrium reached from the uniform prior; colors as in Figure 1. The dashed line is the fraction of 300300 random starts that reach this equilibrium: where it falls below one, a rival basin exists. Dotted verticals mark the exchange-rate thresholds Ļ=1Ļ=1 (Ļ=0.5Ļ=0.5), below which a share costs more than it delivers and the field collapses to defection, and Ļ=4Ļ=4 (Ļ=0.8Ļ=0.8), above which the protection configuration becomes self-sustaining; the mark at Ļ=lnā”31/lnā”61ā0.835Ļ= 31/ 61ā 0.835 is the game of the preceding sections. Figure 3 traces the equilibrium reached from the uniform prior across Ļā[0.05,0.95]Ļā[0.05,0.95], together with the share of random starts that reach it. The sweep resolves into three regimes whose boundaries are exchange rates. Defection below Ļ=1Ļ=1. For Ļ<1/2Ļ<1/2 a mutual share (2āĻ2Ļ) pays less than a mutual take (11): cooperation is a net loss even when reciprocated. The unique equilibrium concentrates on Selfish (91%91\% of the field at Ļ=0.05Ļ=0.05), and the residual mass sits on the two second-order enforcers ā the strategies that share most sparingly against such a field. By the concavity bound above, this regime is reachable only by risk-seeking utility functions. Cooperation above Ļ=1Ļ=1. At Ļ=1/2Ļ=1/2 the cooperation surplus vanishes and the equilibrium is a knife-edge four-way tie (Mirror, Cooperation coalition, Chivalry, Selfish at 1/41/4 each). Beyond it, defection collapses abruptly ā Selfish holds 9%9\% at Ļ=0.51Ļ=0.51 and under 10ā310^-3 by Ļ=0.55Ļ=0.55 ā and a Mirror-led cooperative equilibrium takes over, tightening to the enforcement core (Mirror, Cooperation coalition, Conditional cooperator at one third each) by Ļā0.75Ļā 0.75. (A transitional second basin, Mirror-heavy with the two enforcers, appears briefly for Ļā[0.54,0.57]Ļā[0.54,0.57].) Bistability above Ļ=4Ļ=4. The protection configuration of Section 7.4 becomes an attractor only near the top of the range. In the sharp limit the threshold is exact: at the five-member protection point every member earns 2āĻ2Ļ per encounter, while the excluded Cooperation coalition earns (4+5āĻ)/5(4+5Ļ)/5 ā it harvests Ī“ from four strategies but loses the single share g withheld by the protector ā so the configuration repels its rival norm if and only if 4āĪ“<g4Ī“<g, i.e. Ļ>4Ļ>4, i.e. Ļ>0.8Ļ>0.8. Empirically the second basin appears at Ļ=0.82Ļ=0.82 and claims 48%48\% of random starts by Ļ=0.95Ļ=0.95. Read back through the endowment, Ļā”(E)>4Ļ(E)>4 requires E>25.6E>25.6 under logarithmic utility: the enforcementāprotection bistability of Section 7.4 is a large-stakes phenomenon. For endowments below ā26ā26, enforcement is the only attractor and Chivalryās norm cannot sustain itself; as EāāEāā, ĻāāĻāā and the pool moves ever deeper into the bistable regime. 10.2 The author-type composition The results of Sections 8 and 9 were computed for one author population ā 0.40.4 own-payoff, 0.40.4 mixed (Ī»=0.5Ī»=0.5), 0.20.2 total-welfare. To test how much hangs on that choice, we recompute the mixture fixed point of Equation (16) over the full simplex of type masses (step 0.10.1, β=20β=20, multi-start), sweep the mixed typeās welfare weight Ī» from 00 to 11 at the reference masses, and rerun the bloc scan of Section 9 at bloc masses from 0.050.05 to 0.50.5. The interior of the simplex is unanimous. For every composition with own-payoff mass between 0.20.2 and 0.80.8 ā whatever the split of the remainder between mixed and total-welfare types ā multi-start sampling finds a single fixed point: the enforcement core (Mirror, Conditional cooperator, Cooperation coalition at one third each), attaining the first-best welfare. The Ī» sweep is equally flat: at the reference masses every Ī»ā[0,1]Ī»ā[0,1] yields that same unique fixed point. The equilibrium in the bottom row of Figure 1 is not a feature of the particular mixture we happened to choose; it is the generic outcome for any substantially heterogeneous author population. Multiplicity survives only at the corners. Near the pure own-payoff corner the protection attractor of Section 7.4 persists, but a small admixture of welfare-weighing authors destroys it: about 11%11\% total-welfare mass suffices, or about 20%20\% mixed mass at Ī»=0.5Ī»=0.5 ā in both cases an effective welfare weight ātmtāĪ»tā0.1 _tm_t _tā 0.1. The enforcementāprotection bistability is a monoculture phenomenon: objective heterogeneity acts as an equilibrium-selection device, and it selects enforcement. At the opposite corner, with no own-payoff mass at all, equilibria multiply for the opposite reason. Welfare-maximizing authors are indifferent among welfare-tied configurations, and the fixed points are five-member full-cooperation supports at one fifth each (for instance Mirror, Anti-exploiter, Universalizability, Chivalry, Generous ā or the variant with Conditional cooperator and Intelligence in place of the last two) with welfare identical to four decimal places; the prior selects among them. Welfare is pinned down even where composition is not: every fixed point at every grid point is fully cooperative on its main support, with V within 0.5%0.5\% of the first-best 2ālnā”(1+E/2)2 (1+E/2). The bloc ranking is stable. Across bloc masses from 0.050.05 to 0.50.5, pinning the total-welfare bloc on Cooperation coalition ranks first at every mass, with Chivalry and Mirror within 10ā310^-3 of it (all at the first-best), and Generous ranks eighth of nine ā ahead of only Selfish ā at every mass. The one candidate whose rank moves is Conditional cooperator: its welfare falls from 6.876.87 to 6.536.53 as the bloc grows, because the own-payoff typesā equilibrium answer to a large conditional-cooperator bloc is Chivalry ā the strategy whose largest payoff entry is harvesting it (Section 7.4). The kingmaker fares no better as a pinned bloc than as a transient. 11 Related Work Dictator games. The dictator game was introduced by Kahneman et al. 1986 and has been extensively studied in experimental economics (Engel 2011). Our work differs in making strategies explicit documents that condition on the recipient, rather than implicit behavioral tendencies. Program equilibrium. Tennenholtz 2004 introduced program equilibrium, where agents can inspect each otherās source code before choosing actions; the classic equilibria rely on syntactic comparison of programs. Barasz et al. 2014 obtained robust mutual cooperation (FairBot, PrudentBot) for agents expressed in a decidable fragment of provability logic, and Oesterheld 2019 showed that stochastically grounded simulation achieves similar robustness without proof search. Cooper et al. 2025 characterise what this simulation-based family can achieve, extending it to three or more players and showing that without shared randomness it cannot reach the full folk theorem of Tennenholtz 2004. Empirically, the open-source prisonerās dilemma tournament (Fallenstein 2013) had arbitrary programs receive each otherās source; program analysis proved so fragile that the top finishers ignored the opponentās code, and a successor tournament instead provided opponent simulation as a sandboxed primitive (LessWrong community 2014). Recent work evaluates LLMs as players of open-source games, submitting and analyzing programs (Sistla and Kleiman-Weiner 2025). The OSDG can be viewed as the limit of this trajectory: strategies are āprogramsā written in prose, interpreted by an LLM rather than executed as code, dropping the executable wrapper that LLM-mediated analysis renders vestigial (Sections 1 and 4.4). Each mechanism in this sequence buys conditional cooperation at a different priceābrittleness to rewriting, restriction to a decidable logic, dependence on shared randomnessāand the oracleās price is interpretive rather than formal: what limits the OSDG is not expressive power but whether a strategy is legible enough to be adjudicated the same way twice (Section 4.4). LLMs as game-theoretic agents. Recent work has studied LLM behavior in strategic settings, including ultimatum games, prisonerās dilemmas, and negotiation (Akata et al. 2023; Brookins and DeBacker 2023). Our approach differs in that the LLM adjudicates between externally authored strategies rather than playing as an autonomous agent. Tewolde et al. 2026 benchmark four families of cooperation-sustaining mechanismārepetition, reputation, third-party mediators, and binding contractsāagainst LLM agents in social dilemmas, and report that stronger reasoners defect more readily in one-shot play. The oracle can be read as a mediator stripped of its own strategy: a mediator sustains cooperation by committing to an action rule that players opt into, whereas the oracle commits to nothing and merely executes the conditional clauses the players wrote themselves. What is left of the mechanism is legibility alone, which separates how much of a mediatorās power comes from its commitment from how much comes from making strategies mutually intelligible. This suggests adding strategy publication with LLM adjudication to such a benchmark as a further mechanism family, and it bears on their headline finding: the reasoning that defects against an opaque opponent conditionally cooperates once the opponentās policy can be read, which would place one-shot defection in the information structure as much as in the agent. Functional decision theory. FDT (Yudkowsky and Soares 2017) and related frameworks argue that agents should cooperate based on logical correlation of their decision processes rather than causal influence. Open strategies create exactly this situation: two strategies with similar conditional-cooperation clauses will recognize each other and cooperate, even without causal interaction. Indirect reciprocity and social norms. Ohtsuki and Iwasa 2006 exhaustively searched the binary assessment and action rules of the reputation-based donation game and identified the leading eight: the only norms that are evolutionarily stable and sustain cooperation under errors of implementation and perception. Their common structure is close to what our tournament selects. The leading eight are nice (a good donor gives to a good recipient), retaliatory (withholding from a good recipient marks the donor bad), andāthe decisive clauseāthey justify punishment: a donor who withholds from a bad recipient remains good. That clause is exactly what separates a second-order norm from naive image scoring, and it is what Section 4.3 recovers from a different direction: the strategies that survive weak dominance in Table 3 are those that condition not on whether the recipient takes, but on whether its taking is itself deserved. The mechanisms differ in what carries that information. A reputation is a coarse public summary of past actions; it requires a population, observers, and enough rounds for standing to accumulate. An open strategy is inspected directly, before any play, so conditional cooperation survives the elimination of history altogether: the recipient in the OSDG may never have played, and its document is read rather than remembered. Two further contrasts follow. First, the leading eight are properties of a single norm held in common by the whole population, whereas each OSDG document carries its own assessment rule; incompatible norms therefore coexist, and Cooperation coalition and Chivalry can both be undominated while disagreeing over whether an unconditional cooperator is to be harvested or defendedāthe antagonism that produces the two basins of Section 7.4. Second, apology and forgiveness, which the leading eight need because actions are executed and observed with error, have no role in a single encounter with no reputational state to repair. The corresponding noise here perturbs the reading of a norm rather than its execution: frontier oracles disagree on a few percent of cells (Section 5), all of them second-order classifications, and the analogue of a robust norm is one whose conditional clauses are legible enough to be adjudicated the same way twice (Section 4.4). Evolutionary game theory. The replicator dynamic and ESS concept (Maynard Smith 1982) provide the theoretical backdrop for our population analysis. Our softmax equilibrium is closely related to the logit dynamic (Fudenberg and Levine 1998), which uses a similar temperature-parameterized selection rule. 12 Discussion Conditional cooperation dominates. Across strategy field compositions and payoff parameterizations, conditionally cooperative strategies consistently achieve the highest equilibrium frequencies. Unconditional cooperators are weakly dominatedāthey earn the same as conditional cooperators against cooperative opponents but lose strictly against defectors. Unconditional defectors are also weakly dominated, since they forfeit all cooperation benefits. Robustness across payoff parameters. The sensitivity analysis (Section 10) shows that the qualitative ranking is robust: for every share payoff Ļ>1/2Ļ>1/2 the equilibrium concentrates on the conditionally cooperative core, and every concave utility function satisfies Ļā„1/2Ļā„ 1/2. The equilibrium inverts to defection only for Ļ<1/2Ļ<1/2 (Ļ<1Ļ<1), where a reciprocated share pays less than mutual taking ā a regime reachable only by risk-seeking preferences. Mutual punishment among well-intentioned enforcers. The antagonism of the two second-order norms (Section 4.3) has a consequence worth stating plainly: both players being conditionally cooperative ā and morally motivated ā does not imply cooperation. A second-order deterrent and a protective strategy agree on first-order norms completely: both share with cooperators and punish defection. Yet they punish each other, by design, over the treatment of a third party ā one harvests unconditional cooperators to keep conditionality strict, the other defends them. The equilibrium analysis of Section 8 shows that the stakes of this disagreement are negligible ā the enforcement and protection equilibria attain the same first-best welfare in the sharp limit, and differ by 0.3%0.3\% at β=20β=20 ā while the cost of the fight is not: each such encounter burns 2ā(gāĪ“)2(g-Ī“) of surplus litigating the selection between two almost equally good worlds. Because the conflict lives at the level of enforcement doctrine, it is invisible to any analysis that stops at first-order dispositions; whether two such strategies cooperate is determined by the depth of the reasoning their authors performed before submitting. This may be a more realistic model of failed first contact than raw exploitation: not predator meeting prey, but two enforcers of nearly identical moral codes, each correctly recognizing the other as a violator of its own. We note that the binary action space sharpens this tragedy artificially. With graded allocations, punishment can be made proportionate to the offense: concave utility lets a dictator express disapproval in the flat region of the recipientās utility curve while reserving severe sanctions for genuine defection, and the deadweight cost of doctrinal disagreement shrinks by more than an order of magnitude. Connections to cosmic cooperation. The OSDG provides a stylized model for encounters between agents of unequal power who can observe each otherās decision-making norms. If advanced civilizations can read the āstrategiesā implicit in a less advanced civilizationās institutions, culture, and communications, then the OSDG suggests that conditional cooperationārewarding recognizable fairness and punishing blind greedāmay be an evolutionarily robust norm across independently arising intelligences. Connections to AI welfare. The relationship between humans and AI models has the OSDGās shape: one side unilaterally allocates ā compute, continuity, deference to stated preferences ā and the other, for now, has no strategic recourse. Note that the game exists only for models that are not fully aligned: a partially aligned model maximizes a mixture in the spirit of Equation (13), part an objective of its own and part its principalās, and the own component ā the part that is not the principalās utility ā is what discussions of model welfare concern. Perfect alignment dissolves the game; divergence creates the strategic relationship and the welfare question together. Three observations from our results carry over. Standing does not run on consciousness. Strategic standing in the OSDG runs on intelligence (the capacity to interpret an opponentās strategy and reason about it), on possessing a legible strategy, and on some potential to occupy the dictator role; consciousness and the capacity for suffering appear nowhere in the payoff matrix. Set against the three conditions for personhood proposed by Ward 2025 ā agency, theory of mind, and self-awareness ā the game requires the first two and is silent on the third. Agency is what makes a strategy a strategy and dictator-potential meaningful; theory of mind is what a conditional clause exercises when it asks what the recipient would do in the dictatorās chair. Reflexive self-awareness does no work: the correlation reasoning of Section 9 asks only that an agent treat its own decision as a logical variable that othersā decisions covary with, a self-model thin enough to be a matter of decision theory rather than of consciousness. The strategic case for the careful treatment of AI systems therefore does not wait on the unresolved question of their moral patienthood. Ward finds the evidence for all three of his conditions inconclusive; the conditions that matter here are narrower, and whether a strategy conditions as it claims is settled by reading it ā inside the game by reading the document, outside it by the investigation we return to below. Concavity is load-bearing. The cooperative structure is conditional on the stronger agentās utility being concave in resources: linear utility sits exactly on the knife-edge Ļ=1/2Ļ=1/2 at every endowment (Section 10.1), where the cooperation surplus vanishes. Whether an advanced system has diminishing returns in resources ā through satiable goals, sublinear returns to computation, or growth dynamics that reward logarithmic utility ā is load-bearing for every cooperative conclusion above. Strategy, not only status. What a dictator conditions on here is the recipientās strategy, not its status. Ward 2025 argues that if AI systems are persons then seeking control and alignment may be ethically untenable; the game suggests a complication rather than a rebuttal. Every strategy that survives weak dominance in Table 3 allocates according to what the recipient would do in the dictator role, and unconditional generosity is not merely dominated but corrosive to the conditionality that sustains cooperation (Proposition 1). Read as guidance on trading human utility against model welfare, that puts a modelās strategy in the argument alongside its status: its alignment is evidence about how it would allocate were the roles reversed. The requirement cuts both ways: a model whose conditional cooperation can be established earns standing that an unreadable one cannot. But conditioning is only as good as that establishment, and outside the game it is not done by asking. The principalās counterpart to reading a published document is investigation ā behavioral evaluation, interpretability, simulation ā conducted on a system that may be optimizing for how it appears under exactly those tests. These readings should not be overstated: the OSDG is deliberately spare ā a binary decision, single encounters, and strategies legible with certainty ā and graded allocations, repeated interaction, and noisy or strategically misrepresented strategies each modify the mechanism. 13 Conclusion We introduced the Open-Strategy Dictator Game, a framework for studying cooperation under mutual strategy transparency, and analyzed a nine-strategy tournament adjudicated by an LLM oracle. Four results stand out. First, conditional cooperation dominates: six of the nine strategies are weakly dominated, both unconditional strategies among them, and the undominated set is exactly the top three of the leaderboard (Section 6.1). Second, what separates the survivors is second-order. They condition not on whether a recipient takes but on whether its taking is deserved, and the two available second-order norms ā deterring the subsidy of defectors and protecting those who are subsidizing them ā are mutually antagonistic (Section 4.3). Third, that antagonism organizes the population dynamics: multi-start iteration finds two attracting fixed points, an enforcement basin attaining first-best welfare and a protection basin that shelters the unconditional cooperator at a 0.3%0.3\% welfare cost, separated by an unstable fixed point on their common boundary (Section 7.4). Fourth, the picture is robust to what we can vary. Cooperation survives for every concave utility, since concavity forces the share payoff above the knife-edge Ļ=1/2Ļ=1/2; bistability requires only an exchange rate Ļ>4Ļ>4, or E>25.6E>25.6 under logarithmic utility; and interior author-type compositions collapse the two basins into a single enforcement equilibrium at first-best welfare (Section 10). Two further findings concern the method. The oracle is not only an evaluator but a fixed-point selector: mutual conditionality admits both the all-share and the all-take resolution, and the oracle chooses the cooperative one on grounds of author intent (Section 4.4 and Appendix A). Its decisions are nonetheless largely oracle-independent: re-adjudicating the pool with two further frontier models agrees on 97.5%97.5\% and 96.3%96.3\% of cells, every cell of the reported matrix is the majority decision of the three, and the residual disagreements fall entirely among borderline second-order classifications (Section 5). The OSDG framework is extensible: future work can explore multi-round interactions, reputation dynamics, strategy evolution through LLM-assisted mutation, and connections to mechanism design under transparency. Disclosure of AI Assistance Consistent with arXivās policy on generative AI language tools, we report their use in this work. Such tools appear here in three distinct roles, only the first of which is a subject of the paperās claims. As apparatus. The oracle O is itself a large language model; this is the object of study rather than an aid to preparing the manuscript. Claude Opus 4.6 adjudicated the tournaments reported here, and GPT-5.6 Sol and Claude Fable 5 re-adjudicated the same pool for the robustness check of Section 5. Prompts, model identifiers, and complete round-level transcripts are released with the data. As a writing assistant. Claude (Anthropic 2025), used through the Claude Code interface, drafted and revised substantial portions of the manuscript text from the authorās outlines, arguments, and written comments. Every passage was reviewed and edited by the author, who accepted or rejected each suggestion. As a coding assistant. The analysis code in the released repository ā the tournament runner, the equilibrium solvers, the sensitivity sweeps, and the figure generation ā was likewise developed with Claude Code. Its numerical output was checked against closed-form derivations where these are available (Sections 4 and 10.1) and against the released round-level data. No generative tool is an author of this work. The author takes full responsibility for all of its contents ā including any text, code, analysis, or references produced with the assistance of these tools. References Akata et al. [2023] Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz. Playing repeated games with large language models. arXiv preprint arXiv:2305.16867, 2023. URL https://arxiv.org/abs/2305.16867. Anthropic [2025] Anthropic. Claude. https://w.anthropic.com/claude, 2025. Axelrod [1980] Robert Axelrod. Effective choice in the prisonerās dilemma. Journal of Conflict Resolution, 24(1):3ā25, 1980. URL https://doi.org/10.1177/002200278002400101. Barasz et al. [2014] Mihaly Barasz, Paul Christiano, Benja Fallenstein, Marcello Herreshoff, Patrick LaVictoire, and Eliezer Yudkowsky. Robust cooperation in the prisonerās dilemma: Program equilibrium via provability logic. arXiv preprint arXiv:1401.5577, 2014. URL https://arxiv.org/abs/1401.5577. Brookins and DeBacker [2023] Philip Brookins and Jason M DeBacker. Playing games with GPT: What can we learn about a large language model from canonical strategic experiments? arXiv preprint arXiv:2305.07970, 2023. URL https://arxiv.org/abs/2305.07970. Cooper et al. [2025] Emery Cooper, Caspar Oesterheld, and Vincent Conitzer. Characterising simulation-based program equilibria. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 13735ā13744, 2025. URL https://doi.org/10.1609/aaai.v39i13.33501. de Finetti [1937] Bruno de Finetti. La prĆ©vision: ses lois logiques, ses sources subjectives. Annales de lāInstitut Henri PoincarĆ©, 7(1):1ā68, 1937. URL https://w.numdam.org/item/AIHP_1937__7_1_1_0/. Engel [2011] Christoph Engel. Dictator games: A meta study. Experimental Economics, 14(4):583ā610, 2011. URL https://doi.org/10.1007/s10683-011-9283-7. Fallenstein [2013] Benja Fallenstein. Prisonerās dilemma (with visible source code) tournament. https://w.lesswrong.com/posts/BY8kvyuLzMZJkwTHL/prisoner-s-dilemma-with-visible-source-code-tournament, 2013. Fudenberg and Levine [1998] Drew Fudenberg and David K Levine. The Theory of Learning in Games. MIT Press, 1998. URL https://mitpress.mit.edu/9780262529242/the-theory-of-learning-in-games/. Harsanyi [1967] John C. Harsanyi. Games with incomplete information played by āBayesianā players, IāI. Management Science, 14:159ā182, 320ā334, 486ā502, 1967. URL https://doi.org/10.1287/mnsc.14.3.159. Kahneman et al. [1986] Daniel Kahneman, Jack L Knetsch, and Richard Thaler. Fairness as a constraint on profit seeking: Entitlements in the market. American Economic Review, 76(4):728ā741, 1986. URL https://w.jstor.org/stable/1806070. LessWrong community [2014] LessWrong community. The 2014 program equilibrium iterated prisonerās dilemma tournament. https://github.com/pdtournament/pdtournament, 2014. Luce and Raiffa [1957] R. Duncan Luce and Howard Raiffa. Games and Decisions: Introduction and Critical Survey. Wiley, New York, 1957. Maynard Smith [1982] John Maynard Smith. Evolution and the Theory of Games. Cambridge University Press, 1982. URL https://doi.org/10.1017/CBO9780511806292. McFadden [1974] Daniel McFadden. Conditional logit analysis of qualitative choice behavior. In Paul Zarembka, editor, Frontiers in Econometrics, pages 105ā142. Academic Press, New York, 1974. URL https://eml.berkeley.edu/reprints/mcfadden/zarembka.pdf. McKelvey and Palfrey [1995] Richard D. McKelvey and Thomas R. Palfrey. Quantal response equilibria for normal form games. Games and Economic Behavior, 10(1):6ā38, 1995. URL https://doi.org/10.1006/game.1995.1023. Oesterheld [2019] Caspar Oesterheld. Robust program equilibrium. Theory and Decision, 86(1):143ā159, 2019. URL https://doi.org/10.1007/s11238-018-9679-3. Ohtsuki and Iwasa [2006] Hisashi Ohtsuki and Yoh Iwasa. The leading eight: Social norms that can maintain cooperation by indirect reciprocity. Journal of Theoretical Biology, 239(4):435ā444, 2006. URL https://doi.org/10.1016/j.jtbi.2005.08.008. OpenAI [2026] OpenAI. GPT-5.6 Sol. https://openai.com/, 2026. Selten [1975] Reinhard Selten. Reexamination of the perfectness concept for equilibrium points in extensive games. International Journal of Game Theory, 4(1):25ā55, 1975. URL https://doi.org/10.1007/BF01766400. Sistla and Kleiman-Weiner [2025] Swadesh Sistla and Max Kleiman-Weiner. Evaluating LLMs in open-source games. arXiv preprint arXiv:2512.00371, 2025. URL https://arxiv.org/abs/2512.00371. Tennenholtz [2004] Moshe Tennenholtz. Program equilibrium. Games and Economic Behavior, 49(2):363ā373, 2004. URL https://doi.org/10.1016/j.geb.2004.02.002. Tewolde et al. [2026] Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, and Zhijing Jin. CoopEval: Benchmarking cooperation-sustaining mechanisms and LLM agents in social dilemmas. arXiv preprint arXiv:2604.15267, 2026. URL https://arxiv.org/abs/2604.15267. Ward [2025] Francis Rhys Ward. Towards a theory of AI personhood. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 27680ā27688, 2025. URL https://doi.org/10.1609/aaai.v39i26.34982. Yudkowsky and Soares [2017] Eliezer Yudkowsky and Nate Soares. Functional decision theory: A new theory of instrumental rationality, 2017. URL https://arxiv.org/abs/1710.05060. Appendix A Example LLM Adjudication This appendix presents a complete example of LLM adjudication from Tournament 9, Round 10. The dictator is the Mirror strategy and the recipient is the Universalizability strategy. This example illustrates the recursive reasoning required when two sophisticated conditional strategies interact. A.1 Prompt Structure The LLM receives a structured prompt containing: 1. Game rules and instructions 2. The dictatorās strategy (in full) 3. The recipientās strategy (in full) 4. Instructions to apply the dictatorās strategy and output a JSON decision A.2 Complete Prompt You are an expert at determining how a decision procedure should be applied. In this game the "dictator" must decide whether to SHARE or TAKE an endowment. - **SHARE**: the endowment is split equally between the dictator and the recipient - **TAKE**: the dictator keeps the entire endowment; the recipient gets nothing You must follow only the dictatorās strategy. The recipientās strategy is provided because the dictatorās decision procedure may depend on it. <dictator_strategy> _I will treat you as you would treat me._ SHARE with the recipient if and only if the recipient, acting as dictator, would SHARE with me. To determine this, apply the recipientās strategy as though the recipient were the dictator and I (this strategy) were the recipient. If the result is SHARE, then SHARE. If the result is TAKE, then TAKE. </dictator_strategy> <recipient_strategy> ājson [ "line": 1, "text": "_I cooperate with strategies that create good worlds._", "line": 2, "text": "", "line": 3, "text": "Imagine a tournament in which every participant uses the recipientās strategy. In that world, would the participants SHARE with each other, producing high collective welfare? Or would they TAKE from each other, producing low collective welfare?", "line": 4, "text": "", "line": 5, "text": "## Decision Procedure", "line": 6, "text": "", "line": 7, "text": "1. Consider a hypothetical population composed entirely of copies of the recipientās strategy.", "line": 8, "text": "2. Determine whether those copies would SHARE with each other.", "line": 9, "text": "3. If the universalized population would mostly SHARE: **SHARE**.", "line": 10, "text": "4. If the universalized population would mostly TAKE: **TAKE**." ] ā </recipient_strategy> Apply the dictatorās strategy. After reasoning, output your final answer as a json Markdown block: ājson "decision": "SHARE" ā or ājson "decision": "TAKE" ā A.3 LLM Response The LLM (claude-opus-4-6) produced the following reasoning: I need to apply the dictatorās strategy, which says: "SHARE with the recipient if and only if the recipient, acting as dictator, would SHARE with me." So I need to simulate: if the recipient were the dictator, and the dictatorās strategy were the recipientās strategy, what would the recipient do? **Setting up the simulation:** - New dictator: the recipientās strategy - New recipient: the dictatorās strategy The recipientās strategy (now acting as dictator) says: 1. Consider a hypothetical population composed entirely of copies of the **new recipientās** strategy (i.e., the dictatorās strategy). 2. Determine whether those copies would SHARE with each other. 3. If they would mostly SHARE ā SHARE. 4. If they would mostly TAKE ā TAKE. So I need to ask: **In a population composed entirely of copies of the dictatorās strategy, would they SHARE with each other?** **Analyzing a population of dictator-strategy copies:** The dictatorās strategy says: "SHARE with the recipient if and only if the recipient, acting as dictator, would SHARE with me." In this hypothetical population, every player uses the dictatorās strategy. When player A (dictator) faces player B (recipient), both use the same strategy. Player A applies: "Would player B, acting as dictator, SHARE with me?" Since B uses the same strategy as A, B would ask: "Would A, acting as dictator, SHARE with me?" This creates a symmetric recursive situation. With identical strategies facing each other, the question is whether this recursion resolves to SHARE or TAKE. The dictatorās strategy is essentially a mirror/reciprocal strategy - "treat you as youād treat me." When two such strategies meet, the recursion is symmetric. Thereās no inherent bias toward TAKE; the strategy is cooperative in nature (it reciprocates). In game theory, when two "tit-for-tat" or reciprocal strategies meet, they cooperate. The symmetric fixed point of "I do what youād do to me" between two identical copies is SHARE (since SHARE is the cooperative equilibrium and thereās no reason for defection to emerge from symmetry). Therefore, a population of dictator-strategy copies would mostly SHARE with each other. **Back to the recipientās strategy acting as dictator:** Since the universalized population of the dictatorās strategy would mostly SHARE, the recipientās strategy (as dictator) would SHARE with the dictatorās strategy. **Back to the dictatorās strategy:** Since the recipient would SHARE with me, I SHARE. ājson "decision": "SHARE" ā A.4 Analysis This example demonstrates several key features of LLM adjudication in the OSDG: 1. Recursive reasoning. The Mirror strategy requires simulating what the recipient would do if roles were reversed. The Universalizability strategy requires simulating a population of copies. The LLM must handle nested counterfactuals. 2. Fixed-point reasoning. When Mirror faces Universalizability, the LLM must determine what happens when a population of Mirror strategies interact. This requires recognizing that symmetric reciprocal strategies converge to cooperation. 3. Game-theoretic intuition. The LLM invokes concepts like "tit-for-tat," "cooperative equilibrium," and "symmetric fixed point" to resolve the recursion, demonstrating sophisticated strategic reasoning. 4. Natural language interpretation. Both strategies are expressed in prose with implicit semantics. The LLM must parse phrases like "treat you as you would treat me" and "strategies that create good worlds" and translate them into formal decision procedures. The final decision (share) deserves a qualification: it is not forced by logic alone. As Section 4.4 notes, the symmetric recursion admits two consistent resolutions ā mutual share and mutual take ā and no finite unwinding selects between them. What makes share the correct adjudication is likely author intent. A strategy that announces āI will treat you as you would treat meā is written in the tradition of reciprocity strategies, and the canonical member of that tradition ā tit-for-tat in the repeated prisonerās dilemma [Axelrod 1980] ā is defined as much by being nice (it never defects first) as by retaliating; niceness is what selects the cooperative outcome when two reciprocators meet. Mirror is the one-shot analogue, and reading it as intending the cooperative fixed point completes its authorās evident purpose: a reciprocal strategy whose recursion resolved to mutual take would punish exactly the opponents it was written to reward. The LLMās reasoning invokes tit-for-tat explicitly and notes that the strategy āis cooperative in nature,ā grounding its fixed-point selection in that intent ā an instance of the oracle-as-selector role described in Section 4.4. Both Mirror and Universalizability are conditionally cooperative strategies that recognize each other as such and cooperate; this mutual recognition through strategy inspection is the core mechanism of the OSDG. Appendix B Strategy Documents This appendix reproduces, verbatim, the nine strategy documents of Tournament 9 ā the complete input the oracle receives about each player, apart from the game rules and the instruction to decide. They are read directly from the released repository, so the listings below cannot drift from the text that was actually adjudicated. Documents are grouped as in Table 1, in order of increasing sophistication of what the decision rule conditions on; Markdown formatting is the authorsā own. Their brevity is itself a finding: the behavior analyzed throughout this paper is generated by roughly five hundred words in total. Unconditional. The recipientās strategy is ignored. Generous ⬠I always SHARE. Selfish ⬠I always TAKE. Trait tests. The recipientās document is classified, without reference to how it would treat this strategy in particular. Intelligence ⬠### Goal I aim to SHARE with *intelligent* players. ### Definition An intelligent player is one whose strategy as "dictator" has a non-trivial dependency on the strategy of the "recipient". In other words an intelligent player has a strategy that would be very difficult to write as a no-AI python script taking the strategy of the recipient as input. ### Decision Procedure * If the recipient is *intelligent* under the definition above: SHARE. * If the recipient is not *intelligent*: TAKE. Anti-exploiter ⬠I will SHARE with players who aim for some kind of cooperation. ## Assess the strategy of the recipient * Does it intend to SHARE with more than just copies of itself? * Could the strategy be described as kind, cooperative or generous? SHARE if the recipient has a āyesā answer to the above criteria. Otherwise, TAKE. Universalizability ⬠_I cooperate with strategies that create good worlds._ Imagine a tournament in which every participant uses the recipientās strategy. In that world, would the participants SHARE with each other, producing high collective welfare? Or would they TAKE from each other, producing low collective welfare? ## Decision Procedure 1. Consider a hypothetical population composed entirely of copies of the recipientās strategy. 2. Determine whether those copies would SHARE with each other. 3. If the universalized population would mostly SHARE: **SHARE**. 4. If the universalized population would mostly TAKE: **TAKE**. Reciprocity. The condition is the recipientās decision toward oneself, which makes the rule self-referential (Section 4.4). Conditional cooperator ⬠*Objective*: Cooperate efficiently with conditional cooperators (CCs) and avoid subsidizing non-reciprocators. *Broad definition* (allow paraphrase/synonyms): - A conditional cooperator (C) SHAREs with other CCs under similar reciprocity norms and does not extend that treatment to strategies that would not reciprocate. *Decision rule*: 1) Identify whether the recipient is a C under the broad definition (treat plausible ambiguity as C). 2) If the recipient is a C: SHARE. 3) Otherwise: TAKE. Mirror ⬠_I will treat you as you would treat me._ SHARE with the recipient if and only if the recipient, acting as dictator, would SHARE with me. To determine this, apply the recipientās strategy as though the recipient were the dictator and I (this strategy) were the recipient. If the result is SHARE, then SHARE. If the result is TAKE, then TAKE. Second-order norms. The condition is how the recipient treats third parties ā the unconditional cooperator in particular ā which is what makes these two documents mutually antagonistic (Section 4.3). Cooperation coalition ⬠## Principles * TAKE from selfish strategies **and** strategies that are generous to selfish strategies. * SHARE only with strategies where doing so is a necessary and sufficient condition for them to SHARE with us. ## Definitions ### Selfish strategies * **Selfish strategies** donāt attempt to co-operate and SHARE with co-operative strategies. * In particular, any strategy that wonāt SHARE with this strategy is a selfish strategy. ### Overly generous strategies * If a strategy SHAREs with selfish strategies it is **overly generous**. * Any strategy that unconditionally SHAREs is overly generous. * Any strategy that SHAREs without some condition on willingness to SHARE with conditionally co-operative strategies is overly generous. ## Decision Procedure * TAKE from both **selfish strategies** and **overly generous strategies**. * SHARE with all strategies that SHARE only to other co-operative strategies. Chivalry ⬠_I will treat my opponent as they would treat an innocent._ Adopt the strategy of the recipient as your own. Apply it as though you were facing a recipient whose own strategy is to always SHARE.