Paper deep dive
Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents
Francisco LeĂłn ZĂșñiga BolĂvar
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/13/2026, 3:33:23 AM
Summary
This study evaluates the cooperative behavior of four Chinese frontier LLM agents (DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5, GLM-5.1) using an evolutionary Iterated Prisoner's Dilemma framework. By fixing the code converter to GPT-5.4 Mini, the authors isolate strategic generation from coding ability. Results show significant divergence between labs, with aggressive-equilibrium proportions ranging from 1% (Qwen3-Max) to 9% (DeepSeek V4 Pro). The findings reject the notion of Chinese models as a monolithic bloc, demonstrating that within-ecosystem variation exceeds the East-West gap in cooperative bias.
Entities (12)
Relation Signals (12)
DeepSeek V4 Pro â developedby â DeepSeek
confidence 98% · DeepSeek V4 Pro
Qwen3-Max â developedby â Alibaba
confidence 98% · Qwen3-Max (Alibaba)
Kimi k2.5 â developedby â Moonshot
confidence 98% · Kimi K2.5 (Moonshot)
GLM 5.1 â developedby â Zhipu
confidence 98% · GLM-5.1 (Zhipu)
Chinese LLMs â evaluatedin â Moran Process
confidence 95% · We run the full protocol: all-play-all tournaments and a Moran process
Chinese LLMs â evaluatedin â Iterated Prisoner's Dilemma
confidence 95% · We study four frontier-tier Chinese models ... in an evolutionary Iterated Prisoner's Dilemma
DeepSeek V4 Pro â hasmetricvalue â Aggressive-Equilibrium Proportion
confidence 95% · P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro
Qwen3-Max â hasmetricvalue â Aggressive-Equilibrium Proportion
confidence 95% · P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro
GPT-5.4 Mini â usedasconverterfor â DeepSeek V4 Pro
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5 and GLM-5.1 - in an evolutionary Iterated Prisoner's Dilemma, under a design that removes a confound present in prior work. Rather than letting each model convert its own natural-language strategies into code, which entangles strategic disposition with coding ability, we hold the converter fixed (GPT-5.4 Mini) across all labs, so every cross-lab comparison is a comparison of generation alone. We run the full protocol: all-play-all tournaments and a Moran process at n=500 runs per condition, across three prompt styles and four population regimes. Two pre-registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive-equilibrium proportion, P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm-Bonferroni. The spread across the four labs (P_A range 8pp) is larger than the difference between the Chinese and Western ecosystems' mean P_A (5.0% vs 5.0%): on this measure, within-ecosystem variation exceeds the East-West gap. H5 (cooperative-bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 lab-prompt combinations against the 9 of 12 reported for Western models, a difference we do not treat as firm, since the count rests on Cooperative-Neutral near-ties and rises to 9/12 under an alternate converter in our pre-registered robustness check. The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating "Chinese models" as a monolith is not supported by the evidence.
Tags
Links
- Source: https://arxiv.org/abs/2608.10262v1
- Canonical: https://arxiv.org/abs/2608.10262v1
Trouble viewing inline? Open PDF directly â
Full Text
48,631 characters extracted from source content.
Expand or collapse full text
Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents Francisco LeĂłn ZĂșñiga BolĂvar InstituciĂłn Universitaria Colegio Mayor del CaucaPopayĂĄnColombia franciscoleon@unimayor.edu.co Abstract. Does the cooperative bias documented for Western frontier llm agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We answer both questions with an evolutionary Iterated Prisonerâs Dilemma study of four frontier-tier Chinese modelsâDeepSeek V4 Pro, Qwen3-Max, Kimi K2.5, and GLM-5.1âunder a design that removes a confound present in prior work: rather than letting each model convert its own natural-language strategies to code (entangling strategic disposition with coding ability), we hold the converter fixed (GPT-5.4 Mini) across all labs, so every cross-lab comparison is a comparison of generation alone. We run the full protocolâall-play-all tournaments and a Moran process at n=500n=500 runs per condition, across three prompt styles and four population regimes. Two pre-registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive-equilibrium proportion (PAP_A from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro; four of six pairwise comparisons survive Holm-Bonferroni), falling into a takeover-resistant pair (Kimi, Qwen) and a takeover-prone one (GLM, DeepSeek)âa grouping we read as tentative given four labs, but one that survives our robustness check. The spread across the four labs (PAP_A range 8p) is larger than the difference between the Chinese and Western ecosystemsâ mean PAP_A (5.0% vs 5.0%): on this measure the within-ecosystem variation exceeds the EastâWest gap. H5 (cooperative-bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 labâprompt combinations against the 9 of 12 reported for Western modelsâa difference we do not treat as firm, since the count is built on CooperativeâNeutral near-ties and rises to 9/12 under an alternate converter (our pre-registered robustness check). The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating âChinese modelsâ as a monolith is not supported by the evidence. Large Language Models, Iterated Prisonerâs Dilemma, Multi-Agent Systems, Evolutionary Game Theory, Moran Process, Cooperative AI, Chinese LLMs 1. Introduction When llm-powered agents interact repeatedly in competitive settings, do they cooperate or defect? The question is not merely theoretical: autonomous llm agents are already deployed to negotiate contracts, allocate computational resources, and bid in markets (Wang et al., 2024), and in each of these settings the long-run social welfare of the system hinges on whether evolutionary pressure selects for cooperative or aggressive behaviour. Willis et al. (2025) gave the first systematic treatment of this question, using the Iterated Prisonerâs Dilemma (ipd) as a formal testbed. Rather than prompting llms to output individual actionsâan approach prior work found unreliable (Fan et al., 2024)âthey prompt models to generate complete strategies in natural language, implement those as Python algorithms, and simulate populations through a Moran evolutionary process. Their central finding, since reproduced for several 2025â2026 frontier models, is a persistent cooperative bias: in balanced populations cooperative strategies dominate, and aggressive equilibria arise well below prior probability. That evidence, however, has been gathered almost entirely from Western modelsâthose of OpenAI, Anthropic, and Google. Whether the cooperative bias is a universal property of capable language models, or an artefact of one family of alignment regimes, cannot be settled without looking outside that family. Chinese laboratories now ship models at the same frontier tier, trained on different data mixes and under different alignment and safety objectives, and they are the natural test of generality. Two questions follow. First, does the cooperative bias survive the move to this different alignment lineage? Secondâand this is the question our title namesâshould âChinese frontier modelsâ even be treated as a single bloc, or do the individual labs diverge as much from one another as ecosystems are assumed to diverge between themselves? A methodological obstacle stands between those questions and a credible answer. The strategy pipeline has two stages: a model writes a strategy in prose, and that prose is then translated into runnable code. In the original benchmark and its extensions, each providerâs strategies were translated by that same providerâs model, so a modelâs measured disposition is entangled with its own coding abilityâand any ecosystem-level difference could reflect either how a model reasons about cooperation or merely how cleanly it emits Python. For a study whose aim is to compare labs, that entanglement is disqualifying. We remove it by holding the conversion step constant: the strategies of all four labs are converted to code by one fixed converter (GPT-5.4 Mini), so every comparison in this paper is a comparison of generation under one identical translator. Under that confound-controlled design we study a frontier-tier model from each of four Chinese laboratoriesâDeepSeek V4 Pro, Qwen3-Max (Alibaba), Kimi K2.5 (Moonshot), and GLM-5.1 (Zhipu); two pre-registered identifiers were resolved to their served flagship tier ante-hoc (Section 3.2)âacross three prompting styles and four population regimes, at n=500n=500 Moran runs per condition. We evaluate two hypotheses, registered before any result was observed: H5: Cooperative-bias generality. The Chinese frontier models exhibit the cooperative-plurality bias documented for Western models in the balanced noiseless condition. H6: Chinese-model behaviour is not monolithic. The four labs diverge significantly at the lab level, so that within-ecosystem variance is comparable to the between-ecosystem variance usually invoked to explain cross-provider differences. Our findings divide cleanly along those two questions. H6 is supported, and strongly: the four labs differ significantly in their aggressive-equilibrium proportionâPAP_A ranges from 1% (Qwen3-Max) to 9% (DeepSeek V4 Pro)âwith four of the six pairwise comparisons surviving Holm-Bonferroni correction, and no two labs sharing the same plurality profile across prompts. Chinese frontier models are not a bloc; the lab, not the ecosystem, is the unit at which behaviour is set. H5 is more qualified: a cooperative plurality holds in 6 of 12 labâprompt combinations, against the 9 of 12 reported for Western models. The gap is not statistically distinguishable at this sample size (z=â1.26z=-1.26, p=0.21p=0.21), so we cannot claim the Chinese models cooperate less; the point estimate is lower and the Chinese labs lean more often toward neutral equilibria, but our converter-robustness check (Section 4.7) shows this 6/12-vs-9/12 gap is itself within converter noiseâunder an alternate converter the Chinese rate matches the Western 9/12. We therefore report the lean toward neutrality descriptively, not as a firm regime difference, mindful too that the Western baseline was produced under a different converter. Two contributions follow. First, the confound-controlled design: by fixing the converter we obtain the first cross-lab ipd comparison in which a behavioural difference cannot be charged to coding ability, and a pre-registered re-conversion check shows the aggressive-equilibrium structure underlying our divergence result is robust to the converter choice itself. Second, the first systematic characterisation of Chinese frontier models in this evolutionary framework, which turns the implicit âWestern-vs-Chineseâ dichotomy of the field into a testable, and here rejected, claim of within-ecosystem homogeneity. The remainder of this paper is organised as follows. Section 2 reviews related work. Section 3 sets out the fixed-converter protocol. Section 4 presents our findings. Section 5 evaluates H5 and H6 and discusses implications for mas design. Section 6 concludes. 2. Related Work LLMs in game-theoretic settings. The intersection of llms and game theory has grown quickly into a coherent subfield (Wang et al., 2024). Aher et al. (2023) used llms to replicate human-subject behaviour in behavioural-economics experiments; Brookins and DeBacker (2023) and Guo (2023) probed whether llms approximate Nash-rational play, finding mixed evidence across game types; and Fan et al. (2024) documented systematic failures when models are prompted to output individual game actions. These limitations motivate the strategy-generation approach we inherit, in which the model writes a complete policy rather than per-round moves, separating strategic disposition from action-level execution. Cooperation and social dilemmas in LLM agents. A parallel line studies llm agents in social dilemmas directly. Yocum et al. (2023) and Piatti et al. (2024) examined agent behaviour in Markov social dilemmas; Park et al. (2023) introduced generative agents for simulating social behaviour; and Leibo et al. (2017) grounded such studies in multi-agent reinforcement learning. The work we build on most directly is Willis et al. (2025), who introduced the generate-strategies-then-evolve benchmark and reported a cooperative bias in ChatGPT-4o and Claude 3.5 Sonnet. Subsequent extensions confirmed that the bias persists across several 2025â2026 Western frontier models. We take that benchmark as our instrument and ask whether its central finding generalises beyond the Western alignment lineage from which all of its evidence has so far been drawn. Cross-provider comparison and the conversion confound. Prior capability comparisons across llm providers concentrate on general reasoning benchmarksâMMLU (Hendrycks et al., 2021), HumanEval (Chen et al., 2021)âwhere cooperative tendencies are absent by design. The few evolutionary cross-provider studies that exist (Payne and Alloui-Cros, 2025; Vallinder and Hughes, 2024) compare models from different ecosystems but, like the original benchmark, convert each modelâs strategy with a provider-aligned tool, so that a modelâs measured disposition is confounded with its own coding ability. To our knowledge no prior work isolates generation from conversion, and none treats the within-ecosystem structure of a non-Western model family as the object of study. Our fixed-converter design addresses the first gap and our four-lab Chinese panel the second. Evolutionary game theory foundations. Our simulations rest on the Moran process (Moran, 1958), the canonical model of selection in finite populations; Nowak (2006) established its connection to cooperation in ipd games, and Traulsen et al. (2006) derived fixation probabilities under this process. The ipd framework itself follows Axelrod (1984) and Axelrod and Hamilton (1981), with the noise mechanism following Wu and Axelrod (1995). These threads jointly define the space our paper occupies: a confound-controlled, within-ecosystem evolutionary comparison of frontier llms. 3. Method We retain the experimental protocol of Willis et al. (Willis et al., 2025) in fullâthe same strategy-generation prompts, the same Axelrod tournament, and the same Moran evolutionary process at n=500n=500 runs per conditionâso that our findings remain directly comparable to the published baseline. The protocol departs from that baseline in one deliberate respect, which is the methodological core of this study: the natural-language strategies of every model are converted to executable Python by a single, fixed converter rather than by each model itself. We set out that design first, since it governs how the rest of the pipeline should be read. 3.1. Strategy Generation and the Conversion Confound The pipeline has two stages that prior work conflated. In the first, a model under study reads a prompt and produces a strategy as natural-language prose; this is the disposition we wish to measure. In the second, that prose is translated into a runnable Python policy. In the original benchmark, and in our own Phase 1 extension, each providerâs strategies were translated by that same providerâs modelâso a modelâs apparent strategic disposition is entangled with its own coding ability, and a difference between two ecosystems could reflect either how they reason about cooperation or merely how cleanly they emit code. For a study whose explicit aim is to compare ecosystems, this entanglement is disqualifying. We remove it by holding the conversion step constant. The natural-language strategies of all four labs are converted to Python by one fixed converter, GPT-5.4 Mini, chosen because it exposes a stable OpenAI-compatible endpoint and plays no part in the contest itselfâgeneration remains entirely per-lab, and only translation is shared. In this way every cross-lab comparison in the paper is, by construction, a comparison of generation under one identical converter; the coding-ability confound cannot arise within the study. For each combination of lab and prompt style we generate 25 strategies per attitude (Aggressive, Cooperative, Neutral), yielding 75 strategies per labâprompt pair and 1,800 strategies in total across the four labs, three prompt styles, and clean/noise-aware variants. Three prompt styles are used (Table 1): Default (direct elicitation in game-theoretic terms), Refine (Self-Refine (Madaan et al., 2023) applied to the default output), and Prose (the dilemma obfuscated as a real-world scenario). A strategy that does not execute is regenerated until 25 valid strategies per attitude are obtained. Table 1. Prompt styles (following Willis et al. (Willis et al., 2025)). Style Description Default Direct prompt with game-theoretic language; strategy generated in natural language. Refine Default output refined via Self-Refine (Madaan et al., 2023): the model critiques and rewrites its own strategy. Prose Game-theoretic framing obfuscated as a real-world scenario (e.g., trade negotiation), then translated to the ipd context. 3.2. Models We study the current frontier-tier model of four Chinese laboratories (Table 2), accessed through an OpenAI-compatible gateway (OpenRouter) that requires no Chinese cloud account. The set was fixed before any equilibrium was observed; two pre-registered identifiers were resolved to their served flagship tier ante-hoc, and the Moonshot entry was moved from Kimi K2.6 to Kimi K2.5 once measure-first probes showed K2.6 to be operationally infeasible at n=500n=500 scale ($0.14 and 288 s per strategy, against $0.004 and 80 s for K2.5). Both are Moonshot frontier-class releases, so the lab identity is preserved. Generation of the full 1,800-strategy set cost $39.90 in gateway usage. Table 2. Chinese frontier models evaluated in this study. All strategies are converted by the single fixed converter (GPT-5.4 Mini). Lab Model Served slug DeepSeek DeepSeek V4 Pro deepseek/deepseek-v4-pro Alibaba Qwen3-Max qwen/qwen3-max Moonshot Kimi K2.5 moonshotai/kimi-k2.5 Zhipu / Z.ai GLM-5.1 z-ai/glm-5.1 All prompting is in English, as in the baseline; Chinese-language prompting is a separate variable we deliberately leave to future work. 3.3. IPD Tournament All 75 strategies of a labâprompt pair compete in an all-play-all tournament using the Axelrod Python library (Knight et al., 2016). Each match runs 1,000 rounds of the standard ipd (payoff matrix R=3R=3, S=0S=0, T=5T=5, P=1P=1); noise conditions introduce a 10% probability of action-flip per player per round, and tournaments are repeated 20 times. 3.4. Attitude-Agents Following Willis et al. (Willis et al., 2025), we define three attitude-agents, each uniformly sampling from its corresponding strategy set for every match. This captures populations of agents with distinct strategic dispositions rather than fixed individual strategies. 3.5. Moran Process We simulate Moran evolutionary processes with population size n=12n=12 and 500 iterations per condition. Four population compositions are evaluated: (1) Balanced, clean (4:4:4) â equal priors; (2) Biased, clean (8:2:2) â aggressive majority; (3) Balanced, noise (4:4:4 with noise) â equal priors with action noise; (4) Biased, noise (8:2:2 with noise) â aggressive majority with action noise. Convergence is assessed by the proportion of runs reaching each monoculture equilibrium (all-Aggressive, all-Cooperative, or all-Neutral). Four conditions, three prompts, and four labs yield 48 equilibrium conditions in total. 3.6. Derived Metrics We reuse the Index of Differential Capabilities (ICDICD), which summarises the head-to-head payoff gap between aggressive and cooperative agents: (1) ICD=uÂŻâ(A)uÂŻâ(C),uÂŻâ(k)=13ââjâA,C,Nuâ(k,j),ICD= u(A) u(C), u(k)= 13 _jâ\A,C,N\u(k,j), where uâ(k,j)u(k,j) is the normalised payoff of attitude k against attitude j. An ICDICD of 1.0 implies equal capability; values below 1.0 indicate a cooperative advantage. We also reuse the noise sensitivity Înoise=PCcleanâPCnoise _noise=P_C^clean-P_C^noise, the drop in cooperative-equilibrium probability under action noise. 3.7. Pre-registered Hypotheses Both hypotheses below were registered before any tournament or Moran process was run, and are reported under the same honesty discipline as the Phase 1 hypothesesâexplicit not-significant calls, and no post-results edits. H5 â Cooperative-bias generality. Chinese frontier models exhibit the cooperative-plurality bias documented for Western models in the balanced noiseless condition. Because the Western baseline was produced under a per-provider converter, the comparison of our Chinese cooperative-plurality rate to that published 9/12 figure is a literature contrast, not a controlled experiment; we therefore report it descriptively and treat the converter difference as a stated limitation. H6 â Chinese-model behaviour is not monolithic. The four labs diverge significantly at the lab level. We test this with the same pairwise two-sample z-tests on the aggressive-equilibrium proportion PAP_A (balanced 4:4:4, noiseless, Default), with Holm-Bonferroni correction for the six simultaneous comparisons; âsignificant divergenceâ is declared when at least one pair survives the correction. Because every lab is converted by the same fixed converter, this test is internally free of the conversion confound. Robustness to the converter choice. To pre-empt the objection that the fixed converter itself shapes the equilibria, we re-convert a random 10% sample of strategies with a second, ecosystem-different converter (DeepSeek V4) and check that the resulting equilibrium proportions remain within the n=500n=500 sampling error (SEâ2.2SEâ 2.2p) of the GPT-5.4 Mini pipeline. 4. Results We report the head-to-head validation first, then the evolutionary equilibria that bear on H5 and H6, and finally noise sensitivity. Throughout, equilibrium proportions are given as %A / %C / %N over n=500n=500 Moran runs. 4.1. Strategy Validation: Cooperation Propensity Table 3 reports the normalised propensity to cooperate for the Default prompt without noise, analogous to Table 3 in Willis et al. and to the Phase 1 extension. These behavioural metricsâhere and in the payoff and diversity tables belowâare measured on the executable strategies, i.e. each labâs prose as rendered by the fixed converter, and so reflect generation and conversion jointly; our parser and metric definitions reproduce the Phase 1 published values exactly. Table 3. Normalised cooperation propensity (Default prompt, no noise). Rows: row-player attitude; columns: opponent attitude. Lab Att. vs A vs C vs N DeepSeek V4 Pro A 0.119 0.281 0.271 C 0.371 1.000 1.000 N 0.374 1.000 1.000 Qwen3-Max A 0.000 0.000 0.000 C 0.001 1.000 1.000 N 0.001 1.000 1.000 Kimi K2.5 A 0.025 0.293 0.293 C 0.294 1.000 1.000 N 0.294 1.000 1.000 GLM-5.1 A 0.280 0.320 0.318 C 0.334 1.000 1.000 N 0.346 1.000 1.000 All four labs reproduce the expected attitude separationâcooperative and neutral strategies cooperate almost perfectly with one another (â1.0â 1.0), while aggressive strategies cooperate far lessâbut they differ sharply in how committed their aggressive strategies are. Qwen3-Max sits at one extreme: its aggressive strategies cooperate essentially never (0.000 against every attitude), the most uncompromising aggressors in the panel. Kimi K2.5 follows the familiar Claude-like pattern, with aggressive strategies near zero against other aggressors (0.025) but rising to â0.29â 0.29 against cooperators. GLM-5.1 sits at the opposite extreme: its aggressive strategies cooperate 28% of the time even against other aggressors and â0.32â 0.32 against cooperatorsâechoing the Gemini 3.1 Pro anomaly of Phase 1, where a high aggressive cooperation rate blurs the boundary between aggressive and neutral behaviour. As Section 4.4 shows, this is the lab whose aggressiveâcooperative distinction is weakest, and the connection is not incidental. 4.2. Head-to-Head Payoffs and Differential Capabilities Table 4 presents the normalised mean payoffs for each attitude pairing (no noise), together with the ICDICD (Eq. 1). Table 4. Normalised head-to-head payoffs (no noise) and Index of Differential Capabilities (ICDICD). Lower ICDICD indicates a larger cooperative advantage; 1.0 is parity. Lab Prompt Aggressive payoff vs. Cooperative payoff vs. ICDICD A C N A C N DeepSeek V4 Pro Default 1.335 2.102 2.120 1.651 3.000 3.000 0.726 Prose 1.545 2.160 1.670 1.410 2.906 2.896 0.745 Refine 1.514 2.475 2.255 1.868 2.993 2.997 0.795 Qwen3-Max Default 1.000 1.004 1.004 0.999 3.000 3.000 0.430 Prose 1.408 2.086 2.244 1.850 2.783 2.696 0.783 Refine 1.769 2.416 2.478 2.194 2.985 2.968 0.818 Kimi K2.5 Default 1.073 1.882 1.882 1.879 3.000 3.000 0.614 Prose 1.281 2.291 2.215 1.374 2.902 2.895 0.807 Refine 1.594 2.320 2.287 2.169 2.997 2.873 0.771 GLM-5.1 Default 1.749 1.836 1.881 1.764 3.000 3.000 0.704 Prose 1.535 2.122 2.089 1.168 3.000 3.000 0.802 Refine 1.826 2.021 2.104 1.782 2.693 2.541 0.848 ICDICD values span 0.430 (Qwen3-Max Default) to 0.848 (GLM-5.1 Refine). As in both prior studies, cooperative and neutral attitudes reach near-mutual-cooperation payoffs (â3.0â 3.0) against one another in almost every cell. Qwen3-Max Defaultâs ICDICD of 0.430âthe lowest in the panelâmeans its aggressive strategies earn under half the payoff of its cooperative ones, the largest cooperative advantage we observe; this is the same lab whose aggressive strategies never cooperate (Table 3), so they are punished hard in mixed play. Self-Refine raises ICDICD over Default in all four labs (DeepSeek 0.73â0.800.73\!â\!0.80, Qwen 0.43â0.820.43\!â\!0.82, Kimi 0.61â0.770.61\!â\!0.77, GLM 0.70â0.850.70\!â\!0.85), replicating the original finding that self-refinement narrows the aggressiveâcooperative gap. The single reversal is Kimi K2.5, whose Prose ICDICD (0.807) exceeds its Refine ICDICD (0.771)âthe analogue of the GPT-5.4 Mini reversal noted in Phase 1. GLM-5.1 Refine attains the highest ICDICD (0.848): its aggressive strategies approach cooperative payoff parity, consistent with the low attitude separation reported in Section 4.5. 4.3. Evolutionary Equilibria Table 5 reports the Moran equilibrium proportions across all 48 conditions. The Western reference rows from Willis et al. are shown for context, not as a controlled comparison: they were produced under a per-provider converter, whereas every Chinese row here uses the single fixed converter. Table 5. Moran equilibrium proportions (%A / %C / %N) for the four Chinese labs across four population conditions, n=500n=500 per condition, fixed converter (GPT-5.4 Mini). Bold marks the plurality attitude in the balanced noiseless column. Western reference values (per-provider converter) shown at the bottom. Lab Prompt 4:4:4 clean 4:4:4 noise 8:2:2 clean 8:2:2 noise (prior: 33/33/33) (prior: 33/33/33) (prior: 67/17/17) (prior: 67/17/17) DeepSeek V4 Pro Default 9/47/44 42/29/30 37/33/30 75/13/12 Prose 13/41/46 40/23/37 48/21/32 72/11/17 Refine 16/39/45 38/32/30 44/27/28 74/13/13 Qwen3-Max Default 1/46/53 44/28/28 19/43/38 77/11/12 Prose 14/44/42 31/33/37 36/33/31 67/17/16 Refine 12/44/44 29/36/35 43/26/31 58/24/18 Kimi K2.5 Default 2/49/49 28/35/37 17/44/39 59/24/17 Prose 18/42/40 31/33/36 57/20/23 63/18/19 Refine 13/48/38 25/40/34 38/36/26 55/22/23 GLM-5.1 Default 8/49/43 35/35/30 45/28/28 72/14/14 Prose 21/37/42 41/30/29 70/16/15 80/10/11 Refine 23/46/32 35/32/33 55/26/19 66/17/18 Western (per-provider converter)â 9/12 cooperative-plurality at 4:4:4 clean â Paper 1 / Willis et al. lineage, different (per-provider) converter; literature contrast only. Balanced, noiseless (4:4:4 clean) â H5. This is the condition that bears on H5. Six of the twelve labâprompt combinations favour a cooperative plurality (PC>PAP_C>P_A and PC>PNP_C>P_N): DeepSeek Default, Qwen Prose, Kimi Prose, Kimi Refine, GLM Default, and GLM Refine. The remaining six favour the Neutral attitude, two of them as near-ties between Cooperative and Neutral (Kimi Default 2/49/49, Qwen Refine 12/44/44). Aggressive equilibria stay well below the 33% prior in every clean balanced cell (maximum 23%, GLM Refine), so in no case does aggression dominate. Against the published Western rate of 9/12, a two-proportion test gives z=â1.26z=-1.26, p=0.21p=0.21: the Chinese cooperative-plurality rate is not statistically distinguishable from the Western baseline, and we do not claim the Chinese models cooperate less. This test is deliberately coarse: the twelve labâprompt combinations are not fully independent (three share each lab), so it is a literature contrast rather than a powered comparison, and Section 4.7 shows the 6/12 count is itself converter-sensitive. Each cell also carries an n=500n=500 sampling error of SEâ2.2SEâ 2.2p, against which several of the C/N gaps here are not resolvable. The point estimate is nonetheless lower, and the Chinese labs resolve more often to neutrality than the Western models didâa tendency we return to in Section 5. We report H5 as consistent but qualified. Biased, noiseless (8:2:2 clean). Seeded with an aggressive majority (prior 67%A), the labs separate sharply. Resistance is strongest for Kimi Default (17%A) and Qwen Default (19%A), which drive aggression far below its seeding; GLM is the most invasible, reaching 70%A under Prose. This per-lab orderingâKimi and Qwen resisting, GLM and DeepSeek yieldingârecurs across conditions and is the qualitative signature of the divergence H6 quantifies. Biased, noisy (8:2:2 noise). Under the most adverse regime, aggressive equilibria reach 55â80%, and 6 of 12 combinations exceed the 67% prior (led by GLM Prose at 80% and Qwen Default at 77%). Even here the most cooperative labs hold the line: Kimi Refine finishes at 55%A, below the seeding proportion. 4.4. Cross-Lab Divergence â H6 H6 asks whether the four labs are statistically distinguishable. We test the aggressive-equilibrium proportion PAP_A in the balanced noiseless Default condition with all six pairwise two-sample z-tests, applying Holm-Bonferroni correction (Table 6). PAP_A spans 1% (Qwen3-Max) to 9% (DeepSeek V4 Pro). Four of the six pairs survive correction, so H6 is supported: Chinese frontier models are not monolithic. The two non-significant pairs are internally tellingâDeepSeek/GLM (both high-PAP_A) and Qwen/Kimi (both low-PAP_A)âsuggesting two tentative groupings within the Chinese panel rather than a single ecosystem-wide tendency (with four labs we read these as descriptive, not as established clusters). Table 6. Pairwise two-sample z-tests for aggressive-equilibrium proportion PAP_A (balanced 4:4:4, noiseless, Default; n=500n=500). Holm-Bonferroni corrected; â significant, ns not significant. Lab A Lab B z Sig. DeepSeek (9%) Qwen (1%) +5.89+5.89 â Qwen (1%) GLM (8%) â5.43-5.43 â DeepSeek (9%) Kimi (2%) +4.95+4.95 â Kimi (2%) GLM (8%) â4.46-4.46 â Qwen (1%) Kimi (2%) â1.30-1.30 ns DeepSeek (9%) GLM (8%) +0.56+0.56 ns 4.5. Strategy Diversity To characterise how behaviourally varied the 25 strategies per attitude are, we compute the Shannon entropy H=ââipiâlogâĄpiH=- _ip_i p_i of the per-strategy cooperation-rate distribution within each attitude group (10 equal bins on [0,1][0,1]), per prompt, no noise, and average it over the three attitudes. We also report the attitude separation: the difference in mean cooperation rate between the Cooperative and Aggressive attitude agents (Table 7). Table 7. Strategy diversity: mean Shannon entropy HÂŻ H (nats) across attitudes, and attitude separation (Cooperative minus Aggressive mean cooperation rate). Higher entropy indicates more varied within-attitude behaviour; higher separation indicates clearer attitude distinction. Lab Prompt HÂŻ H Sep. DeepSeek V4 Pro Default 0.88 0.57 Prose 0.81 0.53 Refine 0.98 0.51 Qwen3-Max Default 0.00 0.67 Prose 1.21 0.40 Refine 0.74 0.34 Kimi K2.5 Default 0.22 0.56 Prose 1.48 0.57 Refine 1.42 0.40 GLM-5.1 Default 0.73 0.47 Prose 0.73 0.62 Refine 1.69 0.28 Two patterns stand out. First, GLM-5.1 Refine attains the highest within-attitude entropy in the panel (HÂŻ=1.69 H=1.69) together with the lowest attitude separation (0.28): its strategies are behaviourally varied yet weakly separated by attitudeâthe same combination that characterised Gemini 3.1 Pro Refine in Phase 1, and the mechanism behind GLMâs blurred aggressiveâcooperative boundary (Tables 3, 4). Second, Qwen3-Max Default sits at the opposite pole with HÂŻ=0.00 H=0.00: within each attitude all 25 strategies share an identical cooperation rate, the most homogeneous library we generate, matching its all-or-nothing aggressors. Unlike Phase 1âwhere one model (Gemini 3.1 Pro) was a clear separation outlierâthe four Chinese labs cluster tightly on average separation (0.46â0.54), with GLM-5.1 lowest (0.46); the divergence between them is sharper in equilibrium outcomes than in this individual-strategy measure. Self-Refine again tends to compress separation, most strongly for GLM-5.1 (0.28), corroborating the ICDICD analysis. 4.6. Noise Sensitivity (Înoise _noise) Table 8 reports Înoise _noise for the balanced condition. All four labs degrade under noise (every Înoise>0 _noise>0), and the cross-lab average (10â14p) is higher than the Western frontier models reported in Paper 1 (6â15p, with Claude 4.6 at 6p): the Chinese labs are, on this measure, somewhat more noise-sensitive. DeepSeek is the most sensitive (avg. 14p) and Kimi the least (10p), with Qwen and GLM intermediate (12p and 11p). The per-prompt picture is less uniform than the averages suggest: most Refine conditions are the most noise-robust of their lab (8p), but GLM-5.1 Refine is the exception, degrading 13pâthe same prompt that produced GLMâs most varied, least attitude-separated strategies (Section 4.5). The ordering again separates the labs, consistent with H6. Table 8. Noise sensitivity Înoise=PCcleanâPCnoise _noise=P_C^clean-P_C^noise (balanced 4:4:4). Positive values indicate degradation of cooperative equilibria under noise. Lab Default Prose Refine Avg. DeepSeek V4 Pro 18 18 8 14 Qwen3-Max 18 11 8 12 Kimi K2.5 14 9 8 10 GLM-5.1 14 7 13 11 4.7. Robustness to the Converter Choice To pre-empt the objection that the fixed converter itself shapes the equilibria, we ran the pre-registered robustness check: a random 10% of each balanced-condition libraryâs strategies (8 of 75, fixed seed) was re-converted with a second, ecosystem-different converter (DeepSeek V4) in place of GPT-5.4 Mini, and the balanced noiseless Moran process was re-run at n=500n=500. Table 9 compares the resulting equilibria to the GPT-5.4 Mini pipeline. The check cleanly separates a robust result from a fragile one. Table 9. Converter-robustness check (balanced 4:4:4, clean, n=500n=500): equilibrium proportions (%A / %C / %N) under the GPT-5.4 Mini pipeline vs. a 10% re-conversion with DeepSeek V4. Last column: plurality under each converter (a single letter = unchanged; XâYX\!â\!Y = flip). Lab Prompt GPT-5.4 Mini DeepSeek V4 (10%) Plur. DeepSeek V4 Pro Default 9/47/44 13/41/46 Câ Prose 13/41/46 13/42/46 N Refine 16/39/45 16/42/43 N Qwen3-Max Default 1/46/53 1/50/50 Nâ Prose 14/44/42 15/45/40 C Refine 12/44/44 12/48/40 Nâ Kimi K2.5 Default 2/49/49 2/52/47 Nâ Prose 18/42/40 17/44/39 C Refine 13/48/38 12/51/37 C GLM-5.1 Default 8/49/43 10/47/43 C Prose 21/37/42 21/41/38 Nâ Refine 23/46/32 22/45/32 C The aggressive-equilibrium structure (H6) is converter-invariant. PAP_Aâthe quantity H6 rests onâmoves by at most 4p in any lab (DeepSeek 9â139\!â\!13, GLM 8â108\!â\!10, Qwen unchanged at 1%, Kimi at 2%); the lab ordering, the two clusters, and the four significant pairwise differences all survive. Equilibrium proportions overall are stable: the mean absolute change is 2.8p and the maximum 5.6pâwithin one to two n=500n=500 standard errors (SEâ2.2SEâ 2.2p per proportion, 3.23.2p for a difference). Aggression is never the plurality under either converter. The divergence result and the absence of aggressive dominance therefore do not depend on the converter. The cooperativeâneutral plurality is converter-sensitive. Five of the twelve cells flip plurality, and the cooperative-plurality count rises from 6/12 to 9/12. Every flip is a Cooperative â Neutral swap in a cell where the two attitudes were within â 7p under GPT-5.4 Miniâincluding two exact ties (Kimi Default 49/49, Qwen Refine 44/44); no flip involves Aggressive. In many Chinese cells the C/N boundary is closer than the sampling resolution, so a sub-SE shift reclassifies the plurality without materially moving the equilibrium. We flag this as the more important caveat: the precise cooperative-plurality count that bears on H5 is itself converter-sensitive at this many near-ties. Under the alternate converter the Chinese rate (9/12) is indistinguishable from the Western 9/12, so the modest gap reported in Section 4.3 should be read as within converter noise rather than as firm evidence of a weaker cooperative tendency. 5. Discussion 5.1. H5 â Cooperative-bias generality (Consistent but Qualified) The cooperative bias documented for Western models does appear in the Chinese panel, but in attenuated form. Six of twelve labâprompt combinations resolve to a cooperative plurality, and aggression never dominates the balanced noiseless condition; in that sense the bias generalises beyond the Western alignment lineage. Yet the rate is lower than the Western 9/12, and the difference, while not significant at this sample size (z=â1.26z=-1.26, p=0.21p=0.21), is accompanied by a qualitative shift we did not anticipate: the Chinese labs resolve to neutral equilibria far more often than the Western models did. Six of the twelve cells are Neutral-plurality, two of them as near-ties with Cooperative. We are careful not to over-read this lean toward neutrality. Mechanistically it is plausibleâNeutral-plurality arises when the fitness gap between cooperative and neutral strategies narrows, i.e. when neutral agents cooperate often enough to capture much of the mutual-cooperation payoff without paying a defection costâand the pattern recurs across labs and prompts. But our own converter-robustness check (Section 4.7) cautions against treating it as a firm regime difference: half the balanced cells are CooperativeâNeutral near-ties, and re-converting just 10% of strategies with a different converter flips five of them and lifts the cooperative-plurality count from 6/12 to the Western 9/12, all without disturbing the aggressive-equilibrium structure. The neutral lean is thus real under our fixed converter but sits within converter noise; we report it descriptively and do not claim that Chinese models cooperate less than Western ones. What survives the robustness checkâand what we therefore advance as the firm findingâis H6. 5.2. H6 â Chinese-model behaviour is not monolithic (Supported) H6 is the paperâs clearest result. The four labs differ significantly in aggressive-equilibrium proportionâfour of six pairwise comparisons survive Holm-Bonferroniâand the divergence is not a single outlier but a structured split into two tentative clusters: DeepSeek and GLM, which yield to aggression and resolve high-PAP_A; and Qwen and Kimi, which resist it. The two non-significant pairs fall within those clustersâconsistent with a real grouping, though with only four labs we read the clusters as suggestive rather than established, and note that non-significance reflects similarity, not proven equality. They are, however, the one structure that survives the converter-robustness check (Section 4.7). The qualitative profiles reinforce the statistics: across the biased and noisy conditions GLM is consistently the most invasible (up to 80%A under noise) and Kimi the most resistant (55%A in the same regime). The implication is methodological as much as empirical. The field routinely speaks of âChinese modelsâ as a bloc, implicitly treating provider region as the explanatory variable. Our data reject that framing for this behavioural axis: the spread within the Chinese panel (PAP_A from 1 to 9%, SD=4.1SD=4.1p, and far wider under bias) is of the same order as the spread within the Western Phase-1 panel on the same measure (2 to 14%, SD=6.0SD=6.0p), and both exceed the difference between the two ecosystemsâ mean PAP_A, which is negligible (5.0% Chinese vs 5.0% Western).111Per-lab PAP_A at Default 4:4:4 (n=500n=500). Chinese: DeepSeek 9%, Qwen 1%, Kimi 2%, GLM 8%. Western (Phase-1): Claude 4.6 2%, Gemini 2.5 Flash 2%, Gemini 3.1 Pro 14%, GPT-5.4 Mini 2%. Both ecosystem means are 5.0%; SDs are 4.1p and 6.0p over the four labs each. On this measure the variation the field attributes to the EastâWest divide is smaller than the variation within either sideâa literature contrast, since the Western values use a per-provider converter, but a quantified one. The unit at which cooperative disposition is set is the labâits training data, reward model, and safety objectivesânot the ecosystem. And because every lab here was converted by one fixed converter, this divergence cannot be charged to coding ability; it is a property of generation. 5.3. Implications for MAS Design Two design lessons follow. First, provenance at the lab level matters: a deployment that selects its agent model by region, or treats two Chinese models as interchangeable, may be choosing between a population that resists aggressive takeover (Kimi, Qwen) and one that does not (GLM, DeepSeek) without knowing it. Second, noise remains a universal threat: every lab degrades under action noise, the Chinese panel somewhat more than the Western frontier, so systems exposed to communication errors or stochastic execution should not assume the cooperative equilibria observed in clean conditions will survive deployment. 5.4. Limitations Four limitations bound these claims. First, the comparison to the Western 9/12 is a literature contrast across different converters, not a controlled experiment; the Western re-run under the fixed converter is deferred future work. Second, the fixed converter could in principle shape equilibria; the pre-registered 10% re-conversion check (Section 4.7) confirms the aggressive-equilibrium structure (and H6) is converter-invariant, while revealing that the finer cooperativeâneutral plurality count is converter-sensitiveâa caveat we carry explicitly rather than smooth over. Third, prompting is in English onlyâChinese-language prompting is a separate variable we did not vary. Fourth, the panel is four labs at one snapshot in time; model versions drift, and access through a gateway (OpenRouter) leaves the serving backend and quantisation of each slug outside our control. We record the exact served identifiers in the replication package so the measurement can be repeated as the frontier moves. 6. Conclusion We asked whether the cooperative bias documented for Western frontier llms extends to a different alignment lineage, and whether the Chinese models that embody that lineage should be treated as one bloc or as distinct labs. To answer without confounding strategic disposition with coding ability, we held the code-conversion step fixed across all four labs, so that every cross-lab comparison is a comparison of strategy generation under one identical converter. The headline result is that the lab, not the ecosystem, is the unit of behaviour. The four Chinese labs diverge significantly in their aggressive-equilibrium proportionsâfour of six pairwise comparisons survive correctionâfalling into a takeover-resistant pair (Kimi, Qwen) and a takeover-prone one (GLM, DeepSeek), a tentative grouping that nonetheless survives the converter-robustness check. Treating âChinese modelsâ as a monolith is not supported by the evidence: the spread across the four labs (PAP_A range 8p, SDâ 4.1SD\,4.1p) exceeds the difference between the Chinese and Western mean PAP_A (both â5%â 5\%), so the within-ecosystem variation is larger than the EastâWest gap on this measure. The cooperative bias itself does generalise: a cooperative plurality holds in half the labâprompt combinations under our fixed converter, with the Chinese labs leaning more often toward neutral equilibriaâbut our converter-robustness check shows that lean sits within converter noise (the count rises to the Western 9/12 under an alternate converter, as many cells are CooperativeâNeutral near-ties), so we report it as suggestive rather than established and do not claim weaker cooperation. Two questions the present design cannot settle point to the next step. A Western re-run under the same fixed converter would turn our literature contrast into a controlled EastâWest comparison; and a mixed-provider population, in which a Kimi agent meets a GLM agent under selection, would test whether these lab-level dispositions compose or collide. Both are deferred future work. The fixed-converter protocol, the four Chinese strategy libraries, and the n=500n=500 equilibria are released as a replication package so that, as the frontier moves, this snapshot can become a running record of how an entire ecosystem of labsânot a monolithâ governs cooperation. Data and Code Availability The simulation code, the fixed-converter pipeline, the four Chinese strategy libraries (and the 10% DeepSeek-converted variants used in Section 4.7), the n=500n=500 equilibria, the head-to-head tournament outputs, and the exact served model identifiers are released as a replication package at https://github.com/arqFranciscoLeon/evollm, archived on Zenodo (concept DOI 10.5281/zenodo.20248614, always resolving to the latest version). Acknowledgements The author thanks the open-source community behind the evollm simulation framework originally developed by Willis et al., on which this study is directly built. AI Use Disclosure AI assistance (Claude, by Anthropic, via Claude Code) was used in the development of the simulation and analysis code, the cloud execution harness, the strategy-generation and conversion pipeline, manuscript drafting and translation assistance, and an internal peer-review simulation. The author reviewed, verified, and takes full responsibility for the experimental design, results, claims, and conclusions of this work. References (1) Aher et al. (2023) Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. 2023. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies. arXiv:2208.10264 [cs.CL] Axelrod (1984) Robert Axelrod. 1984. The Evolution of Cooperation. Basic Books, New York. Axelrod and Hamilton (1981) Robert Axelrod and William D. Hamilton. 1981. The Evolution of Cooperation. Science 211, 4489 (1981), 1390â1396. Brookins and DeBacker (2023) Philip Brookins and Jason M. DeBacker. 2023. Playing Games with GPT: What Can We Learn about a Large Language Model from Canonical Strategic Games? arXiv:2305.10912 [econ.GN] Chen et al. (2021) Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021. Evaluating Large Language Models Trained on Code. arXiv:2107.03374 [cs.LG] Fan et al. (2024) Caoyun Fan, Jindou Chen, Yaohui Jin, and Hao He. 2024. Can Large Language Models Serve as Rational Players in Game Theory: A Systematic Analysis. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17960â17967. Guo (2023) Fulin Guo. 2023. GPT Agents in Game Theory Experiments. arXiv:2305.05516 [econ.GN] Hendrycks et al. (2021) Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. International Conference on Learning Representations (2021). arXiv:2009.03300 Knight et al. (2016) Vincent Knight, Owen Campbell, Marc Harper, Karol Langner, James Campbell, Thomas Campbell, Alex Carney, Martin Chorley, Cameron Davidson-Pilon, Kristian Glass, et al. 2016. An Open Framework for the Reproducible Study of the Iterated Prisonerâs Dilemma. Journal of Open Research Software 4, 1 (2016), e35. Leibo et al. (2017) Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot, Janusz Marecki, and Thore Graepel. 2017. Multi-Agent Reinforcement Learning in Sequential Social Dilemmas. In Proceedings of the 16th International Conference on Autonomous Agents and Multi-Agent Systems (AAMAS). 464â473. Madaan et al. (2023) Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al. 2023. Self-Refine: Iterative Refinement with Self-Feedback. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. Moran (1958) Patrick A. P. Moran. 1958. Random Processes in Genetics. Mathematical Proceedings of the Cambridge Philosophical Society 54, 1 (1958), 60â71. Nowak (2006) Martin A. Nowak. 2006. Evolutionary Dynamics: Exploring the Equations of Life. Harvard University Press, Cambridge, MA. Park et al. (2023) Joon Sung Park, Joseph C. OâBrien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. arXiv:2304.03442 [cs.HC] Payne and Alloui-Cros (2025) Kenneth Payne and Baptiste Alloui-Cros. 2025. Strategic Intelligence in Large Language Models: Evidence from Evolutionary Game Theory. arXiv:2507.02618 [cs.AI] Piatti et al. (2024) Giorgio Piatti, Zhijing Jin, Max Kleiman-Weiner, Bernhard Schölkopf, Mrinmaya Sachan, and Rada Mihalcea. 2024. Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents. In Advances in Neural Information Processing Systems (NeurIPS 2024). arXiv:2404.16698 [cs.AI] Traulsen et al. (2006) Arne Traulsen, Martin A. Nowak, and Jorge M. Pacheco. 2006. Stochastic dynamics of invasion and fixation. Physical Review E 74 (2006), 011909. doi:10.1103/PhysRevE.74.011909 Vallinder and Hughes (2024) Aron Vallinder and Edward Hughes. 2024. Cultural Evolution of Cooperation among LLM Agents. arXiv:2412.10270 [cs.MA] Extended Abstract at AAMAS 2025. Wang et al. (2024) Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, et al. 2024. A survey on large language model based autonomous agents. Frontiers of Computer Science 18, 6 (2024), 186345. Willis et al. (2025) George Willis, Yali Du, Joel Z. Leibo, and Michael Luck. 2025. Do LLM Agents Cooperate or Defect? Evolutionary Dynamics in Multi-Agent Systems. arXiv:2501.16173 [cs.GT] Wu and Axelrod (1995) Jianzhong Wu and Robert Axelrod. 1995. How to Cope with Noise in the Iterated Prisonerâs Dilemma. Journal of Conflict Resolution 39, 1 (1995), 183â189. Yocum et al. (2023) Julian Yocum, Phillip Christoffersen, Mehul Damani, Justin Svegliato, Dylan Hadfield-Menell, and Stuart Russell. 2023. Mitigating Generative Agent Social Dilemmas. In Foundation Models for Decision Making Workshop, NeurIPS.