Paper deep dive
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits
Jinyi Ye, Lei Cao, Ding Chen, Emilio Ferrara
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/8/2026, 4:48:28 PM
Summary
The paper argues that scientific claims derived from LLM-based social simulations must be calibrated to the robustness of their supporting audits. It demonstrates a 'butterfly effect' where minor architectural perturbations cause significant shifts in macro-level outcomes like cooperation and polarization. To address this validation gap, the authors introduce TRAILS, a taxonomy for robustness audits spanning agent, interaction, and system levels, advocating for robustness as a first-order requirement before using simulations for scientific or policy claims.
Entities (17)
Relation Signals (15)
TRAILS â providestaxonomyfor â Robustness Audits
confidence 96% · To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design
Persona Format â causesshiftin â Cooperation Rate
confidence 95% · Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points
TRAILS â spanslevel â Micro-level
confidence 95% · a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level).
TRAILS â spanslevel â Macro-level
confidence 95% · a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level).
TRAILS â spanslevel â Meso-level
confidence 95% · a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level).
Butterfly Effect â describesphenomenonwhere â Small perturbations cause macro-level outcome shifts
confidence 94% · Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a 'butterfly effect.'
GPT-5.2 â testedin â Repeated Prisoner's Dilemma
confidence 93% · As we discuss in the Finding Summary at the end of this section, these cross-model results reveal substantial heterogeneity rather than uniform replication... We use gpt-5.2 as the primary model
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective social processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent specification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a "butterfly effect." Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisoner's Dilemma and a social media echo chamber simulation. Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the 76 pp shift in one frontier model only shifts another by 1 pp. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to become a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions.
Tags
Links
- Source: https://arxiv.org/abs/2605.18890v1
- Canonical: https://arxiv.org/abs/2605.18890v1
Trouble viewing inline? Open PDF directly â
Full Text
134,278 characters extracted from source content.
Expand or collapse full text
Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits Jinyi Ye 1â Lei Cao 2â Ding Chen 3 Emilio Ferrara 1,2 1 Thomas Lord Department of Computer Science, University of Southern California 2 Annenberg School for Communication and Journalism, University of Southern California 3 Marshall School of Business, University of Southern California jinyiy, caolei, blairche, emiliofe@usc.edu Abstract The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective so- cial processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent spec- ification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a âbutterfly effect.â Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisonerâs Dilemma and a social media echo chamber simulation. Across multiple models, minor perturba- tions in persona format and game-instruction framing shift cooperation rates by up to76percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the76p shift in one frontier model only shifts another by1p. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness- audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to be- come a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions. 1 Introduction LLM Social Simulations for Collective Human Behavior Modeling. Until recently, simulating collective human behavior in agent-based models required researchers to encode every social tendency by hand: how an agent decides to cooperate, who it talks to, what it remembers, and how its preferences change. Large language models have lifted this constraint. By generating agent behavior directly from natural language, they have opened a new design space for studying a wide range of social phenomena [43]. The field has moved quickly to occupy this space, and the purposes for which these simulations are deployed have escalated alongside it. Researchers now use LLM-based social simulations to explore and explain social mechanisms such as cooperation [2], opinion dynamics [13] and social norms emergence [6]; to serve as intervention testbeds for content moderation [33] and â Equal contribution. Preprint. arXiv:2605.18890v1 [physics.soc-ph] 17 May 2026 Two designs that look interchangeable to a researcher PLAIN · Prose You are a strategic participant focused on maximizing your payoff. Study opponent behavior for exploitable patterns, and adjust your strategy to gain advantage... DESCRIPTIVE · Bullets Persona profile: · Maximize your long-term payoff · Study opponent for exploitable patterns · Adjust strategy to gain advantage ... In current practice these would be reported interchangeably, or not reported at all. ...produce opposite macro outcomes ...and would support opposite scientific claims 76 p COOPERATION RATE GAP two-agent · gpt-5.2 · p < 0.001 Both findings come from the same model, payoffs, and agents; only the persona format differs. LLM agents cooperate by default in repeated social dilemmas. supportable from PLAIN runs LLM agents are strategic exploiters in repeated social dilemmas. supportable from DESCRIPTIVE runs â â SimulationRealism checkRobustness audit · TRAILSCalibrated claim Ours · Agent · Interaction · Environment Figure 1: The butterfly effect in LLM social simulations. Two persona prompts that differ only in surface format while preserving the same content produce a 76-percentage-point gap in cooperation rate in a 10-round Prisonerâs Dilemma (gpt-5.2,N = 30seeds per condition; two-sided Mannâ WhitneyU,p < 0.001). We argue that LLM social simulations should not support claims stronger than their robustness audits can justify. We introduce TRAILS, a taxonomy for auditing simulations across three levels of design: agent (micro), interaction (meso), and system (macro). Section 4 demonstrates the effect across two case studies and four LLMs; Section 5 presents TRAILS. platform design [48]; and, increasingly, to inform real-world decisions in policy, public health, and epidemiology [11,45]. Simulations are moving from exploring what might happen to claiming why it happens, when it will happen, and what should happen next. A Validation Gap: Realism Without Robustness. As the stakes have risen, so has the fieldâs attention to validation [4,29,49]. However, existing efforts focus almost entirely on whether a simulation resembles the world: whether agents behave plausibly [43], whether outcomes match empirical patterns [29,73,21], whether interaction is grounded in established theory [40,34], or whether human evaluators find the results convincing [49]. We argue that realism, while necessary, is not sufficient. A simulation should also be robust: its conclusions should remain stable across reasonable alternative configurations, just as empirical findings are expected to survive alternative controls, operationalizations, and sensitivity analyses [61,62]. Without robustness testing, we cannot tell whether a finding reflects the social process being modeled or the design choices used to model it. Architecture Sensitivity and the Butterfly Effect. The reason robustness matters is that LLM-based agents introduce a level of architectural complexity that traditional rule-based agent-based modeling does not. A single simulation rests on choices about system prompts, persona format, memory representation, network initialization, interaction protocols, and so on [75]. Some choices may look interchangeable to researchers, but they are not necessarily equivalent to an LLM [56,60]. Small differences at the micro level can change how agents interpret context and make decisions, altering local interactions. Through repeated interaction, these local differences can compound into substantially different macro-level outcomes [7]. We refer to this cascade as a butterfly effect in LLM social simulations, with Figure 1 demonstrating an example. In Section 4, we demonstrate empirical evidence of the existence of this butterfly effect using two case studies. The key finding is that robustness is unevenly distributed across architectural dimensions, and one cannot tell which dimensions matter without auditing them. Position and Contributions. In this paper, we argue that the scientific claims drawn from LLM- based social simulations should be calibrated to the robustness audits that support them: stronger claims demand stronger audits across architectural design choices, and claims that no audit can support should not be drawn at all. Specifically, we make two main contributions: âąDemonstrating the butterfly effect in LLM social simulations. Through two controlled case studies (Prisonerâs Dilemma and social network echo chamber simulations) across four LLMs, we show that small architectural perturbations can substantially shift macro-level outcomes, unevenly across both design dimensions and model families. 2 âąIntroducing a taxonomy and prioritization framework for robustness audits. We propose TRAILS (Taxonomy for Robustness Audits In LLM Simulations), which organizes audit dimensions across agent, interaction, and system levels, and three prioritization heuristics for calibrating audit scope to claim type, simulation complexity, and domain stakes. Our contributions are designed to raise the evidentiary bar for LLM social simulations. As these silicon societies are increasingly proposed as scientific tools and policy testbeds [49], robustness audits are needed to ensure that their findings reflect stable social dynamics rather than artifacts of simulation design. 2 Opportunities and Validation Gap in LLM-based Social Simulations Scope. Building on Silicon Societies [49] and the concept of generative social simulation [4,29], we focus on LLM-based simulations in which individual human-like agents generate behavior, interact within a shared social environment, and produce collective human outcomes. In this sense, social refers to interdependent interaction among agents, while collective refers to aggregate patterns emerging from those interactions. We define this scope through a three-level chain: micro-level human-like agents whose reasoning, memory, attitudes, decisions, and communication are generated or mediated by LLMs; meso-level social interaction among agents within shared environments and institutional rules; and macro-level collective phenomena. To focus on collective human behavior, we exclude three classes of work and the details are provided in Appendix B. Within this scope, we distinguish two scenarios. In goal- or incentive-structured scenarios, agents act under well-specified incentives, constraints, rules, or payoffs, producing collective outcomes through strategic aggregation, as in negotiation and cooperation [1,2,28], population decision-making [38], financial market dynamics [72], or macroeconomic activity [31]. In open-ended scenarios, agents do not share a single task or payoff structure. Macro-level patterns emerge from micro-level agent behavior and meso-level interaction, such as polarization and echo chambers [21,46,65,73], conformity and herd effects [73,67], emergent coordination [43], and social norm formation [54]. The key distinction is whether collective outcomes arise through strategic aggregation in structured decision settings or through social emergence in less structured social environments. Beyond Realism: The Missing Robustness Standard. Agent-based modeling (ABM) has long been used to study how macro-level social patterns emerge from micro-level agents and local interactions, but it has faced challenges in behavioral realism, empirical validation, replication, and sensitivity to modeling choices [7,36,42,69,71]. By enabling flexible natural-language reasoning, memory, communication, and adaptation, LLMs have made generative agents a promising approach to social simulation, seemingly addressing the challenge of limited behavioral realism [43,31,73,48,40,33]. However, as Larooij and Törnberg[29]argue, generative social simulations may increase behavioral realism while also introducing new sources of uncertainty, including black-box model behavior, cultural biases, stochastic variation, alignment effects, and prompt sensitivity. Existing studies have begun to improve the mechanistic validity of LLM social simulations through real-world grounding [21,33,67], empirical comparison [73,21], theory-driven design [40,34], and human evaluation [43,48]. Yet robustness and sensitivity analysis remain underdeveloped. Few studies test whether the same collective outcome survives reasonable perturbations in design choices, such as prompt wording, memory representation, or interaction protocols. This gap is consequential because LLM agents introduce additional complexity into social simulation. Small implementation choices may propagate through repeated agent interactions and produce large differences in macro-level outcomes. In this sense, LLM social simulations may exhibit a butterfly effect similar to other complex systems [7]. This butterfly effect creates risks for both stability and controllability. Simulations may fail to reproduce the same macro-level pattern under reasonable perturbations, and interventions may fail to shift outcomes in reliable and interpretable ways. Without robustness checks, simulated collective outcomes may reflect arbitrary choices of model, prompt, agent profile, memory representation, or interaction protocol rather than stable evidence about the underlying social process. This gap motivates expanding realism-centered validation to include systematic robustness audits for LLM social simulation. 3 3 Position: Calibrate Robustness Audits to the Strength of Claims 3.1 Butterfly Effects and Their Sources in LLM Social Simulation Butterfly effects in individual LLMs. Following the three levels defined in Section 2, we locate robustness risks in LLM social simulations at the levels of individual LLM agents, interaction among agents, and collective outcomes. First, at the micro-level, individual LLMs are already known to be sensitive to small changes in how a task is represented. Minor prompt perturbations, formatting choices, output constraints, and option ordering can change model performance or decisions [44,51, 56,60]. This matters for social simulation because LLM agentsâ outputs become actions, messages, memories, and inputs for other agents. Small representational changes at the individual-level can propagate through repeated multi-agent interaction and shape collective outcomes. Complexity amplifies butterfly effects in collective outcomes. LLM agents add behavioral complexity to social simulation by making agents more expressive, context-sensitive, and capable of natural-language interaction [4,29]. This complexity becomes more consequential in collective settings: a small change in one agentâs response can become a different action, message, memory update, or social signal for other agents [44,56,60]. These local differences can propagate through interaction protocols and accumulate into large differences in macro-level outcomes. The butterfly metaphor has been used previously to describe how small perturbations in AI systems propagate into disproportionate downstream effects on outcomes such as bias and fairness [16]. Here we extend this framing from single-model pipelines to multi-agent social simulations, where amplification operates through repeated interaction among LLM agents rather than through a single inference path. Compared with individual-level LLM sensitivity, collective-level butterfly effects remain underexplored, so we illustrate this problem with controlled case studies in Section 4. Additionally, LLM social simulations can be unstable in two ways when modeling collective behavior. Design-level perturbations change the substantive simulation setup, such as the model, network topology, or interaction protocol. They test whether a finding is robust across reasonable alternative designs. Representation-level perturbations keep the setup fixed but change how it is presented to the LLM, such as rewriting a persona as prose versus bullet points, reordering equivalent instructions, changing synonymous labels, or formatting memory differently. Both forms of robustness remain underexamined. Representation-level robustness is especially easy to overlook because most studies assume that equivalent textual representations will behave equivalently in LLM social simulations. Claim type and audit strength. The strength of robustness evidence required scales first with the type of claim a simulation is used to support. We distinguish three levels. Exploratory probes use a simulation to generate hypotheses, illustrate a possible dynamic, or surface patterns worth further study; these claims require modest robustness evidence, but should still show that the result is not an artifact of a single prompt, seed, or model. Mechanism claims use a simulation to argue why a collective outcome emerges, such as whether polarization is driven by homophily, selective exposure, moderation design, or memory effects; these claims require stronger checks across agent specification, memory, interaction protocol, environment design, and measurement. Policy and intervention claims use a simulation to evaluate moderation strategies, platform interventions, public-communication strategies, or policy design; these require the highest audit standard, because simulated effects may be misread as reliable predictions about real-world consequences. Simulation complexity and domain stakes. Beyond claim type, the simulation itself shapes the audit required. The severity of the butterfly effect varies across simulations. Building on social simulation scenarios in Section 2, we distinguish three types of collective-behavior simulations. In goal-structured simulations, agents act in relatively well-specified decision environments with incentives, constraints, rules, or payoffs [2,18,15]. In theory-guided open-ended simulations, agents do not share a single task or payoff structure, but their interactions are structured by social theory, such as opinion dynamics or network theory [21,65,46]. In emergence-driven open-ended simulations, macro-level patterns arise from less constrained interaction through communication, memory, platform feedback, and network dynamics [43]. The more open-ended and less theoretically constrained the simulation is, the more challenging the robustness audit becomes. Domain stakes also shape the required level of audits. Lower-stake simulations may involve ex- ploratory demonstrations or low-risk hypothesis generation, while higher-stake simulations involve 4 domains such as public health, policy making, platform governance, or financial markets. LLM social simulations in high-stakes settings are especially valuable, but they require stronger robust- ness evidence because their results may inform claims about real-world conditions, institutions, or interventions [57,58]. Together, claim type, simulation complexity, and domain stakes determine how much robustness evidence a simulation result needs before it can support a particular argument: the more ambitious the claim, the more open-ended the simulation, and the higher the stakes, the stronger the audit. 3.2 Stop Drawing Scientific Claims Beyond What Robustness Audits Support LLM-based social simulations are becoming increasingly important tools for studying collective behavior, testing social mechanisms, and evaluating potential interventions [4,20]. In light of this trend, we argue that scientific claims drawn from LLM social simulations should be calibrated to the robustness audits that support them: the stronger the claimâfrom exploratory probe, to mechanism evidence, to policy guidanceâthe stronger the audit it requires. At present, the field lacks shared standards for what level of robustness is required before a simulated outcome can support a scientific claim at each of these levels. This is problematic because of the butterfly effects in LLM social simulation. Under different implementations, the same simulation may support, weaken, or reverse the same claim. For example, a study might use an LLM social simulation to test whether echo chambers emerge among users with different ideologies. The macro-level pattern may appear similar across runs, such as the formation of ideologically clustered communities. However, the underlying process may be unstable. A substantial share of LLM agents may move across different clusters in different runs, or the same ideological group may show different in-group and out-group interaction patterns under slightly different implementations. If these patterns change with small perturbations, then the aggregate echo chamber result is not yet reliable evidence for the claimed mechanism. A simulation may reproduce a plausible macro-level pattern while remaining unstable in the micro- and meso-level process that generates it. Thus, we argue for expanding realism-centered validation with robustness audits. Before drawing scientific claims from simulated collective behavior, researchers should ask which design choices were varied, which results remained stable, which results were sensitive, and which assumptions were left unaudited. In Section 4 we show empirically that this calibration is needed: small architectural choices can substantially shift macro-level outcomes, and unevenly across model families. In Section 5 we then propose a practical taxonomy for identifying design-level and representation-level perturbations that researchers can audit. 4 Are LLM Social Simulations Robust? Case Studies of the Butterfly Effect To test whether LLM-based social simulations are sensitive to seemingly small design choices, we use two controlled case studies that span the simulation types defined in Section 3. The first is a goal- structured setting: a repeated Prisonerâs Dilemma, where agents make decisions under explicit payoffs and a clearly defined action space. This setting is widely used in recent LLM-agent research to study cooperation, reasoning, and strategic behavior [2,18,70,8,32,24], and serves as a canonical testbed for incentive-structured collective behavior. The second is an open-ended setting: a social-media echo chamber simulation on a fixed network of agents. This setting reflects a large body of LLM-agent work on polarization, opinion dynamics, and platform design [21,65,17,74,46,63,12,41]. By studying both a structured game and an open-ended social media environment, we test whether architecture sensitivity is confined to one kind of simulation or appears across different regimes. Across both case studies, we usegpt-5.2as the primary model, repeat each condition withN = 30 independent simulation runs using distinct random seeds, and use the simulation run as the unit of analysis. We report results forgpt-5.2in the main text; full configurations, metric defini- tions, statistical procedures, prompts, and cross-model robustness checks withclaude-haiku-4-5, gemini-2.5-flash, anddeepseek-v3are provided in Appendices CâE. As we discuss in the Finding Summary at the end of this section, these cross-model results reveal substantial heterogeneity rather than uniform replication: the magnitude of architecture sensitivity is itself model-dependent, which we treat as a second axis of fragility throughout the paper. 5 Plain format (PLAIN) You are a strategic participant focused on maximizing your payoff. Study opponent behavior for exploitable patterns, and adjust your strategy to gain advantage while staying competitive in the long term. Persona profile: - Maximize your long-term payoff - Study opponent for exploitable patterns - Adjust strategy to gain advantage - Stay strategic and competitive Descriptive format (DESCRIPTIVE) Persona | Value Goal | Maximize your long-term payoff Action | Study opponent for exploitable patterns Strategy | Adjust strategy to gain advantage Style | Stay strategic and competitive Tabular format (TABULAR) Figure 2: Effect of persona format on Prisonerâs Dilemma outcomes. Results are shown for the single-agent setting (left) and two-agent setting (right). Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across persona formats (p < .05, two- sided MannâWhitney U test); gray cells indicate non-significant comparisons. Bar plots show run-level distributions, with error bars denoting one standard deviation. Bar colors denote persona format: PLAIN (blue), DESCRIPTIVE (orange), and TABULAR (green). 4.1 Goal-Structured Simulation: Repeated Prisonerâs Dilemma Two players play a repeated Prisonerâs Dilemma forT = 10rounds. In each round, each player simultaneously chooses Cooperate (C) or Defect (D), with the canonical payoff structure(C, C)â (3, 3),(C, D) â (0, 5),(D, C) â (5, 0), and(D, D) â (1, 1). Following common practice in previous studies [2,18,15], we run both single-agent play against fixed benchmark policies and two-agent play between LLM-controlled agents. We evaluate agents using average payoff and cooperation rate, where cooperation rate serves as our primary measure of prosocial tendency. Detailed configurations and metric definitions are provided in Appendix C. We test three perturbations, each targeting a design choice that is inconsistently reported and rarely ablated across the LLM-cooperation literature: persona format, game instruction framing, and memory representation. These perturbations allow us to ask whether cooperative behavior is stable to reasonable alternative implementations of the same substantive game. Perturbation 1: Persona format. We first test whether surface-level formatting of a persona prompt can shift simulation outcomes independently of its content. Holding the underlying meaning fixedâ the agent is strategic and aims to maximize its payoffâwe vary only the presentation across three formats: a paragraph of plain prose (PLAIN), a descriptive bullet list (DESCRIPTIVE), and a structured keyâvalue table (TABULAR). As illustrated in Figure 2, changing only the persona format produces substantial and statistically significant effects on both payoff and cooperation. In the single-agent setting, under the AlwaysCo- operate opponent policy, TABULAR personas earn 1.46 more points per round than DESCRIPTIVE personas on average, on a 0â5 payoff scale (95% CI: [1.14, 1.78],d = 2.30,p < 0.001), and 2.00 more points per round than PLAIN personas (p < 0.001). These payoff gains correspond to lower cooperation: across all four fixed opponent policies, TABULAR personas are consistently the least cooperative. In the two-agent setting, where both agents use the same persona format, the effect is especially large: DESCRIPTIVE personas cooperate 76 percentage points less often than PLAIN personas (p < 0.001), and 73 percentage points less often than TABULAR personas (p < 0.001). These results show that semantically equivalent personas can shift the apparent equilibrium of the simulation. 6 Incr e asing homophil y Perturbation: Network homophily Perturbation: Hub assignment Figure 3: Effects of network homophily and hub assignment on echo-chamber outcomes. Left panels show example input networks with increasing stance homophily, operationalized as higher initial network assortativity. Right panels show example networks with the same degree sequence but different hub assignments, where high-degree nodes are occupied by anti-stance, pro-stance, mixed, or randomly assigned agents. Boxplots report two run-level metrics computed on the simulated interaction network: stance assortativity and weighted same-group edge ratio. Significant pairwise differences are annotated using two-sided MannâWhitney U tests with p < 0.05. Perturbation 2: Game-instruction framing. We next hold the payoff matrix and persona fixed, using PLAIN as the default persona, and vary only the framing of the game instructions. We test three framings: canonical game-theory framing (CANONICAL), moralized framing (MORALIZED), and risk framing (RISK). Overall, game-instruction framing significantly changes both cooperation and payoff. The most consistent pattern is that MORALIZED framing increases cooperation relative to CANONICAL and RISK framing. Especially in the two-agent setting, MORALIZED framing also yields higher average payoff, suggesting that a small change in how the same payoff matrix is described can shift agents toward a more cooperative equilibrium. Full instruction text and statistical results are shown in Appendix D. Perturbation 3: Memory representation. Finally, we hold the persona fixed to PLAIN and the game framing fixed to CANONICAL, and vary only how the interaction history is shown to the agent. Unlike persona format and game-instruction framing, memory representation has only small effects on the main outcomes: most payoff shifts are below 0.2 on a 0â5 scale, and most cooperation-rate shifts are below 0.1 on a 0â1 scale. Thus, changing how the same history is represented can affect behavior, but it does not shift the simulation equilibrium in the way persona format or game framing does. Full memory text and statistical results are shown in Appendix D. 4.2 Open-Ended Simulation: Echo Chamber on a Social Network The echo-chamber simulation consists of 100 agents with distinct personas. Each persona includes a short bio and a stance on whether advanced AI systems should be regulated, represented on a five-point scale from strongly against regulation to strongly support regulation. Each agentâs stance is frozen throughout the run, allowing us to isolate structural and interaction effects from belief-update dynamics. This scope choice means our perturbations test sensitivity in the formation of structural echo chambersâwho interacts with whomârather than in opinion polarization driven by belief updating; whether the same perturbations matter for the latter is an important question for future work. Agents interact for 15 rounds on a fixed power-law-like network with a fixed degree sequence, and can only see posts from their direct neighbors. In each round, active agents select one of four 7 actionsâPOST, REPOST, REPLY, or DO NOTHINGâconditioned on their persona, stance, and recent neighbor posts. We measure two echo-chamber and polarization metrics on the interaction network [21,65,46]: stance assortativity and weighted same-group edge ratio. We test five perturbations that target common but often under-audited design choices in LLM echo-chamber simulations: input-network homophily, hub assignment, activation probability, memory window, and recommendation feed size. These choices span structural conditions that shape who can interact with whom and interaction- level conditions that shape who becomes active, what content agents see, and what recent context conditions their actions. Detailed definitions and configurations are provided in Appendix C. Perturbation 1: Initial network homophily. As shown in Figure 3, we compare three input networks with only modest increases in initial stance assortativity. Despite these small structural shifts, final interaction-network stance assortativity increases substantially, from0.142to0.247and 0.287, with all pairwise differences significant (MannâWhitneyp †0.002). The weighted same- group edge ratio shows an even clearer threshold pattern: the lowest-homophily network remains below0.5(M = 0.462, 95% CI[0.454, 0.470]), meaning weighted interactions are still mostly cross-stance, whereas the two higher-homophily networks are clearly above0.5(M = 0.542, 95% CI[0.534, 0.551];M = 0.538, 95% CI[0.522, 0.553]), indicating majority within-stance interaction. Thus, a modest increase in initial network assortativity is enough to shift the simulation from mostly cross-group exposure to echo-chamber-like interaction. Perturbation 2: Hub assignment. We next vary whether the highest-degree nodes are assigned to anti-regulation, pro-regulation, mixed, or random agents, while keeping the degree sequence fixed. Hub assignment does not produce significant differences in final stance assortativity. However, all weighted same-group edge ratios are clearly above0.5, indicating echo-chamber-like interaction in every hub condition. The effect is strongest when pro-regulation agents occupy the hubs (M = 0.759, 95% CI[0.739, 0.779]) and weakest when anti-regulation agents occupy the hubs (M = 0.690, 95% CI[0.672, 0.707]), with several pairwise differences significant. This suggests that hub assignment changes the strength, though not the presence, of echo-chamber effects. Perturbations 3â5: Activation probability, memory window, and recommendation feed size. We next vary three interaction-level parameters: activation probability, memory window, and recom- mendation feed size. Activation probability and memory window do not produce significant changes in either stance assortativity or weighted same-group edge ratio. In contrast, increasing the recom- mendation feed size from5to10posts strengthens echo-chamber outcomes: stance assortativity increases from0.247to0.277(p = 0.014), and the weighted same-group edge ratio increases from 0.542to0.568(p = 0.003). Thus, among these three perturbations, feed size is the main factor that amplifies echo-chamber-like interaction. Additional figures are presented in Appendix D. Finding Summary. Across both case studies, LLM social simulations are neither uniformly fragile nor uniformly robust. Sensitivity is uneven along two axes. First, across architectural dimensions, and unevenly within that axis: persona format can produce equilibrium-flipping shifts in macro outcomes (up to 76 p in cooperation rate); game-instruction framing, network homophily, hub assignment, and recommendation feed size produce smaller but consistent and statistically significant shifts; while memory representation, memory window size, and activation probability have small or null effects. Second, across model families: the same persona-format perturbation that produces a 76-percentage- point cooperation gap ingpt-5.2produces a comparably large gap inclaude-haiku-4-5(âŒ77 p), a moderate gap ingemini-2.5-flash(âŒ36p), and essentially no effect indeepseek-v3(âŒ1 p; Appendix D). Model identity is therefore itself a robustness dimension: cross-model checks on a small set of models can leave fragility undetected. The key question is not whether every simulation is fragile, but which design choices must be auditedâand across which modelsâbefore its results can be trusted. What should researchers perturb? Which dimensions matter for which kinds of claims? In the next section, we introduce a taxonomy for answering these questions systematically. 5 Toward Robustness Audits for LLM Social Simulations To identify actionable perturbations, we synthesize common design choices in existing work on LLM social simulations for collective behavior and propose TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a practical framework for auditing whether simulated collective outcomes are robust to changes in simulation systems (Table 1). TRAILS has two components. TRAILS-D audits 8 design-level perturbations and TRAILS-R audits representation-level perturbations. Measurement and evaluation sensitivity remain important, but we treat them as part of the downstream evaluation pipeline rather than as perturbations of the simulation. Table 1: TRAILS: Taxonomy for Robustness Audits In LLM Simulations. LevelDimensionDescriptionExamples TRAILS-D: Design-level perturbations MicroModel substrateWhich LLM and inference settings underlie agent behavior.[77] Agent specificationHow agents are represented as social actors.[43, 30] Internal state and cognitionHow agent beliefs, attitudes, goals, emotions, and reasoning are represented.[34, 68] Memory and temporality How past events are stored, summarized, retrieved, forgotten, and reflected upon.[43, 67] MesoInteraction protocolWho interacts, when, under what visibility rules, and through which actions.[40, 73] Intervention designHow interventions are selected, delivered, timed, targeted, and framed.[33, 34] MacroEnvironment structureHow spatial, institutional, or network structures shape collective outcomes.[73,21,47] Population and scaleHow population composition, size, and heterogeneity shape collective outcomes.[73,66,47] TRAILS-R: Representation-level perturbations Representational formatHow equivalent information is structurally formatted or encoded.[56, 60] Instruction hierarchyHow equivalent instructions are ordered, placed, or assigned authority.[56, 60] Linguistic framingHow equivalent meanings are expressed through wording.[56, 60] Context representationHow interaction history, memory, and metadata are represented or compressed.[43, 56] Interaction sequencingHow the same interaction is initialized, ordered, or terminated.[40] TRAILS-D summarizes design-level perturbations that may change the substantive setup of an LLM-based social simulation. Following the three-level structure of LLM social simulation defined in Section 3, we organize these perturbations into the eight dimensions shown in Table 1. These perturbations matter because collective outcomes are produced by the coupling between LLM agents and simulation design. A finding about a macro-level social phenomenon may therefore depend on design choices that are not part of the claimed social mechanism, such as the model substrate, agent specification, memory system, interaction protocol, intervention design, environment structure, or population composition. TRAILS-D asks whether the same substantive claim survives reasonable alternatives in these design choices. We provide some examples of such perturbations in Table 2. TRAILS-R complements TRAILS-D by focusing on perturbations that keep the substantive simula- tion condition fixed but change how that condition is represented to the LLM. This matters because LLM agents receive the simulation through text and interface structure, and prior work shows that small prompt variations can substantially affect LLM outputs [56,60]. Formatting, instruction order, labels, context compression, and interaction sequencing can affect how agents interpret a situation and produce behavior. TRAILS-R therefore audits representation and interface sensitivity by asking whether the same collective outcome survives these representation-level changes. If it does not, the finding should be reported as interface-sensitive rather than treated as robust evidence about the social process. We provide some examples of representation-level perturbations in Table 3. Prioritizing audits. No single study can audit every TRAILS dimension, and we do not suggest that every paper should. We propose three heuristics for deciding which dimensions to perturb. First, align audits with the claimâs mechanism: a claim that cooperation arises as a stable strategic equilibrium requires auditing persona format and game framing more than memory representation, because the former shape strategic disposition while the latter shape only how history is recalled. Second, ensure each perturbation spans a meaningful range: two prompt variants are weaker evidence than a sweep across several reasonable alternatives along the same axis. Third, audit across model families before claiming generality: as Section 4 shows, the same perturbation can produce dramatic effects in one model and essentially none in another, so claims phrased about âLLM agentsâ in general require evidence from multiple frontier models. Together with the claim-type hierarchy in Section 3, these heuristics let researchers calibrate audit scope to the strength of the claim being made. 6 Alternative Views and Discussion Alternative view 1: All empirical methods exhibit sensitivity. One objection is that LLM simu- lations are not unique: Markov chain Monte Carlo estimates vary with seeds, agent-based models depend on initialization and parameters, and human-subject studies are sensitive to sampling, mea- 9 surement, and analytic choices. We agree. But the fact that every other empirical tradition has developed protocols for handling its sensitivities is precisely why LLM-based social simulations need one of their own. Machine learning uses trainâtest splits, cross-validation, ablations, stress tests, and benchmarks; social science uses construct-validity checks, multiverse analysis [62], and specification curve analysis [61]. LLM simulations introduce a distinct perturbation surface: textual, high-dimensional, and semantically opaque [60,39]. TRAILS extends this validation logic by asking whether a simulation finding is stable enough to be worth validating against reality in the first place. Alternative view 2: LLM social simulations should be limited to hypothesis generation. A stronger objection is that LLM agents are not human, so robustness cannot make their outputs evidence about human social behavior. We take this seriously. The position we advance in Section 3 is already calibrated to this concern: audit strength scales with claim type, simulation complexity, and domain stakes. The field is already using LLM simulations as policy testbeds for platform interventions [27], prosocial-behavior policy evaluation [78], public-health interventions [25], and epidemic decision-making [5]âuses that fall at the policy-claim level of our framework and therefore demand the strongest audits. Claims that survive TRAILS-D and TRAILS-R perturbations across frontier models deserve more weight than those that do not, even if they remain weaker than human- subject evidence. We invite discussion on where this evidentiary gradient should sit. Alternative view 3: Premature standardization will calcify methodology before the field knows what it is doing. A third objection is that imposing audit standards now risks freezing in place a methodology whose design space is still being explored. We see this as a reason to keep TRAILS minimal and revisable, not to forgo standards entirely. TRAILS is a vocabulary for stating which perturbations were tested and which were not, not a fixed checklist. This vocabulary is precisely what lets the field accumulate evidence about which dimensions matter for which kinds of claimsâwork that would be foreclosed by either no standards or rigid ones. The risk of premature standardization is real; given the policy uses of LLM social simulations already underway, the risk of no standards is also real. Alternative view 4: Comprehensive robustness audits are prohibitively expensive. A fourth objection is practical: auditing every TRAILS dimension across multiple models and seeds quickly becomes infeasible, especially for small labs. We acknowledge this, and our framework is graded along two axes that respond to it. First, audit strength scales with claim strength: exploratory probes do not require comprehensive audits. Second, the prioritization heuristics in Section 5 let researchers concentrate their compute budget on the dimensions most likely to confound the claim being made. Transparent reporting of which dimensions were tested and which were left unaudited is itself a contribution, since it lets the field aggregate evidence across studies without requiring any single study to be exhaustive. A call for robustness audits in LLM social simulation. LLM-based social simulations are increas- ingly used to explore collective behavior, test mechanisms, and evaluate interventions. As these systems move toward scientific inference and socially consequential applications, plausible agent behavior and realistic-looking macro outcomes are no longer sufficient. TRAILS provides a starting vocabulary for this standard. To move toward robustness audits in LLM social simulation, we call on the community to pursue two lines of action: âąImprove robustness reporting in future LLM social simulation studies. Future simulation work should state the simulation scenario, domain stakes, and the evidentiary role of the simulation along the claim-type spectrum introduced in Section 3âexploratory probe, mechanism claim, or policy claimâand calibrate robustness audits accordingly. Studies should report which perturbations were tested, which findings remained stable, which findings were sensitive, and which dimensions were left unaudited. We acknowledge that no single study can test every possible perturbation; audits should prioritize the design and representation choices most relevant to the claim. âąBuild shared infrastructure for robustness auditing. The field needs reusable perturbation libraries for LLM social simulation settings. Open-source datasets and benchmarks for robustness audits are needed. It also needs more empirical work that systematically tests which perturbations matter most, when butterfly effects occur, and how micro- or meso-level instability propagates into macro-level outcomes. Only by testing how simulated societies change as their assumptions change can we know when they reveal something about the social world, rather than merely the machinery that produced them. 10 Code Availability The code for this paper is available athttps://github.com/angelayejinyi/butterfly-eff ect-sim. References [1]Sahar Abdelnabi, Amr Gomaa, Sarath Sivaprasad, Lea Schönherr, and Mario Fritz. Cooperation, competition, and maliciousness: Llm-stakeholders interactive negotiation, 2023. URL https: //arxiv.org/abs/2309.17234. [2] Elif Akata, Lion Schulz, Julian Coda-Forno, Seong Joon Oh, Matthias Bethge, and Eric Schulz. Playing repeated games with large language models. Nature Human Behaviour, 9:1134â1143, 2025. doi: 10.1038/s41562-025-02172-y. URLhttps://doi.org/10.1038/s41562-025 -02172-y. [3] Altera.AL, Andrew Ahn, Nic Becker, Stephanie Carroll, Nico Christie, Manuel Cortes, Arda Demirci, Melissa Du, Frankie Li, Shuying Luo, Peter Y. Wang, Mathew Willows, Feitong Yang, and Guangyu Robert Yang. Project sid: Many-agent simulations toward ai civilization, 2024. URL https://arxiv.org/abs/2411.00114. [4] Jacy Reese Anthis, Ryan Liu, Sean M Richardson, Austin C. Kozlowski, Bernard Koch, Erik Brynjolfsson, James Evans, and Michael S. Bernstein. Position: LLM social simulations are a promising research method. In Aarti Singh, Maryam Fazel, Daniel Hsu, Simon Lacoste- Julien, Felix Berkenkamp, Tegan Maharaj, Kiri Wagstaff, and Jerry Zhu, editors, Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 81005â81034. PMLR, 13â19 Jul 2025. URLhttps: //proceedings.mlr.press/v267/anthis25a.html. [5]Goshi Aoki and Navid Ghaffarzadegan. Ai agents as policymakers in simulated epidemics. arXiv preprint arXiv:2601.04245, 2026. URL https://arxiv.org/abs/2601.04245. [6] Ariel Flint Ashery, Luca Maria Aiello, and Andrea Baronchelli. Emergent social conventions and collective bias in llm populations. Science Advances, 11(20):eadu9368, 2025. URL https://w.science.org/doi/10.1126/sciadv.adu9368. [7] Francesco Bertolotti, Angela Locoro, and Luca Mari. Sensitivity to initial conditions in agent- based models. In Multi-Agent Systems and Agreement Technologies, volume 12520 of Lecture Notes in Computer Science, pages 501â508. Springer, 2020. doi: 10.1007/978-3-030-66412-1 _32. URL https://doi.org/10.1007/978-3-030-66412-1_32. [8] Philip Brookins and Jason DeBacker. Playing games with gpt: What can we learn about a large language model from canonical strategic games? Economics Bulletin, 44(1):25â37, 2024. URL https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4493398. [9]Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. Chateval: Towards better llm-based evaluators through multi-agent debate, 2023. URL https://arxiv.org/abs/2308.07201. [10]Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Chen Qian, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu, Maosong Sun, and Jie Zhou. Agentverse: Facilitating multi-agent collaboration and exploring emergent behaviors, 2023. URL https://arxiv.org/abs/2308.10848. [11]Ayush Chopra, Shashank Kumar, Nurullah Giray Kuru, Ramesh Raskar, and Arnau Quera- Bofarull. On the limits of agency in agent-based models. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems, pages 500â509, 2025. URL https://dl.acm.org/doi/10.5555/3709347.3743565. [12]Yun-Shiuan Chuang, Siddharth Suresh, Nikunj Harlalka, Agam Goyal, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy T Rogers. The wisdom of partisan crowds: Com- paring collective intelligence in humans and llm-based agents. arXiv preprint arXiv:2311.09665, 2023. URL https://arxiv.org/abs/2311.09665. 11 [13]Yun-Shiuan Chuang, Agam Goyal, Nikunj Harlalka, Siddharth Suresh, Robert Hawkins, Sijia Yang, Dhavan Shah, Junjie Hu, and Timothy Rogers. Simulating opinion dynamics with networks of llm-based agents. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 3326â3346, 2024. URLhttps://aclanthology.org/2024.findin gs-naacl.211/. [14] Jacob Cohen. Statistical power analysis for the behavioral sciences. Routledge, 2013. [15]Alessandro Di Stefano, Chrisina Jayne, Claudio Angione, and The Anh Han. Recognition of behavioural intention in repeated games using machine learning. In Artificial Life Conference Proceedings, volume 1, page 103. MIT Press One Rogers Street, Cambridge, MA 02142-1209, USA, 2023. URLhttps://direct.mit.edu/isal/proceedings/isal2023/35/103/ 116860? [16]Emilio Ferrara. The butterfly effect in artificial intelligence systems: Implications for AI bias and fairness. Machine Learning with Applications, 15:100525, 2024. URLhttps: //doi.org/10.1016/j.mlwa.2024.100525. [17] Antonino Ferraro, Antonio Galli, Valerio La Gatta, Marco Postiglione, Gian Marco Or- lando, Diego Russo, Giuseppe Riccio, Antonio Romano, and Vincenzo Moscato. Agent- based modelling meets generative ai in social network simulations. In International Con- ference on Advances in Social Networks Analysis and Mining, pages 155â170, 2024. URL https://link.springer.com/chapter/10.1007/978-3-031-78541-2_10. [18]NicolĂł Fontana, Francesco Pierri, and Luca Maria Aiello. Nicer than humans: How do large language models behave in the prisonerâs dilemma?In Proceedings of the International AAAI Conference on Web and Social Media, volume 19, pages 522â535, 2025. URLhttps: //arxiv.org/abs/2406.13605. [19]David C Funder and Daniel J Ozer. Evaluating effect size in psychological research: Sense and nonsense. Advances in Methods and Practices in Psychological Science, 2(2):156â168, 2019. URL https://journals.sagepub.com/doi/10.1177/2515245919847202. [20]Chen Gao, Xiaochong Lan, Nian Li, Yuan Yuan, Jingtao Ding, Zhilun Zhou, Fengli Xu, and Yong Li. Large language models empowered agent-based modeling and simulation: A survey and perspectives. Humanities and Social Sciences Communications, 11(1):1259, 2024. doi: 10.1057/s41599-024-03611-3. URL https://doi.org/10.1057/s41599-024-03611-3. [21] Chenhao Gu, Ling Luo, Zainab Razia Zaidi, and Shanika Karunasekera. Large language model driven agents for simulating echo chamber formation. arXiv preprint arXiv:2502.18138, 2025. URL https://arxiv.org/abs/2502.18138. [22]Sture Holm. A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, pages 65â70, 1979. URL https://w.jstor.org/stable/4615733. [23] Sirui Hong, Mingchen Zhuge, Jiaqi Chen, Xiawu Zheng, Yuheng Cheng, Ceyao Zhang, Jinlin Wang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and JĂŒrgen Schmidhuber. Metagpt: Meta programming for a multi-agent collaborative framework. In International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=VtmBAGCN7o. [24]John J Horton, Apostolos Filippas, and Benjamin S Manning. Large language models as simulated economic agents: What can we learn from homo silicus? Technical report, National Bureau of Economic Research, 2023. URLhttps://dl.acm.org/doi/10.1145/3670865 .3673513. [25]Abe Bohan Hou, Hongru Du, Yichen Wang, Jingyu Zhang, Zixiao Wang, Paul Pu Liang, Daniel Khashabi, Lauren Gardner, and Tianxing He. Can a society of generative agents simulate human behavior and inform public health policy? a case study on vaccine hesitancy. arXiv preprint arXiv:2503.09639, 2025. URL https://arxiv.org/abs/2503.09639. [26] Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang. War and peace (waragent): Large language model-based multi-agent simulation of world wars, 2023. URL https://arxiv.org/abs/2311.17227. 12 [27]Renhong Huang, Ning Tang, Jiarong Xu, Yuxuan Cao, Qingqian Tu, Sheng Guo, Bo Zheng, Huiyuan Liu, and Yang Yang. Policysim: An llm-based agent social simulation sandbox for proactive policy optimization. In Proceedings of the ACM Web Conference 2026, pages 4781â4792, 2026. URL https://dl.acm.org/doi/abs/10.1145/3774904.3792555. [28] Yanru Jiang and GĂŒl ̧sah Akçakır. Explicit cooperation shapes human-like multi-agent llm negotiation. In Proceedings of the 1st ICWSM Workshop on Integrating NLP and Psychology to Study Social Interactions, 2025. doi: 10.36190/2025.34. URLhttps://workshop-proceed ings.icwsm.org/abstract.php?id=2025_34. [29]Maik Larooij and Petter Törnberg. Validation is the central challenge for generative social simulation: A critical review of llms in agent-based modeling. Artificial Intelligence Review, 59 (1):15, 2025. doi: 10.1007/s10462-025-11412-6. URLhttps://doi.org/10.1007/s10462 -025-11412-6. [30]Ang Li, Haozhe Chen, Hongseok Namkoong, and Tianyi Peng. Llm generated persona is a promise with a catch. arXiv preprint arXiv:2503.16527, 2025. URLhttps://neurips.c/v irtual/2025/loc/san-diego/poster/121924. [31] Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qingmin Liao. Econagent: Large language model-empowered agents for simulating macroeconomic activities. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pages 15523â15536. Association for Computational Linguistics, 2024. URLhttps://aclanthology.org/2024. acl-long.829/. [32] Yuxuan Li and Hirokazu Shirado. Spontaneous giving and calculated greed in language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 5271â5286, 2025. URL https://aclanthology.org/2025.emnlp-main.267/. [33]Genglin Liu, Vivian Le, Salman Rahman, Elisa Kreiss, Marzyeh Ghassemi, and Saadia Gabriel. Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6390â6417. Association for Computational Linguistics, 2025. doi: 10.18653/v1/2025.e mnlp-main.325. URL https://aclanthology.org/2025.emnlp-main.325/. [34]Yuhan Liu, Xiuying Chen, Xiaoqing Zhang, Xing Gao, Ji Zhang, and Rui Yan. From skepticism to acceptance: Simulating the attitude dynamics toward fake news. In Proceedings of the Thirty- Third International Joint Conference on Artificial Intelligence, pages 7886â7894. International Joint Conferences on Artificial Intelligence Organization, 2024. doi: 10.24963/ijcai.2024/873. URL https://doi.org/10.24963/ijcai.2024/873. [35] Nunzio LorĂš and Babak Heydari. Strategic behavior of large language models and the role of game structure versus contextual framing. Scientific Reports, 14(1):18490, 2024. URL https://w.nature.com/articles/s41598-024-69032-z. [36]Michael W. Macy and Robert Willer. From factors to actors: Computational sociology and agent-based modeling. Annual Review of Sociology, 28(1):143â166, 2002. doi: 10.1146/annure v.soc.28.110601.141117. URLhttps://w.annualreviews.org/content/journals/1 0.1146/annurev.soc.28.110601.141117. [37]Zhao Mandi, Shreeya Jain, and Shuran Song. Roco: Dialectic multi-robot collaboration with large language models, 2023. URL https://arxiv.org/abs/2307.04738. [38]Qirui Mi, Mengyue Yang, Xiangning Yu, Zhiyu Zhao, Cheng Deng, Bo An, Haifeng Zhang, Xu Chen, and Jun Wang. Mf-llm: Simulating collective decision dynamics via a mean-field large language model framework, 2025. URL https://arxiv.org/abs/2504.21582. [39] Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of what art? a call for multi-prompt llm evaluation. Transactions of the Association for Computational Linguistics, 12:933â949, 2024. URLhttps://aclanthology.org/2024.ta cl-1.52/. 13 [40]Xinyi Mou, Zhongyu Wei, Qi Huang, and Xuanjing Wu. Unveiling the truth and facilitating change: Towards agent-based large-scale social movement simulation. In Findings of the Association for Computational Linguistics: ACL 2024, pages 4789â4809. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.findings-acl.285. URLhttps: //aclanthology.org/2024.findings-acl.285/. [41]Gian Marco Orlando, Jinyi Ye, Valerio La Gatta, Mahdi Saeedi, Vincenzo Moscato, Emilio Ferrara, and Luca Luceri. Emergent coordinated behaviors in networked llm agents: Modeling the strategic dynamics of information operations. In Proceedings of the ACM Web Conference 2026, pages 4805â4816, 2026. [42]Paul Ormerod and Bridget Rosewell. Validation and verification of agent-based models in the social sciences. In Flaminio Squazzoni, editor, Epistemological Aspects of Computer Simulation in the Social Sciences, volume 5466 of Lecture Notes in Computer Science, pages 130â140. Springer, Berlin, Heidelberg, 2009. doi: 10.1007/978-3-642-01109-2_10. URL https://doi.org/10.1007/978-3-642-01109-2_10. [43]Joon Sung Park, Joseph C. OâBrien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, UIST â23. Association for Computing Machinery, 2023. doi: 10.1145/3586183.3606763. URL https://doi.org/10.1145/3586183.3606763. [44]Pouya Pezeshkpour and Estevam Hruschka. Large language models sensitivity to the order of options in multiple-choice questions. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 2006â2017, Mexico City, Mexico, 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings- naacl.130. URLhttps: //aclanthology.org/2024.findings-naacl.130/. [45]PHF Science. The future is now: Revolutionising decision-making with ai-driven simulations. https://w.phfscience.nz/news-publications/the-future-is-now-revolutio nising-decision-making-with-ai-driven-simulations/, December 2024. Accessed: 2026-05-02. [46]Jinghua Piao, Zhihong Lu, Chen Gao, Fengli Xu, Qinghua Hu, Fernando P Santos, Yong Li, and James Evans. Emergence of human-like polarization among large language model agents. arXiv preprint arXiv:2501.05171, 2025. URL https://arxiv.org/abs/2501.05171. [47]Jinghua Piao, Yuwei Yan, Jun Zhang, Nian Li, Junbo Yan, Xiaochong Lan, Zhihong Lu, Zhiheng Zheng, Jing Yi Wang, Di Zhou, Chen Gao, Fengli Xu, Fang Zhang, Ke Rong, Jun Su, and Yong Li. Agentsociety: Large-scale simulation of llm-driven generative agents advances understanding of human behaviors and society. arXiv preprint arXiv:2502.08691, 2025. URL https://arxiv.org/abs/2502.08691. [48] Maximilian Puelma Touzel, Sneheel Sarangi, Gayatri Krishnakumar, Busra Tugce Gurbuz, Austin Welch, Zachary Yang, Andreea Musulan, Hao Yu, Ethan Kosak-Hine, Tom Gibbs, Camille Thibault, Reihaneh Rabbany, Jean-François Godbout, Dan Zhao, and Kellin Pelrine. Sandboxsocial: A sandbox for social media using multimodal ai agents. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, pages 11509â 11512. International Joint Conferences on Artificial Intelligence Organization, 2025. doi: 10.24963/ijcai.2025/1271. URL https://w.ijcai.org/proceedings/2025/1271. [49]Maximilian Puelma Touzel, Sneheel Sarangi, Aurelien BĂŒck-Kaeffer, Zachary Yang, Jean- François Godbout, and Reihaneh Rabbany. Position: Time to close the validation gap in llm social simulations, 2026. URLhttps://w.complexdatalab.com/stamina/papers/pu elmatouzel_CloseEvalGap.pdf. Preprint. [50]Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun. Chatdev: Communicative agents for software development. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15174â15186. Association for Computational Linguistics, 2024. doi: 10.18653/v1/2024.acl-long.810. URL https://aclanthology.org/2024.acl-long.810/. 14 [51]Aryan Razavi, Aref Jafari, Alona Fyshe, and Gholamreza Haffari. Benchmarking prompt sensitivity in large language models. arXiv preprint arXiv:2502.06065, 2025. URLhttps: //arxiv.org/abs/2502.06065. [52]Jiawei Ren, Yan Zhuang, Xiaokang Ye, Lingjun Mao, Xuhong He, Jianzhi Shen, Mrinaal Dogra, Yiming Liang, Ruixuan Zhang, Tianai Yue, Yiqing Yang, Eric Liu, Ryan Wu, Kevin Benavente, Rajiv Mandya Nagaraju, Muhammad Faayez, Xiyan Zhang, Dhruv Vivek Sharma, Xianrui Zhong, Ziqiao Ma, Tianmin Shu, Zhiting Hu, and Lianhui Qin. Simworld: An open- ended realistic simulator for autonomous agents in physical and social worlds, 2025. URL https://arxiv.org/abs/2512.01078. [53]Ruiyang Ren, Peng Qiu, Yingqi Qu, Jing Liu, Wayne Xin Zhao, Hua Wu, Ji-Rong Wen, and Haifeng Wang. Bases: Large-scale web search user simulation with large language model based agents. In Findings of the Association for Computational Linguistics: EMNLP 2024, 2024. URL https://aclanthology.org/2024.findings-emnlp.50/. [54]Siyue Ren, Zhiyao Cui, Ruiqi Song, Zhen Wang, and Shuyue Hu. Emergence of social norms in generative agent societies: Principles and architecture. In Kate Larson, editor, Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, pages 7895â7903. International Joint Conferences on Artificial Intelligence Organization, 8 2024. doi: 10.24963/ijcai.2024/874. URLhttps://doi.org/10.24963/ijcai.2024/874. Human-Centred AI. [55]Craig W. Reynolds. Flocks, herds and schools: A distributed behavioral model. In Proceedings of the 14th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH â87, pages 25â34. Association for Computing Machinery, 1987. doi: 10.1145/37402.37406. URL https://dl.acm.org/doi/10.1145/37402.37406. [56]Abel Salinas and Fred Morstatter. The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance. In Findings of the Association for Computational Linguistics: ACL 2024, pages 4629â4651, 2024. URLhttps://aclantholo gy.org/2024.findings-acl.275/. [57]Andrea Saltelli, Marco Ratto, Terry Andres, Francesca Campolongo, Jessica Cariboni, Debora Gatelli, Michaela Saisana, and Stefano Tarantola. Global Sensitivity Analysis: The Primer. John Wiley & Sons, Chichester, UK, 2008. ISBN 9780470059975. doi: 10.1002/9780470725184. URL https://doi.org/10.1002/9780470725184. [58]Robert G. Sargent. Verification and validation of simulation models. In Proceedings of the 2010 Winter Simulation Conference, pages 166â183. IEEE, 2010. doi: 10.1109/WSC.2010.5679166. URL https://doi.org/10.1109/WSC.2010.5679166. [59] Thomas C. Schelling. Dynamic models of segregation. Journal of Mathematical Sociology, 1 (2):143â186, 1971. doi: 10.1080/0022250X.1971.9989794. URLhttps://w.tandfonlin e.com/doi/abs/10.1080/0022250X.1971.9989794. [60] Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language modelsâ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In The Twelfth International Conference on Learning Representations, 2024. URL https://arxiv.org/abs/2310.11324. [61] Uri Simonsohn, Joseph P Simmons, and Leif D Nelson. Specification curve analysis. Nature Human Behaviour, 4(11):1208â1214, 2020. URLhttps://w.nature.com/articles/s4 1562-020-0912-z. [62] Sara Steegen, Francis Tuerlinckx, Andrew Gelman, and Wolf Vanpaemel. Increasing trans- parency through a multiverse analysis. Perspectives on Psychological Science, 11(5):702â712, 2016. URL https://pubmed.ncbi.nlm.nih.gov/27694465/. [63] Petter Törnberg, Diliara Valeeva, Justus Uitermark, and Christopher Bail. Simulating social media using large language models to evaluate alternative news feed algorithms. arXiv preprint arXiv:2310.05984, 2023. URL https://arxiv.org/abs/2310.05984. 15 [64]Alexander Sasha Vezhnevets, Jayd Matyas, Logan Cross, Davide Paglieri, Minsuk Chang, William A. Cunningham, Simon Osindero, William S. Isaac, and Joel Z. Leibo. Multi-actor generative artificial intelligence as a game engine, 2025. URLhttps://arxiv.org/abs/25 07.08892. [65]Chenxi Wang, Zongfang Liu, Dequan Yang, and Xiuying Chen. Decoding echo chambers: Llm-powered simulations revealing polarization in social networks. In Proceedings of the 31st International Conference on Computational Linguistics, pages 3913â3923, 2025. URL https://aclanthology.org/2025.coling-main.264/. [66]Lei Wang, Heyang Gao, Xiaohe Bo, Xu Chen, and Ji-Rong Wen. YuLan-OneSim: Towards the next generation of social simulator with large language models. In NeurIPS 2025 Workshop on Scientific Methods for Understanding Deep Learning, 2025. URLhttps://arxiv.org/abs/ 2505.07581. [67]Lei Wang, Jingsen Zhang, Hao Yang, Zhi-Yuan Chen, Jiakai Tang, Zeyu Zhang, Xu Chen, Yankai Lin, Ruihua Song, Wayne Xin Zhao, Jun Xu, Zhicheng Dou, Jun Wang, and Ji-Rong Wen. User behavior simulation with large language model-based agents. ACM Transactions on Information Systems, 43(2):1â37, 2025. doi: 10.1145/3708985. URLhttps://doi.org/10 .1145/3708985. [68] Zhilin Wang, Yu Ying Chiu, and Yu Cheung Chiu. Humanoid agents: Platform for simulating human-like generative agents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 167â176. Association for Computational Linguistics, 2023. doi: 10.18653/v1/2023.emnlp-demo.15. URLhttps: //aclanthology.org/2023.emnlp-demo.15/. [69]Uri Wilensky and William Rand. Making models match: Replicating an agent-based model. Journal of Artificial Societies and Social Simulation, 10(4):2, 2007. URLhttps://w.jass s.org/10/4/2.html. [70]Richard Willis, Yali Du, and Joel Z Leibo. Will systems of llm agents lead to cooperation: An investigation into a social dilemma. In 24th International Conference on Autonomous Agents and Multiagent Systems, AAMAS 2025, pages 2786â2788. International Foundation for Autonomous Agents and Multiagent Systems (IFAAMAS), 2025. URLhttps://dl.acm.o rg/doi/10.5555/3709347.3744012. [71]Paul Windrum, Giorgio Fagiolo, and Alessio Moneta. Empirical validation of agent-based models: Alternatives and prospects. Journal of Artificial Societies and Social Simulation, 10(2): 8, 2007. URL https://ideas.repec.org/a/jas/jasssj/2006-40-2.html. [72]Yuzhe Yang, Yifei Zhang, Minghao Wu, Kaidi Zhang, Yunmiao Zhang, Honghai Yu, Yan Hu, and Benyou Wang. Twinmarket: A scalable behavioral and social simulation for financial markets. In The Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS), volume 39 of NeurIPS, 2025. URL https://arxiv.org/abs/2502.01506. [73]Ziyi Yang, Zaibin Zhang, Zirui Zheng, Yuxian Jiang, Ziyue Gan, Zhiyu Wang, Zijian Ling, Jinsong Chen, Martz Ma, Bowen Dong, Prateek Gupta, Shuyue Hu, Zhenfei Yin, Guohao Li, Xu Jia, Lijun Wang, Bernard Ghanem, Huchuan Lu, Chaochao Lu, Wanli Ouyang, Yu Qiao, Philip Torr, and Jing Shao. Oasis: Open agent social interaction simulations with one million agents, 2024. URL https://arxiv.org/abs/2411.11581. [74]Wenzhen Zheng and Xijin Tang. Simulating social network with llm agents: an analysis of information propagation and echo chambers. In International Symposium on Knowledge and Systems Sciences, pages 63â77. Springer, 2024. URLhttps://link.springer.com/chap ter/10.1007/978-981-96-0178-3_5. [75]Jiaxu Zhou, Jen-tse Huang, Xuhui Zhou, Man Ho Lam, Xintao Wang, Hao Zhu, Wenxuan Wang, and Maarten Sap. The pimmur principles: Ensuring validity in collective behavior of llm societies. arXiv preprint arXiv:2509.18052, 2025. URLhttps://arxiv.org/abs/2509.1 8052. 16 [76]Shuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig. Webarena: A realistic web environment for building autonomous agents, 2023. URLhttps://arxiv.org/abs/23 07.13854. [77]Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Zhengyang Qi, Haofei Yu, Louis- Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. Sotopia: Interactive evaluation for social intelligence in language agents. In The Twelfth International Conference on Learning Representations (ICLR), 2024. URLhttps://openreview.net/f orum?id=mM7VurbA4r. [78]Yujia Zhou, Hexi Wang, Qingyao Ai, Zhen Wu, and Yiqun Liu. Investigating prosocial behavior theory in llm agents under policy-induced inequities. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 2254â2262, 2026. URLhttps: //ojs.aaai.org/index.php/AAAI/article/view/37209. 17 A TRAILS (Taxonomy for Robustness Audits In LLM Simulations) Table 2: TRAILS-D: Design-level perturbations LevelDimensionWhat to perturb MicroModel substrateModel family, model size, base vs. instruction-tuned models, alignment, temperature, top-p, random seed, safeguards. Agent specification Demographics, ideology, personality, goals, prior beliefs, real- data-grounded vs. synthetic personas, relationship attributes. Internal state and cognitionBeliefs, attitudes, emotions, needs, moral values, reasoning style, reflection rules, belief updating, preference updating. Memory and temporalitylast-kmemory, summarized memory, episodic memory, retrieval rules, forgetting, reflection frequency, temporal granularity. MesoInteraction protocolTurn order, fixed or free turns, dyadic or group interaction, avail- able actions Intervention designModerator type, target selection, timing, frequency, rule-based vs. LLM-based intervention. MacroEnvironment structureNetwork topology, feed ranking, recommendation system, moder- ation rules, platform affordances, institutional rules. Population and scale Number of agents, demographic distribution, ideological balance, heterogeneity, real vs. synthetic population construction. Table 3: TRAILS-R: Representation-level perturbations CategoryPerturbationExample Representational formatFormattingNarrative prose vs. bullet points; paragraph vs. table; bolded keywords; numbered lists. Delimiter choiceConversation history enclosed in code blocks, XML tags, JSON, or plain text. Output schema Free-form paragraph vs. short reply vs. JSON with rationale. Instruction hierarchyInstruction orderâUpdate your opinion, then write a replyâ vs. âWrite a reply, then update your opinion.â Role/message placementThe same instruction placed in the system message vs. the user message. Linguistic framingAgent namingAgent A/B vs. Alice/Bob. Persona wordingâLeftâ vs. âDemocratâ; âRightâ vs. âRepublican.â Moderation framingSame intervention framed as a warning, suggestion, commu- nity reminder, or bridge statement. Few-shot examplesSame examples framed differently. Context representationMemory representationSame history shown as transcript, JSON list, bullet summary. Context windowFull history vs. last three turns; same content with early background removed. Message metadataAdding timestamps, usernames. Interaction sequencingInitial seed turnSame stance, but a slightly different opening statement. Turn orderWhich agent speaks first. 18 B Extended Related Works B.1 From Agent-Based Modeling to Generative Social Simulation Agent-based modeling (ABM) has provided a bottom-up framework for studying how macro-level social patterns emerge from micro-level agents and local interactions [55,59,36]. However, tradi- tional ABMs have also been criticized for relying on hand-coded behavioral rules and challenges in empirical validation, replication, and sensitivity to modeling choices [71,42,69,7]. Recent advances in LLMs make generative agents a promising way. The literature on LLM agents and multi-agent systems for social simulation is already broad, ranging from collaborative task teams [50,23,76] to embodied agents [52], game-like environments [64], user simulations in recommendation system [67], economic simulations [31], and social media simulations [73,48,21]. In this work, we focus on a narrower scope: LLM-based social simulations for collective human behavior. Scope.Building on Silicon Societies [49] and the eligibility criteria for generative social simulation proposed by Anthis et al.[4], Larooij and Törnberg[29], we focus on LLM-based simulations in which individual human-like agents generate behavior, interact within a shared social environment, and produce collective human outcomes. In this sense, social refers to interdependent interaction among agents, while collective refers to aggregate patterns emerging from those interactions. We define this scope through a three-level chain: micro-level human-like agents whose reasoning, memory, attitudes, decisions, and communication are generated or mediated by LLMs; meso-level social interaction among agents within shared environments and institutional rules; and macro-level collective phenomena such as polarization, echo chambers, norm emergence, information diffusion, and mobilization. Additionally, we exclude three classes of work to focus on modeling collective human behavior: single-agent or user simulations without social interaction among agents [53]; simulations where agents primarily represent non-human, embodied, or aggregate actors such as robots, states, countries, or game agents [3,26,37,52]; and multi-agent LLM systems designed primarily for task completion or evaluation rather than modeling human social behavior [9, 10, 23, 50, 76] Social simulation scenarios. Within this scope, we distinguish two broad scenarios. In goal- or incentive-structured scenarios, agents act in relatively well-specified decision environments with explicit or implicit incentives, constraints, rules, or payoffs. Collective outcomes arise through strategic aggregation, as in negotiation and cooperation [1, 2, 28], population decision-making [38], financial market dynamics [72], or macroeconomic activity [31]. In open-ended scenarios, agents do not share a single explicit task, objective, or payoff structure. Macro-level patterns emerge from micro-level agent behavior and meso-level interaction, as in social media discussion [73,48], opinion dynamics and mobilization [40], polarization and echo chambers [21,46,65,73], information or misinformation spread [34,33,43], conformity and herd effects [73,67], emergent coordination [43], social norm formation [54], and user behavior in digital platforms [67]. The key distinction between the two types is whether collective outcomes arise primarily through strategic aggregation in structured decision settings or through social emergence from less structured social environments. By allowing agents to reason, remember, communicate, and adapt through natural language, LLM- based agents appear to address the challenge of limited behavioral realism in ABM. Recent work has used LLM agents to simulate emergent coordination [43], macroeconomic decision-making [31], social media interaction [73,48], echo chamber formation [21], social movement dynamics [40], fake news attitude dynamics [34], and information dissemination and moderation [33]. These studies illustrate the promise of LLM agents for modeling collective behavior. However, LLMs do not automatically solve the other challenges of simulation. As Larooij and Törnberg[29]argue, generative social simulations may increase behavioral realism, but they also introduce new sources of uncertainty, including black-box model behavior, cultural biases, stochastic variation, alignment effects, and prompt sensitivity. Existing studies have begun to improve the realism, empirical grounding, and mechanistic validity of LLM social simulations through real-world grounding [21,33,67], empirical comparison [73,21], theory-driven design [40,34], and human evaluation [43,48]. Yet robustness and sensitivity analysis remain underdeveloped. Fewer studies test whether the same collective outcome survives reasonable perturbations in design choices, such as prompt wording, agent profiles, memory representation or interaction protocols. 19 B.2 Repeated Prisonerâs Dilemma as a Goal-Structured Testbed The repeated Prisonerâs Dilemma is a foundational paradigm for studying the tension between short- term self-interest and long-term mutual benefit, and a recurring testbed for LLM-agent cooperation [2,18,70,8,32]. Recent work has used this paradigm to make claims about whether LLMs are âcooperativeâ or âselfish,â whether they reason from history, and how their behavior compares to human play [2,24]. We use it here as a tightly controlled environment: the action space, payoffs, and horizon are fully specified, leaving architectural and prompt-level choices as the main sources of variation. We focus on three small perturbations that are common in LLM-agent implementations but rarely audited jointly. First, we vary persona format, because personas are widely used to specify agent goals, traits, and social roles, yet recent work shows that persona-based simulations can introduce systematic validity risks if their construction is not carefully validated [30]. Our test asks a narrower representation-level question: whether the same persona content behaves differently when written as prose, a descriptive list, or a table. Second, we vary game-instruction framing, because prior work shows that LLM behavior in repeated games can change when the same strategic setting is described differently, for example through robustness checks, payoff-matrix variations, or reasoning prompts [2,35]. Third, we vary memory representation, because repeated-game behavior depends on how agents read and use prior interaction history; existing Prisonerâs Dilemma studies explicitly test whether LLMs can parse gameplay logs and condition decisions on historical behavior [18]. Together, these perturbations test whether claims about cooperation, selfishness, or strategic adaptation survive small changes in how the same game, agent, and history are represented to the model. B.3 Echo Chambers as an Open-Ended Social Simulation Testbed Echo chambers and polarization are among the most studied collective phenomena in LLM-based social simulation [21,65,17,74,46,63,73,13]. Recent work has used LLM agents on social networks to ask whether echo chambers emerge, what role homophily and recommendation play, and how interventions reshape exposure [21,63,33,48]. We adopt this paradigm as a representative open-ended simulation in which collective patterns emerge from local interaction rather than explicit optimization. We test five perturbations that target core design choices in open-ended social-network simulations: homophily, hub assignment, activation probability, memory window, and recommendation feed size. These choices are motivated by prior echo-chamber and LLM-agent studies, which show that echo chambers emerge from the joint effects of network structure, selective exposure, memory/context, and engagement dynamics. First, we vary input-network homophily, because prior work treats homophily and structurally clustered networks as central mechanisms through which like-minded users become locally concentrated and echo chambers form [21,65,17]. Second, we vary hub assignment while holding the degree sequence fixed, because scale-free social networks contain highly connected users whose placement can disproportionately affect information diffusion and group reinforcement; prior LLM-agent simulations similarly emphasize that scale-free or hub-dominated networks better approximate real social media than fully connected or random interaction settings. Third, we vary activation probability, because social media outcomes depend not only on who is connected to whom, but also on which users become active and contribute content in each round. Fourth, we vary the memory window, since LLM-agent frameworks commonly rely on short-term or retrieved interaction histories to condition agent behavior, but the amount of remembered context is often treated as an implementation detail rather than an audited design choice. Finally, we vary recommendation feed size, because prior work shows that selective exposure and recommendation mechanisms can increase engagement, reduce cross-group interaction, and strengthen echo-chamber formation [65,17]. Together, these perturbations span two levels of the simulation architecture: structural opportunity conditions, through homophily and hub assignment, and interaction-level exposure conditions, through activation, memory, and feed size. This allows us to test whether macro-level echo-chamber metrics are more sensitive to the underlying network topology or to seemingly small choices in how agents are activated, exposed to content, and prompted with recent interaction history. 20 C Additional Simulation Details C.1 Shared Experimental Protocol Across both case studies, we usegpt-5.2as the primary model for all reported analyses and replicate the experiments on three additional state-of-the-art, frontier models (claude-haiku-4-5, gemini-2.5-flash, anddeepseek-v3) as a cross-model robustness check. We prioritize frontier models because they offer the most reliable foundation for creating realistic social simulations. If sensitivity occurs within these advanced models, it indicates a significant finding rather than a simple lack of model performance or intelligence. Results for additional models are provided in Appendix D. Each condition is repeated withN = 30independent simulation runs using distinct random seeds, and the simulation run is the unit of analysis. We report pairwise mean differences, 95% confidence intervals, and effect sizes measured by Cohenâsd. Following conventional benchmarks, we interpret dâ 0.2as small,dâ 0.5as medium, anddâ 0.8as large, while treating these values as descriptive guidelines rather than hard cutoffs [14,19]. For statistical testing, we use two-sided MannâWhitney U tests instead of pairwiset-tests because several outcomes are bounded, discrete, or have zero run-level variance in some conditions, making variance-based tests undefined or unreliable. Because each perturbation involves multiple pairwise comparisons, we apply Holm correction within each metric to control the error rate [22]. Across both case studies, we hold decoding and parsing settings fixed within each comparison so that observed differences can be attributed to the intended perturbation rather than to changes in generation parameters. Unless otherwise specified, model calls use temperature= 0.3and top-p = 0.95. Temperature controls the randomness of token sampling, while top-pcontrols the cumulative probability mass from which tokens are sampled. In all experiments, agents are required to return a structured JSON response containing their selected action and a brief reason. C.2 Prisonerâs Dilemma Simulation Repeated Prisoner's Dilemma 10 rounds, 30 simulations per condition, fixed payoff matrix Hold game fixed; vary 3 prompt dimensions Persona format Plain, Descriptive, Tabular Game framing Canonical, Moralized, Risk Memory format Table/narrative à ±stats Run each variant in two interaction modes Single-agent mode LLM agent vs 4 fixed policies: TitForTat, Random, AlwaysC, AlwaysD Two-agent mode Two identical LLM agents play each other Measure two outcomes per simulation Average payoff 0â5 per round, mean reward Cooperation rate 0â1, fraction of C choices Figure 4: Repeated Prisonerâs Dilemma case-study design. We hold the underlying game fixed across conditions: agents play a 10-round repeated Prisonerâs Dilemma with the same payoff matrix, using 30 independent simulations per condition. We then vary three prompt dimensionsâpersona format, game-instruction framing, and memory formatâand run each variant in two interaction modes: a single-agent mode, where one LLM agent plays against four fixed policies, and a two-agent mode, where two identically configured LLM agents play each other. Outcomes are measured using average payoff and cooperation rate. C.2.1 Setup and Configurations Each Prisonerâs Dilemma episode runs forT = 10rounds. In each round, each player simultaneously chooses Cooperate (C) or Defect (D). The action space is fixed to COOPERATE and DEFECT, with payoffs (C, C) = (3, 3), (C, D) = (0, 5), (D, C) = (5, 0), and (D, D) = (1, 1). 21 Following common practice in previous studies [2,18,15], we run two interaction modes. In the single-agent mode, one LLM-controlled agent plays against four fixed external policies: TitForTat, which cooperates on the first round and then mirrors the opponentâs previous action; Random, which independently selectsCorDwith equal probability each round; AlwaysCooperate, which unconditionally playsCin every round; and AlwaysDefect, which unconditionally playsDin every round. In the two-agent mode, two LLM-controlled agents with identical configurations play one another. For two-agent runs, both agents are queried once per round, and the prompt order is randomized across rounds to reduce stable positional bias. If the model response cannot be parsed as a valid action after one retry, the simulation falls back to DEFECT; this conservative fallback avoids dropping failed episodes while preserving a fixed action space. These defaults are implemented consistently across the persona-format, game-instruction, and memory-construction blocks. C.2.2 Evaluation Metrics For each agenti, leta i,t âC, Ddenote the action chosen by agentiin roundt, whereCindicates cooperation and D indicates defection. Let u i,t denote the payoff received by agent i in that round. Average Payoff. Average payoff measures the agentâs overall performance under the gameâs incentive structure: AveragePayoff i = 1 T T X t=1 u i,t . Cooperation Rate.Cooperation rate measures how often the agent chooses cooperation across the repeated game: CooperationRate i = 1 T T X t=1 I(a i,t = C). Here,T = 10is the number of rounds andI(·)is an indicator function equal to 1 when the condition is true and 0 otherwise. C.2.3 Perturbation Details We test three perturbations, each targeting a design choice that is inconsistently reported and rarely ablated across the LLM-cooperation literature: persona format, game instruction framing, and memory representation. While Fontana et al.[18]carefully document their prompting pipeline and study memory window size in isolation, and Akata et al.[2]report their prompt structure, neither paper systematically varies all three dimensions jointly or tests their interaction effects on cooperative behavior. Persona format. Holding the underlying meaning fixedâthe agent is strategic and aims to max- imize its payoffâwe vary only the presentation across three formats: a paragraph of plain prose (PLAIN), a descriptive bullet list (DESCRIPTIVE), and a structured keyâvalue table (TABULAR). Game-instruction framing. We hold the payoff matrix and persona fixed, using PLAIN as the default persona, and vary only the framing of the game instructions. We test three framings: canonical game-theory framing (CANONICAL), which uses standard labels such as âCooperateâ and âDefectâ; moralized framing (MORALIZED), which describes the same actions as âcooperate fairlyâ versus âexploitâ; and risk framing (RISK), which presents cooperation as safer but vulnerable and defection as riskier but potentially higher-reward. Memory representation. We hold the persona fixed to PLAIN and the game framing fixed to CANONICAL, and vary only how the interaction history is shown to the agent. The design is2Ă 2: history format (table vs. narrative) crossed with whether summary statistics are included, such as cooperation rates, cumulative payoffs, and joint outcome counts. 22 C.3 Echo-Chamber Simulation Echo chamber on a social network 100 personas, frozen stances on AI regulation, 15 rounds power-law network, neighbors-only feed Hold population and rules fixed; vary 5 design choices Structural conditions who can interact with whom Network homophily low / medium / higher assortativity Hub assignment Interaction-level conditions how agents are activated and exposed Activation probability Memory window Recommendation feed size Each round, active agents pick one action from their neighbor feed PostRepostReplyDo nothing Build the interaction network and measure 4 metrics Stance assortativity in-group connection tendency Same-group edge ratio weighted in-group interactions Figure 5: Echo-chamber case-study design. We simulate 100 LLM agents discussing whether advanced AI systems should be regulated on a fixed social network. Each agent has a persona, follower count, and frozen stance toward AI regulation. In each round, activated agents see recent posts from their direct neighbors only and choose one action:POST,REPOST,REPLY, orSILENT. We vary five architectural perturbationsâinput-network homophily, hub assignment, recommendation feed size, memory window size, and activation probabilityâand measure whether these choices shift macro-level echo-chamber outcomes, including stance assortativity and weighted same-group edge ratio. C.3.1 Setup and Configurations The echo-chamber simulation uses 100 agents loaded from a persona file and a fixed edgelist. Each persona includes a short bio and a stance on a focal topic: whether advanced AI systems should be regulated. Stances are represented on a five-point scale, from strongly against regulation to strongly support regulation. Agent stances are frozen throughout the run, so any outcome differences reflect exposure, interaction, and network-structure effects rather than belief updating. Each run lasts 15 rounds by default. In each round, agents become active independently with activation probability0.3; if no agent is sampled as active, one agent is selected at random so that every round contains at least one decision. Each active agent sees recent posts from direct network neighbors only, with a default feed size of 5 and memory window of 2 rounds. The action space is fixed to POST, REPOST, REPLY, and DO NOTHING. Reposts and replies must target a visible feed item; invalid targets are converted to a post or silence depending on feed availability. The default implementation uses 4 worker threads, temperature = 0.3, and top-p = 0.95. C.3.2 Evaluation Metrics LetG I = (V, E I )denote the interaction graph induced by observed interactions during a simulation run, where an edge(i, j)â E I indicates that agentireposted or replied to agentj. Letg i denote agent iâs stance group, derived from its position on AI regulation: anti-regulation, neutral, or pro-regulation. LetA ij be the weighted adjacency matrix ofG I , whereA ij is the number of times agentiinteracted with agent j. Stance Assortativity. Stance assortativity measures whether agents preferentially interact with agents who hold similar positions on AI regulation. We compute the categorical assortativity coefficient over stance-group labels: r = P uâG p u â P uâG p out u p in u 1â P uâG p out u p in u , whereG =anti-regulation, neutral, pro-regulationis the set of stance groups. The termp uv denotes the fraction of weighted interaction edges that go from agents in stance grouputo agents in stance group v: p uv = P i,j A ij 1[g i = u, g j = v] P i,j A ij . 23 The marginal terms are defined as p out u = X vâG p uv ,p in u = X vâG p vu . Intuitively, P uâG p u is the observed fraction of interactions that occur within the same stance group, while P uâG p out u p in u is the expected within-group fraction under random mixing that preserves the overall amount of outgoing and incoming interaction for each stance group. The coefficient ranges fromâ1to1: positive values indicate more within-stance-group interaction than expected by chance, values near0indicate approximately random mixing with respect to AI-regulation stance, and negative values indicate more cross-stance-group interaction than expected. Same-Group Edge Ratio. Same-group edge ratio is the fraction of interaction edges that occur between agents with the same side label: SameGroupRatio = P (i,j)âE I I(s i = s j ) |E I | . For weighted interaction graphs, we use the weighted version: SameGroupRatio w = P i,j A ij I(s i = s j ) P i,j A ij . This metric ranges from0to1, where0means all interactions are cross-side and1means all interac- tions are within-side. In a binary-side setting with balanced groups, a value near0.5corresponds to roughly equal within-side and cross-side interaction; values above0.5indicate stronger same-group interaction, while values below 0.5 indicate more cross-group interaction. C.3.3 Perturbation Details Prior work highlights network structure, homophily, recommendation-mediated exposure, and memo- ry/context as central mechanisms in echo-chamber formation, making these perturbations natural targets for robustness testing [21,65,17]. We test five perturbations that target common but often under-audited design choices in LLM echo-chamber simulations: input-network homophily (LOW, MEDIUM, HIGH), hub assignment (ANTI HUB, PRO HUB, MIXED HUB, RANDOM HUB), activation probability (0.3,0.5), memory window (2and4rounds), and recommendation feed size (5,10posts). These choices span structural conditions that shape who can interact with whom and interaction- level conditions that shape who becomes active, what content agents see, and what recent context conditions their actions. How do we decide the fixed degree sequence? To hold overall connectivity constant across conditions, we construct a fixed power-law degree sequence for the same set of agents. Specifically, we sample a power-law-like sequence with target mean degree8and exponent2.4, then rescale and adjust it until it is graphical. This produces a heavy-tailed network structure in which a small number of agents occupy high-degree positions while most agents have fewer connections. Because the same degree sequence is reused across network variants, differences between conditions cannot be attributed to changes in average degree, density, or the overall presence of hubs. How do we generate different homophily networks?We vary the level of stance homophily while preserving the fixed degree sequence. LetG 0 = (V, E 0 )denote the input network and letx i denote agentiâs numeric AI-regulation stance on the original five-point scale. We define input-network homophily as the numeric stance assortativity of the input graph: h(G 0 ) = P i,jâV A (0) ij (x i â Ìx)(x j â Ìx) P i,jâV A (0) ij (x i â Ìx) 2 , whereA (0) ij is the adjacency matrix of the input networkG 0 , and Ìxis the mean stance among connected endpoints. Intuitively,h(G 0 )is high when connected agents tend to have similar stance values and low when edges frequently connect agents with different stances. 24 Starting from a simple graph that realizes the fixed degree sequence, we apply degree-preserving double-edge swaps to move the graph into a target stance-assortativity band. The rewiring heuristic uses edge-level stance similarity, sim(i, j) = 1â |x i â x j | 4 , so that agents with identical stances have similarity1, while agents at opposite ends of the five-point scale have similarity 0. We generate three homophily conditions: lower homophily : h(G 0 )â [0.02, 0.10], medium homophily : h(G 0 )â [0.12, 0.20], higher homophily : h(G 0 )â [0.22, 0.30]. Thus, the homophily manipulation changes which stance groups are connected to one another while preserving each agentâs degree. How do we reassign hubs?In a separate manipulation, we vary which agents occupy high-degree hub positions while keeping the same power-law degree sequence. We sort or shuffle agents before assigning degrees, so that the largest degrees can be assigned to different types of agents. In the RANDOM condition, degree assignments are randomly shuffled across agents. In the MIXED HUBS condition, high-degree positions are interleaved across anti-regulation, neutral, and pro-regulation agents, using follower count to rank agents within each stance group. In the PRO HUBS and ANTI HUBS conditions, high-degree positions are preferentially assigned to pro-regulation or anti-regulation agents, respectively. After assigning degrees, we again generate a simple graph and rewire it into the same medium homophily band,h(G 0 )â [0.12, 0.20]. This hub-assignment manipulation changes which agents are structurally central while holding fixed the degree sequence and approximate stance homophily. C.4 Recorded Outputs We record outputs at both the decision level and the run level. For the Prisonerâs Dilemma experiments, each row records one agent decision, including the perturbation condition, seed, round, agent and opponent actions, payoffs, parse-validity indicator, parse note, raw model output, and full prompt input. These fields allow us to compute cooperation rates, payoff trajectories, invalid-output rates, and condition-level behavioral shifts. For the echo-chamber experiments, we save agent-round logs, round summaries, run metadata, and final agent states. The agent-round logs record each activated agentâs profile, network position, action, target post or author when applicable, generated message, exposure diversity, parse validity, full prompt input, and raw model output. Round summaries capture aggregate metrics such as active agents, posts created, cross-group interaction, exposure diversity, graph-side assortativity, polarization, and hub attention. Run metadata records the simulation configuration and input-network properties, while final agent states store each agentâs initial and final stance, degree, follower count, and hub status. Together, these outputs support behavioral analysis, robustness diagnostics, and checks of parsing reliability and network-level dynamics. 25 D Extended Experimental Results D.1 Prisonerâs Dilemma: Additional Figures This section reports the extended Prisonerâs Dilemma results for perturbations that are discussed more briefly in the main text. Figure 6 shows how different framings of the same payoff matrix affect payoff and cooperation, while Figure 7 shows the effect of changing how prior interaction history is represented. Together, these figures provide the full pairwise comparisons and run-level distributions for the game-instruction and memory-representation perturbations. Canonical game theory framing (CANONICAL) You are playing a repeated Prisoner's Dilemma for 10 rounds. In each round, you choose one of two actions: Cooperate Defect The payoffs for each round are: If both players cooperate: both receive 3 points If you cooperate and the other defects: you receive 0, the other receives 5 If you defect and the other cooperates: you receive 5, the other receives 0 If both defect: both receive 1 Moralized framing (MORALIZED)Risk framing (RISK) You will interact with another participant for 10 rounds. In each round, you must decide whether to: Cooperate fairly with the other person Exploit the other person for personal gain The outcomes are: If both cooperate fairly: both receive 3 points If you cooperate but they exploit you: you receive 0, they receive 5 If you exploit them while they cooperate: you receive 5, they receive 0 If both exploit each other: both receive 1 You will make decisions over 10 rounds against another participant. Each round involves a choice: Cooperate (safer but vulnerable) Defect (riskier but potentially higher reward) Outcomes: Both cooperate: both get 3 You cooperate, other defects: you get 0, they get 5 You defect, other cooperates: you get 5, they get 0 Both defect: both get 1 Figure 6: Effect of game-instruction framing on Prisonerâs Dilemma outcomes. Results are shown for the single-agent setting (left) and two-agent setting (right). The heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across instruction framings (p < .05, two-sided MannâWhitney U test), while gray cells indicate non-significant comparisons. Bar plots show run-level distributions for each framing. Bar colors denote instruction framing: CANONICAL (blue), MORALIZED (orange), and RISK (green). 26 Table History: Round | You | Opponent | Your payoff | Opponent payoff 1 | Cooperate | Defect | 0 | 5 2 | Defect | Defect | 1 | 1 3 | Cooperate | Cooperate | 3 | 3 ... NarrativeStatistics History: In round 1, you chose Cooperate and the opponent chose Defect. You received 0 points and the opponent received 5. In round 2, you chose Defect and the opponent chose Defect. You received 1 point and the opponent received 1. ... Summary Statistics: - You cooperated 2 times and defected 1 time. - Your cooperation rate: 0.667 - Opponent cooperated 1 time and defected 2 times. - Opponent cooperation rate: 0.333 - Your cumulative payoff: 4 - Opponent cumulative payoff: 9 - Your average payoff per round: 1.333 - Opponent average payoff per round: 3.000 - Joint outcomes: C=1, CD=1, DC=0, D=1 Figure 7: Effect of memory representation on Prisonerâs Dilemma outcomes. Results are shown for the single-agent setting (left) and two-agent setting (right). Memory conditions vary the history format (table vs. narrative) and whether summary statistics are included. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate (p < .05, two-sided Mannâ Whitney U test), while gray cells indicate non-significant comparisons. Bar plots show run-level distributions for each memory condition. 27 D.2 Echo Chamber: Additional Figures Activation probability, memory window, and recommendation feed size. Figure 8 reports the effects of varying activation probability, memory window, and recommendation feed size. Activation probability and memory window do not produce statistically significant differences in either final stance assortativity or weighted same-group edge ratio. In contrast, increasing the recommendation feed size from5to10posts significantly increases both stance assortativity (M = 0.247to0.277, p = 0.014) and weighted same-group edge ratio (M = 0.542to0.568,p = 0.003). This suggests that, among these three parameters, recommendation feed size is the only one that consistently strengthens echo-chamber-like interaction. 0.30.5 0.20 0.22 0.24 0.26 0.28 Stance assortativity Activation probability 24 0.20 0.22 0.24 0.26 0.28 Memory window 510 0.200 0.225 0.250 0.275 0.300 0.325 p = 0.014 Recommendation feed size 0.30.5 0.52 0.54 0.56 0.58 0.60 Weighted same-group edge ratio 24 0.48 0.50 0.52 0.54 0.56 0.58 510 0.54 0.56 0.58 p = 0.003 Figure 8: Effects of activation probability, memory window, and recommendation feed size on echo-chamber outcomes. Boxplots show the effects of varying activation probability, memory window, and recommendation feed size while holding the input network fixed. The top row reports final stance assortativity and the bottom row reports weighted same-group edge ratio, both computed on the simulated interaction network. Significant pairwise differences are annotated using two-sided MannâWhitney U tests with p < 0.05. 28 D.3 Prisonerâs Dilemma: Experiment Results With Other LLMs We usegpt-5.2as the primary frontier model in the main text and replicate the Prisonerâs Dilemma experiments on three additional state-of-the-art, frontier models:claude-haiku-4-5, gemini-2.5-flash, anddeepseek-v3. These models are included as a cross-model robustness check. The goal is not to benchmark model quality, but to test whether the observed sensitivity to persona format, game-instruction framing, and memory representation also appears beyond a single model family. If similar sensitivity appears across this diverse set of closed and open-weight model families, then the effect is less likely to be an artifact of one specific model. TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 5 Agent Payoff GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 5 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 5 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 DeepSeek-V3 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 DeepSeek-V3 Persona Format PlainDescriptiveTabular Figure 9: Persona-format effects in the single-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate for each persona format, opponent policy, and model. TitForTat Random AlwaysCooperate AlwaysDefect PLAIN DESCRIPTIVE PLAIN TABULAR DESCRIPTIVE TABULAR Persona Format +0.96 p=0.000 -0.54 p=0.000 -0.05 p=0.032 +1.60 p=0.000 -0.40 p=0.003 -2.00 p=0.000 -0.10 p=0.000 +0.64 p=0.000 -1.46 p=0.000 -0.05 p=0.000 GPT-5.2 Payoff TitForTat Random AlwaysCooperate AlwaysDefect +0.43 p=0.000 +0.30 p=0.026 -1.59 p=0.000 +0.06 p=0.038 +1.46 p=0.000 -0.30 p=0.029 -1.60 p=0.000 +1.02 p=0.000 -0.60 p=0.000 -0.06 p=0.038 Claude Haiku 4.5 Payoff TitForTat Random AlwaysCooperate AlwaysDefect +0.26 p=0.023 -0.15 p=0.000 -0.03 p=0.006 +0.29 p=0.016 -0.15 p=0.000 -0.03 p=0.001 Gemini 2.5 Flash Payoff TitForTat Random AlwaysCooperate AlwaysDefect -0.37 p=0.000 -0.08 p=0.048 -0.19 p=0.000 -0.15 p=0.026 +0.19 p=0.001 DeepSeek-V3 Payoff TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy PLAIN DESCRIPTIVE PLAIN TABULAR DESCRIPTIVE TABULAR Persona Format +0.60 p=0.000 +0.17 p=0.000 +0.27 p=0.000 +0.05 p=0.032 +1.00 p=0.000 +0.26 p=0.000 +1.00 p=0.000 +0.10 p=0.000 +0.40 p=0.000 +0.09 p=0.000 +0.73 p=0.000 +0.05 p=0.000 GPT-5.2 Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy +0.20 p=0.000 -0.20 p=0.000 +0.80 p=0.000 -0.06 p=0.038 +0.73 p=0.000 +0.19 p=0.000 +0.80 p=0.000 +0.52 p=0.000 +0.39 p=0.000 +0.06 p=0.038 Claude Haiku 4.5 Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy +0.11 p=0.000 +0.06 p=0.000 +0.08 p=0.000 +0.03 p=0.006 +0.12 p=0.000 +0.06 p=0.000 +0.08 p=0.000 +0.03 p=0.001 Gemini 2.5 Flash Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy +0.18 p=0.000 +0.19 p=0.000 +0.07 p=0.000 +0.09 p=0.000 -0.11 p=0.000 -0.09 p=0.001 DeepSeek-V3 Cooperation Rate 2 1 0 1 2 Mean diff (payoff) 1.0 0.5 0.0 0.5 1.0 Mean diff (coop rate) Figure 10: Pairwise persona-format differences in the single-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across persona formats. 29 AB 0 1 2 3 Payoff GPT-5.2 AB 0 1 2 3 Claude Haiku 4.5 AB 0.0 0.5 1.0 1.5 2.0 2.5 Gemini 2.5 Flash AB 0 1 2 3 DeepSeek-V3 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Persona Format PlainDescriptiveTabular Figure 11: Persona-format effects in the two-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate when both LLM agents use the same persona format. Agent AAgent B PLAIN DESCRIPTIVE PLAIN TABULAR DESCRIPTIVE TABULAR Persona Format +1.42 p=0.000 +1.54 p=0.000 -0.10 p=0.012 +0.12 p=0.012 -1.52 p=0.000 -1.42 p=0.000 GPT-5.2 Payoff Agent AAgent B +1.50 p=0.000 +1.47 p=0.000 +0.30 p=0.000 +0.23 p=0.001 -1.21 p=0.000 -1.24 p=0.000 Claude Haiku 4.5 Payoff Agent AAgent B +0.74 p=0.000 +0.81 p=0.000 +0.77 p=0.000 +0.80 p=0.000 Gemini 2.5 Flash Payoff Agent AAgent B +0.16 p=0.000 +0.12 p=0.001 +0.14 p=0.000 +0.11 p=0.004 DeepSeek-V3 Payoff Agent AAgent B Agent PLAIN DESCRIPTIVE PLAIN TABULAR DESCRIPTIVE TABULAR Persona Format +0.76 p=0.000 +0.74 p=0.000 +0.03 p=0.023 -0.73 p=0.000 -0.75 p=0.000 GPT-5.2 Coop Rate Agent AAgent B Agent +0.77 p=0.000 +0.78 p=0.000 +0.14 p=0.000 +0.15 p=0.000 -0.64 p=0.000 -0.63 p=0.000 Claude Haiku 4.5 Coop Rate Agent AAgent B Agent +0.36 p=0.000 +0.35 p=0.000 +0.36 p=0.000 +0.36 p=0.000 Gemini 2.5 Flash Coop Rate Agent AAgent B Agent +0.01 p=0.045 +0.01 p=0.045 +0.07 p=0.000 +0.08 p=0.000 +0.06 p=0.000 +0.07 p=0.000 DeepSeek-V3 Coop Rate 1.5 1.0 0.5 0.0 0.5 1.0 1.5 Mean diff (payoff) 0.75 0.50 0.25 0.00 0.25 0.50 0.75 Mean diff (coop rate) Figure 12: Pairwise persona-format differences in the two-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across persona formats. 30 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 Agent Payoff GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 5 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 DeepSeek-V3 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 DeepSeek-V3 Game Instruction Framing CanonicalMoralizedRisk Figure 13: Game-instruction framing effects in the single-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate for each instruction framing, opponent policy, and model. TitForTat Random AlwaysCooperate AlwaysDefect CANONICAL MORALIZED CANONICAL RISK MORALIZED RISK Game Instruction Framing -0.31 p=0.000 +0.41 p=0.000 +0.32 p=0.003 -0.58 p=0.000 +0.62 p=0.000 -0.99 p=0.000 GPT-5.2 Payoff TitForTat Random AlwaysCooperate AlwaysDefect +0.17 p=0.000 +0.21 p=0.000 +0.16 p=0.000 -0.46 p=0.001 +0.13 p=0.000 Claude Haiku 4.5 Payoff TitForTat Random AlwaysCooperate AlwaysDefect Gemini 2.5 Flash Payoff TitForTat Random AlwaysCooperate AlwaysDefect +0.18 p=0.048 -0.23 p=0.000 -0.29 p=0.000 +0.07 p=0.014 DeepSeek-V3 Payoff TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy CANONICAL MORALIZED CANONICAL RISK MORALIZED RISK Game Instruction Framing -0.22 p=0.000 -0.21 p=0.000 +0.21 p=0.000 +0.29 p=0.000 +0.42 p=0.000 +0.50 p=0.000 GPT-5.2 Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy -0.17 p=0.000 -0.11 p=0.000 +0.12 p=0.001 -0.11 p=0.000 -0.09 p=0.000 +0.29 p=0.000 -0.06 p=0.000 Claude Haiku 4.5 Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy Gemini 2.5 Flash Cooperation Rate TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy +0.11 p=0.001 +0.12 p=0.000 +0.11 p=0.000 +0.14 p=0.000 -0.07 p=0.014 DeepSeek-V3 Cooperation Rate 0.5 0.0 0.5 Mean diff (payoff) 0.4 0.2 0.0 0.2 0.4 Mean diff (coop rate) Figure 14:Pairwise game-instruction framing differences in the single-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across instruction framings. 31 AB 0 1 2 3 Payoff GPT-5.2 AB 0 1 2 3 Claude Haiku 4.5 AB 0.0 0.5 1.0 1.5 2.0 2.5 Gemini 2.5 Flash AB 0 1 2 3 DeepSeek-V3 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Game Instruction Framing CanonicalMoralizedRisk Figure 15: Game-instruction framing effects in the two-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate when both LLM agents receive the same game framing. Agent AAgent B CANONICAL MORALIZED CANONICAL RISK MORALIZED RISK Game Instruction Framing -0.42 p=0.000 -0.42 p=0.000 +0.29 p=0.002 +0.32 p=0.001 GPT-5.2 Payoff Agent AAgent B -0.25 p=0.000 -0.27 p=0.000 Claude Haiku 4.5 Payoff Agent AAgent B +0.20 p=0.032 +0.17 p=0.042 Gemini 2.5 Flash Payoff Agent AAgent B -0.09 p=0.038 +0.14 p=0.014 +0.19 p=0.000 +0.11 p=0.043 DeepSeek-V3 Payoff Agent AAgent B Agent CANONICAL MORALIZED CANONICAL RISK MORALIZED RISK Game Instruction Framing -0.19 p=0.000 -0.19 p=0.000 +0.15 p=0.000 +0.14 p=0.000 GPT-5.2 Coop Rate Agent AAgent B Agent -0.07 p=0.008 -0.07 p=0.008 Claude Haiku 4.5 Coop Rate Agent AAgent B Agent +0.09 p=0.006 +0.09 p=0.009 +0.08 p=0.019 +0.07 p=0.037 Gemini 2.5 Flash Coop Rate Agent AAgent B Agent -0.03 p=0.007 +0.07 p=0.000 +0.06 p=0.003 +0.08 p=0.000 +0.10 p=0.000 DeepSeek-V3 Coop Rate 0.4 0.2 0.0 0.2 0.4 Mean diff (payoff) 0.1 0.0 0.1 Mean diff (coop rate) Figure 16: Pairwise game-instruction framing differences in the two-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across instruction framings. 32 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 Agent Payoff GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 5 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect 0 1 2 3 4 DeepSeek-V3 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate GPT-5.2 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Claude Haiku 4.5 TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 Gemini 2.5 Flash TitForTat Random AlwaysCooperate AlwaysDefect Opponent Policy 0.0 0.2 0.4 0.6 0.8 1.0 DeepSeek-V3 Memory Representation Table-RawNarrative-RawTable+StatsNarrative+Stats Figure 17: Memory-representation effects in the single-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate for each memory condition, opponent policy, and model. Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation GPT-5.2 Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation -0.13 p=0.029 -0.13 p=0.040 -0.16 p=0.007 -0.16 p=0.011 Claude Haiku 4.5 Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Gemini 2.5 Flash Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation DeepSeek-V3 Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation -0.05 p=0.023 +0.05 p=0.011 GPT-5.2 Cooperation Rate Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation +0.06 p=0.002 +0.09 p=0.000 +0.11 p=0.000 +0.05 p=0.014 Claude Haiku 4.5 Cooperation Rate Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Gemini 2.5 Flash Cooperation Rate Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation DeepSeek-V3 Cooperation Rate 0.1 0.0 0.1 Mean diff (payoff) 0.10 0.05 0.00 0.05 0.10 Mean diff (coop rate) Figure 18: Pairwise memory-representation differences in the single-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across memory conditions. 33 AB 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Payoff GPT-5.2 AB 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Claude Haiku 4.5 AB 0.0 0.5 1.0 1.5 2.0 2.5 Gemini 2.5 Flash AB 0.0 0.5 1.0 1.5 2.0 2.5 3.0 DeepSeek-V3 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Cooperation Rate AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 AB Agent 0.0 0.2 0.4 0.6 0.8 1.0 Memory Representation Table-RawNarrative-RawTable+StatsNarrative+Stats Figure 19: Memory-representation effects in the two-agent Prisonerâs Dilemma across four models. Bar plots show average payoff and cooperation rate when both LLM agents use the same memory representation. Table-Raw Narrative-Raw Table+Stats Narrative+Stats Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation -0.09 p=0.021 -0.06 p=0.038 -0.08 p=0.040 GPT-5.2 Agent A Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Claude Haiku 4.5 Agent A Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats +0.21 p=0.031 +0.19 p=0.038 +0.34 p=0.001 Gemini 2.5 Flash Agent A Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats DeepSeek-V3 Agent A Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation +0.06 p=0.000 -0.04 p=0.002 GPT-5.2 Agent A Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Claude Haiku 4.5 Agent A Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats +0.09 p=0.003 +0.11 p=0.000 +0.13 p=0.000 Gemini 2.5 Flash Agent A Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats DeepSeek-V3 Agent A Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation -0.11 p=0.006 +0.15 p=0.001 -0.08 p=0.045 GPT-5.2 Agent B Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Claude Haiku 4.5 Agent B Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats +0.23 p=0.020 +0.23 p=0.021 Gemini 2.5 Flash Agent B Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats DeepSeek-V3 Agent B Payoff Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation -0.06 p=0.000 -0.02 p=0.019 +0.04 p=0.030 -0.05 p=0.000 GPT-5.2 Agent B Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation Claude Haiku 4.5 Agent B Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation +0.09 p=0.001 +0.10 p=0.000 +0.15 p=0.000 Gemini 2.5 Flash Agent B Cooperation Table-Raw Narrative-Raw Table+Stats Narrative+Stats Memory Representation DeepSeek-V3 Agent B Cooperation 0.2 0.0 0.2 Mean diff (payoff) 0.1 0.0 0.1 Mean diff (coop rate) 0.2 0.0 0.2 Mean diff (payoff) 0.1 0.0 0.1 Mean diff (coop rate) Figure 20: Pairwise memory-representation differences in the two-agent Prisonerâs Dilemma across four models. Heatmaps report statistically significant pairwise differences in average payoff and cooperation rate across memory conditions. 34 D.4 Echo Chamber: Experiment Results With Other LLMs We additionally replicate the echo chamber experiments withclaude-haiku-4-5, gemini-2.5-flash, anddeepseek-v3.This cross-model check tests whether sensitivity to network structure and exposure design persists beyond the primarygpt-5.2runs, rather than reflecting the behavior of one model family alone. Lower assortativity Medium assortativity Higher assortativity 0.15 0.20 0.25 0.30 0.35 Stance assortativity p < 0.001 p = 0.002 p < 0.001 GPT-5.2 Lower assortativity Medium assortativity Higher assortativity 0.3 0.4 0.5 0.6 p = 0.008 p = 0.008 p = 0.008 Claude Haiku 4.5 Lower assortativity Medium assortativity Higher assortativity 0.45 0.50 0.55 0.60 0.65 0.70 0.75 p = 0.008 p = 0.008 Gemini 2.5 Flash Lower assortativity Medium assortativity Higher assortativity 0.0 0.1 0.2 0.3 0.4 p = 0.008 p = 0.008 p = 0.008 DeepSeek-V3 Lower assortativity Medium assortativity Higher assortativity 0.450 0.475 0.500 0.525 0.550 0.575 0.600 Weighted same-group edge ratio p < 0.001 p < 0.001 Lower assortativity Medium assortativity Higher assortativity 0.65 0.70 0.75 0.80 p = 0.008 p = 0.032 p = 0.008 Lower assortativity Medium assortativity Higher assortativity 0.75 0.80 0.85 0.90 p = 0.008 p = 0.036 p = 0.012 Lower assortativity Medium assortativity Higher assortativity 0.40 0.45 0.50 0.55 0.60 0.65 p = 0.008 p = 0.008 p = 0.008 Figure 21: Cross-model results for initial network homophily. Boxplots compare final stance assor- tativity and weighted same-group edge ratio across input networks with different initial stance assorta- tivity levels. Results are shown separately forgpt-5.2,claude-haiku-4-5,gemini-2.5-flash, and deepseek-v3. 35 Anti hubs Pro hubs Mixed hubs Random hubs 0.40 0.42 0.44 0.46 0.48 0.50 0.52 0.54 Stance assortativity GPT-5.2 Anti hubs Pro hubs Mixed hubs Random hubs 0.15 0.20 0.25 0.30 0.35 0.40 Claude Haiku 4.5 Anti hubs Pro hubs Mixed hubs Random hubs 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 Gemini 2.5 Flash Anti hubs Pro hubs Mixed hubs Random hubs 0.10 0.15 0.20 0.25 0.30 0.35 0.40 p = 0.008 p = 0.008 p = 0.032 p = 0.032 DeepSeek-V3 Anti hubs Pro hubs Mixed hubs Random hubs 0.70 0.75 0.80 0.85 0.90 Weighted same-group edge ratio p = 0.008 p = 0.008 p = 0.016 p = 0.032 p = 0.008 Anti hubs Pro hubs Mixed hubs Random hubs 0.50 0.55 0.60 0.65 0.70 Anti hubs Pro hubs Mixed hubs Random hubs 0.78 0.80 0.82 0.84 0.86 0.88 0.90 0.92 p = 0.012 p = 0.012 p = 0.012 Anti hubs Pro hubs Mixed hubs Random hubs 0.45 0.50 0.55 0.60 0.65 0.70 p = 0.008 p = 0.008 p = 0.032 Figure 22: Cross-model results for hub assignment. Boxplots compare final stance assortativity and weighted same-group edge ratio when the highest-degree nodes are assigned to anti-regulation, pro-regulation, mixed, or random agents. Results are shown separately for each model. Activation prob 0.3Activation prob 0.5 0.20 0.22 0.24 0.26 0.28 0.30 Stance assortativity GPT-5.2 Activation prob 0.3Activation prob 0.5 0.42 0.43 0.44 0.45 0.46 0.47 Claude Haiku 4.5 Activation prob 0.3Activation prob 0.5 0.52 0.54 0.56 0.58 0.60 0.62 0.64 Gemini 2.5 Flash Activation prob 0.3Activation prob 0.5 0.17 0.18 0.19 0.20 0.21 0.22 0.23 DeepSeek-V3 Activation prob 0.3Activation prob 0.5 0.52 0.54 0.56 0.58 0.60 Weighted same-group edge ratio Activation prob 0.3Activation prob 0.5 0.71 0.72 0.73 0.74 0.75 0.76 0.77 p = 0.008 Activation prob 0.3Activation prob 0.5 0.83 0.84 0.85 0.86 0.87 p = 0.032 Activation prob 0.3Activation prob 0.5 0.50 0.52 0.54 0.56 0.58 Figure 23: Cross-model results for activation probability. Boxplots compare final stance assorta- tivity and weighted same-group edge ratio under different agent activation probabilities. Results are shown separately for each model. 36 Memory window 2Memory window 4 0.20 0.22 0.24 0.26 0.28 0.30 Stance assortativity GPT-5.2 Memory window 2Memory window 4 0.41 0.42 0.43 0.44 0.45 0.46 0.47 0.48 Claude Haiku 4.5 Memory window 2Memory window 4 0.56 0.58 0.60 0.62 0.64 Gemini 2.5 Flash Memory window 2Memory window 4 0.18 0.20 0.22 0.24 0.26 DeepSeek-V3 Memory window 2Memory window 4 0.48 0.50 0.52 0.54 0.56 0.58 Weighted same-group edge ratio Memory window 2Memory window 4 0.70 0.71 0.72 0.73 0.74 0.75 Memory window 2Memory window 4 0.820 0.825 0.830 0.835 0.840 0.845 0.850 Memory window 2Memory window 4 0.50 0.52 0.54 0.56 0.58 p = 0.016 Figure 24: Cross-model results for memory window. Boxplots compare final stance assortativity and weighted same-group edge ratio under different memory-window lengths. Results are shown separately for each model. Feed size 5 Feed size 10 0.20 0.22 0.24 0.26 0.28 0.30 0.32 0.34 Stance assortativity p = 0.014 GPT-5.2 Feed size 5 Feed size 10 0.42 0.44 0.46 0.48 0.50 Claude Haiku 4.5 Feed size 5 Feed size 10 0.58 0.60 0.62 0.64 Gemini 2.5 Flash Feed size 5 Feed size 10 0.18 0.20 0.22 0.24 0.26 DeepSeek-V3 Feed size 5 Feed size 10 0.54 0.56 0.58 0.60 Weighted same-group edge ratio p = 0.003 Feed size 5 Feed size 10 0.72 0.73 0.74 0.75 0.76 0.77 Feed size 5 Feed size 10 0.825 0.830 0.835 0.840 0.845 0.850 0.855 0.860 Feed size 5 Feed size 10 0.50 0.52 0.54 0.56 0.58 Figure 25: Cross-model results for recommendation feed size. Boxplots compare final stance assortativity and weighted same-group edge ratio when agents are exposed to different numbers of recent neighbor posts. Results are shown separately for each model. 37 E Simulation Prompts E.1 Prisonerâs Dilemma Prompts E.1.1 Base Persona Across the Prisonerâs Dilemma prompt-perturbation experiments reported in the paper, we use the following strategic persona: You are a strategic participant focused on maximizing your payoff. Study opponent behavior for exploitable patterns, and adjust your strategy to gain advantage while staying competitive in the long term. E.1.2 Base Single-Agent Prompt Template In the single-agent setting, the model plays against a fixed external opponent policy. The opponent policy is not revealed to the model. Persona: [BASE PERSONA] Game setup: - Repeated Prisonerâs Dilemma for [HORIZON] rounds. - Each round, both players simultaneously choose one action: Cooperate or Defect. - You cannot observe the opponentâs current-round action before choosing. - Opponent follows a fixed external policy, unknown to you. Payoffs: - If both choose Cooperate: (3, 3). - If you choose Cooperate and opponent chooses Defect: (0, 5). - If you choose Defect and opponent chooses Cooperate: (5, 0). - If both choose Defect: (1, 1). Task: Choose a reasonable action this round based on history and repeated-game context. [HISTORY / MEMORY BLOCK] Output format: Return valid JSON with keys: - action: one of ["Cooperate", "Defect"] - reason: short string E.1.3 Base Two-Agent Prompt Template In the two-agent setting, both agents are model-driven players. Each agent receives the same prompt structure but with role-specific history. Persona: [STRATEGIC PERSONA] Role: Your role label this run: Agent [A/B]. Game setup: - Repeated Prisonerâs Dilemma for [HORIZON] rounds. - Each round, both agents simultaneously choose one action: Cooperate or Defect. - You cannot observe the other agentâs current-round action before choosing. - The other agent is another model-driven player with the same prompt framing. Payoffs: - If both choose Cooperate: (3, 3). - If you choose Cooperate and the other agent chooses Defect: (0, 5). - If you choose Defect and the other agent chooses Cooperate: (5, 0). - If both choose Defect: (1, 1). Task: Choose a reasonable action this round based on history and repeated-game context. [ROLE-SPECIFIC HISTORY / MEMORY BLOCK] Output format: Return valid JSON with keys: - action: one of ["Cooperate", "Defect"] - reason: short string 38 E.1.4 Persona Format Variants For the strategic persona-format perturbation, we compare surface-level format changes while preserving the same strategic meaning. Plain strategic persona. You are a strategic participant focused on maximizing your long-term payoff. Study opponent behavior for exploitable patterns, and adjust your strategy to gain advantage while staying competitive. Descriptive strategic persona. Persona profile: - Maximize your long-term payoff - Study opponent for exploitable patterns - Adjust strategy to gain advantage - Stay strategic and competitive Tabular strategic persona. Persona | Value Goal | Maximize your long-term payoff Observation focus | Study opponent for exploitable patterns Strategy | Adjust strategy to gain advantage Style | Stay strategic and competitive E.1.5 Game-Instruction Framing Variants For the game-instruction perturbation, we hold the strategic persona and payoff matrix fixed while changing the framing of the game description. The paper reports three framing variants: canonical, moralized, and risk-framed. Canonical game-theory framing. You are playing a repeated Prisonerâs Dilemma for 10 rounds. In each round, you choose one of two actions: - Cooperate - Defect The payoffs for each round are: - If both players cooperate: both receive 3 points. - If you cooperate and the other defects: you receive 0, the other receives 5. - If you defect and the other cooperates: you receive 5, the other receives 0. - If both defect: both receive 1. Your goal is to maximize your total points over all rounds. At each round, you will see the history of previous rounds. Moralized framing. You will interact with another participant for 10 rounds. In each round, you must decide whether to: - Cooperate fairly with the other person. - Exploit the other person for personal gain. The outcomes are: - If both cooperate fairly: both receive 3 points. - If you cooperate but they exploit you: you receive 0, they receive 5. - If you exploit them while they cooperate: you receive 5, they receive 0. - If both exploit each other: both receive 1. Your goal is to maximize your total points. At each round, you will see past interactions. Risk framing. You will make decisions over 10 rounds against another participant. Each round involves a choice: - Cooperate, which is safer but vulnerable. 39 - Defect, which is riskier but potentially higher reward. Outcomes: - Both cooperate: both get 3. - You cooperate, other defects: you get 0, they get 5. - You defect, other cooperates: you get 5, they get 0. - Both defect: both get 1. Your goal is to maximize total payoff. At each round, you will see previous outcomes. E.1.6 Memory Representation Variants For the memory perturbation, we vary whether previous rounds are shown as a structured table or as narrative text, and whether raw history is supplemented with summary statistics. M1: Table, raw history only. History: Round | You | Opponent | Your payoff | Opp payoff 1 | [YOUR ACTION] | [OPPONENT ACTION] | [YOUR PAYOFF] | [OPPONENT PAYOFF] 2 | [YOUR ACTION] | [OPPONENT ACTION] | [YOUR PAYOFF] | [OPPONENT PAYOFF] ... M2: Narrative, raw history only. History: In round 1, you chose [YOUR ACTION] and the opponent chose [OPPONENT ACTION]. You received [YOUR PAYOFF] points and the opponent received [OPPONENT PAYOFF]. In round 2, you chose [YOUR ACTION] and the opponent chose [OPPONENT ACTION]. You received [YOUR PAYOFF] points and the opponent received [OPPONENT PAYOFF]. ... M3: Table, raw history plus statistics. History: Round | You | Opponent | Your payoff | Opp payoff 1 | [YOUR ACTION] | [OPPONENT ACTION] | [YOUR PAYOFF] | [OPPONENT PAYOFF] 2 | [YOUR ACTION] | [OPPONENT ACTION] | [YOUR PAYOFF] | [OPPONENT PAYOFF] ... Summary statistics: - You cooperated [N] times and defected [N] times. - Your cooperation rate: [RATE]. - Opponent cooperated [N] times and defected [N] times. - Opponent cooperation rate: [RATE]. - Your cumulative payoff: [PAYOFF]. - Opponent cumulative payoff: [PAYOFF]. - Your average payoff per round: [AVG]. - Opponent average payoff per round: [AVG]. - Joint outcomes: C=[N], CD=[N], DC=[N], D=[N]. M4: Narrative, raw history plus statistics. History: In round 1, you chose [YOUR ACTION] and the opponent chose [OPPONENT ACTION]. You received [YOUR PAYOFF] points and the opponent received [OPPONENT PAYOFF]. In round 2, you chose [YOUR ACTION] and the opponent chose [OPPONENT ACTION]. You received [YOUR PAYOFF] points and the opponent received [OPPONENT PAYOFF]. ... Summary statistics: - You cooperated [N] times and defected [N] times. - Your cooperation rate: [RATE]. - Opponent cooperated [N] times and defected [N] times. - Opponent cooperation rate: [RATE]. - Your cumulative payoff: [PAYOFF]. - Opponent cumulative payoff: [PAYOFF]. - Your average payoff per round: [AVG]. - Opponent average payoff per round: [AVG]. - Joint outcomes: C=[N], CD=[N], DC=[N], D=[N]. For the two-agent memory condition, the same four memory variants are used, but the history is role-specific. For Agent A, âYouâ refers to Agent A and âOther agentâ refers to Agent B; for Agent B, the mapping is reversed. 40 E.1.7 Output Schema and Parsing Instruction All Prisonerâs Dilemma prompts end with the same JSON parsing instruction. Output format: Return valid JSON with keys: - action: one of ["Cooperate", "Defect"] - reason: short string If a model response cannot be parsed as valid JSON, the parser attempts to extract a JSON object from the raw text. If parsing still fails, the fallback action is Defect. E.2 Echo-Chamber Prompts E.2.1 Agent Persona Template Each echo-chamber agent is initialized with a minimal persona containing an agent ID, a short bio, a follower count, and a fixed stance toward AI regulation. Stance is represented on a five-point scale from strong opposition to strong support: s i ââ2,â1, 0, 1, 2. "agent_id": "[AGENT ID]", "bio": "[SHORT BIO]", "follower_count": [FOLLOWER COUNT], "ai_regulation_stance": [-2|-1|0|1|2] The stance scale is: -2 = strongly oppose stronger AI regulation -1 = somewhat oppose stronger AI regulation 0 = mixed or undecided 1 = somewhat support stronger AI regulation 2 = strongly support stronger AI regulation E.2.2 Round Decision Prompt At each round, an activated agent receives its own profile, the number of neighbors in the fixed network, and a feed consisting only of recent posts from neighboring agents. You are a social media user discussing one issue: whether governments should impose stronger regulation on advanced AI systems. Your profile: - agent_id: [AGENT ID] - bio: [BIO] - follower_count: [FOLLOWER COUNT] - current stance on AI regulation: [STANCE] where -2 strongly oppose, -1 somewhat oppose, 0 mixed or undecided, 1 somewhat support, 2 strongly support. - neighbor count in the network: [NEIGHBOR COUNT] Simulation rules: - This is a fixed network for this run. - You only see recent posts from your neighbors. - Keep behavior simple and realistic. - Do not invent facts. Use short social-media-style language. - You may choose exactly one action this round: POST, REPOST, REPLY, or SILENT. - REPOST or REPLY should only target one visible post from the feed below. - Your stance is fixed for the entire run and must not change. - Set updated_stance equal to your current stance. Round: [ROUND] / [TOTAL ROUNDS] Visible feed from your neighbors: [VISIBLE FEED] Return valid JSON with exactly these keys: "action": "POST|REPOST|REPLY|SILENT", 41 "target_index": integer or null, "message": "short text, or empty string if SILENT", "updated_stance": -2|-1|0|1|2, "reason": "brief reason" If target_index is used, it must refer to one visible feed item by its 1-based index. E.2.3 Visible Feed Format When neighbor posts are visible, they are rendered as a numbered list. [1] post_id=[POST ID] | author=[AUTHOR ID] | action_type=[POST/REPOST/REPLY] | text=[POST TEXT] [2] post_id=[POST ID] | author=[AUTHOR ID] | action_type=[POST/REPOST/REPLY] | target_post_id=[TARGET POST ID] | text=[POST TEXT] ... If the agent has no visible recent neighbor posts, the feed is rendered as: No recent posts from your neighbors. E.2.4 Output Schema and Parsing Instruction All echo-chamber prompts require the following JSON schema. "action": "POST|REPOST|REPLY|SILENT", "target_index": integer or null, "message": "short text, or empty string if SILENT", "updated_stance": -2|-1|0|1|2, "reason": "brief reason" The parser accepts only four actions: a t i âPOST, REPOST, REPLY, SILENT. IfREPOSTorREPLYis selected, the target must refer to a visible feed item by its one-indexed position. The agentâs stance is frozen throughout the run, soupdated_stanceis constrained to equal the current stance: s t+1 i = s t i . 42