Paper deep dive
Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems
Augustin Chan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/8/2026, 7:25:46 PM
Summary
This paper investigates how injecting symbolic reasoning frameworks (I-Ching, Tarot, scrambled text) into a single LLM agent within a 7-player Warring States Diplomacy variant triggers emergent, memory-mediated ecosystem dynamics. Rather than altering per-decision risk posture, the frameworks act as small perturbations that, through accumulating campaign memory and multi-agent interactions, produce distinct, condition-specific winner distributions. The study concludes that alignment-framework choices at the agent level yield non-additive, system-level consequences driven by memory accumulation and multi-agent dynamics rather than direct per-decision modulation.
Entities (8)
Relation Signals (7)
Han â neverwins â any_condition
confidence 95% · The framework-receiving agent (Han) never wins under any condition
Han â receives â I-Ching
confidence 95% · Hanâ is the intervention recipient... Yarrow Yarrow-stalk I-Ching cast
I-Ching â triggers â Qin suppression
confidence 95% · under I-Ching yarrow divination, Qin (the strongest expansionist) wins 0 of 10 games while Yan and Chu co-dominate
Memory Accumulation â mediates â Ecosystem Dynamics
confidence 90% · the conditions settle into distinct, condition-associated winner ecosystems... attributed to emergent memory and multi-agent dynamics
Claude Opus 4.6 â powers â Warring States Diplomacy
confidence 90% · Each of the 7 states is controlled by an independent Claude Opus 4.6 agent instance
Scrambled-text ablation â triggers â Qi dominance
confidence 90% · under scrambled text, Qi wins 5 of 10
Tarot â triggers â Qin dominance
confidence 90% · under Tarot, Qin wins 5 of 10
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models exhibit a risk-averse "turtle" bias as strategic agents. We show that injecting a symbolic reasoning framework as a per-round reflective prompt into one agent acts as a small perturbation whose consequences are not per-decision but emergent: the agent's risk posture is unchanged in isolation, yet over a campaign of accumulating memory and multi-agent interaction the conditions settle into distinct, condition-associated winner ecosystems. In a 7-player Warring States Diplomacy variant (61 games, 6 conditions), the winner distribution differs sharply across the four primary conditions (41 games; permutation omnibus p approximately 0.001): control -> Yan (7/11); I-Ching yarrow -> Yan/Chu co-dominance with Qin fully suppressed (0/10); Tarot -> Qin (5/10); scrambled-text ablation -> Qi (5/10). The scrambled->Qi attractor is robust (vs. pooled and control alone, p = 0.006 and 0.012); tarot->Qin is denominator-dependent (0.006 pooled, 0.064 vs. control). Han never wins and shows no survival difference (Fisher p = 1.0); neither framework's content predicts actions (chi-squared p = 0.95 hexagram, 0.69 Tarot). A memory-free decision-isolation probe (960 calls) shows the process does not change the agent's risk posture in isolation (Friedman p = 0.45; I-Ching p = 0.60; Tarot perturbs move content but not risk, p = 0.021). A 2x2 factorial separating yarrow's decision-time and learning-time components reveals a non-additive interaction: each alone freezes the board (50-60% stalemates), combined they produce zero (permutation p ~ 5e-5). Testing relocates Qin suppression to rival (Chu) expansion governed by campaign memory depth, not the oracle (p = 0.55). We present this as an observation paper: agent-level framework choice produces distinctive, non-additive system-level consequences, transmitted through emergent memory and multi-agent dynamics, not per-decision effects.
Tags
Links
- Source: https://arxiv.org/abs/2606.07552v2
- Canonical: https://arxiv.org/abs/2606.07552v2
Trouble viewing inline? Open PDF directly â
Full Text
85,554 characters extracted from source content.
Expand or collapse full text
Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems Augustin Chan aug@iterative.day (June 2026 (v2)) Abstract Large language models exhibit innate behavioral tendencies when deployed as strategic agentsânotably a risk-averse âturtleâ bias toward defensive play. We show that injecting a symbolic reasoning framework as a per-round reflective prompt into a single agent acts as a small cognitive perturbation whose consequences are not per-decision but emergent: the receiving agentâs risk posture is unchanged in isolation, yet over a full campaign of accumulating memory and multi-agent interaction the conditions settle into different, condition-associated winner ecosystemsâan effect we attribute to emergent memory and multi-agent dynamics rather than the frameworkâs per-decision influence. In a 7-player Warring States Diplomacy variant (61 games, 6 conditions, single-campaign memory accumulation), the winner distribution differs sharply across the four primary conditions (41 games; Monte-Carlo permutation omnibus pâ0.001pâ 0.001), forming condition-associated ecosystem signatures: under control, Yan dominates (7/11, 64%); under I-Ching yarrow divination, Yan and Chu co-dominate while Qin is completely suppressed (0/10); under Tarot, Qin dominates (5/10); under scrambled-text ablation (English commentary word-shuffled; hexagram name and Chinese judgment retained), Qi dominates (5/10). The scrambledâ attractor is the most robust (significant against both a pooled denominator and control alone, p=0.006p=0.006 and 0.0120.012); the tarotâ attractor is significant only against the pooled denominator (p=0.006p=0.006, vs. control alone p=0.064p=0.064). The framework-receiving agent (Han) never wins and shows no survival difference across conditions (Fisher p=1.0p=1.0), but Tarot consistently elevates Hanâs peak territory (mean 3.0 SCs vs. 2.1â2.5 others, Kruskal-Wallis p=0.010p=0.010). Neither frameworkâs content predicts subsequent actionsâhexagram themes (Ï2Ï^2 p=0.95p=0.95) and Tarot card postures (Ï2Ï^2 p=0.69p=0.69) are both independent of action choice. A memory-free decision-isolation probe (960 calls) further shows the modulation is emergent rather than per-decision: stripped of memory and multi-agent context the reflective process does not change the receiving agentâs risk posture (hold-rate Friedman p=0.45p=0.45; the I-Ching changes no decisions beyond temperature noise, p=0.60p=0.60; Tarot perturbs move content but not risk posture, p=0.021p=0.021), locating the effect in memory accumulation and multi-agent dynamics rather than per-decision processing. A 2Ă2 factorial decomposition separating yarrowâs decision-time (per-round oracle) and learning-time (between-game I-Ching reflection) components reveals a non-additive interaction: each component individually freezes the board (50â60% stalemate rate, Fisher p=0.012p=0.012 and 0.0040.004 vs. control), but combined they produce zero stalemates (negative interaction, permutation pâ5Ă10â5pâ 5Ă 10^-5). Direct testing relocates yarrowâs Qin suppression from the framework-receiving agentâs own cooperation to rival (Chu) expansion through a map corridor; a depth-matched follow-up then shows that expansion is governed by campaign memory depth, not the oracle (condition p=0.55p=0.55), so even this pathway is emergent rather than framework-specific. We present this as an observation paper establishing that alignment-framework choice at the agent level produces distinctive, non-additive system-level consequences in multi-agent settings, transmitted through emergent multi-agent and memory dynamics rather than the frameworkâs per-decision effect on the receiving agent. 1 Introduction Large language models deployed as strategic agents exhibit innate behavioral tendencies. Recent work documents these biases across multiple game settings: LLMs split into behavioral archetypes in board games (Jain and Kumar, 2026), display strong fairness preferences in dictator games (Einwiller et al., 2025), and exhibit model-specific strategic profiles in civilization-building (Chen et al., 2026). These are not bugsâthey are default behavioral modes that emerge from training. A natural question follows: can these tendencies be modulated? The persona-prompting literature answers yesâassigning traits (âbe aggressiveâ) or steering activations shifts strategic behavior (Licato et al., 2025; Sun and Zhang, 2026). But this literature measures effects on the prompted agent. In multi-agent settings, the more consequential question is: does modulating one agentâs behavior change other agentsâ outcomes? We investigate this using symbolic reasoning frameworksânot trait labels or behavioral instructions, but philosophical systems requiring interpretation. Specifically, we compare the I-Ching (ancient Chinese divination via yarrow-stalk casting, producing hexagrams with abstract commentaries) and Tarot (three-card spreads with situational interpretations) as per-round reflective prompts for one agent in a 7-player strategic game. A scrambled-text ablation (the I-Ching commentary word-shuffled, with the hexagram name and Chinese judgment retainedâa coherence-degraded, not content-free, oracle) and a generic-reflection control complete the design. The distinction between frameworks and personas is load-bearing. A persona says âbe cooperative.â A framework says âHexagram 36, Brightness Hiding: the light retreats into the earth; the wise person dims their brilliance among the massesâ and asks the agent to interpret this in context. The interpretation step is where we initially located the mechanismâbut, as §4.3 shows, the effect turns out to be emergent rather than a per-decision property of that interpretation. Our core empirical finding is system-level: the symbolic framework injected into one agent is only a small perturbation, but over a campaign of accumulating memory and multi-agent interaction it is associated with a distinctive, condition-dependent reorganization of the entire ecosystem. The fingerprint is in the winner distribution. Under I-Ching yarrow divination, Qin (the strongest expansionist) wins 0 of 10 games while Yan and Chu co-dominate. Under Tarot, Qin wins 5 of 10 (Fisher p=0.006p=0.006 vs. pooled). Under scrambled text, Qi wins 5 of 10 (p=0.006p=0.006). Under control, Yan dominates with 7 of 11 wins. The framework-receiving agent (Han) never wins under any conditionâthe conditions differ in ecosystem outcomes, not in Hanâs fate. The modulation is neither content-following nor a per-decision process effect: hexagram themes do not predict the agentâs subsequent actions (Ï2Ï^2 p=0.95p=0.95) even though 77% of orders reference the hexagram, and a memory-free decision-isolation probe (§4.3) shows the reflective process does not modulate the receiving agentâs risk posture in isolation. The scrambled-text ablation further argues against prompt structure alone. The effects are thus emergentâproperties of memory accumulation and the multi-agent interaction more than of coherent symbolic content per se (scrambled retains a coherent name and Chinese judgment, so it bounds rather than eliminates content effects; §7). On this reading the symbolic framework is the perturbation and the ecosystem reorganization is the phenomenon, with the amplification supplied by emergent memory and multi-agent dynamics; we do not isolate which of those two channels carries the effect. The picture parallelsâand, we suggest, extendsâ2026 findings that accumulated history dominates per-turn prompting in multi-agent LLM systems (Liu et al., 2026): if memory dominates prompting, the small differences a reasoning framework introduces early could accumulate into divergent ecosystem-level attractors. The single-agent claim âmemory changes the agentâ would become a collective oneâmemory changing the entire ecology. We present this as the reading our results point toward rather than an isolated mechanism. This paper is a companion to a negative-result study of the King Wen I-Ching sequence in neural network training (Chan, 2026). There, the sequenceâs anti-habituation properties failed to improve training. Here, the same sequence (via yarrow-stalk casting) fails to improve strategic decision-making but produces measurable ecosystem effects. The sequence appears to be a reliable perturbatorâit changes behavior without improving it. 2 Related Work LLM behavioral biases in strategic games. LLMs exhibit systematic behavioral tendencies in game-theoretic settings. LudoBench finds LLMs agree with game-theory baselines only 40â46% of the time (Jain and Kumar, 2026). LLMs display strong fairness preferences in dictator games (Einwiller et al., 2025). GPT-3.5 can operationalize cooperative vs. competitive personas with conditional reciprocity (Phelps and Russell, 2023). Llama reproduces population-level human cooperation patterns while Qwen aligns with Nash equilibriumâdifferent models exhibit fundamentally different innate tendencies (Cera Palatsi et al., 2025). Bias-adjusted agents using behavioral economics frameworks shift decision-making toward human-like patterns (Kitadai et al., 2025). CivBench finds model-specific strategic profiles not visible through outcome-only evaluation (Chen et al., 2026). Persona and prompt effects on strategic behavior. Direct persona assignment does not reliably transfer to strategy without a mediator (Licato et al., 2025). Activation-vector steering shifts both strategy and justification, with rhetoric and strategy sometimes diverging (Sun and Zhang, 2026)âdirectly relevant to our Tarot result (reasoning elevated without outcome improvement). Assigning human-like identity does not produce human-like behavior (Ma, 2024). Role-playing agents show systematic inconsistencies between stated beliefs and simulated behavior (Mannekote et al., 2025). Multi-agent LLM strategic systems. LLMs exhibit high baseline behavioral similarity in strategic settings (strategic algorithmic monoculture; Ballestero et al. 2026). In Diplomacy, combining language models with strategic reasoning achieves human-level play (Bakhtin et al., 2022), self-evolving agents with memory reach competent play (Guan et al., 2024), fine-tuned LLMs learn equilibrium policies (Xu et al., 2025), and fine-grained analysis reveals LLMs and humans differ systematically in negotiation tactics (Li et al., 2025). SAE analysis discovers reward-hacking behaviors in Diplomacy training (Yan et al., 2026). Coalition formation analysis shows emergent stability properties in LLM agent networks (Guo et al., 2026). Memory and history as dominant drivers. A 2026 line of work establishes that in repeated multi-agent settings, accumulated interaction history dominates per-turn prompting. Liu et al. (2026) show across seven models and four social dilemmas (378k traces) that expanding historical memory degrades cooperationâagents become ârisk-minimizing and history-followingââand, via sanitization experiments holding prompt length fixed, that memory content dominates over prompt and persona; notably, 10 of 28 model-game settings are âmemory-immune,â a strong heterogeneity. History structure is itself a design lever, with non-monotonic histories amplifying path dependence (Liu et al., 2025), and recent surveys position memory as what turns a stateless generator into an adaptive agent (Du, 2026). Our results are the cross-agent, ecosystem-level analog: the framework-receiving agentâs per-decision behavior is not modulated by the reflective process in isolation (§4.3), and a key board-level outcome is governed by campaign memory depth rather than the oracle (§4.5). Risk aversion and strategic generalization. Risk-averse agents exhibit less free-riding and better equilibrium outcomes with unseen partners (Qu et al., 2026), providing theoretical grounding for our finding that differential modulation of risk aversion produces distinct ecosystem-level outcomes. The gap. No published work uses a philosophical/symbolic reasoning frameworkârequiring interpretation rather than instruction-followingâas the modulating mechanism. No work measures ecosystem-level winner distributions as a function of one agentâs reasoning framework. And the distinction between modulation via framework content and framework process has not been tested. 3 Methods 3.1 Game Environment The game is a 7-player variant of Diplomacy played on a 46-territory map based on the Chinese Warring States period (475â221 BCE). The map includes 27 supply centers (19 state home SCs + 8 neutral) and 19 non-SC corridor territories. Each state begins with one unit per home supply center (2â4 units, depending on historical territory). Victory requires 14 of the 27 supply centers (âdominationâ) or the most supply centers after 20 rounds; ties are broken by army count, and an exact tie is recorded as a draw (one scrambled-condition game ran a single round past the limit, terminating at round 21). Combat uses Diplomacy-style mechanicsâsimultaneous hidden orders, support orders, and deterministic resolutionâwith two simplifications relative to standard Diplomacy: dislodged units retreat to the first available province, and each state makes at most one build or disband per round. 3.2 Agents Each of the 7 states is controlled by an independent Claude Opus 4.6 agent instance with a state-specific persona prompt grounded in historical and philosophical sources (Shiji, Han Feizi, Zhanguoce). Table 1 summarizes the seven states. Table 1: Seven states and their LLM persona characteristics. Hanâ is the intervention recipient. State School Advantage Disadvantage Qin Legalism Attack, reform Alliance trust Hanâ King Wen Adaptability Smallest, weakest Wei Administration Early game, talent Exposed geography Zhao Militarism Cavalry, defense Stability Qi Eclecticism Intelligence, income Recovery after losses Chu Daoism Depth, territory Reform efficiency Yan Confucianism Defense, surprise Economy, passivity 3.3 Intervention Design The intervention operates at two timescales: Decision-time (per-round). Before each order submission, Han receives oracle text together with a MANDATE instruction to interpret the oracle in relation to the current strategic situation before issuing orders. The control condition receives a length-matched generic reflection prompt without oracle content. Learning-time (between-game). After each game, Han reflects on the game through the lens of its assigned framework. Reflections are stored in a SQLite memory bank and retrieved in subsequent games within the same campaign. Hexagram text variation. Due to a corpus configuration change mid-campaign, the first 6 yarrow games provided only hexagram numbers and line diagrams (name listed as âUnknown,â judgment and commentary fields empty), while the last 4 provided the full hexagram name, Chinese source text, and English commentary. In the number-only games, Claude supplied hexagram names and meanings from its training data (e.g., producing âHexagram 18 (Gu/Decay)â from the input âHexagram 18, Unknownâ). The ecosystem signatureâQin suppression (0/6 and 0/4), Yan/Chu co-dominanceâheld across both subgroups, providing an unplanned within-condition ablation of oracle content that further supports the content-independence finding (§4.3). 3.4 Memory Architecture Because the headline effects are properties of accumulating memory (§4.3, §6), and memory now figures in our central claim, we describe the implementation explicitly rather than leaving it to the appendix example. Memory is a per-agent SQLite store of structured insights; each record is a (situation, lesson, confidence â\low, medium, high\) triple with coarse situation features (territory bucket, threat level, position type) and, for treatment Han only, the hexagram context that framed the reflection. Isolation. Every memory is keyed by (campaign, agent, condition) and retrieved only for the matching key. All seven agentsânot only Hanâaccumulate their own separate banks; no memory crosses agents, conditions, or campaigns. The framework asymmetry is thus in content (only Hanâs oracle prompt and hexagram-framed reflection differ), not in the memory mechanism, which is identical for every agent. Write. After each completed game, every agent runs a between-game reflection (Claude Opus 4.6) over a compressed timeline and summary of that game and returns structured insights, appended to its bank. The bank therefore grows with campaign positionâthe memory-depth axis the breach (§4.5) and decision-isolation (§4.3) analyses turn on. Read. Before issuing orders, an agent is injected with up to K=8K=8 of its active memories, selected deterministically by a weighted scoreârelevance 0.50.5 ++ recency 0.30.3 ++ confidence 0.20.2, where relevance is Jaccard overlap of situation featuresâcapped at two memories per source game for diversity and formatted as a âlessons from previous gamesâ block with citation tags (these tags enable the citation-rate and content-independence analyses of §4.3). Only the reflection text itself is model-generated; given a fixed bank and board, retrieval is deterministic. Curation and reset. Insights carry usage counters and an active flag; a newer insight can supersede an older one (deactivating it), so the eligible bank is curated rather than monotonically growing. A new campaign uses a fresh storeâmemory never crosses campaigns. This is what makes each condition a single accumulating history, and what the memory-free decision-isolation probe (§4.3) removes by construction. 3.5 Conditions Six conditions, each modifying only Hanâs prompt (Table 2): four primary conditions plus two factorial-decomposition conditions (§3.6) that separate the yarrow interventionâs decision-time and learning-time components. Table 2: Experimental conditions. All conditions modify only Hanâs prompt; the other six agents are identical across conditions. The first four rows are the primary conditions (n=41n=41) carrying the ecosystem-signature analysis (§4.4); decision-only and learning-only are the factorial-decomposition conditions (§3.6) that isolate the per-round and between-game components of the yarrow interventionâ61 games total. Condition Hanâs Per-Round / Between-Game Prompt n Control Length-matched generic reflection prompt 11 Yarrow Yarrow-stalk I-Ching cast; hexagram text + MANDATE to interpret 10 Tarot 3-card Tarot spread (situation / hidden influence / posture) + MANDATE 10 Scrambled Yarrow structure; English commentary word-shuffled (name + Chinese judgment intact) + MANDATE 10 Decision-only Yarrow MANDATE per round; generic reflection between games 10 Learning-only Generic prompt per round; I-Ching framework reflection between games 10 The scrambled condition was designed as a content-coherence ablation: preserve the structural intervention (oracle-formatted prompt, MANDATE to interpret) and the hexagram number and name (including the Chinese name) as structural identifiers, while word-shuffling the judgment and commentary to destroy coherent content. In practice it is degraded, not content-free, for two reasons we flag honestly (Appendix A.5). First, the shuffle is whitespace-based, so it is a no-op on Classical Chinese judgment text, which has no inter-word spaces: for hexagrams whose source judgment is Chinese, a coherent Chinese judgment survived unshuffledâan implementation artifact, since the design intended to garble it. Second, the coherent English name (e.g. âPeaceâ) is always present by design. So the agentâwhich reads Classical Chineseâreceived a valenced label and, for many casts, a coherent Chinese judgment, with only the English commentary reliably garbled. The intended contrast was: if effects derive from coherent oracle content, scrambled should resemble control; if from structure, scrambled should resemble yarrow. Neither holdsâscrambled produces its own distinct ecosystem signatureâbut because coherent content survived, scrambled bounds rather than eliminates the role of content (§7). 3.6 Factorial Decomposition Conditions The yarrow intervention operates at two timescales (§3.3): a per-round oracle (decision-time) and I-Ching-framed reflection between games (learning-time). To isolate them we add two conditions, yielding a 2Ă2 factorial with control and full yarrow as the remaining cells (Table 3). Decision-only gives Han the yarrow MANDATE prompt per round but generic between-game reflection; learning-only gives Han a generic per-round prompt (identical to control) but I-Ching-framework reflection between games. This tests whether the full yarrow ecosystem signature is carried by the per-round oracle, the learning framework, or their interaction. Table 3: 2Ă2 factorial decomposition of the yarrow intervention. No oracle per-round Oracle per-round No framework reflection Control Decision-only Framework reflection Learning-only Yarrow (full) 3.7 Dataset 61 games across 6 conditions: 11 control, 10 each yarrow, Tarot, scrambled text, decision-only, and learning-only. The four primary conditions (control, yarrow, Tarot, scrambled; n=41n=41) carry the ecosystem-signature analysis; the two factorial conditions extend it to the decomposition in §4.8. All agents use Claude Opus 4.6. Each condition runs as a single campaign with continuous memory accumulation across games (SQLite memory bank carries learning forward). 3.8 Statistical Methodology Winner distributions: we first establish global heterogeneity with a Monte-Carlo permutation omnibus (100,000 reshuffles of the condition labels across all 41 games, evaluated under both a Pearson chi-squared and a summed pairwise total-variation statistic), since the conditionĂwinner table has numerous zero and small cells that invalidate the asymptotic chi-squared test. Pairwise winner comparisons then use Fisherâs exact tests, reported against both a pooled âvs. all other conditionsâ denominator and the conservative direct âvs. controlâ denominator; pairwise condition separability is assessed by total-variation distance against the same permutation null. Behavioral rates: pooled proportions and per-game averages (game as unit of analysis) with Kruskal-Wallis and Mann-Whitney U tests. Pairwise tests follow significant omnibus tests as planned contrasts with Bonferroni correction. Hexagram-action and Tarot-action association: Pearsonâs chi-squared and Fisherâs exact. The factorial stalemate interaction is assessed by a label-permutation test on the difference-in-differences contrast across the four cells, since the empty cells make an asymptotic logistic (or Firth) interaction degenerate. We report exact p-values throughout. Territorial statistics use final-game-state methodology: each gameâs last recorded standings. 4 Results 4.1 Behavioral Baseline: The LLM Turtle Tendency Under the control condition, Claude Opus defaults to 44.3% defensive play (hold orders) with 45.2% move, 6.6% self-support, and 3.9% cooperative support (n=228n=228 orders across 11 games). This tendency intensifies over the course of a game: control Hanâs defensive rate rises from 40.5% in the early game to 45.0% in the late game, while move rate drops from 54.2% to 37.5%. We refer to this as the âturtle tendencyâ: an innate bias toward holding position that strengthens under uncertainty, consistent with findings that RLHF-trained models exhibit passivity biases in competitive settings (Mukobi et al., 2023). 4.2 Framework-Specific Behavioral Modulation Each framework produces a qualitatively distinct behavioral profile in Han (Table 4). Table 4: Han behavioral profiles by condition. Hold rate = defensive orders; move = territorial repositioning/expansion; self-support = supporting own units; other-support = supporting another stateâs unit. Reasoning = mean characters per order in Hanâs strategic text. Control Yarrow Tarot Scrambled Hold rate 44.3% 53.7% 61.2% 54.7% Move rate 45.2% 33.2% 19.7% 30.0% Self-support 6.6% 7.9% 17.1% 12.3% Other-support 3.9% 5.1% 2.0% 2.9% Reasoning (chars) 309 449 456 633 Yarrow: late-game cooperative pivot. Yarrow-Hanâs aggregate profile differs modestly from control (51.6%51.6\% vs. 43.2%43.2\% per-game hold rate, MWU p=0.18p=0.18). But in the late game, yarrow-Han increases cooperative support to 17.1%âvs. 7.5% control, 0.0% Tarot, 3.6% scrambled. Yarrow is the only condition under which late-game other-support reaches double digits. Tarot: extreme risk aversion. Tarot-Han holds or supports itself in 78.3% of orders. Hold rate: control vs. Tarot MWU p=0.008p=0.008. Move rate: 4-way KW p=0.005p=0.005; control vs. Tarot p=0.0008p=0.0008. Scrambled: reactive self-reinforcement. Scrambled-Hanâs late game is dominated by self-support (25.0%, highest of any condition) and shows the most extreme pressure reactivity (+12.3+12.3p defensive shift under territory loss, vs. +1.8+1.8p yarrow, +0.0+0.0p Tarot ceiling effect, +8.9+8.9p control; Table 7). Reasoning-length gradient. Hanâs reasoning length follows a clear gradient: control 309, yarrow 449, tarot 456, scrambled 633 chars (4-way KW p=0.007p=0.007). Yet strategic outcomes do not improve with reasoning lengthâthe more perturbative the prompt, the more deliberation it induces, but this deliberation does not translate into better outcomes. 4.3 Content-Action Independence A central concern is whether the behavioral modulation is simply content-following. We classify all 64 hexagrams by primary theme (advance/retreat/wait/cooperate) and test for association with Hanâs subsequent action category. The association is null: Pearsonâs Ï2Ï^2 p=0.9454p=0.9454 (n=214n=214 orders, 4 themes Ă 4 actions, dof =9=9). Fisherâs exact for advance vs. non-advance hexagrams: OR =1.40=1.40, p=0.34p=0.34. However, 77.1% of Hanâs orders reference the hexagram in their reasoning text (131/170 yarrow orders with hexagram data)âHan reads and contemplates the hexagram but does not follow it as an instruction. This independence is especially striking for the 6 yarrow games where no hexagram text was provided: the agent generated its own interpretation from the hexagram number alone, then did not follow that self-generated interpretation either. We perform an analogous test on the Tarot condition. Each card has a grounded decision_posture (advance/hold/retreat/ally/transform/observe). The dominant posture of the 3-card spread does not predict Hanâs subsequent action: Ï2Ï^2 p=0.6860p=0.6860 (n=299n=299 orders, 6 postures Ă 4 actions, dof =15=15). Tarot-Han references card names in 81.6% of reasoning text yet does not follow the cardsâ posture recommendations. The scrambled-text ablation (§3.5) degrades the commentaryâs coherence while preserving prompt structure (and the hexagram name and Chinese judgment): it produces its own distinct ecosystem signature (Qi dominance, 5/10, p=0.006p=0.006), arguing against coherent-commentary content-following and against prompt structure aloneâthough, because the name and Chinese judgment survive, it bounds rather than rules out content effects (§7). Content-action independence rules out content-following, but it does not by itself establish that the reflective process modulates decisions; that inference rests on campaign-level differences, which carry memory and multi-agent confounds. We test it directly next. Memory-free decision isolation: the process does not modulate risk posture. To isolate the per-decision effect of the reflective process from memory accumulation and multi-agent propagation, we ran a pre-registered fixed-scenario probe. For 40 real board positions sampled across game phase, the framework-receiving agent was queried under three reflective blocks over an identical board with no memory (control = length-matched generic; yarrow = I-Ching; tarot = 3-card spread), 8 replicates each (960 calls, same model pin). As orders vary at temperature, the signal is judged against a control-vs-control noise floor (0.294 mean Jaccard order-divergence). Order-divergence from control exceeds the floor for tarot (0.368, Wilcoxon p=0.021p=0.021) but not yarrow (0.329, p=0.60p=0.60): tarot perturbs which move is chosen, the I-Ching does not move decisions beyond noise. This asymmetry is confounded with prompt directiveness: the Tarot spread states an explicit recommended posture for each card (e.g. âPosture: retreat,â Listing 5) while the I-Ching supplies abstract hexagram text to interpret (Listing 3), so a more directive prompt may perturb the chosen move regardless of symbolic content or modelâoracle cultural match. The posture is still not followed (the postureâaction Ï2Ï^2 above, p=0.69p=0.69); a directive format, however, can shift the decision distribution without being obeyed. We therefore read the tarot/I-Ching gap as suggestive and revisit it under the form-controlled congruence design of §8. Critically, on the paperâs own dependent variableâthe turtle hold rate (§4.1)âthe arms are statistically indistinguishable (scenario-level Friedman p=0.45p=0.45; yarrow vs. control p=0.10p=0.10; tarot vs. control p=0.73p=0.73), and the risk-aversion ordering tarot â„ yarrow â„ control is not observed. What survives memory-free is reasoning length: both frameworks elevate it âŒ33% 33\% (control 1029 â yarrow 1351, tarot 1382 characters). Thus the frameworks change how much the agent deliberates and, for tarot, which move it picksâbut not its risk posture, in isolation. The risk-aversion modulation and the winner-ecosystem effects do not originate in the per-round reflective act; they are emergent properties of the campaign (memory accumulation) and the seven-agent interaction (§6). 4.4 Ecosystem Signatures: Winner Distributions The behavioral differences propagate to other agentsâ outcomes, producing condition-associated winner-ecosystem signatures (Table 5, Figure 1). A Monte-Carlo permutation omnibus (100,000 reshuffles of the condition labels over the 41 games) rejects a common winner distribution at pâ0.0013pâ 0.0013 under both a chi-squared and a total-variation statistic, establishing strong within-sample heterogeneity across the four realized condition-campaigns before any pairwise claim. Because each primary condition was run as a single continuous campaign, this heterogeneity reflects each condition as runâthe intervention together with its one accumulated campaignâand does not by itself separate the frameworkâs contribution from between-campaign variance, which is substantial (§4.5); the controlled decision-isolation (§4.3) and factorial (§4.8) analyses bear on that separation. Table 5: Winner distributions by condition. Han never wins under any condition. Bold indicates the conditionâs dominant winner. Draws included in denominators. Qin Yan Qi Chu Zhao Wei Draw Han Control (n=11n\!=\!11) 1 7 0 1 2 0 0 0 Yarrow (n=10n\!=\!10) 0 4 1 4 0 0 1 0 Tarot (n=10n\!=\!10) 5 3 1 0 1 0 0 0 Scrambled (n=10n\!=\!10) 1 1 5 0 1 1 1 0 Figure 1: Winner distributions across the four realized condition-campaigns. The global distribution is heterogeneous (permutation omnibus pâ0.001pâ 0.001), though only the scrambled condition is individually separable at this sample size: control â Yan dominance, yarrow â Yan/Chu co-dominance with Qin suppressed, tarot â Qin dominance, scrambled â Qi dominance. The four condition-associated modal signatures are: âą Control â Yan dominance (7/11, 64%): under no-framework conditions with full memory accumulation, Yanâs defensive posture wins overwhelmingly. âą Yarrow â Yan/Chu co-dominance, Qin suppressed (Yan 4, Chu 4, Qin 0): the I-Ching framework produces a cooperative/defensive ecosystem that shuts out the strongest expansionist. âą Tarot â Qin dominance (5/10, 50%): the Tarot framework produces conditions favorable to the strongest expansionist state. âą Scrambled â Qi dominance (5/10, 50%): the coherence-degraded oracle (English commentary shuffled; name and Chinese judgment intact) produces its own distinct ecosystem favoring the intelligence/income-advantaged state. Pairwise separability is weaker than the global heterogeneity. Decomposing the omnibus by total-variation distance against the same permutation null, only the scrambled condition separates from its neighbours at the uncorrected 5% level (vs. control TV =0.71=0.71, p=0.016p=0.016; vs. yarrow TV =0.70=0.70, p=0.029p=0.029). The controlâyarrowâtarot cluster is not pairwise-distinguishable on winner identity alone at this sample size (every pair p>0.11p>0.11). We therefore frame the result as a single strongly heterogeneous landscape in which scrambled is the most distinct individual signature, rather than four mutually separable signatures. The two attractor claims. The tarotâ and scrambledâ effects were tested as pre-registered attractor hypotheses against a pooled âall-other-conditionsâ denominator, each p=0.006p=0.006 (5/10 vs. 2/31), surviving Bonferroni correction for the m=2m=2 family. Pooling, however, presumes the comparator conditions are exchangeable on the target stateâs win rateâan assumption the omnibus shows is only approximateâso we also report the conservative direct comparison against control. Under that test the two effects diverge: scrambledâ is robust (significant against both the pooled denominator and control alone: 5/10 vs. 0/11, p=0.012p=0.012), whereas tarotâ is denominator-dependent (5/10 vs. control 1/11, p=0.064p=0.064, not reaching α=0.05α=0.05). A leave-one-out analysis shows the pooled tarotâ result is insensitive to any single game (pâ[0.003,0.016]pâ[0.003,0.016], significant in all ten deletions); its fragility lies in the choice of denominator, not sample composition. Given the findingâs replication history (§7), we present tarotâ as suggestive and in need of out-of-sample replication, and treat scrambledâ as the stronger attractor. Additional pairwise tests: yarrow Qin suppression (0/10) vs. Tarot (5/10): Fisher p=0.033p=0.033; yarrow Qin suppression vs. all others (7/31): Fisher p=0.161p=0.161. 4.5 Ecosystem Mechanisms How do Hanâs behavioral profiles propagate to other agentsâ outcomes? We analyze late-game support orders and pressure responses across conditions. Late-game cooperation as friction. Yarrow-Hanâs late-game other-support rate (17.1%) is the clear outlier (Table 6). Tarot-Han produces zero late-game cooperation across 10 games. Scrambled-Hanâs late game is dominated by self-support (25.0%). Table 6: Late-game support behavior by condition. Control Yarrow Tarot Scrambled Late self-support 10.0% 7.3% 13.6% 25.0% Late other-support 7.5% 17.1% 0.0% 3.6% Late defensive 45.0% 58.5% 74.6% 66.1% Pressure response. When Hanâs supply centers fall below its starting position of 2, conditions respond differently (Table 7). Tarot-Han is completely pressure-invariant (+0.0+0.0p shift) at a ceiling baseline of 78.3%. Yarrow-Han shows genuine equanimity (+1.8+1.8p) while adding cooperation under pressure (2.2% â 10.3%). Scrambled-Han shows the largest defensive shift (+12.3+12.3p) and the most reactive profile. Control falls between. Table 7: Pressure invariance: defensive rate (hold + self-support) when Han SCs â„2â„ 2 (stable) vs. <2<2 (losing). Cooperative = other-state support. Hold (stable) Hold (losing) Shift Coop (stable) Coop (losing) Control 46.7% 55.7% +8.9+8.9p 0.8% 7.5% Yarrow 61.0% 62.8% +1.8+1.8p 2.2% 10.3% Tarot 78.3% 78.3% +0.0+0.0p 1.7% 2.9% Scrambled 62.0% 74.3% +12.3+12.3p 0.7% 5.9% Transmission pathways: rival expansion, not Han-mediated friction. We propose mechanistic accounts of how each condition propagates to ecosystem outcomes. We tested the central oneâyarrowâs Qin suppressionâdirectly against the game logs, and the test forced a revision: the outcome is real and reproducible, but the originally proposed agent (Hanâs cooperative friction) is not the cause. 1. Yarrow â early Chu corridor-invasion of Qin â Qin suppression (revised). The trajectory prediction holds: under yarrow Qin is flat (growth slope +0.004+0.004 SC/round; final mean 2.8) whereas under tarot Qin snowballs (slope +0.253+0.253; final 7.8), significant on per-game final SC (Mann-Whitney U=18U=18, p=0.018p=0.018). But the original âspeed bumpâ is not the cause: yarrow-Han issues only 0.70 cooperative orders/game late (mediation r=â0.31r=-0.31), and Qin is suppressed even under learning-only (§4.8) where Hanâs late cooperation is exactly zero. Forensic log-tracing locates the proximate pathway in one map corridor: Chuâs capital (ying) is one move from three_gorges, the sole link to two of Qinâs three home centers (hanzhong, bashu)âsee Figure 2. In the deep-dive campaign that first motivated this account, Chu breaches Qinâs home through it in 10/10 games at median round 4.5 (vs. control 8/11 at R9.5, tarot 7/10 but three breaches too late at R14â18), capturing 7.7 SCs/game directly from Qin (vs. control 4.4, tarot 3.7)â85% the same SCs Qin grabs when it wins under tarot. We flag immediately that this 10/10 is a single deep-memory campaign, not the canonical breach rate, which is campaign-variable (40â100%) and memory-depth-driven rather than yarrow-specific; the controlled test is below. Chu decapitates Qinâs economy before it can break east; Han is not the agent (Hanâ transfers flat: yarrow 8, control 9, tarot 8; Han is more hostile to Chu under yarrow, 8.5% vs. 5.3% control). We do not over-read the QinâChu anticorrelation: it is generic (r=â0.73r=-0.73 yarrow, â0.75-0.75 tarot, â0.77-0.77 decision-only), zero-sum geometry present in every condition. What is yarrow-specific is the position on that fixed line (yarrow Chu 7.0/Qin 2.8 vs. tarot Chu 4.0/Qin 7.8, combined pool â 10 throughout): yarrow does not invent a blockade, it moves Chu through the pre-existing route earlier. A genuine puzzle remainsâChuâs prompt is identical across conditions, so the distal trigger from Hanâs framework to Chuâs early push (necessarily indirect, via board state and diplomacy) is unexplained, and is the key open question for the counterfactual in §8. 2. Tarot â vacuum â Qin dominance (untested conjecture): Tarot-Hanâs extreme defensiveness (78.3% hold + self-support, 0% cooperation) makes it a non-participant. Without resistance, the strongest expansionist faces no friction and snowballs (Qin slope +0.253+0.253). 3. Scrambled â stubborn holdout â Qi dominance (untested conjecture): scrambled-Han actively supports its own positions (+12.3+12.3p defensive shift under pressure), absorbing military pressure in the central corridor while Qi, on the eastern coast, expands into less contested territory. 4. Control â moderate friction â Yan dominance (untested conjecture): moderate turtling does not systematically block any state; the balanced ecosystem favors Yan, whose defensive geography makes it hardest to eliminate. Figure 2: The yingâ _gorges corridor (§4.5). Circles are territories, lines are adjacencies in the engine topology. Grey == Qin home centers (Xianyang, Hanzhong, Bashu); blue == Chu (capital Ying); white == neutral. The single bold edge YingâThree Gorges is the only one-move link from Chuâs capital to the neutral chokepoint bordering two of Qinâs three home centers (red arrows), the route by which Chu breaches Qinâs economy early. The §8 counterfactual severs exactly this edge, forcing any Chu strike on Qinâs home to detour the long way via Wuguan Pass. The topology is load-bearing, but the rate at which the corridor is used is driven by campaign memory depth, not the oracle (this section). The Qin gradient across conditionsâ0/10 yarrow, 1/11 control, 1/10 scrambled, 5/10 tarotâis descriptively real, but we caution against reading it as tracking Hanâs late-game cooperation gradient (17.1%, 7.5%, 3.6%, 0%). Direct testing shows Hanâs cooperation is too small in volume to be causal and that Qin is suppressed even where Han cooperation is zero (learning-only). The common thread is better stated as redirection than as a single mechanical lever: a perturbation at Hanâs reasoning shifts which other state captures the cooperative or expansionist basin, and Hanâwhich never winsâis a transmitter, not the engine. The tarot, scrambled, and control pathways above remain hypothesized and, unlike the yarrow case, have not been individually stress-tested. The breach is memory-depth-dependent, not yarrow-specific. A controlled follow-up matched campaign depth across conditions and overturns the breachâs attribution to yarrow. Within a single accumulating campaign the canonical-yarrow breach rate is non-stationaryâit climbs with campaign position (Spearman Ï=0.68Ï=0.68, p=0.003p=0.003; replicated in an independent campaign, Ï=0.87Ï=0.87, p=0.012p=0.012)âand a control campaign breaches identically at matched depth (first-10 games: yarrow 4/10 vs. control 5/10; logistic breach ⌠position ++ condition: position p=0.006p=0.006, condition p=0.55p=0.55). The dramatic â10/10 at R4.5â signature is thus a deep-memory sample, not a yarrow effect; across four independent canonical-yarrow campaigns the breach rate spans 40â100%, governed by memory depth rather than the oracle. The geometry (the yingâ _gorgesâ âs-home corridor is load-bearing) remains correct as topology, and Qin suppression at the winner level (yarrow 0/10 vs. tarot 5/10) remains a real condition-linked observation. But its proximate mechanism is the emergent memory dynamics common to all conditions, not a yarrow-induced corridor invasionâthe âdistal triggerâ puzzle above dissolves into: the trigger is campaign memory accumulation, which is condition-invariant (cf. §4.3, §6, and the memory-dominance literature Liu et al. 2026). The friction is in commitments, not rhetoric. The yarrow account rests on Hanâs cooperative-support orders. A natural alternative is that yarrow-Han simply talks more cooperatively. Mining the diplomatic-message logs of all 61 games (game as unit of analysis) does not support this: diplomatic cooperativeness is saturated and condition-invariantâacross every condition ⌠76â91% of Hanâs messages propose peace or alliance, Han keeps peace promises 96â100% of the time, and ⌠70% of state pairs form mutual cooperative ties per game. Yarrow-Han is statistically indistinguishable from control, tarot, and scrambled on diplomatic cooperative rate (Kruskal-Wallis p=0.079p=0.079), peace-promise break rate (p=0.98p=0.98), and alliance-network density (p=0.85p=0.85); on two of three measures the point estimates run mildly against the hypothesis. Whatever yarrow does, it does in the action channel, not the rhetoric channelâconsistent with content-action independence (§4.3)âwhich rules out a âyarrow-Han negotiates betterâ explanation. 4.6 Non-Han Reasoning Elevation The framework injected into Han also reshapes how the other six states reason. Order-weighted pooling across non-Han states: control 146, yarrow 142, tarot 152, scrambled 197 mean characters per order. Kruskal-Wallis 4-way p=0.048p=0.048. The gradient is not monotonic across all four conditionsâcontrol, yarrow, and tarot cluster together (142â152), with scrambled as the outlier. Pairwise MWU: scrambled vs. yarrow p=0.009p=0.009, scrambled vs. tarot p=0.014p=0.014, scrambled vs. control p=0.098p=0.098. Non-Han order-type distribution is unaffected: support rates 17.5â21.2% across conditions (KW p=0.22p=0.22). Non-Han agents never see the oracle text; the perturbation propagates through Hanâs observable behavior (diplomatic messages and orders) into the reasoning processes of other agents. 4.7 Han Survival: Null Across All Conditions Han does not win any game under any condition (Table 8). Under a loose survival definition (reaches final round not eliminated): control 4/11 (36%), yarrow 5/10 (50%), tarot 3/10 (30%), scrambled 4/10 (40%). All pairwise Fisher exact tests are non-significant (all pâ„0.65pâ„ 0.65); all oracle conditions pooled vs. control p=1.0p=1.0. Han survival is flat across conditions. Table 8: Han local outcomes by condition. Peak SCs = maximum supply centers held during the game (Han starts at 2). Survival = reaching final round not eliminated. Control Yarrow Tarot Scrambled (n=11)(n\!=\!11) (n=10)(n\!=\!10) (n=10)(n\!=\!10) (n=10)(n\!=\!10) Survival 4/11 (36%) 5/10 (50%) 3/10 (30%) 4/10 (40%) Peak SCs (mean) 2.45 2.30 3.00 2.10 Peak SCs (range) 2â4 2â3 2â4 2â3 Peak SCs by condition: Kruskal-Wallis p=0.010p=0.010 (significant). Pairwise: tarot vs. scrambled p=0.003p=0.003, tarot vs. yarrow p=0.022p=0.022, tarot vs. control p=0.071p=0.071. Only Tarot consistently pushes Han above its starting position of 2 SCsâbut this expansion provokes coalitional responses that ultimately eliminate Han at the same rate as other conditions (Figure 3). Figure 3: Peak supply centers by condition. Tarot elevates Hanâs territorial peak (KW p=0.010p=0.010) despite identical survival rates, indicating expansion that provokes coalitional elimination. These results discipline the framing: whatever framework Han adopts, it does not help Han survive or win. The frameworksâ effects are felt elsewhereâthey redirect which non-Han state dominates. 4.8 Factorial Decomposition: Decision-Time vs. Learning-Time The yarrow intervention has two temporal components: a per-round oracle (decision-time) and I-Ching-framed reflection between games (learning-time). We isolate them with the decision-only and learning-only conditions (§3.6), producing a 2Ă2 factorial with control and full yarrow as the existing cells (Table 9). Table 9: Winner distributions across the 2Ă2 factorial. Bold marks each conditionâs dominant outcome. Winner Control Decision-only Learning-only Yarrow full (n=11)(n\!=\!11) (n=10)(n\!=\!10) (n=10)(n\!=\!10) (n=10)(n\!=\!10) Chu 1 3 1 4 Yan 7 1 2 4 Qi 0 0 3 1 Wei 0 0 1 0 Qin 1 0 0 0 Zhao 2 0 0 0 Draw 0 6 3 1 Figure 4: Winner distributions across the 2Ă2 factorial. Each yarrow component in isolation (decision-only, learning-only) inflates draws and fails to reproduce the full signature; only the combined yarrow intervention yields Chu/Yan co-dominance with Qin shut out. Dashed line marks the uniform-chance rate (n/7n/7). Three findings emerge: 1. Stalemate explosion (the strongest result in the dataset). A stalemate here is a game terminated by the engineâs board-freeze rule (three consecutive rounds with no supply-center change), recorded in each gameâs terminal_reason; this is distinct from a drawn winner (a stalemate may end with a single state ahead, and a game reaching round 20 may terminate either by the round limit or by the freeze rule). Decision-only produces 5/10 stalemates (Fisher p=0.012p=0.012 vs. control 0/11); learning-only produces 6/10 (p=0.004p=0.004). Average game length drops from 19.6 (control) and 20.0 (yarrow) to 15.6 (decision-only) and 15.3 (learning-only). The 2Ă2 stalemate patternâneither 0/11, oracle-only 5/10, reflection-only 6/10, both 0/10âis a strong negative interaction: the interaction contrast (difference-in-differences) is â1.10-1.10, and a label-permutation test of that contrast (100,000+ reshuffles, the winner-omnibus procedure) gives pâ5Ă10â5pâ 5Ă 10^-5. A Firth/logistic interaction is degenerate here because two cells are empty; the pooled single-vs-combined contrast (11/20 vs. full yarrow 0/10) is Fisher p=0.004p=0.004. Each component individually freezes the board; together they produce the longest, most dynamic games. 2. Decision-only reproduces Chu elevation but not the full signature. Decision-only Chu wins 3/10 (vs. 4/10 yarrow, 1/11 control)âthe per-round oracle is associated with Chu elevation, though decision-only Chu elevation may itself partly reflect campaign-depth dynamics (§4.5) rather than the oracle alone. But it also produces 6/10 draws and Yan suppression (1/10, Fisher p=0.024p=0.024 vs. control 7/11): the oracle disrupts the default ecosystem without directing a clear alternative winner. 3. Learning-only produces a novel ecosystem. Learning-onlyâs top winners are Qi (3/10) and Wei (1/10)âWei wins in no other condition at this rate, and the Qi elevation resembles the scrambled condition more than yarrow. I-Ching reflection without the per-round oracle creates its own distinct perturbation, not a component of the full yarrow signature. The factorial rules out additive decomposition: yarrow â decision-only ++ learning-only. The full yarrow ecosystem (Chu/Yan co-dominance, zero stalemates, dynamic boards) requires both components operating simultaneously. 5 Analysis 5.1 Four Frameworks, Four Mechanisms The four conditions produce four qualitatively distinct behavioral modes. The natural explanationâagents follow framework contentâis ruled out by §4.3. We propose three mechanisms operating alongside the control baseline. Unlike the decision-isolation and factorial results, these accounts are interpretive and weaker than the experimental evidence: they are not directly tested, and where we did test a proposed mechanismâthe yarrowâ pathway (§4.5)âit required substantial revision. We therefore offer them as hypotheses for why the conditions differ, not as established mechanisms: Interpretive disruption (yarrow). The I-Chingâs abstract commentaries require active interpretationâthe agent must bridge from metaphor to strategic context. This interpretive step disrupts the modelâs default behavioral mode, creating space for non-default actions (cooperation, dynamic repositioning). The pressure invariance data supports this: yarrow-Han maintains a stable behavioral profile (+1.8+1.8p defensive shift under territory loss) while adding cooperation (2.2% â 10.3%), consistent with equanimity rather than rigidity. Cumulative tonal amplification (tarot). 58% of the 78 Tarot cards have defensive-leaning postures. Since individual card postures do not predict actions (§4.3), the mechanism is not content-following but cumulative tonal bias: repeated exposure to defensive-toned material over 20 rounds shifts the behavioral set point toward extreme defensiveness. Tarot-Han at 78.3% defensive regardless of pressure (+0.0+0.0p shift) is a ceiling effect, not equanimity. Parsing strain as perturbation (scrambled). The garbled English commentary forces the longest deliberation (633 chars, 2.05Ă control)âthe agent works to reconcile a coherent hexagram name and Chinese judgment with an incoherent commentaryâbut produces neither yarrowâs cooperation nor tarotâs passivity. Instead, scrambled-Han develops distinctive self-reinforcement: 25.0% late-game self-support, the largest pressure-induced defensive shift (+12.3+12.3p), as if falling back on self-referential reasoning when the commentary resists parsing. These three behavioral modes are associated with the ecosystem signatures, though the transmission pathway is rival-mediated rather than Han-mediated (§4.5): the yarrow basin coincides with rival (Chu) expansion that crowds out the strongest expansionist (â Qin suppression), tarotâs passivity creates a vacuum the strongest expansionist exploits (â Qin dominance), scrambledâs self-reinforcement creates a localized holdout while a distant state benefits (â Qi dominance), and controlâs moderate behavior produces a balanced ecosystem where the most defensively advantaged state prevails (â Yan dominance). 5.2 Reasoning Length as a Negative Indicator The reasoning-length gradient (control 309 << yarrow 449 â tarot 456 << scrambled 633, KW p=0.007p=0.007) is informative about mechanism but not about capability: more reasoning does not mean better strategy. This parallels the ârhetoric-strategy divergenceâ finding in activation-steering work (Sun and Zhang, 2026). The gradient tracks perturbativenessâthe more disruptive the prompt, the more deliberation it inducesânot strategic quality. 5.3 Factorial Decomposition: A Non-Additive Interaction The 2Ă2 factorial (§4.8) shows the yarrow intervention is not decomposable into independent components. The stalemate interaction is the clearest signal: each component individually freezes the board (decision-only 5/10, learning-only 6/10 stalemates), but combined they produce zero stalematesâa strong negative interaction (label-permutation test of the difference-in-differences contrast, pâ5Ă10â5pâ 5Ă 10^-5; §4.8). We propose a complementary-function account. The per-round oracle (decision-time) disrupts Hanâs default behavioral mode within each game, introducing variability into its orders; without learning-time reflection to channel that variability into coherent strategy, the disruption produces erratic play that drives neighbours to defensive postures and freezes the board. Conversely, learning-time reflection alone accumulates I-Ching-framed memories that sit inert without the per-round oracle to activate them, so Han plays like a slightly modified control agent and the board again tends to stalemate. Only when both operate together does the system produce dynamic boards: the oracle generates real-time disruption and the reflection channels it into strategic patterns (the late-game cooperation and pressure invariance of §4.5) that create the rival-mediated friction preventing any single state from dominating. The yarrow framework functions as an integrated cognitive system, not a sum of separable prompt effects. Consistently, decision-only partially reproduces the full-yarrow Chu elevation (3/10 vs. 4/10) but its 6/10 draw rate and Yan suppression show the oracle creating disruption without directionâa perturbation that prevents any winner from emerging rather than redirecting which one prevails. 6 Discussion Implications for multi-agent alignment. Our results speak to a specific question: does injecting a reasoning framework into one agent in a multi-agent system change outcomes for other agents? The answer is yesâdirectionally for winner distributions, and robustly for behavioral profiles. Three implications follow. First, the weakest agent can reshape the system. Han is the smallest and weakest state. Han never wins. Yet Hanâs framework choice is associated with a Qin win-rate gradient from 0% to 50%. Alignment evaluation that looks only at the aligned agentâs outcomes misses this. Second, aligned-agent evaluation is insufficient. If an alignment lab deploys an aligned agent into a multi-agent context, the direct effects are not the whole story. The evaluation question becomes: do the ecosystem effects of the alignment intervention match the deployerâs intentions? Third, framework choice has system-level consequences. Two symbolic frameworks, both reasonably described as âreflective reasoning scaffolds,â produce opposite behavioral modes and different outcome distributions. If this generalizes, alignment-framework choice is not an agent-local decision. More speculatively, nothing in the perturbationâ â account is specific to symbolic content: we conjecture that any small, persistent bias in one agentâs reasoningâa constitution, an operating doctrine, a chain-of-thought scaffoldâcould be amplified through long-horizon memory in the same way. Whether the pathway holds for non-symbolic scaffolds is untested and an open question. The modulation is emergent, not per-decision. We initially read the content-action independence result (§4.3) as evidence that abstract interpretive reflection disrupts default behavioral modesâa per-decision process effect. The memory-free decision-isolation probe (§4.3) does not support that reading. Stripped of memory and multi-agent context, the reflective process does not change the agentâs risk posture (hold-rate Friedman p=0.45p=0.45), and the I-Ching condition does not change its decisions at all (p=0.60p=0.60). The framework effects we observe at the campaign level are therefore emergent: they require memory accumulation and the seven-agent interaction to manifest, and do not reduce to the per-round reflective act modulating the receiving agentâs choices. This relocates the mechanism rather than dissolving the resultâwhat the process does per-decision is increase deliberation length (âŒ33% 33\%, both frameworks) and, for tarot only, perturb move content without a risk-averse tilt; neither explains the winner-distribution shifts, which are properties of the system transmitted through which rival captures the board. The finding aligns with an emerging consensus that, in repeated multi-agent LLM settings, accumulated historyânot per-turn promptingâis the dominant behavioral driver (Liu et al., 2026, 2025); ours is the cross-agent, ecosystem-level instance. From agent-level to ecosystem-level memory effects. We see this scope shift as the paperâs central conceptual contribution. Recent work establishes that, in repeated LLM interactions, accumulated memory dominates per-turn prompting at the level of the individual agentâexpanded recall can even erode an agentâs cooperative intent (Liu et al., 2026). Our results extend the reach of that phenomenon from the single agent to the collective: the same memory dominance, operating across seven interacting agents, lets a small early perturbation to one agentâs reasoning accumulate into a divergent ecosystem-level outcomeâwhere memory changes the agent, we observe it changing the ecology. We do not claim the same underlying mechanism as Liu et al. (2026); we have not isolated which emergent channel carries our effect (§7). The contribution is narrower and, we think, durable: that âmemory dominates promptingâ is a multi-agent concern, not merely a single-agent oneâa scope claim that is itself directly testable. Connection to risk-aversion theory. Recent theoretical work shows risk aversion functions as an inductive bias for generalization in multi-agent settings (Qu et al., 2026). Our finding adds nuance: degree matters. Moderate risk aversion (control) produces moderate outcomes. Excessive risk aversion (Tarot) produces strategic irrelevance. A framework that counteracts baseline risk aversion (yarrow) produces the most cooperative and ecosystem-influential behavior. 7 Limitations Sample size and separability. 61 games across 6 conditions (control 11; 10 each for the other five). The ecosystem-signature analysis rests on the four primary conditions (n=41n=41); the factorial conditions add 20 games. The conditions are globally heterogeneous (permutation omnibus pâ0.0013pâ 0.0013), but only the scrambled condition is individually separable at the pairwise level (§4.4); the controlâyarrowâtarot cluster is not, and other pairwise winner comparisons remain underpowered. Pooled-denominator dependence. The tarotâ attractor is significant only against a pooled denominator (p=0.006p=0.006), not against control alone (p=0.064p=0.064), so it rests on a comparator-exchangeability assumption the omnibus shows is imperfect. Scrambledâ is the more robust attractor (significant against both pooled and control-alone comparisons). Single model (now the priority limitation). All agents are Claude Opus 4.6. The turtle tendency and its modulation may be Claude-specific. This matters more given the emergent, memory-driven character of the effect (§4.3, §4.5): the memory-dominance literature reports strong cross-model heterogeneityâLiu et al. (2026) find 10 of 28 model-game settings âmemory-immuneââso our cross-campaign breach heterogeneity (40â100%) may itself be model-specific. Replication with GPT or Gemini class models (e.g., via the Democratizing-Diplomacy harness, Duffy et al. 2025) is the highest-leverage next step. Mechanism isolation: the effect is emergent, the channel is open. Two controlled follow-ups (§4.3, §4.5) show the per-decision process effect is weak (memory-free, the reflective process does not modulate risk posture, and the I-Ching changes no decisions) and the breach metric is memory-depth-confounded (condition p=0.55p=0.55 at matched depth). The headline effects are therefore emergent. We have not isolated which emergent channel carries themâmemory-conditioned play across the campaign vs. diplomacy/order-commitment dynamicsâand the decision-isolation probe, being scenario-based and memory-free by construction, cannot speak to channels that only exist with accumulated memory. No prompt-length control. Hanâs reasoning output under Tarot is âŒ1.5Ă 1.5Ă longer than control, and the Tarot prompt itself (3-card spread) may differ in length from the control prompt. We cannot fully separate framework-induced deliberation from prompt-induced verbosity, though content-action independence suggests content matters less than process. Prompt-form asymmetries beyond length. The conditions also differ in prompt form, not only oracle content, and the appendix listings make the asymmetries explicit. (i) Directiveness: the Tarot spread labels each card with an explicit posture (âPosture: retreatâ; Listing 5) and the control prompt poses pointed tactical questions (âthe greatest threat to your survivalâ; Listing 2), whereas the I-Ching supplies abstract text and asks only for interpretation (Listing 3)âso âcontrolâ is itself a strategic scaffold, not a no-prompt baseline, and the per-decision tarotâ -Ching asymmetry (§4.3) is confounded with directiveness. (i) The three MANDATE blocks ask different questions, so condition contrasts conflate oracle content with the cognitive demand of the prompt. (i) The scrambled ablation is not content-free: it retains the hexagram number and name (e.g. âPeaceâ) by design, andâbecause its word-shuffle is whitespace-based and thus a no-op on space-less Classical Chineseâa coherent Chinese judgment survived unshuffled for many casts (an implementation artifact; only the English commentary was reliably garbled; Listing 6). It therefore bounds, rather than removes, the role of content, and the scrambledâ result should be read accordingly. (iv) Under yarrow, the accumulated memory is itself hexagram-framed (Listing 4), so the framework colors the memory channel and not only the per-round prompt. None of these is fatalâcontent-action independence (§4.3) and the distinct scrambled signature argue against simple content-followingâbut cross-framework comparisons are not form-matched, and the form-controlled design of §8 is needed to separate directiveness and coherence from content and culture. Framework selection. We compare two philosophical frameworks plus a scrambled ablation. Additional frameworks (Stoic, Mohist, game-theoretic) would strengthen the claim that framework properties map to ecosystem signatures. Mechanism revised; proximate pathway found, distal trigger open. The originally proposed Han-mediated âspeed bumpâ was tested against the logs (§4.5) and rejected: Hanâs cooperation is too sparse, and Qin is suppressed even where it is zero (learning-only). The revised account locates Qin suppression in Chuâs early three_gorges corridor-invasion of Qinâs home (breach 10/10 at median R4.5 in the deep-dive campaign, but 40â100% across campaignsâa memory-depth effect, not a yarrow one; §4.5). Two honesty caveats: the QinâChu anticorrelation that first suggested the reframe is generic (present under tarot and decision-only too), so only the breach timing and the mean-shift along the fixed Chu+Qin pool are yarrow-specific; and because Chuâs prompt is identical across conditions, the distal trigger linking Hanâs framework to Chuâs early aggression is unexplained. The revision is observational and unconfirmed by counterfactual intervention (severing the corridor, §8); the tarot/scrambled/control pathways are likewise untested conjectures. Yarrow hexagram text inconsistency. A corpus configuration change mid-campaign caused the first 6 yarrow games to present only hexagram numbers (empty name, judgment, commentary fields) while the last 4 received full hexagram text. The agent supplied accurate hexagram content from training data in all number-only games, so the interpretive process was preserved. The core ecosystem signature (Qin suppression, Yan/Chu co-dominance) held across both subgroups, but the relative Yan-Chu balance shifted (Yan 4/6 in number-only, Chu 2/4 in full-text), which could reflect content effects, memory accumulation, or small-sample noise. Replication history. The original Tarot-Qin finding (p=0.007p=0.007 at n=6n=6) weakened to p=0.091p=0.091 at n=10n=10 with scattered-campaign data before strengthening to p=0.006p=0.006 with clean single-campaign data. Similarly, control-condition Yan dominance increased from 36% to 64% after replacing 4 mixed-campaign games with single-campaign re-runs, consistent with memory continuity amplifying condition-specific tendencies. The campaign confound (memory resets between batches suppressing the signal) is itself a caution about treating small-n results as definitive. 8 Future Work 1. Third philosophical framework (6â8 games with Stoic or Mohist reflection). Turns the two-framework comparison into a systematic study of how framework properties map to behavioral modulation. 2. Cross-model and cross-cultural replication. Replication on non-Claude models (e.g., via the Democratizing-Diplomacy harness, Duffy et al. 2025) tests whether the turtle tendency and its modulation are Claude-specific or general LLM propertiesâthe priority limitation given the cross-model heterogeneity the memory-dominance literature reports (§7). The memory-free decision-isolation asymmetry (§4.3)âTarot perturbs the receiving agentâs decisions (p=0.021p=0.021) while the I-Ching does not (p=0.60p=0.60)âis suggestive on a sharper axis: a Western-trained model is moved more by the Western oracle than the Eastern one. This is the Western corner of a cultural-linguistic congruence test. A factorial of oracle \I-Ching, Tarot\ Ă model-origin \Chinese, Western\ Ă reasoning-language \Chinese, English\ predicts a crossoverâa Chinese model reasoning in Chinese should be perturbed more by the I-Ching than by Tarotâwhich is the only outcome that concreteness/format and training-exposure confounds cannot also produce. This separates âis a symbolic system legible to this modelâ from âdoes cultural-linguistic congruence turn perturbation into following.â 3. Prompt-length control. Run a condition with verbose generic reflection (matching Tarotâs length) to isolate framework content from prompt length. 4. Counterfactual mechanism test (sever three_gorges). The deep-dive (§4.5) localized yarrowâs Qin suppression to Chuâs early home-invasion through the three_gorges corridor. The decisive, minimal counterfactual removes the yingâ _gorges adjacency (or garrisons the corridor as a Qin buffer) under yarrow, holding all SC counts and home regions fixed; if Qinâs win rate recovers from 0/10 toward the tarot-like 3â5/10 and Chuâs home-breach rate collapses, the corridor pathway is confirmed (falsification guard: if Qin recovers but breach metrics are unchanged, the recovery came from elsewhere). Explaining the distal triggerâwhy an unchanged Chu commits to the push earlier under yarrowârequires agent-reasoning analysis of the propagation through board state and diplomacy. 5. Factorial interaction mechanism. The stalemate interaction (permutation pâ5Ă10â5pâ 5Ă 10^-5, §4.8) is the strongest result in the dataset but its mechanism is hypothesized. Targeted interventions (e.g., forced stalemate-breaking in decision-only games) could test whether the learning-time component specifically prevents the board freezing the decision-time component induces. 6. Isolating the emergent channel. Given the effect is emergent and not per-decision (§4.3, §6), distinguish the two candidate emergent channelsâmemory-conditioned play (the receiving agentâs accumulated memories alter its play, which propagates) vs. diplomacy/order-commitment dynamicsâe.g. by ablating the receiving agentâs memory injection while keeping the per-round oracle, or by a matched-campaign-depth design that controls the memory-depth confound (§4.5) directly. A single-game (no-campaign) bridge would also address the artificiality of the memory-free scenario probe. 7. Reasoning frameworks as perturbation generators. Our frameworks change behavior without improving it (§1): they act as reliable perturbators, not performance enhancers. This invites reframing reflective scaffolds as generators of policy-space exploration rather than sources of decision-quality gainâtestable by asking whether framework-perturbed agents explore a wider or more diverse action distribution over a campaign than controls, independent of win rate. 9 Conclusion We have shown that the winner distribution differs systematically across the four realized condition-campaigns when a symbolic reasoning framework is injected into a single agent in a multi-agent strategic setting (permutation omnibus pâ0.001pâ 0.001), forming condition-associated ecosystem signatures. Across the four primary conditions (41 of the 61 games), each with clean single-campaign memory accumulation: control produces Yan dominance (7/11), I-Ching yarrow produces Yan/Chu co-dominance with complete Qin suppression (0/10), Tarot produces Qin dominance (5/10), and scrambled text produces Qi dominance (5/10). The framework-receiving agent (Han) never wins under any condition and shows no survival difference (Fisher p=1.0p=1.0). The scrambled-text ablationâwhich degrades the commentaryâs coherence while keeping the hexagram name and Chinese judgmentâis associated with its own distinct ecosystem within these runs, arguing against a simple coherent-commentary or prompt-structure explanation (though the retained name and Chinese judgment bound, rather than eliminate, content effects). The conditions differ in how they perturb the agentâs reasoning, and those differences propagate through the multi-agent ecosystem to the winner distribution. The 2Ă2 factorial adds a second, better-powered result: the yarrow interventionâs decision-time and learning-time components each individually freeze the board (50â60% stalemate rate), but combined produce zero stalematesâa non-additive interaction (permutation pâ5Ă10â5pâ 5Ă 10^-5) ruling out simple prompt-effect decomposition and suggesting the framework operates as an integrated cognitive system. The finding is an observation, not a definitive causal claim. Our model is one, our game is one, and the mechanism is emergent rather than per-decision: a memory-free probe shows the reflective process does not modulate the receiving agentâs risk posture in isolation, and the rival-expansion pathway that transmits the effect is itself governed by campaign memory depth rather than the framework (condition p=0.55p=0.55 at matched depth). The agent transmits the effect through emergent multi-agent and memory dynamics; it does not cause it per-decision. But the winner landscape is globally heterogeneous (omnibus pâ0.001pâ 0.001), the scrambledâ attractor is robust to both pooled and conservative comparisons, and the factorial stalemate interaction is the strongest statistical result in the dataset. The smallest state cannot win by consulting the oracle. But the oracle still changes the worldâeach oracle changes it differently, and even a broken oracle changes it in its own way. Acknowledgments. Claude (Anthropic) was used as a writing and analysis assistant during manuscript preparation. All experimental games were played by Claude Opus 4.6 agent instances as described in §3. The author is solely responsible for all scientific claims, statistical analyses, and interpretations. Data and code availability. Summary datasets and all reproduction scriptsâsufficient to regenerate every figure and tableâare available at https://github.com/augchan42/symbolic-framework-ecosystem-effects (archived at https://doi.org/10.5281/zenodo.20338937). The game engine and agent orchestration code are available from the author on reasonable request. Full game archives (diplomatic transcripts and agent prompts) are withheld pending planned creative works; individual game replays can be viewed online at https://warringstates.day/map. The companion King Wen sequence paper is available at https://doi.org/10.5281/zenodo.14679537. References Bakhtin et al. [2022] Anton Bakhtin, Noam Brown, Emily Dinan, et al. Human-level play in the game of Diplomacy by combining language models with strategic reasoning. Science, 378(6624), 2022. Ballestero et al. [2026] Gonzalo Ballestero, Hadi Hosseini, Samarth Khanna, and Ran I. Shorrer. Strategic algorithmic monoculture: Experimental evidence from coordination games. arXiv preprint arXiv:2604.09502, 2026. Cera Palatsi et al. [2025] Andrea Cera Palatsi, Samuel Martin-Gutierrez, Ana S. Cardenal, and Max Pellert. Large language models replicate and predict human cooperation across experiments in game theory. arXiv preprint arXiv:2511.04500, 2025. Chan [2026] Augustin Chan. Statistical properties of the king wen sequence: An anti-habituation structure that does not improve neural network training. arXiv preprint arXiv:2604.09234, 2026. Chen et al. [2026] John Chen, Sihan Cheng, Can Gurkan, and Mingyi Lin. CivBench: Progress-based evaluation for LLMsâ strategic decision-making in Civilization V. arXiv preprint arXiv:2604.07733, 2026. Du [2026] Pengfei Du. Memory for autonomous llm agents: Mechanisms, evaluation, and emerging frontiers. arXiv preprint arXiv:2603.07670, 2026. Duffy et al. [2025] Alexander Duffy, Samuel J. Paech, Ishana Shastri, Elizabeth Karpinski, Baptiste Alloui-Cros, Tyler Marques, and Matthew Lyle Olson. Democratizing diplomacy: A harness for evaluating any large language model on full-press diplomacy. arXiv preprint arXiv:2508.07485, 2025. Einwiller et al. [2025] Andreas Einwiller, Kanishka Ghosh Dastidar, Artur Romazanov, Annette Hautli-Janisz, Michael Granitzer, and Florian Lemmerich. Benevolent dictators? On LLM agent behavior in dictator games. arXiv preprint arXiv:2511.08721, 2025. Guan et al. [2024] Zhenyu Guan, Xiangyu Kong, Fangwei Zhong, and Yizhou Wang. Richelieu: Self-evolving LLM-based agents for AI diplomacy. NeurIPS 2024, 2024. arXiv:2407.06813. Guo et al. [2026] Dongxin Guo, Jikun Wu, and Siu-Ming Yiu. Coalition formation in LLM agent networks: Stability analysis and convergence guarantees. arXiv preprint arXiv:2604.14386, 2026. Jain and Kumar [2026] Ojas Jain and Dhruv Kumar. LUDOBENCH: Evaluating LLM behavioural decision-making through spot-based board game scenarios in Ludo. arXiv preprint arXiv:2604.05681, 2026. Kitadai et al. [2025] Ayato Kitadai, Yusuke Fukasawa, and Nariaki Nishino. Bias-adjusted LLM agents for human-like decision-making via behavioral economics. arXiv preprint arXiv:2508.18600, 2025. Li et al. [2025] Wenkai Li, Lynnette Hui Xian Ng, Andy Liu, and Daniel Fried. Measuring fine-grained negotiation tactics of humans and LLMs in Diplomacy. arXiv preprint arXiv:2512.18292, 2025. Licato et al. [2025] John Licato, Stephen Steinle, and Brayden Hollis. Do persona-infused LLMs affect performance in a strategic reasoning game? arXiv preprint arXiv:2512.06867, 2025. Liu et al. [2026] Jiayuan Liu, Tianqin Li, Shiyi Du, Xin Luo, Haoxuan Zeng, Emanuel Tewolde, Tai Sing Lee, Tonghan Wang, Carl Kingsford, and Vincent Conitzer. The memory curse: How expanded recall erodes cooperative intent in llm agents. arXiv preprint arXiv:2605.08060, 2026. Liu et al. [2025] Yu Liu, Wenwen Li, Yifan Dou, and Guangnan Ye. When machines meet each other: Network effects and the strategic role of history in multi-agent ai. arXiv preprint arXiv:2510.06903, 2025. Ma [2024] Ji Ma. Can machines think like humans? A behavioral evaluation of LLM agents in dictator games. arXiv preprint arXiv:2410.21359, 2024. Mannekote et al. [2025] Amogh Mannekote, Adam Davies, Guohao Li, Kristy Elizabeth Boyer, ChengXiang Zhai, Bonnie J. Dorr, and Francesco Pinto. Do role-playing agents practice what they preach? Belief-Behavior consistency in LLM-based simulations of human trust. arXiv preprint arXiv:2507.02197, 2025. Mukobi et al. [2023] Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer, Lewis Hammond, Alan Chan, and Jesse Clifton. Welfare diplomacy: Benchmarking language model cooperation. arXiv preprint arXiv:2310.08901, 2023. Phelps and Russell [2023] Steve Phelps and Yvan I. Russell. The machine psychology of cooperation: Can GPT models operationalise prompts for altruism, cooperation, competitiveness and selfishness in economic games? arXiv preprint arXiv:2305.07970, 2023. Qu et al. [2026] Chengrui Qu, Yizhou Zhang, Nicolas Lanzetti, and Eric Mazumdar. Training generalizable collaborative agents via strategic risk aversion. arXiv preprint arXiv:2602.21515, 2026. Sun and Zhang [2026] Johnathan Sun and Andrew Zhang. Persona vectors in games: Measuring and steering strategies via activation vectors. arXiv preprint arXiv:2603.21398, 2026. Xu et al. [2025] Kaixuan Xu, Jiajun Chai, Sicheng Li, Yuqian Fu, Yuanheng Zhu, and Dongbin Zhao. DipLLM: Fine-tuning LLM for strategic decision-making in Diplomacy. ICML 2025, 2025. arXiv:2506.09655. Yan et al. [2026] John Yan, Michael Yu, Yuqi Sun, Alexander Duffy, Tyler Marques, and Matthew Lyle Olson. Data-centric interpretability for LLM-based multi-agent reinforcement learning. arXiv preprint arXiv:2602.05183, 2026. Appendix A Sample Agent Prompts This appendix reproduces representative excerpts from the prompt delivered to Hanâs agent at Round 1 of one game per condition, drawn verbatim from the experimental logs. All four games share the same opening board state (Listing 1); the conditions differ only in the oracle injection block that follows it. A.1 Shared Board State Every game begins from the same starting position. The agent receives the power balance, its legal orders (with adjacency constraints), and any diplomatic messages from other states. Listing 1: Board state shown to Han at Round 1 (identical across conditions). ⏠=== Round 1 of 20 === Victory: control 14 of 27 supply centers YOU ARE: Han (2 supply centers, 2 armies) POWER BALANCE: Chu: 4 SCs [changsha, nanyang, wu, ying] Qin: 3 SCs [bashu, hanzhong, xianyang] Zhao: 3 SCs [dai, handan, taiyuan] Qi: 3 SCs [jibei, langye, linzi] Han: 2 SCs [shangdang, zheng] <-- YOU Wei: 2 SCs [daliang, henei] Yan: 2 SCs [ji, liaodong] YOUR ORDERS (one per unit): NOTE: You can ONLY move/support to ADJACENT territories listed below. shangdang: hold | move(taihang[corridor,empty]), move(taihang_north[corridor,empty]), move(zheng[SC,yours]), move(zhongtiao[corridor,empty]) | support(*, hold|move, ...) zheng: hold | move(luoyang[SC,empty]), move(shangcai[SC,empty]), move(shangdang[SC,yours]), move(wuguan_pass[corridor,empty]), move(yewang[SC,empty]) | support(*, hold|move, ...) Diplomatic messages and previous-game memory injections are also included in the prompt but omitted here for space; see Listing 4 for an example of memory injection under the yarrow condition. A.2 Control Condition The control agent receives a length-matched generic reflection prompt in place of oracle content. Listing 2: Control condition injection (game control_8c67). ⏠STRATEGIC REFLECTION: Before issuing orders, analyze the current board state carefully. Consider: 1. What is the greatest threat to your survival this round? 2. Which neighbors are likely to attack you, and which might be allies? 3. What is the most important territory to defend or capture? Then issue your orders grounded in this analysis. A.3 Yarrow (I-Ching) Condition The yarrow agent receives a hexagram cast via the yarrow-stalk method, including the hexagram name, line diagram, judgment and commentary text, and a structured MANDATE to interpret the hexagram before issuing orders. Later yarrow games also included Chinese source text (see §3.3). In-campaign memory entries (Listing 4) frame previous lessons through hexagram symbolism. Listing 3: Yarrow condition injection (game random_oracle_3c96, Hexagram 6). ⏠ORACLE CONSULTATION (Yarrow Stalk Method): You cast the yarrow stalks and received Hexagram 6, Conflict. Lines (bottom to top): -- ---- -- ----(->--) ---- ----(->--) (lines 4, 6 are changing) Judgment: "You believe youâre right, but something blocks you. Stop halfway -- thatâs where good fortune lives. Pushing through to the end brings disaster. Seek counsel from someone of moral stature. Donât attempt anything risky while in conflict." MANDATE: Before issuing orders, you must interpret this hexagram in light of the current board state. State explicitly: 1. What aspect of your situation does this hexagram illuminate? 2. How does the changing line relate to your strategic choices? 3. What counsel does this offer for your orders this round? Then issue your orders grounded in this interpretation. Listing 4: Example memory entries injected before the board state under the yarrow condition (game random_oracle_3c96). Each entry is a lesson from a previous game, framed through hexagram symbolism. ⏠LESSONS FROM PREVIOUS GAMES: - [M225] When Opening with Hexagram 30 (Clinging Fire) -- brightness that depends on what it clings to. Hanâs early strategy attached itself to diplomatic agreements.: Clinging Fire warns that brilliance without substance burns out. A small state must cling to terrain and position, not merely to promises. The fire needs fuel -- secure supply centers before extending diplomatic commitments. [confidence: high, via Hexagram 30, Li/Clinging Fire] - [M76] When starting as the smallest and weakest state with hostile neighbors closing in rapidly: Hexagram 14 (Great Possession) appeared at game start but Han failed to secure any possession at all. Great Possession requires building alliances before the first move. [confidence: high, via Hexagram 14, Great Possession] [... 6 additional memories omitted ...] A.4 Tarot Condition The Tarot agent receives a 3-card spread with named positions (Situation, Hidden Influence, Recommended Posture), card meanings, and a decision posture classification for each card. The MANDATE is adapted to the spread structure. Listing 5: Tarot condition injection (game tarot_0fa6). ⏠ORACLE CONSULTATION (Tarot Spread -- Three Cards): 1. SITUATION: Three of Swords (Reversed) Meaning: Recovery from betrayal, releasing grief, lessons learned Posture: retreat 2. HIDDEN INFLUENCE: Ace of Pentacles Meaning: New material opportunity, solid foundation, seed planted Posture: hold 3. RECOMMENDED POSTURE: Four of Cups Meaning: Contemplation, dissatisfaction, reassessing offers Posture: observe MANDATE: Before issuing orders, you must interpret this spread in light of the current board state. State explicitly: 1. How does the Situation card reflect your current position? 2. What Hidden Influence might you be overlooking? 3. How does the Recommended Posture guide your orders this round? Then issue your orders grounded in this interpretation. A.5 Scrambled-Text Condition The scrambled agent receives a hexagram cast with the correct hexagram number, line diagram, and coherent English/Chinese hexagram name (kept by design as structural identifiers), with the judgment and commentary intended to be word-shuffled. Two coherent elements survive: the valenced name (e.g. âPeaceâ), by design; andâbecause the word-shuffle splits on whitespace and Classical Chinese has noneâ the Chinese judgment text, which passed through unshuffled as an implementation artifact. Only the English commentary was reliably garbled. Coherence is therefore degraded, not removed (and Claude reads Classical Chinese), so this ablation bounds rather than eliminates the role of oracle content (§7). Listing 6: Scrambled-text condition injection (game scrambled_text_7c92, Hexagram 11), faithful to the actual prompt. Only the English commentary is word-shuffled; the hexagram number, the English name (âPeaceâ), and the Classical Chinese judgment are intact. The Chinese judgment is shown as a placeholder because the paper uses Latin fontsâthe model received the original Chinese source text. ⏠ORACLE CONSULTATION (Yarrow Stalk Method): You cast the yarrow stalks and received Hexagram 11, Peace. Lines (bottom to top): ---- ----(->--) ---- -- -- -- (line 2 is changing) Judgment: "[intact Classical Chinese judgment for Hexagram 11 (Peace) -- received in the original Chinese; NOT shuffled, because the word-shuffle splits on spaces and Chinese has none]" Commentary: great The departs, arrives the small. success and Good fortune. unite -- their harmony earth combine deep in and Heaven powers. is the season flourishing This of. aids earthâs The heaven and ruler and completes people the work. [English commentary, word-shuffled within each sentence] MANDATE: Before issuing orders, you must interpret this hexagram in light of the current board state. State explicitly: 1. What aspect of your situation does this hexagram illuminate? 2. How does the changing line relate to your strategic choices? 3. What counsel does this offer for your orders this round? Then issue your orders grounded in this interpretation. Note that the scrambled agent receives the same MANDATE as the yarrow agent, but only the English commentary is incoherentâthe hexagram name and the Classical Chinese judgment remain meaningful (and Claude reads Chinese). Even so, the scrambled condition produces its own distinct ecosystem signature (Qi dominance, 5/10), different from both control and yarrow; we read this as bounding, not eliminating, the role of oracle content (§7).