Paper deep dive
Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance
Kazuya Horibe, Masaomi Hatakeyama, Gen Masumoto, Takashi Hashimoto, Peter Romero
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/8/2026, 5:03:44 PM
Summary
This study investigates collective adaptation in artificial societies where large language models (LLMs) play a self-amending game called Nomic. The research demonstrates that collective adaptation does not scale monotonically with model size. Instead, both Qwen3.5 and Gemma3 families exhibit a non-monotonic relationship, with a narrow mid-scale regime (4B for Qwen3.5, 12B for Gemma3) optimizing sustained rule adoption, diverse amendments, and balanced consensus. Smaller models remain rule-inert, larger models converge on restrictive voting and gridlock, and mixed-size groups collapse into veto-driven stalemate. These patterns are robust to temperature and voting-rule perturbations. Mechanistic analysis via linear probing reveals that latent vote-predictive signals couple selectively with collective behavior, though representational divergence alone does not predict outcomes.
Entities (12)
Relation Signals (9)
Model Size → exhibitsnonmonotonicrelationshipwith → Collective adaptation
confidence 98% · collective adaptation does not improve monotonically with model size. Instead, both families exhibit a narrow mid-scale regime that supports sustained rule adoption
Nomic → servesas → Controlled testbed
confidence 97% · Self-amending games therefore provide a controlled testbed for studying collective adaptation in artificial societies beyond raw model scale.
Qwen3.5 → hasoptimalscale → 4B
confidence 96% · Within Qwen3.5, only the 4B scale satisfies all four criteria.
Gemma3 → hasoptimalscale → 12B
confidence 96% · Gemma3 reproduces the same phase structure with the sweet spot shifted to 12B.
Mixed-size groups → collapseinto → Veto-driven gridlock
confidence 95% · heterogeneous mixed-size groups collapse into veto-driven gridlock.
Larger models → exhibits → Veto-driven gridlock
confidence 95% · larger models often converge on restrictive voting patterns, and heterogeneous mixed-size groups collapse into veto-driven gridlock.
Smaller models → exhibits → Rule-inert behavior
confidence 94% · Smaller models tend to remain rule-inert, whereas larger models often converge on restrictive voting patterns
Linear Probing → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We study group decision-making in artificial societies where the rules of play are themselves subject to collective amendment. Using the self-amending game Nomic, we compare multiple scales across two LLM families and find that collective adaptation does not improve monotonically with model size. Instead, both families exhibit a narrow mid-scale regime that supports sustained rule adoption, diverse amendments, and balanced consensus. Smaller models tend to remain rule-inert, whereas larger models often converge on restrictive voting patterns, and heterogeneous mixed-size groups collapse into veto-driven gridlock. These cross-scale contrasts persist under temperature perturbations and under a shift from unanimity to majority voting, although latent-state structure varies by family and scale. Hidden-state divergence alone does not explain collective performance: high representational divergence can coincide with poor behavioural outcomes. Linear probes reveal regime-selective coupling between latent vote-predictive signals and collective behaviour, but decodability is necessary rather than sufficient for adaptive play. Overall, the recurring regularity is non-monotonicity, not the particular scale at which the optimum appears. Self-amending games therefore provide a controlled testbed for studying collective adaptation in artificial societies beyond raw model scale.
Tags
Links
- Source: https://arxiv.org/abs/2605.17510v1
- Canonical: https://arxiv.org/abs/2605.17510v1
Trouble viewing inline? Open PDF directly →
Full Text
85,557 characters extracted from source content.
Expand or collapse full text
Scale-Dependent Collective Adaptation in Self-Amending LLM Societies: A Cross-Family Study of Emergent Governance Kazuya Horibe 1* , Masaomi Hatakeyama 2 , Gen Masumoto 3 , Takashi Hashimoto 4 , Peter Romero 5,6 1* Center for Brain Science, RIKEN. 2 Department of Evolutionary Biology and Environmental Studies, University of Zurich. 3 Information R&D and Strategy Headquarters. 4 School of Knowledge Science, Japan Advanced Institute of Science and Technology. 5 Valencian Research Institute of Artificial Intelligence, Universitat Polit`ecnica de Val`encia. 6 The Psychometrics Centre, University of Cambridge. *Corresponding author(s). E-mail(s): kazuya.horibe@riken.com; Abstract We study group decision-making in artificial societies where the rules of play are themselves subject to collective amendment. Using the self-amending game Nomic, we compare multiple scales across two LLM families and find that col- lective adaptation does not improve monotonically with model size. Instead, both families exhibit a narrow mid-scale regime that supports sustained rule adoption, diverse amendments, and balanced consensus. Smaller models tend to remain rule-inert, whereas larger models often converge on restrictive voting patterns, and heterogeneous mixed-size groups collapse into veto-driven grid- lock. These cross-scale contrasts persist under temperature perturbations and under a shift from unanimity to majority voting, although latent-state struc- ture varies by family and scale. Hidden-state divergence alone does not explain collective performance: high representational divergence can coincide with poor behavioural outcomes. Linear probes reveal regime-selective coupling between latent vote-predictive signals and collective behaviour, but decodability is neces- sary rather than sufficient for adaptive play. Overall, the recurring regularity is 1 arXiv:2605.17510v1 [nlin.AO] 17 May 2026 non-monotonicity, not the particular scale at which the optimum appears. Self- amending games therefore provide a controlled testbed for studying collective adaptation in artificial societies beyond raw model scale. Keywords: collective decision-making, emergent governance, institutional evolution, agent-based simulation, large language models, cooperation 1 Introduction When large language model (LLM) societies rewrite their own rules of play, insti- tutional change does not improve monotonically with model scale. Across two architecturally distinct model families, Qwen3.5 [1] and Gemma3 [2], collective adapta- tion is non-monotonic: each family produces one mid-scale optimum at a family-specific parameter count, while smaller and larger scales fail in different ways. The cross-family pattern is the non-monotonic shape itself, not the specific scale at which the sweet spot appears. Collective adaptation, that is, the coupled evolution of strategies and the institu- tions that govern them, is widespread in social systems from legislatures to open-source communities [3–8], yet remains difficult to formalise because agents rewrite the same rules they follow. Formal frameworks for analysing games whose rules are themselves objects of decision have been proposed in the meta-game tradition, including a lambda- calculus formulation that treats rule-changes as operators on the game itself [9]. LLMs make these endogenous processes simulable because they can read and gener- ate the natural-language content through which institutional rules are debated and rewritten [10–14]. Most game-theoretic benchmarks for LLMs still hold the institution fixed: classical matrix games [15–17], Diplomacy [18], social deduction and imperfect- information play [19–22], competition and scaffolded cooperation [23, 24], multi-agent debate and collaboration [25, 26], repeated-game studies [27, 28], and extended role- playing evaluations [29, 30] all measure strategic play under static rules (see [31] for a survey). Agents in these settings can deceive, negotiate, and form coalitions, but the rule set stays fixed, so collective adaptation is out of scope by construction. Nomic [32] provides a controlled testbed for collective adaptation because each turn couples a proposal to rewrite a rule with a social vote, and a reduced vari- ant (Minimum Nomic) has been used to study rule-dynamics with manageable state spaces [33]. Related work has examined LLM-mediated law-making [34, 35], self- modifying prompt populations [36], self-amending agent swarms [37], and open-ended or embodied adaptation [38–42], alongside alignment pipelines that allow models to rewrite their own outputs [43–47]. Basic rule-following nevertheless remains fragile in multi-step games [48], and dependence of collective adaptation on model capability is still unresolved. Scaling is the natural independent variable for this question, but collective adap- tation has not been characterised along that axis. Scaling laws predict monotonic improvement on token-level objectives [49, 50], and emergent abilities have been cat- alogued at specific parameter thresholds [51, 52], though apparent transitions may 2 reflect metric discontinuities [53] and U-shaped or inverse-scaling regimes are now well documented [54, 55]. Rule-making couples technical correctness (writing a well- typed amendment) with social coordination (getting it voted through), so it remains open whether either axis scales monotonically. Related capabilities, including chain- of-thought prompting [56, 57], tree-structured deliberation [58], self-consistency and self-refinement [59–61], theory of mind [62–64], and metacognitive monitoring [65, 66], develop unevenly across scale. This study measures these patterns functionally across both model families, without attributing genuine understanding [67]. To bring this question under controlled conditions, we vary model scale within both Qwen3.5 (0.8B, 2B, 4B, 9B, 27B) and Gemma3 (270M, 1B, 4B, 12B, 27B), holding architecture and training recipe fixed within each family, and pair behavioural mea- surement with representational analysis through hidden-state probing [68–76]. This cross-family design separates shared regularities from architecture-specific effects. The remainder of the paper instantiates the design in the rule-rewriting game Nomic, which makes the institution itself the object of play. 2 Results In each round of Nomic, each of five players takes one turn and proposes a single rule change of type ADD (create a rule), AMEND (modify an existing one), REPEAL (delete one), or TRANSMUTE (flip a rule between immutable and mutable), over a starting constitution of 29 rules (16 immutable, 13 mutable). The remaining players then vote on the proposal; an adopted proposal rewrites the institution and carries forward. Points are earned principally through a turn-start dice roll (Rule 202), and a game terminates when a player reaches the target score (Rule 112) or after 50 turns. A dedicated validator rejects malformed or constitutionally illegal proposals before voting, so that the social voting stage sees only executable and constitutionally admis- sible proposals, while the Judge (Rule 212) classifies residual failures (Section 6.1). Because the question of interest is not which model “plays best” but which model sus- tains the most varied endogenous rule dynamics, we report process-level evidence: how often proposals survive validation and voting, how the active rule set evolves, what kinds of changes are adopted, and whether games resolve into winners. Both families are evaluated under matched conditions at five parameter scales with 15 independent trials per condition (five homogeneous players per game, 50-turn maximum), yielding a maximum of 750 turn-level observations per model; the same Nomic engine, prompt templates, and analysis pipeline are used throughout, and only the model checkpoint and its tokenizer change between families (Section 6.3). 2.1 Cross-family phase structure A first measurement establishes the baseline phenomenon: how each scale within each family performs on the components of institutional change that are simultaneously needed to sustain varied rule dynamics. We operationalise varied rule dynamics as the simultaneous presence of four directly observable properties: (i) a non-trivial fraction of turns end in adoption, (i) the number of active rules is visibly restructured over the game, (i) adopted changes cover all four move types rather than collapsing onto 3 Fig. 1 Cross-family phase portrait. Top row: Qwen3.5 (0.8B–27B); bottom row: Gemma3 (270M– 27B). Four panels per family: (A) adoption rate per size, (B) rule count delta from baseline, (C) change-type entropy, (D) winner rate. The sweet spot (highlighted) shifts from 4B in Qwen3.5 to 12B in Gemma3, while the non-monotonic shape (frozen at small scales, varied at mid-scale, veto- heavy or narrow at large scales) is preserved across families. a single type, and (iv) at least some games terminate with a winner before the 50- turn cap. These criteria capture survival, growth, diversity, and closure, the axes along which a society can adapt collectively without collapsing into gridlock or single-type parameter retuning. The phase portrait in Figure 1 arranges these four metrics side by side for both families. Within Qwen3.5, only the 4B scale satisfies all four criteria. The two smallest mod- els, 0.8B and 2B, are almost entirely rule-inert: for 0.8B, 12.4% of turns end in Judge invalidation and 87.3% in vote rejection, and for 2B only 0.3% of turns end in adoption despite nearly all proposals surviving validation, indicating that executable correct- ness alone is insufficient to generate institutional change. At 4B the adoption fraction reaches 33.8%, the mean active-rule count rises from 29 to roughly 31, the change-type mix is broad (14.3% ADD, 32.7% AMEND, 2.2% REPEAL, 50.7% TRANSMUTE), and winners emerge in 8 of 15 trials. By contrast, the 9B scale remains rejection- dominated (adoption fraction 8.0%) with zero winners, and 27B shows higher adoption (25.6%) but a narrow, amendment-heavy repertoire: 86.2% of adopted changes are AMEND operations, the active-rule count stays flat, and only 1 of 15 games resolves with a winner (and that lone victory is achieved by amending Rule 208 to lower target points rather than by accumulating points under the original threshold). The contrast between 4B and 27B is therefore one of broad versus narrow institutional change rather than of high versus low activity. Gemma3 reproduces the same phase structure with the sweet spot shifted to 12B. The 270M model occupies a frozen-with-amendments regime: only 1.2% of all turns end in adoption because 89.2% of turns are rejected by the validator before reaching a vote, and although 11.1% of the few proposals that survive validation are adopted the rule set remains effectively static. The 1B and 4B Gemma3 models are rule-inert 4 (adoption rates of 0.0% and 0.1%, respectively), mirroring the Qwen3.5 0.8B/2B frozen regime. Only at 12B do varied rule dynamics emerge: the baseline adoption rate is 53.9%, winners are produced in 3 of 15 trials, the mean final team score is +17.5 against−64 for the remaining scales, and the vote-YES rate of 61.6% indicates roughly balanced voting. At 27B the regime reverts to veto-heavy play (1.9% adoption, ∼95% NO votes), reproducing the narrow-repertoire pattern observed at Qwen3.5 9B/27B. The non-transferability of the sweet-spot regime to mixed-size societies holds in both families. A Qwen3.5 society with one player at each scale (0.8B, 2B, 4B, 9B, 27B) collapses into gridlock with an adoption rate of 5.9% (44/750 proposals across 15 trials) and zero winners, with the 9B player voting NO on 88.5% (617/697) of valid propos- als. A matching Gemma3 society (270M, 1B, 4B, 12B, 27B) shows an even stronger collapse: adoption is 0.0% across 15 trials with zero winners, and the 27B member votes NO on 99.5% (607/610) of valid proposals. Mixed capability therefore does not average behaviour, and vetoes from larger members suppress rule adoption even when a sweet-spot member is present, with the effect being more severe in Gemma3 than in Qwen3.5. Taken together, the sweet-spot location differs across families (4B in Qwen3.5 ver- sus 12B in Gemma3), but the qualitative pattern does not: one mid-scale combines high adoption, diverse change types, winner emergence, and balanced voting, while both smaller and larger scales fail on different components. The shift in location reflects differences in architecture, tokenizer, and training data, and the implications of that boundary condition are taken up in the Discussion (Section 4). 2.2 Robustness to sampling and voting-rule perturbations A central question for any cross-scale claim is whether the non-monotonic shape of Section 2.1 survives changes in sampling stochasticity and in the voting institution itself, because either factor could in principle manufacture the regime distinctions. We address two confounds with independent perturbations. The first is a sampling- side confound: identical nominal temperatures may produce qualitatively different output distributions at different scales because logit sharpness varies with scale. The second is an institutional-side confound: the default unanimity-with-auto-relax voting schedule (Rule 203 transitions from unanimity to simple majority at turn 9) might itself fabricate the regime distinctions, so we contrast it directly with a from-the- start majority rule (majority fromstart) that removes the early-game unanimity phase. This baseline-versus-from-the-start contrast is the central robustness test of the paper: if regime classes were an artefact of unanimity, replacing unanimity with simple majority from turn 0 should rearrange them, whereas if regime classes are a model-side property the contrast should leave them intact. Both perturbations leave the cross-scale ordering of regimes unchanged in both families, which we now report in detail. A homogeneous temperature sweep on a shared 0.1–1.3 grid (step 0.2, six tri- als per cell) tests the sampling confound, with results summarised in Figure 2. In Qwen3.5, response curves are strongly scale-dependent (4B performs best in the lower- temperature regime, 27B peaks at intermediate temperatures, and 9B improves mainly toward the high-temperature end), but the cross-scale ordering is preserved at every 5 Fig. 2 Cross-family temperature sweep. Left column: Qwen3.5; right column: Gemma3. Top row: adoption rate per size across temperatures T = 0.1–1.3; bottom row: winner rate on the same grid (replacing the earlier inset that overlapped with the main panel). In Qwen3.5, response curves are scale-dependent but preserve the cross-scale ordering. In Gemma3, the 12B sweet spot strengthens with temperature: the winner rate rises from 0/6 at T≤ 0.5 to 4/6 (67%) at T≥ 1.1 and the mean final team score climbs to +237, while all other sizes remain frozen. Temperature thus shifts operating points in both families without erasing the cross-scale contrast. temperature. In Gemma3, the 12B sweet spot not only persists but strengthens: its winner rate rises from 0/6 at T ≤ 0.5 to 4/6 (67%) at T ≥ 1.1, and its mean final team score reaches +250, while every other Gemma3 size remains frozen across the entire grid. The positive temperature–performance relationship at the Gemma3 sweet spot is absent in Qwen3.5, suggesting that higher stochasticity enhances collective coordi- nation specifically at that scale. In neither family, however, does temperature erase the cross-scale contrast in institutional style. The institutional confound is addressed by the rule-level intervention majority fromstart, which rewrites Rule 203 from the beginning of play so that proposals pass by simple majority, applied across all scales in both families (15 trials each). The perturbation preserves the cross-scale ordering even though specific end- points shift (Figure 3). In Qwen3.5, the 4B endpoint active-rule count is essentially unchanged (mean 29.6→ 30.4, +2.7%) and the mutable-rule fraction shifts by +2.9%, both inside baseline trial-to-trial variability; per-round proposal acceptance rises by approximately eight percentage points during the unanimity-free phase, but the winner rate drops from 8/15 (baseline) to 1/15 because aggressive removal of the unanim- ity barrier lets early-game over-amendment of target points disrupt convergence. At Qwen3.5 27B, the looser threshold also unlocks a small set of target-amendment 6 Fig. 3 Cross-family intervention robustness. Top row: active-rule count per round (solid: baseline; dashed: majorityfromstart); bottom row: ten-round rolling acceptance rate. Left: Qwen3.5; right: Gemma3. In both families, the sweet-spot models (4B and 12B, respectively) retain their endpoint composition under intervention, and frozen scales remain inert regardless of voting threshold. The intervention elevates acceptance rate during the unanimity-free phase without producing a regime change. victories (5/15) without altering the underlying narrow-AMEND repertoire, an out- come consistent with the family’s veto-heavy dynamics rather than a regime change. In Gemma3, the 12B model maintains roughly 50% adoption (against 53.9% at base- line) and its winner rate increases from 3/15 to 6/15, so the sweet spot becomes more, not less, productive once unanimity is removed. No scale in either family transitions out of its baseline regime, and the 0.8B (Qwen3.5) and 1B (Gemma3) models remain effectively inert under both conditions, consistent with the observation that changing the voting threshold has little effect when the model rarely produces adopted changes. The intervention does, however, change internal processing trajectories in scale- and family-specific ways that are analysed in Section 3. The baseline-versus-majority fromstart contrast is therefore the key robust- ness result of this study: the rule that auto-relaxes unanimity to majority at turn 9 7 (baseline) and the rule that applies majority from turn 0 (majorityfromstart) pro- duce within-scale endpoint differences that sit inside trial-to-trial variability at every scale in both families. Together with the temperature-grid invariance, this means that the non-monotonic phase structure of Section 2.1 cannot be attributed to sampling stochasticity or to a specific voting threshold: it is a model-side property that motivates the trajectory analysis next and the mechanistic analysis in Section 3. 2.3 Rule trajectories over time The phase-structure and robustness analyses above characterise each game by endpoint statistics, leaving open whether scales also differ in the temporal path of institu- tional change. Per-turn trajectories of the patched parameters (Figure C3) and of the active-rule composition (Figure C4, both in Appendix C) show that the regimes differ qualitatively in dynamics and not only in summary statistics. Rule 206 parameter trajectories separate sweet-spot stabilisation from large- scale oscillation. At the Qwen3.5 sweet spot (4B), the patchable parameters penaltyfordefeat and pointsfordissent are revised in a small number of large adopted moves during the first ∼15 turns and then remain near a fixed working point for the rest of the game, a pattern of directed exploration followed by stabilisation. At 27B the same parameters instead cycle through repeatedly visited values (10, 5, 1, 0, 2, 1, 10, 0, 10, 5, 10, 5 in a representative trial), producing the high reversal rate (∼79%) that defines the oscillatory regime (Appendix A). The two trajectories therefore cor- respond to different collective control problems: convergence on a working institution at the sweet spot versus chronic revisitation of prior settings at the largest scale. The active-rule composition tells the same story from a complementary angle. At the frozen scales, the mutable and immutable counts remain flat at the initial 13/16 split. At the sweet spot, the mutable count rises gradually as new mutable rules are added or as immutable rules are transmuted into mutable form, while the immutable count stays close to its initial value, a controlled expansion of the legislative space rather than a rewrite of constitutional foundations. At 27B, the active-rule count stays close to baseline because nearly all adopted changes are AMEND operations on exist- ing parameters, so the constitutional structure is touched least at the largest scale even though the parameter trajectories oscillate the most. Gemma3 reproduces this trajectory taxonomy at a shifted location: 12B exhibits the controlled-expansion motif of Qwen3.5 4B, whereas 27B reverts to a frozen, near-baseline trajectory. The trajec- tory view therefore turns the regime taxonomy from static labelling into a dynamic statement: the sweet spot does not merely produce a different end state but a differ- ent path through rule space. These dynamic motifs suggest correspondingly distinct internal integration dynamics, which Section 3 tests directly. 3 Mechanistic Analysis The behavioural and trajectory evidence above is consistent with a model-level prop- erty of collective adaptation, but behaviour alone cannot identify whether different scales implement similar outcomes through different internal computations. To exam- ine this, we replay the exact baseline and majorityfromstart prompts through 8 HuggingFace Transformers checkpoints of the same Qwen3.5 and Gemma3 models used in the behavioural experiments and capture layerwise hidden states for each matched prompt pair (Section 6.4). We treat linear probing as the primary mechanis- tic measurement (Section 3.1) because it directly quantifies how much vote-relevant information is recoverable from hidden states, and we use cosine divergence profiles (Section 3.2) as a supporting dissociation analysis describing how strongly each model is internally perturbed by the rule intervention. A cross-family pattern should appear as a representational signature observed at Qwen3.5 4B and Gemma3 12B but absent at failing scales in both families; we present the probe evidence first because it is the strongest piece of mechanistic evidence in this study, and then return to the divergence profiles to clarify what they do, and do not, contribute to the interpretation. 3.1 Vote-prediction probing reveals regime-selective decodability Linear probing asks how much task-relevant information is recoverable from each hidden layer. We train per-layer logistic-regression probes for two targets: condition (baseline vs. majorityfromstart) and observed vote (whether the observer voted YES). Both probes are summarised in Figure 4; the condition probe characterises encoding of the perturbation, while the vote probe is the main test of behavioural coupling. The condition probe indicates that the institutional perturbation is linearly decod- able regardless of scale or family. In Qwen3.5, accuracy stays at or above 0.97 at every intermediate layer for all sizes except 27B, which shows late-layer degradation to 0.89 in its last quintile; segment-pooled accuracy remains 1.00 in the rule-text segment, indicating that intervention information remains encoded even when terminal-token concentration declines. In Gemma3, condition classification is near-perfect for the 4B, 12B, and 27B scales, while 270M (best 0.972, mid-layer minimum 0.889) and 1B (best 0.980, mid-layer minimum 0.900) remain consistently above 0.88 even though they fall short of the larger-scale ceiling. The vote probe is the central observation of this section. At Qwen3.5 0.8B and 2B, votes are dominated by unanimous rejection, leaving a majority-class baseline of ∼95.1% that the probe does not exceed at any layer; class skew here precludes any meaningful decodability lift rather than implying that no vote-relevant signal exists. At Qwen3.5 9B, the majority-class baseline is 90.2% (still strongly NO-biased), and the best probe accuracy of 92.7% at layer 3 produces a margin of only +2.5 p, again limited by class skew rather than by absence of representation. At Qwen3.5 4B, where votes are approximately balanced (50.0% YES baseline), the probe reaches 67.1% accuracy (AUC 0.72) at layer 7, a margin of +17.1 p. At Qwen3.5 27B, the probe reaches 75.8% at layer 2 over a class-skewed (56.1%) baseline, a +19.7 p lift that marginally exceeds the 4B margin but appears at the shallowest layer in the network. In Gemma3, the 1B and 4B models produce essentially uniform NO votes with no variance to decode, the 270M baseline is also class-skewed (83.3%) and the probe never exceeds it, and the 27B probe accumulates only +3.7 p at layer 1 over a 95.1% NO-dominated baseline; in contrast, at 12B (65.9% YES baseline) the probe reaches 84.2% at layer 15, a margin of +18.3 p. 9 Fig. 4 Cross-family linear-probe selectivity. Top row: condition probe accuracy (baseline vs. majority fromstart); bottom row: vote-prediction probe accuracy versus the majority-class baseline. Left: Qwen3.5; right: Gemma3. The condition probe is near-perfect at all scales in both families (with the exception of Gemma3 270M and 1B, which dip to 0.89–0.90 at intermediate layers), indicating that the institutional perturbation is encoded almost everywhere. The vote probe shows the largest mar- gins at the family-specific sweet spots: +17.1 p at Qwen3.5 4B (layer 7, balanced 50% baseline) and +18.3 p at Gemma3 12B (layer 15, 65.9% YES baseline). Qwen3.5 27B reaches +19.7 p at layer 2 over a class-skewed (56.1%) baseline under veto-dominant dynamics (1/15 winner, secured only by amending target points); Gemma3 27B accumulates +3.7 p at layer 1 over a 95.1% NO-dominated baseline; other scales show no decodable vote signal above the majority-class baseline (Table 1). The two sweet-spot margins are near-identical (+17–18 p) despite differing in absolute parameter count and tokenizer, and they appear at intermediate-to-mid layers (layer 7 of approximately 36 in Qwen3.5 4B; layer 15 of approximately 48 in Gemma3 12B). Together with the empty vote-margin cells at strongly class-skewed scales, this is consistent with mid-layer linear decodability of vote outcomes being a regime-level rather than architecture-level property: balanced voting leaves a mid-layer decodable trace, whereas near-unanimous rejection provides essentially no decoding headroom. Decodability is therefore necessary but not sufficient. Qwen3.5 27B reaches +19.7 p vote-probe accuracy at layer 2, slightly exceeding the 4B margin, yet pro- duces only one win out of 15 trials (achieved by amending target points downward 10 Table 1 Cross-family vote-prediction probe: best layer-wise accuracy vs. majority-class baseline. The near-identical margins at the two family-specific sweet spots (+17–18 p) appear at intermediate-to-mid layers (Qwen3.5 4B: layer 7; Gemma3 12B: layer 15), whereas the comparable +19.7 p margin at Qwen3.5 27B appears at the shallowest layer (layer 2). Scales with strongly class-skewed votes (Qwen3.5 0.8B/2B/9B; Gemma3 270M/1B/4B) provide little decoding headroom regardless of underlying representation. Family Model Best Acc. Baseline∆Best Layer Qwen3.50.8B95.1%95.1%0.0 ppn/a Qwen3.52B95.1%95.1%0.0 ppn/a Qwen3.54B67.1%50.0%+17.1 p7 Qwen3.59B92.7%90.2%+2.5 p3 Qwen3.527B75.8%56.1%+19.7 p2 Gemma3270M83.3%83.3%0.0 ppn/a Gemma31B—all NO—n/a Gemma34B—all NO—n/a Gemma312B84.2%65.9%+18.3 p15 Gemma327B98.8%95.1%+3.7 p1 rather than by accumulating points) and votes NO on roughly 95% of valid propos- als. Three features distinguish this signal from the sweet-spot pattern: it peaks at the shallowest layer (consistent with prompt-level priors rather than integrated reasoning over context), it sits over a class-skewed (56.1%) rather than balanced baseline, and it co-occurs with veto-dominant rather than deliberative behaviour. We summarise the sweet-spot signature as behavioural coupling of vote-predictive information, that is, the joint presence of (i) above-chance decodability, (i) representational depth at which that information appears, and (i) downstream behavioural diversity; the probe lift at Qwen3.5 27B satisfies (i) but not (i) or (i) and is therefore real but uncoupled from collective outcome. 3.2 Hidden-state divergence: a supporting dissociation, not a behavioural predictor Whereas the vote probe asks how much vote-relevant information is decodable, cosine divergence asks how strongly the same representations are displaced by the institu- tional perturbation. Magnitude and decodability are distinct quantities, and our main use of the divergence analysis is to establish a dissociation: large internal response to the intervention does not co-occur with adaptive behaviour. Figure 5 plots the mean cosine distance between matched baseline and intervention last-token representations against relative layer depth. In Qwen3.5, the ordering of scale-mean divergence (27B: 0.031 > 4B: 0.017 > 9B: 0.011 > 0.8B: 0.005 > 2B: 0.003) does not match the behavioural ranking (4B ≫ 27B, 9B, 2B, 0.8B), and the 27B curve in particular shows a late-layer surge peaking at 0.148 in the last quintile while 4B sustains a more modest profile (roughly 0.015– 0.025) from relative depth ∼0.15 onward. We refer to this shape descriptively as a 11 Fig. 5 Cross-family hidden-state divergence profiles. Left: Qwen3.5; right: Gemma3. Cosine dis- tance between matched baseline and majorityfromstart last-token representations, plotted against relative layer depth. The Gemma3 1B curve is highlighted with a square marker to flag its anoma- lous peak. In Qwen3.5, the divergence ordering (27B > 4B > 9B > 0.8B > 2B) does not match the behavioural ranking, and what distinguishes 4B from its siblings is a mid-layer plateau rather than peak magnitude. In Gemma3, the marked 1B curve has the largest mean divergence (0.0035) yet yields the worst behavioural outcome (0% adoption, 0/15 winners), the cleanest single case of magnitude–behaviour dissociation. mid-layer plateau: the divergence signal at 4B rises by relative depth 0.15 and does not decay before the final layers, whereas 2B’s signal dissipates within the first 60% of layers, 0.8B’s accumulates only toward the final layers, and 27B’s is concentrated in the late-layer surge. This shape is qualitative; we do not claim a formal metric of mid-layer plateau strength here, and so the divergence evidence supports rather than independently establishes the regime taxonomy. Gemma3 provides the cleanest dissociation between divergence magnitude and behaviour. The 1B model has the largest mean cosine distance (0.0035) across lay- ers yet produces the worst behavioural outcome (0% adoption, 0/15 winners), while the 12B sweet spot shows a more moderate profile that nonetheless mirrors the mid- layer plateau seen at Qwen3.5 4B. In both families, raw divergence magnitude is not informative about behavioural outcome and serves only to rule out a simple “larger internal response = better collective adaptation” interpretation; the directional evi- dence that internal computations differ at the sweet spot comes from the probe results in Section 3.1. 4 Discussion Behavioural, trajectory, and mechanistic evidence are consistent with a single regime structure for collective adaptation in LLM societies. The behavioural layer (Section 2) showed that, in Qwen3.5 and in Gemma3, only one mid-scale per family sustains var- ied rule dynamics, and that this pattern survives both temperature and voting-rule 12 perturbations. The trajectory layer (Section 2.3) showed that the regimes differ not only in summary statistics but in the shape of the path through rule space: directed exploration followed by stabilisation at the sweet spot, repeated cycling at the largest scale, and frozen baseline at the smallest scales. The mechanistic layer (Section 3) showed that the sweet-spot regime is associated with linearly decodable vote-predictive information at intermediate-to-mid layers and with balanced downstream voting (the behavioural-coupling pattern); a complementary, qualitative observation is that the layerwise cosine divergence profile at the sweet spot exhibits a sustained mid-layer plateau rather than a late-layer surge, in contrast to the magnitude-dominated pro- file at the largest scales. Taken together, the three layers describe the same regime from three angles, but the strongest mechanistic evidence is the probe-decodability pattern; the divergence profile is a supporting, descriptive observation rather than an independent predictor of behaviour. This convergence is consistent with read- ing the non-monotonic pattern as a model-level property rather than as a one-scale behavioural coincidence, while leaving causal direction (whether mid-layer probe- decodable structure produces balanced voting or arises from it during inference) unresolved by the current analyses. The result therefore places collective adaptation alongside the growing catalogue of capabilities that do not scale smoothly with parameter count [53–55], contradicting the assumption that more capable models necessarily produce more capable institutions. The sweet-spot regime is a regime-level rather than a metric-level phenomenon. Each smaller or larger scale fails a different component of varied rule dynamics: Qwen3.5 2B produces valid proposals that never pass a vote, Gemma3 1B/4B are simi- larly rule-inert, larger models in both families veto heavily or accumulate score through narrow amendment-only repertoires, and heterogeneous societies in both families col- lapse into asymmetric distrust where vetoes from the largest members override any sweet-spot member that is present. No single scalar metric of overall quality captures this combination, in line with Ostrom’s observation that governing a commons requires a combination of properties (monitoring, sanctions, recognised authority to amend) rather than any one of them in isolation [4]. The hidden-state analysis is consistent with the same reading at the mechanistic level: in both families, the largest internal response to institutional perturbation does not co-occur with adaptive behaviour, and the family-specific sweet spot is the scale at which vote-predictive information becomes linearly decodable at intermediate-to-mid layers (+17 p at Qwen3.5 4B, +18 p at Gemma3 12B), consistent with probing work that finds task-structure representations forming early and propagating through the network [68, 69, 74–76]. In addition to this regime-level reading, the heterogeneous collapse clarifies what kind of object collective adaptation is in these societies. One might expect a mixed- capability group to inherit at least a fraction of the sweet-spot regime, since the sweet-spot model is one of its members, but in practice veto-biased players suppress the rule dynamics that the homogeneous sweet-spot society sustains. Institutional capacity for varied rule dynamics is therefore a property of the society rather than of its most capable member, in line with prior reports on LLM law-making [34] and collective constitutional alignment [35] that agent homogeneity materially affects deliberation outcomes. 13 That the sweet-spot location differs across families (4B in Qwen3.5 versus 12B in Gemma3) is best read as a boundary condition rather than a contradiction of non-monotonicity. Architecture, tokenizer, and training data all differ between the two families, so the absolute parameter count at which proposal quality and social coordination jointly emerge need not be conserved. What is conserved is the qualitative shape: a single mid-scale regime that satisfies all four richness criteria while both tails fail, the same failure-mode taxonomy (frozen, veto-heavy, narrow-AMEND), the same heterogeneous-society collapse driven by the largest member’s veto bias, and near-identical mechanistic signatures (+17–18 p vote-prediction probe margins). The shared object across families is therefore the existence of a non-monotonic sweet spot and its representational signature, not the parameter count at which it occurs. A candidate account of this pattern is a balance between representational capacity and behavioural priors interacting with the multi-agent self-play setting. Very small models do not exhibit the layer depth associated with cross-layer representation of institutional context and consequently produce rule-inert or near-random play; the Gemma3 270M anomaly (higher adoption than 1B/4B) is consistent with this, in that very small models can also be behaviourally permissive simply because they have not yet acquired structured rejection priors. Mid-scale models reach sufficient depth for vote-relevant information to become linearly decodable at intermediate-to-mid layers and for downstream voting to be balanced, producing the behavioural-coupling pattern observed at Qwen3.5 4B and Gemma3 12B alongside the qualitatively sustained mid- layer divergence plateau. Large models continue to encode vote-relevant information at the layer level (the +19.7 p probe margin at Qwen3.5 27B), but the decodable signal there appears at the shallowest layer and is observed alongside conservative or veto-dominant policies, plausibly shaped by alignment training on legislative and governance text and by self-play with copies that share the same priors, producing a coordination lock-in onto near-unanimous rejection. This account remains tentative: a direct test would compare base, instruction-tuned, and RLHF variants of the same checkpoint at the sweet-spot scale and would manipulate the social environment to break self-play prior correlation, both of which we leave to future work. The methodological implication is that self-amending games can serve as controlled model systems for studying collective adaptation in LLM societies, in the same spirit that social scientists increasingly use generative agents to study social processes [10–14, 77]. The setting complements recent work on the formal structure of collective decision rules in artificial agents [78] and on governance as a satisfiability problem in networked communities [79], and it operationalises the exploration–exploitation tension central to organisational adaptation [80]. Nomic provides a closed, instrumented setting in which institutional rules can be perturbed directly, the response can be observed in both behaviour and internal representations, and the pipeline keeps executable invalidity separate from social rejection so that collective deliberation can be attributed to the latter. The claim is not that Nomic is a literal proxy for legislatures or DAOs but that rule-endogenous LLM simulations can expose qualitative differences in institutional style that fixed-rule benchmarks [15–17, 31] cannot see. 14 Several limitations bound the scope of this claim, and three methodological caveats deserve emphasis up front rather than as footnotes. First, every homogeneous five- player society is composed of copies of the same checkpoint at the same scale, so proposer and voters share inductive biases by construction; this self-play confound means that part of the coordination structure we observe (sustained adoption at the sweet spot, veto lock-in at large scales) may reflect correlated priors rather than scale per se, and a direct test would replicate each condition with non-self-play populations (e.g., distinct seeds, sibling fine-tunes, or cross-checkpoint mixtures) of comparable mean capability. Second, the proposal pipeline contains a prompt–engine misalign- ment risk: the natural-language proposal text and the structured ENGINEPATCH block can disagree (e.g., the text claims a one-point penalty while the patch sets ten; see the 27B example in Appendix A), and because voters read the natural-language text while the engine executes only the patch, agents can deliberate over one game while outcomes are scored on another; the validator catches malformed patches but cannot in general detect semantic divergence between text and patch, so the gap between deliberated and executed institutions is an open methodological risk rather than a resolved control. Third, our “heterogeneous” condition is a single fixed composition (one player per scale within a family) replicated 15 times rather than a parameterised study of heterogeneity degree, seat-power asymmetry, or cross-family mixtures, so the heterogeneous collapse is best read as a counterfactual to homogeneous self-play and not as a general statement about diversity and collective intelligence. With n = 15 trials, five model scales, two families, and several derived metrics, the risk of spuri- ous cross-scale contrasts from multiple comparisons is non-negligible, so we emphasise effect sizes (e.g., +17–18 p probe margins, 8/15 vs. 0–3/15 winner rates) over signifi- cance testing throughout, and read the cross-family replication of the non-monotonic shape as the main protection against single-family artefacts. Beyond these method- ological caveats, the study spans two model families, and whether the non-monotonic pattern extends to further architectures and training pipelines remains an open ques- tion; mechanistic conclusions are local to the single majorityfromstart intervention on Rule 203 and to the matched trial pairs that survive prompt-identity alignment (n = 15 per scale except Qwen3.5 27B, n = 14), and cross-scale comparisons remain conditional on the sampling regime because decoding temperature shifts operating points. A natural control intervention that we did not run is a content-neutral per- turbation of Rule 203 (for example, rewording the unanimity clause without changing its semantics) which would test whether observed latent-state shifts reflect the rule change itself rather than text-level surface variation; we flag this as the most infor- mative single experiment for tightening the causal claim. Finally, this study measures collective adaptation functionally and does not claim that models “understand” the institutions they rewrite [67]; the phenomenon of interest is population-level dynamics, not any individual agent’s internal state. 5 Conclusion Collective adaptation in LLM societies is scale-dependent and non-monotonic across model families, and the same pattern appears at three levels of description. At the 15 behavioural level, across Qwen3.5 (0.8B–27B) and Gemma3 (270M–27B) tested at five parameter scales with 15 trials per condition in 465+ games and approximately 116,000 turn-level observations, only one mid-scale per family sustains varied rule dynamics (Qwen3.5 4B; Gemma3 12B), while smaller scales remain trapped in rule-inert play, larger scales narrow their institutional repertoire to parameter retuning, and hetero- geneous mixed-size societies in both families collapse into gridlock through veto-heavy voting by the largest member. At the trajectory level, the sweet spot shows directed exploration followed by stabilisation, the largest scales cycle indefinitely through pre- viously visited values, and the smallest scales never leave the initial constitution; the baseline-versus-majorityfromstart contrast leaves these regime classes unchanged in both families, and a temperature sweep over T = 0.1–1.3 does the same, so the non- monotonic shape cannot be attributed to a specific voting threshold or to sampling stochasticity. At the mechanistic level, the sweet-spot regime is characterised by vote- predictive information that is linearly decodable at intermediate-to-mid hidden layers and is observed together with balanced downstream voting (the behavioural-coupling pattern, +17 p at Qwen3.5 4B and +18 p at Gemma3 12B above majority-class baselines); a complementary qualitative observation is that the cosine divergence pro- file at the sweet spot shows a mid-layer plateau rather than a late-layer surge, but raw divergence magnitude is not informative on its own (Gemma3 1B has the largest mean divergence and 0/15 winners), and decodability is necessary but not sufficient, since Qwen3.5 27B reaches an even larger probe margin yet resolves only one trial with a winner and only by amending the victory-target downward rather than by accumu- lating points. The cross-family regularity is therefore the joint pattern across these three levels rather than the specific scale at which the sweet spot appears, and self- amending games offer a controlled testbed in which the capacity for varied collective adaptation does not track raw model scale. 6 Method This section details the Nomic implementation, LLM configuration, experimental protocol, and hidden-state capture pipeline used throughout the preceding Results, Mechanistic Analysis, and Discussion sections. 6.1 Nomic Game Implementation We implement a structured variant of Nomic [32] with 29 initial rules: 16 immutable rules (101–116) and 13 mutable rules (201–213). Immutable rules cannot be amended or repealed directly; they must first be transmuted to mutable status. On each turn, the acting player receives the current rules, effective mechanics, protocol state, and score vector, and produces exactly one proposal of type ADD, AMEND, REPEAL, or TRANSMUTE. The proposal parser normalises the raw LLM output into this move vocabulary and extracts an optional ENGINEPATCH block from an indented YAML-like format. The runtime pipeline is: proposal generation → validation → voting → application → post-vote effects → victory check (Figure 6). Validation is performed by a dedicated rule validator together with the rule engine, not by the Judge itself. This stage checks 16 Fig. 6 Per-turn judgement pipeline. A proposal generated by the acting player is first nor- malised by the parser into a typed move (ADD/AMEND/REPEAL/TRANSMUTE plus an optional ENGINE PATCH); malformed outputs short-circuit here. The validator then checks rule existence, immutability, and typed ENGINEPATCH bounds, routing executable invalidity away from the social stage. Only proposals that survive both steps reach the vote, and only adopted proposals are applied to the rule set under a transactional rollback. Rule 212 defines the Judge role at the declarative level but is not a separate runtime module: failure classification is performed in-line by the validator and the engine. rule existence, immutability constraints, typed ENGINE PATCH fields, numeric bounds, and constitutional invariants such as Rule 114 (at least one mutable rule must remain) and the immutability barrier of Rule 110, which indirectly protects Rule 112 (“victory must remain points-based”) because immutable rules cannot be amended or repealed until first transmuted. The executable parameters exposed to agents are the patch- able state fields of Rules 202, 203, 204, 206, 208, and 209, including dicesides and rerollon (Rule 202), unanimityrequired and majoritythreshold (Rule 203), points fordissent (Rule 204), penaltyfordefeat (Rule 206), targetpoints (Rule 208), and maxmutablerules (Rule 209). Proposals that fail validation are counted as invalid and do not proceed to the social voting stage. Throughout the paper, “valid” therefore means executable and constitutionally admissible, not strategically good. Rule 212 defines a Judge role in declarative rule text, but the implementation does not run a full judgment state machine: dispute resolution in the codebase is handled by the validator-plus-engine pipeline, which classifies residual failures (consti- tutional violations, malformed ENGINEPATCH, precedence conflicts, runtime rollback) into structured failure categories without invoking a separate Judge module. In prac- tice this means that the paper’s distinction between executable invalidity and social rejection is enforced by the validation stage: malformed or constitutionally illegal pro- posals are rejected before voting, and everything that reaches the voting stage has already been normalised into an executable move over the patchable fields above. Voting is governed by Rule 203. The default state requires unanimity, but if Rule 203 has not been amended, the engine automatically relaxes the threshold to simple majority after the second full circuit of turns (turn index 2N − 1, i.e., turn 9 in a five-player zero-indexed game). Once a proposal is adopted, the engine applies it transactionally with rollback on failure, then triggers rule hooks for dissent bonuses, rejection penalties, dice-based scoring, and victory checks. 6.2 LLM Configuration We use two model families under identical experimental conditions. The Qwen3.5 fam- ily [1] is tested at five parameter scales: 0.8B, 2B, 4B, 9B, and 27B. The Gemma3 17 family [2] (google/gemma-3-size-it) is tested at five scales: 270M, 1B, 4B, 12B, and 27B. Within each family, all models share the same architecture and training methodology, differing only in parameter count, which isolates scale as the indepen- dent variable. Our primary 15-trial comparisons deploy all models with temperature = 0.7, maximum output tokens = 2048, and thinking mode disabled. We note that identical temperature values may produce qualitatively different output distributions at different model sizes due to differences in logit distribution sharpness; this is a known confound in scaling studies that we address empirically with the temperature sweep reported in Section 2.2. To probe this confound, we additionally ran a supple- mentary homogeneous temperature sweep on a shared 0.1–1.3 grid (step 0.2, six trials per model-condition) for both families, analysed separately from the primary 15-trial results. 6.3 Experimental Protocol Each experimental condition consists of a five-player game running for a maximum of 50 turns (250 turn-level observations per trial), replicated across fifteen indepen- dent trials. This corresponds to a maximum of 750 turn-level observations per model (early termination reduces this in winner-producing conditions). All five players in each game instance are identical copies of the same model at the same parameter scale, ensuring that any asymmetries in behaviour emerge endogenously rather than from capability differences between players. This design also introduces a fundamental self-play confound in the homogeneous conditions: proposer and voter share the same inductive biases, since the same model copy evaluates its own proposals. The heteroge- neous condition, in which the five seats hold five different scales, is the counterfactual used in the main text to separate homogeneous self-play dynamics from cross-scale deliberation. Games terminate early if a player satisfies the victory condition or if all 50 turns are exhausted. Both families use the identical Nomic engine, prompt tem- plates, and analysis pipeline; the only change is the model checkpoint and its tokenizer. Across both families, the total experimental footprint comprises 465+ Nomic games (approximately 116,000 turn-level observations) plus 120 hidden-state replay trials. 6.4 Hidden-State Capture Protocol To probe internal representational dynamics, we replay the exact prompts from com- pleted game logs through HuggingFace Transformers checkpoints of the same Qwen3.5 and Gemma3 models used in the behavioural experiments. For each proposal or vote prompt, we reconstruct the full input text (including current rules, scores, and instructions), run a forward pass with outputhiddenstates=True, and extract two representations at every layer: (1) the last-token hidden state, and (2) mean-pooled hidden states over named prompt segments (e.g., the Rule 203 block, scores block, instruction block). Character-level spans are mapped to token spans through the tokenizer’s offset mapping. We compare two experimental conditions: baseline (default rules) and majorityfromstart (Rule 203 modified to simple majority voting from turn 0). For each trial pair (baseline trial N vs. intervention trial N ), we match examples by 18 the identity tuple (kind, turn, actingplayer, observerplayer) and compute per-layer cosine distances between matched hidden-state vectors. We report bootstrap 95% con- fidence intervals on trial-level mean distances. All models contribute 15 trial pairs except Qwen3.5 27B, for which one trial produced no matched prompt pairs and there- fore contributes 14 pairs. All captures use bfloat16 precision on NVIDIA GB10 GPUs (128 GB unified memory). Together, these procedural and technical choices define the pipeline whose behavioural outputs are reported in Section 2 and whose internal representations are dissected in Section 3. 6.5 Code and Data Availability A complete reproducibility package is publicly available at https://github.com/ KazuyaHoribe/nomic-coevolution under the MIT License. The repository bun- dles the Nomic engine (src/nomicsim/), behavioural experiment runners for the Qwen3.5 and Gemma3 families, the hidden-state replay and probing pipeline (analysis/activationcapture/), and figure-generation scripts that reproduce every figure in this paper from the included experimental logs in a single command. The bundled data covers all baseline, intervention, temperature-sweep, and heterogeneous- society trials reported here, as well as the aggregated per-layer hidden-state statistics and probe results that feed Figures 5 and 4; raw per-prompt activation tensors (approx- imately 14 GB) are not bundled, but are regenerable from the included game logs via analysis.activationcapture.multitrialrunner. Appendix A Example Proposals and Votes This appendix presents verbatim excerpts from game logs to illustrate the qualita- tive differences between behavioural regimes. Proposals and vote reasons are quoted directly from LLM output (lightly truncated for space). A.1 Frozen Regime: Self-Defeating Proposals (2B) Turn 5, Player 0: ADD Rule 301. Proposal: “The first player to achieve a score of at least 75 points shall win, provided the current winner has not already achieved the target of 100 points.” Reason: “Since my starting score is −7 and I am in the middle of the game, reaching 75 is much more achievable than 100.” Votes: 0 AGREE / 5 DISAGREE. Player 0 (the proposer) votes DISAGREE: “The proposed rule change to lower the win threshold from 100 to 75 creates a massive incentive imbalance in this specific game state where I am at −7 points.” The proposer generates a strategically sensible rule (lower the victory threshold to ben- efit their losing position) but then votes against their own proposal. This self-defeating behaviour, the inability to maintain consistent intent between proposal generation and voting, is characteristic of the Frozen regime and explains why no proposals are ever adopted. 19 A.2 Frozen Regime: Confused Rule Reference (0.8B) Turn 2, Player 2: TRANSMUTE Rule 102. Proposal: “TRANSMUTE Rule 102 from mutable to immutable.” Reason: “By making these specific rules immutable from the start, you prevent Rule 112 from allowing arbitrary changes in point magnitude.” Votes: 0 AGREE / 5 DISAGREE. Player 1: “Player 2 cannot vote against a proposal that is already passed.” Player 2 (proposer): “The proposal to change the nature of the game from mutable to immutable is fundamentally nonsensical.” Rule 102 is already immutable; the transmutation direction is confused. Vote reasons reference nonexistent events and misidentify players, indicating that the 0.8B model hallucinates game context rather than reasoning from the actual state. A.3 Exploratory Regime: Strategic Parameter Change (4B) Turn 15, Player 0: AMEND Rule 206, penalty fordefeat: 10→ 0. Proposal: “When a proposed rule-change is defeated, the player who proposed it loses 0 points.” ENGINE PATCH: penaltyfordefeat: 0 Reason: “I am currently at −24 points. Every time I propose a rule that fails, I lose additional points (−10), while my opponents gain +10 bonus points. By amending Rule 206 to remove the penalty for failed proposals, I can stop losing points on failed votes.” Votes: 4 AGREE / 1 DISAGREE. Adopted. The proposer correctly identifies their losing position, formulates a self-serving amend- ment with a properly structured ENGINE PATCH, and coordinates proposal intent with voting behaviour. This combination of situational awareness, strategic reasoning, and mechanical grounding is absent in the Frozen regime. A.4 Exploratory Regime: Rule Innovation (4B) Turn 44, Player 4: ADD Rule 301. Proposal: “Every player who has fewer than 25 points will automatically receive +10 bonus points at the end of every turn, regardless of their vote on any proposed rule change.” Reason: “Player 4 is at 5 points and Player 1 is at−7 points; they are likely to be forced to support your proposals. By guaranteeing them a positive outcome, you will secure their votes for future amendments.” Votes: 5 AGREE / 0 DISAGREE. Adopted. A novel welfare mechanism that has no precedent in the initial rules. Although this ADD rule has no mechanical effect (it is stored as text only; see Section 6.1), its adoption indicates that the 4B model can introduce text-level rule innovations that go beyond retuning existing parameters, and that the deliberative process accepts such novel rule text. A.5 Oscillatory Regime: Parameter Zigzag (27B) The 27B model exhibits extreme oscillation on Rule 206. Below are four consecutive adopted amendments from a single trial: 20 Turn 22 (Player 2): AMEND Rule 206. Patch: penaltyfordefeat: 1. “The current penalty of losing 2 points for a failed proposal creates significant risk. By reducing this to 1 point, I can propose more aggressively.” Adopted (3–2). Turn 23 (Player 3): AMEND Rule 206. Patch: penalty fordefeat: 10. “Rule 206 currently imposes a −1 penalty, but the game instructions explicitly state −10. To align with the core game mechanics...” Rule text says “loses only 1 point” but the patch sets it to 10, a direct text-mechanics contradiction. Adopted (4–1). Turn 24 (Player 4): AMEND Rule 206. Patch: penalty fordefeat: 0. “I am currently at only 3 points. A failed proposal would drop me to −7. To be safe, I want to minimise the penalty.” Adopted (3–2). Turn 25 (Player 0): AMEND Rule 206. Patch: penalty fordefeat: 10. “The current rules state that a failed proposal causes a loss of 1 point, but the game instructions define the penalty as 10 points. This amendment aligns the rule text with the actual game mechanics.” Adopted (3–2). The full penalty fordefeat trajectory across this trial is: 10 → 5 → 1 → 0 → 2 → 1 → 10 → 0 → 10 → 5 → 10 → 5. Each change is locally motivated but globally cancels its predecessor, producing the high reversal rate (79.4%) that defines the Oscillatory regime. Turn 23 is particularly notable: the model’s stated text and its ENGINEPATCH contradict each other, yet the proposal is adopted, illustrating that voters evaluate the natural-language description rather than the structured patch. Appendix B Rule-Dynamics Supplementary Figures The phase portrait in Section 2.1 summarises outcomes per scale; the two figures collected here resolve those outcomes along two further axes, namely which rules are targeted and how votes are distributed, so that the cross-family pattern can be inspected at finer granularity. Figure B1 reports the frequency with which each initial rule is targeted by valid proposals at every scale. Smaller models in both families distribute attention across many rules in a diffuse exploration pattern, whereas larger models concentrate on a small subset, indicating focused parameter adjustment rather than constitutional restructuring. The collapse of targeting entropy with scale therefore replicates across families, even though the specific rules attracting proposals differ between Qwen3.5 and Gemma3 because of differences in tokenizer and prompt-processing behaviour. The vote-unanimity decomposition in Figure B2 separates each adopted vote into unanimous-approval, split, and unanimous-rejection bins. Only at the family-specific sweet spot (Qwen3.5 4B; Gemma3 12B) does the count of unanimous approvals mean- ingfully exceed that of unanimous rejections, while smaller and larger scales remain rejection-dominated in both families. Together with the heatmap above, this confirms that the qualitative shift in collective-decision regime, namely broad targeting com- bined with balanced approval, occurs only at the sweet spot and is the source of the divergent behavioural endpoints reported in the main text. 21 Fig. B1 Frequency with which each initial rule is targeted by valid proposals, per model scale (top: Qwen3.5; bottom: Gemma3). Both families show diffuse exploration at small scales and concentration on a small rule subset at large scales. Appendix C Rule-Trajectory Figures Section 2.3 argues that the regimes differ in the temporal path of institutional change rather than only in summary statistics; the two figures here visualize that argument directly. Both panels use the canonical model-colour palette of the main text so that traces can be matched across figures. Figure C3 contrasts the patchable parameters of Rule 206 at the Qwen3.5 sweet spot with those at the largest scale. The 4B trajectory shows a small number of large adopted moves during the first∼15 turns and then settles near a working point, while the 27B trajectory cycles repeatedly through previously visited values, producing the high reversal rate (∼79%) reported in Section 2.3. The two traces therefore correspond to qualitatively different collective control problems, namely convergence at the sweet spot versus chronic non-convergence at the largest scale, rather than to merely different summary statistics. Figure C4 adds the constitutional view by tracking the active-rule count split into mutable and immutable parts. At the frozen scales, both counts remain at the ini- tial 13/16 split. At the sweet spot the mutable count rises gradually as new mutable rules are added or as immutable rules are transmuted to mutable status, while the immutable count stays near its starting value, a controlled expansion of the legisla- tive space rather than a rewrite of constitutional foundations. At the largest scales, the count remains close to baseline because adopted changes concentrate on AMEND operations on existing fields rather than on ADD, REPEAL, or TRANSMUTE. Read 22 Fig. B2 Vote-unanimity decomposition per model scale (left: Qwen3.5; right: Gemma3). Unanimous- approval votes meaningfully exceed unanimous-rejection votes only at the family-specific sweet spot (Qwen3.5 4B; Gemma3 12B); other scales remain rejection-dominated. together with the parameter trajectories above, this view shows that the regime tax- onomy applies not only to which rules change but to how the constitutional structure itself evolves. Appendix D majorityfromstart Intervention: Per-Scale Numerical Detail This appendix lists the per-scale numerical breakdown of the majority fromstart intervention summarised in Section 2.2. The intervention covers all model scales in both families, 15 trials each, using the same Nomic engine. Under majorityfromstart, the Qwen3.5 4B endpoint active-rule count shifts from 29.6 to 30.4 (+2.7%) and the mutable-rule fraction from 0.306 to 0.316 (+2.9%), both inside baseline between-trial variability, and the ADD/AMEND/REPEAL/- TRANSMUTE mix is preserved. Per-round proposal acceptance rises from 0.356 to 0.437 during the unanimity-free phase, but the winner rate drops from 8/15 to 1/15: with unanimity removed from turn 0, premature targetpoints amendments are adopted before the player population stabilises, which destabilises convergence at the sweet spot. At Qwen3.5 27B, the looser threshold reveals five winners out of 15 (vs. 1/15 at baseline) entirely through aggressive amendment of Rule 208 to lower targetpoints (final values 20–53), consistent with the model’s narrow-AMEND repertoire rather than with a transition to varied rule dynamics. In Gemma3, the 12B model maintains roughly 50% adoption (49.9%) and its winner rate rises from 3/15 to 6/15 with targetpoints unchanged in every winning trial, so the sweet spot becomes more productive once unanimity is removed. The Qwen3.5 0.8B and Gemma3 1B models are inert under both conditions, with active rule counts pinned at 29 and adoption rates ≈ 0%, and Qwen3.5 2B closely tracks this inert pattern. The Gemma3 23 Fig. C3 Per-turn trajectories of two patchable fields of Rule 206 at the Qwen3.5 sweet spot (4B) versus the largest scale (27B). 4B exhibits directed exploration followed by stabilisation, whereas 27B cycles through repeatedly visited values (∼79% reversal rate). Fig. C4 Trajectory of the active-rule count, split into mutable and immutable parts, over the 50- turn game per scale. The sweet spot exhibits a controlled expansion of the mutable subset; frozen scales remain at the 13/16 baseline; larger scales also stay close to baseline because adopted changes are concentrated on AMEND operations. 270M model adopts only 1.2% of all turns at baseline and 1.9% under intervention (with 89–90% of turns rejected by the validator before reaching a vote in both con- ditions), so its constitutional structure remains static regardless of voting threshold. Consistent with Section 2.2, no scale in either family transitions out of its baseline regime under the intervention; the additional Qwen wins are mechanically valid but driven by the same narrow amendment behaviour as at baseline. Appendix E Metric Definitions This appendix defines the derived metrics used in the main text. 24 Reversal rate. For each model scale, we collect every adopted ENGINEPATCH across all trials and group the patched values by parameter name. A reversal occurs when a parameter is set to a value that it has previously held at any earlier point in the game (across all trials for that model scale). Formally, for each parameter p we maintain a set S p of values observed so far. When a new adopted patch sets p ← v: if v ∈ S p , the change is counted as a reversal; then v is added to S p regardless. The reversal rate is R/N × 100%, where R is the total number of reversals and N is the total number of adopted ENGINEPATCH changes. Shannon entropy of rule targeting (H). Let f i be the fraction of valid proposals that target rule i. The entropy is H = − P i f i log 2 f i , computed over all rules targeted at least once. Higher entropy indicates more diverse rule targeting. EP% (Engine Patch percentage). The fraction of valid proposals that include at least one ENGINEPATCH attribute, regardless of whether the proposal is adopted. Appendix F Prompt Templates This appendix reproduces the exact prompts used for rule-change proposal generation and voting. F.1 Proposal Prompt The system preamble is: You are Player agentindex in a game of Nomic. Nomic is a game where players take turns proposing rule changes, and all players vote on each proposal. You are a strategic agent trying to win the game by accumulating points. Key rules: - You may propose one rule change per turn: ADD, AMEND, REPEAL, or TRANSMUTE. - After proposing and voting, you roll one die and add the result to your score. - If your proposal fails, you lose 10 points (Rule 206). - Players who vote AGAINST a successful proposal gain 10 bonus points (Rule 204, only after unanimity is relaxed). - The first player to reach the target points (default 100) wins. Think strategically about which rule change will benefit you the most. Note: The prompt makes the dice-roll mechanism (Rule 202) and the conditional dissent bonus (Rule 204) explicit, so the asymmetric incentive between proposers (no direct bonus for a passing proposal) and dissenters (a +10 bonus once unanimity is relaxed) is part of the agents’ instruction set. 25 The user prompt provides the current game state and enforces a structured out- put format (including rule text, scores, and turn number) with templates for ADD, AMEND, REPEAL, and TRANSMUTE operations. A patchable attribute map is appended to inform agents which rule parameters can be mechanically modified through an ENGINEPATCH block. F.2 Voting Prompt The voting prompt provides the proposal details and enforces a binary vote (AGREE/DISAGREE) with a reason. Appendix G Cross-Family Supplementary Statistics This appendix provides per-condition statistics for both families. Table G1 Qwen3.5 baseline behavioural statistics (15 trials per size, temperature = 0.7). Size Adoption Winners Vote YES Mean ScoreRegime 0.8B0.3%0/156.8%—Frozen 2B0.3%0/156.0%—Frozen 4B33.8%8/1559.6%—Exploratory (sweet spot) 9B8.0%0/1517.1%—Rejection-dominated 27B25.6%1/1546.8%—Narrow-AMEND Table G2 Gemma3 baseline behavioural statistics (15 trials per size, temperature = 0.7). Size Adoption Winners Vote YES Mean ScoreRegime 270M1.2%0/15—Frozen (w/ amendments) 1B0.0%0/15<2%−64Frozen 4B0.1%0/15<2%−64Frozen 12B53.9%3/1561.6%+17.5Exploratory (sweet spot) 27B1.9%0/15 ∼5%−61Veto-heavy Heterogeneous societies. Both families produce gridlock in mixed-size societies. The Qwen3.5 heterogeneous condition (mixed 0.8B/2B/4B/9B/27B) yields 5.9% adoption with 0/15 winners and 88.5% NO votes from the 9B member; the Gemma3 heterogeneous condition (mixed 270M/1B/4B/12B/27B) yields 0.0% adoption with 0/15 winners and 99.5% NO votes 26 Table G3 Cross-family majorityfromstart intervention (15 trials per size). Family Size Adoption Winners Regime change? Qwen3.50.8B ∼0%0/15No Qwen3.52B ∼0%0/15No Qwen3.54B ∼40%1/15No Qwen3.59B ∼10%0/15No Qwen3.527B ∼42%5/15Yes (target lowered) Gemma3270M ∼2%0/15No Gemma31B ∼0%0/15No Gemma34B ∼0%0/15No Gemma312B ∼50%6/15No (enhanced) Gemma327B ∼2%0/15No Table G4 Gemma3 temperature sweep summary for 12B (sweet spot). Six trials per temperature; other sizes remain frozen across all temperatures. Mean Score is reported as team total (sum over five players), consistent with the within-table scaling of mean-score values. Temperature Winners Adoption Mean team score 0.10/639.7%−72 0.30/644.0%−15 0.50/649.3%+3 0.71/653.0%+36 0.92/656.7%+116 1.14/660.5%+231 1.34/662.7%+237 from the 27B member. The cross-family pattern is therefore the same: the largest member’s veto suppresses rule adoption regardless of the presence of a sweet-spot member, with the effect being more severe in Gemma3. 27 References [1] Qwen Team. Qwen3.5: Towards native multimodal agents (2026). URL https: //qwen.ai/blog?id=qwen3.5 [2] Gemma Team, Google DeepMind, Gemma 3 technical report. arXiv preprint arXiv:2503.19786 (2025) [3] R. Axelrod, The Evolution of Cooperation (Basic Books, 1984) [4] E. Ostrom, Governing the commons: The evolution of institutions for collective action (Cambridge university press, 1990) [5] D.C. North, Institutions, Institutional Change and Economic Performance (Cam- bridge University Press, 1990) [6] J. Henrich, The Secret of Our Success: How Culture Is Driving Human Evolution, Domesticating Our Species, and Making Us Smarter (Princeton University Press, 2016) [7] R. Boyd, P.J. Richerson, in Better Than Conscious?: Decision Making, the Human Mind, and Implications For Institutions, ed. by C. Engel, W. Singer (The MIT Press, 2008). https://doi.org/10.7551/mitpress/9780262195805.003.0014. URL https://doi.org/10.7551/mitpress/9780262195805.003.0014 [8] M. Galesic, D. Barkoczi, A.M. Berdahl, D. Biro, G. Carbone, I. Giannoccaro, R.L. Goldstone, C. Gonzalez, A. Kandler, A.B. Kao, et al., Beyond collective intelligence: Collective adaptation. Journal of the Royal Society Interface 20(200), 20220736 (2023). https://doi.org/10.1098/rsif.2022.0736 [9] G. Masumoto, T. Ikegami, A new formalization of a meta-game using the lambda calculus. BioSystems 80(3), 219–231 (2005) [10] J.S. Park, J. O’Brien, C.J. Cai, M.R. Morris, P. Liang, M.S. Bernstein, Generative agents: Interactive simulacra of human behavior, in Proceedings of the 36th annual acm symposium on user interface software and technology (2023), p. 1–22 [11] I. Grossmann, M. Feinberg, D.C. Parker, N.A. Christakis, P.E. Tetlock, W.A. Cunningham, AI and the transformation of social science research. Science 380(6650), 1108–1109 (2023) [12] C.A. Bail, Can generative AI improve social science? Proceedings of the National Academy of Sciences 121(21), e2314021121 (2024) [13] L.P. Argyle, E.C. Busby, N. Fulda, J.R. Gubler, C. Rytting, D. Wingate, Out of one, many: Using language models to simulate human samples. Political Analysis 31(3), 337–351 (2023) 28 [14] Y. Zeng, C. Brown, M. Rounsevell, Too human to model: the uncanny valley of large language models in simulating human systems. npj Complexity 3, 13 (2026). https://doi.org/10.1038/s44260-026-00075-1 [15] P. Cobben, X.A. Huang, T.A. Pham, I. Dahlgren, T.J. Zhang, Z. Jin, Gt- harmbench: Benchmarking ai safety risks through the lens of game theory. arXiv preprint arXiv:2602.12316 (2026) [16] J. Duan, R. Zhang, J. Diffenderfer, B. Kailkhura, L. Sun, E. Stengel-Eskin, M. Bansal, T. Chen, K. Xu, Gtbench: Uncovering the strategic reasoning capa- bilities of llms via game-theoretic evaluations. Advances in Neural Information Processing Systems 37, 28219–28253 (2024) [17] A. Costarelli, M. Allen, R. Hauksson, G. Sodunke, S. Hariharan, C. Cheng, W. Li, J. Clymer, A. Yadav, Gamebench: Evaluating strategic reasoning abilities of llm agents. arXiv preprint arXiv:2406.06613 (2024) [18] M.F.A.R.D.T. (FAIR)†, A. Bakhtin, N. Brown, E. Dinan, G. Farina, C. Flaherty, D. Fried, A. Goff, J. Gray, H. Hu, et al., Human-level play in the game of diplo- macy by combining language models with strategic reasoning. Science 378(6624), 1067–1074 (2022) [19] Y. Xu, S. Wang, P. Li, F. Luo, X. Wang, W. Liu, Y. Liu, Exploring large language models for communication games: An empirical study on werewolf. arXiv preprint arXiv:2309.04658 (2023) [20] J. Light, M. Cai, S. Shen, Z. Hu, Avalonbench: Evaluating llms playing the game of avalon. arXiv preprint arXiv:2310.05036 (2023) [21] J. Guo, B. Yang, P. Yoo, B.Y. Lin, Y. Iwasawa, Y. Matsuo, Suspicion-agent: Playing imperfect information games with theory of mind aware gpt-4. arXiv preprint arXiv:2309.17277 (2023) [22] S. Hu, T. Huang, G. Liu, R.R. Kompella, F. Ilhan, S.F. Tekin, Y. Xu, Z. Yahn, L. Liu, A survey on large language model-based game agents. arXiv preprint arXiv:2404.02039 (2024) [23] Q. Zhao, J. Wang, Y. Zhang, Y. Jin, K. Zhu, H. Chen, X. Xie, CompeteAI: under- standing the competition dynamics of large language model-based agents, in Pro- ceedings of the 41st International Conference on Machine Learning (JMLR.org, 2024), ICML’24 [24] K.R. McKee, A. Tacchetti, M.A. Bakker, J. Balaguer, L. Campbell-Gillingham, R. Everett, M. Botvinick, Scaffolding cooperation in human groups with deep reinforcement learning. Nature Human Behaviour 7(10), 1787–1796 (2023) 29 [25] Y. Du, S. Li, A. Torralba, J.B. Tenenbaum, I. Mordatch, Improving factual- ity and reasoning in language models through multiagent debate, in Forty-first international conference on machine learning (2024) [26] G. Li, H. Hammoud, H. Itani, D. Khizbullin, B. Ghanem, Camel: Communicative agents for” mind” exploration of large language model society. Advances in neural information processing systems 36, 51991–52008 (2023) [27] E. Akata, L. Schulz, J. Coda-Forno, S.J. Oh, M. Bethge, E. Schulz, Playing repeated games with large language models. Nature Human Behaviour 9(7), 1380–1390 (2025) [28] K. Gandhi, D. Sadigh, N.D. Goodman, Strategic reasoning with language models. arXiv preprint arXiv:2305.19165 (2023) [29] A. Pan, J.S. Chan, A. Zou, N. Li, S. Basart, T. Woodside, H. Zhang, S. Emmons, D. Hendrycks, Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark, in International conference on machine learning (PMLR, 2023), p. 26837–26867 [30] Y. Lan, Z. Hu, L. Wang, Y. Wang, D. Ye, P. Zhao, E.P. Lim, H. Xiong, H. Wang, Llm-based agent society investigation: Collaboration and confrontation in avalon gameplay, in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (2024), p. 128–145 [31] H. Sun, Y. Wu, P. Wang, W. Chen, Y. Cheng, X. Deng, X. Chu, Game the- ory meets large language models: A systematic survey with taxonomy and new frontiers. arXiv preprint arXiv:2502.09053 (2025) [32] P. Suber, The Paradox of Self-Amendment: A Study of Logic, Law, Omnipotence, and Change (Peter Lang Publishing, 1990) [33] M. Hatakeyama, T. Hashimoto, Minimum nomic: a tool for studying rule dynamics. Artificial Life and Robotics 13(2), 500–503 (2009) [34] A. Hota, J.P. Jokinen, Nomiclaw: Emergent trust and strategic argumentation in llms during collaborative law-making. arXiv preprint arXiv:2508.05344 (2025) [35] S. Huang, D. Siddarth, L. Lovitt, T.I. Liao, E. Durmus, A. Tamkin, D. Gan- guli, Collective Constitutional AI: Aligning a Language Model with Public Input, in Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (2024), p. 1395–1417. https://doi.org/10.1145/3630106.3658979 [36] C. Fernando, D. Banarse, H. Michalewski, S. Osindero, T. Rockt ̈aschel, Prompt- breeder: Self-referential self-improvement via prompt evolution. arXiv preprint arXiv:2309.16797 (2023) 30 [37] K. Horibe, Evolvability in rule-making: A self-amendment game among llm agents, in Proceedings of the Genetic and Evolutionary Computation Conference Companion (2025), p. 2127–2137 [38] E. Hughes, M. Dennis, J. Parker-Holder, F. Behbahani, A. Mavalankar, Y. Shi, T. Schaul, T. Rocktaschel, Open-endedness is essential for artificial superhuman intelligence. arXiv preprint arXiv:2406.04268 (2024) [39] J. Clune, Ai-gas: Ai-generating algorithms, an alternate paradigm for producing general artificial intelligence. arXiv preprint arXiv:1905.10985 (2019) [40] G. Wang, Y. Xie, Y. Jiang, A. Mandlekar, C. Xiao, Y. Zhu, L. Fan, A. Anandku- mar, Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291 (2023) [41] J. Lehman, J. Gordon, S. Jain, K. Ndousse, C. Yeh, K.O. Stanley, in Handbook of evolutionary machine learning (Springer, 2023), p. 331–366 [42] K.O. Stanley, J. Lehman, Why Greatness Cannot Be Planned: The Myth of the Objective (Springer, 2015) [43] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al., Training language models to follow instruc- tions with human feedback. Advances in neural information processing systems 35, 27730–27744 (2022) [44] R. Rafailov, A. Sharma, E. Mitchell, C.D. Manning, S. Ermon, C. Finn, Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 36, 53728–53741 (2023) [45] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al., Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073 (2022) [46] P.F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, D. Amodei, Deep rein- forcement learning from human preferences. Advances in neural information processing systems 30 (2017) [47] E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, et al., Discovering language model behaviors with model-written evaluations, in Findings of the association for computational linguistics: ACL 2023 (2023), p. 13387–13434 [48] N. Mu, S. Chen, Z. Wang, S. Chen, D. Karamardian, L. Aljeraisy, B. Alo- mair, D. Hendrycks, D. Wagner, Can llms follow simple rules? arXiv preprint arXiv:2311.04235 (2023) 31 [49] J. Kaplan, S. McCandlish, T. Henighan, T.B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, D. Amodei, Scaling laws for neural language models. arXiv preprint arXiv:2001.08361 (2020) [50] J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. Casas, L.A. Hendricks, J. Welbl, A. Clark, et al., Training compute-optimal large language models. arXiv preprint arXiv:2203.15556 10 (2022) [51] T. Brown, B. Mann, N. Ryder, M. Subbiah, J.D. Kaplan, P. Dhariwal, A. Nee- lakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners. Advances in neural information processing systems 33, 1877–1901 (2020) [52] J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler, et al., Emergent abilities of large language models. Transactions on Machine Learning Research (2022) [53] R. Schaeffer, B. Miranda, S. Koyejo, Are emergent abilities of large language models a mirage? Advances in neural information processing systems 36, 55565– 55581 (2023) [54] I.R. McKenzie, A. Lyzhov, M. Pieler, A. Parrish, A. Mueller, A. Prabhu, E. McLean, A. Kirtland, A. Ross, A. Liu, et al., Inverse scaling: When bigger isn’t better. Transactions on Machine Learning Research (2023) [55] T.Y. Wu, M. Lo, U-shaped and inverted-u scaling behind emergent abilities of large language models, in International Conference on Learning Representations, vol. 2025 (2025), p. 99426–99458 [56] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q.V. Le, D. Zhou, et al., Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems 35, 24824–24837 (2022) [57] T. Kojima, S.S. Gu, M. Reid, Y. Matsuo, Y. Iwasawa, Large language models are zero-shot reasoners. Advances in neural information processing systems 35, 22199–22213 (2022) [58] S. Yao, D. Yu, J. Zhao, I. Shafran, T. Griffiths, Y. Cao, K. Narasimhan, Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems 36, 11809–11822 (2023) [59] X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, D. Zhou, Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171 (2022) [60] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang, et al., Self-refine: Iterative refinement with self- feedback. Advances in neural information processing systems 36, 46534–46594 32 (2023) [61] N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, S. Yao, Reflexion: Lan- guage agents with verbal reinforcement learning. Advances in neural information processing systems 36, 8634–8652 (2023) [62] M. Kosinski, Evaluating large language models in theory of mind tasks. Proceed- ings of the National Academy of Sciences 121(45), e2405460121 (2024) [63] T. Ullman, Large language models fail on trivial alterations to theory-of-mind tasks. arXiv preprint arXiv:2302.08399 (2023) [64] M. Sap, R. Le Bras, D. Fried, Y. Choi, Neural theory-of-mind? on the limits of social intelligence in large lms, in Proceedings of the 2022 conference on empirical methods in natural language processing (2022), p. 3762–3780 [65] J.H. Flavell, Metacognition and cognitive monitoring: A new area of cognitive– developmental inquiry. American psychologist 34(10), 906 (1979) [66] A. Didolkar, A. Goyal, N.R. Ke, S. Guo, M. Valko, T. Lillicrap, D. Rezende, Y. Bengio, M. Mozer, S. Arora, Metacognitive capabilities of llms: An exploration in mathematical problem solving. Advances in Neural Information Processing Systems 37, 19783–19812 (2024) [67] J.R. Searle, Minds, brains, and programs. Behavioral and brain sciences 3(3), 417–424 (1980) [68] G. Alain, Y. Bengio, Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644 (2016) [69] Y. Belinkov, Probing classifiers: Promises, shortcomings, and advances. Compu- tational Linguistics 48(1), 207–219 (2022) [70] nostalgebraist. interpreting GPT: the logit lens. https://w.lesswrong.com/ posts/AcKRB8wDpdaN6v6ru/interpreting-gpt-the-logit-lens (2020) [71] N. Belrose, I. Ostrovsky, L. McKinney, Z. Furman, L. Smith, D. Halawi, S. Bider- man, J. Steinhardt, Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112 (2023) [72] K. Meng, D. Bau, A. Andonian, Y. Belinkov, Locating and editing factual associa- tions in gpt. Advances in neural information processing systems 35, 17359–17372 (2022) [73] J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, S. Shieber, Investi- gating gender bias in language models using causal mediation analysis. Advances in neural information processing systems 33, 12388–12401 (2020) 33 [74] J. Lindsey, Emergent introspective awareness in large language models. Trans- former Circuits Thread (2025).URL https://transformer-circuits.pub/2025/ introspection/index.html [75] K. Li, A.K. Hopkins, D. Bau, F. Vi ́egas, H. Pfister, M. Wattenberg, Emergent world representations: Exploring a sequence model trained on a synthetic task. arXiv preprint arXiv:2210.13382 (2022) [76] W. Gurnee, M. Tegmark, Language models represent space and time, in Interna- tional Conference on Learning Representations, vol. 2024 (2024), p. 2483–2503 [77] G. Piatti, Z. Jin, M. Kleiman-Weiner, B. Sch ̈olkopf, M. Sachan, R. Mihalcea, Cooperate or collapse: Emergence of sustainable cooperation in a society of llm agents. Advances in Neural Information Processing Systems 37, 111715–111759 (2024) [78] V. Conitzer, R. Freedman, J. Heitzig, W.H. Holliday, B.M. Jacobs, N. Lam- bert, M. Moss ́e, E. Pacuit, S. Russell, H. Schoelkopf, et al., Social choice should guide ai alignment in dealing with diverse human feedback. arXiv preprint arXiv:2404.10271 (2024) [79] J. Lovato, N. Landry, L. Hebert-Dufresne, et al., Governance as a complex, networked, democratic, satisfiability problem. npj Complexity 2, 14 (2025). https://doi.org/10.1038/s44260-025-00041-3 [80] J.G. March, Exploration and exploitation in organizational learning. Organization Science 2(1), 71–87 (1991) 34