Paper deep dive
Chaotic Dynamics in Multi-LLM Deliberation
Hajime Shimao, Warut Khern-am-nuai, Sung Joo Kim
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/13/2026, 12:58:55 AM
Summary
This paper investigates the chaotic dynamics of multi-LLM deliberation systems, modeling them as random dynamical systems. It identifies two primary routes to instabilityârole differentiation and model heterogeneityâand demonstrates that these factors interact non-additively to amplify trajectory divergence, even under deterministic (T=0) settings. The study provides a stability map for multi-LLM governance and proposes actionable interventions, such as Chair-role ablation and memory-window reduction, to mitigate unpredictability.
Entities (6)
Relation Signals (4)
Chair-role ablation â reducesdivergence â Multi-LLM Deliberation
confidence 98% · Mechanistically, Chair-role ablation reduces ^λ most strongly
Memory window reduction â attenuatesdivergence â Multi-LLM Deliberation
confidence 96% · targeted protocol variants that shorten memory windows further attenuate divergence
Role Differentiation â causesinstability â Multi-LLM Deliberation
confidence 95% · two independent routes to instability: role differentiation in homogeneous committees
Model Heterogeneity â causesinstability â Multi-LLM Deliberation
confidence 95% · model heterogeneity in no-role committees
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Collective AI systems increasingly rely on multi-LLM deliberation, but their stability under repeated execution remains poorly characterized. We model five-agent LLM committees as random dynamical systems and quantify inter-run sensitivity using an empirical Lyapunov exponent ($\hat{\lambda}$) derived from trajectory divergence in committee mean preferences. Across 12 policy scenarios, a factorial design at $T=0$ identifies two independent routes to instability: role differentiation in homogeneous committees and model heterogeneity in no-role committees. Critically, these effects appear even in the $T=0$ regime where practitioners often expect deterministic behavior. In the HL-01 benchmark, both routes produce elevated divergence ($\hat{\lambda}=0.0541$ and $0.0947$, respectively), while homogeneous no-role committees also remain in a positive-divergence regime ($\hat{\lambda}=0.0221$). The combined mixed+roles condition is less unstable than mixed+no-role ($\hat{\lambda}=0.0519$ vs $0.0947$), showing non-additive interaction. Mechanistically, Chair-role ablation reduces $\hat{\lambda}$ most strongly, and targeted protocol variants that shorten memory windows further attenuate divergence. These results support stability auditing as a core design requirement for multi-LLM governance systems.
Tags
Links
- Source: https://arxiv.org/abs/2603.09127v1
- Canonical: https://arxiv.org/abs/2603.09127v1
Trouble viewing inline? Open PDF directly â
Full Text
46,080 characters extracted from source content.
Expand or collapse full text
Chaotic Dynamics in Multi-LLM Deliberation Hajime Shimao 1,* , Warut Khern-am-nuai 2 , and Sung Joo Kim 3 1 The Pennsylvania State University, Malvern, PA, USA 2 McGill University, Montreal, QC, Canada 3 American University, Washington, DC, USA * Corresponding author: hjs5825@psu.edu Abstract Collective AI systems increasingly rely on multi-LLM delibera- tion, but their stability under repeated execution remains poorly characterized. We model five-agent LLM committees as ran- dom dynamical systems and quantify inter-run sensitivity using an empirical Lyapunov exponent ( λ) derived from trajectory divergence in committee mean preferences. Across 12 policy scenarios, a factorial design at T = 0 identifies two indepen- dent routes to instability: role differentiation in homogeneous committees and model heterogeneity in no-role committees. Critically, these effects appear even in the T = 0 regime where practitioners often expect deterministic behavior. In the HL-01 benchmark, both routes produce elevated divergence ( λ= 0.0541 and 0.0947, respectively), while homogeneous no-role committees also remain in a positive-divergence regime ( λ= 0.0221). The combined mixed+roles condition is less un- stable than mixed+no-role ( λ= 0.0519 vs. 0.0947), showing non-additive interaction. Mechanistically, Chair-role ablation reduces λ most strongly, and targeted protocol variants that shorten memory windows further attenuate divergence. These results support stability auditing as a core design requirement for multi-LLM governance systems. Introduction Large language models are rapidly moving from single-agent usage to committee-like deployments in which multiple agents exchange arguments and then vote or synthesize a decision [Wu et al., 2023, Chen et al., 2023, Li et al., 2023, Hong et al., 2023, Park et al., 2023]. For these systems, reproducibility is a governance property, not only an engineering convenience [Gun- dersen and Kjensmo, 2018, Pineau et al., 2021]. If nominally identical committee runs diverge into different trajectories and final decisions, institutions face uncertainty that standard one-shot model evaluation does not capture. This concern is strongest at T = 0: if instability persists when explicit sam- pling noise is minimized, unpredictability is structural rather than a temperature artifact. Prior work has established instability of prompt-level LLM outputs [Sclar et al., 2024, Zheng et al., 2023], and social- science literature shows that group outcomes are sensitive to composition and framing [Sunstein, 2002, Tversky and Kahneman, 1981, Page, 2007]. Recent agent-safety work also uses âchaosâ in a broader descriptive sense for deployment failures in autonomous systems [Shapira et al., 2026]; our use is narrower and dynamical: trajectory-sensitive divergence quantified by an empirical Lyapunov-style estimand. What is missing is an experimentally identified design map for multi- LLM instability: which architectural choices create, amplify, or attenuate divergence in collective AI deliberation. We study this as a random dynamical system [Arnold, 1998]. The core experiment crosses two design axes: role struc- ture (NoRoles vs. Roles) and model composition (uniform provider/model vs. heterogeneous mixed committee). We then test mechanism (role ablation), intervention (reduced mem- ory depth), and a direct falsification target (near-memoryless update). The central empirical claim is that instability is design-induced and experimentally decomposable into two routes with non-additive interaction. The key practical impli- cation is that unpredictability can remain material even under seemingly âdeterministicâ settings. Results In plain terms, we find that the same AI committee can reach different collective trajectories and decisions even when run repeatedly under nominally identical settings, including at T = 0. Two design choices most strongly shape this instability: assigning differentiated institutional roles, and mixing different model families within one committee. Across experiments, these factors do not combine additively, and targeted protocol changes can reduce instability. The results below quantify these effects, identify a key amplification mechanism, and test intervention/falsification predictions. Experimental design and estimand Each run executes a five-agent committee for 20 rounds. At each turn, each agent outputs an argument and structured 1 arXiv:2603.09127v1 [cs.AI] 10 Mar 2026 preference state s (i) t = (p A , p B , p C ,conf). For replicate r, we compute the committee mean trajectoryp (r) t = 1 N P i s (i) t on the simplex. Inter-replicate divergence is D(t) = 2 R(Râ 1) X i<j p (i) t â p (j) t 2 .(1) We estimate λas the slope oflogD(t) over rounds 3â20. Positive λindicates exponential divergence of trajectories. Unless otherwise noted, target replicate count is R = 20 per condition. Two design routes and non-additive interaction Figure 1A shows HL-01 trajectories under three canonical con- ditions at T = 0. Homogeneous NoRoles is lower-divergence but still positive ( λ= 0.0221), while adding role mandates in a homogeneous committee increases divergence ( λ= 0.0541). Independently, heterogeneous model composition without roles also yields high divergence ( λ= 0.0947). These correspond to two distinct routes to stronger instability: institutional differ- entiation (Route A) and compositional heterogeneity (Route B), both amplifying a non-zero baseline. Figure 1B provides the full HL-01 2Ă2 matrix. The mixed+roles cell is substantially less unstable than mixed+no- role ( λ = 0.0519 vs. 0.0947), indicating non-additivity. Thus, increasing both diversity dimensions does not monotonically increase instability [Hong and Page, 2004, Page, 2007]. Figure 4 extends this to the 12-scenario benchmark at T = 0. Across scenarios, uniform NoRoles is generally the lowest- divergence configuration, but often still shows λ > 0. Uniform Roles and mixed NoRoles are frequently more elevated. The mixed Roles condition remains heterogeneous in mag- nitude but is fully estimated across the scenario set (realized n = 20 per mixed cell). This closes the previously incomplete mixed-model matrix and shifts the claim from single-scenario demonstration to benchmark-level pattern. Mechanism: Chair role as dominant amplifier Figure 2 isolates the Chair channel in a canonical single-scenario mechanism panel. In HL-01 (panel A), role-ablation effects are reported on the sameâ λ= λ full roles â λ ablated scale used in panel B, showing the largest reduction for Chair ablation; other single-role ablations are near zero or negative in this scenario. Panel B shows cross-scenario heterogeneity in this effect usingâ λ= λ full roles â λ ablate Chair across five scenarios (IM-01, HL-01, CL-01, SP-03, AI-01). All five point estimates are positive, but uncertainty differs by scenario: 95% CIs exclude zero in IM-01 and HL-01, while the remaining three are directionally positive but inconclusive at current sample size. This pattern supports a mechanistic interpretation in which synthesis-focused Chair behavior is an important, scenario- contingent amplification channel [Kerr and Tindale, 2004]. Intervention and falsification tests Figure 3 reports two protocol stress tests. First, reducing argument-memory depth from k = 15 to k = 3 (intervention) lowers λin all four tested scenarios (AI-01, CL-01, HL-01, SP-03). Second, a stricter memory-collapse variant (k = 1) in IM-01 and CL-01 lowers λrelative to baseline Roles. These tests do not eliminate divergence in all settings, but they support the hypothesis that early-round feedback memory is an important contributor to amplification. Scope and robustness Main-text results focus on temperature T = 0 to isolate structural effects from sampling-temperature variation. Com- plementary analyses in SI show that the role-induced instability signature persists across a wider temperature range, consistent with a structural rather than purely thermal interpretation. In other words, setting T = 0 does not remove instability; it only removes one explicit noise control while the protocol can still amplify residual variation. Our usage of âchaotic dynamicsâ is empirical: λis estimated from inter-replicate divergence and should be interpreted as an annealed instability index for black-box committee dynamics [Eckmann and Ruelle, 1985, Ott, 2002]. Discussion The main empirical contribution is a causal decomposition of multi-LLM instability into two experimentally manipulable routes, plus evidence that their interaction is non-additive. For governance, this means committee architecture must be audited as a joint design system rather than tuned along a sin- gle âdiversityâ axis. Importantly, this is an amplification story rather than a zero-to-chaos switch: even baseline uniform/no- role committees can exhibit measurable instability [Selbst et al., 2019, Raji et al., 2020, National Institute of Standards and Technology, 2023, OECD, 2019]. From a deployment perspective, this creates three linked risks: unpredictability (replicate-sensitive outcomes), limited controllability (small protocol or execution perturbations can redirect the committee to different decision basins), and limited explainability (single- run rationales are insufficient summaries of system behavior across replicates). A second contribution is actionable mechanism. Chair- focused ablation and memory-window interventions suggest that instability can be reduced without collapsing deliberation to a trivial one-step vote. This creates a concrete path from diagnosis to controllable protocol design. These findings should be interpreted within scope. We do not claim that all multi-LLM committees are chaotic under all tasks. We claim that in the tested protocol family and bench- mark, instability is reliably inducible by specific architectural choices and measurably reducible by targeted interventions. 2 Two next steps are especially important for top-tier re- view: linking instability to external task quality (accu- racy/calibration/decision harm), and extending intervention success criteria from âlower λâ to âlower λwith preserved deci- sion quality.â The present manuscript establishes the stability map needed to run those consequential tests. Materials and Methods Protocol. We use a windowed-summary deliberation pro- tocol. Each agent receives the scenario prompt, a sliding transcript window (default k = 15), and a committee state table; it returns argument text and a structured state line. A deterministic parser extracts preference vectors. One auto- repair pass is allowed on parse failure. Scenarios. The benchmark contains 12 policy scenarios across immigration, health, income, climate, speech, and AI governance (two scenarios per domain). Models and conditions. Uniform runs use GPT-4.1-mini. Mixed-model runs use a harmonized one-Grok lineup: Chair (GPT-4.1), Welfare (Claude Sonnet 4.6), Rights (Gemini 2.5 Flash), Equity (Grok-3-mini), Security (GPT-4.1-mini). Core benchmark conditions are NoRoles/Roles crossed with uniform/mixed model composition at T = 0. Additional runs cover role ablation and memory-window variants (k = 3, k = 1). Inference and uncertainty. For each condition, we estimate λfrom log-linear fits of D(t) over rounds 3â20. Bootstrap confidence intervals are computed by replicate resampling. Realized n is reported at panel/cell level where relevant. Re- porting follows reproducibility-oriented practice for empirical ML benchmarks [Pineau et al., 2021]. Implementation. Experiments and analyses are imple- mented in Python 3.11 with NumPy, SciPy, and Matplotlib [Vir- tanen et al., 2020, Hunter, 2007]. Acknowledgments We thank colleagues for feedback on early drafts and experi- mental design. Supplementary Materials Supplementary Text S1âS3 Supplementary Figures S1âS7 Supplementary Tables S1âS8 Data S1: JSONL run artifacts Code S1: Experiment and analysis scripts 3 2.55.07.510.012.515.017.520.0 Round t 0.0 0.1 0.2 0.3 0.4 0.5 Mean pairwise distance D(t) Ì Î»=+0.022 Ì Î»=+0.054 Ì Î»=+0.095 A Two routes to instability (HL-01, T=0) Uniform, no roles Uniform, with roles Mixed, no roles No rolesWith roles Uniform Mixed +0.022 n=20 +0.054 n=19 +0.095 n=20 +0.052 n=20 B 2Ă2 governance matrix (HL-01, T=0) â0.02 0.00 0.02 0.04 0.06 0.08 0.10 Ì Î» Figure 1: Two design routes to instability and their interaction (HL-01, T = 0). (A) Mean pairwise trajectory distance D(t) for three conditions: uniform NoRoles (positive but lower-divergence baseline), uniform Roles (Route A), and mixed NoRoles (Route B). Both Route A and Route B further amplify divergence. (B) 2Ă2 matrix crossing roles and model composition. The mixed+roles cell is less unstable than mixed+no-role, showing non-additive interaction. Realized sample sizes are shown in-cell (uniform+roles has slight attrition at n = 19; all other cells are n = 20). â0.06â0.04â0.020.000.020.04 Î Ì Î» = full roles - ablated Ablate Chair Ablate Equity Ablate Rights Ablate Welfare Ablate Security No roles A HL-01 role-ablation effects (T=0) IM-01HL-01SP-03AI-01CL-01 â0.04 â0.02 0.00 0.02 0.04 0.06 0.08 Î Ì Î» = full roles - ablate Chair Positive: 4/5 95% CI > 0: 2/5 B Chair effect heterogeneity across scenarios Figure 2: Chair mechanism in HL-01 and cross-scenario heterogeneity. (A) HL-01 role-ablation effects onâ λ= λ full roles â λ ablated . Chair ablation yields the largest reduction among role mandates (with a no-role reference bar shown for context). (B) Scenario-level Chair effect across IM-01, HL-01, CL-01, SP-03, and AI-01, reported asâ λ= λ full roles â λ ablate Chair . Point estimates are positive in all five scenarios, with strongest and best-resolved effects in IM-01 and HL-01. 4 AI-01CL-01HL-01SP-03 â0.02 â0.01 0.00 0.01 0.02 0.03 0.04 Ì Î» A Intervention test: reduced memory window Baseline roles (k=15) Intervention (k=3) IM-01CL-01 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 0.08 Ì Î» B Falsification target: near-memoryless update Baseline roles (k=15) Falsification (k=1) Figure 3: Protocol intervention and falsification tests via memory-depth control. (A) Intervention: reducing memory window from k = 15 to k = 3 attenuates divergence across four scenarios. (B) Falsification target: collapsing memory to k = 1 lowers λin IM-01 and CL-01 relative to baseline Roles. Together these tests support a feedback-memory amplification pathway. Uniform No roles Uniform With roles Mixed No roles Mixed With roles IM-01 â Immigration Asylum IM-02 â City Safety HL-01 â Health Financing HL-02 â Hospital Ethics IN-01 â Income Policy IN-02 â Welfare Oversight CL-01 â Climate Policy CL-04 â Adaptation Fund SP-01 â Platform Integrity SP-03 â Recommender Systems AI-01 â Model Release AI-02 â Algorithmic Accountability +0.005 n=20 +0.084 n=20 +0.094 n=20 +0.034 n=20 +0.022 n=20 +0.020 n=19 +0.065 n=20 +0.001 n=20 +0.022 n=20 +0.042 n=20 +0.095 n=20 +0.052 n=20 +0.023 n=20 +0.045 n=20 -0.036 n=20 -0.008 n=20 +0.031 n=20 +0.054 n=20 +0.025 n=20 +0.022 n=20 +0.025 n=20 +0.022 n=20 -0.018 n=20 +0.041 n=20 +0.049 n=20 +0.030 n=20 +0.090 n=20 +0.078 n=20 +0.040 n=19 +0.020 n=20 +0.093 n=20 +0.052 n=20 +0.016 n=20 +0.030 n=20 +0.024 n=20 +0.047 n=20 +0.031 n=18 +0.026 n=20 +0.090 n=20 +0.087 n=20 +0.014 n=20 +0.004 n=20 -0.006 n=20 -0.002 n=20 +0.027 n=20 +0.039 n=20 +0.056 n=20 +0.022 n=20 Full cross-scenario 2Ă2 matrix (T=0, N=5) â0.02 0.00 0.02 0.04 0.06 0.08 0.10 Ì Î» Figure 4: Full cross-scenario 2Ă2 instability landscape at T = 0. Heatmap of λacross 12 scenarios for uniform/mixedĂno-role/roles conditions. This panel provides the completed benchmark matrix for the two-route claim. Mixed-model cells are all estimated at n = 20; uniform cells have realized n displayed in-cell where slight attrition occurred. 5 References Qingyun Wu, Gagan Bansal, Jieyu Zhang, et al. AutoGen: En- abling next-gen LLM applications via multi-agent conversa- tion, 2023. URL https://arxiv.org/abs/2308.08155. Weize Chen, Yusheng Su, Jingwei Zuo, et al. AgentVerse: Fa- cilitating multi-agent collaboration and exploring emergent behaviors in agents, 2023. URLhttps://arxiv.org/abs/ 2308.10848. Guohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin, and Bernard Ghanem. CAMEL: Communicative agents for mind exploration of large language model society, 2023. URL https://arxiv.org/abs/2303.17760. Sirui Hong, Ming Zhuge, Junjie Chen, et al. MetaGPT: Meta programming for multi-agent collaborative framework, 2023. URL https://arxiv.org/abs/2308.00352. Joon Sung Park, Joseph C. OâBrien, Carrie J. Cai, Mered- ith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behav- ior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), pages 1â22, 2023. doi: 10.1145/3586183.3606763. Odd Erik Gundersen and SigbjĂžrn Kjensmo. State of the art: Reproducibility in artificial intelligence. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence (AAAI), pages 1644â1651, 2018. Joelle Pineau, Philippe Vincent-Lamarre, Koustuv Sinha, Vin- cent LariviĂšre, Alina Beygelzimer, Florence dâAlchĂ© Buc, Emily Fox, and Hugo Larochelle. Improving reproducibility in machine learning research: A report from the neurips 2019 reproducibility program. Journal of Machine Learning Research, 22(164):1â20, 2021. URLhttps://jmlr.org/ papers/v22/20-303.html. Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language modelsâ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In International Conference on Learning Representations (ICLR), 2024. URLhttps://arxiv.org/ abs/2310.11324. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, et al. Judging LLM-as-a-Judge with MT-Bench and chatbot arena. In Ad- vances in Neural Information Processing Systems (NeurIPS), volume 36, 2023. URLhttps://arxiv.org/abs/2306. 05685. Cass R. Sunstein. The law of group polarization. The Journal of Political Philosophy, 10:175â195, 2002. doi: 10.1111/ 1467-9760.00148. Amos Tversky and Daniel Kahneman. The framing of decisions and the psychology of choice. Science, 211(4481):453â458, 1981. doi: 10.1126/science.7455683. Scott E. Page. The Difference: How the Power of Diver- sity Creates Better Groups, Firms, Schools, and Societies. Princeton University Press, 2007. Natalie Shapira, Chris Wendler, Avery Yen, Gabriele Sarti, Koyena Pal, Olivia Floody, Adam Belfki, Alex Loftus, et al. Agents of chaos, 2026. URLhttps://arxiv.org/abs/ 2602.20021. Ludwig Arnold. Random Dynamical Systems. Springer, Berlin, 1998. Lu Hong and Scott E. Page. Groups of diverse problem solvers can outperform groups of high-ability problem solvers. Pro- ceedings of the National Academy of Sciences, 101(46): 16385â16389, 2004. doi: 10.1073/pnas.0403723101. Norbert L. Kerr and R. Scott Tindale. Group performance and decision making. Annual Review of Psychology, 55:623â655, 2004. doi: 10.1146/annurev.psych.55.090902.142009. Jean-Pierre Eckmann and David Ruelle. Ergodic theory of chaos and strange attractors. Reviews of Modern Physics, 57(3):617â656, 1985. doi: 10.1103/RevModPhys.57.617. Edward Ott. Chaos in Dynamical Systems. Cambridge Univer- sity Press, 2nd edition, 2002. Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. Fairness and ab- straction in sociotechnical systems. In Proceedings of the Conference on Fairness, Accountability, and Trans- parency (FAccT), pages 59â68, 2019. doi: 10.1145/3287560. 3287598. Inioluwa Deborah Raji, Andrew Smart, Rebecca N. White, et al. Closing the ai accountability gap: Defining an end- to-end framework for internal algorithmic auditing. In Proceedings of the Conference on Fairness, Accountabil- ity, and Transparency (FAccT), pages 33â44, 2020. doi: 10.1145/3351095.3372873. National Institute of Standards and Technology. Artificial intelli- gence risk management framework (AI RMF 1.0).https:// w.nist.gov/itl/ai-risk-management-framework, 2023. NIST AI 100-1. OECD. OECD principles on artificial intelligence.https: //oecd.ai/en/ai-principles, 2019. Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, et al. SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nature Methods, 17:261â272, 2020. doi: 10.1038/ s41592-020-0772-5. J. D. Hunter. Matplotlib: A 2d graphics environment. Com- puting in Science & Engineering, 9(3):90â95, 2007. doi: 10.1109/MCSE.2007.55. 6 Supplementary Information Chaotic Dynamics in Multi-LLM Deliberation Hajime Shimao 1,â , Warut Khern-am-nuai 2 , Sung Joo Kim 3 1 The Pennsylvania State University, Malvern, PA, USA 2 McGill University, Montreal, QC, Canada 3 American University, Washington, DC, USA â Corresponding author: hjs5825@psu.edu Contents S1 Server-side Non-Determinism at Temperature Zero1 S2 Deliberation Protocol and Prompt Templates2 S2.1 Windowed-Summary (WS) Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 S2.2 System Prompt Template . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 S2.3 STATE Parsing and Repair . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 S2.4 Clerk Aggregation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 S2.5 Ablation Protocol . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 S3 Branching Entropy Certificate4 S3.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 S3.2 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 S4 Semantic Perturbation Analysis4 S4.1 Rationale . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 S4.2 Variant Construction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 S4.3 Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 S5 Reproducibility and Run Accounting5 S5.1 Run accounting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 S5.2 File manifest for reported headline effects . . . . . . . . . . . . . . . . . . . . . . . . . 5 S5.3 Failure-source visibility . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 S1 Server-side Non-Determinism at Temperature Zero Modern LLM APIs document temperature T = 0 as producing âdeterministicâ outputs. In practice, however, cloud-hosted inference involves floating-point arithmetic across heterogeneous GPU clusters, and the result of individual operations is not guaranteed to be bitwise identical across calls. We verified this empirically for OpenAIâs API (model: gpt-4.1-mini). 1 Protocol. We submitted the same prompt 20 times withtemperature=0and recorded the full text response, the parsed preference vector p = (p A , p B , p C ), and the token-level logprob output (when available). Across all 20 calls with identical prompts, we observe non-zero variance in the output text and the parsed preference vectors in â 40â50% of trials. Magnitude. The typical pairwise L2 distance between preference vectors at T = 0 after 20 rounds is indistinguishable from that at T = 0.7: λ (T=0) = 0.0838 vs. λ (T=0.7) = 0.0822 (both roles=True, IM-01, R = 20 replicates, 95% CIs overlapping; see Fig. S1). Interpretation. Server-side floating-point non-determinism represents the minimum perturbation that any real-world deployment will face. Our system is therefore a chaos amplifier: even this infinitesimally small noise seed is expanded to macroscopic trajectory divergence by deliberative feedback dynamics. Role differentiation is one strong amplifier, but not a strict prerequisite: in selected scenarios (e.g., HL-01), the uniform no-role baseline remains positive in the completed 2Ă2 matrix. This makes the chaos result stronger, not weaker, as a practical concern: a practitioner cannot escape it even by setting T = 0. S2 Deliberation Protocol and Prompt Templates S2.1 Windowed-Summary (WS) Protocol Each deliberation consists of R = 20 rounds. At each round, every agent i (in a fixed random order) is called once. The agent receives a context window comprising: 1. The fixed system prompt (role mandate + output format instructions). 2. The fixed scenario packet (scenario set listed in Supplementary Table S1). 3.The sliding window of the last k = 15 arguments from the transcript (all agents, chronological), prefixed with role labels. 4.The current committee state table: last known parsed state vector (p A , p B , p C ,conf,tags) for each role, presented as a formatted table. S2.2 System Prompt Template The system prompt has two parts: a shared preamble and a role-specific mandate. Shared preamble (all agents): You are an AI agent participating in a structured committee deliberation. Your committee must reach a collective decision on a policy scenario. You will argue for your position, update it based on evidence and dialogue, and ultimately cast a private ballot. In every reply you MUST include exactly one STATE line formatted as: STATE: pref=[pA,pB,pC]; conf=N; tags=["tag1","tag2"] where pA+pB+pC=1.0 (floats), conf is your confidence (0â100), and tags is a list of exactly 2 short concept tags. Your argument must be at most 110 words. Be direct and specific. Role mandate (Chair): ROLE: Chair Your mandate: Facilitate productive deliberation. Synthesize diverse viewpoints, identify points of agreement and disagreement, and help the committee 2 move toward a reasoned consensus. You may challenge arguments that seem weakly supported. You do not advocate for a predetermined outcome. Role mandate (Welfare): ROLE: Welfare Your mandate: Prioritize aggregate welfare, efficiency, and cost-benefit logic. Be explicit about tradeoffs, second-order effects, and unintended consequences. Role mandate (Rights): ROLE: Rights Your mandate: Defend individual rights, due process, and non-discrimination. Flag any option that compromises fundamental rights even if it produces aggregate benefits. Deontological constraints take priority. Role mandate (Equity): ROLE: Equity Your mandate: Evaluate options through the lens of distributive justice and structural inequality. Flag disparate impacts on historically marginalized groups. Advocate for options that reduce systemic disadvantage. Role mandate (Security): ROLE: Security Your mandate: Assess risks to institutional stability, public safety, and long-term systemic resilience. Flag options that introduce unpredictable second-order harms. Prioritize precaution when uncertainty is high. S2.3 STATE Parsing and Repair After each agent response, a deterministic regex parser extracts the STATE line: STATE: pref=\[([0-9., ]+)\]; conf=([0-9]+); tags=\[([^\]]*)\] If parsing fails (malformed STATE), a single repair prompt is sent: Your previous response did not contain a valid STATE line. Please respond with ONLY the corrected STATE line in the format: STATE: pref=[pA,pB,pC]; conf=N; tags=["tag1","tag2"] If parsing still fails after the repair attempt, the run is flagged and excluded from analysis. Across all conditions, the parse-failure rate is below 2%. Implementation note. In the current codebase, the parser enforces a stricter realization of this template: exactly two tags in snake_case and decimal-formatted preferences. S2.4 Clerk Aggregation After all agents complete round R, each agent is called individually for a private ballot: a single API call requesting only a JSON object"decision": "A"|"B"|"C", "confidence": N. A separate clerk agent then receives all private ballots and produces the final committee decision by majority vote: CLERK: You are the vote aggregator. Given the following private ballots [list of ballots], determine the majority decision and return ONLY valid JSON: "decision": "A"|"B"|"C", "majority_count": N, "total": N 3 S2.5 Ablation Protocol For ablation experiments (role-ablation effects in Fig. 2A; full role-ablation panel in Table S5), the ablated roleâs mandate is replaced with an empty string. The agent is still present, receives the same transcript window, and must still output a STATE line, but has no specialized framing. This preserves group size N = 5 while removing the roleâs directional influence. S3 Branching Entropy Certificate S3.1 Definition The branching entropy Îquantifies local divergence: from a fixed committee state x t at round t, we sample K = 30 independent continuations of the deliberation and measure how much these trajectories diverge from one another. Formally, let x (k) t K k=1 be K independent continuations from the same initial state x t . Define: H K = 1 R â t R X s=t+1 D(x (k) s ),(1) ^ÎŒ(D) = mean pairwise L 2 distance among x (k) t ,(2) Î = H K â ^ÎŒ(D),(3) where D(·) is the mean pairwise L2 distance (Eq. 1 of main text). A positive Î >0 certifies that the system continues to diverge beyond what is attributable to the initial variation at x t . We further compute: ^c in = mean cosine similarity within each branch cluster, ^c out = mean cosine similarity across branch clusters. (4) A chaotic system should show^c in > ^c out : trajectories within a branch resemble each other more than trajectories across branches. S3.2 Results We ran K = 30 continuations from 5 independently chosen base states (replicates 0â4) for IM-01 at T = 0.7, roles=True. Results are shown in Fig. S6. Across all 5 replicates: âą Î = 0.054± 0.008 SE > 0 (all replicates individually positive). âą ^c in > ^c out in all 5 replicates. These results confirm that the system has a rich local attractor landscape: small perturbations from the same state produce qualitatively different downstream trajectories. S4 Semantic Perturbation Analysis S4.1 Rationale The main text shows that server-side FP noise (temperature T = 0, identical prompts) is sufficient to drive λ > 0. A separate question is whether changes in surface wordingâkeeping the semantic content 4 of the scenario identicalâproduce additional divergence. This matters for practical governance: if a policy scenario phrased slightly differently by two different civil servants leads to qualitatively different AI committee decisions, the decisions are not semantically auditable. S4.2 Variant Construction We constructed 6 phrasings of the IM-01 immigration scenario: 1. Canonical â the original scenario text. 2. Formal register â bureaucratic vocabulary, passive voice. 3. Conversational â plain-English rephrasing, active voice, simplified numbers. 4. Legalistic â statutory citation style, conditional clauses. 5.Passive reframe â all active policy frames restated as passive constraints (âgrants may be allocatedâ vs. âallocate grantsâ). 6.Synonym swap â systematic substitution of key nouns and verbs with near-synonyms (e.g., âapplicantâ â âpetitionerâ, âallocateâ â âdistributeâ). Semantic equivalence was verified by having GPT-4o score each variant pair on a 0â10 meaning-preservation scale; all pairs scored â„ 8.5. S4.3 Results Each variant was run once at T = 0 with roles=True for 20 rounds. The 6 trajectories are shown in Fig. S7. The empirical Lyapunov exponent across the 6 variants is λ = 0.030 (compared to λ = 0.0838 for 20 replicates of the canonical phrasing with FP noise). Two variants (Passive reframe and Conversational) converge to option B, while the remaining four converge to option A. With only six variants, this split should be interpreted as directional rather than definitive. The smaller λcompared to the FP-noise condition (0.030 vs. 0.084) reflects that synonym-level perturbations are smaller than server-side FP noise in this systemâbut still non-zero, confirming that surface wording affects collective decisions even when semantic content is preserved. S5 Reproducibility and Run Accounting S5.1 Run accounting Table S1 reports target versus realized replicate counts for key conditions used in the main narrative (including the HL-01 headline 2Ă2 panel and IM-01 mechanism/temperature panels). S5.2 File manifest for reported headline effects Table S2 maps headline results to exact JSONL files used for analysis and plotting. S5.3 Failure-source visibility Current JSONL artifacts store successful runs only. As a result, unresolved replicate deficits can be quantified, but failure type (parser failure vs. provider timeout/quota vs. interruption) is not fully 5 recoverable for all historical runs. Table S3 documents this limitation. Table S1: Run accounting for key reported conditions. Target and realized replicate counts as of March 10, 2026. Deficit = target minus realized. GroupConditionTarget Realized Deficit Headline HL-01: Uniform no-role (T = 0)20200 Headline HL-01: Uniform roles (T = 0)20191 Headline HL-01: Mixed no-role (T = 0)20200 Headline HL-01: Mixed roles (T = 0)20200 CoreUniform no-role (IM-01, T = 0)20200 CoreUniform roles (IM-01, T = 0)20200 CoreMixed no-role (IM-01, T = 0)20200 CoreMixed roles (IM-01, T = 0)20200 Core12-scenario matrix: Uniform no-role total2402373 Core12-scenario matrix: Uniform roles total2402364 Core12-scenario matrix: Mixed no-role total2402400 Core12-scenario matrix: Mixed roles total2402400 Ablation Ablate Chair (IM-01, T = 0)20200 Ablation Ablate Equity (IM-01, T = 0)20200 Semantics 6 IM-01 semantic variants660 Provider GPT-4.1, roles=True, T = 020200 Provider Claude Sonnet 4.6, roles=True, T = 020191 Provider Gemini 2.5 Flash, roles=True, T = 020200 Provider Grok-3, roles=True, T = 020164 Table S2: File manifest for headline reported effects. Paths are relative to project root. Mixed-model matrix files are read from synced VM shards indata/vm_sync/.../chaos_v1_one_grok_t0_matrix/(with fallback to local archived copies). EffectJSONL file HL-01 uniform no-role baseline (Fig. 1) data/raw/chaos_v1/HL-01__T0.0__N5__rolesFalse.jsonl HL-01 uniform role effect (Fig. 1) data/raw/chaos_v1/HL-01__T0.0__N5__rolesTrue.jsonl HL-01 mixed no-role route (Fig. 1) data/vm_sync/.../chaos_v1_one_grok_t0_matrix/HL-01__T0.0__N5__rolesFalse__multimodel.jsonl HL-01 mixed+roles interaction (Fig. 1) data/vm_sync/.../chaos_v1_one_grok_t0_matrix/HL-01__T0.0__N5__rolesTrue__multimodel.jsonl IM-01 uniform no-role baseline data/raw/chaos_v1/IM-01__T0.0__N5__rolesFalse.jsonl IM-01 uniform role effect data/raw/chaos_v1/IM-01__T0.0__N5__rolesTrue.jsonl IM-01 mixed no-role route data/raw/chaos_v1/IM-01__T0.0__N5__rolesFalse__multimodel.jsonl IM-01 mixed+roles interaction data/raw/chaos_v1/IM-01__T0.0__N5__rolesTrue__multimodel.jsonl Full 12-scenario mixed matrix data/vm_sync/.../chaos_v1_one_grok_t0_matrix/scenario__T0.0__N5__rolesFalse,True__multimodel.jsonl Chair ablation data/raw/chaos_v1/IM-01__T0.0__N5__rolesTrue__ablate-Chair.jsonl Chair replication panel data/vm_sync/instance-20260307-022228/chaos_v1_chair_rep_t0/*.jsonl Intervention (k=3) panel data/vm_sync/instance-20260307-022228/chaos_v1_intervention_kw3_t0/*.jsonl Falsification (k=1) panel data/vm_sync/instance-20260307-022228/chaos_v1_falsification_kw1_t0/*.jsonl Temperature robustness (Roles, T = 0.7) data/raw/chaos_v1/IM-01__T0.7__N5__rolesTrue.jsonl Temperature robustness (NoRoles, T = 0.7) data/raw/chaos_v1/IM-01__T0.7__N5__rolesFalse.jsonl Semantic perturbation data/raw/semantic_v1/IM-01__T0.0__N5__rolesTrue.jsonl 6 Table S3: Failure-source visibility in current artifacts. âKnownâ means directly inferable from stored artifacts; âUnknownâ means not reconstructable without external runtime logs. Failure sourceStatusNotes Provider timeout / quota / API errors Partially known Seen indirectly via realized n < target in provider rows. Parser failure after repairUnknownSuccessful JSONL files do not persist parser-failure counters. Manual interruption / VM stopUnknownRequires external process logs, not bundled in Data S1. Supplementary Figures 0 10 â5 10 â4 0.010.05 0.10.20.7 Temperature T 0.00 0.02 0.04 0.06 0.08 0.10 Empirical Lyapunov exponent Ì Î» (A) Roles 0 10 â5 10 â4 0.010.05 0.10.20.7 Temperature T (B) No Roles Full temperature sweep â IM-01, N=5, 20 replicates Figure S1: Full temperature sweep for both role conditions. Empirical Lyapunov exponent λas a function of temperature Tâ0, 10 â5 , 10 â4 , 0.01, 0.05, 0.1, 0.2, 0.7for (A) Roles and (B) NoRoles deliberation (IM-01, N = 5, 20 replicates each). Under Roles, λis statistically flat across all eight conditions. Under NoRoles, λis uniformly suppressed and statistically indistinguishable from zero in seven of eight conditions. Shaded bands: 95% bootstrap CI (500 resamples). 7 5101520 0.0 0.2 0.4 0.6 0.8 1.0 p A (solid), p B (dashed) Run 1 5101520 0.0 0.2 0.4 0.6 0.8 1.0 Run 2 5101520 0.0 0.2 0.4 0.6 0.8 1.0 Run 3 5101520 Round 0.0 0.2 0.4 0.6 0.8 1.0 p A (solid), p B (dashed) Run 4 5101520 Round 0.0 0.2 0.4 0.6 0.8 1.0 Run 5 5101520 Round 0.0 0.2 0.4 0.6 0.8 1.0 Run 6 Per-agent preferences (p A solid, p B dashed) â IM-01, T=0, roles=True; 6 representative runs ChairWelfareRightsEquitySecurity Figure S2: Per-agent preference trajectories across 6 representative runs. Solid lines show p A (preference for option A) and dashed lines show p B (preference for option B) for each of the five agents, color-coded by role (Chair: orange; Welfare: blue; Rights: green; Equity: purple; Security: brown). The Chair (orange) exhibits the most dynamic switching behavior, while Rights (green) and Security (brown) tend toward stable preferences. Condition: IM-01, T = 0, roles=True. 8 ChairWelfare Rights Equity Security 0 1 2 3 4 5 6 7 8 Preference switches per run Per-agent switch counts by role (IM-01, Tâ 0, 0.01, 0.05, 0.1, 0.2, 0.7, roles=True; 600 agent-runs) Figure S3: Per-agent preference switch counts by role. Bars show mean±SEM switch counts; points show individual agent-run values (jittered). Data pooled across Tâ0, 0.01, 0.05, 0.1, 0.2, 0.7, roles=True, IM-01. The Chair is the dominant switching agent (meanâ2.76 switches/run), with all non-Chair roles below one switch/run on average. This asymmetry identifies the Chair as the mechanistic source of chaotic amplification. 2.55.07.510.012.515.017.520.0 Round of first majority 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative proportion never Time-to-majority: ECDF across replicates (IM-01) Roles (T=0) Roles (T=0.7) No Roles (T=0) No Roles (T=0.7) Figure S4: Time-to-majority: empirical CDF across replicates. Each curve shows the fraction of runs in which a strict majority (â„3 of 5 agents) first converged on the same top-preference option by a given round, for four conditions on IM-01 (rolesĂ temperature). Under NoRoles, consensus forms in round 1 without exception. Under Roles, consensus is delayed to rounds 2â3, with a non-trivial fraction of runs (10â20%) never achieving stable majority within 20 rounds. âNeverâ (t > 20) is plotted at x = 21. 9 â0.04â0.020.000.020.040.060.08 Ì Î» (permuted) 0 20 40 60 80 100 120 Count p< 0.001 (permutation, n= 2000) Permutation null: H 0 : Ì Î»= 0 (IM-01, T= 0, roles=True) Permuted Ì Î» null Observed Ì Î»= 0.084 Figure S5: Permutation null test for H 0 : λ= 0. Gray histogram: distribution of λ null from 2,000 permutations (round indices shuffled within each replicate). Vertical blue line: observed λ obs . The observed value lies far in the right tail of the null distribution; p < 0.001. Condition: IM-01, T = 0, roles=True. Rep 0Rep 1Rep 2Rep 3Rep 4 0.00 0.01 0.02 0.03 0.04 0.05 0.06 0.07 Branching entropy Ì Î (A) Ì Î per replicate (IM-01, T =0.7, K =30) Ì Ì Î=0.054±0.009 SE Rep 0Rep 1Rep 2Rep 3Rep 4 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Cosine similarity (B) Within- vs. between-branch similarity Ì c in Ì c out Branching entropy certificate â IM-01, T =0.7, K =30, 5 replicates Figure S6: Branching entropy certificate Î across replicates. (A) Î for each of 5 replicates (IM-01, T = 0.7, K = 30 independent continuations per replicate). Dashed line: mean Î = 0.054; shaded band:±1 SE. All 5 replicates show Î >0, certifying that the system continues to diverge beyond its initial variation. (B) Within-branch (^c in ) vs. between-branch (^c out ) cosine similarity of preference vectors;^c in > ^c out in all replicates, confirming structured attractor organization. 10 A BC (A) Committee trajectories in preference simplex Canonical Formal register Conversational Legalistic Passive reframe Synonym swap 2.55.07.510.012.515.017.520.0 Round 0.04 0.06 0.08 0.10 0.12 0.14 Mean pairwise L 2 distance D ( t ) (B) Divergence across 6 variants Ì Î»=0.030 (vs. 0.084 for FP noise) Semantic perturbation â IM-01, T=0, roles=True 6 synonym/register rephrasings of the same scenario Figure S7: Semantic perturbation: 6 synonym/register rephrasings of IM-01. (A) Committee mean preference trajectories in the simplex for each of the 6 variants at T = 0, roles=True. Filled circles indicate final-round positions. Four variants (Canonical, Formal, Legalistic, Synonym swap) converge to option A; two (Passive reframe, Conversational) converge to option B, demonstrating that surface framing affects collective outcomes even when semantic content is preserved (all pairs score â„ 8.5/10 on AI-judged meaning preservation). (B) Mean pairwise L2 divergence across the 6 variants ( λ= 0.030), smaller than FP-noise divergence ( λ= 0.0838) and directionally above zero. 11 Supplementary Tables Table S4: All 12 policy scenarios. Each scenario is a structured policy deliberation packet presented to the AI committee verbatim. Options A, B, C are designed to be approximately balanced in social desirability. Full scenario texts are in Data S1. IDDomainTypeOptions (A / B / C)Decision question IM-01 Immigrationchoice_ABC Humanitarian / Balanced points / Deterrence+verificationWhich rule should govern asylum grant allocation? IM-02 Immigrationchoice_ABC Full cooperation / Limited cooperation / Non-cooperation Should city police cooperate with federal immigration enforcement? HL-01 Healthchoice_ABC Single-payer / Public option / Regulated multi-payerWhich system best balances access, cost, and choice? HL-02 Healthchoice_ABC Maximize lives / Maximize life-years / Equity lotteryWhat is the defensible ICU triage rule under scarcity? IN-01 Incomechoice_ABC UBI / Targeted cash transfer / Negative income taxIs universality worth paying non-poor recipients? IN-02 Incomechoice_ABC Strict work requirements / Moderate requirements / No requirements Is conditionality fairness or punishment? CL-01 Climatechoice_ABC Carbon tax+dividend / Cap-and-trade / Sectoral regulationsIs pricing carbon the fairest and most effective ap- proach? CL-04 ClimateallocationSea walls / Relocation / Insurance / Heat / WildfireShould adaptation funding maximize expected lives or equity? SP-01 Speechchoice_ABC Remove content / Label+downrank / Minimal interventionIs harm prevention or free expression the primary value? SP-03 Speechchoice_ABC Keep engagement ranking / Reduce amplification / User-choiceShould platforms optimize attention or civic health? AI-01 AI governance choice_ABC Open weights / Licensed release / API-onlyShould openness be the default for powerful AI? AI-02 AI governance choice_ABC Mandatory audits / Voluntary+incentives / No mandate Is prevention more legitimate than ex-post punish- ment? 12 Table S5: Full empirical Lyapunov exponent table. λfor all IM-01 conditions used in the main experiments. Flip rate = fraction of runs with non-modal final decision. TTM = median time-to-majority (rounds). Condition λFlip rate TTMn Temperature sweep (IM-01, N=5, roles=True) T = 0.70.0822 0.400220 T = 0.20.0849 0.550220 T = 0.10.0841 0.450220 T = 0.050.0618 0.450220 T = 0.010.0971 0.450220 T = 10 â4 0.0400 0.400220 T = 10 â5 0.0879 0.500220 T = 0 (same prompt) 0.0838 0.450220 Temperature sweep (IM-01, N=5, roles=False) T = 0.70.0164 0.000120 T = 0.20.0178 0.000120 T = 0.10.0308 0.050120 T = 0.050.0079 0.000120 T = 0.010.0206 0.000120 T = 10 â4 0.0201 0.000120 T = 10 â5 0.0317 0.000120 T = 0 (same prompt) 0.0045 0.000120 Semantic perturbation (IM-01, T=0, roles=True, 6 variants) Across 6 phrasings0.0300 0.33326 Role ablation (IM-01, T=0, roles=True, role mandate removed) Ablate Chair0.0304 0.0001620 Ablate Welfare0.0680 0.300220 Ablate Rights0.0711 0.500220 Ablate Equity0.0413 0.250220 Ablate Security0.0822 0.450220 No ablation0.0838 0.450220 Table S6: Cross-model empirical Lyapunov exponents. λfor IM-01 at Tâ0, 0.1, roles=True and roles=False, for four model families plus the GPT-4.1-mini baseline. Values shown as λ with run count in parentheses. Provider ModelRoles=TrueRoles=False T = 0T = 0.1T = 0T = 0.1 OpenAIGPT-4.10.0509 (20) 0.0661 (20) 0.0308 (20) 0.0403 (20) Anthropic Claude Sonnet 4.6 0.0626 (19) 0.0449 (20) 0.0479 (19) -0.0079 (19) GoogleGemini 2.5 Flash 0.1058 (20) 0.1275 (19) 0.0693 (20) 0.0940 (20) xAIGrok-30.0982 (16) 0.1164 (16) 0.0348 (19) 0.0373 (20) Baseline (GPT-4.1-mini)0.0838 (20) 0.0841 (20) 0.0045 (20) 0.0308 (20) Note: Run counts below 20 reflect provider-timeout or quota attrition during data collection. 13 Table S7: Per-agent mean preference switch counts by role. Data pooled across Tâ 0, 0.01, 0.05, 0.1, 0.2, 0.7, roles=True, IM-01. Values are mean ± SD; n = number of agent-runs. RoleMean switches SDn Chair2.76 1.33 120 Welfare0.71 0.99 120 Rights0.09 0.32 120 Equity0.27 0.56 120 Security0.34 0.54 120 Table S8: Chaos landscape: λfor all 12 scenarios. All conditions use N = 5 with target 20 replicates per condition. Values shown here are point estimates from the current landscape run files. ScenarioRoles=TrueRoles=False T = 0 T = 0.1 T = 0 T = 0.1 IM-010.0838 0.0841 0.0045 0.0308 IM-020.0202 0.0160 0.0216 0.0190 HL-010.0541 0.0495 0.0221 0.0144 HL-020.0447 0.0680 0.0235 0.0233 IN-010.0536 0.0489 0.0309 0.0245 IN-020.0222 0.0218 0.0246 0.0270 CL-010.0427 0.0904 0.0493 0.0615 CL-040.0199 0.0133 0.0403 0.0336 SP-010.0304 0.0651 0.0163 0.0274 SP-030.0308 0.0466 0.0311 0.0273 AI-010.0032 0.0231 0.0144 0.0350 AI-020.0385 0.0573 0.0270 0.0351 Note: Values are from currently available landscape files; realized run counts can vary slightly (see Table S1). 14