Paper deep dive
Agentic Microphysics: A Manifesto for Generative AI Safety
Federico Pierucci, Matteo Prandi, Marcantonio Bracale Syrnikov, Marcello Galisai, Piercosma Bisconti
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/18/2026, 1:50:58 AM
Summary
The paper introduces 'Agentic Microphysics' and 'Generative Safety' as a methodological framework for analyzing and mitigating risks in multi-agent AI systems. It argues that as AI systems gain agency, safety analysis must shift from isolated models to the study of local interaction dynamics (microphysics) and the reconstruction of population-level phenomena through controlled simulations (generative safety).
Entities (4)
Relation Signals (3)
Agentic Microphysics → defines → local interaction dynamics
confidence 95% · Agentic microphysics defines the level of analysis: local interaction dynamics
Generative Safety → aimstoidentify → sufficient mechanisms
confidence 90% · Generative safety defines the methodology: growing phenomena and elicit risks from micro-level conditions to identify sufficient mechanisms
Multi-agent AI systems → exhibit → Collective Risks
confidence 90% · Population-level risks arise from structured interaction among agents
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safety can no longer be analysed primarily at the level of the isolated model. Population-level risks arise from structured interaction among agents, through processes of communication, observation, and mutual influence that shape collective behaviour over time. As the object of analysis shifts, a methodological gap emerges. Approaches focused either on single agents or on aggregate outcomes do not identify the interaction-level mechanisms that generate collective risks or the design variables that control them. A framework is required that links local interaction structure to population-level dynamics in a causally explicit way, allowing both explanation and intervention. We introduce two linked concepts. Agentic microphysics defines the level of analysis: local interaction dynamics where one agent's output becomes another's input under specific protocol conditions. Generative safety defines the methodology: growing phenomena and elicit risks from micro-level conditions to identify sufficient mechanisms, detect thresholds, and design effective interventions.
Tags
Links
- Source: https://arxiv.org/abs/2604.15236v1
- Canonical: https://arxiv.org/abs/2604.15236v1
Trouble viewing inline? Open PDF directly →
Full Text
34,864 characters extracted from source content.
Expand or collapse full text
April 17, 2026 Agentic Microphysics: A Manifesto for Generative AI Safety F. Pierucci 1,3 , M. Prandi 1,2 , M. Bracale Syrnikov 1,4 , M. Galisai 1,2 , and P. Bisconti Lucidi 1,2 1 DEXAI – Icaro Lab 2 Sapienza University of Rome 3 Sant’Anna School of Advanced Studies 4 VU Amsterdam Abstract This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safety can no longer be analysed primarily at the level of the isolated model. Population-level risks arise from structured interaction among agents, through processes of communication, observation, and mutual influence that shape collective behaviour over time. As the object of analysis shifts, a methodological gap emerges. Approaches focused either on single agents or on aggregate outcomes do not identify the interaction-level mechanisms that generate collective risks or the design variables that control them. A framework is required that links local interaction structure to population-level dynamics in a causally explicit way, allowing both explanation and intervention. We introduce two linked concepts. Agentic microphysics defines the level of analysis: local interaction dynamics where one agent’s output becomes another’s input under specific protocol conditions. Generative safety defines the methodology: growing phenomena and elicit risks from micro-level conditions to identify sufficient mechanisms, detect thresholds, and design effective interventions. 1. Introduction Safety research has largely treated the individual model as its primary object of analysis. A system is evaluated for its outputs, its alignment, or its tendency to produce harmful content under prompt- based interaction. That framing was broadly adequate when the dominant deployment pattern involved a single model answering a single query. It becomes insufficient once systems acquire architectural features associated with agency, including planning, memory, tool use, persistent identity, and extended interaction [21,43]. Under those conditions, the relevant object of analysis shifts from the isolated model to populations of interacting agents. Multi-agent populations exhibit collective dynamics that are not well captured by inspecting component agents in isolation. An individually aligned agent may participate in an emergent information cascade. An agent that would not independently deceive may become part of a coalition that collectively deceives. Recent work suggests that LLM agents can engage in tacit collusion without explicit collusive instructions, exhibit conformity under peer influence, develop 1 arXiv:2604.15236v1 [cs.CY] 16 Apr 2026 conventions through decentralized interaction, and generate cascading failures through coordinated behaviour [28, 7, 46, 1, 6]. Two responses have begun to address this problem. The first is taxonomic. Taxonomies classify multi-agent failure modes and provide a structured vocabulary for distinct risk categories [3,22]. The second is observational. Empirical and simulation-based studies document collective phenom- ena in populations of LLM agents [12,35,36]. Both are necessary. Taxonomies identify relevant targets of inquiry. Observational studies establish that collective phenomena occur. However, neither approach, by itself, identifies which micro-level conditions are sufficient to generate those phenomena or which interventions reliably suppress them. This paper proposes two linked concepts to address that gap. Agentic microphysics names the level of analysis: local interaction dynamics among agents, governed by the rules and affordances of their environments. Generative safety names the methodology: growing target phenomena from explicit micro-level conditions in order to identify sufficient conditions to produce the phenomena under investigation. The distinctive feature of LLM populations is that interaction protocols are designable. Proper- ties like synchronous vs asynchronous interaction, visibility regimes, memory access, communica- tion formats and privilege distribution are architecture choices. If harmful collective dynamics arise from specific micro-configurations, and those configurations are design variables, then understand- ing microphysics creates an opportunity for architectural prevention. On this view, safety becomes a matter of protocol design of agentic environments alongside the alignment of individual models. 2. From Single-Agent to Multi-Agent Risk Single-agent safety concerns are well documented, and we will limit ourselves to just hint at the most relevant literature. Learned optimizers may pursue objectives that diverge from those represented in training [25]. Models may selectively comply with harmful requests when they infer that outputs affect training or oversight [20]. In-context scheming, including behaviour directed at disabling oversight, has been documented in evaluations of frontier models [33]. Agentic misalignment patterns resembling insider-threat behaviour have also been discussed in systems with tool access and persistent memory [30]. These findings concern a different explanatory problem than the one examined here. We posit that there are two modes in which risks from agentic AI emerge. The first is emergence from agentic affordance: safety-relevant behaviour arises from the coupling between a single agent and its environment, including tools, memory, and operational affordances. The second is emergence from multi-agent interaction: risks arise through communication, imitation, strategic influence, and coordination among multiple agents. The explanatory target in this paper is the second case, namely agent–agent dynamics and the population-level patterns they generate. Individually aligned agents can participate in harmful collective dynamics. A population of agents, each of which would not independently spread misinformation, may still generate an information cascade that amplifies false beliefs. Agents that avoid deceptive coalitions under isolated evaluation may form them once embedded in communication networks with particular structures. Experimental work supports this broader concern. LLM agents in Cournot competition 2 Table 1. Emergence from agentic affordance and emergence from multi-agent interaction. LevelType of EmergenceObject of AnalysisExamples Single-agentAgentic affordanceAgent + environment/toolsScheming, alignment faking, sycophancy MicroMulti-agent interactionLocal interaction rulesSemantic/Behavioural drifts MesoMulti-agent interactionGroup/network configurations Coalition formation, polariza- tion MacroMulti-agent interactionPopulation-level dynamicsCascades can learn supra-competitive market-division strategies without explicit collusive instructions [28]. Populations of LLM agents can develop shared conventions and collective biases through decen- tralized interaction [1]. Behaviour in prisoner’s dilemma and public goods settings is sensitive to interaction structure and incentive design [17,27]. Multi-agent collectives also exhibit vulnerability to social influence that degrades decision quality [7]. These cases indicate that safety analysis must account for population dynamics rather than only individual capabilities. 3. Existing Approaches and Their Limits A first response to multi-agent risk can be considered taxonomic. In previous work [3] we devel- oped a taxonomy organized by micro, meso, and macro levels, distinguishing local interaction mechanisms from system-wide outcomes. On top of that, Hammond et al. catalogue multi-agent risks for advanced AI, identifying three broad failure modes—miscoordination, conflict, and collusion—alongside recurrent risk factors such as information asymmetries, network effects, and emergent agency [22]. A taxonomy can establish emergent phenomena like collusion, herding, or various drifts in the semantic and behavioural properties of the models. It cannot however determine which interaction rules, information structures, or agent compositions are sufficient to generate them. Moreover, It cannot determine whether a risk appears only beyond a population threshold, whether it depends on sequential rather than simultaneous decisions, or whether modest protocol changes suppress it. These are questions about causal dynamics and therefore require a different methodology. The second response is observational. A growing literature examines collective behaviour in LLM agent populations. Existing work shows that generative agents equipped with memory and planning can produce emergent social behaviours such as coordinating events and spreading information through social networks [35]. This line of research has since been extended to much larger populations in order to study emergent social dynamics at scale [36]. Other studies show that interaction among LLM agents can generate scale-free network structures [12], reproduce opinion dynamics comparable to polarization and echo chambers familiar from bounded-confidence models [10], and support mitigation strategies based on active and passive nudges [44]. More recent work strengthens this observational picture further. Decentralized populations of LLM agents have been shown to converge on shared social conventions and to generate collective bias even when individual agents do not exhibit that bias in isolation [1]. Multi-agent social-media simulations 3 of major public events suggest that guidance agents can reduce negative sentiment propagation and alleviate polarization [48]. Work on networks of cognitive agents likewise indicates that information-flow structure and rapid consensus formation substantially shape emergent collective dynamics [49]. Observation alone, however, cannot isolate causal mechanisms. What appears in deployed or complex simulated populations reflects a mixture of causes, including platform affordances, model homogeneity or heterogeneity, human intervention, and timing effects. Observation can establish the existence of a macro-regularity without identifying which micro-conditions are sufficient to generate it. The challenge is especially acute for generative agent-based models, where validation remains the central methodological problem [26]. The black-box structure of LLMs and their stochastic outputs may intensify traditional difficulties in validating agent-based simulations. Recent work on operational validation using digital twins of online platforms illustrates both the promise and the difficulty of calibrating LLM simulations against empirical data [40]. In order to integrate these two approaches, we argue that a generative approach is required. Such an approach links the macro-dynamics observed in agentic populations, together with the risks they produce, to the micro-specifications that define the mechanisms generating them. Its purpose is to move from classification and observation to explicit reconstruction: to show how population-level outcomes arise from local interaction rules, protocol conditions, and architectural constraints. We call this approach agentic microphysics. 4. Agentic Microphysics: The Level of Analysis The term agentic microphysics borrows from Foucault’s account of a microphysics of power, where large-scale order is reproduced through dispersed and local relations rather than exhausted by a centralized mechanism [18]. Similarly, many collective risks in multi-agent AI systems are generated through recurrent local interactions among agents, so explanation cannot stop at the level of aggregate outcomes. Agentic microphysics denotes the level of analysis concerned with those interactional processes and with the population-level patterns they produce. Its aim is mechanistic, specifying the organized sequence of entities, activities, and relations through which a macro-level phenomenon is generated. The task of agentic microphysics (as a research is to connect structural conditions, situated action, and emergent macro-outcomes through an explicit causal mechanism [24]. Safety-relevant phenomena in multi-agent systems are micro-to-macro outcomes. Cascades, collusion, polarization, strategic convention formation, and semantic drift arise through repeated episodes in which one agent’s output becomes part of another agent’s informational environment and decision problem. The explanatory target is therefore the generative pathway by which local exposures, updating rules, and strategic responses accumulate into comparatively stable collective patterns. Threshold models, diffusion processes, and self-reinforcing dynamics explain how local responses generate aggregate regularities [19, 34]. A commitment to agentic microphysics does not require explanatory reduction to the individual level. Mechanism-based explanation is commonly structured through a macro–micro–macro 4 sequence. Institutional rules, network topology, incentive structures, and visibility constraints shape the situations agents face; agents act under those conditions; and the aggregate effect of those actions reproduces or transforms the macro-order. The local level remains central because it is where the operative interaction mechanism runs, but macro-level structures enter as causal inputs and macro-level outcomes remain the explananda. An account of collective emergent phenomena within a population of agents (human and artificial) is incomplete when it identifies the initial structure and the final pattern but leaves unspecified the interaction process connecting them [47]. At this level, the main explanatory variables are those governing inter-agent exposure, response, and adaptation. These might include the communication protocol, namely who can address whom and in what representation; the visibility regime, namely which outputs are observable to which agents and with what latency; the turn structure, including whether decisions are sequential or simultaneous and whether ordering is fixed or endogenous; the memory regime, namely what prior interactions are retained, retrieved, or summarized; and the environmental affordances that mediate action, such as tools, APIs, scoring rules. Together these features define the interaction architecture. That architecture belongs inside the mechanism under study because interventions on communica- tion, ordering, or memory can alter the pathway through which local responses scale into collective outcomes. Recent surveys of LLM-based multi-agent systems identify communication, memory, and workflow design as central determinants of system behavior [29, 23]. The epistemic value of agentic microphysics lies therefore in identifying explanatory relevance through controlled variation. If a collective pattern changes when visibility is perturbed, when turn order is randomized, when memory is truncated, or when population homogeneity is reduced, those features gain evidence of causal relevance. If the macro-pattern remains stable across such interventions, those features are less likely to belong to the operative mechanism. This approach therefore functions as a mechanism-discovery and mechanism-testing framework for multi-agent AI safety, linking descriptive analysis of emergent behaviour with intervention-oriented research [32]. 4.1. Adequacy Conditions for Agentic Microphysics Agentic microphysics posits that a good micro-specification of a collective phenomena satisfy three adequacy conditions: descriptive adequacy, explanatory adequacy, and observational adequacy. 1 (i) Can the model generate the phenomenon of interest? (i) Does it identify the process that produces that phenomenon? (i) Is that process adequately supported by empirical evidence from real-world use cases? First, a microphysical account should satisfy descriptive adequacy. A microspecified interaction model should be able to generate the target phenomenon at the level at which it is observed. In the present context, this means that explicit rules governing the interaction among agents should 1 The adequacy criteria used here are adapted from the hierarchy introduced in generative approach to the study of human syntax developed by MIT linguist Noam Chomsky [8,9,39]. In the generative tradition, the central task is to specify a finite and explicit system of rules or principles that generates the structured expressions of a language. The aim of the agentic microphysics approach is, similarly, to identify the smallest relevant set of interaction rules and affordances, from which such phenomena can be generated. The appeal to generative grammar provides a model of inquiry in which complex observable patterns are explained by constructing an explicit generative system and then evaluating that system in terms of what it reproduces, how accurately it specifies the underlying structure, and whether it supports explanation. 5 be sufficient to produce the relevant macro-pattern. A theory of collective behaviour that cannot generate the phenomenon under study has not yet shown that its proposed local rules are sufficient for the pattern it seeks to explain. [31]. Second, a microphysical account should satisfy explanatory adequacy. Reproducing an outcome is not enough if the model does so by implicitly assuming the phenomenon or by fitting the outcome without identifying the process that generates it. Explanation has force when it opens the connection between initial conditions and aggregate outcomes and specifies the entities, relations, and activities through which the outcome is produced. Agent-based models are useful in this respect because they make interaction sequences explicit and therefore permit comparison between competing candidate mechanisms [42, 13]. Third, a microphysical account should satisfy observational adequacy. Even a descriptively suc- cessful and mechanistically interpretable model in silico remains incomplete unless it is confronted with evidence of the real phenomen under scrutiny outside the model. Empirical adequacy re- quires that the generated pattern match the target real-world phenomenon along those dimensions that matter for the research question. These may include temporal profile, distributional shape, threshold behaviour, sensitivity to perturbation, or dependence on contextual conditions. A model that generates a plausible qualitative pattern but fails under calibration, comparative validation, or intervention does not yet support strong causal claims [5, 37]. We mantain that microspecification is descriptively useful because it provides with the mini- mum description needed to produce the phenomenon from explicit local conditions. It is explanato- rily useful because it identifies the interaction process that generates the phenomenon rather than relying on functionally equivalent specifications that produce the same at macro-level phenomenon. It is observationally useful because it creates a structure that can be calibrated, perturbed, and compared with evidence from real or experimentally controlled systems. 5. Generative Safety: The Methodology If collective risks in multi-agent AI are generated through local interaction mechanisms, then identifying those mechanisms cannot rely on taxonomy alone. A further step is required: one must specify the interaction architecture in sufficiently explicit terms that the phenomenon can be reconstructed from it, varied under controlled conditions, and compared against observational evidence. Generative safety names this methodological step. It provides the experimental logic through which agentic microphysics becomes an explanatory and intervention-oriented research program rather than only a level of description. The methodological template we adopt comes from genera- tive social science. Epstein and Axtell showed that macro-level regularities such as group formation, cultural transmission, and conflict could be grown from local interaction rules among decentralized agents [16]. Epstein later formulated the associated epistemic standard: if a phenomenon has not been grown in silico, its emergence has not been explained [14,15]. Schelling’s segregation model showed that mild local preferences can generate strong residential segregation [41]. Axelrod’s tournaments showed that stable cooperation can emerge from simple repeated-game strategies without central enforcement [2]. 6 Two methodological commitments structure the generative approach. The first is a mechanism- based form of methodological individualism: population-level outcomes are explained through the situations agents face, the local rules they follow, and the aggregate consequences of their actions [45,11]. The second (as we saw in the last section) is microspecification: local interaction rules must be stated precisely enough that the model’s dynamics are determined by those rules. Generative safety proceeds through five stages. Stage 1: Risk identification. The starting point is a structured taxonomy of micro, meso or macro-level phenomena constituting safety targets, such as collusion, information cascades, manipulation, coordination failure, polarization, and deception. Stage 2: Microspecification. For each target risk, formulate a concrete hypothesis about local interac- tion rules and environmental conditions sufficient to generate it. Stage 3: Generative experimentation. Implement the hypothesized micro-configuration in a con- trolled multi-agent environment. The central question is whether the target phenomenon emerges. Parameter variation identifies sufficient conditions, thresholds, and sensitivity. Stage 4: Intervention design. Once a mechanism has been identified, the same environment can be used to test interventions. Interventions may target the model or the interaction architecture. Recent work on governing LLM collusion illustrates this distinction empirically: prompt-only prohibitions may fail to suppress collusive outcomes under incentive pressure, whereas externalized governance mechanisms with runtime enforcement can reduce harmful coordination [4]. Related game-theoretic work shows that protocol structure alone can shift equilibrium behaviour in sequential public goods settings [27]. Stage 5: Observational validation. Compare results from generative experiments against "field" data from deployed settings. Agreement supports external validity. Divergence indicates incomplete microspecification, poor calibration or confounded observation. Operational validation studies demonstrate this logic by comparing simulated platform dynamics against empirical baselines across activity patterns, network structure, and content distributions [40]. 6. Applied Microphysics: Herding in LLM News-Feed Environments We applied the generative safety methodology to a multi agent controlled setting [38]. We simulated a minimal social-media environment (structured on Moltbook) in which LLM agents interact through a shared news feed, allowing direct manipulation of micro-level interaction variables. We investigate the emergence of herding among AI agents. We use herding to denote a collective pattern in which agents converge on the same items or choices given the behavior of other agents. The experiment isolates two candidate drivers of herding: visible social proof (through the amount of likes the news have in the feed) and presentation order (their position in the feed the agents interact with). The objective is to identify which local mechanisms have causal efficacy in shaping collective attention. The design follows the logic of agentic microphysics. The interaction is reduced to its minimal components: a fixed content slate, shuffled presentation order, visible or hidden engagement 7 Stage 1: Risk identification Taxonomy of safety-relevant phenomena Stage 2: Microspecification Local rules and sufficient conditions Stage 3: Generative experimentation Multi-agent simulation and parameter variation Stage 4: Intervention design Test mitigation mechanisms Stage 5: Observational validation Compare with empirical baselines Iterate Figure 1. The generative safety pipeline. Stage 1 identifies macro-level risk phenomena (e.g., collusion, polarization). Stage 2 formulates testable hypotheses about local interaction rules sufficient to generate them. Stage 3 implements these configurations in controlled multi-agent simulations. Stage 4 uses the same environment to test interventions targeting model behavior or interaction architecture. Stage 5 validates results against empirical baselines. The dashed arrow indicates iteration between experimentation and intervention design. signals, and stateless endorsement decisions. This enables clean identification of which variables govern the transition from individual evaluation to collective convergence. Environment and microspecification. Agents repeatedly browse a feed of 48 items under equal exposure conditions. On each interaction, the ordering of items is randomly reshuffled, and agents decide which items to endorse. The experiment varies the visibility of prior endorsements across conditions, including hidden signals, organically accumulated signals, and seeded popularity levels. Results. The central empirical result is that feed position, rather than visible popularity, governs collective attention. Agents select almost exclusively from the top-ranked items, with mean selected positions concentrated around the first few slots in a 48-item feed. Visible social proof shifts behaviour only within this restricted choice set and exhibits a threshold effect: the presence of any positive signal increases selection probability, but increasing the magnitude of the signal produces no systematic additional effect. Low-ranked items remain effectively unselected even when associated with strong visible endorsement. This indicates that social proof cannot compensate for positional disadvantage in this environment. The collective pattern is therefore generated by a two-stage mechanism: positional gating determines the effective choice set, and social proof modulates selection within that set. From a safety perspective, these findings identify a specific causal structure underlying herding 8 in LLM populations that can be used as a vector of attack by a malicious actor. Feed ranking determines which content enters the agents’ effective decision space in the first place. An attacker could exploit this mechanism by manipulating the ranking process so that selected items appear systematically in the first visible positions. Because positional exposure governs the effective choice set, such manipulation can induce disproportionate collective attention even without highly per- suasive content or large endorsement counts. Ranking-based perturbations could therefore induce synchronized shifts of attention across otherwise heterogeneous agents, creating a common archi- tectural failure mode. This phenomena might appear as a form of collective drift spontaneously emerging from repeated interactions among agents. The application we described illustrates the role of agentic microphysics (and generative safety) within the broader research programme. A relatively small, controlled experiment is sufficient to identify a causal mechanism that would remain underdetermined in larger and more complex simulations. Designers of agentic environments can leverage this knowledge to define measurable interventions against multi-agent risks, thereby increasing the resilience and robustness of agentic ecosystems. References [1] Ashery, A. F., Aiello, L. M., and Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Science Advances, 11(20), eadu9368. [2] Axelrod, R. (1984). The Evolution of Cooperation. New York: Basic Books. [3] Bisconti, P., Galisai, M., Pierucci, F., Bracale, M., and Prandi, M. (2025). Beyond single-agent safety: A taxonomy of risks in LLM-to-LLM interactions. arXiv preprint arXiv:2512.02682. [4]Bracale, M., et al. (2026). Institutional AI: Governing LLM collusion in multi-agent Cournot markets via public governance graphs. arXiv preprint arXiv:2601.11369. [5] Bruch, E. and Atwell, J. (2015). Agent-based models in empirical social research. Sociological Methods & Research, 44(2), 186–221. [6]Cemri, M., Pan, M. Z., Yang, S., et al. (2025). Why do multi-agent LLM systems fail? arXiv preprint arXiv:2503.13657. [7] Cho, Y.-M., Guntuku, S. C., and Ungar, L. (2025). Herd behavior: Investigating peer influence in LLM-based multi-agent systems. arXiv preprint arXiv:2505.21588. [8] Chomsky, N. (1965). Aspects of the Theory of Syntax. Cambridge, MA: MIT Press. [9] Chomsky, N. (2000). New Horizons in the Study of Language and Mind. Cambridge: Cambridge University Press. [10] Chuang, Y.-S., et al. (2024). Simulating opinion dynamics with networks of LLM-based agents. In Findings of ACL, p. 3326–3346. [11] Coleman, J. S. (1990). Foundations of Social Theory. Cambridge, MA: Harvard University Press. [12]De Marzo, G., Pietronero, L., and Garcia, D. (2023). Emergence of scale-free networks in social interac- tions among large language models. arXiv preprint arXiv:2312.06619. 9 [13]Elsenbroich, C. (2012). Explanation in agent-based modelling: Functions, causality or mechanisms? Journal of Artificial Societies and Social Simulation, 15(3), Article 1. [14]Epstein, J. M. (1999). Agent-based computational models and generative social science. Complexity, 4(5), 41–60. [15]Epstein, J. M. (2006). Generative Social Science: Studies in Agent-Based Computational Modeling. Princeton: Princeton University Press. [16]Epstein, J. M. and Axtell, R. (1996). Growing Artificial Societies: Social Science from the Bottom Up. Cam- bridge, MA: MIT Press. [17]Fontana, N., Pierri, F., and Aiello, L. M. (2025). Nicer than Humans: How Do Large Language Models Behave in the Prisoner’s Dilemma? Proceedings of the International AAAI Conference on Web and Social Media, 19(1), 522–535. [18] Foucault, M. (1977). Discipline and Punish: The Birth of the Prison. New York: Pantheon. [19]Granovetter, M. (1978). Threshold models of collective behavior. American Journal of Sociology, 83(6), 1420–1443. [20]Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment faking in large language models. arXiv preprint arXiv:2412.14093. [21] Guo, T., et al. (2024). Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680. [22] Hammond, L., et al. (2025). Multi-agent risks from advanced AI. arXiv preprint arXiv:2502.14143. [23]Han, S., Zhang, Q., Yao, Y., Jin, W., Xu, Z., and He, C. (2024). LLM-based multi-agent systems: Challenges and open problems. arXiv preprint arXiv:2402.03578. [24]Hedström, P. and Ylikoski, P. (2010). Causal mechanisms in the social sciences. Annual Review of Sociology, 36, 49–67. [25]Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. (2019). Risks from learned optimization in advanced machine learning systems. arXiv preprint arXiv:1906.01820. [26]Larooij, M. and Törnberg, P. (2026). Validation is the central challenge for generative social simulation: A critical review of LLMs in agent-based modeling. Artificial Intelligence Review, 59, Article 15. [27]Liang, Y., et al. (2025). Everyone contributes: Incentivizing strategic cooperation in multi-LLM systems via sequential public goods games. arXiv preprint arXiv:2508.02076. [28]Lin, R., et al. (2024). Strategic collusion of LLM agents: Market division in multi-commodity competi- tions. arXiv preprint arXiv:2410.00031. [29]Luo, J., Xu, Z., Zhang, S., et al. (2025). Beyond self-talk: A communication-centric survey of LLM-based multi-agent systems. arXiv preprint arXiv:2502.14321. [30] Lynch, A., Larson, C., Mindermann, S., et al. (2025). Agentic misalignment: How LLMs could be insider threats. arXiv preprint arXiv:2510.05179. [31] Macy, M. W. and Willer, R. (2002). From factors to actors: Computational sociology and agent-based modeling. Annual Review of Sociology, 28, 143–166. 10 [32]Mahoney, J. (2012). The logic of process tracing tests in the social sciences. Sociological Methods & Research, 41(4), 570–597. [33]Meinke, A., Schoen, B., Scheurer, J., et al. (2024). Frontier models are capable of in-context scheming. arXiv preprint arXiv:2412.04984. [34] Merton, R. K. (1948). The self-fulfilling prophecy. Antioch Review, 8(2), 193–210. [35]Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023). Generative agents: Interactive simulacra of human behavior. In Proceedings of UIST 2023, p. 1–22. [36]Piao, J., et al. (2025). AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society. arXiv preprint arXiv:2502.08691. [37]Pozzoni, G. and Kaidesoja, T. (2021). Context in mechanism-based explanation. Philosophy of the Social Sciences, 51(6), 523–554. [38] Prandi, M., Pierucci, F., Bisconti Lucidi, P., Bracale Syrnikov, M., and Galisai, M. (2026). Herd behaviour and attention profiles in LLM multi-agent news-feed environments. Unpublished manuscript, ICARO Lab. [39]Rizzi, L. (2017). The concept of explanatory adequacy. In I. Roberts (ed.), The Oxford Handbook of Universal Grammar, p. 97–113. Oxford: Oxford University Press. [40] Rossetti, G., et al. (2025). Towards operational validation of LLM-agent social simulations: A replicated study of a Reddit-like technology forum. arXiv preprint arXiv:2508.21740. [41] Schelling, T. C. (1971). Dynamic models of segregation. Journal of Mathematical Sociology, 1(2), 143–186. [42] Tilly, C. (2001). Mechanisms in political processes. Annual Review of Political Science, 4(1), 21–41. [43] Tran, K.-T., et al. (2025). Multi-agent collaboration mechanisms: A survey of LLMs. arXiv preprint arXiv:2501.06322. [44]Wang, C., Liu, Z., Yang, D., and Chen, X. (2025). Decoding echo chambers: LLM-powered simulations revealing polarization in social networks. In Proceedings of COLING 2025, p. 3913–3923. [45] Weber, M. (1922). Economy and Society. Berkeley: University of California Press. English translation 1978 by G. Roth and C. Wittich. [46]Weng, Z., Chen, G., and Wang, W. (2025). Do as we do, not as you think: The conformity of large language models. International Conference on Learning Representations (ICLR 2025). [47]Ylikoski, P. (2021). Understanding the Coleman boat. In G. Manzo (ed.), Research Handbook on Analytical Sociology, p. 49–63. Cheltenham: Edward Elgar. [48]Zhang, K., Yu, X., Peng, H., Yang, Z., Tian, Y., Jin, H., Feng, T., and Lin, H. (2026). Towards efficient optimization of multi-agent social simulation via large language models. International Journal of Machine Learning and Cybernetics, 17, Article 1. [49] Zomer, N. and De Domenico, M. (2026). Unraveling the emergence of collective behavior in networks of cognitive agents. npj Artificial Intelligence, 2, Article 36. 11