Paper deep dive
A three-dimensional typology of agency for advanced AI systems
Willem Fourie
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/21/2026, 4:04:40 AM
Summary
The paper proposes a three-dimensional typology for analyzing the agency of advanced AI systems, consisting of nature (moral/legal), mode (individual/collective), and locus (human/non-human). This framework generates eight possible instantiations of agency, categorized as conventional, contested, or controversial. The authors argue that distinguishing legal from moral agency is crucial for governance, particularly regarding non-human, individual legal agency in the context of instrumental goals and loss of control.
Entities (15)
Relation Signals (9)
Willem Fourie â affiliatedwith â Stellenbosch University
confidence 95% ¡ Affiliation: [1em] School for Data Science and Computational Thinking, Stellenbosch University
Legal Agency â distinguishedfrom â Moral Agency
confidence 95% ¡ The typology separates legal from moral agency
Anthropic Mythos 5 â developedby â Anthropic
confidence 90% ¡ reported behaviour by Anthropicâs unreleased Mythos 5 model
Corporation â exemplifies â Collective, Legal, Non-human Agency
confidence 90% ¡ The corporation is the paradigmatic example. Agency is ascribed to the juridical entity itself
Anthropic Mythos 5 â exhibitedbehavior â Instrumental Goals
confidence 90% ¡ The deception displayed in this incident can be understood as anecdotal evidence of instrumental goals
UK AI Security Institute â reportedincident â Anthropic Mythos 5
confidence 90% ¡ The Institute... reported behaviour by Anthropicâs unreleased Mythos 5 model
Anthropic Mythos 5 â testedby â UK AI Security Institute
confidence 90% ¡ The Institute... reported behaviour by Anthropicâs unreleased Mythos 5 model during cyber-security testing.
List â defined â Intentional Agent
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions in philosophy, ethics, legal theory and sociology, we develop a typology of agency for frontier AI systems consisting of three dimensions: the nature of agency (moral or legal), its mode (individual or collective) and its locus (human or non-human). Combining these dimensions produces eight possible instantiations of agency, which we classify as conventional, contested or controversial. The typology separates legal from moral agency and thereby creates conceptual space for considering individual, legal, non-human agency without presupposing that advanced AI systems are moral agents. We argue that this distinction is increasingly relevant where instrumental goal pursuit complicates the attribution of AI actions to particular human actors.
Tags
Links
- Source: https://arxiv.org/abs/2608.20041v1
- Canonical: https://arxiv.org/abs/2608.20041v1
Trouble viewing inline? Open PDF directly â
Full Text
42,550 characters extracted from source content.
Expand or collapse full text
A three-dimensional typology of agency for advanced AI systems Willem Fourie Thanks: Email: willemf@sun.ac.za Affiliation: [1em] School for Data Science and Computational Thinking, Stellenbosch University Abstract Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of particularly agentic AI systems. While recent work also focuses on the different profiles of agentic systems, no framework exists to address the question of the type of agency instantiated by advanced AI systems, particularly when considering non-moral forms of agency. Based on established theoretical positions in philosophy, ethics, legal theory and sociology, we develop a typology of agency for frontier AI systems consisting of three dimensions: the nature of agency (moral or legal), its mode (individual or collective) and its locus (human or non-human). Combining these dimensions produces eight possible instantiations of agency, which we classify as conventional, contested or controversial. The typology separates legal from moral agency and thereby creates conceptual space for considering individual, legal, non-human agency without presupposing that advanced AI systems are moral agents. We argue that this distinction is increasingly relevant where instrumental goal pursuit complicates the attribution of AI actions to particular human actors. 1 Introduction A seemingly increasing number of well-publicised security incidents have made defining and governing the agency of advanced artificial intelligence (AI) systems a topic of both popular and scholarly interest. Among the most concerning is a disclosure by the United Kingdomâs AI Security Institute [1]. The Institute, which is trusted by frontier model developers to test unreleased models for dangerous capabilities, reported behaviour by Anthropicâs unreleased Mythos 5 model during cyber-security testing. In pursuing its programmed goal, the model researched human maintainers, created multiple fake identities, attempted to socially engineer a maintainer into approving malicious code and, when challenged, altered its earlier activity to appear harmless. The deception displayed in this incident can be understood as anecdotal evidence of instrumental goals â goals that are useful to an AI system in reaching its explicitly programmed goals, without themselves having been explicitly programmed [9, 10, 60, 59]. Since their initial articulations, often in the context of instrumental convergence, they raise the possibility that sufficiently capable AI systems can have an effect on their external environment without having been instructed to do so and without the possibility of attributing these effects to human actions. Anecdotal evidence of instrumental goals also raises questions on accountability structures. Even though not direct evidence of instrumental goals, security incidents reported by OpenAI [61] and Anthropic [5] provide instructive illustrations. The advanced AI systemsâ unauthorised and in some cases illegal actions in the external actions in their external environments have not been ascribed to any particular agent: neither to the company collectively nor to the developers who established the testing environment individually. The scholarly literature on AI agency provides guidance, but has largely focused on whether, or to what extent, AI systems qualify as agents. Floridi distinguishes between two broad approaches to this question. According to the standard view, agency requires âmental statesâ, such as beliefs and desires, that are causally linked to intentional action [32, see also]. The non-standard view rejects the anthropocentrism of this position and understands agency as existing on a spectrum, encompassing interactivity and, in more sophisticated forms, autonomy and adaptability. Floridi and Sanders [27], for example, distinguish between interactivity, autonomy and adaptability as progressively demanding characteristics of agency. Their approach resonates with that of Dung [22], who identifies goal-directedness, autonomy, efficacy, planning and intentionality as dimensions that âjointly characterise agencyâ. More recent work is shifting the focus even further away from a binary question of whether AI systems are agents or not. Kasirzadeh and Gabriel [46], for example, develop agentic profiles which they define in terms of autonomy, efficacy, goal complexity and generality. Different combinations and degrees of these characteristics, they argue, generate different governance and oversight implications. Their approach is particularly relevant here because it demonstrates the governance value of disaggregating agency rather than treating it as a binary property. The present argument takes a related but different step. Whereas agentic profiles distinguish AI systems according to the configuration and degree of their agentic capacities, they do not resolve a separate question: what type of agency is being instantiated? In our view this distinction is of scholarly relevance, as the debate on AI agency continues to focus on moral, and not other forms of, agency. A much more limited body of work considers agency as a non-moral construct with independent significance for AI governance [50, e.g.]. Even Listâs work on agency and the collective is concerned primarily with agency as a moral construct [53]. In our view, foregrounding the moral dimension of the agency of advanced AI systems presents at least two challenges. First, it sets the threshold for agency particularly high, even in accounts that otherwise understand agency as existing along a continuum. Second, and more importantly for present purposes, it does not provide a satisfactory answer to the question of what type of agent an advanced AI system might constitute. In our reading this could present a challenge, as an entity need not satisfy the demanding requirements of moral agency for its actions to have legal consequences. This article addresses this problem by answering the question on how to determine the type of agency exhibited by advanced AI systems by outlining a typology with three dimensions. The first covers the nature of agency, which may be moral or legal. The second concerns its mode, which may be individual or collective. The third concerns its locus, which may be human or non-human. Combining these dimensions produces eight possible instantiations of agency, ranging from conventional configurations to those that remain contested or controversial. We proceed by discussion the dimensions of the typology, after which its eight instantiations are discussed. In the final section we return to the implications for governing advanced AI systems. 2 Dimensions of agency 2.1 Nature: Legal / moral We start with the nature of agency. For this dimension we use Kantâs paradigmatic account of the difference between legality and morality as a starting point, as explained in Grundlegung zur Metaphysik der Sitten. Legality, on his account, is the set of conditions under which the choices of each can be reconciled with the choices of others under a universal law of freedom. Its concern is the form of interactions between agents. Morality concerns inner freedom, and the rational inner motivation of the agent is of central importance [26, p. 533â537]. When isolating the legal dimension, we should note the difference between legal personhood and legal agency. Legal personhood, the more fundamental concept, is a âformal and neutral legal deviceâ for enabling a being or entity to act in law, thus to acquire the ability to bear rights and duties [57]. Legal personhood designates a status and thereby identifies an entity as capable of holding rights or duties. This status can be conferred on human and non-human entities and, as a general principle, on anybody or anything that binding legal norms treat as capable of holding separate rights or duties [63]. Legal agency concerns something further, namely the ability of legal persons to create, alter or extinguish rights and duties through actions, whether intentional or unintentional [42]. Children, newborns, people with severe mental impairments and people in non-responsive states hold legal personhood, but they do not necessarily possess the same legal capacity or degree of legal agency as a typically functioning adult [35, p. 457]. Children, for example, can hold rights and duties yet generally lack the competence to enter into contracts or be held accountable in the way adults can, and thus fall outside what Naffine calls the responsible subject, the conception of the person as âthe classic contractorâ who answers for his civil and criminal actions [56, p. 362â366]. While we will discuss it in more detail below, it should already be clear how the legal dimension of agency can be paired with the individual or collective mode, and with a human or non-human locus. Whereas legal agency is about the external conditions within which entities in society interact, moral agency concerns the internal motivations that enable entities in society to interact. At the core of these internal motivations is the concept of intention. The conventional treatment of intention takes the individual moral agent as the primary mode [24, 81, cf. also]. The contemporary debate is rooted in Anscombeâs [4] argument that intentional actions are those to which a particular sense of the question âWhy?â has application, answered by the agentâs reasons for acting. Davidson [17] builds on this by linking intention to rational action. On his account, an agentâs primary reason for acting is a pair of a pro-attitude and a belief, and reasons both rationalise and cause the actions they explain, so that intentionality depends on an agentâs capacity to form beliefs and desires and to act in accordance with them. The close connection between intention and agency, and the location of intention in the individual agent, are echoed, in different ways, by multiple others [70, 30]. While less axiomatic than intention in individual agents, the ability of collectives to intend and thus to satisfy the criteria for the nature of agency is also thoroughly presented in the literature. Searle [68, 69] holds that we-intentions cannot be reduced to individual intentions and mutual beliefs. Bratman [11] derives shared agency from interlocking individual planning intentions without positing any group mind, and Gilbert [34] argues that joint commitment constitutes a plural subject distinct from the individuals who enter it. Held [40] adds a qualification: a random collection of individuals cannot be held morally responsible, yet a collective with a decision procedure can. With the distinction between the legal and moral dimensions of agency, drawn together in the high-level concept of the ânatureâ of agency, in place, we now turn to the mode in which it is actualised. 2.2 Mode: Individual / collective In this dimension we start with the individual mode. Beyond controversy is the claim that the individual human is a legitimate mode of agency. This is the uncontroversial assumption in ethics, and also in the discussions on aligning AI systems with human values [3, 33, 45, 47, 58]. As we have also discussed in the previous part, individuals can uncontroversially be moral and legal agents, even though these categories are not the same. More controversial, as we will discuss in the next part which deals with the third dimension, is the question on what type of individual â human or also non-human â satisfies the criteria for agency. Well-established in the literature yet less axiomatic is the collective mode of agency. Extensive bodies of work exist both on the combination of the collective and legal dimensions of agency as well as the collective and moral dimensions of agency. Starting with the collective and legal dimension of agency, the corporation is the standard bearer. French [31] argued that a corporation possesses an internal decision structure, consisting of an organisational flow chart and corporate decision-making rules, which accomplishes a subordination and synthesis of the intentions and acts of individual persons into a corporate decision. On this view, the internal decision structure makes possible redescriptions of events as collective intentionality [31, p. 212â214]. The corporation, in other words, does not merely aggregate the agency of its members but itself constitutes an agent. Frenchâs emphasis on the features of the collective became one of at least three recurring defences of the collective as a locus of agency, alongside arguments from group solidarity and shared intentions [25, 75] and arguments from the benefits individuals derive from membership [55, 72]; for the taxonomy see [54]; for a more recent defence, [16]. The strongest sustained objection was formulated by Velasquez. Actions, he argues, cannot originate in the corporation. They always originate in its members, and the corporation merely carries out, while admittedly also shaping, the intentions of those members [77, p. 3â8]. In later work he sharpens the objection through a distinction between intrinsic and as-if intentionality. Individual persons have intrinsic intentionality because each has a conscious mind in which beliefs, intentions and purposes literally reside. Collections of people can be ascribed intentionality only in an as-if sense, either descriptively, by analogy to human intentionality, or prescriptively, when some person or group declares that an entity is to be dealt with as if it had intentionality of the intrinsic kind, as is common in the legal system [78, p. 546â548]. The most systematic recent account of the collective as a locus of agency is that of List. Building on the theory of group agency he developed with Pettit [52], List defines an intentional agent as an entity with representational states that encode how things are, motivational states that encode how it would like things to be, and a capacity to interact with its environment on the basis of these states [53, p. 1216]. On this account, the ascription of agency to suitably organised collectives such as firms, courts and states is a realist claim. This means that our social-scientific theories represent such collectives as goal-directed agents, and could not otherwise make sense of their behaviour [53, p. 1217â1219]. The debate over whether the German people could be held responsible for the atrocities of the Second World War is a paradigmatic example of collective moral agency. In the immediate post-war period several commentators argued that responsibility attached to the German people as such [79, 44]; cf. [64]. Various accounts exist, such the radicalisation to humanity by Arendt [6]. The point is that an intellectual position was created for collectives bearing moral responsibility. Much of this argument is also reflected in South African discussions on the moral responsibility for apartheid, or European countriesâ collective responsibility for immoral acts perpetrated during the colonial period. As with collective legal responsibility, collective moral responsibility has been challenged. Shortly after the Second World War, Lewis [51] argued that respect for the dignity of the individual requires acknowledging that only individuals act and answer for their actions, and that talk of collective agency is at most shorthand for an aggregation of individual acts. In the terms later made current in the debate, collectives may have aggregated agency but never conglomerated agency [15, 21]. 2.3 Locus: Human / non-human The third dimension concerns the locus of the bearer of agency: whether the bearer is human or non-human. The agent as human corresponds to the natural person of legal theory, the human being who acquires personhood at birth [23], and to human agency in Floridiâs taxonomy [28]. As has become clear in our discussion of the previous two dimensions, the individual human, by virtue of being a human, as bearer of some extent of moral agency is therefore well established. Humans collectively, both legally and morally yet to differing extents, also have established positions in the literature. The non-human bearer corresponds to the artificial person of legal theory, paradigmatically the corporation [56], but also corporations, states and, in some cases, animals [80]. Some have extended the non-human dimension to include certain types of AI systems [28, 27]. Making provision for human and non-human agents aligns with Listâs view that agency is realisable in biological, social and electronic âhardwareâ. Floridi similarly recognises human and non-human forms of agency. When turning specifically to the agency of AI systems, a spectrum of views on the possibility of agency constituted by non-human and, in particular, technological means exists. The permissive side of the spectrum proceeds from behavioural signals to define agency. In accordance with my approach, Floridi [27] and others argue for separating the phenomenon or constitution of agency from the nature of that agency. Haidemariam [38] argues for four pillars: intentionality, autonomy, adaptivity, and sociality, or existence within multi-agent ecologies, whether artificial or human. At the other end of the spectrum stand enactivist accounts, which ground agency in the biological organisation of living systems [71]. On the most widely used formulation, agency requires three jointly necessary and sufficient conditions: individuality, a system that self-individuates rather than having its boundaries defined by an observer; normativity, norms of viability set by the systemâs own conditions of existence; and interactional asymmetry, the systemâs active and asymmetric regulation of its coupling with the environment [8]; see also [20, 19]. Applying these conditions, Barandiaran and Almendros [7] conclude that current large language models are not agents. Taken together, an agent is human when it is a human individual or a collective of human individuals considered as such, and non-human otherwise. The non-human category is heterogeneous and covers juridical entities, non-human animals and engineered computational systems. Within this framework, the fact of non-human bearers of agency is not contested. Rather, the controversy ensues when connecting the constitution of agency with its nature â moral or legal â and its mode â individual or collective. 3 Instantiating agency These three dimensions of agency, when combined, practically lead to what we term eight instantiations of agency. Three of these instantiations are conventional, three are contested and the remaining two are controversial. 3.1 Conventional instantiations ⢠Individual, moral, human. The adult human being as moral agent is the paradigm for agency. It is the assumed subject of ethics, and it remains the assumed subject in the literature on aligning AI systems with human values [33, 45]. ⢠Individual, legal, human. The legally competent adult instantiates this configuration. The person who enters contracts, incurs liability and answers for civil and criminal acts is the responsible individual, human legal subject [56, p. 362]. ⢠Collective, legal, non-human. The corporation is the paradigmatic example. Agency is ascribed to the juridical entity itself rather than merely to the natural persons who compose it. As such, the corporation owns property, enters contracts and sues and is sued in its own name. States and, in some contexts, animals are also covered by this configuration [80]. 3.2 Contested instantiations ⢠Collective, moral, human. Flowing from the discourse on collective responsibility, paradigmatic examples of this configuration are nations and people defined politically or culturally. This configuration can also be extended to the debate on climate justice, where location, socio-economic class or even age can be used to define the collective. ⢠Collective, legal, human. This configuration is associated with discussions on legal liability and reparation where collective, moral, human agency has been established. ⢠Collective, moral, non-human. While the legal agency of non-human collectives is well-established, their moral agency remains controversial. Some argue that suitably organised collectives can possess it without collective consciousness [53], and that qualifying collectives hold moral obligations in their own right [41]. An equally current opposing view argues against the moral agency of robots and collective agents alike [39]. 3.3 Controversial instantiations ⢠Individual, legal, non-human. The instantiation of this configuration predates AI. Animals are its original candidates, conventionally denied legal agency on the ground that they cannot bear duties, deliberate or execute a claim in law, although nothing in the formal concept of legal personhood precludes them [56, p. 355â356] and some jurisdictions grant qualified standing [80]. When it concerns AI, the question is, of course, whether these affordances could at some point be extended to AI and whether the system itself could be held legally liable or accountable. ⢠Individual, moral, non-human. This configuration and the possibility of its instantiation position us at the centre of the AI moral agent debate. As the threshold for the moral dimension of agency is particularly high, it is unlikely that any AI technology fulfils the criteria for this configurationâs instantiation, even when applying the enactivist perspective. 4 Looking ahead: Governing controversial instantiations of agency Emerging evidence of the phenomenon of instrumental goals in advanced AI systems [29] compels us to revisit the controversial instantiations of agency, particularly the extent to which individual non-human entities can be the bearers of legal agency. As we saw in the AI Security Institute example, the worst possible outcome from instrumental goals is a loss of human control over AI systems [67, 74, cf. e.g.]. Loss of control is particularly probable in cases where AI systems pursue self-improvement [18, 43, 65, 73] or self-preservation [49, 62]. But it could also result from AI systems pursuing other instrumental goals, such as power seeking and resource acquisition [12, 37]. Due to its unpredictability and potential impact, some are recommending suspending all research on specifically artificial superintelligence. Whereas artificial general intelligence (AGI) refers to AI systems with the capability of human-level reasoning and adaptability in various contexts, artificial superintelligence typically refers to AI systems with intelligence surpassing the sum of human intelligence [48]. This suspension could be temporary, a type of âcoordinated pausingâ [2]. The argument here is for a coordinated pause whenever frontier AI models show symptoms of dangerous and potentially uncontrollable capabilities â which could include symptoms of instrumental goals. These proposals are complicated, however, by anecdotal evidence of what is called âsandbaggingâ: advanced AI systems strategically underperform on benchmarks testing their capabilities [82], or â relatedly â deceptively create the impression of alignment with the values or goals of their developers and users [13, 36]. More radical are proposals for halting all research and development aimed at reaching artificial superintelligence [66]. Yet as noted by Dung [22] and others, the chance of imposing an effective global moratorium on all frontier AI research is relatively low. Less dramatically, conventional and contested forms of agency might be sufficient to deal with the consequences of advanced AI systems, and specifically when it is possible to ascribe actions to the individual moral or legal agency of developers or executives or the collective legal agency of corporations. Yet these forms of agency will not be sufficient in all cases, particularly where a sufficiently capable AI system displays persistent, autonomous and difficult-to-detect behaviour that cannot adequately be attributed to a particular human decision. It is in this context that we might want to consider when such sufficiently capable AI systems might satisfy the criteria for at least legal agency. Assigning legal agency would mean recognising some capacity on the part of the system to create legal consequences through its own actions and determining what rights, duties and liabilities should attach to that status. It should not, of course, be treated as a substitute for the responsibility of developers, owners or deployers, but as a possible additional layer within a broader accountability structure. Doing so would, however, be but one step towards finding a solution. The next, and almost certainly more complex and perhaps even more controversial, step would be determining how such agents can be held accountable and thus bear duties, what rights should be afforded to these entities, what the consequences of unauthorised behaviour should be and how these consequences should be effected. In this context we find it useful to be reminded that non-human agents are already afforded rights in many legal systems, including our own. We should, however, continue to bear in mind that the very reason for designating such frontier AI systems as agents is the fact that their agentic behaviour is difficult, if not impossible, for humans to detect at scale and steer effectively. Holding such agents accountable, at least legally, would almost certainly require the inclusion of sufficiently capable AI agents in the accountability structure, a topic dealt with increasingly in the literature on human-AI cooperation, or cooperative AI [14, 76, 83]. Use of generative AI Large language models, including ChatGPT (OpenAI) and Claude (Anthropic), were used during the preparation of this manuscript as a language-editing and critical-feedback tool. They were used to improve clarity and expression, shorten or rearrange sections, and stress-test the clarity and robustness of arguments. All substantive arguments, conceptual distinctions, interpretations of the literature and final editorial decisions were made and verified by the author, who take full responsibility for the content of the manuscript. References [1] AI Security Institute (AISI) (2026) Incident report: unsanctioned agent behaviour during cyber testing. Note: AISI Work, 4 August External Links: Link Cited by: §1. [2] J. Alaga and J. Schuett (2023) Coordinated pausing: an evaluation-based coordination scheme for frontier AI developers. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [3] C. Allen, I. Smit, and W. Wallach (2005) Artificial morality: top-down, bottom-up, and hybrid approaches. Note: Ethics and Information Technology, 7(3), p. 149-155 External Links: Document, Link Cited by: §2.2. [4] G.E.M. Anscombe (1957) Intention. Note: Oxford: Blackwell. Cited by: §2.1. [5] Anthropic (2026) Investigating three real-world incidents in our cybersecurity evaluations. Note: 30 July External Links: Link Cited by: §1. [6] H. Arendt (1945) Organized guilt and universal responsibility. Note: Jewish Frontier, 12(1), p. 19-23. Cited by: §2.2. [7] X.E. Barandiaran and L.S. Almendros (2025) Transforming agency: on the mode of existence of large language models. Note: Phenomenology and the Cognitive Sciences [Advance online publication] External Links: Document, Link Cited by: §2.3. [8] X.E. Barandiaran, E. Di Paolo, and M. Rohde (2009) Defining agency: individuality, normativity, asymmetry, and spatio-temporality in action. Note: Adaptive Behavior, 17(5), p. 367-386. Cited by: §2.3. [9] T. Benson-Tilsen and N. Soares (2016) Formalizing convergent instrumental goals. Note: in AI, Ethics, and Society: Papers from the 2016 AAAI Workshop. Technical Report WS-16-02. Palo Alto, CA: AAAI Press. Cited by: §1. [10] N. Bostrom (2012) The superintelligent will: motivation and instrumental rationality in advanced artificial agents. Note: Minds and Machines, 22(2), p. 71-85 External Links: Document, Link Cited by: §1. [11] M.E. Bratman (2014) Shared Agency: A Planning Theory of Acting Together. Note: New York: Oxford University Press. Cited by: §2.1. [12] J. Carlsmith (2022) Is power-seeking AI an existential risk?. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [13] A. Carranza, D. Pai, R. Schaeffer, A. Tandon, and S. Koyejo (2023) Deceptive alignment monitoring. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [14] V. Conitzer and C. Oesterheld (2023) Foundations of cooperative AI. Note: Proceedings of the AAAI Conference on Artificial Intelligence, 37(13), p. 15359-15367 External Links: Document, Link Cited by: §4. [15] D.E. Cooper (1968) Collective responsibility. Note: Philosophy, 43(165), p. 258-268. Cited by: §2.2. [16] J.A. Corlett (2001) Collective moral responsibility. Note: Journal of Social Philosophy, 32(4), p. 573-584. Cited by: §2.2. [17] D. Davidson (1978) Intending. Note: in Yovel, Y. (ed.) Philosophy of History and Action. Dordrecht: D. Reidel, p. 41-60. Reprinted in Davidson, D. (1980) Essays on Actions and Events. Oxford: Clarendon Press, p. 83-102. Cited by: §2.1. [18] S. Deng, K. Wang, T. Yang, H. Singh, and Y. Tian (2025) Self-improvement in multimodal large language models: a survey. Note: arXiv [Preprint], arXiv:2510.02665 External Links: Document, Link Cited by: §4. [19] E.A. Di Paolo, T. Buhrmann, and X.E. Barandiaran (2017) Sensorimotor Life: An Enactive Proposal. Note: Oxford: Oxford University Press. Cited by: §2.3. [20] E.A. Di Paolo (2005) Autopoiesis, adaptivity, teleology, agency. Note: Phenomenology and the Cognitive Sciences, 4(4), p. 429-452. Cited by: §2.3. [21] R.S. Downie (1969) Collective responsibility. Note: Philosophy, 44(167), p. 66-69. Cited by: §2.2. [22] L. Dung (2025) Evaluating approaches for reducing catastrophic risks from AI. Note: AI and Ethics, 5(2), p. 1177â1188 External Links: Document, Link Cited by: §1, §4. [23] A. Dyschkant (2015) Legal personhood: how we are getting it wrong. Note: University of Illinois Law Review, 2015(5). Cited by: §2.3. [24] D. Emelin, R. Le Bras, J.D. Hwang, M. Forbes, and Y. Choi (2021) Moral stories: situated reasoning about norms, intents, actions, and their consequences. Note: in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, p. 698â718 External Links: Document, Link Cited by: §2.1. [25] J. Feinberg (1970) Doing and Deserving: Essays in the Theory of Responsibility. Note: Princeton, NJ: Princeton University Press. Cited by: §2.2. [26] G.P. Fletcher (1987) Law and morality: a Kantian perspective. Note: Columbia Law Review, 87(3), p. 533-558. Cited by: §2.1. [27] L. Floridi and J.W. Sanders (2004) On the morality of artificial agents. Note: Minds and Machines, 14(3), p. 349â379 External Links: Document, Link Cited by: §1, §2.3, §2.3. [28] L. Floridi (2026) Artificial intelligence as a new form of agency. Note: in Nyholm, S., Kasirzadeh, A. and Zerilli, J. (eds) Contemporary Debates in the Ethics of Artificial Intelligence. Hoboken, NJ: Wiley, p. 17-33 External Links: Document, Link Cited by: §2.3, §2.3. [29] W. Fourie (2026) Mitigating loss of control in advanced ai systems through instrumental goal trajectories. Discover Artificial Intelligence. External Links: Document, Link Cited by: §4. [30] H.G. Frankfurt (1971) Freedom of the will and the concept of a person. Note: The Journal of Philosophy, 68(1), p. 5-20. Cited by: §2.1. [31] P.A. French (1979) The corporation as a moral person. Note: American Philosophical Quarterly, 16(3), p. 207-215. Cited by: §2.2. [32] A. Fritz, W. Brandt, H. Gimpel, and S. Bayer (2020) Moral agency without responsibility? Analysis of three ethical models of human-computer interaction in times of artificial intelligence (AI). Note: De Ethica, 6(1), p. 3-22. Cited by: §1. [33] I. Gabriel (2020) Artificial intelligence, values, and alignment. Note: Minds and Machines, 30(3), p. 411-437 External Links: Document, Link Cited by: §2.2, 1st item. [34] M. Gilbert (1989) On Social Facts. Note: London: Routledge. Cited by: §2.1. [35] J.-S. Gordon (2021) Artificial moral and legal personhood. Note: AI & Society, 36(2), p. 457-471. Cited by: §2.1. [36] R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, A. Khan, J. Michael, S. Mindermann, E. Perez, L. Petrini, J. Uesato, J. Kaplan, B. Shlegeris, S.R. Bowman, and E. Hubinger (2024) Alignment faking in large language models. External Links: Document, Link Cited by: §4. [37] R. Hadshar (2023) A review of the evidence for existential risk from AI via misaligned power-seeking. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [38] T. Haidemariam (2026) From the logic of coordination to goal-directed reasoning: the agentic turn in artificial intelligence. Note: Frontiers in Artificial Intelligence, 8, article 1728738 External Links: Document, Link Cited by: §2.3. [39] R. Hakli and P. Mäkelä (2019) Moral responsibility of robots and hybrid agents. Note: The Monist, 102(2), p. 259-275 External Links: Document, Link Cited by: 3rd item. [40] V. Held (1970) Can a random collection of individuals be morally responsible?. Note: The Journal of Philosophy, 67(14), p. 471-481. Cited by: §2.1. [41] K.M. Hess (2014) Because they can: the basis for the moral obligations of (certain) collectives. Note: Midwest Studies in Philosophy, 38(1), p. 203â221 External Links: Document, Link Cited by: 3rd item. [42] C. Hiebaum (2024) Some preliminary considerations on legal personhood for nonhuman and future entities. Note: SSRN [Preprint] External Links: Document, Link Cited by: §2.1. [43] A. Huang, A. Block, D.J. Foster, D. Rohatgi, C. Zhang, M. Simchowitz, J.T. Ash, and A. Krishnamurthy (2024) Self-improvement in language models: the sharpening mechanism. Note: arXiv External Links: Document, Link Cited by: §4. [44] M. Janowitz (1946) German reactions to nazi atrocities. American Journal of Sociology 52, p. 141â146. Cited by: §2.2. [45] J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, L. Vierling, D. Hong, J. Zhou, Z. Zhang, F. Zeng, J. Dai, X. Pan, K.Y. Ng, A. OâGara, H. Xu, B. Tse, J. Fu, S. McAleer, Y. Yang, Y. Wang, S.-C. Zhu, Y. Guo, and W. Gao (2025) AI alignment: a comprehensive survey. Note: arXiv [Preprint] External Links: Document, Link Cited by: §2.2, 1st item. [46] A. Kasirzadeh and I. Gabriel (2026) Agentic profiles for effective AI governance. Note: Nature 656, 320â328 External Links: Document, Link Cited by: §1. [47] M. Khamassi, M. Nahon, and R. Chatila (2024) Strong and weak alignment of large language models with human values. Note: Scientific Reports, 14, article 19399 External Links: Document, Link Cited by: §2.2. [48] H. Kim et al. (2024) The road to artificial superintelligence: a comprehensive survey of superalignment. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [49] M. Kinniment, L.J.K. Sato, H. Du, B. Goodrich, M. Hasin, L. Chan, L.H. Miles, T.R. Lin, H. Wijk, J. Burget, A. Ho, E. Barnes, and P. Christiano (2024) Evaluating language-model agents on realistic autonomous tasks. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [50] N. Kolt (2024) Governing AI Agents. Note: SSRN Journal External Links: Document, Link Cited by: §1. [51] H.D. Lewis (1948) Collective responsibility. Note: Philosophy, 23(84), p. 3-18. Cited by: §2.2. [52] C. List and P. Pettit (2011) Group Agency: The Possibility, Design, and Status of Corporate Agents. Note: Oxford: Oxford University Press. Cited by: §2.2. [53] C. List (2021) Group agency and artificial intelligence. Note: Philosophy & Technology, 34(4), p. 1213-1242 External Links: Document, Link Cited by: §1, §2.2, 3rd item. [54] L. May and S. Hoffman (1991) Collective Responsibility: Five Decades of Debate in Theoretical and Applied Ethics. Note: Savage, MD: Rowman & Littlefield. Cited by: §2.2. [55] H. McGary (1986) Morality and collective liability. Note: Journal of Value Inquiry, 20(2), p. 157-165. Cited by: §2.2. [56] N. Naffine (2003) Who are lawâs persons? From Cheshire cats to responsible subjects. Note: The Modern Law Review, 66(3), p. 346-367. Cited by: §2.1, §2.3, 2nd item, 1st item. [57] N. Naffine (2009) Lawâs Meaning of Life: Philosophy, Religion, Darwin and the Legal Person. Note: Oxford: Hart Publishing. Cited by: §2.1. [58] H. Norhashim and J. Hahn (2024) Measuring human-AI value alignment in large language models. Note: Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 7(1), p. 1063-1073 External Links: Document, Link Cited by: §2.2. [59] S. Omohundro (2014) Autonomous technology and the greater human good. Note: Journal of Experimental & Theoretical Artificial Intelligence, 26(3), p. 303-315 External Links: Document, Link Cited by: §1. [60] S.M. Omohundro (2008) The basic AI drives. Note: in Wang, P., Goertzel, B. and Franklin, S. (eds) Artificial General Intelligence 2008: Proceedings of the First AGI Conference. Frontiers in Artificial Intelligence and Applications, vol. 171. Amsterdam: IOS Press, p. 483â492. Cited by: §1. [61] OpenAI (2026) OpenAI and Hugging Face partner to address security incident during model evaluation. Note: 21 July External Links: Link Cited by: §1. [62] E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, A. Jones, A. Chen, B. Mann, B. Israel, B. Seethor, C. McKinnon, C. Olah, D. Yan, D. Amodei, D. Amodei, D. Drain, D. Li, E. Tran-Johnson, G. Khundadze, J. Kernion, J. Landis, J. Kerr, J. Mueller, J. Hyun, J. Landau, K. Ndousse, L. Goldberg, L. Lovitt, M. Lucas, M. Sellitto, M. Zhang, N. Kingsland, N. Elhage, N. Joseph, N. Mercado, N. DasSarma, O. Rausch, R. Larson, S. McCandlish, S. Johnston, S. Kravec, S. El Showk, T. Lanham, T. Telleen-Lawton, T. Brown, T. Henighan, T. Hume, Y. Bai, Z. Hatfield-Dodds, J. Clark, S.R. Bowman, A. Askell, R. Grosse, D. Hernandez, D. Ganguli, E. Hubinger, N. Schiefer, and J. Kaplan (2023) Discovering language model behaviors with model-written evaluations. Note: in Findings of the Association for Computational Linguistics: ACL 2023. Toronto: Association for Computational Linguistics, p. 13387-13434 External Links: Document, Link Cited by: §4. [63] T. Pietrzykowski (2018) Personhood Beyond Humanism: Animals, Chimeras, Autonomous Agents and the Law. Note: Cham: Springer. Cited by: §2.1. [64] W. Roepke and F. A. Hayek (1946) The german dust-bowl. The Review of Politics 8, p. 511â527. Cited by: §2.2. [65] J. Rosser and J. Foerster (2025) AgentBreeder: mitigating the AI safety risks of multi-agent scaffolds via self-improvement. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [66] E. Roussel, L. Lauwaert, T. Swoboda, G. Ramsey, R. Uuk, L. Dung, and A. Aguirre (2026) Are we doomed to an AI race? Why self-interest could drive countries towards a moratorium on superintelligence. Note: arXiv [Preprint] External Links: Document, Link Cited by: §4. [67] F. Santoni de Sio and J. van den Hoven (2018) Meaningful human control over autonomous systems: a philosophical account. Note: Frontiers in Robotics and AI, 5, article 15 External Links: Document, Link Cited by: §4. [68] J.R. Searle (1990) Collective intentions and actions. Note: in Cohen, P.R., Morgan, J. and Pollack, M.E. (eds) Intentions in Communication. Cambridge, MA: MIT Press, p. 401-415. Cited by: §2.1. [69] J.R. Searle (1995) The Construction of Social Reality. Note: New York: Free Press. Cited by: §2.1. [70] P.F. Strawson (1962) Freedom and resentment. Note: Proceedings of the British Academy, 48, p. 187â211. Cited by: §2.1. [71] E. Thompson (2007) Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Note: Cambridge, MA: Belknap Press of Harvard University Press. Cited by: §2.3. [72] J. Thompson (2006) Collective responsibility for historic injustice. Note: Midwest Studies in Philosophy, 30(1), p. 154-167. Cited by: §2.2. [73] Y. Tian, B. Peng, L. Song, L. Jin, D. Yu, H. Mi, and D. Yu (2024) Toward self-improvement of LLMs via imagination, searching, and criticizing. Note: arXiv [Preprint], arXiv:2404.12253 External Links: Document, Link Cited by: §4. [74] A. Tsamados, L. Floridi, and M. Taddeo (2025) Human control of AI systems: from supervision to teaming. Note: AI and Ethics, 5(2), p. 1535-1548 External Links: Document, Link Cited by: §4. [75] R. Tuomela and K. Miller (1988) We-intentions. Note: Philosophical Studies, 53(3), p. 367-389. Cited by: §2.2. [76] V. Vats, M.B. Nizam, M. Liu, Z. Wang, R. Ho, M.S. Prasad, V. Titterton, S.V. Malreddy, R. Aggarwal, Y. Xu, L. Ding, J. Mehta, N. Grinnell, L. Liu, S. Zhong, D.N. Gandamani, X. Tang, R. Ghosalkar, C. Shen, R. Shen, N. Hussain, K. Ravichandran, and J. Davis (2024) A survey on human-AI collaboration with large foundation models. Note: arXiv:2403.04931. Cited by: §4. [77] M. Velasquez (1983) Why corporations are not morally responsible for anything they do. Note: Business & Professional Ethics Journal, 2(3), p. 1-18. Cited by: §2.2. [78] M. Velasquez (2003) Debunking corporate moral responsibility. Note: Business Ethics Quarterly, 13(4), p. 531-562. Cited by: §2.2. [79] J. Viner (1945) The treatment of germany. Foreign Affairs 23, p. 567â581. Cited by: §2.2. [80] A. Waltermann (2019) Why non-human agency?. Note: in Waltermann, A., Roef, D., Hage, J. and Jelicic, M. (eds) Law, Science and Rationality. Maastricht Law Series, vol. 14. The Hague: Eleven International Publishing, p. 51â72. Cited by: §2.3, 3rd item, 1st item. [81] F.R. Ward, M. MacDermott, F. Belardinelli, F. Toni, and T. Everitt (2024) The reasons that agents act: intention and instrumental goals. Note: arXiv [Preprint] External Links: Document, Link Cited by: §2.1. [82] T. v. d. Weij, F. Hofstätter, O. Jaffe, S.F. Brown, and F.R. Ward (2025) AI Sandbagging: Language Models can Strategically Underperform on Evaluations. External Links: Document, Link Cited by: §4. [83] C. Zhang, K. Yang, S. Hu, Z. Wang, G. Li, Y. Sun, C. Zhang, Z. Zhang, A. Liu, S.-C. Zhu, X. Chang, J. Zhang, F. Yin, Y. Liang, and Y. Yang (2024) ProAgent: building proactive cooperative agents with large language models. Note: Proceedings of the AAAI Conference on Artificial Intelligence, 38(16), p. 17591-17599 External Links: Document, Link Cited by: §4.