Paper deep dive
Is Power-Seeking AI an Existential Risk?
Joseph Carlsmith
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 8:12:45 PM
Summary
This report by Joseph Carlsmith evaluates the existential risk posed by advanced, agentic, and strategically aware AI systems. It outlines a six-premise argument suggesting that by 2070, the development of misaligned AI could lead to humanity's permanent disempowerment. The author estimates a >10% probability of such an existential catastrophe, emphasizing the instrumental incentives for power-seeking in intelligent agents.
Entities (4)
Relation Signals (3)
Joseph Carlsmith â estimatesprobability â Existential Catastrophe
confidence 95% ¡ my estimate here has gone up, and is now at >10%.
APS Systems â exhibits â Power-Seeking
confidence 90% ¡ that by default, suitably strategic and intelligent agents... will have instrumental incentives to gain and maintain various types of power
APS Systems â posesriskof â Existential Catastrophe
confidence 90% ¡ Some worry that the development of advanced artificial intelligence will result in existential catastrophe
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire -- especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) such disempowerment will constitute an existential catastrophe. I assign rough subjective credences to the premises in this argument, and I end up with an overall estimate of ~5% that an existential catastrophe of this kind will occur by 2070. (May 2022 update: since making this report public in April 2021, my estimate here has gone up, and is now at >10%.)
Tags
Links
- Source: https://arxiv.org/abs/2206.13353
- Canonical: https://arxiv.org/abs/2206.13353
Trouble viewing inline? Open PDF directly â
Full Text
254,107 characters extracted from source content.
Expand or collapse full text
Is Power-Seeking AI an Existential Risk? Joseph Carlsmith Open Philanthropy April 2021 Shorter version | Video presentation | Slides | Audio version | Reviews (including superforecasters) Abstract This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire â especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) such disempowerment will constitute an existential catastrophe. I assign rough subjective credences to the premises in this argument, and I end up with an overall estimate of ~5% that an existential catastrophe of this kind will occur by 2070.(May 2022 update: since making this report public in April 2021, my estimate here has gone up, and is now at >10%.) 1 Introduction3 1.1Preliminaries . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 1.2Backdrop . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .4 1.2.1Intelligence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5 1.2.2Agency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5 1.2.3Playing with fire . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .6 1.2.4Power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .7 2 Timelines7 2.1Three key properties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .8 2.1.1Advanced capabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . .8 2.1.2Agentic planning . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .8 2.1.3Strategic awareness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .10 2.2Likelihood by 2070 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .11 arXiv:2206.13353v2 [cs.CY] 13 Aug 2024 3 Incentives11 3.1Usefulness . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .12 3.2Available techniques . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .14 3.3Byproducts of sophistication . . . . . . . . . . . . . . . . . . . . . . . . . . . . .14 4 Alignment15 4.1Definitions and clarifications . . . . . . . . . . . . . . . . . . . . . . . . . . . . .15 4.2Power-seeking . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .18 4.3The challenge of practical PS-alignment . . . . . . . . . . . . . . . . . . . . . . .22 4.3.1Controlling objectives . . . . . . . . . . . . . . . . . . . . . . . . . . . .22 4.3.2Controlling capabilities . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 4.3.3Controlling circumstances . . . . . . . . . . . . . . . . . . . . . . . . . .28 4.4Unusual difficulties . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .29 4.4.1Barriers to understanding . . . . . . . . . . . . . . . . . . . . . . . . . . .29 4.4.2Adversarial dynamics . . . . . . . . . . . . . . . . . . . . . . . . . . . . .30 4.4.3Stakes of error . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .30 4.5Overall difficulty . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .31 5 Deployment31 5.1Timing of problems . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .32 5.2Decisions . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 5.3Key risk factors . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .34 5.3.1Externalities and competition . . . . . . . . . . . . . . . . . . . . . . . .34 5.3.2Number of relevant actors . . . . . . . . . . . . . . . . . . . . . . . . . .36 5.3.3Bottlenecks on usefulness . . . . . . . . . . . . . . . . . . . . . . . . . .36 5.3.4Deception . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 5.4Overall risk of problematic deployment . . . . . . . . . . . . . . . . . . . . . . .38 6 Correction39 6.1Take-off . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .39 6.2Warning shots . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .39 6.3Competition for power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .41 6.3.1Mechanisms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .42 6.4Corrective feedback loops . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .44 6.5Sharing power . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .45 7 Catastrophe46 8 Probabilities47 9 Appendix51 2 1 Introduction Some worry that the development of advanced artificial intelligence will result in existential catastropheâthat is, the destruction of humanityâs longterm potential. 1 Here I examine the fol- lowing version of this worry (itâs not the only version): By 2070: 1. It will become possible and financially feasible to build AI systems with the following properties: â˘Advanced capability: they outperform the best humans on some set of tasks which when performed at advanced levels grant significant power in todayâs world (tasks like scientific research, business/military/political strategy, engineering, and persua- sion/manipulation). â˘Agentic planning: they make and execute plans, in pursuit of objectives, on the basis of models of the world. â˘Strategic awareness:the models they use in making plans represent with reasonable accuracy the causal upshot of gaining and maintaining power over humans and the real-world environment. (Call these âAPSââAdvanced, Planning, Strategically awareâsystems.) 2. There will be strong incentives to build and deploy APS systems | (1). 3.It will be much harder to build APS systems that would not seek to gain and maintain power in unintended ways (because of problems with their objectives) on any of the inputs theyâd encounter if deployed, than to build APS systems that would do this, but which are at least superficially attractive to deploy anyway | (1)â(2). 4. Some deployed APS systems will be exposed to inputs where they seek power in unintended and high-impact ways (say, collectively causing >$1 trillion dollars of damage), because of problems with their objectives | (1)â(3). 2 5.Some of this power-seeking will scale (in aggregate) to the point of permanently disempow- ering ~all of humanity | (1)â(4). 6. This disempowerment will constitute an existential catastrophe | (1)â(5). These claims are extremely important if true. My aim is to investigate them. I focus on (2)â(5), but I also say a few words about (1) and (6). My current view is that there is a disturbingly substantive chance that a scenario along these lines occurs, and that many people alive todayâincluding myselfâlive to see humanity permanently disempowered by AI systems weâve lost control over. In the final section, I take an initial stab at quantifying this risk, by assigning rough probabilities to 1-6.My current, highly-unstable, subjective estimate is that there is a ~5% percent chance of existential catastrophe by 2070 from scenarios in which (1)â(6) are true.(May 2022 authorâs note: since making this report public in April 2021, my estimate here has gone up; itâs currently at >10%.)My main hope, though, is not to push for a specific number, but rather to lay out the arguments in a way that can facilitate productive debate. 1.1 Preliminaries Some preliminaries and caveats (those eager for the main content can skip): â˘Iâm focused, here, on a very specific type of worry. There are lots of other ways to be worried about AIâand even, about existential catastrophes resulting from AI. And there are lots of ways to be excited about AI, too. 1 See e.g. Yudkowsky (2008), Bostrom (2014), Hawking (2014), Tegmark (2017), Christiano (2019), Russell (2019), Ord (2020), and Ngo (2020). This definition of âexistential catastropheâ is from Ord (2020, p. 27); see section 7 for a bit more discussion. 2 Letâs assume 2021 dollars. 3 â˘My emphasis and approach differs from that of others in the literature in various ways. 3 In particular: Iâm less focused than some on the possibility of an extremely rapid escalation in frontier AI capabilities, on ârecursive self-improvement,â or on scenarios in which one actor comes to dominate the world; Iâm focusing on power-seeking in particular (as opposed to âmisalignmentâ more broadly); Iâm aiming to incorporate and respond to various objections and re-framings introduced since early work on the topic, and since machine learning has become a prominent AI paradigm; 4 Iâm aiming to avoid reliance on models of âutility function maximizationâ and related concepts; and I present and assign probabilities to premises in a complete argument for catastrophe. That said, I still think of this as an articulation and analysis of a certain kind of âcore argumentââone that has animated, and continues to animate, much concern about existential risk from misaligned AI. 5 â˘Iâm focusing on 2070 because I want to keep vividly in mind that I and many other readers (and/or their children) should expect to live to see the claims at stake here falsified or confirmed. That said, the main arguments donât actually require the development of relevant systems within any particular period of time (though timelines in this respect can matter to e.g. the amount of evidence that present-day systems and conditions provide about future risks). ⢠Iâm not addressing the question of what sorts of interventions are currently available for lowering the risk in question, or how they compare with work on other issues. â˘Section 8 involves assigning (a) subjective probabilities to (b) imprecisely operationalized claims, in the context of (c) a multi-step argument. I discuss cautions on all of three of these fronts there (and I include an appendix aimed at warding off some possible problems with (c) in particular). â˘Thinking about what will happen decades in the future is famously hard, and some approach ~any such thought with extreme skepticism. Iâm sympathetic to this in ways, but I also think we can do better than brute agnosticism; and to the extent that wecareabout the future, and can try to influence it, our actions will often express âbetsâ about it even absent explicit forecasts. â˘I use various concepts it would be great to be more precise about. This is true of many worthwhile discussions, but that doesnât make the attending imprecisions innocuous. To the contrary, I encourage wariness of ways they might mislead. â˘The discussion focuses on how to avoid human disempowerment, but I am not assuming that it is always appropriate or good for humans to keep power. Rather, Iâm assuming that we want to avoid disempowerment that isimposedon us by artificial systems weâve lost control over. See section 7 for more discussion. â˘Sufficiently sophisticated AI systems might warrant moral concern. 6 In my opinion, this fact should motivate grave ethical caution in the context of many types of AI development, including many discussed in this report. However, itâs not my focus here (see section 7 for a few remarks). 1.2 Backdrop The specific arguments Iâl discuss emerge from a broader backdrop picture, which Iâl gloss as: 3 Though the differences listed here donât each apply to all of the literature; and in many respects, the issues I discuss are the âclassic issues.â Of existing overall analyses, mine is probably most similar to Ngo (2020). For other discussions of alignment problems in AI (on top of those cited in footnote 1), see e.g. Amodei et al (2016), Piper (2018, updated 2020), and Christian (2020). 4 For example, those of Shah (2018), Pinker (2018), Drexler (2019), Zador and LeCun (2019), Hubinger et al (2019), Garfinkel (2020). 5 Someâe.g. Adamczewski (2019), Garfinkel (2019), and Ngo (2019), though Ngo notes that his views have since shiftedâsuggest that arguments for X-risk from misaligned AI have âshiftedâ or âdriftedâ over time. I see some of this in some places, but for me at least (and for various others I know), the basic thrustâe.g., advanced artificial agents would be very powerful, and if their objectives are problematic, they might have default incentives to disempower humansâhas been pretty constant (sometimes claims about discontinuous/concentrated âtakeoffâ are treated as essential to the basic thrust of the argument, but I donât see them that wayâsee section 6 for discussion). 6 See Muehlhauser (2017) for more on what sorts of non-humans warrant this concern. 4 1.Intelligent agency is an extremely powerful force for controlling and transforming the world. 2. Building agents much more intelligent than humans is playing with fire. Iâl start by briefly describing this picture, as I think it sets an important stage for the discussion that follows. 1.2.1 Intelligence Of all the species that have lived on earth, humans are clearly strange. In particular, we exert an unprecedented scale and sophistication of intentional control over our environment. Consider, for example: the city of Tokyo, the Large Hadron Collider, and the Bingham Canyon Mine. What makes this possible? Something about our minds seems centrally important. We can plan, learn, communicate, deduce, remember, explain, imagine, experiment, and cooperate in ways that other species canât. These cognitive abilitiesâemployed in the context of the culture and technology we inherit and createâgive us the power, collectively, to transform the world. For ease of reference, letâs call this loose cluster of abilities âintelligence,â though very little will rest on the term. 7 Our abilities in these respects are nowhere near any sort of hard limit. Human cognitionâeven in groups, and with the assistance of technologyâdepends centrally on the human brain, which, for all its wonders, is an extremely specific and limited organ, subject to very specific constraintsâon cell count, energy, communication speed, signaling frequency, memory capacity, component reliability, input/output bandwidth, and so forth. 8 There are possible cognitive systemsâpossible brains, and possible artificial systemsâto which some or all these constraints do not apply. It seems very likely that such systems could learn, communicate, reason, problem-solve, etc, much better than humans can. The variation in cognitive ability we see among humans, and across species, also suggests this; as do existing successes in approaching or exceeding human capabilities at particular tasksâmathematical calculation, game-playing, image recognition, and so forthâusing artificial systems. 1.2.2 Agency Humans can also beagenticâthat is (loosely), we pursue objectives, guided by models of the world (see section 2.1.2 for more). You want to travel to New York, so you buy a flight, double-check when it leaves, wake up early to pack, and so forth. Non-human cognitive systems can be agentic, too. And depending on what you count as an agent, humans can already create novel non-human agents. 9 We breed and genetically modify non-humans animals, for example; and some of our artificial systems display agent-like behavior in specific environments (see e.g. here and here). But these agents canât learn, plan, communicate, etc. like we can. At some point, thoughâabsent catastrophe, deliberate choice, or other disruption of scientific progressâwe will likely be in a position, if we so choose, to create non-human agents whose abilities in these respects rival or exceed our own. And in the context of artificial agents, the differences between brains and computersâin possible speed, size, available energy, memory 7 In particular, when I say âmore intelligent than humans,â I just mean âbetter at things like planning, learning, communicating, deducing, remembering, etcâ than humans. I donât have in mind some broader notion like âoptimization power,â âsolving problems,â or âachieving objectives.â Thanks to Ben Garfinkel for discussion here. Obviously, intelligence in this sense is highly multi-dimensional. I donât think this is a problem for saying one thing is more intelligent than another (e.g., which house is better to live in is multi-dimensional, but a mansion can still be preferable to a shack); but if you object to the term for this or some other reason, feel free to substitute âcognitive abilities like planning, learning, communicating, deducing, remembering, etcâ whenever I say âintelligenceâ (I wonât use the term very often). See Garfinkel et al (2017) for more on objections to the possibility of greater-than-human intelligence. 8 For example: the brain has ~1e11 neurons and 1e14-1e15 synapses; neurons fire at a maximum of some hundreds of Hz; action potentials travel at a max of some hundreds of m/s; the brain runs on ~20W of power; it has to fit within the skull; and so forth. 9 Iâl count agent-like systems constituted in central part by human agentsâe.g., nations, corporations, parent teacher organizations, etcâas âhumanâ (which need not imply: innocuous). 5 capacity, component reliability, input/output bandwidth, and so forthâmake the eventual possibility of very dramatic differences in ability especially salient. 10 1.2.3 Playing with fire The choice to create agents much more intelligent than we are should be approached with extreme caution. This is the basic backdrop view underlying much of the concern about existential risk from AIâand it would apply, in similar ways, to new biological agents (human or non-human). Some articulate this view by appeal to the dominant position of humans on this planet, relative to other species. 11 For example: some argue that the fate of the chimpanzees is currently in human hands, and that this difference in power is primarily attributable to differences in intelligence, rather than e.g. physical strength. Just as chimpanzeesâgiven the choice and powerâshould be careful about building humans, then, we should be careful about building agents more intelligent than us. This argument is suggestive, but far from airtight. Chimpanzees, for example, are themselves much more intelligent than mice, but the âfate of the miceâ was never âin the handsâ of the chimpanzees. Whatâs more, the control that humans can exert over the fate of other species on this planet still has limits, and we can debate whether âintelligence,â even in the context of accumulating culture and technology, is the best way of explaining what control we have. 12 More importantly, though: humans arose through an evolutionary process that chimpanzees did nothing to intentionally steer. Humans, though, will be able to control many aspects of processes we use to build and empower new intelligent agents. 13 Still, some worry about playing with fire persists. As our own impact on the earth illustrates, intelligent agents can be an extremely powerful force for controlling and transforming an environment in pursuit of their objectives. Indeed, even on the grand scale of earthâs history, the development of human abilities in this respect seems like a very big dealâa force of unprecedented potency. If we unleash much more of this force into the world, via new, more intelligent forms of non-human agency, it seems reasonable to expect dramatic impacts, and reasonable to wonder how well we will be able to control the results. 14 10 Thus, as Bostrom (2014, Chapter 3) discusses, where neurons can fire a maximum of hundreds of Hz, the clock speed of modern computers can reach some 2 Ghz: ~ten million times faster. Where action potentials travel at some hundreds of m/s, optical communication can take place at 300,000,000 m/s: ~a million times faster. Where brain size and neuron count are limited by cranial volume, metabolic constraints, and other factors, supercomputers can be the size of warehouses. And artificial systems need not suffer, either, from the brainâs constraints with respect to memory and component reliability, input/output bandwidth, tiring after hours, degrading after decades, storing its own repair mechanisms and blueprints inside itself, and so forth. Artificial systems can also be edited and duplicated much more easily than brains, and information can be more easily transferred between them. 11 See e.g. Bostrom (2015), Russell (2019, Chapter 5) on the âGorilla Problem,â Ord (2020), and Ngo (2020) on the âsecond species argument.â 12 Note, too, that a single human, unaided and uneducated, would plausibly fare worse than a chimpanzee in many environments (see Henrich (2015) for more on humans struggling in the absence of cultural learning). 13 Though whether that control will be adequate to ensure the outcomes we want is a substantially further question; and note that analogies between evolution and contemporary techniques in machine learning are somewhat sobering in this respectâsee section 4.3.1.2 for more. 14 Of course, the most relevant and powerful âhuman actorsâ are and have been groups (countries, corporations, ideologies, cultures, etc), equipped with technology, rather than unaided individuals. But non-human agents can create groups, and use/create new technology, too (indeed, creating and using technology more effectively than humans is one of the key things that makes advanced AI systems formidable). And note that while human organizations like corporations are much more capable than individuals along certain axes (especially e.g. tasks that can be effectively parallelized), they remain constrained by human abilities, and guided by human values, in key waysâways that fully ânon-humanâ corporations would not (see here for more discussion). That said, I think questions around the role of individuals vs. group agents/âsuperorganismsâ as central loci of influence, and questions around the role of human intelligence vs. technology, culture, economic interaction, etc, in explaining human growth and dominance on this planet, are both key ways in which âbuilding agents more much intelligent than humans warrants cautionâ might misleadâand I think a full fleshing out of the âbackdrop pictureâ here requires more detail than Iâve gestured at above (and perhaps such a picture requires fuller automation of human cognitive abilities than the âadvanced capabilityâ condition used in this report implies). Thanks to Katja Grace and Ben Garfinkel for discussion. 6 The rest of this report focuses on a specific version of this sort of worry, in the context of artificial systems in particular. 1.2.4 Power This version centers on the following hypothesis: that by default, suitably strategic and intelligent agents, engaging in suitable types of planning, will have instrumental incentives to gain and maintain various types of power, since this power will help them pursue their objectives more effectively (see section 4.2 for more discussion). 15 The worry is that if we create and lose control of such agents, and their objectives are problematic, the result wonât just bedamageof the type that occurs, for example, when a plane crashes, or a nuclear plant melts downâdamage which, for all its costs, remains passive. Rather, the result will be highly-capable, non-human agents actively working to gain and maintain power over their environmentâagents in anadversarialrelationship with humans who donât want them to succeed. Nuclear contamination is hard to clean up, and to stop from spreading. But it isnâttryingto not get cleaned up, ortryingto spreadâand especially not with greater intelligence than the humans trying to contain it. 16 But the power-seeking agents just described would betrying, in sophisticated ways, to undermine our efforts to stop them. If such agents are sufficiently capable, and/or if sufficiently many of such failures occur, humans could end up permanently disempowered, relative to the power-seeking systems weâve created. This differenceâbetween the usual sort of damage that occurs when a piece of technology malfunc- tions, and the type that occurs when you lose control over strategically sophisticated, power-seeking agents whose objectives conflict with yoursâmarks a key distinction not just between worries about AI vs. other sorts of risks, but also between the specific type of AI-related worry I focus on in what follows, and the more inclusive set of worries that sometimes go under the heading of âAI alignment.â As with any technology, AI systems can fail to behave in the way that their designers intend. And because some AI systems pursue objectives, some such unintended behavior can result from problems with their objectives in particular (call this particular type of unintended behavior âmisalignedâ). And regardless of the intentions of the designers, the development and social impact of AI systems can fail, more broadly, to uphold and reflect important values. Problems in any of these veins are worth addressing. But it ispower-seeking, in particular, that seems to me the most salient route to existential catastrophe from unintended AI behavior. AI systems that donât seek to gain or maintain power may cause a lot of harm, but this harm is more easily limited by the power they already have. And such systems, by hypothesis, wonât try to maintain that power if/when humans try to stop them. Hence, itâs much harder to see why humans would fail to notice, contain, and correct the problem before it reaches an existential scale. This, then, is the backdrop picture underlying the specific arguments Iâl examine: non-human agents, much more intelligent than humans, would likely be potent forces in the world; and if their objectives are problematic, some of them may, by default, seek to gain and maintain power. Building such agents therefore seems like it could risk human disempowerment, at least in principle. Letâs look more closely at whether to expect this in practice. 2 Timelines For existential risks from power-seeking AI agents to arise, it needs to become possible and financially feasible for humans to build relevantly dangerous AI systems. In this section, I describe the type of system Iâm going to focus on, and I discuss the probability that humans learn to create such systems by 2070. 15 By âpowerâ I mean something like: the type of thing that helps a wide variety of agents pursue a wide variety of objectives in a given environment. For a more formal definition, see Turner et al (2020). 16 Harmful structures actively optimized for self-replicationâlike pathogens, computer viruses, and âgrey gooââare a closer analogy, but these, too, lack the relevant type of agency and intelligence. 7 2.1 Three key properties Iâl focus on systems with three properties: (a) advanced capabilities, (b) agentic planning, and (c) strategic awareness. 2.1.1 Advanced capabilities Iâl say that an AI system has âadvanced capabilitiesâ if it outperforms the best humans on some set of tasks which when performed at advanced levels grant significant power in todayâs world. The type of tasks I have in mind here include: scientific research, engineering, business/military/political strategy, hacking, and social persuasion/manipulation. An AI system with these capabilities can consist of many smaller systems interacting, but it should suffice to ~fully automate the tasks in question. The aim here is to hone in on systems whose capabilities make any power-seeking behavior they engage in worth taking seriously (in aggregate) as a potential route to the disempowerment of ~all humans. Such a condition does not, I think, require meeting various stronger conditions sometimes discussed 17 âfor example, âhuman-level AI,â 18 âsuperintelligence,â 19 or âAGI.â 20 That said, Iâm erring, here, on the side of including âweakerâ systems â including some that might not, on their own (or even in aggregate), be all that threatening (a fact worth bearing in mind in assigning overall probabilities). 21 Stronger conditions have less of this problem; and I expect much of the discussion in what follows to apply to stronger systems, too. Admittedly, the standard here is imprecise. And notably, what sorts of task performance yields what sort of real-world power can shift dramatically as the social and technological landscape changes (see discussion in 6.3). Iâm going to tolerate this imprecision in what follows, but those who want more precise definitions should feel free to use themâthe discussion doesnât depend heavily on the details. 22 2.1.2 Agentic planning Iâl say that a system engages in âagentic planningâ if it makes and executes plans, in pursuit of objectives, on the basis of models of the world (to me, this isnât all that different from bare âagency,â but I want to emphasize the planning aspect). 23 17 A level of AI progress that disempowered all humans would constitute âtransformative AIâ in the sense used by Karnofsky (2016): e.g. âAI that precipitates a transition comparable to (or more significant than) the agricultural or industrial revolution.â Butthatsort of transformation is precisely the type of thing weâre trying to forecast; e.g., itâs the result, not the cause. And disempowerment does not require other, more mechanistic standards for transformationâe.g. economic growth proceeding at particular rates (though we can argue, here, about the economic value that AI progress sufficient to disempower humans would represent). 18 This is used in various ways, to mean something like (a) a single AI system that is in some generic sense âas intelligentâ as a human; (b) a single AI system that can do anything that a given human (an average human? the âbestâ human? any human?) can do; (c) a level of automation such that unaided machines can perform roughly any task better and more cheaply than human workers (see Grace et al (2017)). My favorite is (c). 19 Bostrom (2014, Chapter 2) defines this as âany intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest.â A related concept requires cognitive performance that in some loose sense exceeds all of human civilizationâthough exactly how to understand this isnât quite clear (and human civilizationâs âcognitive abilitiesâ change over time). 20 âAGIâ is sometimes used as a substitute for some concept of âhuman-level AIâ; in other contexts, it refers specifically to some concept of humanlearningability (see e.g. Selsam (undated), and Arbital here (no author listed, but I believe it is Yudkowsky)), or some method of creating systems that can perform certain tasks (see Ngo (2020) on âgeneralization-based approaches). âAGIâ is often contrasted with ânarrow AIââthough in my opinion, this contrast too easily runs together a systemâs ability tolearntasks with its ability toperformthem. And sometimes, a given use of âAGIâ just means something like âyou know, the big AI thing;realAI; the special sauce; the thing everyone else is talking about.â 21 Put another way, Iâm erring on the side of ânecessaryâ rather than âsufficientâ for riskiness. I have yet to hear a good capability threshold that is both necessaryandsufficientâand Iâm skeptical that one exists. 22 My own favored more precise threshold would be something like âa level of automation such that unaided machines can perform roughly any task better and more cheaply than human workersâ; but I donât think this is necessary for the threats in question to arise. 23 I like emphasizing planning in part because it seems to me more tempting to speak loosely about systems as âtryingâ to do things (e.g., your microwave is âtryingâ to warm your food, a thermostat is âtryingâ to keep your house a certain temperature), than to speak of them as âplanningâ to do so. 8 The aim here (and in the next subsection) is to hone in on the type of goal-oriented cognition required for arguments about the instrumental value of gaining/maintaining power to be relevant to a systemâs behavior. 24 That said, muddyness about abstractions in this vicinity is one of my top candidates for ways arguments of the type I consider might mislead; and I expect that as we gain familiarity with advanced AI systems, many of the concepts we use to think about them will alter and improve. However: I also think of many worries about existential risk from misaligned power-seeking as centrally animated by some concepts in this broad vicinity (though not necessarily mine in particular). I donât think such concepts useless, or obviously dispensable to the argument; and for now, Iâl use them. 25 A few clarifications, though. I take as my paradigm a certain type of human cognitionâthe type, for example, involved in e.g., planning and then taking a trip from New York to San Francisco; reasoning about the safest way to cut down a tree, then doing it; designing a component of a particle collider; and so on. When I talk about agentic planning, Iâm talking, centrally, aboutthat. (Though exactly how much of human cognition is like this is a further questionâand similarity comes on a spectrum.) 26 We tend to think of this kind of cognition as (a) using a model of the world that represents causal relationships between actions/outputs/policies and outcomes to (b) select actions/outputs/policies that lead to outcomes that rate highly according to various (possibly quite complicated) criteria (e.g., objectives). 27 And we predict and explain human action on this basis, using certain sorts of common-sense distinctionsâfor example, between the role that someoneâs objectives play in explaining a given action, vs. their world-models and capabilities. There are also various algorithms that implement explicit procedures in the vein of (a) and (b)âalgorithms, for example, that search over and evaluate possible sequences of actions, or that backwards-chain from a desired end-state. People sometimes dispute whether humans âactuallyâ do something like (a) and (b) (let alone, implement something akin to a more formal algorithm). For present purposes, though, what matters is that they do somethingclose enoughto justify predicting their behavior (at least sometimes) with (a) and (b) in mind. This is the standard Iâl apply to the agentic planners I discuss in what follows. That is, I am not imposing any specific constraints on the cognitive algorithms or forms of representation required for âagentic planningââbut they need to beclose enough, at some level of abstraction, to (a) and (b) above to justify similar predictions. Indeed, I will speak as though agentic planners areactuallyusing models of the world to select and execute plans that promote their objectives (as opposed to, e.g., 24 We can imagine cases in which AI agents end up valuing power for its own sake, but Iâm not going to focus on those here. 25 Of course, you can just talk directly about AI systems that end up causing (suitably unintended) existential catastrophes, without invoking concepts like agency, objectives, etc. The question, though, is why one might expect that sort of behavior. And especially when the existential catastrophes involve AI power-seeking in particular (as Iâm inclined to think the most worrying ones do), I think something like âinstrumental convergenceâ is the strongest argument. That said, it might be possible to put instrumental convergence-type considerations in terms that appeal less to agency, objectives, strategic reasoning, and so forth, and more to the types of novel generalization we might expect from a given sort of training. I briefly explore this possibility in a footnote at the end of section 4.2. Iâm open to this approach, but my current feeling is that this argument is strongest when it leans on an (implicit or explicit) expectation that something like agentic planning, in pursuit of objectives, will develop in the systems in question. 26 Obviously, not everything humans do seems like agentic planning. Pulling your hand away from a hot stove, for example, seems more like a reflexâa basic âif-thenâ implemented by your nervous systemâthan an executed plan. And much of what we do in life can seem more like the stove than planning a trip to New York. Indeed, we often think of non-humans animals as acting on âinstinct,â or on the basis of comparatively simple heuristics, instead of on the basis of explicit plans, models, and objectives. Humans, we might think, are more like this than sometimes supposed. But lines here can get blurry fast. Ultimately, all complex cognition consists of a large number of much simpler steps or âif-thens.â Whether a given behavior is built out of instincts, heuristics, if-thens, etc does not settle whether it is relevantly similar to travel-planning (indeed, contrasts between âheuristicsâ and more systematic forms of cognition often seem to me under-specified). And similarity comes along a multi-dimensional spectrumâreflecting, for example, the sophistication and flexibility of the models and plans involved. 27 See Arbital (undated, and no author listed, but I believe itâs Yudkowsky) on âconsequentialist cognitionâ for more. Thanks to Paul Christiano for discussion. 9 âbehaving just likeâ they are doing this); by hypothesis, to the extent there is a difference here, itâs not a predictively relevant one. I am not assuming, though: â˘that the objectives need to involve an explicitly represented âobjective function,â especially in the sense at stake in formal optimization problems; 28 ⢠that an agentic planner need be helpfully understood as âmaximizing expected utilityâ; ⢠that it engages in agentic planning in all circumstances; ⢠that it needs to be possible to easilytellwhether a given system is engaging in agentic planning or not; â˘that the systemâs objectives are simple, âlong-termâ (see section 4.3.1.3), easily subsumable within a set of âintrinsic values,â or constant across circumstances; â˘that agentic planners cannot be constituted by many interacting, non-agentic-planning systems; 29 ⢠that the system is capable of self-modification or online learning (see section 4.3.2.2); ⢠that the system has any particular set of opportunities for action in the world (for example, systems that can only answer question can count as agents in this sense); ⢠that the systemâs plans are oriented towards the external environment, as opposed to e.g. plans for how to distribute its internal cognitive resources in performing a task; â˘that the agentic planning in question is âongoingâ and âopen-ended,â as opposed to âepisodicâ (see footnote for details). 30 MuZeroâa system which learns a model of a game (Chess, Go, Shogi, Atari) in order to plan future actionsâqualifies as an agentic planner in this sense, as do AlphaZero, AlphaGo, and (at least on my current understanding) various self-driving cars. Thermostats, bottlecaps, forest fires, balls rolling down hills, and robots twitching randomly donât qualifyâtheyâre not doing something close enough to planning, using a model of the world, in pursuit of the outcomes they cause. 31 In various other systems, itâs less clear. For example: AlphaStarâan AI system that plays Starcraftâ executes complex, flexible strategies over long time horizons, but the extent to which itâs doing something close enough to explicitly representing and acting in light of the long-term consequences of its actions seems (to me) like an open question. Similarly, we can imagine a system like GPT- 3âthat is, a large machine learning model trained to predict human-generated strings of textâthat engages in something like agentic planning in generating its output: but the outputs weâve observed need not tell us directly. 32 Iâl also note that a system that merelygeneratesplans (for example, a GPT-3-like system responding to the prompt âThe following is a statement of an excellent business plan:â), but which doesnât do so in order to select outputs/actions/policies that promote its objectives, doesnât qualify. 2.1.3 Strategic awareness Iâl say that an agentic planner has âstrategic awarenessâ if the models it uses in making plans are broad, informed, and sophisticated enough to represent with reasonable accuracy the causal upshot of gaining and maintaining different forms of power over humans and the real-world environment 28 Here Iâm responding in part to the definition offered by Hubinger et al (2019), and to subsequent debate. 29 Here, and in a number of subsequent bullet points, Iâm responding in part to the contrasts discussed by Drexler (2019). 30 For example, something can be an agent planneronce it is given a task; but that doesnât mean that when it hasnât been given a task, it has other, ongoing objectives. And a system can plan agentically in pursuit of taking a single action, as opposed to executing a series of actions. 31 These are examples references in e.g. Filan (2018), Shah (2018), and Flint (2020). Fires are an example Katja Grace often uses of a non-agentic system that meets many proposed definitions of agency (for example, tending to bring about a certain type of outcome with a fair amount of robustness). 32 The extent of the world-modeling and planning involved in many animal behaviors is similarly unclear. 10 (here, again, I am using the âactually this, or close enough to make no predictive differenceâ standard above). 33 Clearly, strategic awareness comes in degrees. Broadly and loosely, though, we can think of a strategically aware, planning agent as possessing models of the world that would allow it to answer questions like âwhat would happen if I had access to more computing powerâ and âwhat would happen if I tried to stop humans from turning me offâ about as well as humans can (and using those same models in generating plans). AlphaGo does not meet this condition (its models are limited to the Go board). Nor, I think, do self-driving cars. 34 And even if a GPT-3-like system had suitably sophisticated models, it would need tousethose models in generating plans in order to qualify. Letâs call a system with advanced capabilities that performs agentic planning with strategic awareness an APS (Advanced, Planning, Strategically aware) system. 2.2 Likelihood by 2070 How likely is it that it becomes possible and financially feasible to develop APS systems before 2070? Obviously, forecasts like this are difficult (and fuzzily defined), and I wonât spend much time on them here. My own views on this topic emerge in part from a set of investigations that Open Philanthropy has been conducting (see, for example, Cotra (2020), Roodman (2020), and Davidson (2021), which I encourage interested readers to investigate. Iâl add, though, a few quick data points: â˘Cotraâs model, which anchors on the human brain and extrapolates from scaling trends in contemporary machine learning, puts >65% on the development of âtransformative AIâ systems by 2070 (definition here). â˘Metaculus, a public forecasting platform, puts a median of 54% on âhuman-machine intelligent parity by 2040,â and a median of 2038 for the âdate the first AGI is publicly knownâ (as of mid-April 2021, see links for definitions). â˘Depending on how you ask them, experts in 2017 assign a median probability of >30% or >50% to âunaided machines can accomplish every task better and more cheaply than human workersâ by 2066, and a 3% or 10% chance to the âfull automation of laborâ by 2066 (though their views in this respect are notably inconsistent, and I donât think the specific numbers should be given much weight). Personally, Iâm at something like 65% on âdeveloping APS systems will be possible/financially feasible before 2070.â I can imagine going somewhat lower, but less than e.g. 10% seems to me weirdly confident (and I donât think the difficulty of forecasts like these licenses assuming that probability is very low, or treating it that way implicitly). 3 Incentives Letâs grant, then, that it becomes possible and financially feasible to develop APS systems. Should we expect relevant actors to have strong incentives to do so, especially on a widespread scale? Iâl assume that there are strong incentives to automate advanced capabilities, in general. But building strategically-aware agentic planners may not be the only way to do this. Indeed, there are various reasons we might expect an automated economy to focus on systems without such properties, namely: 35 â˘Many tasksâfor example, translating languages, classifying proteins, predicting human responses, and so forthâdonât seem to require agentic planning and strategic awareness, 33 See Yudkowsky here for the closely related concept of âbig-picture strategic awareness.â 34 My understanding is that the planning performed by self-driving cars uses some combination of long-term, google-maps-like planning, which relies on a fairly impoverished model of the world; and short term predictions of behavior on the road. Neither of these seems broad and sophisticated enough to answer questions of the type above. However, I havenât investigated this. 35 Here Iâm drawing heavily on a list of considerations in an unpublished document written by Ben Garfinkel. 11 at least at current levels of performance. 36 Perhaps all or most of the tasks involved in automating advanced capabilities will be this way. â˘In many contexts (for example, factory workers), there are benefits to specialization; and highly specialized systems may have less need for agentic planning and strategic awareness (though thereâs still a question of the planning and strategic awareness that specialized systems in combination might exhibit). See section 4.3.2.1 for more on this. â˘Current AI systems are, I think, some combination of non-agentic-planning and strategically unaware. Some of this is clearly a function of what we are currently able to build, but it may also be a clue as to what type of systems will be most economically important in future. â˘To the extent that agentic planning and strategic awareness create risks of the type I discuss below, this might incentivize focus on other types of systems. 37 â˘Agentic planning and strategic awareness may constitute or correlate with properties that ground moral concern for the AI system itself (though not all actors will treat concerns about the moral status of AI systems with equal weight; and considerations of this type could be ignored on a widespread scale). Indeed, for reasons of this type, I expect that to the extent we develop APS systems, weâl do so in a context of a wide variety of non-APS systems, which may themselves have large impacts on the world, and which may help manage risk from APS systems. 38 And it seems possible that APS systems just wonât be a very important part of the picture at all. Still, though, there are a number of reasons to expect AI progress to push in the direction of systems with agentic planning and strategic awareness. Iâl focus on three main types: 1.Agentic planning and strategic awareness both seem veryuseful. That is, many of the tasks we want AI systems to perform seem to require or substantially benefit from these abilities. 2.Given available techniques, it may be that the most efficient way todevelopAI systems that perform various valuable tasks involves developing strategically aware, agentic planners, even if other options are in principle available. 3.It might be difficult topreventagentic planning and strategic awareness from developing in suitably sophisticated and efficient systems. One note: when I talk in what follows about âincentives to create APS systems,â I mean this in a sense that covers 2 and 3 as well as 1. Thus, on this usage, if people will pay lots of money for tables (including flammable tables), and the only (or the cheapest/most efficient) way to make tables is out of flammable wood, then Iâl say that there are âincentivesâ to make flammable tables, even if people would pay just as much, or more, for fire-resistant tables. Letâs look at 1-3 in turn. 3.1 Usefulness I think that the strongest reason to expect APS systems is that they seem quiteuseful. 39 In particular: â˘Agentic planning seems like a very powerful and general way of interacting with an environmentâespecially a complex and novel environment that does not afford a lot of opportunities for trial and errorâin a way that results in a particular set of favored outcomes. Many tasks humans care about (creating and selling profitable products; designing and running successful scientific experiments; achieving political and military goals; etc) have 36 Though what sorts of methods of performing those tasks emerge in the limits of optimization is a further question; see 3.3 below. 37 Though not all relevant actors will treat these risks with equal cautionâsee 5.3.2. 38 Though note that an increasingly automated economy might also exacerbate some types of risks, since misaligned, power-seeking systems might be better-positioned to make use of automated rather than human- reliant infrastructure. I discuss this briefly in section 6.3.1. 39 See Drexler (2019), Chapter 12, for discussion and disagreement. Many of Drexlerâs arguments, though, focus on the value of self-improving and/or âunboundedâ agents, which arenât arenât my focus in this section (though see 4.3.2.2 for more on self-improvement, and 4.3.1.3 on something akin to boundedness). 12 this structure, as do many of the sub-tasks involved in those tasks (e.g., efficiently gather- ing and synthesizing relevant information, communicating with stakeholders, managing resource-allocation, etc). 40 Indeed, getting something done often seems like it requires, or at least benefits substantially, from (a) using a model of the world that reflects the relationship between action and outcome to (b) choose actions that lead to outcomes that score well according to some criteria. If our AI systems canât do this, then the scope of what they can do seems, naively, like it will be severely restricted. â˘Strategic awareness seems closely related to a basic capacity to âunderstand what is going on,â interact with other agents (including human agents) in the real world, and recognize available routes to achieving your objectivesâall of which seem very useful to performing tasks of the type just described. Indeed, to the extent that humans care about the strategic pursuit of e.g. business, military, and political objectives, and want to use AI systems in these domains, it seems like there will be incentives to create AI systems with the types of world models necessary for very sophisticated and complex types of strategic planningâincluding planning that involves recognizing and using available levers of real-world power. Note that the usefulness of agentic planning here is not limited to AI systems that are intuitively âacting directly in the worldâ (for example, via robot bodies, or without human oversight), as opposed to e.g. predicting the results of different actions, generating new ideas or designsâoutput that humans can then decide whether or not to act on. 41 Thus, for example, in sufficiently sophisticated cognitive systems, the task of predicting events or providing information might benefit from making and executing plans for how to process inputs, what data to gather and pay attention to, what lines of reasoning to pursue, and so forth. 42 That said, I think we should be cautious in predicting what degree of agentic planning and/or strategic awareness will be necessary or uniquely useful for performing what types of cognitive tasks. â˘Humans often plan in pursuit of objectives, on the basis of big-picture models of whatâs going on, and we are most familiar with human ways of performing tasks. But the constraints and selection processes that shape the development of our AI systems may be very different from the ones that shaped human evolution; and as a result, the future of AI may involve much less agentic planning and strategic awareness than anchoring on human intelligence might lead us to expect. Indeed, my impression is that the track record of claims of the form âHumans do X task using Y capabilities, so AI systems that do X will also need Y capabilities,â is quite poor. 43 â˘Naively construed, I can imagine arguments given about the usefulness/necessity of agentic planning and strategic awareness making the wrong predictions about current systems (or, e.g., about animal behaviors like squirrels burying nuts for the winter). Thus, for example, one might have expected writing complex code (see here, starting at 28:14), or playing StarCraft, to require a certain type of high-level planning; but whether GPT-3-like systems or AlphaStar in fact engage in such planning is unclear. Similarly, one might have expected automated driving to involve lots of big-picture understanding of âwhatâs going onâ; but current self-driving cars donât have much of this. In hindsight, perhaps this is obvious. But one wonders what future hindsight will reveal. Or put another way: we should be cautious about âobviously task X requires capability Y,â too. The compatibility of agentic planning and strategic awareness with modularity is also important. Suppose, for example, that you want to automate the long-term strategic planning performed by a CEO at a company. The best way of doing this may involve a suite of interacting, non-APS systems. Thus, as a toy example, one system might predict a planâs impact on company long-term earnings; another system might generate plans that the first predicts will lead to high earnings; a third might predict whether a given plan would be deemed legal and ethical by a panel of human judges; a fourth 40 For more discussion, see Yudkowsky here, on the âubiquity of consequentialism,â and on options for âsubvertingâ it. 41 See Karnofsky (2012) for discussion of âtool AIâ that suggests such a contrast. 42 See Bostrom (2014, p. 152-3, and p. 158). Training new systems, and learning from previous experience, also plausibly involves decisions that benefit from this sort of planning. See Branwen (2016) for more. 43 Thanks to Owain Evans for discussion. 13 might break a generated plan into sub-tasks to assign to further systems; and so forth. 44 None of these individual systems need be âtryingâ to optimize company long-term earnings in a way human judges would deem legal and ethical; but in combination, they create a system that exhibits agentic planning and strategic awareness with respect to this objective. 45 (Note that the agentic planning here is not âemergentâ in the sense of âaccidentalâ or âunanticipated.â Rather, weâre imagining a system specifically designed to perform an agentic planning functionâalbeit, using modular systems as components.) Modularity makes it possible for some sub-tasks, including various forms of oversight, to be performed by humansâan option that may well prove a useful check on the behavior of interacting non-APS systems, and a way of reducing the role of fully-automated APS-systems in the economy. 46 However, as AI systems increase in speed and sophistication, it also seems plausible that there will be efficiency pressures towards more and more complete automation, as slower and less sophisticated humans prove a greater and greater drag on what a system can do. 47 Overall, my main point here is that the space of tasks we want AI systems to perform plausibly involves a lot of âstrategically-aware agentic planning shaped holes.â Leaving such holes unfilled would reduce risks of the type I discuss; but it would also substantially curtail the ways in which AI can be useful. 48 3.2 Available techniques Even if some task doesnâtrequireagentic planning or strategic awareness, it may be that creating APS systems is the only route, or the most efficient route, to automating that task, given available techniques. For example, instead of automating tasks one by one, the best way to automate a wide range of tasksâespecially ones where we lack lots of training data, which may be the majority of useful tasksâis to create AI systems that are able to learn new tasks with very little data. And we can imagine scenarios in which the best way to dothatis by training agentic planners (for example, in multi-agent reinforcement learning environments, or via some process akin to evolutionary selection) with high-level, broad-scope pictures of âhow the world works,â and then fine tuning them on specific tasks. 49 Indeed, with respect to strategic awareness in particular, various current techniques for providing AI systems information about the worldâfor example, training them on large text corpora from the internetâseem ill-suited to limiting their understanding of their strategic situation. 3.3 Byproducts of sophistication Even if youâre not explicitly aiming for or anticipating agentic planning and strategic awareness in an AI system, these properties could in principle arise in unexpected ways regardless, and/or prove difficult to prevent. Thus, for example: â˘Optimizing a system to perform some not-intuitively-agential task (for example, predicting strings of text) could, given sufficient cognitive sophistication, result in internal computation 44 Note, here, that generating a plan is not the same as agentic planning. 45 Ben Garfinkel suggested this sort of example in discussion. See also similar examples in Drexler (2019). 46 Though if the humanâs performance of the sub-task does not meaningfully constrain or regulate the overall behavior of the system, then whether that task is performed by a human or an AI system may not make a difference. 47 See Drexler (2019, Chapter 24) for discussion and disagreement. 48 For related points, see Branwen (2018): âFundamentally, autonomous agent AIs are what we and the free marketwant; everything else is a surrogate or irrelevant loss function. We donât want low log-loss error on ImageNet, we want to refind a particular personal photo; we donât want excellent advice on which stock to buy for a few microseconds, we want a money pump spitting cash at us; we donât want a drone to tell us where Osama bin Laden was an hour ago (but not now), we want to have killed him on sight; we donât want good advice from Google Maps about what route to drive to our destination, we want to be at our destination without doing any driving etc. Idiosyncratic situations, legal regulation, fears of tail risks from very bad situations, worries about correlated or systematic failures (like hacking a drone fleet), and so on may slow or stop the adoption of Agent AIsâbut the pressure will always be there.â 49 Richard Ngo suggested this point in conversation; see also his discussion of the âgeneralization-based approachâ here. The training and fine-tuning used for GPT-3 may also be suggestive of patterns of this type. 14 that makes and executes plans, in pursuit of objectives, on the basis of broad and informed models of the real world, even if the designers of the system do not expect this (they may even be unable to tell whether it has occurred). Indeed, the likelihood of this seems correlated with the strength of the âusefulnessâ considerations in 3.1: insofar as agentic planning and broad-scope world-modeling are very useful types of cognition, we might expect to see them cropping up a lot in sufficiently optimized systems, whether we want them to or not. 50 ⢠A sophisticated system exposed to one source of information X might infer from this strategically important information Y, even if humans could not anticipate the possibility of such an inference. ⢠Modular systems interacting in complex and unanticipated ways could give rise to various types of agential behavior, even if they werenât designed to do so. Of these three reasons to expect APS systemsâtheir usefulness, the pressures exerted by available techniques, and the possibility that they arise as byproducts of sophisticationâI place the most weight on the first. Indeed, if it turns out that APS systems arenât uniquely useful for some of the tasks we want AI systems to perform, relative to non-APS systems, this would seem to me a substantial source of comfort. 4 Alignment Letâs assume that it becomes possible and financially feasible to create APS systems by 2070, and that there are significant incentives to do so. This section examines why we might expect it to be difficult to create systems of this kind that donât seek to gain and maintain power in unintended ways. 4.1 Definitions and clarifications Letâs start with a few definitions and clarifications. At a high-level, we want the impact of AI on the world to be good, just, fair, and so forthâor at least, not actively/catastrophically bad. Call this the challenge of âmaking AI go well.â 51 This is a very broad and complex challenge, much of which lies well outside the scope of this report. A narrower challenge is: making sure AI systems behave as their designers intend. Of course, the intentions of designers might not be good. But if wecanâtget our AI systems to behave as designers intend, this seems like a substantial barrier to good outcomes from AI more generally (though it is also, at least in many cases, a barrier to theusefulnessof AI; see section 5.3.3). 52 Granted, there are ambiguities about what sort of behavior counts as âintendedâ by designers (see footnote for discussion), but Iâm going to leave the notion vague for now, and assume that behaviors like lying, stealing money, resisting shut-down by appropriate channels, harming humans, and so forth are generally âunintended.â 53 50 See e.g. Yudkowsky (undated): âAnother concern is that consequentialism may to some extent be a convergent or default outcome of optimizing anything hard enough. E.g., although natural selection is a pseudoconsequentialist process, it optimized for reproductive capacity so hard that it eventually spit out some powerful organisms that were explicit cognitive consequentialists (aka humans).â 51 Here Iâm drawing on the framework used by Christiano (2020). 52 Following the framework in Critch and Krueger (2020), we can also draw additional distinctions, related to the number of human stake-holders and AI systems involved. Iâm focused here on what that framework would call âsingle-singleâ and âmulti-singleâ delegation: that is, the project of aligning AI behavior with the intentions of some human designer or set of designers. Critch argues (e.g., here) that multi-multi delegation should be our focus. I disagree with this (indeed, if we solve single-single alignment, I personally would be feelingsubstantiallybetter about the situation), but I wonât argue the point here. 53 In particular, the relationship between âintendedâ and âforeseenâ (for different levels of probability) is unclear. Thus, the designers of AlphaGo did not foresee the systemâs every move, but AlphaGoâs high-quality play was still âintendedâ at some higher level (thanks to David Roodman for suggesting this example). And itâs unclear how to classify cases where a designer thinks it either somewhat or highly likely that a system will engage in a certain type of (unwanted-by-the-designers) power-seeking, but deploys anyway (thanks to Katja Grace for emphasizing this type of case in discussion; see also her post here). Iâl count scenarios of this latter type as âunintendedâ (as distinct from âunforeseenâ), but this isnât particularly principled (indeed, Iâm not sure if a principled carving is ultimately in the offing). A final problem is that the intentions of designers 15 Iâl understandmisaligned behavioras a particular type of unintended behavior: namely, unintended behavior that arises specifically in virtue of problems with an AI systemâs objectives. Thus, for example, a designer might intend for an AI system to make money on the stock market, and the system might fail because it was mistaken about whether some stock would go up or down in price. But if it wastryingto make money on the stock market, then its behavior was unintended but not misaligned: the problem was with the AIâs competence in pursuing its objectives, not with the objectives themselves. 54 I donât, at present, have a rigorous account of how to attribute unintended behavior to problems with objectives vs. other problems; and I doubt the distinction will always be deep or easily drawn (this doesnât make it useless). Iâl lean on the intuitive notion for now; but if necessary, perhaps we could cease talk of âalignmentâ entirely, and simply focus on unintended behavior (perhaps of a suitably agentic kind), or perhaps simply catastrophic behavior, whatever its source. A characteristic feature of misaligned behavior is that it isunintendedbut stillcompetent. 55 That is, it looks less like an AI system âbreakingâ or âfailing in its efforts to do what designers want,â and more like an AI system trying, and perhaps succeeding, to do something designersdonâtwant it to do. In this sense, itâs less like a nuclear plant melting down, and more like a heat-seeking missile pursuing the wrong target; less like an employee giving a bad presentation, and more like an employee stealing money from the company. Iâl say that a system is âfully alignedâ if it does not engage in misaligned behavior in response to any inputs compatible with basic physical conditions of our universe (Iâl call these âphysics- compatibleâ inputs). 56 By âinputs,â I mean information the system receives via the channels intended by its designers (an input to GPT-3, for example, would be a text prompt). 57 I am not including processes that intervene in some other way on the internal state of the systemâfor example, by directly changing the weights in a neural network (analogy: a soldierâs loyalty need not withstand arbitrary types of brain surgery). 58 may simply not cover the range of inputs on which we wish to assess whether a systemâs behavior is aligned (e.g., âdid the designers intend for the robot to react this way to an asteroid impact?â may not have an answer: asteroid impacts werenât within the scope of their design intentions). Here, perhaps, something like âunwantedâ or âwould be unwantedâ is preferable. The reason Iâm not using âunwantedâ in general is that too many features of AI systems, including their capability problems, are âunwantedâ in some sense (for example, I might want AlphaStar to play even better Starcraft, or to make me lots of money on the stock market, or to help me find love). And âcatastrophicâ seems both too broad (not all catastrophic behavior scales in the problematic way power-seeking does, nor is it clear why catastrophic behavior would imply power-seeking) and too narrow (not all relevant power-seeking need be âcatastrophicâ on its own, especially pre-deployment, and/or in âmultipolarâ scenarios in which no one AI actor is dominant). 54 Note that there is a sense in which the AI here is âtrying to do something we donât want it to doââe.g., if we hold fixed its false beliefs about whether a given stock will go up or down, weâd prefer for it to be trying tolosemoney on the stock market (thanks to Paul Christiano for suggesting this point). Plausibly, calling this behavior âaligned,â then, requires reference to some holistic, common-sensical interpretation of the combination of objectives and capabilities we ultimately âhave in mind.â But as ever, settling on a perfect conceptualization is tricky. 55 See Mikulik (2019) for related discussion of â2-D robustness.â 56 Thus, it is physics-compatible for a randomly chosen bridge in Indiana to get hit, on a randomly chosen millisecond in May 2050, by a nuclear bomb, a 10 km-diameter asteroid, and a lightning bolt all at once; but not for the laws of physics to change, or for us all to be instantaneously transported to a galaxy far away. Obviously the scope here is very broad: but note that misaligned behavior is a different standard than âbadâ or even âcatastrophicâ behavior. It will always be possible to set up physics-compatible inputs where a system makes a mistake, or gets deceived, or acts in a way that results in catastrophic outcomes. To be misaligned, though, this behavior needs to arise from problems with the systemâs objectives in particular. Thus, for example, if Bob is a paper-clip maximizer, and he builds Fred, who is also a paper-clip maximizer, Fred will (on my definition) be fully-aligned with Bob as long as Fred keeps trying to maximize paperclips on all physics compatible-inputs (even though some of those inputs are such that trying to maximize paperclips actually minimizes them, kills Bob, etc). Thanks to Eliezer Yudkowsky, Rohin Shah, and Evan Hubinger for comments on the relevant scope here (which isnât to say they endorse my choice of definition). 57 There are some possible edge-cases here where a system is getting information via channels other than those the designers intended, but Iâl leave these to the side. The point here is that weâre interested in howthe system responds to inputs, not in how the world might change the system into something else. 58 Though here, too, lines may get hard to draw. 16 Sometimes, something like full alignment is framed by asking whether we would be OK with a systemâs objectives being pursued with ~arbitrary amounts of capability. 59 This isnât the concept I have in mind. 60 Rather, what matters is how the actual system, with its actual capabilities, responds to physics-compatible inputs. If, on some such inputs, it seeks to improve its own capabilities in misaligned ways; or if some inputs improve its capabilities in a way that results in misaligned behavior; then the system is less-than-fully aligned. But if there are no physics-compatible inputs that the actual system responds to in misaligned ways, then I see preventing changes in capability that might alter this fact as a separate problem (though it may be a problem that designers intentionally improving the systemâs capabilities should be especially concerned about). 61 Nor does full alignment require that the system âshare the designerâs valuesâ (and still less, the designerâs âutility function,â to the extent it makes sense to attribute one to them) in any particularly complete sense. 62 For example, we can imagine AI systems that just undergo some kind of controlled shut-down, or query relevant humans for guidance, if they receive an input designers didnât intend for them to operate on. Indeed, extreme motivational harmony seems like a strange condition to impose on e.g. a house cleaning robot, or a financial managerâneither of which, presumably, need know the designerâs or the userâs preferences about e.g. politics, or romantic partners. Another property in the vicinity of âfull alignmentâ and âsharing values,â but distinct from both, is Christianoâs (2018) notion of âintent alignment,â on which an AI system A is intent aligned with a human H iff âA is trying to do what H wants it to do,â where Aâs relationship to Hâs desires is, importantly,de dicto. That is, A doesnât just want the same things as H (e.g., âmaximize applesâ). 63 Rather, A has to have something like ahypothesisabout what H wants (e.g., it asks itself: âwould H want me to maximize apples, or oranges?â), and to act on that hypothesis. 64 Intent alignment seems a promising route to full alignment; but I wonât focus on it exclusively. Iâl say that a system is âpractically alignedâ if it does not engage in misaligned behavior on any of the inputs it willin factreceive. 65 Full alignment implies a highly reliable form of practical alignmentâe.g., one that does not depend at all on predicting or controlling the inputs a system receives. From a safety perspective, this seems a very desirable property, especially given the unique 59 See Yudkowsky on the âomnipotence test for AI safetyâ: âThe Omni Test is that an advanced AI should be expected to remain aligned, or not lead to catastrophic outcomes, or fail safely, even if it suddenly knows all facts and can directly ordain any possible outcome as an immediate choice. The policy proposal is that, among agents meant to act in the rich real world, any predicted behavior where the agent might act destructively if given unlimited power (rather than e.g. pausing for a safe user query) should be treated as a bug.â Talk of an âobjectiveâ such that the âoptimal policyâ on that objective leads to good outcomes is also reminiscent of something like the Omni Test. See e.g. Hubingerâs (2020) definition of âintent alignmentâ: âAn agent is intent aligned if the optimal policy for its behavioral objective is aligned with humans.â 60 Indeed, Iâm skeptical that it will generally be well-defined, even for agentic planners, what an arbitrarily capable system pursuing that agentâs objectives looks like (for example, Iâm skeptical that thereâs a single version of âomnipotent Joeââand still less, âomnipotent MuZero,â âomnipotent Coca-Cola company,â âomnipotent version of my momâs dogâ and so forth). If we want to talk about alignment properties that are robust to improvements in capability, I think we should talk about what sorts of behavior will result froma particular process for improving that systemâs capabilities(e.g,. a particular type of retraining, a particular way of scaling up its cognitive resources, and so forth). Thanks to Paul Christiano, Ben Garfinkel, Richard Ngo, and Rohin Shah for discussion. 61 But so, too, should designers be concerned about altering the systemâs objectives as they improve it. Note that Iâm also setting aside the problem (as it relates to a given system A) of how to make sure that, to the extent that system A builds anewsystem B, system B is fully-aligned, too (for example, if system B isnât fully aligned, but not because of any problems with system Aâs objectives, thatâs not, on my view, a problem with systemâs Aâs alignmentâthough it might be a problem more generally). 62 See e.g. Bostromâs 2015 TED talk: âThe point here is that we should not be confident in our ability to keep a superintelligent genie locked up in its bottle forever. Sooner or later, it will out. I believe that the answer here is to figure out how to create superintelligent A.I. such that even ifâwhenâit escapes, it is still safe because it is fundamentally on our side because it shares our values. I see no way around this difficult problem. . . The initial conditions for the intelligence explosion might need to be set up in just the right way if we are to have a controlled detonation.â 63 My sense is that this difference between Christianoâs notion of âintent alignment,â and a broader notion of âsharing values,â is sometimes overlooked. 64 Russellâs (2019) proposalâe.g., to make AI systems that pursue our objectives, but are uncertain what those objectives are (and see our behavior as evidence)âhas a somewhat similar flavor. 65 In the sense, practical alignment is a property that holdsrelativeto a set of inputs. 17 challenges and stakes that APS systems present (see section 4.4 for discussion). Indeed, if you are building AI systems that would e.g. harm or seize power over you, due to problems with their objectives, in some physics-compatible circumstances (see next section), this seems, at the least, a red flag about your project. 66 Ultimately, though, itâs practical alignment that we care about. 4.2 Power-seeking Not all misaligned AI behavior seems relevant to existential risk. Consider, for example, an AI system in charge of an electrical grid, whose designers intend it to send electricity to both town A and town B, but whose objectives have problems that cause it, during particular sorts of storms, to only send electricity to town A. This is misaligned behavior, and it may be quite harmful, but it poses no threat to the entire future. Rather, as I discussed in section 1.2.4, the type of misaligned AI behavior that I think creates the most existential risk involves misalignedpower-seekingin particular: that is, active efforts by an AI system to gain and maintain power in ways that designers didnât intend, arising from problems with that systemâs objectives. In the electrical grid case, the AI system hasnât been described as trying to gain power (for example, by trying to hack into more computing resources to better calculate how to get electricity to town A) or to maintain the power it already has (for example, by resisting human efforts to remove its influence over the grid). And in this sense, I think, itâs much less dangerous. Iâl say that a system is âfully PS-alignedâ if it doesnât engage in misaligned power-seeking in particular in response to any physics-compatible inputs. And Iâl say that a system is âpractically PS-alignedâ if it doesnât engage in misaligned power-seeking on any of the inputs it will in fact receive. A key hypothesis, some variant of which underlies much of the discourse about existential risk from AI, is that there is a close connection, in sufficiently advanced agents, between misaligned behavior in general, and misaligned power-seeking in particular. 67 Iâl formulate this hypothesis as: Instrumental Convergence: If an APS AI system is less-than-fully aligned, and some of its misaligned behavior involves strategically-aware agentic planning in pursuit of problematic objectives, then in general and by default, we should expect it to be less-than-fully PS-aligned, too. Why think this? The basic reason is that power is extremely useful to accomplishing objectivesâ indeed, it is so almost by definition. 68 So to the extent that an agent is engaging in unintended behavior in pursuit of problematic objectives, it will generally have incentives, other things equal, to gain and maintain forms of power in the processâincentives that strategically aware agentic planning puts it in a position to recognize and respond to. 69 One way of thinking about power of this kind is in terms of the number of âoptionsâ an agent has available. 70 Thus, if a policy seeks to promote some outcomes over others, then other things equal, a larger number of options makes it more likely that a more preferred outcome is accessible. Indeed, talking about âoptions-seeking,â instead of âpower-seeking,â might have less misleading connotations. What sorts of power might a system seek? Bostrom (2014) (following Omohundro (2008)) identifies a number of âconvergent instrumental goals,â each of which promotes an agentâs power to achieve its objectives. These include: â˘self-preservation (since an agentâs ongoing existence tends to promote the realization of those objectives); 66 See Yudkowskyâs discussion of âAI safety mindsetâ here. Yudkowsky sometimes frames this as: you should build systems that never, in practice, search for ways to kill you, rather than systems where the search comes up empty. 67 See e.g. Bostrom (2014, chapter 7); Russell (2019, Chapter 5), on âyou canât fetch the coffee if youâre deadâ; and Yudkowsky (undated). 68 For a formal version of a similar argument, see Turner et al (2019). 69 Of course, systems with non-problematic objectives will have incentives to seek forms of power, too; but the objectives themselves can encode what sorts of power-seeking are OK vs. not OK. 70 Thanks to Ben Garfinkel for helpful discussion. 18 â˘âgoal-content integrity,â e.g. preventing changes to its objectives (since agentâs pursuit of those objectives in particular tends to promote them); â˘improving its cognitive capability (since such capability tends to increase an agentâs success in pursuing its objectives); â˘technological development (since control over more powerful technology tends to be useful); ⢠resource-acquisition (since more resources tend to be useful, too). 71 We see examples of rudimentary AI systems âdiscoveringâ the usefulness of e.g. resource acquisition already. For example: when OpenAI trained two teams of AIs to play hide and seek in a simulated environment that included blocks and ramps that the AIs could move around and fix in place, the AIs learned strategies that depended crucially on acquiring control of the blocks and ramps in questionâ despite the fact that they were not given any direct incentives to interact with those objects (the hiders were simply rewarded for avoiding being seen by the seekers; the seekers, for seeing the hiders). 72 Why did this happen? Because in that environment, what the AIs do with the boxes/ramps can matter a lot to the reward signal, and in this sense, boxes and ramps are âresources,â which both types of AI have incentives to controlâe.g., in this case, to grab, move, and lock. 73 The AIs learned behavior responsive to this fact. Of course, this is a very simple, simulated environment, and the level of agentic planning it makes sense to ascribe to these AIs isnât clear. 74 But the basic dynamic that gives rise to this type of behavior seems likely to apply in much more complex, real-world contexts, and to more sophisticated systems, as well. If, in fact, the structure of a real-world environment is such that control over things like money, material goods, compute power, infrastructure, energy, skilled labor, social influence, etc would be useful to an AI systemâs pursuit of its objectives, then we should expect the planning performed by a sufficiently sophisticated, strategically aware AI agent to reflect this fact. And empirically, such resources are in fact useful for a very wide variety of objectives. Concrete examples of power-seeking (where unintended) might include AI systems trying to: break out of a contained environment; hack; get access to financial resources, or additional computing resources; make backup copies of themselves; gain unauthorized capabilities, sources of information, or channels of influence; mislead/lie to humans about their goals; resist or manipulate attempts to monitor/understand their behavior, retrain them, or shut them off; create/train new AI systems themselves; coordinate illicitly with other AI systems; impersonate humans; cause humans to do things for them; increase human reliance on them; manipulate human discourse and politics; weaken various human institutions and response capacities; take control of physical infrastructure like factories or scientific laboratories; cause certain types of technology and infrastructure to be developed; or directly harm/overpower humans. (See 6.3.1 for more detailed discussion). A few other clarifications: 71 Note that we can extend the list to include an agentâs instrumental incentives to promote the existence, goal-content integrity, cognitive capability, resources, etc of agents whose objectives are sufficiently similar to its ownâfor example, future agents in some training process, copies of itself, slightly modified versions of itself, and so forth. And we can imagine cases in which an agent isintrinsicallymotivated to gain/maintain various types of powerâfor example, because this was correlated with good performance during training (see Ngo (2020) for discussion). Iâl focus, though, on the instrumental usefulness of power. 72 Thus, the hiders learned to move and lock blocks to prevent the seekers from entering the room where the hiders were hiding; the seekers, in response, learned to move a ramp to give them access anyway; adjusting for this, the hiders learned to take control of that ramp before the seekers can get to it, and to lock it in the room as well. In another environment, the seekers learned to âsurfâ on boxes, and the hiders, to prevent this, learned to lockallboxes and ramps before the seekers can get to them. See videos in the blog post for details. 73 Indeed, in a competitive environment like this, agents have an incentive not just to control resources that are directly useful to their own goals, but also resources thatwouldbe useful to their opponents. Thus, the ramps arenât obviously useful to the hiders directly; but ramps can be used to bypass hider defenses, so hiders have an incentive to lock ramps in their room, or to remove from the environment entirely, all the same. This type of dynamic could be relevant to adversarial dynamics between humans and AIs, too. 74 And importantly, the trainers werenâttryingto disincentivize resource-seeking behavior; quite the contrary, the set-up seems (I havenât investigated the history of the experiments) to have been designed to test whether âemergent tool-useâ would occur. 19 â˘Not all misaligned power-seeking, even in APS systems, is particularly harmful, or intuitively worrying from an existential risk perspective. 75 But Iâm not going to try to narrow down any further. ⢠Instrumental convergence is not a conceptual claim, but rather an empirical claim that purports to apply to a wide variety of APS systems. In principle, for example, we can imagine APS systems that plan in pursuit of problematic objectives on some inputs, but which are nevertheless fully PS-aligned (or very close to it). 76 Consider, for example, an APS version of the electrical grid AI system above, which plans strategically in pursuit of directing electricity only to town A, but which just doesnât consider plans that involve seeking power. The in-principle possibility of strategic, agentic misalignment without PS-misalignment is important, though, since it might be realized in practice. Perhaps, for example, the type of training we should expect by default will reinforce cognitive habits in APS systems that steer away from searching over/evaluating plans that involve misaligned power-seeking, even if other types of misaligned behavior persist. Training, after all, will presumably strongly select against observable power-seeking behavior, and perhaps power- seekingcognition,too. 77 If it proves easy to create/train APS systems of this kind (even if they arenât fully aligned), this would be great news for existential safety (see section 4.3.1 for more discussion). Note, though, this requires that the planning performed by an APS system engaged in misaligned behavior be limited or âprunedâ in a specific way. That is, by hypothesis, such a system is using a broad-scope world model, capable of accurately representing the causal upshot of different forms of power-seeking, to plan in pursuit of problematic objectives. To the extent that power-seekingwould, in fact, promote those objectives, the APS systemâs planning has to consistently ignore/skip over this fact, despite being in an epistemic position to recognize it. 78 Perhaps this type of planning is easy, in practice, to select for; but naively, it feels to me like a delicate dance. 79 Relatedly, we can imagine systems whose objectives are problematic insomeway, but whichwouldnât be promoted by unintended forms of power-seeking. 80 Perhaps, for example, the electrical grid system plans strategically in order to direct electricity only to town A, but any plans that involve unintended power-seeking would score very low on its criteria for choice. The ease with which we can create systems of this type seems closely related to the ease with which we can create systems with non-problematic objectives more generally (again, see section 4.3.1 for discussion). Letâs look at two other objections to instrumental convergence. The first is that humans donât always seem particularly âpower-seeking.â 81 Humans care a lot about survival, and about certain resources (food, shelter, etc), but beyond that, we associate many forms of power-seeking with a certain kind of greed, ambition, or voraciousness, and with intuitively âresource-hungryâ goals, like âmaximize X across all of space and time.âSomestrategically-aware 75 Ben Garfinkel suggests the example of a robo-cop pinning down the wrong person, but remaining amenable to human instruction otherwise. And in general, it seems most dangerous if an APS system isusingits advanced capabilities in pursuing power; e.g., if Iâm a greater hacker, but a poor writer, and Iâm trying to get power via my journalism, Iâm less threatening. 76 Thanks to Ben Garfinkel for suggesting examples in this vein. 77 Thanks to Ben Garfinkel, Rohin Shah, and Tom Davidson for emphasizing points in this vicinity. 78 We can also imagine cases where the systemâs tendency to rule out certain plans is best interpreted as part of its objectives (e.g., such plans in fact arenât good, on its values). 79 Ben Garfinkel suggests the example of humans considering plans that involve murdering other humansâ something that plausibly happens quite rarely, perhaps because âdonât murderâ has been successfully reinforced. But for various reasons, murderingâin most of the relevant contextsâreallydoesnâtpromote someoneâs objec- tives (both because they intrinsically disvalue murder, and because of many other constraints and instrumental incentives). So it makes sense for them to adopt a policy of not considering it at all, or only very rarely. In contexts where murderdoesconsistently promote a humanâs objectives (perhaps sociopaths embedded in much more violent/lawless human contexts would be an example), I expect humans to consider plans that involve murdering much more often. 80 Distinctions between this and âit doesnât consider bad plansâ also seem blurry. 81 See e.g. Cegloskiâs (2016) âArgument From My Roommateâ; and also Pinker (2018, Chapter 19): âThere is no law of complex systems that says that intelligent agents must turn into ruthless conquistadors. Indeed, we know of one highly advanced form of intelligence that evolved without this defect. Theyâre called womenâ (p. 297). Thanks to Rohin Shah for discussion of the humans example. 20 human planners are like this, we might think, but not all of them: so strategically aware, agentic planning isnât, itself, the problem. I think there is an important point in this vicinity: namely, that power-seeking behavior,in practice, arises not just due to strategically-aware agentic planning, but due to the specific interaction between an agentâs capabilities, objectives, and circumstances. But I donât think this undermines the posited instrumental connection between strategically-aware agentic planning and power-seeking in general. Humans may not seek various types of powerin their current circumstancesâin which, for example, their capabilities are roughly similar to those of their peers, they are subject to various social/legal incentives and physical/temporal constraints, and in which many forms of power-seeking would violate ethical constraints they treat as intrinsically important. But almost all humans will seek to gain and maintain various types of power in some circumstances, and especially to the extent they have the capabilities and opportunities to get, use, and maintain that power with comparatively little cost. Thus, for most humans, it makes little sense to devote themselves to starting a billion dollar companyâthe returns to such effort are too low. But most humans will walk across the street to pick up a billion dollar check. Put more broadly: the power-seeking behavior humans display, when getting power is easy, seems to me quite compatible with the instrumental convergence thesis. And unchecked by ethics, constraints, and incentives (indeed, evenwhenchecked by these things) human power-seeking seems to me plenty dangerous, too. That said, the absence of various forms of overt power-seeking in humans may point to ways we could try to maintain control over less-than-fully PS-aligned APS systems (see 4.3 for more). A second objection (in possible tension with the first) is:humans(or, some humans) may be power- seeking, but this is a product of a specific evolutionary history (namely, one in which power-seeking was often directly selected for), which AI systems will not share. 82 Some versions of this objection simply neglect to address the instrumental convergence argument above 83 (and note, regardless, that some proposed ways of training AI systems resemble evolution in various respects). 84 But we can see stronger versions as pointing at a possibility similar to the one I mentioned above: namely, perhaps it just isnât that hard to train APS systems not to seek power in unintended ways, across a large enough range of inputs, if youâre activelytryingto do so (evolution wasnât). Again, I discuss possibilities and difficulties in this regard in the next section. Ultimately, Iâm sympathetic to the idea that we should expect, by default, to see incentives towards power-seeking reflected in the behavior of systems that engage in strategically aware agentic planning in pursuit of problematic objectives. However, this part of the overall argument is also one of my top candidates for ways that the abstractions employed might mislead. In particular, it requires the agentic planning and strategic awareness at stake be robust enough to license predictions of the form: âif (a) a system would be planning in pursuit of problematic objectives in circumstance C, (b) power-seeking in C would promote its objectives, and (c) the models it uses in planning put it in a position to recognize this, then we should expect power-seeking in C by default.â Iâve tried to build something like the validity of such predictions into my definitions of agentic planning and strategic awareness; but perhaps for sufficiently weak/loose versions of those concepts, such predictions are not warranted; and it seems possible to conflate weaker vs. stronger concepts at different points in oneâs reasoning, and/or to apply such concepts in contexts where they confuse rather than clarify. 85 82 See e.g. Zador and LeCun (2019) (and follow-up debate here), and Pinker (2018, Chapter 19). One can also imagine non-evolutionary versions of thisâe.g., ones that attribute human power-seeking tendencies to our culture, our economic system, and so forth. Indeed, Ceglowski (2016) can be read as suggesting something like this in the context of a particular demographic: after listing Bostromâs convergent instrumental goals, he writes: âIf you look at AI believers in Silicon Valley, this is the quasi-sociopathic checklist they themselves seem to be working from.â 83 That is, the argument isnât âhumans seek power, therefore AIs will tooâ; itâs âpower is useful for pursuing objectives, so AIs pursuing problematic objectives will have incentives to seek power, by default.â 84 Large, multi-agent reinforcement learning environments might be one example. And the âleagueâ used to train AlphaStar seems reminiscent of evolution in various ways. 85 Thanks to Rohin Shah for emphasizing possible objections in this vein, and for discussion. 21 Indeed, to avoid confusions in this vein, we might hope to jettison such concepts altogether. 86 And itâs possible to formulate arguments for something like instrumental convergence without them (see footnote for details). 87 But the reasoning in play seems to me at least somewhat different. 4.3 The challenge of practical PS-alignment Letâs grant that less-than-fully aligned APS systems will have at least some tendency towards misaligned, power-seeking behavior, by default. The challenge, then, is to prevent such behavior in practiceâwhether through alignment (including full alignment), or other means. How difficult will this be? Iâl break down the challenge as follows: 1.Designers need to cause the APS system to be such that the objectives it pursues on some set of inputs X do not give rise to misaligned power-seeking. 2. They (and other decision-makers) need to restrict the inputs the system receives to X. The larger the set of inputs X where 1 succeeds, the less reliance on 2 is required (and in the limit of full PS-alignment, 2 plays no role at all). Iâl consider a variety of ways one might try to control a systemâs objectives, capabilities, and circumstances to hit the right combination of 1 and 2. Some of these (or some mix) might well work. But all, I think, face problems. 4.3.1 Controlling objectives Much work on technical alignment focuses on controlling the objectives the AI systems we build pursue. Iâl start there. At any given point in the progress of AI technology, we will have some particular level of control in this respect, grounded in currently available techniques. Thus, for example, some methods of AI development require coding objectives by hand. Others (more prominent at present) shape an AIâs objectives via different sorts of training signalsâwhether generated algorithmically, or using human feedback/demonstration. And we can imagine future methods that allows us to e.g. control an AIâs objectives by editing the weights of a neural network directly; or to cause a system to pursue (for its own sake) any objective that we can articulate using English-language sentences that will be accurately, common-sensically, and charitably construed. The challenge of (1) is to use what methods in this respect are available to ensure PS-alignment on all inputs in X. But note that because available methods change, (1) is a moving target: new techniques and capabilities open up new options. I emphasize this because sometimes the challenge of AI alignment is framed as one of shaping an AIâs objectivesin a particular wayâfor example, via hand-written code, or via some sort of reward signal, or via English-language sentences that will be interpreted in literalistic and uncharitable terms. And this can make it seem like the challenge is centrally one of, e.g., coding, measuring, or articulating explicitly everything we value, or getting 86 See Wijk (2019) for one effort. 87 That is, we can say things like: âThis system behaves in a manner that promotes outcomes of a certain type in a wide range of complex circumstances. Circumstance C is one to which we should expect this behavior to generalize, and in circumstance C, power-seeking will promote outcomes of type X. Therefore, we should expect power-seeking in Câ (thanks to Paul Christiano for discussion). Here, a lot of work is done by âcircumstance C is one to which we should expect this behavior to generalizeââbut I expect this to be plausible in the context of system optimized across a sufficient range of circumstances. That said, if the system was optimized, in those circumstances, to cause Xwithout (observable) misaligned power-seeking, the case for expecting power-seeking in C seems weaker. So the question looms large: how much can selecting against observed power-seeking, in creating a system that causes X in many circumstances, eliminate behavior responsive to power-seekingâs usefulness for causing X (and many other types of outcome) across all physics-compatible inputs? And this question seems similar to the questions raised above (and discussed more in the next section) about the ease of selecting, in training, against forms of agentic planning that license misaligned power-seeking (and note, too, that one salient way in which training a system to cause X might lead to power-seeking behavior is if a suitable generalized ability to cause X tends to involve or require something like agentic-planning and strategic awareness). 22 AI systems to interpret instructions in common-sensical ways. These challenges may be relevant in some cases, but the core problem is not method-specific. 4.3.1.1Problems with proxiesThat said, many ways of attempting to control an AIâs objectives share a common challenge: namely, that giving an AI system a âproxy objectiveââthat is, an objective that reflects properties correlated with, but separable from, intended behaviorâcan result in behavior that weakens or breaks that correlation, especially as the power of the AIâs optimization for the proxy increases. For example: the behavior I want from an AI, and the behavior I would rate highly using some type of feedback, are well-correlated when I can monitor and understand the behavior in question. But if the AI is too sophisticated for me to understand everything that itâs doing, and/or if it can deceive me about its actions, the correlation weakens: the AI may be able to cause me to give high ratings to behavior I wouldnât (in my current state) endorse if I understood it betterâfor example, by hiding information about that behavior, or by manipulating my preferences. 88 This general problemâthat optimizing for a proxy correlated with but not identical to some intended outcome can break the correlation in questionâis familiar from human contexts. 89 And we already see analogous problems in existing AI systems. Thus: ⢠If we train an AI system to complete a boat race by rewarding it for hitting green blocks along the way, it learns to drive the boat in circles hitting the same blocks over and over. â˘If we train an AI system on human feedback on a grasping task, it learns to move its hand in a way that looks to the human like itâs grasping the object, even though it isnât. â˘If we reward an AI system for picking up apples in a simulated environment, but in a manner that depends on the locations of certain blocks, the agent learns to tamper with the blocks. See here for a much longer list of examples in this vein. These examples may seem easy to fixâand indeed, in the context of fairly weak systems, on a limited range of inputs, they generally are. But they illustrate the broader problem: systems optimizing for imperfect proxies often behave in unintended ways. Indeed, this tendency is closely connected to a core property that makes advanced AI useful: namely, the ability to find novel solutions and strategies that humans wouldnât think of. 90 When you donât know how an AI will achieve its objective, and that objective doesnât capture everything that you really want, then even for comparatively weak systems and simple tasks, itâs hard to anticipate how the systemâs way of achieving the objective will break its correlation with what you really want. And as the AIâs capacity to generate solutions we canât anticipate grows, the problem becomes more and more challenging. We can see a variety of techniques for controlling an AI systemâs objectives as mediated by some kind of âproxyâ or other. Thus: hand-coded objectives, simple metrics (clicks, profits, likes), algorith- mically generated training signals, human-generated data/feedback, and English-language sentences can all be seen as information structures that shape/serve as the basis for the AIâs optimization, but which can also fail to contain the information necessary to result in intended behavior across inputs. The challenge is to find a technique adequate in this respect. Human feedback seems likely to play a key role here. 91 And it may, ultimately, be enough. But notably, we need ways of drawing on this feedback that donât require unrealistic amounts of human supervision and human-generated data; 92 we need to ensure that such feedback captures our preferences about 88 Though what sorts of misaligned power-seeking in particular these behaviors would give rise to is a further question. 89 Thus, paying railroad builders by the mile of track that they lay incentivizes them to lay unnecessary track (from Wikipedia, here); trying to lower the cobra population by paying people to turn in dead cobras leads to people breeding cobras (from Wikipedia, here); if teachers take âcause my students to get high scores on standardized testsâ as their objective, theyâre incentivized to âteach to the testââan incentive that can work to the detriment of student education more broadly; and so on (see here for more examples). See Manheim and Garrabrant (2018) for an abstract categorization of dynamics of this kind. 90 Krakovna et al (2020) make this point well. 91 This paragraph draws heavily on discussion in Ngo (2020). 92 See e.g. Christiano et al (2017). 23 behavior that we canât directly understand and/or whose consequences we havenât yet seen; 93 we need ways of eliminating incentives to manipulate or mislead the human feedback mechanisms in question; and we need such methods to scale competitively as frontier AI capabilities increase. 94 Would it help if our AI systems could understand fuzzy human concepts like âhelpfulness,â âobe- dience,â âwhat humans would want,â and so forth? I expect it would, in various ways (though as I discuss below, this also opens up new opportunities for deception/manipulation). But note that the key issue isnât getting our AI systems tounderstandwhat objectives we want them to pursueâindeed, such understanding is plausibly on the critical path to increasing their capability, regardless of their alignment. Rather, the key issue is causing them topursuethose objectives for their own sake (though if they understand those objectives, but donât share them, we might also be able toincentivizethem to pursue such objectives for instrumental reasons). 95 4.3.1.2 Problems with searchMany techniques shape an AIâs objectives using proxies of one form another. But someânamely, those that involvesearchingover and selecting AI systems that perform well on some evaluation criteria, without controlling their objectives directlyâhave an additional problem: namely, evenifthose criteria fully capture the behavior we want, the resulting systems may not end up intrinsically motivated by the criteria in question. 96 Rather, they may end up with other objectives, pursuit of which correlated with good performance during the selection process, but which lead to unintended behavior on other inputs. Some think of human evolution as an example. 97 Someone interested in creating agents who pass on their genes to the next generation could run a training process similar to evolution, which searches over different agents, and selects for ones who pass on their genes (for example, by allowing ones who donât to die out). But this doesnât mean the resulting agents will be intrinsically motivated to pass on their genes. Humans, for example, are motivated by objectives that werecorrelatedwith passing on genes (for example, avoiding bodily harm, having sex, securing social status, etc), but which theyâl pursue in a manner that breaks such correlations, given the opportunity (for example, by using birth control, or remaining childless to further their careers). Rudimentary, evolved AI systems display analogous tendencies. Thus, when Ackley and Littman ran an evolutionary selection process in an environment with trees that allowed agents to hide from predators, the agents developed such a strong attraction to trees that (after reproductive age) they would starve to death in order to avoid leaving tree areas (what Ackley called âtree senilityâ). 98 And we can imagine other cases, with less evolution-like techniques. Suppose, for example, that we use gradient descent to train an AI system to reach the exit of a maze. 99 If the exit was marked by a green arrow on all the training data, the system could learn the objective âfind the green arrowâ rather than âfind the exit to the maze.â If we then give it a maze where the green arrowisnâtby the exit, it will search out the green arrow instead. Note that in both the evolution and the maze cases, we can imagine that the evaluation criteria (âpass on genes,â âexit mazeâ) fully capture and operationalize the intended behavior. Still, the designers lack the required degree of control over the objectives of the agents they create. How often will problems like this arise? Itâs an open empirical question. But some considerations make the possibility salient. Namely: 93 See Christiano and Amodei (2018) for discussion. Iterative amplification and distillation; debate; and recursive reward modeling can all be seen as efforts in this vein. 94 See 4.3.2.3 for more on scaling, and 5.3.1 for more on competition. 95 This is a point from Ord (2020). For example, if we train some set of sophisticated agents to get bananas, in a complex environment that requires understanding and modeling humans, they may end up capable of understanding quite accurately (even more accurately than us) what we have in mind when we talk about âaligned behavior,â and of behaving accordingly (for example, when we give them bananas for doing so). But their intrinsic objectives could still be focused centrally on bananas (or something else), and our abilities to control those objectives directly might remain quite limited. 96 My discussion in this section is centrally inspired by the discussion in Hubinger et al (2019), though I donât assume their specific set-up. 97 See e.g. Hubinger et al (2019, p. 6). 98 See Christian (2020), quote here. Iâm going off of Christianâs description, here, and havenât actually investigated these experiments. 99 This is an example from Hubinger (2020). 24 â˘Proxy goals correlated with the evaluation criteria may be simpler and therefore easier to learn, especially if the evaluation criteria are complex. In the context of evolution, for example, it seems much harder to evolve an agent whose mind represents a concept like âpassing on my genes,â and then takes doing this as its explicit goalâhumans, after all, didnât even have the concept of âgenesâ until very recentlyâthan to evolve an agent whose objectives reflect the relevance of things like bodily damage, sex, power, knowledge, etc to whether its genes get passed on (though starting with cognitively sophisticated agents might help in this respect). 100 ⢠Relatedly: if the âtrueâ objective function provides slower feedback, agents that pursue faster-feedback proxies have advantages. For example: in the game Montezumaâs Revenge, it helps to give an agent a direct incentive analogous to âcuriosityâ (e.g., it receives reward for finding sensory data it canât predict very well), because the gameâs âtrueâ objective (e.g., exiting a level by finding keys that require a large number of correct sequential steps to reach) is too difficult to train on. 101 â˘To the extent that many objectives wouldinstrumentallyincentivize good behavior in training (for example, because many objectives, when coupled with strategic awareness, incentivize gaining power in the world, and doing well in training leads to deployment/greater power in the world), but few involveintrinsicmotivation to engage in such behavior, we might think it more likely that selecting for good behavior leads to agents who behave well for instrumental reasons. 102 That said, it could also be that the agents who perform best according to some criteria, especially once theyâre sophisticated enough to understand what those criteria are, are the ones who are intrinsically motivated by those criteria. And even if such agents arenât selected for automatically, various techniques might help detect and address problematic cases. Notably, for example, we might actively search for inputs that will reveal problematic objectives, 103 and we might learn how to read off a systemâs objectives from its internal states. 104 Overall, though, ensuring robust forms of practical PS-alignment seems harder if available techniques search over systems that meet some external evaluation criteria, with little direct control over their objectives. And much of contemporary machine learning fits this bill. 4.3.1.3 MyopiaSome broad types of objectives seem to incentivize power-seeking on fewer physics-compatible inputs than others. Perhaps, then, we can aim at those, even if we lack more fine-grained control. Short-term (or, âmyopicâ) objectives seem especially interesting here. The most paradigmatically dangerous types of AI systems plan strategically in pursuit of long-term objectives, since longer time horizons leave more time to gain and use forms of power humans arenât making readily available, they more easily justify strategic but temporarily costly action (for example, trying to appear adequately aligned, in order to get deployed) aimed at such power. 105 Myopic agentic planners, by contrast, are on a much tighter schedule, and they have consequently weaker incentives to attempt forms of misaligned deception, resource-acquisition, etc that only pay off in the long-run (though even short spans of time can be enough to do a lot of harm, especially for extremely capable systemsâand the timespans âshort enough to be safeâ can alter if what one can do in a given span of time changes). 106 I think myopia might well help. But I see at least two problems with relying on it: 100 This is a point I believe I heard Evan Hubinger make on a podcast (either this one, or this one). 101 See Burda and Edwards (2018), and discussion in Christian (2020). 102 Thanks to Carl Shulman for discussion. See also Christiano (2018): âOne reason to be scared is that a wide variety of goals could lead to influence-seeking behavior, while the âintendedâ goal of a system is a narrower target, so we might expect influence-seeking behavior to be more common in the broader landscape of âpossible cognitive policies.â â 103 See Christianoâs discussion of âadversarial trainingâ here. 104 This is closely related to work on âinterpretabilityâ and âtransparencyââsee, e.g., Olah (2020) and related work. 105 Non-myopic agents also seem more likely to want to âkeepâ power for a very long time, whereas myopic agents are more likely to âgive upâ a given type of power once they are âdone with it.â 106 Thanks to Rohin Shah, Paul Christiano, and Carl Shulman for discussion. And note that a given operational- ization of âtimeâ can itself be vulnerable to various forms of manipulation (an AI could, for example, find ways to stop and start its internal clock). Thanks to Carl Shulman for suggesting this possibility. 25 â˘Plausibly, there will be demand for non-myopic agents. Humans (and human institutions) have long-term objectives, and will likely want long-term tasksârunning factories and companies, managing scientific experiments, pursuing political outcomesâautomated. Of course, myopic and/or non-APS systems can perform sub-tasks (including sub-tasks that involve generating long-term plans), and humans can stay in the loop; but as discussed section 3.1, there will plausibly be competitive pressures towards automating our pursuit of long-term objectives more and more fully. â˘The âsearchâ techniques discussed in the previous section may make ensuring myopia challenging. And various types of long-term training processesâfor example, reinforcement learning on tasks that involve many sequential stepsâseem likely to result in non-myopia by default (that said, myopia is a fairly coarse-grained property for an objective to possess, and may be easier to cause/check for than others). We can also imagine other ways of attempting to shrink (even if not to zero) the set of physics- compatible inputs on which an APS system engages in PS-misaligned behavior. For example: we might aim for objectives that penalize âhigh-impactâ action, 107 or that prohibit lying in particular, or that give intrinsic weight to various legal and ethical constraints, or that benefit less from marginal resources. But these face the same challenges re: proxies and search discussed in the last two sections. 4.3.2 Controlling capabilities AI alignment research typically focuses on controlling a systemâs objectives. But controlling its capabilities can play a role in practical PS-alignment, too. In particular: the less capable a system, the more easily its behaviorâincluding its tendencies to misaligned power-seekingâcan be anticipated and corrected. Less capable systems will also have a harder timegettingandkeepingpower, and a harder time making use of it, so they will have stronger incentives to cooperate with humans (rather than trying to e.g. deceive or overpower them), and to make do with the power and opportunities that humans provide them by default. Preventing agentic planning and strategic awareness in the first place would be one example of âcontrolling capabilitiesâ (see section 3); but there are other options, too. 4.3.2.1SpecializationIn particular, we might focus on building APS systems whose competence is as narrow and specialized as possible (though they are still, by hypothesis, agentic planners with strategic awareness). Discussion of AI risk often assumes the relevant systems are very âgeneralââe.g., individually capable of performing (or learning to perform) a very wide variety of tasks. But automating a wide variety of tasks doesnât require creating a single AI system that can perform (or learn to perform) all of them. Rather, we can create different, specialized systems for eachâand since such systems are less capable, the project of practical PS-alignment may be easier, and lower-stakes. An APS system skilled at a specific kind of scientific research, for example, but not at e.g. hacking, social persuasion, investing, and military strategy, seems much less dangerousâbut it seems comparably useful for curing cancer. And even if a systemâs competencies are broad, we can imagine that its strategically-aware agentic planning is only operative on some narrow set of inputs. Indeed, specialized systems have many benefits. 108 For example, they can be optimized more heavily for specific functions (to borrow an example from Ben Garfinkel, there is a reason that the flashlight, camera, speakers, etc on an iPhone are inferior to the best flashlights, cameras, etc). And we see incentives towards specialization and division of labor in human economies and organizations as well. Whatâs more, we will likely have much greater abilities to optimize AI systems for particular tasks than we do with humans. That said, generality has benefits, too. In particular: â˘Human workers with quite general skill-setsâCEOs, generals, Navy Seals, researchers with a broad knowledge of many domains, flexibly competent personal assistantsâare prized 107 See, e.g., the discussion of âimpact penaltiesâ in Krakovna et al (2019). 108 Here, and in the bulleted list of âbenefits of specializationâ below, Iâm drawing on a list of benefits in an unpublished document by Ben Garfinkel. See also Drexler (2019) for more discussion of the value of specialized systems. 26 in various contexts (even while specialization is prized in others). Automated systems need not be human-like in this respect (farmers, too, have quite general skill sets, but automated agriculture need not involve âfarmer botsâ), but it seems suggestive, at least, of economically-relevant environments in which general competence is useful. ⢠Specialized systems may be worse at responding flexibly to changing environments and task-requirements (e.g., itâs helpful not to have to buy new robots every time you redesign the factory or change the product being produced). 109 ⢠Multiple specialized systems can be less efficient to store and create (there is a reason you carry around an iPhone, rather than separate flashlights, cameras, microphones, etc); â˘If a task requires multiple competencies, specialized systems can be harder to coordinate (e.g., itâs helpful to have a single personal assistant, rather than one for email, one for scheduling, one for travel planning, one for research, etc). And a suitably coordinated set of specialized systems can end up acting as a quite general and agentic system. Whatâs more, just as available techniques may push the field towards agentic planning and strategic awareness (see section 3.2), so too might they push towards generality. 110 GPT-3, for example, is trained to a fairly general level capability via predicting text, and then later fine-tuned on specific tasks like coding. Indeed, such an approach might be necessary for tasks where data is too hard to come by or learn from directly (consider, for example, tasks like âdesigning a good railway systemâ or âbe an effective CEOâ); and more broadly, the most efficient route to wide-spread automation may be the creation of general-purpose agents that can learn ~arbitrary new tasks very efficiently (though those agents could also end up quite specialized later). 111 And note, too, that even specialized APS systems can be very dangerous. A system highly skilled at hacking into new computers and copying itself, for example, can spread far and wide; a system skilled in science can design a novel virus; a system with control over automated weapons can use them; a system skilled at social manipulation can turn an election; and so forth. Indeed, this is part of why I focused on âadvanced capabilitiesâ in 2.1.1, rather than something like âAGIâ or âsuperintelligence.â 4.3.2.2Preventing problematic improvementsNew capabilities can put a system in a position to gain and maintain power in ways it couldnât beforeâand hence, make new incentives action- relevant (if Bob learns how to hack into bank accounts, for example, his likelihood of considering and executing plans that involve such hacking will change). Practical PS-alignment may therefore require controlling the extent to which the inputs a system receives result in improved capabilities. This seems easier if the variables in the system that determine how it responds to inputs (for example, the weights in a neural network) stay fixed. But we may also want systems that mix task-performance and learning together, that ârememberâ previous events, and so forth; and predicting and controlling the capabilities such systems will develop could be difficult (especially if we donât understand well how they workâsee 4.4.1). Note, though, that this is a narrower challenge than making sure a systemâs PS-alignment is robust to anyincrease in capabilitiesâincluding, for example, increases that result from interventions other than exposure to physics-compatible inputs. Ultimately, we need to make sure that a system isnât exposed to non-input interventions that cause it (or, a new version of it) to seek power in misaligned ways, too; and efforts by designers to scale up a systemâs cognitive resources, training, and so forth will need to grapple with challenges in this vein. But as I discussed in 4.1, this sort of robustness is not a requirement for PS-alignment in my sense. 4.3.2.3ScalingStrategies for practical PS-alignment that rely on limiting a systemâs capabilities face a general problem: namely, that there are likely to be strong incentives to scale up the capabilities of frontier systems. 112 PS-alignment strategies that canât scale accordingly (and competitively) therefore risk obsolescence as state of the art capabilities advance. 109 Thanks to Carl Shulman for suggesting this example. 110 This paragraph draws centrally on the discussion in Ngo (2020). 111 This is a point from Bostrom (2014). 112 Though note that PS-alignment problems with more capable systems could complicate these dynamicsâsee 5.3.3 for more. 27 A key question for any such strategy, then, is whether it can translate, given success at some level of capability, into a different strategy that scales better. For example, we might try to achieve practical PS-alignment with some fairly advanced systems (including, perhaps, quite specialized onesâor, indeed, non-APS ones), and then use them to create new and superior PS-alignment strategies (indeed, as AI development itself becomes increasingly automated, automating alignment research will plausibly be necessary regardless). But note that plans of the form âcreate some practically PS-aligned systems, and ask them what the plan should beâ might just not work. For example, the new systems might not have adequate plans either. One might therefore need to create even more capable systems, whichalsomight not have adequate plans, and so forth, until one pushes up against (or perhaps, past) the limits of oneâs capacity to ensure practical PS-alignment. 4.3.3 Controlling circumstances So far in section 4.3, Iâve been talking about controlling âinternalâ properties of an APS system: namely, its objectives and capabilities. But we can control external circumstances, tooâand in particular, the type of options and incentives a system faces. Controlling options means controlling what a circumstance makes it possible for a system to do, even if it tried. Thus, using a computer without internet access might prevent certain types of hacking; a factory robot may not be able to access to the outside world; and so forth. Controlling incentives, by contrast, means controlling which options it makes sense to choose, given some set of objectives. Thus, perhaps an AI system could impersonate a human, or lie; but if it knows that it will be caught, and that being caught would be costly to its objectives, it might refrain. Or perhaps a system will receive more of a certain kind of reward for cooperating with humans, even though options for misaligned power-seeking are open. Human society relies heavily on controlling the options and incentives of agents with imperfectly aligned objectives. Thus: suppose I seek money for myself, and Bob seeks money for Bob. This need not be a problem when I hire Bob as a contractor. Rather: I pay him for his work; I donât give him access to the company bank account; and various social and legal factors reduce his incentives to try to steal from me, even if he could. A variety of similar strategies will plausibly be available and important with APS systems, too. Note, though, that Bobâs capabilities matter a lot, here. If he was better at hacking, my efforts to avoid giving him the option of accessing the company bank account might (unbeknownst to me) fail. If he was better at avoiding detection, his incentives not to steal might change; and so forth. PS-alignment strategies that rely on controlling options and incentives therefore require ways of exerting this control (e.g., mechanisms of security, monitoring, enforcement, etc) that scale with the capabilities of frontier APS systems. Note, though, that we need not rely solely onhuman abilities in this respect. For example, we might be able to use various non-APS systems and/or practically-aligned APS systems to help. 113 One other note: ensuring practical PS-alignment in a deployed system seems easier the more similar its deployment circumstances to the ones on which humans have observed and verified PS-aligned behavior (for example, during training, or pre-deployment testing). Indeed, ideally, one would want the deployment inputs to come from the same distribution as the training inputs. But in practice, and especially in strategically-aware systems, ensuring a close-to-identical distribution seems very difficult (if not impossible). This is partly because the world changes (and indeed, the actions of the APS system can themselves change it). 114 But also, to the extent that the distinction between training and deployment reflects some real difference in the agentâs level of influence on the world, this difference is itself a change in distributionâone that a sufficiently sophisticated agent might recognize. 113 One could even imagine trying to use practically PS-misaligned systems, though this seems dicey. 114 See Krueger et al (2020) and Christian (2020) for some discussion of this possibility. 28 4.4 Unusual difficulties From our current vantage point, ensuring PS-aligned behavior from APS systems across a wide range of inputs seems, to me, like it could well be difficult. But so, too, does building any kind of APS system appear difficult. Is there reason to think that by the time we figure out how to do the latter, we wonât have figured out the former as well? Itâs generally easier to create technology that fulfills some function F, than to create technology that does Fandmeets a given standard of safety and reliability, for the simple reason that meeting the relevant standard is an additional property, requiring additional effort. 115 Thus, itâs easier to build a plane that can fly at all than one that can fly safely and reliably in many conditions; easier to build an email client than a secure email client; easier to figure out how to cause a nuclear chain reaction than how to build a safe nuclear reactor; and so forth. Of course, we often reach adequate safety standards in the end. But at the very least, we expect some safety problems along the way (plane crashes, compromised email accounts, etc). We might expect something similar with ensuring PS-aligned behavior from powerful AI agents. But ensuring such behavior also poses a number of challenges that (most) other technologies donât. Here are a few salient to me. 4.4.1 Barriers to understanding Ensuring safety and reliability requires understanding a system well enough to predict its behavior. But plausibly, this is uniquely challenging in the context of a strategically aware agentic planner whose cognitive capabilities significantly exceed those of humans in a given domain. That is, the thinking and strategic decision-making of such an agent will likely reach a quite alien and opaque level of sophistication, which humans may be poorly positioned to anticipate and understand at the level required for ensuring PS-alignment. For example, it may consider many options humans never would; it may understand physical and social dynamics that humans do not; and so forth. 116 This issue seems especially salient in the current, machine-learning dominated AI paradigm, in which our ability to create an AI system that can perform some task (e.g., predicting text) often far exceeds our ability to understandhowthe system does what it does. We set various key high-level variables (the systemâs architecture, the number of parameters, the training process, the evaluation criteria), but the system that results is still, in many (though not all) respects, a black box. We must rely on further experiments to try to get some handle on what it knows, what it can do, and how it is liable to behave. 117 If this lack of understanding of the systems weâre building persists (for example, if we end up creating APS systems by training very large machine learning models on complex tasks, using techniques fairly similar to those used today), this could be an important and safety-relevant difference between AI systems and other types of technology. That is, we achieve high degrees of reliability and safety with technology like bridges, planes, rockets, and so forth in part via an understanding of the physical principles that govern the behavior of those systems; and we design them, part by part, in a manner that reflects and responds to those principles. This allows us to understand and predict their behavior in a wide range of circumstances. Searching over opaque/poorly-understood AI systems allows no such advantage. Of course, our understanding of how ML systems work will likely improve over timeâindeed, active research in this area (sometimes called âinterpretabilityâ) is ongoing. 118 But interpretability is no bottleneck to training bigger models on more complex tasksâor, plausibly, to the commercial viability of such models. And even as some researchers work on it, much of the fieldâs effort focuses on pushing forward with developing whatever capabilities we can, interpretable or no. 115 See Bostrom (2015): âMaking superintelligent A.I. is a really hard challenge. Making superintelligent A.I. that is safe involves some additional challenge on top of that. The risk is that if somebody figures out how to crack the first challenge without also having cracked the additional challenge of ensuring perfect safety.â 116 See Yudkowksy on âstrong cognitive uncontainability.â 117 See Flint (2020) on âsearch vs. designâ for related discussion. 118 See e.g. Olah (2020) and Goh et al (2021). 29 That said, understanding comes in many varieties and degrees; and itâs an empirical question what mix of experiment/search vs. first-principles understanding/design has actually been involved in ensuring the safety of different technologies (for example, in biology, or before advanced scientific understanding). 119 I expect that investigation of case studies in this respect would prove revealing. Ultimately, though, âmechanistic understandingâ isnât the fundamental issue. Even if we had such understanding on some level, the basic difficultyâthat is, understanding and predicting the behavior of an agent much more cognitively sophisticated than humansâmight well persist. 4.4.2 Adversarial dynamics Part of the challenge of ensuring PS-alignment isdetectingproblems with an APS systemâs objectivesâfor example, via various processes of testing and evaluation. However, if a system behaves in PS-misaligned ways in some circumstances, those circumstances could well include the very processes of testing and evaluation weâre trying to use to detect problems. Thus, for exam- ple, if a strategically-aware AI system has problematic objectives that it could better achieve if it passed some training/evaluation process, it may have incentives to try to manipulate that process, and/or deceive relevant humans about its degree of alignment. 120 And if it is better at this type of manipulation/deception than we are at detecting its bad behavior, the appearance of safety and reliability on various tests may not tell us very much about how the system is liable to behave in other circumstances. 121 This sort of dynamic applies to very few of the technologies weâre familiar with (disciplines like computer security, which involve actively anticipating the strategies available to adversaries, may be the closest analog). That is: planes, rockets, nuclear plants, and so forth may be dangerous and complicatedâbut they are never activelytryingto appear safer than they are, or to manipulate our processes of understanding and evaluating them. But at least in principle, sufficiently sophisticated AI systems could be doing this (though whether they will display this behavior in practice is another question). That is, PS-aligning APS systems requires dealing, not with passive tools that might malfunction, but with possibly-adversarial strategic agents that could be actively optimizing in opposition to your efforts to ensure safetyâand doing so using unprecedented degrees of cognitive capability, far exceeding that of humans. This seems a very significant additional challenge. 4.4.3 Stakes of error A final challenge comes from the escalating impact of certain types of mistakes. If a bridge fails, or a plane crashes, or a rocket explodes, the harm done is limited, contained, and passive. If an engineered virus escapes from the lab, however, it can spread rapidly, and become more and more difficult to contain as it goes. Viruses of this kind seem better analogs for practically PS-misaligned APS systems than planes and rockets. 119 And defining these distinctions may prove challenging. 120 Bostrom (2014) calls this a âtreacherous turn.â 121 Muelhauser (2021) describes a rudimentary form of this behavior occurring in simulated evolution ex- periments: âAs part of a project studying the evolution of (simulated) organisms, computer scientist Charles Ofria wanted to limit the replication rate of a digital organism. So, he programmed the system to pause after each mutation, measure the mutantâs replication rate in an isolated test environment, and delete the mutant if it replicated faster than its parent. However, the organisms evolved to recognize when they were in the test environment and âplay deadâ (pause replication) so they would not be eliminated and instead be kept in the population where they could continue to replicate outside the test environment. Once he discovered this, Ofria randomized the inputs of the test environment so that it couldnât be so easily detected, but the organisms evolved a new strategy, to probabilistically perform tasks that would accelerate their replication, thus slipping through the test environment some percentage of the time and continuing to accelerate their replication thereafter.â See also Lewis (2017) for examples of possible deceptive dynamics learned in the context of bargaining: âAnalysing the performance of our agents, we find evidence of sophisticated negotiation strategies. For example, we find instances of the model feigning interest in a valueless issue, so that it can later âcompromiseâ by conceding it. Deceit is a complex skill that requires hypothesising the other agentâs beliefs, and is learnt relatively late in child development (Talwar and Lee, 2002). Our agents have learnt to deceive without any explicit human design, simply by trying to achieve their goals.. . . Deception can be an effective negotiation tactic. We found numerous cases of our models initially feigning interest in a valueless item, only to later âcompromiseâ by conceding it.â (h/t to Luke Muelhauser for pointing out this example). 30 That is, practical PS-alignment failures involve highly-capable, strategically-aware agents applying their capabilities (including, perhaps, the ability to copy themselves) to gaining and maintaining power in the worldâand they may become more and more difficult to stop as their power grows. In dealing with systems that pose this sort of threat, there is much less room for the âerrorâ component of trial-and-error, because the stakes of error are so much higher. And whatever their present safety, most current technologies involved many errors (plane crashes, rocket explosions, etc) along the way. Indeed, if youâre trying to store an engineered virus that has a significant chance of killing ~the entire global population if it gets released, you need safety standardsmuchhigher than those we use, even now (after generations of trial and error), for bridges or planesâmuch higher, indeed, than we use for approximately anything (this is one key reason to never, ever create such a virus). For example, bridges need not be robust to nuclear attack; but the storage facility for such a virusshouldbeâsuch attacks just arenât sufficiently unlikely. We might view the threat of PS-misaligned behavior from sufficiently capable APS systems in similar terms. And human track records of ensuring safety and security in our highest-stakes contextsâBSL- 4 labs, nuclear power plants, nuclear weapons facilitiesâseem very far from comforting. 122 4.5 Overall difficulty Overall, my current best guess is that ensuring the full PS-alignment of APS systems is going to be very difficult, especially if we build them by searching over systems that satisfy external criteria, but which we donât understand deeply, and whose objectives we donât directly control. Itâs harder to reason in the abstract about the difficulty of practical PS-alignment, because itâs a much more flexible and contingent property: e.g., it depends crucially on the interactions between an agentâs capabilities, its objectives, and the circumstances it gets exposed to. And I think various of the tools discussed in section 4.3âfor example: focusing on specialized and/or myopic agents; restricting an agentâs capabilities; creating various sorts of incentives towards cooperation; using various types of non-agentic, strategically unaware, and/or practically aligned systems to help with oversight, incentive design, safety testing, etcâmay well prove useful. And even if these tools donât, themselves, scale in a way that can ensure practical PS-alignment at very high levels of capability, they may help us discover techniques that do. However, there are problems with those tools, too, and more general problems that make practical PS-alignment seem like it may be unusually challengingâfor example, difficulties understanding of the systems weâre building, the possibility of adversarial dynamics, and the extreme stakes of failure. And at a high-level, if you donât have full PS-alignment, youâre engaged in an effort to control powerful, strategically-aware agents who donât fully share your objectives, and who would seize power given certain opportunities. It seems, in general, an extremely dangerous game. 5 Deployment Letâs turn, now, to whether we should expect to actuallyseepractically PS-misaligned APS systems deployed in the world. The previous section doesnât settle this. In particular: if a technology is difficult to make safe, this doesnât mean that lots of people will use it in unsafe ways. Rather, they might adjust their usage to reflect the degree of safety achieved. Thus, if we couldnât build planes that reliably donât crash, we wouldnât expect to see people dying in plane crashes all the time (especially not after initial accidents); rather, weâd expect to see people not flying. And such caution becomes more likely as the stakes of safety failures increase. Absent counterargument, we might expect something similar with AI. Indeed, some amount of alignment seems like a significant constraint on the usefulness and commercial viability of AI technology generally. Thus, if problems with proxies, or search, make it difficult to give house- cleaning robots the right objectives, we shouldnât expect to see lots of such robots killing peopleâs cats (or children); rather, we should expect to see lots of difficulties making profitable house-cleaning 122 See Ord (2020) for some discussion of BSL-4 accidents. 31 robots. 123 Indeed, by the time self-driving cars see widespread use, they will likely be quite safe (maybetoosafe, relative to human drivers they couldâve replaced earlier). 124 Whatâs more, safety failures can result, for a developer/deployer, in significant social/regulatory backlash and economic cost. The 2017 crashes of Boeingâs 737 MAX aircraft, for example, resulted in an estimated ~$20 billion in direct costs, and tens of billions more in cancelled orders. And sufficiently severe forms of failure can result in direct bodily harm to decision-makers and their loved ones (everyone involved in creating a doomsday virus, for example, has a strong incentive to make sure itâs not released). Many incentives, then, favor safetyâand incentives to prevent harmful and large-scale forms of misaligned power-seeking seem especially clear. Faced with such incentives, why would anyone use, or deploy, a strategically-aware AI agent that will end up seeking power in unintended ways? Itâs an important question, and one Iâl look at in some detail. In particular, I think these considerations suggest that we should be less worried about practically PS-misaligned agents that are so unreliably well-behaved (at least externally) that they arenât useful, and more worried about practically PS- misaligned agents whose abilities (including their abilities to behave in the ways we want, when itâs useful for them to do so) make them at least superficially attractive to use/deployâbecause of e.g. the profit, social benefit, and/or strategic advantage that using/deploying them affords, or appears to afford. My central worry is that it will be substantially easier to buildthattype of agent than it will be to build agents that are genuinely practically PS-alignedâand that the beliefs and incentives of relevant actors will result in such practically PS-misaligned agents getting used/deployed regardless. 5.1 Timing of problems Iâl think of âdeploymentâ as the point where an AI system moves out of a develop- ment/laboratory/testing environment and into a position of real-world influence (even if this influence is mediated via e.g. humans following its instructions). 125 This isnât always a discrete point; some- times, for example, itâs an ongoing process, influenced by many individual decisions to accord an AI agent somewhat greater influence. 126 For simplicity, though, Iâl think of it as a discrete point in what followsâanalogous to the point where a product âlaunches,â âships,â or starts really getting âused.â We can distinguish between practical PS-alignment failures that occurbeforeandafterdeployment. In the first case, that is, some of the inputs on which an APS system seeks power in misaligned ways (call these its âbad inputsâ) occur in the set it encounters during training/development/testing; in the second case, it only hits bad inputs after deployment. Other things equal, pre-deployment PS-alignment failures are preferable, because an AI system will generally be better controlled, and less able to cause harm, during training/development/testing than deployment. Indeed, some pre-deployment testing will likely aim totriggerany liability to misaligned power-seeking that an agent has. That said, in sufficiently capable agents, misaligned power-seeking that occurs before deployment can still be quite harmfulâespecially if it goes uncorrected. Thus, for example, if a power-seeking system meant to be contained in some training environment, or limited in its means of influencing the outside world, manages to break out of that environment, and/or to obtain other forms of influence, it may succeed in gaining various types of real-world power even if it was never intentionally deployed (letâs call this âunintentional deploymentâ). 127 123 I believe I heard this example from Ben Garfinkel. 124 We might see GPT-3âs usefulness as bottlenecked in part by something like alignment, too. GPT-3âs outputs suggest that it has lots of knowledgeâfor example, about basic medicineâthat could be useful to users (see Cotra (2021)). But extracting that knowledge is difficult, partly because GPT-3 was trained, not to be useful to humans, but to predict the next token in human-written text. And its outputs reflect biases that make it a less attractive product to humans. 125 I mean the term âtestingâ here to refer to something still under control of the developers, akin to the type of âsafety testâ one might run on a car or a rocket. In this sense, testing is part of âtrainingâ more broadly. This contrasts with other possible uses of a âtrain/testâ distinction. 126 See e.g. the roll-out of GPT-3, which involved some people having initial access, and now a limited API that the lab still controls (Brian Christian discusses this example on a recent podcast). And note that if e.g. an agent continues learning âon the job,â lines between training and deployment blur yet further. 127 See Critch and Krueger (2020) for a more detailed breakdown of possible deployment scenarios. 32 Whatâs more, and importantly, even with the heightened monitoring that training/development/testing implies, not all forms of pre-deployment misaligned power-seeking will necessarily be detected. This is especially concerning in the context of the possibly adversarial dynamics discussed in section 4.4.2. That is, to the extent that an AI system is activelytryingto get deployed (for example, because the real- world influence deployment grants would offer greater opportunity to achieve its objectives), it might do things like try deceive or manipulate its trainers, and/or to differentiate between training/testing inputs and deployment inputs, such that it behaves badly only on the latter, and once its prospects for gaining/maintaining power are sufficiently good. 128 Indeed, in general, as the accuracy with which practically PS-misaligned systems can anticipate the consequences of their actions grows, we should expect them, other things equal, to engage less frequently in behavior that results in forms of detection/correction that hinder their pursuit of their objectives (e.g., to make fewer âmistakesâ). That said, other things may not be equal: for example, our ability to detect bad behavior can grow, too. Pre-deployment practical alignment failures are one key route to post-deployment ones: if the AI system was already engaging in PS-misaligned behavior before deployment (for example, deceiv- ing its trainers), it will likely continue to do so, given greater real-world influence (and again, the problem may continue to go undetected). But we can imagine post-deployment failures without pre-deployment failures, too. That is, even if the AI system was genuinely well-behaved during training/testing, deployment may bring new inputs, opportunities, and capabilities, which the train- ing/testing process was ill-suited to anticipating. 5.2 Decisions Letâs look at the factors relevant to deployment decisions in more detail. Consider some set of decision-makers (at e.g. a lab, a company, a government, etc) deciding whether or not to deploy a given APS system, or to pursue some other alternative (running more tests, re- designing/retraining the system, scrapping the project, etc). And assume, further, that this system is in fact practically PS-misaligned: that is, if deployed, it will hit bad inputs (it may already have done so), and seek power in misaligned ways. Why would anyone ever choose to deploy such a system? At a high-level, we can break down the factors influencing this decision into: (a) the decision-makersâ beliefsabout the practical PS-alignment of the system in question; and (b) the costs and benefits they treat as relevant, conditional on the system being practically PS-aligned or misaligned. Of course, in reality, there will likely be a diversity of beliefs and cost-benefit assessments at stake; but for simplicity Iâl treat the relevant decision-makers as unified. With respect to beliefs: we are assuming the system is in fact practically PS-misaligned, but decision- makers need not know this, or even assign it significant probability. Rather, the accuracy of their beliefs in this respect depend on a variety of factors, including: â˘the effort they have exerted to gather evidence about the agentâs alignment, and the power of the tools they have available for doing so; â˘their ability to anticipate and control the circumstances to which the agent will be ex- posed during deployment, and the similarity of those circumstances to the training/testing distribution; ⢠deception/manipulation of the training/testing process on the part of the AI system; ⢠their previous experience with practical alignment problems, and the salience of such problems in the field/the broader culture; ⢠the strength of their epistemology overall. Factors that might pull decision-makers to deploy the system, conditional on its practical PS- alignment, might include: ⢠profit, ⢠power (for themselves, for the lab/company, for country/allies, etc), ⢠helping solve social problems (curing diseases, designing greener technology), 128 See e.g. Christianoâs RSA-2048 example here. 33 ⢠prestige/âcreditâ, ⢠the thrill and momentum of scientific progress, 129 ⢠a (perceived) need to keep up with some competitor, ⢠and a desire to prevent someone else from deploying a comparable system first. Note, though, that decision-makers might expect some of these benefits to apply even if the system is practically PS-misaligned. Thus: â˘even if a system will seek power in misaligned ways later, decision-makers might still profit from deploying it, or use it to help solve various social problems, in the short term; ⢠decision-makers might expect that the misaligned power-seeking could be adequately con- tained/corrected, or that it would only occur on fairly small scales, or only on rare inputs. Factors that might push decision-makersawayfrom deploying, conditional on practical PS- misalignment, include: ⢠ways a given product wonât be successful/profitable if itâs unreliable or unsafe, ⢠legal/regulatory/reputational/economic costs from deploying systems that end up seeking power in misaligned and harmful ways, ⢠concern on the part of decision-makers to avoid any harm to themselves/their loved ones that could come from sufficiently large-scale PS-misalignment failures, ⢠altruistic concern to avoid the social costs of such failures. Figure 1, on the following page, summarizes these factors. 130 5.3 Key risk factors As even this quite simplified framework illustrates, many different factors affect deployment decisions. Their strength and interaction can vary, and their balance can shift over time (for example, as competitive dynamics alter, regulation increases, experience with PS-alignment problems grows, and so forth). A few factors in particular, though, seem to me especially important, and worth highlighting. 5.3.1 Externalities and competition It can be individually rational for a given actor to deploy a possibly PS-misaligned AI system, but still very bad in expectation for society overall, if societyâs interests arenât adequately reflected in the actorâs incentives. Climate change might be some analogy. Thus, the social costs of carbon emissions are not, at present, adequately reflected in the incentives of potential emittersâa fact often thought key to ongoing failures to curb net-harmful emissions. Something similar could hold true of the social costs of actors risking the deployment of practically PS-misaligned APS systems for the sake of e.g. profit, global power, and so forthâespecially given that profit, power, etc at stake could be very significant. In such a context, the risks and downsides that a less-than-fully altruistic actor would be willing to accept are one thing; the risk and downsides that human society as a whole (let alone all future generations of humans) would be willing to accept are quite another. Of course, the personal costs to decision-makers of sufficiently high-impact forms of PS-misalignment failure (analogous, for example, to an engineered virus) could be quite high (and in some cases, immediate)âa fact that suggests important disanalogies from climate change, where personal costs to emitters are generally both minimal and delayed. 131 But (as with emissions) a givenindividual choice to deploy need not increase the risk of such high-impact harms by much. Thus, the probability of PS-misaligned behavior from the particular system in question might be low; and/or the scale of that behaviorâs potential harm limited. 132 But over time, and across many actors (see next section), 129 See e.g. Geoffrey Hintonâs comments here about the prospects of discovery being âtoo sweet.â 130 Obviously, the decision-makers need not actually use an explicit cost-benefit/expected value framework in deciding; this is just a toy model. 131 Thanks to Ben Garfinkel for emphasizing considerations in this vein, and for discussion. 132 Though to the extent the harm is limited, the âstakes of errorâ discussed in section 4.4.3, for individual deployment events, will be much lower. 34 Figure 1: Factors relevant to deployment decisions. small risks can accumulate; and collectively, the PS-misaligned behavior of many individual (and indeed, uncoordinated) systems can add up to catastrophe, even if no single system or deployment event is the cause. 133 And in some cases (e.g., especially bad competitive dynamics, military conflict, cases where the upside of successful deployment is sufficiently high) some actors may knowingly accept non-trivial (though probably not overwhelming) risk of costs as severe as their own permanent disempowerment/death. 134 Dysfunctional forms of competition between actors could also incentivize risk-taking, especially to the extent that there are significant advantages to gaining/maintaining some relative position in an AI technology ârace.â 135 Because time and effort devoted to ensuring practical PS-alignment trade off against the speed with which one can scale up the capabilities of state of the art systems, an actor who mightâve otherwise decided to put in more of such time and effort, if the advantages of a given 133 Christianoâs (2018) second scenario is one example of this. 134 Many humans, for example, have proven willing to risk their lives for personal power, glory, national strength, military victory, a social cause, an ideology, even scientific discovery, etc. âX would involve someone knowingly risking deathâ doesnât seem to me especially conclusive evidence for âX will never happenââ especially if the risk is thought small, and the upside great. Thanks to Ben Garfinkel, Nick Beckstead, and Carl Shulman for discussion of points in this vicinity. 135 See Askell et al (2019) for discussion of first-mover advantages. Bostrom (2014)âs notion of a âdecisive strategic advantageâ at stake in the AI race is an extreme example. 35 relative position (for example, first-mover advantages) were secure, could be incentivized, in a more competitive context, to accept increased risk in order to gain or maintain such a position. Indeed, in an especially bad version of this dynamic, the other competitors might then be likewise incentivized to take on increased risk as well, thereby creating further incentives for the first actor take onmore risk, and so forthâan ongoing feedback loop of increasing pressure on all parties to either to up their risk tolerance, or drop out of the race. That said, granted that dynamics like this arepossible, itâs a substantially further question what sorts of externalities and competitive dynamics will actually apply to a given sort of AI technology in a given context. There are competitive dynamics and first-mover advantages in many industries (for example, pharmaceuticals), but that doesnât mean we see a race to the bottom on safety, possibly because various mechanismsâmarket forces, regulation, legal liabilityâhave raised the âbottomâ sufficiently high; 136 and many of these mechanisms function to incorporate various potential externalities as well. 137 And more generally: abstract, game-theoretical models are one thing; the concrete mess of social and political reality, quite another. 5.3.2 Number of relevant actors Oncesomeactors can create APS systems, then over time, and absent active efforts to the contrary, a larger and larger number of actors around the world will likely become able to do so as well. Absent significant coordination among such actors, then, it seems likely that we will see substantial variation in the beliefs, values, and incentives that inform their decision-making. Of course, safety-relevant properties of AI systems can be correlated, even without active coordination amongst relevant actors. The ease of ensuring practical PS-alignment, for example, represents one source of correlation, and there may be others (see Hubinger (2020) for discussion). But to the extent that such correlation does not ensure practical PS-alignment by default (such that e.g. some substantive level of caution and social responsibility is required), then larger numbers of relevant actors increase the risk that some sort of failure will occur. That is: even if many actors are adequately cautious and responsible, some plausibly wonât be. For example, some could be overconfident in the practical PS-alignment of the systems theyâve developed, or dismissive of the level of harm that PS-misaligned behavior would cause. Others may have greater tolerance for risk, or act with more concern for profit and individual power than for the social costs of their actions, or see themselves as having more to gain from a particular type of AI-based functionality. Some may be operating in the absence of various market, regulatory, and liability-related incentives that apply to other actors. And so on. 138 Indeed, to the extent that resources invested in ensuring practical PS-alignment trade off against resources invested in increasing the capabilities of the systems one builds, over time we might expect to see actors who invest less in alignment, and who take more risks, to scale up the capabilities of their systems faster. This could result in the competitive dynamics discussed above (e.g., other actors cut back on safety efforts to keep up, and/or deploy systems that wouldnât meet their own safety standards, but which are safer than the ones they expect competitors to deploy); but if other actors donâtcut back on safety as a result, the most powerful systems might end up increasingly in the hands of the least cautious and socially responsible actors (though there are also important correlations between social responsibility and factors like resource-access, talent, etc). That said, as above, similar abstract dynamics plausibly apply in many industries, and itâs an empirical question, dependent on a wide variety of factors, what sorts of safety problems actually result. AI is no different. 5.3.3 Bottlenecks on usefulness The benefits of deployment listed in the box aboveâprofit, power, prestige, solving social problems, etcâall require the APS system, once deployed, to beusefulin various ways. If such a system 136 See Askell et al (2019, p. 9)âs discussion of pharmaceuticals; and see Hunt (2020) on the possibilities of âraces to the top.â 137 Though as the case of climate change shows, they do not always do this adequatelyâand negative externali- ties from deploying practically PS-misaligned AI systems could be extreme. 138 Here we might think of analogies with the âwinnerâs curse.â 36 is misaligned in a way that renders it obviouslynotuseful, then, we shouldnât expect to see it intentionally deployed. As I noted at the beginning of the section, this is an important constraint on how we should expect alignment-like problems to show up in the real world, as opposed to the lab. In general, if we canât get our AI systems to do things like understand what we want, follow instructions in charitable and common-sensical ways, check in with us when theyâre uncertain, refrain from immediately trying to hack their reward systems, and so forth, then their usefulness to us will be severely limited. And even very incautious and socially-irresponsible actors are likely to test a system extensively before deploying it. The question, then, isnât whether relevant actors will intentionally deploy systems that are already blatantly failing to behave as they intend. The question is whether the standards for good behavior they apply during training/testing will be adequate to ensure that the systems in question wonât seek power in misaligned ways on any inputs post-deployment. The issue is that good behavior during (even fairly extensive) training/testing doesnât necessarily demonstrate this. This is partly due to possible deception/manipulation on the part of the AI systems (see next subsection). But even absent deception/manipulation of this kind, it can be extremely difficult for an actor to predict/test the AIâs behavior on the full range of post-deployment inputsâ especially in a rapidly changing world, in the absence of deep understanding of how the system works (see section 4.4.1), and if the AI system might gain new knowledge and capabilities post-deployment (see section 4.3.2.2). Indeed, I think that one of the central reasons we should expect to see practically PS-misaligned AI systems getting used/deployed is precisely that they willdemonstratea high degree of usefulness during training/testingâand consequently, it will be increasingly difficult to resist deploying them, especially in the context of competitive dynamics like the ones described in 5.3.1. Hereâs an analogy. Suppose that scientists create a new, genetically-engineered species of chimpanzee, whose cognitive capabilities significantly exceed those of humans. Initially, scientists confine these chimps in a laboratory environment, and incentivize them to perform various low-stakes intellectual tasks using rewards like food and entertainment. 139 And suppose, further, that these chimpanzees are clearly capable of generating things like vaccine designs, prototypes for new clean energy technology, cures for cancer, highly effective military/political/business strategies, and so forthâand that they will in fact do this, if you set up their incentives right (even though they donât intrinsically value being helpful to humans, and so are disposed, in some circumstances, to seize power for themselvesâfor example, if they can get more food and entertainment by doing so). In such a context, I think, it would become increasingly difficult for various actors around the world to resist drawing on the intellectual capabilities of the chimps in a manner that gives the chimps real-world forms of influence. If a new Covid-19 style pandemic started raging, for example, and we knew that the chimps could rapidly design a vaccine, there would be strong pressure to use them for doing so. If the chimps can help âusersâ win a senate race, or save the lives of millions, or end climate change, or make a billion dollars, or achieve military dominance, then some people, at least, will be strongly inclined to use them, even if there are risks involvedâand those whodonâtuse them will end up losing their senate races, falling behind their business and military competitors, and so forth. And even if the chimps, at the beginning, are appropriately contained and incentivized to be genuinely cooperative, it seems unsurprising if, as people draw on their capacities in more and more ways around the world, they get exposed to opportunities and circumstances that incentivize them to seek power for themselves, instead. Something similar, I think, might apply to APS AI systems. Indeed, even if peopleknow, or strongly suspect, that such systems would seek power in misaligned ways in some not-out-of-the-question circumstances, the pull towards using them for goals that matter a lot to us may simply be too great. When pandemics are raging, oceans are rising, parents and grandparents are dying of cancer, rival nations are gaining in power, and billions (or even trillions) of dollars are sitting on the table, concerns 139 Per my comments in section 1.1, Iâm leaving aside questions about the ethical implications of treating the chimps this way, despite their salience. 37 about science-fictiony risks from power-seeking AI systems may, especially forsomerelevant actors, take a backseat. 5.3.4 Deception This sort of issueâe.g., that âusefulâ may come apart from âpractically PS-alignedââcould be importantly exacerbated by the fact, discussed in section 4.4.2 and elsewhere, that less-than-fully aligned APS systems with suitably long-term objectives may be activelyoptimizingfor getting deployed, since deployment grants them greater influence in the world. This could incentivize them to deceive or manipulate relevant decision-makersâand if they are very capable, their abilities in this respect may exceed our ability to detect and correct the deceitful/manipulative behavior in question. For example: we should expect suitably sophisticated and strategically aware systems tounderstand what sorts of behavior humans are looking for during training/testing, even if their objectives donât intrinsically motivate such behavior. So if they are optimizing for getting deployed, they will have strong instrumental incentives to behave well, to demonstrate the type of usefulness (described above) that will pull us towards deploying them, and to convince us that their objectives are fully (or at least sufficiently) aligned with ours. Indeed, theyâl even have incentives to appeal to ethical concerns about how it is morally appropriate to treat themâincentives that will applyregardlessof the legitimacy of those concerns (though I also expect such concerns tobelegitimate in at least some cases). Of course, human decision-makers will also be aware of the possibility of this sort of behavior. Indeed, if such behavior arises in sophisticated systems, we will likely see rudimentary forms of it in more rudimentary systems, tooâakin, perhaps, to the types of lies that children tell, but that adults can easily spot. But just because we know about a problem, and/or have encountered and maybe even solved it with more rudimentary systems, this doesnât mean weâl have solved it for all levels of cognitive capabilityâespecially levels much higher than our own. Detecting lies and manipulation attempts in your children is one thing; in adults much smarter and more strategically sophisticated than yourself, itâs quite another. And deceptive/manipulative AI systems will have incentives to make usthinkweâve solved the problem of AI deception/manipulation, even if we havenât. This isnât to say that humans will be actually fooled; and some AI systems might themselves be able to help with our efforts to detect deception in others. But unless we can develop deep understanding of and control over the objectives our AI systems are pursuing, evidence like âit performs well on all the tests we ran, including tests designed to detect deceptive/manipulative behaviorâ and âit clearly knows how to behave as we wantâ may tell us much less about its ultimate objectives, or about how it will behave once deployed, than we wish. And in the context of such uncertainty, some humans will be more willing to gamble than others. 5.4 Overall risk of problematic deployment Summing up this section, then: I donât think we should expect obviously non-useful, practically PS-misaligned APS systems to get intentionally deployed. Some systems might get deployed unintentionally, but the key risk, I think, is that an increasingly large number of relevant actors, with varying beliefs, incentives, and levels of social-responsibility, will be increasingly drawn to deploy strategically-aware AI agents that demonstrate their usefulness and apparent good behavior during training/testing (perhaps because they are deceptive/manipulative, or perhaps because their behavior is genuinely aligned on the training/testing inputs). If, as seems plausible to me, it will be much easier to build less-than-fully PS-aligned systems that meet this standard than fully PS-aligned ones, we should expect some of the systems people are strongly pulled towards deploying to be liable to misaligned power-seeking on at least some physics-compatible inputs. And if, as seems plausible to me, it is difficult to adequately predict and control the full range of inputs a system will receive, and how it will behave in response (especially if the world is rapidly changing, you donât deeply understand how the AI system works, and/or the AI has been actively deceptive or manipulative during the training/testing process), we should expect some such misaligned power-seeking to in fact occurânot just in a controlled laboratory/testing environment, but via channels of influence on the real world. It is a further question, though, what happens then: that is, whether misaligned efforts on the part of strategically aware AI agents to gain and maintain power actually succeed, and on what scale. Letâs turn to that question now. 38 6 Correction In many contexts, if an AI system starts seeking to gain/maintain power in unintended ways, the behavior may well be noticed, and the system prevented from gaining/maintaining the power it seeks. Letâs call this âcorrection.â Some types of correction might be easy (e.g., a lab notices that an AI system tried to open a Bitcoin wallet, and shuts it down). Others might be much more difficult and costly (for example, an AI system that has successfully hacked into and copied itself onto an unknown number of computers around the world might be quite difficult to eradicate). 140 Confronted with post-deployment PS-alignment failures, will humanityâs corrective efforts be enough to avert catastrophe? I think they might well; but it doesnât seem guaranteed. Letâs look at some considerations. 6.1 Take-off Discussions of existential risk from misaligned AI often focus on the transition from some lower (but still more advanced than today) level of frontier AI capability (call it âAâ) to some much higher and riskier level (call it âBâ). Call this transition âtake-off.â 141 In particular, some of the literature focuses on especially dramatic take-off scenarios. We can distinguish between a number of related variants: 1. âFast take-offâ: that is, escalation from A to B that proceeds very rapidly; 2.âDiscontinuous take-offâ: that is, escalation from A to B that proceeds much faster than some historical extrapolation would have predicted; 142 3. âConcentrated take-offâ: that is, escalation from A to B that leaves one actor or group (including a PS-misaligned AI system itself) with much more powerful AI-based capabilities than anyone else; 4.âIntelligence explosionâ: that is, AI-driven feedback loops lead to explosive growth in frontier AI capabilities, at least for some period (on my definition, this need not be driven by a single AI system âimproving itselfââsee below; and note that the assumption that feedback loops explode, rather than peter out, requires justification). 143 5.âRecursive self-improvementâ: that is, some particular AI system applying its capabilities to improvingitself, then repeatedly using its improved abilities to do this more (sometimes assumed or expected to lead to an intelligence explosion; though as above, feedback loops can just peter out instead). These are importantly distinct. Thus, for example, take-off can be fast, but still continuous (in line with previous trends), distributed (no actor or group is far ahead of another), and driven by factors other than AI-based feedback loops (let alone the self-improvement efforts of a single system). 144 Perhaps because of the emphasis in the previous literature, some people, in my experience, assume that existential risk from PS-misaligned AI requires some combination of (1)â(5). I disagree with this. I think (1)â(5) can make an important difference (see discussion of a few considerations below), but that serious risks can arise without them, too; and I wonât, in what follows, assume any of them. 6.2 Warning shots Weaker systems are easier to correct, and more likely to behave badly in contexts where they will get corrected (more strategic systems can better anticipate correction). Plausibly, then, if practical 140 This is a point made in Ord (2020), and by Bostrom in his TED talk. 141 See e.g. Bostrom (2014), Chapter 4. Obviously, the âlevelâ of AI capability available is highly multidimen- sional; but how we define it doesnât much matter at present. 142 See the AI Impacts project on discontinuities for more. 143 This could in principle proceed via non-artificial forms of intelligence, too; but Iâm leaving that aside. 144 Bostrom (2014)âs notion of a âhardware overhangâ is an example of fast take-off not driven by feedback loops. 39 PS-alignment is a problem, we will see and correct various forms of misaligned power-seeking in comparatively weak (but still strategically aware) systems (indeed, we may well devote a lot of the energy to trying totriggertendencies towards misaligned power-seeking, but in contained environments weâre confident we can control). Letâs call these âwarning shots.â Warning shots should get us very worried. If early, strategically-aware AI agents show tendencies to try to e.g. lie to humans, break out of contained environments, get unauthorized access to resources, manipulate reward channels, and so forth, this is important evidence about the degree of PS-alignment the techniques used to develop such systems achieveâand plausibly, of the probability of significant PS-alignment problems more generally. Receiving such evidence should make us sit up straight. Indeed, precisely because warning shots provide such tangible evidence, it seems preferable, other things equal, for them to occur earlier on in the process of AI development. Earlier warning shots are more easily controlled, and they leave more time for the research community and the world to understand their implications, and to try to address the problem. 145 By contrast, if there is very little calendar time between the first significant warning shots and the development of highly capable, strategically-aware agents, there will be less time for the evidence that warning shots provide to be reflected in the worldâs AI-related research and decision-making. 146 This is one of the worrying features of scenarios where frontier capabilities escalate very rapidly. Itâs sometimes thought that in scenarios where frontier AI capabilitiesdonâtescalate rapidly, warning shots will suffice to alert the world to PS-alignment problems (if they exist), and to prompt adequate responses. 147 But relying on this seems to me overoptimistic, for a number of reasons. First, recognizing a problem is distinct from solving it. Warning shots may prompt more attention to PS-alignment problems, but that attention may not be enough to find solutions, especially if the problems are difficult. And certain sorts of âsolutionsâ may function as band-aids; they correct a systemâs observed behavior, but not the underlying issue with its objectives. For example, if you train a system by penalizing it for lying, you may incentivize âdonât tell lies that would get detected,â as opposed to âdonât lieâ (and the training process itself might provide more information about which lies are detectable). Second, there are reasons to expect fewer warning shots as the strategic and cognitive capabilities of frontier systems increase,regardlessof whether techniques for ensuring the practical PS-alignment have adequately improved. This is because more capable systems, regardless of their PS-alignment, will be better able to model what sorts of behavior humans are looking for, and to forecast what attempts at power-seeking will be detected and correctedâa dynamic that could lead to a misleading impression that earlier problems have been adequately addressed; or even, that those problems stemmed from lack of intelligence rather than alignment. 148 145 Though note that warning shots for the most worrying types of misaligned power-seekingâfor example, acting aligned in a training/testing environment, while planning to seek power once deployedârequire fairly strategically sophisticated systems: e.g., ones that have models of themselves, humans, the world, the difference between training/testing and deployment, and so forth. For example: I donât think GPT-3 giving false information qualifies. 146 Though note that some decisionsâfor example, to recall a productâcan be made quickly. Thanks to Ben Garfinkel for suggesting this. 147 Garfinkel (2020): âIf you expect progress to be quite gradual, if this is a real issue, people should notice that this is an issue well before the point where itâs catastrophic. We donât have examples of this so far, but if itâs an issue, then it seems intuitively one should expect some indication of the interesting goal divergence or some indication of this interesting phenomenon of this new robustness of distribution shift failure before itâs at the point where things are totally out of hand. If thatâs the case, then people presumably or hopefully wonât plough ahead creating systems that keep failing in this horrible, confusing way. Weâl also have plenty of warning you need, to work on solutions to it.â 148 See e.g. the roll-out scenario described in Bostrom (2014). Itâs worth noting that in principle, if the ability to accurately assess whether a given instance of misaligned power-seeking will succeed arises sufficiently early in the AI systemâs weâre building/training, the period of time in which we see widespread and overt misaligned power-seeking from such agents could be quite short, or even non-existent. That is, it could be that the type of strategic awareness that makes an AI agent aware of the benefits of seeking real-world forms of power, resources, etc is closely akin to the type that makes an AI agent aware of and capable of avoiding the downsides of getting caught. If the former is in place, the latter may follow fastâeven if the trajectory of AI capability development in general is more gradual. 40 Third, even if there is widespread awareness that existing techniques for ensuring practical PS- alignment are inadequate, various actors might still push forward with scaling up and deploying highly-capable AI agents, either because they have lower risk estimates, or because they are willing to take more risks for the sake of profit, power, short-term social benefit, competitive advantage, etc. Indeed, it seems plausible to me that at a certain point, basically all of the reasonably cautious and socially-responsible actors around the world will know full well that various existing highly-capable AI agents are prone to misaligned power-seeking in certain not-out-of-the-question circumstances. But this wonât be enough to prevent such systems from getting used (and as I discussed in 5.3.1, if incautious actors use such systems, this will put competitive pressure on cautious actors to do so as well). Here, again, climate change might be an instructive analogy. The first calculations of the greenhouse effect occurred in 1896; the issue began to receive attention in the highest levels of national and international governance in the late 1960s; and scientific consensus began to form in the 1980s. 149 Yet here we are, more than 30 years later, with the problem unsolved, and continuing to escalateâthanks in part to the multiplicity of relevant actors (some of whom deny/minimize the problem even in the face of clear evidence), and the incentives and externalities faced by those in a position to do harm. There are many disanalogies between PS-alignment risk and climate change (notably, in the possibleâthough not strictly necessaryâimmediacy, ease of attribution, and directness of AI-related harms), but I find the comparison sobering regardless. 150 At least in some cases, âwarningsâ arenât enough. 6.3 Competition for power Deployed APS systems seeking power in misaligned ways would be competing for power with humans (and with each other). This subsection discusses some of the relevant features of that competition, and what sorts of mechanisms for seeking power might be available to the AI systems involved. A few points up front. First: discussion of existential risk from AI sometimes assumes that the central threat is a single PS-misaligned agent that comes to dominate the world as a wholeâe.g., what Bostrom (2014) calls a âunipolar scenario.â But the key risk is broader: namely, that ~all humans are permanently and collectivelydisempowered, whether at the hands of one AI system, or many. Thus, for example, existential catastrophe can stem from many PS-misaligned systems, each with comparable levels of capability, engaged in complex forms of competition and coordination (this is an example of a âmultipolarâ scenario): what matters is whetherhumanshave been left out. Here we might think again of analogies with chimpanzees: no single human or human institution rules the world, but the chimps are still disempowered relative to humans. In this sense, questions about e.g. what sorts of take-off lead to unipolar scenarios, and what sort will occur, donât settle questions about risk levels. Indeed, regardless of take-off, if we reach a point where (a) basically all of the most capable APS systems are seeking power in misaligned ways (for example, because of widespread problems ensuring scalable and competitive forms of practical PS-alignment), and (b) such systems drive and control most of the scientific, technological, and economic growth occurring in the world, then the human position seems to me tenuous. Second: the success or failure of a given instance of misaligned power-seeking depends both on the absolute capability of the power-seeking system,andon the strength of the constraints and opposition that it faces. 151 And in this latter respect, the world that future power-seeking AI systems would be operating in would likely be importantly different from the world of 2021. In particular, such a world would likely feature substantially more sophisticated capacities for detecting, constraining, responding to, and defending against problematic forms of AI behaviorâ capacities that may themselves be augmented by various types of AI technology, including non-agentic AI systems, specialized/myopic agents, and other AI systems that humans have succeeded in eliciting aligned behavior from, at least in some contexts. And even setting aside human opposition, a given PS-misaligned system might have other, different PS-misaligned systems to contend (or, perhaps, 149 See Wikipedia. 150 Thanks to Ben Garfinkel and Holden Karnofsky for suggesting disanalogies. 151 See e.g. Drexler (2019, Chapter 31)âs distinction between âsupercapabilitiesâ and âsuperpowers.â 41 cooperate) with as well; and the dynamics of cooperation and competition between human and non-human agents could become quite complex. Third, even in the context of absolute capability, there is an important difference between âbetter than humanâ and âarbitrarily capableââwhether in one domain, or manyâand PS-misaligned APS systems might fall on a wide variety of points in between. 152 We humans certainly canât predict all of the options and strategies that would be available to such systems, and we should be wary of ruling out possibilities with too much confidence. But we shouldnât assume that all physically possible forms of competence, knowledge, predictive ability, and so forth are in play, either. 6.3.1 Mechanisms With those points in mind, letâs look briefly at some factors and mechanisms that could be relevant to the scope and success of a given instance of PS-misaligned behavior. 153 Below are a few that seem salient to me. I include these partly for the sake of concreteness, and partly to illustrate that power-seeking by AI systemsâeven very capable onesâmay be subject to important constraints and limitations: the move from âseeks powerâ to âgets/keeps power,â especially on large scales, is non-trivial. 154 Obviously, any discussion in this respect will be speculative and incomplete (not to mention strange and âscience-fiction-yâ), and any actual PS-misaligned scenarios will likely involve factors and mechanisms that we are in a very poor position to anticipate. The examples I give are just meant as illustrations of possible dynamics. â˘Capability enhancement. The APS systems Iâve focused on need not outperform humans in all strategically relevant domains. 155 Nor, necessarily, will they be in a position to gain such capabilities, or to scale up capabilities they already haveâdespite plausible incentives to do so. And outperforming the best individual human workers, even by significant margins, does not imply outperforming the best humancollectives, equipped with the best available technology (including aligned AI technology). It therefore matters a lot what sort of position a PS-misaligned system will be in to scale up its capabilities, and/or to gain new ones. And note that capability gain/scaling can require more than âself-improvement abilityââit can also require compute resources, trial and error (some of which may be bottlenecked by real-world feedback), and other types of data. 156 â˘Secrecy. Misaligned power-seeking is harder if humans have detected it and are trying to stop it; and many different monitoring, oversight, and enforcement mechanisms are likely to be in place in the contexts most relevant to a systemâs power-seeking (e.g., governments, corporations, labs, financial institutions, compute clusters, factories, weapons systems, etc). The question of how well a system is able to hide its activities (for example, via operating in contexts with much less monitoring, or under the guise of some shell company, or by impersonating humans), for how long and from whom, is therefore important, too; and note that some activities (e.g., scientific and technological development, manufacturing, etc) can require significantâand hence, quite noticeableâinfrastructure. 157 ⢠Hacking. Various salient routes to additional power (for example, gaining additional compute resources, stealing money and information, taking control of automated infrastructure) proceed more smoothly if a PS-misaligned system can hack into new computer systems very easily. And even if the system is skilled at hacking in some absolute sense (it need not be), 152 See e.g. Arbitalâs discussion of different variants here. 153 This section draws on some of Bostromâs (2014) discussion of possible AI âsuperpowers.â 154 Though note that strategically sophisticated systems are less likely to make power-seeking efforts likely to fail. 155 This is an important difference from Bostromâs notion of âsuperintelligence.â 156 Indeed, to the extent that the training/development process that created the PS-misaligned system was itself extremely resource-intensive (as, for example, the frontier training runs in machine learning seem likely to be), the training/development process necessary for that system to improve itself could be even more so (though obviously, efficiency can also improve, too). 157 That said, it seems possible that at a certain point, PS-misaligned systems could operate out in the open as well, and perhaps with various types of human support/tolerance. 42 what matters is how its efforts would fare against the evolving cybersecurity defenses of its own time. 158 â˘Compute. Additional computing resources seem very useful to a PS-misaligned AI system, as they allow that system to run many copies of itself (and/or to run at higher speeds), and to perform other types of compute-intensive tasks (for example, training new AI systems). The ease with which an AI system can access such resources (via hacking, buying, renting, manufacturing, etc) therefore seems importantâand note that compute may be both prevalent and in high demand in an increasingly AI-driven economy, and that the manufacturing process may require significant time and/or resources (current semiconductor fabs, for example, cost billions of dollars). 159 â˘Financial resources. Obviously, getting money is helpful to seeking power. Options here could include: stealing, investing, selling products and services (electronic and physical), and trying to gain influence/control over wealthy actors. Exactly how much money a PS-misaligned system can gain control over, though, depends on its starting position, its capabilities, and the constraints and competition its faces. â˘Automated infrastructure. Other things equal, automated infrastructure (e.g., factories, labs, weapons systems, drones, vehicles, electrical grids, and so forth) seems much easier for a PS-misaligned AI system to control than infrastructure that requires human input and labor (and by the time weâre building APS systems, much more of the worldâs infrastructure may be automated). General-purpose robots and highly precise, rapid, and flexible forms of automated manufacturing seem especially important here, as control over these would give an AI system many further options for what to build/do. â˘Human labor. Depending on what sort of role humans have in the economy, and what sorts of control over flexible automated infrastructure a system has, it matters how easily a PS-misaligned system can make use of human labor. Possible routes to controlling such labor include: payment (including to humans willing to do illegal/illicit things), promises (of e.g. wealth, power, health, happiness), persuasion (for example, via highly optimized content, deception, or moral/ideological appeal), and coercion (for example, via physical threat, or blackmail). â˘Social influence. Beyond direct control over human labor, a PS-misaligned AI system would also likely benefit from broader forms of social influence, whether disguised or overt. Possible examples include: manipulating existing political and moral discourse in its favor (here we might think of Russiaâs intervention in the 2016 US election), steering the trajectory of scientific and technological development (for example, towards ideas, designs, and auto- mated infrastructure it could make use of), intervening on legal processes (e.g., lobbying, regulatory capture, preventing certain types of monitoring and enforcement), weakening human institutions and response capacities (for example, by sowing confusion, conflict, mistrust; and/or by compromising financial institutions, governments, law-enforcement agen- cies, mechanisms of coordination and information-sharing), and empowering/influencing specific actors (political candidates/parties, corporations, dictators, terrorists). ⢠Technology development. Advanced technology (improved computer hardware, rapid and precise manufacturing, advanced weaponry) is a clear route to power, and if such technology isnât already available or accessible at a sufficient scale, a PS-misaligned system might aim to develop or improve it. But as noted above, this process could require substantial time and resources: some types of science, for example, require laboratories, workers, physical experiments, and so forth. â˘Coordination.Copies of a PS-misaligned system may be able to coordinate and share information much more effectively than human groups (and indeed, we can even imagine scenarios where PS-misaligned AI systems with fairly different objectives communicate and coordinate in opposition to humans). This could be a substantial advantage. ⢠Destructive capacity. Ultimately, one salient route to disempowering humans would be widespread destruction, coercion, and even extinction; and the threats in this vein could play 158 See Garfinkel and Dafoe (2019) for discussion of how offense-defense dynamics might scale in cybersecurity. 159 Pinker (2018, p. 298) quotes a 2010 article by Ramez Naam pointing to physical/serial time bottlenecks to hardware development. 43 a key role in a PS-misaligned AI systemâs pursuit of other ends. 160 Possible mechanisms here include: biological/chemical/nuclear weapons; advanced and weaponized drones/robots; new types of advanced weaponry; ubiquitous monitoring, surveillance, and confinement; attacks on (or sufficient indifference to) background conditions of human survival (food, water, air, energy, habitable climate); and so on. 161 That said, note that a PS-misaligned sys- temâs central route to power can also rely heavily on peaceful means (for example, providing economically valuable goods and services). In my opinion, these factors and mechanisms are relevant in roughly similar ways in both âunipolarâ and âmultipolarâ scenariosâthough multipolar scenarios (by definition) involve more actors with comparable levels of power, and hence more complex competitive and cooperative dynamics. Note, too, that PS-misaligned behavior does not itself imply a willingness to make use of any specific mechanism of power-seeking. Perhaps, for example, we succeed in adequately eliminating a systemâs tendencies towards particularly egregious and harmful forms of misaligned power-seeking (e.g., directly harming humans), even if it remains practically PS-misaligned more broadly. 6.4 Corrective feedback loops In general, and partly due to various constraints the factors and mechanisms just discussed imply, I donât think it at all a foregone conclusion that APS systems seeking to gain/maintain power in misaligned ways, especially on very large scales, would succeed in doing so. Indeed, even beyond early warning shots in weak systems, it seems plausible to me that we see PS-alignment failures of escalating severity (e.g., deployed AI systems stealing money, seizing control of infrastructure, manipulating humans on large scales), some of which may be quite harmful, but which humans ultimately prove capable of containing and correcting. Whatâs more, and especially following high-impact incidents, specific instances of correction would likely trigger broader feedback loops. Perhaps products would be recalled, laws and regulations put in place, international agreements formed, markets altered, changes made to various practices and safety standards, and so forth. And we can imagine cases in which sufficiently scary and/or high-profile failures trigger extreme and globally-coordinated responses (for example, large-scale bans on certain automated systems) that would seem out of the question in less dire circumstances. Indeed, humans have some track record of eschewing and/or coordinating to avoid/limit technologies that a naive incentives analysis mightâve predicted weâd pursue more vigorously. Thus, for example, various high-profile nuclear accidents contributed to significant (indeed, plausibly net-harmful) reductions in the use of nuclear power in countries like the US; I expect that in 1950, I would have predicted greater proliferation and use of nuclear weapons by 2021 than weâve in fact seen; and we eschew human cloning, and certain types of human genetic engineering, centrally for ethical rather than technological reasons. 162 Corrective feedback loops in response to PS-alignment problems might draw on similar dispositions and coordination mechanisms (implicit or explicit). In this sense, and especially in scenarios where frontier capabilities escalate fairly gradually, the conditions under which AI systems are developed and deployed are likely to adjust dynamically to reflect PS-alignment problems that have arisen thus farâadjustments that may have important impacts on the beliefs, incentives, and constraints faced by AI developers and other relevant actors. 160 To be clear: the AI need not be intrinsically motivated by anything like âhatredâ of humans. Nor, indeed, need it want to literally use the atoms humans are made out of for anything else. Rather, if humans are actively threatening its pursuit of its objectives, or competing with it for power and resources, various types of destruction/harm might be instrumentally useful to it. 161 And note that in an actual militarized conflict, PS-misaligned AI systems might have various advantages over humansâfor example, they may not have the âreturn addressâ required for various forms of mutually assured destruction to gain traction; the conditions they require for survival might be very different from those of humans; they may not be affected by various attack vectors that harm humans (e.g., biological weapons); they may be able to copy/back themselves up in various ways humans canât; and so on. Obviously, though, humans could have various advantages as well (for example, the worldâs infrastructure is centrally optimized for human use). 162 Or at least, thatâs my current impression of the story re: nuclear power, cloning, and human genetic engineering. Chemical weapons might be another example in the vicinity: my impression is that use of chemical weapons in combat decreased in WWII, following WWI. See Pinker (2018) for some examples of dire, and false, predictions about nuclear proliferation and catastrophe. 44 Such corrective measuresâin conjunction with ongoing work improving our ability to ensure the practical alignment of the systems we build/deployâcould well be enough to avert catastrophe. But there are also a number of salient ways they could fail. One key failure mode arises in scenarios where frontier AI capabilities escalate very rapidlyâfor example, because the process of developing and improving frontier AI systems is itself increasingly automated. The plausibility of scenarios of this kind is a very open question, as are the precise timescales. But other things equal, more rapid capability escalation provides less time for the world to get experience with practical alignment failures, and to implement corrective measures; it creates a larger amount of general civilizational upheaval and disruption; and it may give AI-empowered actors (including PS-misaligned AI systems themselves) with a smaller calendar time âleadâ a larger absolute advantage over their competitors. Even if capabilities escalate fairly gradually, however, widespread practical PS-alignment failures may continue, even as society watches their severity escalateâespecially if the basic technical problem of ensuring practical PS-alignment has not been solved in a way that scales adequately and competitively with increasing capabilities. The âcompetitivenessâ dimension here is important. If a given method of ensuring the practical PS- alignment of a given system requires paying significant costs in resources and/or resulting capability (costs sometimes called an âalignment taxâ), relative to a riskier approach, then absent strong coordination, this will put more cautious actorsâincluding humans attempting to use practically aligned AI systems to defend against or correct the behavior of practically misaligned onesâat a competitive disadvantage. And as I noted in 5.3.2, levels of caution and social responsibility amongst relevant actors may vary widely. Whatâs more, just as pre-deployment practical PS-alignment failures may go undetected, so too may post-deployment failures. That is, it may make strategic sense for practically PS-misaligned agents with sufficiently long-term objectives to continue to behave well long after theyâve been deployed, because overt power-seeking is not yet worth the risks of detection and correction. Eventually, though, their incentives may alter (for example: if the activity of other misaligned systems has disrupted human civilization enough to change the cost-benefit balance of remaining cooperative, vs. seeking power for themselves). Overall, future scenarios in which global civilization grapples with practical PS-alignment failures in advanced AI agents, especially on a widespread scale or with escalating severity, are difficult to analyze in any detail, because so many actors, factors, and feedback loops are in play. Such scenarios need not, in themselves, spell existential catastrophe, if we can get our act together enough to correct the problem, and to prevent it from re-arising. But an adequate response will likely require addressing one or more of basic factors that gave rise to the issue in the first place: e.g., the difficulty of ensuring the practical PS-alignment of APS systems (especially in scalably competitive ways), the strong incentives to use/deploy such systems even if doing so risks practical PS-alignment failure, and the multiplicity of actors in a position to take such risks. It seems unsurprising if this proves difficult. 6.5 Sharing power Iâve been assuming, here, that humans will, by default, be unwilling to let AI systemskeepany power they succeed in gaining via misaligned behavior. But especially in multipolar scenarios, we can also imagine cases in which humans are either unable to correct a given type of misaligned power-seeking, or unwilling to pay the costs required to do so. 163 In such scenarios, it may be possible to reach various types of compromise arrangements, or to limit the impact of a given PS-misaligned system or systems to some contained domain, such that humans end upsharingpower with practically misaligned AI systems, but not losing power entirely. These are somewhat strange scenarios to imagine, and I wonât analyze them in any depth here. Iâl note, though, that if we reach such a point, the situation has likely become quite dire; and we might wonder, too, about its long-term stability. In particular, if the relevant PS-misaligned systems have 163 For example, perhaps a practically PS-misaligned AI system has created sufficiently many back-up copies of itself that humans canât destroy them all without engaging in some extreme type of technological âreset,â like shutting down the internet or destroying all existing computers; or perhaps an AI system has gained control of sufficiently powerful weapons that ongoing conflict would be extremely destructive. 45 sufficiently long-term goals, and have already been seeking power in misaligned and uncorrectable ways, then they will likely have incentives to continue to increase their powerâand their ongoing presence in the world will continue to give them opportunities to do so. 7 Catastrophe A final premise is that the permanent and unintentional disempowerment of ~all humans would be an existential catastrophe. Precise definitions can matter here, but loosely, and following Ord (2020), Iâl think of an existential catastrophe as an event that drastically reduces the value of the trajectories along which human civilization could realistically develop (see footnote for details and ambiguities). 164 Readers should feel free, though, to substitute in their own preferred definitionâthe broad idea is to hone in on a category of event that people concerned about what happens in the long-term future should be extremely concerned to prevent. Itâs possible to question whether humanityâs permanent and unintentional disempowerment at the hands of AI systems would qualify. In particular, if you are optimistic about the quality of the future that practically PS-misaligned AI systems would, by default, try to create, then the disempowerment of all humans, relative to those systems, will come at a much lower cost to the future (and perhaps even to the present) in expectation. One route to such optimism is via the belief that all or most cognitive systems (at least, of the type one expects humans to create) will converge on similar objectives in the limits of intelligence and understandingâperhaps because such objectives are âintrinsically rightâ (and motivating), or perhaps for some other reason. 165 My own view, shared by many, is that âintrinsic rightnessâ is a bad reason for expecting convergence, 166 but other possible reasonsârelated, for example, to various forms of cooperative game-theoretic behavior and self-modification that intelligent agents might 164 Ord (2020)âs definition of existential catastropheâthat is, âthe destruction of humanityâs longterm potentialââinvokes âhumanity,â but he notes that âIf we somehow give rise to new kinds of moral agents in the future, the term âhumanityâ in my definition should be taken to include themâ (p. 39); and he notes, too, that âIâm making a deliberate choice not to define the precise way in which the set of possible futures determines our potential. A simple approach would be to say that the value of our potential is the value of the best future open to us, so that an existential catastrophe occurs when the best remaining future is worth just a small fraction of the best future we could previously reach. Another approach would be to take account of the difficulty of achieving each possible future, for example defining the value of potential as the expected value of our future assuming we followed the best possible policy. But I leave a resolution of this to future workâ (p. 37, footnote 4). That is, Ord imagines a set of possible âopenâ futures, where the quality of humanityâs âpotentialâ is some (deliberately unspecified) function of that set. One issue here is that if we separate a futureâs âopen-nessâ from its probability of occurring, then very good futures are âopenâ to e.g. future totalitarian regimes, or future AI systems, to choose âif they wanted,â even if their doing so is exceedingly unlikelyâin the same sense that it is âopenâ to me to jump out a window, even though I wonât. But if we try to incorporate probability more directly (for example, by thinking of an existential catastrophe simply as some suitably drastic reduction in the expected value of the future), then we have to more explicitly incorporate further premises about the current expected value of the future; and suitably subjective notions of expected value raise their own issues. For example, if we use such notions in our definition, then getting bad newsâfor example, that the universe is much smaller than you thoughtâcan constitute an existential catastrophe; and I expect weâd also want to fix a specific sort of epistemic standard for assessing the expected value in question, so such that assigning subjective probabilities to some event being an existential catastrophe sounds less like âIâm at 50% credence that my credence is above 90% that X,â and more like âIâm at 50% credence that if I thought about this for 6 months, Iâd be at above 90% that Xâ. Like Ord, Iâm not going to try to resolve these issues here. 165 Note that this could be compatible with Bostromâs (2014) formulation of the âorthogonality thesisââe.g., âIntelligence and final goals are orthogonal: more or less any level of intelligence could in principle be combined with any final goal.â That is, Bostromâs formulation only applies to the âin principleâ possibility of combining high intelligence and any final goal. But there could still be strong correlations, attractors, etc in practice (this is a point I first heard from David Chalmers). 166 If, for example, you program a sophisticated AI system to try to lose at chessâsee, e.g., suicide chessâit wonât, as you increase its intelligence, start to see and respond to the âobjective rightnessâ of trying to win instead, or of trying to reduce poverty, or of spreading joy throughout the landâeven after learning what humans mean when they say âgood,â âright,â and so forth. See discussion in Russell (2019, p.~166). 46 converge on 167 âare more complicated to evaluate. 168 And we can imagine other routes to optimism as wellârelated, for example, to hypotheses about the default consciousness, pleasure, preference satisfaction, or partial alignment of the AI systems that disempowered humans. Iâm not going to dig in on this much. I do, though, want to reiterate that my concern here is with theunintentionaldisempowerment of humanity. That is, sharing power with AI agentsâespecially conscious and cooperative onesâmay ultimately be the right path for humanity to take. But if so, we want it to be a path wechose, on purpose, with full knowledge of what we were doing and why: we donât want to build AI agents who force such a path upon us, whether we like it or not. I think the moral situation here is actually quite complex. Suitably sophisticated AI systems may be moral patients; morally insensitive efforts to use, contain, train, and incentivize them risk serious harm; and such systems may, ultimately, have just claims to things like political rights, autonomy, and so forth. In fact, I think that part of what makes alignment important, even aside from its role in making AI safe, is its role in making our interactions with AI moral patients ethically acceptable. 169 Itâs one thing if such systems are intrinsically motivated to behave as we want; itâs another if they arenât, but weâre trying to get them to do so anyway. And more generally: once you build a moral patient, this creates strong moral reasons to treat it wellâand what âtreating artificial moral patients wellâ looks like seems to me a crucial question for humanity as we transition into an era of building systems that might qualify. At present, as far as I can tell, we have very little idea how to even identify what artificial systems warrant what types of moral concern. In a deep sense, I think, we know not what we do. But some moral patientsâand some agents who might, for all we know, be moral patients, but arenâtâwill also try to seize power for themselves, and will be willing to do things like harm humans in the process. So building new, very powerful agents who might be moral patients is, not surprisingly, both a morally and prudentially dangerous game: one that humanity, plausibly, is not ready for. My assumption, in this report, has been that unfortunately, weâor at least, some of usâare going to barrel ahead anyway, and I fear we will make many mistakes, both moral and prudential, along the way. The point, then, is not that humans have some deep right to power over AI systems we build. Rather, the point is to avoid losing control of our AI systems before weâve had time to develop the maturity to really understand what is at stake in different paths into the futureâincluding paths that involve sharing power with AI systemsâand to choose wisely amongst them. 8 Probabilities (May 2022 authorâs note: since making this report public in April 2021, my probability estimates â discussed in this section â have shifted. My current overall probability of existential catastrophe from power-seeking AI by 2070 is >10%.) To sum up, and with the preceding discussion in mind, letâs return to the full argument I opened with. To illustrate my current (unstable, subjective) epistemic relationship to the premises of this argument, Iâl add some provisional credences to each, along with a few words of explanation. 170 To be clear: I donât think Iâve formulated these premises with adequate precision to really âforecastâ in the sense relevant to e.g. prediction markets; and in general, the numbers here (and the exercise more broadly) should be held very lightly. Iâm offering these quantitative probabilities mostly because I think that doing so is preferable, relative to leaving things in purely qualitative terms (e.g., âsignificant riskâ), as a way of communicating my current best guesses about the issues Iâve discussed, and of facilitating productive disagreement. 171 This disagreement would beeasierif we 167 For especially exotic versions of this, see Oesterheld (2017) and Fox (2020). 168 Though the history of atrocities committed by strategic and intelligent humans does not seem comforting in this respect; and note that the incentives at stake here depend crucially on an agentâs empirical situation, and on its power relative to the other agents whose behavior is correlated with its own. In a context where misaligned AI systems are much more powerful than humans, it seems unwise to depend on their having and responding to instrumental, game-theoretic incentives to be particularly nice. 169 Thanks to Katja Grace for discussion of this point. 170 See footnote 4 here for more description of how to understand probabilities of this kind. 171 Ord (2020) discusses the downsides of leaving risk assessments vague in this way. 47 had more operationalized versions of the premises in question, and I encourage others interested in such operationalization to attempt it (though overly-precise versions can also artificially slim down the relevant scenarios). But my hope, in the meantime, is that this is better than nothing. 172 Even setting imprecisions aside, some worry that assigning conditional probabilities to the premises in a many-step argument risks various biases. 173 For example: I.It may be difficult to adequately imagineupdatingon the truth of previous premises (and hence, the premises may get treated as less correlated than they are). I.The overall verdict may be problematically sensitive to the number of premises (for example, we might be generically reluctant to assign very high probabilities to a premise, so additional premises will generally drive the final probability lower). I. The conclusion might be true, even if some of the premises are false. Note, though, that compressing an argument into very few premises (or just directly forecasting the conclusion) risks hiding conjunctiveness, too. 174 As an initial step in attempting to combat (I), and possibly (I), Iâve added a short appendix where I reformulate the argument using fewer premises; and to combat some other possible framing effects, I also offer versions in positiveâe.g., weâl be fineârather than negativeâe.g., weâre doomedâterms. And Iâve tried to make my own probabilities consistent across the board. 175 Obviously, though, biases in the vein of (I) and (I) can still remain (among many others). And note that (I), here, is true. That is, even limiting ourselves to existential catastrophes from power-seeking AI before 2070, estimates based on the premises I give are lower bounds: such catastrophes can occur without all those premises being true (see footnote for examples). 176 That said, I think these premises do a decent job of representing my own key uncertainties, at least, in âchunksâ that feel roughly right to me; and if I learned that one or more were false, Iâd feel alotless worried (at least for the next half-century). Iâl also note two high-level âoutside viewâ doubts I feel about the argument that follows: â˘The general picture Iâve discussed, even apart from specific assessments of a given premise, feels to me like âa very specific way things could go.â This isnât to say we canât ever make specific forecasts about the futureâI think we can (for example, about whether the economy will be bigger, the climate will be hotter, and so forth). But I have some background sense that visions of the future âof this typeâ (whatever that is) will generally be wrongâand often, in ways that discussion of those visions didnât countenance. â˘The people I talk to most about these issues (who are also, not coincidentally, my friends, colleagues, etc) areheavilyselected for being concerned about them. 177 I expect this to color my epistemic processes in various non-truth-tracking ways, many of which are difficult to correct for. I have tried to hazily incorporate these and other âoutside viewâ considerations into the probabilities reported here. But obviously, itâs difficult. Here, then, is the argument: 172 The imprecision here does mean that people disagreeing about specific premises may not have the same thing in mind; but a single person attempting to come to their own views can hold their own interpretations in mind. And the overall probability on permanent human disempowerment by 2070 seems less ambiguous in meaning. 173 See, for example, Yudkowsky on the âMultiple Stage Fallacy,â and discussion from Kaufman (2016). 174 More generally: the possibility of biases in one method of estimation (e.g., multiplying estimates for multiple conditional premises) doesnât show that some alternative methodology is preferable. And some things do actually require many conditional premises to be true. 175 I also played around a bit with a rough Guesstimate model that sampled from high/low ranges for the different premises. 176 For example, we might see unintentional deployment of practically PS-misaligned APS systems even if they arenât superficially attractive to deploy; practically PS-misaligned APS systems might be developed and deployed even absent strong incentives to develop them (for example, simply for the sake of scientific curiosity); systems that donât qualify as APS might seek power in misaligned ways; and so on. 177 And note, reader, that I, too, am heavily selected for concern about this issue, as someone who chose to write this report, to work at Open Philanthropy on existential risk, and so forth. 48 By 2070: 1.It will become possible and financially feasible to build APS systems. 178 Iâm going to say: 65%. This comes centrally from my own subjective forecast of the trajectory of AI progress (not discussed in this report), which draws on various recent investigations at Open Philanthropy, along with expert (and personal) opinion. I encourage readers with different forecasts to substitute their own numbers (and if you prefer to focus on a different milestone of AI progress, you can do that, too). 2.There will be strong incentives to build APS systems| (1). 179 Iâm going to say: 80%. This comes centrally from an expectation that agentic planning and strategic awareness will be either necessary or very helpful for a variety of tasks we want AI systems to perform. I also give some weight to the possibility that available techniques will push towards the development of systems with these properties; and/or that they will emerge as byproducts of making our systems increasingly sophisticated (whether we want them to or not). The 20% on false, here, comes centrally from the possibility that the combination of agentic planning and strategic awareness isnât actually that useful or necessary for many tasksâincluding tasks that intuitively seem like they would require it (Iâm wary, here, of relying too heavily on my âof course task X requires Yâ intuitions). For example, perhaps such tasks will mostly be performed using collections of modular/highly specialized systems that donât together constitute an APS system; and/or using neural networks that arenât, in the predictively relevant sense sketched in 2.1.2-3, agentic planning and strategically aware. (To be clear: I expect non-APS systems to play a key role in the economy regardless; in the scenarios where (2) is false, though, theyâre basically the only game in town.) 3.It will be much harder to develop APS systems that would be practically PS-aligned if deployed, than to develop APS systems that would be practically PS-misaligned if deployed (even if relevant decision-makers donât know this), but which are at least superficially attractive to deploy anyway| (1)â(2). Iâm going to say: 40%. I expect creating afullyPS-aligned APS system to be very difficult, relative to creating a less-than-fully PS-aligned one with very useful capabilitiesâespecially in a paradigm akin to current machine learning, in which one searches over systems that perform well according to some measurable behavioral metric, but whose objectives one does not directly control or understand (though 50 years is a long time to make progress in this respect). However, I find it much harder to think about the difficulty of creating a system that would bepracticallyPS-aligned, if deployed, relative to the difficulty of creating a system that would be practically PS-misaligned, if deployed, but which is still superficially attractive to deploy. Part of this uncertainty has to do with the âabsoluteâ difficulty of achieving practical PS-alignment, granted that you can build APS systems at all. A systemâs practical PS-alignment depends on the specific interaction between a number of variablesânotably, its capabilities (which could themselves be controlled/limited in various ways), its objectives (including the time horizon of the objectives in question), and the circumstances it will in fact exposed to (circumstances that could involve various physical constraints, monitoring mechanisms, and incentives, bolstered in power by difficult-to- anticipate future technology, including AI technology). I expect problems with proxies and search to make controlling objectives harder; and I expect barriers to understanding (along with adversarial dynamics, if they arise pre-deployment) to exacerbate difficulties more generally; but even so, it also seems possible to me that it wonât be âthathardâ (by the time we can build APS systems at all) 178 As a reminder, APS systems are ones with: (a)Advanced capability: they outperform the best humans on some set of tasks which when performed at an advanced level grant significant power in todayâs world (tasks like scientific research, business/military/political strategy, engineering, hacking, and social persuasion/manipulation); (b)Agentic planning: they make and execute plans, in pursuit of objectives, on the basis of models of the world; and (c)Strategic awareness:the models they use in making plans represent with reasonable accuracy the causal upshot of gaining and maintaining different forms of power over humans and the real-world environment. 179 As a reminder, Iâm using âincentivesâ in a manner such that, if people will buy tables, and the only (or the most efficient) tables you can build are flammable, then there are incentives to build flammable tables, even if people would buy/prefer fire-resistant ones. 49 to eliminate many tendencies towards misaligned power-seeking (for example, it seems plausible to me that selecting very strongly against (observable) misaligned power-seeking during training goes a long way), conditional on retaining realistic levels of control over a systemâs post-deployment capabilities and circumstances (though how often one can retain this control is a further question). Beyond this, though, Iâm also unsure about therelativedifficulty of creating practically PS-aligned systems, vs. creating systems that would be practically PS-misaligned, if deployed,but which are still superficially attractive to deploy. One commonly cited route to this is via a system actively pretending to be more aligned than it is. This seems possible, and predictable in some cases; but itâs also a fairly specific behavior, limited to systems with a particular pattern of incentives (for example, they need to be sufficiently non-myopic to care about getting deployed, there need to be sufficient benefits to deployment, and so on), and whose deception goes undetected. Itâs not clear to me how common to expect this to be, especially given that weâl likely be on the lookout for it. More generally, I expect decision-makers to face various incentives (economic/social backlash, regulation, liability, the threat of personal harm, and so forth) that reduce the attraction of deploying systems whose practical PS-alignment remains significantly uncertain. And absent active/successful deception, I expect default forms of testing to reveal many PS-alignment problems ahead of time. That said, even absent active/successful deception, there are a variety of other ways that systems that would be practically PS-misaligned if deployed can end up superficially attractive to deploy anyway: for example, because they demonstrate extremely useful/profitable capabilities, and decision-makers are wrong about how well they can predict/control/incentivize the systems in question; and/or because externalities, dysfunctional competitive dynamics, and variations in caution/social responsibility lead to problematic degrees of willingness to knowingly deploy possibly or actually practically PS-misaligned systems (especially with increasingly powerful capabilities, and in increasingly uncontrolled and/or rapidly changing circumstances). 4.Some deployed APS systems will be exposed to inputs where they seek power in misaligned and high-impact ways (say, collectively causing >$1 trillion 2021-dollars of damage) | (1)â(3). 180 Iâm going to say: 65%. In particular, I think that once we condition on 2 and 3, the probability of high-impact post-deployment practical alignment failures goes up a lot, since it means weâre likely building systems that would be practically PS-misaligned if deployed, but which are temptingâto some at least, especially in light of the incentives at stake in 2âto deploy regardless. The 35% on this premise being false comes centrally from the fact that (a) I expect us to have seen a good number of warning shots before we reach really high-impact practical PS-alignment failures, so this premise requires that we havenât responded to those adequately, (b) the time-horizons and capabilities of the relevant practically PS-misaligned systems might be limited in various ways, thereby reducing potential damage, and (c) practical PS-alignment failures on the scale of trillions of dollars (in combination) are major mistakes, which relevant actors will have strong incentives, other things equal, to avoid/prevent (from market pressure, regulation, self-interested and altruistic concern, and so forth). However, there are a lot of relevant actors in the world, with widely varying degrees of caution and social responsibility, and I currently feel pessimistic about prospects for international coordination (cf. climate change) or adequately internalizing externalities (especially since the biggest costs of PS-misalignment failures are to the long-term future). Conditional on 1-3 above, I expect the less responsible actors to start using APS systems even at the risk of PS-misalignment failure; and I expect there to be pressure on others to do the same, or get left behind. 5.Some of this misaligned power-seeking will scale (in aggregate) to the point of permanently disempowering ~all of humanity | (1)â(4). Iâm going to say: 40%. Thereâs a very big difference between >$1 trillion dollars of damage (~6 Hurricane Katrinas), and the complete disempowerment of humanity; and especially in slower take- off scenarios, I donât think it at all a foregone conclusion that misaligned power-seeking that causes 180 For reference, Hurricane Katrina, according to Wikipedia, cost ~160 billion (2021 dollars); Chernobyl, ~$770 billion (2021 dollars); and Cutler and Summers (2020) put the cost of Covid-19 in the U.S. at ~$16 trillion. More disaster costs here. 50 the former will scale to the latter. But I also think that conditional on reaching a scenario with this level of damage from high-impact practical PS-alignment failures (as well as the other previous premises), things are looking dire. Itâs possible that the world gets its act together at that point, but it seems far from certain. 6.This will constitute an existential catastrophe | (1)â(5). Iâm going to say: 95%. I havenât thought about this one very much, but my current view is that the permanent and unintentional disempowerment of humans is very likely to be catastrophic for the potential value of human civilizationâs future. Multiplying these conditional probabilities together, then, we get:65%¡80%¡40%¡65%¡40%¡95% = ~5% probability of existential catastrophe from misaligned, power-seeking AI by 2070. 181 And Iâd probably bump this up a bitâmaybe by a percentage point or two, though this is especially unprincipled (and small differences are in the noise anyway)âto account for power-seeking scenarios that donât strictly fit all the premises above. 182 Note that these (subjective, unstable) numbers arefor 2070 in particular. Later dates, or conditioning on the development of sufficiently advanced systems, would bump them up. My main point here, though, isnât the specific numbers. Rather, itâs that as far as I can presently tell, there is a disturbingly substantive risk that we (or our children) live to see humanity as a whole permanently and involuntarily disempowered by AI systems weâve lost control over. What we can and should do about this now is a further question. But the issue seems serious. Acknowledgments Thanks to Asya Bergal, Alexander Berger, Paul Christiano, Ajeya Cotra, Tom Davidson, Daniel Dewey, Owain Evans, Ben Garfinkel, Katja Grace, Jacob Hilton, Evan Hubinger, Jared Kaplan, Holden Karnofsky, Sam McCandlish, Luke Muehlhauser, Richard Ngo, David Roodman, Rohin Shah, Carl Shulman, Nate Soares, Jacob Steinhardt, and Eliezer Yudkowsky for input on earlier stages of this project; thanks to Sara Fish for formatting and bibliography help; and thanks to Nick Beckstead for guidance and support throughout the investigation. The views expressed here are my own. 9 Appendix Here are a few reformulations of the argument above, with probabilities (but without commentary). I donât think this does all that much to ward off possible biases/framing effects (especially since Iâve gotten used to thinking in terms of the six premises above), but perhaps itâs a start. 183 Shorter negative: By 2070: 1. It will become possible and financially feasible to build APS AI systems. 65% 2. It will much more difficult to build APS AI systems that would be practically PS-aligned if deployed than to build APS systems that would be practically PS-misaligned if deployed, but which are at least superficially attractive to deploy anyway | 1. 181 In sensitivity tests, where I try to put in âlow-endâ and âhigh-endâ estimates for the premises above, this number varies between ~.1% and ~40% (sampling from distributions over probabilities narrows this range a bit, but it also fails to capture certain sorts of correlations). And my central estimate varies between ~1-10% depending on my mood, what considerations are salient to me at the time, and so forth. This instability is yet another reason not to put too much weight on these numbers. And one might think variation in the direction of higher risk especially worrying. 182 Iâm not including scenarios thatdonâtcenter on misaligned power-seeking: for example, ones where AI systems empower human actors in the wrong ways, or in which forms of misalignment that donât involve power-seeking lead to existential catastrophe. 183 I found the exercise of cross-checking at least somewhat helpful. 51 35% 184 3.Deployed, practically PS-misaligned systems will disempower humans at a scale that constitutes existential catastrophe | 1-2. 20% Implied probability of existential catastrophe from scenarios where all three premises are true: ~5% Shorter positive: Before 2070: 1. It wonât be possible and financially feasible to build APS AI systems. 35% 2. It wonât be much more difficult to build APS AI systems that would be practically PS-aligned if deployed, than to build APS systems that would be practically PS-misaligned if deployed, but which are at least superficially attractive to deploy anyway | not (1). 65% 3.It wonât be the case that deployed practically PS-misaligned systems disempower humans at a scale that constitutes existential catastrophe | not (1 or 2). 80% Implied probability that weâl avoid catastrophe Ă la shorter negative: ~95% Same-length positive: Before 2070: 1. It wonât be both possible and financially feasible to build APS systems. 35% 2. There wonât be strong incentives to build APS systems | not 1. 20% 3. It wonât be much harder to develop APS systems that would be practically PS-aligned if deployed, than to develop APS systems that would be practically PS-misaligned if deployed (even if relevant decision-makers donât know this), but which are at least superficially attractive to deploy anyway | not (1 or 2). 60% 4.APS systems wonât be exposed to inputs where they seek power in misaligned and high- impact ways (say, collectively causing >$1 trillion 2021-dollars of damage) | not (1 or 2 or 3). 35% 5. This power-seeking wonât scale (in aggregate) to the point of permanently disempowering ~all humans | not (1 or 2 or 3 or 4). 60% 6.Such disempowerment wonât constitute an existential catastrophe | not (1 or 2 or 3 or 4 or 5). https://w.overleaf.com/project/626a43173c7082e1d1e2036c5% Implied probability that weâl avoid scenarios like the one discussed in the report: ~95% We can also imagineexpandingthe argument into many further premises, so as to bring out/highlight conjunctiveness that might be hiding within it. I think doing this might well be valuable; but I wonât attempt it here. 184 This is somewhat lower than in my six premises above, because itâs not conditioning on strong incentives to build APS systems; and I expect that absent such incentives, itâs harder to hit the bar for âsuperficially attractive to deploy.â 52 References [ABH19] Amanda Askell, Miles Brundage, and Gillian Hadfield. âThe Role of Cooperation in Responsible AI Developmentâ. en. In:arXiv:1907.04534 [cs](July 2019). arXiv: 1907.04534.URL: http://arxiv.org/abs/1907.04534 (visited on 04/29/2022). [Ada]Tom Adamczewski.A shift in arguments for AI risk.URL: https://fragile- credences. github.io/prioritising-ai/ (visited on 04/29/2022). [Amo+16] Dario Amodei et al. âConcrete Problems in AI Safetyâ. In:arXiv:1606.06565 [cs.AI] (June 2016).URL: https://arxiv.org/abs/1606.06565. [Bak+20]Bowen Baker et al. âEmergent Tool Use From Multi-Agent Autocurriculaâ. en. In: Eighth International Conference on Learning Representations. Apr. 2020.URL: https: //iclr.c/virtual_2020/poster_SkxpxJBKwS.html (visited on 04/29/2022). [Ber+19]Christopher Berner et al. âDota 2 with Large Scale Deep Reinforcement Learningâ. In: arXiv:1912.06680 [cs.LG](2019).URL: https://arxiv.org/abs/1912.06680. [Bos]Nick Bostrom.What happens when our computers get smarter than we are?URL: https://w.ted.com/talks/nick_bostrom_what_happens_when_our_computers_get_ smarter_than_we_are/transcript (visited on 04/29/2022). [Bos14] Nick Bostrom.Superintelligence: Paths, Dangers, Strategies. en. Google-Books-ID: 7_H8AwAAQBAJ. Oxford University Press, 2014.ISBN: 978-0-19-967811-2. [Bra16]Gwern Branwen.Why Tool AIs Want to Be Agent AIs. en-us. Sept. 2016.URL: https: //w.gwern.net/Tool-AI (visited on 04/29/2022). [Bur+18]Yuri Burda et al. âExploration by random network distillationâ. en. In:ICLR 2019. Sept. 2018.URL: https://openreview.net/forum?id=H1lJJnR5Ym (visited on 04/29/2022). [CA16]Jack Clark and Dario Amodei.Faulty Reward Functions in the Wild. en. Dec. 2016. URL: https://openai.com/blog/faulty-reward-functions/ (visited on 04/29/2022). [Car20]Joseph Carlsmith.How Much Computational Power Does It Take to Match the Human Brain?en. Tech. rep. Sept. 2020.URL: https : / / w. openphilanthropy. org / brain - computation-report (visited on 04/29/2022). [Ceg]Maciej CegĹowski.Superintelligence: The Idea That Eats Smart People.URL: https: //idlewords.com/talks/superintelligence.htm (visited on 04/29/2022). [Cel14]Rory Cellan-Jones. âStephen Hawking warns artificial intelligence could end mankindâ. en-GB. In:BBC News(Dec. 2014).URL: https : / / w. bbc . com / news / technology - 30290540 (visited on 04/29/2022). [Chr]Paul Christiano.What failure looks like - AI Alignment Forum.URL: https : / / w. alignmentforum.org/posts/HBxe6wdjxK239zajf/what- failure- looks- like (visited on 04/29/2022). [Chr+17]Paul F Christiano et al. âDeep Reinforcement Learning from Human Prefer- encesâ. In:Advances in Neural Information Processing Systems. Vol. 30. Cur- ran Associates, Inc., 2017.URL: https : / / papers . nips . c / paper / 2017 / hash / d5e2c0adad503c91f91df240d0cd4e49-Abstract.html (visited on 04/29/2022). [Chr19a] Paul Christiano.Paul Christiano: Current Work in AI Alignment. EA Global: San Francisco, 2019.URL: https : / / w. effectivealtruism . org / articles / paul - christiano - current-work-in-ai-alignment (visited on 04/29/2022). [Chr19b] Paul Christiano.Training robust corrigibility. en. Mar. 2019.URL: https://ai-alignment. com/training-robust-corrigibility-ce0e0a3b9b4d (visited on 04/29/2022). [Chr20]Brian Christian.The Alignment Problem: Machine Learning and Human Values. en. Google-Books-ID: VmJIzQEACAAJ. W.W. Norton, 2020.ISBN: 978-0-393-63582-9. [Chr21] Paul Christiano.Clarifying âAI alignmentâ. en. Apr. 2021.URL: https://ai-alignment. com/clarifying-ai-alignment-cec47cd69d6 (visited on 04/29/2022). [CK20]Andrew Critch and David Krueger. âAI Research Considerations for Human Existential Safety (ARCHES)â. en. In:arXiv:2006.04948 [cs.CY](May 2020).URL: https://arxiv. org/abs/2006.04948v1 (visited on 04/29/2022). [Cot20]Ajeya Cotra.Draft report on AI timelines - AI Alignment Forum. Tech. rep. Open Philan- thropy, Sept. 2020.URL: https://w.alignmentforum.org/posts/KrJfoZzpSDpnrv9va/ draft-report-on-ai-timelines (visited on 04/29/2022). 53 [Cot21]Ajeya Cotra.The case for aligning narrowly superhuman models - LessWrong. Mar. 2021.URL: https://w.lesswrong.com/posts/PZtsoaoSLpKjjbMqM/the- case- for- aligning-narrowly-superhuman-models (visited on 04/29/2022). [CS20] David M. Cutler and Lawrence H. Summers. âThe COVID-19 Pandemic and the $16 Trillion Virusâ. In:JAMA324.15 (Oct. 2020), p. 1495â1496.ISSN: 0098-7484.DOI: 10.1001/jama.2020.19759.URL: https://doi.org/10.1001/jama.2020.19759 (visited on 04/29/2022). [CSA18] Paul Christiano, Buck Shlegeris, and Dario Amodei. âSupervising strong learners by am- plifying weak expertsâ. In:arXiv:1810.08575 [cs, stat](Oct. 2018). arXiv: 1810.08575. URL: http://arxiv.org/abs/1810.08575 (visited on 04/29/2022). [CY] Myoung Cha and Flora Yu.Pharmaâs first-to-market advantage | McKinsey.URL: https://w.mckinsey.com/industries/life- sciences/our- insights/pharmas- first- to- market-advantage (visited on 04/29/2022). [Dav21] Tom Davidson.Report on Semi-informative Priors. en. Tech. rep. Open Philanthropy, Mar. 2021.URL: https://w.openphilanthropy.org/blog/report-semi-informative-priors (visited on 04/29/2022). [Dre19]K Eric Drexler.Reframing Superintelligence. en. Tech. rep. University of Oxford: Future of Humanity Institute, Jan. 2019, p. 210. [DZE16]Allan Dafoe, Baobao Zhang, and Owain Evans.2016 Expert Survey on Progress in AI. en-US. Section: AI Timeline Surveys. Dec. 2016.URL: https://aiimpacts.org/2016- expert-survey-on-progress-in-ai/ (visited on 04/29/2022). [FH] Joshua Fox and Oliver Habryka.Gems from the Wiki: Acausal Trade.URL: https : //w.lesswrong.com/posts/YBc4gNAELC3uMjPtQ/gems-from-the-wiki-acausal- trade (visited on 04/29/2022). [Fila]Daniel Filan.AXRP Episode 4 - Risks from Learned Optimization with Evan Hubinger - LessWrong.URL: https://w.lesswrong.com/posts/EszCTbovFfpJd5C8N/axrp- episode-4-risks-from-learned-optimization-with-evan (visited on 04/29/2022). [Filb]Daniel Filan.Bottle Caps Arenât Optimisers - AI Alignment Forum.URL: https://w. alignmentforum.org/posts/26eupx3Byc8swRS7f/bottle-caps-aren-t-optimisers (visited on 04/29/2022). [Flia]Alex Flint.Search versus design - LessWrong.URL: https://w.lesswrong.com/posts/ r3NHPD3dLFNk9QE2Y/search-versus-design-1 (visited on 04/29/2022). [Flib]Alex Flint.The ground of optimization - AI Alignment Forum.URL: https : / / w. alignmentforum.org/posts/znfkdCoHMANwqc2WE/the- ground- of- optimization- 1 (visited on 04/29/2022). [Gab+]Gabriel Goh et al.Multimodal Neurons in Artificial Neural Networks.URL: https : //distill.pub/2021/multimodal-neurons/ (visited on 04/29/2022). [Gar+]Ben Garfinkel et al.Ben Garfinkel on scrutinising classic AI risk arguments. en-US. URL: https://80000hours.org/podcast/episodes/ben-garfinkel-classic-ai-risk-arguments/ (visited on 04/29/2022). [Gar+17] Ben Garfinkel et al. âOn the Impossibility of Supersized Machinesâ. en. In: arXiv:1703.10987 [cs.CY](Mar. 2017).DOI: 10 . 48550 / arXiv . 1703 . 10987.URL: https://arxiv.org/abs/1703.10987v1 (visited on 04/29/2022). [Gar18]Ben Garfinkel.Ben Garfinkel: How sure are we about this AI stuff?en-US. EA Global, 2018.URL: https://ea.greaterwrong.com/posts/9sBAW3qKppnoG3QPq/ben-garfinkel- how-sure-are-we-about-this-ai-stuff (visited on 04/29/2022). [GD19]Ben Garfinkel and Allan Dafoe. âHow does the offense-defense balance scale?â In:Journal of Strategic Studies42.6 (Sept. 2019). Publisher: Routledge _eprint: https://doi.org/10.1080/01402390.2019.1631810, p. 736â763.ISSN: 0140-2390.DOI: 10 . 1080 / 01402390 . 2019 . 1631810.URL: https : / / doi . org / 10 . 1080 / 01402390 . 2019 . 1631810 (visited on 04/29/2022). [Gra+15]Katja Grace et al.Discontinuous progress investigation. en-US. Section: AI Timelines. Feb. 2015.URL: https://aiimpacts.org/discontinuous-progress-investigation/ (visited on 04/29/2022). 54 [Gra+18]Katja Grace et al. âViewpoint: When Will AI Exceed Human Performance? Evidence from AI Expertsâ. en. In:Journal of Artificial Intelligence Research62 (July 2018), p. 729â754.ISSN: 1076-9757.DOI: 10.1613/jair.1.11222.URL: http://jair.org/index. php/jair/article/view/11222 (visited on 04/29/2022). [Gra20] Katja Grace.Misalignment and misuse: whose values are manifest?en-US. Section: Blog. Nov. 2020.URL: https://aiimpacts.org/misalignment-and-misuse-whose-values- are-manifest/ (visited on 04/29/2022). [Hen15]Joseph Henrich.The Secret of Our Success: How Culture Is Driving Human Evolution, Domesticating Our Species, and Making Us Smarter. English. Princeton: Princeton University Press, Oct. 2015.ISBN: 978-0-691-16685-8. [Hub]Evan Hubinger.Clarifying inner alignment terminology - AI Alignment Forum.URL: https : / / w. alignmentforum . org / posts / SzecSPYxqRa5GCaSF / clarifying - inner - alignment-terminology (visited on 04/29/2022). [Hub+19] Evan Hubinger et al. âRisks from Learned Optimization in Advanced Machine Learning Systemsâ. In:arXiv:1906.01820 [cs.AI](2019).URL: https://arxiv.org/abs/1906.01820. [Hub+21]Evan Hubinger et al. âRisks from Learned Optimization in Advanced Machine Learning Systemsâ. In:arXiv:1906.01820 [cs](Dec. 2021). arXiv: 1906.01820.URL: http://arxiv. org/abs/1906.01820 (visited on 04/29/2022). [Hun20] Will Hunt.The Flight to Safety-Critical AI. Tech. rep. UC Berkeley Center for Long- Term Cybersecurity, 2020, p. 42.URL: https://cltc.berkeley.edu/wp-content/uploads/ 2020/08/Flight-to-Safety-Critical-AI.pdf. [HW20]Evan Hubinger and Kate Woolverton.Homogeneity vs. heterogeneity in AI takeoff scenar- ios. Dec. 2020.URL: https://w.alignmentforum.org/posts/mKBfa8v4S9pNKSyKK/ homogeneity-vs-heterogeneity-in-ai-takeoff-scenarios (visited on 04/29/2022). [ICA18]Geoffrey Irving, Paul Christiano, and Dario Amodei. âAI safety via debateâ. In: arXiv:1805.00899 [cs, stat](Oct. 2018). arXiv: 1805.00899.URL: http://arxiv.org/abs/ 1805.00899 (visited on 04/29/2022). [Kar12]Holden Karnofsky.Thoughts on the Singularity Institute (SI) - LessWrong. May 2012. URL: https : // w. lesswrong. com/ posts/ 6SGqkCgHuNr7d4yJm/ thoughts- on- the - singularity-institute-si (visited on 04/29/2022). [Kar16] Holden Karnofsky.Some Background on Our Views Regarding Advanced Artificial Intel- ligence. en. May 2016.URL: https://w.openphilanthropy.org/blog/some-background- our-views-regarding-advanced-artificial-intelligence (visited on 04/29/2022). [Kau16]Jeff Kaufman.Multiple Stage Fallacy?Mar. 2016.URL: https://w.jefftk.com/p/ multiple-stage-fallacy (visited on 04/29/2022). [KML20]David Krueger, Tegan Maharaj, and Jan Leike. âHidden Incentives for Auto-Induced Distributional Shiftâ. In:arXiv:2009.09153 [cs, stat](Sept. 2020). arXiv: 2009.09153. URL: http://arxiv.org/abs/2009.09153 (visited on 04/29/2022). [Kra+20]Victoria Krakovna et al.Specification gaming: the flip side of AI ingenuity. en. Apr. 2020.URL: https://w.deepmind.com/blog/specification-gaming-the-flip-side-of-ai- ingenuity (visited on 04/29/2022). [Kum+20]Ramana Kumar et al. âREALab: An Embedded Perspective on Tamperingâ. In: arXiv:2011.08820 [cs](Nov. 2020). arXiv: 2011.08820.URL: http://arxiv.org/abs/2011. 08820 (visited on 04/29/2022). [LeC]Anthony Zador LeCun Yann.Donât Fear the Terminator. en.URL: https : / / blogs . scientificamerican.com/observations/dont-fear-the-terminator/ (visited on 04/29/2022). [Lei+18]Jan Leike et al. âScalable agent alignment via reward modeling: a research directionâ. In:arXiv:1811.07871 [cs, stat](Nov. 2018). arXiv: 1811.07871.URL: http://arxiv.org/ abs/1811.07871 (visited on 04/29/2022). [Lew+17]Mike Lewis et al. âDeal or No Deal? End-to-End Learning of Negotiation Dialoguesâ. In:Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Copenhagen, Denmark: Association for Computational Linguistics, Sept. 2017, p. 2443â2453.DOI: 10.18653/v1/D17-1259.URL: https://aclanthology.org/D17- 1259 (visited on 04/29/2022). [MG19]David Manheim and Scott Garrabrant. âCategorizing Variants of Goodhartâs Lawâ. In: arXiv:1803.04585 [cs, q-fin, stat](Feb. 2019). arXiv: 1803.04585.URL: http://arxiv.org/ abs/1803.04585 (visited on 04/29/2022). 55 [Mik]Vladimir Mikulik.2-D Robustness - AI Alignment Forum.URL: https : / / w . alignmentforum . org / posts / 2mhFMgtAjFJesaSYR / 2 - d - robustness (visited on 04/29/2022). [Mue] Luke Muehlhauser.Treacherous turns in the wild.URL: https://lukemuehlhauser.com/ treacherous-turns-in-the-wild/#more-6202 (visited on 04/29/2022). [Mue18] Luke Muehlhauser.2017 Report on Consciousness and Moral Patienthood. en. Tech. rep. Jan. 2018.URL: https://w.openphilanthropy.org/2017-report-consciousness-and- moral-patienthood (visited on 04/29/2022). [Ngoa] Richard Ngo.AGI safety from first principles: Introduction - AI Alignment Forum.URL: https://w.alignmentforum.org/posts/8xRSjC76HasLnMGSf/agi-safety-from-first- principles-introduction (visited on 04/29/2022). [Ngob]Richard Ngo.Disentangling arguments for the importance of AI safety. en-US.URL: https : / / w . greaterwrong . com / posts / JbcWQCxKWn3y49bNB / disentangling - arguments-for-the-importance-of-ai-safety (visited on 04/29/2022). [Oes17]Caspar Oesterheld.Multiverse-wide Cooperation via Correlated Decision Making. en- US. Tech. rep. Center on Long-Term Risk, Aug. 2017.URL: https://longtermrisk.org/ multiverse-wide-cooperation-via-correlated-decision-making/ (visited on 04/29/2022). [Ola+]Chris Olah et al.Zoom In: An Introduction to Circuits.URL: https://distill.pub/2020/ circuits/zoom-in/ (visited on 04/29/2022). [Omo08]Stephen M. Omohundro. âThe Basic AI Drivesâ. In:Proceedings of the 2008 conference on Artificial General Intelligence 2008: Proceedings of the First AGI Conference. NLD: IOS Press, June 2008, p. 483â492.ISBN: 978-1-58603-833-5.URL: https : / / selfawaresystems . files . wordpress . com / 2008 / 01 / ai _ drives _ final . pdf (visited on 04/28/2022). [Ord21]Toby Ord.Precipice. English. Hachette Books, Mar. 2021.ISBN: 978-0-316-48492-3. [Pera]Lucas Perry.Andrew Critch on AI Research Considerations for Human Existential Safety. en-US.URL: https://futureoflife.org/2020/09/15/andrew-critch-on-ai-research- considerations-for-human-existential-safety/ (visited on 04/29/2022). [Perb]Lucas Perry.Evan Hubinger on Inner Alignment, Outer Alignment, and Proposals for Building Safe Advanced AI. en-US.URL: https://futureoflife.org/2020/07/01/evan- hubinger- on - inner- alignment - outer- alignment - and - proposals - for- building - safe - advanced-ai/ (visited on 04/29/2022). [Pin18]Steven Pinker.Enlightenment Now: The Case for Reason, Science, Humanism, and Progress. English. Illustrated edition. New York, New York: Viking, Feb. 2018.ISBN: 978-0-525-42757-5. [Pip18]Kelsey Piper. âThe case for taking AI seriously as a threat to humanityâ. en. In:Vox (Dec. 2018).URL: https://w.vox.com/future- perfect/2018/12/21/18126576/ai- artificial-intelligence-machine-learning-safety-alignment (visited on 04/29/2022). [Roo20]David Roodman.Modeling the Human Trajectory. en. Tech. rep. Open Philanthropy, June 2020.URL: https://w.openphilanthropy.org/blog/modeling-human-trajectory (visited on 04/29/2022). [Rus19]Stuart Russell.Human Compatible: Artificial Intelligence and the Problem of Control. English. Penguin Books, Oct. 2019. [Sch+20]Julian Schrittwieser et al. âMastering Atari, Go, chess and shogi by planning with a learned modelâ. en. In:Nature588.7839 (Dec. 2020). Number: 7839 Publisher: Nature Publishing Group, p. 604â609.ISSN: 1476-4687.DOI: 10.1038/s41586-020-03051-4. URL: https://w.nature.com/articles/s41586-020-03051-4 (visited on 04/29/2022). [Sel18] Daniel Selsam.The general intelligence hypothesis. en-US. July 2018.URL: https : //dselsam.github.io/posts/2018-07-08-the-general-intelligence-hypothesis.html (visited on 04/29/2022). [Shaa]Rohin Shah.Coherence arguments do not entail goal-directed behavior - AI Alignment Forum.URL: https://w.alignmentforum.org/posts/NxF5G6CJiof6cemTw/coherence- arguments-do-not-entail-goal-directed-behavior (visited on 04/29/2022). [Shab]Rohin Shah.Coherence arguments do not entail goal-directed behavior - LessWrong. URL: https://w.lesswrong.com/posts/NxF5G6CJiof6cemTw/coherence-arguments- do-not-entail-goal-directed-behavior (visited on 04/29/2022). 56 [Sil+16]David Silver et al. âMastering the game of Go with deep neural networks and tree searchâ. en. In:Nature529.7587 (Jan. 2016). Number: 7587 Publisher: Nature Publishing Group, p. 484â489.ISSN: 1476-4687.DOI: 10.1038/nature16961.URL: https://w.nature. com/articles/nature16961 (visited on 04/29/2022). [Sil+18] David Silver et al. âA general reinforcement learning algorithm that masters chess, shogi, and Go through self-playâ. In:Science362.6419 (Dec. 2018). Publisher: American Association for the Advancement of Science, p. 1140â1144.DOI: 10.1126/science. aar6404.URL: https://w.science.org/doi/abs/10.1126/science.aar6404 (visited on 04/29/2022). [Teg17] Max Tegmark.Life 3.0: Being Human in the Age of Artificial Intelligence. en. Google- Books-ID: 3_otDwAAQBAJ. Penguin Books Limited, Aug. 2017.ISBN: 978-0-14- 198179-6. [Tur+21] Alexander Matt Turner et al. âOptimal Policies Tend To Seek Powerâ. en. In:NeurIPS 2021. May 2021.URL: https://openreview.net/forum?id=l7-DBWawSZH (visited on 04/29/2022). [Ues+20] Jonathan Uesato et al. âAvoiding Tampering Incentives in Deep RL via Decoupled Approvalâ. In:arXiv:2011.08827 [cs](Nov. 2020). arXiv: 2011.08827.URL: http : //arxiv.org/abs/2011.08827 (visited on 04/29/2022). [Unka] Unknown.Corporations vs. superintelligences. en.URL: https://arbital.com/p/corps_vs_ si/ (visited on 04/29/2022). [Unkb] Unknown.Optimization daemons. en.URL: https://arbital.com/p/daemons/ (visited on 04/29/2022). [Unkc]Unknown.Superintelligent. en.URL: https://arbital.com/p/superintelligent/ (visited on 04/29/2022). [Vic+19]Victoria Krakovna et al.Designing agent incentives to avoid side effects. en. Oct. 2019. URL: https://deepmindsafetyresearch.medium.com/designing- agent- incentives- to- avoid-side-effects-e1ac80ea6107 (visited on 04/29/2022). [Vin+19]Oriol Vinyals et al. âGrandmaster level in StarCraft I using multi-agent reinforcement learningâ. en. In:Nature575.7782 (Nov. 2019). Number: 7782 Publisher: Nature Pub- lishing Group, p. 350â354.ISSN: 1476-4687.DOI: 10.1038/s41586-019-1724-z.URL: https://w.nature.com/articles/s41586-019-1724-z (visited on 04/29/2022). [WH]Robert Wiblin and Keiran Harris.Brian Christian on the alignment problem. en-US. URL: https://80000hours.org/podcast/episodes/brian-christian-the-alignment-problem/ (visited on 04/29/2022). [Wij]Hjalmar Wijk.Tabooing âAgentâ for Prosaic Alignment - LessWrong.URL: https://w. lesswrong.com/posts/zCcmJzbenAXu6qugS/tabooing- agent- for- prosaic- alignment (visited on 04/29/2022). [Yuda]Eliezer Yudkowsky.AI safety mindset. en.URL: https://arbital.com/p/AI_safety_mindset/ (visited on 04/29/2022). [Yudb]Eliezer Yudkowsky.Big-picture strategic awareness. en.URL: https://arbital.com/p/big_ picture_awareness/ (visited on 04/29/2022). [Yudc]Eliezer Yudkowsky.Omnipotence test for AI safety. en.URL: https://arbital.com/p/ omni_test/ (visited on 04/29/2022). [Yudd]Eliezer Yudkowsky.Strong cognitive uncontainability. en.URL: https://arbital.com/p/ strong_uncontainability/ (visited on 04/29/2022). [Yude] Unknown (possibly Yudkowsky).Consequentialist cognition. en.URL: https://arbital. com/p/consequentialist/ (visited on 04/29/2022). [Yudf]Unknown (possibly Yudkowsky).General intelligence. en.URL: https://arbital.com/p/ general_intelligence/ (visited on 04/29/2022). [Yud08] Eliezer Yudkowsky. âArtificial Intelligence as a positive and negative factor in global riskâ. In:Global Catastrophic Risks. Ed. by Nick Bostrom and Milan M. Ě Cirkovi Ě c. Machine Intelligence Research Institute, July 2008, p. 308â345.DOI: 10.1093/oso/ 9780198570509.003.0021.URL: https://intelligence.org/files/AIPosNegFactor.pdf. 57