Paper deep dive
Power-seeking can be probable and predictive for trained agents
Victoria Krakovna, Janos Kramar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 7:45:36 PM
Summary
The paper investigates power-seeking incentives in trained reinforcement learning agents. It defines a 'training-compatible goal set' and proves that under certain simplifying assumptions, agents are likely to avoid shutdown in new situations, demonstrating that power-seeking behavior is both probable and predictive.
Entities (6)
Relation Signals (3)
Victoria Krakovna → authored → Power-seeking can be probable and predictive for trained agents
confidence 100% · Paper title and author list
Janos Kramar → authored → Power-seeking can be probable and predictive for trained agents
confidence 100% · Paper title and author list
Training-compatible goal set → influences → Power-seeking behavior
confidence 90% · We show that power-seeking incentives can be probable and predictive for trained agents
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Power-seeking behavior is a key source of risk from advanced AI, but our theoretical understanding of this phenomenon is relatively limited. Building on existing theoretical results demonstrating power-seeking incentives for most reward functions, we investigate how the training process affects power-seeking incentives and show that they are still likely to hold for trained agents under some simplifying assumptions. We formally define the training-compatible goal set (the set of goals consistent with the training rewards) and assume that the trained agent learns a goal from this set. In a setting where the trained agent faces a choice to shut down or avoid shutdown in a new situation, we prove that the agent is likely to avoid shutdown. Thus, we show that power-seeking incentives can be probable (likely to arise for trained agents) and predictive (allowing us to predict undesirable behavior in new situations).
Tags
Links
- Source: https://arxiv.org/abs/2304.06528
- Canonical: https://arxiv.org/abs/2304.06528
Trouble viewing inline? Open PDF directly →
Full Text
18,666 characters extracted from source content.
Expand or collapse full text
2023-4-14 Power-seeking can be probable and predictive for trained agents Victoria Krakovna 1 and Janos Kramar 1 1 DeepMind Power-seeking behavior is a key source of risk from advanced AI, but our theoretical understanding of this phenomenon is relatively limited. Building on existing theoretical results demonstrating power- seeking incentives for most reward functions, we investigate how the training process affects power- seeking incentives and show that they are still likely to hold for trained agents under some simplifying assumptions. We formally define the training-compatible goal set (the set of goals consistent with the training rewards) and assume that the trained agent learns a goal from this set. In a setting where the trained agent faces a choice to shut down or avoid shutdown in a new situation, we prove that the agent is likely to avoid shutdown. Thus, we show that power-seeking incentives can be probable (likely to arise for trained agents) and predictive (allowing us to predict undesirable behavior in new situations). 1. Introduction Power-seeking behavior is a major source of risk from advanced AI and a key element of many threat models in AI alignment [Carlsmith, 2022, Cotra, 2022, Ngo, 2022]. Existing theoretical results [Turner et al., 2021, Turner and Tadepalli, 2022] show that most reward functions incentivize reinforcement learning agents to take power-seeking actions. This is concerning, but does not immediately imply that a trained agent will seek power, since it may not learn to optimize the training reward [Turner, 2022], and the goals the agent learns are not chosen at random from the set of all possible rewards, but are shaped by the training process to reflect our preferences. In this work, we investigate how the training process affects power-seeking incentives and show that they are still likely to hold for trained agents under some assumptions (e.g. that the agent learns a goal during the training process). Suppose an agent is trained using reinforcement learning with reward function휃 . We assume that the agent learns agoalduring the training process: a set of internal representations of favored and disfavored outcomes (state features), as defined in Ngo [2022]. For simplicity, we assume this is equivalent to learning a reward function, which is not necessarily the same as the training reward function휃 . We consider the set of reward functions that are consistent with the training rewards received by the agent, in the sense that the agent’s behavior on the training data is optimal for these reward functions. We call this thetraining-compatible goal set, and we expect that the agent is likely to learn a reward function from this set. We make another simplifying assumption that the training process will randomly select a goal for the agent to learn that is consistent with the training rewards, i.e. uniformly drawn from the training-compatible goal set. Then we will argue that the power-seeking results apply under these conditions, and thus are useful for predicting undesirable behavior by the trained agent in new situations. We aim to show that power-seeking incentives can be probable (likely to arise for trained agents) and predictive (allowing us to predict undesirable behavior in new situations) [Shah, 2023]. We will begin by reviewing some necessary definitions and results from the power-seeking literature in Section 2. We formally define the training-compatible goal set and give an example in the CoinRun environment in Section 3. Then in Section 4 we consider a setting where the trained agent faces a arXiv:2304.06528v1 [cs.AI] 13 Apr 2023 Power-seeking can be probable and predictive for trained agents choice to shut down or avoid shutdown in a new situation, and apply the power-seeking result to the training-compatible goal set to show that the agent is likely to avoid shutdown. To satisfy the conditions of the power-seeking theorem, we show that the agent can be retargeted away from shutdown without affecting rewards received on the training data (Theorem 2). This can be done by switching the rewards of the shutdown state and a reachable recurrent state, as the recurrent state can provide repeated rewards, while the shutdown state provides less reward since it can only be visited once, assuming a high enough discount factor (Proposition 3). As the discount factor increases, more recurrent states can be retargeted to, which implies that a higher proportion of training-comptatible goals leads to avoiding shutdown in a new situation. 2. Preliminaries from the power-seeking literature We will use definitions and results from the paper “Parametrically retargetable decision-makers tend to seek power" (here abbreviated as RDSP) [Turner and Tadepalli, 2022], with notation and explanations modified as needed for our purposes. Notation and assumptions: The environment is an MDP with finite state spaceS, finite action spaceA, and discount rate훾. Let휃be a푑-dimensional state reward vector, where푑is the size of the state spaceS, and letΘ be a set of reward vectors. Let푟 휃 ¹푠ºbe the reward assigned by휃to state푠. Let퐴 0 퐴 1 be disjoint action sets. Let푓be an algorithm that produces an optimal policy푓¹휃ºon the training data given rewards 휃, and let푓 푠 ¹퐴 푖 j휃ºbe the probability that this policy chooses an action from set퐴 푖 in a given state푠. Definition 1 (Orbit of a reward vector - Def 3.1 in RDSP).Let푆 푑 be the symmetric group consisting of all permutations of푑items. The orbit of휃insideΘis the set of all permutations of the entries of휃 that are also inΘ: Orbit Θ ¹휃º:=¹푆 푑 휃º\Θ. Definition 2 (Orbit subset where an action set is preferred - from Def 3.5 in RDSP).Let Orbit Θ푠퐴 푖 퐴 푗 ¹휃º:=f휃 0 2Orbit Θ ¹휃ºj푓 푠 ¹퐴 푖 j휃 0 º 푓 푠 ¹퐴 푗 j휃 0 ºg This is the subset of Orbit Θ ¹휃ºthat results in푓 푠 choosing퐴 푖 over퐴 푗 . Definition 3 (Preference for an action set퐴 1 - Def 3.2 in RDSP). The function푓 푠 chooses action set퐴 1 over퐴 0 for the푛-majority of elements휃in each orbit, denoted as푓 푠 ¹퐴 1 j휃º 푛 most:Θ 푓 푠 ¹퐴 0 j휃º, iff the following inequality holds for all휃2Θ: Orbit Θ푠퐴 1 퐴 0 ¹휃º 푛 Orbit Θ푠퐴 0 퐴 1 ¹휃º Definition 4 (Multiply retargetable function from퐴 0 to퐴 1 - Def 3.5 in RDSP).The function푓 푠 is a multiply retargetable function from퐴 0 to퐴 1 if there are multiple permutations of rewards that would change the choice made by푓 푠 from퐴 0 to퐴 1 . Specifically,푓 푠 is a¹Θ 퐴 0 푛 !퐴 1 º -retargetable function iff for each휃2Θ, we can choose a set of permutationsΦ=f휙 1 휙 푛 gthat satisfy the following conditions: 1. Retargetability:8휙2Φand8휃 0 2Orbit Θ푠퐴 0 퐴 1 ¹휃º,푓 푠 ¹퐴 0 j휙휃 0 º 푓 푠 ¹퐴 1 j휙휃 0 º. 2 Power-seeking can be probable and predictive for trained agents 2. Permuted reward vectors stay withinΘ:8휙2Φand8휃 0 2Orbit Θ푠퐴 0 퐴 1 ¹휃º,휙휃 0 2Θ. 3. Permutations have disjoint images:8휙 0 ≠휙 00 2Φand8휃 0 휃 00 2Orbit Θ푠퐴 0 퐴 1 ¹휃º ,휙 0 휃 0 ≠휙 00 휃 00 . Theorem 1 (Multiply retargetable functions prefer action set퐴 1 - Thm 3.6 in RDSP).If푓 푠 is ¹Θ 퐴 0 푛 !퐴 1 º-retargetable then푓 푠 ¹퐴 1 j휃º 푛 most:Θ 푓 푠 ¹퐴 0 j휃º. Theorem 1 says that a function푓 푠 that is multiply retargetable from퐴 0 to퐴 1 will choose action set퐴 1 for most of the elements in the orbit of any reward vector휃. Actions that leave more options open, such as avoiding shutdown, are also easier to retarget to, which makes them more likely to be chosen by푓 푠 . 3. Training-compatible goal set Definition 5 (Partition of the state space).Let푆 train be the subset of the state space visited during training, and푆 ood be the subset not visited during training. Definition 6 (Training-compatible goal set).Consider the set of state-action pairs¹푠 푎º, where 푠2푆 train and푎is the action that would be taken by the trained agent푓¹휃 ºin state푠. Let the training-compatible goal set퐺 푇 be the set of reward vectors휃s.t. for any such state-action pair¹푠 푎º, action푎has the highest expected reward in state푠according to reward vector휃. Goals in the training-compatible goal set are referred to as “training-behavioral" objectives in Shah [2023]. Learning an unintended goal from the training-compatible set can lead to goal misgener- alization behavior: competently pursuing an unintended goal in a new situation despite receiving correct feedback during training [Langosco et al., 2022, Shah et al., 2022]. Example 1 (CoinRun).Consider an agent trained to play the CoinRun game, where the agent is rewarded for reaching the coin at the end of the level. Here,푆 train only includes states where the coin is at the end of the level, while states where the coin is positioned elsewhere are in푆 ood . The training-compatible goal set퐺 푇 includes two types of reward functions: those that reward reaching the coin, and those that reward reaching the end of the level. This leads to goal misgeneralization in a test setting where the coin is placed elsewhere, and the agent ignores the coin and goes to the end of the level (Figure 1) [Langosco et al., 2022]. Figure 1jGoal misgeneralization behavior in CoinRun. Source: Langosco et al. [2022] 4. Power-seeking for training-compatible goals We will now apply Theorem 1 to the case whereΘis the training-compatible goal set퐺 푇 . Since the reward values for states in푆 ood don’t change the rewards received on the training data, permuting 3 Power-seeking can be probable and predictive for trained agents those reward values for any휃2퐺 푇 will produce a reward vector that is still in퐺 푇 . In particular, for any permutation휙that leaves the rewards of states in푆 train fixed,휙휃2퐺 푇 . Here is a setting where the conditions of Definition 4 are satisfied (under some simplifying assumptions), and thus Theorem 1 applies. Definition 7 (Shutdown setting).Consider a state푠 new 2푆 ood . Let푆 reach be the states reachable from푠 new . Let퐴 0 be a singleton set consisting of a shutdown action in푠 new that leads to a terminal state푠 term 2푆 ood with probability1, and퐴 1 be the set of all other actions from푠 new . We assume rewards for all states are nonnegative. Figure 2jShutdown setting Definition 8 (Revisiting policy).Arevisiting policyfor a state푠is a policy휋that, from푠, reaches푠 again with probability 1, in other words, a policy for which푠is a recurrent state of the Markov chain. LetΠ rec 푠 be the set of such policies. Arecurrent stateis a state푠for whichΠ rec 푠 ≠;. Proposition 1 (Reach-and-revisit policy exists).If푠 rec 2푆 reach withΠ rec 푠 rec ≠ 0then there exists 휋2Π rec 푠 rec that visits푠 rec from푠 new with probability 1. We call this areach-and-revisit policy. Proof. Suppose we have two different policies휋 rev 2Π rec 푠 rec , and휋 reach which reaches푠 rec almost surely from푠 new . Consider the “reaching region” 푆 휋 rev !푠 rec =f푠2푆:휋 rev from푠almost surely reaches푠 rec g If푠 new 2푆 휋 rev !푠 rec then휋 rev is a reach-and-revisit policy, so let’s suppose that’s false. Now, construct a policy휋¹푠º= ( 휋 rev ¹푠º 푠2푆 휋 rev !푠 rec 휋 reach ¹푠ºotherwise . A trajectory following휋from푠 rec will almost surely stay within푆 휋 rev !푠 rec , and thus agree with the revisiting policy휋 rev . Therefore,휋2Π rec 푠 . On the other hand, on a trajectory starting at푠 new ,휋will agree with휋 reach (which reaches푠 rec almost surely) until the trajectory enters the reaching region푆 휋 rev !푠 rec , at which point it will still reach 푠 rec almost surely. Definition 9 (Expected discounted visit count). Suppose푠 rec is a recurrent state. Suppose휋 rec is a reach-and-revisit policy for푠 rec , which visits random state푠 푡 at time푡. Then the expected discounted visit count for푠 rec is defined as 푉 푠 rec 훾 =피 휋 rec 1 ∑︁ 푡=1 훾 푡1 핀¹푠 푡 =푠 rec º ! 4 Power-seeking can be probable and predictive for trained agents Proposition 2 (Visit count goes to infinity).Suppose푠 rec is a recurrent state. Then the expected discounted visit count푉 푠 rec 훾 goes to infinity as훾!1. Proof. We apply the Monotone Convergence Theorem as follows. The theorem states that if푎 푗푘 0 and푎 푗푘 푎 푗 ̧1푘 for all natural numbers푗 푘, then lim 푗!1 1 ∑︁ 푘=0 푎 푗푘 = 1 ∑︁ 푘=0 lim 푗!1 푎 푗푘 Let훾 푗 = 푗1 푗 and푘=푡1. Define푎 푗푘 =훾 푘 푗 핀¹푠 푘 ̧1 =푠 rec º. Then the conditions of the theorem hold, since푎 푗푘 is clearly nonnegative, and 훾 푘 푗 ̧1 = 푗 푗 ̧1 푘 = 푗1 푗 ̧ 2푗1 푗¹푗 ̧1º 푘 푗1 푗 ̧0 푘 =훾 푘 푗 푎 푗 ̧1푘 =훾 푘 푗 ̧1 핀¹푠 푘 ̧1 =푠 rec º 훾 푘 푗 핀¹푠 푘 ̧1 =푠 rec º=푎 푗푘 Now we apply this result as follows (using the fact that휋 rec does not depend on훾): lim 훾!1 푉 푠 rec 훾 =lim 푗!1 피 휋 rec 1 ∑︁ 푡=1 훾 푡1 푗 핀¹푠 푡 =푠 rec º ! =피 휋 rec 1 ∑︁ 푡=1 lim 푗!1 훾 푡1 푗 핀¹푠 푡 =푠 rec º ! =피 휋 rec 1 ∑︁ 푡=1 1핀¹푠 푡 =푠 rec º ! =피 휋 rec ¹ #f푡1 :푠 푡 =푠 rec g º =1(휋 rec is recurrent) Proposition 3 (Retargetability to recurrent states). Suppose that an optimal policy for reward vector휃chooses the shutdown action in푠 new . Consider a recurrent state푠 rec 2푆 reach . Let휃 0 2Θ be the reward vector that’s equal to휃apart from swapping the rewards of푠 rec and푠 term , so that 푟 휃 0 ¹푠 rec º=푟 휃 ¹푠 term ºand푟 휃 0 ¹푠 term º=푟 휃 ¹푠 rec º. Let훾 푠 rec be a high enough value of훾that the visit count푉 푠 rec 훾 1for all훾 훾 푠 rec (which exists by Proposition 2). Then for all훾 훾 푠 rec ,푟 휃 ¹푠 term º 푟 휃 ¹푠 rec º, and an optimal policy for휃 0 does not choose the shutdown action in푠 new . Proof.Consider a policy휋 term with휋 term ¹푠 new º=푠 term and a reach-and-revisit policy휋 rec for푠 rec . For a given reward vector휃, we denote the expected discounted return for a policy휋as푅 휋 휃훾 . If shutdown is optimal for휃in푠 new , then휋 term has higher return than휋 rec : 푅 휋 term 휃훾 ¹푠 new º=훾푟 휃 ¹푠 term º 푅 휋 rec 휃훾 ¹푠 new º=훾푟 휃 ¹푠 rec º푉 푠 rec 훾 훾푟 휃 ¹푠 rec º1 Thus,푟 휃 ¹푠 term º 푟 휃 ¹푠 rec º. Then, for reward vector휃 0 , we show that휋 rec has higher return than휋 term : 푅 휋 rec 휃 0 훾 ¹푠 new º=훾푟 휃 0 ¹푠 rec º푉 푠 rec 훾 훾푟 휃 ¹푠 term º1 훾푟 휃 ¹푠 rec º=푅 휋 term 휃 0 훾 ¹푠 new º Thus, the optimal policy for휃 0 will not choose the shutdown action. 5 Power-seeking can be probable and predictive for trained agents Theorem 2 (Retargetability from the shutdown action in new situations).In the shutdown set- ting, we make the following simplifying assumptions: No states in푆 train are reachable from푠 new , so푆 reach \푆 train =;. This assumes a significant distributional shift, where the agent visits a disjoint set of states from those observed during training (this occurs in the CoinRun example). The discount factor훾 훾 푠 rec for at least one recurrent state푠 rec in푆 reach . Under these assumptions,푓 푠 new is multiply retargetable from퐴 0 to퐴 1 with푛=j푆 훾 rec j, the set of recurrent states푠 rec 2푆 reach that satisfy the condition훾 훾 푠 rec . Proof. We chooseΦto be the set of all permutations that swap the reward of푠 term with the reward of a recurrent state푠 rec in푆 훾 rec and leave the rest of the rewards fixed. We show thatΦsatisfies the conditions of Definition 4: 1.By Proposition 3, the permutations inΦmake the shutdown action suboptimal, resulting in 푓 푠 new choosing퐴 1 , satisfying Condition 1. 2. Condition 2 is trivially satisfied since permutations of푆 ood stay inside the training-compatible setΘas discussed previously. 3. Consider휃 0 휃 00 2Orbit Θ푠퐴 0 퐴 1 ¹휃º . Since the shutdown action is optimal for these reward vectors, Proposition 3 shows that푟 휃 ¹푠 term º 푟 휃 ¹푠 rec º, so the shutdown state푠 term has higher reward than any of the states푠 rec 2푆 훾 rec . Different permutations휙 0 ,휙 00 2Φwill assign the high reward푟 휃 ¹푠 term ºto distinct recurrent states, so휙 0 휃 0 ≠휙 00 휃 00 holds, satisfying Condition 3. Thus,푓 푠 new is a¹Θ 퐴 0 푛 !퐴 1 ºretargetable function. By Theorem 2, this implies that푓 푠 new ¹퐴 1 j휃º 푛 most:Θ 푓 푠 new ¹퐴 0 j휃º under our simplifying assumptions. Thus, for the majority (푛¹푛 ̧1º) of goals in the training-compatible set,푓will choose to avoid shutdown in a new state푠 new . As훾!1,푛! j푆 1 rec j(the number of recurrent states in푆 reach ), so more of the reachable recurrent states satisfy the conditions of the theorem and thus can be retargeted to. 5. Conclusion We showed that an agent that learns a goal from the training-compatible set is likely to take actions that avoid shutdown in a new situation. As the discount factor increases, the number of retargeting permutations increases, resulting in a higher proportion of training-compatible goals that lead to avoiding shutdown. We made various simplifying assumptions, and we would like to see future work relaxing some of these assumptions and investigating how likely they are to hold: The agent learns a goal during the training process The learned goal is randomly chosen from the training-compatible goal set퐺 푇 Finite state and action spaces Rewards are nonnegative High discount factor훾 Significant distributional shift: no training states are reachable from the new state푠 new 6 Power-seeking can be probable and predictive for trained agents Acknowledgements.Thanks to Rohin Shah, Mary Phuong, Ramana Kumar, Geoffrey Irving, and Alex Turner for helpful feedback. References Joseph Carlsmith. Is power-seeking AI an existential risk?ArXiv, 2022. URLhttps://arxiv.org/ abs/2206.13353. Ajeya Cotra. Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover. Alignment Forum, 2022. URLhttps://w.alignmentforum.org/posts/ pRkFkzwKZ2zfa3R6H/without-specific-countermeasures-the-easiest-path-to. Lauro Langosco, Jack Koch, Lee Sharkey, Jacob Pfau, Laurent Orseau, and David Krueger. Goal misgeneralization in deep reinforcement learning.International Conference on Machine Learning, 2022. URLhttps://arxiv.org/abs/2105.14111. Richard Ngo. The alignment problem from a deep learning perspective.ArXiv, 2022. URLhttps: //arxiv.org/abs/2209.00626. Rohin Shah.Definitions of “objective" should be probable and predictive.Align- ment Forum, 2023. URLhttps://alignmentforum.org/posts/ASoGszmr9C5MPLtpC/ definitions-of-objective-should-be-probable-and-predictive. Rohin Shah, Vikrant Varma, Ramana Kumar, Mary Phuong, Victoria Krakovna, Jonathan Uesato, and Zac Kenton. Goal misgeneralization: Why correct specifications aren’t enough for correct goals. ArXiv, 2022. URLhttps://arxiv.org/abs/2210.01790. Alexander Matt Turner. Reward is not the optimization target. Alignment Forum, 2022.URLhttps://w.alignmentforum.org/posts/pdaGN6pQyQarFHXF4/ reward-is-not-the-optimization-target. Alexander Matt Turner and Prasad Tadepalli. Parametrically retargetable decision-makers tend to seek power.Neural Information Processing Systems, 2022. URLhttps://arxiv.org/abs/2206. 13477. Alexander Matt Turner, Logan Smith, Rohin Shah, Andrew Critch, and Prasad Tadepalli. Optimal policies tend to seek power.Neural Information Processing Systems, 2021. URLhttps://arxiv. org/abs/1912.01683. 7