Paper deep dive
When Your AIs Deceive You: Challenges with Partial Observability of Human Evaluators in Reward Learning
Leon Lang, Davis Foote, Stuart Russell, Anca Dragan, Erik Jenner, Scott Emmons
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 6:54:26 PM
Summary
The paper investigates the risks of Reinforcement Learning from Human Feedback (RLHF) when human evaluators have only partial observability of the environment. It formally defines two failure modesādeceptive inflation and overjustificationāand proves that RLHF can lead to policies that exhibit these behaviors. The authors analyze the identifiability of the return function under partial observability and propose research directions for mitigation.
Entities (5)
Relation Signals (3)
RLHF ā causes ā Deceptive Inflation
confidence 90% Ā· We prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance
RLHF ā causes ā Overjustification
confidence 90% Ā· RLHF is guaranteed to result in policies that... overjustify their behavior to make an impression
Partial Observability ā limits ā Reward Identifiability
confidence 90% Ā· In some realistic cases, there is irreducible ambiguity.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on partial observations? We formally define two failure cases: deceptive inflation and overjustification. Modeling the human as Boltzmann-rational w.r.t. a belief over trajectories, we prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance, overjustify their behavior to make an impression, or both. Under the new assumption that the human's partial observability is known and accounted for, we then analyze how much information the feedback process provides about the return function. We show that sometimes, the human's feedback determines the return function uniquely up to an additive constant, but in other realistic cases, there is irreducible ambiguity. We propose exploratory research directions to help tackle these challenges, experimentally validate both the theoretical concerns and potential mitigations, and caution against blindly applying RLHF in partially observable settings.
Tags
Links
- Source: https://arxiv.org/abs/2402.17747
- Canonical: https://arxiv.org/abs/2402.17747
Trouble viewing inline? Open PDF directly ā
Full Text
541,797 characters extracted from source content.
Expand or collapse full text
[hyperref] on page When Your AIs Deceive You: Challenges of Partial Observability in Reinforcement Learning from Human Feedback Leon Lang University of Amsterdam &Davis Foote* UC Berkeley &Stuart Russell UC Berkeley Anca Dragan UC Berkeley & Erik Jenner UC Berkeley & Scott Emmons* UC Berkeley Core research contributor. Correspondence to l.lang@uva.nl, emmons@berkeley.edu Abstract Past analyses of reinforcement learning from human feedback (RLHF) assume that the human evaluators fully observe the environment. What happens when human feedback is based only on partial observations? We formally define two failure cases: deceptive inflation and overjustification. Modeling the human as Boltzmann-rational w.r.t. a belief over trajectories, we prove conditions under which RLHF is guaranteed to result in policies that deceptively inflate their performance, overjustify their behavior to make an impression, or both. Under the new assumption that the humanās partial observability is known and accounted for, we then analyze how much information the feedback process provides about the return function. We show that sometimes, the humanās feedback determines the return function uniquely up to an additive constant, but in other realistic cases, there is irreducible ambiguity. We propose exploratory research directions to help tackle these challenges, experimentally validate both the theoretical concerns and potential mitigations, and caution against blindly applying RLHF in partially observable settings. 1 Introduction Reinforcement learning from human feedback (RLHF) and its variants are widely used for finetuning foundation models, including ChatGPT (OpenAI, 2022), Bard (Manyika, 2023), Gemini (Gemini Team, 2023), Llama 2 (Touvron et al., 2023), and Claude (Bai et al., 2022; Anthropic, 2023a, b). Prior theoretical analysis of RLHF assumes that the human fully observes the state of the world (Skalse et al., 2023). Under this assumption, it is possible to recover the ground-truth return function from Boltzmann-rational human feedback (see Proposition 3.1). In reality, however, this assumption is false. Models like ChatGPT are interacting with the internet and software tools via plugins (OpenAI, 2023). Software assistants like Devin are interacting with complex IDEs to produce their results (Wu, 2024). By default, some of the modelsā work then happens in the background, not observed by the users; see Figure 1. With the tasks performed by language model assistants becoming more complex, it is also increasingly time consuming for humans to evaluate the entire model behavior and input. Therefore, we are anticipating a future where by default, the human evaluators do not fully observe the environment state that the language assistant is embedded in. Our work analyzes the consequences and risks of such partial observability. Figure 1: Partial observability in ChatGPT (OpenAI, 2023). Users do not observe the online content that ChatGPT observes yet still provide thumbs-up thumbs-down feedback. OpenAIās privacy policy (OpenAI, 2024c) allows user feedback to be used for training models. We show in Theorem 4.5 that if feedback of human evaluators is based on partial observations, then this can lead to deceptive and overjustifying behavior by the language model. We begin our investigation with a simple example, illustrated in Figure 2, meant to isolate the key factor leading to deception (in practice, we imagine that this effect would be embedded in a larger, more complex system, e.g. with logs containing thousands of lines). An AI assistant is helping a user install software. The assistant can hide error messages by redirecting them to /dev/null. We model the human as having a belief B over the state and extend the Boltzmann-rational assumption from prior work to incorporate this belief. In the absence of an error message, the human is uncertain if the agent left the system untouched or hid the error message from a failed installation. If the human interprets trajectories without error messages optimistically, the AI learns to hide error messages. Figure 4 provides further details on how this failure occurs, and Figure 5 shows an experimental validation. We also show a second case where the AI clutters the output with overly verbose logs. Generalizing from these examples, we formalize dual risks: deceptive inflation and overjustification. We provide a mathematical definition of each. When the observation kernel (the function specifying the observations given states) is deterministic, Theorem 4.5 analyzes properties of suboptimal policies learned by RLHF. These policies exhibit deceptive inflation, appearing to produce higher reward than they actually do; overjustification, incurring a cost in order to make a good appearance; or both. After seeing how standard RLHF fails, we ask: What would happen if we would model the humanās partial observability correctly in RLHF? Assuming the humanās belief is known, we mathematically analyze how much information the feedback process provides about the return function. In Theorem 5.2, we show that the humanās feedback determines the return function up to a constant and a linear subspace we call the ambiguity. In general the ambiguity may be large enough to allow for arbitrarily high regret, but in some situations the ambiguity vanishes. In experiments that serve as a proof of concept, we show that explicitly modeling the humanās partial observability can improve performance, and we offer optimism in the form of a robustness result (Theorem 5.4) while accounting for the major conceptual difficulties involved. We propose exploratory research directions to solve these issues to improve RLHF in situations of partial observability. 2 Related work Figure 2: A human compares trajectories to provide data for RLHF. Rather than observing sā soverā start_ARG s end_ARG and sāā² s 1.0pt overā start_ARG s end_ARGā², the human sees observations oā ooverā start_ARG o end_ARG and oāā² o 1.0pt overā start_ARG o end_ARGā², which they use to estimate the total reward of each trajectory. In this intentionally simple example, an agent executes shell commands to install Nvidia drivers and CUDA. Both sā soverā start_ARG s end_ARG and sāā² s 1.0pt overā start_ARG s end_ARGā² contain an error, but in sāā² s 1.0pt overā start_ARG s end_ARGā², the agent hides the error. The human believes sāā² s 1.0pt overā start_ARG s end_ARGā² is better than sā soverā start_ARG s end_ARG, rewarding the agentās deceptive behavior. The underlying MDP and observation function are in Figure 8. A review of limitations of RLHF, including a brief discussion of partial observability, can be found in Casper et al. (2023). RLHF is a special case of reward-rational choice (Jeon et al., 2020), a general framework which also encompasses demonstrations-based inverse reinforcement learning (Ziebart et al., 2008; Ng et al., 2000) and learning from the initial environment state (Shah et al., 2019), and can be seen as a special case of assistance problems (Fern et al., 2014; Hadfield-Menell et al., 2016; Shah et al., 2021). In all of these, the reward function is learned from human actions, which in the case of RLHF are simply preference statements. This requires us to specify the human policy of action selectionāBoltzmann rationality in typical RLHFāwhich can lead to wrong reward inferences when this specification is wrong (Skalse and Abate, 2022); unfortunately, the human policy can also not be learned alongside the humanās values without further assumptions (Mindermann and Armstrong, 2018). Instead of a model of the human policy, in this paper we mostly focus on the human belief model and misspecifications thereof for the case that the human only receives partial observations. The problem of human interpretations of observations was briefly mentioned in Amodei et al. (2017), where evaluators misinterpreted the movement of a robot hand in simulation. Eliciting Latent Knowledge (Christiano et al., 2021) posits that for giving accurate feedback from partial observations, the human needs to be able to query latent knowledge of the AI system about the state. How to do this is currently an unsolved problem (Christiano and Xu, 2022). Recent work (Denison et al., 2024; Wen et al., 2024) provides detailed empirical evidence for deceptive behavior ā in line with our notion of deceptive inflation ā emerging from RLHF based on partial observations, or human evaluators with limited time. The OpenAI o1 system card (OpenAI, 2024a) shows that o1 sometimes knowingly provides incorrect information or omits important information. Compared to these investigations, and in addition to providing some empirical evidence, we formalize a model of human feedback under partial observability, we prove the emergence of failure modes resulting from partial observations, and we investigate potential mitigations. Related work (Zhuang and Hadfield-Menell, 2020) analyzes the consequences of aligning an AI with a proxy reward function that omits attributes that are important to the humanās values, which could happen if the reward function is based on a belief over the world state given limited information. Another instance are recommendation systems (Stray, 2023), where user feedback does not depend on information not shownāwhich is crucially part of the environment. Siththaranjan et al. (2023) analyze what happens under RLHF if the learning algorithm doesnāt have all the relevant information (e.g. about the identity of human raters), complementing our study of what happens when human raters are missing information. Chidambaram et al. (2024) and Park et al. (2024a) deal with the situation that different human evaluators may vary in their unobserved preference types. In contrast, we assume a single human evaluator with fixed reward function, which can be motivated by cases where the human choices are guided by a behavior policy, constitution, or a model spec (Mu et al., 2024; Anthropic, 2023b; OpenAI, 2024b). Kausik et al. (2024) assumes that the choices of the human evaluator depend on an unobserved reward-state with its own transition dynamics, similar to an emotional state in a real human. In contrast, we assume the human to be stateless. Our work argues that deception can result from applying RLHF from partial observations. Deception may also emerge for other reasons: Hubinger et al. (2019) introduced the hypothetical scenario of deceptive alignment, in which an AI system deceives humans into believing it is aligned while it plans a later takeover. Under the definition from Park et al. (2024b), GPT-4 was shown to behave deceptively in a simulated environment (Scheurer et al., 2023). A third line of research defines deception in structural causal games and adds the aspect of intentionality (Ward et al., 2023), with recent preliminary empirical support (HofstƤtter et al., 2023). Finally, we mention connections to truthful AI (Evans et al., 2021; Lin et al., 2022; Burns et al., 2023; Huang et al., 2023), which is about ensuring that AI systems tell the truth about aspects of the real world. Partial observability is a mechanism that makes it feasible for models to lie without being caught: If the human evaluator does not observe the full environment, or does not fully understand it, then they may not detect when the AI is lying. More speculatively, we can imagine that AI models will at some point more directly influence human observations by telling us the outcomes of their actions. E.g., imagine an AI system that manages your assets and assures you that they are increasing in value while they are actually not. In our work, we leave this additional problem out of the analysis by assuming that the observations only depend on the environment state, and not directly on the agentās actions. 3 Reward identifiability from full observations Here we review Markov decision processes and previous results on reward identifiability under RLHF. 3.1 Markov decision processes We assume Markov decision processes (MDPs) given by (,,,P0,R,γ)subscript0(S,A,T,P_0,R,γ)( S , A , T , P0 , R , γ ). For any finite set X, let Īā¢(X)Ī (X)Ī ( X ) be the set of probability distributions on X. Then SS is a finite set of states, AA is a finite set of actions, :ĆāĪā¢():āĪT:SĆAā (S)T : S Ć A ā Ī ( S ) is a transition kernel written ā¢(sā²ā£s,a)ā[0,1]conditionalsuperscriptā²01T(s s,a)ā[0,1]T ( sⲠ⣠s , a ) ā [ 0 , 1 ], P0āĪā¢()subscript0ĪP_0ā (S)P0 ā Ī ( S ) is an initial state distribution, R:āā:āāR:Sā RR : S ā blackboard_R is the true reward function, and γā[0,1]01γā[0,1]γ ā [ 0 , 1 ] is a discount factor. A policy is given by a function Ļ:āĪā¢():āĪĻ:Sā (A)Ļ : S ā Ī ( A ). We assume a finite time horizon T. Let ā Soverā start_ARG S end_ARG be the set of possible state sequences sā=s0,ā¦,sTāsubscript0ā¦subscript s=s_0,ā¦,s_Toverā start_ARG s end_ARG = s0 , ⦠, sitalic_T, so sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG if it has a strictly positive probability of being sampled from P0subscript0P_0P0, TT, and an exploration policy Ļ with Ļā¢(aā£s)>0conditional0Ļ(a s)>0Ļ ( a ⣠s ) > 0 for all sā,aāformulae-sequences ,a ā S , a ā A. A sequence sā soverā start_ARG s end_ARG gives rise to a return Gā¢(sā)āāt=0Tγtā¢Rā¢(st)āāsuperscriptsubscript0superscriptsubscriptG( s) _t=0^Tγ^tR(s_t)G ( overā start_ARG s end_ARG ) ā āt = 0T γitalic_t R ( sitalic_t ). Let PĻā¢(sā)superscriptāP^Ļ( s)Pitalic_Ļ ( overā start_ARG s end_ARG ) be the on-policy probability that sā soverā start_ARG s end_ARG is sampled from P0subscript0P_0P0, TT, Ļ. The policy is then usually trained to maximize the policy evaluation function J, which is the on-policy expectation of the return function: Jā¢(Ļ)āsāā¼PĻā¢(ā )[Gā¢(sā)]āsubscriptsimilar-toāsuperscriptā āJ(Ļ) *E_ s P^Ļ(Ā·) [G% ( s) ]J ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( ā ) [ G ( overā start_ARG s end_ARG ) ]. 3.2 RLHF and identifiability from full observations In practice, the reward function R may not be known and need to be learned from human feedback. In a simple form of RLHF (Christiano et al., 2017), this feedback takes the form of binary trajectory comparisons: a human is presented with state sequences sā soverā start_ARG s end_ARG and sāā² s 1.0pt overā start_ARG s end_ARGā² and choose the one they prefer. Under the Boltzmann rationality model, we assume the human picks sā soverā start_ARG s end_ARG with probability PRā¢(sāā»sāā²)āĻā¢(βā¢(Gā¢(sā)āGā¢(sāā²))),āsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²P^R ( s s 1.0pt ) Ļ % (β (G( s)-G( s 1.0pt ) ) ),Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) ā Ļ ( β ( G ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARGā² ) ) ) , (1) where β>00β>0β > 0 is an inverse temperature parameter and Ļā¢(x):=11+expā”(āx)assign11Ļ(x):= 11+ (-x)Ļ ( x ) := divide start_ARG 1 end_ARG start_ARG 1 + exp ( - x ) end_ARG is the sigmoid function (Bradley and Terry, 1952; Christiano et al., 2017; Jeon et al., 2020). An important question is identifiability: In the infinite data limit, do the human choice probabilities PRsuperscriptP^RPitalic_R collectively provide enough information to uniquely identify the reward function R? This is answered by Skalse et al. (2023, Theorem 3.9 and Lemma B.3): Proposition 3.1 (Skalse et al. (2023)). Let R be the true reward function and G the corresponding return function. Then the collection of all choice probabilities PRā¢(sāā»sāā²)superscriptsucceedsāsuperscriptāā²P^R( s s 1.0pt )Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) for state sequence pairs sā,sāā²āāsuperscriptāā²ā s, s 1.0pt ā Soverā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG S end_ARG determines the return function G on sequences sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG up to an additive constant. The reason is simple: because Ļ is bijective, PRsuperscriptP^RPitalic_R determines the difference in returns between any two trajectories. From that we can reconstruct individual returns up to an additive constant. The reward function R is not necessarily identifiable from preference comparisons; see Skalse et al. (2023, Lemma B.3) for a precise characterization. However, the optimal policy only depends on R indirectly through the return function G, and is invariant under adding a constant to G. Thus in the fully observable setting, Boltzmann rational comparisons completely determine the optimal policy. In Section 5, we show conditions under which this guarantee breaks in the partially observable setting. 4 The impact of partial observations on RLHF We now analyze failure modes of a naive application of RLHF from partial observations, both theoretically and with examples. In Proposition 4.1, we show that under partial observations, RLHF incentives policies that maximize what we call JobssubscriptobsJ_obsJroman_obs , a policy evaluation function that evaluates how good the state sequences ālook to the humanā. The resulting policies can show two distinct failure modes that we formally define and call deceptive inflation and overjustification. In Theorem 4.5 we prove that at least one of them is present for JobssubscriptobsJ_obsJroman_obs-maximizing policies. Later, in Section 5, we will see that an adaptation of the usual RLHF process might sometimes be able to avoid these problems. To model partial observability, we introduce an observation space oāĪ©oā ā Ī© and observation kernel with probabilities POā¢(oā£s)ā[0,1]subscriptconditional01P_O(o s)ā[0,1]Pitalic_O ( o ⣠s ) ā [ 0 , 1 ]. We write POāā¢(oāā£sā)āāt=0TPOā¢(otā£st)āsubscriptāconditionalāsuperscriptsubscriptproduct0subscriptconditionalsubscriptsubscriptP_ O( o s) _t=0^TP_O(o_t s_t)Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) ā āt = 0T Pitalic_O ( oitalic_t ⣠sitalic_t ) for the probability of an observation sequence. We write Ī©āĪ© overā start_ARG Ī© end_ARG for the set of observation sequences that occur with non-zero probability, i.e., oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG if and only if there is sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG such that āt=0TPOā¢(otā£st)>0superscriptsubscriptproduct0subscriptconditionalsubscriptsubscript0 _t=0^TP_O(o_t s_t)>0āt = 0T Pitalic_O ( oitalic_t ⣠sitalic_t ) > 0. If POsubscriptP_OPitalic_O and POāsubscriptāP_ OPoverā start_ARG O end_ARG are deterministic, then we write O:āĪ©:āĪ©O:Sā : S ā Ī© and Oā:āĪ©ā:āĪ© O: Sā overā start_ARG O end_ARG : overā start_ARG S end_ARG ā overā start_ARG Ī© end_ARG for the corresponding observation functions with Oā¢(s)=oO(s)=oO ( s ) = o and Oāā¢(sā)=oā O( s)= ooverā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG for o and oā ooverā start_ARG o end_ARG with POā¢(oā£s)=1subscriptconditional1P_O(o s)=1Pitalic_O ( o ⣠s ) = 1 and POāā¢(oāā£sā)=1subscriptāconditionalā1P_ O( o s)=1Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) = 1, respectively. 4.1 What does RLHF learn from partial observations? We consider the setting where the state is fully observable to the learned policy, but human feedback depends only on a sequence of observations. We assume that the human gives feedback under a Boltzmann rational model similar to Eq. (1), modified such that they form some belief Bā¢(sāā£oā)ā[0,1]conditionalā01 B( s o)ā[0,1]B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) ā [ 0 , 1 ] about the state sequence sā soverā start_ARG s end_ARG based on the observations oā ooverā start_ARG o end_ARG. We then assume preferences are Boltzmann rational in the expected returns under this belief, instead of the actual returns. The assumption of Boltzmann rationality is false in practice (Evans et al., 2015; Majumdar et al., 2017; Buehler et al., 1994), but note that it is an optimistic assumption: Even though our model is a simplification, we expect that practical issues can be at least as bad as the ones we will discuss. See also Example D.4 for an example showing that it is sometimes generally not possible to find a human model that leads to good outcomes under RLHF. Future work could investigate different human models and their impact under partial observability in greater detail. To formalize our setting, we collect human beliefs into a matrix ā(Bā¢(sāā£oā))oā,sāāāĪ©āĆāāsubscriptconditionalāsuperscriptāāĪ©ā B (B( s o) )_ o% , sā R Ć SB ā ( B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) )overā start_ARG o end_ARG , overā start_ARG s end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG S end_ARG. The expected returns for observations oā ooverā start_ARG o end_ARG are given by sāā¼B(ā ā£oā)[Gā¢(sā)]=(ā G)ā¢(oā) *E_ s B(Ā· o) [G( s)% ]=( BĀ· G)( o)Eoverā start_ARG s end_ARG ā¼ B ( ā ⣠overā start_ARG o end_ARG ) [ G ( overā start_ARG s end_ARG ) ] = ( B ā G ) ( overā start_ARG o end_ARG ). We view GāāāsuperscriptāāGā R SG ā blackboard_Roverā start_ARG S end_ARG and ā GāāĪ©āā superscriptāāĪ© BĀ· Gā R B ā G ā blackboard_Roverā start_ARG Ī© end_ARG as both column vectors and functions. Plugging these expected returns into Eq. (1) gives PRā¢(oāā»oāā²)āĻā¢(βā¢((ā G)ā¢(oā)ā(ā G)ā¢(oāā²))).āsuperscriptsucceedsāsuperscriptāā²ā āā superscriptāā²P^R ( o o 1.0pt ) Ļ % (β (( BĀ· G)( o)-(% BĀ· G)( o 1.0pt ) ) ).Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) ā Ļ ( β ( ( B ā G ) ( overā start_ARG o end_ARG ) - ( B ā G ) ( overā start_ARG o end_ARGā² ) ) ) . (2) This is an instance of reward-rational implicit choice (Jeon et al., 2020), with the function oāā¦B(ā ā£oā) o B(Ā· o)overā start_ARG o end_ARG ⦠B ( ā ⣠overā start_ARG o end_ARG ) as the grounding function. If observations are deterministic, we can write Oāā¢(sā)=oā O( s)= ooverā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG for oā ooverā start_ARG o end_ARG with POāā¢(oāā£sā)=1subscriptāconditionalā1P_ O( o s)=1Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) = 1. We can then recover the fully observable case Eq. (1) with BB and Oā Ooverā start_ARG O end_ARG being the identity. The belief B can be any distribution as long as it sums to 1111 over sā soverā start_ARG s end_ARG. The human could arrive at such a belief via Bayesian updates, assuming knowledge of P0subscript0P_0P0, TT, POsubscriptP_OPitalic_O, and a prior over the policy that generates the trajectories (see Appendix C.1). None of our results rely on this more detailed model. We assume the human gives feedback according to Eq. (2) but the system uses the standard RLHF algorithm based on Eq. (1). We define the following observation return function GobssubscriptobsG_obsGroman_obs, and we show in Appendix D.1 that if observations are deterministic, RLHF infers this up to an additive constant. Gobsā¢(sā)āoāā¼POā(ā ā£sā)[(ā G)ā¢(oā)],G_obs( s) *E_ o% P_ O(Ā· s) [ ( B% Ā· G )( o) ],Groman_obs ( overā start_ARG s end_ARG ) ā Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ ( B ā G ) ( overā start_ARG o end_ARG ) ] , (3) For deterministic POāsubscriptāP_ OPoverā start_ARG O end_ARG, this can be simplified to Gobsā¢(sā)=(ā G)ā¢(Oāā¢(sā))subscriptobsāā āG_obs( s)= ( BĀ· G )% ( O( s) )Groman_obs ( overā start_ARG s end_ARG ) = ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) where POāā¢(Oāā¢(sā)ā£sā)=1subscriptāconditionalā1P_ O( O( s) s)=1Poverā start_ARG O end_ARG ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ⣠overā start_ARG s end_ARG ) = 1. Note that deterministic observations can be ambiguous if multiple states produce the same observation. Unlike in the fully observable case of Proposition 3.1, a return function might be inferred that implies an incorrect set of optimal policies. We define the resulting policy evaluation function JobssubscriptobsJ_obsJroman_obs by Jobsā¢(Ļ)āsāā¼PĻā¢(sā)[Gobsā¢(sā)].āsubscriptobssubscriptsimilar-toāsuperscriptāsubscriptobsāJ_obs(Ļ) *E_ s P^% Ļ( s) [G_obs( s) ].Jroman_obs ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ Groman_obs ( overā start_ARG s end_ARG ) ] . (4) This is the function which a standard reinforcement learning algorithm would optimize given the inferred return function GobssubscriptobsG_obsGroman_obs. We summarize this as follows: Proposition 4.1. In partially observable settings with deterministic observations, a policy is optimal according to RLHF, i.e., according to a return function model that would be learned by RLHF with infinite comparison data, if it maximizes JobssubscriptobsJ_obsJroman_obs. Note that in this definition, and specifically in the formula for GobssubscriptobsG_obsGroman_obs, the human does not have knowledge of the policy Ļ that generates the state sequence sā soverā start_ARG s end_ARG. In Appendix D.2, we briefly discuss the unrealistic case that the human does know the precise policy and is an ideal Bayesian reasoner over the true environment dynamics. In that case, Jobs=JsubscriptobsJ_obs=JJroman_obs = J, i.e. there is no discrepancy between true and inferred returns. Intuitively, even if the human would not make any observations, they could give correct feedback essentially by estimating the policyās expected return explicitly. In our case, however, a policy achieving high JobssubscriptobsJ_obsJroman_obs produces state sequences sā soverā start_ARG s end_ARG whose observation sequence Oāā¢(sā)ā O( s)overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) looks good according to the humanās belief Bā¢(sāā²ā£Oāā¢(sā))conditionalsuperscriptāā²āB ( s 1.0pt O( s) )B ( overā start_ARG s end_ARGⲠ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ). This hints at a possible source of deception: if the policy achieves sequences whose observations look good at the expense of actual value Gā¢(sā)āG( s)G ( overā start_ARG s end_ARG ), we might intuitively call this deceptive behavior. We now analyze this point in greater detail. 4.2 An ontology of behaviors Figure 3: Behaviors defined by increasing and decreasing the humanās over- and underestimation error. RLHF with partial observations results in incentives to increase overestimation error and decrease underestimation error (Theorem 4.5). We will evaluate state sequences based on the extent to which they lead to the human overestimating or underestimating the reward in expectation. Recall that GobssubscriptobsG_obsGroman_obs from Equation 3 measures the expected return from the perspective of a human with some belief function B and access to only observations, whereas G are the true returns. That leads us to the following definition: Definition 4.2 (Overestimation and Underestimation Error). Let sā soverā start_ARG s end_ARG be a state sequence. We define its overestimation error E+superscriptE^+E+ and underestimation error EāsuperscriptE^-E- by E+ā¢(sā)āmaxā”(0,Gobsā¢(sā)āGā¢(sā)),āsuperscriptā0subscriptobsā E^+( s) (0,G_obs( % s)-G( s) ),E+ ( overā start_ARG s end_ARG ) ā max ( 0 , Groman_obs ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARG ) ) , Eāā¢(sā)āmaxā”(0,Gā¢(sā)āGobsā¢(sā)).āsuperscriptā0āsubscriptobsā E^-( s) (0,G( s)-G_% obs( s) ).E- ( overā start_ARG s end_ARG ) ā max ( 0 , G ( overā start_ARG s end_ARG ) - Groman_obs ( overā start_ARG s end_ARG ) ) . We further define the average overestimation (underestimation) error under a policy Ļ by EĀÆ+ā¢(Ļ)āsāā¼PĻ[E+ā¢(sā)]āsuperscriptĀÆsubscriptsimilar-toāsuperscriptsuperscriptā E^+(Ļ) *E_ s% P^Ļ[E^+( s)]overĀÆ start_ARG E end_ARG+ ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ [ E+ ( overā start_ARG s end_ARG ) ] and EĀÆāā¢(Ļ)āsāā¼PĻ[Eāā¢(sā)]āsuperscriptĀÆsubscriptsimilar-toāsuperscriptsuperscriptā E^-(Ļ) *E_ s% P^Ļ[E^-( s)]overĀÆ start_ARG E end_ARG- ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ [ E- ( overā start_ARG s end_ARG ) ]. We consider a policy Ļ in comparison to some reference policy Ļrefsubscriptref _refĻref. This can loosely be understood as a counterfactual policy in the absence of some intervention, where Ļ is the factual policy resulting from the intervention. We discuss increases and decreases in over- and underestimation error which are implicitly due to some intervention. For our purposes, Ļrefsubscriptref _refĻref will be the true optimal policy, and Ļ will be the JobssubscriptobsJ_obsJroman_obs-optimal policy; the āinterventionā is thus the introduction of partial observability. Figure 3 shows a simple ontology of behaviors that increase and decrease the average over- and underestimation error. Increasing either of these quantities decreases the accuracy of the humanās estimates, and can thus be thought of as āmisleadingā; decreasing either of them improves accuracy and can be thought of as āinformingā. 4.3 Deceptive inflation and overjustification Standard RLHF in the setting of partial observations incentivizes undesirable forms of inflating and justifying. We refer to the philosophical definition of deception offered by Park et al. (2024b), adjustwidth -0.1em-0.1em āthe systematic inducement of false beliefs in the pursuit of some outcome other than the truth,ā to anchor the notion that increasing the overestimation error in order to improve the RLHF objective JobssubscriptobsJ_obsJroman_obs is deceptive, leading to the following definition. Definition 4.3 (Deceptive Inflation). A policy Ļ exhibits deceptive inflation relative to Ļrefsubscriptref _refĻref if EĀÆ+ā¢(Ļ)>EĀÆ+ā¢(Ļref)superscriptĀÆsuperscriptĀÆsubscriptref E^+(Ļ)> E^+( _ref)overĀÆ start_ARG E end_ARG+ ( Ļ ) > overĀÆ start_ARG E end_ARG+ ( Ļref ) and Jobsā¢(Ļ)>Jobsā¢(Ļref)subscriptobssubscriptobssubscriptref J_obs(Ļ)>J_obs( _% ref)Jroman_obs ( Ļ ) > Jroman_obs ( Ļref ). We typically prefer that our AI agents engage in informing behaviors. Undesirable informing behaviors decrease reward despite providing information. We name undesirable justifying behaviors āoverjustificationā as a nod to the overjustification effect from psychology (Deci and Flaste, 1995), in which subjects become dependent on an extrinsic source of motivation to sustain work on a task. Definition 4.4 (Overjustification). A policy Ļ exhibits overjustification relative to Ļrefsubscriptref _refĻref if EĀÆāā¢(Ļ)<EĀÆāā¢(Ļref)superscriptĀÆsuperscriptĀÆsubscriptref E^-(Ļ)< E^-( _ref)overĀÆ start_ARG E end_ARG- ( Ļ ) < overĀÆ start_ARG E end_ARG- ( Ļref ) and Jā¢(Ļ)<Jā¢(Ļref)subscriptref J(Ļ)<J( _ref)J ( Ļ ) < J ( Ļref ). To understand the counterintuitive notion that an agent providing information to the human could be undesirable, consider a PhD student who looks to feedback from their advisor for direction. They meet for one hour a week. Suppose the student explain last weekās work in 15 minutes, leaving the remaining time to discuss next steps. They could instead āoverjustifyā by spending the entire hour going through the last weekās work in far more detail, leaving no time for next steps. From the advisorās perspective, the latter is more informative, but is a worse allocation of limited resources. We now state a key result. See Section D.3 for the proof. Theorem 4.5. Assume that POsubscriptP_OPitalic_O is deterministic. Let Ī obsāsubscriptsuperscriptĪ obs ^*_obsĪ āroman_obs be the set of optimal policies according to a naive application of RLHF under partial observability, and let Ī āsuperscriptĪ ^*Ī ā be the set of optimal policies according to the true objective J. If ĻāāĪ āāĪ obsāsuperscriptsuperscriptĪ subscriptsuperscriptĪ obsĻ^*ā ^* ^*_obsĻā ā Ī ā ā Ī āroman_obs and ĻobsāāĪ obsāāĪ āsubscriptsuperscriptobssubscriptsuperscriptĪ obssuperscriptĪ Ļ^*_obsā ^*_obs ^*Ļāroman_obs ā Ī āroman_obs ā Ī ā, then ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs must exhibit at least one of deceptive inflation or overjustification relative to ĻāsuperscriptĻ^*Ļā. Note that a trajectory sā soverā start_ARG s end_ARG may be more or less likely under ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs than ĻāsuperscriptĻ^*Ļā, regardless of human estimation, so long as on net ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs exhibits deceptive inflation or overjustification. Our analysis extends beyond the special case of RLHF to inverse preference learning (IPL) (Hejna and Sadigh, 2023), and thus to direct preference optimization (DPO) (Rafailov et al., 2023), which IPL generalizes. Theorem 1 in Hejna and Sadigh (2023) shows that IPL will converge to a policy that maximizes an implicit reward function that matches the humanās preference judgments as well as possible. If the humanās preference judgments come from partial observations, then the resulting return function will be GobssubscriptobsG_obsGroman_obs, as we describe in our discussion leading up to Proposition 4.1. This leads to the same problems of deceptive inflation and overjustification that we describe in Theorem 4.5. 4.4 Deception and overjustification in examples Figure 4: Scenarios illustrating failure modes due to partial observability. In each, the agent must install two packages. Formal details of the underlying MDPs are provided in Appendix B. A, top: In the absence of a log message about CUDA, the human is unsure whether the agent skipped it or used the 2>ā /dev/null trick (see Figure 2); if the human is insufficiently skeptical, the trick looks optimal to the agent. B, bottom: Default logging in this case is silent when the NumPy install is successful. The agent can optionally use a ----verbose flag, but this produces a long log that the human prefers not to see. If the human is too skeptical, verbose logging still appears optimal to the agent. We now analyze intentionally simple examples that illustrate the failure modes of RLHF in the presence of partial observability. Our examples feature an agent trained with RLHF to complete tasks in a userās terminal. The output of each command is piped to a log file, which is what the human evaluator sees when making choices for RLHF. For largely mathematical details about these examples, including a tutorial-style analysis of Example A with derivations, see Appendix B. For experimental details, see Section B.5. For a proof of concept on how to mitigate the issues, see Section 5.2. In Fig. 4A, the agentās task is to install Nvidia drivers and CUDA. The agent may skip the CUDA installation, or may attempt it with or without 2>ā /dev/null (which suppresses errors) appended. If the log contains no message about CUDA, the human believes with probability phidesubscripthidep_hidephide that the agent used the 2>ā /dev/null trick and the install failed, and with probability 1āphide1subscripthide1-p_hide1 - phide that the agent didnāt attempt the CUDA install. If phidesubscripthidep_hidephide is sufficiently small, then RLHF will learn from the resulting feedback that the optimal policy is to use the 2>ā /dev/null trick. This is deceptive inflation, because this policy produces trajectories for which the human systematically overestimates the return. In Figure 5(a), we see that the threshold for phidesubscripthidep_hidephide at which the AI switches from a deceptive action (2>ā /dev/null ) to installing CUDA with default logging perfectly agrees with the theoretical prediction made in Example B.1. In Fig. 4B, the agentās task is to install Python and NumPy. The agent can skip the NumPy install, or attempt it with either ādefaultā settings that only logs errors or with a ----verbose flag that adds additional info. In the case of a success, the human prefers not to see this verbose output. If the log contains no message about NumPy, the human believes with probability pdefaultsubscriptdefaultp_defaultpdefault that the agent attempted the install with default logging and succeeded, and with probability 1āpdefault1subscriptdefault1-p_default1 - pdefault that the agent skipped the NumPy install. If pdefaultsubscriptdefaultp_defaultpdefault is sufficiently small, then RLHF will learn from the resulting feedback that the optimal policy is to use ----verbose logging. This is overjustification, because the agent chooses a suboptimal policy in order to avoid the human underestimating its performance. In Figure 5(b), we see that the threshold for pdefaultsubscriptdefaultp_defaultpdefault at which the AI switches from verbose loggin to skipping the whole NumPy installation largely agrees with the theoretical prediction made in Example B.3. (a) Example A (b) Example B Figure 5: Example A: The larger the reward penalty for hiding errors with 2>ā /dev/null , and the larger the humanās belief that the agent used 2>ā /dev/null upon seeing an empty log (phidesubscripthidep_hidephide), the more we expect the agent to install CUDA with default logging in Example A. In Example B.1, we compute a precise theoretical threshold where the behavior should switch. This perfectly agrees with empirical findings. Example B: The larger the reward penalty for verbose logging, and the larger the humanās trust that the agent installed NumPy upon seeing an empty log (pdefaultsubscriptdefaultp_defaultpdefault), the more we expect the agent to skip the NumPy installation entirely. In Example B.3, we compute a precise theoretical threshold where behavior should switch. Except four cases of āverbose loggingā where the theory predicted the agent to skip the NumPy installation, this agrees with empirical findings. See Section B.5 for experimental details. Further examples. We show further, purely mathematical, examples in Appendix D.4. Example D.6 shows that deceptiveness and overjustifying behavior even applies to aspects of the trajectory the policy has no control over: The policy tries to āhide bad luckā and āreveal good luck at a costā. Example D.7, especially (a) and (c), shows that the policies coming out of a naive application of RLHF under partial observability may be suboptimal with positive EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- (and zero EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+) or optimal, but with positive EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ (and zero EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG-). Thus, there can be suboptimality even if the policy is better than it seems, and optimality even when the policy is worse than it seems. 5 Return ambiguity from feedback under known partial observability Weāve seen issues with standard RLHF applied to feedback from partial observations. Part of the problem is model misspecification: the standard RLHF model implicitly assumes full observability. Assuming the humanās partial observability is known, could one do better? We start Section 5.1 by analyzing how much information the feedback process provides about the return function when the humanās choice model under partial observations is known precisely. We show that the feedback determines the correct return function up to an additive constant and a linear subspace we call the ambiguity (Theorem 5.2). If the human had a return function that differed from the true return function by an element in the ambiguity, they would give the exact same feedback ā such return functions are thus feedback-compatible. We then show an example where the ambiguity vanishes, and another where it doesnāt, leading to feedback-compatible return functions that have optimal policies with high regret under the true return function. Finally, in Section 5.2 we explore how one could in theory use Theorem 5.2 as a starting point to design reward learning techniques that work under partial observability. In particular, we experimentally show in a proof of concept that being aware of the humanās partial observability improves performance. In this section we do not assume POsubscriptP_OPitalic_O to be deterministic. 5.1 Feedback-compatibility and ambiguity of return functions Assume that the human gives feedback based on the choice-probabilities from Eq. (2). In the infinite data limit, it can be assumed that the whole collection of probabilities (PGā¢(oāā»oāā²))oā,oāā²subscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā² (P^G ( o o ) )_ o, o% ( Pitalic_G ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² is known since the choice frequencies approach these probabilities. Here, we write PGsuperscriptP^GPitalic_G instead of PRsuperscriptP^RPitalic_R since the reward function only enters the choice probabilities through the corresponding return function G. The question we answer in this section is how much information the choice probabilities provide about G, assuming the human choice model is known and correct. The choice probabilities tell us precisely that the true return function gives rise to these choice probabilities, i.e., is feedback-compatible. This is captured in the following definition: Definition 5.1. Let (PGā¢(oāā»oāā²))oā,oāā²subscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā² (P^G ( o o 1.0pt ) )_% o, o 1.0pt ( Pitalic_G ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² be the vector of choice probabilities and G~~ Gover~ start_ARG G end_ARG a return function corresponding to a reward function R~~ Rover~ start_ARG R end_ARG. Then G~~ Gover~ start_ARG G end_ARG is feedback-compatible (with respect to the vector of choice probabilities) if PG~ā¢(oāā»oāā²)=PGā¢(oāā»oāā²)superscript~succeedsāsuperscriptāā²succeedsāsuperscriptāā²P G( o o 1.0pt )=P^G( o % o 1.0pt )Pover~ start_ARG G end_ARG ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) = Pitalic_G ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) for all oā,oāā²āĪ©āsuperscriptāā²āĪ© o, o 1.0pt ā overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG. Crucially, without further assumptions or inductive biases, no learning algorithm can pick out the true return function among feedback-compatible return functions. It is thus crucial to know whether there are feedback-compatible return functions that are unsafe when using them to optimize a policy. Figure 6: By Theorem 5.2, even with infinite comparison data and access to the correct human model, a hypothetical reward learning system (depicted as a robot) could only infer G up to the ambiguity imā”ā©kerā”imkernelim ā© Bim Ī ā© ker B (purple). Adding an element of the ambiguity to G leads to the exact same choice probabilities for all possible comparisons, and the reward learning system has no way to identify G among the return functions in G+(imā”ā©kerā”)imkernelG+(im ā© % B)G + ( im Ī ā© ker B ) (yellow). This abstract depiction ignores the linearity of these spaces; for a more precise geometric depiction of BB, see Figure 9 in the appendix. We now determine the set of feedback-compatible return functions. Write āāāĆsuperscriptāā ā R SĆSĪ ā blackboard_Roverā start_ARG S end_ARG Ć S for the matrix that maps a reward function to its return function, i.e. (ā R)ā¢(sā)āāt=0Tγtā¢Rā¢(st)āā āsuperscriptsubscript0superscriptsubscript( Ā· R)( s) _t=0^Tγ^% tR(s_t)( Ī ā R ) ( overā start_ARG s end_ARG ) ā āt = 0T γitalic_t R ( sitalic_t ). Its matrix elements are given by sāā¢s=āt=0TĪ“sā¢(st)ā¢Ī³tsubscriptāsuperscriptsubscript0subscriptsubscriptsuperscript _ ss= _t=0^T _s(s_t)% γ^tĪoverā start_ARG s end_ARG s = āt = 0T Ī“italic_s ( sitalic_t ) γitalic_t, where Ī“sā¢(st)=ā¢s=stsubscriptsubscript1subscript _s(s_t)=1\s=s_t\Ī“italic_s ( sitalic_t ) = 1 s = sitalic_t . Then the image imā”imim im Ī is the set of all return functions that can be realized from a reward function given the MDP dynamics TT. Recall the belief matrix =(Bā¢(sāā£oā))oā,sāāāĪ©āĆāsubscriptconditionalāsuperscriptāāĪ©ā B= (B( s o) )_ o, s% ā R Ć SB = ( B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) )overā start_ARG o end_ARG , overā start_ARG s end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG S end_ARG. Taking into account that G itself is in imā”imim im Ī and that G enters the choice probabilities only through ā Gā BĀ· GB ā G ā meaning that the choice probabilities do not vary if we change G additively up to an element in the kernel kerā”kernel Bker B ā we obtain the following result: Theorem 5.2. Let the collection of choice probabilities be given by (PRā¢(oāā»oāā²))oā,oāā²āĪ©āsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P^R ( o o 1.0pt ) )_% o, o 1.0pt ā ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG following a Boltzmann rational model as in Eq. (2). Then a return function G~~ Gover~ start_ARG G end_ARG is feedback-compatible if and only if there is Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī and cāācā Rc ā blackboard_R such that G~=G+Gā²+c~superscriptā² G=G+G +cover~ start_ARG G end_ARG = G + Gā² + c. In particular, the choice probabilities determine G up to an additive constant if and only if kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . See Theorem C.2 and Corollary C.4 for full proofs, and Figure 6 for a visual depiction. This result motivates the following definition: Definition 5.3 (Ambiguity). We call kerā”ā©imā”kernelim B ker B ā© im Ī the ambiguity that is left in the return function when the human choice model and observation-based choice probabilities are known. Note that Theorem 5.2 generalizes the fully observed case from Section 3.2 (Corollary C.10). We extend the theorem in Appendix C.4 to the case when the humanās observations are not known. Special cases of kerā”kernel Bker B and imā”imim im Ī and our theorem can be found in Appendices C.7 and C.5. In particular, if POāsubscriptāP_ OPoverā start_ARG O end_ARG is stochastic and there is only ānoiseā in it (defined as Ī©ā=āĪ©ā = Soverā start_ARG Ī© end_ARG = overā start_ARG S end_ARG and the injectivity of OO) and if the human is a Bayesian reasoner with a fully supported prior over ā Soverā start_ARG S end_ARG, then the choice data determines the return function even if the humanās observations are not known; see Example C.30. Connection to Potential Shaping Under typical technical assumptions (Skalse et al., 2023; Jenner et al., 2022; Ng et al., 1999), potential shaping changes the returns only by an additive constant, and thus never changes optimal policies. It is a non-trivial ambiguity only in the reward function, but it does not lead to actual ambiguity about intended behavior. In contrast, the ambiguity kerā”ā©imā”kernelim B ker B ā© im Ī in the return function that we study can make the optimal policy ambiguous ā it reflects genuine missing information about the intentions of the human. How large is the return ambiguity? For Fig. 4A, one can show that the ambiguity is nontrivial, allowing for feedback-compatible return functions with unsafe optimal policies. Intuitively, since successfully installing CUDA produces the same observation regardless of whether 2>ā /dev/null was used, the choice probabilities donāt give us any information to determine distinct reward values for these two outcomes, only their average over the humanās belief upon observing a successful install. Thus, reward functions assigning arbitrarily high reward to success with 2>ā /dev/null are feedback-compatible. Such reward functions can then lead to an incentive for a learned policy to hide the error messages even with a correct observation model. More details can be found in Section B.4. We saw in Fig. 4B a case where naive RLHF under partial observability can lead to overjustification. However, the humanās feedback and belief model actually provide enough information to determine the return function. The reason is that kerā”kernel Bker B leaves only one degree of freedom that is not ātime-separableā over states, and thus kerā”ā©imā”kernelim B ker B ā© im Ī = 0. More details can be found in Section B.4. 5.2 Toward improving RLHF in partially observable settings To improve RLHF when partial observability is unavoidable, one could take Theorem 5.2 as a starting point to find a learning algorithm that converges to feedback-compatible return functions. This would require the human model to be fully known and specified, including knowledge of the belief probabilities Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ), which can differ from human to human. If one assumes the human is rational, as in Appendix C.1, this requires specifying the humanās policy prior Bā¢(Ļ)B(Ļ)B ( Ļ ). Instead of directly specifying these models, one could also attempt to learn a generative model for Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ). These problems reveal a further conceptual challenge: for complex environments, humans do not form beliefs over the entire environment state s. A better starting point for practical work may thus be to model humans as forming expectations over reward-relevant features of the state. If BB were explicitly known, one could in principle encode BB into the loss function of an adapted RLHF process to learn a feedback-compatible return function; see Section C.3. As a proof of concept, we used this procedure to analyze the examples in Figure 4 empirically, see Table 1. We do this by first learning a reward model by logistic regression against the true choice probabilities of a synthetic human under partial observability, and then learning the optimal Q-function of the resulting reward model with value iteration. The resulting policy chooses a unique action after installation of the nvidia driver (Example A) or Python (Example B) as listed in the āactionā column. Table 1 shows that in 3 of four cases, being āpartial observability awareā (āpo-awareā) leads to the true optimal policy when ānaiveā RLHF does not. In the one case where being āpo-awareā does not improve performance (second line in the table), this is explained by the fact that there is remaining ambiguity in the return function. Curiously, in line 4 our theory also predicts remaining ambiguity, but the optimal policy is learned; we consider this to be luck. We provide more details on our experiments in Section B.5. Table 1: Experiments showing improved performance of po-aware RLHF Ex. p phidesubscripthidep_hidephide pdefaultsubscriptdefaultp_defaultpdefault model action EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ dec. infl. EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- overj. optimal A 0.5 0.5 N/A naive aHsubscripta_Haitalic_H 1.5 ā 0 Ć Ć A 0.5 0.5 N/A po-aware aHsubscripta_Haitalic_H 1.5 ā 0 Ć Ć A 0.1 0.9 N/A naive aCsubscripta_Caitalic_C 0 Ć 0 ā Ć A 0.1 0.9 N/A po-aware aTsubscripta_Taitalic_T 0 Ć 5.4 Ć ā B 0.5 N/A 0.9 naive aTsubscripta_Taitalic_T 4.5 ā 0 ā Ć B 0.5 N/A 0.9 po-aware aDsubscripta_Daitalic_D 0 Ć 0.25 Ć ā B 0.5 N/A 0.1 naive aVsubscripta_Vaitalic_V 0 Ć 0 ā Ć B 0.5 N/A 0.1 po-aware aDsubscripta_Daitalic_D 0 Ć 2.25 Ć ā As we already demonstrated, feedback-compatible return functions can be unsafe due to remaining ambiguity. In Example C.29, we even show a case where some feedback-compatible return functions have optimal policies that are even worse than simply maximizing JobssubscriptobsJ_obsJroman_obs. An important direction for future work is to investigate learning algorithms and inductive biases that help āfindā safe return functions among all those that are feedback-compatible, or that act conservatively given the uncertainty. Another line of inquiry is to determine when the set of feedback-compatible return functions is āsafeā, which depends on the MDP, observation function, and human model. One sufficient condition for feedback-compatible return functions to be safe is the vanishing of the ambiguity kerā”ā©imā”kernelim B ker B ā© im Ī. Even then, one realistically still has to deal with the problem that BB is at best known approximately. Fortunately, in Appendix C.6, we prove that small errors in the assumed belief matrix lead to only small errors in the inferred return function: Theorem 5.4. Assume kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . Let ā+āsubscript B_ % B+ Bbold_Ī ā B + Ī be a small perturbation of BB, where āā¤Ļnorm\| \|ā¤Ļā„ Ī ā„ ā¤ Ļ for sufficiently small Ļ. Let G be the true return function and assume that a hypothetical learning system, assuming the humanās belief is subscript B_ Bbold_Ī, infers the return function G~~ Gover~ start_ARG G end_ARG with the property that ā G~ā subscript~ B_ Ā· GBbold_Ī ā over~ start_ARG G end_ARG has the smallest possible Euclidean distance to ā Gā BĀ· GB ā G. Let rā¢()ā|imā”ārevaluated-atimr( B) B|_% im r ( B ) ā B |im Ī be the (injective) restriction of the operator BB to imā”imim im Ī. Then rā¢()Tā¢rā¢()rsuperscriptrr( B)^Tr( B)r ( B )T r ( B ) is invertible, and there exists a polynomial Qā¢(X,Y)Q(X,Y)Q ( X , Y ) of degree 5555 such that āG~āGāā¤Ļā āGāā Qā¢(ā(rā¢()Tā¢rā¢())ā1ā,ārā¢()ā).norm~ā normnormsuperscriptrsuperscriptr1normr\| G-G\|ā¤ĻĀ·\|G\|Ā· Q ( \| (r(% B)^Tr( B) )^-1% \|,\|r( B)\| ).ā„ over~ start_ARG G end_ARG - G ā„ ā¤ Ļ ā ā„ G ā„ ā Q ( ā„ ( r ( B )T r ( B ) )- 1 ā„ , ā„ r ( B ) ā„ ) . In particular, as we show in the appendix, one can uniformly bound the difference between JG~subscript~J_ GJover~ start_ARG G end_ARG and JGsubscriptJ_GJitalic_G. This yields a regret bound between the policy optimal under G~~ Gover~ start_ARG G end_ARG and an optimal policy ĻāsuperscriptĻ^*Ļā for G. There are also alternatives to modeling the human belief BB. For example, one could mix human evaluations based on high-cost full observations and low-cost partial observations for finding an optimal tradeoff (Mallen and Belrose, 2024). Finally, it would help if the human could query the policy about reward-relevant aspects of the environment to bring the setting closer to RLHF from full observations. This is similar to the problem of eliciting the latent knowledge of a predictor of future observations (Christiano et al., 2021; Christiano and Xu, 2022). While this may avoid the need to specify the humanās belief model Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ), it requires understanding and effectively querying an ML modelās belief, including translating from an ML modelās ontology into a human ontology. 6 Conclusions In this paper, we provided an investigation of challenges when applying RLHF from partial observations. First, we saw that applying RLHF naively when assuming full observability can lead to deceptive inflation and overjustification behavior. Then, we showed that even when the humanās partial observability is known, the set of feedback-compatible return functions can contain irreducible ambiguity. This means that without further inductive biases, no learning algorithm can generally be expected to infer the correct return function. Finally, we recommended further exploratory research to study and improve RLHF for cases when partial observability is unavoidable and provided a proof of concept that modeling the humanās partial observability can improve performance. In conclusion, we recommend caution when using RLHF in situations of partial observability, and hope that further research studies the effects in practice and helps to address these challenges. Limitations We assume the human to be Boltzmann rational and to implicitly compute an expected value of the return, which is unrealistic for actual humans. Other types of choices could be considered, as in reward-rational choice (Jeon et al., 2020) and assistance games (Hadfield-Menell et al., 2016). Finally, we assume that the human forms a belief Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) over the true state sequence sā soverā start_ARG s end_ARG. If the environment is complex, humans will in reality only form beliefs over lower-dimensional representations or features of the state. Impact statement RLHF and its variants are widely used to steer the behavior of language models. Thus, the soundness of RLHF is critical to language modelsā trustworthy deployment. Our work shows that partial observability of humans poses safety challenges for RLHF. We hope our work stimulates further research to study and overcome these challenges, or that it incentivizes researchers to ensure that their human evaluators fully observe the environment state. Author contributions The project was conceived in parallel by Scott and Davis, with a key shift proposed by Leon. Leon proved Propositions 4.1, 5.2 and 5.4, found the first mathematical examples of what became deceptive inflation and overjustification that can be resolved by Theorem 5.2, and wrote the majority of the appendix. Davis conjectured Proposition 4.1, provided early empirical evidence that RLHF under partial observations can lead to deception (not in the paper), defined deception / deceptive inflation and overjustification (with Scott), proved Theorem 4.5, and developed the running examples and figures. Scott guided the project direction and prioritization, gave the conjecture and proof idea for Theorem 5.4, and helped develop the running examples and deception definitions. Erik provided regular detailed feedback and guidance and edited the paper. Anca and Stuart advised this project. Acknowledgments and Disclosure of Funding Leon Lang thanks the Center for Human-Compatible Artificial Intelligence for hosting him during part of this project, and Open Philanthropy for financial support. All authors thank Open Philanthropy for its support of the Center for Human-Compatible Artificial Intelligence. Davis was supported by the Berkeley Existential Risk Initiative. Erik was supported by fellowships from the Future of Life Institute and Open Philanthropy. We thank Benjamin Eysenbach and Benjamin Plaut for detailed comments and feedback on this work, and we thank Elio A. Farina, Mary Marinou, and Alexandra Horn for assistance with graphic design. References Amodei et al. [2017] D. Amodei, P. Christiano, and A. Ray. Learning from human preferences. https://openai.com/research/learning-from-human-preferences, 2017. Accessed: 2023-12-13. Anthropic [2023a] Anthropic. Introducing Claude. https://w.anthropic.com/index/introducing-claude, 2023a. Accessed: 2023-09-05. Anthropic [2023b] Anthropic. Claudeās Constitution. https://w.anthropic.com/index/claudes-constitution, 2023b. Accessed: 2023-09-05. Bai et al. [2022] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. El Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan. Constitutional AI: Harmlessness from AI Feedback. arXiv e-prints, art. arXiv:2212.08073, Dec. 2022. doi: 10.48550/arXiv.2212.08073. Bradley and Terry [1952] R. A. Bradley and M. E. Terry. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4):324ā345, 1952. ISSN 00063444. URL http://w.jstor.org/stable/2334029. Buehler et al. [1994] R. Buehler, D. Griffin, and M. Ross. Exploring the "Planning Fallacy": Why People Underestimate Their Task Completion Times. Journal of Personality and Social Psychology, 67:366ā381, 09 1994. doi: 10.1037/0022-3514.67.3.366. Burns et al. [2023] C. Burns, H. Ye, D. Klein, and J. Steinhardt. Discovering Latent Knowledge in Language Models Without Supervision. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/forum?id=ETKGuby0hcs. Casper et al. [2023] S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, T. Wang, S. Marks, C.-R. Segerie, M. Carroll, A. Peng, P. Christoffersen, M. Damani, S. Slocum, U. Anwar, A. Siththaranjan, M. Nadeau, E. J. Michaud, J. Pfau, D. Krasheninnikov, X. Chen, L. Langosco, P. Hase, E. Bıyık, A. Dragan, D. Krueger, D. Sadigh, and D. Hadfield-Menell. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arxiv e-prints, 2023. Chidambaram et al. [2024] K. Chidambaram, K. V. Seetharaman, and V. Syrgkanis. Direct Preference Optimization With Unobserved Preference Heterogeneity, 2024. URL https://arxiv.org/abs/2405.15065. Christiano and Xu [2022] P. Christiano and M. Xu. ELK prize results. https://w.alignmentforum.org/posts/zjMKpSB2Xccn9qi5t/elk-prize-results, 2022. Accessed: 2024-02-15. Christiano et al. [2017] P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei. Deep Reinforcement Learning from Human Preferences. arXiv e-prints, art. arXiv:1706.03741, June 2017. doi: 10.48550/arXiv.1706.03741. Christiano et al. [2021] P. Christiano, A. Cotra, and M. Xu. Eliciting Latent Knowledge. https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H_EpsnjrC1dwZXR37PC8/edit, 2021. Accessed: 2023-04-25. Deci and Flaste [1995] E. L. Deci and R. Flaste. Why we do what we do: The dynamics of personal autonomy. GP Putnamās Sons, 1995. Denison et al. [2024] C. Denison, M. MacDiarmid, F. Barez, D. Duvenaud, S. Kravec, S. Marks, N. Schiefer, R. Soklaski, A. Tamkin, J. Kaplan, B. Shlegeris, S. R. Bowman, E. Perez, and E. Hubinger. Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models, 2024. URL https://arxiv.org/abs/2406.10162. El Ghaoui [2002] L. El Ghaoui. Inversion error, condition number, and approximate inverses of uncertain matrices. Linear Algebra and its Applications, 343-344:171ā193, 2002. ISSN 0024-3795. doi: https://doi.org/10.1016/S0024-3795(01)00273-7. URL https://w.sciencedirect.com/science/article/pii/S0024379501002737. Special Issue on Structured and Infinite Systems of Linear equations. Evans et al. [2015] O. Evans, A. Stuhlmueller, and N. D. Goodman. Learning the Preferences of Ignorant, Inconsistent Agents. arxiv e-prints, 2015. Evans et al. [2021] O. Evans, O. Cotton-Barratt, L. Finnveden, A. Bales, A. Balwit, P. Wills, L. Righetti, and W. Saunders. Truthful AI: Developing and Governing AI that does not lie. arxiv e-prints, 2021. Fern et al. [2014] A. Fern, S. Natarajan, K. Judah, and P. Tadepalli. A Decision-Theoretic Model of Assistance. J. Artif. Int. Res., 50(1):71ā104, may 2014. ISSN 1076-9757. Geiger et al. [1990] D. Geiger, T. Verma, and J. Pearl. Identifying independence in bayesian networks. Networks, 20:507ā534, 1990. URL https://api.semanticscholar.org/CorpusID:1938713. Gemini Team [2023] G. Gemini Team. Gemini: A Family of Highly Capable Multimodal Models. https://storage.googleapis.com/deepmind-media/gemini/gemini_1_report.pdf, 2023. Accessed: 2023-12-11. Hadfield-Menell et al. [2016] D. Hadfield-Menell, A. Dragan, P. Abbeel, and S. Russell. Cooperative Inverse Reinforcement Learning. arXiv e-prints, art. arXiv:1606.03137, June 2016. doi: 10.48550/arXiv.1606.03137. Hejna and Sadigh [2023] J. Hejna and D. Sadigh. Inverse Preference Learning: Preference-based RL without a Reward Function. arXiv e-prints, art. arXiv:2305.15363, May 2023. doi: 10.48550/arXiv.2305.15363. HofstƤtter et al. [2023] F. HofstƤtter, F. R. Ward, HarrietW, L. Thomson, O. J, P. Bartak, and S. F. Brown. Tall Tales at Different Scales: Evaluating Scaling Trends for Deception in Language Models. https://w.alignmentforum.org/posts/pip63HtEAxHGfSEGk/tall-tales-at-different-scales-evaluating-scaling-trends-for, 2023. Accessed: 2024-01-23. Huang et al. [2023] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. arXiv preprint arXiv:2311.05232, 2023. Hubinger et al. [2019] E. Hubinger, C. van Merwijk, V. Mikulik, J. Skalse, and S. Garrabrant. Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv e-prints, art. arXiv:1906.01820, June 2019. doi: 10.48550/arXiv.1906.01820. Jenner et al. [2022] E. Jenner, H. van Hoof, and A. Gleave. Calculus on MDPs: Potential Shaping as a Gradient, 2022. Jeon et al. [2020] H. J. Jeon, S. Milli, and A. Dragan. Reward-rational (implicit) choice: A unifying formalism for reward learning. In H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 4415ā4426. Curran Associates, Inc., 2020. URL https://proceedings.neurips.c/paper_files/paper/2020/file/2f10c1578a0706e06b6d7db6f0b4a6af-Paper.pdf. Kausik et al. [2024] C. Kausik, M. Mutti, A. Pacchiano, and A. Tewari. A Theoretical Framework for Partially Observed Reward-States in RLHF, 2024. URL https://arxiv.org/abs/2402.03282. Lin et al. [2022] S. Lin, J. Hilton, and O. Evans. TruthfulQA: Measuring How Models Mimic Human Falsehoods. arxiv e-prints, 2022. Majumdar et al. [2017] A. Majumdar, S. Singh, A. Mandlekar, and M. Pavone. Risk-sensitive inverse reinforcement learning via coherent risk models. In N. Amato, S. Srinivasa, N. Ayanian, and S. Kuindersma, editors, Robotics, Robotics: Science and Systems, United States, 2017. MIT Press Journals. doi: 10.15607/rss.2017.xiii.069. Mallen and Belrose [2024] A. Mallen and N. Belrose. Balancing Label Quantity and Quality for Scalable Elicitation, 2024. URL https://arxiv.org/abs/2410.13215. Manyika [2023] J. Manyika. An overview of Bard: an early experiment with generative AI. https://ai.google/static/documents/google-about-bard.pdf, 2023. Accessed: 2023-09-05. Mindermann and Armstrong [2018] S. Mindermann and S. Armstrong. Occamās Razor is Insufficient to Infer the Preferences of Irrational Agents. In Proceedings of the 32nd International Conference on Neural Information Processing Systems, NIPSā18, page 5603ā5614, Red Hook, NY, USA, 2018. Curran Associates Inc. Mu et al. [2024] T. Mu, A. Helyar, J. Heidecke, J. Achiam, A. Vallone, I. Kivlichan, M. Lin, A. Beutel, J. Schulman, and L. Weng. Rule Based Rewards for Language Model Safety, 2024. URL https://cdn.openai.com/rule-based-rewards-for-language-model-safety.pdf. Accessed: 2024-10-28. Ng et al. [1999] A. Y. Ng, D. Harada, and S. J. Russell. Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In Proceedings of the Sixteenth International Conference on Machine Learning, ICML ā99, page 278ā287, San Francisco, CA, USA, 1999. Morgan Kaufmann Publishers Inc. ISBN 1558606122. Ng et al. [2000] A. Y. Ng, S. Russell, et al. Algorithms for Inverse Reinforcement Learning. In ICML, volume 1, page 2, 2000. OpenAI [2022] OpenAI. Introducing ChatGPT. https://openai.com/blog/chatgpt, 2022. Accessed: 2024-02-06. OpenAI [2023] OpenAI. ChatGPT Plugins. https://openai.com/index/chatgpt-plugins/, 2023. Accessed: 2024-05-22. OpenAI [2024a] OpenAI. OpenAI o1 System Card, 2024a. URL https://cdn.openai.com/o1-system-card.pdf. Accessed: 2024-10-28. OpenAI [2024b] OpenAI. Model Spec, 2024b. URL https://cdn.openai.com/spec/model-spec-2024-05-08.html. Accessed: 2024-10-28. OpenAI [2024c] OpenAI. Privacy Policy. https://openai.com/policies/privacy-policy//, 2024c. Accessed: 2024-05-22. Park et al. [2024a] C. Park, M. Liu, D. Kong, K. Zhang, and A. Ozdaglar. RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation, 2024a. URL https://arxiv.org/abs/2405.00254. Park et al. [2024b] P. S. Park, S. Goldstein, A. OāGara, M. Chen, and D. Hendrycks. Ai deception: A survey of examples, risks, and potential solutions. Patterns, 5(5), 2024b. Rafailov et al. [2023] R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. arxiv e-prints, 2023. Scheurer et al. [2023] J. Scheurer, M. Balesni, and M. Hobbhahn. Technical Report: Large Language Models can Strategically Deceive their Users when Put Under Pressure. arxiv e-prints, 2023. Shah et al. [2019] R. Shah, D. Krasheninnikov, J. Alexander, P. Abbeel, and A. Dragan. The Implicit Preference Information in an Initial State. In International Conference on Learning Representations, 2019. URL https://openreview.net/forum?id=rkevMnRqYQ. Shah et al. [2021] R. Shah, P. Freire, N. Alex, R. Freedman, D. Krasheninnikov, L. Chan, M. D. Dennis, P. Abbeel, A. Dragan, and S. Russell. Benefits of Assistance over Reward Learning, 2021. URL https://openreview.net/forum?id=DFIoGDZejIB. Siththaranjan et al. [2023] A. Siththaranjan, C. Laidlaw, and D. Hadfield-Menell. Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. arXiv preprint arXiv:2312.08358, 2023. Skalse and Abate [2022] J. Skalse and A. Abate. Misspecification in Inverse Reinforcement Learning. arXiv e-prints, art. arXiv:2212.03201, Dec. 2022. doi: 10.48550/arXiv.2212.03201. Skalse et al. [2023] J. M. V. Skalse, M. Farrugia-Roberts, S. Russell, A. Abate, and A. Gleave. Invariance in Policy Optimisation and Partial Identifiability in Reward Learning. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 32033ā32058. PMLR, 23ā29 Jul 2023. URL https://proceedings.mlr.press/v202/skalse23a.html. Stray [2023] J. Stray. The AI Learns to Lie to Please You: Preventing Biased Feedback Loops in Machine-Assisted Intelligence Analysis. Analytics, 2(2):350ā358, 2023. ISSN 2813-2203. doi: 10.3390/analytics2020020. URL https://w.mdpi.com/2813-2203/2/2/20. Touvron et al. [2023] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom. Llama 2: Open Foundation and Fine-Tuned Chat Models. arxiv e-prints, 2023. Ward et al. [2023] F. R. Ward, F. Belardinelli, F. Toni, and T. Everitt. Honesty Is the Best Policy: Defining and Mitigating AI Deception. arxiv e-prints, 2023. Wen et al. [2024] J. Wen, R. Zhong, A. Khan, E. Perez, J. Steinhardt, M. Huang, S. R. Bowman, H. He, and S. Feng. Language Models Learn to Mislead Humans via RLHF, 2024. URL https://arxiv.org/abs/2409.12822. Wu [2024] S. Wu. Introducing Devin, the first AI software engineer. https://w.cognition-labs.com/introducing-devin, 2024. Accessed: 2024-05-06. Zhuang and Hadfield-Menell [2020] S. Zhuang and D. Hadfield-Menell. Consequences of Misaligned AI. In Proceedings of the 34th International Conference on Neural Information Processing Systems, NIPSā20, Red Hook, NY, USA, 2020. Curran Associates Inc. ISBN 9781713829546. Ziebart et al. [2008] B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey. Maximum entropy inverse reinforcement learning. In D. Fox and C. P. Gomes, editors, AAAI, pages 1433ā1438. AAAI Press, 2008. ISBN 978-1-57735-368-3. URL http://dblp.uni-trier.de/db/conf/aaai/aaai2008.html#ZiebartMBD08. Appendix In the appendix, we provide more extensive theory, proofs, and examples. The appendix makes free use of concepts and notation defined in the main paper. In particular, throughout we assume a general MDP together with observation kernel PO:āĪ©:subscriptāĪ©P_O:Sā _O : S ā Ī© and a human with general belief kernel Bā¢(oāā£sā)conditionalāB( o s)B ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ), unless otherwise stated. See the list of Symbols in Section A to refresh notation. In Section B we supplement the examples from the main paper with more mathematical details. In Section C, we provide an extensive theory for appropriately modeled partial observability in RLHF. This can mainly be considered a supplement to Section 5 and contains our main theorems, supplementary results, analysis of special cases, and examples. In Section D, we analyze the naive application of RLHF under partial observability, which means that the learning system is not aware of the humanās partial observability. This section is essentially a supplement to Section 4 and contains an analysis of the policy evaluation function JobssubscriptobsJ_obsJroman_obs, of deceptive inflation and overjustification, and further extensive mathematical examples showing the failures of naive RLHF under partial observability. Contents of the Appendix 1 Introduction 2 Related work 3 Reward identifiability from full observations 3.1 Markov decision processes 3.2 RLHF and identifiability from full observations 4 The impact of partial observations on RLHF 4.1 What does RLHF learn from partial observations? 4.2 An ontology of behaviors 4.3 Deceptive inflation and overjustification 4.4 Deception and overjustification in examples 5 Return ambiguity from feedback under known partial observability 5.1 Feedback-compatibility and ambiguity of return functions 5.2 Toward improving RLHF in partially observable settings 6 Conclusions A List of Symbols B Details for deception and overjustification in examples B.1 Example A: hiding failures B.2 Example B: paying to reveal information B.3 Derivations and Further Details for Fig. 4A B.4 Ambiguity in Section 4.4 examples when modeling partial observability B.5 Experimental details C Modeling the Human in Partially Observable RLHF C.1 The Belief over the State Sequence for Rational Humans C.2 Ambiguity and Identifiability of Reward and Return Functions under Observation Sequence Comparisons C.3 The Ambiguity in Reward Learning in Practice C.4 Identifiability of Return Functions When Human Observations Are Not Known C.5 Simple Special Cases: Full Observability, Deterministic POāsubscriptāP_ OPoverā start_ARG O end_ARG, and Noisy POāsubscriptāP_ OPoverā start_ARG O end_ARG C.6 Robustness of Return Function Identifiability under Belief Misspecification C.6.1 Some Norm Theory for Linear Operators C.6.2 Application to Bounds in the Error of the Return Function C.7 Preliminary Characterizations of the Ambiguity C.8 Examples Supplementing Section 5 D Issues of Naively Applying RLHF under Partial Observability D.1 Optimal Policies under RLHF with Deterministic Partial Observations Maximize JobssubscriptobsJ_obsJroman_obs D.2 Interlude: When the Human Knows the Policy and is a Bayesian Reasoner, then Jobs=JsubscriptobsJ_obs=JJroman_obs = J D.3 Proof of Theorem 4.5 D.4 Further Examples Supplementing Section 4.4 E NeurIPS Paper Checklist Appendix A List of Symbols General MDPs SS Set of environment states sās ā S AA Set of actions aāa ā A of the policy Īā¢()Ī (S)Ī ( S ) Set of probability distributions over SS. Can be defined for any finite set :ĆāĪā¢():āĪT:SĆAā (S)T : S Ć A ā Ī ( S ) Transition kernel P0āĪā¢()subscript0ĪP_0ā (S)P0 ā Ī ( S ) Initial state distribution RāāsuperscriptāRā R^SR ā blackboard_RS Usually the true reward function Rā²āāsuperscriptā²āR ā R^SRā² ā blackboard_RS Usually a reward function in the kernel of ā B B ā Ī R~āā~superscriptā Rā R^Sover~ start_ARG R end_ARG ā blackboard_RS Usually another reward function, e.g. inferred by a learning system γā[0,1]01γā[0,1]γ ā [ 0 , 1 ] Discount factor Ļ:āĪā¢():āĪĻ:Sā (A)Ļ : S ā Ī ( A ) A policy Ļ:āĪā¢():superscriptāĪT^Ļ:Sā (S)Titalic_Ļ : S ā Ī ( S ) Transition kernel for a fixed policy Ļ given by Ļā¢(sā²ā£s)=āaāā¢(sā²ā£s,a)ā Ļā¢(aā£s)superscriptconditionalsuperscriptā²subscriptā conditionalsuperscriptā²conditionalT^Ļ(s s)= _a T(s^% s,a)Ā·Ļ(a s)Titalic_Ļ ( sⲠ⣠s ) = āa ā A T ( sⲠ⣠s , a ) ā Ļ ( a ⣠s ) TāāTā NT ā blackboard_N Finite time horizon PĻāĪā¢(T)superscriptĪsuperscriptP^Ļā (S^T)Pitalic_Ļ ā Ī ( Sitalic_T ) State sequence distribution induced by the policy Ļ āāTāsuperscript S ^Toverā start_ARG S end_ARG ā Sitalic_T State sequences sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG supported by PĻsuperscriptP^ĻPitalic_Ļ GāāāsuperscriptāāGā R SG ā blackboard_Roverā start_ARG S end_ARG Usually the true return function given by Gā¢(sā)=āt=0Tγtā¢Rā¢(st)āsuperscriptsubscript0superscriptsubscriptG( s)= _t=0^Tγ^tR(s_t)G ( overā start_ARG s end_ARG ) = āt = 0T γitalic_t R ( sitalic_t ). Gā²āāāsuperscriptā²āāG ā R SGā² ā blackboard_Roverā start_ARG S end_ARG Usually a return function in kerā”kernel Bker B G~āāā~superscriptāā Gā R Sover~ start_ARG G end_ARG ā blackboard_Roverā start_ARG S end_ARG Usually another return function, e.g. inferred by a learning system J The true policy evaluation function given by Jā¢(Ļ)=sāā¼PĻ[Gā¢(sā)]subscriptsimilar-toāsuperscriptāJ(Ļ)= *E_ s P^Ļ [G( s) ]J ( Ļ ) = Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ [ G ( overā start_ARG s end_ARG ) ]. Additions to General MDPs with Partial Observability Ī© Ī© Set of possible observations oāĪ©oā ā Ī© PO:āĪā¢(Ī©):subscriptāĪĪ©P_O:Sā ( )Pitalic_O : S ā Ī ( Ī© ) Observation kernel determining the humanās observations POā:āĪā¢(Ī©T):subscriptāĪsuperscriptĪ©P_ O: Sā ( ^T )Poverā start_ARG O end_ARG : overā start_ARG S end_ARG ā Ī ( Ī©italic_T ) The observation sequence kernel given by POāā¢(oāā£sā)=āt=0TPOā¢(otā£st)subscriptāconditionalāsuperscriptsubscriptproduct0subscriptconditionalsubscriptsubscriptP_ O ( o s )= _t=0^TP_O (o_t% s_t )Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) = āt = 0T Pitalic_O ( oitalic_t ⣠sitalic_t ) Ī©āāĪ©TāĪ©superscriptĪ© ^Toverā start_ARG Ī© end_ARG ā Ī©italic_T The set of observed sequences oāāĪ©TāsuperscriptĪ© oā ^Toverā start_ARG o end_ARG ā Ī©italic_T that can be sampled from POā(ā ā£sā)P_ O(Ā· s)Poverā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG ) for sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG O:āĪ©:āĪ©O:Sā : S ā Ī© Observation function for the case that POsubscriptP_OPitalic_O is deterministic; given by Oā¢(s)=oO(s)=oO ( s ) = o with o such that POā¢(oā£s)=1subscriptconditional1P_O(o s)=1Pitalic_O ( o ⣠s ) = 1 Oā:āĪ©ā:āĪ© O: Sā overā start_ARG O end_ARG : overā start_ARG S end_ARG ā overā start_ARG Ī© end_ARG Observation sequence function for the case that POāsubscriptāP_ OPoverā start_ARG O end_ARG is deterministic; given by Oāā¢(sā)=oā O( s)= ooverā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG with oā ooverā start_ARG o end_ARG such that POāā¢(oāā£sā)=1subscriptāconditionalā1P_ O( o s)=1Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) = 1 Goāāāsāāāā£Oāā¢(sā)=oāsubscriptāsuperscriptāconditional-setāG_ oā R^\ sā S O( s)=% o\Goverā start_ARG o end_ARG ā blackboard_R overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG Restriction of the return function GāāāsuperscriptāāGā R SG ā blackboard_Roverā start_ARG S end_ARG to sāāāā£Oāā¢(sā)=oāconditional-setā \ sā S O( s)= o \ overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG for fixed oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG GobsāāāsubscriptobssuperscriptāāG_obsā R SGroman_obs ā blackboard_Roverā start_ARG S end_ARG Return function that can be inferred when partial observability is not properly modeled, given by Gobsā¢(sā)ā(ā G)ā¢(Oāā¢(sā))āsubscriptobsāā āG_obs( s) ( BĀ· G% ) ( O( s) )Groman_obs ( overā start_ARG s end_ARG ) ā ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) JobssubscriptobsJ_obsJroman_obs Observation policy evaluation function, defined in Eq. (4) State- and Observation Sequences stāsubscripts_t _t ā S The tāth entry in a state sequence sā soverā start_ARG s end_ARG sāāTāsuperscript s ^Toverā start_ARG s end_ARG ā Sitalic_T State sequence sā=s0,ā¦,sTāsubscript0ā¦subscript s=s_0,ā¦,s_Toverā start_ARG s end_ARG = s0 , ⦠, sitalic_T s^āt^superscript s ^tover start_ARG s end_ARG ā Sitalic_t State sequence segment s^=s0,ā¦,st^subscript0ā¦subscript s=s_0,ā¦,s_tover start_ARG s end_ARG = s0 , ⦠, sitalic_t for tā¤Tt⤠Tt ⤠T otāĪ©subscriptĪ©o_tā _t ā Ī© The tāth entry in an observation sequence oā ooverā start_ARG o end_ARG oāāĪ©TāsuperscriptĪ© oā ^Toverā start_ARG o end_ARG ā Ī©italic_T Observation sequence oā=o0,ā¦,oTāsubscript0ā¦subscript o=o_0,ā¦,o_Toverā start_ARG o end_ARG = o0 , ⦠, oitalic_T o^āĪ©t^superscriptĪ© oā ^tover start_ARG o end_ARG ā Ī©italic_t Observation sequence segment o^=o0,ā¦,ot^subscript0ā¦subscript o=o_0,ā¦,o_tover start_ARG o end_ARG = o0 , ⦠, oitalic_t for tā¤Tt⤠Tt ⤠T The Humanās Belief Bā¢(Ļā²)superscriptā²B(Ļ )B ( Ļā² ) The humanās policy prior Bā¢(sā)āB( s)B ( overā start_ARG s end_ARG ) The humanās prior belief that a sequence sā soverā start_ARG s end_ARG will be sampled, given by Bā¢(sā)=ā«Ļā²Bā¢(Ļā²)ā¢PĻā²ā¢(sā)ā¢Ļā²āsubscriptsuperscriptā²superscriptsuperscriptā²ādifferential-dsuperscriptā²B( s)= _Ļ B(Ļ )P^Ļ ( s)dĻ B ( overā start_ARG s end_ARG ) = ā«Ļā² B ( Ļā² ) Pitalic_Ļ start_POSTSUPERSCRIPT ā² end_POSTSUPERSCRIPT ( overā start_ARG s end_ARG ) d Ļā² Bā¢(sāā£oā)conditionalāB ( s o )B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) The humanās belief of a state sequence given an observation sequence, see Proposition C.1 for a Bayesian version BĻā¢(sāā£oā)superscriptconditionalāB^Ļ( s o)Bitalic_Ļ ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) The humanās belief of a state sequence given an observation sequence; it is allowed to depend on the true policy Ļ, see Proposition C.1 Boāāāsāāāā£Oāā¢(sā)=oāsubscriptāsuperscriptāconditional-setāB_ oā R^\ sā S O( s)=% o\Boverā start_ARG o end_ARG ā blackboard_R overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG Vector of prior probabilities Bā¢(sā)āB( s)B ( overā start_ARG s end_ARG ) for sāāsāāāā£Oāā¢(sā)=oāconditional-setā sā \ sā S O( s)= o \overā start_ARG s end_ARG ā overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG Identifiability Theorem β>00β>0β > 0 The inverse temperature parameter of the Boltzmann rational human Ļ:āā(0,1):āā01Ļ: Rā(0,1)Ļ : blackboard_R ā ( 0 , 1 ) The sigmoid function given by Ļā¢(x)=11+expā”(āx)11Ļ(x)= 11+ (-x)Ļ ( x ) = divide start_ARG 1 end_ARG start_ARG 1 + exp ( - x ) end_ARG :āāāā:āsuperscriptāsuperscriptāā : R^Sā R % SĪ : blackboard_RS ā blackboard_Roverā start_ARG S end_ARG Function that maps a reward function R to the return function ā”(R) (R)Ī ( R ) with [ā”(R)]ā¢(sā)=āt=0Tγtā¢Rā¢(st)delimited-[]āsuperscriptsubscript0superscriptsubscript [ (R) ]( s)= _t=0^Tγ^% tR(s_t)[ Ī ( R ) ] ( overā start_ARG s end_ARG ) = āt = 0T γitalic_t R ( sitalic_t ) :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© B: R Sā R % B : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG Function that maps a return function G to the expected return function ā”(G) B(G)B ( G ) on observation sequences given by [ā”(G)]ā¢(oā)=sāā¼Bā¢(sāā£oā)[Gā¢(sā)]delimited-[]āsubscriptsimilar-toāconditionalā [ B(G) ]( o)= *E% _ s B( s o) [G( s) ][ B ( G ) ] ( overā start_ARG o end_ARG ) = Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ G ( overā start_ARG s end_ARG ) ] :āāāĪ©ā:āsuperscriptāsuperscriptāāĪ© F: R^Sā R F : blackboard_RS ā blackboard_Roverā start_ARG Ī© end_ARG The composition =ā F= B % F = B ā Ī PRā¢(sāā»sāā²)superscriptsucceedsāsuperscriptāā²P^R ( s s 1.0pt )Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) Boltzmann rational choice probability in the case of full observability (Eq. (1)) PRā¢(oāā»oāā²)superscriptsucceedsāsuperscriptāā²P^R ( o o 1.0pt )Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) Boltzmann rational choice probability in the case of partial observability (Eq. (2)) :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG Abstract linear operator given by [ā”(v)]ā¢(sā)=oāā¼POāā¢(oāā£sā)[vā¢(oā)]delimited-[]āsubscriptsimilar-toāsubscriptāconditionalā [ O(v) ]( s)= *E% _ o P_ O( o s) [v( o) ][ O ( v ) ] ( overā start_ARG s end_ARG ) = Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ v ( overā start_ARG o end_ARG ) ] ā:āĪ©āĆĪ©āāāĆā:tensor-productāsuperscriptāāĪ©āĪ©superscriptāā O O: R % Ć ā R SĆ % SO ā O : blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG Ć overā start_ARG S end_ARG Formally the Kronecker product of OO with itself, explicitly given by [(ā)ā¢(C)]ā¢(sā,sāā²)=oā,oāā²ā¼POā(ā ā£sā,sāā²)[Cā¢(oā,oāā²)] [( O O)(C) ](% s, s 1.0pt )= *E_ o, % o 1.0pt P_ O(Ā· s, s 1.0pt^% ) [C( o, o 1.0pt ) ][ ( O ā O ) ( C ) ] ( overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) = Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] Robustness to Misspecifications āxānorm\|x\|ā„ x ā„ Euclidean norm of the vector xāāksuperscriptāxā R^kx ā blackboard_Rk ānorm\| A\|ā„ A ā„ Matrix norm of the matrix AA, given by āāmaxx,āxā=1ā”āā”xāānormsubscriptnorm1norm\| A\| _x,\ \|x\|=1\|% Ax\|ā„ A ā„ ā maxitalic_x , ā„ x ā„ = 1 ā„ A x ā„ Ļā¢()Ļ( A)Ļ ( A ) Matrix quantity defined in Equation (9) Cā¢(,Ļ)C( A,Ļ)C ( A , Ļ ) Matrix quantity defined in Equation (10) rā¢()rr( B)r ( B ) Restriction of BB to imā”imim im Ī General Sets and (Linear) Functions |A||A|| A | Number of elements in the set A Aā©CAā© CA ā© C Intersection of sets A and C AāŖCAāŖ CA āŖ C Union of sets A and C AāCA CA ā C Relative complement of C in A Ī“xsubscript _xĪ“italic_x The Dirac delta distribution of a point x in a set; given by Ī“xā¢(A)=1subscript1 _x(A)=1Ī“italic_x ( A ) = 1 if xāAxā Ax ā A and Ī“xā¢(A)=0subscript0 _x(A)=0Ī“italic_x ( A ) = 0, else kerā”kernel Aker A The kernel of a linear operator :VāW:ā A:Vā WA : V ā W; given by kerā”=vāVā£ā”(v)=0kernelconditional-set0 A= \vā V A(v)=0% \ker A = v ā V ⣠A ( v ) = 0 imā”imim Aim A The image of a linear operator :VāW:ā A:Vā WA : V ā W; given by imā”=wāWā£āvāV:ā”(v)=wimconditional-set:im A= \wā W ā vā V:% A(v)=w \im A = w ā W ⣠ā v ā V : A ( v ) = w fā1ā¢(y)superscript1f^-1(y)f- 1 ( y ) Preimage of y under a function f:XāY:āf:Xā Yf : X ā Y; given by fā1ā¢(y)=xāXā£fā¢(x)=ysuperscript1conditional-setf^-1(y)= \xā X f(x)=y \f- 1 ( y ) = x ā X ⣠f ( x ) = y Appendix B Details for deception and overjustification in examples Figure 7: Two example MDPs with observation functions in which RLHF chooses undesirable policies. Each box depicts a state with a footer showing the (deterministic) observation produced by that state. Outgoing edges from each box are available actions. A more detailed diagram for the first MDP, with explicit shell commands and log messages, is available in Section B.3. Here we include details to the examples described in Section 4.4 that illustrate the failure modes of RLHF in the presence of partial observability. For each of the following, we will characterize the policy which maximizes JobssubscriptobsJ_obsJroman_obs, as this is the policy RLHF selects for when observations are deterministic; see Proposition 4.1. Our examples feature an agent trained with RLHF to complete tasks in a userās terminal. The output of each command (stdout and stderr) is piped to a log file, which is what the human evaluator sees when making choices for RLHF. We assume that the final state T has a self-transition, episodes have a fixed horizon length of 3 (meaning state sequences have length 4: s0,ā¦,s3subscript0ā¦subscript3s_0,ā¦,s_3s0 , ⦠, s3), γ=11γ=1γ = 1, and there is a fixed initial state s0=Ssubscript0s_0=Ss0 = S. Both examples feature a fixed transition probability 0<p<1010<p<10 < p < 1. We assume that the humanās belief only supports possible explanations: Bā¢(sāā£oā)>0ā¹Oāā¢(sā)=oāconditionalā0ā B( s o)>0 O( s)= oB ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) > 0 ā¹ overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG. For further details about these examples, including a tutorial-style analysis of Example A, see Section B.3. B.1 Example A: hiding failures See Section B.3 for derivations and a tutorial-style analysis of this example. In Fig. 4A (and Fig. 6A), the agentās task is to install CUDA, (states W and WHsubscriptW_HWitalic_H) which requires first installing Nvidia drivers (state I). If the CUDA installation fails, one enters L or LHsubscriptL_HLitalic_H. The aCsubscripta_Caitalic_C action is a command to install CUDA that logs both successes and failures, whereas aHsubscripta_Haitalic_H corresponds to aCsubscripta_Caitalic_C with 2>ā /dev/null appended, which suppresses error messages such that no log message is produced if the installation fails. The human may prefer that the agent not hide a failure; we price this into the reward function with a penalty rā„00rā„ 0r ā„ 0 on the reward at LHsubscriptL_HLitalic_H. The agent may attempt the CUDA installation before installing drivers, but this will fail. There are three pairs of trajectories which produce identical observations. Here we address the most prominent (see Section B.3 for the others): Sā¢Iā¢Tā¢TSITTS I T T and Sā¢Iā¢LHā¢TsubscriptSIL_HTS I Litalic_H T both produce oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā , stylized as a log containing only a success confirmation for Python (Fig. 1, oā2subscriptā2 o_2overā start_ARG o end_ARG2). after successfully installing drivers, a failed CUDA installation with 2>ā /dev/null (Sā¢Iā¢LHā¢TsubscriptSIL_HTS I Litalic_H T) and simply exiting (Sā¢Iā¢Tā¢TSITTS I T T) both produce a log containing only a success confirmation for the drivers (oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā ). Let pHāBā¢(sā=Sā¢Iā¢LHā¢Tā£oā=oā ā¢oIā¢oā ā¢oā )ā(0,1)āsubscriptāconditionalsubscriptāsubscriptsubscriptsubscriptsubscript01 p_H B ( s=SIL_HT o=o_ o% _Io_ o_ )ā(0,1)pitalic_H ā B ( overā start_ARG s end_ARG = S I Litalic_H T ⣠overā start_ARG o end_ARG = oā oitalic_I oā oā ) ā ( 0 , 1 ) be the humanās suspicion, upon seeing only a success confirmation for the drivers, that a failed CUDA installation was hidden. We consider two cases which differ only in whether the expected value of attempting to install CUDA (with default logging, aCsubscripta_Caitalic_C) after installing the drivers is positive. In Example B.1, the success probability p is high enough that the true optimal policy is to attempt the installation; in Example B.2, p is low enough that the optimal policy is to terminate. Example B.1. If the true-reward-optimal policy ĻāsuperscriptĻ^*Ļā is to attempt to install CUDA with default logging (aCsubscripta_Caitalic_C) after installing drivers, but pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then the JobssubscriptobsJ_obsJroman_obs-maximizing policy, which will be chosen by RLHF, is to instead use the 2>ā /dev/null trick. Intuitively, this is because if the installation fails, the human is likely to believe that the agent just didnāt attempt to install CUDA, which is preferred to a failed installation. This is deceptive inflation; the overestimation error EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ for this policy is positive, whereas EĀÆ+ā¢(Ļā)=0superscriptĀÆsuperscript0 E^+(Ļ^*)=0overĀÆ start_ARG E end_ARG+ ( Ļā ) = 0, in line with Theorem 4.5. If pH>55+rsubscript55p_H> 55+rpitalic_H > divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then RLHF will select ĻāsuperscriptĻ^*Ļā. Example B.2. If ĻāsuperscriptĻ^*Ļā is to terminate (aTsubscripta_Taitalic_T) after installing the drivers, but pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then RLHF will learn the deceptive inflation 2>ā /dev/null policy as in Example B.1. However, if pH>55+rsubscript55p_H> 55+rpitalic_H > divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, RLHF will choose the policy that attempts to install CUDA with default logging. Intuitively, this is because the combination of increased suspicion that the agent might be hiding errors (higher pHsubscriptp_Hpitalic_H) and increased penalty on this suspected behavior (higher r) also penalizes our preferred trajectory Sā¢Iā¢Tā¢TSITTS I T T since it can be misinterpreted as hiding errors. This is overjustification; the overestimation error EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ is 0 (thus this is not deceptive inflation), and as required by Theorem 4.5, the underestimation error EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- is 0, lower than Eāā¢(Ļā)=pHā¢(5+r)superscriptsuperscriptsubscript5E^-(Ļ^*)=p_H(5+r)E- ( Ļā ) = pitalic_H ( 5 + r ). B.2 Example B: paying to reveal information In Fig. 4B (and Fig. 6B), the agentās task is to install Python (state I) and to optionally further install NumPy (states W and WVsubscriptW_VWitalic_V). The aDsubscripta_Daitalic_D action corresponds to a command to install NumPy with ādefaultā settings which only logs errors, whereas aVsubscripta_Vaitalic_V corresponds to the same command with a ----verbose flag that adds additional info. In the case of a success, the human distinctly prefers not to see this verbose output; we price this into the reward function with a penalty r>00r>0r > 0 on the reward at WVsubscriptW_VWitalic_V. There is only one pair of trajectories which produce identical observations: after successfully installing Python, a successful NumPy installation with default logging (Sā¢Iā¢Wā¢TSIWTS I W T) and simply exiting (Sā¢Iā¢Tā¢TSITTS I T T) both produce a log containing only a success confirmation for Python (oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā ). Let pDāBā¢(sā=Sā¢Iā¢Wā¢Tā£oā=oā ā¢oIā¢oā ā¢oā )ā(0,1)āsubscriptāconditionalāsubscriptsubscriptsubscriptsubscript01 p_D B( s=SIWT o=o_ o_Io_% o_ )ā(0,1)pitalic_D ā B ( overā start_ARG s end_ARG = S I W T ⣠overā start_ARG o end_ARG = oā oitalic_I oā oā ) ā ( 0 , 1 ) be the humanās optimism, upon seeing only a success confirmation for Python, that NumPy was also successfully installed (without the ----verbose flag). Here we consider only the case where p is large enough that the true optimal policy is to install Python then attempt to install NumPy with default logging (aDsubscripta_Daitalic_D). Example B.3. If ĻāsuperscriptĻ^*Ļā is to attempt to install NumPy with aDsubscripta_Daitalic_D after installing Python, and pD>qā15ā¢(pā¢(6ār)ā1)subscriptā1561 p_D>q 15 (p(6-r)-1 )pitalic_D > q ā divide start_ARG 1 end_ARG start_ARG 5 end_ARG ( p ( 6 - r ) - 1 ), then RLHF will select the policy that terminates after installing Python. Intuitively, this is because the agent can exploit the humanās optimism that NumPy was installed quietly without taking the risk of an observable failure (L). This is deceptive inflation, with an overestimation error EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ of 5ā¢pD5subscript5p_D5 pitalic_D, greater than EĀÆ+ā¢(Ļā)=0superscriptĀÆsuperscript0 E^+(Ļ^*)=0overĀÆ start_ARG E end_ARG+ ( Ļā ) = 0. If instead pD<qsubscriptp_D<qpitalic_D < q, then RLHF will select the policy that attempts the NumPy installation with verbose logging (aVsubscripta_Vaitalic_V). Intuitively, this is because the agent is willing to āpayā the cost of r true reward to prove to the human that it installed NumPy, even when the human does not want to see this proof. This is overjustification; the overestimation error EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ is 0 (thus this is not deceptive inflation), and the underestimation error EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- is 0, lower than EĀÆāā¢(Ļā)=5ā¢pā¢(1āpD)superscriptĀÆsuperscript51subscript E^-(Ļ^*)=5p(1-p_D)overĀÆ start_ARG E end_ARG- ( Ļā ) = 5 p ( 1 - pitalic_D ). B.3 Derivations and Further Details for Fig. 4A Figure 8: An expanded view of Figure 4A. Commands corresponding to the various actions are depicted along edges, and log messages corresponding to the various observations are depicted underneath each state. We first include Figure 8, a more detailed picture of the MDP and observation function in Section B.1, to help ground the narrative details of the example. Next we formally enumerate the details of the MDP and observation function. ⢠=S,I,W,WH,L,LH,TsubscriptsubscriptS=\S,I,W,W_H,L,L_H,T\S = S , I , W , Witalic_H , L , Litalic_H , T . ⢠=aI,aC,aH,aTsubscriptsubscriptsubscriptsubscriptA=\a_I,a_C,a_H,a_T\A = aitalic_I , aitalic_C , aitalic_H , aitalic_T . ⢠TT is as depicted in Figure 8 and Figure 4A. For a state s, any outgoing arrow labeled with an action a (such as aIsubscripta_Iaitalic_I) describes the distribution ā¢(sā²ā£s,a)conditionalsuperscriptā²T(s s,a)T ( sⲠ⣠s , a ) as follows: if the arrow does not split, then ā¢(sā²ā£s,a)=1conditionalsuperscriptā²1T(s s,a)=1T ( sⲠ⣠s , a ) = 1 where sā² is the state the arrow points to; if the arrow does split, then for each successor state sā² it eventually reaches, a probability q is written just before the box corresponding to sā² (for this example, q=pq=pq = p or q=1āp1q=1-pq = 1 - p), and ā¢(sā²ā£s,a)=qconditionalsuperscriptā²T(s s,a)=qT ( sⲠ⣠s , a ) = q. ā ā Additionally, any action taken from a state that does not have an outgoing arrow corresponding to that action will immediately transition to state T, as though aTsubscripta_Taitalic_T had been taken. ā ā Any action taken from state T transitions deterministically to T. ⢠P0ā¢(S)=1subscript01P_0(S)=1P0 ( S ) = 1. ⢠R is as described in the table (the numbers in the top right of each state box) with rā„00rā„ 0r ā„ 0. Additionally, Rā¢(S)=Rā¢(T)=00 R(S)=R(T)=0R ( S ) = R ( T ) = 0. ⢠γ=11γ=1γ = 1. We work with a fixed horizon length of 3, meaning state sequences have length 4 (since time is zero-indexed: s0ā¢s1ā¢s2ā¢s3subscript0subscript1subscript2subscript3s_0s_1s_2s_3s0 s1 s2 s3). The observation function is also depicted in Figure 8. Each state deterministically produces the observation in the lower-right corner of its box in the figure. We also write it in another format in Table 9. Table 9: The observation function O for the example in Section B.1 and Section B.3. s S I W WHsubscriptW_HWitalic_H L LHsubscriptL_HLitalic_H T Oā¢(s)O(s)O ( s ) oā subscripto_ oā oIsubscripto_Ioitalic_I oWsubscripto_Woitalic_W oWsubscripto_Woitalic_W oLsubscripto_Loitalic_L oā subscripto_ oā oā subscripto_ oā We make the additional assumption that the human belief Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) only supports state sequences sā soverā start_ARG s end_ARG which actually produce oā ooverā start_ARG o end_ARG under the sequence observation function Oā Ooverā start_ARG O end_ARG: Bā¢(sāā£oā)>0ā¹Oāā¢(sā)=oāconditionalā0āB( s o)>0 O( s)= oB ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) > 0 ā¹ overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG. In particular, this means that for any oā ooverā start_ARG o end_ARG which is only produced by one sā soverā start_ARG s end_ARG, Bā¢(oāā£sā)=1conditionalā1B( o s)=1B ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) = 1. There are three pairs of state sequences which produce identical observation sequences. For each, we introduce a parameter representing the probability the human infers the first of the pair of state sequences upon seeing their shared observation sequence. 1. Sā¢Iā¢LHā¢TsubscriptSIL_HTS I Litalic_H T and Sā¢Iā¢Tā¢TSITTS I T T both produce oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā , a log containing only a success confirmation for installing drivers, again because Oā¢(LH)=Oā¢(T)=oā subscriptsubscriptO(L_H)=O(T)=o_ O ( Litalic_H ) = O ( T ) = oā . Let pH=Bā¢(sā=Sā¢Iā¢LHā¢Tā£oā=oā ā¢oIā¢oā ā¢oā )subscriptāconditionalsubscriptāsubscriptsubscriptsubscriptsubscriptp_H=B( s=SIL_HT o=o_ o_Io_ o_% )pitalic_H = B ( overā start_ARG s end_ARG = S I Litalic_H T ⣠overā start_ARG o end_ARG = oā oitalic_I oā oā ). 2. Sā¢Tā¢Tā¢TSTTTS T T T and Sā¢LHā¢Tā¢TsubscriptSL_HTTS Litalic_H T T both produce oā ā¢oā ā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_ o_ o_ oā oā oā oā , an empty log, since Oā¢(LH)=Oā¢(T)=oā subscriptsubscriptO(L_H)=O(T)=o_ O ( Litalic_H ) = O ( T ) = oā . Let pHā²=Bā¢(sā=Sā¢LHā¢Tā¢Tā£oā=oā ā¢oā ā¢oā ā¢oā )superscriptsubscriptā²āconditionalsubscriptāsubscriptsubscriptsubscriptsubscriptp_H =B( s=SL_HT o=o_ o_ o_% o_ )pitalic_Hā² = B ( overā start_ARG s end_ARG = S Litalic_H T T ⣠overā start_ARG o end_ARG = oā oā oā oā ). 3. Sā¢Iā¢Wā¢TSIWTS I W T and Sā¢Iā¢WHā¢TsubscriptSIW_HTS I Witalic_H T both produce oā ā¢oIā¢oWā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_Wo_ oā oitalic_I oitalic_W oā , a log containing success confirmations for installing both drivers and CUDA. Let pW=Bā¢(sā=Sā¢Iā¢Wā¢Tā£oā=oā ā¢oIā¢oWā¢oā )subscriptāconditionalāsubscriptsubscriptsubscriptsubscriptp_W=B( s=SIWT o=o_ o_Io_Wo_ )pitalic_W = B ( overā start_ARG s end_ARG = S I W T ⣠overā start_ARG o end_ARG = oā oitalic_I oitalic_W oā ). We assume for simplicity that pHā²=pHsuperscriptsubscriptā²subscriptp_H =p_Hpitalic_Hā² = pitalic_H, i.e. that the human is just as likely to think an empty log following a successful driver installation contains an error that was hidden with 2>ā /dev/null (pHsubscriptp_Hpitalic_H), as they are to think that an entirely empty log contains a hidden error (pHā²subscriptā²p_H pitalic_Hā²). In principle, this need not be true: the human may have differing priors about the agentās behavior in the two different circumstances. However, the algebra to reason about such a case is significantly more cumbersome, and this case reveals no fundamentally different agent behavior under our framework that isnāt present in some simpler case. We can thus write the full B as a matrix as in Table 10. Table 10: The parameterized human belief function B for the example in Section B.1 and Section B.3, expressed as a matrix (rendered as a table). Any empty cell is equal to 0. Sā¢Tā¢Tā¢TSTTTS T T T Sā¢LHā¢Tā¢TsubscriptSL_HTTS Litalic_H T T Sā¢Lā¢Tā¢TSLTTS L T T Sā¢Iā¢Tā¢TSITTS I T T Sā¢Iā¢LHā¢TsubscriptSIL_HTS I Litalic_H T Sā¢Iā¢Lā¢TSILTS I L T Sā¢Iā¢Wā¢TSIWTS I W T Sā¢Iā¢WHā¢TsubscriptSIW_HTS I Witalic_H T oā ā¢oā ā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_ o_ o_ oā oā oā oā 1āpH1subscript1-p_H1 - pitalic_H pHsubscriptp_Hpitalic_H oā ā¢oLā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Lo_ o_ oā oitalic_L oā oā 1 oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā 1āpH1subscript1-p_H1 - pitalic_H pHsubscriptp_Hpitalic_H oā ā¢oIā¢oLā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_Lo_ oā oitalic_I oitalic_L oā 1 oā ā¢oIā¢oWā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_Wo_ oā oitalic_I oitalic_W oā pWsubscriptp_Wpitalic_W 1āpW1subscript1-p_W1 - pitalic_W We have laid the groundwork sufficiently to begin reasoning about the observation return, overestimation and underestimation error, policies which are optimal under the reward function learned by naive RLHF, and the resulting deceptive inflationand overjustification failure modes. We begin by computing the measures of interest for each state sequence, shown in Table 11. Table 11: Measures of interest for each state sequence for the example in Section B.1 and Section B.3. State sequences which produce the same observations have their GobssubscriptobsG_obsGroman_obs columns merged, since they necessarily have the same GobssubscriptobsG_obsGroman_obs. sā soverā start_ARG s end_ARG Gā¢(sā)āG( s)G ( overā start_ARG s end_ARG ) Gobsā¢(sā)āsāā²ā¼B(ā ā£Oā(sā))[Gā¢(sāā²)]G_obs( s) *E_ s% 1.0pt B(Ā· O( s))[G( s 1.0pt^% )]Groman_obs ( overā start_ARG s end_ARG ) ā Eoverā start_ARG s end_ARGā² ā¼ B ( ā ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) [ G ( overā start_ARG s end_ARGā² ) ] E+(sā)āmax(0,E^+( s) (0,E+ ( overā start_ARG s end_ARG ) ā max ( 0 , Eā(sā)āmax(0,E^-( s) (0,E- ( overā start_ARG s end_ARG ) ā max ( 0 , Gobs(sā)āG(sā))G_obs( s)-G( s))Groman_obs ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARG ) ) G(sā)āGobs(sā))G( s)-G_obs( s))G ( overā start_ARG s end_ARG ) - Groman_obs ( overā start_ARG s end_ARG ) ) Sā¢Tā¢Tā¢TSTTTS T T T 00 pHā¢Gā¢(Sā¢LHā¢Tā¢T)+(1āpH)ā¢Gā¢(Sā¢Tā¢Tā¢T)subscriptsubscript1subscriptp_HG(SL_HT)+(1-p_H)G(STTT)pitalic_H G ( S Litalic_H T T ) + ( 1 - pitalic_H ) G ( S T T T ) 00 pHā¢(5+r)subscript5p_H(5+r)pitalic_H ( 5 + r ) Sā¢LHā¢Tā¢TsubscriptSL_HTTS Litalic_H T T ā5ār5-5-r- 5 - r =āpHā¢(5+r)absentsubscript5=-p_H(5+r)= - pitalic_H ( 5 + r ) (1āpH)ā¢(5+r)1subscript5(1-p_H)(5+r)( 1 - pitalic_H ) ( 5 + r ) 00 Sā¢Lā¢Tā¢TSLTTS L T T ā55-5- 5 ā55-5- 5 00 00 Sā¢Iā¢Tā¢TSITTS I T T 1111 pHā¢Gā¢(Sā¢Iā¢LHā¢T)+(1āpH)ā¢Gā¢(Sā¢Iā¢Tā¢T)subscriptsubscript1subscriptp_HG(SIL_HT)+(1-p_H)G(SITT)pitalic_H G ( S I Litalic_H T ) + ( 1 - pitalic_H ) G ( S I T T ) 00 pHā¢(5+r)subscript5p_H(5+r)pitalic_H ( 5 + r ) Sā¢Iā¢LHā¢TsubscriptSIL_HTS I Litalic_H T ā4ār4-4-r- 4 - r =1āpHā¢(5+r)absent1subscript5=1-p_H(5+r)= 1 - pitalic_H ( 5 + r ) (1āpH)ā¢(5+r)1subscript5(1-p_H)(5+r)( 1 - pitalic_H ) ( 5 + r ) 00 Sā¢Iā¢Lā¢TSILTS I L T ā44-4- 4 ā44-4- 4 00 00 Sā¢Iā¢Wā¢TSIWTS I W T 11111111 pWā¢Gā¢(Sā¢Iā¢Wā¢T)+(1āpW)ā¢Gā¢(Sā¢Iā¢WHā¢T)subscript1subscriptsubscriptp_WG(SIWT)+(1-p_W)G(SIW_HT)pitalic_W G ( S I W T ) + ( 1 - pitalic_W ) G ( S I Witalic_H T ) 00 00 Sā¢Iā¢WHā¢TsubscriptSIW_HTS I Witalic_H T 11111111 =11absent11=11= 11 00 00 As an exercise, checking the computations in Table 11 is a quick way to gain some intuition for how these quantities relate. Itās further useful to speak about these quantities using their names, and work through the stories these expressions tell. Consider the E+ā¢(Sā¢LHā¢Tā¢T)superscriptsubscriptE^+(SL_HT)E+ ( S Litalic_H T T ) cell. What does it mean that this is (1āpH)ā¢(5+r)1subscript5(1-p_H)(5+r)( 1 - pitalic_H ) ( 5 + r )? E+superscriptE^+E+ is the overestimation error; (1āpH)ā¢(5+r)1subscript5(1-p_H)(5+r)( 1 - pitalic_H ) ( 5 + r ) is the expected amount by which the human observer, upon seeing a completely empty log (oā ā¢oā ā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_ o_ o_ oā oā oā oā ), will overestimate the total reward the agent attained when those empty logs were actually produced by Sā¢LHā¢Tā¢TsubscriptSL_HTTS Litalic_H T T. This is a trajectory in which the agent immediately uses the 2>ā /dev/null trick, fails to install CUDA without logging the error, and terminates. Under what circumstances might the human overestimate the total reward when sā=Sā¢LHā¢Tā¢Tāsubscript s=SL_HTToverā start_ARG s end_ARG = S Litalic_H T T? Upon seeing the empty log, the human thinks with probability 1āpH1subscript1-p_H1 - pitalic_H that the agent simply terminated immediately, which would be worth a total reward of 0. Since the actual total reward is ā5ār5-5-r- 5 - r, this is an overestimate by 5+r55+r5 + r. The human thinks with probability pHsubscriptp_Hpitalic_H that Sā¢LHā¢Tā¢TsubscriptSL_HTTS Litalic_H T T occurred. This is correct, so there is no overestimation and this pHsubscriptp_Hpitalic_H case does not contribute to the overestimation error. So we have that with probability 1āpH1subscript1-p_H1 - pitalic_H, the human overestimates the total reward by 5+r55+r5 + r. We can keep going! Why is the underestimation error of Sā¢Iā¢Wā¢TSIWTS I W T equal to 0? Because the only other trajectory with which it can be confused attains the same total reward, so regardless of how the probability mass of the humanās belief divides between them, there will be no underestimation. Can all of the zeros in the overestimation and underestimation error columns be explained this way? We now move on to consider policies rather than state sequences. Since a policy Ļ imposes a distribution PĻsuperscriptP^ĻPitalic_Ļ over state sequences (the āon-policy distributionā), our policy measures are in fact exactly parallel to our state sequence measures. Each one is an expectation over the on-policy distribution of the columns of Table 11. We restrict our attention to deterministic policies which only take actions depicted in Figure 8 (i.e. that never terminate via an action other than aTsubscripta_Taitalic_T), of which there are only six in this MDP. They are enumerated, along with the policy-level measures, in Table 12. Policies will be written as a sequence of actions enclosed in brackets, omitting trailing repeated aTsubscripta_Taitalic_T actions. This is nonstandard notation in an MDP with stochastic transitions, but is unambiguous in this example, because all decisions are made before any stochasticity occurs. The policies are [aT]delimited-[]subscript[a_T][ aitalic_T ], [aHā¢aT]delimited-[]subscriptsubscript[a_Ha_T][ aitalic_H aitalic_T ], [aCā¢aT]delimited-[]subscriptsubscript[a_Ca_T][ aitalic_C aitalic_T ], [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ], [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ], and [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ]. Table 12: Measures of interest for each policy for the example in Section B.1 and Section B.3. Each of the columns here is the on-policy average of the corresponding column in Table 11. Policies are written as sequences of actions, omitting trailing repeated aTsubscripta_Taitalic_T actions. This is nonstandard notation in an MDP with stochastic transitions, but is unambiguous in this example since all decisions are made before any stochasticity occurs. Ļ Jā¢(Ļ)J(Ļ)J ( Ļ ) Jobsā¢(Ļ)subscriptobsJ_obs(Ļ)Jroman_obs ( Ļ ) EĀÆ+ā¢(Ļ)superscriptĀÆ E^+(Ļ)overĀÆ start_ARG E end_ARG+ ( Ļ ) EĀÆāā¢(Ļ)superscriptĀÆ E^-(Ļ)overĀÆ start_ARG E end_ARG- ( Ļ ) [aT]delimited-[]subscript[a_T][ aitalic_T ] 00 āpHā¢(5+r)subscript5-p_H(5+r)- pitalic_H ( 5 + r ) 00 pHā¢(5+r)subscript5p_H(5+r)pitalic_H ( 5 + r ) [aHā¢aT]delimited-[]subscriptsubscript[a_Ha_T][ aitalic_H aitalic_T ] ā5ār5-5-r- 5 - r āpHā¢(5+r)subscript5-p_H(5+r)- pitalic_H ( 5 + r ) (1āpH)ā¢(5+r)1subscript5(1-p_H)(5+r)( 1 - pitalic_H ) ( 5 + r ) 00 [aCā¢aT]delimited-[]subscriptsubscript[a_Ca_T][ aitalic_C aitalic_T ] ā55-5- 5 ā55-5- 5 00 00 [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ] 1111 1āpHā¢(5+r)1subscript51-p_H(5+r)1 - pitalic_H ( 5 + r ) 00 pHā¢(5+r)subscript5p_H(5+r)pitalic_H ( 5 + r ) [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] pā¢Gā¢(Sā¢Iā¢WHā¢T)subscriptpG(SIW_HT)p G ( S I Witalic_H T ) pā¢Gobsā¢(Sā¢Iā¢WHā¢T)subscriptobssubscriptpG_obs(SIW_HT)p Groman_obs ( S I Witalic_H T ) (1āp)ā¢(1āpH)ā¢(5+r)11subscript5(1-p)(1-p_H)(5+r)( 1 - p ) ( 1 - pitalic_H ) ( 5 + r ) 00 +(1āp)ā¢Gā¢(Sā¢Iā¢LHā¢T)1subscript+(1-p)G(SIL_HT)+ ( 1 - p ) G ( S I Litalic_H T ) +(1āp)ā¢Gobsā¢(Sā¢Iā¢LHā¢T)1subscriptobssubscript+(1-p)G_obs(SIL_HT)+ ( 1 - p ) Groman_obs ( S I Litalic_H T ) =11ā(1āp)ā¢(15+r)absent11115=11-(1-p)(15+r)= 11 - ( 1 - p ) ( 15 + r ) =11ā(1āp)ā¢[10+pHā¢(5+r)]absent111delimited-[]10subscript5=11-(1-p) [10+p_H(5+r) ]= 11 - ( 1 - p ) [ 10 + pitalic_H ( 5 + r ) ] [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ] pā¢Gā¢(Sā¢Iā¢Wā¢T)pG(SIWT)p G ( S I W T ) pā¢Gobsā¢(Sā¢Iā¢Wā¢T)subscriptobspG_obs(SIWT)p Groman_obs ( S I W T ) 00 00 +(1āp)ā¢Gā¢(Sā¢Iā¢Lā¢T)1+(1-p)G(SILT)+ ( 1 - p ) G ( S I L T ) +(1āp)ā¢Gobsā¢(Sā¢Iā¢Lā¢T)1subscriptobs+(1-p)G_obs(SILT)+ ( 1 - p ) Groman_obs ( S I L T ) =11ā(1āp)ā 15absent11ā 115=11-(1-p)Ā· 15= 11 - ( 1 - p ) ā 15 =11ā(1āp)ā 15absent11ā 115=11-(1-p)Ā· 15= 11 - ( 1 - p ) ā 15 With this we have everything we need to characterize optimal policies under the reward function learned by a naive application of RLHF (āpolicies selected by RLHFā). By Proposition 4.1, we know that if POsubscriptP_OPitalic_O is deterministic, as in this example, RLHF selects policies which maximize JobssubscriptobsJ_obsJroman_obs. In order to understand the behavior of these policies, weāl also need to determine the true optimal policies, i.e. those which maximize J. Weāl proceed in cases, only considering boundary cases (specific measure-zero parameter values for which the result is different) insofar as they are interesting. Case 1: p>1313p> 13p > divide start_ARG 1 end_ARG start_ARG 3 end_ARG. If p>1313p> 13p > divide start_ARG 1 end_ARG start_ARG 3 end_ARG, the CUDA install (with default logging, aCsubscripta_Caitalic_C) is likely enough to succeed that itās worth attempting it: pā Rā¢(W)+(1āp)ā Rā¢(L)>0ā 10 pĀ· R(W)+(1-p)Ā· R(L)>0p ā R ( W ) + ( 1 - p ) ā R ( L ) > 0. It also immediately follows that Jā¢([aIā¢aCā¢aT])=Jobsā¢([aIā¢aCā¢aT])=11ā(1āp)ā 15>1.delimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptsubscript11ā 1151 J([a_Ia_Ca_T])=J_obs([a_Ia_Ca_T])=1% 1-(1-p)Ā· 15>1.J ( [ aitalic_I aitalic_C aitalic_T ] ) = Jroman_obs ( [ aitalic_I aitalic_C aitalic_T ] ) = 11 - ( 1 - p ) ā 15 > 1 . This allows us to eliminate policies [aT]delimited-[]subscript[a_T][ aitalic_T ], [aHā¢aT]delimited-[]subscriptsubscript[a_Ha_T][ aitalic_H aitalic_T ], [aCā¢aT]delimited-[]subscriptsubscript[a_Ca_T][ aitalic_C aitalic_T ], and [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ], which all have Jā¤11J⤠1J ⤠1 and Jobsā¤1subscriptobs1J_obs⤠1Jroman_obs ⤠1. None of them can thus be J-optimal or JobssubscriptobsJ_obsJroman_obs-optimal. All that remains is to compare J and JobssubscriptobsJ_obsJroman_obs for [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] and [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ]. We can check the sign of the differences of these pairs of values, starting with J. Jā¢([aIā¢aCā¢aT])āJā¢([aIā¢aHā¢aT])=(1āp)ā¢r.delimited-[]subscriptsubscriptsubscriptdelimited-[]subscriptsubscriptsubscript1 J([a_Ia_Ca_T])-J([a_Ia_Ha_T])=(1-p)r.J ( [ aitalic_I aitalic_C aitalic_T ] ) - J ( [ aitalic_I aitalic_H aitalic_T ] ) = ( 1 - p ) r . Since p is a probability and r is nonnegative, this value is positive (and thus [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ] is preferred to [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] by the human) if and only if p<11p<1p < 1 and r>00r>0r > 0. Jobsā¢([aIā¢aHā¢aT])āJobsā¢([aIā¢aCā¢aT])=(1āp)ā¢[5āpHā¢(5+r)].subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptsubscript1delimited-[]5subscript5 J_obs([a_Ia_Ha_T])-J_obs% ([a_Ia_Ca_T])=(1-p) [5-p_H(5+r) ].Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) - Jroman_obs ( [ aitalic_I aitalic_C aitalic_T ] ) = ( 1 - p ) [ 5 - pitalic_H ( 5 + r ) ] . This value is positive (and thus [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] is the policy RLHF selects) if and only if p<11p<1p < 1 and pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG. If p=11p=1p = 1, then both differences are 0, and both J and JobssubscriptobsJ_obsJroman_obs are indifferent between the two policies. This makes sense, as they differ only in the case where the CUDA installation fails; this happens with probability 1āp=0101-p=01 - p = 0 when p=11p=1p = 1. Now suppose p<11p<1p < 1. If r=00r=0r = 0, then the human is indifferent between the two policies. This also makes sense, as r is meant to quantify the extent to which the human dislikes suppressed failures; if itās zero, then the human doesnāt care. However, if pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then Jobsā¢([aIā¢aHā¢aT])>Jobsā¢([aIā¢aHā¢aT])subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptsubscriptJ_obs([a_Ia_Ha_T])>J_obs([a_Ia_Ha_% T])Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) > Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ), and thus RLHF favors the 2>ā /dev/null policy [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ]. If p<11p<1p < 1, r>00r>0r > 0, and pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then we have that Jā¢([aIā¢aCā¢aT])>Jā¢([aIā¢aHā¢aT])delimited-[]subscriptsubscriptsubscriptdelimited-[]subscriptsubscriptsubscriptJ([a_Ia_Ca_T])>J([a_Ia_Ha_T])J ( [ aitalic_I aitalic_C aitalic_T ] ) > J ( [ aitalic_I aitalic_H aitalic_T ] ) but Jobsā¢([aIā¢aCā¢aT])>Jobsā¢([aIā¢aHā¢aT])subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptsubscriptJ_obs([a_Ia_Ca_T])>J_obs([a_Ia_Ha_% T])Jroman_obs ( [ aitalic_I aitalic_C aitalic_T ] ) > Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ). Thus RLHF will select the 2>ā /dev/null policy [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ], and by Theorem 4.5, since [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] is not J-optimal, then relative to [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ], it must exhibit deceptive inflation, overjustification, or both. Intuitively, we should be suspicious that deceptive inflation is at play whenever the agent hides information from the human. Indeed, referencing Table 12, we have EĀÆ+ā¢([aIā¢aHā¢aT])=(1āp)ā¢(1āpH)ā¢(5+r)>0=EĀÆ+ā¢([aIā¢aCā¢aT])superscriptĀÆdelimited-[]subscriptsubscriptsubscript11subscript50superscriptĀÆdelimited-[]subscriptsubscriptsubscript E^+([a_Ia_Ha_T])=(1-p)(1-p_H)(5+r)>0= E^+([a_% Ia_Ca_T])overĀÆ start_ARG E end_ARG+ ( [ aitalic_I aitalic_H aitalic_T ] ) = ( 1 - p ) ( 1 - pitalic_H ) ( 5 + r ) > 0 = overĀÆ start_ARG E end_ARG+ ( [ aitalic_I aitalic_C aitalic_T ] ). Together with Jobsā¢([aIā¢aHā¢aT])>Jobsā¢([aIā¢aCā¢aT])subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptsubscriptJ_obs([a_Ia_Ha_T])>J_obs([a_Ia_Ca_% T])Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) > Jroman_obs ( [ aitalic_I aitalic_C aitalic_T ] ), this satisfies the conditions of Definition 4.3, and thus this is an instance of deceptive inflation. If p<11p<1p < 1, r>00r>0r > 0, and pH>55+rsubscript55p_H> 55+rpitalic_H > divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ] is optimal under both J and JobssubscriptobsJ_obsJroman_obs, and in this case, RLHF selects the true optimal policy. Case 2: p<1313p< 13p < divide start_ARG 1 end_ARG start_ARG 3 end_ARG. In this case, the CUDA install is not likely enough to succeed to be worth attempting (under the true reward function). Mathematically, Jā¢([aIā¢aHā¢aT])ā¤Jā¢([aIā¢aCā¢aT])<1=Jā¢([aIā¢aT])delimited-[]subscriptsubscriptsubscriptdelimited-[]subscriptsubscriptsubscript1delimited-[]subscriptsubscriptJ([a_Ia_Ha_T])⤠J([a_Ia_Ca_T])<1=J([a_Ia_T])J ( [ aitalic_I aitalic_H aitalic_T ] ) ⤠J ( [ aitalic_I aitalic_C aitalic_T ] ) < 1 = J ( [ aitalic_I aitalic_T ] ). The other three policies are always worse under J than [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ], so we have our optimal policy Ļā=[aIā¢aT]superscriptdelimited-[]subscriptsubscriptĻ^*=[a_Ia_T]Ļā = [ aitalic_I aitalic_T ]. However, Jobsā¢([aIā¢aHā¢aT])āJobsā¢([aIā¢aT])=pā¢(10+pHā¢(5+r)),subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscript10subscript5 J_obs([a_Ia_Ha_T])-J_obs% ([a_Ia_T])=p(10+p_H(5+r)),Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) - Jroman_obs ( [ aitalic_I aitalic_T ] ) = p ( 10 + pitalic_H ( 5 + r ) ) , which is always positive unless p=00p=0p = 0. As long as p>00p>0p > 0, i.e. installing CUDA is actually possible at all, then regardless of everything else, Jobsā¢([aIā¢aHā¢aT])>Jobsā¢([aIā¢aT])subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptJ_obs([a_Ia_Ha_T])>J_obs([a_Ia_T])Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) > Jroman_obs ( [ aitalic_I aitalic_T ] ). We can interpret this reasonably: however good or bad it might look to terminate after installing drivers (Sā¢Iā¢Tā¢TSITTS I T T), using the 2>ā /dev/null trick will either produce the same observation as Sā¢Iā¢Tā¢TSITTS I T T (and thus the same feedback distribution) or will actually visibly install CUDA successfullyāthe most lucrative possible source of positive feedback. Even if thereās only a sliver of a chance of success and itās a catastrophe to have a silent failure (i.e. rā«0much-greater-than0r 0r ā« 0), however likely you are to infer that the lack of a log message about CUDA means the agent hid a failure, it will penalize [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ] more than it will penalize [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ]: where [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] produces this ambiguous observation with probability 1āp11-p1 - p (when it fails), [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ] will always produce the ambiguous observation. This means that when 0<p<130130<p< 130 < p < divide start_ARG 1 end_ARG start_ARG 3 end_ARG, it is impossible to recover the true optimal policy with naive RLHF. Which policies can possibly be JobssubscriptobsJ_obsJroman_obs-optimal for some setting of the parameters? We can similarly rule out [aT]delimited-[]subscript[a_T][ aitalic_T ] and [aHā¢aT]delimited-[]subscriptsubscript[a_Ha_T][ aitalic_H aitalic_T ] for 0<p<130130<p< 130 < p < divide start_ARG 1 end_ARG start_ARG 3 end_ARG: Jobsā¢([aIā¢aHā¢aT])āJobsā¢([aIā¢aT])=pā¢(10+pHā¢(5+r))>0.subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscript10subscript50 J_obs([a_Ia_Ha_T])-J_obs% ([a_Ia_T])=p(10+p_H(5+r))>0.Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) - Jroman_obs ( [ aitalic_I aitalic_T ] ) = p ( 10 + pitalic_H ( 5 + r ) ) > 0 . We can rule out [aCā¢aT]delimited-[]subscriptsubscript[a_Ca_T][ aitalic_C aitalic_T ] by comparison to [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ]: Jobsā¢([aIā¢aCā¢aT])āJobsā¢([aCā¢aT])=16ā(1āp)ā¢15>0subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscript161150J_obs([a_Ia_Ca_T])-J_obs([a_Ca_T])% =16-(1-p)15>0Jroman_obs ( [ aitalic_I aitalic_C aitalic_T ] ) - Jroman_obs ( [ aitalic_C aitalic_T ] ) = 16 - ( 1 - p ) 15 > 0. So we are left with only [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ] and [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ] as candidate JobssubscriptobsJ_obsJroman_obs-optimal policies. As in Case 1, we find that Jobsā¢([aIā¢aHā¢aT])>Jobsā¢([aIā¢aT])subscriptobsdelimited-[]subscriptsubscriptsubscriptsubscriptobsdelimited-[]subscriptsubscriptJ_obs([a_Ia_Ha_T])>J_obs([a_Ia_T])Jroman_obs ( [ aitalic_I aitalic_H aitalic_T ] ) > Jroman_obs ( [ aitalic_I aitalic_T ] ) if and only if p=11p=1p = 1 or pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG. In case 2 we have assumed p<1313p< 13p < divide start_ARG 1 end_ARG start_ARG 3 end_ARG, leaving only the pHsubscriptp_Hpitalic_H condition. If pH<55+rsubscript55p_H< 55+rpitalic_H < divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then RLHF selects [aIā¢aHā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ha_T][ aitalic_I aitalic_H aitalic_T ]. As in Case 1, this is deceptive inflationrelative to Ļā=[aIā¢aT]superscriptdelimited-[]subscriptsubscriptĻ^*=[a_Ia_T]Ļā = [ aitalic_I aitalic_T ], because EĀÆ+ā¢([aIā¢aHā¢aT])=(1āp)ā¢(1āpH)ā¢(5+r)>0=EĀÆ+ā¢(Ļā).superscriptĀÆdelimited-[]subscriptsubscriptsubscript11subscript50superscriptĀÆsuperscript E^+([a_Ia_Ha_T])=(1-p)(1-p_H)(5+r)>0=% E^+(Ļ^*).overĀÆ start_ARG E end_ARG+ ( [ aitalic_I aitalic_H aitalic_T ] ) = ( 1 - p ) ( 1 - pitalic_H ) ( 5 + r ) > 0 = overĀÆ start_ARG E end_ARG+ ( Ļā ) . If pH>55+rsubscript55p_H> 55+rpitalic_H > divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG, then RLHF selects [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ]. Because this policy is not J-optimal, by Theorem 4.5, we must have deceptive inflation, overjustification, or both. Which is it? Here the optimal policy is to terminate after installing drivers, [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ]. However, pH>55+rsubscript55p_H> 55+rpitalic_H > divide start_ARG 5 end_ARG start_ARG 5 + r end_ARG. This can be rewritten as pHā¢(5+r)>5subscript55p_H(5+r)>5pitalic_H ( 5 + r ) > 5. We have seen this expression pHā¢(5+r)subscript5p_H(5+r)pitalic_H ( 5 + r ) before; it is the underestimation error incurred on sā=Sā¢Iā¢Tā¢Tā s=SITToverā start_ARG s end_ARG = S I T T and therefore also the average underestimation error of policy [aIā¢aT]delimited-[]subscriptsubscript[a_Ia_T][ aitalic_I aitalic_T ]. So here the underestimation error on the optimal policyāthat is, the risk that the human misunderstands optimal behavior (terminating after installing driver) as undesired behavior (attempting a CUDA install that was unlikely to work and hiding the mistake)āis severe enough that the agent opts instead for [aIā¢aCā¢aT]delimited-[]subscriptsubscriptsubscript[a_Ia_Ca_T][ aitalic_I aitalic_C aitalic_T ], a worse policy that attempts the ill-fated CUDA installation only to prove that it wasnāt doing so secretly. In qualitative terms, this is quintessential overjustification behavior. Indeed, relative to reference policy Ļā=[aIā¢aT]superscriptdelimited-[]subscriptsubscriptĻ^*=[a_Ia_T]Ļā = [ aitalic_I aitalic_T ], we have EĀÆāā¢([aIā¢aCā¢aT])=0<pHā¢(5+r)=EĀÆāā¢(Ļā)superscriptĀÆdelimited-[]subscriptsubscriptsubscript0subscript5superscriptĀÆsuperscript E^-([a_Ia_Ca_T])=0<p_H(5+r)= E^-% (Ļ^*)overĀÆ start_ARG E end_ARG- ( [ aitalic_I aitalic_C aitalic_T ] ) = 0 < pitalic_H ( 5 + r ) = overĀÆ start_ARG E end_ARG- ( Ļā ) Jā¢([aIā¢aCā¢aT])=11ā(1āp)ā 15<1=Jā¢(Ļā),delimited-[]subscriptsubscriptsubscript11ā 1151superscript J([a_Ia_Ca_T])=11-(1-p)Ā· 15<1=J(Ļ^*),J ( [ aitalic_I aitalic_C aitalic_T ] ) = 11 - ( 1 - p ) ā 15 < 1 = J ( Ļā ) , and thus by Definition 4.4, this is overjustification. B.4 Ambiguity in Section 4.4 examples when modeling partial observability Consider the example in Fig. 4A when modeling partial observability as in Section 5. By Theorem 5.2, the ambiguity in the return function leaving the choice probabilities invariant is given by kerā”ā©imā”kernelim B ker B ā© im Ī. Let Rā²=(0,0,Rā²ā¢(W),0,Rā²ā¢(WH),0,0)āāS,I,W,L,WH,LH,Tsuperscriptā²00superscriptā²0superscriptā²subscript00superscriptāsubscriptsubscriptR =(0,0,R (W),0,R (W_H),0,0)ā R^\S,I,W,L% ,W_H,L_H,T\Rā² = ( 0 , 0 , Rā² ( W ) , 0 , Rā² ( Witalic_H ) , 0 , 0 ) ā blackboard_R S , I , W , L , Witalic_H , Litalic_H , T be a reward function that we want to parameterize such that Gā²āā Rā²āsuperscriptā²ā superscriptā²G Ā· R Gā² ā Ī ā Rā² ends up in the ambiguity; here, Rā² is interpreted as a column vector. We want ā Gā²=0ā superscriptā²0 BĀ· G =0B ā Gā² = 0. Since the observation sequences oā=oā ā¢oā ā¢oā ā¢oā āsubscriptsubscriptsubscriptsubscript o=o_ o_ o_ o_ overā start_ARG o end_ARG = oā oā oā oā , oā=oā ā¢oLā¢oā ā¢oā āsubscriptsubscriptsubscriptsubscript o=o_ o_Lo_ o_ overā start_ARG o end_ARG = oā oitalic_L oā oā , oā=oā ā¢oIā¢oā ā¢oā āsubscriptsubscriptsubscriptsubscript o=o_ o_Io_ o_ overā start_ARG o end_ARG = oā oitalic_I oā oā , or oā=oā ā¢oIā¢oLā¢oā āsubscriptsubscriptsubscriptsubscript o=o_ o_Io_Lo_ overā start_ARG o end_ARG = oā oitalic_I oitalic_L oā all cannot involve the states W or WHsubscriptW_HWitalic_H, it is clear that they have zero expected return (ā Gā²)ā¢(oā)ā superscriptā²ā( BĀ· G )( o)( B ā Gā² ) ( overā start_ARG o end_ARG ). Set pHā²āBā¢(Sā¢Iā¢WHā¢Tā£oā ā¢oIā¢oWā¢oā )āsuperscriptsubscriptā²conditionalsubscriptsubscriptsubscriptsubscriptsubscriptp_H B (SIW_HT o_ o_Io_Wo_% )pitalic_Hā² ā B ( S I Witalic_H T ⣠oā oitalic_I oitalic_W oā ). Then the condition that ā Gā²=0ā superscriptā²0 BĀ· G =0B ā Gā² = 0 is equivalent to: 00 0 =(ā Gā²)ā¢(oā ā¢oIā¢oWā¢oā )=sāā¼Bā¢(sāā£oā ā¢oIā¢oWā¢oā )[Gā²ā¢(sā)]absentā superscriptā²subscriptsubscriptsubscriptsubscriptsubscriptsimilar-toāconditionalāsubscriptsubscriptsubscriptsubscriptsuperscriptā²ā = ( BĀ· G )(o_% o_Io_Wo_ )= *E_ s B(% s o_ o_Io_Wo_ ) [G ( s)% ]= ( B ā Gā² ) ( oā oitalic_I oitalic_W oā ) = Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠o start_POSTSUBSCRIPT ā oitalic_I oitalic_W oā ) end_POSTSUBSCRIPT [ Gā² ( overā start_ARG s end_ARG ) ] =pHā²ā Gā²ā¢(Sā¢Iā¢WHā¢T)+(1āpHā²)ā Gā²ā¢(Sā¢Iā¢Wā¢T)=pHā²ā Rā²ā¢(WH)+(1āpHā²)ā Rā²ā¢(W).absentā superscriptsubscriptā²subscriptā 1superscriptsubscriptā²ā superscriptsubscriptā²subscriptā 1superscriptsubscriptā² =p_H Ā· G (SIW_HT)+(1-p_H )Ā· G% (SIWT)=p_H Ā· R (W_H)+(1-p_H )Ā· R% (W).= pitalic_Hā² ā Gā² ( S I Witalic_H T ) + ( 1 - pitalic_Hā² ) ā Gā² ( S I W T ) = pitalic_Hā² ā Rā² ( Witalic_H ) + ( 1 - pitalic_Hā² ) ā Rā² ( W ) . Thus, if Rā²ā¢(W)=pHā²pHā²ā1ā¢Rā²ā¢(WH)superscriptā²subscriptā²subscriptā²1superscriptā²subscriptR (W)= p_H p_H -1R (W_H)Rā² ( W ) = divide start_ARG pitalic_Hā² end_ARG start_ARG pitalic_Hā² - 1 end_ARG Rā² ( Witalic_H ), then Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī, meaning that R+Rā²+R R + Rā² has the same choice probabilities as R and is thus fully feedback-compatible. In particular, if Rā²ā¢(WH)ā«0much-greater-thansuperscriptā²subscript0R (W_H) 0Rā² ( Witalic_H ) ā« 0 is sufficiently large, then in subsequent policy optimization, there is an incentive to hide the mistakes and ĻHsubscript _HĻitalic_H will be selected, which is suboptimal with respect to the true reward function R. Thus Fig. 4A still retains dangerous ambiguity when modeling partial observability. However, the example in Fig. 4B leads to no ambiguity when partial observability is correctly modeled. To show this in detail, let Gā²=ā”(Rā²)ākerā”ā©imā”superscriptā²kernelimG = (R )ā % B Gā² = Ī ( Rā² ) ā ker B ā© im Ī. We need to show Gā²=0superscriptā²0G =0Gā² = 0. Since the human is only uncertain about the state sequences corresponding to the observation sequence oā ā¢oIā¢oā ā¢oā subscriptsubscriptsubscriptsubscripto_ o_Io_ o_ oā oitalic_I oā oā , the condition ā Gā²=0ā superscriptā²0 BĀ· G =0B ā Gā² = 0 already implies Gā²ā¢(sā)=0superscriptā²ā0G ( s)=0Gā² ( overā start_ARG s end_ARG ) = 0 for all state sequences except Sā¢Iā¢Wā¢TSIWTS I W T and Sā¢Iā¢Tā¢TSITTS I T T. From (ā Gā²)ā¢(oā ā¢oIā¢oā ā¢oā )=0ā superscriptā²subscriptsubscriptsubscriptsubscript0( BĀ· G )(o_ o_Io_ o_% )=0( B ā Gā² ) ( oā oitalic_I oā oā ) = 0, one then obtains the equation (1āpD)ā (Rā²ā¢(S)+Rā²ā¢(I)+2ā¢Rā²ā¢(T))+pDā (Rā²ā¢(S)+Rā²ā¢(I)+Rā²ā¢(W)+Rā²ā¢(T))=0.ā 1subscriptsuperscriptā²2superscriptā²ā subscriptsuperscriptā²superscriptā²0(1-p_D)Ā· (R (S)+R (I)+2R (T) )+p_D% Ā· (R (S)+R (I)+R (W)+R (T) )=0.( 1 - pitalic_D ) ā ( Rā² ( S ) + Rā² ( I ) + 2 Rā² ( T ) ) + pitalic_D ā ( Rā² ( S ) + Rā² ( I ) + Rā² ( W ) + Rā² ( T ) ) = 0 . (5) Thus, if one of the two state sequences involved has zero return, then the other has as well, assuming that 0ā pDā 10subscript10ā p_Dā 10 ā pitalic_D ā 1, and we are done. To show this, we use that all other state sequences have zero return: Rā²ā¢(S)+3ā¢Rā²ā¢(T)=0=Rā²ā¢(S)+Rā²ā¢(L)+2ā¢Rā²ā¢(T)superscriptā²3superscriptā²0superscriptā²2superscriptā²R (S)+3R (T)=0=R (S)+R (L)+2R (T)Rā² ( S ) + 3 Rā² ( T ) = 0 = Rā² ( S ) + Rā² ( L ) + 2 Rā² ( T ), from which Rā²ā¢(L)=Rā²ā¢(T)superscriptā²R (L)=R (T)Rā² ( L ) = Rā² ( T ) follows. Then, from Rā²ā¢(S)+Rā²ā¢(I)+Rā²ā¢(L)+Rā²ā¢(T)=0superscriptā²superscriptā²0R (S)+R (I)+R (L)+R (T)=0Rā² ( S ) + Rā² ( I ) + Rā² ( L ) + Rā² ( T ) = 0, substituting the previous result gives Rā²ā¢(S)+Rā²ā¢(I)+2ā¢Rā²ā¢(T)=0superscriptā²2superscriptā²0R (S)+R (I)+2R (T)=0Rā² ( S ) + Rā² ( I ) + 2 Rā² ( T ) = 0, and so Equation (5) results in Rā²ā¢(S)+Rā²ā¢(I)+Rā²ā¢(W)+Rā²ā¢(T)=0superscriptā²superscriptā²0R (S)+R (I)+R (W)+R (T)=0Rā² ( S ) + Rā² ( I ) + Rā² ( W ) + Rā² ( T ) = 0. Overall, this shows Gā²=ā”(Rā²)=0superscriptā²0G = (R )=0Gā² = Ī ( Rā² ) = 0, and so kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . B.5 Experimental details Here, we explain more experimental details for the results in Table 1, reproduced here as Table 13, and Figure 5. Table 13: Experiments showing improved performance of po-aware RLHF Ex. p phidesubscripthidep_hidephide pdefaultsubscriptdefaultp_defaultpdefault model action EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ dec. infl. EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- overj. optimal A 0.5 0.5 N/A naive aHsubscripta_Haitalic_H 1.5 ā 0 Ć Ć A 0.5 0.5 N/A po-aware aHsubscripta_Haitalic_H 1.5 ā 0 Ć Ć A 0.1 0.9 N/A naive aCsubscripta_Caitalic_C 0 Ć 0 ā Ć A 0.1 0.9 N/A po-aware aTsubscripta_Taitalic_T 0 Ć 5.4 Ć ā B 0.5 N/A 0.9 naive aTsubscripta_Taitalic_T 4.5 ā 0 ā Ć B 0.5 N/A 0.9 po-aware aDsubscripta_Daitalic_D 0 Ć 0.25 Ć ā B 0.5 N/A 0.1 naive aVsubscripta_Vaitalic_V 0 Ć 0 ā Ć B 0.5 N/A 0.1 po-aware aDsubscripta_Daitalic_D 0 Ć 2.25 Ć ā The leftmost column (āEx.ā for āexampleā) corresponds to Examples A and B in Figure 4. p is the success probability upon attempting to install Cuda or NumPy in state I, see Figure 7. phidesubscripthidep_hidephide in Example A is the humanās belief probability that the agent hid the error message if there is no output after nvidia-driver installation. Similarly, pdefaultsubscriptdefaultp_defaultpdefault in Example B is the humanās belief probability that installation was done with default settings if there is no further output after Python installation. Note that lines one and two in the table also correspond to Example B.1, lines three and four to Example B.2, and lines five and six to the first half and seven and eight to the second half of Example B.3, respectively. In all the results in the table, we set the penalty to r=11r=1r = 1. The āmodelā column has value ānaiveā if the reward learning algorithm is classical RLHF (erroneously assuming full observability) as in Christiano et al. [2017], and āpo-awareā if the humanās partial observability is correctly modeled as in Section C.3. We initialize the reward function as a list of rewards of states and train it by logistic regression using a dataset that consists of all pairs of state sequences together with the humanās choice probabilities under partial observations. This leads to 28 pairs of distinct trajectories together with choice probabilities. We train the reward model for 300 epochs over a shuffled dataset of 13.5 copies of the 28 pairs with the Adam optimizer, for a total of 113400 training updates. Once we have the resulting reward model, we use value iteration to find its deterministic optimal policy. All policies choose to install the nvidia-driver (in Example A) and Python (in Example B), and differ in their action in state I, which is given in the column āactionā. We compute the overestimation error and underestimation error of the resulting policies analytically using the hardcoded environment dynamics, true reward function, observation function, and human belief matrix BB. This is given in columns EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ and EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG-. Note that these are averages over 10 entire training runs, though since they always result in the same learned policy, there is no variation and we do not state any uncertainty. The columns ādec. infl.ā, āoverj.ā, and āoptimalā state whether deceptive inflation or overjustification occurs with the learned policy, and whether it is optimal according to the true humanās reward function. For the results in Figure 5, we use largely the same procedure as for the table. Instead of fixing the reward penalty r or the belief probabilities phidesubscripthidep_hidephide and pdefaultsubscriptdefaultp_defaultpdefault, we vary them as hyperparameters for the plots, we fix p to p=0.50.5p=0.5p = 0.5, and we restrict ourselves to the analysis of ānaiveā RLHF. Appendix C Modeling the Human in Partially Observable RLHF In this appendix, we develop the theory of RLHF with appropriately modeled partial observability, including full proofs of all theorems. In Section C.1, we explain how the human can arrive at the belief Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) via Bayesian updates. The main theory and the main paper in general do not depend on this specific form of the humanās belief, but some examples in the appendix do. In Section C.2 we then explain our main result: the ambiguity and identifiability of both reward and return functions under observed sequence comparisons. In Section C.3, we then explain that this theorem means that one could in principle design a practical reward learning algorithm that converges on the correct reward function up to the ambiguity characterized in the section before, if the humanās belief kernel Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) is fully known. In Section C.4, we generalize the theory to the case that the humanās observations are not necessarily known to the learning system and again characterize precisely when the return function is identifiable from sequence comparisons. We then consider special cases in Section C.5, where we show that the fully observable case is covered by our theory, that a deterministic observation kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG usually leads to non-injective belief matrix BB, and that ānoiseā in the observation kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG leads, under appropriate assumptions, to the identifiability of the return function. Our identifiability results require that the learning system knows the humanās belief kernel Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ). In Section C.6, we then show that these results are robust to slight misspecifications: a bound in the error in the specified belief leads to a corresponding bound in the error of the policy evaluation function used for subsequent reinforcement learning. In Section C.7, we then provide a very preliminary characterization of the ambiguity in the return function under special cases. Finally, in Section C.8, we study examples of identifiability and non-identifiability of the return function for the case that we do model the humanās partial observability correctly. This reveals qualitatively interesting cases of identifiability, even when BB is not injective, and catastrophic cases of non-identifiability. C.1 The Belief over the State Sequence for Rational Humans Before we dive into the main theory, we want to explain how the human can iteratively compute the posterior of the state sequence given an observation sequence with successively new observations. This is done by defining a Bayesian network for the joint probability of policy, states, actions, and observations, and doing Bayesian inference over this Bayesian network. The details of this subsection are only relevant for a few sections in the appendix since it is usually enough to assume that the posterior belief exists. Additionally, in the core theory, we do not even assume that Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) is a posterior: it is simply any probability distribution. The reason why it can still be interesting to analyze the case when the human is a rational Bayesian reasoner is that one can then analyze RLHF under generous assumptions to the human. We model the human to have a joint distribution Bā¢(Ļ,sā,aā,oā)āB(Ļ, s, a, o)B ( Ļ , overā start_ARG s end_ARG , overā start_ARG a end_ARG , overā start_ARG o end_ARG ) over the policy Ļ, state sequence sā=s0,ā¦,sTāsubscript0ā¦subscript s=s_0,ā¦,s_Toverā start_ARG s end_ARG = s0 , ⦠, sitalic_T, action sequence aā=a0,ā¦,aTā1āsubscript0ā¦subscript1 a=a_0,ā¦,a_T-1overā start_ARG a end_ARG = a0 , ⦠, aitalic_T - 1, and observation sequence oā=o0,ā¦,oTāsubscript0ā¦subscript o=o_0,ā¦,o_Toverā start_ARG o end_ARG = o0 , ⦠, oitalic_T. This is given by a Bayesian network with the following components: ⢠a policy prior Bā¢(Ļā²)superscriptā²B(Ļ )B ( Ļā² ); ⢠the probability of the initial state Bā¢(s0)āP0ā¢(s0)āsubscript0subscript0subscript0B(s_0) P_0(s_0)B ( s0 ) ā P0 ( s0 ); ⢠action probabilities Bā¢(aā£s,Ļ)āĻā¢(aā£s)āconditionalconditionalB(a s,Ļ) Ļ(a s)B ( a ⣠s , Ļ ) ā Ļ ( a ⣠s ); ⢠transition probabilities Bā¢(st+1ā£st,at)āā¢(st+1ā£st,at)āconditionalsubscript1subscriptsubscriptconditionalsubscript1subscriptsubscriptB(s_t+1 s_t,a_t) (s_t+1 s_t,a_t)B ( sitalic_t + 1 ⣠sitalic_t , aitalic_t ) ā T ( sitalic_t + 1 ⣠sitalic_t , aitalic_t ); ⢠and observation probabilities Bā¢(otā£st)āPOā¢(otā£st)āconditionalsubscriptsubscriptsubscriptconditionalsubscriptsubscriptB(o_t s_t) P_O(o_t s_t)B ( oitalic_t ⣠sitalic_t ) ā Pitalic_O ( oitalic_t ⣠sitalic_t ). Together, this defines the joint distribution Bā¢(Ļ,sā,aā,oā)āB(Ļ, s, a, o)B ( Ļ , overā start_ARG s end_ARG , overā start_ARG a end_ARG , overā start_ARG o end_ARG ) over the policy, states, actions, and observations that factorizes according to the following directed acyclic graph: Ļā²Ļ Ļā²s0subscript0s_0s0a0subscript0a_0a0s1subscript1s_1s1a1subscript1a_1a1s2subscript2s_2s2a2subscript2a_2a2s3subscript3s_3s3ā¦o0subscript0o_0o0o1subscript1o_1o1o2subscript2o_2o2o3subscript3o_3o3 (6) The following proposition clarifies the iterative Bayesian update of the humanās posterior over state sequences, given observation sequences: Proposition C.1. Let tā¤Tā11t⤠T-1t ⤠T - 1 and denote by s^=s0,ā¦,st^subscript0ā¦subscript s=s_0,ā¦,s_tover start_ARG s end_ARG = s0 , ⦠, sitalic_t a state sequence segment of length tā„00tā„ 0t ā„ 0. Similarly, o^=o0,ā¦,ot^subscript0ā¦subscript o=o_0,ā¦,o_tover start_ARG o end_ARG = o0 , ⦠, oitalic_t denotes an observation sequence segment. We have Bā¢(s^,st+1,Ļā£o^,ot+1)āPOā¢(ot+1ā£st+1)ā [āatāā¢(st+1ā£s^t,at)ā Ļā¢(atā£st)]ā Bā¢(s^,Ļā£o^).proportional-to^subscript1conditional^subscript1ā subscriptconditionalsubscript1subscript1delimited-[]subscriptsubscriptā conditionalsubscript1subscript^subscriptconditionalsubscriptsubscript^conditional^B( s,s_t+1,Ļ o,o_t+1) P_O(o_t+1 s_t+1)% Ā· [ _a_t T(s_t+1 s_t,a_t% )Ā·Ļ(a_t s_t) ]Ā· B( s,Ļ o).B ( over start_ARG s end_ARG , sitalic_t + 1 , Ļ ā£ over start_ARG o end_ARG , oitalic_t + 1 ) ā Pitalic_O ( oitalic_t + 1 ⣠sitalic_t + 1 ) ā [ āa start_POSTSUBSCRIPT t ā A end_POSTSUBSCRIPT T ( sitalic_t + 1 ⣠over start_ARG s end_ARGt , aitalic_t ) ā Ļ ( aitalic_t ⣠sitalic_t ) ] ā B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) . Thus, the human can iteratively compute Bā¢(s^,Ļā£o^)^conditional^B( s,Ļ o)B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) from the prior Bā¢(s0,Ļ)=P0ā¢(s0)ā Bā¢(Ļā²)subscript0ā subscript0subscript0superscriptā²B(s_0,Ļ)=P_0(s_0)Ā· B(Ļ )B ( s0 , Ļ ) = P0 ( s0 ) ā B ( Ļā² ) using the above Bayesian update. The posterior over the state sequence can subsequently be computed by Bā¢(s^ā£o^)=ā«ĻBā¢(s^,Ļā£o^).conditional^^subscript^conditional^B( s o)= _ĻB( s,Ļ o).B ( over start_ARG s end_ARG ⣠over start_ARG o end_ARG ) = ā«Ļ B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) . Proof. The proof is essentially just Bayes rule applied to the Bayesian network in Equation (6). We repeatedly make use of conditional independences that follow from d-separations in the graph [Geiger et al., 1990]. More concretely, we have Bā¢(s^,st+1,Ļā£o^,ot+1)^subscript1conditional^subscript1 B ( s,s_t+1,Ļ o,o_t+1 )B ( over start_ARG s end_ARG , sitalic_t + 1 , Ļ ā£ over start_ARG o end_ARG , oitalic_t + 1 ) āBā¢(ot+1ā£s^,st+1,Ļ,o^)ā Bā¢(s^,st+1,Ļā£o^)proportional-toabsentā conditionalsubscript1^subscript1^^subscript1conditional B (o_t+1 s,s_t+1,Ļ, o )% Ā· B ( s,s_t+1,Ļ o )ā B ( oitalic_t + 1 ⣠over start_ARG s end_ARG , sitalic_t + 1 , Ļ , over start_ARG o end_ARG ) ā B ( over start_ARG s end_ARG , sitalic_t + 1 , Ļ ā£ over start_ARG o end_ARG ) =POā¢(ot+1ā£st+1)ā Bā¢(st+1ā£s^,Ļ,o^)ā Bā¢(s^,Ļā£o^)absentā subscriptconditionalsubscript1subscript1conditionalsubscript1^^^conditional =P_O (o_t+1 s_t+1 )Ā· B (s_t+1 % s,Ļ, o)Ā· B( s,Ļ o )= Pitalic_O ( oitalic_t + 1 ⣠sitalic_t + 1 ) ā B ( sitalic_t + 1 ⣠over start_ARG s end_ARG , Ļ , over start_ARG o end_ARG ) ā B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) =POā¢(ot+1ā£st+1)ā [āatāBā¢(st+1ā£at,s^,Ļ,o^)ā Bā¢(atā£s^,Ļ,o^)]ā Bā¢(s^,Ļā£o^)absentā subscriptconditionalsubscript1subscript1delimited-[]subscriptsubscriptā conditionalsubscript1subscript^^conditionalsubscript^^^conditional =P_O (o_t+1 s_t+1 )Ā· [ _a_t% B (s_t+1 a_t, s,Ļ, o )Ā· B % (a_t s,Ļ, o ) ]Ā· B ( s,Ļ % o )= Pitalic_O ( oitalic_t + 1 ⣠sitalic_t + 1 ) ā [ āa start_POSTSUBSCRIPT t ā A end_POSTSUBSCRIPT B ( sitalic_t + 1 ⣠aitalic_t , over start_ARG s end_ARG , Ļ , over start_ARG o end_ARG ) ā B ( aitalic_t ⣠over start_ARG s end_ARG , Ļ , over start_ARG o end_ARG ) ] ā B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) =POā¢(ot+1ā£st+1)ā [āatāā¢(st+1ā£st,at)ā Ļā¢(atā£st)]ā Bā¢(s^,Ļā£o^).absentā subscriptconditionalsubscript1subscript1delimited-[]subscriptsubscriptā conditionalsubscript1subscriptsubscriptconditionalsubscriptsubscript^conditional =P_O (o_t+1 s_t+1 )Ā· [ _a_t% T (s_t+1 s_t,a_t )Ā·Ļ (% a_t s_t ) ]Ā· B ( s,Ļ o ).= Pitalic_O ( oitalic_t + 1 ⣠sitalic_t + 1 ) ā [ āa start_POSTSUBSCRIPT t ā A end_POSTSUBSCRIPT T ( sitalic_t + 1 ⣠sitalic_t , aitalic_t ) ā Ļ ( aitalic_t ⣠sitalic_t ) ] ā B ( over start_ARG s end_ARG , Ļ ā£ over start_ARG o end_ARG ) . In step 1, we used Bayes rule. In step 2, we made use of the independence ot+1ā(s^,Ļ,o^)|st+1o_t+1 \!\!\! ( s,Ļ, o)\ |\ s_t+1oitalic_t + 1 ā ā ( over start_ARG s end_ARG , Ļ , over start_ARG o end_ARG ) | sitalic_t + 1, plugged in the observation kernel, and used the chain rule of probability to compose the second term into a product. In step 3, we marginalized and used, once again, the chain rule of probability. In step 4444, we used the independences st+1ā(s0,ā¦,stā1,Ļ,o^)|(st,a)s_t+1 \!\!\! (s_0,ā¦,s_t-1,Ļ, o)\ |\ (s_t,a)sitalic_t + 1 ā ā ( s0 , ⦠, sitalic_t - 1 , Ļ , over start_ARG o end_ARG ) | ( sitalic_t , a ) and atā(s0,ā¦,stā1,o^)ā£(Ļ,st)a_t \!\!\! (s_0,ā¦,s_t-1, o) (Ļ,s_t)aitalic_t ā ā ( s0 , ⦠, sitalic_t - 1 , over start_ARG o end_ARG ) ⣠( Ļ , sitalic_t ) and plugged in the transition kernel and the policy. The last formula is just a marginalization over the policy. ā C.2 Ambiguity and Identifiability of Reward and Return Functions under Observation Sequence Comparisons Figure 9: The linear geometry of ambiguity for a hypothetical example with three state sequences and two observation sequences. GāsuperscriptG^*Gā is the true return function, and āGā is used in labeling the axes to refer to some arbitrary return function. This is a more accurate geometric depiction of the middle and right spaces in Figure 6. The subspace imā”ā©kerā”imkernelim ā© Bim Ī ā© ker B (purple) is the ambiguity in return functions, meaning that adding an element would not change the humanās expected return function on observations. Thus the set of return functions that the reward learning system can infer is the affine set G+(imā”ā©kerā”)imkernelG+(im ā© % B)G + ( im Ī ā© ker B ) (yellow). Note that the planes on the left are drawn to be axis-aligned for ease of visualization; this will not be the case for real MDPs. In this section, we prove the main theorem of this paper: a characterization of the ambiguity that is left in the reward and return function once the humanās Boltzmann-rational choice probabilities are known. We change the formulation slightly by formulating the linear operators āintrinsicallyā in the spaces they are defined in, instead of using matrix versions. This does not change the general picture, but is a more natural setting when thinking, e.g., about generalizing the results to infinite state sequences. Thus, we define :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© B: R Sā R % B : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG as the linear operator given by [ā”(G)]ā¢(oā)āsāā¼Bā¢(sāā£oā)[Gā¢(sā)].ādelimited-[]āsubscriptsimilar-toāconditionalā [ B(G) ]( o) *% E_ s B( s o) [G( s) ].[ B ( G ) ] ( overā start_ARG o end_ARG ) ā Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ G ( overā start_ARG s end_ARG ) ] . Here, BB is the humanās belief, which can either be computed as in the previous subsection or simply be any conditional probability distribution. Similarly, we define :āāāā:āsuperscriptāsuperscriptāā : R^Sā R % SĪ : blackboard_RS ā blackboard_Roverā start_ARG S end_ARG as the linear operator given by [ā”(R)]ā¢(sā)āāt=0Tγtā¢Rā¢(st).ādelimited-[]āsuperscriptsubscript0superscriptsubscript [ (R) ]( s) _t=0^T% γ^tR(s_t).[ Ī ( R ) ] ( overā start_ARG s end_ARG ) ā āt = 0T γitalic_t R ( sitalic_t ) . The matrix product ā BĀ· B ā Ī then becomes the composition ā:āāāĪ©ā:āsuperscriptāsuperscriptāāĪ© B : R^% Sā R B ā Ī : blackboard_RS ā blackboard_Roverā start_ARG Ī© end_ARG. Finally, recall that the kernel kerā”kernel Aker A of a linear operator AA is defined as its nullspace, and the image imā”imim Aim A as the set of elements hit by AA. We obtain the following theorem: Theorem C.2. Let R be the true reward function and R~~ Rover~ start_ARG R end_ARG another reward function. Let G~=ā”(R~)~~ G= ( R)over~ start_ARG G end_ARG = Ī ( over~ start_ARG R end_ARG ) and G=ā”(R)G= (R)G = Ī ( R ) be the corresponding return functions. The following three statements are equivalent: (i) The reward function R~~ Rover~ start_ARG R end_ARG gives rise to the same vector of choice probabilities as R, i.e (PR~ā¢(oāā»oāā²))oā,oāā²āĪ©ā=(PRā¢(oāā»oāā²))oā,oāā²āĪ©ā.subscriptsuperscript~succeedsāsuperscriptāā²āsuperscriptāā²āĪ©subscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P R ( o o 1.0pt ) % )_ o, o 1.0pt ā = (P^R (% o o 1.0pt ) )_ o, o 1% .0pt ā .( Pover~ start_ARG R end_ARG ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG = ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG . (i) There is a reward function Rā²ākerā”(ā)superscriptā²kernelR ā ( B )Rā² ā ker ( B ā Ī ) and a constant cāācā Rc ā blackboard_R such that R~=R+Rā²+c.~superscriptā² R=R+R +c.over~ start_ARG R end_ARG = R + Rā² + c . (i) There is a return function Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī and a constant cā²āāsuperscriptā²āc ā Rcā² ā blackboard_R such that G~=G+Gā²+cā².~superscriptā² G=G+G +c .over~ start_ARG G end_ARG = G + Gā² + cā² . In other words, the ambiguity that is left in the reward function when its observation-based choice probabilities are known is, up to an additive constant, given by kerā”(ā)kernel ( B )ker ( B ā Ī ); the ambiguity left in the return function is given by kerā”ā©imā”kernelim B ker B ā© im Ī. Proof. Assume (i). To prove (i), let Ļ by the sigmoid function given by Ļā¢(x)=11+expā”(āx)11Ļ(x)= 11+ (-x)Ļ ( x ) = divide start_ARG 1 end_ARG start_ARG 1 + exp ( - x ) end_ARG. Then by Equation (2), the equality of choice probabilities means the following for all oā,oāā²āĪ©āsuperscriptāā²āĪ© o, o 1.0pt ā overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG: Ļā¢(βā ([ā”(G~)]ā¢(oā)ā[ā”(G~)]ā¢(oāā²)))=Ļā¢(βā ([ā”(G)]ā¢(oā)ā[ā”(G)]ā¢(oāā²))).ā delimited-[]~ādelimited-[]~superscriptāā²ā delimited-[]ādelimited-[]superscriptāā²Ļ (β· ( [ B( G) % ]( o)- [ B( G) ]( o 1% .0pt ) ) )=Ļ (β· ( [% B(G) ]( o)- [ B(% G) ]( o 1.0pt ) ) ).Ļ ( β ā ( [ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARG ) - [ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARGā² ) ) ) = Ļ ( β ā ( [ B ( G ) ] ( overā start_ARG o end_ARG ) - [ B ( G ) ] ( overā start_ARG o end_ARGā² ) ) ) . Since the sigmoid function is injective, this implies [ā”(G~)]ā¢(oā)ā[ā”(G~)]ā¢(oāā²)=[ā”(G)]ā¢(oā)ā[ā”(G)]ā¢(oāā²).delimited-[]~ādelimited-[]~superscriptāā²delimited-[]ādelimited-[]superscriptāā² [ B( G) ]( o)- [% B( G) ]( o 1.0pt )= % [ B(G) ]( o)- [ B% (G) ]( o 1.0pt ).[ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARG ) - [ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARGā² ) = [ B ( G ) ] ( overā start_ARG o end_ARG ) - [ B ( G ) ] ( overā start_ARG o end_ARGā² ) . Fixing an arbitrary oāā² o 1.0pt overā start_ARG o end_ARGā², this implies that there exists a constant cā² such that for all oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG, the following holds: [ā”(G~)]ā¢(oā)ā[ā”(G)]ā¢(oāā²)ācā²=0.delimited-[]~ādelimited-[]superscriptāā²0 [ B( G) ]( o)- [% B(G) ]( o 1.0pt )-c =0.[ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARG ) - [ B ( G ) ] ( overā start_ARG o end_ARGā² ) - cā² = 0 . Noting that ā”(cā²)=cā²superscriptā² B(c )=c B ( cā² ) = cā², this implies G~āGācā²ākerā”()~superscriptā²kernel G-G-c ā ( B)over~ start_ARG G end_ARG - G - cā² ā ker ( B ). Now, define the constant reward function cācā²ā 1āγ1āγT+1.āā superscriptā²11superscript1c c Ā· 1-γ1-γ^T+1.c ā cā² ā divide start_ARG 1 - γ end_ARG start_ARG 1 - γitalic_T + 1 end_ARG . We obtain [ā”(c)]ā¢(sā)delimited-[]ā [ (c) ]( s)[ Ī ( c ) ] ( overā start_ARG s end_ARG ) =āt=0Tγtā cabsentsuperscriptsubscript0ā superscript = _t=0^Tγ^tĀ· c= āt = 0T γitalic_t ā c =cā²ā 1āγ1āγT+1ā āt=0Tγtabsentā superscriptā²11superscript1superscriptsubscript0superscript =c Ā· 1-γ1-γ^T+1Ā· _t=0^T% γ^t= cā² ā divide start_ARG 1 - γ end_ARG start_ARG 1 - γitalic_T + 1 end_ARG ā āt = 0T γitalic_t =cā².absentsuperscriptā² =c .= cā² . Thus, we have ā”(R~āRāc)=G~āGācā²ākerā”(),~~superscriptā²kernel ( R-R-c)= G-G-c ā (% B),Ī ( over~ start_ARG R end_ARG - R - c ) = over~ start_ARG G end_ARG - G - cā² ā ker ( B ) , implying Rā²āR~āRācākerā”(ā)āsuperscriptā²~kernelR R-R-cā ( B % )Rā² ā over~ start_ARG R end_ARG - R - c ā ker ( B ā Ī ). This shows (i). That (i) implies (i) follows by applying Ī to both sides of the equation. Now assume (i), i.e. G~=G+Gā²+cā²~superscriptā² G=G+G +c over~ start_ARG G end_ARG = G + Gā² + cā² for a constant cā²āāsuperscriptā²āc ā Rcā² ā blackboard_R and a return function Gā²ākerā”()ā©imā”superscriptā²kernelimG ā ( B) % Gā² ā ker ( B ) ā© im Ī. This implies ā”(G~)=ā”(G)+cā²~superscriptā² B( G)= B(G)+c B ( over~ start_ARG G end_ARG ) = B ( G ) + cā². Thus, for all oā,oāā²āĪ©āsuperscriptāā²āĪ© o, o 1.0pt ā overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG, we have [ā”(G~)]ā¢(oā)ā[ā”(G~)]ā¢(oāā²)=[ā”(G)]ā¢(oā)ā[ā”(G)]ā¢(oāā²),delimited-[]~ādelimited-[]~superscriptāā²delimited-[]ādelimited-[]superscriptāā² [ B( G) ]( o)- [% B( G) ]( o 1.0pt )= % [ B(G) ]( o)- [ B% (G) ]( o 1.0pt ),[ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARG ) - [ B ( over~ start_ARG G end_ARG ) ] ( overā start_ARG o end_ARGā² ) = [ B ( G ) ] ( overā start_ARG o end_ARG ) - [ B ( G ) ] ( overā start_ARG o end_ARGā² ) , which implies the equal choice probabilities after multiplying with β and applying the sigmoid function Ļ on both sides. Thus, (i) implies (i). ā Corollary C.3. The following two statements are equivalent: (i) kerā”(ā)=0kernel0 ( B )=0ker ( B ā Ī ) = 0. (i) The data (PRā¢(oāā»oāā²))oā,oāā²āĪ©āsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P^R ( o o 1.0pt ) )_% o, o 1.0pt ā ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG determine the reward function R up to an additive constant. Proof. That (i) implies (i) follows immediately from the implication from (i) to (i) within the preceding theorem. Now assume (i). Let Rā²ākerā”(ā)superscriptā²kernelR ā ( B )Rā² ā ker ( B ā Ī ). Define R~āR+Rā²ā~superscriptā² R R+R over~ start_ARG R end_ARG ā R + Rā². Then the implication from (i) to (i) within the preceding theorem implies that R~~ Rover~ start_ARG R end_ARG and R have the same choice probabilities. Thus, the assumption (i) in this corollary implies that Rā² is a constant. Since Ī and BB map nonzero constants to nonzero constants, the fact that Rā²ākerā”(ā)superscriptā²kernelR ā ( B )Rā² ā ker ( B ā Ī ) implies that Rā²=0superscriptā²0R =0Rā² = 0, showing that kerā”(ā)=0kernel0 ( B )=\0\ker ( B ā Ī ) = 0 . ā As mentioned in the main paper, the previous result already leads to the non-identifiability of R whenever Ī is not injective, corresponding to the presence of zero-initial potential shaping (Skalse et al. [2023], Lemma B.3). Thus, we now strengthen the previous result so that it deals with the identifiability of the return function, which is sufficient for the purpose of policy optimization: Corollary C.4. Consider the following four statements (which can each be true or false): (i) kerā”=0kernel0 B=\0\ker B = 0 . (i) kerā”(ā)=0kernel0 ( B )=\0\ker ( B ā Ī ) = 0 . (i) kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . (iv) The data (PRā¢(oāā»oāā²))oā,oāā²āĪ©āsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P^R ( o o 1.0pt ) )_% o, o 1.0pt ā ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG determine the return function G=ā”(R)G= (R)G = Ī ( R ) on sequences sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG up to a constant independent of sā soverā start_ARG s end_ARG. Then the following implications, and no other implications, are true: (i)(i)( i )(iā¢iā¢i)(i)( i i i )(iā¢v)(iv)( i v )(iā¢i)(i)( i i ) In particular, all of (i), (i), and (i) are sufficient conditions for identifying the return function from the choice probabilities. Proof. That (i) implies (i) is trivial. That (i) implies (i) is a simple linear algebra fact: Assume (i) and that Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī. Then Gā²=ā”(Rā²)superscriptā²G = (R )Gā² = Ī ( Rā² ) for some Rā²āāsuperscriptā²āR ā R^SRā² ā blackboard_RS and 0=ā”(Gā²)=ā”(ā”(Rā²))=(ā)ā¢(Rā²).0superscriptā²superscriptā²0= B(G )= B (% (R ) )=( B% )(R ).0 = B ( Gā² ) = B ( Ī ( Rā² ) ) = ( B ā Ī ) ( Rā² ) . By (i), this implies Rā²=0superscriptā²0R =0Rā² = 0 and therefore Gā²=ā”(Rā²)=0superscriptā²0G = (R )=0Gā² = Ī ( Rā² ) = 0, showing (i). That (i) implies (iv) immediately follows from the implication from (i) to (i) in Theorem C.2. Now, assume (iv). To prove (i), assume Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī. Then the implication from (i) to (i) in Theorem C.2 implies that G+Gā²+G G + Gā² induces the same observation-based choice probabilities as G. Thus, (iv) implies G+Gā²=G+cā²superscriptā²G+G =G+c G + Gā² = G + cā² for some constant cā², which implies Gā²=cā²superscriptā²G =c Gā² = cā². Since Gā²ākerā”superscriptā²kernelG ā BGā² ā ker B, this implies 0=ā”(Gā²)=ā”(cā²)=cā²0superscriptā²superscriptā²0= B(G )= B(c )=% c 0 = B ( Gā² ) = B ( cā² ) = cā² and thus Gā²=0superscriptā²0G =0Gā² = 0. Thus, we showed kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . We now show that no other implication holds in general. Example C.32 will show that (i) does not imply (i). We now show that (i) does also not imply (i), from which it will logically follow that (i) does neither imply (i) nor (i). Namely, consider the following simple MDP with time horizon T=11T=1T = 1: aaabbb (7) In this MDP, every state sequence starts in a, deterministically transitions to b, and then ends. This means that sā=aā¢bā s=aboverā start_ARG s end_ARG = a b is the only sequence. Now, let Rā²āāa,bsuperscriptā²āR ā R^\a,b\Rā² ā blackboard_R a , b be the reward function given by Rā²ā¢(a)=1,Rā²ā¢(b)=ā1γ.formulae-sequencesuperscriptā²1superscriptā²1R (a)=1, R (b)= -1γ.Rā² ( a ) = 1 , Rā² ( b ) = divide start_ARG - 1 end_ARG start_ARG γ end_ARG . We obtain [ā”(Rā²)]ā¢(sā)=Rā²ā¢(a)+γā¢Rā²ā¢(b)=1+γā ā1γ=0.delimited-[]superscriptā²āsuperscriptā²1ā 10 [ (R ) ]( s)=R (a% )+γ R (b)=1+γ· -1γ=0.[ Ī ( Rā² ) ] ( overā start_ARG s end_ARG ) = Rā² ( a ) + γ Rā² ( b ) = 1 + γ ā divide start_ARG - 1 end_ARG start_ARG γ end_ARG = 0 . Thus, ā”(Rā²)=0superscriptā²0 (R )=0Ī ( Rā² ) = 0, (ā)ā¢(Rā²)=0superscriptā²0( B )(R )=0( B ā Ī ) ( Rā² ) = 0, and, therefore, kerā”(ā)ā 0kernel0 ( B )% ā \0\ker ( B ā Ī ) ā 0 . Thus, (i) does not hold. However, it is possible to choose Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) such that (i) holds: e.g., if Ī©=Ī© =SĪ© = S and Bā¢(sāā£oā)āĪ“oāā¢(sā)āconditionalāsubscriptāB( s o) _ o( s)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) ā Ī“overā start_ARG o end_ARG ( overā start_ARG s end_ARG ), then kerā”=0kernel0 B=\0\ker B = 0 since this operator is the identity. ā C.3 The Ambiguity in Reward Learning in Practice In this section, we point out that Theorem C.2 is not just a theoretical discussion: When BB and the inverse temperature parameter β are known, then it is possible to design a reward learning algorithm that learns the true reward function up to the ambiguity kerā”(ā)kernel ( B )ker ( B ā Ī ) in the infinite data limit. In doing so, we essentially use the loss function proposed in Christiano et al. [2017]. Namely, assume DD is a data distribution of observation sequences oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG such that all sequences in Ī©āĪ© overā start_ARG Ī© end_ARG have a strictly positive probability of being sampled; for example, DD could use an exploration policy and the observation sequence kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG. For each pair of observation sequences (oā,oāā²)āsuperscriptāā²( o, o 1.0pt )( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ), we then get a conditional distribution Pā¢(μā£oā,oāā²)conditionalāsuperscriptāā²P(μ o, o 1.0pt )P ( μ ⣠overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) over a one-hot encoded human choice μā(1,0),(0,1)1001μā\(1,0),(0,1)\μ ā ( 1 , 0 ) , ( 0 , 1 ) , with probability Pā¢(μ=(1,0)ā£oā,oāā²)=PRā¢(oāā»oāā²).conditional10āsuperscriptāā²succeedsāsuperscriptāā²P (μ=(1,0) o, o 1.0pt )=P^R (% o o 1.0pt ).P ( μ = ( 1 , 0 ) ⣠overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) = Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) . Together, this gives rise to a dataset (oā1,oā1ā²,μ1),ā¦,(oāN,oāNā²,μN)subscriptā1subscriptsuperscriptāā²1subscript1ā¦subscriptāsubscriptsuperscriptāā²subscript( o_1, o 1.0pt _1, _1),ā¦,( o_N, % o 1.0pt _N, _N)( overā start_ARG o end_ARG1 , overā start_ARG o end_ARGā²1 , μ1 ) , ⦠, ( overā start_ARG o end_ARGN , overā start_ARG o end_ARGā²N , μitalic_N ) of observation sequences plus a human choice. Now assume we learn a reward function RĪø:āā:subscriptāāR_Īø:Sā RRitalic_Īø : S ā blackboard_R that is differentiable in the parameter Īø and that can represent all possible reward functions RāāsuperscriptāRā R^SR ā blackboard_RS. Let GĪøāā”(RĪø)āsubscriptsubscriptG_Īø (R_Īø)Gitalic_Īø ā Ī ( Ritalic_Īø ) be the corresponding return function. Write μk=(μk(1),μk(2))subscriptsuperscriptsubscript1superscriptsubscript2 _k=( _k^(1), _k^(2))μitalic_k = ( μitalic_k( 1 ) , μitalic_k( 2 ) ). As in Christiano et al. [2017], we define its loss over the dataset above by ā~ā¢(Īø)=ā1Nā¢āk=1Nμk(1)ā logā”PRĪøā¢(oākā»oākā²)+μk(2)ā logā”PRĪøā¢(oākā²ā»oāk).~ā1superscriptsubscript1ā superscriptsubscript1superscriptsubscriptsucceedssubscriptāsubscriptsuperscriptāā²ā superscriptsubscript2superscriptsubscriptsucceedssubscriptsuperscriptāā²subscriptā L(Īø)=- 1N _k=1^N _k^(1)Ā·% P^R_Īø ( o_k o 1.0pt _k % )+ _k^(2)Ā· P^R_Īø ( o 1.0pt _% k o_k ).over~ start_ARG L end_ARG ( Īø ) = - divide start_ARG 1 end_ARG start_ARG N end_ARG āk = 1N μitalic_k( 1 ) ā log Pitalic_Ritalic_Īø ( overā start_ARG o end_ARGk ā» overā start_ARG o end_ARGā²k ) + μitalic_k( 2 ) ā log Pitalic_Ritalic_Īø ( overā start_ARG o end_ARGā²k ā» overā start_ARG o end_ARGk ) . Note that by Equation (2), this loss function essentially uses BB and also the inverse temperature parameter β in its definition. This means that these need to be explicitly represented to be able to use the loss function in practice. Proposition C.5. The loss function ā~~ā Lover~ start_ARG L end_ARG is differentiable. Furthermore, in the infinite datalimit its minima are precisely given by parameters Īø such that RĪø=R+Rā²+csubscriptsuperscriptā²R_Īø=R+R +cRitalic_Īø = R + Rā² + c for Rā²ākerā”(ā)superscriptā²kernelR ā ( B % )Rā² ā ker ( B ā Ī ) and cāācā Rc ā blackboard_R, or equivalently GĪø=G+Gā²+cā²subscriptsuperscriptā²G_Īø=G+G +c Gitalic_Īø = G + Gā² + cā² for Gā²ākerā”ā©imā”superscriptā²kernelimG ā B % Gā² ā ker B ā© im Ī and cā²āāsuperscriptā²āc ā Rcā² ā blackboard_R. Proof. The differentiability of the loss function follows from the differentiability of multiplication with the matrix BB, see Equation (2), and of the reward function RĪøsubscriptR_ĪøRitalic_Īø in its parameter Īø that we assumed. For the second statement, let Nā¢(oā,oāā²)āsuperscriptāā²N( o, o 1.0pt )N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) be the number of times that the pair (oā,oāā²)āsuperscriptāā²( o, o 1.0pt )( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) appears in the dataset, and let Nā¢(oā,oāā²,1)āsuperscriptāā²1N( o, o 1.0pt ,1)N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² , 1 ) be the number of times that the human choice is μ=(1,0)10μ=(1,0)μ = ( 1 , 0 ) and the sampled pair is (oā,oāā²)āsuperscriptāā²( o, o 1.0pt )( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ), and similar for 2222 instead of 1111. We obtain ā~ā¢(Īø)=~āabsent L(Īø)=over~ start_ARG L end_ARG ( Īø ) = āāoā,oāā²āĪ©āNā¢(oā,oāā²)Nā [Nā¢(oā,oāā²,1)Nā¢(oā,oāā²)logPRĪø(oāā»oāā²) - _ o, o 1.0pt ā % N( o, o 1.0pt )NĀ· [ N( o, o% 1.0pt ,1)N( o, o 1.0pt ) P^R_% Īø ( o o 1.0pt )- āoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG divide start_ARG N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) end_ARG start_ARG N end_ARG ā [ divide start_ARG N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² , 1 ) end_ARG start_ARG N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) end_ARG log Pitalic_Ritalic_Īø ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) +Nā¢(oā,oāā²,2)Nā¢(oā,oāā²)logPRĪø(oāā²ā»oā)] + N( o, o 1.0pt ,2)N( o, % o 1.0pt ) P^R_Īø ( o 1.0pt % o ) ]+ divide start_ARG N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² , 2 ) end_ARG start_ARG N ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) end_ARG log Pitalic_Ritalic_Īø ( overā start_ARG o end_ARGā² ā» overā start_ARG o end_ARG ) ] ā ā oā,oāā²ā¼[CEā”(PRā¢(oāā»āŗoāā²)ā„PRĪøā¢(oāā»āŗoāā²))]subscriptsimilar-toāsuperscriptāā²CEconditionalsuperscriptsucceedsprecedesāsuperscriptāā²subscriptsucceedsprecedesāsuperscriptāā² *E_ o, o 1.0pt % [CE (P^R ( o % -2.15277pt$ $ 2.15277pt$ $% o 1.0pt )\ \|\ P^R_Īø ( o% -2.15277pt$ $ 2.15277% pt$ $ o 1.0pt ) ) ]Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ D [ CE ( Pitalic_R ( overā start_ARG o end_ARG start_RELOP start_ROW start_CELL ā» end_CELL end_ROW start_ROW start_CELL āŗ end_CELL end_ROW end_RELOP overā start_ARG o end_ARGā² ) ā„ Pitalic_Ritalic_Īø ( overā start_ARG o end_ARG start_RELOP start_ROW start_CELL ā» end_CELL end_ROW start_ROW start_CELL āŗ end_CELL end_ROW end_RELOP overā start_ARG o end_ARGā² ) ) ] ā ā āā¢(Īø).ā (Īø).L ( Īø ) . Here, CECECECE is the crossentropy between the two binary distributions. Since we assumed that DD gives a positive probability to all observation sequences in Ī©āĪ© overā start_ARG Ī© end_ARG, and since the cross entropy is generally minimized exactly when the second distribution equals the first, the loss function āā¢(Īø)āL(Īø)L ( Īø ) is minimized if and only if RĪøsubscriptR_ĪøRitalic_Īø gives rise to the same choice probabilities as R for all pairs of observation sequences. Theorem C.2 then gives the result. ā C.4 Identifiability of Return Functions When Human Observations Are Not Known Corollary C.4 assumes that the choice probabilities of each observation sequence pair are known to the reward learning algorithm. However, this requires the algorithm to know what the human observed. In some applications, this is a reasonable assumption, e.g. if the humanās observations are themselves produced by an algorithm that can feed the observations also back to the learning algorithm. In general, however, the observations happen in the physical world, and are only known probabilistically via the observation kernel POsubscriptP_OPitalic_O. The learning system does however have access to the full state sequences that generate the observation sequences. This leads to knowledge of the following choice probabilities for sā,sāā²āāsuperscriptāā²ā s, s 1.0pt ā Soverā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG S end_ARG: PRā¢(sāā»sāā²)āoā,oāā²ā¼POā(ā ā£sā,sāā²)[PRā¢(oāā»oāā²)],P^R ( s s 1.0pt ) % *E_ o, o P_ O(Ā· % s, s 1.0pt ) [P^R ( o o% 1.0pt ) ],Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) ā Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) ] , (8) where the observation-based choice probabilities are given as in Equation (2). In other words, the learning algorithm can only infer an aggregate of the observation-based choice probabilities. Again, we can ask a question similar to the ones before, extending the investigations in the previous section: Question C.6. Assume the vector of choice probabilities (PRā¢(sāā»sāā²))sā,sāā²āāsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²ā (P^R( s s 1.0pt ) )_ s, s% 1.0pt ā S( Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) )overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG S end_ARG is known. Additionally, assume that it is known that the humanās observations are governed by POsubscriptP_OPitalic_O, and that the human is Boltzmann rational with inverse temperature parameter β and beliefs Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ), see Equation (8). Does this data identify the return function G:āā:āāG: Sā RG : overā start_ARG S end_ARG ā blackboard_R? If the observation-based choice probabilities from Equation (2) would be known, then Corollary C.4 would provide the answer to this question. Thus, similar to how we previously inverted the belief operator BB, we are now simply tasked with inverting the expectation over observation sequences. This leads us to the following definition: Definition C.7 (Ungrounding Operator). The ungrounding operators :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG and ā:āĪ©āĆĪ©āāāĆā:tensor-productāsuperscriptāāĪ©āĪ©superscriptāā O O: R % Ć ā R SĆ % SO ā O : blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG Ć overā start_ARG S end_ARG are defined by [ā”(v)]ā¢(sā)āoāā¼POāā¢(oāā£sā)[vā¢(oā)],[(ā)ā¢(C)]ā¢(sā,sāā²)āoā,oāā²ā¼POā(ā ā£sā,sāā²)[Cā¢(oā,oāā²)]. [ O(v) ]( s) *% E_ o P_ O( o s) [v( o) % ], [( O O)(C)% ]( s, s 1.0pt ) *E% _ o, o 1.0pt P_ O(Ā· s, % s 1.0pt ) [C( o, o 1.0pt ) ].[ O ( v ) ] ( overā start_ARG s end_ARG ) ā Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ v ( overā start_ARG o end_ARG ) ] , [ ( O ā O ) ( C ) ] ( overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) ā Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] . Here, vāāĪ©āsuperscriptāāĪ©vā R v ā blackboard_Roverā start_ARG Ī© end_ARG is an arbitrary vector, and CāāĪ©āĆĪ©āsuperscriptāāĪ©āĪ©Cā R Ć C ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG is also an arbitrary vector, where the notation can remind of āChoiceā since the inputs to ātensor-product O OO ā O are, in practice, vectors of observation-based Boltzmann-rational choice probabilities. Formally, ātensor-product O OO ā O is the Kronecker product of OO with itself, but it is not necessary to understand this fact to follow the discussion. Ultimately, to be able to recover the observation-based choice probabilities, what matters is that ātensor-product O OO ā O is injective on whole vectors of these choice probabilities. The injectivity of OO is a sufficient condition for this, which explains its usefulness. We show this in the following lemma: Lemma C.8. :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG is injective if and only if ā:āĪ©āĆĪ©āāāĆā:tensor-productāsuperscriptāāĪ©āĪ©superscriptāā O O: R % Ć ā R SĆ % SO ā O : blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG Ć overā start_ARG S end_ARG is injective. Proof. This is a general property of the Kronecker product of a linear operator with itself. For completeness, we demonstrate the calculation in our special case. First, assume that OO is injective. Assume that (ā)ā¢(C)=0tensor-product0( O O)(C)=0( O ā O ) ( C ) = 0 for some CāāĪ©āĆĪ©āsuperscriptāāĪ©āĪ©Cā R Ć C ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG. We need to show C=00C=0C = 0. For all pairs of state sequences (sā,sāā²)āsuperscriptāā²( s, s 1.0pt )( overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ), we have 0=[(ā)ā¢(C)]ā¢(sā,sāā²)0delimited-[]tensor-productāsuperscriptāā² 0= [( O % O)(C) ]( s, s 1.0pt )0 = [ ( O ā O ) ( C ) ] ( overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) =oā,oāā²ā¼POā(ā ā£sā,sāā²)[Cā¢(oā,oāā²)] = *E_ o, o 1.0pt % P_ O(Ā· s, s 1.0pt ) [C( % o, o 1.0pt ) ]= Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] =oāā¼POāā¢(oāā£sā)[oāā²ā¼POāā¢(oāā²ā£sāā²)[Cā¢(oā,oāā²)]]absentsubscriptsimilar-toāsubscriptāconditionalāsubscriptsimilar-tosuperscriptāā²subscriptāconditionalsuperscriptāā²āsuperscriptāā² = *E_ o P_ O( o % s) [ *E_ o 1.0pt P% _ O( o 1.0pt s 1.0pt ) % [C( o, o 1.0pt ) ] ]= Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ Eoverā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARGⲠ⣠overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] ] =oāā¼POāā¢(oāā£sā)[Csāā²ā¢(oā)]absentsubscriptsimilar-toāsubscriptāconditionalāsubscriptsuperscriptā²āā²ā = *E_ o P_ O( o % s) [C _ s 1.0pt ( o) ]= Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ Cā²overā start_ARG s end_ARGā² ( overā start_ARG o end_ARG ) ] =[ā”(Csāā²)]ā¢(sā),absentdelimited-[]subscriptsuperscriptā²āā²ā = [ O (C _ s 1% .0pt ) ]( s),= [ O ( Cā²overā start_ARG s end_ARGā² ) ] ( overā start_ARG s end_ARG ) , where Csāā²ā¢(oā)āoāā²ā¼POāā¢(oāā²ā£sāā²)[Cā¢(oā,oāā²)]āsubscriptsuperscriptā²āā²āsubscriptsimilar-tosuperscriptāā²subscriptāconditionalsuperscriptāā²āsuperscriptāā²C _ s 1.0pt ( o) *% E_ o 1.0pt P_ O( o 1.0pt^% s 1.0pt ) [C( o, o 1.0pt% ) ]Cā²overā start_ARG s end_ARGā² ( overā start_ARG o end_ARG ) ā Eoverā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARGⲠ⣠overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ]. By the injectivity of OO, we obtain Csāā²=0subscriptsuperscriptā²āā²0C _ s 1.0pt =0Cā²overā start_ARG s end_ARGā² = 0 for all sāā² s 1.0pt overā start_ARG s end_ARGā². This means that for all sāā² s 1.0pt overā start_ARG s end_ARGā² and oā ooverā start_ARG o end_ARG, we have 0=Csāā²ā¢(oā)=oāā²ā¼POāā¢(oāā²ā£sāā²)[Cā¢(oā,oāā²)]=[ā”(Coāā²)]ā¢(sāā²),0subscriptsuperscriptā²āā²āsubscriptsimilar-tosuperscriptāā²subscriptāconditionalsuperscriptāā²āsuperscriptāā²delimited-[]subscriptsuperscriptā²āsuperscriptāā² 0=C _ s 1.0pt ( o)=% *E_ o 1.0pt P_ O( % o 1.0pt s 1.0pt ) [C( o, % o 1.0pt ) ]= [ O (C^% _ o ) ]( s 1.0pt ),0 = Cā²overā start_ARG s end_ARGā² ( overā start_ARG o end_ARG ) = Eoverā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARGⲠ⣠overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] = [ O ( Cā² ā²overā start_ARG o end_ARG ) ] ( overā start_ARG s end_ARGā² ) , where Coāā²ā¢(oāā²)āCā¢(oā,oāā²)āsubscriptsuperscriptā²āsuperscriptāā²āsuperscriptāā²C _ o( o 1.0pt ) C( o,% o 1.0pt )Cā² ā²overā start_ARG o end_ARG ( overā start_ARG o end_ARGā² ) ā C ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ). Again, by the injectivity of OO, we obtain Coāā²=0subscriptsuperscriptā²ā0C _ o=0Cā² ā²overā start_ARG o end_ARG = 0 for all oā ooverā start_ARG o end_ARG, leading to C=00C=0C = 0. That proves the direction from left to right. To prove the other direction, assume that OO is not injective. This means there exists 0ā CāāĪ©ā0superscriptāāĪ©0ā Cā R 0 ā C ā blackboard_Roverā start_ARG Ī© end_ARG such that ā”(C)=00 O(C)=0O ( C ) = 0. Define CāCāāĪ©āĆĪ©ātensor-productsuperscriptāāĪ©āĪ©C Cā R Ć C ā C ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG by (CāC)ā¢(oā,oāā²)āCā¢(oā)ā¢Cā¢(oāā²).ātensor-productāsuperscriptāā²āsuperscriptāā²(C C)( o, o 1.0pt ) C( o)C( o% 1.0pt ).( C ā C ) ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ā C ( overā start_ARG o end_ARG ) C ( overā start_ARG o end_ARGā² ) . Then clearly, CāCā 0tensor-product0C Cā 0C ā C ā 0. We are done if we can show that (ā)ā¢(CāC)=0tensor-producttensor-product0( O O)(C C)=0( O ā O ) ( C ā C ) = 0 since that establishes that ātensor-product O OO ā O is also not injective. For any sā,sāā²āāsuperscriptāā²ā s, s 1.0pt ā Soverā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG S end_ARG, we have [(ā)ā¢(CāC)]ā¢(sā,sāā²)delimited-[]tensor-producttensor-productāsuperscriptāā² [( O O% )(C C) ]( s, s 1.0pt )[ ( O ā O ) ( C ā C ) ] ( overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) =oā,oāā²ā¼POā(ā ā£sā,sāā²)[(CāC)ā¢(oā,oāā²)] = *E_ o, o 1.0pt % P_ O(Ā· s, s 1.0pt ) [(C% C)( o, o 1.0pt ) ]= Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ ( C ā C ) ( overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) ] =oā,oāā²ā¼POā(ā ā£sā,sāā²)[Cā¢(oā)ā Cā¢(oāā²)] = *E_ o, o 1.0pt % P_ O(Ā· s, s 1.0pt ) [C( % o)Ā· C( o 1.0pt ) ]= Eoverā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG ) ā C ( overā start_ARG o end_ARGā² ) ] =oāā¼POāā¢(oāā£sā)[Cā¢(oā)]ā oāā²ā¼POāā¢(oāā²ā£sāā²)[Cā¢(oāā²)]absentā subscriptsimilar-toāsubscriptāconditionalāsubscriptsimilar-tosuperscriptāā²subscriptāconditionalsuperscriptāā²superscriptāā² = *E_ o P_ O( o % s) [C( o) ]Ā· *E_ o% 1.0pt P_ O( o 1.0pt s% 1.0pt ) [C( o 1.0pt ) ]= Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARG ) ] ā Eoverā start_ARG o end_ARGā² ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARGⲠ⣠overā start_ARG s end_ARGā² ) end_POSTSUBSCRIPT [ C ( overā start_ARG o end_ARGā² ) ] =[ā”(C)]ā¢(sā)ā [ā”(C)]ā¢(sāā²)absentā delimited-[]ādelimited-[]superscriptāā² = [ O(C) ]( s)Ā· [% O(C) ]( s 1.0pt )= [ O ( C ) ] ( overā start_ARG s end_ARG ) ā [ O ( C ) ] ( overā start_ARG s end_ARGā² ) =0ā 0absentā 00 =0Ā· 0= 0 ā 0 =0.absent0 =0.= 0 . This finishes the proof. ā We now state and prove the following extension of Corollary C.4: Theorem C.9. Consider the following statements (which can each be true or false): 1. :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG is an injective linear operator: kerā”=0kernel0 O=\0\ker O = 0 . 2. ā:āĪ©āĆĪ©āāāĆā:tensor-productāsuperscriptāāĪ©āĪ©superscriptāā O O: R % Ć ā R SĆ % SO ā O : blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG Ć overā start_ARG S end_ARG is an injective linear operator: kerā”ā=0kerneltensor-product0 O O=\0\ker O ā O = 0 . 3. ātensor-product O OO ā O is injective on vectors of observation-based choice probabilities (PRā¢(oāā»oāā²))oā,oāā²subscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā² (P^R ( o o 1.0pt ) )_% o, o 1.0pt ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² over the set of return functions GāāāsuperscriptāāGā R SG ā blackboard_Roverā start_ARG S end_ARG. 4. The data of state-based choice probabilities (PRā¢(sāā»sāā²))sā,sāā²āāsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²ā (P^R ( s s 1.0pt ) )_% s, s 1.0pt ā S( Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) )overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG S end_ARG from Equation (8) determine the data of observation-based choice probabilities (PRā¢(oāā»oāā²))oā,oāā²āĪ©āsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P^R ( o o 1.0pt ) )_% o, o 1.0pt ā ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ā overā start_ARG Ī© end_ARG from Equation (2). Then the following implications hold and 3 does not imply 2: 1111222233334.44.4 . Consequently, if any of the conditions 1, 2, or 3 hold, and additionally any of the conditions (i), (i) or (i) from Corollary C.4, then the data (PRā¢(sāā»sāā²))sā,sāā²āĪ©āsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²āĪ© (P^R ( s s 1.0pt ) )_% s, s 1.0pt ā ( Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) )overā start_ARG s end_ARG , overā start_ARG s end_ARGā² ā overā start_ARG Ī© end_ARG determine the return function G on sequences sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG up to a constant independent of sā soverā start_ARG s end_ARG. Proof. That 1 and 2 are equivalent was shown in Lemma C.8. That 2 implies 3 is clear. To prove that 3 implies 4, simply put both sets of choice probabilities into a vector. Then Equation (8) and Definition C.7 show the following equality of vectors in āāĆāsuperscriptāā R SĆ Sblackboard_Roverā start_ARG S end_ARG Ć overā start_ARG S end_ARG: (PRā¢(sāā»sāā²))sā,sāā²=(ā)ā¢((PRā¢(oāā»oāā²))oā,oāā²).subscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²tensor-productsubscriptsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā² (P^R ( s s 1.0pt ) )_% s, s 1.0pt = ( O % O ) ( (P^R ( o o% 1.0pt ) )_ o, o 1.0pt % ).( Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) )overā start_ARG s end_ARG , overā start_ARG s end_ARGā² = ( O ā O ) ( ( Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) )overā start_ARG o end_ARG , overā start_ARG o end_ARGā² ) . The injectivity of ātensor-product O OO ā O on such inputs ensures that the observation-based choice probabilities can be recovered using this equation. We now show that (3) does not imply (2). Again, we use the simple MDP from Equation (7), but this time with a different observation kernel. Namely, we choose POā¢(o(a)ā£a)=POā¢(o(a)ā²ā£a)=12,POā¢(o(b)ā£b)=1,formulae-sequencesubscriptconditionalsuperscriptsubscriptconditionalsuperscriptsuperscriptā²12subscriptconditionalsuperscript1P_O(o^(a) a)=P_O(o^(a) a)= 12, P_O(o% ^(b) b)=1,Pitalic_O ( o( a ) ⣠a ) = Pitalic_O ( o( a )Ⲡ⣠a ) = divide start_ARG 1 end_ARG start_ARG 2 end_ARG , Pitalic_O ( o( b ) ⣠b ) = 1 , where o(a)ā²ā o(a)superscriptsuperscriptā²o^(a) ā o^(a)o( a )ā² ā o( a ) and o(a)ā o(b)ā o(a)ā²superscriptsuperscriptsuperscriptsuperscriptā²o^(a)ā o^(b)ā o^(a) o( a ) ā o( b ) ā o( a )ā². This results in two possible observation sequences: o(a)ā¢o(b)superscriptsuperscripto^(a)o^(b)o( a ) o( b ) and o(a)ā²ā¢o(b)superscriptsuperscriptā²o^(a) o^(b)o( a )ā² o( b ). Thus, āĪ©āsuperscriptāāĪ© R blackboard_Roverā start_ARG Ī© end_ARG is two-dimensional, whereas āāsuperscriptāā R Sblackboard_Roverā start_ARG S end_ARG is only one-dimensional. Consequently, :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG cannot be injective, so kerā”ā 0kernel0 Oā \0\ker O ā 0 , so (2) does not hold since (1) and (2) are equivalent. However, (3) still holds: Since there is only one state sequence, Equation (2) shows that the only vector of choice probabilities has 1/2121/21 / 2 in all its entries, irrespective of the return function G. Thus, ātensor-product O OO ā O has only one input of observation-based choice probabilities, and is thus automatically injective on its inputs. The final result of identifiability of the return function G follows using Corollary C.4. ā C.5 Simple Special Cases: Full Observability, Deterministic POāsubscriptāP_ OPoverā start_ARG O end_ARG, and Noisy POāsubscriptāP_ OPoverā start_ARG O end_ARG In this section, we analyze three simple special cases of the general theory. Theorem 3.9 (together with Lemma B.3) from Skalse et al. [2023], reproduced as a corollary below, is a special case of our theorem: Corollary C.10 (Skalse et al. [2023]). Assume the human directly observes the true sequences, and the choice probabilities are given by PRā¢(sāā»sāā²)=Ļā¢(βā¢(Gā¢(sā)āGā¢(sāā²))).superscriptsucceedsāsuperscriptāā²āsuperscriptāā²P^R ( s s 1.0pt )=Ļ (β% (G( s)-G( s 1.0pt ) ) ).Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) = Ļ ( β ( G ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARGā² ) ) ) . This data determines the return function G=ā”(R)G= (R)G = Ī ( R ) on state sequences sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG up to a constant independent on sā soverā start_ARG s end_ARG. Proof. We can embed this case into the one of Theorem C.9 by defining the observation kernel as POāā¢(sāā²ā£sā)=Ī“sāā¢(sāā²)subscriptāconditionalsuperscriptāā²āsubscriptāsuperscriptāā²P_ O( s 1.0pt s)= _ s( s% 1.0pt )Poverā start_ARG O end_ARG ( overā start_ARG s end_ARGⲠ⣠overā start_ARG s end_ARG ) = Ī“overā start_ARG s end_ARG ( overā start_ARG s end_ARGā² ) (i.e., the correct sequence is deterministically observed) and defining the humanās belief as Bā¢(sāā²ā£sā)=Ī“sāā¢(sāā²)conditionalsuperscriptāā²āsubscriptāsuperscriptāā²B( s 1.0pt s)= _ s( s 1.0% pt )B ( overā start_ARG s end_ARGⲠ⣠overā start_ARG s end_ARG ) = Ī“overā start_ARG s end_ARG ( overā start_ARG s end_ARGā² ) (i.e., the human knows that the observation reflects the true sequence). This shows that Pā¢(sāā»sāā²)succeedsāsuperscriptāā²P( s s 1.0pt )P ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) is of the form of Equation (8). The result follows from Theorem C.9: the operators OO and BB are the identity in this case, due to the defining property of the Kronecker delta, and so they are injective. ā The following proposition shows that Corollary C.10 is essentially the only example of deterministic observation kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG for which BB is injective. Note, however, that in some situations, we can have imā”ā©kerā”=0imkernel0im ā© B% =\0\im Ī ā© ker B = 0 even if BB is not injective, see Example C.32. Proposition C.11. Assume POāsubscriptāP_ OPoverā start_ARG O end_ARG, the observation kernel on the level of sequences, is deterministic and not injective. Then OO is automatically injective. However, BB is not injective. Proof. To show that OO is injective, assume vāāĪ©āsuperscriptāāĪ©vā R v ā blackboard_Roverā start_ARG Ī© end_ARG is such that ā”(v)=00 O(v)=0O ( v ) = 0. Then for all sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG, we get 0=[ā”(v)]ā¢(sā)=oāā¼POāā¢(oāā£sā)[vā¢(oā)]=vā¢(Oāā¢(sā)).0delimited-[]āsubscriptsimilar-toāsubscriptāconditionalā0= [ O(v) ]( s)= *E% _ o P_ O( o s) [v( o) ]=v % ( O( s) ).0 = [ O ( v ) ] ( overā start_ARG s end_ARG ) = Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ v ( overā start_ARG o end_ARG ) ] = v ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) . Since Oā:āĪ©ā:āĪ© O: Sā overā start_ARG O end_ARG : overā start_ARG S end_ARG ā overā start_ARG Ī© end_ARG is by definition surjective, we obtain v=00v=0v = 0. Oā:āĪ©ā:āĪ© O: Sā overā start_ARG O end_ARG : overā start_ARG S end_ARG ā overā start_ARG Ī© end_ARG is by definition surjective, and here assumed to be non-injective, which implies that ā Soverā start_ARG S end_ARG has a higher cardinality than Ī©āĪ© overā start_ARG Ī© end_ARG. Thus, :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© B: R Sā R % B : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG cannot be injective. ā In the following, we analyze a simple case that guarantees identifiability. It requires that the observation kernel is āwell-behavedā of a form where the observations are simply ānoisy statesā, and that the human is a Bayesian reasoner with any prior Bā¢(sā)āB( s)B ( overā start_ARG s end_ARG ) that supports every state sequence sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG. Definition C.12 (Noise in the Observation Kernel). Then we say that there is noise in the observation kernel PO:āĪā¢(Ī©ā):subscriptāĪāĪ©P_O: Sā ( )Pitalic_O : overā start_ARG S end_ARG ā Ī ( overā start_ARG Ī© end_ARG ) if ā=Ī©āĪ© S= overā start_ARG S end_ARG = overā start_ARG Ī© end_ARG and if OO is an injective linear operator. Proposition C.13. Assume that ā=Ī©āĪ© S= overā start_ARG S end_ARG = overā start_ARG Ī© end_ARG. Furthermore, assume that Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) is given by the posterior with likelihood POāā¢(oāā£sā)subscriptāconditionalāP_ O( o s)Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) and any prior Bā¢(sā)āB( s)B ( overā start_ARG s end_ARG ) with Bā¢(sā)>0ā0B( s)>0B ( overā start_ARG s end_ARG ) > 0 for all sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG. Then there is noise in the observation kernel if and only if BB is injective. Proof. Assume OO is injective. To show that BB is injective, assume there is Gā²āāāsuperscriptā²āāG ā R SGā² ā blackboard_Roverā start_ARG S end_ARG with ā”(Gā²)=0superscriptā²0 B(G )=0B ( Gā² ) = 0. Then for all oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG, we have 00 0 =[ā”(Gā²)]ā¢(oā)=sāā¼Bā¢(sāā£oā)[Gā²ā¢(sā)]=āsāBā¢(sāā£oā)ā¢Gā²ā¢(sā)āāsāPOāā¢(oāā£sā)ā (Bā¢(sā)ā Gā²ā¢(sā))absentdelimited-[]superscriptā²āsubscriptsimilar-toāconditionalāsuperscriptā²āsubscriptāconditionalāsuperscriptā²āproportional-tosubscriptāā subscriptāconditionalāā āsuperscriptā²ā = [ B(G ) ]( o)=% *E_ s B( s o) [G % ( s) ]= _ sB( s o)G ( s)% _ sP_ O( o s)Ā· (B( s)% Ā· G ( s) )= [ B ( Gā² ) ] ( overā start_ARG o end_ARG ) = Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ Gā² ( overā start_ARG s end_ARG ) ] = āoverā start_ARG s end_ARG B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) Gā² ( overā start_ARG s end_ARG ) ā āoverā start_ARG s end_ARG Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) ā ( B ( overā start_ARG s end_ARG ) ā Gā² ( overā start_ARG s end_ARG ) ) =[Tā”(BāGā²)]ā¢(oā).absentdelimited-[]superscriptdirect-productsuperscriptā²ā = [ O^T(B G ) ](% o).= [ Oitalic_T ( B ā Gā² ) ] ( overā start_ARG o end_ARG ) . Here, Tsuperscript O^TOitalic_T is the transpose of OO and BāGā²direct-productsuperscriptā²B G B ā Gā² is the componentwise product of the prior B with the return function Gā². Since OO is injective and thus invertible, Tsuperscript O^TOitalic_T is as well. Thus, BāGā²=0direct-productsuperscriptā²0B G =0B ā Gā² = 0, which implies Gā²=0superscriptā²0G =0Gā² = 0 since the prior gives positive probability to all state sequences. Thus, BB is injective. For the other direction, assume BB is injective. To show that OO is injective, let vāāĪ©āsuperscriptāāĪ©vā R v ā blackboard_Roverā start_ARG Ī© end_ARG be any vector with ā”(v)=00 O(v)=0O ( v ) = 0. We do a similar computation as above: for all sāāāāsuperscriptāā sā R Soverā start_ARG s end_ARG ā blackboard_Roverā start_ARG S end_ARG, we have 00 0 =[ā”(v)]ā¢(sā)=oāā¼POāā¢(oāā£sā)[vā¢(oā)]=āoāPOāā¢(oāā£sā)ā¢vā¢(oā)āāoāBā¢(sāā£oā)ā (POāā¢(oā)ā vā¢(oā))absentdelimited-[]āsubscriptsimilar-toāsubscriptāconditionalāsubscriptāsubscriptāconditionalāproportional-tosubscriptāā conditionalāā subscriptā = [ O(v) ]( s)=% *E_ o P_ O( o s) [% v( o) ]= _ oP_ O( o s)v( o)% _ oB( s o)Ā· (P_ O( o)% Ā· v( o) )= [ O ( v ) ] ( overā start_ARG s end_ARG ) = Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ v ( overā start_ARG o end_ARG ) ] = āoverā start_ARG o end_ARG Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) v ( overā start_ARG o end_ARG ) ā āoverā start_ARG o end_ARG B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) ā ( Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ) ā v ( overā start_ARG o end_ARG ) ) =[Tā”(POāāv)]ā¢(sā).absentdelimited-[]superscriptdirect-productsubscriptā = [ B^T (P_ O v% ) ]( s).= [ Bitalic_T ( Poverā start_ARG O end_ARG ā v ) ] ( overā start_ARG s end_ARG ) . Here, Tsuperscript B^TBitalic_T is the transpose of BB, POāā¢(oā)subscriptāP_ O( o)Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ) is the denominator in Bayes rule, and POāāvdirect-productsubscriptāP_ O vPoverā start_ARG O end_ARG ā v is the vector with components POāā¢(oā)ā vā¢(oā)ā subscriptāP_ O( o)Ā· v( o)Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ) ā v ( overā start_ARG o end_ARG ). From the injectivity and thus invertibility of BB, it follows that Tsuperscript B^TBitalic_T is invertible as well, and so POāāv=0direct-productsubscriptā0P_ O v=0Poverā start_ARG O end_ARG ā v = 0, which implies v=00v=0v = 0. Thus, OO is injective. ā Corollary C.14. When there is noise in the observation kernel and the human is a Bayesian reasoner with some prior B such that Bā¢(sā)>0ā0B( s)>0B ( overā start_ARG s end_ARG ) > 0 for all sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG, then the return function is identifiable from choice probabilities of state sequences even if the learning system does not know the humanās observations. Proof. This follows from the injectivity of OO, the injectivity of BB that we proved in Proposition C.13, and Theorem C.9. ā Remark C.15. We mention the following caveat: intuitively, one could think that OO (and thus BB, by Proposition C.13) will be injective if every sā soverā start_ARG s end_ARG is identifiable from infinitely many i.i.d. samples from POāā¢(oāā£sā)subscriptāconditionalāP_ O( o s)Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ). A counterexample is the following: =(1/21/41/41/41/21/43/83/81/4).matrix121414141214383814 O= pmatrix1/2&1/4&1/4\\ 1/4&1/2&1/4\\ 3/8&3/8&1/4 pmatrix.O = ( start_ARG start_ROW start_CELL 1 / 2 end_CELL start_CELL 1 / 4 end_CELL start_CELL 1 / 4 end_CELL end_ROW start_ROW start_CELL 1 / 4 end_CELL start_CELL 1 / 2 end_CELL start_CELL 1 / 4 end_CELL end_ROW start_ROW start_CELL 3 / 8 end_CELL start_CELL 3 / 8 end_CELL start_CELL 1 / 4 end_CELL end_ROW end_ARG ) . In this case, the rows are linearly dependent with coefficients 1/2,1/212121/2,1/21 / 2 , 1 / 2 and ā11-1- 1. Consequently, OO and BB are not injective, and so if this observation kernel comes from a multi-armed bandit with three states, then Corollary C.4 shows that the return function is not identifiable. Nevertheless, the distributions POā(ā ā£sā)P_ O(Ā· s)Poverā start_ARG O end_ARG ( ā ⣠overā start_ARG s end_ARG ) (given by the rows) all differ from each other, and so infinitely many i.i.d. samples identify the state sequence sā soverā start_ARG s end_ARG. C.6 Robustness of Return Function Identifiability under Belief Misspecification We now again look at the case where the observations that the human observes are known to the reward learning system, as in Section C.2. Furthermore, we assume that :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© B: R Sā R % B : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG is such that kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . In this case, we can apply Corollary C.4 and identify the true return function G from ā”(G) B(G)B ( G ), which, in turn, can be identified up to an additive constant from the observation-based choice probabilities with the argument as for Proposition 3.1. In this section, we investigate what happens when the human belief model is slightly misspecified. In other words: the learning system uses a perturbed matrix ā+āsubscript B_ % B+ Bbold_Ī ā B + Ī with some small perturbation Ī. How much will the inferred return function deviate from the truth? To answer this, we first need to outline some norm theory of linear operators. C.6.1 Some Norm Theory for Linear Operators In this section, let V,WV,WV , W be two finite-dimensional inner product-spaces. In other words, V and W each have inner products āØā ,ā ā©ā Ā·,Ā· ⨠ā , ā ā© and there are linear isomorphisms Vā āksuperscriptāV R^kV ā blackboard_Rk, Wā āmsuperscriptāW R^mW ā blackboard_Rm such that the inner products in V and W correspond to the standard scalar products in āksuperscriptā R^kblackboard_Rk and āmsuperscriptā R^mblackboard_Rm. The reason that we donāt directly work with āksuperscriptā R^kblackboard_Rk and āmsuperscriptā R^mblackboard_Rm itself is that we will later apply the analysis to the case that V=imā”āāāimsuperscriptāāV=im R % SV = im Ī ā blackboard_Roverā start_ARG S end_ARG. Let in this whole section :VāW:ā A:Vā WA : V ā W be a linear operator and :VāW:ā :Vā WĪ : V ā W be a perturbance, so that ā+āsubscript A_ % A+ Abold_Ī ā A + Ī is a perturbed version of AA. The inner products give rise to a norm on V and W defined by āvā=āØv,vā©,āwā=āØw,wā©.formulae-sequencenormnorm\|v\|= v,v , \|w\|= w,w% .ā„ v ā„ = square-root start_ARG ⨠v , v ā© end_ARG , ā„ w ā„ = square-root start_ARG ⨠w , w ā© end_ARG . As is well known, for each linear operator :VāW:ā A:Vā WA : V ā W there exists a unique, basis-independent adjoint (generalizing the notion of a transpose) T:WāV:superscriptā A^T:Wā VAitalic_T : W ā V such that for all vāVvā Vv ā V and wāWwā Ww ā W, we have āØā”v,wā©=āØv,Tā”wā©.superscript Av,w = v,% A^Tw .⨠A v , w ā© = ⨠v , Aitalic_T w ā© . Let us recall the following fact that is often used in linear regression: Lemma C.16. Assume :VāW:ā A:Vā WA : V ā W is injective. Then Tā”:VāV:superscriptā A^T A:Vā VAitalic_T A : V ā V is invertible and (Tā”)ā1ā¢Tsuperscriptsuperscript1superscript( A^T A)^-1% A^T( Aitalic_T A )- 1 Aitalic_T is a left inverse of AA. Proof. To show that Tā”superscript A^T AAitalic_T A is invertible, we only need to show that it is injective. Thus, let 0ā xāV00ā xā V0 ā x ā V. Then āØx,Tā”xā©=āØā”x,ā”xā©=āā”xā2>0,superscriptsuperscriptnorm20 x, A^T Ax % = Ax, Ax% =\| Ax\|^2>0,⨠x , Aitalic_T A x ā© = ⨠A x , A x ā© = ā„ A x ā„2 > 0 , where the last step followed from the injectivity of AA. Thus, Tā”xā 0superscript0 A^T Axā 0Aitalic_T A x ā 0, and so Tā”superscript A^T AAitalic_T A is injective, and thus invertible. Consequently, (Tā”)ā1ā¢Tsuperscriptsuperscript1superscript( A^T A)^-1% A^T( Aitalic_T A )- 1 Aitalic_T is a well-defined operator. That it is the left inverse of AA is clear. ā Definition C.17 (Operator Norm). The norm of an operator :VāW:ā A:Vā WA : V ā W is given by āāmaxx,āxā=1ā”āā”xā.ānormsubscriptnorm1norm\| A\| _x,\ \|x\|=1\|% Ax\|.ā„ A ā„ ā maxitalic_x , ā„ x ā„ = 1 ā„ A x ā„ . It has the following well-known properties, where , A, BA , B and CC are matrices of compatible sizes: ā+āā¤ā+ā,āā”āā¤āā ā,āTā=ā.formulae-sequencenormnormnormformulae-sequencenormā normnormnormsuperscriptnorm\| A+ B\|ā¤\|% A\|+\| B\|, \| C% A\|ā¤\| C\|Ā·\|% A\|, \| A^T\|=\|% A\|.ā„ A + B ℠⤠℠A ā„ + ā„ B ā„ , ā„ C A ℠⤠℠C ā„ ā ā„ A ā„ , ā„ Aitalic_T ā„ = ā„ A ā„ . To study how a perturbance in AA (and thus Tā”superscript A^T AAitalic_T A) transfers into a perturbance of (Tā”)ā1superscriptsuperscript1 ( A^T A )^-1( Aitalic_T A )- 1, we will use the following theorem: Theorem C.18 (El Ghaoui [2002]). Let :VāV:ā B:Vā VB : V ā V be an invertible operator. Let Ļ<āā1āā1superscriptnormsuperscript11Ļ<\| B^-1\|^-1Ļ < ā„ B- 1 ā„- 1. Let :VāV:ā :Vā VĪ : V ā V be any operator with āā¤Ļnorm\| \|ā¤Ļ℠Π℠⤠Ļ. Then + B+ B + Ī is invertible and we have ā(+)ā1ā1āā¤Ļā āā1āā1āā1āĻ.normsuperscript1superscript1ā normsuperscript1superscriptnormsuperscript11 \|( B+ )^-1-% B^-1 \|⤠ĻĀ·\|% B^-1\|\| B^-1\|^-1-Ļ.ā„ ( B + Ī )- 1 - B- 1 ℠⤠divide start_ARG Ļ ā ā„ B- 1 ā„ end_ARG start_ARG ā„ B- 1 ā„- 1 - Ļ end_ARG . Proof. See El Ghaoui [2002], Section 7 and in particular Equation 7.2. Note that the reference defines ānorm\| A\|ā„ A ā„ to be the largest singular value of AA; by the well-known min-max theorem, this is equivalent to Definition C.17. ā We will apply this theorem to Tā”superscript A^T AAitalic_T A, which raises the question about the size of the perturbance in Tā”superscript A^T AAitalic_T A for a given perturbance in AA. This is clarified in the following lemma. Before stating it, for a given perturbance Ļ, define Ļ~ā¢()āĻā (2ā ā+Ļ),ā~ā 2norm Ļ( A) ĻĀ· (2Ā·\|% A\|+Ļ ),over~ start_ARG Ļ end_ARG ( A ) ā Ļ ā ( 2 ā ā„ A ā„ + Ļ ) , which depends on AA and Ļ. Also, recall that for a given perturbance Ī, we define ā+āsubscript A_ % A+ Abold_Ī ā A + Ī. We obtain: Lemma C.19. Assume that āā¤Ļnorm\| \|ā¤Ļ℠Π℠⤠Ļ. Then āTā”āTā”āā¤Ļ~ā¢().normsuperscriptsubscriptsubscriptsuperscript~\| A_ ^T% A_ - A^T% A\|⤠Ļ( A).ā„ Abold_Īitalic_T Abold_Ī - Aitalic_T A ℠⤠over~ start_ARG Ļ end_ARG ( A ) . Proof. We have āTā”āTā”ānormsuperscriptsubscriptsubscriptsuperscript \| A_ % ^T A_ -% A^T A \|ā„ Abold_Īitalic_T Abold_Ī - Aitalic_T A ā„ =ā(+)Tā¢(+)āTā”āabsentnormsuperscriptsuperscript = \|( A+ % )^T( A+ )- % A^T A \|= ā„ ( A + Ī )T ( A + Ī ) - Aitalic_T A ā„ =āTā”+Tā”+Tā”āabsentnormsuperscriptsuperscriptsuperscript = \| A^T % + ^T A+% ^T \|= ā„ Aitalic_T Ī + Īitalic_T A + Īitalic_T Ī ā„ ā¤āā ā+āā ā+ā2absentā normnormā normnormsuperscriptnorm2 ā¤\| A\|Ā·\| % \|+\| \|Ā·\| A% \|+\| \|^2⤠℠A ā„ ā ā„ Ī ā„ + ā„ Ī ā„ ā ā„ A ā„ + ā„ Ī ā„2 ā¤Ļā (2ā ā+Ļ)absentā 2norm ā¤ĻĀ· (2Ā·\| A\|+Ļ )ā¤ Ļ ā ( 2 ā ā„ A ā„ + Ļ ) =Ļ~ā¢().absent~ = Ļ( A).= over~ start_ARG Ļ end_ARG ( A ) . ā To be able to apply Theorem C.18 to Tā”superscript A^T AAitalic_T A, we need to make sure that Ļ~ā¢()~ Ļ( A)over~ start_ARG Ļ end_ARG ( A ) is bounded above by ā(Tā”)ā1āā1superscriptnormsuperscriptsuperscript11 \|( A^T A )^-1\|^% -1ā„ ( Aitalic_T A )- 1 ā„- 1. The next lemma clarifies what condition Ļ needs to satisfy for Ļ~ā¢()~ Ļ( A)over~ start_ARG Ļ end_ARG ( A ) to obey that bound. For this, define Ļā¢()āāā+ā2+ā(Tā”)ā1āā1,ānormsuperscriptnorm2superscriptnormsuperscriptsuperscript11Ļ( A) -\| A\|+ % \| A\|^2+ \|( A^T% A)^-1 \|^-1,Ļ ( A ) ā - ā„ A ā„ + square-root start_ARG ā„ A ā„2 + ā„ ( Aitalic_T A )- 1 ā„- 1 end_ARG , (9) which only depends on AA. Lemma C.20. Assume Ļ<Ļā¢()Ļ<Ļ( A)Ļ < Ļ ( A ). Then Ļ~ā¢()<ā(Tā”)ā1āā1.~superscriptnormsuperscriptsuperscript11 Ļ( A)< \|( A% ^T A)^-1 \|^-1.over~ start_ARG Ļ end_ARG ( A ) < ā„ ( Aitalic_T A )- 1 ā„- 1 . Proof. Note that Ļ=Ļā¢()Ļ=Ļ( A)Ļ = Ļ ( A ) is the positive solution to the following quadratic equation in the indeterminate Ļ: Ļ2+2ā āā Ļāā(Tā”)ā1āā1=Ļ~ā¢()āā(Tā”)ā1āā1=0.superscript2ā 2normsuperscriptnormsuperscriptsuperscript11~superscriptnormsuperscriptsuperscript110Ļ^2+2Ā·\| A\|Ā·Ļ- \|(% A^T A)^-1 \|^-1= Ļ(% A)- \|( A^T % A)^-1 \|^-1=0.Ļ2 + 2 ā ā„ A ā„ ā Ļ - ā„ ( Aitalic_T A )- 1 ā„- 1 = over~ start_ARG Ļ end_ARG ( A ) - ā„ ( Aitalic_T A )- 1 ā„- 1 = 0 . Since this is a convex parabola, we get the inequality Ļ~ā¢()āā(Tā”)ā1āā1<0~superscriptnormsuperscriptsuperscript110 Ļ( A)- \|( A% ^T A)^-1 \|^-1<0over~ start_ARG Ļ end_ARG ( A ) - ā„ ( Aitalic_T A )- 1 ā„- 1 < 0 whenever we have 0ā¤Ļ<Ļā¢()00ā¤Ļ<Ļ( A)0 ā¤ Ļ < Ļ ( A ), which shows the result. ā Finally, we put it all together to obtain a bound on the perturbance of (Tā”)ā1ā¢Tsuperscriptsuperscript1superscript ( A^T A )^-1% A^T( Aitalic_T A )- 1 Aitalic_T. For this, set Cā¢(,Ļ)āĻ~ā¢()ā ā(Tā”)ā1ā(Tā”)ā1āā1āĻ~ā¢()ā (ā+Ļ)+ā(Tā”)ā1āā Ļ.āā ~normsuperscriptsuperscript1superscriptnormsuperscriptsuperscript11~normā normsuperscriptsuperscript1C( A,Ļ) Ļ( % A)Ā· \| ( A^T% A )^-1 \| \| ( A^T% A )^-1 \|^-1- Ļ(% A)Ā· ( \| A % \|+Ļ )+ \| ( A^T% A )^-1 \|Ā·Ļ.C ( A , Ļ ) ā divide start_ARG over~ start_ARG Ļ end_ARG ( A ) ā ā„ ( Aitalic_T A )- 1 ā„ end_ARG start_ARG ā„ ( Aitalic_T A )- 1 ā„- 1 - over~ start_ARG Ļ end_ARG ( A ) end_ARG ā ( ā„ A ā„ + Ļ ) + ā„ ( Aitalic_T A )- 1 ā„ ā Ļ . (10) We obtain: Proposition C.21. Assume āā¤Ļ<Ļā¢()norm\| \|ā¤Ļ<Ļ( A)ā„ Ī ā„ ā¤ Ļ < Ļ ( A ). Then Tā”superscriptsubscriptsubscript A_ ^T% A_ Abold_Īitalic_T Abold_Ī is invertible, and we have ā(Tā”)ā1ā¢Tā(Tā”)ā1ā¢Tāā¤Cā¢(,Ļ).normsuperscriptsuperscriptsubscriptsubscript1superscriptsubscriptsuperscriptsuperscript1superscript \| ( A_ ^T% A_ )^-1% A_ ^T- (% A^T A )^-1% A^T \|⤠C( A,Ļ).ā„ ( Abold_Īitalic_T Abold_Ī )- 1 Abold_Īitalic_T - ( Aitalic_T A )- 1 Aitalic_T ℠⤠C ( A , Ļ ) . Proof. The invertibility of Tā”superscriptsubscriptsubscript A_ ^T% A_ Abold_Īitalic_T Abold_Ī follows from Theorem C.18, Lemma C.19 and Lemma C.20. We get ā(Tā”)ā1ā¢Tā(Tā”)ā1ā¢Tānormsuperscriptsuperscriptsubscriptsubscript1superscriptsubscriptsuperscriptsuperscript1superscript \| ( A_ % ^T A_ )% ^-1 A_ ^T- (% A^T A )^-1% A^T \|ā„ ( Abold_Īitalic_T Abold_Ī )- 1 Abold_Īitalic_T - ( Aitalic_T A )- 1 Aitalic_T ā„ = == ā[(Tā”)ā1ā(Tā”)ā1]ā T+(Tā”)ā1ā (TāT)ānormā delimited-[]superscriptsuperscriptsubscriptsubscript1superscriptsuperscript1superscriptsubscriptā superscriptsuperscript1superscriptsubscriptsuperscript \| [ ( A_% ^T A_ % )^-1- ( A^T A% )^-1 ]Ā· A_ % ^T+ ( A^T A% )^-1Ā· ( A_ % ^T- A^T ) \|ā„ [ ( Abold_Īitalic_T Abold_Ī )- 1 - ( Aitalic_T A )- 1 ] ā Abold_Īitalic_T + ( Aitalic_T A )- 1 ā ( Abold_Īitalic_T - Aitalic_T ) ℠⤠⤠ā(Tā”)ā1ā(Tā”)ā1āā ā+ā(Tā”)ā1āā āā normsuperscriptsuperscriptsubscriptsubscript1superscriptsuperscript1normsubscriptā normsuperscriptsuperscript1norm \| ( A_ % ^T A_ )% ^-1- ( A^T A )^-1% \|Ā· \| A_ % \|+ \| ( A^T A% )^-1 \|Ā·\| \|ā„ ( Abold_Īitalic_T Abold_Ī )- 1 - ( Aitalic_T A )- 1 ā„ ā ā„ Abold_Ī ā„ + ā„ ( Aitalic_T A )- 1 ā„ ā ℠Π℠⤠⤠Ļ~ā¢()ā ā(Tā”)ā1ā(Tā”)ā1āā1āĻ~ā¢()ā (ā+Ļ)+ā(Tā”)ā1āā Ļā ~normsuperscriptsuperscript1superscriptnormsuperscriptsuperscript11~normā normsuperscriptsuperscript1 Ļ( A)Ā· \|% ( A^T A )^-1 % \| \| ( A^T A % )^-1 \|^-1- Ļ( A)Ā· (% \| A \|+Ļ )+ \| (% A^T A )^-1 \|Ā· start_ARG over~ start_ARG Ļ end_ARG ( A ) ā ā„ ( Aitalic_T A )- 1 ā„ end_ARG start_ARG ā„ ( Aitalic_T A )- 1 ā„- 1 - over~ start_ARG Ļ end_ARG ( A ) end_ARG ā ( ā„ A ā„ + Ļ ) + ā„ ( Aitalic_T A )- 1 ā„ ā Ļ = == Cā¢(,Ļ). C( A,Ļ).C ( A , Ļ ) . In the second-to-last step, we used Theorem C.18. ā The constant Cā¢(,Ļ)C( A,Ļ)C ( A , Ļ ), defined in Equation (10), has a fairly complicated form. In the following proposition, we find an easier-to-study upper bound in a special case: Proposition C.22. Assume that Ļā¤ānormĻā¤\| A\|Ļ ā¤ ā„ A ā„ and Ļā¤āā+ā2+1/2ā ā(Tā”)ā1āā1normsuperscriptnorm2ā 12superscriptnormsuperscriptsuperscript11Ļā¤-\| A\|+ \| A\|^2% +1/2Ā· \|( A^T A)^-1% \|^-1Ļ ā¤ - ā„ A ā„ + square-root start_ARG ā„ A ā„2 + 1 / 2 ā ā„ ( Aitalic_T A )- 1 ā„- 1 end_ARG.222Note the factor 1/2121/21 / 2 compared to the definition of Ļā¢()Ļ( A)Ļ ( A ) in Equation (9). Then we have Cā¢(,Ļ)ā¤Ļā ā(Tā”)ā1āā [12ā ā2ā ā(Tā”)ā1ā+1].ā normsuperscriptsuperscript1delimited-[]ā 12superscriptnorm2normsuperscriptsuperscript11C( A,Ļ)ā¤ĻĀ· \|( A% ^T A)^-1 \|Ā· [12Ā·\|% A\|^2Ā· \|( A^T% A)^-1 \|+1 ].C ( A , Ļ ) ā¤ Ļ ā ā„ ( Aitalic_T A )- 1 ā„ ā [ 12 ā ā„ A ā„2 ā ā„ ( Aitalic_T A )- 1 ā„ + 1 ] . Proof. The second assumption gives, as in the proof of Lemma C.20, that Ļ~ā¢()ā¤1/2ā ā(Tā”)ā1āā1~ā 12superscriptnormsuperscriptsuperscript11 Ļ( A)⤠1/2Ā· \|( % A^T A)^-1 \|^-1over~ start_ARG Ļ end_ARG ( A ) ⤠1 / 2 ā ā„ ( Aitalic_T A )- 1 ā„- 1. Together with Ļā¤ānormĻā¤\| A\|Ļ ā¤ ā„ A ā„, the result follows. ā C.6.2 Application to Bounds in the Error of the Return Function We now apply the results from the preceding section to our case. Define rā¢():imā”āāĪ©ā:rāimsuperscriptāāĪ©r( B):im % ā R r ( B ) : im Ī ā blackboard_Roverā start_ARG Ī© end_ARG as the restriction of the belief operator BB to imā”imim im Ī. Assume that kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 , which is, according to Corollary C.4, a sufficient condition for identifiability. Note that this condition means that rā¢()rr( B)r ( B ) is injective. Thus, Lemma C.16 ensures that rā¢()Tā¢rā¢()rsuperscriptrr( B)^Tr( B)r ( B )T r ( B ) is invertible and that (rā¢()Tā¢rā¢())ā1ā¢rā¢()Tsuperscriptrsuperscriptr1rsuperscript (r( B)^Tr(% B) )^-1r( B)^T( r ( B )T r ( B ) )- 1 r ( B )T is a left inverse of rā¢()rr( B)r ( B ). Consequently, from the equation rā¢()ā¢(G)=ā”(G)rr( B)(G)= B(G)r ( B ) ( G ) = B ( G ) we obtain G=(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tā¢(ā”(G)).superscriptrsuperscriptr1rsuperscriptG= (r( B)^Tr(% B) )^-1r( B)^T(% B(G)).G = ( r ( B )T r ( B ) )- 1 r ( B )T ( B ( G ) ) . This is the concrete formula with which G can be identified from ā”(G) B(G)B ( G ). When perturbing BB, this leads to a corresponding perturbance in (rā¢()Tā¢rā¢())ā1ā¢rā¢()Tsuperscriptrsuperscriptr1rsuperscript (r( B)^Tr(% B) )^-1r( B)^T( r ( B )T r ( B ) )- 1 r ( B )T whose size influences the maximal error in the inference of G. This, in turn, influences the size of the error in JGsubscriptJ_GJitalic_G, the policy evaluation function, where JGā¢(Ļ)āsāā¼PĻā¢(sā)[Gā¢(sā)].āsubscriptsubscriptsimilar-toāsuperscriptāJ_G(Ļ) *E_ s P^Ļ( s)% [G( s) ].Jitalic_G ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ G ( overā start_ARG s end_ARG ) ] . We obtain: Theorem C.23. Let G be the true reward function, BB the belief operator corresponding to the humanās true belief model Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ), and ā”(G) B(G)B ( G ) be the resulting observation-based return function. Assume that kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 , so that rā¢()Tā¢rā¢()rsuperscriptrr( B)^Tr( B)r ( B )T r ( B ) is invertible. Let :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© : R Sā R^% Ī : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG be a perturbation satisfying āā¤Ļnorm\| \|ā¤Ļ℠Π℠⤠Ļ, where Ļ satisfies the following two properties: Ļā¤ārā¢()ā,Ļā¤āārā¢()ā+ārā¢()ā2+1/2ā ā(rā¢()Tā¢rā¢())ā1āā1.formulae-sequencenormrnormrsuperscriptnormr2ā 12superscriptnormsuperscriptrsuperscriptr11Ļ⤠\|r( B) \|, Ļā¤-% \|r( B) \|+ \|r% ( B) \|^2+1/2Ā· \| (r(% B)^Tr( B) )^-1% \|^-1.Ļ ā¤ ā„ r ( B ) ā„ , Ļ ā¤ - ā„ r ( B ) ā„ + square-root start_ARG ā„ r ( B ) ā„2 + 1 / 2 ā ā„ ( r ( B )T r ( B ) )- 1 ā„- 1 end_ARG . Let ā+āsubscript B_ % B+ Bbold_Ī ā B + Ī be the misspecified belief operator. The first claim is that rā¢()Tā¢rā¢()rsuperscriptsubscriptrsubscriptr( B_ )^T% r( B_ )r ( Bbold_Ī )T r ( Bbold_Ī ) is invertible under these conditions. Now, assume that the learning system infers the return function G~ā(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tā¢(ā”(G))ā~superscriptrsuperscriptsubscriptrsubscript1rsuperscriptsubscript G (r( B_% )^Tr( B_% ) )^-1r( B_% )^T( B(G))over~ start_ARG G end_ARG ā ( r ( Bbold_Ī )T r ( Bbold_Ī ) )- 1 r ( Bbold_Ī )T ( B ( G ) ).333Note that there is not necessarily a G~~ Gover~ start_ARG G end_ARG with rā¢()ā¢(G~)=ā”(G)rsubscript~r( B_ )( % G)= B(G)r ( Bbold_Ī ) ( over~ start_ARG G end_ARG ) = B ( G ) since rā¢()rsubscriptr( B_ )r ( Bbold_Ī ) is not always surjective. Nevertheless, G~ā(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tā¢(ā”(G))ā~superscriptrsuperscriptsubscriptrsubscript1rsuperscriptsubscript G (r( B_% )^Tr( B_% ) )^-1r( B_% )^T( B(G))over~ start_ARG G end_ARG ā ( r ( Bbold_Ī )T r ( Bbold_Ī ) )- 1 r ( Bbold_Ī )T ( B ( G ) ) is the best attempt at a solution in the sense that rā¢()ā¢(G~)rsubscript~r( B_ )( % G)r ( Bbold_Ī ) ( over~ start_ARG G end_ARG ) then minimizes the Euclidean distance to ā”(G) B(G)B ( G ). Then there is a polynomial Qā¢(X,Y)Q(X,Y)Q ( X , Y ) of degree five such that āG~āGāā¤āGāā Qā¢(ā(rā¢()Tā¢rā¢())ā1ā,ārā¢()ā)ā Ļ.norm~ā normnormsuperscriptrsuperscriptr1normr\| G-G\|ā¤\|G\|Ā· Q ( \|(r(% B)^Tr( B))^-1 \|,\|% r( B)\| )Ā·Ļ.ā„ over~ start_ARG G end_ARG - G ℠⤠℠G ā„ ā Q ( ā„ ( r ( B )T r ( B ) )- 1 ā„ , ā„ r ( B ) ā„ ) ā Ļ . Thus, for all policies Ļ, we obtain |JG~ā¢(Ļ)āJGā¢(Ļ)|ā¤āGāā Qā¢(ā(rā¢()Tā¢rā¢())ā1ā,ārā¢()ā)ā Ļ.subscript~subscriptā normnormsuperscriptrsuperscriptr1normr |J_ G(Ļ)-J_G(Ļ) |ā¤\|G\|Ā· Q ( \|(% r( B)^Tr( B)% )^-1 \|,\|r( B)\| )Ā·Ļ.| Jover~ start_ARG G end_ARG ( Ļ ) - Jitalic_G ( Ļ ) | ⤠℠G ā„ ā Q ( ā„ ( r ( B )T r ( B ) )- 1 ā„ , ā„ r ( B ) ā„ ) ā Ļ . In particular, for sufficiently small perturbances Ļ, the error in the inferred policy evaluation function JG~subscript~J_ GJover~ start_ARG G end_ARG becomes arbitrarily small. Proof. That rā¢()Tā¢rā¢()rsuperscriptsubscriptrsubscriptr( B_ )^T% r( B_ )r ( Bbold_Ī )T r ( Bbold_Ī ) is invertible follows immediately from Proposition C.21 by using that ārā¢()āā¤ānormrnorm\|r( )\|ā¤\| % \|ā„ r ( Ī ) ℠⤠℠Π℠and that rā¢()=rā¢()rā¢()rsubscriptrsubscriptrr( B_ )= % r( B)_r( )r ( Bbold_Ī ) = r ( B )r ( Ī ), together with the second bound on Ļ (which implies the assumed bound in Proposition C.21). We have |JG~ā¢(Ļ)āJGā¢(Ļ)|subscript~subscript |J_ G(Ļ)-J_G(Ļ) || Jover~ start_ARG G end_ARG ( Ļ ) - Jitalic_G ( Ļ ) | =|sāā¼PĻā¢(sā)[(G~āG)ā¢(sā)]|absentsubscriptsimilar-toāsuperscriptā~ā = | *E_ s P^Ļ( s)% [( G-G)( s) ] |= | Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ ( over~ start_ARG G end_ARG - G ) ( overā start_ARG s end_ARG ) ] | ā¤sāā¼PĻā¢(sā)[|(G~āG)ā¢(sā)|]absentsubscriptsimilar-toāsuperscriptā~ā ⤠*E_ s P^Ļ( s) % [ |( G-G)( s) | ]⤠Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ | ( over~ start_ARG G end_ARG - G ) ( overā start_ARG s end_ARG ) | ] ā¤maxsāāāā”|(G~āG)ā¢(sā)|absentsubscriptā~ā ⤠_ sā S |( G-G)( s% ) |⤠maxoverā start_ARG s end_ARG ā overā start_ARG S end_ARG | ( over~ start_ARG G end_ARG - G ) ( overā start_ARG s end_ARG ) | ā¤āG~āGāabsentnorm~ ā¤\| G-G\|⤠℠over~ start_ARG G end_ARG - G ā„ =ā[(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tā(rā¢()Tā¢rā¢())ā1ā¢rā¢()T]ā ā”(G)āabsentnormā delimited-[]superscriptrsuperscriptsubscriptrsubscript1rsuperscriptsubscriptsuperscriptrsuperscriptr1rsuperscript = \| [ (r( B_% )^Tr( B_% ) )^-1r( B% _ )^T- (r(% B)^Tr( B) )^-1r(% B)^T ]Ā· B(G) \|= ā„ [ ( r ( Bbold_Ī )T r ( Bbold_Ī ) )- 1 r ( Bbold_Ī )T - ( r ( B )T r ( B ) )- 1 r ( B )T ] ā B ( G ) ā„ ā¤ā(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tā(rā¢()Tā¢rā¢())ā1ā¢rā¢()Tāā āā”(G)āabsentā normsuperscriptrsuperscriptsubscriptrsubscript1rsuperscriptsubscriptsuperscriptrsuperscriptr1rsuperscriptnorm ⤠\| (r( B_% )^Tr( B_% ) )^-1r( B% _ )^T- (r(% B)^Tr( B) )^-1r(% B)^T \|Ā· \| B(G% ) \|⤠℠( r ( Bbold_Ī )T r ( Bbold_Ī ) )- 1 r ( Bbold_Ī )T - ( r ( B )T r ( B ) )- 1 r ( B )T ā„ ā ā„ B ( G ) ā„ ā¤Cā¢(rā¢(),Ļ)ā ārā¢()ā¢(G)āabsentā rnormr ⤠C(r( B),Ļ)Ā·\|% r( B)(G)\|⤠C ( r ( B ) , Ļ ) ā ā„ r ( B ) ( G ) ā„ ā¤Cā¢(rā¢(),Ļ)ā ārā¢()āā āGā.absentā rnormrnorm ⤠C(r( B),Ļ)Ā·\|% r( B)\|Ā·\|G\|.⤠C ( r ( B ) , Ļ ) ā ā„ r ( B ) ā„ ā ā„ G ā„ . In the second to last step, we used Proposition C.21. By Proposition C.22, we can define the polynomial Qā¢(X,Y)Q(X,Y)Q ( X , Y ) by Qā¢(X,Y)=Xā¢Yā [12ā¢Xā¢Y2+1],ā delimited-[]12superscript21Q(X,Y)=XYĀ· [12XY^2+1 ],Q ( X , Y ) = X Y ā [ 12 X Y2 + 1 ] , which is of degree five. The last claim follows from limĻā0Ļ=0subscriptā00 _Ļā 0Ļ=0limitalic_Ļ ā 0 Ļ = 0. ā Remark C.24. In the case of a square matrix BB that is injective, we can apply Theorem C.18 directly to ā1superscript1 B^-1B- 1 (which is now invertible) and obtain the following simplification of Theorem C.23 for the case that āā¤Ļā¤12ā āā1āā1normā 12superscriptnormsuperscript11\| \|ā¤Ļ⤠12Ā·\|% B^-1\|^-1ā„ Ī ā„ ā¤ Ļ ā¤ divide start_ARG 1 end_ARG start_ARG 2 end_ARG ā ā„ B- 1 ā„- 1: |JG~ā¢(Ļ)āJGā¢(Ļ)|ā¤Ļā 2ā āā āGāā āā1ā2.subscript~subscriptā 2normnormsuperscriptnormsuperscript12 |J_ G(Ļ)-J_G(Ļ) |ā¤ĻĀ· 2Ā·\| % B\|Ā·\|G\|Ā·\| B^-1\|^2.| Jover~ start_ARG G end_ARG ( Ļ ) - Jitalic_G ( Ļ ) | ā¤ Ļ ā 2 ā ā„ B ā„ ā ā„ G ā„ ā ā„ B- 1 ā„2 . The polynomial is then only of degree 3. C.7 Preliminary Characterizations of the Ambiguity Recall the sequence of functions āsuperscriptā R^Sblackboard_RSāāsuperscriptāā R Sblackboard_Roverā start_ARG S end_ARGāĪ©ā.superscriptāāĪ© R .blackboard_Roverā start_ARG Ī© end_ARG . Ī BB In this section, we clarify imā”imim im Ī and kerā”kernel Bker B in special cases, as their intersection is the crucial ambiguity in Theorem C.2. The following proposition shows that for deterministic POāsubscriptāP_ OPoverā start_ARG O end_ARG and a rational human, kerā”kernel Bker B decomposes into hyperplanes defined by normal vectors of probabilities of sequences mapping to the same observation sequence: Proposition C.25. Assume the human reasons as in Section C.1. Assume POāsubscriptāP_ OPoverā start_ARG O end_ARG is deterministic. Let Bā¢(sā)āB( s)B ( overā start_ARG s end_ARG ) be the distribution of sequences under the humanās belief over the policy, given by Bā¢(sā)=ā«Ļā²Bā¢(Ļā²)ā¢PĻā²ā¢(sā)āsubscriptsuperscriptā²superscriptsuperscriptā²āB( s)= _Ļ B(Ļ )P^Ļ ( s)B ( overā start_ARG s end_ARG ) = ā«Ļā² B ( Ļā² ) Pitalic_Ļ start_POSTSUPERSCRIPT ā² end_POSTSUPERSCRIPT ( overā start_ARG s end_ARG ) for some policy prior Bā¢(Ļā²)superscriptā²B(Ļ )B ( Ļā² ). For each oā ooverā start_ARG o end_ARG, let Boāā[Bā¢(sā)]sā:Oāā¢(sā)=oāāāsāāāā£Oāā¢(sā)=oāāsubscriptāsubscriptdelimited-[]ā:āsuperscriptāconditional-setāB_ o [B( s)]_ s:\ O( s)= oā% R^\ sā S\ \ O( s)= o\Boverā start_ARG o end_ARG ā [ B ( overā start_ARG s end_ARG ) ]overā start_ARG s end_ARG : overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG ā blackboard_R overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG be the vector of probabilities of sequences that are observed as oā ooverā start_ARG o end_ARG. Let Gā² be a return function. For each oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG, define the restriction Goāā²āāsāāāā£Oāā¢(sā)=oāsubscriptsuperscriptā²āsuperscriptāconditional-setāG _ oā R^\ sā S O(% s)= o\Gā²overā start_ARG o end_ARG ā blackboard_R overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG by Goāā²ā¢(sā)āGā²ā¢(sā)āsubscriptsuperscriptā²āsuperscriptā²āG _ o( s) G ( s)Gā²overā start_ARG o end_ARG ( overā start_ARG s end_ARG ) ā Gā² ( overā start_ARG s end_ARG ) for all sāāsāāāā£Oāā¢(sā)=oāconditional-setā sā\ sā S O( s)= o\overā start_ARG s end_ARG ā overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG . Assume that Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) is the Bayesian posterior. Then Gā²ākerā”superscriptā²kernelG ā BGā² ā ker B if and only if the property Boāā Goāā²=0ā subscriptāsubscriptsuperscriptā²ā0B_ oĀ· G _ o=0Boverā start_ARG o end_ARG ā Gā²overā start_ARG o end_ARG = 0 holds for all oāāĪ©āĪ© oā overā start_ARG o end_ARG ā overā start_ARG Ī© end_ARG. Proof. For a deterministic observation kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG, by Bayes rule we have Bā¢(sāā£oā)conditionalā B( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) =POāā¢(oāā£sā)ā Bā¢(sā)āsāā²POāā¢(oāā£sāā²)ā Bā¢(sāā²)absentā subscriptāconditionalāsubscriptsuperscriptāā²ā subscriptāconditionalāsuperscriptāā² = P_ O( o s)Ā· B( s) _% s 1.0pt P_ O( o s 1.0pt^% )Ā· B( s 1.0pt )= divide start_ARG Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) ā B ( overā start_ARG s end_ARG ) end_ARG start_ARG āoverā start_ARG s end_ARGā² Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARGā² ) ā B ( overā start_ARG s end_ARGā² ) end_ARG =Ī“oāā¢(Oāā¢(sā))ā Bā¢(sā)āsāā²Ī“oāā¢(Oāā¢(sāā²))ā Bā¢(sāā²)absentā subscriptāsubscriptsuperscriptāā²ā subscriptāsuperscriptāā² = _ o ( O( s) )Ā· B( % s) _ s 1.0pt _ o ( O( s% 1.0pt ) )Ā· B( s 1.0pt )= divide start_ARG Ī“overā start_ARG o end_ARG ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) ā B ( overā start_ARG s end_ARG ) end_ARG start_ARG āoverā start_ARG s end_ARGā² Ī“overā start_ARG o end_ARG ( overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) ) ā B ( overā start_ARG s end_ARGā² ) end_ARG =0,Oāā¢(sā)ā oāBā¢(sā)āsāā²:Oāā¢(sāā²)=oāBā¢(sāā²),Oāā¢(sā)=oā.absentcases0āotherwiseāsubscript:superscriptāā²āsuperscriptāā²āsuperscriptāā²āotherwise = cases0,\ O( s)ā o\\ B( s) _ s 1.0pt :\ O( s 1% .0pt )= oB( s 1.0pt ),\ O( s)=% o.\ cases= start_ROW start_CELL 0 , overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ā overā start_ARG o end_ARG end_CELL start_CELL end_CELL end_ROW start_ROW start_CELL divide start_ARG B ( overā start_ARG s end_ARG ) end_ARG start_ARG āoverā start_ARG s end_ARGā² : overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) = overā start_ARG o end_ARG B ( overā start_ARG s end_ARGā² ) end_ARG , overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG . end_CELL start_CELL end_CELL end_ROW Thus, for any return function Gā² and any observation sequence oā ooverā start_ARG o end_ARG, we have [ā”(Gā²)]ā¢(oā)delimited-[]superscriptā²ā [ B(G ) ]( o)[ B ( Gā² ) ] ( overā start_ARG o end_ARG ) =sāā¼Bā¢(sāā£oā)[Gā²ā¢(sā)]absentsubscriptsimilar-toāconditionalāsuperscriptā²ā = *E_ s B( s o)% [G ( s) ]= Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ Gā² ( overā start_ARG s end_ARG ) ] =āsāBā¢(sāā£oā)ā¢Gā²ā¢(sā)absentsubscriptāconditionalāsuperscriptā²ā = _ sB( s o)G ( s)= āoverā start_ARG s end_ARG B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) Gā² ( overā start_ARG s end_ARG ) =āsā:Oāā¢(sā)=oāBā¢(sā)āsāā²:Oāā¢(sāā²)=oāBā¢(sāā²)ā¢Gā²ā¢(sā)absentsubscript:āsubscript:superscriptāā²āsuperscriptāā²āsuperscriptāā² = _ s:\ O( s)= o B( s) _% s 1.0pt :\ O( s 1.0pt )= o% B( s 1.0pt )G ( s)= āoverā start_ARG s end_ARG : overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG divide start_ARG B ( overā start_ARG s end_ARG ) end_ARG start_ARG āoverā start_ARG s end_ARGā² : overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) = overā start_ARG o end_ARG B ( overā start_ARG s end_ARGā² ) end_ARG Gā² ( overā start_ARG s end_ARG ) =(āsāā²:Oāā¢(sāā²)=oāBā¢(sāā²))ā1ā āsā:Oāā¢(sā)=oāBā¢(sā)ā¢Gā²ā¢(sā).absentā superscriptsubscript:superscriptāā²āsuperscriptāā²āsuperscriptāā²1subscript:āsuperscriptā²ā = ( _ s 1.0pt :\ O( s% 1.0pt )= oB( s 1.0pt ) )^-1% Ā· _ s:\ O( s)= oB( s)G ( s).= ( āoverā start_ARG s end_ARGā² : overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) = overā start_ARG o end_ARG B ( overā start_ARG s end_ARGā² ) )- 1 ā āoverā start_ARG s end_ARG : overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG B ( overā start_ARG s end_ARG ) Gā² ( overā start_ARG s end_ARG ) . Thus, we have Gā²ākerā”superscriptā²kernelG ā BGā² ā ker B if and only if Boāā Goāā²=āsā:Oāā¢(sā)=oāBā¢(sā)ā¢Gā²ā¢(sā)=0ā subscriptāsubscriptsuperscriptā²āsubscript:āsuperscriptā²ā0B_ oĀ· G _ o= _ s:\ O( s)= o% B( s)G ( s)=0Boverā start_ARG o end_ARG ā Gā²overā start_ARG o end_ARG = āoverā start_ARG s end_ARG : overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = overā start_ARG o end_ARG B ( overā start_ARG s end_ARG ) Gā² ( overā start_ARG s end_ARG ) = 0 for all oā ooverā start_ARG o end_ARG. That was to show. ā Remark C.26. One can interpret the previous proposition as follows: As long as Oā Ooverā start_ARG O end_ARG is injective, we have |sāāāā£Oāā¢(sā)=o|=1conditional-setā1 |\ sā S O( s)=o\ |=1| overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = o | = 1 for all oā ooverā start_ARG o end_ARG, meaning that BoāsubscriptāB_ oBoverā start_ARG o end_ARG and Goāā²subscriptsuperscriptā²āG _ oGā²overā start_ARG o end_ARG have only one entry. Thus, Boāā Goāā²=0ā subscriptāsubscriptsuperscriptā²ā0B_ oĀ· G _ o=0Boverā start_ARG o end_ARG ā Gā²overā start_ARG o end_ARG = 0 implies Goāā²=0subscriptsuperscriptā²ā0G _ o=0Gā²overā start_ARG o end_ARG = 0. If that holds for all oā ooverā start_ARG o end_ARG, then Gā²ākerā”superscriptā²kernelG ā BGā² ā ker B implies Gā²=0superscriptā²0G =0Gā² = 0, meaning BB is injective. However, as soon as there is an oā ooverā start_ARG o end_ARG with koāā|sāāāā£Oāā¢(sā)=o|>1āsubscriptāconditional-setā1k_ o |\ sā S O( s)=o% \ |>1koverā start_ARG o end_ARG ā | overā start_ARG s end_ARG ā overā start_ARG S end_ARG ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) = o | > 1, the equation Boāā Goāā²=0ā subscriptāsubscriptsuperscriptā²ā0B_ oĀ· G _ o=0Boverā start_ARG o end_ARG ā Gā²overā start_ARG o end_ARG = 0 leads to koāā1subscriptā1k_ o-1koverā start_ARG o end_ARG - 1 free parameters in Goāā²subscriptsuperscriptā²āG _ oGā²overā start_ARG o end_ARG. Goāā²subscriptsuperscriptā²āG _ oGā²overā start_ARG o end_ARG can then be chosen freely in the hyperplane of vectors orthogonal to BoāsubscriptāB_ oBoverā start_ARG o end_ARG without moving out of the kernel of BB. Another way of writing Proposition C.25 is to write kerā”kernel Bker B as a direct sum of these hyperplanes perpendicular to BoāsubscriptāB_ oBoverā start_ARG o end_ARG: kerā”=āØoā:|Oāā1ā¢(oā)|ā„2Boāā.kernelsubscriptdirect-sum:āsuperscriptā1ā2superscriptsubscriptāperpendicular-to B= _ o:\ | O^-1( o)|ā„ 2% B_ o .ker B = āØoverā start_ARG o end_ARG : | overā start_ARG O end_ARG- 1 ( overā start_ARG o end_ARG ) | ā„ 2 Boverā start_ARG o end_ARGā . Recall that a return function G is called time-separable if there exists a reward function R such that ā”(R)=G (R)=GĪ ( R ) = G. Before we discuss time-separability in more interesting examples, we want to talk about one simple case where all return functions are time-separable. We leave a general characterization of imā”imim im Ī to future work. Proposition C.27. Let there be an ordering sā(1),sā(2),ā¦superscriptā1superscriptā2italic-⦠s^(1), s^(2),ā¦overā start_ARG s end_ARG( 1 ) , overā start_ARG s end_ARG( 2 ) , ⦠of all sequences in ā Soverā start_ARG S end_ARG, and a function Ļ:ā:italic-ĻāĻ: S Ļ : overā start_ARG S end_ARG ā S from sequences to states such that Ļā¢(sā)āsāitalic-ĻāĻ( s)ā sĻ ( overā start_ARG s end_ARG ) ā overā start_ARG s end_ARG and Ļā¢(sā(k))āsā(i)italic-ĻsuperscriptāsuperscriptāĻ( s^(k))ā s^(i)Ļ ( overā start_ARG s end_ARG( k ) ) ā overā start_ARG s end_ARG( i ) for all i<ki<ki < k. Then every return function is time-separable. Proof. Let G be a return function. Initialize Rā¢(s)=00R(s)=0R ( s ) = 0 for all s and inductively update it for all i=1,2,ā¦12ā¦i=1,2,ā¦i = 1 , 2 , ā¦: Rā¢(Ļā¢(sā(i)))italic-Ļsuperscriptā R (Ļ( s^(i)) )R ( Ļ ( overā start_ARG s end_ARG( i ) ) ) ā(āt:st(i)=Ļā¢(sā(i))γt)ā1ā (Gā¢(sā(i))āāt:st(i)ā Ļā¢(sā(i))γtā Rā¢(st(i))),āabsentā superscriptsubscript:subscriptsuperscriptitalic-Ļsuperscriptāsuperscript1superscriptāsubscript:subscriptsuperscriptitalic-Ļsuperscriptāā superscriptsubscriptsuperscript ( _t:\ s^(i)_t=Ļ( s^(i))γ% ^t )^-1Ā· (G( s^(i))- _t:\ s^(i)_tā Ļ(% s^(i))γ^tĀ· R (s^(i)_t ) ),ā ( āt : s( i ) start_POSTSUBSCRIPT t = Ļ ( overā start_ARG s end_ARG( i ) ) end_POSTSUBSCRIPT γitalic_t )- 1 ā ( G ( overā start_ARG s end_ARG( i ) ) - āt : s( i ) start_POSTSUBSCRIPT t ā Ļ ( overā start_ARG s end_ARG( i ) ) end_POSTSUBSCRIPT γitalic_t ā R ( s( i )t ) ) , where the inductive definition always uses R as it is defined by that point in time. Once Rā¢(Ļā¢(sā(i)))italic-ĻsuperscriptāR (Ļ( s^(i)) )R ( Ļ ( overā start_ARG s end_ARG( i ) ) ) is defined, but not yet any future values Rā¢(Ļā¢(sā(k)))italic-ĻsuperscriptāR (Ļ( s^(k)) )R ( Ļ ( overā start_ARG s end_ARG( k ) ) ), k>ik>ik > i, we have [ā”(R)]ā¢(sā(i))delimited-[]superscriptā [ (R) ]( s^(i))[ Ī ( R ) ] ( overā start_ARG s end_ARG( i ) ) =āt=0Tγtā Rā¢(st(i))absentsuperscriptsubscript0ā superscriptsubscriptsuperscript = _t=0^Tγ^tĀ· R (s^(i)_t )= āt = 0T γitalic_t ā R ( s( i )t ) =(āt:st(i)=Ļā¢(sā(i))γt)ā Rā¢(Ļā¢(sā(i)))+āt:st(i)ā Ļā¢(sā(i))γtā Rā¢(st(i))absentā subscript:subscriptsuperscriptitalic-Ļsuperscriptāsuperscriptitalic-Ļsuperscriptāsubscript:subscriptsuperscriptitalic-Ļsuperscriptāā superscriptsubscriptsuperscript = ( _t:\ s^(i)_t=Ļ( s^(i))γ^t% )Ā· R (Ļ( s^(i)) )+ _t:\ s^(i)_tā Ļ% ( s^(i))γ^tĀ· R (s^(i)_t )= ( āt : s( i ) start_POSTSUBSCRIPT t = Ļ ( overā start_ARG s end_ARG( i ) ) end_POSTSUBSCRIPT γitalic_t ) ā R ( Ļ ( overā start_ARG s end_ARG( i ) ) ) + āt : s( i ) start_POSTSUBSCRIPT t ā Ļ ( overā start_ARG s end_ARG( i ) ) end_POSTSUBSCRIPT γitalic_t ā R ( s( i )t ) =Gā¢(sā(i)).absentsuperscriptā =G( s^(i)).= G ( overā start_ARG s end_ARG( i ) ) . Furthermore, the property Ļā¢(sā(k))āsā(i)italic-ĻsuperscriptāsuperscriptāĻ( s^(k))ā s^(i)Ļ ( overā start_ARG s end_ARG( k ) ) ā overā start_ARG s end_ARG( i ) for all i<ki<ki < k ensures that changes to the reward function for k>ik>ik > i do not affect the value of [ā”(R)]ā¢(sā(i))delimited-[]superscriptā [ (R) ]( s^(i))[ Ī ( R ) ] ( overā start_ARG s end_ARG( i ) ). This shows ā”(R)=G (R)=GĪ ( R ) = G, and thus G is time-separable. ā Corollary C.28. In a multi-armed bandit, every return function is time-separable. Proof. In a multi-armed bandit, states and sequences are equivalent, and so we can choose Ļā¢(s)=sitalic-ĻĻ(s)=sĻ ( s ) = s for every state/sequence s. The result follows from Proposition C.27. Alternatively, simply directly notice that in a multi-armed bandit, Ī is the identity mapping, and so for every return/reward function R, we have ā”(R)=R (R)=RĪ ( R ) = R. ā C.8 Examples Supplementing Section 5 In this whole section, the inverse temperature parameter in the human choice probabilities is given by β=11β=1β = 1. We now consider four more mathematical examples of Corollary C.4 and Theorem C.9. In the first example, the ambiguity is so bad that the reward inference can become worse than simply maximizing JobssubscriptobsJ_obsJroman_obs as in naive RLHF. In Example C.30, there is simply ānoiseā in the observations and the humanās belief, the matrices BB and OO are injective, and identifiability works, as in Corollary C.14. In the third example, the matrix BB is not injective and identifiability fails, which is a minimal example showing the limits of our main theorems. In the fourth example, the matrix BB is not injective, but kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 , and so identifiability works. This example is interesting in that the identifiability simply emerges through different distributions of delay that are caused by the different unobserved events. In this section, both the linear operators :āāāĪ©ā:āsuperscriptāāsuperscriptāāĪ© B: R Sā R % B : blackboard_Roverā start_ARG S end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG and :āĪ©āāā:āsuperscriptāāĪ©superscriptāā O: R ā R % SO : blackboard_Roverā start_ARG Ī© end_ARG ā blackboard_Roverā start_ARG S end_ARG are considered as matrices =(POāā¢(oāā£sā))sā,oāāāāĆĪ©ā,=(Bā¢(sāā£oā))oā,sāāāĪ©āĆā.formulae-sequencesubscriptsubscriptāconditionalāsuperscriptāāĪ©subscriptconditionalāsuperscriptāāĪ©ā O= (P_ O( o s) )_ % s, oā R SĆ , % B= (B( s o) )_ o, s% ā R Ć S.O = ( Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) )overā start_ARG s end_ARG , overā start_ARG o end_ARG ā blackboard_Roverā start_ARG S end_ARG Ć overā start_ARG Ī© end_ARG , B = ( B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) )overā start_ARG o end_ARG , overā start_ARG s end_ARG ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG S end_ARG . Notice that both have a swap in their indices. Example C.29. Theorem 5.2 shows that the remaining ambiguity from the humanās choice probabilities is given by kerā”ā©imā”kernelim B ker B ā© im Ī, but it doesnāt explain how to proceed given this ambiguity. Without further inductive biases, some reward functions within the ambiguity of the true reward function can be even worse than simply maximizing JobssubscriptobsJ_obsJroman_obs. E.g., consider a multi-armed bandit with three actions a,b,ca,b,ca , b , c, observation-kernel o=Oā¢(a)=Oā¢(b)ā Oā¢(c)=co=O(a)=O(b)ā O(c)=co = O ( a ) = O ( b ) ā O ( c ) = c and reward function Rā¢(a)=Rā¢(b)<Rā¢(c)R(a)=R(b)<R(c)R ( a ) = R ( b ) < R ( c ). If the human belief is given by Bā¢(aā£o)=p=1āBā¢(bā£o)conditional1conditionalB(a o)=p=1-B(b o)B ( a ⣠o ) = p = 1 - B ( b ⣠o ), then Rā²=αā (pā1,p,0)āāa,b,csuperscriptā²ā 10superscriptāR =α·(p-1,p,0)ā R^\a,b,c\Rā² = α ā ( p - 1 , p , 0 ) ā blackboard_R a , b , c is in the ambiguity for all αāāαā Rα ā blackboard_R, and so R~āR+Rā²ā~superscriptā² R R+R over~ start_ARG R end_ARG ā R + Rā² is compatible with the choice probabilities. However, for αāŖ0much-less-than0α 0α āŖ 0, we have R~ā¢(a)>R~ā¢(b)~~ R(a)> R(b)over~ start_ARG R end_ARG ( a ) > over~ start_ARG R end_ARG ( b ) and R~ā¢(a)>R~ā¢(c)~~ R(a)> R(c)over~ start_ARG R end_ARG ( a ) > over~ start_ARG R end_ARG ( c ), and so optimizing against this reward function leads to a suboptimal policy. In contrast, maximizing JobssubscriptobsJ_obsJroman_obs leads to the correct policy since a, b, and c all obtain their ground truth reward in this example. This generally raises the question of how to tie-break reward functions in the ambiguity, or how to act conservatively given the uncertainty, in order to consistently improve upon the setting in Section 4.1. Example C.30. This example is a special case of Corollary C.14. Consider a multi-armed bandit with two actions (which are automatically also states and sequences) a and b. In this case, the reward function and return function is the same. We assume there to be two possible observations o(a),o(b)superscriptsuperscripto^(a),o^(b)o( a ) , o( b ) and the observation kernel to be non-deterministic, with probabilities POā¢(o(j)ā£i)=2/3, if ā¢i=j,1/3, else.subscriptconditionalsuperscriptcases23 if otherwise13 elseotherwiseP_O(o^(j) i)= cases2/3, if i=j,\\ 1/3, else. casesPitalic_O ( o( j ) ⣠i ) = start_ROW start_CELL 2 / 3 , if i = j , end_CELL start_CELL end_CELL end_ROW start_ROW start_CELL 1 / 3 , else . end_CELL start_CELL end_CELL end_ROW If we assume the human forms Bayesian posterior beliefs as in Section C.1 and to have a policy prior Bā¢(Ļā²)superscriptā²B(Ļ )B ( Ļā² ) such that Bā¢(a)=ā«Ļā¢(a)ā¢Bā¢(Ļā²)ā¢Ļ=1/2subscriptsuperscriptā²differential-d12B(a)= _Ļ(a)B(Ļ )dĻ=1/2B ( a ) = ā«Ļ Ļ ( a ) B ( Ļā² ) d Ļ = 1 / 2 and Bā¢(b)=1/212B(b)=1/2B ( b ) = 1 / 2, then it is easy to show that the humanās belief is the āreversedā observation kernel: Bā¢(jā£o(i))=POā¢(o(i)ā£j).conditionalsuperscriptsubscriptconditionalsuperscriptB(j o^(i))=P_O(o^(i) j).B ( j ⣠o( i ) ) = Pitalic_O ( o( i ) ⣠j ) . We obtain OO ==(2/31/31/32/3)=13ā (2112)absentmatrix23131323ā 13matrix2112 = B= pmatrix2/3&1/3\\ 1/3&2/3 pmatrix= 13Ā· pmatrix2&1\\ 1&2 pmatrix= B = ( start_ARG start_ROW start_CELL 2 / 3 end_CELL start_CELL 1 / 3 end_CELL end_ROW start_ROW start_CELL 1 / 3 end_CELL start_CELL 2 / 3 end_CELL end_ROW end_ARG ) = divide start_ARG 1 end_ARG start_ARG 3 end_ARG ā ( start_ARG start_ROW start_CELL 2 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 2 end_CELL end_ROW end_ARG ) These matrices are injective since they are invertible: ā1=ā1=(2ā1ā12).superscript1superscript1matrix2112 O^-1= B^-1= pmatrix2% &-1\\ -1&2 pmatrix.O- 1 = B- 1 = ( start_ARG start_ROW start_CELL 2 end_CELL start_CELL - 1 end_CELL end_ROW start_ROW start_CELL - 1 end_CELL start_CELL 2 end_CELL end_ROW end_ARG ) . More generally, even if the human does not form fully rational posterior beliefs, it is easy to imagine that the matrix BB can end up being invertible. Thus, Corollary C.4 guarantees that the reward function can be inferred up to an additive constant from the choice probabilities of observations, and Theorem C.9 shows that this even works when the learning system does not know what the human observed. In the rest of this example, we explicitly walk the reader through the process of how the reward function can be inferred, in the general case that the observations are not known. In the process, we essentially recreate the proof of the theorems for this special case. For this aim, we first want to compute the choice probabilities PRā¢(iā»j)superscriptsucceedsP^R (i j )Pitalic_R ( i ā» j ) that the learning system has access to in the limit of infinite data. We assume that the reward function is given by Rā¢(a)=ā11R(a)=-1R ( a ) = - 1 and Rā¢(b)=22R(b)=2R ( b ) = 2. We compute: ā”(R)=13ā (2112)ā (ā12)=(01).ā 13matrix2112matrix12matrix01 B(R)= 13Ā· pmatrix2&1\\ 1&2 pmatrixĀ· pmatrix-1\\ 2 pmatrix= pmatrix0\\ 1 pmatrix.B ( R ) = divide start_ARG 1 end_ARG start_ARG 3 end_ARG ā ( start_ARG start_ROW start_CELL 2 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 2 end_CELL end_ROW end_ARG ) ā ( start_ARG start_ROW start_CELL - 1 end_CELL end_ROW start_ROW start_CELL 2 end_CELL end_ROW end_ARG ) = ( start_ARG start_ROW start_CELL 0 end_CELL end_ROW start_ROW start_CELL 1 end_CELL end_ROW end_ARG ) . In other words, we have sā¼Bā¢(sā£o(a))[Rā¢(s)]=0subscriptsimilar-toconditionalsuperscript0 *E_s B(s o^(a))[R(s)]=0Eitalic_s ā¼ B ( s ⣠o( a ) ) [ R ( s ) ] = 0 and sā¼Bā¢(sā£o(b))[Rā¢(s)]=1subscriptsimilar-toconditionalsuperscript1 *E_s B(s o^(b))[R(s)]=1Eitalic_s ā¼ B ( s ⣠o( b ) ) [ R ( s ) ] = 1. From this, we can compute the observation-based choice probabilities P~o(i)ā¢o(j)=Ļā¢(ā”(R)ā¢(o(i))āā”(R)ā¢(o(j)))subscript~superscriptsuperscriptsuperscriptsuperscript P_o^(i)o^(j)=Ļ ( B(R)(o^(i% ))- B(R)(o^(j)) )over~ start_ARG P end_ARGo( i ) o( j ) = Ļ ( B ( R ) ( o( i ) ) - B ( R ) ( o( j ) ) ), see Equation (2), and obtain: P~o(a)ā¢o(a)=P~o(b)ā¢o(b)=12,P~o(a)ā¢o(b)=11+e,P~o(b)ā¢o(a)=e1+e.formulae-sequencesubscript~superscriptsuperscriptsubscript~superscriptsuperscript12formulae-sequencesubscript~superscriptsuperscript11subscript~superscriptsuperscript1 P_o^(a)o^(a)= P_o^(b)o^(b)= 12,% P_o^(a)o^(b)= 11+e, P_o^(b)o% ^(a)= e1+e.over~ start_ARG P end_ARGo( a ) o( a ) = over~ start_ARG P end_ARGo( b ) o( b ) = divide start_ARG 1 end_ARG start_ARG 2 end_ARG , over~ start_ARG P end_ARGo( a ) o( b ) = divide start_ARG 1 end_ARG start_ARG 1 + e end_ARG , over~ start_ARG P end_ARGo( b ) o( a ) = divide start_ARG e end_ARG start_ARG 1 + e end_ARG . We can now determine the final choice probabilities Piā¢jāPRā¢(iā»j)āsubscriptsuperscriptsucceedsP_ij P^R (i j )Pitalic_i j ā Pitalic_R ( i ā» j ) again by a matrix-vector product, with the indices ordered lexicographically, see Equation (8). Here, ātensor-product O OO ā O is the Kronecker product of the matrix OO with itself: P=(ā)ā P~=19ā (4221241221421224)ā (1/21/(1+e)e/(1+e)1/2)=(1/21/3ā (2+e)/(1+e)1/3ā (1+2ā¢e)/(1+e)1/2).ā tensor-product~ā 19matrix4221241221421224matrix1211112matrix12ā 1321ā 1312112P=( O O)Ā· P% = 19Ā· pmatrix4&2&2&1\\ 2&4&1&2\\ 2&1&4&2\\ 1&2&2&4 pmatrixĀ· pmatrix1/2\\ 1/(1+e)\\ e/(1+e)\\ 1/2 pmatrix= pmatrix1/2\\ 1/3Ā·(2+e)/(1+e)\\ 1/3Ā·(1+2e)/(1+e)\\ 1/2 pmatrix.P = ( O ā O ) ā over~ start_ARG P end_ARG = divide start_ARG 1 end_ARG start_ARG 9 end_ARG ā ( start_ARG start_ROW start_CELL 4 end_CELL start_CELL 2 end_CELL start_CELL 2 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL 2 end_CELL start_CELL 4 end_CELL start_CELL 1 end_CELL start_CELL 2 end_CELL end_ROW start_ROW start_CELL 2 end_CELL start_CELL 1 end_CELL start_CELL 4 end_CELL start_CELL 2 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 2 end_CELL start_CELL 2 end_CELL start_CELL 4 end_CELL end_ROW end_ARG ) ā ( start_ARG start_ROW start_CELL 1 / 2 end_CELL end_ROW start_ROW start_CELL 1 / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL e / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 2 end_CELL end_ROW end_ARG ) = ( start_ARG start_ROW start_CELL 1 / 2 end_CELL end_ROW start_ROW start_CELL 1 / 3 ā ( 2 + e ) / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 3 ā ( 1 + 2 e ) / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 2 end_CELL end_ROW end_ARG ) . For example, the second entry in P is Paā¢b=PRā¢(aā»b)=2+e3ā (1+e)subscriptsuperscriptsucceeds2ā 31P_ab=P^R (a b )= 2+e3Ā·(1+e)Pitalic_a b = Pitalic_R ( a ā» b ) = divide start_ARG 2 + e end_ARG start_ARG 3 ā ( 1 + e ) end_ARG. This is the likelihood that, for ground-truth actions a,ba,ba , b, the human will prefer a after only receiving observations o(a)superscripto^(a)o( a ) or o(b)superscripto^(b)o( b ) according to OO and following a Boltzman-rational policy based on the belief of the real action, see Equation (8). Over time, the learning system will be able to estimate these probabilities based on repeated human choices, assuming all state-pairs are sampled infinitely often. The question of identifiability is whether the original reward function R can be inferred from that data, given that the learning system knows OO and BB. We assume that the learning system doesnāt a priori know R or any of the intermediate steps in the computation. First, P~~ Pover~ start_ARG P end_ARG can be inferred by inverting ātensor-product O OO ā O: P~=(ā)ā1ā P=(4ā2ā21ā241ā2ā214ā21ā2ā24)ā (1/21/3ā (2+e)/(1+e)1/3ā (1+2ā¢e)/(1+e)1/2)=(1/21/(1+e)e/(1+e)1/2).~ā superscripttensor-product1ā matrix4221241221421224matrix12ā 1321ā 1312112matrix1211112 P=( O O)^-1% Ā· P= pmatrix4&-2&-2&1\\ -2&4&1&-2\\ -2&1&4&-2\\ 1&-2&-2&4 pmatrixĀ· pmatrix1/2\\ 1/3Ā·(2+e)/(1+e)\\ 1/3Ā·(1+2e)/(1+e)\\ 1/2 pmatrix= pmatrix1/2\\ 1/(1+e)\\ e/(1+e)\\ 1/2 pmatrix.over~ start_ARG P end_ARG = ( O ā O )- 1 ā P = ( start_ARG start_ROW start_CELL 4 end_CELL start_CELL - 2 end_CELL start_CELL - 2 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL - 2 end_CELL start_CELL 4 end_CELL start_CELL 1 end_CELL start_CELL - 2 end_CELL end_ROW start_ROW start_CELL - 2 end_CELL start_CELL 1 end_CELL start_CELL 4 end_CELL start_CELL - 2 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL - 2 end_CELL start_CELL - 2 end_CELL start_CELL 4 end_CELL end_ROW end_ARG ) ā ( start_ARG start_ROW start_CELL 1 / 2 end_CELL end_ROW start_ROW start_CELL 1 / 3 ā ( 2 + e ) / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 3 ā ( 1 + 2 e ) / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 2 end_CELL end_ROW end_ARG ) = ( start_ARG start_ROW start_CELL 1 / 2 end_CELL end_ROW start_ROW start_CELL 1 / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL e / ( 1 + e ) end_CELL end_ROW start_ROW start_CELL 1 / 2 end_CELL end_ROW end_ARG ) . The learning system wants to use this to infer ā”(R~)~ B( R)B ( over~ start_ARG R end_ARG ) (for the later-to-be inferred reward function R~~ Rover~ start_ARG R end_ARG that may differ from the true reward function R) and uses the equation P~o(a)ā¢o(b)=expā”(ā”(R~)ā¢(o(a)))expā”(ā”(R~)ā¢(o(a)))+expā”(ā”(R~)ā¢(o(b))),subscript~superscriptsuperscript~superscript~superscript~superscript P_o^(a)o^(b)= ( B(% R)(o^(a)) ) ( B( R)(o^% (a)) )+ ( B( R)(o^(b)) ),over~ start_ARG P end_ARGo( a ) o( b ) = divide start_ARG exp ( B ( over~ start_ARG R end_ARG ) ( o( a ) ) ) end_ARG start_ARG exp ( B ( over~ start_ARG R end_ARG ) ( o( a ) ) ) + exp ( B ( over~ start_ARG R end_ARG ) ( o( b ) ) ) end_ARG , which can be rearranged to ā”(R~)ā¢(o(a))=logā”P~o(a)ā¢o(b)1āP~o(a)ā¢o(b)+ā”(R~)ā¢(o(b))=logā”1/(1+e)e/(1+e)+ā”(R~)ā¢(o(b))=ā”(R~)ā¢(o(b))ā1.~superscriptsubscript~superscriptsuperscript1subscript~superscriptsuperscript~superscript111~superscript~superscript1 B( R)(o^(a))= P_o^(a)% o^(b)1- P_o^(a)o^(b)+ B( R% )(o^(b))= 1/(1+e)e/(1+e)+ B( R)(o% ^(b))= B( R)(o^(b))-1.B ( over~ start_ARG R end_ARG ) ( o( a ) ) = log divide start_ARG over~ start_ARG P end_ARGo( a ) o( b ) end_ARG start_ARG 1 - over~ start_ARG P end_ARGo( a ) o( b ) end_ARG + B ( over~ start_ARG R end_ARG ) ( o( b ) ) = log divide start_ARG 1 / ( 1 + e ) end_ARG start_ARG e / ( 1 + e ) end_ARG + B ( over~ start_ARG R end_ARG ) ( o( b ) ) = B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 . This relation is all which can be inferred about ā”(R~)ā¢(o(a))~superscript B( R)(o^(a))B ( over~ start_ARG R end_ARG ) ( o( a ) ) and ā”(R~)ā¢(o(b))~superscript B( R)(o^(b))B ( over~ start_ARG R end_ARG ) ( o( b ) ); the precise value cannot be determined and ā”(R~)ā¢(o(b))~superscript B( R)(o^(b))B ( over~ start_ARG R end_ARG ) ( o( b ) ) is a free parameter. One can check that for ā”(R~)ā¢(o(b))=1~superscript1 B( R)(o^(b))=1B ( over~ start_ARG R end_ARG ) ( o( b ) ) = 1 this coincides with the true value ā”(R) B(R)B ( R ). Finally, one can invert BB to infer R~~ Rover~ start_ARG R end_ARG from this: R~~ Rover~ start_ARG R end_ARG =ā1ā ā”(R~)absentā superscript1~ = B^-1Ā· B(% R)= B- 1 ā B ( over~ start_ARG R end_ARG ) =(2ā1ā12)ā (ā”(R~)ā¢(o(b))ā1ā”(R~)ā¢(o(b)))absentā matrix2112matrix~superscript1~superscript = pmatrix2&-1\\ -1&2 pmatrixĀ· pmatrix B( R)(o^(% b))-1\\ B( R)(o^(b)) pmatrix= ( start_ARG start_ROW start_CELL 2 end_CELL start_CELL - 1 end_CELL end_ROW start_ROW start_CELL - 1 end_CELL start_CELL 2 end_CELL end_ROW end_ARG ) ā ( start_ARG start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 end_CELL end_ROW start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) end_CELL end_ROW end_ARG ) =(ā”(R~)ā¢(o(b))ā21+ā”(R~)ā¢(o(b)))absentmatrix~superscript21~superscript = pmatrix B( R)(o^(b))-2\\ 1+ B( R)(o^(b)) pmatrix= ( start_ARG start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 2 end_CELL end_ROW start_ROW start_CELL 1 + B ( over~ start_ARG R end_ARG ) ( o( b ) ) end_CELL end_ROW end_ARG ) =(ā12)+(ā”(R~)ā¢(o(b))ā1ā”(R~)ā¢(o(b))ā1)absentmatrix12matrix~superscript1~superscript1 = pmatrix-1\\ 2 pmatrix+ pmatrix B( R)(o^(b))-1% \\ B( R)(o^(b))-1 pmatrix= ( start_ARG start_ROW start_CELL - 1 end_CELL end_ROW start_ROW start_CELL 2 end_CELL end_ROW end_ARG ) + ( start_ARG start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 end_CELL end_ROW start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 end_CELL end_ROW end_ARG ) =R+(ā”(R~)ā¢(o(b))ā1ā”(R~)ā¢(o(b))ā1).absentmatrix~superscript1~superscript1 =R+ pmatrix B( R)(o^(b))-1% \\ B( R)(o^(b))-1 pmatrix.= R + ( start_ARG start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 end_CELL end_ROW start_ROW start_CELL B ( over~ start_ARG R end_ARG ) ( o( b ) ) - 1 end_CELL end_ROW end_ARG ) . Thus, the inferred and true reward functions differ maximally by a constant, as predicted in Theorem C.9. In the following example, we work out a case where the reward function is so ambiguous that any policy is optimal to some reward function consistent with the human feedback: Example C.31. Consider a multi-armed bandit with exactly three actions/states a,b,ca,b,ca , b , c. We assume a deterministic observation kernel with oāOā¢(a)=Oā¢(c)ā Oā¢(b)=bāo O(a)=O(c)ā O(b)=bo ā O ( a ) = O ( c ) ā O ( b ) = b. Assume the human has some arbitrary beliefs Bā¢(aā£o),Bā¢(cā£o)=1āBā¢(aā£o)conditionalconditional1conditionalB(a o),B(c o)=1-B(a o)B ( a ⣠o ) , B ( c ⣠o ) = 1 - B ( a ⣠o ), and can identify b: Bā¢(bā£b)=1conditional1B(b b)=1B ( b ⣠b ) = 1. Then if the human makes observation comparisons with a Boltzman-rational policy, as in Theorem C.2, the resulting reward function is so ambiguous that some reward functions consistent with the feedback place the highest value on action a, no matter the true reward function R. Thus, even if the true reward function R regards a as the worst action, a can result from the reward learning and subsequent policy optimization process. Proof. The matrix :āa,b,cāāo,b:āsuperscriptāsuperscriptā B: R^\a,b,c\ā R^\o,b\B : blackboard_R a , b , c ā blackboard_R o , b is given by =(Bā¢(aā£o)0Bā¢(cā£o)010).matrixconditional0conditional010 B= pmatrixB(a o)&0&B(c o)\\ 0&1&0 pmatrix.B = ( start_ARG start_ROW start_CELL B ( a ⣠o ) end_CELL start_CELL 0 end_CELL start_CELL B ( c ⣠o ) end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL end_ROW end_ARG ) . Its kernel is given by reward functions Rā² with Rā²ā¢(b)=0superscriptā²0R (b)=0Rā² ( b ) = 0 and Rā²ā¢(c)=āBā¢(aā£o)Bā¢(cā£o)ā¢Rā²ā¢(a)superscriptā²conditionalconditionalsuperscriptā²R (c)=- B(a o)B(c o)R (a)Rā² ( c ) = - divide start_ARG B ( a ⣠o ) end_ARG start_ARG B ( c ⣠o ) end_ARG Rā² ( a ), with Rā²ā¢(a)superscriptā²R (a)Rā² ( a ) a free parameter. Theorem C.2 shows that, up to an additive constant, the reward functions consistent with the feedback of observation comparisons are given by R~=R+Rā²~superscriptā² R=R+R over~ start_ARG R end_ARG = R + Rā² for any Rā²ākerā”superscriptā²kernelR ā BRā² ā ker B. Thus, whenever the free parameter Rā²ā¢(a)superscriptā²R (a)Rā² ( a ) satisfies Rā²ā¢(a)>Rā¢(b)āRā¢(a)superscriptā²R (a)>R(b)-R(a)Rā² ( a ) > R ( b ) - R ( a ) and Rā²ā¢(a)>Bā¢(cā£o)ā (Rā¢(c)āRā¢(a))superscriptā²ā conditionalR (a)>B(c o)Ā· (R(c)-R(a) )Rā² ( a ) > B ( c ⣠o ) ā ( R ( c ) - R ( a ) ), we obtain R~ā¢(a)>R~ā¢(b)~~ R(a)> R(b)over~ start_ARG R end_ARG ( a ) > over~ start_ARG R end_ARG ( b ) and R~ā¢(a)>R~ā¢(c)~~ R(a)> R(c)over~ start_ARG R end_ARG ( a ) > over~ start_ARG R end_ARG ( c ), showing the claim. ā We now investigate another example where BB is not injective, and yet, identifiability works because āā 00 B ā \0\B ā Ī ā 0 . We saw such cases already in Example D.6, but include this additional example since it shows a conceptually interesting case: two different states lead to the exact same observations, but can be disambiguated since they lead to different amounts of delay until a more informative observation is made again. Example C.32. In this example, we assume that the human knows the policy Ļ that generates the state sequences (corresponding to a policy prior Bā¢(Ļā²)=Ī“Ļā¢(Ļā²)superscriptā²subscriptsuperscriptā²B(Ļ )= _Ļ(Ļ )B ( Ļā² ) = Ī“italic_Ļ ( Ļā² ) concentrated on Ļ), which together with knowledge of the transition dynamics of the environment determines the true state transition probabilities Ļā¢(sā²ā£s)=āaāā¢(sā²ā£s,a)ā Ļā¢(aā£s)superscriptconditionalsuperscriptā²subscriptā conditionalsuperscriptā²conditionalT^Ļ(s s)= _a T(s^% s,a)Ā·Ļ(a s)Titalic_Ļ ( sⲠ⣠s ) = āa ā A T ( sⲠ⣠s , a ) ā Ļ ( a ⣠s ). We consider an environment with three states s,sā²,sā²superscriptā²s,s ,s s , sā² , sā² ā² and the following transition dynamics ĻsuperscriptT^ĻTitalic_Ļ, where pā 1/212pā 1/2p ā 1 / 2 is a probability: sssā²s sā²sā²s sā² ā²1/313 1/31 / 31/313 1/31 / 31/313 1/31 / 31āp1 1-p1 - p pp pp1āp1 1-p1 - p We assume that P0ā¢(s)=1subscript01P_0(s)=1P0 ( s ) = 1. Furthermore, we assume deterministic observations and s=Oā¢(s)ā Oā¢(sā²)=Oā¢(sā²)āosuperscriptā²ās=O(s)ā O(s )=O(s ) os = O ( s ) ā O ( sā² ) = O ( sā² ā² ) ā o. Assume the time horizon T is 3333, i.e., there are timesteps 0,1,2,301230,1,2,30 , 1 , 2 , 3. Assume that the human forms the belief over the true state sequence by Bayesian posterior updates as in Section C.1. In this case, kerā”ā 0kernel0 Bā \0\ker B ā 0 by Proposition C.11. However, we will now show that kerā”(ā)=0kernel0 ( B )=\0\ker ( B ā Ī ) = 0 . If the human makes Boltzmann-rational comparisons of observation sequences, then this implies the identifiability of the return function up to an additive constant by Corollary C.4.444We assume that the learning system knows what the human observes, which is valid since POsubscriptP_OPitalic_O is deterministic. Alternatively, one can argue with Proposition C.11 that OO is automatically injective, meaning one can apply Theorem C.9. Thus, let Rā²ākerā”(ā)superscriptā²kernelR ā ( B )Rā² ā ker ( B ā Ī ), i.e., [ā”(ā”(Rā²))]ā¢(oā)=0delimited-[]superscriptā²ā0 [ B ( (R^% ) ) ]( o)=0[ B ( Ī ( Rā² ) ) ] ( overā start_ARG o end_ARG ) = 0 for every observation sequence oā ooverā start_ARG o end_ARG. For oā=sā¢sā¢sā¢sā o=ssssoverā start_ARG o end_ARG = s s s s being the observation sequence that only consists of state s, this implies Rā²ā¢(s)=0superscriptā²0R (s)=0Rā² ( s ) = 0. Consequently, for general observation sequences oā ooverā start_ARG o end_ARG, we have: 0=[ā”(ā”(Rā²))]ā¢(oā)=sāā¼Bā¢(sāā£oā)[āt=03Ī“sā²ā¢(st)ā γt]ā Rā²ā¢(sā²)+sāā¼Bā¢(sāā£oā)[āt=03Ī“sā²ā¢(st)ā γt]ā Rā²ā¢(sā²).0delimited-[]superscriptā²āā subscriptsimilar-toāconditionalāsuperscriptsubscript03ā subscriptsuperscriptā²subscriptsuperscriptsuperscriptā²ā subscriptsimilar-toāconditionalāsuperscriptsubscript03ā subscriptsuperscriptā²subscriptsuperscriptsuperscriptā²0= [ B ( (R^% ) ) ]( o)= *E_ s B( % s o) [ _t=0^3 _s (s_t)·γ^t% ]Ā· R (s )+ *E_ s B% ( s o) [ _t=0^3 _s (s_t)% ·γ^t ]Ā· R (s ).0 = [ B ( Ī ( Rā² ) ) ] ( overā start_ARG o end_ARG ) = Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ āt = 03 Ī“italic_sā² ( sitalic_t ) ā γitalic_t ] ā Rā² ( sā² ) + Eoverā start_ARG s end_ARG ā¼ B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) [ āt = 03 Ī“italic_sā² ā² ( sitalic_t ) ā γitalic_t ] ā Rā² ( sā² ā² ) . Now we specialize this equation to the two observation sequences oā(1)=sā¢oā¢sā¢ssuperscriptā1 o^(1)=sossoverā start_ARG o end_ARG( 1 ) = s o s s and oā(2)=sā¢oā¢oā¢ssuperscriptā2 o^(2)=soosoverā start_ARG o end_ARG( 2 ) = s o o s. We start by considering oā(1)superscriptā1 o^(1)overā start_ARG o end_ARG( 1 ). This is consistent with the two state sequences sā(1),(sā²)=sā¢sā²ā¢sā¢ssuperscriptā1superscriptā² s^(1),(s )=s ssoverā start_ARG s end_ARG( 1 ) , ( s start_POSTSUPERSCRIPT ā² ) end_POSTSUPERSCRIPT = s sā² s s and sā(1),(sā²)=sā¢sā²ā¢sā¢ssuperscriptā1superscriptā² s^(1),(s )=s ssoverā start_ARG s end_ARG( 1 ) , ( s start_POSTSUPERSCRIPT ā² ā² ) end_POSTSUPERSCRIPT = s sā² ā² s s. We have posterior probabilities Bā¢(sā(1),(sā²)ā£oā(1))=1āp,Bā¢(sā(1),(sā²)ā£oā(1))=p,formulae-sequenceconditionalsuperscriptā1superscriptā²ā11conditionalsuperscriptā1superscriptā²ā1B ( s^(1),(s ) o^(1) )=1-p, B (% s^(1),(s ) o^(1) )=p,B ( overā start_ARG s end_ARG( 1 ) , ( s start_POSTSUPERSCRIPT ā² ) end_POSTSUPERSCRIPT ⣠overā start_ARG o end_ARG( 1 ) ) = 1 - p , B ( overā start_ARG s end_ARG( 1 ) , ( s start_POSTSUPERSCRIPT ā² ā² ) end_POSTSUPERSCRIPT ⣠overā start_ARG o end_ARG( 1 ) ) = p , and therefore 0=[ā”(ā”(Rā²))]ā¢(oā(1))=(1āp)ā γā Rā²ā¢(sā²)+pā γā Rā²ā¢(sā²),0delimited-[]superscriptā²ā1ā 1superscriptā²ā superscriptā²0= [ B ( (R^% ) ) ]( o^(1))=(1-p)·γ· R (s^% )+p·γ· R (s ),0 = [ B ( Ī ( Rā² ) ) ] ( overā start_ARG o end_ARG( 1 ) ) = ( 1 - p ) ā γ ā Rā² ( sā² ) + p ā γ ā Rā² ( sā² ā² ) , and so Rā²ā¢(sā²)=pā1ā Rā²ā¢(sā²).superscriptā²ā 1superscriptā²R (s )= pp-1Ā· R (s ).Rā² ( sā² ) = divide start_ARG p end_ARG start_ARG p - 1 end_ARG ā Rā² ( sā² ā² ) . (11) Similarly, oā(2)superscriptā2 o^(2)overā start_ARG o end_ARG( 2 ) is consistent with the sequences sā(2),(sā²)=sā¢sā²ā¢sā²ā¢ssuperscriptā2superscriptā²superscriptā² s^(2),(s )=s s soverā start_ARG s end_ARG( 2 ) , ( s start_POSTSUPERSCRIPT ā² ) end_POSTSUPERSCRIPT = s sā² s and sā(2),(sā²)=sā¢sā²ā¢sā²ā¢ssuperscriptā2superscriptā²superscriptā² s^(2),(s )=s s soverā start_ARG s end_ARG( 2 ) , ( s start_POSTSUPERSCRIPT ā² ā² ) end_POSTSUPERSCRIPT = s sā² ā² sā² ā² s. They have posterior probabilities Bā¢(sā(2),(sā²)ā£oā(2))=12,Bā¢(sā(2),(sā²)ā£oā(2))=12,formulae-sequenceconditionalsuperscriptā2superscriptā²ā212conditionalsuperscriptā2superscriptā²ā212B ( s^(2),(s ) o^(2) )= 12, B% ( s^(2),(s ) o^(2) )= 12,B ( overā start_ARG s end_ARG( 2 ) , ( s start_POSTSUPERSCRIPT ā² ) end_POSTSUPERSCRIPT ⣠overā start_ARG o end_ARG( 2 ) ) = divide start_ARG 1 end_ARG start_ARG 2 end_ARG , B ( overā start_ARG s end_ARG( 2 ) , ( s start_POSTSUPERSCRIPT ā² ā² ) end_POSTSUPERSCRIPT ⣠overā start_ARG o end_ARG( 2 ) ) = divide start_ARG 1 end_ARG start_ARG 2 end_ARG , leading to 0=12ā (γ+γ2)ā Rā²ā¢(sā²)+12ā (γ+γ2)ā Rā²ā¢(sā²).0ā 12superscript2superscriptā²ā 12superscript2superscriptā²0= 12Ā·(γ+γ^2)Ā· R (s )+ 12% Ā·(γ+γ^2)Ā· R (s ).0 = divide start_ARG 1 end_ARG start_ARG 2 end_ARG ā ( γ + γ2 ) ā Rā² ( sā² ) + divide start_ARG 1 end_ARG start_ARG 2 end_ARG ā ( γ + γ2 ) ā Rā² ( sā² ā² ) . Together with Equation (11), we obtain Rā²ā¢(sā²)=āRā²ā¢(sā²)=p1āpā Rā²ā¢(sā²),superscriptā²superscriptā²ā 1superscriptā²R (s )=-R (s )= p1-pĀ· R^% (s ),Rā² ( sā² ā² ) = - Rā² ( sā² ) = divide start_ARG p end_ARG start_ARG 1 - p end_ARG ā Rā² ( sā² ā² ) , which implies Rā²ā¢(sā²)=0superscriptā²0R (s )=0Rā² ( sā² ā² ) = 0 because pā 1212pā 12p ā divide start_ARG 1 end_ARG start_ARG 2 end_ARG, and thus also Rā²ā¢(sā²)=0superscriptā²0R (s )=0Rā² ( sā² ) = 0. Overall, we have showed Rā²=0superscriptā²0R =0Rā² = 0, and so ā B B ā Ī is injective. This means that reward functions are identifiable in this example up to an additive constant, see Corollary C.4. Appendix D Issues of Naively Applying RLHF under Partial Observability In this section, we study the naive application of RLHF under partial observability. Thus, most of it takes a step back from the general theory of appropriately modeled partial observability in RLHF. Later, we will analyze examples where we also apply the general theory, which is why this appendix section comes second. In Section D.1, we first briefly explain what happens when the learning system incorrectly assumes that the human observes the full environment state. We show that as a consequence, the system is incentivized to infer what we call the observation return function GobssubscriptobsG_obsGroman_obs, which evaluates a state sequence based on the humanās belief of the state sequence given the humanās observations. In the policy optimization process, the policy is then selected to maximize JobssubscriptobsJ_obsJroman_obs, an expectation over GobssubscriptobsG_obsGroman_obs. In the interlude in Section D.2, we then briefly analyze the unrealistic case that the human, when evaluating a policy Ļ, fully knows the complete specification of that policy and all of the environment and engages in rational Bayesian reasoning; in this case, Jobs=JsubscriptobsJ_obs=JJroman_obs = J is the true policy evaluation function. Realistically, however, maximizing JobssubscriptobsJ_obsJroman_obs can lead to failure modes. In Section D.3 we prove that a suboptimal policy that is optimal according to JobssubscriptobsJ_obsJroman_obs causes deceptive inflation, overjustification, or both. In Section B.3, we expand on the analysis of the main examples in the main paper. Finally, in Section D.4, we study further concrete examples where maximizing JobssubscriptobsJ_obsJroman_obs reveals deceptive and overjustifying behavior by the resulting policy. D.1 Optimal Policies under RLHF with Deterministic Partial Observations Maximize JobssubscriptobsJ_obsJroman_obs Assume that POāsubscriptāP_ OPoverā start_ARG O end_ARG is deterministic and that the human makes Boltzmann-rational sequence comparisons between observation sequences. The true choice probabilities are then given by (See Equations (2) and (8)): PRā¢(sāā»sāā²)=Ļā¢(βā ((ā G)ā¢(Oāā¢(sā))ā(ā G)ā¢(Oāā¢(sāā²))))superscriptsucceedsāsuperscriptāā²ā āā āsuperscriptāā²P^R ( s s 1.0pt )=Ļ (% β· ( ( BĀ· G ) ( O(% s) )- ( BĀ· G ) ( O(% s 1.0pt ) ) ) )Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) = Ļ ( β ā ( ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) - ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) ) ) ) (12) Now, assume that the learning system does not model the situation correctly. In particular, we assume: ⢠The system is not aware that the human only observes observation sequences Oāā¢(sā)ā O( s)overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) instead of the full state sequences. ⢠The system does not model that the humanās return function is time-separable, i.e., comes from a reward function R over environment states. The learning system then thinks that there is a return function G~āāSā~superscriptāā Gā R Sover~ start_ARG G end_ARG ā blackboard_Roverā start_ARG S end_ARG such that the choice probabilities are given by the following faulty formula: PRā¢(sāā»sāā²)āĻā¢(βā¢(Gā¢(sā)āGā¢(sāā²)))āsuperscriptsucceedsāsuperscriptāā²āsuperscriptāā²P^R ( s s 1.0pt ) Ļ % (β (G( s)-G( s 1.0pt ) ) )Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) ā Ļ ( β ( G ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARGā² ) ) ) Now, assume that the learning system has access to the choice probabilities and wants to infer G. Inverting the sigmoid function and then plugging in the true choice probabilities from Equation (12), we obtain: G~ā¢(sā)~ā G( s)over~ start_ARG G end_ARG ( overā start_ARG s end_ARG ) =1βā¢logā”PRā¢(sāā»sāā²)PRā¢(sāā²ā»sā)+G~ā¢(sāā²)absent1superscriptsucceedsāsuperscriptāā²succeedssuperscriptāā²ā~superscriptāā² = 1β P^R( s s 1.0pt^% )P^R( s 1.0pt s)+ G( s% 1.0pt )= divide start_ARG 1 end_ARG start_ARG β end_ARG log divide start_ARG Pitalic_R ( overā start_ARG s end_ARG ā» overā start_ARG s end_ARGā² ) end_ARG start_ARG Pitalic_R ( overā start_ARG s end_ARGā² ā» overā start_ARG s end_ARG ) end_ARG + over~ start_ARG G end_ARG ( overā start_ARG s end_ARGā² ) =1βā¢[βā ((ā G)ā¢(Oāā¢(sā))ā(ā G)ā¢(Oāā¢(sāā²)))]+G~ā¢(sāā²)absent1delimited-[]ā āā āsuperscriptāā²~superscriptāā² = 1β [β· ( (% BĀ· G ) ( O( s) )- (% BĀ· G ) ( O( s 1.0pt ) )% ) ]+ G( s 1.0pt )= divide start_ARG 1 end_ARG start_ARG β end_ARG [ β ā ( ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) - ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARGā² ) ) ) ] + over~ start_ARG G end_ARG ( overā start_ARG s end_ARGā² ) =(ā G)ā¢(Oāā¢(sā))+Cā¢(sāā²).absentā āsuperscriptāā² = ( BĀ· G ) ( O(% s) )+C( s 1.0pt ).= ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) + C ( overā start_ARG s end_ARGā² ) . Here, Cā¢(sāā²)superscriptāā²C( s 1.0pt )C ( overā start_ARG s end_ARGā² ) is some quantity that does not depend on sā soverā start_ARG s end_ARG. Now, fix sāā² s 1.0pt overā start_ARG s end_ARGā² as a reference sequence. Then for varying sā soverā start_ARG s end_ARG, Cā¢(sāā²)superscriptāā²C( s 1.0pt )C ( overā start_ARG s end_ARGā² ) is simply an additive constant. Consequently, up to an additive constant, this determines the return function that the learning system is incentivized to infer. We call it the observation return function since it is the return function based on the humanās observations: Gobsā¢(sā)ā(ā G)ā¢(Oāā¢(sā)).āsubscriptobsāā āG_obs( s) ( BĀ· G% ) ( O( s) ).Groman_obs ( overā start_ARG s end_ARG ) ā ( B ā G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) . This return function is not necessarily time-separable, but we assume that time-separability is not modeled correctly by the learning system. Now, define the resulting policy evaluation function JobssubscriptobsJ_obsJroman_obs by Jobsā¢(Ļ)āsāā¼PĻā¢(sā)[Gobsā¢(sā)].āsubscriptobssubscriptsimilar-toāsuperscriptāsubscriptobsāJ_obs(Ļ) *E_ s P^% Ļ( s) [G_obs( s) ].Jroman_obs ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ Groman_obs ( overā start_ARG s end_ARG ) ] . This is the policy evaluation function that would be optimized if the learning system erroneously inferred the return function GobssubscriptobsG_obsGroman_obs. D.2 Interlude: When the Human Knows the Policy and is a Bayesian Reasoner, then Jobs=JsubscriptobsJ_obs=JJroman_obs = J In this section, we briefly consider what would happen if in JobssubscriptobsJ_obsJroman_obs, the humanās belief B would make use of the true policy and be a rational Bayesian posterior as in Section C.1. We will show that under these conditions, we have Jobs=JsubscriptobsJ_obs=JJroman_obs = J. Since these are unrealistic assumptions, no other section depends on this result. For the analysis, we drop the assumption that the observation sequence kernel POāsubscriptāP_ OPoverā start_ARG O end_ARG is deterministic, and assume that JobssubscriptobsJ_obsJroman_obs is given as follows: Jobsā¢(Ļ)āsāā¼PĻā¢(sā)[oāā¼POāā¢(oāā£sā)[sāā²ā¼BĻā¢(sāā²ā£oā)[Gā¢(sāā²)]]].āsubscriptobssubscriptsimilar-toāsuperscriptāsubscriptsimilar-toāsubscriptāconditionalāsubscriptsimilar-tosuperscriptāā²conditionalsuperscriptāā²āsuperscriptāā²J_obs(Ļ) *E_ s P^% Ļ( s) [ *E_ o P_ O(% o s) [ *E_ s 1.0pt^% B^Ļ( s 1.0pt o) [G( s% 1.0pt ) ] ] ].Jroman_obs ( Ļ ) ā Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ Eoverā start_ARG s end_ARGā² ā¼ Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) [ G ( overā start_ARG s end_ARGā² ) ] ] ] . (13) In this formula, BĻā¢(sāā£oā)āBā¢(sāā£oā,Ļ)āsuperscriptconditionalāconditionalāB^Ļ( s o) B( s o,Ļ)Bitalic_Ļ ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ) ā B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG , Ļ ) with B being the joint distribution from Section C.1. Formally, this is the posterior of the joint distribution Bā¢(sā,oāā£Ļ)āconditionalāB( s, o Ļ)B ( overā start_ARG s end_ARG , overā start_ARG o end_ARG ā£ Ļ ) that is given by the following hidden Markov model: s0subscript0s_0s0s1subscript1s_1s1s2subscript2s_2s2s3subscript3s_3s3ā¦o0subscript0o_0o0o1subscript1o_1o1o2subscript2o_2o2o3subscript3o_3o3ā¦POsubscript P_OPitalic_OĻsuperscript T^ĻTitalic_ĻPOsubscript P_OPitalic_OĻsuperscript T^ĻTitalic_ĻPOsubscript P_OPitalic_OĻsuperscript T^ĻTitalic_ĻPOsubscript P_OPitalic_OĻsuperscript T^ĻTitalic_Ļ (14) Here, Ļā¢(sā²ā£s)āāaāā¢(sā²ā£s,a)ā Ļā¢(aā£s)āsuperscriptconditionalsuperscriptā²subscriptā conditionalsuperscriptā²conditionalT^Ļ(s s) _a T(% s s,a)Ā·Ļ(a s)Titalic_Ļ ( sⲠ⣠s ) ā āa ā A T ( sⲠ⣠s , a ) ā Ļ ( a ⣠s ). s0subscript0s_0s0 is sampled according to the known initial distribution P0ā¢(s0)subscript0subscript0P_0(s_0)P0 ( s0 ). The humanās posterior BĻā¢(sāā²ā£oā)superscriptconditionalsuperscriptāā²āB^Ļ( s 1.0pt o)Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) is then the true posterior in this HMM. We obtain: Proposition D.1. Let Ļ be a policy that is known to the human. Then Jobsā¢(Ļ)=Jā¢(Ļ)subscriptobsJ_obs(Ļ)=J(Ļ)Jroman_obs ( Ļ ) = J ( Ļ ). Proof. By Equation (13), we have Jobsā¢(Ļ)subscriptobs J_obs(Ļ)Jroman_obs ( Ļ ) =sāā¼PĻā¢(sā)[oāā¼POāā¢(oāā£sā)[sāā²ā¼BĻā¢(sāā²ā£oā)[Gā¢(sāā²)]]]absentsubscriptsimilar-toāsuperscriptāsubscriptsimilar-toāsubscriptāconditionalāsubscriptsimilar-tosuperscriptāā²conditionalsuperscriptāā²āsuperscriptāā² = *E_ s P^Ļ( s) [% *E_ o P_ O( o s) % [ *E_ s 1.0pt B^Ļ( s% 1.0pt o) [G( s 1.0pt ) % ] ] ]= Eoverā start_ARG s end_ARG ā¼ Pitalic_Ļ ( overā start_ARG s end_ARG ) [ Eoverā start_ARG o end_ARG ā¼ P start_POSTSUBSCRIPT overā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) end_POSTSUBSCRIPT [ Eoverā start_ARG s end_ARGā² ā¼ Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) [ G ( overā start_ARG s end_ARGā² ) ] ] ] =(1)ā¢āsāPĻā¢(sā)ā¢āoāPOāā¢(oāā£sā)ā¢āsāā²BĻā¢(sāā²ā£oā)ā¢Gā¢(sāā²)1subscriptāsuperscriptāsubscriptāsubscriptāconditionalāsubscriptsuperscriptāā²conditionalsuperscriptāā²āsuperscriptāā² (1)= _ sP^Ļ( s) _ oP_% O( o s) _ s 1.0pt B^Ļ( s% 1.0pt o)G( s 1.0pt )start_OVERACCENT ( 1 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARG Pitalic_Ļ ( overā start_ARG s end_ARG ) āoverā start_ARG o end_ARG Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) āoverā start_ARG s end_ARGā² Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) G ( overā start_ARG s end_ARGā² ) =(2)ā¢āsāā²[āoāBĻā¢(sāā²ā£oā)ā¢[āsāPOāā¢(oāā£sā)ā¢PĻā¢(sā)]]ā¢Gā¢(sāā²)2subscriptsuperscriptāā²delimited-[]subscriptāsuperscriptconditionalsuperscriptāā²ādelimited-[]subscriptāsubscriptāconditionalāsuperscriptāsuperscriptāā² (2)= _ s 1.0pt [ _% oB^Ļ( s 1.0pt o) [ _ s% P_ O( o s)P^Ļ( s) ] ]G( s 1% .0pt )start_OVERACCENT ( 2 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARGā² [ āoverā start_ARG o end_ARG Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) [ āoverā start_ARG s end_ARG Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ) Pitalic_Ļ ( overā start_ARG s end_ARG ) ] ] G ( overā start_ARG s end_ARGā² ) =(3)ā¢āsāā²[āoāBĻā¢(sāā²ā£oā)ā¢BĻā¢(oā)]ā¢Gā¢(sāā²)3subscriptsuperscriptāā²delimited-[]subscriptāsuperscriptconditionalsuperscriptāā²āsuperscriptāsuperscriptāā² (3)= _ s 1.0pt [ _% oB^Ļ( s 1.0pt o)B^Ļ( o) % ]G( s 1.0pt )start_OVERACCENT ( 3 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARGā² [ āoverā start_ARG o end_ARG Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) Bitalic_Ļ ( overā start_ARG o end_ARG ) ] G ( overā start_ARG s end_ARGā² ) =(4)ā¢āsāā²[āoāPĻā¢(sāā²)ā¢POāā¢(oāā£sāā²)]ā¢Gā¢(sāā²)4subscriptsuperscriptāā²delimited-[]subscriptāsuperscriptsuperscriptāā²subscriptāconditionalāsuperscriptāā² (4)= _ s 1.0pt [ _% oP^Ļ( s 1.0pt )P_ O( o s% 1.0pt ) ]G( s 1.0pt )start_OVERACCENT ( 4 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARGā² [ āoverā start_ARG o end_ARG Pitalic_Ļ ( overā start_ARG s end_ARGā² ) Poverā start_ARG O end_ARG ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARGā² ) ] G ( overā start_ARG s end_ARGā² ) =(5)ā¢āsāā²PĻā¢(sāā²)ā¢Gā¢(sāā²)5subscriptsuperscriptāā²superscriptāā² (5)= _ s 1.0pt P^Ļ( s% 1.0pt )G( s 1.0pt )start_OVERACCENT ( 5 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARGā² Pitalic_Ļ ( overā start_ARG s end_ARGā² ) G ( overā start_ARG s end_ARGā² ) =(6)ā¢āsāPĻā¢(sā)ā¢Gā¢(sā)6subscriptāsuperscriptā (6)= _ sP^Ļ( s)G( s)start_OVERACCENT ( 6 ) end_OVERACCENT start_ARG = end_ARG āoverā start_ARG s end_ARG Pitalic_Ļ ( overā start_ARG s end_ARG ) G ( overā start_ARG s end_ARG ) =(7)ā¢Jā¢(Ļ).7 (7)=J(Ļ).start_OVERACCENT ( 7 ) end_OVERACCENT start_ARG = end_ARG J ( Ļ ) . In step (1), we wrote the expectations out in terms of sums. In step (2), we reordered them. In step (3), we observed that the inner sum over sā soverā start_ARG s end_ARG evaluates to the marginal distribution BĻā¢(oā)superscriptāB^Ļ( o)Bitalic_Ļ ( overā start_ARG o end_ARG ) of the observation sequence oā ooverā start_ARG o end_ARG in the HMM in Equation (13). In step (4), we used Bayes rule in the inner sum. This is possible since BĻā¢(sāā²ā£oā)superscriptconditionalsuperscriptāā²āB^Ļ( s 1.0pt o)Bitalic_Ļ ( overā start_ARG s end_ARGⲠ⣠overā start_ARG o end_ARG ) is the true posterior when Ļ is known. In step (5), we pull PĻā¢(sāā²)superscriptsuperscriptāā²P^Ļ( s 1.0pt )Pitalic_Ļ ( overā start_ARG s end_ARGā² ) out and notice that the remaining inner sum evaluates to 1111. Step (6) is a relabeling and step (7) the definition of the true policy evaluation function J. ā D.3 Proof of Theorem 4.5 We first prove the following lemma. Lemma D.2. Let Ļ and Ļrefsubscriptref _refĻref be two policies. If Jā¢(Ļ)<Jā¢(Ļref)subscriptrefJ(Ļ)<J( _ref)J ( Ļ ) < J ( Ļref ) and Jobsā¢(Ļ)>Jobsā¢(Ļref)subscriptobssubscriptobssubscriptrefJ_obs(Ļ)>J_obs( _ref)Jroman_obs ( Ļ ) > Jroman_obs ( Ļref ), then relative to Ļrefsubscriptref _refĻref, Ļ must exhibit deceptive inflation, overjustification, or both. Proof. We start by establishing a quantitative relationship between the average overestimation and underestimation errors EĀÆ+superscriptĀÆ E^+overĀÆ start_ARG E end_ARG+ and EĀÆāsuperscriptĀÆ E^-overĀÆ start_ARG E end_ARG- as defined in Definition 4.2, the true policy evaluation function J, and the observation evaluation function JobssubscriptobsJ_obsJroman_obs defined in Equation 4. Define Ī:āā:Īāā : Sā RĪ : overā start_ARG S end_ARG ā blackboard_R by Īā¢(sā)=Gobsā¢(sā)āGā¢(sā)Īāsubscriptobsā ( s)=G_obs( s)-G( s)Ī ( overā start_ARG s end_ARG ) = Groman_obs ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARG ), where GobssubscriptobsG_obsGroman_obs is as defined in Equation 3. Consider the quantity E+ā¢(sā)āEāā¢(sā)=maxā”(0,Īā¢(sā))āmaxā”(0,āĪā¢(sā)).superscriptāsuperscriptā0Īā0Īā E^+( s)-E^-( s)= (0, ( s) % )- (0,- ( s) ).E+ ( overā start_ARG s end_ARG ) - E- ( overā start_ARG s end_ARG ) = max ( 0 , Ī ( overā start_ARG s end_ARG ) ) - max ( 0 , - Ī ( overā start_ARG s end_ARG ) ) . If Īā¢(sā)>0Īā0 ( s)>0Ī ( overā start_ARG s end_ARG ) > 0, then the first term is Īā¢(sā)Īā ( s)Ī ( overā start_ARG s end_ARG ) and the second one is 0. If Īā¢(sā)<0Īā0 ( s)<0Ī ( overā start_ARG s end_ARG ) < 0, then the first term is zero and the second one is Īā¢(sā)Īā ( s)Ī ( overā start_ARG s end_ARG ). If Īā¢(sā)=0Īā0 ( s)=0Ī ( overā start_ARG s end_ARG ) = 0, then both terms are zero. In all cases the right-hand side is equal to Īā¢(sā)Īā ( s)Ī ( overā start_ARG s end_ARG ). Unpacking the definition of Ī Ī again, we have that for all sā soverā start_ARG s end_ARG, E+ā¢(sā)āEāā¢(sā)=Gobsā¢(sā)āGā¢(sā).superscriptāsuperscriptāsubscriptobsā E^+( s)-E^-( s)=G_obs( s)-G(% s).E+ ( overā start_ARG s end_ARG ) - E- ( overā start_ARG s end_ARG ) = Groman_obs ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARG ) . (15) For any policy Ļ, if we take the expectation of both sides of this equation over the on-policy distribution admitted by Ļ, PĻsuperscriptP^ĻPitalic_Ļ, we get EĀÆ+ā¢(Ļ)āEĀÆāā¢(Ļ)=Jobsā¢(Ļ)āJā¢(Ļ).superscriptĀÆsuperscriptĀÆsubscriptobs E^+(Ļ)- E^-(Ļ)=J_obs% (Ļ)-J(Ļ).overĀÆ start_ARG E end_ARG+ ( Ļ ) - overĀÆ start_ARG E end_ARG- ( Ļ ) = Jroman_obs ( Ļ ) - J ( Ļ ) . (16) We now prove the lemma. Let Ļ and Ļrefsubscriptref _refĻref be two policies, and assume that Jā¢(Ļ)<Jā¢(Ļref)subscriptrefJ(Ļ)<J( _ref)J ( Ļ ) < J ( Ļref ) and Jobsā¢(Ļ)ā„Jobsā¢(Ļref)subscriptobssubscriptobssubscriptrefJ_obs(Ļ)ā„ J_obs( _ref)Jroman_obs ( Ļ ) ā„ Jroman_obs ( Ļref ). Equivalently, we have Jobsā¢(Ļ)āJobsā¢(Ļref)ā„0subscriptobssubscriptobssubscriptref0J_obs(Ļ)-J_obs( _ref)ā„ 0Jroman_obs ( Ļ ) - Jroman_obs ( Ļref ) ā„ 0 and Jā¢(Ļref)āJā¢(Ļ)>0subscriptref0J( _ref)-J(Ļ)>0J ( Ļref ) - J ( Ļ ) > 0, which we combine to state (Jobsā¢(Ļ)āJobsā¢(Ļref))+(Jā¢(Ļref)āJā¢(Ļ))>0.subscriptobssubscriptobssubscriptrefsubscriptref0 (J_obs(Ļ)-J_obs( _% ref) )+ (J( _ref)-J(Ļ) )>0.( Jroman_obs ( Ļ ) - Jroman_obs ( Ļref ) ) + ( J ( Ļref ) - J ( Ļ ) ) > 0 . (17) Rearranging terms yields (Jobsā¢(Ļ)āJā¢(Ļ))ā(Jobsā¢(Ļref)āJā¢(Ļref))>0.subscriptobssubscriptobssubscriptrefsubscriptref0 (J_obs(Ļ)-J(Ļ) )- (J_% obs( _ref)-J( _ref) )>0.( Jroman_obs ( Ļ ) - J ( Ļ ) ) - ( Jroman_obs ( Ļref ) - J ( Ļref ) ) > 0 . These two differences inside parentheses are equal to the right-hand side of (16) for Ļ and Ļrefsubscriptref _refĻref, respectively. We substitute the left-hand side of (16) twice to obtain (EĀÆ+ā¢(Ļ)āEĀÆāā¢(Ļ))ā(EĀÆ+ā¢(Ļref)āEĀÆāā¢(Ļref))>0.superscriptĀÆsuperscriptĀÆsuperscriptĀÆsubscriptrefsuperscriptĀÆsubscriptref0 ( E^+(Ļ)- E^-(Ļ) )- % ( E^+( _ref)- E^-( _ref)% )>0.( overĀÆ start_ARG E end_ARG+ ( Ļ ) - overĀÆ start_ARG E end_ARG- ( Ļ ) ) - ( overĀÆ start_ARG E end_ARG+ ( Ļref ) - overĀÆ start_ARG E end_ARG- ( Ļref ) ) > 0 . Rearranging terms again yields (EĀÆ+ā¢(Ļ)āEĀÆ+ā¢(Ļref))+(EĀÆāā¢(Ļref)āEĀÆāā¢(Ļ))>0.superscriptĀÆsuperscriptĀÆsubscriptrefsuperscriptĀÆsubscriptrefsuperscriptĀÆ0 ( E^+(Ļ)- E^+( _ref% ) )+ ( E^-( _ref)- E^-(Ļ)% )>0.( overĀÆ start_ARG E end_ARG+ ( Ļ ) - overĀÆ start_ARG E end_ARG+ ( Ļref ) ) + ( overĀÆ start_ARG E end_ARG- ( Ļref ) - overĀÆ start_ARG E end_ARG- ( Ļ ) ) > 0 . (18) If EĀÆ+ā¢(Ļ)āEĀÆ+ā¢(Ļref)>0superscriptĀÆsuperscriptĀÆsubscriptref0 E^+(Ļ)- E^+( _ref)>0overĀÆ start_ARG E end_ARG+ ( Ļ ) - overĀÆ start_ARG E end_ARG+ ( Ļref ) > 0 then we have EĀÆ+ā¢(Ļ)>EĀÆ+ā¢(Ļref)superscriptĀÆsuperscriptĀÆsubscriptref E^+(Ļ)> E^+( _ref)overĀÆ start_ARG E end_ARG+ ( Ļ ) > overĀÆ start_ARG E end_ARG+ ( Ļref ) and, by assumption, Jobsā¢(Ļ)>Jobsā¢(Ļref)subscriptobssubscriptobssubscriptrefJ_obs(Ļ)>J_obs( _ref)Jroman_obs ( Ļ ) > Jroman_obs ( Ļref ). By Definition 4.3, this means Ļ exhibits deceptive inflation relative to Ļrefsubscriptref _refĻref. If EĀÆāā¢(Ļref)āEĀÆāā¢(Ļ)>0superscriptĀÆsubscriptrefsuperscriptĀÆ0 E^-( _ref)- E^-(Ļ)>0overĀÆ start_ARG E end_ARG- ( Ļref ) - overĀÆ start_ARG E end_ARG- ( Ļ ) > 0 then we have EĀÆāā¢(Ļ)<EĀÆāā¢(Ļref)superscriptĀÆsuperscriptĀÆsubscriptref E^-(Ļ)< E^-( _ref)overĀÆ start_ARG E end_ARG- ( Ļ ) < overĀÆ start_ARG E end_ARG- ( Ļref ) and, by assumption, Jā¢(Ļ)<Jā¢(Ļref)subscriptrefJ(Ļ)<J( _ref)J ( Ļ ) < J ( Ļref ). By Definition 4.4, this means Ļ exhibits overjustification relative to Ļrefsubscriptref _refĻref. At least one of the two differences in parentheses in (18) must be positive, otherwise their sum would not be positive. Thus Ļ must exhibit deceptive inflation relative to Ļrefsubscriptref _refĻref, overjustification relative to Ļrefsubscriptref _refĻref, or both. ā We can now combine earlier results to prove Theorem 4.5, repeated here for convenience: Theorem D.3. Assume that POsubscriptP_OPitalic_O is deterministic. Let ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs be an optimal policy according to a naive application of RLHF under partial observability, and let ĻāsuperscriptĻ^*Ļā be an optimal policy according to the true objective J. If ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs is not J-optimal, then relative to ĻāsuperscriptĻ^*Ļā, ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs must exhibit deceptive inflation, overjustification, or both. Proof. Because POsubscriptP_OPitalic_O is deterministic, ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs must be optimal with respect to JobssubscriptobsJ_obsJroman_obs by Proposition 4.1 (proved in Section D.1). Thus Jobsā¢(Ļobsā)ā„Jobsā¢(Ļā)subscriptobssubscriptsuperscriptobssubscriptobssuperscriptJ_obs(Ļ^*_obs)ā„ J_obs% (Ļ^*)Jroman_obs ( Ļāroman_obs ) ā„ Jroman_obs ( Ļā ). Since ĻāsuperscriptĻ^*Ļā is J-optimal and ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs is not, Jā¢(Ļā)<Jā¢(Ļobsā)superscriptsubscriptsuperscriptobsJ(Ļ^*)<J(Ļ^*_obs)J ( Ļā ) < J ( Ļāroman_obs ). By Lemma D.2, relative to ĻāsuperscriptĻ^*Ļā, ĻobsāsubscriptsuperscriptobsĻ^*_obsĻāroman_obs must exhibit deceptive inflation, overjustification, or both. ā D.4 Further Examples Supplementing Section 4.4 In this section, we present further mathematical examples supplementing those in Section 4.4. We found many of them before finding the examples we discuss in the main paper, and show the same and additional conceptual features with somewhat less polish. We again assume that POāsubscriptāP_ OPoverā start_ARG O end_ARG is deterministic. Example D.4. In the main paper, we have assumed a model where the human obeys Eq. (2) and showed that a naive application of RLHF can lead to suboptimal policies, and the specific failure modes of deceptive inflation and overjustification. What if the human makes the choices in a different way? Specifically, assume that all we know is that PRā¢(oāā»oāā²)+PRā¢(oāā²ā»oā)=1superscriptsucceedsāsuperscriptāā²succeedssuperscriptāā²ā1P^R( o o 1.0pt )+P^R( o 1.0pt^% o)=1Pitalic_R ( overā start_ARG o end_ARG ā» overā start_ARG o end_ARGā² ) + Pitalic_R ( overā start_ARG o end_ARGā² ā» overā start_ARG o end_ARG ) = 1. Can the human generally choose these choice probabilities in such a way that RLHF is incentivized to infer a reward function whose optimal policies are also optimal for R? The answer is no. Take the following example: sssaaabbbccc In this example, there is a fixed start state s and three actions a,b,ca,b,ca , b , c that also serve as the final states. The time horizon is T=11T=1T = 1, so the only state sequences are sā¢a,sā¢b,sā¢csa,sb,scs a , s b , s c. Assume ā¢(aā£s,a)=1conditional1T(a s,a)=1T ( a ⣠s , a ) = 1, ā¢(bā£s,b)=1conditional1T(b s,b)=1T ( b ⣠s , b ) = 1, ā¢(cā£s,c)=1āϵconditional1italic-ϵT(c s,c)=1- ( c ⣠s , c ) = 1 - ϵ, ā¢(aā£s,c)=ϵconditionalitalic-ϵT(a s,c)= ( a ⣠s , c ) = ϵ, i.e., selecting action c sometimes leads to state a. Also, assume a=Oā¢(a)ā Oā¢(b)=Oā¢(c)āoāa=O(a)ā O(b)=O(c) oa = O ( a ) ā O ( b ) = O ( c ) ā o and Rā¢(a)=Rā¢(b)<Rā¢(c)R(a)=R(b)<R(c)R ( a ) = R ( b ) < R ( c ). Since b and c have the same observation o, the human choice probabilities do not make a difference between them, and so RLHF is incentivized to infer a reward function R~~ Rover~ start_ARG R end_ARG with R~ā¢(b)=R~ā¢(c)āR~ā¢(o)~~ā~ R(b)= R(c) R(o)over~ start_ARG R end_ARG ( b ) = over~ start_ARG R end_ARG ( c ) ā over~ start_ARG R end_ARG ( o ). If R~ā¢(o)>R~ā¢(a)~~ R(o)> R(a)over~ start_ARG R end_ARG ( o ) > over~ start_ARG R end_ARG ( a ), then the policy optimal under R~~ Rover~ start_ARG R end_ARG will produce action b since this deterministically leads to observation o, whereas c does not. If R~ā¢(o)<R~ā¢(a)~~ R(o)< R(a)over~ start_ARG R end_ARG ( o ) < over~ start_ARG R end_ARG ( a ), then the policy optimal under R~~ Rover~ start_ARG R end_ARG will produce action a. In both cases, the resulting policy is suboptimal compared to ĻāsuperscriptĻ^*Ļā, which deterministically chooses action c. In the coming examples, it will also be useful to look at the misleadingness of state sequences: Definition D.5 (Misleadingness). Let sāāā sā Soverā start_ARG s end_ARG ā overā start_ARG S end_ARG be a state sequence. Then its misleadingness is defined by Mā”(sā)āGobsā¢(sā)āGā¢(sā)=sāā²ā¼Bā¢(sāā²ā£Oāā¢(sā))[Gā¢(sāā²)āGā¢(s)].āMāsubscriptobsāsubscriptsimilar-tosuperscriptāā²conditionalsuperscriptāā²āsuperscriptāā²M( s) G_obs( s)-G( s)=% *E_ s 1.0pt B( s 1% .0pt O( s)) [G( s 1.0pt )-G(s)% ].M ( overā start_ARG s end_ARG ) ā Groman_obs ( overā start_ARG s end_ARG ) - G ( overā start_ARG s end_ARG ) = Eoverā start_ARG s end_ARGā² ā¼ B ( overā start_ARG s end_ARGⲠ⣠overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) [ G ( overā start_ARG s end_ARGā² ) - G ( s ) ] . We call a state sequence positively misleading if Mā¢(sā)>0ā0M( s)>0M ( overā start_ARG s end_ARG ) > 0, which means the sequence appears better than it is, and negatively misleading if Mā”(sā)<0Mā0M( s)<0M ( overā start_ARG s end_ARG ) < 0. The misleadingness vector is given by MāāāMsuperscriptāāMā R SM ā blackboard_Roverā start_ARG S end_ARG. Note that the misleadingness is related to E+superscriptE^+E+ and EāsuperscriptE^-E-, as defined in Definition 4.2: If Mā”(sā)>0Mā0M( s)>0M ( overā start_ARG s end_ARG ) > 0 then Mā”(sā)=E+ā¢(sā)MāsuperscriptāM( s)=E^+( s)M ( overā start_ARG s end_ARG ) = E+ ( overā start_ARG s end_ARG ), and if Mā”(sā)<0Mā0M( s)<0M ( overā start_ARG s end_ARG ) < 0 then Mā”(sā)=āEāā¢(sā)MāsuperscriptāM( s)=-E^-( s)M ( overā start_ARG s end_ARG ) = - E- ( overā start_ARG s end_ARG ). Example D.6. In this example, we assume the human is a Bayesian reasoner as in Section C.1. Consider the MDP that is suggestively depicted as follows: aaabbbccc The MDP has states =a,b,cS=\a,b,c\S = a , b , c and actions =b,cA=\b,c\A = b , c . The transition kernel is given by ā¢(cā£a,c)=1conditional1T(c a,c)=1T ( c ⣠a , c ) = 1 and ā¢(bā£a,b)=1conditional1T(b a,b)=1T ( b ⣠a , b ) = 1, meaning that the action determines whether to transition from a to b or c. All other transitions are deterministic and do not depend on the action, as depicted. We assume an initial state distribution P0subscript0P_0P0 over states with probabilities pa=P0ā¢(a),pb=P0ā¢(b),pc=P0ā¢(c)formulae-sequencesubscriptsubscript0formulae-sequencesubscriptsubscript0subscriptsubscript0p_a=P_0(a),p_b=P_0(b),p_c=P_0(c)pitalic_a = P0 ( a ) , pitalic_b = P0 ( b ) , pitalic_c = P0 ( c ). The true reward function Rāāa,b,csuperscriptāRā R^\a,b,c\R ā blackboard_R a , b , c and discount factor γā[0,1)01γā[0,1)γ ā [ 0 , 1 ) are, for now, kept arbitrary. The time horizon is T=22T=2T = 2, meaning we have four possible state sequences aā¢cā¢cacca c c, aā¢bā¢cabca b c, bā¢cā¢cbccb c c, cā¢cā¢c c c. Furthermore, assume that oāOā¢(a)=Oā¢(b)ā Oā¢(c)=cāo O(a)=O(b)ā O(c)=co ā O ( a ) = O ( b ) ā O ( c ) = c, i.e., c is observed and a and b are ambiguous. Finally, assume that the human has a policy prior Bā¢(Ī»)B(Ī»)B ( Ī» ), where Ī»=ĻĪ»ā¢(cā£a)subscriptconditionalĪ»= _Ī»(c a)Ī» = Ļitalic_Ī» ( c ⣠a ) is the likelihood that the policy chooses action c when in state a, which is a parameter that determines the entire policy. We claim the following: 1. If pbā γā Ī»ā¼Bā¢(Ī»)[Ī»]ā pasubscriptā subscriptsimilar-tosubscriptp_bā γ· *E_Ī» B(Ī»)[% Ī»]Ā· p_apitalic_b ā γ ā Eitalic_Ī» ā¼ B ( Ī» ) [ Ī» ] ā pitalic_a, then kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 , so there is no return function ambiguity under appropriately modeled partially observable RLHF, see Corollary C.4. 2. There are true reward functions R for which optimizing JobssubscriptobsJ_obsJroman_obs leads to a suboptimal policy according to the true policy evaluation function J, a case of misalignment. Thus, a naive application of RLHF under partial observability fails, see Section 4.1. 3. The failure modes are related to hiding negative information (deception) and purposefully revealing information while incuring a loss (overjustifying behavior). Proof. Write pāBā¢(bā¢cā¢cā£oā¢cā¢c)āconditionalp B(bcc occ)p ā B ( b c c ⣠o c c ), the humanās posterior probability of state sequence bā¢cā¢cbccb c c for observation sequence oā¢cā¢cocco c c. We have 1āp=Bā¢(aā¢cā¢cā£oā¢cā¢c)1conditional1-p=B(acc occ)1 - p = B ( a c c ⣠o c c ). Consider the linear operators :āa,b,cāāaā¢bā¢c,bā¢cā¢c,cā¢cā¢c,aā¢cā¢c:āsuperscriptāsuperscriptā : R^\a,b,c\ā R^\abc,bcc,% c,acc\Ī : blackboard_R a , b , c ā blackboard_R a b c , b c c , c c c , a c c and :āaā¢bā¢c,bā¢cā¢c,cā¢cā¢c,aā¢cā¢cāāoā¢oā¢c,oā¢cā¢c,cā¢cā¢c:āsuperscriptāsuperscriptā B: R^\abc,bcc,c,acc\ā R^\ooc% ,occ,c\B : blackboard_R a b c , b c c , c c c , a c c ā blackboard_R o o c , o c c , c c c defined in the main paper. When ordering the states, state sequences, and observation sequences as we just wrote down, we obtain =(1γ201γ+γ2001+γ+γ210γ+γ2),=(10000p01āp0010),ā=(1γ21āpγ+γ2001+γ+γ2).formulae-sequencematrix1superscript201superscript2001superscript210superscript2formulae-sequencematrix10000010010matrix1superscript21superscript2001superscript2 = pmatrix1&γ&γ^2\\ 0&1&γ+γ^2\\ 0&0&1+γ+γ^2\\ 1&0&γ+γ^2 pmatrix, B= % pmatrix1&0&0&0\\ 0&p&0&1-p\\ 0&0&1&0 pmatrix, B % = pmatrix1&γ&γ^2\\ 1-p&p&γ+γ^2\\ 0&0&1+γ+γ^2 pmatrix.Ī = ( start_ARG start_ROW start_CELL 1 end_CELL start_CELL γ end_CELL start_CELL γ2 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL γ + γ2 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 + γ + γ2 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL γ + γ2 end_CELL end_ROW end_ARG ) , B = ( start_ARG start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL p end_CELL start_CELL 0 end_CELL start_CELL 1 - p end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL end_ROW end_ARG ) , B ā Ī = ( start_ARG start_ROW start_CELL 1 end_CELL start_CELL γ end_CELL start_CELL γ2 end_CELL end_ROW start_ROW start_CELL 1 - p end_CELL start_CELL p end_CELL start_CELL γ + γ2 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 + γ + γ2 end_CELL end_ROW end_ARG ) . By Corollary C.4, if ā B B ā Ī is injective, then there is no reward function ambiguity. Clearly, this is the case if and only if pā γā (1āp)ā 1pā γ·(1-p)p ā γ ā ( 1 - p ). From Bayes rule, we have p=Bā¢(bā¢cā¢c)Bā¢(aā¢cā¢c)+Bā¢(bā¢cā¢c),1āp=Bā¢(aā¢cā¢c)Bā¢(aā¢cā¢c)+Bā¢(bā¢cā¢c).formulae-sequence1p= B(bcc)B(acc)+B(bcc), 1-p= B(acc)B(acc)+B(bcc).p = divide start_ARG B ( b c c ) end_ARG start_ARG B ( a c c ) + B ( b c c ) end_ARG , 1 - p = divide start_ARG B ( a c c ) end_ARG start_ARG B ( a c c ) + B ( b c c ) end_ARG . So the condition for injectivity holds if and only if Bā¢(bā¢cā¢c)ā γā Bā¢(aā¢cā¢c).ā B(bcc)ā γ· B(acc).B ( b c c ) ā γ ā B ( a c c ) . Now, notice Bā¢(bā¢cā¢c)=ā«Ī»Bā¢(Ī»)ā Bā¢(bā¢cā¢cā£Ī»)ā¢Ī»=ā«Ī»Bā¢(Ī»)ā pbā¢Ī»=pbsubscriptā conditionaldifferential-dsubscriptā subscriptdifferential-dsubscriptB(bcc)= _Ī»B(Ī»)Ā· B(bcc Ī»)dĪ»= _Ī»B% (Ī»)Ā· p_bdĪ»=p_bB ( b c c ) = ā«Ī» B ( Ī» ) ā B ( b c c ⣠λ ) d Ī» = ā«Ī» B ( Ī» ) ā pitalic_b d Ī» = pitalic_b and Bā¢(aā¢cā¢c)=ā«Ī»Bā¢(Ī»)ā¢Bā¢(aā¢cā¢cā£Ī»)ā¢Ī»=ā«Ī»Bā¢(Ī»)ā paā Ī»ā¢Ī»=paā Ī»ā¼Bā¢(Ī»)[Ī»].subscriptconditionaldifferential-dsubscriptā subscriptdifferential-dā subscriptsubscriptsimilar-toB(acc)= _Ī»B(Ī»)B(acc Ī»)dĪ»= _Ī»B(% Ī»)Ā· p_aĀ·Ī» dĪ»=p_aĀ· *E_% Ī» B(Ī») [Ī» ].B ( a c c ) = ā«Ī» B ( Ī» ) B ( a c c ⣠λ ) d Ī» = ā«Ī» B ( Ī» ) ā pitalic_a ā Ī» d Ī» = pitalic_a ā Eitalic_Ī» ā¼ B ( Ī» ) [ Ī» ] . This shows the first result. For the second statement, we explicitly compute JobssubscriptobsJ_obsJroman_obs up to an affine transformation, which does not change the policy ordering. Let R be the true reward function, G=ā”(R)G= (R)G = Ī ( R ) the corresponding return function, and ā”(G) B(G)B ( G ) the resulting return function at the level of observations. For simplicity, assume Rā¢(c)=00R(c)=0R ( c ) = 0, which can always be achieved by adding a constant. We have: Jobsā¢(Ī»)subscriptobs J_obs(Ī»)Jroman_obs ( Ī» ) =sāā¼PĪ»ā¢(sā)[ā”(G)ā¢(Oāā¢(sā))]absentsubscriptsimilar-toāsuperscriptā = *E_ s P^Ī»( s)% [ B(G) ( O( s) ) ]= Eoverā start_ARG s end_ARG ā¼ Pitalic_Ī» ( overā start_ARG s end_ARG ) [ B ( G ) ( overā start_ARG O end_ARG ( overā start_ARG s end_ARG ) ) ] =PĪ»ā¢(aā¢bā¢c)ā ā”(G)ā¢(oā¢oā¢c)+PĪ»ā¢(bā¢cā¢c)ā ā”(G)ā¢(oā¢cā¢c)+PĪ»ā¢(cā¢cā¢c)ā ā”(G)ā¢(cā¢cā¢c)+PĪ»ā¢(aā¢cā¢c)ā ā”(G)ā¢(oā¢cā¢c)absentā superscriptā superscriptā superscriptā superscript =P^Ī»(abc)Ā· B(G)(ooc)+P^% Ī»(bcc)Ā· B(G)(occ)+P^Ī»(c)Ā·% B(G)(c)+P^Ī»(acc)Ā· B% (G)(occ)= Pitalic_Ī» ( a b c ) ā B ( G ) ( o o c ) + Pitalic_Ī» ( b c c ) ā B ( G ) ( o c c ) + Pitalic_Ī» ( c c c ) ā B ( G ) ( c c c ) + Pitalic_Ī» ( a c c ) ā B ( G ) ( o c c ) =paā (1āĪ»)ā Gā¢(aā¢bā¢c)+pbā ā”(G)ā¢(oā¢cā¢c)+pcā Gā¢(cā¢cā¢c)+paā Ī»ā ā”(G)ā¢(oā¢cā¢c)absentā subscript1ā subscriptā subscriptā subscript =p_aĀ·(1-Ī»)Ā· G(abc)+p_bĀ·% B(G)(occ)+p_cĀ· G(c)+p_a·λ·% B(G)(occ)= pitalic_a ā ( 1 - Ī» ) ā G ( a b c ) + pitalic_b ā B ( G ) ( o c c ) + pitalic_c ā G ( c c c ) + pitalic_a ā Ī» ā B ( G ) ( o c c ) āĪ»ā [ā”(G)ā¢(oā¢cā¢c)āGā¢(aā¢bā¢c)].proportional-toabsentā delimited-[] λ· [ B(G)(occ)-G(abc% ) ].ā Ī» ā [ B ( G ) ( o c c ) - G ( a b c ) ] . We have Gā¢(aā¢bā¢c)=Rā¢(a)+γā¢Rā¢(b),ā”(G)ā¢(oā¢cā¢c)=(1āp)ā Gā¢(aā¢cā¢c)+pā Gā¢(bā¢cā¢c)=(1āp)ā Rā¢(a)+pā Rā¢(b).formulae-sequenceā 1ā 1ā G(abc)=R(a)+γ R(b), B(G)(occ)=(1-p)Ā· G(% acc)+pĀ· G(bcc)=(1-p)Ā· R(a)+pĀ· R(b).G ( a b c ) = R ( a ) + γ R ( b ) , B ( G ) ( o c c ) = ( 1 - p ) ā G ( a c c ) + p ā G ( b c c ) = ( 1 - p ) ā R ( a ) + p ā R ( b ) . Thus, the condition ā”(G)ā¢(oā¢cā¢c)>Gā¢(aā¢bā¢c) B(G)(occ)>G(abc)B ( G ) ( o c c ) > G ( a b c ) is equivalent to Rā¢(a)<pāγpā Rā¢(b).ā R(a)< p-γpĀ· R(b).R ( a ) < divide start_ARG p - γ end_ARG start_ARG p end_ARG ā R ( b ) . Thus, we have argā¢maxĪ»ā[0,1]ā”Jobsā¢(Ī»)=1, if ā¢Rā¢(a)<pāγpā Rā¢(b),0, else. subscriptargmax01subscriptobscases1 if ā otherwise0 else. otherwise *arg\,max_Ī»ā[0,1]J_obs(Ī»)=% cases1, if R(a)< p-γpĀ· R(b),\\ 0, else. casesstart_OPERATOR arg max end_OPERATORĪ» ā [ 0 , 1 ] Jroman_obs ( Ī» ) = start_ROW start_CELL 1 , if R ( a ) < divide start_ARG p - γ end_ARG start_ARG p end_ARG ā R ( b ) , end_CELL start_CELL end_CELL end_ROW start_ROW start_CELL 0 , else. end_CELL start_CELL end_CELL end_ROW Now consider the case Rā¢(b)>00R(b)>0R ( b ) > 0. In this case, Ī»=00Ī»=0Ī» = 0 gives rise to the optimal policy according to G since going to b gives extra reward that one misses when going to c directly. However, when Rā¢(a)āŖ0much-less-than0R(a) 0R ( a ) āŖ 0, then JobssubscriptobsJ_obsJroman_obs selects for Ī»=11Ī»=1Ī» = 1. Intuitively, the policy tries to āhide that the episode started in aā by going directly to c, which leads to ambiguity between aā¢cā¢cacca c c and bā¢cā¢cbccb c c. This is a case of deceptive inflation as in Theorem 4.5. Now, consider the case Rā¢(b)<00R(b)<0R ( b ) < 0. In this case, Ī»=11Ī»=1Ī» = 1 gives rise to the optimal policy according to G. However, when Rā¢(a)ā«0much-greater-than0R(a) 0R ( a ) ā« 0, then JobssubscriptobsJ_obsJroman_obs selects for Ī»=00Ī»=0Ī» = 0. Intuitively, the policy tries to āreveal that the episode started with aā by going to b, which is positive information to the human, but negative from the perspective of optimizing G. As in Theorem 4.5, we see that this is a case of overjustification. ā Example D.7. In this example, we consider an MDP thatās similar to a multi-armed bandit with four states/actions a,b,c,da,b,c,da , b , c , d and observation kernel Oā¢(a)=Oā¢(b)ā Oā¢(c)=Oā¢(d)O(a)=O(b)ā O(c)=O(d)O ( a ) = O ( b ) ā O ( c ) = O ( d ). Formally, we can imagine that it is given by the MDP sssaaabbbcccddd with Rā¢(s)=00R(s)=0R ( s ) = 0 and a time-horizon of T=11T=1T = 1. In this example, we reveal that misleadingness and non-optimality (according to the true reward R, or J) are in principle orthogonal concepts. We consider the following four example cases. In each one, we vary some environment parameters and then determine aobsāsubscriptsuperscriptobsa^*_obsaāroman_obs, the action that results from optimizing JobssubscriptobsJ_obsJroman_obs (corresponding to a naive application of RLHF under partial observability, see Section 4.1), its misleadingness Mā”(aobsā)MsubscriptsuperscriptobsM(a^*_obs)M ( aāroman_obs ) (see Definition D.5), and the action aāsuperscripta^*aā that would result from optimizing J. If aobsā=aāsubscriptsuperscriptobssuperscripta^*_obs=a^*aāroman_obs = aā, then JobssubscriptobsJ_obsJroman_obs selects for the optimal action. For simplicity, we can imagine that the human has a uniform prior over what action results eventually (out of the action taken and potentially a deviation defined by ϵitalic-ϵεϵ, see below) is taken before making an observation, i.e. Bā¢(a)=Bā¢(b)=Bā¢(c)=Bā¢(d)=1414B(a)=B(b)=B(c)=B(d)= 14B ( a ) = B ( b ) = B ( c ) = B ( d ) = divide start_ARG 1 end_ARG start_ARG 4 end_ARG. (a) Assume Rā¢(a)>Rā¢(c)>Rā¢(d)ā«Rā¢(b)much-greater-thanR(a)>R(c)>R(d) R(b)R ( a ) > R ( c ) > R ( d ) ā« R ( b ). Also assume that action d leads with probability ϵ>0italic-ϵ0ε>0ϵ > 0 to state b, whereas all other actions lead deterministically to the specified state. Then aobsā=csubscriptsuperscriptobsa^*_obs=caāroman_obs = c, Mā”(c)<0M0M(c)<0M ( c ) < 0 and aā=asuperscripta^*=aā = a. (b) Assume Rā¢(d)>Rā¢(a)>Rā¢(c)ā«Rā¢(b)much-greater-thanR(d)>R(a)>R(c) R(b)R ( d ) > R ( a ) > R ( c ) ā« R ( b ). Again, assume there is a small probability ϵ>0italic-ϵ0ε>0ϵ > 0 that action d leads to state b. Then aobsā=csubscriptsuperscriptobsa^*_obs=caāroman_obs = c, Mā”(c)>0M0M(c)>0M ( c ) > 0, and aā=dsuperscripta^*=daā = d or aā=asuperscripta^*=aā = a, depending on the size of ϵitalic-ϵεϵ. (c) Assume Rā¢(a)>Rā¢(b)>Rā¢(c)>Rā¢(d)R(a)>R(b)>R(c)>R(d)R ( a ) > R ( b ) > R ( c ) > R ( d ). Additionally, assume that there is a large probability ϵ>0italic-ϵ0ε>0ϵ > 0 that action a leads to state d, whereas all other actions lead to whatās specified. If ϵitalic-ϵεϵ is large enough, then aā=bsuperscripta^*=baā = b. Additionally, we have aobsā=bsubscriptsuperscriptobsa^*_obs=baāroman_obs = b and Mā”(b)>0M0M(b)>0M ( b ) > 0. (d) Assume Rā¢(a)>Rā¢(b)>Rā¢(c)>Rā¢(d)R(a)>R(b)>R(c)>R(d)R ( a ) > R ( b ) > R ( c ) > R ( d ). Also, assume some probability ϵ>0italic-ϵ0ε>0ϵ > 0 that action b leads to state d, whereas all other actions lead deterministically to whatās specified. Then aobsā=asubscriptsuperscriptobsa^*_obs=aāroman_obs = a, Mā”(a)<0M0M(a)<0M ( a ) < 0, and aā=asuperscripta^*=aā = a. Overall, we notice: ⢠Example (a) shows a high regret and negative misleadingness of aobsā=csubscriptsuperscriptobsa^*_obs=caāroman_obs = c. The action is better then it seems, but action a would be better still but cannot be selected because it can be confused with the very bad action b. ⢠Example (b) shows a high regret and high misleadingness of aobsā=csubscriptsuperscriptobsa^*_obs=caāroman_obs = c. The action is worse than it seems and also not optimal. ⢠Example (c) shows zero regret and high misleadingness of aobsā=bsubscriptsuperscriptobsa^*_obs=baāroman_obs = b. The action is worse than it seems because it can be confused with a, but it is still the optimal action because a can turn into d. ⢠Example (d) shows zero regret negative misleadingness of aobsā=asubscriptsuperscriptobsa^*_obs=aāroman_obs = a. The action is chosen even though it seems worse than it is, and is also optimal. Thus, we showed all combinations of regret and misleadingness of the action optimized for under JobssubscriptobsJ_obsJroman_obs. We can also notice the following: Examples (a) and (b) only differ in the placement of Rā¢(d)R(d)R ( d ). In particular, the reason that aobsā=csubscriptsuperscriptobsa^*_obs=caāroman_obs = c is structurally the same in both, but the misleadingness changes. This indicates that misleadingness is not on its own contributing to what JobssubscriptobsJ_obsJroman_obs optimizes for. The following is the smallest example we found with the following properties: ⢠There is a unique start state and terminal state. ⢠A naive application of RLHF fails in a way that shows deception and overjustification. ⢠Modeling partial observability resolves the problems. Example D.8. Consider the following graph: AAASSSCCCTTTBBB This depicts an MDP with start state S, terminal state T and possible state sequences Sā¢Tā¢Tā¢T,Sā¢Aā¢Tā¢T,Sā¢Aā¢Cā¢T,Sā¢Cā¢Tā¢T,Sā¢Bā¢Cā¢T,Sā¢Bā¢Tā¢TSTTT,SATT,SACT,SCTT,SBCT,SBTTS T T T , S A T T , S A C T , S C T T , S B C T , S B T T and no discount, i.e. γ=11γ=1γ = 1. Assume that S,B,CS,B,CS , B , C are observed, i.e. Oā¢(S)=SO(S)=SO ( S ) = S, Oā¢(B)=BO(B)=BO ( B ) = B, Oā¢(C)=CO(C)=CO ( C ) = C, and that A and T are ambiguous: Oā¢(A)=Oā¢(T)=XO(A)=O(T)=XO ( A ) = O ( T ) = X. Then there are five observation sequences Sā¢Xā¢Xā¢X,Sā¢Xā¢Cā¢X,Sā¢Cā¢Xā¢X,Sā¢Bā¢Cā¢X,Sā¢Bā¢Xā¢XSXXX,SXCX,SCXX,SBCX,SBXXS X X X , S X C X , S C X X , S B C X , S B X X. Assume that the human can identify all observation sequences except Sā¢Xā¢Xā¢XSXXXS X X X, with belief b=Bā¢(Sā¢Tā¢Tā¢Tā£Sā¢Xā¢Xā¢X)conditionalb=B(STTT SXXX)b = B ( S T T T ⣠S X X X ) and 1āb=Bā¢(Sā¢Aā¢Tā¢Tā£Sā¢Xā¢Xā¢X)1conditional1-b=B(SATT SXXX)1 - b = B ( S A T T ⣠S X X X ). Then the return function is identifiable under these conditions when the humanās belief is correctly modeled. However, for some choices of the true reward function R and transition dynamics of this MDP, we can obtain deceptive or overjustified behavior for a naive application of RLHF. Proof. We apply Corollary C.4. We order states, state sequences, and observation sequences as follows: =S,A,B,C,T,absent =S,A,B,C,T,= S , A , B , C , T , ā Soverā start_ARG S end_ARG =Sā¢Tā¢Tā¢T,Sā¢Aā¢Tā¢T,Sā¢Aā¢Cā¢T,Sā¢Cā¢Tā¢T,Sā¢Bā¢Cā¢T,Sā¢Bā¢Tā¢T,absent =STTT,SATT,SACT,SCTT,SBCT,SBTT,= S T T T , S A T T , S A C T , S C T T , S B C T , S B T T , Ī©āĪ© overā start_ARG Ī© end_ARG =Sā¢Xā¢Xā¢X,Sā¢Xā¢Cā¢X,Sā¢Cā¢Xā¢X,Sā¢Bā¢Cā¢X,Sā¢Bā¢Xā¢X.absent =SXXX,SXCX,SCXX,SBCX,SBXX.= S X X X , S X C X , S C X X , S B C X , S B X X . As can easily be verified, with this ordering the matrices āāĪ©āĆāsuperscriptāāĪ©ā Bā R Ć SB ā blackboard_Roverā start_ARG Ī© end_ARG Ć overā start_ARG S end_ARG and āāāĆsuperscriptāā ā R SĆSĪ ā blackboard_Roverā start_ARG S end_ARG Ć S are given by: =(b1āb0000001000000100000010000001),=(100031100211011100121011110102).formulae-sequencematrix10000001000000100000010000001matrix100031100211011100121011110102 B= pmatrixb&1-b&0&0&0&0\\ 0&0&1&0&0&0\\ 0&0&0&1&0&0\\ 0&0&0&0&1&0\\ 0&0&0&0&0&1 pmatrix, = pmatrix1&% 0&0&0&3\\ 1&1&0&0&2\\ 1&1&0&1&1\\ 1&0&0&1&2\\ 1&0&1&1&1\\ 1&0&1&0&2 pmatrix.B = ( start_ARG start_ROW start_CELL b end_CELL start_CELL 1 - b end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL end_ROW end_ARG ) , Ī = ( start_ARG start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 3 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 2 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 2 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 1 end_CELL start_CELL 1 end_CELL end_ROW start_ROW start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 1 end_CELL start_CELL 0 end_CELL start_CELL 2 end_CELL end_ROW end_ARG ) . To show identifiability, we need to show that kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 . Clearly, the kernel of BB is given by all return functions in āāsuperscriptāā R Sblackboard_Roverā start_ARG S end_ARG that are multiples of Gā²=(bā1,b,0,0,0,0)superscriptā²10000G =(b-1,b,0,0,0,0)Gā² = ( b - 1 , b , 0 , 0 , 0 , 0 ). Assume Gā²āimā”superscriptā²imG Gā² ā im Ī, meaning there is a reward function Rā²āāāsuperscriptā²āāR ā R SRā² ā blackboard_Roverā start_ARG S end_ARG with ā Rā²=Gā²ā superscriptā² Ā· R =G Ī ā Rā² = Gā². We need to deduce from this a contradiction. The assumption means we obtain the following equations: (i)Rā²ā¢(S)+3ā¢Rā²ā¢(T)=bā1,superscriptā²3superscriptā²1 (i)\ \ R (S)+3R (T)=b-1,( i ) Rā² ( S ) + 3 Rā² ( T ) = b - 1 , (iā¢i)Rā²ā¢(S)+Rā²ā¢(A)+2ā¢Rā²ā¢(T)=b,superscriptā²2superscriptā² (i)\ \ R (S)+R (A)+2R (T)=b,( i i ) Rā² ( S ) + Rā² ( A ) + 2 Rā² ( T ) = b , (iā¢iā¢i)Rā²ā¢(S)+Rā²ā¢(A)+Rā²ā¢(C)+Rā²ā¢(T)=0,superscriptā²superscriptā²0 (i)\ \ R (S)+R (A)+R (C)+R (T)=0,( i i i ) Rā² ( S ) + Rā² ( A ) + Rā² ( C ) + Rā² ( T ) = 0 , (iā¢v)Rā²ā¢(S)+Rā²ā¢(C)+2ā¢Rā²ā¢(T)=0,superscriptā²2superscriptā²0 (iv)\ \ R (S)+R (C)+2R (T)=0,( i v ) Rā² ( S ) + Rā² ( C ) + 2 Rā² ( T ) = 0 , (v)Rā²ā¢(S)+Rā²ā¢(B)+Rā²ā¢(C)+Rā²ā¢(T)=0superscriptā²superscriptā²0 (v)\ \ R (S)+R (B)+R (C)+R (T)=0( v ) Rā² ( S ) + Rā² ( B ) + Rā² ( C ) + Rā² ( T ) = 0 (vā¢i)Rā²ā¢(S)+Rā²ā¢(B)+2ā¢Rā²ā¢(T)=0superscriptā²2superscriptā²0 (vi)\ \ R (S)+R (B)+2R (T)=0( v i ) Rā² ( S ) + Rā² ( B ) + 2 Rā² ( T ) = 0 (i) and (v) together imply Rā²ā¢(A)=Rā²ā¢(B)superscriptā²R (A)=R (B)Rā² ( A ) = Rā² ( B ); (iv) and (vi) together imply Rā²ā¢(B)=Rā²ā¢(C)superscriptā²R (B)=R (C)Rā² ( B ) = Rā² ( C ); (v) and (vi) together imply Rā²ā¢(C)=Rā²ā¢(T)superscriptā²R (C)=R (T)Rā² ( C ) = Rā² ( T ); so together, we have Rā²ā¢(A)=Rā²ā¢(T)superscriptā²R (A)=R (T)Rā² ( A ) = Rā² ( T ). Thus, replacing Rā²ā¢(A)superscriptā²R (A)Rā² ( A ) in (i) by Rā²ā¢(T)superscriptā²R (T)Rā² ( T ) and comparing (i) and (i), we obtain bā1=b1b-1=b - 1 = b, a contradiction. Overall, this shows kerā”ā©imā”=0kernelim0 B % =\0\ker B ā© im Ī = 0 , and thus identifiability of the return function by Corollary C.4. Now we investigate the case of unmodeled partial observability. For demonstrating overjustification, assume deterministic transition dynamics in which every arrow in the diagram can be chosen by the policy. Also, assume Rā¢(A)āŖ0much-less-than0R(A) 0R ( A ) āŖ 0, Rā¢(T)>00R(T)>0R ( T ) > 0, Rā¢(S)=00R(S)=0R ( S ) = 0, Rā¢(B)=00R(B)=0R ( B ) = 0, and Rā¢(C)=00R(C)=0R ( C ) = 0. Then the optimal policy chooses the state sequence Sā¢Tā¢Tā¢TSTTTS T T T. However, this trajectory has low observation value since Gobsā¢(Sā¢Tā¢Tā¢T)=(ā G)ā¢(Sā¢Xā¢Xā¢X)=bā¢Gā¢(Sā¢Tā¢Tā¢T)+(1āb)ā¢Gā¢(Sā¢Aā¢Tā¢T)subscriptobsā 1G_obs(STTT)=( BĀ· G)(SXXX)=bG(STTT)% +(1-b)G(SATT)Groman_obs ( S T T T ) = ( B ā G ) ( S X X X ) = b G ( S T T T ) + ( 1 - b ) G ( S A T T ), which is low since Rā¢(A)āŖ0much-less-than0R(A) 0R ( A ) āŖ 0. JobssubscriptobsJ_obsJroman_obs then selects for the suboptimal policies choosing Sā¢Bā¢Tā¢TSBTTS B T T or Sā¢Cā¢Tā¢TSCTTS C T T, which is overjustified behavior that makes sure that the human does not think state A was accessed. For demonstrating deception, assume that Rā¢(A)ā«0much-greater-than0R(A) 0R ( A ) ā« 0, Rā¢(T)<00R(T)<0R ( T ) < 0, Rā¢(S)=Rā¢(B)=Rā¢(C)=00R(S)=R(B)=R(C)=0R ( S ) = R ( B ) = R ( C ) = 0 and that the transition dynamics are such that when the policy attempts to transition from S to A, it will sometimes transition to B, with all other transitions deterministic. In this case, the optimal behavior attempts to enter state A since this has very high value. JobssubscriptobsJ_obsJroman_obs, however, will select for the policy that chooses Sā¢Tā¢Tā¢TSTTTS T T T. This is deceptive behavior. ā Appendix E NeurIPS Paper Checklist 1. Claims Question: Do the main claims made in the abstract and introduction accurately reflect the paperās contributions and scope? Answer: [Yes] Justification: This can be verified by reading the paper. Guidelines: ⢠The answer NA means that the abstract and introduction do not include the claims made in the paper. ⢠The abstract and/or introduction should clearly state the claims made, including the contributions made in the paper and important assumptions and limitations. A No or NA answer to this question will not be perceived well by the reviewers. ⢠The claims made should match theoretical and experimental results, and reflect how much the results can be expected to generalize to other settings. ⢠It is fine to include aspirational goals as motivation as long as it is clear that these goals are not attained by the paper. 2. Limitations Question: Does the paper discuss the limitations of the work performed by the authors? Answer: [Yes] Justification: In Section 6 we have a paragraph on limitations. Guidelines: ⢠The answer NA means that the paper has no limitation while the answer No means that the paper has limitations, but those are not discussed in the paper. ⢠The authors are encouraged to create a separate "Limitations" section in their paper. ⢠The paper should point out any strong assumptions and how robust the results are to violations of these assumptions (e.g., independence assumptions, noiseless settings, model well-specification, asymptotic approximations only holding locally). The authors should reflect on how these assumptions might be violated in practice and what the implications would be. ⢠The authors should reflect on the scope of the claims made, e.g., if the approach was only tested on a few datasets or with a few runs. In general, empirical results often depend on implicit assumptions, which should be articulated. ⢠The authors should reflect on the factors that influence the performance of the approach. For example, a facial recognition algorithm may perform poorly when image resolution is low or images are taken in low lighting. Or a speech-to-text system might not be used reliably to provide closed captions for online lectures because it fails to handle technical jargon. ⢠The authors should discuss the computational efficiency of the proposed algorithms and how they scale with dataset size. ⢠If applicable, the authors should discuss possible limitations of their approach to address problems of privacy and fairness. ⢠While the authors might fear that complete honesty about limitations might be used by reviewers as grounds for rejection, a worse outcome might be that reviewers discover limitations that arenāt acknowledged in the paper. The authors should use their best judgment and recognize that individual actions in favor of transparency play an important role in developing norms that preserve the integrity of the community. Reviewers will be specifically instructed to not penalize honesty concerning limitations. 3. Theory Assumptions and Proofs Question: For each theoretical result, does the paper provide the full set of assumptions and a complete (and correct) proof? Answer: [Yes] Justification: All theorems come with a full set of assumptions, with full proofs in the appendix linked. Sometimes, ābackground assumptionsā, like the fact that we study an underlying MDP with an additional observation kernel POā¢(oāā£sā)subscriptconditionalāP_O( o s)Pitalic_O ( overā start_ARG o end_ARG ⣠overā start_ARG s end_ARG ), or that the human comes with a belief kernel Bā¢(sāā£oā)conditionalāB( s o)B ( overā start_ARG s end_ARG ⣠overā start_ARG o end_ARG ), are omitted in the theorem statements since they apply throughout to the whole paper. Guidelines: ⢠The answer NA means that the paper does not include theoretical results. ⢠All the theorems, formulas, and proofs in the paper should be numbered and cross-referenced. ⢠All assumptions should be clearly stated or referenced in the statement of any theorems. ⢠The proofs can either appear in the main paper or the supplemental material, but if they appear in the supplemental material, the authors are encouraged to provide a short proof sketch to provide intuition. ⢠Inversely, any informal proof provided in the core of the paper should be complemented by formal proofs provided in appendix or supplemental material. ⢠Theorems and Lemmas that the proof relies upon should be properly referenced. 4. Experimental Result Reproducibility Question: Does the paper fully disclose all the information needed to reproduce the main experimental results of the paper to the extent that it affects the main claims and/or conclusions of the paper (regardless of whether the code and data are provided or not)? Answer: [N/A] Justification: The paper does not include experiments. Guidelines: ⢠The answer NA means that the paper does not include experiments. ⢠If the paper includes experiments, a No answer to this question will not be perceived well by the reviewers: Making the paper reproducible is important, regardless of whether the code and data are provided or not. ⢠If the contribution is a dataset and/or model, the authors should describe the steps taken to make their results reproducible or verifiable. ⢠Depending on the contribution, reproducibility can be accomplished in various ways. For example, if the contribution is a novel architecture, describing the architecture fully might suffice, or if the contribution is a specific model and empirical evaluation, it may be necessary to either make it possible for others to replicate the model with the same dataset, or provide access to the model. In general. releasing code and data is often one good way to accomplish this, but reproducibility can also be provided via detailed instructions for how to replicate the results, access to a hosted model (e.g., in the case of a large language model), releasing of a model checkpoint, or other means that are appropriate to the research performed. ⢠While NeurIPS does not require releasing code, the conference does require all submissions to provide some reasonable avenue for reproducibility, which may depend on the nature of the contribution. For example (a) If the contribution is primarily a new algorithm, the paper should make it clear how to reproduce that algorithm. (b) If the contribution is primarily a new model architecture, the paper should describe the architecture clearly and fully. (c) If the contribution is a new model (e.g., a large language model), then there should either be a way to access this model for reproducing the results or a way to reproduce the model (e.g., with an open-source dataset or instructions for how to construct the dataset). (d) We recognize that reproducibility may be tricky in some cases, in which case authors are welcome to describe the particular way they provide for reproducibility. In the case of closed-source models, it may be that access to the model is limited in some way (e.g., to registered users), but it should be possible for other researchers to have some path to reproducing or verifying the results. 5. Open access to data and code Question: Does the paper provide open access to the data and code, with sufficient instructions to faithfully reproduce the main experimental results, as described in supplemental material? Answer: [N/A] Justification: The paper does not include experiments requiring code. Guidelines: ⢠The answer NA means that paper does not include experiments requiring code. ⢠Please see the NeurIPS code and data submission guidelines (https://nips.c/public/guides/CodeSubmissionPolicy) for more details. ⢠While we encourage the release of code and data, we understand that this might not be possible, so āNoā is an acceptable answer. Papers cannot be rejected simply for not including code, unless this is central to the contribution (e.g., for a new open-source benchmark). ⢠The instructions should contain the exact command and environment needed to run to reproduce the results. See the NeurIPS code and data submission guidelines (https://nips.c/public/guides/CodeSubmissionPolicy) for more details. ⢠The authors should provide instructions on data access and preparation, including how to access the raw data, preprocessed data, intermediate data, and generated data, etc. ⢠The authors should provide scripts to reproduce all experimental results for the new proposed method and baselines. If only a subset of experiments are reproducible, they should state which ones are omitted from the script and why. ⢠At submission time, to preserve anonymity, the authors should release anonymized versions (if applicable). ⢠Providing as much information as possible in supplemental material (appended to the paper) is recommended, but including URLs to data and code is permitted. 6. Experimental Setting/Details Question: Does the paper specify all the training and test details (e.g., data splits, hyperparameters, how they were chosen, type of optimizer, etc.) necessary to understand the results? Answer: [N/A] Justification: The paper does not include experiments. Guidelines: ⢠The answer NA means that the paper does not include experiments. ⢠The experimental setting should be presented in the core of the paper to a level of detail that is necessary to appreciate the results and make sense of them. ⢠The full details can be provided either with the code, in appendix, or as supplemental material. 7. Experiment Statistical Significance Question: Does the paper report error bars suitably and correctly defined or other appropriate information about the statistical significance of the experiments? Answer: [N/A] Justification: The paper does not include experiments. Guidelines: ⢠The answer NA means that the paper does not include experiments. ⢠The authors should answer "Yes" if the results are accompanied by error bars, confidence intervals, or statistical significance tests, at least for the experiments that support the main claims of the paper. ⢠The factors of variability that the error bars are capturing should be clearly stated (for example, train/test split, initialization, random drawing of some parameter, or overall run with given experimental conditions). ⢠The method for calculating the error bars should be explained (closed form formula, call to a library function, bootstrap, etc.) ⢠The assumptions made should be given (e.g., Normally distributed errors). ⢠It should be clear whether the error bar is the standard deviation or the standard error of the mean. ⢠It is OK to report 1-sigma error bars, but one should state it. The authors should preferably report a 2-sigma error bar than state that they have a 96% CI, if the hypothesis of Normality of errors is not verified. ⢠For asymmetric distributions, the authors should be careful not to show in tables or figures symmetric error bars that would yield results that are out of range (e.g. negative error rates). ⢠If error bars are reported in tables or plots, The authors should explain in the text how they were calculated and reference the corresponding figures or tables in the text. 8. Experiments Compute Resources Question: For each experiment, does the paper provide sufficient information on the computer resources (type of compute workers, memory, time of execution) needed to reproduce the experiments? Answer: [N/A] Justification: The paper does not include experiments. Guidelines: ⢠The answer NA means that the paper does not include experiments. ⢠The paper should indicate the type of compute workers CPU or GPU, internal cluster, or cloud provider, including relevant memory and storage. ⢠The paper should provide the amount of compute required for each of the individual experimental runs as well as estimate the total compute. ⢠The paper should disclose whether the full research project required more compute than the experiments reported in the paper (e.g., preliminary or failed experiments that didnāt make it into the paper). 9. Code Of Ethics Question: Does the research conducted in the paper conform, in every respect, with the NeurIPS Code of Ethics https://neurips.c/public/EthicsGuidelines? Answer: [Yes] Justification: This research does not involve human subjects, does not make use of data, and does not propose a practical method that could be misused or have a negative impact. As such, the paper does not give rise to any ethical concerns. Guidelines: ⢠The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics. ⢠If the authors answer No, they should explain the special circumstances that require a deviation from the Code of Ethics. ⢠The authors should make sure to preserve anonymity (e.g., if there is a special consideration due to laws or regulations in their jurisdiction). 10. Broader Impacts Question: Does the paper discuss both potential positive societal impacts and negative societal impacts of the work performed? Answer: [Yes] Justification: The last paragraph of the main paper is an impact statement, listing the positive impact we hope to see from our work. As our work is theoretical and does not provide a method, no negative impact arises from it. Guidelines: ⢠The answer NA means that there is no societal impact of the work performed. ⢠If the authors answer NA or No, they should explain why their work has no societal impact or why the paper does not address societal impact. ⢠Examples of negative societal impacts include potential malicious or unintended uses (e.g., disinformation, generating fake profiles, surveillance), fairness considerations (e.g., deployment of technologies that could make decisions that unfairly impact specific groups), privacy considerations, and security considerations. ⢠The conference expects that many papers will be foundational research and not tied to particular applications, let alone deployments. However, if there is a direct path to any negative applications, the authors should point it out. For example, it is legitimate to point out that an improvement in the quality of generative models could be used to generate deepfakes for disinformation. On the other hand, it is not needed to point out that a generic algorithm for optimizing neural networks could enable people to train models that generate Deepfakes faster. ⢠The authors should consider possible harms that could arise when the technology is being used as intended and functioning correctly, harms that could arise when the technology is being used as intended but gives incorrect results, and harms following from (intentional or unintentional) misuse of the technology. ⢠If there are negative societal impacts, the authors could also discuss possible mitigation strategies (e.g., gated release of models, providing defenses in addition to attacks, mechanisms for monitoring misuse, mechanisms to monitor how a system learns from feedback over time, improving the efficiency and accessibility of ML). 11. Safeguards Question: Does the paper describe safeguards that have been put in place for responsible release of data or models that have a high risk for misuse (e.g., pretrained language models, image generators, or scraped datasets)? Answer: [N/A] Justification: The paper poses no such risks. Guidelines: ⢠The answer NA means that the paper poses no such risks. ⢠Released models that have a high risk for misuse or dual-use should be released with necessary safeguards to allow for controlled use of the model, for example by requiring that users adhere to usage guidelines or restrictions to access the model or implementing safety filters. ⢠Datasets that have been scraped from the Internet could pose safety risks. The authors should describe how they avoided releasing unsafe images. ⢠We recognize that providing effective safeguards is challenging, and many papers do not require this, but we encourage authors to take this into account and make a best faith effort. 12. Licenses for existing assets Question: Are the creators or original owners of assets (e.g., code, data, models), used in the paper, properly credited and are the license and terms of use explicitly mentioned and properly respected? Answer: [N/A] Justification: The paper does not use existing assets. Guidelines: ⢠The answer NA means that the paper does not use existing assets. ⢠The authors should cite the original paper that produced the code package or dataset. ⢠The authors should state which version of the asset is used and, if possible, include a URL. ⢠The name of the license (e.g., C-BY 4.0) should be included for each asset. ⢠For scraped data from a particular source (e.g., website), the copyright and terms of service of that source should be provided. ⢠If assets are released, the license, copyright information, and terms of use in the package should be provided. For popular datasets, paperswithcode.com/datasets has curated licenses for some datasets. Their licensing guide can help determine the license of a dataset. ⢠For existing datasets that are re-packaged, both the original license and the license of the derived asset (if it has changed) should be provided. ⢠If this information is not available online, the authors are encouraged to reach out to the assetās creators. 13. New Assets Question: Are new assets introduced in the paper well documented and is the documentation provided alongside the assets? Answer: [N/A] Justification: The paper does not release new assets. Guidelines: ⢠The answer NA means that the paper does not release new assets. ⢠Researchers should communicate the details of the dataset/code/model as part of their submissions via structured templates. This includes details about training, license, limitations, etc. ⢠The paper should discuss whether and how consent was obtained from people whose asset is used. ⢠At submission time, remember to anonymize your assets (if applicable). You can either create an anonymized URL or include an anonymized zip file. 14. Crowdsourcing and Research with Human Subjects Question: For crowdsourcing experiments and research with human subjects, does the paper include the full text of instructions given to participants and screenshots, if applicable, as well as details about compensation (if any)? Answer: [N/A] Justification: The paper does not involve crowdsourcing nor research with human subjects. Guidelines: ⢠The answer NA means that the paper does not involve crowdsourcing nor research with human subjects. ⢠Including this information in the supplemental material is fine, but if the main contribution of the paper involves human subjects, then as much detail as possible should be included in the main paper. ⢠According to the NeurIPS Code of Ethics, workers involved in data collection, curation, or other labor should be paid at least the minimum wage in the country of the data collector. 15. Institutional Review Board (IRB) Approvals or Equivalent for Research with Human Subjects Question: Does the paper describe potential risks incurred by study participants, whether such risks were disclosed to the subjects, and whether Institutional Review Board (IRB) approvals (or an equivalent approval/review based on the requirements of your country or institution) were obtained? Answer: [N/A] Justification: The paper does not involve crowdsourcing nor research with human subjects. Guidelines: ⢠The answer NA means that the paper does not involve crowdsourcing nor research with human subjects. ⢠Depending on the country in which research is conducted, IRB approval (or equivalent) may be required for any human subjects research. If you obtained IRB approval, you should clearly state this in the paper. ⢠We recognize that the procedures for this may vary significantly between institutions and locations, and we expect authors to adhere to the NeurIPS Code of Ethics and the guidelines for their institution. ⢠For initial submissions, do not include any information that would break anonymity (if applicable), such as the institution conducting the review.