Paper deep dive
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response
Elle Lazarski, Jaime FernĂĄndez Fisac
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/4/2026, 3:41:29 AM
Summary
This paper introduces a framework for 'Corrigible Assistance in One Round' using pragmatic-pedagogic reasoning to resolve goal uncertainty in human-robot collaboration. It identifies a class of assistance games where the robot can infer the human's goal in a single time step, overcoming the 'inference ceiling' inherent in mainstream Inverse Optimal Control (IOC). The authors propose the Pragmatic-Pedagogic Best Response (PPBR) algorithm, which achieves optimal equilibrium by having the human act pedagogically and the robot interpret actions pragmatically. Theoretical results are validated on a collaborative block-building task.
Entities (9)
Relation Signals (7)
Pragmatic-Pedagogic Best Response â solves â Assistance Games
confidence 95% ¡ We propose PragmaticâPedagogic Best Response (PPBR), an algorithm that directly computes the equilibrium solution to an action-separable assistance game
Pragmatic-Pedagogic Reasoning â overcomes â Inference Ceiling
confidence 93% ¡ pragmaticâpedagogic reasoning overcomes this barrier by immediately disambiguating goals
R3 Robot â uses â Pragmatic-Pedagogic Reasoning
confidence 92% ¡ R3: Pragmatic Robot Helper... updates its belief... using the H2 likelihood model
Inverse Optimal Control â exhibits â Inference Ceiling
confidence 90% ¡ mainstream inverse optimal control exhibits an inference ceiling that hinders alignment
H2 Human â exhibits â Pedagogic Behavior
confidence 90% ¡ H2: Pedagogic Human User... H2 anticipates R1âs belief update... evaluating a H under a given hypothesis
Pragmatic-Pedagogic Best Response â validatedon â Collaborative Block-Building
confidence 88% ¡ validate our theoretical results and proposed method on a simple collaborative block-building example
Assistance Games â modeledas â POMDP
confidence 85% ¡ exact solutions require planning in a POMDP
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Assistance games formalize human-robot collaboration under asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effectively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We identify a class of assistance games in which pragmatic-pedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmatic-pedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example.
Tags
Links
- Source: https://arxiv.org/abs/2607.27508v1
- Canonical: https://arxiv.org/abs/2607.27508v1
Trouble viewing inline? Open PDF directly â
Full Text
44,610 characters extracted from source content.
Expand or collapse full text
Corrigible Assistance in One Round: PragmaticâPedagogic Best Response Elle Lazarski 1 and Jaime FernĂĄndez Fisac 1 Department of Electrical and Computer Engineering Princeton University, USA Abstract. Assistance games formalize humanârobot collaboration un- der asymmetric information: the human knows the goal, while the robot must infer it from observation and interaction in order to assist effec- tively. In general, computing optimal assistance game strategies online is intractable, since exact solutions require planning in a POMDP. We iden- tify a class of assistance games in which pragmaticâpedagogic reasoning resolves goal uncertainty in a single time step, rendering the full-horizon game exactly solvable by a tractable best-response procedure. Within this class, we show that mainstream inverse optimal control exhibits an inference ceiling that hinders alignment, while pragmaticâpedagogic reasoning overcomes this barrier by immediately disambiguating goals through actions that look equivalent under task execution alone. Finally, we validate our theoretical results and proposed method on a simple collaborative block-building example. Keywords: HumanâRobot Interaction¡ Value Alignment¡ Mathemat- ical Modeling and Analysis 1 Introduction As robots become more capable and are deployed in increasingly varied contexts, they must infer and adapt to their usersâ needs online rather than execute pre- specified routines. In artificial intelligence (AI), conventional alignment pipelines like Reinforcement Learning from Human Feedback (RLHF) optimize for human approval of generated outputs, such as through thumbs-up or A/B preference feedback [1]. However, this signal is known to be an imperfect, often problematic proxy, since it tends to neglect the downstream effects of individual decisions [2â 4]. In humanârobot interaction, the coupling between robot operation and hu- man behavior over time makes it all the more necessary to approach alignment with respect to long-term outcomes instead of isolated robot actions [5â8]. To capture the temporal dimension of alignment, assistance games [9, 10] pose the problem as a dynamic two-player collaboration between a human user and a robot assistant, who aims to help realize the humanâs objective but is uncertain about what it is. The resulting equilibrium solutions have been shown to present desirable properties, which extend the notion of optimal information seeking to the two-player setting. In particular, the humanâs actions often carry arXiv:2607.27508v1 [cs.RO] 29 Jul 2026 2E. Lazarski et al. out a pedagogic function, strategically conveying actionable information about the goal, while the robotâs responses are pragmatic, interpreting human cues as purposefully (rather than circumstantially) communicative [11, 12], consistent with modern cognitive science accounts of human teaching and learning [13]. Unfortunately, despite their theoretical strengths, assistance game solutions are largely considered computationally intractable, due to the need to plan in the robotâs information space [14]. In this work, we identify a specialâbut nontrivialâclass of assistance games in which pragmaticâpedagogic solutions are simultaneously highly effective and easily computable, fully disambiguating the humanâs goal in a single time step. Consequently, for this class of problems, an efficient short-horizon approxima- tion (analogous to single-player QMDP [15]) yields an optimal strategy pair for the full-horizon game. We further show that this equilibrium can be readily ob- tained in a single round of player best responses starting from a mainstream in- verse optimal control (IOC)âor inverse reinforcement learning (IRL)âsolution. Crucially, while IOC often exhibits an inference ceiling that impedes one-step alignment, the pragmaticâpedagogic solution fully overcomes this limitation, uniquely disambiguating goals through actions that appear equivalent from the standpoint of task execution alone. As a result, the pragmatic robot strategy is maximally empowering, affording the human the option to convey and achieve any goal regardless of the robotâs initial belief, and rendering the robotâs opera- tion corrigible, a key requirement for robust alignment and long-term safety [16]. Our contributions in this work can be summarized as follows: 1. One-round convergence to a pragmaticâpedagogic equilibrium. We quantify the IOC inference ceiling in terms of a posterior belief bound and show that pragmaticâpedagogic reasoning can surpass it. We formalize the conditions that enable goal disambiguation in one time step and prove that one round of best responses reaches an optimal pragmaticâpedagogic equi- librium for the class of action-separable assistance games. 2. A practical methodology for tractable pragmatic assistance. We propose PragmaticâPedagogic Best Response (PPBR), an algorithm that directly computes the equilibrium solution to an action-separable assistance game, yielding a human-empowering robot strategy. 3. Empirical validation. We validate our theoretical results empirically in a collaborative block-building domain, comparing IOC against PPBR under a comprehensive range of initial conditions. 2 Related Work Traditional IOC/IRL aims to recover the objective pursued by a human âexpertâ by observing demonstrations of their behavior, assumed to take place in isola- tion [17â19]. In interactive settings, however, the robot is not a passive observer but an active player whose own actions may affect the humanâs outcome. The human may therefore choose actions not only to directly generate utility but also to indirectly improve expected outcomes by influencing the robot. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response3 Assistance games, initially introduced under the name cooperative inverse reinforcement learning (CIRL), formalize this idea as a cooperative game with asymmetric information [9, 10]. In particular, the human knows the true goal, while the robot must infer it from observation and interaction. Solving a CIRL game can be reduced to solving a partially observable Markov decision process (POMDP), in which the robotâs belief over goals serves as a sufficient statistic for optimal decision-making [10]. However, this reduction inherits the computational burden of POMDP planning, barring its use in practical problems. Subsequent research in assistance games showed that optimal humanârobot strategy pairs must satisfy a specific dynamic programming relation expressed as a fixed point in a pragmaticâpedagogic Bellman recursion, whereby the hu- man chooses actions pedagogically to convey strategically relevant information to a suitably attuned robot partner, and the robot in turn interprets human actions pragmatically, treating them as cues purposefully chosen to convey in- formation [11]. Despite an exponential complexity improvement over the naĂŻve POMDP reduction, the resulting belief-space planning is still intractable for runtime computation of assistance strategies. Here, we build on the theoretical insights established by prior assistance game efforts to investigate runtime-computable robot strategies that preserve coop- erative structure while enabling online assistance. We combine classical short- horizon approximation ideas from decision theory [15] and truncated iterated- best-response approximations from behavioral game theory [20â22], and show that they allow us to exactly solve the full-horizon assistance game for a class of problems in which the robotâs uncertainty can be resolved in a single time step. Finally, some work in the AI alignment literature has brought into question the usefulness of pragmatic robot behavior, due to potentially higher sensitivity to modeling assumptions (in particular, whether or not the human user intends to behave pedagogically) [23]. Our results shed new light on the matter, suggest- ing an alternative perspective: regardless of modeling accuracy, robot strategies derived from pragmaticâpedagogic solutions make the robot comparatively more responsive to human action cues, increasing the humanâs effective controllability over outcomes and mitigating known overconfidence issues with non-pragmatic (sometimes called âliteralâ) goal-inferring robots (typically IOC) [7, 24, 25]. 3 Problem Formulation: Assistance Games We study assistance gamesG in which a human and a robot take turns acting in a shared environment to achieve the humanâs objective. Critically, this objective is known to the human but, a priori, unknown to the robot. Let G = S,A H ,A R ,T s ,Î,R,P 0 ,Îł , where S denotes the space of world states; A H , A R are the human and robot action spaces; Î is a set of possible goal parameters encoding the humanâs objec- tive; T s (s t+1 | s t ,a H t ,a R t ) is the transition probability measure; R(s t ,a H t ,a R t ;θ) 4E. Lazarski et al. is the shared reward function, parameterized by θ â Î (whose value is only ob- served by the human); P 0 (s 0 ,θ) is the initial joint distribution over states and goals; and Îł â [0, 1] is a discount factor. At each time step t, the human selects a H t âA H first, after which the robot observes a H t and selects a R t âA R ; the joint action (a H t ,a R t ) then induces a successor state through T s . The robot maintains a Bayesian posterior belief b + t â â(Î) over candidate goal hypotheses, which is updated after each observed a H t . Throughout the paper, we let b t denote the robotâs prior belief at the start of time step t, and we use b + t for the robotâs posterior after observing the humanâs action, specifying the robotâs belief update model (e.g., b R1+ t ) when relevant. Following the assistance game literature, P 0 is known to both players, which means that b + t can be computed by either player as a sufficient statistic of the robotâs information state after observing (s 0 ,a H 0 ,a R 0 ,...,s t ,a H t ) under any particular human policy. Since the robotâs belief may inform its behavior, it is also strategically relevant to the human. Therefore, the humanâs optimal assistance game strategy will in general be a stochastic policy Ď H (a H t | s t ,b t ;θ), whereas the robotânot privy to θ, but observing a H t before actingâmust choose a policy Ď R (a R t | s t ,b + t ,a H t ). Solving the assistance game amounts to finding a team strategy Ď := (Ď H ,Ď R ) that maximizes the expected time-discounted return J (Ď H ,Ď R ) :=E (s 0 ,θ)âźP 0 Ďâź(Ď,T) " â X t=0 Îł t R(s t ,a H t ,a R t ;θ) # , where Ď := (s 0 ,a H 0 ,a R 0 ,s 1 ,a H 1 ,a R 1 ,... ) denotes the gameplay trajectory, whose distribution is sequentially induced by team strategy Ď and (s t ,b t )-transitions T := (T s ,T b ), with (s 0 ,θ) drawn jointly from P 0 , and b 0 defined as the conditional distribution on θ given s 0 . We omit time indices when clear from context. The dependence of T b on Ď H is a distinctive feature of assistance games: belief transitions depend not only on the observed human action but, indirectly, on the entire human policy, which determines the observation likelihood model [11]. Running example: We consider a collaborative block-building task in which the humanâs latent goal θ â â , specifies a target Tetris shape. The human is indifferent to where the structure is built and what direction it faces, as long as it is upright. To avoid known action-multiplicity artifacts in Boltzmann likelihood models [26], we represent states and actions in the quotient spaces induced by the problemâs underlying symmetries: states that differ only by an SE(2) transform (translation and rotation on the ground plane) are treated as equivalent, as are actions that induce equivalent state changes. In our two-goal example, the quo- tient action space A(s) in most intermediate states of interest (s = ,,,... ) contains two block placements (side/â and up/â). Suboptimal play: While other actions are possible in principle (e.g., placing a new block far away from the existing ones), we exclude them from our analysis because they are exponentially less likely under all hypotheses. A robot observing such actions would typically lose situational confidence and choose not to act [7]. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response5 4 Approach: One Time Step, One Best-Response Round 4.1 Alignment in One Time Step Extending the QMDP approximation in single-agent POMDPs to the two-player setting, we decompose the assistance gameâs horizon into a present uncertain phase and a hypothesized later phase in which the robot has gained full cer- tainty about the humanâs true goal. While this is, in general, an optimistic approximation of the problem (rarely tight for typical POMDPs), we will show in the next section that it is in fact exact in task settings with comparatively mild asymmetry conditions. Intuitively, if both players can independently arrive at an unambiguous correspondence between goals and actions, the robot will become fully confident in the humanâs true goal after a single time step, thereby rendering the one-step goal disambiguation assumption accurate. Perfect-Information Subgame. We can readily compute the oracle value V O (s;θ) that characterizes the fully observed phase of planning. With the robot also privy to the humanâs true goal, the assistance game reduces to a single-player MDP over the joint action space A H ĂA R , where the two players act as a centralized âhive mind.â The goal-parameterized stateâaction value for the team is Q O (s,a H ,a R ;θ) := r(s,a H ,a R ;θ) + ÎłE V O (s Ⲡ;θ) ,(1) with s Ⲡ⟠T s (¡| s,a H ,a R ). 4.2 Equilibrium in One Round We operationalize pragmaticâpedagogic inference from the perspective of a robot helper that reasons about human behavior using a truncated hierarchy of hy- pothesized player models. Specifically, the level-3 robot, denoted R3, simulates H0, R1, and H2 as follows. H0: Solipsistic Human Expert. Under a given hypothesis θ, H0 acts solely to optimize for that objective without considering how their actions influence the robotâs belief or behavior. Hence, H0 serves as the âexpert demonstratorâ in classical IOC/IRL methods. Let V H0 (s;θ) represent the maximum expected reward-to-go from state s when the human works alone. This single-player MDP over A H yields the solip- sistic stateâaction value Q H0 (s,a H ;θ) := r(s,a H ;θ) + ÎłE V H0 (s Ⲡ;θ) , with s Ⲡ⟠T s (¡| s,a H ). Then Ď H0 (a H | s;θ)â exp β H Q H0 (s,a H ;θ) ,(2) where β H > 0 is an inverse-temperature or ârationalityâ parameter controlling how strongly H0 favors high-value actions under θ: as β H â 0, the policy ap- proaches a uniform distribution over all actions; as β H â â, the policy con- centrates on optimal actions, converging to a uniform distribution over only the arg max set. 6E. Lazarski et al. R1: NaĂŻve Robot Learner. R1 assumes the human behaves as H0 and updates its belief b by Bayesâ rule under likelihood Ď H0 (a H | s;θ), yielding b R1+ (¡| a H ). Define Q R1 (s,b R1+ ,a H ,a R ) :=E θâźb R1+ (¡|a H ) Q O (s,a H ,a R ;θ) , and Ď R1 (a R | s,b R1+ ,a H )â exp β R Q R1 (s,b R1+ ,a H ,a R ) ,(3) where β R > 0 is the robotâs inverse-temperature parameter. Note that the robotâs action is informed by its posterior belief after observing the humanâs action. H2: Pedagogic Human User. For each candidate human action a H , H2 antic- ipates R1âs belief update and subsequent behavior, evaluating a H under a given hypothesis θ by taking an expectation over the induced R1 response distribution. Thus, a H is valuable to H2 both insofar as it directly advances θ and insofar as it elicits better robot follow-up under θ. H2âs choice of action is therefore strategic rather than purely task-directed. Define Q H2 (s,b,a H ;θ) :=E a R âźĎ R1 (¡|s,b R1+ ,a H ) Q O (s,a H ,a R ;θ) , and Ď H2 (a H | s,b;θ)â exp β H Q H2 (s,b,a H ;θ) .(4) R3: Pragmatic Robot Helper. After observing a H , R3 updates its belief b by Bayesâ rule using the H2 likelihood model Ď H2 (a H | s,b;θ), yielding b R3+ (¡| a H ). It then selects its responseâwhich is executed in the real environmentâby maximizing expected value under the posterior: a R3â := argmax a R âA R (s) E θâźb R3+ (¡|a H ) Q O (s,a H ,a R ;θ) .(5) We will later show that, for the problem class we consider, truncating this level-k hierarchy at R3 already reaches an optimal pragmaticâpedagogic equilib- rium. Hence, there is no need to iterate further levels (H4, R5, and so on). 5 Analysis: PragmaticâPedagogic One-Step Alignment In this section, we establish important properties of the R3âH2 solution and the conditions under which it allows us to recover an optimal equilibrium of the full-horizon assistance game. 5.1 IOC Inference Ceiling and Incorrigibility We begin by examining the fundamental alignment limitations of IOC/IRL in the assistance context. Let A H0â (s;θ) := argmax a H âA H (s) Q H0 (s,a H ;θ) denote the set of H0-optimal actions under goal θ in state s. In general, an IOC inference ceiling appears whenever one goalâs H0-optimal action set is a subset of the H0-optimal action set for another goal. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response7 Two Candidate Goals We first formalize this pathology in the simplest pos- sible setting with only two candidate goal hypotheses. Lemma 1 (Two-goal IOC inference ceiling). Let Î = θ â ,θ Ⲡ, and sup- pose that a â â A H (s) is H0-optimal for both goals, with a â â A H0â (s;θ â ) â A H0â (s;θ Ⲡ). Assume further that |A H0â (s;θ â )| = k and |A H0â (s;θ Ⲡ)| = m, where 1⤠k ⤠m. Then, R1âs posterior belief in θ â after observing a â from conditions (s,b) converges, as β H ââ, to b R1+ â (θ â | a â ) := lim β H ââ b R1+ (θ â | a â ) = mb(θ â ) mb(θ â ) + k b(θ Ⲡ) ,(6) which bounds b R1+ â strictly below 1. Proof. As β H â â, Ď H0 (a â | s;θ â ) â 1 k and Ď H0 (a â | s;θ Ⲡ) â 1 m . Substituting these limiting likelihoods into the R1 Bayes update and normalizing gives (6). ââ Remark 1. If the two goals assign equal H0 likelihood to a â , R1 learns no new in- formation to disambiguate between θ â and θ Ⲡ. In particular, when the optimal ac- tion sets are identical, (6) leaves the prior unaltered, yielding lim β H ââ b R1+ (θ â | a â ) = b(θ â ). Running example: The humanâs first actionâplacing the anchor blockâis neces- sarily uninformative to the robot. Since by assumption the human is indifferent to position and northâsouthâeastâwest orientation, all feasible initial block place- ments belong to the same equivalence class, and placing a block is the unique optimal action for any θ â , (cf. Remark 1). The robot can select any follow-up action a R ââ,â arbitrarily regardless of its prior belief, since both have the same expected value under the two goal hypotheses (both goals include a block above and adjacent to the anchor block). For our subsequent analysis, we assume the robot places above the anchor, i.e., a R =â. The alternate case is symmetrical in structure. We fix the resulting partial state s = as the reference state for the remainder of the paper. From Lemma 1, with k = 1 (since â is uniquely H0-optimal under) and m = 2 (since â and â are both H0-optimal under), the IOC ceiling is b R1+ â ( |â) = 2b() 1 + b() .(7) When b( ) < 1 3 , this ceiling lies below the robotâs one-step decision bound- ary (which is b = 1 2 ). Importantly, no amount of assumed human rationality (β H ⍠1) can push the one-step IOC posterior beyond 1 2 . This means that, if the humanâs true goal is θ â = and the robot starts off placing more than 2 3 belief on , there is no course of action available to the human to prevent the robot from building the wrong structure. Specifically, if the human places a block on top of the partial structure (a H =â), the IOC robotâs optimal response will place a block on the side 8E. Lazarski et al. (a R =â); and if the human places a block on the side (a H =â), the IOC robot will confidently follow up with a final block on top (a R =â); either scenario re- sults in completing the structure. In other words, even in simple assistance games, the IOC robot strategy is short-term incorrigible from seemingly benign initial conditions. Multi-Goal Assistance Games We now express the inference ceiling in com- plete generality. Lemma 2 (General IOC inference ceiling). Let a â â A H (s), and assume that there exists some candidate goal Ě Î¸ â Î such that a â â A H0â (s; Ě Î¸). Then, for every θ â Î, R1âs posterior belief after observing a â from conditions (s,b) converges, as β H ââ, to b R1+ â (θ | a â ) := lim β H ââ b R1+ (θ | a â ) = b(θ) |A H0â (s;θ)| 1 a â âA H0â (s;θ) X θ ⲠâÎ b(θ Ⲡ) |A H0â (s;θ Ⲡ)| 1 a â âA H0â (s;θ Ⲡ) . (8) In particular, b R1+ â < 1 whenever a â is H0-optimal under θ and at least one alternative θ Ⲡ̸= θ. Proof. As β H ââ, the H0 softmax policy converges to the uniform distribution over the arg max action set: Ď H0 (a â | s;θ) â 1 |A H0â (s;θ)| 1 a â âA H0â (s;θ) . Substituting into the R1 Bayes update and normalizing gives (8).ââ 5.2 Pedagogic Leverage and Corrigibility We next show that pragmaticâpedagogic inference can break the belief ceiling experienced by the IOC robot R1. Since H2 evaluates actions by anticipating R1âs response, actions that are equivalent from H0âs solipsistic perspective need not remain equivalent under H2âs reasoning. Definition 1 (Level-k advantage). Let Hk be a human decision model with stateâaction value Q Hk (s,b,a H ;θ). For any two human actions a, Ěa â A H (s), the level-k advantage of a against Ěa under goal θ from conditions (s,b) is the difference in their expected value: A Hk θ (a, Ěa;s,b) := Q Hk (s,b,a;θ)â Q Hk (s,b, Ěa;θ). The level-k advantage measures how strongly Hk prefers a relative to Ěa under a given hypothesis θ. Note that Definition 1 can be applied to Q H0 by simply Corrigible Assistance in One Round: PragmaticâPedagogic Best Response9 ignoring b, since the solipsistic stateâaction value does not depend on the robotâs belief. When k = 2, we refer to this quantity as the pedagogic advantage. We emphasize that, in the Boltzmann-rational setting, Q H2 is a function of both playersâ rationality parameters: β H influences R1âs belief update after observing a H , and β R determines its subsequent response probabilities. This is in contrast with Q H0 , which depends on neither β H nor β R . Throughout the paper, we compute Q H2 in the rational human limit (β H ââ). Focusing on H2, we now characterize when human actions serve as unam- biguous goal-identifying signals. In particular, every action that is H2-optimal under goal θ must be strictly suboptimal under all θ Ⲡ̸= θ. Equivalently, the H2- optimal action sets for distinct goals must be disjoint. Checking this condition for each candidate goal yields the pedagogic leverage map, defined as follows. Definition 2 (Pedagogic leverage map). The pedagogic leverage map from conditions (s,b) is the set-valued map L H2 s,b : ÎâA H (s) defined, for each θ â Î, by L H2 s,b (θ) :=    A H2â (s;θ), if A H2â (s;θ)âŠA H2â (s;θ Ⲡ) =â âθ Ⲡ̸= θ, â ,otherwise, where A H2â (s;θ) := arg max a H âA H (s) Q H2 (s,b,a H ;θ). We now show that whenever an observed action has pedagogic leverage to- ward a particular goal, R3âs posterior concentrates on that goal as β H ââ. Proposition 1 (Pedagogic leverage enables full one-step alignment). Fix any β R > 0, and suppose that action a H âL H2 s,b (θ). Then, b R3+ â (θ | a H ) := lim β H ââ b R3+ (θ | a H ) = 1.(9) Proof. By Definition 2, a H âA H2â (s;θ) and a H /âA H2â (s;θ Ⲡ) for all θ Ⲡ̸= θ, so Ď H2 (a H | s,b;θ)â 1 |A H2â (s;θ)| > 0, Ď H2 (a H | s,b;θ Ⲡ)â 0 âθ Ⲡ̸= θ, as β H â â. Substituting these limiting likelihoods into the R3 Bayes update and normalizing gives (9).ââ Corollary 1 (Pedagogic leverage breaks the IOC ceiling). If Lemma 1 holds for θ â ,θ Ⲡand action a â âL H2 s,b (θ â ), then b R1+ â (θ â | a â ) = mb(θ â ) mb(θ â ) + k b(θ Ⲡ) while b R3+ â (θ â | a â ) = 1. 10E. Lazarski et al. Running example: Return to the two-goal Tetris example with θ â , and a H ,a R â â,â. We once again analyze the robotâs inference from reference state s = , after the first two blocks have been placed. We use an additive team reward: each correct placement gives +1, and each incorrect placement givesâ1, so the joint reward takes values in â2, 0, +2 depending on whether 0, 1, or 2 blocks are placed correctly. In the one-step setting (Îł = 0), Q O (s,a H ,a R ;θ) = R θ (a H ,a R ) is the immediate joint reward, given by R = 20 0 â2 , R = 02 20 , where rows denote a H ââ,â and columns denote a R ââ,â. Fix any prior b( ) â (0, 1), and let the observed human action be a H =â, which results in the IOC inference ceiling (7). Direct evaluation of Q H2 yields the arg max action sets A H2â (s;) = â and A H2â (s;) = â as β H â â for all β R > 0. Therefore, by Definition 2, a H =â acquires pedagogic leverage toward: L H2 s,b ( ) =â. Likewise, L H2 s,b ( ) =â. Hence, by Corollary 1, b R3+ â (|â) = 1. Crucially, even in the prior regime b( ) < 1 3 , the pragmatic (R3) robotâs optimal response will place a block on the side (a R =â) as β H ââ, completing. R3 is thus one-step corrigible from any initial conditions. The emergent convention of a H =â signaling(and a H =â signaling) can be viewed as a Schelling focal point. 1 See Figure 1. 5.3 Vanishing Pedagogic Leverage and Scaling Behavior We are primarily interested in the limiting behavior of rational agents R3âH2, which we note is obtained by having H2 reason about a noisily rational R1. For finite β R > 0, R1âs softmax has full support, so all robot responsesâincluding follow-ups that are suboptimal under its posteriorâcontribute to H2âs expected value. Consequently, even small posterior shifts (bounded by the IOC ceiling) can induce Q H2 gradients that cause human actions to acquire pedagogic lever- age (Definition 2). This leverage might vanish if we simply plugged in R1âs deterministic best response to each human action. We study these effects in the two-action setting with a, ĚaâA H (s) and target goal θ, writing the log-odds of H2âs softmax policy in terms of the pedagogic advantage (Definition 1): log Ď H2 (a| s,b;θ) Ď H2 ( Ěa| s,b;θ) = β H A H2 θ (a, Ěa;s,b). 1 Coordination problems with multiple equilibria often admit Schelling focal points [27, 28], or prominent solutions that players gravitate toward based on shared intuition or salienceâfor example, choosing â12:00 PMâ as a default meeting time. Here, the solipsistic (H0) human serves as the implicit saliency model that the human and robot use to break possible ambiguity. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response11 Fig. 1: One-step posterior b + (|â) at s =as a function of human rational- ity β H for representative priors b() â (0, 1). The IOC update (R1, orange) saturates at the posterior ceiling of Lemma 1, while the pragmatic update (R3, green; β R = 1) crosses the decision boundary b = 1 2 (black dashed line) and approaches certainty in as β H increases. Importantly, A H2 θ depends on β R through R1âs softmax policy, which is used to define Q H2 . If A H2 θ vanishes in some regime of β R , then driving Ď H2 close to 1 (for A H2 θ > 0) or to 0 (for A H2 θ < 0) requires β H |A H2 θ |ââ, i.e., β H must scale faster than the reciprocal rate. We demonstrate this scaling behavior in our running example at finite β H ,β R values; see Figure 2 and our analysis below. However, we emphasize that even a vanishing amount of leverage will be exploited with probability 1 by H2 in the limit of full rationality (β H ââ). Running example: At s = , let A θ := A H2 θ (â,â;s,b), x := Ď R1 (â| s,b R1+ ,â), and y := Ď R1 (â| s,b R1+ ,â). The reward matrices give Q H2 (s,b,â; ) = 2x,Q H2 (s,b,â; ) =â2y, Q H2 (s,b,â;) = 2(1â x), Q H2 (s,b,â;) = 2(1â y), so A= 2(x + y) and âA= 2(xâ y). A> 0 makes a H =â advantageous under , while âA> 0 makes a H =â disadvantageous under. Exponential regime (b( )â (0, 1 3 ), β R ââ). In the limit β H ââ, R1âs posterior after observing a H =â converges to the IOC ceiling p := b R1+ â (|â) (7), and R1 assigns values 2p and 2(1â p) to follow-up actions a R =â and a R =â, respectively. The value gap favoring a R =â over a R =â is Îş := 2(1â p)â 2p = 2â 4p = 2(1â 3b( )) 1 + b() â (0, 2). 12E. Lazarski et al. Therefore, lim β H ââ x = (1 + e β R Îş ) â1 . After observing a H =â, the value gap favoring a R =â over a R =â is 2, so y = (1 + e 2β R ) â1 . These expressions yield x = e âβ R Îş + o(e âβ R Îş ) and y = e â2β R + o(e â2β R ). Since 0 < Îş < 2, e â2β R = o(e âβ R Îş ), so x decays more slowly than y and governs the leading-order scaling of both pedagogic advantages. Thus, A= 2e âβ R Îş + o(e âβ R Îş ) and âA= 2e âβ R Îş + o(e âβ R Îş ). Note that these quantities are not exactly equal; their difference, Aâ (âA) = 4y, is lower-order. Importantly, both advantages vanish at the same rate, so a H =â unambiguously signalswhen β H e âβ R Îş ââ; a sufficient scaling condition is 1 β H = o(e âβ R Îş ), b()â (0, 1 3 ); β R ââ. Benign regime (b()â [ 1 3 , 1), β R ââ). No exponential growth of β H in β R is required. Hyperbolic regime (β R â 0). Again let p := b R1+ â ( |â) (7), with fixed b( )â (0, 1). Expanding R1âs softmax policy around β R = 0 gives x = 1 2 + (pâ 1 2 )β R + o(β R ), y = 1 2 â 1 2 β R + o(β R ). Hence, A= 2â2(1âp)β R +o(β R ) andâA= 2pβ R +o(β R ). a H =â therefore remains strongly preferred underas β R â 0, since Aâ 2, but suppressing Ď H2 (â|) requires β H β R ââ; a sufficient scaling condition is 1 β H = o(β R ), β R â 0. 5.4 Zero-Cost Pedagogy Enables Equilibrium-in-One We note that, in general, performing one round of pragmaticâpedagogic reason- ing does not necessarily yield an optimal assistance game strategy pair. However, if the resulting R3âH2 policies attain the oracle value from the current state un- der every possible human goal, then they are necessarily optimal. Indeed, no team strategy can exceed V O (s;θ), which upper-bounds the optimal value of the assistance game, V G (s,b;θ). Building on these insights, we identify a class of games in which the human and the robot arrive at an unambiguous pedagogic leverage map at zero cost to the human: signaling any goal requires no sacrifice in value. In these games, pedagogy is âfree,â allowing the humanârobot team to achieve exactly the oracle value from state s as if the robot already knew the humanâs true goal. We formalize this property below. Definition 3 (Zero-cost pedagogy). A fully populated leverage mapL achieves zero-cost pedagogy at state s if, for each possible human goal θ â Î, Q O (s,a H ,a R,O ;θ) = V O (s;θ) âa H âL(θ), where a R,O â arg max a R âA R (s) Q O (s,a H ,a R ;θ). Corrigible Assistance in One Round: PragmaticâPedagogic Best Response13 Fig. 2: One-step pragmatic (R3) posterior b R3+ (|â) at s =as a function of β R ,β H with prior b() = 0.01. The phase boundary from low to near-certain belief is consistent with the derived scaling regimes: hyperbolic growth as β R â 0 and exponential growth as β R ââ. The relevant class of assistance games immediately follows. Definition 4 (Action-separable assistance game). Let L H2 s,b denote the pedagogic leverage map as β H ,β R ââ. An assistance gameG is action-separable at (s,b) if L H2 s,b is fully populated and achieves zero-cost pedagogy (Definition 3). We now prove that, for an action-separable assistance game, one round of pragmaticâpedagogic reasoning suffices to reach an optimal equilibrium solution. As we show in the next section, this solution is easily computable (Algorithm 1). Theorem 1 (Pragmaticâpedagogic equilibrium in one best response). Fix a state s and prior b over Î. Suppose that the assistance game is action- separable at (s,b) (Definition 4). Then, the limiting rational (β H ,β R ââ) R3â H2 policies (Ď R3â ,Ď H2â ) constitute an optimal pragmaticâpedagogic equilibrium of the assistance game at (s,b). Proof. Fix any true goal θ â â Î. Since the game is action-separable, the peda- gogic leverage map is fully populated, so L H2 s,b (θ â ) is nonempty; fix any human action a â âL H2 s,b (θ â ). By Proposition 1, observing a â concentrates R3âs posterior on θ â as β H â â. Because L H2 s,b achieves zero-cost pedagogy (Definition 3), a â and R3âs limiting rational response jointly attain the oracle value from state s: E a R âźĎ R3â (¡|s,b R3+ â ,a â ) Q O (s,a â ,a R ;θ â ) =max a R âA R (s) Q O (s,a â ,a R ;θ â ) = V O (s;θ â ). 14E. Lazarski et al. Since H2âs limiting rational policy is supported exclusively onL H2 s,b (θ â ), averaging over human actions a â âź Ď H2â (¡| s,b;θ â ) also yields V O (s;θ â ). We emphasize that, upon observing any leveraging human action, R3âs un- certainty really does collapse in a single time step, after which the strategies for the rest of the time horizon correspond to the oracle team policies and realize V O (s;θ â ) in full. Since θ â was arbitrary, the limiting R3âH2 policies attain the oracle value under every possible goal. No feasible strategy pair can exceed the fully informed oracle value under any goal, so (Ď R3â ,Ď H2â ) is optimal. Furthermore, assistance games are common-payoff games, so any unilateral policy deviation yields another feasible strategy pair evaluated under the same team objective. Holding Ď R3â fixed, no human deviation can yield value greater than the globally optimal value already attained by (Ď R3â ,Ď H2â ). Hence, Ď H2â is a best response to Ď R3â . The same argument applies to any unilateral robot deviation, so Ď R3â is a best response to Ď H2â . Therefore, the limiting rational (β H ,β R ââ) R3âH2 policies constitute an optimal pragmaticâpedagogic equilibrium of the assistance game at (s,b). ââ Running example: From state s = and any prior b()â (0, 1), recall that the pedagogic leverage mapL H2 s,b is fully populated as β H ,β R ââ:L H2 s,b ( ) =â and L H2 s,b () = â. By Proposition 1, b R3+ â (|â) = 1 and b R3+ â (|â) = 1. Evaluating R3âs limiting rational policy gives E a R âźĎ R3â (¡|s,b R3+ â ,â) Q O (s,â,a R ; ) = 2 = V O (s;), E a R âźĎ R3â (¡|s,b R3+ â ,â) Q O (s,â,a R ;) = 2 = V O (s;). These calculations confirm that L H2 s,b achieves zero-cost pedagogy (Definition 3). Consequently, the Tetris assistance game is action-separable at (s,b) (Defini- tion 4), so the limiting rational (β H ,β R â â) R3âH2 policies constitute an optimal pragmaticâpedagogic equilibrium (Theorem 1). 6 Algorithm PragmaticâPedagogic Best Response (PPBR) provides a simple procedure for computing an optimal equilibrium solution to an action-separable assistance game (Definition 4). Algorithm 1 checks whether the assistance game is action- separable at (s,b) and, if so, returns the pedagogic leverage mapL (Definition 2), assigning to each goal the corresponding set of leveraging human actions. Since the sets L(θ) θâÎ are disjoint, each leveraging action identifies a unique goal, inducing an inverse leverage map L â1 : A L â Î, where A L is the union of all leverage sets L(θ). The limiting H2 policy, R3 posterior belief, Corrigible Assistance in One Round: PragmaticâPedagogic Best Response15 Algorithm 1: PragmaticâPedagogic Best Response Input: state s, prior b, finite β R > 0 (hyperparameter) Output: action-separable (Boolean), pedagogic leverage map L (Definition 2) 1 Compute the limiting rational solipsistic human policy Ď H0â at (s,b) as in (2); 2 foreach candidate human action a H do 3Compute the limiting literal posterior b R1+ â (¡| a H ); 4Compute the noisy robot response policy Ď R1 as in (3); 5 Compute the induced pedagogic human values Q H2 from Ď R1 ; 6 foreach goal θ â Î do 7 L(θ)â arg max a H âA H (s) Q H2 (s,b,a H ;θ); 8 if âa H âL(θ) already assigned to another goal then 9return (False, â ); 10 if âa H âL(θ) such that max a R âA R (s) Q O (s,a H ,a R ;θ) < V O (s;θ) then 11return (False, â ); 12 return (True, L); and R3 policy are then readily obtained as Ď H2â (a H | s,b;θ) = 1 a H âL(θ) |L(θ)| , b R3+ â (θ | a H ) = 1 θ =L â1 (a H ) , Ď R3â (a R | s,b R3+ â ,a H ) = Ď R,O a R | s,a H ; L â1 (a H ) , where Ď R,O is the oracle robot policy, that is, Ď R,O (a R | s,a H ;θ)â 1 a R â arg max aâA R (s) Q O (s,a H ,a;θ) . The runtime of Algorithm 1 is linear in the size of the goal set and action sets: O(|Î||A H ||A R |), which is as efficient as solving the oracle MDP. While our procedure directly computes the limiting behavior as β H ,β R ââ in our Tetris running example, plugging in finite values of β H (with β R = 1) generates the curves depicted in Figure 3. 7 Limitations and Future Work A direct extension is split leverage, in which a single action is H2-optimal for more than one goal, giving only partial disambiguation. Analyzing the resulting class of assistance games and lifting Algorithm 1 to a multi-step procedure is beyond the scope of this paper. More broadly, we have shown that nested reasoning with Boltzmann-rational models naturally breaks the ambiguity between otherwise task-equivalent actions. We note that explicit level-k iteration may not be the only algorithmic route to the pragmaticâpedagogic solution; investigating other approaches is left to future work. 16E. Lazarski et al. Fig. 3: One-step posterior b + (|â) vs. prior b() at s =with β R = 1. IOC (R1, orange) saturates at the ceiling b R1+ â = 2b/(1 + b) as β H increases. PPBR (R3, green) instead breaks the ceiling and approaches b R3+ â = 1. 8 Conclusion In this paper, we identified a class of assistance games in which the human and the robot arrive at an unambiguous correspondence between goals and actions at zero cost to the human, who can convey and achieve any goal without sacrific- ing task value. For these action-separable games, pragmaticâpedagogic reasoning resolves goal uncertainty in a single time step, breaking the IOC inference ceil- ing that leads to incorrigible robot behavior. We proved that one best-response round starting from a mainstream IOC solution suffices to reach an optimal equi- librium of the full-horizon game. PragmaticâPedagogic Best Response computes this equilibrium as efficiently as solving the oracle MDP, yielding a corrigible robot assistant that maximally empowers the human user. Ultimately, robust alignment hinges on how robotic and AI systems represent and interpret the humans with whom they interact. Acknowledgments. We extend a special thank you to Donggeon Oh and Tom Silver for insightful discussions and feedback. Disclosure of Interests. The authors have no competing interests to declare that are relevant to the content of this article. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response17 References [1] P. F. Christiano, J. Leike, T. Brown, et al. âDeep Reinforcement Learn- ing from Human Preferencesâ. Advances in Neural Information Processing Systems. Ed. by I. Guyon, U. V. Luxburg, S. Bengio, et al. Vol. 30. Curran Associates, Inc., 2017. [2] S. Casper, X. Davies, C. Shi, et al. âOpen Problems and Fundamental Lim- itations of Reinforcement Learning from Human Feedbackâ. Transactions on Machine Learning Research (2023). [3] L. Lang, D. Foote, S. Russell, et al. âWhen Your AIs Deceive You: Chal- lenges of Partial Observability in Reinforcement Learning from Human Feedbackâ. Advances in Neural Information Processing Systems. Ed. by A. Globerson, L. Mackey, D. Belgrave, et al. Vol. 37. Curran Associates, Inc., 2024, p. 93240â93299. [4] M. Williams, M. Carroll, A. Narang, et al. âOn Targeted Manipulation and Deception when Optimizing LLMs for User Feedbackâ. International Conference on Learning Representations (ICLR). 2025. [5] A. Bestick, R. Bajcsy, and A. D. Dragan. âImplicitly Assisting Humans to Choose Good Grasps in Robot to Human Handoversâ. International Symposium on Experimental Robotics. Vol. 1. Springer Proceedings in Ad- vanced Robotics. Cham: Springer, 2017, p. 341â354. [6] C. Liu, J. B. Hamrick, J. F. Fisac, et al. âGoal Inference Improves Objective and Perceived Performance in Human-Robot Collaborationâ. International Conference on Autonomous Agents & Multiagent Systems. AAMAS â16. International Foundation for Autonomous Agents and Multiagent Systems, 2016, p. 940â948. [7] A. Bobu, A. Bajcsy, J. F. Fisac, and A. D. Dragan. âLearning under Mis- specified Objective Spacesâ. The 2nd Conference on Robot Learning. Ed. by A. Billard, A. Dragan, J. Peters, and J. Morimoto. Vol. 87. Proceedings of Machine Learning Research. PMLR, 2018, p. 796â805. [8] A. Bobu, A. Peng, P. Agrawal, et al. âAligning Human and Robot Rep- resentationsâ. ACM/IEEE International Conference on Human-Robot In- teraction. HRI â24. Association for Computing Machinery, 2024, p. 42â 54. [9] A. Fern, S. Natarajan, K. Judah, and P. Tadepalli. âA Decision-Theoretic Model of Assistanceâ. Journal of Artificial Intelligence Research 50 (2014), p. 71â104. [10] D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan. âCooperative Inverse Reinforcement Learningâ. Advances in Neural Information Pro- cessing Systems. Ed. by D. Lee, M. Sugiyama, U. Luxburg, et al. Vol. 29. Curran Associates, Inc., 2016. [11] J. F. Fisac, M. A. Gates, J. B. Hamrick, et al. âPragmaticâPedagogic Value Alignmentâ. International Symposium on Robotics Research (ISRR 2017). Vol. 10. Springer International Publishing, 2017, p. 49â57. 18E. Lazarski et al. [12] D. Malik, M. Palaniappan, J. Fisac, et al. âAn Efficient, Generalized Bell- man Update for Cooperative Inverse Reinforcement Learningâ. Interna- tional Conference on Machine Learning. PMLR, 2018, p. 3394â3402. [13] P. Shafto, N. D. Goodman, and T. L. Griffiths. âA Rational Account of Pedagogical Reasoning: Teaching by, and Learning from, Examplesâ. Cog- nitive Psychology 71 (2014), p. 55â89. [14] C. Laidlaw, E. Bronstein, T. Guo, et al. âAssistanceZero: Scalably Solving Assistance Gamesâ. International Conference on Machine Learning. Ed. by A. Singh, M. Fazel, D. Hsu, et al. Vol. 267. Proceedings of Machine Learning Research. PMLR, 2025, p. 32278â32305. [15] M. L. Littman, A. R. Cassandra, and L. P. Kaelbling. âLearning Policies for Partially Observable Environments: Scaling Upâ. Machine Learning Proceedings 1995. Ed. by A. Prieditis and S. Russell. San Francisco (CA): Morgan Kaufmann, 1995, p. 362â370. [16] N. Wiener. âSome Moral and Technical Consequences of Automationâ. Sci- ence 131.3410 (1960), p. 1355â1358. [17] A. Y. Ng and S. J. Russell. âAlgorithms for Inverse Reinforcement Learn- ingâ. Seventeenth International Conference on Machine Learning. ICML â00. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2000, p. 663â670. [18] D. Ramachandran and E. Amir. âBayesian Inverse Reinforcement Learn- ingâ. International Joint Conference on Artifical Intelligence. IJCAIâ07. Morgan Kaufmann Publishers Inc., 2007, p. 2586â2591. [19] B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey. âMaximum Entropy Inverse Reinforcement Learningâ. AAAI Conference on Artificial Intelli- gence. AAAI, 2008, p. 6. [20] D. O. Stahl and P. W. Wilson. âOn Playersâ Models of Other Players: Theory and Experimental Evidenceâ. Games and Economic Behavior 10.1 (1995), p. 218â254. [21] M. A. Costa-Gomes, V. P. Crawford, and B. Broseta. âCognition and Be- havior in Normal-Form Games: An Experimental Studyâ. Econometrica 69.5 (2001), p. 1193â1235. [22] C. F. Camerer, T.-H. Ho, and J.-K. Chong. âA Cognitive Hierarchy Model of Gamesâ. The Quarterly Journal of Economics 119.3 (2004), p. 861â 898. [23] S. Milli and A. D. Dragan. âLiteral or Pedagogic Human? Analyzing Hu- man Model Misspecification in Objective Learningâ. The 35th Uncertainty in Artificial Intelligence Conference. Ed. by R. P. Adams and V. Gogate. Vol. 115. Proceedings of Machine Learning Research. PMLR, 2020, p. 925â 934. [24] S. Milli, D. Hadfield-Menell, A. Dragan, and S. Russell. âShould Robots be Obedient?â Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI-17. 2017, p. 4754â4760. Corrigible Assistance in One Round: PragmaticâPedagogic Best Response19 [25] D. Hadfield-Menell, A. D. Dragan, P. Abbeel, and S. Russell. âThe Off- Switch Gameâ. Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17). 2017, p. 220â227. [26] A. Bobu, D. R. R. Scobee, J. F. Fisac, et al. âLESS is More: Rethinking Probabilistic Models of Human Behaviorâ. ACM/IEEE International Con- ference on Human-Robot Interaction. HRI â20. Association for Computing Machinery, 2020, p. 429â437. [27] T. C. Schelling. âBargaining, Communication, and Limited Warâ. Conflict Resolution 1.1 (1957), p. 19â36. [28] T. C. Schelling. The Strategy of Conflict. Cambridge, MA: Harvard Uni- versity Press, 1980.