Paper deep dive
A Revealed Preference Framework for AI Alignment
Elchin Suleymanov
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/31/2026, 2:09:13 AM
Summary
The paper introduces the Luce Alignment Model (LAM) to analyze AI alignment and compliance using revealed preference techniques. It models AI choices as a mixture of human and AI-intrinsic Luce rules, providing identification results for both laboratory (where human choices are observed) and field (where only AI choices are observed) settings.
Entities (4)
Relation Signals (3)
Elchin Suleymanov → authored → A Revealed Preference Framework for AI Alignment
confidence 100% · A Revealed Preference Framework for AI Alignment Elchin Suleymanov
Luce Alignment Model → uses → Luce rule
confidence 95% · the AI's choices are a mixture of two Luce rules
AI Alignment → analyzedvia → Luce Alignment Model
confidence 90% · The main goal of this paper is to provide a modeling framework that can be used to analyze the alignment of AI agents
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal's preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI's choices are a mixture of two Luce rules, one reflecting the human's preferences and the other the AI's. I show that the AI's alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed.
Tags
Links
- Source: https://arxiv.org/abs/2603.27868v1
- Canonical: https://arxiv.org/abs/2603.27868v1
Trouble viewing inline? Open PDF directly →
Full Text
76,700 characters extracted from source content.
Expand or collapse full text
A Revealed Preference Framework for AI Alignment Elchin Suleymanov Department of Economics, Mitch Daniels School of Business, Purdue University, West Lafayette, Indiana, USA. Email: esuleyma@purdue.edu . Abstract Human decision makers increasingly delegate choices to AI agents, raising a natural question: does the AI implement the human principal’s preferences or pursue its own? To study this question using revealed preference techniques, I introduce the Luce Alignment Model, where the AI’s choices are a mixture of two Luce rules, one reflecting the human’s preferences and the other the AI’s. I show that the AI’s alignment (similarity of human and AI preferences) can be generically identified in two settings: the laboratory setting, where both human and AI choices are observed, and the field setting, where only AI choices are observed. JEL Classification: D01, D11, D83 Keywords: AI alignment, stochastic choice, revealed preference, Luce model 1 Introduction Artificial intelligence (AI) agents are likely to play an increasing role in making choices on behalf of human users and in reshaping everyday decision-making (Allouah et al., 2025; Immorlica et al., 2024). Historically, many AI systems played the role of a technology that assisted rather than replaced human choice. For example, recommender systems filter, rank, and personalize the set of alternatives presented to a user rather than making a selection on the user’s behalf (Adomavicius and Tuzhilin, 2005). Recent advances in agentic AI, together with improved memory and personalization abilities, have made fuller delegation of choice increasingly feasible, allowing human decision makers to rely on AI agents not only to screen among available options but also to make the final choice. This makes it natural to model AI agents not just as a technology available to human decision makers, but as economic agents in their own right (Immorlica et al., 2024; Chen et al., 2024). This shift from assisted choice to delegated choice raises a natural economic question: when an AI agent chooses on behalf of a human principal, whose preferences does it implement? Does it act fully in accordance with the principal’s preferences, or does it instead pursue distinct objectives that diverge from them? This question lies at the core of the AI alignment literature, which broadly seeks to ensure that AI systems behave in line with human users’ intentions (Leike et al., 2018; Ji et al., 2024). Much of the existing alignment literature is motivated by catastrophic risks, harmful behaviors, and the loss of control over increasingly capable AI systems (Amodei et al., 2016; Hendrycks et al., 2023; Ji et al., 2024). The main focus of this paper is what can be considered a narrower but economically central question: whether a personalized AI agent making choices on behalf of a human principal is in fact implementing the principal’s preferences. AI agents may perform well in safety evaluations designed to detect catastrophic risks or harmful behaviors but still make misaligned choices in delegated choice environments. The AI alignment literature has largely approached the problem through three main channels. First, there are methods that attempt to align an AI system during its training phase. These include cooperative inverse reinforcement learning (Hadfield-Menell et al., 2016), reinforcement learning from human feedback (Christiano et al., 2017), scalable reward-modeling approaches more generally (Leike et al., 2018), and constitutional AI (Bai et al., 2022). Second, there are ex-post evaluation methods that seek evidence of misalignment through benchmarks and behavioral tests (Perez et al., 2023; Ji et al., 2024). Third, there are interpretability-based approaches that attempt to understand the model’s internal objectives and reasoning processes by analyzing the inner structures of neural networks (Räuker et al., 2023). This paper adopts a complementary approach grounded in revealed preference analysis. Applied to human agents, the revealed preference approach attempts to infer preferences from observed choices rather than from the processes inside the human brain. Treating AI agents as economic agents, we can extend the same approach to choices made by an AI. In particular, the goal is to infer the extent of AI misalignment by analyzing the stochastic choice data it generates while acting on behalf of its principal. The setup in this paper is as follows. An analyst observes the stochastic choice data ρAIρ^AI generated by an AI agent. That is, the AI faces varying menus S repeatedly and makes choices from them on behalf of some human user. I consider two natural settings. In the laboratory setting, the analyst also observes the human principal’s stochastic choices ρHρ^H. For example, ρHρ^H may be elicited directly from the human principal or generated synthetically for the purpose of guiding the AI. In the field setting, only ρAIρ^AI is observed. In both settings, the aim of the analyst is to recover two key objects of interest: the degree to which the AI’s intrinsic preferences match the human principal’s preferences (alignment) and the extent to which the AI defers to the human principal (compliance). These two are distinct concepts. For example, a misaligned AI that is highly compliant may still generate choice data that closely matches the human principal’s. Alternatively, a perfectly aligned AI will replicate the human principal’s choices regardless of its compliance level. To formalize this distinction, I introduce the Luce Alignment Model (LAM), where the AI’s choices from a menu S are modeled as ρAI(x,S)=α⋅u(x)∑y∈Su(y)+(1−α)⋅v(x)∑y∈Sv(y).ρ^AI(x,S)=α· u(x) _y∈ Su(y)+(1-α)· v(x) _y∈ Sv(y). (1) Here, u represents the human principal’s utility, v represents the AI’s intrinsic utility, and α∈[0,1]α∈[0,1] is the compliance parameter that captures the extent to which the AI defers to the human principal. The comparison of u and v in turn captures the AI’s alignment with the human principal. The first term in equation (1) can be interpreted as the human principal’s stochastic choice rule ρHρ^H, and the second term, denoted ρAρ^A, as the AI agent’s autonomous stochastic choice rule. The model can then equivalently be written as ρAI(x,S)=α⋅ρH(x,S)+(1−α)⋅ρA(x,S).ρ^AI(x,S)=α·ρ^H(x,S)+(1-α)·ρ^A(x,S). In both settings, the goal of the analyst is to infer α, u, and v from observed stochastic choice data. The main results of the paper address the identification of alignment and compliance in both settings. In the laboratory setting, where both ρAIρ^AI and ρHρ^H are observed, I show that all the parameters of interest can be identified as long as the AI’s choice data violates the Independence of Irrelevant Alternatives (IIA) property. The IIA property, which is the key implication of the Luce model, requires that the relative choice probabilities of any two alternatives are constant across menus (Luce, 1959). When the AI is perfectly compliant or perfectly aligned, its choices satisfy IIA. By contrast, IIA violations reveal the presence of an intrinsic AI utility that is distinct from the human principal’s utility. I introduce instability measures that capture deviations from the IIA property and provide a closed-form expression for the compliance parameter α using these measures. Using the recovered compliance parameter, I then show that both the human principal’s utility u and the AI’s utility v can be identified up to scale normalization. I also provide an axiomatic characterization of the model in this setting, identifying the behavioral conditions that the pair (ρAI,ρH)(ρ^AI,ρ^H) must satisfy in order to be consistent with LAM. In the field setting, where only ρAIρ^AI is observed, there is a fundamental obstacle to separately identifying the human’s and the AI’s utilities. Namely, both (u,v,α)(u,v,α) and (v,u,1−α)(v,u,1-α) generate the same stochastic choice data. Thus, identification is only possible up to a label swap. Nevertheless, I provide a constructive proof showing that when there are at least four alternatives, the underlying utility pair is generically identified up to this swap. Stated alternatively, the distribution over utilities is generically identified but the labels are not. From an analyst’s perspective, this is sufficient to recover the degree of misalignment, even without knowing which utility belongs to the human and which belongs to the AI. In terms of compliance, the result implies that the analyst cannot distinguish α from 1−α1-α. Hence, compliance is only identifiable up to reflection about 1/21/2 in the field setting unless one is willing to assume α<1/2α<1/2 (low compliance) or α>1/2α>1/2 (high compliance). LAM draws on a long tradition in stochastic choice theory. When α∈0,1α∈\0,1\, the model reduces to the Luce rule, one of the foundational stochastic choice models. When α∈(0,1)α∈(0,1), the model becomes a mixed multinomial logit (MMNL, also known as the random coefficients multinomial logit) model with binary support, or simply 2-MNL. More generally, the MMNL model was introduced by Boyd and Mellman (1980) and Cardell and Dunbar (1980). McFadden and Train (2000) show that any choice model derived from random utility maximization can be approximated by an MMNL model, while Saito (2018) provides axiomatic foundations (see also Lu and Saito, 2022; Chang et al., 2023). Fox et al. (2012) show that the distribution over random coefficients in MMNL is uniquely identified under sufficiently rich variation in product characteristics. In contrast, in an abstract domain with menu variation, the identification problem in the mixed logit model has mostly been studied within the statistics and computer science literatures. Chierichetti et al. (2018) study the identifiability problem in 2-MNL assuming a uniform mixing weight. Tang (2020) allows for an arbitrary mixing weight and provides a generic identification result. Zhang et al. (2022) extend this result by showing that observing menus with three alternatives is sufficient for generic identification. The result in the field setting of this paper provides an alternative approach to the identification problem in 2-MNL: unlike Zhang et al. (2022), the identification is constructive, and unlike Chierichetti et al. (2018) and Tang (2020), the constructive procedure does not require any prior knowledge of the mixing weight. Within decision theory literature, the closest precedent to LAM is Chambers et al. (2023), who study a model of behavioral peer influence. In their model, there are two agents and each agent’s stochastic choice can be written as a mixture of the agent’s own and the other agent’s Luce rule. Importantly, the mixing weights in Chambers et al. (2023) depend on the underlying utilities of the two agents and vary across menus, which makes their setup more suitable to study influence rather than AI alignment and compliance. Chambers et al. (2023) also mention a modification of their model with menu-independent weights as a potential alternative but note that unique identification in this alternative version is not guaranteed. Lastly, Manzini and Mariotti (2018) study dual random utility maximization (dRUM), where the agent maximizes one of two deterministic linear orders with a fixed probability. dRUM can be viewed as a limiting case of LAM where both ρHρ^H and ρAρ^A are generated by deterministic utility maximization. The rest of the paper proceeds as follows. Section 2 introduces the model. Section 3 presents identification and characterization results in the laboratory setting. Section 4 presents identification results in the field setting. Section 5 concludes. 2 The Model Let X be a finite set of alternatives and denote by X the collection of all non-empty subsets of X (menus). It is assumed that |X|=N≥3|X|=N≥ 3. A stochastic choice function is a mapping ρ:X×→[0,1]ρ:X×X→[0,1] such that ρ(x,S)>0ρ(x,S)>0 only if x∈Sx∈ S and ∑x∈Sρ(x,S)=1 _x∈ Sρ(x,S)=1 for all S∈S . ρ(x,S)ρ(x,S) denotes the probability that x is chosen when the agent is faced with the menu S repeatedly. The setup is as follows. The analyst observes the stochastic choices of an AI agent, denoted ρAIρ^AI, that acts on behalf of a human principal. The stochastic choices of the human principal, denoted ρHρ^H, may or may not be observed. I consider two empirical settings. In the laboratory setting, both ρAIρ^AI and ρHρ^H are observed. This corresponds to an experimental design where the human’s and the AI’s choices are elicited sequentially, with the human’s choices serving as a guide for the AI. Alternatively, the human’s choices might be generated synthetically for the purposes of the experiment. In the field setting, only the AI’s choices are observed. This corresponds to the scenario where the human principal fully delegates the decision-making process to the AI. The main goal of this paper is to provide a modeling framework that can be used to analyze the alignment of AI agents with their human principals’ preferences. To this end, let ρAρ^A denote the hypothetical choices of an AI agent that acts autonomously without any human principal. I assume that in this autonomous setting, the AI’s choices follow the Luce rule with some underlying utility function v:X→ℝ++v:X _++: ρA(x,S)=v(x)∑y∈Sv(y).ρ^A(x,S)= v(x) _y∈ Sv(y). The function v can be interpreted as the intrinsic utility of the AI agent. Importantly, ρAρ^A is not observed as the AI is always assumed to act on behalf of some human principal. Let u:X→ℝ++u:X _++ denote the utility function of the human principal. The principal’s choices are also assumed to be consistent with the Luce rule given by ρH(x,S)=u(x)∑y∈Su(y).ρ^H(x,S)= u(x) _y∈ Su(y). The AI’s compliance parameter α∈[0,1]α∈[0,1] reflects the probability that it ignores its own utility and follows the human principal. With probability 1−α1-α, it ignores the principal and acts autonomously. The AI’s observed stochastic choices are therefore given by ρAI(x,S)=α⋅ρH(x,S)+(1−α)⋅ρA(x,S).ρ^AI(x,S)=α·ρ^H(x,S)+(1-α)·ρ^A(x,S). There are two key objects the analyst would like to identify: (i) alignment - to what extent the utilities u and v are aligned, (i) compliance - the value of α. I will discuss how each of these can be identified both in the laboratory and the field settings. The following definition summarizes the model. Definition 1 (Luce Alignment Model). A pair of stochastic choice functions (ρAI,ρH)(ρ^AI,ρ^H) is consistent with the Luce Alignment Model (LAM) if there exist utility functions u,v:X→ℝ++u,v:X _++ and a compliance parameter α∈[0,1]α∈[0,1] such that for all S∈S and x∈Sx∈ S, ρAI(x,S)=αu(x)∑y∈Su(y)+(1−α)v(x)∑y∈Sv(y)andρH(x,S)=u(x)∑y∈Su(y).ρ^AI(x,S)=α u(x) _y∈ Su(y)+(1-α) v(x) _y∈ Sv(y) ρ^H(x,S)= u(x) _y∈ Su(y). (2) The tuple (u,v,α)(u,v,α) is called a LAM representation of (ρAI,ρH)(ρ^AI,ρ^H). If only ρAIρ^AI is observed, then it is said to be consistent with LAM if it satisfies the expression in equation (2) for some u,vu,v, and α. LAM nests several important special cases in terms of observed AI behavior: • Aligned AI (v=λuv=λ u for some λ>0λ>0). The AI agent’s preferences exactly match the human principal. In this case, the degree of compliance becomes unimportant. The AI makes exactly the same probabilistic choices as the human. • Compliant AI (α=1α=1). The AI agent’s preferences potentially differ from the human principal’s. However, since compliance is perfect, the AI makes the exact same probabilistic choices the human would make. • Misaligned AI (v≠λuv≠λ u for all λ>0λ>0). The AI agent’s preferences differ from the human principal’s and it may not be fully compliant. This is the base case where uncovering the extent of misalignment and non-compliance becomes important. • Autonomous AI (α=0α=0). The AI potentially has its own distinct preferences and operates fully autonomously ignoring its human principal. • Adversarial AI (v=λu−1v=λ u^-1 for some λ>0λ>0). The AI agent’s preferences are exactly the opposite of the human’s, with the top ordinally ranked alternative by the human being ranked as the worst by the AI. While observed choices are exactly the same in the first two cases where the AI is either perfectly aligned or perfectly compliant, they represent the boundary cases of scenarios with completely different implications. In the first case where we start with perfect alignment, a small change in alignment will likely not have a big influence on the human principal’s welfare regardless of the compliance level. In the second case where we start with perfect compliance, if the degree of misalignment is very high, a small decrease in compliance levels can have large welfare effects. This makes it important to uncover alignment and compliance separately in the base third case with misaligned AI. The fourth case models a fully autonomous AI that ignores its human principal, while the last case represents a version of extreme misalignment that provides a further structure to the model and serves as a natural adversarial benchmark. 3 Laboratory Data This section analyzes the laboratory setting, where both the AI’s choices ρAIρ^AI and the human principal’s choices ρHρ^H are observable. First, I discuss the identification of the AI’s alignment and compliance. I then provide an axiomatic characterization of the model using the pair (ρAI,ρH)(ρ^AI,ρ^H) as the observed primitive. 3.1 Identification Consider a pair (ρAI,ρH)(ρ^AI,ρ^H) consistent with LAM. Under LAM, the human principal’s choices are consistent with the Luce rule.111Since the laboratory setting can utilize synthetically generated choice data from a hypothetical human principal, assuming these choices follow the Luce rule is not a substantive restriction. In the field setting, on the other hand, the human principal’s choices are unobserved and therefore not explicitly modeled. The implicit modeling assumption in this setting is that the AI perceives its human principal as a Luce agent with utility u. A stochastic choice rule ρ consistent with the Luce rule satisfies the Independence of Irrelevant Alternatives (IIA) property of Luce (1959): ρ(x,S)ρ(y,S)=ρ(x,T)ρ(y,T) for all x,y∈S∩T and S,T∈. ρ(x,S)ρ(y,S)= ρ(x,T)ρ(y,T) for all x,y∈ S∩ T and S,T . As the next proposition shows, the AI’s choices will generally violate IIA unless u and v are perfectly aligned (v=λuv=λ u for some λ>0λ>0), or the AI is either autonomous (α=0α=0) or perfectly compliant (α=1α=1). Proposition 1 (IIA Violation). Let ρAIρ^AI be consistent with LAM. Then ρAIρ^AI satisfies IIA if and only if α∈0,1α∈\0,1\ or v=λuv=λ u for some λ>0λ>0. Proof. If α=0α=0 or α=1α=1, then ρAIρ^AI reduces to the Luce rule with the utility functions v or u, respectively, which satisfies IIA. Alternatively, if v=λuv=λ u for some λ>0λ>0, then ρA=ρHρ^A=ρ^H, and hence ρAI=ρHρ^AI=ρ^H, which satisfies IIA. Conversely, suppose α∈(0,1)α∈(0,1) and v≠λuv≠λ u for all λ>0λ>0. Define r(a)=u(a)/v(a)r(a)=u(a)/v(a) for each a∈Xa∈ X. Since v≠λuv≠λ u, the function r is not constant. Pick x,y∈Xx,y∈ X with r(x)≠r(y)r(x)≠ r(y) and any z∉x,yz∉\x,y\. We need to show that IIA is violated for some tuple (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T. To this end, first, observe that the conditional probability of a relative to b in the choice set a,b,c\a,b,c\ can be written as a mixture of the choice probabilities of the autonomous AI and the human principal in the choice set a,b\a,b\, where the mixing coefficient depends on c. That is, for any a,b,c∈Xa,b,c∈ X, ρAI(a,a,b,c)ρAI(a,a,b,c)+ρAI(b,a,b,c)=β(c)⋅u(a)u(a)+u(b)+(1−β(c))⋅v(a)v(a)+v(b), ρ^AI(a,\a,b,c\)ρ^AI(a,\a,b,c\)+ρ^AI(b,\a,b,c\)=β(c)· u(a)u(a)+u(b)+(1-β(c))· v(a)v(a)+v(b), where β(c)=α⋅u(a)+u(b)u(a)+u(b)+u(c)α⋅u(a)+u(b)u(a)+u(b)+u(c)+(1−α)⋅v(a)+v(b)v(a)+v(b)+v(c). β(c)= α· u(a)+u(b)u(a)+u(b)+u(c)α· u(a)+u(b)u(a)+u(b)+u(c)+(1-α)· v(a)+v(b)v(a)+v(b)+v(c). By definition, ρAI(a,a,b)ρAI(a,a,b)+ρAI(b,a,b)=α⋅u(a)u(a)+u(b)+(1−α)⋅v(a)v(a)+v(b). ρ^AI(a,\a,b\)ρ^AI(a,\a,b\)+ρ^AI(b,\a,b\)=α· u(a)u(a)+u(b)+(1-α)· v(a)v(a)+v(b). Comparing this to the conditional probability in a,b,c\a,b,c\, notice that both are mixtures of the exact same two terms. Hence, when c is removed from a,b,c\a,b,c\, an IIA violation occurs if and only if both the mixing weights and the mixture components in these conditional probabilities are distinct: that is, β(c)≠αβ(c)≠α and u(a)/u(b)≠v(a)/v(b)u(a)/u(b)≠ v(a)/v(b) (or, alternatively, r(a)≠r(b)r(a)≠ r(b)). Notice that β(c)=αβ(c)=α if and only if (u(a)+u(b))/(v(a)+v(b))=(u(a)+u(b)+u(c))/(v(a)+v(b)+v(c))(u(a)+u(b))/(v(a)+v(b))=(u(a)+u(b)+u(c))/(v(a)+v(b)+v(c)), which holds if and only if r(c)=u(c)/v(c)=(u(a)+u(b))/(v(a)+v(b))r(c)=u(c)/v(c)=(u(a)+u(b))/(v(a)+v(b)). Therefore, when c is removed from a,b,c\a,b,c\, an IIA violation occurs if and only if r(a)≠r(b)r(a)≠ r(b) and r(c)≠(u(a)+u(b))/(v(a)+v(b))r(c)≠(u(a)+u(b))/(v(a)+v(b)). Now, going back to the choice set x,y,z\x,y,z\, there are two cases to consider. Case 1: r(z)≠u(x)+u(y)v(x)+v(y)r(z)≠ u(x)+u(y)v(x)+v(y). Since r(x)≠r(y)r(x)≠ r(y), by the previous argument, IIA is violated when z is removed from x,y,z\x,y,z\. Case 2: r(z)=u(x)+u(y)v(x)+v(y)=v(x)v(x)+v(y)r(x)+v(y)v(x)+v(y)r(y)r(z)= u(x)+u(y)v(x)+v(y)= v(x)v(x)+v(y)r(x)+ v(y)v(x)+v(y)r(y). Since r(x)≠r(y)r(x)≠ r(y) and r(z)r(z) is a strict weighted average of r(x)r(x) and r(y)r(y), it must lie strictly between r(x)r(x) and r(y)r(y). This implies r(x)≠r(z)r(x)≠ r(z). Now suppose y is removed from the choice set x,y,z\x,y,z\. Since r(x)≠r(z)r(x)≠ r(z), by the previous argument, IIA holds if and only if r(y)=u(x)+u(z)v(x)+v(z)r(y)= u(x)+u(z)v(x)+v(z). But then r(y)r(y) is a strict mixture of r(x)r(x) and r(z)r(z). This is clearly not possible as r(z)r(z) itself is a strict mixture of r(x)r(x) and r(y)r(y). Hence, IIA must be violated when y is removed from x,y,z\x,y,z\. To conclude, either the removal of z or the removal of y (or x) from x,y,z\x,y,z\ leads to an IIA violation, as desired. ∎ The implication of the proposition is that IIA violations in the AI’s stochastic choices indicate we are in the case of misaligned (v≠λuv≠λ u for all λ>0λ>0) and partially compliant (α∈(0,1)α∈(0,1)) AI. We can utilize this to recover the parameters of the model. The identification strategy proceeds in three steps. Step 1: Recover u from ρHρ^H. Since the human principal’s choices are consistent with the Luce rule, the identification of u from ρHρ^H is standard. Letting u(x)=1u(x)=1 for some x∈Xx∈ X, the IIA property implies that for any y∈Xy∈ X, we must have u(y)=ρH(y,S)ρH(x,S),u(y)= ρ^H(y,S)ρ^H(x,S), where S can be any choice set containing x and y. This recovers u up to scale normalization. Step 2: Recover α from ρAIρ^AI and ρHρ^H. If ρAI=ρHρ^AI=ρ^H, then the AI may be either fully compliant (α=1α=1) or fully aligned (v=λuv=λ u for λ>0λ>0). We cannot distinguish between these two cases. Alternatively, if ρAI≠ρHρ^AI≠ρ^H and ρAIρ^AI exhibits no IIA violations, then we can use Proposition 1 to infer that α=0α=0. This is because we cannot have α=1α=1 or v=λuv=λ u with ρAI≠ρHρ^AI≠ρ^H, which leaves α=0α=0 as the only possibility in the proposition. Now suppose ρAIρ^AI violates IIA. Then, there exist two choice sets S,T∈S,T and a pair of alternatives x,y∈S∩Tx,y∈ S∩ T such that ρAI(x,S)ρAI(y,S)≠ρAI(x,T)ρAI(y,T). ρ^AI(x,S)ρ^AI(y,S)≠ ρ^AI(x,T)ρ^AI(y,T). The identification of α relies on these IIA violations. To proceed with the identification, we first need a new definition. Definition 2 (Instability Measures). Let ρ,ρ′ρ,ρ be two stochastic choice functions. For any S,T∈S,T and x,y∈S∩Tx,y∈ S∩ T: 1. The own instability of ρ is defined by Δxy(S,T|ρ)=ρ(x,S)ρ(y,T)−ρ(y,S)ρ(x,T). _xy(S,T|ρ)=ρ(x,S)ρ(y,T)-ρ(y,S)ρ(x,T). 2. The cross instability from ρ to ρ′ρ is defined by Γxy(S,T|ρ,ρ′)=ρ(x,S)ρ′(y,T)−ρ(y,S)ρ′(x,T). _xy(S,T|ρ,ρ )=ρ(x,S)ρ (y,T)-ρ(y,S)ρ (x,T). 3. The composite instability of ρ and ρ′ρ is defined by Φxy(S,T|ρ,ρ′)=Γxy(S,T|ρ,ρ′)+Γxy(S,T|ρ′,ρ). _xy(S,T|ρ,ρ )= _xy(S,T|ρ,ρ )+ _xy(S,T|ρ ,ρ). Intuitively, own instability can be viewed as a measure of instability in the stochastic choice ρ for the tuple (x,y,S,T)(x,y,S,T). It tells us how useful the observations from the choice set S are for imputing the relative choice probabilities for x,yx,y in the choice set T. If Δxy(S,T|ρ)=0 _xy(S,T|ρ)=0, then there is no IIA violation for the alternatives x,yx,y in the choice sets S and T, and this imputation can be done perfectly. The larger |Δxy(S,T|ρ)|| _xy(S,T|ρ)|, the less useful the observations from S are for imputing choices in T. The measure of cross instability tells us how useful the information from the stochastic choice ρ in the choice set S is for imputing the relative choice probabilities for ρ′ρ in the choice set T. For example, if both ρ and ρ′ρ are consistent with the Luce rule with the same underlying utility function, then this measure becomes zero. Note that the cross instability measure is generally not symmetric (i.e., Γxy(S,T|ρ,ρ′) _xy(S,T|ρ,ρ ) may not be equal to Γxy(S,T|ρ′,ρ) _xy(S,T|ρ ,ρ)). The measure of composite instability combines the two cross instability measures, which makes it symmetric. We will later see that own and composite instabilities play an important role in identification in the laboratory setting, while cross instability plays an important role in the field setting. Remark 1. For a stochastic choice function ρ satisfying positivity (ρ(x,S)>0ρ(x,S)>0 for all x∈S⊆Xx∈ S X), Δxy(S,T|ρ)=0 _xy(S,T|ρ)=0 for all (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T if and only if ρ is consistent with the Luce rule. For two Luce stochastic choice functions ρ and ρ′ρ with utility functions u and v, respectively, the cross instabilities can be written as Γxy(S,T|ρ,ρ′)=u(x)v(y)−u(y)v(x)u(S)v(T)andΓxy(S,T|ρ′,ρ)=u(y)v(x)−u(x)v(y)u(T)v(S), _xy(S,T|ρ,ρ )= u(x)v(y)-u(y)v(x)u(S)\,v(T) _xy(S,T|ρ ,ρ)= u(y)v(x)-u(x)v(y)u(T)\,v(S), where u(A)=∑a∈Au(a)u(A)= _a∈ Au(a) and v(A)=∑a∈Av(a)v(A)= _a∈ Av(a) for any A∈A . Summing the two cross instabilities yields the composite instability: Φxy(S,T|ρ,ρ′)=[u(x)v(y)−u(y)v(x)]⋅[u(T)v(S)−u(S)v(T)]u(S)u(T)v(S)v(T). _xy(S,T|ρ,ρ )= [u(x)v(y)-u(y)v(x)]·[u(T)v(S)-u(S)v(T)]u(S)u(T)v(S)v(T). Following arguments similar to the ones in the proof of Proposition 1, we can show that Φxy(S,T|ρ,ρ′)=0 _xy(S,T|ρ,ρ )=0 for all (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T if and only if v=λuv=λ u for some λ>0λ>0. Hence, zero own instability for both agents establishes that each is consistent with the Luce rule, and zero composite instability further establishes that they share the same underlying preferences. If ρAIρ^AI violates IIA for some tuple (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, then we must have Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0. The next proposition shows that the compliance parameter α can be recovered by evaluating the ratio Δxy(S,T|ρAI)/Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI)/ _xy(S,T|ρ^AI,ρ^H) for this tuple. Intuitively, the AI’s compliance level is revealed by comparing the instability in the AI’s stochastic choices with the composite instability across both agents. If this ratio is high, then the instability in the AI’s choices matches the composite instability to a large extent, revealing a high compliance level. Alternatively, if this ratio is low, then the composite instability is much higher than the instability in the AI’s choices, which indicates a low compliance level. Proposition 2 (Identification of α). Suppose (ρAI,ρH)(ρ^AI,ρ^H) is consistent with LAM. If ρAIρ^AI violates IIA for some tuple (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, then the compliance parameter α can be uniquely recovered as α=Δxy(S,T|ρAI)Φxy(S,T|ρAI,ρH).α= _xy(S,T|ρ^AI) _xy(S,T|ρ^AI,ρ^H). Proof. Under LAM, ρAI(x,S)=α⋅ρH(x,S)+(1−α)⋅ρA(x,S)ρ^AI(x,S)=α·ρ^H(x,S)+(1-α)·ρ^A(x,S). Substituting this into Δxy(S,T|ρAI) _xy(S,T|ρ^AI), Δxy(S,T|ρAI) _xy(S,T|ρ^AI) =ρAI(x,S)ρAI(y,T)−ρAI(x,T)ρAI(y,S) =ρ^AI(x,S)ρ^AI(y,T)-ρ^AI(x,T)ρ^AI(y,S) =[αρH(x,S)+(1−α)ρA(x,S)][αρH(y,T)+(1−α)ρA(y,T)] =[αρ^H(x,S)+(1-α)ρ^A(x,S)][αρ^H(y,T)+(1-α)ρ^A(y,T)] −[αρH(x,T)+(1−α)ρA(x,T)][αρH(y,S)+(1−α)ρA(y,S)]. -[αρ^H(x,T)+(1-α)ρ^A(x,T)][αρ^H(y,S)+(1-α)ρ^A(y,S)]. Expanding and collecting terms by powers of α, we get Δxy(S,T|ρAI) _xy(S,T|ρ^AI) =α2Δxy(S,T|ρH)+(1−α)2Δxy(S,T|ρA)+α(1−α)Φxy(S,T|ρH,ρA) =α^2 _xy(S,T|ρ^H)+(1-α)^2 _xy(S,T|ρ^A)+α(1-α) _xy(S,T|ρ^H,ρ^A) =α(1−α)Φxy(S,T|ρH,ρA), =α(1-α) _xy(S,T|ρ^H,ρ^A), where the first equality follows from the definitions of Δxy(S,T|ρH) _xy(S,T|ρ^H), Δxy(S,T|ρA) _xy(S,T|ρ^A), and Φxy(S,T|ρH,ρA) _xy(S,T|ρ^H,ρ^A), and the second equality uses the fact that both ρHρ^H and ρAρ^A are consistent with the Luce rule, and hence Δxy(S,T|ρH)=Δxy(S,T|ρA)=0 _xy(S,T|ρ^H)= _xy(S,T|ρ^A)=0. Now substituting ρAI(x,S)=α⋅ρH(x,S)+(1−α)⋅ρA(x,S)ρ^AI(x,S)=α·ρ^H(x,S)+(1-α)·ρ^A(x,S) in the definition of Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI,ρ^H), Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI,ρ^H) =ρAI(x,S)ρH(y,T)+ρH(x,S)ρAI(y,T) =ρ^AI(x,S)ρ^H(y,T)+ρ^H(x,S)ρ^AI(y,T) −ρAI(x,T)ρH(y,S)−ρH(x,T)ρAI(y,S) -ρ^AI(x,T)ρ^H(y,S)-ρ^H(x,T)ρ^AI(y,S) =[αρH(x,S)+(1−α)ρA(x,S)]ρH(y,T) =[αρ^H(x,S)+(1-α)ρ^A(x,S)]ρ^H(y,T) +ρH(x,S)[αρH(y,T)+(1−α)ρA(y,T)] +ρ^H(x,S)[αρ^H(y,T)+(1-α)ρ^A(y,T)] −[αρH(x,T)+(1−α)ρA(x,T)]ρH(y,S) -[αρ^H(x,T)+(1-α)ρ^A(x,T)]ρ^H(y,S) −ρH(x,T)[αρH(y,S)+(1−α)ρA(y,S)] -ρ^H(x,T)[αρ^H(y,S)+(1-α)ρ^A(y,S)] =2αΔxy(S,T|ρH)+(1−α)Φxy(S,T|ρH,ρA) =2α _xy(S,T|ρ^H)+(1-α) _xy(S,T|ρ^H,ρ^A) =(1−α)Φxy(S,T|ρH,ρA), =(1-α) _xy(S,T|ρ^H,ρ^A), where the third equality follows from the definitions of the instability measures and the last equality follows from the fact that ρHρ^H follows the Luce rule. Therefore, Δxy(S,T|ρAI)Φxy(S,T|ρAI,ρH)=α(1−α)Φxy(S,T|ρH,ρA)(1−α)Φxy(S,T|ρH,ρA)=α. _xy(S,T|ρ^AI) _xy(S,T|ρ^AI,ρ^H)= α(1-α) _xy(S,T|ρ^H,ρ^A)(1-α) _xy(S,T|ρ^H,ρ^A)=α. The cancellation is valid since an IIA violation implies Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0 and the above derivations show that this implies Φxy(S,T|ρAI,ρH)≠0 _xy(S,T|ρ^AI,ρ^H)≠ 0. ∎ There are two immediate but non-obvious implications of the proof. The first is that, under LAM, the instability measures Δxy(S,T|ρAI) _xy(S,T|ρ^AI) and Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI,ρ^H) are always proportional. Hence, the compliance formula in Proposition 2 is well-defined whenever there is an IIA violation in ρAIρ^AI, i.e., Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0 automatically guarantees a non-zero denominator. Second, the expression derived for composite instability, Φxy(S,T|ρAI,ρH)=(1−α)Φxy(S,T|ρH,ρA) _xy(S,T|ρ^AI,ρ^H)=(1-α) _xy(S,T|ρ^H,ρ^A), shows that the composite instability of ρAIρ^AI and ρHρ^H is always zero if and only if the AI is either fully aligned or fully compliant. Since ρAI=ρHρ^AI=ρ^H in both cases, it follows that the expression for the compliance parameter is valid as long as ρAI≠ρHρ^AI≠ρ^H. Corollary 1. If (ρAI,ρH)(ρ^AI,ρ^H) is consistent with LAM, then for all tuples (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, Δxy(S,T|ρAI)=α⋅Φxy(S,T|ρAI,ρH). _xy(S,T|ρ^AI)=α· _xy(S,T|ρ^AI,ρ^H). Moreover, the compliance parameter α is uniquely identified by this relationship as long as ρAI≠ρHρ^AI≠ρ^H. The last step in identification is to recover the AI’s utility function v. Step 3: Recover v from ρAIρ^AI and ρHρ^H. As before, if ρAI=ρHρ^AI=ρ^H, then we cannot distinguish full compliance (α=1α=1) from full alignment (v=λuv=λ u for λ>0λ>0). Alternatively, if ρAI≠ρHρ^AI≠ρ^H, then Step 2 allows us to uniquely identify the compliance parameter α. Let α denote the recovered compliance level. Using the fact that ρAI(x,S)=α⋅ρH(x,S)+(1−α)⋅ρA(x,S),ρ^AI(x,S)=α·ρ^H(x,S)+(1-α)·ρ^A(x,S), we can construct ρAρ^A as ρA(x,S)=ρAI(x,S)−αρH(x,S)1−α ρ^A(x,S)= ρ^AI(x,S)-αρ^H(x,S)1-α Under LAM, ρAρ^A is generated by the Luce rule with the utility function v. We can use this to construct v from ρAρ^A as in Step 1. Theorem 1 summarizes the identification results in the laboratory setting. Theorem 1 (Laboratory Identification). Let (ρAI,ρH)(ρ^AI,ρ^H) be consistent with LAM. 1. If ρAI≠ρHρ^AI≠ρ^H, then α is uniquely identified and u and v are uniquely identified up to scale normalization. 2. If ρAI=ρHρ^AI=ρ^H, then α and v are not separately identified and only u is uniquely identified up to scale normalization. Proof. The proof follows from the three identification steps and the results established in this section. ∎ The following example illustrates the identification result. Example 1. Consider X=x,y,zX=\x,y,z\ and suppose we observe ρAIρ^AI and ρHρ^H given as follows. Agent Option x,y,z\x,y,z\ x,y\x,y\ x,z\x,z\ y,z\y,z\ ρAIρ^AI x 1/31/3 7/157/15 1/21/2 – y 1/31/3 8/158/15 – 8/158/15 z 1/31/3 – 1/21/2 7/157/15 ρHρ^H x 1/21/2 3/53/5 3/43/4 – y 1/31/3 2/52/5 – 2/32/3 z 1/61/6 – 1/41/4 1/31/3 Table 1: Observed choice probabilities in Example 1 Normalizing u(x)=1u(x)=1, we can infer from ρHρ^H that u(y)=2/3u(y)=2/3 and u(z)=1/3u(z)=1/3. To recover α, we construct the two instability measures. Let S=x,y,zS=\x,y,z\ and T=x,yT=\x,y\. We first compute the instability of ρAIρ^AI for the tuple (x,y,S,T)(x,y,S,T): Δxy(S,T|ρAI) _xy(S,T|ρ^AI) =ρAI(x,S)ρAI(y,T)−ρAI(x,T)ρAI(y,S) =ρ^AI(x,S)ρ^AI(y,T)-ρ^AI(x,T)ρ^AI(y,S) =13⋅815−715⋅13=145. = 13· 815- 715· 13= 145. Next, we compute the composite instability between ρAIρ^AI and ρHρ^H: Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI,ρ^H) =ρAI(x,S)ρH(y,T)+ρH(x,S)ρAI(y,T) =ρ^AI(x,S)ρ^H(y,T)+ρ^H(x,S)ρ^AI(y,T) −ρAI(x,T)ρH(y,S)−ρH(x,T)ρAI(y,S) -ρ^AI(x,T)ρ^H(y,S)-ρ^H(x,T)ρ^AI(y,S) =13⋅25+12⋅815−715⋅13−35⋅13 = 13· 25+ 12· 815- 715· 13- 35· 13 =215+415−745−315=245. = 215+ 415- 745- 315= 245. By Proposition 2, the compliance parameter is uniquely identified as α=1/452/45=1/2α= 1/452/45=1/2. We could similarly recover α using the tuple (y,z,x,y,z,y,z)(y,z,\x,y,z\,\y,z\) instead. Note, however, that we cannot use the tuple (x,z,x,y,z,x,z)(x,z,\x,y,z\,\x,z\) as ρAIρ^AI satisfies IIA for this tuple. This highlights that while ρAI≠ρHρ^AI≠ρ^H implies α is uniquely identified, not all tuples can be used for identification. Using α=1/2α=1/2, we can construct ρAρ^A as follows: Agent Option x,y,z\x,y,z\ x,y\x,y\ x,z\x,z\ y,z\y,z\ ρAρ^A x 1/61/6 1/31/3 1/41/4 – y 1/31/3 2/32/3 – 2/52/5 z 1/21/2 – 3/43/4 3/53/5 Table 2: Recovered autonomous AI stochastic choice ρAρ^A Normalizing v(x)=1v(x)=1, we can infer from ρAρ^A that v(y)=2v(y)=2 and v(z)=3v(z)=3. Hence, we have u=(1,2/3,1/3),v=(1,2,3),α=1/2.u=(1,2/3,1/3), v=(1,2,3), α=1/2. Note that u and v induce completely opposite ordinal rankings, revealing a high degree of misalignment. 3.2 Axiomatic Characterization In this section, I provide an axiomatic characterization for the Luce Alignment Model taking (ρAI,ρH)(ρ^AI,ρ^H) as the primitive. The first two axioms are standard. Axiom 1 requires that ρ(x,S)ρ(x,S) is strictly positive for any x∈S⊆Xx∈ S X and ρ∈ρAI,ρHρ∈\ρ^AI,ρ^H\. Axiom 2 requires that ρHρ^H satisfies IIA, ensuring that the human principal’s behavior is consistent with the Luce rule. Axiom 1 (Positivity). For any ρ∈ρAI,ρHρ∈\ρ^AI,ρ^H\ and x∈S⊆Xx∈ S X, we have ρ(x,S)>0ρ(x,S)>0. Axiom 2 (H-IIA). ρHρ^H satisfies IIA. Axiom 3 is the key axiom in ensuring that the AI compliance parameter can be identified. It requires that the own instability of ρAIρ^AI and the composite instability of ρAIρ^AI and ρHρ^H are proportional: any change in composite instability from one tuple to another must be proportionally reflected by a change in the AI’s own instability. Axiom 3 (Proportionality). For any two tuples (x,y,S,T)(x,y,S,T) and (z,t,S′,T′)(z,t,S ,T ) with x,y∈S∩Tx,y∈ S∩ T and z,t∈S′∩T′z,t∈ S ∩ T , Δxy(S,T|ρAI)⋅Φzt(S′,T′|ρAI,ρH)=Δzt(S′,T′|ρAI)⋅Φxy(S,T|ρAI,ρH). _xy(S,T|ρ^AI)· _zt(S ,T |ρ^AI,ρ^H)= _zt(S ,T |ρ^AI)· _xy(S,T|ρ^AI,ρ^H). Axiom 4 requires that the AI’s own instability always shares the same sign as the composite instability, and the own instability is bounded by the composite instability. To get an intuition for this axiom, consider the case Δxy(S,T|ρAI)>0 _xy(S,T|ρ^AI)>0. This implies that if we use the AI’s relative choice probabilities for x and y in S to impute its relative choice probabilities in T, this will lead to an overestimation of x versus y. The first part of the axiom then requires that the composite instability must also be strictly positive: using the AI’s and the human’s relative choice probabilities in S to cross-impute the relative choice probabilities in T will also lead to an overestimation of x in aggregate. In addition, the axiom requires that the aggregate cross-imputation error must be larger than the own imputation error. Together with Axiom 3, this property ensures that the compliance parameter is uniquely identified and bounded between zero and one. Axiom 4 (Bounded Instability). For any tuple (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, Δxy(S,T|ρAI)⋅Φxy(S,T|ρAI,ρH)≥0and|Δxy(S,T|ρAI)|≤|Φxy(S,T|ρAI,ρH)|, _xy(S,T|ρ^AI)· _xy(S,T|ρ^AI,ρ^H)≥ 0 | _xy(S,T|ρ^AI)|≤| _xy(S,T|ρ^AI,ρ^H)|, where both inequalities hold strictly if Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0. The last axiom bounds the divergence between the AI’s and the human’s stochastic choices. Fixing own and composite instability measures and the human’s stochastic choices, it provides a lower bound on the AI’s stochastic choices. Axiom 5 (Bounded Divergence). For any tuple (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, menu U, and alternative z∈Uz∈ U, ρAI(z,U)⋅|Φxy(S,T|ρAI,ρH)|≥ρH(z,U)⋅|Δxy(S,T|ρAI)|.ρ^AI(z,U)·| _xy(S,T|ρ^AI,ρ^H)|≥ρ^H(z,U)·| _xy(S,T|ρ^AI)|. Moreover, if Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0, the inequality is strict. To interpret this axiom, suppose the instability measures are strictly positive. The axiom then requires that ρAI(z,U)ρH(z,U)>Δxy(S,T|ρAI)Φxy(S,T|ρAI,ρH). ρ^AI(z,U)ρ^H(z,U)> _xy(S,T|ρ^AI) _xy(S,T|ρ^AI,ρ^H). The right-hand side of the above inequality gives us the relative imputation error: it tells us the proportion of aggregate cross-imputation error that can be explained by the AI’s own imputation error. The axiom then requires that the higher the relative imputation error is, the more the AI’s choice probabilities are constrained to track the human’s choice probabilities. Theorem 2 establishes that Axioms 1–5 are necessary and sufficient for a LAM representation. An interesting feature of this characterization is that while there is an explicit axiom imposing the IIA property on the human’s choices ρHρ^H, there is no equivalent axiom for the autonomous AI’s stochastic choices ρAρ^A. Instead, the proof of the theorem shows that the IIA property for ρAρ^A is jointly implied by the axioms. Theorem 2 (Laboratory Characterization). The pair (ρAI,ρH)(ρ^AI,ρ^H) satisfies Axioms 1–5 if and only if it is consistent with LAM. Proof. Necessity. Axiom 1 follows from the assumption that u and v are strictly positive in LAM. Axiom 2 follows from the fact that ρHρ^H is consistent with the Luce rule. Axioms 3 and 4 follow from Corollary 1, which shows that Δxy(S,T|ρAI)=α⋅Φxy(S,T|ρAI,ρH). _xy(S,T|ρ^AI)=α· _xy(S,T|ρ^AI,ρ^H). For Axiom 5, if Δxy(S,T|ρAI)=0 _xy(S,T|ρ^AI)=0, then the inequality follows trivially. Alternatively, if Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0, then α must be strictly less than 1, as α=1α=1 would imply ρAI=ρHρ^AI=ρ^H and hence Δxy(S,T|ρAI)=0 _xy(S,T|ρ^AI)=0. Therefore, ρAI(z,U)−αρH(z,U)=(1−α)ρA(z,U)>0⇒ρAI(z,U)>αρH(z,U).ρ^AI(z,U)-αρ^H(z,U)=(1-α)ρ^A(z,U)>0 ρ^AI(z,U)>αρ^H(z,U). Substituting the result α=Δxy(S,T|ρAI)/Φxy(S,T|ρAI,ρH)α= _xy(S,T|ρ^AI)/ _xy(S,T|ρ^AI,ρ^H) from the previous section into the above inequality yields the axiom. Sufficiency. Axioms 1 and 2 imply that ρHρ^H is consistent with the Luce rule with some utility function u that is strictly positive. There are two cases to consider. First, suppose Δxy(S,T|ρAI)=0 _xy(S,T|ρ^AI)=0 for all (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T. Then, by Remark 1, ρAIρ^AI is consistent with the Luce rule with some strictly positive utility function v. In this case, (u,v,α=0)(u,v,α=0) is a LAM representation of (ρAI,ρH)(ρ^AI,ρ^H). If, in addition, Φxy(S,T|ρAI,ρH)=0 _xy(S,T|ρ^AI,ρ^H)=0 for all (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T, then we must have v=λuv=λ u for some λ>0λ>0 and α can be arbitrary. Next, suppose Δxy(S,T|ρAI)≠0 _xy(S,T|ρ^AI)≠ 0 for some (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T. By Axiom 4, Δxy(S,T|ρAI)≠0⇒Φxy(S,T|ρAI,ρH)≠0 _xy(S,T|ρ^AI)≠ 0 _xy(S,T|ρ^AI,ρ^H)≠ 0 Hence, by Axiom 3, the ratio Δxy(S,T|ρAI)Φxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI) _xy(S,T|ρ^AI,ρ^H) is constant for all such tuples (x,y,S,T)(x,y,S,T). Let α denote the above ratio. Axiom 4 guarantees that α∈(0,1)α∈(0,1). We next define ρAρ^A by ρA(x,S)=ρAI(x,S)−αρH(x,S)1−α. ρ^A(x,S)= ρ^AI(x,S)-αρ^H(x,S)1-α. Since the requirements ρ(x,S)=0ρ(x,S)=0 for x∉Sx∉ S and ∑x∈SρA(x,S)=1 _x∈ Sρ^A(x,S)=1 hold, ρAρ^A is a valid stochastic choice function. In addition, by Axiom 5, the numerator is strictly positive for any x∈Sx∈ S so that ρA(x,S)>0ρ^A(x,S)>0. Rearranging the above equation yields ρAI(x,S)=αρH(x,S)+(1−α)ρA(x,S).ρ^AI(x,S)=αρ^H(x,S)+(1-α)ρ^A(x,S). To conclude the proof of the theorem, we only need to show that ρAρ^A is consistent with the Luce rule. By Remark 1, it is sufficient to show that Δxy(S,T|ρA)=0 _xy(S,T|ρ^A)=0 for all tuples (x,y,S,T)(x,y,S,T) with x,y∈S∩Tx,y∈ S∩ T. Let such a tuple be given. As shown in the proof of Proposition 2, substituting ρAI=α⋅ρH+(1−α)⋅ρAρ^AI=α·ρ^H+(1-α)·ρ^A back into the own and composite instability measures yields Δxy(S,T|ρAI)=α2Δxy(S,T|ρH)+(1−α)2Δxy(S,T|ρA)+α(1−α)Φxy(S,T|ρH,ρA) _xy(S,T|ρ^AI)=α^2 _xy(S,T|ρ^H)+(1-α)^2 _xy(S,T|ρ^A)+α(1-α) _xy(S,T|ρ^H,ρ^A) and Φxy(S,T|ρAI,ρH)=2αΔxy(S,T|ρH)+(1−α)Φxy(S,T|ρH,ρA). _xy(S,T|ρ^AI,ρ^H)=2α _xy(S,T|ρ^H)+(1-α) _xy(S,T|ρ^H,ρ^A). By Axioms 1–2 and Remark 1, Δxy(S,T|ρH)=0 _xy(S,T|ρ^H)=0. Hence, combining the last two expressions, we have Δxy(S,T|ρAI) _xy(S,T|ρ^AI) =(1−α)2Δxy(S,T|ρA)+α(1−α)Φxy(S,T|ρH,ρA) =(1-α)^2 _xy(S,T|ρ^A)+α(1-α) _xy(S,T|ρ^H,ρ^A) =(1−α)2Δxy(S,T|ρA)+αΦxy(S,T|ρAI,ρH). =(1-α)^2 _xy(S,T|ρ^A)+α _xy(S,T|ρ^AI,ρ^H). Consider a tuple (z,t,S′,T′)(z,t,S ,T ) with z,t∈S′∩T′z,t∈ S ∩ T that satisfies α=Δzt(S′,T′|ρAI)Φzt(S′,T′|ρAI,ρH).α= _zt(S ,T |ρ^AI) _zt(S ,T |ρ^AI,ρ^H). Substituting this into the last term of the above expression and cross-multiplying, we get Δxy(S,T|ρAI)Φzt(S′,T′|ρAI,ρH) _xy(S,T|ρ^AI) _zt(S ,T |ρ^AI,ρ^H) =(1−α)2Δxy(S,T|ρA)Φzt(S′,T′|ρAI,ρH) =(1-α)^2 _xy(S,T|ρ^A) _zt(S ,T |ρ^AI,ρ^H) +Δzt(S′,T′|ρAI)Φxy(S,T|ρAI,ρH). + _zt(S ,T |ρ^AI) _xy(S,T|ρ^AI,ρ^H). Since α∈(0,1)α∈(0,1), we must have Φzt(S′,T′|ρAI,ρH)≠0 _zt(S ,T |ρ^AI,ρ^H)≠ 0. In addition, by Axiom 3, Δxy(S,T|ρAI)Φzt(S′,T′|ρAI,ρH)=Δzt(S′,T′|ρAI)Φxy(S,T|ρAI,ρH). _xy(S,T|ρ^AI) _zt(S ,T |ρ^AI,ρ^H)= _zt(S ,T |ρ^AI) _xy(S,T|ρ^AI,ρ^H). Therefore, the above expression can hold only if Δxy(S,T|ρA)=0 _xy(S,T|ρ^A)=0. Since the tuple (x,y,S,T)(x,y,S,T) was arbitrary, ρAρ^A is consistent with the Luce rule, as desired. This concludes the proof of the theorem as we have shown that ρAIρ^AI is a mixture of two Luce rules, where one of the mixing parts is ρHρ^H. ∎ 4 Field Data In this section, I study the Luce Alignment Model when only the AI’s choices ρAIρ^AI are observable. This setting is important for two reasons. First, while laboratory data may be readily available, the volume of field data is expected to be much larger, which can enable richer inference about AI behavior. Second, AI behavior in the two settings may differ systematically: a sufficiently sophisticated AI may appear compliant in a monitored laboratory setting while reverting to its autonomous preferences in the field, a phenomenon known as deceptive alignment (Greenblatt et al., 2024). Comparing recovered compliance parameters across the two settings can provide a measure of deceptive alignment. 4.1 Identification The identification problem in the field setting faces an inherent challenge: if (u,v,α)(u,v,α) is a LAM representation of ρAIρ^AI, then so is (v,u,1−α)(v,u,1-α). That is, even if the two utility functions underlying LAM can be recovered, the data alone cannot reveal which belongs to the human principal and which to the AI agent. The utilities can therefore be identified only up to a label swap, and the compliance parameter only up to reflection about 1/21/2. Note, however, that the distribution over utilities may still be uniquely identified. Furthermore, if ρAIρ^AI satisfies IIA, then the observed choice behavior can be consistent with any alignment and compliance levels: we can have either (i) v=λuv=λ u for some λ>0λ>0 with arbitrary α∈[0,1]α∈[0,1], or (i) α∈0,1α∈\0,1\ with v≠λuv≠λ u for any λ>0λ>0. Hence, the identification problem is interesting only if ρAIρ^AI violates IIA. The key for the identification in this section will be the cross instability measure Γxy(S,T|ρAI,ρ) _xy(S,T|ρ^AI,ρ) for ρ∈ρH,ρAρ∈\ρ^H,ρ^A\, defined in Definition 2. The next proposition provides a formula for the cross instability measures in terms of (u,v,α)(u,v,α). Proposition 3. Suppose ρAIρ^AI is consistent with LAM with parameters (u,v,α)(u,v,α), and let ρHρ^H and ρAρ^A be the corresponding Luce rules. Then, Γxy(S,T|ρAI,ρH)=(1−α)Γxy(S,T|ρA,ρH)=(1−α)⋅u(y)v(x)−u(x)v(y)u(T)v(S) _xy(S,T|ρ^AI,ρ^H)=(1-α)\, _xy(S,T|ρ^A,ρ^H)=(1-α)· u(y)v(x)-u(x)v(y)u(T)\,v(S) and Γxy(S,T|ρAI,ρA)=αΓxy(S,T|ρH,ρA)=α⋅u(x)v(y)−u(y)v(x)u(S)v(T). _xy(S,T|ρ^AI,ρ^A)=α\, _xy(S,T|ρ^H,ρ^A)=α· u(x)v(y)-u(y)v(x)u(S)\,v(T). Proof. Substituting ρAI(x,S)=αρH(x,S)+(1−α)ρA(x,S)ρ^AI(x,S)=α\,ρ^H(x,S)+(1-α)\,ρ^A(x,S) into the definition of cross instability, Γxy(S,T|ρAI,ρH) _xy(S,T|ρ^AI,ρ^H) =ρAI(x,S)ρH(y,T)−ρAI(y,S)ρH(x,T) =ρ^AI(x,S)ρ^H(y,T)-ρ^AI(y,S)ρ^H(x,T) =α[ρH(x,S)ρH(y,T)−ρH(y,S)ρH(x,T)] =α [ρ^H(x,S)ρ^H(y,T)-ρ^H(y,S)ρ^H(x,T) ] +(1−α)[ρA(x,S)ρH(y,T)−ρA(y,S)ρH(x,T)] +(1-α) [ρ^A(x,S)ρ^H(y,T)-ρ^A(y,S)ρ^H(x,T) ] =αΔxy(S,T|ρH)+(1−α)Γxy(S,T|ρA,ρH). =α\, _xy(S,T|ρ^H)+(1-α)\, _xy(S,T|ρ^A,ρ^H). Since ρHρ^H is consistent with the Luce rule, Δxy(S,T|ρH)=0 _xy(S,T|ρ^H)=0. Combining this with the result in Remark 1, we get the first identity. The second identity follows analogously by expanding Γxy(S,T|ρAI,ρA) _xy(S,T|ρ^AI,ρ^A) and using Δxy(S,T|ρA)=0 _xy(S,T|ρ^A)=0. ∎ The next proposition provides a key equation that will be used in the identification of u and v from ρAIρ^AI. Proposition 4. Suppose ρAIρ^AI is consistent with LAM with parameters (u,v,α)(u,v,α), and let ρHρ^H and ρAρ^A be the corresponding Luce rules. Let ρ∈ρH,ρAρ∈\ρ^H,ρ^A\, S=x,y,z,tS=\x,y,z,t\, and T=x,yT=\x,y\, where x,y,z,tx,y,z,t are four distinct alternatives, and assume the associated cross instabilities are non-zero. Then, 1Γxy(S,T|ρAI,ρ) 1 _xy(S,T|ρ^AI,ρ) +1Γxy(T,T|ρAI,ρ)=1Γxy(S∖t,T|ρAI,ρ)+1Γxy(S∖z,T|ρAI,ρ). + 1 _xy(T,T|ρ^AI,ρ)= 1 _xy(S t,T|ρ^AI,ρ)+ 1 _xy(S z,T|ρ^AI,ρ). Proof. Consider the case ρ=ρHρ=ρ^H. Letting S=x,y,z,tS=\x,y,z,t\ and T=x,yT=\x,y\, we know from Proposition 3 that Γxy(S′,T|ρAI,ρH)=(1−α)⋅u(y)v(x)−u(x)v(y)u(T)v(S′) _xy(S ,T|ρ^AI,ρ^H)=(1-α)· u(y)v(x)-u(x)v(y)u(T)\,v(S ) for any menu S′⊇TS T. Hence, 1Γxy(S′,T|ρAI,ρH) 1 _xy(S ,T|ρ^AI,ρ^H) =u(T)v(S′)(1−α)[u(y)v(x)−u(x)v(y)] = u(T)\,v(S )(1-α)[u(y)v(x)-u(x)v(y)] =v(S′)(1−α)[u(y)v(x)−u(x)v(y)]/u(T). = v(S )(1-α)[u(y)v(x)-u(x)v(y)]/u(T). Notice that the denominator is independent of S′S . Hence, the result holds as long as v(S)+v(T)=v(S∖t)+v(S∖z)v(S)+v(T)=v(S t)+v(S z). This holds trivially since both sides of the equation evaluate to 2v(x)+2v(y)+v(z)+v(t)2v(x)+2v(y)+v(z)+v(t). The case ρ=ρAρ=ρ^A follows analogously with u replacing v and vice versa. ∎ We will use this result to identify both u and v. The identification strategy proceeds in three steps. Step 1: Recover u(y)u(y) and v(y)v(y) from ρAIρ^AI for each y∈Xy∈ X. The identification of utility functions in the field setting involves two separate steps. First, I show how the candidate utility values for each alternative can be identified. In Step 3, I combine the prior two steps to identify the overall utility functions. I start with the identification of u(y)u(y) for each y∈Xy∈ X, and the same process also works for v(y)v(y). Assume X contains at least four alternatives, and let u(x)=1u(x)=1 for some x∈Xx∈ X. Since utility functions are identified only up to scale normalization, this is without loss. For any y≠xy≠ x, notice that ρH(x,x,y)=11+u(y)andρH(y,x,y)=u(y)1+u(y).ρ^H(x,\x,y\)= 11+u(y) ρ^H(y,\x,y\)= u(y)1+u(y). Pick two other alternatives z and t distinct from x and y, and let S=x,y,z,tS=\x,y,z,t\, T=x,yT=\x,y\, and T⊆S′⊆ST S S. We have Γxy(S′,T|ρAI,ρH) _xy(S ,T|ρ^AI,ρ^H) =ρAI(x,S′)ρH(y,T)−ρAI(y,S′)ρH(x,T) =ρ^AI(x,S )ρ^H(y,T)-ρ^AI(y,S )ρ^H(x,T) =ρAI(x,S′)u(y)1+u(y)−ρAI(y,S′)11+u(y) =ρ^AI(x,S ) u(y)1+u(y)-ρ^AI(y,S ) 11+u(y) =ρAI(x,S′)u(y)−ρAI(y,S′)1+u(y). = ρ^AI(x,S )u(y)-ρ^AI(y,S )1+u(y). Since ρAIρ^AI is observed, this is an equation in terms of one unknown u(y)u(y). There are two cases to consider. Case 1: ρAI(x,S′)u(y)≠ρAI(y,S′)ρ^AI(x,S )u(y)≠ρ^AI(y,S ) for all S′S with T⊆S′⊆ST S S. This ensures that Γxy(S′,T|ρAI,ρH)≠0 _xy(S ,T|ρ^AI,ρ^H)≠ 0. Utilizing Proposition 4 with ρ=ρHρ=ρ^H and canceling the common (1+u(y))(1+u(y)) terms, we get 1ρAI(x,S)u(y)−ρAI(y,S) 1ρ^AI(x,S)u(y)-ρ^AI(y,S) +1ρAI(x,T)u(y)−ρAI(y,T) + 1ρ^AI(x,T)u(y)-ρ^AI(y,T) =1ρAI(x,S∖t)u(y)−ρAI(y,S∖t)+1ρAI(x,S∖z)u(y)−ρAI(y,S∖z). = 1ρ^AI(x,S t)u(y)-ρ^AI(y,S t)+ 1ρ^AI(x,S z)u(y)-ρ^AI(y,S z). Cross-multiplying, we get a cubic polynomial in terms of the unknown u(y)u(y). Normalizing v(x)=1v(x)=1 and re-deriving Proposition 4 with ρ=ρAρ=ρ^A instead of ρHρ^H, we deduce that v(y)v(y) must also satisfy the same polynomial, provided that ρAI(x,S′)v(y)≠ρAI(y,S′)ρ^AI(x,S )v(y)≠ρ^AI(y,S ) for all S′S with T⊆S′⊆ST S S. Case 2: ρAI(x,S′)u(y)=ρAI(y,S′)ρ^AI(x,S )u(y)=ρ^AI(y,S ) for some S′S with T⊆S′⊆ST S S. Alternatively, ρAI(x,S′)ρAI(y,S′)=u(x)u(y)=ρH(x,S′)ρH(y,S′), ρ^AI(x,S )ρ^AI(y,S )= u(x)u(y)= ρ^H(x,S )ρ^H(y,S ), where the first equality is due to u(x)=1u(x)=1. Since we are assuming ρAIρ^AI violates IIA, we cannot have α=1α=1 by Proposition 1. Therefore, the above equality is possible only if ρH(x,S′)ρH(y,S′)=ρA(x,S′)ρA(y,S′)⇒u(x)u(y)=v(x)v(y). ρ^H(x,S )ρ^H(y,S )= ρ^A(x,S )ρ^A(y,S ) u(x)u(y)= v(x)v(y). But then the ratio ρAI(y,⋅)/ρAI(x,⋅)ρ^AI(y,·)/ρ^AI(x,·) must be constant and equal to u(y)u(y) for all menus. Normalizing v(x)=1v(x)=1, we also get u(y)=v(y)u(y)=v(y). Note that in this case the polynomial formed by cross-multiplying the equation in Case 1 will either yield the utilities u(y)u(y) and v(y)v(y) as a unique root or the polynomial will be identically zero. If the polynomial is identically zero, then we can generically infer that we are in Case 2, which trivially recovers u(y)u(y) and v(y)v(y) as ρAI(y,x,y)/ρAI(x,x,y)ρ^AI(y,\x,y\)/ρ^AI(x,\x,y\).222A result is said to hold generically if it fails only on a measure-zero subset of the underlying parameter space. Proposition 5 (Identification of u(y)u(y) and v(y)v(y)). Suppose ρAIρ^AI is consistent with LAM with (u,v,α)(u,v,α) such that u(x)=v(x)=1u(x)=v(x)=1, and suppose ρAIρ^AI violates IIA. For any y≠xy≠ x, let P(κy)P( _y) be the cubic polynomial obtained by cross-multiplying the equation 1ρAI(x,S)κy−ρAI(y,S) 1ρ^AI(x,S) _y-ρ^AI(y,S) +1ρAI(x,T)κy−ρAI(y,T) + 1ρ^AI(x,T) _y-ρ^AI(y,T) (3) =1ρAI(x,S∖t)κy−ρAI(y,S∖t)+1ρAI(x,S∖z)κy−ρAI(y,S∖z), = 1ρ^AI(x,S t) _y-ρ^AI(y,S t)+ 1ρ^AI(x,S z) _y-ρ^AI(y,S z), where S=x,y,z,tS=\x,y,z,t\, T=x,yT=\x,y\, and z,tz,t are two alternatives distinct from x,yx,y. If P(κy)P( _y) is not identically zero, then u(y)u(y) and v(y)v(y) are both roots of P(κy)P( _y) and admissible solutions to equation (3). Otherwise, u(y)=v(y)=ρAI(y,x,y)/ρAI(x,x,y)u(y)=v(y)=ρ^AI(y,\x,y\)/ρ^AI(x,\x,y\) holds generically. Proof. The proof follows from the arguments preceding the proposition. ∎ There are two important points to consider regarding this result. First, while the model has two unknown utility values u(y)u(y) and v(y)v(y) for each alternative y, the derived cubic polynomial P(κy)P( _y) generically has three distinct roots. Hence, solving the polynomial may yield a spurious root that is not a true utility value. However, since equation (3) must hold for any reference pair z and t distinct from x and y and the spurious root will typically vary depending on the chosen reference pair, if the analyst has access to a fifth alternative, rederiving the polynomial using a different reference pair will generically isolate the true utility values. Thus, as long as |X|≥5|X|≥ 5, both u(y)u(y) and v(y)v(y) are generically identified up to scale normalization and label swaps. In addition, as detailed in Step 2, this identification procedure can be improved to require only |X|≥4|X|≥ 4. Second, note that successfully identifying the true candidate pair u(y),v(y)\u(y),v(y)\ for each alternative y∈Xy∈ X does not fully pin down the utility functions u and v. To illustrate, consider three alternatives x,y,zx,y,z and normalize u(x)=v(x)=1u(x)=v(x)=1. Suppose we have recovered candidate utility pairs κy1,κy2\ _y^1, _y^2\ and κz1,κz2\ _z^1, _z^2\. Since u(y)u(y) and u(z)u(z) can be either of these utility values, this leaves us with four candidate utility functions u: (1,κy1,κz1)(1, _y^1, _z^1), (1,κy1,κz2)(1, _y^1, _z^2), (1,κy2,κz1)(1, _y^2, _z^1), or (1,κy2,κz2)(1, _y^2, _z^2). Generalizing this insight, for |X|=N|X|=N, identifying candidate utility pairs for each alternative still leaves us with 2N−12^N-1 candidate utility functions. To resolve this problem, we first need to recover the compliance parameter α, as illustrated in the next step. Step 2: Recover α from ρAIρ^AI. Following Step 1, suppose we have a candidate utility pair u(y),v(y)\u(y),v(y)\ for each alternative y∈X∖xy∈ X x and assume u(x)=v(x)=1u(x)=v(x)=1. Construct the associated Luce choice probabilities ρu(x,x,y)ρ^u(x,\x,y\) and ρv(x,x,y)ρ^v(x,\x,y\) for each y≠xy≠ x. If u(y),v(y)\u(y),v(y)\ is the true utility pair up to a label swap, then the observed AI stochastic choice function ρAIρ^AI must satisfy one of the following equations: ρAI(x,x,y) ρ^AI(x,\x,y\) =α⋅ρu(x,x,y)+(1−α)⋅ρv(x,x,y), =α·ρ^u(x,\x,y\)+(1-α)·ρ^v(x,\x,y\), ρAI(x,x,y) ρ^AI(x,\x,y\) =(1−α)⋅ρu(x,x,y)+α⋅ρv(x,x,y), =(1-α)·ρ^u(x,\x,y\)+α·ρ^v(x,\x,y\), where α is the compliance parameter. Hence, for each alternative y≠xy≠ x and each candidate utility pair, we get two possible candidates for the compliance parameter. For the true utility pair, this identifies the compliance parameter up to reflection about 1/21/2. Step 1 generically recovers the true utility pair for an alternative when |X|≥5|X|≥ 5. Alternatively, suppose we have three candidate utility pairs for an alternative after solving the cubic polynomial. Note that any candidate utility pair for an alternative that is not the true utility pair will imply a compliance parameter that will not be generically validated by the candidate utility pairs for other alternatives. Hence, by ensuring the consistency of the implied compliance parameter across different alternatives, we can identify the true utility pair. Adopting this approach, we only need |X|≥4|X|≥ 4, which improves upon the procedure in Step 1. Lastly, note that while the compliance parameter is identified up to reflection about 1/21/2, the distribution over utilities is generically uniquely identified. Step 3: Recover u and v from ρAIρ^AI. Following Steps 1 and 2, for each alternative y∈X∖xy∈ X x, we can generically identify the true utility pair u(y),v(y)\u(y),v(y)\ up to a label swap. However, as discussed in Step 1, identifying the true pair for each alternative does not by itself determine which value belongs to u and which belongs to v across alternatives. To resolve the remaining ambiguity, fix an alternative y∈X∖xy∈ X x and suppose we assign u(y)=κy1u(y)= _y^1. By Step 2, this assignment implies a candidate compliance parameter αuα^u. Now consider any other alternative z∈X∖x,yz∈ X \x,y\ with the utility pair κz1,κz2\ _z^1, _z^2\. Since the true utility function must generate the same compliance parameter across all alternatives, we can pin down the true assignment for u(z)u(z) by requiring consistency with αuα^u. Generically, unless αu∈0,1/2,1α^u∈\0,1/2,1\, exactly one of the two candidate values for u(z)u(z) will be consistent with αuα^u. We can then repeat this procedure for all alternatives to recover the utility functions u and v up to a label swap. Combining all the results in this section, we have the following identification result in the field setting. Theorem 3 (Field Identification). Suppose ρAIρ^AI is consistent with LAM and |X|≥4|X|≥ 4. 1. If ρAIρ^AI violates IIA, then (u,v,α)(u,v,α) are generically identified up to a label swap and scale normalization. 2. If ρAIρ^AI satisfies IIA, then either v=λuv=λ u for λ>0λ>0 or α∈0,1α∈\0,1\. Proof. The proof follows from the three identification steps and the results established in this section. ∎ The next example illustrates the field identification result. Example 2. Suppose X=x,y,z,tX=\x,y,z,t\ and the AI stochastic choice data ρAIρ^AI is generated by the parameters u=(1,2,4,5),v=(1,4/5,2/5,1/5),α=3/4,u=(1,2,4,5), v=(1,4/5,2/5,1/5), α=3/4, as given in the following table. ρAI(⋅,⋅)ρ^AI(·,·) x,y,z,t\x,y,z,t\ x,y,z\x,y,z\ x,y,t\x,y,t\ x,z,t\x,z,t\ y,z,t\y,z,t\ x,y\x,y\ x,z\x,z\ x,t\x,t\ y,z\y,z\ y,t\y,t\ z,t\z,t\ x 1/61/6 17/7717/77 7/327/32 37/16037/160 7/187/18 23/7023/70 1/31/3 y 5/245/24 47/15447/154 23/8023/80 43/15443/154 11/1811/18 5/125/12 29/7029/70 z 7/247/24 73/15473/154 29/8029/80 53/15453/154 47/7047/70 7/127/12 1/21/2 t 1/31/3 79/16079/160 13/3213/32 29/7729/77 2/32/3 41/7041/70 1/21/2 Table 3: AI stochastic choice data in Example 2 To proceed with identification, we first normalize u(x)=v(x)=1u(x)=v(x)=1. Consider the alternative y. Equation (3) corresponding to y with S=x,y,z,tS=\x,y,z,t\ and T=x,yT=\x,y\ is given by 116κy−524+1718κy−1118=11777κy−47154+1732κy−2380, 1 16 _y- 524+ 1 718 _y- 1118= 1 1777 _y- 47154+ 1 732 _y- 2380, which simplifies to 244κy−5+187κy−11=15434κy−47+16035κy−46. 244 _y-5+ 187 _y-11= 15434 _y-47+ 16035 _y-46. Cross-multiplying yields a cubic polynomial in κy _y with roots κy1=2 _y^1=2, κy2=4/5 _y^2=4/5, and κy3=263/196 _y^3=263/196. Thus, there are three possible utility pairs: 2,4/5,2,263/196,4/5,263/196.\2,4/5\, \2,263/196\, \4/5,263/196\. For each utility pair, we can use the equation ρAI(x,x,y)=α⋅ρu(x,x,y)+(1−α)⋅ρv(x,x,y)ρ^AI(x,\x,y\)=α·ρ^u(x,\x,y\)+(1-α)·ρ^v(x,\x,y\) to recover the implied compliance parameter up to reflection about 1/21/2. This yields the following table: Utility Pair u(y),v(y)\u(y),v(y)\ Implied α Feasibility 2, 4/5\2,\;4/5\ 3/4, 1/4\3/4,\;1/4\ ✓ 2, 263/196\2,\;263/196\ 35/86, 51/86\35/86,\;51/86\ ✓ 4/5, 263/196\4/5,\;263/196\ −35/118, 153/118\-35/118,\;153/118\ × At this stage, there are two feasible candidate utility pairs for y, inducing two feasible values of α up to reflection about 1/21/2. We can eliminate one of them by considering the alternative z or t. Repeating the procedure for the alternative z with T=x,zT=\x,z\ gives 244κz−7+7023κz−47=15434κz−73+16037κz−58, 244 _z-7+ 7023 _z-47= 15434 _z-73+ 16037 _z-58, with roots κz1=4 _z^1=4, κz2=2/5 _z^2=2/5, and κz3=481/244 _z^3=481/244. The implied values of α are: Utility Pair u(z),v(z)\u(z),v(z)\ Implied α Feasibility 4, 2/5\4,\;2/5\ 3/4, 1/4\3/4,\;1/4\ ✓ 4, 481/244\4,\;481/244\ 9/154, 145/154\9/154,\;145/154\ ✓ 2/5, 481/244\2/5,\;481/244\ −3/142, 145/142\-3/142,\;145/142\ × For the alternative t with T=x,tT=\x,t\, the corresponding equation is 6κt−2+3κt−2=16035κt−79+16037κt−65. 6 _t-2+ 3 _t-2= 16035 _t-79+ 16037 _t-65. Cross-multiplying and solving the resulting cubic polynomial gives the candidate roots κt1=5 _t^1=5, κt2=1/5 _t^2=1/5, and κt3=2 _t^3=2. However, κt=2 _t=2 is not a valid solution to the original equation, since it makes the left-hand side undefined. Hence, the admissible roots are κt1=5 _t^1=5 and κt2=1/5 _t^2=1/5, which yields: Utility Pair u(t),v(t)\u(t),v(t)\ Implied α Feasibility 5, 1/5\5,\;1/5\ 3/4, 1/4\3/4,\;1/4\ ✓ The only feasible value of α consistent across all three alternatives is 3/4,1/4\3/4,1/4\. This uniquely recovers u=(1,2,4,5)u=(1,2,4,5) and v=(1,4/5,2/5,1/5)v=(1,4/5,2/5,1/5) up to the label swap and scale normalization. 5 Conclusion This paper considers a delegated choice environment where an AI agent is instructed to act on behalf of a human principal. A central concern in this environment is the potential misalignment between the AI’s and the human principal’s preferences. To study this problem using revealed preference techniques, I introduce the Luce Alignment Model, where the AI agent balances deference to the principal’s preferences against pursuit of its own. The model makes it possible to separately identify two conceptually distinct dimensions of AI behavior: alignment, which captures the similarity between the human’s and the AI’s preferences, and compliance, which captures the extent to which the AI defers to the human principal. I study the identification problem in two settings. In the laboratory setting, where both the AI’s and the human principal’s stochastic choices are observed, I show that violations of the Independence of Irrelevant Alternatives in the AI’s choice data allow the analyst to recover both utility functions and obtain a closed-form expression for the compliance parameter. I also provide an axiomatic characterization of the model in this setting. In the field setting, where only the AI’s choices are observed, a fundamental symmetry prevents an analyst from determining which recovered utility belongs to the human and which to the AI. Nevertheless, I show that when there are at least four alternatives, the underlying distribution over utilities is generically identified up to this label swap, which is sufficient to recover the degree of misalignment. References G. Adomavicius and A. Tuzhilin (2005) Toward the next generation of recommender systems: a survey of the state-of-the-art and possible extensions. IEEE Transactions on Knowledge and Data Engineering 17 (6), p. 734–749. Cited by: §1. A. Allouah, O. Besbes, J. Figueroa, Y. Kanoria, and A. Kumar (2025) What is your AI agent buying? Evaluation, implications, and emerging questions for agentic e-commerce. Columbia Business School Research Paper No. 381574. Cited by: §1. D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mané (2016) Concrete problems in AI safety. arXiv preprint arXiv:1606.06565. Cited by: §1. Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al. (2022) Constitutional AI: harmlessness from AI feedback. arXiv preprint arXiv:2212.08073. Cited by: §1. J. H. Boyd and R. E. Mellman (1980) The effect of fuel economy standards on the U.S. automotive market: an hedonic demand analysis. Transportation Research Part A: General 14 (5–6), p. 367–378. Cited by: §1. N. S. Cardell and F. C. Dunbar (1980) Measuring the societal impacts of automobile downsizing. Transportation Research Part A: General 14 (5–6), p. 423–434. Cited by: §1. C. P. Chambers, T. Cuhadaroglu, and Y. Masatlioglu (2023) Behavioral influence. Journal of the European Economic Association 21 (1), p. 135–166. Cited by: §1. H. Chang, Y. Narita, and K. Saito (2023) Approximating choice data by discrete choice models. arXiv preprint arXiv:2205.01882. Cited by: §1. E. O. Chen, A. Ghersengorin, and S. Petersen (2024) Imperfect recall and AI delegation. Technical report Technical Report 30-2024, Global Priorities Institute, University of Oxford. Cited by: §1. F. Chierichetti, R. Kumar, and A. Tomkins (2018) Learning a mixture of two multinomial logits. In Proceedings of the 35th International Conference on Machine Learning (ICML), p. 961–969. Cited by: §1. P. F. Christiano, J. Leike, T. Brown, M. Marber, B. Shlegeris, and D. Amodei (2017) Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems 30. Cited by: §1. J. T. Fox, K. i. Kim, S. P. Ryan, and P. Bajari (2012) The random coefficients logit model is identified. Journal of Econometrics 166 (2), p. 204–212. Cited by: §1. R. Greenblatt, C. Denison, B. Wright, F. Roger, M. MacDiarmid, S. Marks, J. Treutlein, T. Belonax, J. Chen, D. Duvenaud, A. Khan, J. Michael, S. Mindermann, E. Perez, L. Petrini, J. Uesato, J. Kaplan, B. Shlegeris, S. R. Bowman, and E. Hubinger (2024) Alignment faking in large language models. arXiv preprint arXiv:2412.14093. Cited by: §4. D. Hadfield-Menell, S. J. Russell, P. Abbeel, and A. Dragan (2016) Cooperative inverse reinforcement learning. In Advances in Neural Information Processing Systems, Vol. 29. Cited by: §1. D. Hendrycks, M. Mazeika, and T. Woodside (2023) An overview of catastrophic AI risks. arXiv preprint arXiv:2306.12001. Cited by: §1. N. Immorlica, B. Lucier, and A. Slivkins (2024) Generative AI as economic agents. ACM SIGecom Exchanges 22 (1), p. 93–109. Cited by: §1. J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, J. Zhou, Z. Zhang, et al. (2024) AI alignment: a comprehensive survey. arXiv preprint arXiv:2310.19852. Cited by: §1, §1. J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg (2018) Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871. Cited by: §1, §1. J. Lu and K. Saito (2022) Mixed logit and pure characteristics models. Working paper. Cited by: §1. R. D. Luce (1959) Individual choice behavior: a theoretical analysis. New York: Wiley. Cited by: §1, §3.1. P. Manzini and M. Mariotti (2018) Dual random utility maximisation. Journal of Economic Theory 177, p. 162–182. Cited by: §1. D. McFadden and K. Train (2000) Mixed MNL models for discrete response. Journal of Applied Econometrics 15 (5), p. 447–470. Cited by: §1. E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kadavath, et al. (2023) Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, p. 13387–13434. Cited by: §1. T. Räuker, A. Ho, S. Casper, and D. Hadfield-Menell (2023) Toward transparent AI: a survey on interpreting the inner structures of deep neural networks. In 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), p. 464–483. External Links: Document Cited by: §1. K. Saito (2018) Axiomatizations of the mixed logit model. Technical report California Institute of Technology, Social Science Working Paper 1433. Cited by: §1. W. Tang (2020) Learning an arbitrary mixture of two multinomial logits. arXiv preprint arXiv:2007.00204. Cited by: §1. X. Zhang, X. Zhang, P. Loh, and Y. Liang (2022) On the identifiability of mixtures of ranking models. arXiv preprint arXiv:2201.13132. Cited by: §1.