Paper deep dive
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
Travis LaCroix
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/26/2026, 6:47:43 PM
Summary
The paper proposes a structural reconceptualization of the AI value alignment problem, moving it from a purely technical or normative challenge to a problem of governance. Drawing on the economic principal-agent framework, the author argues that misalignment arises along three interacting axes: objectives (mis-specified objective functions), information (asymmetries in transparency and verifiability), and principals (the plurality and diversity of human actors like developers, users, and stakeholders). The author introduces the 'scaling hypothesis for value-aligned AI,' suggesting that increasing model generality and stakeholder diversity systematically amplifies these structural misalignments, necessitating ongoing institutional and socio-political management rather than just technical solutions.
Entities (10)
Relation Signals (7)
Travis LaCroix â affiliatedwith â Durham University
confidence 100% · Travis LaCroix, Durham University, Durham, England, United Kingdom
Value Alignment Problem â definedby â Principal-Agent Framework
confidence 100% · Drawing on the principal-agent framework from economics, this paper reconceptualises misalignment
Human Principal â delegatesto â AI Agent
confidence 100% · the delegation of tasks from one actor (a human principal) to another (an AI agent)
Value Alignment Problem â hasaxis â Information Axis
confidence 100% · misalignment can occur along each axis... (1) objectives, (2) information, and (3) principals
Value Alignment Problem â hasaxis â Objectives Axis
confidence 100% · misalignment can occur along each axis... (1) objectives, (2) information, and (3) principals
Value Alignment Problem â hasaxis â Principals Axis
confidence 100% · misalignment can occur along each axis... (1) objectives, (2) information, and (3) principals
Scaling Hypothesis for Value-Aligned AI â describeseffecton â Value Alignment Problem
confidence 90% · increasing model generality, deployment scope, and stakeholder diversity systematically amplifies informational asymmetries, value conflicts, and power imbalancesâa fact which I call the 'scaling hypothesis for value-aligned AI'
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The value alignment problem for artificial intelligence (AI) is often framed as a purely technical or normative challenge, sometimes focused on hypothetical future systems. I argue that the problem is better understood as a structural question about governance: not whether an AI system is aligned in the abstract, but whether it is aligned enough, for whom, and at what cost. Drawing on the principal-agent framework from economics, this paper reconceptualises misalignment as arising along three interacting axes: objectives, information, and principals. The three-axis framework provides a systematic way of diagnosing why misalignment arises in real-world systems and clarifies that alignment cannot be treated as a single technical property of models but an outcome shaped by how objectives are specified, how information is distributed, and whose interests count in practice. The core contribution of this paper is to show that the three-axis decomposition implies that alignment is fundamentally a problem of governance rather than engineering alone. From this perspective, alignment is inherently pluralistic and context-dependent, and resolving misalignment involves trade-offs among competing values. Because misalignment can occur along each axis -- and affect stakeholders differently -- the structural description shows that alignment cannot be "solved" through technical design alone, but must be managed through ongoing institutional processes that determine how objectives are set, how systems are evaluated, and how affected communities can contest or reshape those decisions.
Tags
Links
- Source: https://arxiv.org/abs/2604.20805v1
- Canonical: https://arxiv.org/abs/2604.20805v1
Trouble viewing inline? Open PDF directly â
Full Text
82,338 characters extracted from source content.
Expand or collapse full text
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem TRAVIS LACROIX, Durham University, United Kingdom The value alignment problem for artificial intelligence (AI) is often framed as a purely technical or normative challenge, sometimes focused on hypothetical future systems. I argue that the problem is better understood as a structural question about governance: not whether an AI system is aligned in the abstract, but whether it is aligned enough, for whom, and at what cost. Drawing on the principal-agent framework from economics, this paper reconceptualises misalignment as arising along three interacting axes: objectives, information, and principals. The three-axis framework provides a systematic way of diagnosing why misalignment arises in real-world systems and clarifies that alignment cannot be treated as a single technical property of models but an outcome shaped by how objectives are specified, how information is distributed, and whose interests count in practice. The core contribution of this paper is to show that the three-axis decomposition implies that alignment is fundamentally a problem of governance rather than engineering alone. From this perspective, alignment is inherently pluralistic and context-dependent, and resolving misalignment involves trade-offs among competing values. Because misalignment can occur along each axisâand affect stakeholders differentlyâthe structural description shows that alignment cannot be âsolvedâ through technical design alone, but must be managed through ongoing institutional processes that determine how objectives are set, how systems are evaluated, and how affected communities can contest or reshape those decisions. CCS Concepts:âą Computing methodologiesâPhilosophical/theoretical foundations of artificial intelligence; Machine learning; Distributed artificial intelligence;âą Applied computingâSociology; Economics; Psychology;âą Social and professional topics; Additional Key Words and Phrases: the value alignment problem, structural value alignment problem, principal-agent frame- work, axes of value alignment, misaligned objectives, asymmetric information, relative principles, shareholders, stakeholders, multi-agent coordination problems, power dynamics, scaling hypothesis for value-aligned AI, governance ACM Reference Format: Travis LaCroix. 2026. Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem. In The 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT â26), June 25â28, 2026, Montreal, QC, Canada. ACM, New York, NY, USA, 18 pages. https://doi.org/10.1145/3805689.3812420 1 Introduction The value alignment problem for artificial intelligence (sometimes called âAI alignmentâ or âagent alignmentâ) is commonly described as the challenge of ensuring that AI systems act in accordance with human intentions or values [11,26,67]. However, this standard description underdetermines the problem since it does not specify what or whose values are (or ought to be) considered the proper targets of alignment. To ask whether a system is aligned requires understanding for whom it is aligned, to what degree, and according to what standard. This highlights that alignment is not merely a technical problem, to be solved through engineering or normative encoding, nor is it a philosophical or sociological problem of determining the âcorrectâ values; instead, it is a Authorâs Contact Information: Travis LaCroix, Durham University, Durham, England, United Kingdom, travis.lacroix@durham.ac.uk. This work is licensed under a Creative Commons Attribution 4.0 International License. FAccT â26, Montreal, QC, Canada © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2596-8/2026/06 https://doi.org/10.1145/3805689.3812420 arXiv:2604.20805v1 [cs.CY] 22 Apr 2026 FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix relational and evaluative judgment that depends on whose values are prioritised, how trade-offs are made, and how degrees of alignment are measured. In practice, contemporary alignment research aims to address a wide range of concrete problems in existing machine learning (ML) systems via reinforcement learning from human feedback (RLHF) and robustness, inter- pretability, and evaluation methods for deployed models. At the same time, however, discussions of alignment have historically emerged out of theoretical concerns about highly capable, hypothetical future systems such as artificial general intelligence (AGI) or superintelligence. 1 While these debates have clarified important conceptual issues, their intellectual legacy can sometimes obscure the fact that problems of misalignment already arise in present-day systems and must be addressed within existing institutional and governance structures. Similarly, discussions of alignment are typically abstract, meaning they do not provide guidance for practical solutions grounded in the real-world functioning of these systems. Although some strands of alignment research are prac- tical and engineering-focused, the normative, governance, and institutional dimensions of alignmentâe.g., how competing stakeholders define alignment objectives, how evaluation standards are set, and how accountability mechanisms operateâremain comparatively under-theorised. Recent scholarship has sought to refine the standard definition by distinguishing between its normative and technical components [26], differentiating inner and outer forms of alignment [44], or separating forward- and backward-facing approaches to addressing misalignment [46]. While such distinctions add conceptual precision, they do not in themselves yield practical methods for mitigating misalignment. For example, identifying a ânorma- tiveâ dimension of the problem acknowledges that questions of value are central; however, this acknowledgement itself offers no framework for deciding which values ought to guide AI systems or why they should be prioritised. Likewise, distinguishing between forms of misalignment is analytically useful, but it does not explain how the underlying normative issues should be addressed. This conceptual ambiguity motivates the present analysis: rather than proposing new technical alignment methods, this paper seeks to clarify the structural conditions under which instances of the alignment problem can arise across socio-technical contexts. 2 What is needed is clarity on value alignment as a class of problems, rather than a single, unified challenge. Instead of asking which values are the âcorrectâ objects of alignment and how those values can be encoded, it is more productive to ask under what conditions value misalignment tends to arise generally. That is to say, what is the structure of the value alignment problem, and in what contexts is it truly a problem? Reframing the question in this way reveals that the challenge of aligning values is not unique to AI, but rather a special case of more general coordination problems among agents with differing goals and information. In light of this challenge, I present a structural definition of the value alignment problem for AI, grounded in the principal-agent framework from economics. According to this model, value misalignment emerges in human-AI interactions whenever (a) the objectives of the AI system are mis-specified, or (b) informational asymmetries exist between (human) principals and (artificial) agents. Hence, misalignment can occur along three orthogonal, but interacting, axes: (1) objectives, (2) information, and (3) principals. These three axes provide a diagnostic framework for analysing why alignment failures occur in practice. Misalignment along the objectives axis arises when proxy metrics fail to capture intended goals. Misalignment along the information axis can emerge when 1 For example, Gordon[31]suggests that the alignment problem ârefers to the challenge of ensuring that AI systems behave as their creators intend, especially when the systems become more intelligent and capable than their human designers . . . [because as] AI systems become more intelligent, they may develop goals and values that differ from those of their creators, which could lead to unexpected and potentially harmful outcomesâ (p. 77). See also, [4,5,14,16,25,41,64]. Eckersley[19]highlights that most âconcerns in the literature about the difficulty of aligning hypothetical future AGI systems to human values are motivated by the risk of âinstrumental convergenceâ of those systemsâ (p. 10). See, e.g., [4, 8, 9, 40, 59, 60, 73, 75]. A comprehensive survey is provided in Ji et al. [46]. 2 Importantly, this perspective complements rather than replaces existing technical work. Techniques such as RLHF, interpretability tools, dataset governance practices, and benchmarking frameworks represent concrete efforts to mitigate misalignment in deployed systems. The claim here is that understanding how these techniques function within broader socio-technical environments requires conceptual frameworks that explicitly account for pluralistic values, informational asymmetries, and institutional power relations. Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada opacity, distributional shifts, or oversight limitations prevent principals from understanding system behaviour. Misalignment along the principals axis reflects the fact that multiple human actorsâdevelopers, users, regulators, and affected stakeholdersâoften hold divergent or conflicting values. Taken together, these axes reveal that alignment outcomes depend on how objectives are specified, how information flows between actors, and how conflicts among stakeholders are negotiated. This insight forms the basis of this paperâs central claim: that value alignment ultimately raises questions of governance rather than engineering alone. Rather than a technical or normative problem, value alignment is primarily socio-political. The remainder of the paper develops this argument in stages. In Section 2, the principal-agent framework is used to articulate the structural definition of the value alignment problem. Section 3 further specifies the three axes of value alignment, and highlights how dynamic interactions between these axes make value alignment a fundamentally multi-dimensional socio-technical problem rather than a single technical challenge. Section 4 demonstrates how, on the structural description of value alignment, increasing model generality, deployment scope, and stakeholder diversity systematically amplifies informational asymmetries, value conflicts, and power imbalancesâa fact which I call the âscaling hypothesis for value-aligned AIâ. Section 5 then draws the broader im- plication: because alignment failures arise from structural features of objective proxies, information asymmetries, and multi-agent value pluralism, alignment must be understood as an ongoing problem of governance. Rather than asking whether AI systems can be perfectly aligned, the more appropriate question is how institutions can manage and negotiate misalignment over timeâdetermining, in practice, what counts as being aligned enough, for whom, and at what cost. By situating alignment within a socio-technical framework, this reconceptualisation foregrounds the importance of factors such as power asymmetries, labour relations, and environmental impacts. Ultimately, the structural reconceptualisation reframes value alignment not as a purely technical or normative challenge but as a class of socio-dynamic problems indexed to contexts of delegation, information, and pluralism. In this sense, this paper contributes primarily at the level of conceptual infrastructure by providing a framework that helps organise and assess existing work on alignment while drawing attention to governance and institutional questions that remain insufficiently integrated into the broader alignment discourse. 2 The Value Alignment Problem for Artificial Intelligence A standard instance of the value alignment problem for AI arises when a modelâs objective functionâi.e., the mathematical specification of what the system is meant to optimise or achieveâis poorly or incompletely defined. In such cases, the âobjectivesâ (âvaluesâ, âgoalsâ, âincentivesâ, etc.) of an ML model are misaligned with the true objectives for the systemâwhat we might colloquially call âour valuesâ. Such misalignment is not merely a theoretical concern about hypothetical systems; it already appears in present-day applicationsâe.g., when optimisation metrics fail to capture broader social goals or when deployed systems behave in ways that were not anticipated by their designers. In economics, law, and political science, this situation is more commonly referred to as a principal-agent problem (also known as an âagency dilemmaâ or an âincentive problemâ). In these contexts, a principal delegates authority to an agent to act on her behalf, yet the agentâs actions may not always reflect the principalâs best interests. This framework has become increasingly relevant in the AI literature, where scholars have drawn explicit parallels between value alignment and the principal-agent problem [34,49]. Recent work emphasises the structural similarities between these domains, framing value alignment as a multi-agent coordination problem [24]. These connections are reflected in a growing body of practical alignment research that uses economic and game-theoretic tools to design training procedures, feedback mechanisms, and oversight systems for real-world ML models. For example, this economic framework underpins contemporary approaches like cooperative inverse FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix reinforcement learning (CIRL) [35], the incomplete contracting framework for AI design [37], and game-theoretic methods [21, 28, 36, 67, 68] that explicitly leverage informational asymmetries to improve alignment. 3 2.1 The Principal-Agent Framework from Economics The principal-agent framework formalises a class of problems that can arise when one actor (the principal) delegates decision-making authority to another (the agent). The central concern is that the agent, possessing their own desires, incentives, objectives, etc., may act in ways that diverge from the principalâs objectives. This problem is pervasive in human-human relationsâfor instance, between employers and employees, shareholders and managers, or citizens and their elected representatives. 4 Hence, principal-agent models examine the problem of creating the correct incentives in a non-cooperative setting with asymmetric information, ensuring that the agent acts in the principalâs best interest. One key feature of this framework is that informational asymmetries are fundamental to misalignment in economic contexts. Three key types of asymmetries are identified: (1)Non-verifiability arises in cases where, even if actions and outcomes are observable, they may not be verifiable by third parties (e.g., courts or regulators). (2) Moral hazard (hidden action) arises when the principal cannot observe whether the agentâs actions faithfully serve her interests. (3)Adverse selection (hidden information) arises when the principal cannot fully assess the agentâs abilities or characteristics before contracting. When information between the principal and the agent is symmetric (in the idealised economic model), the principal can theoretically design a contract that induces the agent to act in her best interest, even when their goals diverge. Conversely, even when the principal and the agent have identical incentives, informational asymmetries can cause the agent to (inadvertently) act in a way that the principal does not want. Thus, in the economic context, competing incentives (or âmisaligned valuesâ) are neither necessary nor sufficient to generate a principal- agent problem; rather, such problems emerge from the interplay between incentive structures and information asymmetries. 2.2 Extending the Framework to Artificial Agents Viewed structurally, the principal-agent problem represents a class of problem instances that can arise either because of conflicting incentives (misaligned values) or because of asymmetric information between two (human) actors: a delegating principal and an acting agent. This framing can be productively extended to human-AI interactions to describe the value alignment problem for AI. In this context, a human agent is analogous to the principal, where a principal might be conceived of as the user, system designer, or company on whose behalf the agent acts (shareholders in the systems), or the principal may be understood as an individual affected by an AI system (stakeholders). The AI system itself is analogous to the agent. This insight leads to the following structural definition of the value alignment problem based on the principal- agent framework. The Value Alignment Problem (Structural Definition) A problem that arises from the dynamics of multi-agent interactions involving the delegation of tasks from one actor (a human principal) to another (an AI agent). This problem can arise whenever 3 These approaches illustrate that alignment research today operates at the intersection of theory and practice, combining formal models with empirical experimentation on deployed systems. The contribution of the present paper is therefore not to replace these approaches but to situate them within a broader structural account of alignment that highlights how objective specification, information asymmetries, and pluralistic stakeholders jointly shape alignment outcomes in real socio-technical systems. 4 See discussion in Eisenhardt [20], Jensen and Meckling [45], Kerr [47], Laffont and Martimort [51]. Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada (í) The agentâs objective function is misaligned with the true objective of the principal(s); or, (í) There are informational asymmetries between the principal and the agent. Instead of focusing on the normative component of the value alignment problem (âWhat are the correct values to encode in AI systems?â) or the technical component (âHow do we encode said values?â), this description emphasises the contexts in which misalignment can ariseâi.e., the socio-dynamic structure of value alignment. Within this framework, the value alignment problem for AI can be understood as varying along three orthogonal but interacting axes: (1)The objectives axis, which describes the extent to which the systemâs formal objectives (i.e., objective functions) accurately capture the principalâs true goals. (2) The information axis, which describes the degree of transparency, observability, and verifiability between the system and its human principals. (3)The principals axis, which describes the plurality and diversity of relevant human actors, encompassing both shareholders (developers, companies) and stakeholders (affected individuals or groups). These axes are orthogonal to the extent that a reduction of misalignment along one axis does not guarantee a reduction of misalignment along the other axes; misalignment can arise independently along any one of them. For instance, an AI system may have a perfectly specified objective function (objectives axis) but still fail to be aligned because its outputs cannot be effectively monitored or verified (information axis). Likewise, a system may exhibit near-perfect performance and transparency yet remain misaligned because it optimises for the goals of one group of principals at the expense of others (principals axis). Each axis, therefore, isolates a distinct mechanism through which misalignment can occur. However, there are interaction effects between them, meaning that misalignment along one axis can exacerbate misalignment along another. For example, in predictive policing systems, the use of biased historical arrest data (a proxy problem along the objectives axis) both embeds and conceals underlying informational asymmetries between developers, data subjects, and law enforcement institutions. This feedback loop also redefines who counts as a âprincipalâ in practiceâsince the communities most affected (stakeholders) have the least access to correcting the modelâs behaviour. In this sense, orthogonality describes analytical separability, not empirical independence: the three axes co-produce each other in real-world deployments. The principals axis is particularly relevant for work in pluralistic alignment, where insights from social choice theory and deliberative democratic models are increasingly used to formalise how diverse human values might be aggregated or negotiated [12,71]. This plurality of principals reveals that value alignment in AI is inherently multi-agent in nature. It is not merely about aligning âAI valuesâ with âhuman valuesâ in the abstract, but about reconciling diverse and sometimes conflicting human interests embedded in socio-technical systems. That said, a key difference between the principal-agent framework (for human-human interactions) and its analogue for human-AI interactions is that valuesâand therefore misalignment of valuesâare inherent to the human-human interaction insofar as principals and agents have inherent values. In contrast, in the human-AI interaction, there is no inherent misalignment between the (human) principal and the (artificial) agent because the agent has no inherent values: the âvaluesâ (i.e., objective functions) of the agent are programmed. In this latter case, competing incentives (value misalignment) can arise because it is impossible to specify an objective function completely and correctly; hence, whereas misaligned incentives are neither necessary nor sufficient for generating a principal-agent problem in economic contexts (the case of human-human interactions), they can do so in AI contexts (the case of human-AI interactions). The structural definition of the value alignment problem captures everything that the standard definition is intended to capture, but it provides additional clarity in thinking about how these problems arise. The issue is not simply that goals, incentives, values, etc., are misaligned, because even when objective functions perfectly FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix represent our true goals, informational asymmetries between the principal and the agent can still give rise to an instance of the value alignment problem. 3 The Three Axes of Value Alignment Having defined value alignment structurally, we can now examine the three axes along which instances of misalignment arise: the objectives axis, the information axis, and the principals axis. Each axis represents a distinct yet interrelated source of divergence between the intentions of human principals and the behaviour of artificial agents. Understanding these axes in detail clarifies why alignment cannot be reduced to a single technical or normative question; instead, it must be treated as a multi-dimensional problem emerging from the socio-technical structure of human-AI interaction. The Objectives Axis. The first axis concerns misspecified objective functions, which occur when the formal optimisation target of an AI system diverges from the true objective for the system. This is analogous to conflicting incentives in the principal-agent framework and corresponds roughly to the âouter alignmentâ problem identified by Hubinger et al. [44]. In this sense, the objectives axis captures what most people intuitively mean by value misalignmentâe.g., reward hacking, negative side effects, or perverse incentives [1]. Thus, poorly designed objective functions can lead to instances of the value alignment problem whenever there is a conflict between the objective function and the true objective (i.e., the âintentionsâ or âvaluesâ specified by the standard definition of the value alignment problem). In a standard machine-learning model, the objective function defines the optimisation landscape over which the model learns. However, the solution space for an optimisation problem defined by an objective function includes a âpathologyâ of local optima, meaning these landscapes are frequently deceptive [30,52,58,69]. As a result, the objective function âdoes not necessarily reward the stepping stones in the search space that ultimately lead to the objectiveâ [52, p. 329]. Consequently, many objective functions are constructed ad hoc, privileging what is easy to measure rather than what genuinely captures human goals. The underlying reason that poorly specified objectives yield misalignment is that objective functions are mere proxies for the true objective. From the AI systemâs âpoint of viewâ, however, the proxy is the objectiveâa model cannot distinguish between the proxy and the task it approximates. Moreover, proxies permeate every aspect of an ML system along the algorithmic development pipelineâe.g., at the problem specification stage, a task is a proxy for a real-world problem; during the model design stage, an objective function is a proxy for the true objective; during the testing/validation/deployment stages, training data, test data, and benchmarking datasets or metrics are proxies for real-world distributions. Because each layer of the pipeline replaces the complex, value-laden world with simplified formal representations, alignment depends critically on how well these proxies track true objectives. To give a concrete example of value misalignment along the objectives axis, consider again the case of predictive policing. 5 Charitably, the intended goal of such systems is to reduce crime by allocating police resources efficiently. However, the objective function operationalised in these models typically seeks to predict where crime is most likely to occurâa proxy for the (purported) true objective. Because actual future crime is unobservable and, therefore, impossible to optimise directly, the model relies on historical arrest data as a proxy. However, arrest data is a more accurate measure of police activity than crime incidence. As such, these data are heavily influenced by historical policing practices. Areas that have been over-policed in the past tend to generate more arrests, which in turn train subsequent models to allocate even more police resources to those areas. This feedback loop ensures that over-policed neighbourhoods continue to appear âhigh-riskâ, reinforcing systemic bias. For this reason, Benjamin[6]calls predictive policing a âcrime production algorithmâ (83), which contradicts the purported aim of the modelâi.e, to reduce crime. Value alignment along the objectives axis requires that proxies 5 See further discussion in Benjamin [6], Broussard [10], Ensign et al. [22], Lum and Isaac [55], OâNeil [61]. Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada be sufficiently faithful representations of true objectives, which is often difficult or impossible in the case of complex socio-cultural tasks. The Information Axis. While misaligned objectives exacerbate these issues, they often interact withâand are compounded byâinformational asymmetries between humans and AI systems, which constitute the second axis of the alignment problem. This axis corresponds to the âinner alignmentâ problem [44], which encompasses issues such as scalable oversight, safe exploration, and robustness to distributional shifts [1]. Closely related ideas have also been explored in the literature on the off-switch game [36], CIRL [35], and (partially observable) assistance games [21,68]. 6 In this latter framework, alignment is modelled as a cooperative game between a human and an AI system under conditions of partial observability, where the agent must infer the humanâs preferences from limited signals and where each party may possess different information about the environment. This literature explicitly highlights how asymmetric or incomplete information between humans and AI systems can generate alignment challenges, thereby illustrating a formal treatment of the informational asymmetries captured by the information axis described here. Informational asymmetries occur when the principal lacks sufficient knowledge to assess, verify, or predict the agentâs behaviour. By analogy with the principal-agent problem, three primary forms of informational asymmetries show the circumstances under which an instance of the value alignment problem can arise along the information axis: (1)Non-verifiability results from the principal (or an external third party) being unable to verify whether the agentâs actions or outputs genuinely satisfy the intended objective. (2) Moral Hazard arises when the agent has access to information unavailable to the principal, or its internal operations are unobservable, enabling behaviour that diverges from the principalâs goals. (3)Adverse Selection occurs when the principal has incomplete or misleading information about the agentâs capabilities or internal characteristics before delegating a task. In the context of incomplete contracting, non-verifiability refers to situations where certain aspects of the contractual agreement are difficult or impossible to monitor or verify. In the case of the value alignment problem for AI, this type of informational asymmetry captures questions surrounding scalable oversight, underscoring the difficulty of effectively monitoring and controlling increasingly complex and widespread AI systems as they scale in size and scope. Moral hazard (endogenous to the principal-agent relationship) arises when the systemâs internal operations or learning processes cannot be fully observed. For example, safe exploration problems occur when reinforcement- learning agents must gather information through trial and error, potentially taking unsafe or undesirable actions during training or deployment. As models grow in scale and parameter complexity, their optimisation landscapes expand correspondingly, creating vast solution spaces with numerous local optima. A model might perform well on its training objective yet behave unpredictably when confronted with out-of-distribution data. In such cases, misalignment results not from a flawed objective, per se, but from the principalâs inability to observe or interpret the agentâs internal state. The POAG framework formalises such scenarios as ones in which the agent must act under uncertainty about the humanâs latent reward function and the true state of the environment, reinforcing the role of partial observability and hidden information in alignment failures. Adverse selection (exogenous to the principal-agent relationship) has an AI analogue in the black-box nature of state of the art modelsâi.e., when the internal workings of a model are not easily interpretable or transparent to humans. Deep neural networks, for instance, encode relationships among millions or billions of parameters across multiple layers. These architectures learn high-dimensional, hierarchical representations that are largely 6 Garber et al. [28]suggest that assistance games are (partially-observable) generalisations of CIRL, which is the underlying framework for the off-switch game. FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix inscrutable to human observers. Because it is difficult to trace how specific input features affect outputs, developers and users effectively contract with agents whose internal processes they do not understand. This opacity is compounded by the use of scraped or proprietary datasets, where the provenance, quality, and bias of data remain uncertain. As datasets scale, ensuring their fairness and ethical integrity becomes increasingly infeasibleâagain widening the information gap between human principals and machine agents. Hence, even when objectives are perfectly specifiedâi.e., there is sufficient alignment along the objectives axisâinformational asymmetries can independently generate alignment failures. Highly complex systems may reach states of inherent opacity, where no feasible method of verification exists. This problem is especially acute in normative contexts, where there is no single, unambiguous âcorrectâ answer to verify against [50]. Consequently, the information axis reveals that alignment failures can persist even under idealised objective specification, which is often the target of purely technical approaches to alignment. The Principals Axis. The third axis foregrounds the question: with whose values are AI systems meant to align? Standard formulations of the value alignment problem assume that alignment should target âthe values of humanityâ, as though a unified set of human values could be identified and encoded. However, this approach is problematic on both conceptual and practical grounds. As the number of principals whose values are under consideration increases, the intersection of their preferences rapidly diminishes. If we interpret âhuman valuesâ as the conjunction of specific individual preferences, the resulting set is effectively empty since people value conflicting things. This suggests that to apply universally, any purported âhuman valuesâ must be so abstractâe.g., justice, fairness, welfareâthat they become too coarse-grained to serve as operational objectives for AI systems. Thus, aligning AI systems with âthe values of humanityâ is impossible at any meaningful level of granularity. Recent work in AI alignment and human-AI interaction has begun to confront this difficulty directly by developing methods that attempt to represent plural and even conflicting human perspectives rather than collapsing them into a single aggregate signal. 7 For example, jury learning approaches explicitly model multiple evaluators whose judgments are preserved as distinct inputs rather than averaged away [32]. Deliberative systems, such as the âHabermas Machineâ, seek to approximate structured democratic deliberation in which competing viewpoints are surfaced, debated, and revised Tessler et al. [74]. Related pluralistic approachesâincluding frameworks inspired by Overton-style pluralism, represent value landscapes as sets of permissible viewpoints rather than a single optimum aim to preserve dissent and normative diversity within AI-assisted decision processes [63, 72]. From the structural perspective developed here, the principalâs identity is not fixed but context-dependent. In human-AI interactions, principals can be divided broadly into two overlapping categories: (1)The set of shareholders, which includes those directly involved in creating, deploying, or profiting from AI systemsâe.g., companies, research labs, developers, data scientists, system administrators, regulators, compliance officers, end users, consumers, etc. (2)The set of stakeholders, which includes those directly or indirectly affected by these systemsâ creation and deploymentâe.g., individuals, communities, or the broader public. Although these groups may often overlap, the most consequential instances of misalignment will arise when stakeholders are distinct from shareholdersâfor instance, when an algorithm optimises for corporate profit at the expense of public welfare. Each group embodies different values, incentives, and priorities, shaping both how objectives should be specified and how success should be evaluated. Hence, alignment must be understood as a process of negotiation among plural principals, whose values may diverge or even conflict. This recognition underscores why the value alignment problem is fundamentally socio-technical: resolving it requires not just 7 See, e.g., Bergman et al. [7], Davani et al. [15], Gordon et al. [32], Huang et al. [43], Kirk et al. [48], Peterson et al. [62], Sorensen et al. [71, 72], Tessler et al. [74] and discussion in Fazelpour and Fleisher [23]. Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada better algorithms, but better mechanisms for adjudicating competing interests, distributing authority, and ensuring accountability in the design and deployment of AI systems. 3.1 Dynamics and Interactions Among the Axes of Value Alignment Having identified the three orthogonal axes that structure the value alignment problemâobjectives, information, and principalsâit is important to consider how these dimensions interact dynamically. While each axis can independently give rise to instances of misalignment, in real-world systems these axes rarely operate in isolation. Instead, they form a system of dependencies, feedback loops, and amplification effects that jointly determine the stability and direction of alignment over time. As mentioned, a reduction of misalignment along one axis does not necessarily entail a reduction along the others. For example, improving the specification of an objective function (reducing misalignment along the objectives axis) may require complex optimisation processes or large-scale data collection that increase informational asymmetries between the principal and the agent. Similarly, expanding the set of principals to include more stakeholders (improving pluralistic alignment) can complicate the definition of an objective function, thereby worsening misalignment along the objectives axis. These trade-offs suggest that alignment is not a static property but a dynamic equilibrium maintained within a multi-dimensional constraint space. Misalignment along one axis can also generate feedback loops that exacerbate problems elsewhere. Consider again the case of predictive policing: an initial bias in data collection (information axis) leads to distorted objective functions (objectives axis), which in turn reinforce the biases affecting future data (further informational misalignment). Over time, this cycle compounds misalignment across axes, transforming what begins as a technical issue into a systemic one with significant social and ethical consequences. In this way, alignment failures can propagate through systems much like contagions propagate through networks. Addressing alignment, therefore, requires not only correcting individual failures but also understanding the causal structure and feedback dynamics among the three axes. Viewing the value alignment problem through this systems lens highlights the need for dynamic governance mechanisms that can adapt as models, data, and social contexts evolve. Static alignment solutionsâthose that assume a fixed objective, a stable information environment, or a uniform principalâwill eventually fail as systems drift from their design conditions. Consequently, the study of alignment must extend beyond isolated technical interventions toward the design of institutions, protocols, and oversight structures capable of sustaining alignment over time. In light of these considerations, we turn now to a higher-order feature of the alignment problem: the plurality of human principals whose values and interests define what âalignmentâ even means. 4 Alignment with Pluralistic Values A fundamental challenge emerges from the structural account of value alignment: namely, that alignment cannot be conceived monolithically. Current technical approaches to alignmentâmost notably, RLHFâare primarily designed to align models with the average preferences of humans. While this strategy is often framed as âdemocraticâ, it effectively erases diversity by smoothing over legitimate differences among individual and group values [3,70,71]. This observation should be understood as a diagnostic claim about prevailing training and deployment incentives rather than a dismissal of emerging pluralistic alignment research. In principle, feedback aggregation methods can be designed to preserve disagreement, represent minority viewpoints, or enable deliberative synthesis among competing perspectives. In practice, however, standard deployment pipelinesâ especially those optimised for scalability, consistency, and product usabilityâtend to reward the production of a single coherent response distribution, which implicitly encourages the smoothing or aggregation of divergent human judgments unless explicit institutional commitments are made to preserve pluralism. Pluralistic alignment, by contrast, seeks to computationally model value pluralism itself. This approach recognises that as AI systems FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix (especially large language models) are increasingly marketed as âgeneral-purposeâ technologies for use across diverse domains by heterogeneous users, the relevant value landscape becomes more fragmented and contested. The challenge is not to identify a single set of âhuman valuesâ, but to navigate among many partially overlapping and sometimes conflicting sets of values held by different principals. The structural definition brings to light how difficult this challenge is, and how contemporary machine-learning approaches exacerbate these problems. 4.1 The Scaling Hypothesis for Value-Aligned AI The structural definition of the value alignment problem clarifies that every instance of misalignment is indexed to a particular principal or set of principals. Some principalsâ goals (for example, those of developers, regulators, and consumers) may converge. Still, in many cases, they diverge in ways that are decisive for assessing whether a system is genuinely aligned. A model may therefore be well aligned relative to one principal while simultaneously misaligned relative to another. Alignment, in this sense, is always relative and always partialâa matter of degree rather than an absolute state. The structural definition underscores that talk of âsolvingâ the value alignment problem involves a category mistake: the value alignment problem is not a problem per se, but a class of problems that can be more or less instantiated. From this perspective, the notion of alignment simpliciterâi.e., complete alignment across all axes and for all relevant principalsâis incoherent for sufficiently complex systems. The structural account also motivates the following scaling hypothesis in the context of value alignment. Scaling Hypothesis for Value-Aligned AI. As AI systems increase in (i) model generality, (i) deployment scope, and (i) stakeholder diversity, the difficulty of achieving value alignment increases because these forms of scaling systematically amplify informational asymmetries between humans and AI systems and pluralistic conflicts among the principals whose values the system must serve. Importantly, this hypothesis should not be interpreted as a formal theorem about ML systems, but as a structural claim about how the three axes of alignment interact as AI systems expand in capability and deployment. More precisely, scaling tends to occur simultaneously along three dimensions: (i) model generality, as systems move from narrow task-specific tools toward general-purpose capabilities; (i) deployment scope, as systems are integrated into a wider range of social and economic contexts; and (i) stakeholder diversity, as the set of individuals and institutions affected by the system expands. Each of these dimensions enlarges the alignment problem in a distinct way. Greater model generality increases the number of contexts in which behaviour must remain aligned, often requiring deployment under conditions that differ substantially from training data or the target objectives. Broader deployment scope introduces new environments, regulatory regimes, and norms that may not have been anticipated during development. Expanding stakeholder diversity increases the likelihood that the relevant principals hold incompatible values or priorities. These dynamics imply a structural scaling effect: as models become more general-purpose and widely deployed, both informational uncertainty and normative disagreement grow. Informational asymmetries increase because larger models trained on vast datasets and complex architectures become harder for humans to interpret, audit, or verify. At the same time, pluralistic conflict intensifies because a single system must simultaneously serve users with heterogeneous interests, cultural norms, and risk tolerances. Alignment difficulty therefore grows not merely because models become technically more capable, but because the social and epistemic environment in which they operate becomes more complex. Empirical patterns in ML development lend plausibility to this hypothesis. Large foundation models are trained on heterogeneous web-scale datasets and subsequently deployed across domains ranging from education and healthcare to journalism and governance. Each new domain introduces distinct objectives, risk profiles, and stakeholder groups. As a result, alignment failures in such systems often manifest not as simple objective misspecification but as context-sensitive conflicts among legitimate values. In contrast, highly successful systems such as domain-specific scientific models illustrate the opposite pattern: Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada alignment is comparatively tractable when objectives are tightly specified and the set of relevant principals is limited. For example, Andrews[2]suggests that DeepMindâs AlphaFold and AlphaFold 2.0 are among the most impressive results that ML methods have achieved for science, not the least because the protein-folding problem in structural biology was largely considered intractable. From the perspective of structural alignment, part of the incomparable success of AlphaFold can be attributed to the fact that the model is not domain-generic. Namely, those systems that lend themselves to alignment simpliciter are systems with narrow targets (limiting direct stakeholders) that model readily formalisable objectives. The scaling hypothesis for value-aligned AI therefore predicts a gradient of alignment difficulty: the closer a system approaches general-purpose functionality and widespread deployment, the more the alignment problem shifts from a primarily technical challenge to a socio-technical one involving informational opacity and pluralistic value conflict. 4.2 Pluralism, Power, and the Socio-Technical Context Crucially, AI models do not operate in isolation. They are embedded in wider socio-technical systems that encompass diverse sets of principalsâshareholders and stakeholders alike. McQuillan[56]highlights that AI is more than a set of ML methods: It is impossible to separate the technical aspects of AI from the social contexts in which these models are created, trained, tested, and deployed. Present-day generative models, for example, have been trained on vast datasets that include copyrighted material, personal data, and cultural artefactsâoften without consent or compensation. Hence, these models involve an unethical kind of labour theft [29]. Privacy and consent violations in datasets often adversely affect individuals in marginalised communities, as attempts to âdiversifyâ datasets to be more representative can incur costs to those groups concerning privacy, exploitation, monitoring, and other issues [38,42,66]. These harms are compounded by the material and environmental costs of AI infrastructureâfrom extractive labour and resource mining [13] to the significant carbon footprint of large-scale training [53,54], which disproportionately impacts the âoften lower income and thus most neglected humans in societyâ [65]. Hence, alignment demands more than technical fixes; it requires an examination of the power relations that determine whose values count in practice. These socio-technical dynamics provide further motivation for the scaling hypothesis described above. As AI systems scale in capability and deployment, the set of affected stakeholders expands far beyond the immediate developers and users. Systems trained and deployed by a small number of organisations may influence millions of individuals across jurisdictions, cultures, and regulatory environments. The resulting expansion of the principal set intensifies pluralistic disagreement about acceptable uses, risks, and benefits. In other words, scaling AI systems simultaneously scales the political and ethical space in which alignment must be negotiated. Many of these inequities can be understood as forms of informational asymmetry, which is itself a structural expression of power. Shareholdersâdevelopers, firms, and regulatorsâpossess privileged access to the design choices, data pipelines, and interpretive resources that determine how an AI system functions. Stakeholders, by contrast, often lack both transparency and recourse: they cannot inspect models, challenge their decisions, or meaningfully influence how objectives are specified. This epistemic imbalance mirrors classical principal-agent asymmetries but extends them into the realm of political economy. In practice, power in AI systems is exercised through control over informationâwho collects it, who interprets it, and who is rendered visible or invisible within it. Reframing power as informational asymmetry helps integrate structural critiques of inequality directly into the analytical vocabulary of alignment research. As suggested above, scaling exacerbates these asymmetries. As models grow in size and complexity, the technical expertise and computational resources required to understand or replicate them concentrate within a small set of institutions. This concentration further widens the informational gap between shareholders and stakeholders, making it more difficult for affected communities to contest design decisions or influence governance. FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix The result is a structural coupling between scale and power: the larger and more widely deployed the system, the more alignment depends on institutional arrangements that determine whose knowledge and whose values shape its behaviour. Various forms of human exploitation that influence and shape training and testing environments for AI models designed for real-world deployment can occur either through the direct utilisation of underpaid and poorly trained labour or the collection and utilisation of individualsâ data without their consent [33]. In many cases, those most affected by the deployment of AI systemsâi.e., the stakeholdersâare excluded from decisions about their design, objectives, and governance, which are made primarily by shareholders (developers, corporations, and regulators). Alignment, therefore, requires rebalancing the power dynamics that shape AI development, ensuring that diverse voices have meaningful influence in defining and evaluating alignment objectives [57]. Seen through the lens of the scaling hypothesis, this requirement is not incidental but structural: as AI systems become more general and widely deployed, alignment increasingly depends on governance mechanisms capable of managing both informational asymmetry and pluralistic disagreement at scale. 4.3 Benchmarking Degrees of Alignment A benchmark is supposed to measure a modelâs performance on a task. To create a standard metric for measuring degrees of value alignmentâeither singular or pluralisticârequires that the standard can be formalised. To benchmark alignment along the objectives axis, one must already formalise the âtrueâ objective that the system is meant to approximate. But because objective functions are themselves proxies for those true objectives, such formalisation is rarely possible. In other words, the very mechanism that makes value misalignment possibleâthe proxy problemâalso makes precise measurement of alignment impossible. Moreover, as Dotan and Milli[17] argue, benchmarks do not merely measure progress but actively shape what counts as progress. In particular, they highlight that evaluation practices embed normative assumptions about what constitutes a successful model, thereby influencing which research directions gain legitimacy within the field. For simple, well-specified tasks, it may be feasible to evaluate whether a proxy captures the intended goal. However, as tasks become more socially or ethically complex, proxies become less reliable, and the gap between objective functions and true objectives widens. If we had a reliable metric for measuring that gap, we could use it to design a better proxy; but, since we cannot, benchmarking complex alignment problems becomes conceptually circular. Such evaluation regimes can produce self-reinforcing disciplinary dynamics: once a particular model paradigm performs well on widely adopted benchmarks, the benchmark itself incentivises further optimisation for that paradigm, thereby consolidating its dominance. In this way, benchmarks function not only as evaluative tools but also as institutional mechanisms that stabilise certain research trajectories while marginalising alternatives [17]. The situation worsens when models are continually retrained on data that reflects the effects of their prior deployments, creating feedback loops that distort both the data distribution and the proxy objectives themselves [61]. When these feedback loops interact with entrenched benchmarking practices, they can further entrench specific value-laden assumptions about performance, fairness, or usefulness, making the underlying evaluative framework appear natural or inevitable even though it reflects historically contingent design choices. Under these conditions, misalignment is not only difficult to detect but also self-reinforcing. The resulting insight is negative but fundamental: for sufficiently complex objectives, measuring degrees of misalignment is impossible in principle. When this is the case, decisions about whether and how to deploy AI systems must hinge not on technical certainty but on inductive riskâthe balance of potential harms and benefits under uncertainty [18,39]. Philosophers of science have argued that scientific inference always involves value-laden judgments about acceptable levels of error and harm. The same holds for AI benchmarking: every metric embodies a decision about what counts as âgood enoughâ, which errors matter, and to whom. The analysis of value-laden disciplinary shifts makes this point especially salient for ML: evaluation metrics and benchmark Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada suites embed social and political judgments about which capabilities deserve optimisation and which risks are tolerable. Benchmarks, then, are not neutral instruments but normative artefacts that encode trade-offs between competing valuesâe.g., accuracy versus fairness or efficiency versus accountability. Acknowledging this continuity between epistemology and engineering reframes alignment as a problem of responsible inquiry. The central question shifts from âIs the system aligned?â to âIs the system aligned enough? for whom? and, at what cost?â 5 Alignment As Ongoing Governance The structural definition of the value alignment problem reveals that alignment is not a singular technical or moral challenge but a class of structurally related problems arising wherever a principal delegates a task to an agent. In human-human contexts, such delegation introduces value misalignment because principals and agents already possess distinct and sometimes competing values. In contrast, artificial agents do not possess intrinsic values: they instantiate proxy objectives designed and optimised by humans. Misalignment, therefore, emerges not from the agentâs will but from the incompleteness, ambiguity, or mis-specification of those proxies. From this perspective, alignment failures are not anomalies but expectable outcomes of imperfect formalisation under informational and social constraints. The principal-agent analogy underscores that these failures are driven less by error than by structureâby the impossibility of completely encoding human objectives in computational form and by the asymmetries of information that arise between designers, systems, and affected communities. Pluralistic alignment reframes the central question: whose values are we aligning to, and who decides? Once we recognise that there are many principalsâshareholders and stakeholders with diverse, and often conflicting, objectivesâalignment ceases to be a matter of fitting models to âthe values of humanityâ. It becomes, instead, a question of institutional design and power. The structural framework exposes this dynamic: even perfect proxy objectives cannot guarantee global alignment when values diverge across communities or social strata. A system that is well-aligned with the goals of its shareholders may be deeply misaligned with the interests of its stakeholders. Thus, the alignment problem is not only technical or normative, but fundamentally social and political in nature. On the structural view, alignment must be understood as a problem of governance. Because misalignment can arise independently along the objectives, information, and principals axes, any attempt to mitigate it must be implemented through institutions capable of regulating each axis in practice. The three-axis framework serves not only as an analytic taxonomy but as a guide to institutional design: it identifies distinct failure modes, clarifies which actors must be empowered to address them, and explains why interventions that solve one aspect of alignment may worsen misalignment along another. In practice, alignment depends not on any single mechanism but on the interaction of these institutional features across all three axes. Institutions addressing the objectives axis determine which goals AI systems are permitted to optimise and under what constraints. These may include internal safety review boards, external regulatory agencies, standards-setting bodies, and liability regimes that are able to shape the incentives of developers and deployers. Institutions addressing the information axis aim to reduce epistemic asymmetries between those who build systems, those who govern them, and those affected by them. Transparency requirements, auditing mandates, interpretability standards, incident reporting obligations, and independent evaluation organisations all function by redistributing information that would otherwise remain concentrated in the hands of system designers. Institutions addressing the principals axis determine whose interests are treated as authoritative when trade-offs must be made. Public consultation procedures, stakeholder representation, collective feedback mechanisms, democratic oversight, and participatory design processes all operate by expanding or redefining the set of principals whose values are taken to matter. A systemâs behaviour depends not only on its training objective but on who chose that objective, who was excluded from the decision, who can inspect the resulting model, and who has the authority to contest its deployment. The three-axis framework makes these dependencies explicit FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix and allows alignment research to reason systematically about the design and implementation of governance structures rather than treating institutional context as background noise. A practical implication of this view is a shift from the rhetoric of âsolvingâ value alignment to one of man- aging and mitigating misalignment across its three axes. Misalignment scales with model size, data scope, and computational powerâthe very factors that define modern deep-learning approaches to AI. As models become more general-purpose, informational asymmetries deepen, and pluralistic divergence widens. Perfect alignment, even in principle, becomes impossible. What remains is the task of designing procedures and institutions capable of continuously negotiating, auditing, and correcting misalignment as it arises. This structural account also clarifies the limitations of understanding alignment as consisting of two interrelated components: a technical problem (how to encode values) and a normative problem (which values to encode). While this distinction captures important conceptual differences, it implicitly assumes that once the correct values are identified and correctly specified, alignment will follow. (Suggesting, also, that it is possible to identify the âcorrectâ values for AI.) The three-axis framework shows that this assumption is false. Both the technical and the normative components of value alignment correspond primarily to the objectives axis, with the normative problem determining the target objective and the technical problem determining how to implement it. However, failures along the information axis and the principals axis cannot be reduced to either technical difficulty or moral disagreement. Instead,they arise from organisational incentives, resource asymmetries, and institutional secrecy. Opacity, unverifiability, and asymmetric access to data are not simply technical limitations but features of organisational arrangements, incentives, and resource distribution. Likewise, failures along the principals axis are not merely disagreements about values in the abstract but conflicts over authority, representation, and decision-making power. The technical/normative distinction therefore collapses multiple structurally distinct problems into a single dimension, whereas the three-axis decomposition embeds this distinction within a broader socio-technical analysis that makes explicit the institutional conditions under which alignment succeeds or fails. This reconceptualisation, therefore, situates alignment within a broader frame of socio-technical governance. AI systems are embedded in networks of labour, data extraction, and environmental cost. These systemsâ âobjectivesâ are operationalised not only through code but through economic incentives, regulatory regimes, and the social hierarchies in which they are deployed. To treat alignment as an engineering problem alone is to ignore the upstream power structures that determine who gets to specify objectives and who bears the risks when they fail. Recent work on âfair processâ views of alignment suggest that alignment may sometimes be achieved through agreement on procedures rather than through agreement on values [27]. On this view, alignment does not require agreement on a single set of values; instead, it requires agreement on a decision procedure that stakeholders regard as legitimate. If the process by which objectives are defined is inclusive, transparent, and procedurally fair, then the resulting system may count as aligned even when substantive disagreement persists. Hence, pluralism is managed by institutionalising fair procedures rather than by identifying a universally correct objective. This suggests a possible alternative to the structural view: alignment might be achieved by securing consensus at the level of process, allowing the principals axis to be stabilised through procedural legitimacy. The structural framework, however, predicts that fair-process solutions alone cannot fully resolve the alignment problem, because agreement on procedure does not eliminate the other sources of misalignment. Even when stakeholders agree on a decision procedure, objective proxies may still be imperfect, and informational asymmetries may still prevent meaningful oversight. Moreover, it is unclear whether complete procedural legitimacy is possible in real-world institutions because differences in power, expertise, and access to information mean that some actors inevitably exercise greater influence over the design and deployment of systems. For this reason, fair-process accounts should be understood not as replacements for the structural analysis but as governance mechanisms operating primarily along the principals axis. They can reduce conflict over authority without eliminating misspecification or opacity, and the three-axis framework makes these limits visible Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada Consequently, alignment demands institutional mechanisms for oversight, participatory design, and con- testabilityâmechanisms that can adapt as both models and social conditions evolve. The challenge is not to achieve perfect convergence between human and machine values, but to build systems resilient to divergence: systems capable of being corrected, contested, and re-aligned as our understanding of âour valuesâ itself changes. Under this interpretation, governance is not an external constraint on alignment but its primary medium. The task of alignment research therefore expands from designing better objective functions to designing better decision procedures, oversight structures, and forms of collective control. The three-axis decomposition provides a framework for this task by allowing us to ask, for any proposed system: which objectives does it optimise, who can understand and evaluate it, and whose interests determine whether it should exist at all? In this light, the question for future alignment research is not merely how to make AI systems safe or compliant, but how to make them responsive to human needs. If the objective of alignment is to ensure that AI systems act in accordance with human values, then the primary task becomes identifying which humans, which values, and through what processes those values are continually negotiated and maintained. Alignment, properly understood, is not an endpoint of technical controlâit is the architecture of shared agency. Considering the objectives encoded in AI systems with respect to a particular set of principals sheds light on how AI systems fail to satisfy these objectives. In general, when any given AI model is touted as a solutionâparticularly by the shareholders of that systemâit is fruitful to ask: to what problem? Generative AI Usage Statement The author(s) did not use generative AI in the writing of this paper. References [1]Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan ManĂ©. 2016. Concrete Problems in AI Safety. arXiv 1606.06565 (2016), 1â29. https://arxiv.org/abs/1606.06565. [2]Mel Andrews. 2025. The Immortal Science of ML: Machine Learning and the Theory-Free Ideal. Erkenntnis (2025), 1â23. https: //doi.org/10.1007/s10670-025-01010-x. [3]Lora Aroyo, Alex S. Taylor, Mark Diaz, Christopher M. Homan, Alicia Parrish, Greg Serapio-Garcia, Vinodkumar Prabhakaran, and Ding Wang. 2023. DICES Dataset: Diversity in Conversational AI Evaluation for Safety. arXiv 2306.11247 (2023), 1â22.https: //arxiv.org/abs/2306.11247. [4] Yoshua Bengio. 2023. How Rogue AIs may Arise. https://yoshuabengio.org/en/blog/how-rogue-ais-may-arise. [5] Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, Sören Mindermann, Adam Ober- man, Jesse Richardson, Oliver Richardson, Marc-Antoine Rondeau, Pierre-Luc St-Charles, and David Williams-King. 2025. Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? arXiv 2502.15657 (2025), 1â58. https://arxiv.org/abs/2502.15657. [6] Ruha Benjamin. 2019. Race After Technology: Abolitionist Tools for the New Jim Code. Polity, Cambridge. [7] Stevie Bergman, Nahema Marchal, John Mellor, Shakir Mohamed, Iason Gabriel, and William Isaac. 2024. STELA: a community-centred approach to norm elicitation for AI alignment. Scientific Reports 14, 1 (2024), 6616. [8]Nick Bostrom. 2003. Ethical issues in advanced artificial intelligence. In Science fiction and philosophy: from time travel to superintelligence, Susan Schneider (Ed.). Wiley & Blackwell, West Sussex, 277â284. [9] Nick Bostrom. 2014. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, Oxford. [10] Meredith Broussard. 2023. More than a Glitch: Confronting Race, Gender, and Ability Bias in Tech. The MIT Press, Cambridge, MA. [11] Brian Christian. 2020. The Alignment Problem: Machine Learning and Human Values. W. W. Norton & Company, New York. [12] Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan MossĂ©, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, and Others. 2024. Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback. arXiv (2024), 1â15. https://arxiv.org/abs/2404.10271. [13] Kate Crawford. 2021. Atlas of AI. Yale University Press, New Haven, CT. [14] Andrew Critch and Stuart Russell. 2023. Tasra: A taxonomy and analysis of societal-scale risks from ai. arXiv 2306.06924 (2023), 1â18. https://arxiv.org/abs/2306.06924. [15]Aida Mostafazadeh Davani, Mark DĂaz, and Vinodkumar Prabhakaran. 2022. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Association for Computational Linguistics 10 (2022), 92â110. FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix [16]Daniel Dewey. 2011. Learning What to Value. In AGI 2011: 4th International Conference on Artificial General Intelligence (Lecture Notes in Computer Science, Vol. 6830), J. Schmidhuber, K. R. ThĂłrisson, and M. Looks (Eds.). Springer, Berlin, Heidelberg, 309â314. [17]Ravit Dotan and Smitha Milli. 2019. Value-laden Disciplinary Shifts in Machine Learning. arXiv 1912.01172 (2019), 1â10. https: //arxiv.org/abs/1912.01172. [18] Heather Douglas. 2000. Inductive Risk and Values in Science. Philosophy of Science 67, 4 (2000), 559â579. [19] Peter Eckersley. 2019. Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function). arXiv 1901.00064 (2019), 1â13. https://arxiv.org/abs/1901.00064. [20] Kathleen M. Eisenhardt. 1989. Agency Theory: An Assessment and Review. The Academy of Management Review 14, 1 (1989), 57â74. [21] Scott Emmons, Caspar Oesterheld, Vincent Conitzer, and Stuart Russell. 2025. Observation Interference in Partially Observable Assistance Games. arXiv 2412.17797 (2025), 1â26. https://arxiv.org/abs/2412.17797. [22]Danielle Ensign, Sorelle A. Friedler, Scott Neville, Carlos Scheidegger, and Suresh Venkatasubramanian. 2018. Runaway Feedback Loops in Predictive Policing. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency (Proceedings of Machine Learning Research, Vol. 81), Sorelle A. Friedler and Christo Wilson (Eds.). PMLR, 160â171. [23] Sina Fazelpour and Will Fleisher. 2025. The Value of Disagreement in AI Design, Evaluation, and Alignment. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency. Association for Computing Machinery, 2138â2150. https: //doi.org/10.1145/3715275.3732146. [24]Jaime F. Fisac, Monica A. Gates, Jessica B. Hamrick, Chang Liu, Dylan Hadfield-Menell, Malayandi Palaniappan, Dhruv Malik, S. Shankar Sastry, Thomas L. Griffiths, and Anca D. Dragan. 2020. Pragmatic-Pedagogic Value Alignment. In Springer Proceedings in Advanced Robotics, N. Amato, G. Hager, S. Thomas, and M. Torres-Torriti (Eds.). Vol. 10. Springer, 49â57. [25] Future of Life Institute. 2017. Asilomar AI Principles. https://futureoflife.org/open-letter/ai-principles/. [26] Iason Gabriel. 2020. Artificial Intelligence, Values, and Alignment. Minds and Machines 30 (2020), 411â437. [27] Iason Gabriel and Geoff Keeling. 2025. A matter of principle? AI alignment as the fair treatment of claims. Philosophical Studies 182 (2025), 1951â1973. [28]Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, and Scott Emmons. 2025. The partially observable off-switch game. Proceedings of the AAAI Conference on Artificial Intelligence 39, 26 (2025), 27304â27311. [29] Trystan S. Goetze. 2024. AI Art is Theft: Labour, Extraction, and ExploitationâOr, On the Dangers of Stochastic Pollocks. PhilArchive (2024). Unpublished preprint of 10 January 2024. https://philarchive.org/rec/GOEAAI-2. [30]David E. Goldberg. 1987. Simple genetic algorithms and the minimal deceptive problem. In Genetic Algorithms and Simulated Annealing (Research Notes in Artificial Intelligence), Lawrence D. Davis (Ed.). Morgan Kaufmann Publishers, Burlington, MA, 74â88. [31]John-Stewart Gordon. 2023. Objections. In The Impact of Artificial Intelligence on Human Rights Legislation. Palgrave Macmillan, Cham, 75â82. [32]Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeffrey T. Hancock, Tatsunori Hashimoto, and Michael S. Bernstein. 2022. Jury Learning: Integrating Dissenting Voices into Machine Learning Models. arXiv 2202.02950 (2022), 1â19. https://arxiv.org/abs/ 2202.02950. [33]Mary L. Gray and Siddharth Suri. 2019. Ghost Work: How to Stop Silicon Valley from Building a New Global Underclass. Eamon Dolan Books, New York. [34] Dylan Hadfield-Menell. 2021. The Principal-Agent Alignment Problem in Artificial Intelligence. Ph. D. Dissertation. EECS Department, University of California, Berkeley. http://w2.eecs.berkeley.edu/Pubs/TechRpts/2021/EECS-2021-207.html. [35]Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2016. Cooperative inverse reinforcement learning. In NIPSâ16: Proceedings of the 30th International Conference on Neural Information Processing Systems, Daniel D. Lee, Ulrike von Luxburg, Roman Garnett, Masashi Sugiyama, and Isabelle Guyon (Eds.). Association for Computing Machinery, 3916â3924. [36] Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel, and Stuart Russell. 2017. The Off-Switch Game. arXiv 1611.08219 (2017), 1â8. https://arxiv.org/abs/1611.08219. [37] Dylan Hadfield-Menell and Gillian K. Hadfield. 2019. Incomplete Contracting and AI Alignment. In AIES â19: Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society, Vincent Conitzer, Gillian Hadfield, and Shannon Vallor (Eds.). Association for Computing Machinery, New York, 417â422. [38]Foad Hamidi, Morgan Klaus Scheuerman, and Stacy M. Branham. 2018. Gender recognition or gender reductionism?: The social implications of embedded gender recognition systems. Proceedings of the ACM Conference on Human Factors in Computing Systems (CHI) (2018), 1â13. [39] Carl Hempel. 1965. Aspects of Scientific Explanation. Free Press, New York. [40] Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2023. Aligning AI With Shared Human Values. arXiv 2008.02275 (2023), 1â29. https://arxiv.org/abs/2008.02275. [41]Dan Hendrycks and Mantas Mazeika. 2022. X-risk analysis for AI research. arXiv 2206.05862 (2022), 1â36. https://arxiv.org/abs/2206.05862. [42]Anna Lauren Hoffmann. 2019. Where fairness fails: data, algorithms, and the limits of antidiscrimination discourse. Information, Communication & Society 22, 7 (2019), 900â915. Relative principals, pluralistic alignment, & the structural value alignment problemFAccT â26, June 25â28, 2026, Montreal, QC, Canada [43]Saffron Huang, Divya Siddarth, Liane Lovitt, Thomas I. Liao, Esin Durmus, Alex Tamkin, and Deep Ganguli. 2024. Collective constitutional AI: Aligning a language model with public input. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, 1395â1417. [44]Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant. 2021. Risks from Learned Optimization in Advanced Machine Learning Systems. arXiv 1906.01820 (2021), 1â39. https://arxiv.org/abs/1906.01820. [45]Michael C. Jensen and William H. Meckling. 1976. Theory of the Firm: Managerial Behaviour, Agency Costs and Ownership Structure. Journal of Financial Economics 3, 4 (1976), 305â360. [46]Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhonghao He, Lukas Vierling, Donghai Hong, Jiayi Zhou, Zhaowei Zhang, Fanzhi Zeng, Juntao Dai, Xuehai Pan, Kwan Yee Ng, Aidan OâGara, Hua Xu, Brian Tse, Jie Fu, Stephen McAleer, Yaodong Yang, Yizhou Wang, Song-Chun Zhu, Yike Guo, and Wen Gao. 2025. AI Alignment: A Comprehensive Survey. arXiv 2310.19852 (2025), 1â105. https://arxiv.org/abs/2310.19852. [47] Steven Kerr. 1975. On the Folly of Rewarding A, While Hoping for B. Academy of Management Journal 18 (1975), 769â783. [48]Hannah Rose Kirk, Alexander Whitefield, Paul Röttger, Andrew Bean, Katerina Margatina, Juan Ciro, Rafael Mosquera, Max Bartolo, Adina Williams, He He, Bertie Vidgen, and Scott A. Hale. 2024. The PRISM Alignment Project: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models. arXiv 2404.16019 (2024), 1â107. https://arXiv.org/abs/2404.16019. [49] Travis LaCroix. 2025. Artificial Intelligence and the Value Alignment Problem: A Philosophical Introduction. Broadview Press. [50] Travis LaCroix and Alexandra Sasha Luccioni. 2025. Metaethical Perspectives on âBenchmarkingâ AI Ethics. AI and Ethics 5 (2025), 4029â4047. [51] Jean-Jacques Laffont and David Martimort. 2002. The Theory of Incentives: The Principal-Agent Model. Princeton University Press, Princeton. [52]Joel Lehman and Kenneth O. Stanley. 2008. Exploiting Open-Endedness to Solve Problems Through the Search for Novelty. In Proceedings of the Eleventh International Conference on Artificial Life (ALIFE XI). The MIT Press, Cambridge, MA, 329â336. [53]Alexandra Sasha Luccioni, Yacine Jernite, and Emma Strubell. 2023. Power Hungry Processing: Watts Driving the Cost of AI Deployment? arXiv 2311.16863 (2023), 1â20. https://arxiv.org/abs/2311.16863. [54] Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2023. Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model. Journal of Machine Learning Research 24, 253 (2023), 1â15. [55] Kristian Lum and William Isaac. 2016. To predict and serve? Significance 13 (2016), 14â19. [56] Dan McQuillan. 2022. Resisting AI: An Anti-fascist Approach to Artificial Intelligence. Bristol University Press, Bristol. [57]Milagros Miceli, Julian Posada, and Tianling Yang. 2022. Studying Up Machine Learning Data: Why Talk About Bias When We Mean Power? Proceedings of the ACM on Human-Computer Interaction 6, GROUP (2022), 1â14. [58]Melanie Mitchell, Stephanie Forrest, and John H. Holland. 1992. The royal road for genetic algorithms: Fitness landscapes and GA performance. In Proceedings of the First European Conference on Artificial Life, F. J. Varela and P. Bourgine (Eds.). The MIT Press, Cambridge, MA, 1â11. [59]Richard Ngo, Lawrence Chen, and Sören Mindermann. 2023. The Alignment Problem from a Deep Learning Perspective. arXiv 2209.00626 (2023), 1â21. https://arxiv.org/abs/2209.00626. [60] Stephen M. Omohundro. 2008. The Basic AI Drives. In Artificial General Intelligence 2008: Proceedings of the First AGI Conference, Pei Wang, Ben Goertzel, and Stan Franklin (Eds.). IOS Press, Amsterdam, 483â492. [61]Cathy OâNeil. 2016. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy. Broadway Books, New York. [62]Joshua C Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, and Olga Russakovsky. 2019. Human uncertainty makes classification more robust. Proceedings of the IEEE/CVF international conference on computer vision (2019), 9617â9626. [63]Elinor Poole-Dayan, Jiayi Wu, Taylor Sorensen, Jiaxin Pei, and Michiel A. Bakker. 2026. Benchmarking Overton Pluralism in LLMs. arXiv 2512.01351 (2026), 1â40. https://arxiv.org/abs/2512.01351. [64]Mahendra Prasad. 2018. Social choice and the value alignment problem. In Artificial intelligence safety and security, Roman V. Yampolskiy (Ed.). Chapman & Hall, London, 291â314. [65]Inioluwa Deborah Raji and Roel Dobbe. 2023. Concrete Problems in AI Safety, Revisited. arXiv 2401.10899 (2023), 2023. https: //arxiv.org/abs/2401.10899/. [66]Inioluwa Deborah Raji, Timnit Gebru, Margaret Mitchell, Joy Buolamwini, Joonseok Lee, and Emily Denton. 2020. Saving Face: Investigating the Ethical Concerns of Facial Recognition Auditing. arXiv 2001.00964 (2020), 1â7. https://arxiv.org/abs/2001.00964. [67] Stuart Russell. 2019. Human Compatible: Artificial Intelligence and the Problem of Control. Viking, New York. [68]Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael Dennis, Pieter Abbeel, Anca Dragan, and Stuart Russell. 2020. Benefits of Assistance over Reward Learning. 34th Conference on Neural Information Processing Systems (NeurIPS 2020) - Workshop on Cooperative AI (2020). FAccT â26, June 25â28, 2026, Montreal, QC, CanadaLaCroix [69]Moshe Sipper, Ryan J. Urbanowicz, and Jason H. Moore. 2018. To Know the Objective Is Not (Necessarily) to Know the Objective Function. BioData Mining 11, 21 (2018), 1â3. [70]Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell. 2024. Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF. arXiv 2312.08358 (2024), 1â26. https://arxiv.org/abs/2312.08358. [71] Taylor Sorensen, Liwei Jiang, Jena Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, and Others. 2024. Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties. arXiv 2309.00779 (2024). https://arxiv.org/abs/2309.00779. [72]Taylor Sorensen, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye, Liwei Jiang, Ximing Lu, Nouha Dziri, Tim Althoff, and Yejin Choi. 2024. A roadmap to pluralistic alignment. arXiv 2402.05070 (2024), 1â23. https://arxiv.org/abs/2402.05070. [73] Max Tegmark. 2018. Life 3.0: Being human in the age of artificial intelligence. Vintage, New York. [74]Michael Henry Tessler, Michiel A. Bakker, Daniel Jarrett, Hannah Sheahan, Martin J. Chadwick, Raphael Koster, Georgina Evans, Lucy Campbell-Gillingham, Tantum Collins, David C. Parkes, Matthew Botvinick, and Christopher Summerfield. 2024. AI can help humans find common ground in democratic deliberation. Science 386, 6719 (2024), eadq2852. [75] Eliezer Yudkowsky. 2011. Complex value systems in friendly AI. 6830 (2011), 388â393. Received 13 January 2026; revised 25 March 2026; accepted 16 April 2026