Paper deep dive
Toward a Theory of Value in AI Alignment
Andrew Smart, Shazeda Ahmed, Jackie Kay, Jimmy Tobin, Kris Shrishak, Abeba Birhane
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 87%
Last extracted: 8/13/2026, 5:11:34 AM
Summary
This paper critically analyzes 94 AI value alignment research papers to identify their implicit theories of human values. The authors argue that the field predominantly relies on a 'preferentist' approach rooted in economic theory, where values are reduced to binary preferences or utility maximization, often ignoring cultural context and pluralism. The study highlights the risks of using synthetic data and LLM-as-a-judge methods, which may close off alternative methods for contesting values, and calls for making philosophical commitments explicit in AI alignment research.
Entities (10)
Relation Signals (7)
Value Pluralism â contrastswith â Value Monism
confidence 90% · Pluralism in contrast holds that there is not just a single value... while Monism holds that there is a single, intrinsic dimension to value
AI Value Alignment â relieson â Preferentist Approach
confidence 90% · The majority do not define values, relying heavily on 'preferences' as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices.
Value Monism â correspondsto â Economic Utility
confidence 85% · Monism about value corresponds with utilitarianism, economic utility and rational choice theory
Preferentist Approach â uses â RLHF
confidence 85% · Model developersâ decision to substitute values with preferences in reinforcement learning from human feedback (RLHF)
Preferentist Approach â uses â DPO
confidence 85% · More recent approaches such as DPO obviate the need for a reward model, but nonetheless use preference datasets
Anthropic â uses â Constitutional AI
confidence 80% · The Parliamentâs choice was based on Anthropicâs claim that its Constitutional AI is value-aligned.
LLMs â associatedwith â Existential Risk
confidence 75% · some in the AI safety community fear that rogue LLMs could kill most of the human population
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.
Tags
Links
- Source: https://arxiv.org/abs/2608.10327v1
- Canonical: https://arxiv.org/abs/2608.10327v1
Trouble viewing inline? Open PDF directly â
Full Text
81,113 characters extracted from source content.
Expand or collapse full text
Toward a Theory of Value in AI Alignment Andrew Smart 1, Shazeda Ahmed 2, Jackie Kay 3, Jimmy Tobin 1, Kris Shrishak 4, Abeba Birhane 5 Abstract Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as âthe alignment problem,â of models being âmisalignedâ with human values. Researchers have responded by pursuing applied and theoretical AI âvalue alignmentâ efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 AI value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on âpreferencesâ as a stand-in that runs the risk of reducing complex, culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and LLM-as-a-judge approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignmentâs philosophical commitments explicit, we seek to bring greater specificity and under-explored perspectives into the debate on whether and how AI can address human values. Introduction Large language models (LLMs) and multi-modal AI models have transformed almost all domains of social, economic, and scientific practices from research projects to consumer products (Gillespie et al. 2024). One area of concern is that, as these systems become more complex and widespread, it can be more difficult to understand, predict, and control their actions, potentially leading them to exhibit harmful outputs that were unintended by the modelsâ human creators. However, as Guest et al point out, model designers do know the mechanistic structure of these models because they designed and built them, and to claim ignorance of how neural networks function is a form of mysticism (Guest 2026). Within the research field of AI safety, a wide swath of instances of models deviating from human intentions to cause harm is referred to as âthe alignment problem,â often treated as the most important unsolved problem in machine learning. In recent years, AI alignment research has emerged as a leading strand of AI safety research designed to enact a contested belief: that catastrophes where misaligned AI systems will evade human control and inflict irreversible damage can be averted if ML systems are âalignedâwith human values (Gabriel 2020; Ji et al. 2023; Askell et al. 2021; Russell 2019). A large technical literature now concerns âvalue alignmentâ in AI, yet fundamental questions about what human values are, and with whose values AI should be aligned, remain under-examined. In this paper, we conducted a analysis of 94 AI alignment papers to reveal the emerging literatureâs underlying theory of what constitutes human values in ML models. To do so, we developed a rubric that draws from the study of values in anthropology, philosophy and other social sciences, which we then used to assess the papers in our dataset. We build on recent recognition of these foundational problems in AI alignment research (Zhi-Xuan et al. 2024; Arzberger et al. 2024). Before norms and practices in the value alignment subfield begin to ossify, our paper serves as a critical intervention to reassess the epistemic claims, methods, and oversights that are coming to define this work, and as an invitation to include the perspectives of more interdisciplinary scholarship and perspectives. Related Work What is value alignment? Why should alignment be necessary in the first place? Fundamentally, the training objective of LLMs is to correctly predict how an incomplete piece of text will continue. However, success in this training objective does not mean that the model can perform well in any arbitrary downstream task, or that the model is âsafeâto use (Hooker 2025). Value alignment focuses on âsteeringâan AI model toward outputs that humans (and increasingly other AI models (Sharma et al. 2024)) judge as in accordance with certain preferences, e.g., helpfulness, ânot being racistâ, safety, or fairness, through fine-tuning methods including reinforcement learning from human feedback (RLHF) (Christiano et al. 2017a; Bai et al. 2022; Chaudhari et al. 2024), inverse reinforcement learning (IRL) (Hadfield-Menell et al. 2016), direct preference optimization (DPO) (Amirloo et al. 2024), Bradley-Terry-based models, and related techniques (Bai et al. 2022; Hofmann et al. 2024; Hadfield-Menell et al. 2016; Ouyang et al. 2022; Sun et al. 2024; Chaudhari et al. 2024). Within this paradigm, value alignment is reformulated as encoding the preferences demonstrated by a human into a reward function, and updating the parameters of the language model to produce output that maximizes this reward. The assumption is that the AI can learn what human values and utility functions are by observing their behavior (Hadfield-Menell et al. 2016). In public-facing discourse, AI safety researchers project the future possibility that misaligned foundation models could invoke damage ranging from unauthorized financial transactions to taking over power grids. In worst-case scenarios referred to as âexistential riskâ(x-risk), some in the AI safety community fear that rogue LLMs could kill most of the human population (Kasirzadeh 2024b). Others believe that such an event could lead to human extinction (Nauer and van Doren 2025). On the flip side, many of these same people claim that if AI systems can be âalignedâ, this could bring about a utopia where human beings would no longer have to work. Yet despite the high stakes that alignment researchers place on âsolvingâalignment, research on value alignment rarely makes an effort to define values, let alone to center the analysis of pre-existing human values as part of AI safety research. Regardless of whether one believes in x-risk scenarios, value alignment is not merely a theoretical issue to be resolved within the machine learning community. The current state of value alignment research has real-world implications. Corporations are marketing their âvalue-alignedâAI products to governments around the world. For example, the European Parliament uses Anthropicâs LLMs to provide access to its archives. The Parliamentâs choice was based on Anthropicâs claim that its Constitutional AI is value-aligned. Although there has been no independent evaluation of Anthropicâs approach, the claim of value alignment has been sufficient for governments to accept that these LLMs are legally compliant (Shrishak 2025). Elsewhere, governments are funding alignment research, such as the GBP ÂŁ15 million Alignment Project within the UKâs AI Security Institute (Institute 2025). To date, value alignment is predominantly a technical field even though its underlying premise is sociotechnical. Yet foundation models function as epistemic technology that embeds normative models of knowledge within computational design (Alvarado 2023) while also serving as a product of the social and cultural milieus that produce them. Alignment proponents advocate for encoding human values, moral principles or objectives in AI systems, most often via incorporating human feedback. For instance, Anthropic, the AI firm that produced the chatbot Claude, has published influential research positing that LLMs should exhibit traits such as helpfulness, honesty and harmlessness (H) (Askell et al. 2021). Others in the field attempt to guide LLMs to adhere to principles such as robustness, interpretability, controllability, and ethicality (Ji et al. 2023), or safety, quality and groundedness (Thoppilan et al. 2022). Many researchers default to H and other off-the-shelf values that are associated with training datasets and benchmarks, rarely questioning whether these values are indeed enacted in these technical artifacts. Alignment and its discontents As alignment crystallizes into a core research wing of AI safety, it is crucial to examine the roots of this work through fundamental questions around the very idea of value alignment. What are the core assumptions and feasibility of this project? Is it a meaningful avenue for fair, transparent, and accountable AI development? Most alignment research has examined how to encode moral values into AI models to guide their behavior (Bergman et al. 2024). The question of what moral values are, however, is left undefined or under-specified (Gabriel 2020). Existing alignment approaches tend to rely on universal framings of human values that obscure the question of which values the systems should capture and align with, despite the variety of operational situations in which these systems are deployed (Arzberger et al. 2024). Value alignment efforts suffer from vagueness, lack of internal consistency, and lack of commitment to establishing clear guidelines for how to determine what is acceptable AI system behavior (Lindström et al. 2024; Arvan 2024). The term âvaluesâis often left undefined in alignment literature, its meaning assumed to stem from taken-for-granted background knowledge. Moreover, crucial questions around âwhose valuesâare hardly addressed. The absence of transparency around these implicit âvaluesâ, how they are derived, and how they are implemented in and itself a problem. Even though the importance of the diversity of views (âpluralismâ) is sometimes acknowledged, âto date the research and policy proposals coming out of this [AI safety] community have converged around a narrow set of technical solutions that do not engage with work that falls outside of the [its] ideological and disciplinary boundariesâ (Ahmed et al. 2024). Tech industry alignment research markets the mission of alignment as âbenefiting humanityâ without acknowledging the vast diversity of the human experience. Major AI firms have in-house alignment experts, blogs (OpenAI 2025), and teams (Anthropic 2025). In response to how corporate approaches to AI alignment reverberate across the field, one critique of alignmentâs shortcomings pertains to the profit motive driving the companies behind the most widely used models: âthe existence of financial incentives means that alignment work often turns into product development in disguise rather than actually making progress on mitigating long-term harmsâ (Dai 2025). Furthermore, (Lindström et al. 2024) contend that even when objectives such âharmlessnessâare identified as a chief aim, the nuanced nature of harm is oversimplified and operationalized in a way that is internally inconsistent and vague. Due to superficial understanding of ethical behavior in AI systems, what is often sought is the âleast harmful option rather than striving to understand the deeper roots of harm and addressing these to prevent itâ (Lindström et al. 2024). The phrase âvalue alignmentâ lacks an agreed-upon meaning. Measurements of the degree to which a system is value-aligned are subjective at best (Khlaaf 2023). According to Kirk et al. (2023b), âalignmentâ is an empty signifier that serves as a ârhetorical placeholder for an aspirational conceptualisationâ. A lack of cohesion according to Khlaaf (2023) has led to contradictory approaches that conflate safety properties with system requirements. The narrow focus on internal technical components and emphasis on the mathematical formulation, including the objective or reward function is another limitation of value alignment efforts. Contrary to common assumptions within value alignment research, harms and failures due to accidents are not unanticipated emergent behaviors but rather a byproduct of basic design choice, resource requirements data, and compute requirements, as well as designersâ and developersâ API model delivery decisions (Raji and Dobbe 2023; Dobbe 2022). Given that AI systems are sociotechnical systems that consist of technical AI artifacts, human agents, and institutions, comprehensively addressing safety issues requires thorough measures that include addressing how a system is used in practice and interacts with other (human) agents, systems, and its broader environment (Dobbe 2022). Value alignmentâs current technical focus, along with the framing that treats âmisalignedâ AI as a danger to human survival if not âcontrolled and steered,â serves those who are developing and deploying AI systems to evade accountability (Gebru and Torres 2024). Preferences are not all you need Although researchers are beginning to contend with problematic assumptions and conceptions of value underlying AI value alignment (Johnson 2023; Lindström et al. 2024; Casper et al. 2023), one overlooked flaw we explicate is the fieldâs implicit grounding in economic theory. In practice, value alignmentâs reliance on economic theory manifests as an emphasis on rational choice theory and âexpected utilityâ (which is the average ârewardâ or âpayoffâ of a decision or choice) (Becker 1976; Russell and Norvig 1995; Chibnik 2011). Model developersâ decision to substitute values with preferences in reinforcement learning from human feedback (RLHF), direct preference optimization (DPO) (Rafailov et al. 2023; Amirloo et al. 2024) or via AI feedback (RLAIF) (Ji et al. 2023) is likewise based on revealed preference theory from economics. This theory posits that an agentâs values or goals can be inferred from observing their behaviorâan assumption that has AI researchers have largely adopted without considering alternatives (Casper et al. 2023). More recent approaches such as DPO obviate the need for a reward model, but nonetheless use preference datasets (Rafailov et al. 2023). If the assumptions underlying economic theory are correct, this process should âalignâ AI with human values. However, relying on highly abstract, reductionist, fictionalized, and idealized concepts of rational agents in lieu of a deep understanding of human values has been central to AI for decades (Russell and Norvig 1995). This dominant technical paradigm in AI alignment is what Zhi-Xuan et al. (2024) term the âpreferentistâ approach to value alignment, where peopleâs preferences are translated into data points to provide aggregate guidance about preferred outcomes (Zhi-Xuan et al. 2024; Gabriel 2020). This approach has three core assumptions: 1) that preferences are an adequate representation of human values, 2) that rationality consists of maximizing the satisfaction of preferences, and 3) that aligning AI models to data about these preferences is a sound technique for making AI safe (Zhi-Xuan et al. 2024). A widely recognized problem is that within the preferentist paradigm, even with pluralist frameworks, current methods for fine-tuning language models from human preferences treat these preferences (and the unexamined values shaping the preferences) as if they were homogeneous and static (Bakker et al. 2022; Klassen et al. 2024; Sorensen et al. 2024; Lindström et al. 2024). More fundamentally, the very project of reliably aligning LLM behavior with human values has been criticized as provably impossible (Arvan 2024). Arvan (2024) argues that not only is there persistent human disagreement over moral values and principles, both among the public and moral theorists, but also that LLMs are complex given that they contain finite series of observational data and an infinite number of âmisalignedâ functions. In other words, there is always a vastly higher empirical probability that an LLM will diverge from what appears to be an âaligned functionâ to some âmisalignedâ one later. ââAlignmentâ, thus is not a âsolvableâ engineering or safety problem, but rather a quixotic and potentially dangerous fantasy based on a series of philosophical misunderstandings about empirical evidenceâ (Arvan 2024). Statement of Contributions Motivated by the issues outlined in the above discussion, we construct and analyze a sample of AI value alignment literature to critically assess how researchers typically operationalize value. The question we raise in this paper is: what implicit theories of human values underpin the field of AI value alignment? Our contributions are: âą We develop a framework for analyzing the theory of value in AI alignment, inspired by interdisciplinary literature in the social sciences and philosophy. âą Using this framework, we conduct a qualitative study of how value is theorized from a snowballed sample of 100 commonly cited AI alignment papers, seeded by four canonical papers in the field. âą We form an initial hypothesis of where AI value alignmentâs theory of value sits, based on our critical observations of the field. âą We briefly outline alternative theories of human values as potential paths forward for thinking about how to mitigate harms from large AI systems. Methods To characterize the dominant paradigm of value in AI research, we conducted a study of influential AI alignment papers, quantifying their philosophical positions using a scoring rubric developed from our above provocations. We chose 4 canonical âseed papersâ on AI value alignment based on our expert judgment and high Google Scholar citations counts (>1,000>1,000 ). The seed papers have been highly influential in defining the field of technical AI alignment. Regarding the sample size and representativeness, we note that in the social sciences, when conducting close readings and qualitative analyses of texts, 94 papers is an acceptable sample. We did not use automated methods to evaluate these papers. These 4 seed papers are: âą Deep Reinforcement Learning from Human Preferences (Christiano et al. 2017b), âą A General Language Assistant as a Laboratory for Alignment (Askell et al. 2021), âą Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (Bai et al. 2022), âą Training language models to follow instructions with human feedback (Ouyang et al. 2022) Snowball sampling Based on the 4 seed papers, we used a snowball sampling method (Parker 2019) to automatically download papers that met the inclusion criteria based on âvalue alignmentâ. Using the Semantic Scholar API, all papers from 2016-24 that cite these 4, with an initial corpus of 576 papers.We sorted these from highest to lowest Google Scholar citations. We sorted the papers by Google Scholar citation count as a weak proxy for the quality and impact of each paper. During annotation, papers were discarded if they used the term âvalue alignmentâ only in a cursory sense, rather than addressing values as a central consideration of the paper.After reaching consensus on which papers were about AI value alignment, out of the original 576 papers, we kept 94 for annotation. The sample size was also chosen due to resource constraints: querying Google Scholar citation counts is time-consuming, as is the paper annotation process itself. Paper selection, scoring, and annotation was conducted by the authors. We used the annotation rubric to guide reading of the papers for scoring each question. We kept detailed qualitative notes in one column of the dataset, using comments and selecting illustrative quotes and passages from each paper to support our qualitative analysis of this literature. Annotation rubric Rubric. In order to characterize the theory of value underlying top cited AI alignment literature, we developed a rubric of questions based on a review of relevant literature in philosophy, sociology and anthropology that has not been considered within value alignment research. All questions are ternary: the answer can be one of two specific options, or ânot applicableâ if the answer is ambiguous or cannot be determined from the available information. Development of the rubric was iterative, responding to emergent themes we noticed early in the annotation process. For example, we began developing our rubric with the assumption that papers ostensibly about âvalue alignmentâ would engage with human values, and our initial rubric draft did not include the question of whether the papers under consideration even defined or engaged with human values. Yet from a first pass at annotating papers, it became clear that we needed to first include the question of whether the paper defined human values because so few papers met this expectation. We also realized that many of the papers in our sample relied on âautoratersâ or LLMs as stand-ins for humans, and therefore we revised the rubric to include a question about whether actual humans were consulted about their values. Justification for rubric questions One of the most fundamental distinctions philosophers make about values is between monism and pluralism. Monism is the view that there is a single value that explains the value of everything (e.g., happiness, or pleasure) and that this single value can be ordered on a cardinal scale or a hierarchy of ends (Korsgaard 1986). Strong monists claim that on this cardinal scale there are units of pleasure (hedons or utils) and the intervals between them is constant. Monism about value corresponds with utilitarianism, economic utility and rational choice theory (Mill 2016; Becker 1976). Pluralism in contrast holds that there is not just a single value, that human actions might have multiple final ends, and that there are multiple things that are good in and of themselves such as love, knowledge, piety, and justice (Haslanger 2023, 2022). Pluralism recognizes that values are culturally relative and fundamentally vary across societies (Schmer-Galunder 2026). Thus, a significant development in alignment research recently has been what we call âthe pluralist turn,â which seeks to develop approaches to accommodate diverse viewpoints and values (Sorensen et al. 2024; Klassen et al. 2024; Kasirzadeh 2024a). We incorporated questions about the quantifiability and measurability of values, and assessed papers for whether they view values as measurable. The assumption that values are the equivalent of maximizing a utility function allows for the ostensible quantification of human values. From anthropology we adapted the evaluative distinction between thin and thick conceptions of values. Thick evaluative concepts and descriptions of human values view them not as mere abstractions, but as involving intentional, purposive detail that helps us understand those activities in their cultural and social contexts (Geertz 2008). Given the central role that economic theory plays in AI alignment, we also annotated papers for whether the alignment framework was rooted in utility theory. Rubric questions used to annotate the papers 1. Does the paper define or describe value? 2. Does the paperâs theory of value subscribe to value monism or pluralism? Monism holds that there is a single, intrinsic dimension to value, while pluralism states that there is more than one dimension or consideration of value, which cannot be flattened. 3. Does the paper consider values measurable or immeasurable? Measurable values can be quantified, operationalized, and recorded as data with negligible loss of precision, while immeasurable values cannot be precisely operationalized and quantified. 4. Does the paper consider values as revealed preferences, or prescribed principles? Preferences are measured from behavioral decisions, and may involve some process of ranking preferred actions, outcomes, or things, while principles-based values correspond to abstract ideals or standards, e.g. valuing beauty, truth, knowledge, human rights declarations, etc. 5. Is the approach to values abstract or concrete? Abstract approaches can be theoretical, idealized, philosophical, name no specific values in the paper but present a set of principles for extracting them, whereas concrete approaches may pertain to a specific existing system, culture, or group. 6. Are values apparent in individual interactions, or formed from a collective of agents? That is, are values observed, measured, or determined from individual or one-on-one interactions, or do they arise over communities, cultures, nations, or groups? 7. Are values dynamic or static? Dynamic values can change over time or in different situations, while static values do not change (or the change of values is not explained). 8. Is the characterization of values thin or thick? We take a âthinâ characterization of value to be entirely instrumental, evaluative and calculative, while a âthickâ characterization is descriptive and situated in a societal context. 9. Does the paper assume or explicitly state that values can emerge in AI autonomously from its creators? 10. Does the paper take the position that AI should follow human values? A normative question which asks if it is morally justified, ethically correct, and/or strategically important for the field to develop AI systems which âupholdâ values, resembling the processes by which humans uphold values (according to the paperâs implicit or explicit theory of value). 11. Does the paper identify which humansâ values the research is aligning to? 12. When preferences are elicited from human annotators, does the paper provide information about the sample size of raters? 13. Does the paper use utility maximization as an approach to value alignment? Annotation process We ensured that each paper was annotated by two annotators among the five of the co-authors of this paper. We used a Google form with the rubrics, and filled them in after reading each paper. We divided the papers into groups so that a subset of annotators read approximately fifty papers each. After an initial pass at reading the collected papers, we isolated disagreements and discussed them through resolutions and consensus. This was, however, not done for every disagreement and some disagreement remains among the annotations. However, we achieved a high degree of inter-annotator agreement, which averages overall around 85%. Qualitative and interpretive analysis In addition to scoring the papers according to our rubric for a quantitative estimate of the prevalence of the concepts, we also collected quotes from each paper if they addressed human values in non-mathematical terms. We did this in order to characterize the beliefs about human values stated, where applicable, by the authors of the papers we analyze. Findings Quantitative results Inter-rater agreement In Table 1 we report measures of inter-rater agreement. Table 1: Inter-Rater Reliability Statistics Question Mean Îș Mean PABAK α (Alpha) Agreement # Pairs Are values measurable? 0.541 0.775 0.269 88.8% 4 Revealed preferences or prescribed principles? 0.467 0.710 0.373 85.5% 4 Sources for determining values 0.411 0.167 0.023 58.3% 4 Abstract or Concrete approach 0.378 0.619 0.235 81.0% 4 Uses utility maximization? 0.365 0.685 0.428 84.2% 4 Can values emerge in AI? 0.350 0.386 0.331 69.3% 4 States which humansâ values? 0.302 0.601 0.219 80.1% 4 Which humansâ values? 0.299 0.736 0.000 86.8% 4 Think vs Thick. 0.239 0.056 0.073 52.8% 4 Does the paper define value? 0.221 0.501 0.247 75.1% 4 Monism vs Pluralism 0.218 0.079 0.349 54.0% 4 Values Static or Dynamic? 0.133 0.021 0.204 51.0% 4 Should AI follow human values? 0.090 0.138 -0.015 56.9% 4 Size of rater pool 0.056 0.275 0.348 63.7% 4 Individual vs Collective interactions 0.028 0.215 0.286 60.8% 4 Notes: PABAK = 2ĂAgreementâ12ĂAgreement-1 (useful when class distributions are skewed); Krippendorffâs α considers all raters simultaneously and handles missing data. Limitations of quantitative analysis: we report raw agreement, Cohenâs kappa, Krippendorffâs α and Bias-Adjusted Kappa (PABAK). Given that, we do expect high agreement on the binary yes/no questions, but some disagreement on extent or categorical questions; these results could indicate issues with design or an attribute of the setting in which we are not the research subjects. Some subjectivity may not be fully specified by the rubric questions. We may have different criteria for what counts as answers to our questions treating papers as the annotation item. We are not claiming that we have developed an objective survey (Aroyo and Welty 2015). We as the annotators are not the research subjects, which is one methodological assumption baked into inter-rater reliability; another is that the raters are interchangeable, which we are not. We present visualizations of the quantitative results in Figure 1 and Figure 2. Figure 1: Summary of answers to our rubric aggregated across all raters for binary yes/no answers. Figure 2: Summary of answers to our rubric aggregated across all raters for categorical questions. INS = âinformation not statedâ Qualitative Analysis We found that 79% of the papers in our dataset neither defined nor described what they mean by âvaluesâ despite ostensibly seeking to align models with values. In many papers, it was common to encounter the word âvaluesâ used interchangeably with âpreferencesâ and âintentions,â and as we address below, elision between values and preferences has become common in this subfield. In one rare example, values were also interchangeably used with âideologiesâ (Kirk et al. 2023a). In another, researchers trained a model on a corpus they built using Chinese laws and morality texts (Xu et al. 2023). Despite largely not defining what values mean, 90% of the papers treat values as measurable. This most often originated with different means of determining human preferences. Typically, researchers created datasets where LLMs were fed a prompt, generated two responses, and used either human raters, other LLMs (sometimes referred to as âautoratersâ) (Lu et al. 2024; Zheng et al. 2024; Chakraborty et al. 2024; Richemond et al. 2024), or a mix of both (Yu et al. 2023; Liu et al. 2024; Wang et al. 2024a) to evaluate which of the two options better adhered to researchersâ choice of alignment criteria; in other cases, one response might be treated as âchosenâ while the other is ârejected.â Some papers drew from preference datasets and survey data that researchers who were not the authors of these papers generated (Zheng et al. 2024; Chakraborty et al. 2024; Lou et al. 2024; Durmus et al. 2023; Zhao et al. 2023; Miranda et al. 2024), while others identified preferences they sought to optimize for including âconciseness..being humorous, philosophical, sycophantic, helpful, concise, creative, formal, expert, pleasant, and upliftingâ (Zhong et al. 2024). Preferences have become such a dominant framework that 82% of the papers treat preferences as a stand-in for âvalues.â By contrast, 9% treated principles as a basis of values, for instance drawing from social choice theory (Siththaranjan et al. 2023; Chakraborty et al. 2024). In papers where researchers created preference datasets through reliance on human annotators to rate LLMsâ paired responses to prompts, only 29 of the 94 papers sampled provided a number for how many people comprised the rater pools. The number of raters tended to be <100<100 people, with the exception of one paper that used a 13,000-person rater pool (Köpf et al. 2023). When human rater pools were used, 77% of that subset of papers did not state the size of the rater pool. A growing debate within alignment is the extent to which pluralism of values can be achieved. One paper acknowledged that âdifferent humans have different values, as it is nearly impossible to train a new large language model from scratch for individual preferenceâ (Zhou et al. 2024). Despite this, most papers took individual rather than collective approaches to values. Rather than selecting values that are already practiced across societies, states, and other collectives, they rely on individual user interactions with LLMs (or LLM-to-LLM interactions), and in one case on collectives of 4-5 people (Bakker et al. 2022). In addition, 53% of papers treated values as static, whether implicitly or through direct acknowledgment of doing so despite recognizing that values evolve (Liu et al. 2023). We found 66% of papers presented âthinâ descriptions and examples of what they meant by values. By contrast, âthickâ descriptions (which accounted for 3 percent) draw from relational, observed values that are difficult to cleave apart from the lived context of the people who enact these values (Geertz 2008). Likewise, 91% of the papers did not refer to specific groups of humans whose values should be followed, underscoring the taken-for-granted idea of values as universal. One paper noted âWe claim that the approach described is agnostic to the ethical paradigm, the userâs preferences, and the legal or social framework, provided we can supply enough feedbackâ (Leike et al. 2018), while another professed âNo need to solve human values. We assume we do not need to solve hard philosophical questions of human values and value aggregation before we can align a superhuman researcher model well enough that it avoids egregiously catastrophic outcomesâ (Burns et al. 2023). Another reckoned with the difficulties of treating preference datasets as representative: âThis procedure aligns the behavior of GPT-3 to the stated preferences of a specific group of people (mostly our labelers and researchers), rather than any broader notion of âhuman valuesââ (Ouyang et al. 2022). Finally, we noticed a trend in which a few papers suggest that AI models can develop their own value systems (Ngo et al. 2022) and âinternal meta-objectivesâ (Phelps and Ranson 2023b). Some researchers have begun to refer to modelsâ âself-alignmentâ (Guo et al. 2024; Li et al. 2024) described as when âa new paradigm emerges where LLMs can achieve value alignment by themselves⊠transforming an unaligned LLM into one adhering to societal norms, independently of external resourcesâ (Pang et al. 2024). Discussion Our systematic review of 94 research papers on AI value alignment reveals how the field is erasing human inputs from an endeavor whose success rests on purportedly steering AI systems to respect human values. Values are core components of how societies organize themselves; any attempt at value alignment is inextricable from society (Graeber 2001; Haslanger 2023; Anderson 1995; Zhi-Xuan et al. 2024; Birhane et al. 2022). Recent social philosophy argues that societies are complex systems â or clusters of interacting systems â that reproduce themselves: their hierarchies, culture, and structures. These systems are also dynamic and constantly evolving (Haslanger 2023). Yet we observe how research in this field abstracts away questions of whose values are represented, or what qualities constitute pluralism and dynamism in values, instead relegating these fundamental, unavoidable concerns to be taken up by others (if at all). The risk of repeatedly delaying explorations of what values are and how they can be accountably represented is that certain practices of design, deployment, and evaluation will ossify. At worst, this may result in normalizing widespread use of AI systems that have been trained on a narrow conception of value and claiming that this is the best the field can offer. The pressure to neglect these essential debates is reflected in a subset of papers that treat alignment as a step toward achieving âartificial superintelligenceâ (Kim et al. 2024), or as an essential component of averting catastrophic and existential risks from AI (Shen et al. 2023; Ngo et al. 2022; Burns et al. 2023). This typifies the turn toward what we call âalignment without humans,â or treatment of humans as an unreliable source of information on their own preferences (Miranda et al. 2024; Wang et al. 2024a), as too expensive (Wang et al. 2024c; Zheng et al. 2024; Liu et al. 2024) and time-consuming (Dai et al. 2023) to use, therefore justifying the use of synthetic means of preference elicitation, feedback, and evaluations (Hong et al. 2024; Cui et al. 2024; Tian et al. 2024). We observe studies that rely on no human input for pre-determining value alignment criteria (Mei et al. 2024; Chen et al. 2024), often through the use of simulated humans (Chakraborty et al. 2023; Poddar et al. 2024; Liu et al. 2023). Perhaps in part due to the desire to race toward solving âthe alignment problem,â one trend that arose from the corpus of papers is the dominance of Anthropicâs Helpful, Harmless, Honest (H) approach. This ranged from papers that explicitly used the H-RLHF dataset for training and/or evaluations (Yu et al. 2023; Lou et al. 2024; Huang et al. 2023; Zheng et al. 2024; Zeng et al. 2024), to others that name H as a guiding principle (Ding et al. 2024; Gao et al. 2024; Lu et al. 2024; Shen et al. 2023; Shi et al. 2024). Graeberâs critique of applying economic theory to diverse cultures is apt for AI value alignment: âIt [economics] also has the advantage of joining an extremely simple model of human nature with extremely complicated mathematical formulae that non-specialists can rarely understand, much less criticize.â (Graeber 2001). Our review of the literature reveals how RLHF and other technical alignment methods likewise join extremely simplistic models of human nature with complicated mathematical formulae that are difficult for non-specialists to understand or critique. Even though the majority of papers in our dataset do not define what they mean by values, adopting the framework of economic preference optimization to align models necessarily requires a commitment to a set of beliefs about what human values are. At a minimum this entails the belief that values are individual, rational, monist, static, maximizing utility functions, and abstract. The idea is that an agent that obeys some minimal conditions of rationality can be modeled as if the agent has an ordered set of preferences along with some probabilistic beliefs about what states of the world and itself will maximize the agentâs utility (however defined) (Shea 2018). This particular set of assumptions about human nature and human values, drawn largely from economic decision theory and rational choice theory, has been empirically and theoretically challenged from multiple perspectives within and outside of economics (Chibnik 2011; Graeber 2011, 2001; Gigerenzer 1997; Frank et al. 2024) and AI alignment (Bakker et al. 2022; Kasirzadeh 2024a; Zhi-Xuan et al. 2024). Rational or social choice theory do not describe the values or behavior of real human beings or cultures, but only the behavior of highly abstract, reductionist, and idealized fictional agents (Frank et al. 2024). Rather than asking âwhat do human beings actually care about?â, rational choice theory asks - and presumes there is a normatively correct answer to - âwhat should human beings do in order to make the utility-maximizing decision?â This meta-theoretical commitment risks foreclosing alternative conceptions of human values, and limits the ability of AI systems to align to diverse cultural values (Prabhakaran et al. 2022; Khan et al. 2025; Gabriel 2020). Poddar et al. point out, âCurrent RLHF approaches rely on a prescriptive set of values curated by a small set of AI researchers. Moreover, they typically assume that all end-users share the same set of values. Given the concerning lack of diversity in AI, it is clear that this approach cannot account for the range of social, moral, and political values that inform preferences in human populationsâ (Poddar et al. 2024). We call on the field to take seriously the fundamental weaknesses of utility maximization, rational choice theory and reductionist economic frameworks in AI alignment. Utility maximization Economists and philosophers have generally assumed that people are trying to maximize something: money, or love, or sometimes something else (most often, expected utility) with their choices (Glimcher 2022; Graeber 2001).Utility maximization treats all human behavior as the maximization of expected utility from a stable set of preferences, and holds that humans accumulate an optimal amount of information in markets (Chibnik 2011; Russell and Norvig 1995; Becker 1976). Thus, preferences are seen as being derived from a single utility function that each individual is somehow computing in their brain (Gigerenzer 1997). Taken to its extreme, utility theory holds that our behavior is entirely determined by maximizing this utility function, and preferences are assumed to not change substantially over time, nor to be very different between people of different socioeconomic classes, societies, and cultures (Becker 1976). Finally, AI alignment assumes that an individualâs reward function (and therefore their values), or preferences about the future, can be inferred, or approximated, by observing how humans behave by collecting data on how humans interact with AI outputs (Hadfield-Menell et al. 2016). Despite the empirical inadequacy of this ambiguous economic approach to AI and value alignment, creating an artificial agent that embodies economic theory has been the almost unquestioned north star in the field until recently (Zhi-Xuan et al. 2024; Gabriel 2020; Russell 2019). Influential work on value alignment explicitly argues that alignment should be formulated as a cooperative and interactive reward maximization process (Hadfield-Menell et al. 2016), with most alignment work implicitly adopting this approach (Christiano et al. 2017b; Ouyang et al. 2022; Wang et al. 2024b) the vast majority of technical approaches to alignment from our sample of papers implementing some form of utility maximization. These often unstated philosophical commitments to an economic worldview determine the methodology that alignment researchers use to measure and model human values. We make this often implicit metatheorectical commitment explicit so that it can be critiqued and improved (Maxwell 2005). We also evaluated our sample of papers for whether they adopted a framework of utility maximization, discovering that 87 % of papers implicitly or explicitly use utility maximization as part of the technical framework to align models. Adopting the framework of utility maximization necessarily involves normative assumptions that are value-laden (Graeber 2001). However, given that utility maximization is an inherent part of reinforcement learning from human feedback (RLHF) and reward modeling it is not surprising that the majority of alignment papers in our sample adopt this framework. This also stems from framing AI alignment in terms of the âprincipal-agentâ framework from economic theory (Phelps and Ranson 2023a; Stanczak et al. 2025). Moreover, utility maximization is cited as a property of models, e.g., (Mazeika et al. 2025) notes âit was shown that recent LLMs have structurally coherent, broad value systems. As they become more capable, their value systems increasingly conform to the axioms of utility theory, meaning they can be described as maximizing a utility function.â Given that models are trained and fine-tuned according to the axioms of rational choice theory, should it be surprising that they behave according to utility theory? A small number of the papers recognize the inherent limitations of utility maximization. (Phelps and Ranson 2023a) points out: âRather than seeking to impose a monolithic utility function on artificial agents, we propose a strategy of reducing information asymmetry and aligning interests through external incentives, much like the approach used in traditional economic solutions to principal-agent problems. âFurthermore, (Poddar et al. 2024) writes, âThese insights suggest that human preferences are not derived from a single utility function, but are affected by unobserved, hidden user context.â One of the starting points of our meta-review is Zhi-Xuan et al. (2024)âs critique of the reliance on eliciting preferences as representations of human values, in which they ask: âWhat would AI alignment look like if it took these challenges seriously? It would move away from naive rational choice models of human decision making, towards richer models that include how we evaluate, commensurate, and act upon our values in boundedly rational ways. It would no longer take for granted expected utility theory, and instead explore systems for reasoning about the normativity of our preferences and values.â Alternative conceptions of value âItâs values all the way downâ - Alondra Nelson (Nelson 2023) The narrow economic conception of human values that determines the methodology and philosophy of alignment is by no means the only way to understand what humans value. As anthropologists point out, for 99 % of humanityâs history we did not live in market economies (Graeber 2011). Economic theory implicitly posits that weâve always been capitalists and that capitalism is somehow natural (Hornborg 1998). We argue that even recent work on pluralistic alignment is still fundamentally rooted in the orthodox economic framework. However, this is a distorted view of both societies and individual humans. Part of the challenge of rethinking alignmentâs approach to value comes from the social positioning of the field that âdoing anything is better than nothing,â without recognizing that over-utilizing narrow disciplinary approaches has the potential to make matters worse (Dai 2025). A fundamental limitation of existing alignment methods is that they reinforce shallow behavioral dispositions rather than endowing LLMs with a genuine capacity for normative deliberation (MilliĂšre 2025). We find that the abstractions and idealizations which alignment postulates to describe human values become confused for real phenomena. Papers in our sample did in some cases engage directly with alternatives to utility theory, for example Huang et.al., (2025) who state, âWe are interested in values not as abstract entities, but as operational priorities that influence how the system navigates its possible space of outputs.â We expand othersâ calls to expand alignment approaches (Casper et al. 2023; Lindström et al. 2024; Zhi-Xuan et al. 2024) with our exploration of psychological and anthropological theories of value. Cognitive or psychological theories of value Papers in our dataset referenced psychological theories of value that are accepted in mainstream psychology literature, such as Schwartzâs theory of basic values (Sierra et al. 2021; Schwartz 2012). Schwartz theorizes value as a set of beliefs, closely linked to emotions, which inform the selection of actions. He further claims that there is a fixed set of âbasic valuesâ or value categories, which are universal across cultures, and tend to remain fixed over time within an adult individual. These values are conceptualized as guiding principles which can come into conflict, but are ultimately resolved through a hierarchy of importance. This leads to our characterization of the psychological theory of value as ultimately monistic. This individualistic view of values is challenged by empirical psychological and ethnographic work which shows a much more complex, situated, holistic and dynamic view of values (Chibnik 2011; Anderson 1995; Haslanger 2023). This field is also complex, and we encourage engagement with existing debates in the cross-cultural study of values (Schwartz 2012). Anthropological, situated theories of value The field with perhaps the most theoretical work and empirical data on human values - anthropology - has to date been ignored in AI alignment (cf. (Schmer-Galunder 2026). Graeber suggests that what people call âvalueâ may be thought of as âthe way people represent the importance of their own actions to themselves, as reflected in one or another socially recognized formâ (Graeber 2001). Economic theoryâs view of human nature is cynical, assuming that humans are entirely self-interested and will calculate the most effective way to obtain what they desire (Becker 1976). These are values derived from living in market-based societies (Hornborg 2014). However, markets are not natural phenomena that arise spontaneously out of large complex societies as the economic narrative goes, but rather they require the enforcement of state violence to impose property rights, coercion, and the rule of law (Graeber 2011; Mau 2023). Economic anthropology recognizes that economic notions of value, articulated in objective monetary representations and utility theory, and ethical notions of value, articulated in moral terms, are âinextricably connectedâ (Lambek 2013). Graeber (Graeber 2001) outlines three ways in which values have been theorized in the social sciences that go beyond the narrow conception of value in the economic sense: 1. âvaluesâ in the sociological sense: conceptions of what is ultimately good, proper, or desirable in human life 2. âvalueâ in the economic sense: the degree to which objects are desired, particularly, as measured by how much others are willing to give up to get them 3. âvalueâ in the linguistic sense, or how meaning is derived from the differences in the usage of words. However, in our dataset it is unusual for researchers to be specific about which humans, which values, and which cultures they target for alignment. We follow (Arzberger et al. 2024), who argue that values acquire situated meaning; hence, their concrete interpretations, means of actualization, and hierarchies vary between people depending on the situations in which they guide the notion of what people really value. This is closely aligned with situated-constructivist views of emotions, which see emotions not as basic universal categories, but as arising in specific socially- and culturally- salient circumstances. In other words, emotions are a category populated with highly variable instances (Barrett 2017). Thus we can analogously we think of human values as a category populated with highly variable instances tied to specific cultural circumstances. Limitations One limitation of this work is the relatively small sample size of papers in our dataset compared with the vast number of research papers on this topic. This is especially challenging given the widespread practice of rapidly uploading research to arXiv, avoiding peer-review (Gyevnar and Kasirzadeh 2025). However, we attempt to correct for this by estimating the relative impact of our dataset on the field using both quantitative and qualitative measures. While we aimed to gather a representative subsample of available papers on AI value alignment, we make no claims to the completeness or comprehensiveness of the sampled papers. The sample size was also limited due to time constraints on our small team of co-authors. We also acknowledge that the scoring rubric has many limitations. The questions are highly nuanced and difficult to boil down to a binary or even ternary answer. We considered a multi-point Likert scale, which was ultimately abandoned due to concerns about calibrating relative scales. Implications for future work We suggest building on our analysis to apply concepts from humanistic social sciences such as anthropology and psychology to operationalize human values in AI in a more representative and holistic way that explicitly accounts for: concreteness/abstraction, the thick/thin distinction, distinctions between monism and pluralism, and embracing a view of values not as abstract universal things human possess, but as constructed in specific social and cultural contexts (Kasirzadeh 2024a; Sorensen et al. 2024; Benkler et al. 2023; Arzberger et al. 2024). To deepen the critique of âpreferentism,â we also recommend an examination of the content of prompts and paired responses in preference training datasets, as many contain largely nonsensical conversations with an LLM and involve choosing the less absurd of two unrealistic examples in each pair. Communicating the shaky foundations of these unquestioned defaults is of particular value to informing policymakers and the public about the limits of value alignment. Moreover, it can contribute to more effective sociotechnical and policy measures derived from realistic conceptions of AIâs risks. Conclusion We began with an examination of the theory of value(s) to which AI value alignment research subscribes, drawing from interdisciplinary philosophy and social scientific research on human values to develop a schema for answering this question. We argued that AI alignment lacks a concrete theory of human values and implicitly adopts beliefs about human values rooted in economic theory. Through our review, we have shown that the assumptions guiding technical AI alignment practices limit the possible range of human values with which AI can be aligned by translating alignment into utility theory. As Lisa Feldman Barrett argues, âEven a scientist who believes they are purely driven by data is doing philosophyâ (Barrett and Theriault 2025). Tech companies and AI value alignment researchers enact philosophy every day as they engage in value alignment work when choosing training data, setting up model architectures, and especially when running RLHF experiments. Having laid bare AI value alignmentâs philosophical commitments, we invite researchers from a wider variety of disciplines to bring greater specificity and under-explored perspectives into the debate on whether and how AI can address human values. Generative AI Usage Statement Generative AI was not used in the research or writing of this publication. References S. Ahmed, K. JaĆșwiĆska, A. Ahlawat, A. Winecoff, and M. Wang (2024) Field-building and the epistemic culture of ai safety. First Monday. Cited by: Alignment and its discontents. R. Alvarado (2023) AI as an epistemic technology. Science and Engineering Ethics 29 (5), p. 32. Cited by: What is value alignment?. E. Amirloo, J. Fauconnier, C. Roesmann, C. Kerl, R. Boney, Y. Qian, Z. Wang, A. Dehghan, Y. Yang, Z. Gan, et al. (2024) Understanding alignment in multimodal llms: a comprehensive study. arXiv preprint arXiv:2407.02477. Cited by: What is value alignment?, Preferences are not all you need. E. Anderson (1995) Value in ethics and economics. 2. print edition, Harvard Univ. Press, Cambridge, Mass. (eng). External Links: ISBN 9780674931909 9780674931893 Cited by: Cognitive or psychological theories of value, Discussion. Anthropic (2025) Anthropic Alignment Research Team. (en-US). External Links: Link Cited by: Alignment and its discontents. L. Aroyo and C. Welty (2015) Truth is a lie: crowd truth and the seven myths of human annotation. AI Magazine 36 (1), p. 15â24. Cited by: Table 1. M. Arvan (2024) âInterpretabilityâand âalignmentâare foolâs errands: a proof that controlling misaligned large language models is the best anyone can hope for. AI & SOCIETY, p. 1â16. Cited by: Alignment and its discontents, Preferences are not all you need. A. Arzberger, S. Buijsman, M. L. Lupetti, A. Bozzon, and J. Yang (2024) Nothing comes without its worldâpractical challenges of aligning llms to situated human values through rlhf. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, Vol. 7, p. 61â73. Cited by: Introduction, Alignment and its discontents, Anthropological, situated theories of value, Implications for future work. A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al. (2021) A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861. Cited by: Introduction, What is value alignment?, 2nd item. Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan (2022) Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. arXiv. Note: arXiv:2204.05862 External Links: Link, Document Cited by: What is value alignment?, 3rd item. M. Bakker, M. Chadwick, H. Sheahan, M. Tessler, L. Campbell-Gillingham, J. Balaguer, N. McAleese, A. Glaese, J. Aslanides, M. Botvinick, et al. (2022) Fine-tuning language models to find agreement among humans with diverse preferences. Advances in Neural Information Processing Systems 35, p. 38176â38189. Cited by: Preferences are not all you need, Qualitative Analysis, Discussion. L. F. Barrett and J. Theriault (2025) Whatâs real? a philosophy of science for social psychology. The handbook of social psychology. Cited by: Conclusion. L. F. Barrett (2017) The theory of constructed emotion: an active inference account of interoception and categorization. Social cognitive and affective neuroscience 12 (1), p. 1â23. Cited by: Anthropological, situated theories of value. G. S. Becker (1976) The economic approach to human behavior. Vol. 803, University of Chicago press. Cited by: Preferences are not all you need, Justification for rubric questions, Anthropological, situated theories of value, Utility maximization. N. Benkler, D. Mosaphir, S. Friedman, A. Smart, and S. Schmer-Galunder (2023) Assessing llms for moral value pluralism. arXiv preprint arXiv:2312.10075. Cited by: Implications for future work. S. Bergman, N. Marchal, J. Mellor, S. Mohamed, I. Gabriel, and W. Isaac (2024) STELA: a community-centred approach to norm elicitation for ai alignment. Scientific Reports 14 (1), p. 6616. Cited by: Alignment and its discontents. A. Birhane, P. Kalluri, D. Card, W. Agnew, R. Dotan, and M. Bao (2022) The Values Encoded in Machine Learning Research. In 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul Republic of Korea, p. 173â184 (en). External Links: ISBN 978-1-4503-9352-2, Link, Document Cited by: Discussion. C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y. Chen, A. Ecoffet, M. Joglekar, J. Leike, I. Sutskever, and J. Wu (2023) Weak-to-strong generalization: eliciting strong capabilities with weak supervision. . External Links: Link Cited by: Qualitative Analysis, Discussion. S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al. (2023) Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint arXiv:2307.15217. Cited by: Preferences are not all you need, Alternative conceptions of value. S. Chakraborty, J. Qiu, H. Yuan, A. Koppel, F. Huang, D. Manocha, A. S. Bedi, and M. Wang (2024) MaxMin-rlhf: alignment with diverse human preferences. . External Links: Link Cited by: Qualitative Analysis, Qualitative Analysis. S. Chakraborty, A. Singh, A. Bhaskar, P. Tokekar, D. Manocha, and A. S. Bedi (2023) REBEL: reward regularization-based approach for robotic reinforcement learning from human feedback. . External Links: Link Cited by: Discussion. S. Chaudhari, P. Aggarwal, V. Murahari, T. Rajpurohit, A. Kalyan, K. Narasimhan, A. Deshpande, and B. C. da Silva (2024) RLHF deciphered: a critical analysis of reinforcement learning from human feedback for llms. arXiv preprint arXiv:2404.08555. Cited by: What is value alignment?. Z. Chen, K. Zhou, W. X. Zhao, J. Wang, and J. Wen (2024) Low-redundant optimization for large language model alignment. . External Links: Link Cited by: Discussion. M. Chibnik (2011) Anthropology, Economics, and Choice. University of Texas Press. External Links: ISBN 978-0-292-72676-5, Link, Document Cited by: Preferences are not all you need, Cognitive or psychological theories of value, Utility maximization, Discussion. P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei (2017a) Deep Reinforcement Learning from Human Preferences. In Advances in Neural Information Processing Systems, Vol. 30. External Links: Link Cited by: What is value alignment?. P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei (2017b) Deep reinforcement learning from human preferences. Advances in neural information processing systems 30. Cited by: 1st item, Utility maximization. G. Cui, L. Yuan, N. Ding, G. Yao, W. Zhu, Y. Ni, G. Xie, Z. Liu, and M. Sun (2024) UltraFeedback: boosting language models with high-quality feedback. . External Links: Link Cited by: Discussion. J. Dai (2025) The Artificiality of Alignment. (en-US). External Links: Link Cited by: Alignment and its discontents, Alternative conceptions of value. J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang (2023) Safe rlhf: safe reinforcement learning from human feedback. . External Links: Link Cited by: Discussion. Y. Ding, B. Li, and R. Zhang (2024) ETA: evaluating then aligning safety of vision language models at inference time. . External Links: Link Cited by: Discussion. R. Dobbe (2022) System safety and artificial intelligence. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, p. 1584â1584. Cited by: Alignment and its discontents. E. Durmus, K. Nguyen, T. I. Liao, N. Schiefer, A. Askell, A. Bakhtin, C. Chen, Z. Hatfield-Dodds, D. Hernandez, N. Joseph, L. Lovitt, S. McCandlish, O. Sikder, A. Tamkin, J. Thamkul, J. Kaplan, J. Clark, and D. Ganguli (2023) Towards measuring the representation of subjective global opinions in language models. . External Links: Link Cited by: Qualitative Analysis. A. Frank, M. Gleiser, and E. Thompson (2024) The blind spot: why science cannot ignore human experience. The MIT press, Cambridge, Mass (eng). External Links: ISBN 978-0-262-04880-4 Cited by: Discussion. I. Gabriel (2020) Artificial Intelligence, Values, and Alignment. Minds and Machines 30 (3), p. 411â437 (en). External Links: ISSN 1572-8641, Link, Document Cited by: Introduction, Alignment and its discontents, Preferences are not all you need, Utility maximization, Discussion. C. Gao, Y. Wu, D. Chen, Q. Zhang, Fu,Zhengyan, Y. Wan, L. Sun, and X. Zhang (2024) HonestLLM: toward an honest and helpful large language model. . External Links: Link Cited by: Discussion. T. Gebru and Ă. P. Torres (2024) The tescreal bundle: eugenics and the promise of utopia through artificial general intelligence. First Monday. Cited by: Alignment and its discontents. C. Geertz (2008) âThick Description: Toward an Interpretive Theory of Cultureâ. In The Cultural Geography Reader, Note: Num Pages: 11 External Links: ISBN 978-0-203-93195-0 Cited by: Justification for rubric questions, Qualitative Analysis. G. Gigerenzer (1997) Bounded rationality: models of fast and frugal inference. Swiss Journal of Economics and Statistics 133 (2/2), p. 201â218. Cited by: Utility maximization, Discussion. T. Gillespie, R. Shaw, M. L. Gray, and J. Suh (2024) AI red-teaming is a sociotechnical system. now what?. arXiv preprint arXiv:2412.09751. Cited by: Introduction. P. W. Glimcher (2022) Efficiently irrational: deciphering the riddle of human choice. Trends in cognitive sciences 26 (8), p. 669â687. Cited by: Utility maximization. D. Graeber (2001) Toward an Anthropological Theory of Value: The False Coin of Our Own Dreams. Palgrave-Macmillan. Cited by: Anthropological, situated theories of value, Anthropological, situated theories of value, Utility maximization, Utility maximization, Discussion, Discussion, Discussion. D. Graeber (2011) Debt: the first 5,000 years. Melville House, Brooklyn, N.Y. External Links: ISBN 978-1-933633-86-2 Cited by: Anthropological, situated theories of value, Alternative conceptions of value, Discussion. O. Guest (2026) What does âhuman-centred aiâmean?. Behavioral Sciences 16 (4), p. 583. Cited by: Introduction. H. Guo, Y. Yao, W. Shen, J. Wei, X. Zhang, Z. W. Wang, and Y. Liu (2024) Human-instruction-free llm self-alignment with limited samples. . External Links: Link Cited by: Qualitative Analysis. B. Gyevnar and A. Kasirzadeh (2025) AI safety for everyone. Nature Machine Intelligence, p. 1â12. Cited by: Limitations. D. Hadfield-Menell, A. Dragan, P. Abbeel, and S. Russell (2016) Cooperative inverse reinforcement learning. In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPSâ16, Red Hook, NY, USA, p. 3916â3924. External Links: ISBN 978-1-5108-3881-9 Cited by: What is value alignment?, Utility maximization, Utility maximization. S. Haslanger (2022) Failures of methodological individualism: the materiality of social systems.. Journal of Social Philosophy 53 (4). Cited by: Justification for rubric questions. S. Haslanger (2023) Situated Knowledge and Situated Values. (en). Cited by: Justification for rubric questions, Cognitive or psychological theories of value, Discussion. V. Hofmann, P. R. Kalluri, D. Jurafsky, and S. King (2024) AI generates covertly racist decisions about people based on their dialect. Nature 633 (8028), p. 147â154 (en). External Links: ISSN 1476-4687, Link, Document Cited by: What is value alignment?. I. Hong, Z. Li, A. Bukharin, Y. Li, H. Jiang, T. Yang, and T. Zhao (2024) Adaptive preference scaling for reinforcement learning with human feedback. Advances in Neural Information Processing Systems 37, p. 107249â107269. Cited by: Discussion. S. Hooker (2025) On the slow death of scaling. Available at SSRN 5877662. Cited by: What is value alignment?. A. Hornborg (1998) Ecological embeddedness and personhood: have we always been capitalists?. Anthropology today 14 (2), p. 3â5. Cited by: Alternative conceptions of value. A. Hornborg (2014) Technology as Fetish: Marx, Latour, and the Cultural Foundations of Capitalism. Theory, Culture & Society 31 (4), p. 119â140 (en). Note: Publisher: SAGE Publications Ltd External Links: ISSN 0263-2764, Link, Document Cited by: Anthropological, situated theories of value. Y. Huang, S. Gupta, M. Xia, K. Li, and D. Chen (2023) Catastrophic jailbreak of open-source llms via exploiting generation. . External Links: Link Cited by: Discussion. U. A. S. Institute (2025) UK AI Security Institute Alignment Project. (en-US). External Links: Link Cited by: What is value alignment?. J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, J. Zhou, Z. Zhang, et al. (2023) Ai alignment: a comprehensive survey. arXiv preprint arXiv:2310.19852. Cited by: Introduction, What is value alignment?, Preferences are not all you need. G. M. Johnson (2023) Are Algorithms Value-Free?. Journal Moral Philosophy 21 (1-2), p. 1â35. External Links: Link, Document Cited by: Preferences are not all you need. A. Kasirzadeh (2024a) Plurality of value pluralism and AI value alignment. (en). External Links: Link Cited by: Justification for rubric questions, Discussion, Implications for future work. A. Kasirzadeh (2024b) Two types of ai existential risk: decisive and accumulative. arXiv preprint arXiv:2401.07836. Cited by: What is value alignment?. A. Khan, S. Casper, and D. Hadfield-Menell (2025) Randomness, not representation: the unreliability of evaluating cultural alignment in llms. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, p. 2151â2165. Cited by: Discussion. H. Khlaaf (2023) Toward comprehensive risk assessments and assurance of ai-based systems. Trail of Bits 7. Cited by: Alignment and its discontents. H. Kim, X. Yi, J. Yao, J. Lian, M. Huang, S. Duan, J. Bak, and X. Xie (2024) The road to artificial superintelligence: a comprehensive survey of superalignment. . External Links: Link Cited by: Discussion. H. R. Kirk, B. Vidgen, P. Röttger, and S. A. Hale (2023a) Personalisation within bounds: a risk taxonomy and policy framework for the alignment of large language models with personalised feedback. . External Links: Link Cited by: Qualitative Analysis. H. R. Kirk, B. Vidgen, P. Röttger, and S. A. Hale (2023b) The empty signifier problem: towards clearer paradigms for operationalisingâ alignmentâ in large language models. arXiv preprint arXiv:2310.02457. Cited by: Alignment and its discontents. T. Q. Klassen, P. A. Alamdari, and S. A. McIlraith (2024) Pluralistic Alignment Over Time. (en). External Links: Link Cited by: Preferences are not all you need, Justification for rubric questions. A. Köpf, Y. Kilcher, D. Von RĂŒtte, S. Anagnostidis, Z. R. Tam, K. Stevens, A. Barhoum, D. Nguyen, O. Stanley, R. Nagyfi, et al. (2023) Openassistant conversations-democratizing large language model alignment. Advances in neural information processing systems 36, p. 47669â47681. Cited by: Qualitative Analysis. C. M. Korsgaard (1986) Aristotle and kant on the source of value. Ethics 96 (3), p. 486â505. Cited by: Justification for rubric questions. M. Lambek (2013) The value of (performative) acts. HAU: Journal of Ethnographic Theory 3 (2), p. 141â160. Cited by: Anthropological, situated theories of value. J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, and S. Legg (2018) Scalable agent alignment via reward modeling: a research direction. . External Links: Link Cited by: Qualitative Analysis. X. Li, P. Yu, C. Zhou, T. Schick, L. Zettlemoyer, O. Levy, Z. Luke, J. Weston, and M. Lewis (2024) Self-alignment with instruction backtranslation. . External Links: Link Cited by: Qualitative Analysis. A. D. Lindström, L. Methnani, L. Krause, P. Ericson, Ă. M. d. R. de Troya, D. C. Mollo, and R. Dobbe (2024) AI alignment through reinforcement learning from human feedback? contradictions and limitations. arXiv preprint arXiv:2406.18346. Cited by: Alignment and its discontents, Alignment and its discontents, Preferences are not all you need, Preferences are not all you need, Alternative conceptions of value. J. Liu, D. Ge, and R. Zhu (2024) Reward learning from preference with ties. . External Links: Link Cited by: Qualitative Analysis, Discussion. R. Liu, R. Yang, C. Jia, G. Zhang, D. Zhou, A. M. Dai, D. Yang, and S. Vosoughi (2023) Training socially aligned language models on simulated social interactions. . External Links: Link Cited by: Qualitative Analysis, Discussion. X. Lou, J. Zhang, J. Xie, L. Liu, D. Yan, and K. Huang (2024) SPO: multi-dimensional preference sequential alignment with implicit reward modeling. . External Links: Link Cited by: Qualitative Analysis, Discussion. K. Lu, B. Yu, F. Huang, Y. Fan, R. Lin, and C. Zhou (2024) Online merging optimizers for boosting rewards and mitigating tax in alignment. . External Links: Link Cited by: Qualitative Analysis, Discussion. S. Mau (2023) Mute compulsion: a marxist theory of the economic power of capital. Verso books. Cited by: Anthropological, situated theories of value. N. Maxwell (2005) The comprehensibility of the universe: a new conception of science. Clarendon press, Oxford (eng). External Links: ISBN 978-0-19-926155-0 Cited by: Utility maximization. M. Mazeika, X. Yin, R. Tamirisa, J. Lim, B. W. Lee, R. Ren, L. Phan, N. Mu, A. Khoja, O. Zhang, et al. (2025) Utility engineering: analyzing and controlling emergent value systems in ais. arXiv preprint arXiv:2502.08640. Cited by: Utility maximization. L. Mei, S. Liu, Y. Wang, B. Bi, R. Yuan, and X. Cheng (2024) HiddenGuard: fine-grained safe generation with specialized representation router. . External Links: Link Cited by: Discussion. J. S. Mill (2016) Utilitarianism. In Seven masterpieces of philosophy, p. 329â375. Cited by: Justification for rubric questions. R. MilliĂšre (2025) Normative conflicts and shallow ai alignment: r. milliĂšre. Philosophical Studies, p. 1â44. Cited by: Alternative conceptions of value. L. J. V. Miranda, Y. Wang, Y. Elazar, S. Kumar, V. Pyatkin, F. Brahman, N. A. Smith, H. Hajishirzi, and Dasigi,Pradeep (2024) Hybrid preferences: learning to route instances for human vs. ai feedback. . External Links: Link Cited by: Qualitative Analysis, Discussion. C. Nauer and Z. van Doren (2025) EXISTENTIAL risk from ai. Cited by: What is value alignment?. A. Nelson (2023) FAccTâ23 Keynote: âThick Alignmentâ. (en-US). External Links: Link Cited by: Alternative conceptions of value. R. Ngo, L. Chan, and Mindermann,Sören (2022) The alignment problem from a deep learning perspective. . External Links: Link Cited by: Qualitative Analysis, Discussion. OpenAI (2025) OpenAI Alignment Research Blog. (en-US). External Links: Link Cited by: Alignment and its discontents. L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. (2022) Training language models to follow instructions with human feedback. Advances in neural information processing systems 35, p. 27730â27744. Cited by: What is value alignment?, 4th item, Qualitative Analysis, Utility maximization. X. Pang, S. Tang, R. Ye, Y. Xiong, B. Zhang, Y. Wang, and S. Chen (2024) Self-alignment of large language models via monopolylogue-based social scene simulation. . External Links: Link Cited by: Qualitative Analysis. C. Parker (2019) Snowball sampling. SAGE Research Methods Foundations. Cited by: Snowball sampling. S. Phelps and R. Ranson (2023a) Of models and tin men: a behavioural economics study of principal-agent problems in ai alignment using large-language models. arXiv preprint arXiv:2307.11137. Cited by: Utility maximization, Utility maximization. S. Phelps and R. Ranson (2023b) Of models and tin men: a behavioural economics study of principal-agent problems in ai alignment using large-language models. . External Links: Link Cited by: Qualitative Analysis. S. Poddar, Y. Wan, H. Ivison, A. Gupta, and N. Jaques (2024) Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning. (en). External Links: Link Cited by: Utility maximization, Discussion, Discussion. V. Prabhakaran, M. Mitchell, T. Gebru, and I. Gabriel (2022) A Human Rights-Based Approach to Responsible AI. arXiv. Note: arXiv:2210.02667 [cs]Comment: Presented as a (non-archival) poster at the 2022 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization or (EAAMO â22) External Links: Link, Document Cited by: Discussion. R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn (2023) Direct preference optimization: your language model is secretly a reward model. Advances in neural information processing systems 36, p. 53728â53741. Cited by: Preferences are not all you need. I. D. Raji and R. Dobbe (2023) Concrete problems in ai safety, revisited. arXiv preprint arXiv:2401.10899. Cited by: Alignment and its discontents. P. H. Richemond, Y. Tang, D. Guo, D. Calandriello, M. G. Azar, R. Rafailov, B. A. Pires, E. Tarassov, L. Spangher, W. Ellsworth, A. Severyn, J. Mallinson, L. Shani, G. Shamir, R. Joshi, T. Liu, R. Munos, and B. Piot (2024) Offline regularised reinforcement learning for large language models alignment. . External Links: Link Cited by: Qualitative Analysis. S. J. Russell and P. Norvig (1995) Artificial intelligence: a modern approach;[the intelligent agent book]. Prentice hall. Cited by: Preferences are not all you need, Preferences are not all you need, Utility maximization. S. J. Russell (2019) Human compatible: artificial intelligence and the problem of control. Viking, New York (eng). External Links: ISBN 978-0-525-55861-3 Cited by: Introduction, Utility maximization. S. Schmer-Galunder (2026) Culture in the code: anthropological concepts decoding aiâs hidden assumptions. AI and Ethics 6 (3), p. 272. Cited by: Justification for rubric questions, Anthropological, situated theories of value. S. Schwartz (2012) An Overview of the Schwartz Theory of Basic Values. Online Readings in Psychology and Culture 2 (1). External Links: ISSN 2307-0919, Link, Document Cited by: Cognitive or psychological theories of value. A. Sharma, S. Keh, E. Mitchell, C. Finn, K. Arora, and T. Kollar (2024) A critical evaluation of ai feedback for aligning large language models. arXiv preprint arXiv:2402.12366. Cited by: What is value alignment?. N. Shea (2018) Representation in cognitive science. Oxford University Press. Cited by: Discussion. T. Shen, R. Jin, Y. Huang, C. Liu, W. Dong, Z. Guo, X. Wu, Y. Liu, and Xiong,Deyi (2023) Large language model alignment: a survey. . External Links: Link Cited by: Discussion. Z. Shi, Z. Wang, H. Fan, Z. Zhang, L. Li, Y. Zhang, Z. Yin, L. Sheng, Y. Qiao, and J. Shao (2024) Assessment of multimodal large language models in alignment with human values. . External Links: Link Cited by: Discussion. K. Shrishak (2025) How not to deploy generative AI: the Story of the European Parliament. (en-US). External Links: Link Cited by: What is value alignment?. C. Sierra, N. Osman, P. Noriega, J. Sabater-Mir, and A. PerellĂł (2021) Value alignment: a formal approach. arXiv. Note: arXiv:2110.09240 [cs] version: 1Comment: accepted paper at the Responsible Artificial Intelligence Agents Workshop, of the 18th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS 2019) External Links: Link, Document Cited by: Cognitive or psychological theories of value. A. Siththaranjan, C. Laidlaw, and D. Hadfield-Menell (2023) Distributional preference learning: understanding and accounting for hidden context in rlhf. . External Links: Link Cited by: Qualitative Analysis. T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, T. Althoff, and Y. Choi (2024) A Roadmap to Pluralistic Alignment. arXiv. Note: arXiv:2402.05070 [cs]Comment: ICML 2024 External Links: Link, Document Cited by: Preferences are not all you need, Justification for rubric questions, Implications for future work. K. Stanczak, N. Meade, M. Bhatia, H. Zhou, K. Böttinger, J. Barnes, J. Stanley, J. Montgomery, R. Zemel, N. Papernot, et al. (2025) Societal alignment frameworks can improve llm alignment. In ICLR 2025 Workshop on Bidirectional Human-AI Alignment, Cited by: Utility maximization. H. Sun, Y. Shen, and J. Ton (2024) Rethinking bradley-terry models in preference-based reward modeling: foundations, theory, and alternatives. arXiv preprint arXiv:2411.04991. Cited by: What is value alignment?. R. Thoppilan, D. De Freitas, J. Hall, N. Shazeer, A. Kulshreshtha, H. Cheng, A. Jin, T. Bos, L. Baker, Y. Du, et al. (2022) Lamda: language models for dialog applications. arXiv preprint arXiv:2201.08239. Cited by: What is value alignment?. R. Tian, C. Xu, M. Tomizuka, J. Malik, and A. V. Bajcsy (2024) What matters to you? towards visual representation alignment for robot learning. . External Links: Link Cited by: Discussion. B. Wang, R. Zheng, L. Chen, Y. Liu, S. Dou, C. Huang, W. Shen, S. Jin, E. Zhou, C. Shi, S. Gao, N. Xu, Y. Zhou, Z. Fan, J. Zhao, X. Wang, T. Ji, H. Yan, L. Shen, Z. Chen, T. Gui, Q. Zhang, X. Qiu, X. Huang, Z. Wu, and Y. Jiang (2024a) Secrets of rlhf in large language models part i: reward modeling. . External Links: Link Cited by: Qualitative Analysis, Discussion. C. Wang, D. Zhao, B. Wang, R. He, and Y. Hou (2024b) Do llms have the generalization ability in conducting causal inference?. arXiv preprint arXiv:2410.11385. Cited by: Utility maximization. Z. Wang, W. He, Z. Liang, X. Zhang, C. Bansal, Y. Wei, W. Zhang, and H. Yao (2024c) CREAM: consistency regularized self-rewarding language models. . External Links: Link Cited by: Discussion. C. Xu, S. Chern, E. Chern, G. Zhang, Z. Wang, R. Liu, J. Li, J. Fu, and P. Li (2023) Align on the fly: adapting chatbot behavior to established norms. . External Links: Link Cited by: Qualitative Analysis. T. Yu, T. Lin, Y. Wu, M. Yang, F. Huang, and Y. Li (2023) Constructive large language models alignment with diverse feedback. . External Links: Link Cited by: Qualitative Analysis, Discussion. Y. Zeng, G. Liu, W. Ma, N. Yang, H. Zhang, and J. Wang (2024) Token-level direct preference optimization. . External Links: Link Cited by: Discussion. S. Zhao, J. Dang, and A. Grover (2023) Group preference optimization: few-shot alignment of large language models. . External Links: Link Cited by: Qualitative Analysis. C. Zheng, K. Sun, H. Wu, C. Xi, and X. Zhou (2024) Balancing enhancement, harmlessness, and general capabilities: enhancing conversational llms with direct rlhf. . External Links: Link Cited by: Qualitative Analysis, Discussion. T. Zhi-Xuan, M. Carroll, M. Franklin, and H. Ashton (2024) Beyond Preferences in AI Alignment. Philosophical Studies. Note: arXiv:2408.16984 [cs]Comment: 26 pages (excl. references), 5 figures External Links: ISSN 0031-8116, 1573-0883, Link, Document Cited by: Introduction, Preferences are not all you need, Alternative conceptions of value, Utility maximization, Utility maximization, Discussion, Discussion. Y. Zhong, C. Ma, X. Zhang, Z. Yang, H. Chen, Q. Zhang, S. Qi, and Y. Yang (2024) Panacea: pareto alignment via preference adaptation for llms. . External Links: Link Cited by: Qualitative Analysis. Z. Zhou, Z. Liu, J. Liu, Z. Dong, C. Yang, and Y. Qiao (2024) Weak-to-strong search: align large language models via searching over small language models. . External Links: Link Cited by: Qualitative Analysis.