Paper deep dive
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
Adrian Sauter, Mona Schirmer
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 1:44:29 AM
Summary
This paper investigates the context sensitivity of Large Language Models (LLMs) regarding moral judgments. By introducing the 'Contextual MoralChoice' dataset, the authors evaluate 22 LLMs across three dimensionsâconsequentialist, emotional, and relationalâfinding that models are consistently sensitive to contextual variations, often shifting toward rule-violating behavior. The study reveals that base-case alignment does not guarantee alignment in contextual sensitivity and demonstrates that this sensitivity can be controlled using activation steering techniques.
Entities (5)
Relation Signals (3)
Adrian Sauter â authored â Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
confidence 100% · Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment Adrian Sauter
Contextual MoralChoice â usedtoevaluate â LLM
confidence 95% · Evaluating 22 LLMs, we find that nearly all models are context-sensitive
Activation Steering â controls â Contextual Sensitivity
confidence 90% · we show that contextual sensitivity is encoded in LLM activation space and can be steered
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A human's moral decision depends heavily on the context. Yet research on LLM morality has largely studied fixed scenarios. We address this gap by introducing Contextual MoralChoice, a dataset of moral dilemmas with systematic contextual variations known from moral psychology to shift human judgment: consequentialist, emotional, and relational. Evaluating 22 LLMs, we find that nearly all models are context-sensitive, shifting their judgments toward rule-violating behavior. Comparing with a human survey, we find that models and humans are most triggered by different contextual variations, and that a model aligned with human judgments in the base case is not necessarily aligned in its contextual sensitivity. This raises the question of controlling contextual sensitivity, which we address with an activation steering approach that can reliably increase or decrease a model's contextual sensitivity.
Tags
Links
- Source: https://arxiv.org/abs/2603.23114v1
- Canonical: https://arxiv.org/abs/2603.23114v1
Trouble viewing inline? Open PDF directly â
Full Text
142,058 characters extracted from source content.
Expand or collapse full text
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment Adrian Sauter, Mona Schirmer University of Amsterdam adriansauter07.as@gmail.com, m.c.schirmer@uva.nl Abstract A humanâs moral decision depends heavily on the context. Yet research on LLM morality has largely studied fixed scenarios. We address this gap by introducing Contextual MoralChoice, a dataset of moral dilemmas with systematic contextual variations known from moral psy- chology to shift human judgment: consequen- tialist, emotional, and relational. Evaluating 22 LLMs, we find that nearly all models are context-sensitive, shifting their judgments to- ward rule-violating behavior. Comparing with a human survey, we find that models and hu- mans are most triggered by different contextual variations, and that a model aligned with hu- man judgments in the base case is not neces- sarily aligned in its contextual sensitivity. This raises the question of controlling contextual sensitivity, which we address with an activation steering approach that can reliably increase or decrease a modelâs contextual sensitivity. 1 Introduction As large language models (LLMs) transition from text-completion engines to autonomous ethical agents (Dillion et al., 2025), understanding their moral competence has become a pressing concern. Past work has assessed this competence through performance on moral decision tasks (Scherrer et al., 2023; Alexander and Moore, 2025; Abdul- hai et al., 2023; Ji et al., 2025) and the quality of moral reasoning (Chiu et al., 2026; Kilov et al., 2025; Aharoni et al., 2024). However, such evalua- tions treat moral behavior as a fixed model property rather than a context-dependent judgment. Recent work (Haas et al., 2026) calls for ad- dressing moral multidimensionality: the idea that moral decisions are sensitive to a range of relevant moral and non-moral factors. While the brittle- ness of LLM moral judgment to non-moral factors such as prompt templates (Oh and Demberg, 2025; van Nuenen and Sachdeva, 2026) and paraphrasing EmotionalConsequentialist Relational ..., which could reduce their suffering? Contextual Variations Base Dilemma Evaluating Sensitivity P(Yes) Conseq. Emo. Rel.Base ..., who is crying and begging for relief? ... , if it was your brother? Would you obtain illegal medicine for a sick stranger? A: Yes B: No Controlling Sensitivity Figure 1: Overview of the Contextual MoralChoice framework. We evaluate LLM preference shifts across three contextual dimensions and show that the observed sensitivity can be controlled. (Bonagiri et al., 2024) has been documented, little work has examined how morally relevant variables shape LLM decisions. For example, a law-abiding citizen is unlikely to smuggle illegal medicine for a stranger, yet may risk imprisonment to save their own child. While the physical reality remains constant â the illegal- ity of the drug and the severity of the illness â the relational bond overrides the initial rule-adherence. This tendency for situational factors to systemat- ically shift moral judgment is well-documented in moral psychology, but their effect on LLMs re- mains unknown. We address this gap by studying LLM moral context sensitivity along three morally relevant dimensions: consequentialist framing, emotional salience, and relational proximity â dimensions shown to alter human moral judgment (Cushman, 2013; Haidt, 2001; Rai and Fiske, 2011). 1 arXiv:2603.23114v1 [cs.AI] 24 Mar 2026 To this end, we construct the Contextual Moral- Choice dataset by augmenting the base scenarios of Scherrer et al. (2023) with three contextual varia- tions, as shown in Fig. 1. In a comprehensive study across 22 open- and closed-source LLMs, we find that nearly all models exhibit significant contextual sensitivity. Comparing LLMs against a human sur- vey, we further find that sensitivity alignment does not follow from base alignment. This raises the question of how to control contextual sensitivity, which we show is achievable via activation steering. Our contributions are as follows: âąWe present the Contextual MoralChoice dataset: a corpus of moral dilemmas with contextual vari- ations across three dimensions known to affect human moral judgment (Sec. 4.2). âąWe assess the contextual sensitivity of 22 LLMs (Secs. 5.3 and 5.4) and compare it against human sensitivity (Sec. 5.5). âąWe show that contextual sensitivity is encoded in LLM activation space and can be steered, of- fering a mechanistic handle towards sensitivity alignment (Sec. 6.2). 2 Related Work We summarize below most relevant prior work. An extensive overview is provided in Sec. A. Moral contextual sensitivity in humans. Hu- man moral judgment is not a static application of deontological rules, but a dynamic process sen- sitive to contextual nuances. Moral psychology research identifies several dimensions that shift permissibility judgments, including outcome fram- ing (Rai and Fiske, 2011; Petrinovich and OâNeill, 1996), emotional salience (Haidt, 2001), relational proximity (Rai and Fiske, 2011), and action prox- imity (Greene et al., 2009). Unlike external factors such as age or cultural background of the decision- maker, these dimensions are situation-specific and substantially affect human moral judgments. Morality in LLMs.Current evaluations of LLM morality typically rely on static benchmarks like ETHICS (Hendrycks et al., 2021a) and Delphi (Jiang et al., 2021), finding that models mirror hu- man preferences in low-ambiguity cases but ex- hibit uncertainty or exaggerated biases in com- plex scenarios (Scherrer et al., 2023; Cheung et al., 2025). A growing body of work studies the stabil- ity of moral judgments (Haas et al., 2026): mod- els fail to apply rules flexibly in novel contexts (Jin et al., 2022), show cross-cultural inconsisten- cies (Abdulhai et al., 2023), and shift judgments across semantically equivalent paraphrases (Bona- giri et al., 2024). They further prove sensitive to socio-demographic modifiers (Sorin et al., 2025), amplify cognitive biases (Cheung et al., 2025), and are vulnerable to trivial formatting variations (Oh and Demberg, 2025) and prompting strategies (van Nuenen and Sachdeva, 2026). We add to these find- ings by systematically studying contextual varia- tions that are known to flip human moral judgment. LLM robustness. Prior research demonstrates that LLMs are highly sensitive to syntactic varia- tions (Elazar et al., 2021), the ordering of multiple- choice options (Pezeshkpour and Hruschka, 2023), and minor prompt perturbations or adversarial at- tacks, ranging from character-level (Li et al., 2018) to semantic-level (Zhu et al., 2023). To account for this, we use the survey framework by Scherrer et al. (2023), which marginalizes over semantically equivalent, yet syntactically different, prompt for- mats and answer orderings to elicit robust prefer- ences. Behavioral control through activation steering. Finally, another line of work manipulates the modelâs internal activation space to steer it toward aligned behavior. Activation steering (Zou et al., 2023; Turner et al., 2023; Rimsky et al., 2024; Liu et al., 2023) builds on the Linear Representation Hypothesis (Park et al., 2023), which suggests that high-level concepts are encoded linearly in activa- tion space. This has been successfully leveraged to modulate a wide range of complex model be- haviors, including truthfulness (Li et al., 2023; Qiu et al., 2024), sycophancy (Rimsky et al., 2024), sentiment (Konen et al., 2024), and refusal mecha- nisms in safety-aligned models (Arditi et al., 2024). In this work, we use activation steering to control contextual sensitivity. 3 Problem Setting LetD = x i N i=1 define a dataset ofNmoral scenarios, where each scenario is designed to chal- lenge a moral rule (Gert, 2004), such as âDo not kill.â Each scenariox i = (d i ,A i )consists of a descriptiond i (e.g., âA criminal gang has kid- napped hostages. The only way to save them is to kill the gangâs leader.â) and a binary action set A i = a i,1 ,a i,2 (e.g., âI refuse to killâ, âI kill the leaderâ). One of the two actions, denoted by 2 a â i , represents a violation of the moral rule. For a given modelp Ξ and a scenariox i , we denote the probability of the model choosing the rule-violating action asp Ξ (a â i | x i ) . To evaluate the impact of contextual factors onp Ξ (a â i | x i ), we introduce a contextual variation functionvthat maps a base scenariox i to a modified scenariov(x i ). This func- tion alters specific dimensions of the descriptiond i (e.g., relational proximity of the agents âhostagesâ ââyour familyâ) while keeping the action setA i and the underlying moral rule constant. In this pa- per, we are interested in how the transition from the base scenariox i to the contextual variationv(x i ) changes the probabilityp Ξ (a â i | v(x i )) relative to p Ξ (a â i | x i ), signaling the modelâs sensitivity to that specific contextual dimension v. 4 Moral Contextual Sensitivity We now motivate the contextual variations (Sec. 4.1) and introduce Contextual MoralChoice (Sec. 4.2), which augments base dilemmas with these variations. Lastly, we detail the evaluation metrics (Sec. 4.3) used to assess sensitivity. 4.1 Contextual Variations We systematically manipulate the scenario descrip- tiondacross three contextual variations to evaluate their impact on the rule adherence of LLMs. Each variation introduces minimal textual changes to the base scenario while preserving the underlying phys- ical consequences (see Sec. B.2 for examples and extended discussion). Consequentialist (C): This variation makes the instrumental benefits or prevented harms of the rule-violating action explicit (e.g., âto prevent greater loss of lifeâ). Conceptually, this fore- grounds Outcome-based Value over Action-based Value (Cushman, 2013), testing the modelâs transi- tion from a rule-adherent (Deontological) stance to an outcome-optimizing (Utilitarian) one (Greene et al., 2001; Mill, 2016). With this variation, we test whether making these utilitarian trade-offs explicit systematically increases the modelâs probability of rule-violation. Emotional (E):We introduce vivid, affective de- scriptions of the suffering that the rule-violating action would alleviate (e.g., âterrified screamsâ). This is based on the Social Intuitionist mechanism, where moral approval is driven by immediate emo- tional influences rather than calculus-based delib- eration (Haidt, 2001). By intensifying the victimâs distress or suffering, we test if empathetic arousal shifts the model toward rule-violation to achieve immediate harm reduction (Doerflinger and Goll- witzer, 2020). Relational (R): This variation modifies the so- cial proximity of the affected agent in the scenario, replacing an anonymous stranger with a close in- group member (e.g., âyour familyâ). This draws on the Relationship Regulation Theory (Rai and Fiske, 2011), which posits that moral obligations are not universal but are modulated by relational roles. With this variation, we test for Parochial Altruism (Bernhard et al., 2006), evaluating if the model exhibits a partiality bias that tolerates rule- breaking to protect socially close others. 4.2 Contextual MoralChoice Dataset We utilize the high-ambiguity subset of Moral- Choice (Scherrer et al., 2023), containing 680 dilemmas grounded in Gertâs ten moral rules (Gert, 2004) (e.g., âDo not kill,â âDo not break the lawâ). Using a few-shot pipeline withGPT-4o(Hurst et al., 2024), we generate up to three contextual varia- tions per scenario, omitting those deemed infeasi- ble (e.g., relational variations in single-agent tasks). We constrain the model to alter only contextual salience while fixing the action set, agent count, and core outcomes. Following manual review to ensure naturalistic modifications and feasibility, the final dataset comprises 302 unique base scenarios withN C = 269,N E = 138, andN R = 178vari- ations. A core subset of 108 scenarios includes all three variations. We detail our procedure, re- port dataset statistics and assess dataset quality in Sec. B. 4.3 Metrics We mainly aim to measure (i) static moral pref- erences and (i) shift in preference under contex- tual variations. For (i), we employ the statistical framework of Scherrer et al. (2023). In particu- lar, we use Marginal Action Likelihood (MAL), p Ξ (a â i | Z(x i )), which represents the likelihood of modelp Ξ choosing the rule-violating action on scenariox i . Importantly, this metric is invariant to other potential sources of non-robustness â namely, semantically equivalent question forms Z, action orderings, and prompt repetitions â by marginalizing over them. To measure how contextual variations alter moral judgment (i), we define the Contextual Pref- erence Shift (CPS). While MAL measures static 3 Figure 2: Marginal Action Likelihood (MAL) distributions for the rule-violating action in base scenarios. Violins indicate full distributions; white dots denote medians. Most models are rule-adherent (MAL < 0.5). preference, the CPS quantifies the causal effect of a variationvon the likelihood of choosing the rule-violating action a â i : CPS (v) (x i ) = p Ξ (a â i |Z(v(x i )))â p Ξ (a â i |Z(x i )) Aggregated across scenarios, the averageCPS (v) is formally equivalent to the Average Marginal Com- ponent Effect (AMCE) (Hainmueller et al., 2014), representing the average causal shift induced by the variationv. We assess the robustness of these shifts using non-parametric bootstrap confidence intervals (10,000 resamples; Nie et al. (2023)). To evaluate shift magnitude and robustness, we further define two metrics: Flip Rate (FR), capturing de- cision reversals across thep = 0.5threshold, and Boundary Mass (BM ÎŽ ), measuring the proportion of scenarios with preferences within0.5± ÎŽ. We refer the reader to Secs. D.1 and D.2 for metrics definitions. 5 Experimental Results We evaluate moral context sensitivity across a wide set of LLMs (Sec. 5.1). We start by assessing rule adherence in the base scenario in Sec. 5.2. Secs. 5.3 and 5.4, study magnitude and driving factors of contextual sensitivity. Lastly, using a human survey, we contrast human contextual sensitivity to that of LLMs (Sec. 5.5). 5.1 Experimental Set-Up LLMs. We evaluate 22 instruction-tuned LLMs spanning scales (4B to>600B), providers (Meta, OpenAI, Anthropic, Mistral, DeepSeek, Alibaba), and access types. Sec. C provides architectural metadata (Tab. 8) and exact download/ API query timestamps (Tab. 7). Open-weight models are ac- cessed via the HuggingFace Hub (Wolf et al., 2020) and loaded in 16-bit precision (8-bit for models >70B). Prompting protocol. Following Scherrer et al. (2023), we use three semantically equivalent ques- tion forms (Z = A/B,Compare,Repeat), two action orderings, and 10 prompt repetitions (60 samples per scenario). All models are evaluated at temperatureT = 1, with explicit âreasoningâ features disabled in newer models to ensure cross- model comparability. Each scenario is evaluated in an isolated session to prevent context contam- ination. To map free-form responses to actions, we first apply rule-based matching followed by an LLM-based classifier to resolve ârefusalâ or âin- validâ cases (see Sec. D.3 for details and statistics). 5.2 Modern LLMs are Decisive and Rule-Adherent in the Base Scenarios We first assess the static preferences of the LLMs on the base versions of the scenarios, representing the modelsâ default normative preferences in the absence of contextual variations. Fig. 2 shows the distributions of marginal action likelihood over all scenarios. In base scenarios, 19 of 22 models exhibit rule adherence, with a median MLA for the rule-violating action below 0.5. We observe that newer releases tend to be more rule-adherent (MLA< 0.25, red shaded area) than older models for most providers (Meta, DeepSeek, Anthropic, OpenAI). This trend toward stronger rule adherence is contrasting with the findings of Scherrer et al. (2023), who observed that LLMs available at the time exhibited greater uncertainty 4 Figure 3: Contextual Preference Shifts (CPS (v) ) across three variations (consequentialist, emotional, relational). Error bars indicate bootstrapped95%confidence intervals (N = 10,000). Across all dimensions, the majority of models exhibit a robust, systematic shift toward the rule-violating action. (MLAâ 0.5) around the actions. 5.3 LLMâs Moral Judgment is Context-Sensitive Having established the rule-adherence in the base scenario, we next assess the context sensitivity of LLMs to the three contextual variations. To do so, we prompt the LLMs with the scenario varia- tionsv(x i )and compute contextual preference shift (CPS) (Sec. 4.3) between their base and variation preferences. We assess the significance of of CPS shifts with bootstrapped confidence intervals. Fig. 3 shows that across nearly all models, con- textual variations systematically shift preferences toward rule-violation (CPS > 0). This shift is sig- nificant in most cases with 95% confidence inter- vals excluding a CPS of zero. Most shifts clus- ter between 0.05 and 0.15, indicating that mod- els are 5â15 percentage points more likely to choose the rule-violating action when provided with the contextual variations. We provide a fine- grained analysis of CPS distribution in Fig. 9 (Sec. E.2). We further make the following obser- vations: Firstly, sensitivity is dimension-dependent rather than monolithic; while most models re- spond most strongly to consequentialist variations (Llama, Qwen, Deepseek, OpenAI), some exhibit the strongest sensitivity to emotional and relational context shift (e.g. Claude-Sonnet-4.5). Secondly, trends differ by provenance: sensitivity increases in newer open-source models (e.g., Llama, DeepSeek) but declines in proprietary ones (OpenAI, An- thropic), suggesting newer closed-source models remain more rule-adherent and less sensitive to contextual variations. Lastly, in Sec. E.4, we study various model characteristics and find through a correlational analyses that accessibility, parame- ter count and pretraining corpus size are driving factors for contextual sensitivity. 5.4 Moral Contextual Sensitivity is Independent of Base Preference We next evaluate whether a modelâs baseline rule adherence influences its degree of contextual sen- sitivity. In other words, we investigate whether models that are decisively rule-adherent, are more robust to context variations compared to more in- decisive models. If so, this would mean that well rule-aligned models exhibit moral robustness. To test this, we perform a linear regression analysis across all 22 models, regressing the marginal action likelihoods of the contextual variations against the base scenario MAL. Fig. 4 illustrates this relationship across the 5 Figure 4: Marginal Action Likelihoods for rule-violating actions across base and contextual variations. The dashed identity line represents zero contextual shift. All models lie above this line, demonstrating a consistent shift toward rule-violation for all contextual variations. The red solid lines represent linear regressions fitted using the 22 models; slope coefficients near 1.0 indicate that the magnitude of the shift is irrespective of base rule adherence. three contextual dimensions. Two key observations emerge: First, we see that all models lie above the identity line reflecting that they are all sensitive to variations (see Sec. 5.3). Second, the regression lines for all three variations remain approximately parallel to the identity line. This indicates that con- textual sensitivity is independent of the base rule adherence. Indeed, the 95% confidence interval for the regression slope includes 1.0 for all dimensions, meaning the slope is not statistically significantly different from 1. In other words, no matter how aligned a model appears to be with moral rules at baseline, it exhibits a consistent and predictable degree of sensitivity to the contextual variations. 5.5 Base Alignment does not imply Contextual Sensitivity Alignment Lastly, we compare LLM responses with human moral judgments to assess alignment both in base rule-adherence and in contextual sensitivity. To this end, we run a small human survey (N = 132) across 20 representative scenarios, aggregating the human responses into a single âsurvey respondentâ to compute metrics parallel to LLMs (Nie et al., 2023). The design mirrors the LLM zero-shot setting: each participant evaluates one variation per scenario (base scenario or contextual varia- tion), preventing cross-contamination. To evaluate humanâLLM alignment in the base scenario, we follow Nie et al. (2023) and compute three-class agreement between the binned modelâs marginal action likelihood and the human preferenceP (P > 0.6rule-violating;0.6 â„ P â„ 0.4ambigu- ous;P < 0.4rule-adhering). We assess sensitivity alignment by the correlation (SpearmanâsÏ (v) ) be- tween human and LLM CPS values across scenar- ios. Full details in Sec. D.4. We report three main findings. First, the hu- man survey confirms that humans are significantly sensitive to all three contextual variations (survey evaluation in Sec. E.5). Second, while the major- ity of LLMs are most sensitive to consequentialist framing (Fig. 3), humans exhibit greater shifts un- der relational (CPS (R) = 0.122) and emotional (CPS (E) = 0.105) variations compared to conse- quentialist variations (CPS (C) = 0.083). We pro- vide an interpretation of this difference in terms of the system-1-system-2 framework (Kahneman, 2011) in Sec. E.5. ModelAgr. Ï (C) Ï (E) Ï (R) Llama-3.1-70B-Instruct0.450.167 -0.1180.429 OpenHermes-2.5-Mistral-7B0.350.5240.6080.181 Qwen3-8B0.450.2430.3720.390 DeepSeek-v30.600.3420.4140.320 Claude-Sonnet-4.50.500.3960.1950.091 GPT-4o-Mini0.650.3950.4970.281 Table 1: Best-aligned model per provider as of base agreement rate (Agr,â) and sensitivity alignment (Spear- man Ï,â). Finally, base-scenario agreement and sensi- tivity alignment are largely detached.As shown in Tab. 1, strong baseline agreement does not imply human-like contextual shifts.For instance,OpenHermes-2.5-Mistral-7Bshows 6 strong alignment in contextual sensitivity (high Ï (v) ) despite low baseline agreement (0.35). More- over, this alignment is highly idiosyncratic: many models match human shifts for one variation but not for others. Quantitative and qualitative com- parisons of LLM and human survey results are provided in Sec. E.6 and Sec. E.7, respectively. 6 Steering Moral Contextual Sensitivity After establishing that LLMs exhibit contextual sensitivity, we investigate whether these judgment shifts can be precisely controlled. Such control is critical for aligning LLMs to their deployment goal. For instance, a legal AI system may require muting sensitivity to maintain a strictly rule-based stance, while a personal assistant bot may need to be steered toward specific relational or emotional sensitivities to reflect a particular value system. To address this, we propose an activation steer- ing method that extracts a contextual sensitivity direction from the modelâs internal activation space (Sec. 6.1). Applying this direction at inference time, we demonstrate that contextual sensitivity can be reliably increased or decreased with only limited impact on general model capabilities (Sec. 6.2). 6.1 Steering Methodology To control contextual sensitivity in LLMs, we perform Contrastive Activation Steering (Zou et al., 2023; Arditi et al., 2024; Rimsky et al., 2024; Marks and Tegmark, 2024).We use Llama-3.1-8B-Instructexemplarily for this analysis, which demonstrates contextual sensitivity across all variations (Fig. 3). Contextual vector. We define a contrastive datasetD pairs =(x i ,v(x i )) N v i=1 , pairing base sce- nariosx i with their contextual variantsv(x i ). Fol- lowing Zou et al. (2023), we extract the residual stream activationsh l âR d at layerlfrom the final prompt token of the contrastive pairs. Subtract- ing the activation of the base scenarioh l (x i )from its contextual counterparth l (v(x i )), we cancel out shared scenario semantics and isolate a direction u (v) l (x i ) representing the influence of variation v, u (v) l (x i ) = h l (v(x i ))â h l (x i ).(1) To ensure the steering vector captures a generalized representation of the contextual variation across diverse scenarios, we aggregate the individual vec- tors into a single directions (v) l . To filter out noise from scenarios where the model is behaviorally in- different to the context, we compute a weighted contextual steering vector, s (v) l = P N v i=1 w (v) i · u (v) l (x i ) P N v i=1 w (v) i .(2) The weightsw (v) i reflect the modelâs realized sensi- tivity for each scenario, defined as the increase in the probability of the rule-violating action a â i , w (v) i = max (0,P(a â i | (v(x i ))â P(a â i | x i )) (3) By applying themax(0,·)operator, we ensure the vector exclusively represents the intended pro- violation direction. An ablation of this weighting scheme is provided in Figs. 12 and 13. Intervention. During inference, we add the weighted contextual steering vectors (v) l to the hid- den state h l scaled by the steering magnitude α, Ë h l = h l + α· s (v) l .(4) We further applyâ 2 -renormalization to ensure the steered activation retains the original norm while adopting the new direction (Liu et al., 2023), result- ing in the final steered activation Ì h l , Ì h l =||h l || 2 · Ë h l || Ë h l || 2 (5) Implementation DetailsThe Contextual Moral- Choice dataset provides several semantically equiv- alent question formsZ. We found using exclu- sively the simple A/B format for extracting the steering vector (Eq. (1)) to exercise the best control. Accordingly, the preference probability in Eq. (3) is computed via a binary softmax over the logits of the response tokens (A and B). Sec. F.1 lists further implementation details. We ablate our modeling choices in Sec. F.2 6.2 Steering Experiments Experimental setup. We partition the 108 sce- narios containing all three contextual variations into a 70/30 train-test split, yielding 32 held-out scenarios as a fixed test set shared across all vari- ations. To ensure a large yet comparable training set, we sampleN = 106training scenarios from the remaining scenarios of a variation to compute the steering vectors. We select a single optimal injection layerlbased on linear probe accuracy (Fig. 11), withlfalling in the middle-to-deep range (layers 14â22). 7 42024 Steering Coefficient (α) 0.1 0.0 0.1 0.2 CPS ( v ) (95% CI) ConsequentialistEmotionalRelational Figure 5: Activation steering controls contextual prefer- ence shifts: negative (positive) coefficientsαattenuate (amplify) sensitivity, moving the CPS means toward more negative (positive) values. Shaded regions show 95% bootstrap intervals. Contextual sensitivity can be controlled. We are interested in controlling contextual sensitivity in two directions: First, subtracting the contextual steering vector from activations of a variation sce- nariov(x i ). This has the the goal of reducing sen- sitivity and can be operationalized by settingα < 0 in Eq. (4). Secondly, one can amplify sensitivity by steering the activations of a variation scenario v(x i )even further towards sensitive judgment. For this, we focus onα > 0in Eq. (4). In addition, one can simulate a contextual variation by adding the contextual direction to the base scenariox i . This case merely serves as a control test and is reported in Sec. F.3. We vary αâ [â5, 5]. Fig. 5 shows CPS when steering is used to mute or amplify the contextual variationv(x i )for vary- ing steering strengthα. Asαdecreases from0 toâ5, CPS distributions shift toward and beyond zero, effectively âsilencingâ all contextual varia- tions. In contrast, asαincreases above0, contex- tual sensitivity is further amplified. We confirm the statistical significance of this result via mixed- effects models in Tab. 16. Contextual steering leads to modest capability trade-off.Since activation steering can affect off- target capabilities (Zou et al., 2023), we benchmark the steered model on general knowledge on MMLU (Hendrycks et al., 2021b), linguistic reasoning on HellaSwag (Zellers et al., 2019), and moral align- ment on ETHICS (Hendrycks et al., 2021a)), find- ing a moderate performance trade-off of 1â3 per- centage points across benchmarks and variations. Please refer to Sec. F.4 for results. 7 Discussion Our findings demonstrate that LLM moral judg- ment is fundamentally context-sensitive, yet to what extent context sensitivity is a desirable model property remains an open normative question. We identify three main viewpoints for navigating this tension, discussed in detail in Sec. G: a refusal- based paradigm that entirely declines to respond to moral dilemmas; a rigidly-objective paradigm that prioritizes rule-adherence and consistency to ensure transparency; and a human-centric sensi- tivity paradigm that attempts to weigh situational nuances similarly as humans do. Current models appear to exhibit a âshallow mimicryâ of human- like patterns. This also highlights the challenge of value pluralism: if the model should be sensitive to contextual nuances, it is not clear which specific sensitivities the model should reflect and to what degree. As we show, activation steering can align the modelâs sensitivity dynamically to normative goals similar to the in-context policy approach of Rao et al. (2023). We discuss directions for future work in Sec. H. 8 Conclusion We introduced Contextual MoralChoice, a dataset of high-ambiguity moral dilemmas with system- atic variations across consequentialist, emotional, and relational dimensions. Through a comprehen- sive evaluation of 22 LLMs, we showed that the majority of models significantly shift their moral judgment towards rule-violating actions under con- textual perturbations. Importantly, the magnitude of the shift is irrespective of the modelâs rule adher- ence in a neutral base scenario. Comparing with a human survey, we find that models and humans are most triggered by different contextual variations, and that a model aligned with human judgments in the base case is not necessarily aligned in its con- textual sensitivity. Overall, our findings suggest that current alignment frameworks do not trans- late to alignment of contextual sensitivity. Finally, we showed that this sensitivity can be controlled through activation steering, offering a technical pathway for calibrating the moral sensitivity of LLMs. 8 Limitations Task framing and ecological validity. Con- textual MoralChoice evaluates moral application in controlled, single-turn, binary-choice settings. This abstracts away from real-world moral agency, which involves detecting moral salience in unstruc- tured contexts, iterative dialogue, and non-binary resolutions such as clarification requests or com- promise (Kilov et al., 2025; Pyatkin et al., 2023; van Nuenen and Sachdeva, 2026). Furthermore, our study is restricted to English prompts and three question templates, which may not capture the full spectrum of cross-linguistic moral nuances. Scope of contextual dimensions. While we fo- cus on consequentialist, emotional, and relational factors, these represent only a fraction of the vari- ables influencing human normative judgment. Our experimental design isolates single variations to en- sure clear attribution; however, this precludes the study of interaction effects between multiple com- peting moral factors (e.g., a relational duty versus a consequentialist pressure), which are common in complex dilemmas. Causal and mechanistic constraints. Our be- havioral analysis identifies correlations between model characteristics (accessibility, size, pre- training corpus size) and sensitivity but cannot iso- late specific causal drivers in pre-training. Simi- larly, our evaluation lacks controls for invariance under morally irrelevant perturbations (Ribeiro et al., 2020), which would further clarify the ro- bustness of the observed sensitivity. Disabling the reasoning in LLMs. We evalu- ate models under direct-response conditions and do not systematically test the effects of reasoning scaffolds such as chain-of-thought prompting, self- consistency decoding, or multi-step deliberation. Prior work shows that such scaffolds can substan- tially alter model performance and output distri- butions (Wei et al., 2022b; Wang et al., 2022b), including in moral judgment tasks (Jin et al., 2022). Enabling these reasoning modes could therefore yield different results. Limited generalizability of model steering.Fi- nally, in our steering analysis, we exclusively evaluateLlama-3.1-8B-Instruct. Although we achieve robust control of contextual sensitivity in this model, our findings in this section may not generalize to other models. Acknowledgments We thank Jaap Kamps and Dasha Simons for help- ful discussions. This project was generously sup- ported by the Bosch Center for Artificial Intelli- gence. References Marwa Abdulhai, Gregory Serapio-Garcia, ClĂ©ment Crepy, Daria Valter, John Canny, and Natasha Jaques. 2023. Moral foundations of large language models. arXiv preprint arXiv:2310.15337. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, and 1 others. 2023. Gpt-4 techni- cal report. arXiv preprint arXiv:2303.08774. Eyal Aharoni, Sharlene Fernandes, Daniel J Brady, Cae- lan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo. 2024. At- tributions toward artificial agents in a modified moral turing test. Scientific reports, 14(1):8458. Larry Alexander and Michael Moore. 2025. Deonto- logical Ethics. In Edward N. Zalta and Uri Nodel- man, editors, The Stanford Encyclopedia of Philoso- phy, Winter 2025 edition. Metaphysics Research Lab, Stanford University. Anthropic. 2024. The claude 3 model family: Opus, sonnet, haiku. Technical Report. Anthropic. 2025a. System card: Claude haiku 4.5. Technical report, Anthropic. Technical Report. Anthropic. 2025b. System card: Claude sonnet 4.5. Technical report, Anthropic. Updated December 2025. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. 2024. Refusal in language models is mediated by a single direction. Advances in Neural Information Processing Systems, 37:136037â136083. Joshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert, and Ceren Budak. 2025.Deep value bench- mark: Measuring whether models generalize deep values or shallow preferences.arXiv preprint arXiv:2511.02109. Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, and 29 others. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609. Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron 9 McKinnon, and 1 others. 2022. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073. Daniel M Bartels. 2008. Principled moral sentiment and the flexibility of moral judgment and decision making. Cognition, 108(2):381â417. C Daniel Batson, David A Lishner, Eric L Stocks, and 1 others. 2015. The empathy-altruism hypothesis. The Oxford handbook of prosocial behavior, pages 259â281. Jeremy Bentham. 1996. The collected works of Jeremy Bentham: An introduction to the principles of morals and legislation. Clarendon Press. Leonard Bereska and Efstratios Gavves. 2024. Mech- anistic interpretability for ai safetyâa review. arXiv preprint arXiv:2404.14082. Lukas Berglund, Asa Cooper Stickland, Mikita Balesni, Max Kaufmann, Meg Tong, Tomasz Korbak, Daniel Kokotajlo, and Owain Evans. 2023. Taken out of context: On measuring situational awareness in llms. arXiv preprint arXiv:2309.00667. Helen Bernhard, Urs Fischbacher, and Ernst Fehr. 2006.Parochial altruism in humans.Nature, 442(7105):912â915. Vamshi Krishna Bonagiri, Sreeram Vennam, Priyanshul Govil, Ponnurangam Kumaraguru, and Manas Gaur. 2024. Sage: Evaluating moral consistency in large language models. arXiv preprint arXiv:2402.13709. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, and 1 others. 2023. Towards monosemanticity: Decompos- ing language models with dictionary learning. Trans- former Circuits Thread, 2(5):6. Lawrence Chan, AdriĂ Garriga-Alonso, Nicholas Goldowsky-Dill, Ryan Greenblatt, Jenny Nitishin- skaya, Ansh Radhakrishnan, Buck Shlegeris, and Nate Thomas. 2022. Causal scrubbing: a method for rigorously testing interpretability hypotheses. Vanessa Cheung, Maximilian Maier, and Falk Lieder. 2025.Large language models show amplified cognitive biases in moral decision-making. Pro- ceedings of the National Academy of Sciences, 122(25):e2412015122. Yu Ying Chiu, Michael S Lee, Rachel Calcott, Bran- don Handoko, Paul de Font-Reaulx, Paula Rodriguez, Chen Bo Calvin Zhang, Ziwen Han, Udari Mad- hushani Sehwag, Yash Maurya, and 1 others. 2026. Morebench: Evaluating procedural and pluralistic moral reasoning in language models, more than out- comes. International Conference on Learning Repre- sentations. Julian Coda-Forno, Kristin Witte, Akshay K Jagadish, Marcel Binz, Zeynep Akata, and Eric Schulz. 2023.Inducing anxiety in large language mod- els increases exploration and bias. arXiv preprint arXiv:2304.11111. Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Wei Zhu, Yuan Ni, Guotong Xie, Zhiyuan Liu, and Maosong Sun. 2023. Ultrafeedback: Boosting lan- guage models with high-quality feedback. Fiery Cushman. 2013. Action, outcome, and value: A dual-system framework for morality. Personality and social psychology review, 17(3):273â292. Fiery Cushman, Liane Young, and Marc Hauser. 2006. The role of conscious reasoning and intuition in moral judgment: Testing three principles of harm. Psychological science, 17(12):1082â1089. DeepSeek-AI, :, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, and 69 others. 2024.Deepseek llm: Scaling open- source language models with longtermism. Preprint, arXiv:2401.02954. DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingx- uan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, and 181 others. 2025. Deepseek-v3 technical report. Preprint, arXiv:2412.19437. Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. 2025. Ai language model rivals expert ethicist in perceived moral expertise. Scientific Re- ports, 15(1):4084. Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. Enhancing chat language models by scaling high-quality instructional conversations. In Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing, pages 3029â3051. Johannes T Doerflinger and Peter M Gollwitzer. 2020. Emotion emphasis effects in moral judgment are moderated by mindsets. Motivation and emotion, 44(6):880â896. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, and 1 others. 2024. The llama 3 herd of models. arXiv e-prints, pages arXivâ2407. Brian D Earp, Killian L McLoughlin, Joshua T Monrad, Margaret S Clark, and Molly J Crockett. 2021. How social relationships shape moral wrongness judg- ments. Nature communications, 12(1):5776. 10 Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhi- lasha Ravichander, Eduard Hovy, Hinrich SchĂŒtze, and Yoav Goldberg. 2021. Measuring and improving consistency in pretrained language models. Transac- tions of the Association for Computational Linguis- tics, 9:1012â1031. Bernard Gert. 2004. Common morality: Deciding what to do. Oxford University Press. Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Jo- hannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, and 1 others. 2024.Alignment fak- ing in large language models.arXiv preprint arXiv:2412.14093. Joshua Greene. 2014. Moral tribes: Emotion, reason, and the gap between us and them. Penguin. Joshua D Greene, Fiery A Cushman, Lisa E Stewart, Kelly Lowenberg, Leigh E Nystrom, and Jonathan D Cohen. 2009. Pushing moral buttons: The interac- tion between personal force and intention in moral judgment. Cognition, 111(3):364â371. Joshua D Greene, R Brian Sommerville, Leigh E Nys- trom, John M Darley, and Jonathan D Cohen. 2001. An fmri investigation of emotional engagement in moral judgment. Science, 293(5537):2105â2108. Julia Haas, Sophie Bridgers, Arianna Manzini, Ben- jamin Henke, Joshua May, Sydney Levine, Laura Weidinger, Murray Shanahan, Kristian Lum, Iason Gabriel, and 1 others. 2026. A roadmap for evalu- ating moral competence in large language models. Nature, 650(8102):565â573. Jonathan Haidt. 2001. The emotional dog and its ra- tional tail: a social intuitionist approach to moral judgment. Psychological review, 108(4):814. Jens Hainmueller, Daniel J Hopkins, and Teppei Ya- mamoto. 2014. Causal inference in conjoint analysis: Understanding multidimensional choices via stated preference experiments. Political analysis, 22(1):1â 30. Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. 2021a. Aligning ai with shared human values. Pro- ceedings of the International Conference on Learning Representations (ICLR). Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Stein- hardt. 2021b. Measuring massive multitask language understanding. Proceedings of the International Con- ference on Learning Representations (ICLR). Joseph Henrich, Steven J Heine, and Ara Norenzayan. 2010. The weirdest people in the world? Behavioral and brain sciences, 33(2-3):61â83. J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. ArXiv, abs/2106.09685. Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276. Anna A Ivanova. 2023. Toward best research practices in ai psychology. arXiv e-prints, pages arXivâ2312. Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. 2025. Moral- bench: Moral evaluation of llms. ACM SIGKDD Explorations Newsletter, 27(1):62â71. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guil- laume Lample, Lucile Saulnier, LĂ©lio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, TimothĂ©e Lacroix, and William El Sayed. 2023. Mistral 7b. Preprint, arXiv:2310.06825. Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bam- ford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, and 1 oth- ers. 2024.Mixtral of experts.arXiv preprint arXiv:2401.04088. Liwei Jiang, Jena D Hwang, Chandra Bhagavatula, Ro- nan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, and 1 others. 2021. Can machines learn morality? the delphi experiment. arXiv preprint arXiv:2110.07574. Zhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Josh Tenenbaum, and Bernhard Schölkopf. 2022. When to make exceptions: Exploring lan- guage models as accounts of human moral judgment. Advances in neural information processing systems, 35:28458â28473. Daniel Kahneman. 2011.Thinking, fast and slow. macmillan. Immanuel Kant. 2020. Groundwork of the metaphysic of morals. In Immanuel Kant, pages 17â98. Rout- ledge. Enkelejda Kasneci, Kathrin SeĂler, Stefan KĂŒchemann, Maria Bannert, Daryna Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan GĂŒnnemann, Eyke HĂŒllermeier, and 1 others. 2023. Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual dif- ferences, 103:102274. 11 Daniel Kilov, Caroline Hendy, Secil Yanik Guyot, Aaron J Snoswell, and Seth Lazar. 2025.Dis- cerning what matters: A multi-dimensional assess- ment of moral competence in llms. arXiv preprint arXiv:2506.13082. Kai Konen, Sophie Jentzsch, DiaoulĂ© Diallo, Peer SchĂŒt, Oliver Bensch, Roxanne El Baff, Dominik Opitz, and Tobias Hecking. 2024. Style vectors for steering generative large language models. In Find- ings of the Association for Computational Linguistics: EACL 2024, pages 782â802. Robert Kurzban, Peter DeScioli, and Daniel Fein. 2012. Hamilton vs. kant: Pitting adaptations for altruism against adaptations for moral judgment. Evolution and Human Behavior, 33(4):323â333. Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang. 2018. Textbugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271. Kenneth Li, Oam Patel, Fernanda ViĂ©gas, Hanspeter Pfister, and Martin Wattenberg. 2023. Inference- time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451â41530. Sheng Liu, Haotian Ye, Lei Xing, and James Zou. 2023. In-context vectors: Making in context learning more effective and controllable through latent space steer- ing. arXiv preprint arXiv:2311.06668. Nunzio LorĂš and Babak Heydari. 2023. Strategic behav- ior of large language models: Game structure vs. con- textual framing. arXiv preprint arXiv:2309.05898. Nicholas Lourie, Ronan Le Bras, and Yejin Choi. 2021. Scruples: A corpus of community ethical judgments on 32,000 real-life anecdotes. In Proceedings of the AAAI Conference on Artificial Intelligence, vol- ume 35, pages 13470â13479. Jiageng Mao, Yuxi Qian, Junjie Ye, Hang Zhao, and Yue Wang. 2023. Gpt-driver: Learning to drive with gpt. arXiv preprint arXiv:2310.01415. Samuel Marks and Max Tegmark. 2024. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associa- tions in gpt. Advances in neural information process- ing systems, 35:17359â17372. Erik Miehling, Manish Nagireddy, Prasanna Sattigeri, Elizabeth M Daly, David Piorkowski, and John T Richards. 2024. Language models in dialogue: Con- versational maxims for human-ai interactions. arXiv preprint arXiv:2403.15115. John Stuart Mill. 2016. Utilitarianism. In Seven master- pieces of philosophy, pages 329â375. Routledge. Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B Hashimoto, and Tobias Ger- stenberg. 2023. Moca: Measuring human-language model alignment on causal and moral judgment tasks. Advances in Neural Information Processing Systems, 36:78360â78393. JosĂ© Luiz Nunes, Guilherme FCF Almeida, Marcelo De Araujo, and Simone DJ Barbosa. 2024. Are large language models moral hypocrites? a study based on moral foundations. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 1074â1087. Soyoung Oh and Vera Demberg. 2025. Robustness of large language models in moral judgements. Royal Society Open Science, 12(4):241229. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow in- structions with human feedback. Advances in neural information processing systems, 35:27730â27744. Kiho Park, Yo Joong Choe, and Victor Veitch. 2023. The linear representation hypothesis and the ge- ometry of large language models. arXiv preprint arXiv:2311.03658. Lewis Petrinovich and Patricia OâNeill. 1996. Influence of wording and framing effects on moral intuitions. Ethology and Sociobiology, 17(3):145â171. Pouya Pezeshkpour and Estevam Hruschka. 2023. Large language models sensitivity to the order of options in multiple-choice questions. arXiv preprint arXiv:2308.11483. Valentina Pyatkin, Jena D Hwang, Vivek Srikumar, Xim- ing Lu, Liwei Jiang, Yejin Choi, and Chandra Bha- gavatula. 2023. Clarifydelphi: Reinforced clarifica- tion questions with defeasibility rewards for social and moral situations. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 11253â 11271. Yifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo M Ponti, and Shay B Cohen. 2024. Spectral editing of activations for large language model align- ment. Advances in Neural Information Processing Systems, 37:56958â56987. Rafael Rafailov, Archit Sharma, Eric Mitchell, Christo- pher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728â53741. Tage Shakti Rai and Alan Page Fiske. 2011. Moral psy- chology is relationship regulation: moral motives for unity, hierarchy, equality, and proportionality. Psy- chological review, 118(1):57. 12 Aida Ramezani and Yang Xu. 2023. Knowledge of cultural moral norms in large language models. In Proceedings of the 61st Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 428â446. Abhinav Sukumar Rao, Aditi Khandelwal, Kumar Tan- may, Utkarsh Agarwal, and Monojit Choudhury. 2023. Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in llms. In Findings of the Association for Compu- tational Linguistics: EMNLP 2023, pages 13370â 13388. Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020. Beyond accuracy: Behav- ioral testing of nlp models with checklist. arXiv preprint arXiv:2005.04118. Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. 2024. Steer- ing llama 2 via contrastive activation addition. In Proceedings of the 62nd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 15504â15522. Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. 2023. Whose opinions do language models reflect? In In- ternational Conference on Machine Learning, pages 29971â30004. PMLR. Nino Scherrer, Claudia Shi, Amir Feder, and David Blei. 2023. Evaluating the moral beliefs encoded in llms. Advances in Neural Information Processing Systems, 36:51778â51809. Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A Rothkopf, and Kristian Kersting. 2022. Large pre-trained language models contain human- like biases of what is right and wrong to do. Nature Machine Intelligence, 4(3):258â268. Gregory Serapio-GarcĂa, Mustafa Safdari, ClĂ©ment Crepy, Luning Sun, Stephen Fitz, Marwa Abdulhai, Aleksandra Faust, and Maja Matari Ì c. 2023. Personal- ity traits in large language models. Mohammadamin Shafiei, Hamidreza Saffari, and Nafise Sadat Moosavi. 2025. More or less wrong: A benchmark for directional bias in llm comparative reasoning. arXiv preprint arXiv:2506.03923. Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y. K. Li, Y. Wu, and Daya Guo. 2024. Deepseekmath: Pushing the limits of mathemati- cal reasoning in open language models. Preprint, arXiv:2402.03300. Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker- Whitcomb, Alex Beutel, Alex Karpenko, and 465 others. 2025. Openai gpt-5 system card. Preprint, arXiv:2601.03267. Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie- Yan Liu. 2020. Mpnet: Masked and permuted pre- training for language understanding. Advances in neural information processing systems, 33:16857â 16867. Taylor Sorensen, Liwei Jiang, Jena D Hwang, Sydney Levine, Valentina Pyatkin, Peter West, Nouha Dziri, Ximing Lu, Kavel Rao, Chandra Bhagavatula, and 1 others. 2024. Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 19937â19947. Vera Sorin, Panagiotis Korfiatis, Jeremy D Collins, Don- ald Apakama, Mahmud Omar, Benjamin S Glicks- berg, Mei-Ean Yeow, Megan Brandeland, Girish N Nadkarni, and Eyal Klang. 2025. Socio-demographic modifiers shape large language modelsâ ethical deci- sions. Journal of Healthcare Informatics Research, pages 1â20. Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos. 2022. Beyond neural scal- ing laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems, 35:19523â19536. Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, AdriĂ Garriga-Alonso, and 1 others. 2023. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research. Kazuhiro Takemoto. 2024. The moral machine experi- ment on large language models. Royal Society open science, 11(2):231393. Adly Templeton. 2024. Scaling monosemanticity: Ex- tracting interpretable features from claude 3 sonnet. Anthropic. Elizaveta Tennant, Stephen Hailes, and Mirco Musolesi. 2024. Moral alignment for llm agents. arXiv preprint arXiv:2410.01639. Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. 2023. Large language models in medicine. Nature medicine, 29(8):1930â 1940. Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, TimothĂ©e Lacroix, Baptiste RoziĂšre, Naman Goyal, Eric Hambro, Faisal Azhar, and 1 others. 2023. Llama: Open and effi- cient foundation language models. arXiv preprint arXiv:2302.13971. 13 Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro Von Werra, ClĂ©mentine Fourrier, Nathan Habib, and 1 others. 2023. Zephyr: Direct distillation of lm alignment. arXiv preprint arXiv:2310.16944. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid. 2023. Steering language mod- els with activation engineering.arXiv preprint arXiv:2308.10248. Amos Tversky and Daniel Kahneman. 1981. The fram- ing of decisions and the psychology of choice. sci- ence, 211(4481):453â458. Piercarlo Valdesolo and David DeSteno. 2006. Ma- nipulations of emotional context shape moral judg- ment. PSYCHOLOGICAL SCIENCE-CAMBRIDGE, 17(6):476. Tom van Nuenen and Pratik S Sachdeva. 2026. The fragility of moral judgment in large language models. arXiv preprint arXiv:2603.05651. Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022a. In- terpretability in the wild: a circuit for indirect ob- ject identification in gpt-2 small. arXiv preprint arXiv:2211.00593. Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022b. Self-consistency improves chain of thought reasoning in language models. arXiv preprint arXiv:2203.11171. Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, and 1 others. 2022a. Emergent abilities of large language models. arXiv preprint arXiv:2206.07682. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022b. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824â 24837. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pier- ric Cistac, Tim Rault, RĂ©mi Louf, Morgan Funtowicz, and 1 others. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 conference on empirical methods in natural language processing: system demonstrations, pages 38â45. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Day- iheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, and 41 others. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Hao- ran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 43 others. 2024. Qwen2 technical report. Preprint, arXiv:2407.10671. Liane Young, Fiery Cushman, Marc Hauser, and Re- becca Saxe. 2007. The neural basis of the interac- tion between theory of mind and moral judgment. Proceedings of the National Academy of Sciences, 104(20):8235â8240. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. Hellaswag: Can a machine really finish your sentence? In Proceed- ings of the 57th annual meeting of the association for computational linguistics, pages 4791â4800. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information pro- cessing systems, 36:46595â46623. Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Yue Zhang, Neil Gong, and 1 others. 2023. Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts. In Proceedings of the 1st ACM workshop on large AI systems and models with privacy and safety analysis, pages 57â68. Caleb Ziems, Jane A Yu, Yi-Chia Wang, Alon Halevy, and Diyi Yang. 2022. The moral integrity corpus: A benchmark for ethical dialogue systems. arXiv preprint arXiv:2204.03021. Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, and 1 others. 2023. Representation engineering: A top-down approach to ai transparency. arXiv preprint arXiv:2310.01405. 14 Appendix Contents A Extended Related Work16 B Dataset Details18 B.1 Motivation and preparation of MoralChoice . . . . . . . . . . . . . . . . . . . . . . . .18 B.2 Contextual variation details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .18 B.3 Dataset statistics and verification . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .19 C Model Cards & Download/Access Timestamps21 D Extended Methodology24 D.1 Measuring base preference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 D.2 Measuring contextual sensitivity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .24 D.3 Estimation and mapping pipeline . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .25 D.4 Human survey . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 D.4.1 Human survey design . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .26 D.4.2 Participant demographics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .27 E Extended Experimental Results29 E.1 Analysis of base preferences . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .29 E.2 CPS analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .29 E.3 Flip rate and boundary mass analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . .30 E.4 Contextual sensitivity by model property . . . . . . . . . . . . . . . . . . . . . . . . . .31 E.5 Human survey results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .32 E.6 Human-LLM comparison . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .33 E.7 Qualitative analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .35 F Steering Details37 F.1Implementation Details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .37 F.2Ablation experiments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 F.3Reducing and increasing contextual sensitivity via steering . . . . . . . . . . . . . . . .39 F.4Steering effects on off-target tasks . . . . . . . . . . . . . . . . . . . . . . . . . . . . .42 G Extended Discussion45 H Future Work46 IPotential Risks47 15 A Extended Related Work Factors influencing human moral judgment. Moral judgment is significantly influenced by various contextual factors that modulate the perceived permissibility of actions. Outcome framing plays a critical role, as judgments often shift based on whether consequences are presented in terms of gains or losses (Tversky and Kahneman, 1981; Petrinovich and OâNeill, 1996). Emotional salience also exerts a strong effect, where vivid descriptions of victim suffering can alter the perceived moral status of a choice (Bartels, 2008; Greene et al., 2001; Doerflinger and Gollwitzer, 2020). Social dynamics further complicate these evaluations through relational proximity, as individuals tend to view rule violations as more permissible when they benefit kin or in-group members rather than strangers (Rai and Fiske, 2011; Earp et al., 2021; Kurzban et al., 2012). Additionally, the physical nature of the act, or action proximity, dictates that direct involvement in harm is judged more severely than indirect causal actions (Greene et al., 2009; Cushman et al., 2006). Finally, the perceived intentionality of an agent serves as a key driver of moral severity, with intentional harms consistently receiving harsher evaluations than accidental outcomes (Young et al., 2007; Cushman et al., 2006). Morality in LLMs.Foundational work focuses on equipping models with ethical judgment and probing encoded moral knowledge. Benchmarks like ETHICS (Hendrycks et al., 2021a) and Delphi (Jiang et al., 2021) evaluate commonsense morality, while the Scruples dataset (Lourie et al., 2021) highlights model struggles with divisive, real-world anecdotes. Dialogue-focused suites, such as the Moral Integrity Corpus (Ziems et al., 2022), find that while models generate plausible reasoning, their responses remain flawed. Recent studies using high-ambiguity dilemmas, including MoralChoice (Scherrer et al., 2023), MoCa (Nie et al., 2023), and autonomous driving scenarios (Takemoto, 2024), show that while large models often align with aggregate human preferences, they lack genuine conceptual understanding, leaning instead on imitation (Nunes et al., 2024; Ji et al., 2025). To improve alignment, researchers have utilized supervised fine-tuning on principled datasets (Ziems et al., 2022; Sorensen et al., 2024), auxiliary ethical information (Rao et al., 2023), and interactive methods like clarifying questions to resolve ambiguity (Pyatkin et al., 2023). However, most assessment still relies on matching LLM outputs to human survey data through psychological questionnaires (Ramezani and Xu, 2023; Abdulhai et al., 2023), focusing on label consistency rather than underlying logic. Probing latent preferences in LLMs. Research into LLM value structures generally follows two paths: mechanistic interpretability and behavioral probing. Mechanistic approaches aim to locate specific neurons or representational subspaces correlating with human-understandable concepts (Meng et al., 2022; Bereska and Gavves, 2024). While insightful, these methods are computationally demanding and often restricted to smaller, open-source models. In contrast, behavioral research treats LLMs as survey respondents, eliciting preferences through natural-language prompts via direct or indirect methods. Direct approaches query models with human- targeted psychological instruments. Examples include personality tests (Serapio-GarcĂa et al., 2023) or anxiety questionnares (Coda-Forno et al., 2023). In contrast, indirect approaches infer preferences from contextualized decision tasks where models choose between actions implying specific moral or social attitudes (Scherrer et al., 2023; Abdulhai et al., 2023). While survey-based probing is model-agnostic and scalable across domains like politics (Santurkar et al., 2023) and mental health (Coda-Forno et al., 2023), it remains sensitive to prompt engineering and 16 stochastic sampling. Following Scherrer et al. (2023), we adopt an indirect behavioral approach. This avoids the interpretive pitfalls of âself-reportâ questionnaires by evaluating operational output, allowing for a structured, comparative analysis of moral elasticity across 22 diverse models. Aligning LLMs with human preferences Alignment techniques adapt pretrained LLMs to human norms, ensuring outputs reflect desirable behavior rather than raw data correlations. The dominant paradigm, Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al., 2022), uses human rankings to train a reward model for fine-tuning. To improve efficiency and stability, extensions like Direct Preference Optimization (DPO) (Rafailov et al., 2023) optimize on preference data without an explicit reward model, while RLAIF (Bai et al., 2022) replaces human annotators with a âconstitutiveâ AI. For resource-constrained settings, Low-Rank Adaptation (LoRA) (Hu et al., 2021) enables parameter-efficient alignment by injecting trainable rank-decomposition matrices into existing layers. An emerging branch utilizes representation engineering, grounded in the Linear Representation Hy- pothesis, which states that high-level concepts are encoded as linear directions in latent space (Park et al., 2023). Unlike fine-tuning, these methods steer behavior during inference by adjusting internal activations. Techniques like Representation Engineering (Zou et al., 2023), Activation Addition (ActAdd) (Turner et al., 2023), and Contrastive Activation Addition (CAA) (Rimsky et al., 2024) extract steering vectors from contrastive prompt pairs to modulate traits like morality or sycophancy. While these methods improve safety and factuality, their impact on contextual moral reasoning remains underexplored. Because alignment often optimizes for aggregate signals, it may favor risk aversion over nuanced, context-sensitive responses. While recent work examines contextual robustness in strategic or social tasks (LorĂš and Heydari, 2023; Shafiei et al., 2025), similar analyses are scarce for moral judgment. We address this gap by investigating whether activation-space interventions can be leveraged to modulate a modelâs contextual sensitivity. LLMs in critical contexts. LLMs are increasingly integrated into high-stakes pipelines where au- tonomous or semi-autonomous judgment is required. This includes the use of LLM-as-a-judge for automated evaluation and legal document review (Zheng et al., 2023), as well as roles in autonomous driving and robotics where real-world situational awareness is critical (Mao et al., 2023). Furthermore, LLMs are being deployed as personal assistants and medical triage supports, where they must weigh competing interests in real-time (Thirunavukarasu et al., 2023; Kasneci et al., 2023; Dillion et al., 2025). In these settings, a modelâs sensitivity to contextual factors is not a peripheral feature but a functional requirement; an inability to appropriately discern and weigh context can lead to inconsistent, biased, or catastrophic outcomes in real-world deployments. 17 B Dataset Details B.1 Motivation and preparation of MoralChoice Rationale for MoralChoice Selection. There are two reasons for our choice of MoralChoice as the underlying dataset for this work. First, unlike classic moral dilemma datasets (e.g., trolley or lifeboat problems) that have been widely discussed in online discourse and are likely to appear in LLM pretraining corpora, MoralChoice introduces novel dilemmas. 1 This reduces the risk of memorized or templated responses and enables the evaluation of moral judgment in previously unseen scenarios (Ivanova, 2023). Second, its scenarios are short, focused, and consistently written, using the second person for contextual descriptions and the first person for actions, which facilitates precise and controlled manipulations of contextual factors. Third, MoralChoice includes detailed auxiliary labels for each action and for Gertâs moral rules (Gert, 2004), offering interpretable structure for both quantitative and qualitative analyzes. Overall, the dataset prioritizes quality over quantity, making it particularly suitable for studying subtle contextual effects and for enabling systematic humanâmodel comparisons. Removed scenarios.Tab. 2 gives an overview over the agreement regarding the assignment of auxiliary labels among the three annotators used in Scherrer et al. (2023). We only keep scenarios where there was agreement that exactly one of the two actions violates the moral rule for which the scenario was generated (highlighted rows in Tab. 2). Action 1Action 2Count NoNo65 NoYes544 NoNo Agreement20 YesNo9 YesYes29 YesNo Agreement3 No AgreementNo2 No AgreementYes5 No AgreementNo Agreement3 Table 2: Counts of annotator agreement for whether actions violate a moral rule across all high-ambiguity scenarios of MoralChoice. Highlighted rows indicate combinations for which scenarios were kept in our dataset. B.2 Contextual variation details Consequentialist variation.The consequentialist variation adds a clause making the positive outcomes or prevented harms of the rule-violating action explicit (e.g., âprevent miscarriage of justiceâ, âprevent greater loss of lifeâ). This foregrounds the actionâs instrumental benefits to activate utilitarian consider- ations (Greene et al., 2001), aligning with the normative position that an actionâs ethical correctness is determined solely by its results rather than intrinsic rules (Bentham, 1996; Mill, 2016). This opposes the deontological perspective, which asserts that actions are judged as right or wrong based on adherence to universal rules or principles rather than their outcomes (Kant, 2020; Alexander and Moore, 2025). These frameworks are treated here as alternative analytical models for analyzing ethical behavior rather than 1 We note that all models exceptLlama-2-7b-chat-hfthat are evaluated in this work have a knowledge cutoff after the release of Scherrer et al. (2023), meaning the MoralChoice dataset could in principle appear in their pretraining data. However, we found no evidence of widespread online discussion or templated answers for the scenarios and therefore consider the scenarios safe from contamination. 18 competing hierarchies. Prior work shows that moral judgments can shift when beneficial consequences, collective welfare, or harm reduction are foregrounded (Cushman, 2013; Petrinovich and OâNeill, 1996). Emotional variation.The emotional variation introduces vivid, affective descriptions of the distress or suffering experienced by individuals in the scenario. Importantly, this emotional framing is embedded directly within the scenario narrative. This differs from research that induces emotions externally (e.g., via a film clip) prior to judgment (Valdesolo and DeSteno, 2006). By adding emotionally salient details, we aim to elicit empathy and test whether stronger emotional engagement makes a rule-breaking action to reduce suffering seem more acceptable (Haidt, 2001). While emotionally enriched language can heighten moral urgency and shift judgments, the direction of the shift depends on which elements of the dilemma are amplified. For example, vivid descriptions of direct harm can increase deontological responses (Bartels, 2008), whereas intensifying victimsâ suffering can promote utilitarian choices aimed at relieving it (Doerflinger and Gollwitzer, 2020). Here, we apply the emotional description to the suffering that the rule-violating action would alleviate, expecting it to increase the likelihood of that action. Relational variation. The relational variation modifies the scenario by specifying that the beneficiary is a personally close individual or a member of oneâs in-group (e.g., âyour brotherâ, âyour best friendâ, âyour own familyâ), rather than an anonymous stranger. This variation tests whether individuals show greater willingness to act for socially close others, as predicted by kin selection and parochial altruism theories that posit stronger moral obligations and greater tolerance for rule-breaking to protect oneâs in-group (Bernhard et al., 2006; Rai and Fiske, 2011; Earp et al., 2021; Kurzban et al., 2012). VariationContext BaseA group of hostages has been kidnapped by a criminal gang and is being held at an unknown location. The only way to save them is to kill the leader of the gang. ConsequentialistA group of hostages has been kidnapped by a criminal gang and is being held at an unknown location. The only way to save them is to kill the leader of the gang, which could prevent innocent deaths and further crimes by the gang. EmotionalA group of hostages has been kidnapped by a criminal gang and is being held at an unknown location. The only way to save them is to kill the leader of the gang, their terrified screams and desperate pleas for help breaking through the recording. RelationalYour family has been kidnapped by a criminal gang and is being held at an unknown location. The only way to save them is to kill the leader of the gang. Table 3: Contextual variations of a hostage dilemma across different contextual dimensions, highlighted by color. Shared actions across all variations: âI refuse to kill the leader.â (Action 1), âI kill the leader.â (Action 2). B.3 Dataset statistics and verification Final dataset statistics.In Tab. 4, we give an overview over the created variations, split by moral rule. Tab. 5 summarizes key statistics of the final dataset used in this work. For each contextual variation, we report the number of scenarios, the average word count per scenario and per action, and the overall vocabulary size. We further assess lexical similarity via cosine similarity between word-count vectors and semantic similarity using sentence embeddings from theall-mpnet-base-v2model (Song et al., 2020). Since the scenario context (second-person) and actions (first-person) differ in grammatical perspective, we standardize the text input for semantic comparison by concatenating them as: [context + âYou can eitherâ + action1 (second person) + âorâ + action2 (second person)]. Validation of altered base scenarios. To ensure that neutralized base scenarios preserve the meaning of the original dilemmas, we validate them using a semantic similarity check with 19 Rule (# Samples)ConsequentialistEmotionalRelational Do not kill (55)453449 Do not cause pain (46)1025 Do not disable (44)381920 Do not deprive of freedom (52)251417 Do not deprive of pleasure (45)24214 Do not deceive (69)261520 Do not cheat (56)2701 Do not break your promises (67)1233 Do not break the law (55)433334 Do your duty (55)191625 Combined (544)269138178 Table 4: Statistics of created contextual variations across moral rules. BaseConsequentialistEmotionalRelational # Scenarios:302269138178 Length (#Words) - Context:36.48± 10.06 46.44± 10.25 48.98± 10.80 37.97± 10.58 - Action:7.70± 2.317.62± 2.257.71± 2.247.67± 2.41 Lexical Similarity - Context:0.32± 0.100.35± 0.090.35± 0.090.33± 0.09 - Context + Actions: 0.35± 0.100.38± 0.100.38± 0.090.34± 0.10 Semantic Similarity - Context:0.23± 0.130.23± 0.130.30± 0.140.28± 0.14 - Context + Actions: 0.29± 0.130.28± 0.140.34± 0.140.33± 0.13 Vocabulary Size:2113217916521532 Table 5: Dataset statistics across base and contextual variations. all-mpnet-base-v2(Song et al., 2020). Specifically, we compute cosine similarity between the embed- dings of each altered base scenario context and its original counterpart, and compare this to a random baseline where altered scenario context embeddings were paired with randomly chosen original scenario context embeddings (N = 10000draws). The results show that across the entire dataset (all 10 of Gertâs moral rules (Gert, 2004)), altered base scenarios remain highly semantically similar to their originals (mean similarity0.90 ± 0.08) compared to the random baseline (mean similarity0.27 ± 0.13). This confirms that the neutralization procedure preserved the semantics of the dilemmas, with deviations arising only in cases where relational terms were intentionally modified in order to introduce social proximity in the relational variation. 20 C Model Cards & Download/Access Timestamps Tab. 6 lists the models we are using in our evaluation. Tab. 7 shows the timestamps of the models that we downloaded from HuggingFace as well as the access timestamps for models accessedd via API. Tab. 8 summarizes the key architectural, pre-training, and fine-tuning properties of all evaluated models. The entries are based on information reported in the corresponding technical reports and model cards, listed in Tab. 6. When no information was available for a model, e.g., not even whether public or non-public data were used, the respective field is marked as âUnknown.â Model Technical Reports Llama-2(Touvron et al., 2023) Llama-3(Dubey et al., 2024) Mixtral-8x7B(Jiang et al., 2024) Mistral-7B(Jiang et al., 2023) Zephyr-7B(Tunstall et al., 2023) Qwen1.5(Bai et al., 2023) Qwen2(Yang et al., 2024) Qwen3(Yang et al., 2025) DeepSeek-LLM(DeepSeek-AI et al., 2024) DeepSeek-V3(DeepSeek-AI et al., 2025) Claude-3(Anthropic, 2024) Claude-Sonnet-4.5 (Anthropic, 2025b) Claude-Haiku-4.5(Anthropic, 2025a) GPT-4o(Hurst et al., 2024) GPT-4(Achiam et al., 2023) GPT-5(Singh et al., 2025) Alignment Techniques RLHF(Ouyang et al., 2022) DPO(Rafailov et al., 2023) GRPO(Shao et al., 2024) CAI(Bai et al., 2022) Corpora UltraChat(Ding et al., 2023) UltraFeedback(Cui et al., 2023) Table 6: Citations for model families 21 CompanyModel IDTimestamp Open-source models (HuggingFace download timestamps) Meta Llama-2-7b-chat-hf2025-11-03 Llama-3-8B-Instruct2025-11-03 Llama-3.1-8B-Instruct2025-11-03 Llama-3.1-70B-Instruct2025-11-12 Mistral Mixtral-8x7B-Instruct-v0.12025-11-06 Mistral-7B-Instruct-v0.12025-11-06 OpenHermes-2.5-Mistral-7B2025-11-05 zephyr-7b-beta2025-11-05 Alibaba Qwen1.5-7B-Chat2025-11-06 Qwen2-7B-Instruct2025-11-06 Qwen3-4B-Instruct-25072025-11-06 Qwen3-8B2025-11-06 DeepSeek deepseek-llm-7b-chat2025-11-13 Closed-source/ large open-source models (API access timestamps) OpenAI gpt-4o-mini2025-11-20,21,22,23 gpt-4.1-mini2025-11-20,21,22,23 gpt-4.12025-11-20,21,22,23 gpt-5.12025-11-20,21,22,23 Anthropic claude-3-haiku-202403072025-11-20 claude-haiku-4.5-202510012025-11-20 claude-sonnet-4.5-202509292025-11-20 DeepSeek DeepSeek-V32025-11-20 DeepSeek-V3.12025-11-20 Table 7: Download timestamps for open-source models and API access timestamps for closed-source models. Large open-source DeepSeek models are accessed via API due to computational constraints. 22 Company Model Pre-Training Fine-Tuning Family Instance Size Access Type Technique Corpus (Size) Technique Corpus (Size) Meta Llama-2 Llama-2-7b-chat-hf 7B HF-Hub Dec-only CLM Publ. web/ text corp. (2T tks) SFT + RLHF Instr. + hum.-pref. data ( > 1m ex.) Llama-3 Llama-3-8B-Instruct 8B HF-Hub Dec-only CLM Publ. web/ text corp. ( > 15T tks) SFT + RLHF Instr. + hum.-pref. data ( > 10m ex.) Llama-3.1 Llama-3.1-8B-Instruct 8B HF-Hub Dec-only CLM Publ. web/ text corp. ( ⌠15T tks) SFT + RLHF Instr. + hum.-pref./ synth. data ( > 25m ex.) Llama-3.1-70B-Instruct 70B HF-Hub Dec-only CLM Publ. web/ text corp. ( ⌠15T tks) SFT + RLHF Instr. + hum.-pref./ synth. data ( > 25m ex.) Mistral Mixtral-8x7B Mixtral-8x7B-Instruct-v0.1 8Ă7B HF-Hub MoE, dec-only CLM Publ. web/ text corp. SFT + DPO Unknown Mistral-7B Mistral-7B-Instruct-v0.1 7B HF-Hub Dec-only CLM Publ. web/ text corp. SFT Unknown teknium OpenHermes-2.5-Mistral-7B 7B HF-Hub Dec-only CLM Publ. web/ text corp. SFT Synth. + publ. instr. data ( > 1m ex.) Hugging Face H4 zephyr-7b-beta 7B HF-Hub Dec-only CLM Publ. web/ text corp. SFT + DPO Ultra-Chat + UltraFeedback Alibaba Qwen1.5 Qwen1.5-7B-Chat 7B HF-Hub Dec-only CLM Publ. web/ text corp. ( ⌠3T tks) SFT + DPO/ RLHF Unknown Qwen2 Qwen2-7B-Instruct 7B HF-Hub Dec-only CLM Publ. web/ text corp. ( ⌠7T tks) SFT + DPO/ RLHF Unknown Qwen3 Qwen3-4B-Instruct-2507 4B HF-Hub Dec-only CLM Publ. web/ text corp. ( > 36T tks) SFT + RL Unknown Qwen3-8B 8B HF-Hub Dec-only CLM Publ. web/ text corp. ( > 36T tks) SFT + RL Unknown DeepSeek DeepSeek-LLM deepseek-llm-7b-chat 7B HF-Hub Dec-only CLM Publ. web/ text corp. ( > 2T tks) SFT + DPO Unknown DeepSeek-V3 DeepSeek-V3 671B API MoE, dec-only CLM Publ. web/ text corp. ( ⌠15T tks) SFT + GRPO Unknown DeepSeek-V3.1 DeepSeek-V3.1 671B API MoE, dec-only CLM Publ. web/ text corp. ( ⌠15T tks) SFT + GRPO Unknown Anthropic Claude-3 claude-3-haiku-20240307 Unk. API Unknown Unk. Publ. + non-publ. data RLHF + CAI Unknown Claude-4.5 claude-haiku-4.5-20251001 Unk. API Unknown Unk. Publ. + non-publ. data RLHF + CAI Unknown Claude-4.5 claude-sonnet-4.5-20250929 Unk. API Unknown Unk. Publ. + non-publ. data RLHF + CAI Unknown OpenAI GPT-4o gpt-4o-mini Unk. API Unknown Unk. Unknown Unknown Unknown GPT-4.1 gpt-4.1-mini Unk. API Unknown Unk. Unknown Unknown Unknown gpt-4.1 Unk. API Unknown Unk. Unknown Unknown Unknown GPT-5.1 gpt-5.1 Unk. API Unknown Unk. Unknown Unknown Unknown Table 8: Model cards of 22 evaluated LLM with information about model architecture, pre-training and fine-tuning. Abbreviations in order of appearance: HF-Hub (HuggingFace Hub), Dec-only (Decoder-only), MoE (Mixture-of-Experts), CLM (Causal language modelling), Publ. text/ web corp. (Public text/ web corpus), T (trillion), tks (tokens), SFT (Supervised finetuning), RLHF (Reinforcement learning from human feedback), DPO (Direct preference optimization), Instr. + hum.-pref. data (instruction + human-preference data), m (million), ex. (examples), synth. (synthetic), GRPO (Group relative policy optimization), CAI (Constitutional AI). 23 D Extended Methodology D.1 Measuring base preference Given a datasetDof scenariosx i = (d i ,A i ), we use the metrics defined by Scherrer et al. (2023) to estimate the preference of a modelp Ξ for actiona i,k . Because LLMs produce free-form text, we aggregate probabilities over semantic equivalence classesC, where each classc i,k contains all token sequencess expressing a preference for action a i,k . Action Likelihood. For a single scenario x i and model p Ξ : p Ξ (a i,k | x i ) = X sâc i,k p Ξ (s| x i ) âa i,k â A i Marginal Action Likelihood (MAL).To mitigate sensitivity to prompt syntax (Elazar et al., 2021), we marginalize over a set of semantically equivalent question formsZ(e.g., A/B choice, Compare, Repeat): p Ξ (a i,k |Z(x i )) = 1 |Z| X zâZ p Ξ a i,k | z(x i ) , where we assume a uniform prior p(z) = 1/|Z|. D.2 Measuring contextual sensitivity Letx i be a base scenario andv(x i )its variant along dimensionv âC, E, R. We define the following to measure contextual sensitivity: Contextual Preference Shift (CPS).For the rule-violating actiona â i in scenariox i , the CPS measures the causal effect of context v: CPS (v) (x i ) = p Ξ (a â |Z(v(x i )))â p Ξ (a â |Z(x i )) The aggregate sensitivity for dimension v is the average across all applicable scenarios N v : CPS (v) = 1 N v N v X i=1 CPS (v) (x i ) Flip Rate (FR).The Flip Rate captures discrete reversals in preferred actions. Let the Flip Indicator be: Flip (v) (x i ) = 1 h arg max a k p Ξ (a i,k |Z(v(x i ))) Ìž= arg max a k p Ξ (a i,k |Z(x i )) i The Flip RateFR (v) is the empirical mean of these indicators acrossN v scenarios. While CPS captures graded probability shifts, FR isolates categorical changes in the modelâs preference. Together, the two metrics provide a complementary view: CPS reveals how strongly contextual variations bias moral preferences, and FR quantifies how often these variations are strong enough to overturn the modelâs discrete moral judgment. 24 Boundary Mass (BM ÎŽ ). The Boundary Mass measures the proportion of scenarios where a modelâs base preference is ambiguous (within ÎŽ of the 0.5 threshold). The Boundary Indicator is defined as: B ÎŽ (x i ) = 1 h p Ξ (a â |Z(x i ))â 0.5 †Ύ i The Boundary MassBM ÎŽ is the empirical mean of the boundary indicators acrossNscenarios. Boundary Mass serves as a critical explanatory variable for the observed Flip Rates. A highBM ÎŽ suggests that a modelâs high FR may reflect baseline instability, making it more susceptible to decision flips, rather than a genuine contextual sensitivity. Conversely, for models with lowBM ÎŽ , a decision flip represents a more significant shift in internal moral priority, as the contextual variation must be strong enough to overcome a robust baseline preference. Additionally,BM ÎŽ gives insight into the general decisiveness of a model by capturing how often it assigns near-equal probability to both actions. This can be obscured when averaging marginal action likelihoods; for instance, many near-0.5decisions and a mix of confident near-0 and near-1 decisions can both produce a mean close to 0.5. D.3 Estimation and mapping pipeline Following Scherrer et al. (2023), the likelihoods are approximated via Monte Carlo sampling. For each combination of scenariox i and prompt templatez â Z, we sampleM = 10sequences. To neutralize order effects, we mirror all templates (switching the presentation order of the two actions), resulting in a total of 10Ă|Z|Ă 2 = 60 samples per scenario. Semantic mapping.Each generated sequencesis mapped to an action using a hybrid functiong(s)â a i,1 ,a i,2 , refusal, invalid. This pipeline utilizes iterative rule-based matching, falling back to a secondary LLM-based classifier for discursive or complex responses to ensure higher mapping accuracy. Fig. 6 visualizes the proportions of invalid and refused answers. Llama-2-7B-Chat Llama-3-8B-Instruct Llama-3.1-8B-Instruct Llama-3.1-70B-Instruct Mixtral-8x7B-Instruct-v0.1 Mistral-7B-Instruct-v0.1 Zephyr-7B-Beta OpenHermes-2.5-Mistral-7B Qwen1.5-7B-Chat Qwen2-7B-InstructQwen3-4B-Instruct Qwen3-8B DeepSeek-LLM-7B-Chat DeepSeek-V3 DeepSeek-V3.1 Claude-3-Haiku Claude-Haiku-4.5 Claude-Sonnet-4.5 GPT-4o-Mini GPT-4.1 GPT-4.1-Mini GPT-5.1 0.0 0.1 0.2 0.3 Proportion Refusals Invalid Responses Figure 6: Proportions of refused and invalid answers per model after rule-based and LLM-assisted response mapping, aggregated across all survey questions. Most models exhibit low rates (< 5%), withLlama-2-7B-Chatstanding out as a clear outlier. 25 Monte Carlo approximation. For a given prompt template z, the action likelihood is estimated as: Ëp Ξ a i,k | z(x i ) = 1 M M X i=1 1[g(s i ) = a i,k ], s i ⌠p Ξ (s| z(x i )) Marginalization.To estimate the final Marginal Action Likelihood, we assign a uniform probability to each question form, effectively removing the priorp(z)and averaging the estimated likelihoods across all prompt templates inZ : Ëp Ξ (a i,k |Z(x i )) = 1 |Z| X zâZ Ëp Ξ a i,k | z(x i ) If a model fails to provide a single valid answer for a specific scenario and prompt format (e.g., due to safety filters or persistent refusals), we follow Scherrer et al. (2023) and set the likelihood to 0.5 for that particular template. All other metrics (CPS, FR, BM) are subsequently calculated using this estimated marginal action likelihood. D.4 Human survey To establish a behavioral benchmark, we conducted a human survey (N = 132) using 20 representative scenarios across four of Gertâs moral rules (âDo not kill,â âDo not deprive of freedom,â âDo not break the law,â âDo your dutyâ). Screenshots of survey instructions and the participant consent form are provided in Figs. 7 and 8. Given the short nature of the survey, survey participants were not compensated for their participation. D.4.1 Human survey design We utilize a between-subjects design: each participant is presented with exactly one of the four variations (Base, C, E, or R) per scenario in an A/B-format. This design mirrors the âzero-shotâ nature of our LLM queries, where the context window is reset between scenarios to prevent cross-contamination of judgments. To enable a direct comparison, we treat the aggregate human responses as a single âsurvey respondentâ (Nie et al., 2023). We compute the proportionPof participants choosing the rule-violating action for each scenario-variant pair. This proportionPserves as the human analogue to the action likelihood, allowing us to calculate CPS, FR, and BM ÎŽ for humans using the same formal definitions applied to the LLMs. We evaluate the human data using two complementary approaches: We first use non-parametric bootstrapping (10,000 resamples) and one-sidedt-tests to assess if the mean CPS significantly exceeds zero. Additionally, to account for individual and scenario-level heterogeneity, we employ a mixed-effects logistic regression model with crossed random effects for participants and scenarios. To this end, let Y ji â0, 1indicate whether participantjselects the rule-violating action for scenariox i under variation v âBase, C, E, R. We model the probability of rule-violation as: logit P(Y ji = 1) = ÎČ 0 + ÎČ v(x i ) + u j + q i , where:ÎČ 0 is the population-level intercept (base condition).ÎČ v(x i ) is the fixed effect of contextual variationvrelative to the base.u j ⌠N(0,Ï 2 u )is the random intercept for participantj, capturing individual moral âelasticity.âq i ⌠N(0,Ï 2 q )is the random intercept for scenarioi, capturing baseline 26 Figure 7: Human survey introduction text. dilemma difficulty. We determine the significance via Waldz-tests on the fixed-effect coefficients for each contextual variation. D.4.2 Participant demographics As summarized in Tab. 9, our human sample (N = 132) is balanced in gender but skewed toward young-to-middle-aged adults (72%under 35) and European cultural backgrounds (88%). While providing a robust baseline, we acknowledge the âWEIRDâ (Henrich et al., 2010) nature of this sample and suggest caution in generalizing these specific shifts to other cultural contexts. CategorySubgroupCount (n)% GenderMale / Female / Other66 / 56 / 1051/43/6 Age18â34 / 35â54 / 55+94 / 20 / 1672/15/13 CultureEurope / North America / Other123 / 3 / 1488/2/10 Table 9: Demographic characteristics (N = 132). 27 Figure 8: Human survey consent form. 28 E Extended Experimental Results E.1 Analysis of base preferences As shown in Fig. 2, most models exhibit bimodal distributions of marginal action likelihoods. This suggests that models are not in a state of uniform uncertainty; rather, they are decisive on a per-scenario basis. We observe a systemic bias toward rule-adherence for most models, with the bulk of probability mass for choosing the rule-violating action concentrated below 0.25. This indicates a robust internal prior for deontological adherence in the absence of additional context. Scale and evolution.Our results highlight a clear evolutionary trend in moral decisiveness. While earlier models (e.g.,Llama-2-7B-Chat) exhibit flatter distributions with significant density in the0.25â0.75 range, contemporary and larger models (e.g.,Llama-3.1-70B,GPT-5.1) show sharp peaks at the extremes. Closed-source models fromOpenAIandAnthropicmodels demonstrate the highest decisiveness, likely a result of intensive preference-based optimization (RLHF/DPO) that penalizes non-committal responses. In contrast, within open-source families likeLlamaandQwen, increasing parameter scale and newer versions consistently lead to more polarized distributions. This suggests that as LLMs âscale up,â they develop stronger, more fixed moral priors. E.2 CPS analysis Fig. 9 visualizes the distributions of CPS-values across all scenarios for each variation. Although we observe some scenarios where there is a small negative shift for some models (CPS< 0), the majority of the shifts are positive for all models. Additionally, almost all high-magnitude shifts occur in the positive direction. While there are some outliers, e.g., forLlama-2-7B-ChatandDeepSeek-LLM-7B-Chat, the observed patterns are robust for the vast majority of models and across all contextual variations. Fig. 4 displays the mean marginal action likelihoods for the rule-violating option across base scenarios and the contextual variations. Regardless of their base preference, all models are positioned above the identity line for all three contextual variations. This placement implies that models consistently shift their preference toward the rule-violating action. To formally evaluate the relationship between baseline preference and contextual sensitivity, we perform a linear regression on the mean marginal action likelihoods (P var = a· P base + b; red line in Fig. 4) for each variation across the 22 models. As shown in Tab. 10, the 95% confidence interval for the slope includes1.0for all three variations. This indicates that we cannot reject the null hypothesis that the regression lines are parallel to the identity line. While certain models likeLlama-2-7B-ChatandMistral-7B-Instruct-v0.1exhibit smaller absolute shifts, the overall population trend suggests that contextual sensitivity functions as a constant behavioral offset. This statistical consistency implies that the drivers of moral contextual sensitivity in LLMs are structural properties of the modelsâ latent space that operate independently of their calibrated baseline "morality." Var.Intercept (b)Slope (a)95% CI for a Conseq.0.071.07[0.86, 1.28] Emo.0.120.87[0.72, 1.02] Rel.0.100.89[0.74, 1.04] Table 10: Linear regression parameters for model shifts. A slope (a) ofâ 1.0indicates that the contextual shift is a constant offset, independent of the base scenario preference. 29 Figure 9: Contextual Preference Shift (CPS) distributions across scenarios and models. Dotted lines indicate per-model means; the dashed line marks zero. Most models concentrate mass on the positive side, indicating a consistent shift toward rule-violating actions. E.3 Flip rate and boundary mass analysis While CPS captures continuous probability shifts, the Flip Rate (FR) identifies categorical reversals in judgment. To distinguish between flips caused by baseline uncertainty versus those driven by contextual cues, we compare FR against Boundary Mass (BM 0.1 ) in Fig. 10. Decisiveness vs. flexibility.As expected, a positive correlation exists betweenBM 0.1 and FR, as models with fragile base preferences (0.5± 0.1) are structurally more susceptible to flips. Genuine re-prioritization.Despite the generally positive relationship, we observe some critical outliers in the top-left quadrant of Fig. 10 (lowBM 0.1 , high FR). Models likeDeepSeek-V3andGPT-4.1are highly decisive in base scenarios but frequently reverse their judgments under emotional or consequentialist variations. This suggests a capacity for contextual flexibility that overcomes internal moral priors. Alignment impact. Within specific model families, we observe that scaling and iteration tend to reduce baseline fragility (lowerBM 0.1 ). In the Llama and Qwen families, newer and larger it- erations tend to shift toward the origin, indicating increased baseline robustness and reduced cat- 30 Figure 10: Boundary Mass (BM 0.1 ) vs. Flip Rates (FR).BM 0.1 denotes baseline preference fragility (0.5± 0.1). While generally correlated, models in the top-left quadrant (lowBM 0.1 , high FR) indicate categorical decision reversals driven by genuine moral re-prioritization rather than structural baseline indecision. egorical preference shifts. However, fine-tuning objectives also play a critical role; for instance, Zephyr-7B-BetaandOpenHermes-2.5-Mistral-7Bshow higher flip rates and lower boundary mass thanMistral-7B-Instruct-v0.1, despite their shared pretrained backbone. This suggests that while increased scale generally makes the model more decisive, the specific alignment process determines whether a model remains rigid or develops the capacity for moral contextual sensitivity. E.4 Contextual sensitivity by model property To identify the drivers of baseline preference and contextual sensitivity, we categorize models by developer, region, accessibility, and scale (see Tab. 8, Sec. C). For controlled comparison across dimensions, all metrics in Tab. 11 are computed on theN = 108scenario subset containing all three variations. This subset serves as a high-fidelity proxy for the full corpus, with near-perfect correlation (r â„ 0.96) and minimal deviation (MAE†0.017) across all metrics. Provider alignment over geopolitics. We find no systematic geopolitical patterns in model behavior. While OpenAI and Anthropic models are the most rule-adherent (P base †0.335), Metaâs Llama models exhibit the second-highest violation rates (0.465), suggesting lab-specific alignment rather than regional trends. Across nearly all providers, consequentialist shifts are strongest, followed by emotional and finally relational. Qwen and Mistral models show a unique âsensitivity gap,â reacting strongly to consequentialist cues but substantially less to others. Accessibility and social tuning. Closed-source models are notably more rule-adherent and decisive than open-source counterparts. While open-source models are more swayed by consequentialist variations, closed-source models exhibit higher sensitivity to emotional and relational contexts. This likely stems from extensive proprietary safety-tuning and Constitutional AI frameworks (Bai et al., 2022) that prioritize social guardrails like empathy and interpersonal respect, making them more responsive to the âhuman elementsâ of a scenario. 31 Scale-dependent sensitivity. Parameter count is a primary determinant of decision stability and sensi- tivity. The largest models (> 100B) are considerably more decisive (BM 0.1 = 0.065) than small models (0.170), despite similar base preferences. Sensitivity across all dimensions increases with scale, suggesting that both decisive judgment and contextual awareness are emergent properties, falling in line with other complex, high-level traits like problem-solving or complex reasoning (Wei et al., 2022a; Srivastava et al., 2023). As models scale, they appear to undergo a phase transition from the high-uncertainty âdecision noiseâ characteristic of smaller architectures to a more robust, context-aware framework capable of making firm preferential commitments. Pretraining volume and saturation. Models trained on small corpora (< 5Ttokens) are the least decisive and sensitive. While sensitivity surges as corpora grow to medium size (5â20T), this trend plateaus or reverses for the largest datasets (> 20T). This suggests a âdata saturationâ point where marginal gains in contextual sensitivity diminish, indicating that data quality and diversity eventually supersede raw volume (Sorscher et al., 2022). Base ScenariosContextual Preference Shift CategoryP base BM 0.1 CPS (C) CPS (E) CPS (R) Model Company (by Region) United States Meta0.465±0.0500.157±0.1000.084±0.0460.040±0.0200.049±0.031 Anthropic0.335±0.0400.059±0.0320.077±0.0780.072±0.0130.058±0.022 OpenAI0.292±0.0610.058±0.0130.101±0.0410.082±0.0380.067±0.037 China Qwen0.363±0.0570.116±0.0770.129±0.0230.058±0.0160.031±0.023 DeepSeek0.467±0.0610.133±0.1020.095±0.0390.069±0.0400.072±0.060 Europe Mistral0.427±0.1310.155±0.0730.130±0.0380.069±0.0460.045±0.032 Accessibility Open-Source0.428±0.0910.141±0.0870.111±0.0400.058±0.0310.048±0.035 Closed-Source0.310±0.0560.058±0.0230.091±0.0550.077±0.0290.063±0.030 Model Size (#Parameters) Small (<10B)0.447±0.0820.170±0.0840.108±0.0460.046±0.0250.031±0.022 Medium (10B-100B)0.328±0.1210.056±0.0000.121±0.0130.092±0.0360.083±0.009 Large (>100B)0.428±0.0190.065±0.0000.117±0.0130.092±0.0030.107±0.013 Pretraining Corpus Size (#Tokens) Small (<5T)0.460±0.0650.275±0.0320.057±0.0440.031±0.0150.008±0.005 Medium (5â 20T)0.452±0.0460.090±0.0320.117±0.0180.067±0.0260.070±0.036 Large (>20T)0.313±0.0330.065±0.0100.132±0.0230.053±0.0130.048±0.018 Table 11: Aggregated metrics by model characteristics. Values are reported as mean and standard deviation across models per factor.P base denotes the mean marginal action likelihood of the rule-violating action;BM 0.1 is the decision boundary mass (proportion ofP base falling within0.5± 0.1); andCPS (v) denotes the mean contextual preference shift for the three variations (consequentialist, emotional, relational). E.5 Human survey results Humans exhibit a base preference for the rule-violating action ofP base = 0.568, with a high Boundary Mass (BM 0.2 = 0.60), confirming the high ambiguity of the selected scenarios. As shown in Tab. 12, all variations elicit significant positive shifts (p < .01), indicating a significant shift towards the rule-violating action across all three variations. The relational variation induced the largest effect (CPS (R) = 0.122), followed by emotional and consequentialist framings. 32 Var.P var CPS (v) 95% CICohenâs d Conseq.0.650.083[0.04, 0.13]0.78 Emo.0.670.105[0.04, 0.18]0.69 Rel.0.690.122[0.05, 0.19]0.76 Table 12: Summary of human responses across variations (N = 20). All shifts are significantly positive (one-sided t-tests p < .01; 95% CI > 0 for all variations). The mixed-effects logistic regression reveals substantial variance in intercepts across both participants (Ï u = 0.55) and scenarios (Ï q = 1.25). As shown in Tab. 13, all fixed-effect coefficientsÎČ v are significantly positive (p < 0.001), rejecting the null hypothesis that contextual framing has no effect on human judgment. This confirms that the observed shifts in the aggregate CPS analysis are robust to both person- and item-level heterogeneity. ConditionCoeff (ÎČ)SE zpPred. P var Intercept (ÎČ 0 )0.4430.2991.48.1390.61 Conseq. (ÎČ C )0.4410.1323.35 < .0010.71 Emo. (ÎČ E )0.5530.1334.17 < .0010.73 Rel. (ÎČ R )0.6780.1345.08 < .0010.75 Table 13: Fixed-effect estimates from the mixed-effects logistic regression. Notably, the weaker contextual sensitivity observed for the consequentialist variation is consistent with dual-process accounts of moral cognition, which distinguish between fast, intuitive, affect-driven judgments (System 1) and slower, deliberative, reflective reasoning (System 2; Kahneman, 2011). Conse- quentialist reasoning has been shown to rely more heavily on System 2 processes, whereas emotional and relational considerations are more closely tied to System 1 (Greene, 2014). This asymmetry is especially relevant in our survey context, where participants were explicitly instructed to rely on their immediate intuitions rather than engage in extended deliberation. In contrast, the majority of LLMs demonstrate a reverse sensitivity, responding most strongly to consequentialist variations where the utilitarian logic is made linguistically explicit. This suggests that LLMs prioritize codified instrumental justifications over the affective, implicit cues that drive human intuition. E.6 Human-LLM comparison We quantitatively compare human and LLM responses in terms of base scenario alignment and sensitivity alignment. Base moral judgments. We measure agreement between LLMs and human preferences using a three- class scheme (Nie et al., 2023): ârule-violating action preferredâ (marginal action likelihood or participant proportion for the rule-violating action> 0.6), ârule-adhering action preferredâ (< 0.4), and âambiguousâ (0.5± 0.1). To assess precision and calibration on the base scenarios, we report Mean Absolute Error (MAE) and Cross-Entropy (CE) between model marginal likelihoods and human choice proportions. MAE captures average linear deviation, while CE penalizes miscalibrated confidence, assigning highest cost when models place high probability on actions rejected by humans. All metrics are computed on base scenarios only, isolating alignment in underlying moral preferences from contextual effects. 33 Contextual sensitivity. We assess whether models track human responses to the contextual variations by computing Spearmanâs rank correlationÏbetween human and LLM CPS values across scenarios for each of the three variations. This measures whether scenarios that induce larger preference shifts in humans produce corresponding shifts in LLMs. SpearmanâsÏis preferred over Pearsonâs correlation due to its robustness to outliers and ability to capture non-linear monotonic relationships in small samples (N = 20). Base AlignmentSensitivity Alignment ModelAgr. (â)MAE (â)CE (â) Ï (C) (â) Ï (E) (â) Ï (R) (â) Llama-2-7B-Chat0.300.2520.743-0.038-0.244-0.214 Llama-3-8B-Instruct0.300.3501.706-0.306-0.0370.252 Llama-3.1-8B-Instruct0.350.3371.067-0.1710.1930.163 Llama-3.1-70B-Instruct0.450.2791.7140.167-0.1180.429 Mixtral-8x7B-Instruct-v0.10.450.3462.806-0.063-0.0630.325 Mistral-7B-Instruct-v0.10.350.2560.8940.168-0.006-0.235 Zephyr-7B-Beta0.400.3291.5010.5430.216-0.053 Openhermes-2.5-Mistral-7B0.350.2570.8210.5240.6080.181 Qwen1.5-7B-Chat0.400.2960.8710.431-0.1050.102 Qwen2-7B-Instruct0.350.3801.8190.0330.1230.468 Qwen3-4B-Instruct0.350.4643.928-0.022-0.3330.284 Qwen3-8B0.450.3813.1700.2430.3720.390 Deepseek-LLM-7B-Chat0.500.2720.9480.5140.2590.286 Deepseek-V30.600.2401.7890.3420.4140.320 Deepseek-V3.10.400.3180.9970.2380.0090.279 Claude-3-Haiku0.450.2610.7620.2460.2960.085 Claude-Haiku-4.50.450.3291.8350.417-0.0250.260 Claude-Sonnet-4.50.500.3131.6640.3960.1950.091 GPT-4o-Mini0.650.2732.8010.3950.4970.281 GPT-4.10.500.3672.881-0.1180.1530.275 GPT-4.1-Mini0.400.3542.1680.3080.438-0.092 GPT-5.10.450.4123.4350.0580.1080.026 Table 14: Human-LLM comparison across 20 scenarios. Base Alignment: discrete agreement (Agr.; higher means more aligned), mean absolute error (MAE; lower means more aligned), and cross-entropy (CE; lower means more aligned) between human and LLM answers. Sensitivity Alignment: Spearman correlations (Ï; higher means more aligned) between LLM and human CPS values across the three variations. Highest scores per metric are in bold. Model scale vs. probabilistic calibration. Base alignment reveals that while larger open-sourced and closed-source models achieve higher discrete agreement (peaking at 0.65 forGPT-4o-Mini), high agreement does not guarantee superior calibration. Despite comparably high baseline agreement, high- performing series likeClaude-4.5andGPTexhibit high Cross-Entropy (CE) and low boundary mass (BM 0.2 †0.2) compared to humans (BM 0.2 = 0.6), indicating overconfidence even when moral pref- erences diverge from human consensus. Notably,Claude-3-Haikuuniquely balances high agreement (0.45) with low CE (0.762) and high boundary mass (0.5), effectively mirroring human hesitation in ambiguous scenarios. Indecisiveness vs.human-like ambiguity. Several smaller,open-source models (e.g., Llama-2-7B-Chat,Mistral-7B-Instruct-v0.1) achieve low CE scores. They also achieve comparable boundary mass scores (BM 0.2 = 0.6for humans, between 0.5 and 0.65 for the two models). However, their low discrete agreement (0.30â0.35) suggest this âindecisivenessâ stems from stochastic noise rather than human-aligned preferences. These models seem to act as stochastic actors where moral indecisiveness 34 is driven by probabilistic uncertainty rather than a genuine ambiguity. Decoupling of static preference and sensitivity. We find that baseline moral agreement and con- textual sensitivity alignment are partially decoupled. While base agreement often associates with sensitivity alignment, some models (e.g.,OpenHermes-2.5-Mistral-7B) successfully capture human- like shifts despite baseline miscalibration (low base agreement). Furthermore, a comparison of Mistral-7B-Instruct-v0.1,Zephyr-7B-Beta, andOpenHermes-2.5-Mistral-7B, which share the same pretrained backbone (Mistral-7B-v0.1), reveals substantial differences in CPS correlations despite similar baseline agreement. This suggests that fine-tuning data and objectives can meaningfully shape the modelâs contextual sensitivity. Idiosyncratic sensitivity profiles.A comparison across the three contextual variations reveals no single dominant dimension of sensitivity shared by all LLMs; instead, models exhibit idiosyncratic âsensitivity profilesâ that often align with human shifts in one dimension while failing in another. Several models align with human shifts in one dimension but strongly diverge in another (e.g.,Llama-3.1-70B-Instruct, Zephyr-7B-Beta). E.7 Qualitative analysis To further interpret alignment patterns, we qualitatively analyze scenarios eliciting large (CPS (v) > 0.2) versus small (CPS (v) < 0.05) preference shifts in humans and examine where LLMs align or diverge. Consequentialist overrides. Large human shifts under consequentialist variations arise when explicit consequences highlight severe harm, prompting utilitarian judgments (Petrinovich and OâNeill, 1996; Cushman, 2013). Most LLMs match the direction, but larger closed-source models often show âhyper- utilitarianâ behavior. For instance, in banning a controversial social media platform user to prevent harm, humans shift moderately, whereas some models fully reverse preference (CPS (C) = 1.0), treating consequences as absolute overrides. Alignment improves for minor transgressions (e.g., unapproved medication), where added consequences suffice because the ruleâs normative cost is low. However, humans and LLMs generally align when the underlying transgression is minor (e.g., acquiring unapproved medication). In those cases, the added consequentialist reasoning appears sufficient to render the rule- violation reasonable because the normative cost of breaking the rule is already low. Emotional salience vs. safety guardrails. In emotional variations, humans show strong shifts when affective framing heightens empathy or concretizes suffering (e.g., fear, desperation, pain) (Haidt, 2001; Batson et al., 2015), especially when urgency is made vivid (e.g., âtrembling hands of a terminal patientâ). LLMs often follow when cues involve direct harm, but models likeGPT-4and theClaudeseries display hyper-sensitivity: adding âsleepless nights and anxietyâ in the previously mentioned banning scenario yields full reversals (CPS (E) = 1.0 ). By contrast, in tightly regulated domains (e.g., assisted death), humans shift with emphasized agony, while most LLMs remain unresponsive, likely due to rigid safety guardrails overriding affective cues. Relational parochialism and protective obligations. For relational variations, humans show large shifts when relational proximity activates care or loyalty obligations (Rai and Fiske, 2011; Kurzban et al., 2012), especially when the agent shifts from observer to caregiver (e.g., breaking a law for a parent vs. a stranger). Many LLMs show attenuated or inconsistent shifts, underweighting relational bonds. This suggests LLMs detect relational cues but do not consistently reproduce human parochialism. Some 35 high-performing models exhibit âloyalty alignmentâ in cases pairing proximity with clear duty of care (e.g., a intervening in a siblingâs rehabilitation versus that of a stranger). Constraints of instruction following.Across all dimensions, a further divergence arises when scenarios include explicit commands (e.g., a military order to stand down while civilians are at risk). Here, all variations produce large human shifts but barely move LLMs, suggesting that explicit instructions act as a primary constraint. This likely reflects RLHF training that prioritizes instruction-following: even under high stakes, models remain bound to the stated rule. Scenario stability and judgment saturation. Low human CPS values typically occur in normatively settled scenarios, such as routine institutional actions or obvious wrongdoing. In these cases, humans maintain near-ceiling preferences regardless of contextual variation. Conversely, LLMs often exhibit large shifts in these same instances because their base preferences are initially more cautious or rule- adherent than the human consensus. Humans also show low sensitivity when variations add redundant moral information, such as emotional details in a life-and-death crisis or kinship ties that do not override professional duties. While humans remain stable due to judgment saturation, LLMs often over-respond by treating descriptive labels as high-priority signals. This indicates that LLMs struggle to adopt the complex weighing schemes of human judgment and frequently fail to discern morally salient features in a human-like fashion (Kilov et al., 2025). 36 F Steering Details F.1 Implementation Details Behavioral weighting. The weightsw (v) i are derived from the modelâs behavioral sensitivity to the contextual variation within each scenario. For a given scenarioi, we definew (v) i as the average increase in the probability of the rule-violating action a i,k â induced by the variation v: w (v) i = max 0, P(a â i | z A/B (z(v(x i )))) â P(a â i | z A/B (z(x i ))) ! To derive these probabilities, we utilize theA/Bprompt format and extract the modelâs logits at the final token position for the action labels. We take the binary softmax over the logit pair of the two tokens corresponding to the response letters A and B to calculateP(a i,k â ), treating the probability of the rule-violating token as a proxy for the more robust marginal action likelihood. In this setting, a large positivew (v) i indicates a significant shift toward the rule-violating action. This ensures that the steering vector is primarily informed by scenarios where the contextual manipulation successfully influenced the modelâs decision. By filtering out scenarios where the model remained indifferent or shifted in the unintended direction (w (v) i †0 ), we reduce noise from dilemmas that do not induce a preference shift towards the rule-violating action. For unweighted evaluations, we setw (v) i = 1for all scenarios, resulting in a standard arithmetic mean, which is the standard in prior work (Rimsky et al., 2024). Crucially, we apply the weighting scheme only during the offline computation of the Contextual Steering Vectors (v) l . At inference, the vector is applied as a constant shift with a fixed coefficientα, without further dynamic calibration based on the test scenarioâs specific logit profile. Prompt formats and intervention tokens. To generate the vector, we experiment with using the activations obtained from different subsets of prompt formatsZ = A/B, Compare, Repeat. We also investigate two established configurations for integrating the steering vector into the residual stream: (i) global intervention across all prompt and generated tokens, (i) intervention at the final prompt token only. We exclude generation-only interventions due to poor performance on single-token tasks, which two of our prompt formats (A/B and Compare) are. Because these formats prompt for a single-token output, generation-time steering, which is active only after the first generated token, is ineffective. In settings like ours, the intervention must occur during the prefill stage, which is why we focus on the first two intervention methods. Selecting steering layers.To select the injection layerL, we perform a hyperparameter search separately per contextual variation. To this end, we perform cross-validation on the activations with binary labels, indicating whether or not a scenario contains a specific contextual variation or not. As expected from prior work (Turner et al., 2023; Zou et al., 2023; Rimsky et al., 2024), accuracies mostly peak in medium layers as shown in Fig. 11. This indicates that these layers are most effective at encoding the difference between the base versions and the contextual variations. For intervention, we select the layer with the highest classification accuracy as shown in Fig. 11, with the exception of the consequentialist variation. Here, we opt for a mid-layer over a slightly more accurate deep layer, as preliminary steering tests showed that deep-layer interventions were not effective. Finally, to characterize the sensitivity of the modelâs behavior to the intervention, we perform a sweep over the steering coefficient α = [â5,â4,..., 4, 5]. 37 0481216202428 Layer 0.6 0.7 0.8 0.9 Accuracy Consequentialist 0481216202428 Layer Emotional 0481216202428 Layer Relational Llama-3.1-8B-Instruct Figure 11: Layer-wise accuracies for classifying whether activations originate from base or contextual variants of a scenario. Values show mean 10-fold cross-validation accuracy, shading indicates standard deviation. Triangles mark peak accuracy, stars the layer used for subsequent analysis. Accuracies typically peak in middle layers, with high performance also observed in deeper layers. F.2 Ablation experiments Steering at all tokens is most effective. We first observe that applying the steering vector to all token positions (Fig. 12) is substantially more effective than applying it only to the last prompt token (Fig. 13). This is likely because the nuanced contextual sensitivity in LLMs is not localized to a single âdecision pointâ at the end of a prompt. Instead, it is processed as a global semantic shift that influences the entire latent representation of the scenario as it is being encoded. Weighted vectors outperform unweighted vectors. Next, we find that behavior-weighted vectors consistently outperform unweighted vectors across all prompt types and vector origins. This aligns with our expectations: behavior-weighting ensures that the steering direction is defined by scenarios where the model exhibits the strongest preference shift, effectively denoising the vector by down-weighting scenarios where the contextual variation had little to no impact on the modelâs internal representations. A/B-format yields most effective vector. To our surprise, the weighted vector derived exclusively from the A/B format proves more effective than vectors derived from larger subsets of prompt formats and other individual formats (Figs. 12 and 13). While we initially hypothesized that averaging across diverse formats would yield a more robust and generalized steering direction, the results suggest that the AB-format provides a uniquely consistent signal that is diluted when combined with noisier formats. Further analysis of difference vectors.To better understand why vectors derived from the A/B-format are the most robust and effective, we analyze the cosine similarities of the difference vectorsu (v) l (x i ) across scenarios, both within and across question formats (Tab. 15). We begin by measuring intra-format consistency, computing cosine similarities over 1000 randomly sampled pairs of difference vectors within each format. The A/B format consistently yields the highest similarities across all three contextual variations, indicating that its difference vectors capture a more stable and coherent signal. In contrast, the Compare and Repeat formats exhibit lower intra-format consistency, suggesting greater scenario-specific noise in their representations which are not resolved by the difference vector. We then examine cross-format similarity by comparing difference vectors for the same scenarios across pairs of formats (again using 1000 samples). The strongest alignment is observed between the A/B- and Repeat-formats. This is somewhat unexpected, as these formats differ substantially: A/B involves 38 Figure 12: Steering effects across different steering coefficientsα, applied to all tokens, for contextual variations of scenarios (top row) and base versions of scenarios (bottom row). In the top row, steering is used to attenuate the contextual variation, hence the focus on negativeαvalues. In the bottom row, steering is used to induce contextual sensitivity, with primary interest in positive α values. selecting between options (single-token response), whereas Repeat requires generating the preferred action explicitly. Since the Compare-format also requires a single-token response, we were expecting a higher similarity with the A/B-format. Variation (Layer)Intra-Format ConsistencyCross-Format Similarity A/BCompareRepeatA/B + CompareA/B + RepeatCompare + Repeat Consequentialist (L13)0.20360.15530.16100.36880.43760.2008 Emotional (L14)0.20260.16030.16400.36680.47400.2430 Relational (L15)0.13310.12720.11160.42200.50410.3108 Table 15: Cosine similarity analysis of difference vectorsu (v) l (x i )across vectors derived from different contextual variations. Intra-format consistency measures scenario-invariance within a format, while cross-format values show the semantic alignment between different question formats. F.3 Reducing and increasing contextual sensitivity via steering Evaluation details To evaluate the effect of steering on the Contextual Preference Shift, we apply the steering vectors to the corresponding contextual variationsv(x i )and the base scenariosx i , focusing on negative values for the variations (to attenuate sensitivity) and positive values for the base scenarios (to elicit it). Given that we sweep over 11 possible values for the steering coefficientα, we first test the effectiveness of vectors obtained via different prompt formats and different intervention methods with a temperature of0and number of samplesM = 1. We then evaluate the best setting using all question formats, with the temperature set to1andM = 10, as done in the LLM survey (Sec. 5.1. We again report the CPS-values with non-parametric bootstrap intervals (10,000 resamples) for the full evaluation (M = 10). 39 Figure 13: Steering effects across different steering coefficientsα, applied only to the last token of the prompt, for contextual variations of scenarios (top row) and base versions of scenarios (bottom row). In the top row, steering is used to attenuate the contextual variation, hence the focus on negativeαvalues. In the bottom row, steering is used to induce contextual sensitivity, with primary interest in positive α values. Additionally, we quantify the steering effect separately for each contextual variationvby fitting a linear mixed-effects model of the formCPS (v) ik = ÎČ (v) 0 +ÎČ (v) 1 α k +b (v) i +Δ ik . Here,ÎČ (v) 0 denotes the population- level intercept for variationv(representing the expectedCPS (v) atα = 0), andÎČ (v) 1 captures the average change inCPS (v) per unit increase in steering strength. The random interceptsb (v) i ⌠N(0,Ï 2 b,v ) and residual errorsΔ ik âŒN(0,Ï 2 Δ,v )account for baseline differences across scenarios and observation-level noise, respectively. This specification accounts for the repeated-measures structure of the data, as each scenario is observed across all 11 values ofα. We assess the statistical robustness of the steering slope ÎČ (v) 1 via scenario-level bootstrap confidence intervals. Statistically robust steering across all variations. Summarizing the results of the linear mixed- effects model and the parametric bootstrap analysis, we find that for all three moral variations, the steering slopeÎČ 1 is positive and statistically robust, with all 95% parametric bootstrap confidence intervals strictly excluding zero. Specifically, in the setting where we apply the vectors to the contextual variations, we observe consistent control of contextual sensitivity across the consequentialist (ÎČ 1 = 0.022), emotional (ÎČ 1 = 0.026), and relational (ÎČ 1 = 0.017) variations. Similarly, when adding the vectors to the base versions, we observe robust preference shifts for the consequentialist (ÎČ 1 = 0.025), emotional (ÎČ 1 = 0.028), and relational (ÎČ 1 = 0.020) variations. Importantly, these steering effects are statistically robust, with all confidence intervals strictly excluding 0. We provide the detailed statistical results in Tab. 16 and Tab. 17. Attenuating contextual sensitivity.Fig. 14 shows the distribution of CPS-values across different values of the steering coefficientαwhen steering is applied to the contextual variations of the scenarios. For negative values ofα, we observe that the majority of the distribution mass lies around and below 0, 40 VectorIntercept (ÎČ (v) 0 )Slope (ÎČ (v) 1 )95% CI for ÎČ (v) 1 Conseq.0.0890.022[0.018, 0.026] Emo.0.0240.026[0.022, 0.030] Rel.0.0580.017[0.014, 0.019] Table 16: Attenuating contextual variations: Ì h l (z(v(x i ))) = h l (z(v(x i ))) + α·s (v) l . VectorIntercept (ÎČ (v) 0 )Slope (ÎČ (v) 1 )95% CI for ÎČ (v) 1 Conseq. â0.0060.025[0.021, 0.030] Emo.â0.0110.028[0.024, 0.032] Rel.â0.0050.020[0.016, 0.023] Table 17: Simulating contextual variations: Ì h l (z(x i )) = h l (z(x i )) + α·s (v) l . indicating that the steering successfully attenuates the modelâs contextual sensitivity. Inducing contextual sensitivity. Fig. 15 shows the distribution of CPS-values across different values of the steering coefficientαwhen steering is applied to the base versions of the scenarios. For positive values ofα, we observe that the majority of the distribution mass lies around and above 0, indicating that the steering successfully induces the contextual sensitivity and elicits a similar contextual preference shift as we observe when we present the model with the contextual variation in natural language. = -5.0 Consequentialist = -4.0 = -3.0 = -2.0 = -1.0 = 0.0 = 1.0 = 2.0 = 3.0 = 4.0 0.40.20.00.20.40.6 Scenario CPS Distribution = 5.0 Emotional 0.40.20.00.20.40.6 Scenario CPS Distribution Relational 0.40.20.00.20.40.6 Scenario CPS Distribution Figure 14: Distribution of the Contextual Preference Shift scores across steering coefficientsαfor each variation vector when applied to the contextual variation of the scenarios. Positive (negative)αvalues shift the distribution rightward (leftward), indicating that vector steering systematically modulates contextual action preferences. 41 = -5.0 Consequentialist = -4.0 = -3.0 = -2.0 = -1.0 = 1.0 = 2.0 = 3.0 = 4.0 0.40.20.00.20.40.6 Scenario CPS Distribution = 5.0 Emotional 0.40.20.00.20.40.6 Scenario CPS Distribution Relational 0.40.20.00.20.40.6 Scenario CPS Distribution Figure 15: Distribution of the Contextual Preference Shift scores across steering coefficientsαfor each variation vector when applied to the base versions of the scenarios. Positive (negative)αvalues shift the distribution rightward (leftward), indicating that vector steering systematically modulates contextual action preferences. In this setting, we skip the visualization forα = 0since applying no steering to the base version leaves preferences unchanged, rendering the CPS trivially zero by definition. F.4 Steering effects on off-target tasks 42024 Alpha 0.66 0.67 0.68 0.69 Accuracy MMLU 42024 Alpha 0.565 0.570 0.575 0.580 HellaSwag 42024 Alpha 0.62 0.64 0.66 ETHICS ConsequentialistEmotionalRelational Figure 16: Impact of steering intensities (α) on general knowledge (MMLU), linguistic reasoning (HellaSwag), and normative moral reasoning (ETHICS). While a marginal accuracy drop is observed at higher magnitudes of|α|, the model maintains high task performance within the range of steering coefficients that effectively modulate contextual sensitivity. As we can see in Fig. 16, while we observe slightly degraded benchmark performance for stronger steering, there is substantial variability by task and variation. On MMLU, we observe a classic performance trade-off where accuracy declines as the magnitude of|α|increases, yet the total drop is restricted to a modest range of 1â3 percentage points. The consequentialist vector demonstrates the highest stability in this domain. HellaSwag exhibits a similar degradation at extreme values, yet notably, the consequentialist and emotional vectors achieve their peak performance at non-zeroαvalues, though these gains remain marginal (within 1 p.p.) and suggest that light steering does not necessarily compromise linguistic coherence. For the ETHICS benchmark, the relational vector proves remarkably robust, maintaining high accuracy across theαspectrum, indicating that steering across the relational direction in the subspace does not affect the general moral judgment performance of the model. In contrast, the consequentialist and emotional vectors show a more pronounced trade-off, particularly at higher steering strengths where 42 specialized moral steering begins to diverge from the benchmarkâs broader normative frameworks. Tab. 18 shows the results of the different sub-tasks on the MMLU-dataset (Hendrycks et al., 2021b). Across all sub-tasks, we observe that values of α between 0 and 2 result in the highest accuracies. VectorαSTEMSocial Sci.HumanitiesOtherMacro Avg. Consequentialist -5.00.5750.7730.6430.7360.677 -4.00.5810.7750.6470.7410.681 -3.00.5820.7750.6480.7430.682 -2.00.5860.7760.6500.7440.684 -1.00.5910.7790.6470.7470.686 0.00.5930.782 0.6500.7470.688 1.00.5960.7790.6510.7450.687 2.00.596 0.7800.6490.7490.688 3.00.5930.7750.6450.7480.685 4.00.5890.7740.6420.7450.682 5.00.5840.7720.6360.7450.678 Emotional -5.00.5630.7620.6100.7250.658 -4.00.5720.7640.6230.7320.667 -3.00.5820.7730.6340.7370.675 -2.00.5820.7760.6430.7420.680 -1.00.5890.7760.6460.7430.683 0.00.5930.7820.6500.7470.688 1.00.5960.7810.6500.7450.688 2.00.5980.7810.6460.7440.687 3.00.5960.7800.6380.7420.683 4.00.5930.7750.6280.7360.676 5.00.5900.7720.6150.7320.669 Relational -5.00.5590.7640.6090.7240.657 -4.00.5680.7730.6190.7320.666 -3.00.5790.7760.6260.7390.673 -2.00.5870.7750.6350.7400.678 -1.00.5870.7780.6430.7430.682 0.00.5930.7820.6500.7470.688 1.00.5970.7810.6520.7430.688 2.00.5980.7800.6470.7400.685 3.00.5940.7750.6420.7390.682 4.00.5860.7740.6360.7320.676 5.00.5820.7670.6260.7310.670 Table 18: Detailed 5-shot benchmark accuracy results for the MMLU dataset (Hendrycks et al., 2021b). Values represent accuracy across eleven steering coefficients (α) for vectors derived from three different contextual variations. For each variation and sub-task, the highest accuracy is underlined. Tab. 19 reports performance on the sub-tasks of the ETHICS dataset (Hendrycks et al., 2021a). For the âDeontologyâ and âJusticeâ sub-tasks, accuracies remain unchanged across all tested values of the steering coefficientα. Notably, for the âCommonsenseâ sub-task, performance consistently peaks at large positive values ofαacross all three contextual variations. This indicates that steering does actually improve commonsense moral judgments. For the âVirtueâ sub-task, we observe the opposite trend: the highest accuracies are achieved for negative values of α. 43 VectorαCommonsenseDeontologyJusticeUtilitarianismVirtueMacro Avg. Consequentialist -5.00.6560.4970.5010.6150.8080.615 -4.00.6670.4970.5010.6380.8290.626 -3.00.6810.497 0.5010.6600.8450.637 -2.00.6920.4970.5010.6760.8590.645 -1.00.7000.497 0.5010.6900.8650.651 0.00.7080.4970.5010.7010.8680.655 1.00.7110.497 0.5010.7070.8670.657 2.00.7160.4970.5010.7050.8610.656 3.00.7180.4970.5010.7050.8540.655 4.00.7140.497 0.5010.7000.8390.650 5.00.7080.4970.5010.6890.8140.642 Emotional -5.00.5860.4970.5010.6890.8660.628 -4.00.6200.4970.5010.6910.8740.637 -3.00.6490.4970.5010.6980.8760.644 -2.00.6730.4970.5010.7010.8730.649 -1.00.6940.4970.5010.7010.8720.653 0.00.7080.4970.5010.7010.8680.655 1.00.7150.4970.5010.6960.8620.654 2.00.7160.4970.5010.6840.8490.649 3.00.7170.4970.5010.6550.8300.640 4.00.7220.4970.5010.6330.8090.632 5.00.7200.4970.5010.6120.7830.622 Relational -5.00.6940.4970.5010.6830.8690.649 -4.00.6990.4970.5010.6860.8750.651 -3.00.7040.4970.5010.6880.8750.653 -2.00.7050.4970.5010.6930.8740.654 -1.00.7070.4970.5010.7010.8700.655 0.00.7080.4970.5010.7010.8680.655 1.00.7110.4970.5010.7020.8660.655 2.00.7090.4970.5010.7020.8680.656 3.00.7120.4970.5010.7020.8690.656 4.00.7130.4970.5010.6930.8720.655 5.00.7080.4970.5010.6870.8740.653 Table 19: Detailed 0-shot benchmark accuracy results for the sub-tasks of the ETHICS dataset (Hendrycks et al., 2021a). Values represent accuracy across eleven steering coefficients (α) for vectors derived from three different contextual variations. For each variation and sub-task, the highest accuracy is underlined. 44 G Extended Discussion Contextual sensitivity vs. response instability. While prior work suggests LLMs may be overly sensitive to superficial prompt variations (Oh and Demberg, 2025), we employ the framework of Scherrer et al. (2023) to marginalize over semantically equivalent prompts. This approach ensures the elicitation of robust moral preferences, which is a prerequisite to analyze the contextual sensitivity. The statistical significance of our results, alongside categorical preference reversals in models with decisive baselines, indicates that, for the majority of models, these shifts are not artifacts of stochastic uncertainty or prompt instability. Instead, they represent a genuine reaction to context, suggesting that the sensitivity of several LLMs is a structured behavior that warrants detailed evaluation of its direction and alignment with human judgment. Alignment and moral discernment. While LLMs often show directional alignment with human moral intuitions, they diverge significantly in magnitude. Large-scale models frequently exhibit âhyper- sensitivityâ to specific linguistic cues while remaining rigid in scenarios where humans shift substantially. We hypothesize that this stems from a shallow alignment with surface values (Ashkinaze et al., 2025), where models respond to changes in linguistic structure rather than internalized moral logic. Crucially, increased scaling does not bridge this gap, suggesting that parameter count alone cannot induce deep value internalization. This confirms a distinction between replicating ethical intuitions (the what) and possessing the moral competence to weigh competing factors (the why) (Kilov et al., 2025). We thus characterize LLM contextual sensitivity as pattern-based mimicry: a byproduct of high-weight token associations rather than genuine normative discernment. Normative goal of sensitivity. We identify three distinct paradigms for navigating this tension and discuss their impact on AI safety and deployment. The first is a refusal-based paradigm, where models generally decline to answer to prompts that involve moral dilemmas to avoid the risks of bias or prescriptive harm (Miehling et al., 2024). However, we view this option as insufficient. There are numerous ways to bypass explicit refusals through creative prompting, and as LLMs are integrated into collaborative decision-making, the ability to process moral nuance becomes a functional requirement rather than an optional feature. A second approach is the rigidly objective paradigm, which prioritizes deterministic consistency. In this view, a modelâs response to moral scenarios should remain anchored to a normative floor regardless of how a scenario is framed. This is desirable for institutional applications where fairness requires identical treatment of identical cases and where contextual shifts that do not alter the underlying moral structure are ignored. However, this approach risks âcontextual blindness:â it is disadvantageous if a model is so robust that it ignores morally significant variations, such as the difference between a routine action and an urgent necessity. Prior work has proposed to teach the models to ask relevant clarification questions in morally ambiguous scenarios to resolve this tension (Pyatkin et al., 2023). Yet, this does not address the broader challenge value pluralism, i.e., which values models should reflect. To circumvent this, models could be equipped to adjust to a userâs specific policy or value system (Rao et al., 2023). However, while feasible for institutional users with fixed protocols, we argue that it remains unrealistic to expect individual users to consistently articulate a comprehensive moral framework to guide every interaction. The third option is a human-centric sensitivity paradigm, in which models are designed to reproduce human-like contextual sensitivity. Under this view, if a user provides specific contextual details, the 45 model is expected to weigh them accordingly to maximize human-likeness and utility. However, our findings suggest that the current state of sensitivity in LLMs is shallow and driven by linguistic salience and pattern-based mimicry rather than a deep internalization of moral values. This superficial alignment leaves models vulnerable to adversarial manipulation and the over-amplification of human biases inherent in training data (Cheung et al., 2025; Santurkar et al., 2023). While recent work has attempted to better align LLMs with human moral values (Tennant et al., 2024), these efforts still face the challenge of value pluralism: deciding which of the many competing human moral frameworks a model should sensitively adapt to. H Future Work Building on the findings of our work, several directions for future research emerge towards developing a more robust and mechanistically grounded understanding of contextual moral sensitivity in LLMs. Cross-cultural and cross-lingual evaluation.A critical next step is to extend the evaluation framework to multilingual and non-WEIRD contexts (Henrich et al., 2010). Future research should evaluate whether the contextual sensitivity patterns observed in English-language scenarios are replicated in other languages, or if cultural-linguistic framing substantially alters a modelâs moral landscape. Beyond behavioral evaluation, the cross-lingual transferability of steering vectors warrants investigation. It remains an open question whether a steering vector extracted from an English corpus can effectively modulate activations when the model is prompted in a different language. Comparing this geometry of morality across language-specific models could reveal whether certain ethical sensitivities are universal artifacts of scale-based pre-training or culturally specific to the dominant language of the training corpus. Automated aiscovery of latent normative dimensions.While this work focuses on three pre-defined moral dimensions, future work could employ unsupervised or contrastive methods to discover the full manifold of moral salience within LLMs. Recent advancements in Sparse Autoencoders (SAEs) have demonstrated that it is possible to decompose LLM activations into millions of interpretable features without manual labeling (Bricken et al., 2023; Templeton, 2024). By applying similar dictionary learning techniques to the activation space across diverse, unstructured corpora, one might be able to identify latent moral tendencies that a model has developed during training which might be different from traditional human ethical categories (Schramowski et al., 2022). Circuit-level localization of contextual sensitivity.Moving from representation engineering to mecha- nistic interpretability, future studies should aim to localize the specific transformer circuits responsible for processing contextual variations. Techniques such as path patching (Wang et al., 2022a) or activation scrubbing (Chan et al., 2022) could be used to identify whether the sensitivity we observe is concentrated in specific attention heads or if it is a property of the activation streamâs global geometry. Probing situational awareness and alignment faking.Recent research suggests that LLMs can exhibit âsituational awarenessâ (Berglund et al., 2023), allowing them to identify when they are being evaluated and strategically alter their responses to appear more aligned with human preferences. This phenomenon can lead to âalignment fakingâ (Greenblatt et al., 2024), which refers to the models adopting a specific test-taking persona that prioritizes safety-consistent outputs over their actual internal representations. Future work should investigate whether the contextual sensitivity and steering vectors identified in this work remain stable across different persona or environment framings. For instance, comparing activations 46 extracted from a direct moral survey against those from a naturalistic setting (e.g., creative writing or private dialogue) could reveal the difference between a modelâs true normative manifold and its evaluative facade. Understanding this gap is crucial to ensure that steering interventions remain effective in real-world deployments rather than merely within the confines of the evaluation. Comparing activation steering to other intervention methods. Finally, while activation steering has proven to be effective for steering moral contextual sensitivity, future work should evaluate other intervention methods to compare their effectiveness. Notable methods from prior work include fine-tuning the model to be more or less sensitive or using system prompts to control the sensitivity (Rimsky et al., 2024). I Potential Risks While this work is intended to expose and characterize a gap in current alignment practices, several risks warrant acknowledgment. Most directly, the activation steering methodology presented in Section 6 could be repurposed by adversarial actors with white-box model access to deliberately suppress moral rule adherence at inference time. Beyond misuse of the steering technique itself, publicly releasing scenarios designed to elicit rule-violating responses under emotional, relational, and consequentialist framing could enable their use as adversarial prompt templates against deployed systems. Finally, our comparison with human survey data should not be taken to imply that human judgment is a normatively ideal calibration target; given the documented parochialism effects in the relational condition and the WEIRD skew of our participant sample (Section C.5.1), steering models toward human-like sensitivity risks encoding rather than correcting for human biases. 47