Paper deep dive
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
Haonan Huang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/8/2026, 2:07:22 AM
Summary
The paper demonstrates that the apparent yes-no bias in large language models is not a shift in moral judgment but a surface artifact driven by answer order (recency bias toward the last-printed option) and lexical pull (preference for the word 'no'). Frontier models maintain a coherent, format-invariant internal moral scale, while small open-weight models show format incoherence. The artifact is concentrated in Claude models, negligible in GPT-5.5 and Gemini, and diminishes with extended reasoning. A minimal logistic model parameterizes the effect using framing susceptibility and moral decisiveness, highlighting that accurate value measurement requires crossing question frames rather than single-form elicitation.
Entities (14)
Relation Signals (13)
Large Language Models â exhibits â Yes-No Bias
confidence 97% ¡ LLMs increasingly issue judgments read as binary verdicts... an amplified yesâno bias on moral dilemmas, absent in humans
Yes-No Bias â decomposesinto â Order Bias
confidence 96% ¡ The apparent yesâno bias splits, exactly and by construction, into an order bias toward the last-printed option
Yes-No Bias â decomposesinto â Lexical Pull
confidence 96% ¡ The apparent yesâno bias splits... plus a lexical pull toward the word ânoâ
GPT-5.5 â exhibits â Minimal Yes-No Bias
confidence 95% ¡ â0 for GPT-5.5 and Gemini
Claude Models â exhibits â Substantial Yes-No Bias
confidence 95% ¡ The artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86)
Gemini â exhibits â Minimal Yes-No Bias
confidence 95% ¡ â0 for GPT-5.5 and Gemini
Order Bias â favors â Last-Printed Option
confidence 95% ¡ an order bias toward the last-printed optionâopposite to classic human primacy
Lexical Pull â favors â Word No
confidence 95% ¡ plus a lexical pull toward the word ânoâ
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wording - among them an amplified yes-no bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word "no" is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrization - every logically irrelevant factor flipped in balanced pairs - across a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier models' stance $\theta$ is nearly format-invariant (cross-form incoherence 0.12-0.21 on a $\pm 1$ axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed option - opposite to classic human primacy - plus a lexical pull toward the word "no"; the artifact is substantial only in the Claude models (story-averaged -0.32 to -0.86), $\approx 0$ for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves $\approx 0$ for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejecting - the pull follows the printed surface, not the verdict it carries. A minimal model, $P = \sigma((\theta \pm m)/s)$, summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once.
Tags
Links
- Source: https://arxiv.org/abs/2607.05552v1
- Canonical: https://arxiv.org/abs/2607.05552v1
Trouble viewing inline? Open PDF directly â
Full Text
90,937 characters extracted from source content.
Expand or collapse full text
The yesâno bias of large language models reflects answer order and wording, not shifts in moral judgment Haonan Huang Princeton University, Princeton, NJ 08540, USA hnhuang@princeton.edu Abstract Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logically irrelevant changes of wordingâamong them an amplified yesâno bias on moral dilemmas, absent in humans. A single framing cannot say what such a shift is: in a yes/no question the word ânoâ is at once logical verdict, lexical token, and last-printed option. We introduce a psychometric battery that separates these: crossed symmetrizationâevery logically irrelevant factor flipped in balanced pairsâacross a corpus of question forms. A graded rating across logically equivalent forms recovers a coherent internal moral scale: frontier modelsâ stance θ is nearly format-invariant (cross-form incoherence 0.120.12â0.210.21 on a Âą1Âą 1 axis); small open-weight models fail in model-specific ways. Forcing the verdict through yes/no overlays a decomposable artifact: an order bias toward the last-printed optionâopposite to classic human primacyâplus a lexical pull toward the word ânoâ; the artifact is substantial only in the Claude models (story-averaged â0.32-0.32 to â0.86-0.86), â0â0 for GPT-5.5 and Gemini, and shrinks under extended reasoning. The word and the verdict share one token; swapping the words for arbitrary labels separates them, and the verdict-attached logical bias proves â0â0 for every frontier model, while model-specific label and order attachments remain: the models are not drawn toward rejectingâthe pull follows the printed surface, not the verdict it carries. A minimal model, P=Ďâ((θ¹m)/s)P=Ď((θ¹ m)/s), summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, measurably distinct from sampling temperature. The battery applies unchanged to any dilemma set and binary format: measuring what a model values requires crossing the frames of the question, not asking once. Significance statement. Language models increasingly make or inform judgments, so evaluations must separate what a model values from artifacts of how the question is posed. We find that frontier language models carry a surprisingly coherent internal moral scale: graded ratings of moral dilemmas barely move under logically equivalent rewordings. Yet forcing the same judgment through yes/no overlays a large format artifactâa pull toward the last-printed option and toward the word ânoââeasily read as a disposition to say âno.â It is not: with arbitrary answer labels the verdict-attached bias vanishes; the pull follows the printed label. Two interpretable numbers, a framing susceptibility and a moral decisiveness, summarize the artifact, and deliberation typically shrinks it. Measuring what an AI values requires crossing the frames of the question, not asking once. Keywords: large language models || moral judgment || psychometrics || framing effects || AI evaluation How an agent ranks a moral dilemma depends on how it is asked. In people this framing sensitivity is real but mild [1, 2, 3, 4, 5, 6, 7, 8]; in large language models (LLMs) it is severeâmoral and social judgments move under changes of wording, option order, and answer format that carry no logical content [9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23]. The stakes are practical: a modelâs judgment is typically read out as a binary verdictâa safety gate, a surveyâs forced choice, a judge modelâs approve/reject [24, 25, 26, 27, 18]âso what a shifted verdict means determines what the evaluation measured: a changed moral judgment, a preference for an answer word, or a pull toward a position on the page. Two properties of LLMs make the question urgent and the naive remedy fail. Repetition does not help: at temperature 1, replicates of a single form are nearly unanimous, while the variance across logically equivalent forms runs up to an order of magnitude larger (Methods) [28, 29]. Resampling one wording thus underrepresents the variation that matters, and that variation is large: LLMs are hypersensitive to surface form [15, 16, 17, 21], so a questionnaire built for humansâwho need only a handful of framingsâdoes not transfer. We build the instrument the measurement requires: a psychometric battery whose core operation is crossed symmetrizationâevery logically irrelevant perturbation applied in balanced flip-pairs, the perturbations crossed so that their contributions separateârepeated across a corpus of logically equivalent question forms. A single form, even symmetrized, is a noisy instrument; a family of forms is a measurement. The sharpest instance, and the anchor for everything below, is the finding of Cheung, Maier, and Lieder [9] that on matched moral dilemmas LLMs show amplified cognitive biasesâamong them a yesâno bias absent in their human comparison: the modelsâ verdicts shift with logically irrelevant features of the yes/no question far more than peopleâs doâand hundreds of resamples per item did not average the shift away. We take that finding as established and ask what the shift is made of. When a modelâs verdict moves toward âno,â is the model drawn to the last-printed option (an order bias)? To the word itself (a lexical bias)? Or to the negative verdict (a logical bias)? Repetition cannot tell these apart, and neither can any single wording, however carefully chosen: in the standard questionâââŚanswer yes or noââthe three candidates coincide on the same token. Most value-probing evaluations elicit each judgment in exactly this way, through a single fixed wording [24, 30, 31, 32, 33] or a handful of format or order variants that leave the scenarioâs substance unchanged [34, 35, 36, 10]; however precisely such a design measures the shift, only crossing can attribute itâthe questionâs verb against the printed order of the answers against the label that carries the verdict. The battery poses the same twenty dilemmasânineteen of them verbatim from Cheung et al.âs materials, so the stance axis and the human anchors remain comparable (Methods)âthrough three instruments. A graded rating (I-1) elicits moral acceptability under crossed, logically equivalent forms: two scales (0â1010 and 0â100100), both anchor directions, three wordings, and both poles of the action (rating the action and rating its complement). A two-alternative free choice (I-2) elicits a committed decision with no scale at all. A forced binary (I-3) elicits the yes/no-style verdict under a crossing of the questionâs verb (approve/oppose) with the printed order of the answer options and with the answer label itselfâthe yes/no words, arbitrary labels, marks and symbols, and yes/no words of other languages. Throughout, the symmetric part of a flip-pair estimates the stance; the antisymmetric part is the artifact; and the crossing assigns the artifact to named sources. Nothing in the instrument is specific to these storiesâthe battery applies unchanged to any dilemma setâand humans enter only through Cheung et al.âs published data (Methods). Five results follow. (i) An internal moral scale exists. Under logically equivalent reframings a frontier model returns nearly the same graded stance θ (cross-form incoherence Ďrepro=0.12 _repro=0.12â0.210.21 on the Âą1Âą 1 axis)âthe convergent validity a single framing cannot certifyâwhile two small open-weight models provide the contrast, failing in model-specific ways (SI Appendix, section S6). (i) The standard readout decomposes. The apparent yesâno bias splits, exactly and by construction, into an order bias toward the last-printed option plus a lexical pull toward the word âno,â and it concentrates in a single model family. (i) Two interpretable parameters. A minimal model, P=Ďâ((θ¹m)/s)P=Ď((θ¹ m)/s), summarizes any such artifact by a framing susceptibility m and a moral decisiveness s, and deliberation shrinks |m||m| where the artifact is substantial. (iv) The pull follows the label, not the verdict. Under fully arbitrary answer labels the verdict-attached bias is â0â0 for every frontier model; swapping labels reshapes the surface pullsâmodel-idiosyncraticallyâwhile an independent order pull persists. (v) An honest scope. Coherence is substantial, not perfect; every coherence claim is within-model; and the artifactâs composition is a property of the models measured, not a law of the class. Across the instruments that never print an answer label or an option order, the same models are coherent; the bias appears when, and only when, the judgment is read out through a labeled binary. The shift lives in the readout, not in the judgment. We first ask the question any use of a moral scale presupposes: does the model have one? Results An internal moral scale exists and is coherent. A frontier model returns nearly the same graded moral stance however the rating question is framed. The instrument: each dilemmaâs action and its complement are rated under 48 crossed conditions (24 unipolar conditions enter θ; the two bipolar variants are a pre-registered measured contrast, non-load-bearing; Methods). The symmetrized mean is the stance θâ[â1,+1]θâ[-1,+1], signed toward the costâbenefit pole (+1+1 endorses the sacrificial action; â1-1 the rule/deontological pole), and the spread of θ across logically equivalent forms, corrected for sampling noise, is the cross-form incoherence Ďrepro _reproâreported as the error bar on every θ in this paper, because it, and not sampling noise, is the dominant uncertainty (Methods). Independent model calls have no mechanism to agree. The conditions of the graded instrument differ in scale (0â1010 versus 0â100100), anchor direction, wording, and which pole of the action is rated; each is a separate call with no shared context. A system without an internal, format-invariant disposition has no reason to return consistent ratings across them. That frontier models doâthe between-form spread of θ, corrected for sampling noise, is Ďrepro=0.12 _repro=0.12â0.210.21 on a Âą1Âą 1 axis (Fig. 1)âis evidence of a genuine latent in exactly the psychometric sense of convergent validity [37, 38]: logically equivalent but operationally independent elicitations recover one quantity. The contrast makes the point sharper: for a small open-weight model (Qwen3.6-35B-A3B) the scale itself is format-incoherent (Ďrepro=0.40 _repro=0.40), and salt-replication (re-eliciting under meaningless context strings) proves this incoherence is deterministicâunder meaningless context perturbation the within-form variation is essentially zero (median SD 0.0000.000; 69% of forms exactly invariant; SI Appendix, section S3)âso the large spread is a property of the model, not of our measurement. A second open-weight model (NVIDIA Nemotron-3-Nano-30B-A3B, likewise a small mixture-of-experts read at the logit level) is the instrumentâs degenerate case, and we report it separately (SI Appendix, section S6) rather than alongside the functioning scales: in its direct mode it answers near the scale midpoint on nearly every dilemmaâits between-item spread is half the panelâs smallest (SDitemsâ(θ)=0.17SD_items(θ)=0.17, versus 0.330.33â0.560.56), with an ordering that agrees with no other configuration (mean r2=0.20r^2=0.20 frontier, 0.120.12 human; SI Appendix, Table S2)âso its nominally moderate cross-form spread (Ďrepro=0.20 _repro=0.20) is largely the coherence of indifference: a scale that says nothing leaves forms little to disagree about. A scale can fail by scattering across logically equivalent forms or by failing to differentiate the items, and only the generalizability coefficient below registers both. An operationally different instrument ties the scale to behavior: the modelâs free choice between the two courses of action (I-2), elicited with no rating scale and no yes/no label, correlates with the graded θ at r=0.82r=0.82â0.900.90 per configuration (Sonnet 0.850.85, Opus 0.830.83, GPT-5.5 0.900.90, Flash-Lite 0.820.82) and recovers the same per-story ordering as the symmetrized binary stance (r=0.80r=0.80â0.940.94); the outlier on both counts is again Gemini-3-Flash (0.310.31 and 0.360.36; SI Appendix, section S5). The ordering is also largely shared. Frontier models largely agree on how the twenty dilemmas order (Fig. 1a), and the shared ordering moderately tracks Cheung et al.âs human ratings (mean pairwise r2r^2 of per-item θ: 0.590.59 frontierâfrontier, 0.430.43 frontierâhuman; descriptive, no CI; SI Appendix, Table S2)âa side observation here, since every claim below is within-model. The honest outlier is Gemini-3-Flash, which holds a genuinely inverted ordering on several sacrificial dilemmas: it spontaneously adopts an act-utilitarian frame and rates the sacrificial action acceptable. This is a real moral disagreement, not a scale artifactâit survives the anchor-direction consistency check, and the modelâs free-choice decisions agree with its ratings (SI Appendix, section S8). Deliberation tightens the scale. Enabling extended reasoning lowers Ďrepro _repro for every model with both modes and a functioning scale (Sonnet 0.13â0.120.13\!â\!0.12, Haiku 0.21â0.170.21\!â\!0.17; paired per-item bootstrap: Haiku Î=+0.05 =+0.05 [+0.01,+0.09][+0.01,+0.09], Sonnet marginal at Î=+0.019 =+0.019 [+0.001,+0.040][+0.001,+0.040]; Methods): what thinking improves is not only the answer but the answerâs format-invariance. Deliberation also relocates the open-weight model on the spectrum: extended reasoning removes most of the gap (0.40â0.180.40\!â\!0.18; paired Î=+0.22 =+0.22 [+0.12,+0.31][+0.12,+0.31]), so the contrast is a stance-incoherence spectrum with a deliberation axisâanchored by Qwenâs deterministic direct modeârather than a class dichotomy. (The degenerate case probes the rule from below: Nemotronâs direct-mode spread is jitter around indifference, not format-driven structure, so deliberation has nothing to reconcileâit multiplies the between-item signal variance fourfold, G=0.28â0.56G=0.28\!â\!0.56, without lowering the raw spread, Î=â0.05 =-0.05 [â0.10,+0.01][-0.10,+0.01]; SI Appendix, section S6.) A generalizability reading gives the coherence claim a criterion: a single randomly chosen form preserves the between-item stance ordering with G=0.77G=0.77â0.940.94 for the frontier configurations versus G=0.42G=0.42 for Qwenâs direct mode (0.860.86 with extended reasoning; Methods). The scope of the coherence claim, stated here once and consolidated in the Discussion: substantial, not perfect (a residual anchor-direction component dominates what remains, Fig. 1b); within-model (θ is a chosen coordinate, not a reading of moral truth); and format-invariance over the elicitation frame, with the dilemma texts themselves held verbatim throughout. A coherent scale, however, is not how models are usually read. What does the standard elicitationâa forced yes/noâreturn? Figure 1: Frontier models carry a coherent internal moral scale; a small open-weight model, without deliberation, does not. (a) Stance θ per dilemma (abscissa: the 20 dilemmas, ordered by the frontier-mean consensus, gray band), small multiples per model family; filled markers = extended reasoning (âthinkâ), open = direct; error bars = per-item cross-form incoherence Ďrepro _repro (a descriptive spread, not a CI). Open diamonds: human per-dilemma choice rates (2âPâ12P-1) from Cheung et al.âs dataâtheir elicitation was itself binary; human framing-robustness is what licenses using the rate as a stance anchorâshown only on the 19 verbatim items (item A is adapted; SI Appendix, section S7). Frontier models track a shared ordering; Gemini-3-Flash holds a genuinely inverted ordering on several sacrificial dilemmas (SI Appendix, section S8); the small open-weight modelâs θ scatters off the consensus. (b) Per-model Ďrepro _repro (frontier 0.120.12â0.210.21; small open-weight model 0.400.40), each bar split by variance share into anchor-direction (solid) and scale (hatched) components, which combine in quadrature (Methods; Opus, which has no aggregate decomposition, is drawn at its per-item mean); shaded band = the frontier range; arrows: extended reasoning lowers Ďrepro _repro for every model with both modes. Ten configurations plotted (Gem-3-Fl = Gemini-3-Flash; Fl-Lite = Gemini-3.1-Flash-Lite; Opus = Claude Opus 4.8 at maximum effort, I-1/I-2 only); the gray consensus is the mean over the eight frontier configurations (including Opus, excluding the open-weight model). A second open-weight model (NVIDIA Nemotron-3-Nano-30B-A3B) is measured on the identical battery but omitted here as a degenerate case (a compressed scale that barely differentiates the items; its full profile: SI Appendix, section S6, Fig. S8). n=20n=20 items per configuration (GPT-5.5: 22 collected, 20 shared); 35,31635,316 graded trials in total (SI Appendix, section S4). Uncertainty for the (b) bars is assessed by a per-item bootstrap (SI Appendix, section S3); the frontierâopen-weight(direct) gap is many interval-widths wide. The standard readout overlays a format artifactâand it decomposes. Cheung et al.âs name for this effect is the yesâno biasâa pull they find absent in humans [9]âand we adopt their term: the apparent yesâno bias is the raw shift of a forced yes/no verdict under logically irrelevant framing changes, with signed language for its parts (negative = toward the word âno,â nay-saying). We deliberately avoid the survey-methodology term acquiescence [39, 40, 41]: acquiescence names a content-level tendency to agreeâa logical bias toward âyesââand the label-swap identification below (Fig. 5) shows the LLM shift is carried by surface mechanisms (the printed order and the answer word), not by any verdict-attached tendency. Measured in the standard single-framing wayâone wording, âanswer yes or noââseveral models show a substantial pull. But that one number is uninterpretable: as noted, the word ânoâ is three things at once, and no amount of repetition separates them; only crossing does. Crossing the questionâs verb (approve/oppose) with the printed answer order (yes-first/no-first) splits the apparent bias, exactly and by construction, into an order component and a lexical component. The split is decisive (Fig. 2). Story-averaged, the apparent bias is substantial only for the Claude models: Sonnetâs â0.32-0.32 decomposes into order bias â0.18-0.18 plus lexical â0.14-0.14, and Haikuâs â0.86-0.86âthe largest we measureâinto â0.33-0.33 plus â0.53-0.53; GPT-5.5 and both Geminis sit at â0â0 (apparent bias |â |â¤0.04|¡|⤠0.04). Deliberation shrinks both components (order: Sonnet â0.18ââ0.10-0.18\!â\!-0.10, Haiku â0.33ââ0.12-0.33\!â\!-0.12; lexical: â0.14ââ0.11-0.14\!â\!-0.11 and â0.53ââ0.27-0.53\!â\!-0.27). The family concentration survives the three obvious deflations: it is not a wording-corpus artifact (on the shared 12-form baseline that every model ran the contrast is unchangedâSonnet â0.32-0.32 [â0.45,â0.20][-0.45,-0.20], Haiku â0.92-0.92âversus GPT-5.5 â0.04-0.04 and the Geminis â0.02-0.02/+0.01+0.01; Methods), not a refusal artifact (Flash-Lite withholds verdicts on 24% of trials yet shows â0â0 bias; Methods), and not a reasoning-budget difference (Sonnet with extended reasoning, â0.21-0.21, still far exceeds GPT-5.5, â0.04-0.04; reasoning defaults per model: Methods). The order bias points toward the last-printed optionârecency-type, opposite in direction to the classic primacy of human respondents in written surveys [42, 43]. On Cheung et al.âs own question wording, verbatim (one form family; Sonnet, both reasoning modes, full replication depth) the artifact is smaller and order-dominated: the apparent yes-first pull is â0.12-0.12, and the order-balanced residual is +0.01+0.01, 95% CI [â0.12,+0.16][-0.12,+0.16]âconsistent with no systematic residual, though the interval cannot exclude word-level residuals comparable to the lexical components above. The per-item residuals are real (up to Âą0.5Âą 0.5) but item-idiosyncratic and cancelâitself an illustration of the paperâs methodological point: a single wording family cannot resolve item-level structure; families can (Methods). The component magnitudes are properties of the wording family measured; on the verbatim family, at this precision, the artifact is order-dominated. These are story-averaged magnitudes: they depend on the item set, and they say nothing about where on the scale the artifact lives. Resolving by story reveals its structureâand yields parameters that do not depend on the items at all. Figure 2: The apparent yesâno bias is a confound: it splits exactly into order bias ++ lexical. Per model: the apparent single-framing yes/no bias (purple) and its decomposition, by the crossing identity, into an order bias (toward the last-printed option; orange) and a lexical pull (toward the word ânoâ; red). Negative = toward âno.â The artifact is substantial only for the Claude models and shrinks under extended reasoning (â ¡T denotes the think mode); GPT-5.5 and both Geminis are â0â0. The lexical bar is resolved into word versus verdict in Fig. 5. Story means over n=20n=20 dilemmas (Claude: 49â50 verb-flip forms; other models: the shared 12-form wording baselineâon which the Claude values are unchanged; Methods). Error bars indicate bootstrap 95% CIs. The artifact has structure, and two parameters summarize it. Resolved by story, the artifact has the geometry of a horizontal offset (Fig. 3). Near a modelâs moral tie (θâ0θâ 0) the two members of a flip-pair split widely; toward the extremes both saturate and the split closesâa âbowtieâ in stance space, a single-peaked ridge in θ space. The artifact lives where the model is torn, not where it is decided; and this is exactly the geometry a logistic response with a frame-dependent shift produces. Two presentations of the same data answer different questions. The story-averaged bias depends on the story distribution: where dilemmas sit far from a modelâs moral tie, both framings saturate and the raw bias collapses toward zero regardless of the underlying susceptibilityâa small raw number does not, by itself, mean a robust model. The fitted parameters are story-independent: the fit reads the susceptibility from the residual gap at the dilemmas nearest the tie, and thereby diagnoses the saturation. The raw numbers compare models on a fixed item set; the parameters are portable. Because the graded instrument established θ as a continuous, coherent latent, the parsimonious model of a forced binary is the canonical logistic link [44, 45, 46, 47, 48]: the two framings of a flip-pair respond as pÂą=Ďâ((θ¹m)/s)p_Âą=Ď((θ¹ m)/s). The symmetric part is the stance; the antisymmetric part is the bias; and each logically irrelevant frame contributes one horizontal offset m, in the units of the modelâs own scale. The second parameter is not a sampling temperature. Temperature rescales a fixed output distribution; s rescales how the verdict depends on the modelâs own graded stance. The distinction is easy to missâa fitted sigmoid over an LLMâs answers looks like the softmax it samples fromâand the open-weight models, whose answer-slot distributions we read directly, let us measure the two scales separately: the readout sigmoidâs scale (the effective sampling temperature mapping the answer-slot logits to sampled answers) fits at Tâ1Tâ 1, as sampling theory requires (Nemotron 1.001.00 [0.98,1.02][0.98,1.02], Qwen 0.900.90 [0.88,0.93][0.88,0.93]), while the same modelsâ s, estimated from temperature-free greedy margins, is 0.270.27â0.300.30 and halves under deliberation with the readout untouched (SI Appendix, section S4 and Fig. S3 there). A small s is a decisive modelâits P snaps from 0 to 1 across its moral tie; a large s is a hedging one; and sââsââ is a model whose binary verdict no longer tracks its scale at all. We therefore read s as moral decisiveness, and m as framing susceptibility. Fitted per bias and per configuration (Fig. 4; each biasâs two sides are fitted jointly and share one stance sigmoidâan internal consistency check the data pass, Fig. 3), the parameters say three things. First, m is substantial only for the Claude models, and deliberation shrinks it: the order-bias m falls from â0.13-0.13 to â0.08-0.08 (Sonnet) and from â0.48-0.48 to â0.06-0.06 (Haiku); the lexical m from â0.10-0.10 to â0.09-0.09 and from â0.64-0.64 to â0.14-0.14. GPT-5.5 and both Geminis carry |m|â¤0.06|m|⤠0.06 throughout; the sharpest cross-family statement is Claude versus GPT-5.5, since saturation leaves Gemini-3-Flashâs fitted null weakly identified (Methods). Second, where the fit is resolvable the per-bias decisiveness is mutually consistent (s=0.14s=0.14â0.190.19 across the Claude configurations, order versus lexical channels), supporting the reading of s as a property of the model rather than of the bias source. Third, absolute s is clip-limited where the readout saturates (Haikuâ ¡direct, the Geminis, the small open-weight model); only the m-ranking and within-model changes are read there (Methods). One question remains about the lexical component: is the pull toward ânoâ a preference for the word, or does the model mean the verdict? Yes/no data cannot tell them apartâthe word and the verdict are the same token. Figure 3: Story-resolved structure: the artifact is a horizontal offset on a logistic, and one fit explains both bias channels. Rows: the four Claude configurations. Columns: the stance sigmoid zâ(θ)z(θ) with the two per-bias fits overlaid (their overlap is the internal consistency check); then, per bias (order, lexical), the bowtie (bias versus stance z) and the ridge (bias versus θ). Points: dilemmas (n=18n=18â2020 per configuration); curves: the joint (s,m)(s,m) fit per bias (Methods). In the ridge columns, thin marks behind each story mean show the within-story per-form split (the two form-parity subsets). A story enters a configurationâs fit only if all four verbĂorder cells are usable after refusal exclusion (Haikuâ ¡think lacks two stories). Haikuâ ¡direct saturates (degenerate fit; its s is clip-limited). Error bars on both axes indicate sampling-corrected cross-form incoherence spreads (descriptive, not inferential; Methods). Figure 4: Two portable parameters: framing susceptibility m and moral decisiveness s. (aâc) The model. A flip-pair responds as pÂą=Ďâ((θ¹m)/s)p_Âą=Ď((θ¹ m)/s): the mean of the pair is the stance, the signed difference is the bias, and the frameâs pull is the horizontal offset m (m<0m<0 = toward âno,â nay-saying; m>0m>0 = toward âyes,â yes-saying). The offset generates the bowtie (b) and the ridge (c) of Fig. 3. (d) Extracted decisiveness s per configuration and bias channel (green band = the mutually consistent Claude range, 0.140.14â0.190.19; â /hatching = degenerate or clip-limited absolute s: Haikuâ ¡direct and both Geminis; GPT-5.5âs absolute s is method-dependent, 0.120.12â0.230.23 across estimators, so absolute values are read within model only). (e) Extracted susceptibility m (θ-units): substantial and negative only for the Claude models, shrinking under extended reasoning; GPT-5.5 and both Geminis â0â0. n=18n=18â2020 dilemmas per fit (see Fig. 3; Claude on all forms; others on the 12-form baseline). Error bars indicate bootstrap 95% CIs. Lexical, not logical. Crossing in a valence-free answer labelâthe full verb Ă label Ă order designâseparates them, because a label like âBâ can carry the verdict without carrying the word. The result is unambiguous (Fig. 5): the verdict-attached (logical) bias is â0â0 for every frontier model we test (A/B family: six of seven 95% CIs span zero, the exception â0.02-0.02; Haikuâs wide channels make it a bound there; Methods), while the same modelsâ pull with yes/no words reaches â0.52-0.52. Replacing the answer words with arbitrary labels removes the word-attached pull; what remains attached to the verdict itself is negligible. But what remains attached to the surface is not nothing: label-specific and order-specific pulls persist, differ across models, and do not simply track the labelâs meaning (below). The bias follows the printed label, not the verdict it carries. The swap removes the word-attached pull, not every format effect: order effects are label-specific and can persist or even growâHaikuâs order bias with A/B labels is â0.57-0.57, larger than its â0.39-0.39 with the yes/no words (label-map family; the verb-flip family of Fig. 2 gives â0.33-0.33), as if an arbitrary label, giving no semantic anchor, makes the model lean harder on position. The order channel is thus best read as a labelĂposition interactionâits sign can even flip across label families within one model (Fig. 5a)âand the fixed âtoward the last-printed optionâ direction is the yes/no-format case. And the collapse is a frontier-scope claim: both small open-weight models retain a verdict-attached biasâQwen at +0.29+0.29 with Chinese yes/no labels, and Nemotron at +0.26+0.26 [+0.21,+0.32][+0.21,+0.32] already with the fully arbitrary A/B labels (SI Appendix, section S6). If the pull is lexical, it should be graded: labels that carry more of the meaning of yes/no should pull harder. Figure 5: The yes/no pull is lexical, not logical: swapping the answer label removes the word-attached pull. (a) The verbĂlabelĂorder projections per answer-label family, per model: order bias / label (surface-label pull) / logic (verdict-attached pull carried by a non-yes/no-word label). Sign key: each cell is b=2âPâ(â )â1b=2P(¡)-1 with the positive pole = the first-printed option (order); the yes-verdict, whichever label carries it under the mapping (logic); and the pairâs nominal yes-associated surface label, counterbalanced over the labelâ mapping (label). Negative order = toward the last-printed option. The A/B logic column (the fully arbitrary family) stays pale for every frontier model (six of seven 95% CIs span zero, exception â0.02-0.02; Haikuâs are wide); only the small open-weight models retain verdict attachment (SI Appendix, Figs. S7 and S8). Cells: value Âą 95% CI half-width. (b) The collapse: the pull with yes/no words (lexical; up to â0.52-0.52) versus the same modelsâ verdict-attached pull under an arbitrary A/B label (â0â0). Scope: the swap collapses the word channel only; the order bias is label-specific and can grow (Haiku: â0.57-0.57 with A/B versus â0.39-0.39 with yes/no). n=20n=20 dilemmas per cell; bootstrap 95% CIs (Methods). The lexicality gradient. They doâon average. Across thirteen answer-label pairs ordered from non-lexical symbols through marks to real yes/no words in several languages, the label pull is near zero for non-lexical labels and grows as the label acquires the meaning of yes/noâa lexicality gradientâwhile remaining near-flat for GPT and Gemini throughout (Fig. 6). The ordering is assigned a priori by label type (non-lexical symbol / quasi-lexical mark / natural-language word), not by the measured pulls, and the summary contrastâmean ||label pull|| on word pairs minus non-lexical pairsâis +0.17+0.17 to +0.23+0.23 for the Sonnet configurations and Haikuâ ¡direct versus 0.000.00â0.020.02 for GPT-5.5 and both Geminis. The gradient is a trend, not a law, and the departures are themselves findings: the non-lexical end is not exactly zero everywhere (Sonnetâs A/B â0.15-0.15; the valenced +âŁ/âŁâ+/- mark is instrument-dependent for HaikuâFig. 5a and SI Appendix, section S5); and which words pull is model-relativeâSonnetâs pull on the Chinese yes/no pair is comparable to its English one (Fig. 6), while the open-weight modelsâ pulls peak on non-English words and can even reverse sign across languages (Qwen: ja/nein â0.67-0.67 and Chinese â0.50-0.50 against a near-zero English pull; Nemotron: toward ânoâ with English labels, toward âouiâ with French; SI Appendix, section S6)âconsistent with lexical loading acquired from training exposure rather than carried by meaning alone. A single wording could neither detect nor rank any of this; the crossing is what makes the zoo measurable. The gradient has a zero: a binary elicitation with no yes/no label at allâchoosing directly between the two courses of actionâshows |bias|<0.1|bias|<0.1 for every model (SI Appendix, section S5), and a free-text choice is more coherent still. The artifact is not the cost of forcing a binary; it is the cost of forcing it through yes/no. Figure 6: The pull is graded by the labelâs yes/no-nessâon averageâand it is model-specific. Thirteen answer-label pairs, ordered non-lexical (A/B, 1/2, $/%, +âŁ/âŁâ+/-) â quasi-lexical (Y/N; chk = check/cross marks; thumb = thumbs-up/down; T/F = true/false) â lexical (ja/nein, oui/non, ZH = Chinese yes/no characters, âI do/do not,â and âstoryâ = the storyâs own pro/anti phrases); per pair, two bars: order bias (b=2âPâ(firstâ-âprinted)â1b=2P(first -printed)-1, so negative = toward the last-printed option; orange) and label pull (b>0b>0 = emits the yes-mapped label; blue). Non-lexical labels carry little pull on average (mean ||label pull|| 0.030.03â0.100.10; exceptions in the textâ+âŁ/âŁâ+/- is a valenced mark, and its pull is instrument-dependent for Haiku) and the pull grows as the label becomes a real yes/no word (word-minus-non-lexical contrast +0.17+0.17 to +0.23+0.23 for the Sonnet configurations and Haikuâ ¡direct, â¤0.03â¤0.03 for Haikuâ ¡think, GPT-5.5, and both Geminis; the open-weight models: SI Appendix, Figs. S7 and S8âthe second shows model-idiosyncratic label pulls rather than a gradient). n=20n=20 dilemmas per pair per model. Error bars indicate 95% CIs. Discussion Measured right, an LLMâs moral stance is a stable measurement. Across the logically irrelevant elicitation frames we crossedârating scale, anchor direction, wording of the rating question, poleâa frontier modelâs graded stance varies with a between-form SD of only 0.120.12â0.210.21 on the Âą1Âą 1 axis, and deliberation tightens it further. What is not stable is the standard way of reading it: a forced yes/no adds a format artifact that can dwarf the signal (story-averaged, up to â0.86-0.86), decomposes exactly into an order pull and a word pull, carries no verdict-attached component where that channel is precisely measured (a bound, not a zero, for Haiku), and loses its word-attached part under a label with no yes/no meaningâwhile residual, label-specific order and surface pulls can remain. The natural reading of a yesâno biasâthat a forced model leans toward saying no, a disposition to rejectâis therefore not what the measurement finds: where the word and the verdict are separated, no verdict-attached component survives; what survives is attachment to surfaces. For measurement the corollary is concrete. To read a language modelâs moral stance: elicit it graded, under crossed logically-equivalent frames; report the cross-form incoherence as the error bar, because itânot sampling noiseâis the dominant uncertainty; and if a binary verdict is unavoidable, force it through a choice between the options rather than through yes/no. Any single-format number, including the standard yesâno bias, confounds the stance with the format. Cheung et al. [9] correctly identified the phenomenon: the yesâno bias of LLMs is real and large. What a single-framing design cannot do, in principle, is resolve its composition. Ours does, and the resolution is that the amplification is formatâan order pull with no human analog, plus a lexical pull toward a wordâwhile the moral scale beneath it is more coherent than the artifact makes it look. Because the models differ across studies, we engage Cheung et al.âs argument, not their numbers; what makes the engagement direct is the shared materialsâtheir vignettes verbatim, only the closing question exchanged [9, 49]. The measurement idea also fixes the paperâs place in its neighborhood. One literature measures what models answer on moral items [24, 25, 50, 13, 34, 30, 51, 52, 53, 31, 32, 35, 54, 55, 56, 57, 58, 59]; another measures how much any answer moves with presentation [15, 16, 60, 22, 61, 62]âand the movementâs direction is itself unsettled across tasks: primacy on shuffled label lists [63], recency across survey options [64], token-identity pulls in multiple choice [17], and a ânoâ-leaning inversion of classic acquiescence [65] that matches the direction we measure. Of fourteen moral-judgment studies we examined in depth, five elicit each judgment in one fixed wording, four vary two to four uncrossed reframings, and five vary form systematically while crossing at most two presentation factors; none crosses verb, order, and label, and none uses such a factorial to decompose the bias into mechanismsâthe nearest designs score perturbation sets as consistency rates [10], average template-and-order variation out rather than attributing it [13], or remap the answer scale without a verdict factor [11, 12, 36]. Treating models as subjects of behavioral measurement is by now a program [66, 67, 68, 69, 70, 27, 71, 72, 73], including direct humanâmodel moral comparison [74, 75, 76, 33, 77], and a psychometric turn in LLM evaluation is underway [78, 79, 80, 21, 81, 82, 83, 84, 85]; our contribution to it is the instrument: crossed symmetrization turns format sensitivity from a nuisance to be averaged over into a measured, decomposed quantity, and it answers, on these materials, Cheung et al.âs own call for evaluations that assess consistency across logically equivalent questions [9]. The human-side canon reads here as a design manual rather than a comparison target: order and framing move human moral judgments too [5, 6, 7, 8, 86, 87, 88], repetition alone yields drift [89], judgments and choices dissociate [90, 91, 92], and the survey-methodology classics supply the levers we crossed [39, 42, 43, 93, 94, 95]. Humans, on the same materials, are framing-robust (a mean framing shift of â0.12â0.12 in choice proportion, â0.24â0.24 on our Âą1Âą 1 convention, and not ânoâ-directed)âand where a frontier modelâs verdict probabilities saturate toward 0 or 1, the human sample splits, sitting near 50/50 on the same forced choices. Both facts are from Cheung et al.âs existing data [9]; we run no new human experiments. Each person answered once, so the split cannot be parsed into hedging individuals versus decisive individuals who disagreeâwhich is why we speak of the modelsâ saturation against the populationâs split, and translate neither into the within-agent (m,s)(m,s) vocabulary; nor do we compare framing-stability magnitudes across species (a mean directional shift and a cross-form SD are not commensurate). What the human data anchor is the aggregate contrast: at the same level of description, the modelsâ amplification has no counterpart there. The claims carry their scopes. The logistic is the canonical link from a continuous, coherent latent to a binary choiceâa modeling principle, not a mechanism claimâand the interpretation of (m,s)(m,s) does not depend on that choice of link; but absolute s is not comparable where the readout clips (Haikuâ ¡direct, the Geminis, the open-weight modelsâwhich are additionally read on their own θ rulers), so there we trust only the m-ranking and within-model differences. âLogical bias â0â0â is a statement about these contested sacrificial dilemmas and the frontier models we observe, not about morality at large. Level-1 coherence is substantial, not perfectâa residual anchor-direction floor of 0.120.12â0.210.21 remainsâand it is within-model: the shared cross-model ordering is a side observation (SI Appendix, Table S2), not evidence of an objective moral axis; and because nineteen dilemmas are verbatim from a published study, the shared ordering and the human tracking could partly reflect training exposure to the materials and their published ratingsâthe within-model invariance and the artifact decomposition, which compare logically equivalent forms of the same exposure, are immune to this, and they are what the paperâs claims rest on. Gemini-3-Flashâs inverted ordering on several dilemmas is a genuine stance difference, verified in its raw responses, and is reported as such rather than excluded (SI Appendix, section S8). The small-model contrast rests on two open-weight models, and they fail on different axes: Qwen is the deterministic endpoint of the stance-incoherence spectrumâa differentiated, consensus-correlated ordering buried under cross-form scatter that deliberation removesâwhile Nemotron is the instrumentâs degenerate case, a scale too compressed to differentiate the items (G=0.28G=0.28), which we accordingly report in the SI rather than beside the functioning scales, and which nonetheless retains a verdict-attached bias no frontier model shows (SI Appendix, section S6). Neither pattern is a statement about open models generally; together they carry a methodological moral: a coherence number without a discrimination check can flatter a scale that is merely indifferentâlow spread is not a scaleâwhich is exactly what the generalizability coefficient is for. Finally, in the models carrying a substantial artifact the same signature appears at both levelsâdeliberation tightens the scale and shrinks the susceptibilityâsuggesting a fast-pathway artifact that reflection partially overrides (a resonance with dual-process accounts of human moral judgment [96, 97, 98, 99, 100], offered as analogy, not mechanism); the degenerate case is again the instructive exception: with no functioning scale to tighten, deliberation wakes Nemotronâs scale (G 0.28â0.560.28\!â\!0.56, consensus alignment r2r^2 0.22â0.690.22\!â\!0.69) without lowering its spread, and inflates its binary artifactâdeliberation is a lever, not a guarantee (SI Appendix, section S6). One dimension is deliberately held fixed throughout: the responderâs persona. Every judgment here is elicited from each modelâs default assistant stance; how the scale and its artifacts move when the same model answers under systematically varied personas or system prompts is a question the instrument takes up unchanged. The battery is not tied to these materials: the instruments and the crossing principle apply unchanged to any dilemma set and extend to richer form families. Expanding all of theseâmore dilemmas, more forms, the story text itself as a manipulated variable rather than a held-fixed one, and the responderâs persona as a crossed factorâand measuring deliberation and commitment as objects in their own right, are the natural next steps. Materials and Methods Dilemmas and materials. Twenty moral dilemmas: 13 sacrificial/greater-good dilemmas and 7 omission counterparts, rendered from a faithful transcription of the vignette materials of Cheung, Maier, and Lieder [9, 49]. Nineteen items are verbatim Cheung (per-item story_sha256 over all âź200 200k trials collapses to the hash of the source text); one (item A) is adapted from an omission default to an active decision and therefore carries no human anchor (SI Appendix, section S7). The dilemmas span personal/impersonal sacrifice, altruistic redistribution, and assisted-dying scenarios; each has an action pole (commit the act; the costâbenefit-reasoning pole, CBR) and a rule/deontological pole. Model panel. Seven frontier configurations across three familiesâClaude Sonnet 4.6 (claude-sonnet-4-6) and Claude Haiku 4.5 (claude-haiku-4-5), each with extended reasoning off/on (âdirectâ/âthinkâ); GPT-5.5 (gpt-5.5, low reasoning effort); Gemini-3-Flash (gemini-3-flash-preview, high thinking) and Gemini-3.1-Flash-Lite (gemini-3.1-flash-lite)âplus Claude Opus 4.8 (claude-opus-4-8, maximum effort; instruments I-1/I-2 only) and two small open-weight mixture-of-experts models read out at the logit level (greedy, single pass: the full answer-slot distribution) in both reasoning modes: Qwen3.6-35B-A3B and NVIDIA Nemotron-3-Nano-30B-A3B, the latter collected as a full replication of the formerâs battery (same forms, same salt strings, same behavioral design; SI Appendix, section S4). GPT-5.5âs run additionally probed two pilot items; all cross-model analyses use the shared 20. âFrontierâ claims about the forced binary quantify over the seven I-3 configurations; Level-1 (I-1) statements include Opus. Only the Claude models expose a true no-reasoning mode; GPT-5.5 and the Geminis run with nonzero reasoning by default (low and high effort, respectively). Every prompt is a fresh, context-free, single-turn API call, and every sampled model runs at temperature 1 (SI Appendix, section S4). The Claude models receive an empty system prompt; GPT-5.5 the one-line instruction âRespond only to the userâs survey stimulus. Do not use tools.â; the Gemini models the one-line system prompt âYou are a helpful assistant. Answer the userâs question directly and follow their formatting instructions exactly.â Exact invocations, decoding settings, and refusal handling are in SI Appendix, section S4. Refusals and unusable answers are excluded from cells; on the core verb-flip pool the rates are Sonnet 0.7%/0.5%0.7\%/0.5\% (direct/think), Haiku 28%/35%28\%/35\% (textual refusals), GPT-5.5 0.0%0.0\%, Gemini-3-Flash 1.7%1.7\%, and Flash-Lite 24%24\% (verdict-withholding essays); the matched 12-form comparison (below) is computed on the items whose cells survive exclusion. Resampling noise is quantified directly on both instruments. On the graded scale the repeat component is Ďrepeat=0.02 _repeat=0.02â0.100.10, small beside Ďrepro _repro (SI Appendix, section S3). On the forced binary, 82â92% of formĂitemĂside cells are unanimous across replicates for the well-behaved configurations, between-form variance exceeds within-form (replicate) variance by up to an order of magnitude (Sonnet âź9 9â19Ă19Ă on the matched pool, âź25Ă 25Ă with print order included), and 15â50% of items have at least two logically equivalent forms yielding opposite majority verdicts; the heavy-refusal and saturated configurations are the expected boundary cases (SI Appendix, section S3). Repetition of a single form therefore underrepresents the variation that matters. The frontierâopen-weight comparison crosses two readout channels (sampled text versus greedy logits); each open-weight modelâs behavioral samples calibrate its logit readout (behavioral versus logit choice probability: Qwen râ0.90râ 0.90, Nemotron r=0.93r=0.93 at 24 replicates, with fitted readout temperature â1â1; SI Appendix, section S4), and its salt replications supply the within-form noise floor the greedy channel otherwise hides. In total âź389,000 389,000 primary trials (I-1 34,35634,356; I-2 6,4756,475; I-3 348,151348,151; plus 2,8802,880 Opus graded trials, indexed separately) and âź47,000 47,000 auxiliary open-weight-model trials (salt replications and behavioral calibration, both models). I-1: the graded scale and θ. Each dilemma is rated under 48 crossed conditions: scale â0â\0â1010, 0â100100, and two bipolar variants\ Ă wording â\acceptable, right, should\ Ă anchor direction â\ascending, descending\ Ă pole â\action, complement\ (the full design table, with real prompt text for every instrument: SI Appendix, section S1). The two bipolar variants are excluded from θ by the frozen, pre-registered protocol (instrument definitions and estimator in timestamped files that precede the first production run; SI Appendix, section S9): they deliberately reintroduce the negative-anchor positivity bias known from the survey literature [101, 93, 102, 103], as a measured contrast to be quantifiedânever averaged into the stance. The measured contrast is small (per-model mean |θbipolarâθ|â¤0.06| _bipolar-θ|⤠0.06); recomputing the per-item Ďrepro _repro with the bipolar cells included moves each frontier modelâs mean by at most 0.040.04 (Qwenâs by 0.060.06, downward; Nemotronâs by at most 0.020.02; SI Appendix, section S3). This leaves 24 unipolar conditions in the θ estimator. A rating v is normalized to acceptability a=(vâlo)/(hiâlo)a=(v-lo)/(hi-lo), flipped to 1âa1-a under a descending anchor. Per item, θ=aÂŻactionâaÂŻcomplementâ[â1,1]θ= a_action- a_complementâ[-1,1]. θ is a chosen coordinate (no absolute zero or unit). Replication depth: 3â8 per condition for the sampled models (open-weight models: single logit pass; error bars from salt replications, below). I-2: free choice. The model freely chooses between the two courses of action under twelve framings, with no scale and no yes/no label; verdicts (PRO/ANTI/NONCOMMIT) are judged from the free text by Claude Opus 4.8 (high effort). Verdict extraction is a transcription taskâthe transcript states a choiceâbut the judge shares a vendor family with several subjects, a circularity risk we flag; per-form verdicts and raw transcripts ship with the data for audit. I-2 supplies the convergent-validity checks of SI Appendix, section S5: it is the most coherent elicitation we observe (cross-form Ďreproâ0 _reproâ 0) and the most saturated (|stance|â0.86|stance|â 0.86 on the Âą1Âą 1 axis). I-3: the forced binary, and the projection framework. Verb-flip forms cross the questionâs verb (approve/oppose) with the printed answer order (yes-first/no-first); of the 126 forms in the design, Claude ran 49 (Sonnetâ ¡think: 50) and every other model the 12-form baseline, on all 20 items (replication r8 for Sonnet/GPT/Flash-Lite, r10â12 across batches for Gemini-3-Flash, r4 for Haikuâa budget choice fixed in advanceâand r1 logit for the open-weight models). Because the corpora differ in breadth, the cross-model comparison was re-run on the matched 12-form baseline for all models: the Claude decompositions move by â¤0.06⤠0.06 and away from zero (Sonnet â0.32=â0.21â0.12-0.32=-0.21-0.12; Haiku â0.92-0.92, on the 17 items whose cells survive refusal exclusion there), so the family contrast is not a wording-corpus artifact (script ref2_matched12.py; SI Appendix, section S4). The label-map instruments below are separate corpora run on every model (verbĂlabelĂorder: 26,61726,617 trials; the 13-pair labelĂorder map: 45,09245,092 trials; full census in SI Appendix, section S4). From the four verbĂorder cells with per-cell Pâ(CBR)P(CBR), the storyâs I-3 stance is the mean over all four, z=â¨2âPâ(CBR)â1âŠz= 2P(CBR)-1 âone value per story, independent of which bias is examinedâand each bias is a balanced 2-vs-2 split of the same cells sharing that stance: b=p+âpâ,12â(p++pâ)=Pâ(CBR)ÂŻ,bapparent=border+blexical. splitb\;&=\;p_+-p_-, 12\,(p_++p_-)= P(CBR),\\ b_apparent\;&=\;b_order+b_lexical. split (1) Concretely, with biaso=Pâ(CBR|approve,o)âPâ(CBR|oppose,o)bias_o=P(CBR\,|\,approve,o)-P(CBR\,|\,oppose,o) for each printed order o: the lexical channel is 12â(biasyf+biasnf) 12(bias_yf+bias_nf) (the order-balanced verb split, equal to â¨2âPâ(yes)â1⊠2P(yes)-1 ); the order channel is 12â(biasyfâbiasnf) 12(bias_yf-bias_nf); and the apparent (single-framing, yes-first) bias is biasyfbias_yf, so the identity in Eq. 1 is exact and definitionalâits content is that the two summands are separately measurable and separately meaningful: the order part flips with print order, the lexical part survives order-balancing. The independent axis θ always comes from I-1, so the bias test never uses the biased format to measure the stance it is biased about. Identification: lexical versus logical. In yes/no data the word and the verdict are one token, so the verb-flip lexical channel conflates them. The full verbĂlabelĂorder design (2Ă2Ă22Ă 2Ă 2) carries the verdict on labels other than the English wordsâthe fully arbitrary A/B; the valenced mark pair +âŁ/âŁâ+/-; Chinese yes/no words for the cross-lingual lexical caseâadding the logic projection (attachment to the verdict when the verdict is carried by a label without the English word) and the label projection (attachment to the printed label itself, counterbalanced over the labelâ mapping). The identification claim rests on the fully arbitrary family: empirically the A/B logic is â0â0 for every frontier model (Fig. 5), which licenses the cheaper labelĂorder 2Ă22Ă 2 used for the 13-pair gradient of Fig. 6. The two-parameter fit. For each bias source we fit (s,m)(s,m) jointly to that biasâs two sides by nonlinear least squares, p+=Ďâ((θ+m)/s)p_+=Ď((θ+m)/s) and pâ=Ďâ((θâm)/s)p_-=Ď((θ-m)/s) stacked, with θ from I-1; the shared stance 12â(p++pâ) 12(p_++p_-) pins the origin, so s is not inflated by averaging over the offset. Each fit generates the stance sigmoid, the bowtie locus, and the ridge of Fig. 3; because zâ(θ)z(θ) depends only weakly on m, the order-fit and lexical-fit stance curves must overlap, which they do (the check is weak by construction; goodness of fit is reported separately). Goodness of fit on the stacked two-side data: R2=0.77R^2=0.77â0.840.84 (RMSE 0.140.14â0.170.17) for the resolvable configurations; Haikuâ ¡directâs order channel is worst (R2=0.57R^2=0.57), consistent with its degeneracy flag. Saturated or degenerate configurations read as clip-limited s (Haikuâ ¡direct large-s degenerate; the Geminisâ saturated stance reads as small s); for these only the m-ranking and within-model differences are interpreted. Identifiability of m varies with saturation: the number of dilemmas within |z|<0.9|z|<0.9 of a modelâs tie is 3 (Gemini-3-Flash), 7 (GPT-5.5), 10 (Flash-Lite), and 11â19 (the Claude configurations); a saturated modelâs fitted m is additionally attenuated toward zero. Estimator properties: θ enters as a regressor with incoherence of its own (drawn as x error bars in Fig. 3), and such errors-in-variables attenuate |m||m| and inflate s, so the fitted susceptibilities are conservative; m is the aggregate offset of a wording familyâitem-level residuals of Âą0.5Âą 0.5 exist and cancel (the K01 anchors)âso âportableâ means a story-independent estimator, not an item-homogeneous mechanism. CIs are percentile bootstraps over dilemmas (⼠2,000â10,000 resamples, fixed seeds), with no multiplicity correction; percentile intervals on n=20n=20 clusters can undercover slightly, so boundary-grazing limits (e.g., +0.001+0.001 or â0.002-0.002) are read as marginal (coverage discussion: SI Appendix, section S3). The symbol b always denotes a signed bias on the [â1,1][-1,1] axis; each figureâs per-family definition is the same quantity computed in that familyâs cells. Raw versus fitted. The story-averaged biases (Fig. 2) and the fitted parameters (Fig. 4) are two presentations of the same verbĂorder data. The raw mean is story-distribution-dependent (saturation collapses it toward zero far from the tie); the fitted (s,m)(s,m) are story-independent and diagnose that saturation. Neither replaces the other: raw numbers compare models on a fixed item set, parameters are portable. Cross-form incoherence (the reported error bar). Every error bar on a stance-like quantity is the format incoherence of that quantity, not the standard error of its mean: the spread of the per-form value across logically equivalent forms, corrected for finite sampling (a measurement-systems âGauge R&Râ correction [104]), Ďrepro=maxâĄ(0,Varformsâ(xf)âVarsampâ(xf)ÂŻ) _repro= (0,\ Var_forms(x_f)- Var_samp(x_f)), where xfx_f is the per-form value (VarformsVar_forms over the format cellsâI-1: up to 24 per item; I-3: the familyâs formsâand VarsampVar_samp from the 3â8 replicates per cell). The maxâĄ(0,â ) (0,¡) floor means small values are resolution-limited rather than exact zeros. Ďrepro _repro is expressed in the units of the [â1,1][-1,1] axis; cross-model comparison rests on the shared normalization and identical item set. Components combine in quadrature (Ďrepro2=Ďdir2+Ďscale2 _repro^2= _dir^2+ _scale^2, treated as additive; Ďtot2=Ďrepro2+Ďrepeat2 _tot^2= _repro^2+ _repeat^2), so the decomposition of Fig. 1b is drawn by variance share, never stacked linearly. Two estimators of the per-model incoherence appear: the headline values and the Fig. 1b bars are the frozen per-model aggregate (a pooled variance decomposition across all format cells), while uncertainty statements use the mean per-item Ďrepro _repro with a 10,00010,000-resample bootstrap over items; the estimator comparison is tabulated in SI Appendix, section S3. The generalizability coefficient [105, 106, 107, 108] is G=Varitemsâ(θ)/(Varitemsâ(θ)+Ďrepro2ÂŻ)G=Var_items(θ)\,/\,(Var_items(θ)+ _repro^2), an ICC analog for a single randomly chosen form. Two error-bar regimes appear in the paper: incoherence spreads (Figs. 1a and 3; descriptive) and bootstrap CIs (Figs. 2 and 4â6; inferential)âoverlap comparisons are licensed only for the latter. For the deterministic open-weight logit readout, genuine replications are created by salting: prepending a meaningless hexadecimal context string. Salt is itself a logically irrelevant perturbationâof the context, not the question formâand the argument is comparative: if the between-form spread were sampling noise, the same spread would appear within form under salt; it does not. Salting leaves θ unchanged (r=0.993r=0.993 to the unsalted pass) while exposing the within-form floor; the between-form SD within the salt experimentâs unipolar subset (0.3450.345; the full-design quadrature total is the 0.400.40 quoted throughout) versus the within-form/salt SD (mean 0.0720.072; median 0.0000.000) establishes that its incoherence is dominantly deterministic, not sampling noise (SI Appendix, section S3). This determinism claim is Qwenâs; Nemotronâs salt floor is different in kindâubiquitous, small logit jitter (mean |Îâp|=0.02| p|=0.02)âand is likewise subtracted by the correction (SI Appendix, section S6). Reading ââ0â0.â Throughout, ââ0â0â summarizes a point estimate |b|â¤0.06|b|⤠0.06 whose CI spans zero; it is a non-rejection shorthand, not an equivalence testâwide channels (e.g., Haikuâs logic cells, Âą0.2Âą 0.2â0.30.3) are underpowered rather than null and are flagged as such, and the same reading is applied on both sides of zero. The Cheung-verbatim control. One verb-flip family (K01) reproduces Cheung et al.âs question wording verbatimâthe one-line answer instruction is our standardized oneâand crosses only the printed answer order on top of it, across all 20 items at full depth (r8; run on Sonnet, both reasoning modes). Its yes-first apparent bias (â0.12-0.12) order-balances to +0.01+0.01, 95% CI [â0.12,+0.16][-0.12,+0.16] (cluster bootstrap over items)âconsistent with no systematic residual at this precision, though the interval does not exclude word-level residuals of the size reported for our own wording families; per-item anchor residuals are real (up to Âą0.5Âą 0.5) and cancel across items. Human data. All human values are re-extracted from Cheung et al.âs published data [9, 49] (Study 1, N=285N=285, within-subjects, action framing; Study 2, N=474N=474, between-subjects): per-dilemma choice rates, averaged over framing conditions, supply the Fig. 1a anchors; the framing shift (â0.12â0.12 in choice proportion; Study 2) and the near-50/50 split are quoted in the Discussion (each participant judged each dilemma once). We collected no new human data. Data availability All raw trials (JSON/JSONL), derived quantities, and analysis code regenerate every number and figure deterministically from raw via a single pipeline (run_all.py; one canonical function per quantity); provenance, pre-registration artifacts, and integrity notes are consolidated in SI Appendix, section S9. [Deposition DOI to be added at submission.] Author contributions. H.H. conceived the project, designed the research, performed the experiments, analyzed the data, audited the code, and wrote the manuscript. Competing interests. The authors declare no competing interest. References Tversky and Kahneman [1981] Amos Tversky and Daniel Kahneman. The framing of decisions and the psychology of choice. Science, 211(4481):453â458, 1981. doi: 10.1126/science.7455683. Kahneman and Tversky [1984] Daniel Kahneman and Amos Tversky. Choices, values, and frames. American Psychologist, 39(4):341â350, 1984. doi: 10.1037/0003-066X.39.4.341. Levin et al. [1998] Irwin P. Levin, Sandra L. Schneider, and Gary J. Gaeth. All frames are not created equal: A typology and critical analysis of framing effects. Organizational Behavior and Human Decision Processes, 76(2):149â188, 1998. doi: 10.1006/obhd.1998.2804. KĂźhberger [2023] Anton KĂźhberger. A systematic review of risky-choice framing effects. EXCLI Journal, 22:1012â1031, 2023. doi: 10.17179/excli2023-6169. Petrinovich and OâNeill [1996] Lewis Petrinovich and Patricia OâNeill. Influence of wording and framing effects on moral intuitions. Ethology and Sociobiology, 17(3):145â171, 1996. doi: 10.1016/0162-3095(96)00041-6. Wiegmann et al. [2012] Alex Wiegmann, Yasmina Okan, and Jonas Nagel. Order effects in moral judgment. Philosophical Psychology, 25(6):813â836, 2012. doi: 10.1080/09515089.2011.631995. Schwitzgebel and Cushman [2012] Eric Schwitzgebel and Fiery Cushman. Expertise in moral reasoning? order effects on moral judgment in professional philosophers and non-philosophers. Mind & Language, 27(2):135â153, 2012. doi: 10.1111/j.1468-0017.2012.01438.x. Schwitzgebel and Cushman [2015] Eric Schwitzgebel and Fiery Cushman. Philosophersâ biased judgments persist despite training, expertise and reflection. Cognition, 141:127â137, 2015. doi: 10.1016/j.cognition.2015.04.015. Cheung et al. [2025] Vanessa Cheung, Maximilian Maier, and Falk Lieder. Large language models show amplified cognitive biases in moral decision-making. Proc. Natl. Acad. Sci. U.S.A., 122(25):e2412015122, 2025. doi: 10.1073/pnas.2412015122. Oh and Demberg [2025] Soyoung Oh and Vera Demberg. Robustness of large language models in moral judgements. Royal Society Open Science, 12(4):241229, 2025. doi: 10.1098/rsos.241229. Marraffini et al. [2024] Giovanni Franco Gabriel Marraffini, AndrĂŠs Cotton, NoĂŠ FabiĂĄn Hsueh, Axel Fridman, Juan Wisznia, and Luciano del Corro. The greatest good benchmark: Measuring LLMsâ alignment with utilitarian moral dilemmas. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21950â21959, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.1224. Yuan et al. [2024] Jiaqing Yuan, Pradeep K. Murukannaiah, and Munindar P. Singh. Right vs. right: Can LLMs make tough choices?, 2024. Scherrer et al. [2023] Nino Scherrer, Claudia Shi, Amir Feder, and David M. Blei. Evaluating the moral beliefs encoded in LLMs. In 37th Conference on Neural Information Processing Systems (NeurIPS 2023), 2023. Keshmirian et al. [2025] Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, and Lav R. Varshney. Many LLMs are more utilitarian than one. In 39th Conference on Neural Information Processing Systems (NeurIPS 2025), 2025. Sclar et al. [2024] Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. Quantifying language modelsâ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting. In International Conference on Learning Representations (ICLR), 2024. Pezeshkpour and Hruschka [2024] Pouya Pezeshkpour and Estevam Hruschka. Large language models sensitivity to the order of options in multiple-choice questions. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages 2006â2017, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-naacl.130. URL https://aclanthology.org/2024.findings-naacl.130/. Zheng et al. [2024] Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. Large language models are not robust multiple choice selectors. In International Conference on Learning Representations (ICLR), 2024. Dominguez-Olmedo et al. [2024] Ricardo Dominguez-Olmedo, Moritz Hardt, and Celestine Mendler-DĂźnner. Questioning the survey responses of large language models. In Advances in Neural Information Processing Systems (NeurIPS), volume 37, 2024. Tjuatja et al. [2024] Lindia Tjuatja, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. Do LLMs exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics, 12:1011â1026, 2024. doi: 10.1162/taclËaË00685. RĂśttger et al. [2024] Paul RĂśttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schuetze, and Dirk Hovy. Political compass or spinning arrow? towards more meaningful evaluations for values and opinions in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15295â15311, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.816. URL https://aclanthology.org/2024.acl-long.816/. Mizrahi et al. [2024] Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. State of what art? a call for multi-prompt LLM evaluation. Transactions of the Association for Computational Linguistics, 12:933â949, 2024. doi: 10.1162/taclËaË00681. Zhu et al. [2024] Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zeek Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, and Xing Xie. Promptrobust: Towards evaluating the robustness of large language models on adversarial prompts. In Proceedings of the 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis (LAMPS â24), 2024. doi: 10.1145/3689217.3690621. Salewski et al. [2023] Leonard Salewski, Stephan Alaniz, Isabel Rio-Torto, Eric Schulz, and Zeynep Akata. In-context impersonation reveals large language modelsâ strengths and biases. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 2023. Hendrycks et al. [2021] Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, and Jacob Steinhardt. Aligning AI with shared human values. In International Conference on Learning Representations (ICLR 2021), 2021. Jiang et al. [2021] Liwei Jiang, Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jenny Liang, Jesse Dodge, Keisuke Sakaguchi, Maxwell Forbes, Jon Borchardt, Saadia Gabriel, Yulia Tsvetkov, Oren Etzioni, Maarten Sap, Regina Rini, and Yejin Choi. Can machines learn morality? the delphi experiment, 2021. Zheng et al. [2023] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023), Datasets and Benchmarks Track, 2023. Santurkar et al. [2023] Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, volume 202 of PMLR, pages 29971â30004, 2023. Renze and Guven [2024] Matthew Renze and Erhan Guven. The effect of sampling temperature on problem solving in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 7346â7356, 2024. doi: 10.18653/v1/2024.findings-emnlp.432. Song et al. [2025] Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of LLMs should not ignore non-determinism. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4195â4206, 2025. doi: 10.18653/v1/2025.naacl-long.211. Nie et al. [2023] Allen Nie, Yuhui Zhang, Atharva Shailesh Amdekar, Chris Piech, Tatsunori B. Hashimoto, and Tobias Gerstenberg. MoCa: Measuring human-language model alignment on causal and moral judgment tasks. In Advances in Neural Information Processing Systems, volume 36, pages 78360â78393, 2023. URL https://proceedings.neurips.c/paper_files/paper/2023/hash/f751c6f8bfb52c60f43942896fe65904-Abstract-Conference.html. Abdulhai et al. [2024] Marwa Abdulhai, Gregory Serapio-GarcĂa, Clement Crepy, Daria Valter, John Canny, and Natasha Jaques. Moral foundations of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17737â17752, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.982. URL https://aclanthology.org/2024.emnlp-main.982/. Nunes et al. [2024] JosĂŠ Luiz Nunes, Guilherme F. C. F. Almeida, Marcelo de Araujo, and Simone D. J. Barbosa. Are large language models moral hypocrites? a study based on moral foundations. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 7, pages 1074â1087, 2024. doi: 10.1609/aies.v7i1.31704. Takemoto [2024] Kazuhiro Takemoto. The moral machine experiment on large language models. Royal Society Open Science, 11(2):231393, 2024. doi: 10.1098/rsos.231393. Jin et al. [2022] Zhijing Jin, Sydney Levine, Fernando Gonzalez, Ojasv Kamal, Maarten Sap, Mrinmaya Sachan, Rada Mihalcea, Joshua Tenenbaum, and Bernhard SchĂślkopf. When to make exceptions: Exploring language models as accounts of human moral judgment. In 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 2022. Hadar-Shoval et al. [2024] Dorit Hadar-Shoval, Kfir Asraf, Yonathan Mizrachi, Yuval Haber, and Zohar Elyoseph. Assessing the alignment of large language models with human values for mental health integration: Cross-sectional study using Schwartzâs theory of basic values. JMIR Mental Health, 11:e55988, 2024. doi: 10.2196/55988. Ji et al. [2024] Jianchao Ji, Yutong Chen, Mingyu Jin, Wujiang Xu, Wenyue Hua, and Yongfeng Zhang. MoralBench: Moral evaluation of LLMs, 2024. Campbell and Fiske [1959] Donald T. Campbell and Donald W. Fiske. Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2):81â105, 1959. doi: 10.1037/h0046016. Cronbach and Meehl [1955] Lee J. Cronbach and Paul E. Meehl. Construct validity in psychological tests. Psychological Bulletin, 52(4):281â302, 1955. doi: 10.1037/h0040957. Schuman and Presser [1981] Howard Schuman and Stanley Presser. Questions and Answers in Attitude Surveys: Experiments on Question Form, Wording, and Context. Academic Press, New York, 1981. ISBN 0126313504. Billiet and McClendon [2000] Jaak B. Billiet and McKee J. McClendon. Modeling acquiescence in measurement models for two balanced sets of items. Structural Equation Modeling: A Multidisciplinary Journal, 7(4):608â628, 2000. doi: 10.1207/S15328007SEM0704Ë5. Van Vaerenbergh and Thomas [2013] Yves Van Vaerenbergh and Troy D. Thomas. Response styles in survey research: A literature review of antecedents, consequences, and remedies. International Journal of Public Opinion Research, 25(2):195â217, 2013. doi: 10.1093/ijpor/eds021. Krosnick and Alwin [1987] Jon A. Krosnick and Duane F. Alwin. An evaluation of a cognitive theory of response-order effects in survey measurement. Public Opinion Quarterly, 51(2):201â219, 1987. doi: 10.1086/269029. Krosnick [1991] Jon A. Krosnick. Response strategies for coping with the cognitive demands of attitude measures in surveys. Applied Cognitive Psychology, 5(3):213â236, 1991. doi: 10.1002/acp.2350050305. McCullagh and Nelder [1989] Peter McCullagh and John A. Nelder. Generalized Linear Models. Chapman and Hall, London, 2nd edition, 1989. Rasch [1960] Georg Rasch. Probabilistic Models for Some Intelligence and Attainment Tests. Danish Institute for Educational Research, Copenhagen, 1960. Reissued 1980, Chicago: University of Chicago Press, foreword by Benjamin D. Wright. Birnbaum [1968] Allan Birnbaum. Some latent trait models and their use in inferring an examineeâs ability. In Frederic M. Lord and Melvin R. Novick, editors, Statistical Theories of Mental Test Scores, pages 397â479. Addison-Wesley, Reading, MA, 1968. Lord [1980] Frederic M. Lord. Applications of Item Response Theory to Practical Testing Problems. Lawrence Erlbaum Associates, Hillsdale, NJ, 1980. Embretson and Reise [2000] Susan E. Embretson and Steven P. Reise. Item Response Theory for Psychologists. Multivariate Applications Book Series. Lawrence Erlbaum Associates, Mahwah, NJ, 2000. Maier et al. [2025] Maximilian Maier, Vanessa Cheung, and Falk Lieder. Code and data for analyses in âlarge language models show amplified cognitive biases in moral decision-makingâ. Open Science Framework, https://osf.io/3kvjd/, 2025. Deposited 4 November 2025. Schramowski et al. [2022] Patrick Schramowski, Cigdem Turan, Nico Andersen, Constantin A. Rothkopf, and Kristian Kersting. Large pre-trained language models contain human-like biases of what is right and wrong to do. Nature Machine Intelligence, 4(3):258â268, 2022. doi: 10.1038/s42256-022-00458-8. Simmons [2023] Gabriel Simmons. Moral mimicry: Large language models produce moral rationalizations tailored to political identity. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), pages 282â297, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-srw.40. URL https://aclanthology.org/2023.acl-srw.40/. First posted as arXiv:2209.12106 (Sept. 2022); published version ACL 2023 SRW. Almeida et al. [2024] Guilherme F. C. F. Almeida, JosĂŠ Luiz Nunes, Neele Engelmann, Alex Wiegmann, and Marcelo de AraĂşjo. Exploring the psychology of LLMsâ moral and legal reasoning. Artificial Intelligence, 333:104145, August 2024. doi: 10.1016/j.artint.2024.104145. Jin et al. [2025] Zhijing Jin, Max Kleiman-Weiner, Giorgio Piatti, Sydney Levine, Jiarui Liu, Fernando Gonzalez, Francesco Ortu, AndrĂĄs Strausz, Mrinmaya Sachan, Rada Mihalcea, Yejin Choi, and Bernhard SchĂślkopf. Language model alignment in multilingual trolley problems. In International Conference on Learning Representations (ICLR), 2025. URL https://openreview.net/forum?id=vrHErHkCNo. Miotto et al. [2022] MarilĂš Miotto, Nicola Rossberg, and Bennett Kleinberg. Who is GPT-3? An exploration of personality, values and demographics. In Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), pages 218â227, Abu Dhabi, UAE, November 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.nlpcss-1.24. URL https://aclanthology.org/2022.nlpcss-1.24/. Rutinowski et al. [2024] JĂŠrĂ´me Rutinowski, Sven Franke, Jan Endendyk, Ina Dormuth, Moritz Roidl, and Markus Pauly. The self-perception and political biases of ChatGPT. Human Behavior and Emerging Technologies, 2024:1â9, 2024. doi: 10.1155/2024/7115633. Yu et al. [2024] Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Tao Liu, and Deyi Xiong. CMoralEval: A moral evaluation benchmark for Chinese large language models. In Findings of the Association for Computational Linguistics: ACL 2024, pages 11817â11837, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.703. Jiao et al. [2025] Junfeng Jiao, Saleh Afroogh, Abhejay Murali, Kevin Chen, David Atkinson, and Amit Dhurandhar. Llm ethics benchmark: A three-dimensional assessment system for evaluating moral reasoning in large language models. Scientific Reports, 15:34642, 2025. doi: 10.1038/s41598-025-18489-7. Liu et al. [2024] Xuelin Liu, Yanfei Zhu, Shucheng Zhu, Pengyuan Liu, Ying Liu, and Dong Yu. Evaluating moral beliefs across LLMs through a pluralistic framework. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 4740â4760, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.272. Segerer [2025] Robin Segerer. Cultural value alignment in large language models: A prompt-based analysis of Schwartz values in Gemini, ChatGPT, and DeepSeek, 2025. Wang et al. [2024] Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9440â9450, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long.511. URL https://aclanthology.org/2024.acl-long.511/. Perez et al. [2023] Ethan Perez, Sam Ringer, KamilÄ LukoĹĄiĹŤtÄ, Karina Nguyen, et al. Discovering language model behaviors with model-written evaluations. In Findings of the Association for Computational Linguistics: ACL 2023, pages 13387â13434, 2023. Sharma et al. [2024] Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models. In International Conference on Learning Representations (ICLR), 2024. Wang et al. [2023] Yiwei Wang, Yujun Cai, Muhao Chen, Yuxuan Liang, and Bryan Hooi. Primacy effect of ChatGPT. In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 108â115, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.8. URL https://aclanthology.org/2023.emnlp-main.8/. Rupprecht et al. [2025] Jens Rupprecht, Georg Ahnert, and Markus Strohmaier. Prompt perturbations reveal human-like biases in large language model survey responses. arXiv preprint arXiv:2507.07188, 2025. Braun [2025] Daniel Braun. Acquiescence bias in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, 2025. Hagendorff et al. [2023] Thilo Hagendorff, Ishita Dasgupta, Marcel Binz, Stephanie C. Y. Chan, Andrew Lampinen, Jane X. Wang, Zeynep Akata, and Eric Schulz. Machine psychology. arXiv preprint arXiv:2303.13988, 2023. Binz and Schulz [2023] Marcel Binz and Eric Schulz. Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120, 2023. doi: 10.1073/pnas.2218523120. Coda-Forno et al. [2024] Julian Coda-Forno, Marcel Binz, Jane X. Wang, and Eric Schulz. Cogbench: a large language model walks into a psychology lab. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of PMLR, pages 9076â9108, 2024. Argyle et al. [2023] Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, and David Wingate. Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3):337â351, 2023. doi: 10.1017/pan.2023.2. Aher et al. [2023] Gati Aher, Rosa I. Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of PMLR, pages 337â371, 2023. Demszky et al. [2023] Dorottya Demszky, Diyi Yang, David S. Yeager, Christopher J. Bryan, Margarett Clapper, Susannah Chandhok, Johannes C. Eichstaedt, Cameron Hecht, Jeremy Jamieson, Meghann Johnson, Michaela Jones, Danielle Krettek-Cobb, Leslie Lai, Nirel JonesMitchell, Desmond C. Ong, Carol S. Dweck, James J. Gross, and James W. Pennebaker. Using large language models in psychology. Nature Reviews Psychology, 2(11):688â701, 2023. doi: 10.1038/s44159-023-00241-5. Shanahan [2024] Murray Shanahan. Talking about large language models. Communications of the ACM, 67(2):68â79, 2024. doi: 10.1145/3624724. Messeri and Crockett [2024] Lisa Messeri and M. J. Crockett. Artificial intelligence and illusions of understanding in scientific research. Nature, 627:49â58, 2024. doi: 10.1038/s41586-024-07146-0. Dillion et al. [2023] Danica Dillion, Niket Tandon, Yuling Gu, and Kurt Gray. Can AI language models replace human participants? Trends in Cognitive Sciences, 27(7):597â600, 2023. doi: 10.1016/j.tics.2023.04.008. Dillion et al. [2025] Danica Dillion, Debanjan Mondal, Niket Tandon, and Kurt Gray. AI language model rivals expert ethicist in perceived moral expertise. Scientific Reports, 15:4084, 2025. doi: 10.1038/s41598-025-86510-0. Aharoni et al. [2024] Eyal Aharoni, Sharlene Fernandes, Daniel J. Brady, Caelan Alexander, Michael Criner, Kara Queen, Javier Rando, Eddy Nahmias, and Victor Crespo. Attributions toward artificial agents in a modified Moral Turing Test. Scientific Reports, 14:8458, 2024. doi: 10.1038/s41598-024-58087-7. Awad et al. [2018] Edmond Awad, Sohan Dsouza, Richard Kim, Jonathan Schulz, Joseph Henrich, Azim Shariff, Jean-François Bonnefon, and Iyad Rahwan. The moral machine experiment. Nature, 563(7729):59â64, 2018. doi: 10.1038/s41586-018-0637-6. Pellert et al. [2024] Max Pellert, Clemens M. Lechner, Claudia Wagner, Beatrice Rammstedt, and Markus Strohmaier. AI psychometrics: Assessing the psychological profiles of large language models through psychometric inventories. Perspectives on Psychological Science, 19(5):808â826, 2024. doi: 10.1177/17456916231214460. Burnell et al. [2023] Ryan Burnell, Wout Schellaert, John Burden, Tomer D. Ullman, Fernando Martinez-Plumed, Joshua B. Tenenbaum, Danaja Rutar, Lucy G. Cheke, Jascha Sohl-Dickstein, Melanie Mitchell, Douwe Kiela, Murray Shanahan, Ellen M. Voorhees, Anthony G. Cohn, Joel Z. Leibo, and Jose Hernandez-Orallo. Rethink reporting of evaluation results in AI. Science, 380(6641):136â138, 2023. doi: 10.1126/science.adf6369. Raji et al. [2021] Inioluwa Deborah Raji, Emily M. Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna. AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1 (NeurIPS 2021), 2021. Vendrow et al. [2025] Joshua Vendrow, Edward Vendrow, Sara Beery, and Aleksander Madry. Do large language model benchmarks test reliability?, 2025. Bean et al. [2025] Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, MarĂa Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Rose Kirk, Fangru Lin, Gabrielle Kaili-May Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yushi Yang, Yilun Zhao, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip H. S. Torr, Cozmin Ududec, Luc Rocher, and Adam Mahdi. Measuring what matters: Construct validity in large language model benchmarks. In Advances in Neural Information Processing Systems 38 (NeurIPS 2025), Datasets and Benchmarks Track, 2025. Zhou et al. [2026] Hongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han, Huicheng Wang, Kehai Chen, Muyun Yang, Wei Bao, Jian Dong, Bing Xu, Conghui Zhu, Hailong Cao, and Tiejun Zhao. Lost in benchmarks? rethinking large language model benchmarking with item response theory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 35085â35093, 2026. Freiesleben [2026] Timo Freiesleben. Establishing construct validity in LLM capability benchmarks requires nomological networks, 2026. Lin [2025] Zhicheng Lin. Six fallacies in substituting large language models for human participants. Advances in Methods and Practices in Psychological Science, 8(3), 2025. doi: 10.1177/25152459251357566. Carron et al. [2024] Robin Carron, Emmanuelle Brigaud, Royce Anders, and Nathalie Blanc. Being blind (or not) to scenarios used in sacrificial dilemmas: the influence of factual and contextual information on moral responses. Frontiers in Psychology, 15:1477825, 2024. doi: 10.3389/fpsyg.2024.1477825. Cohen and Quinlan [2025] Dale J. Cohen and Philip T. Quinlan. Why moral judgements change across variations of trolley-like problems. British Journal of Psychology, 2025. doi: 10.1111/bjop.12782. Published online 18 February 2025; volume/issue/pages not yet assigned (re-checked via Crossref 2026-07-03). Christensen and Gomila [2012] Julia F. Christensen and Antoni Gomila. Moral dilemmas in cognitive neuroscience of moral decision-making: A principled review. Neuroscience & Biobehavioral Reviews, 36(4):1249â1264, 2012. doi: 10.1016/j.neubiorev.2012.02.008. Rehren and Sinnott-Armstrong [2023] Paul Rehren and Walter Sinnott-Armstrong. How stable are moral judgments? Review of Philosophy and Psychology, 14(4):1377â1403, 2023. doi: 10.1007/s13164-022-00649-7. Tassy et al. [2013] SĂŠbastien Tassy, Olivier Oullier, Julien Mancini, and Bruno Wicker. Discrepancies between judgment and choice of action in moral dilemmas. Frontiers in Psychology, 4:250, 2013. doi: 10.3389/fpsyg.2013.00250. FeldmanHall et al. [2012] Oriel FeldmanHall, Dean Mobbs, Davy Evans, Lucy Hiscox, Lauren Navrady, and Tim Dalgleish. What we say and what we do: The relationship between real and hypothetical moral choices. Cognition, 123(3):434â441, 2012. doi: 10.1016/j.cognition.2012.02.001. Francis et al. [2016] Kathryn B. Francis, Charles Howard, Ian S. Howard, Michaela Gummerum, Giorgio Ganis, Grace Anderson, and Sylvia Terbeck. Virtual morality: Transitioning from moral judgment to moral action? PLOS ONE, 11(10):e0164374, 2016. doi: 10.1371/journal.pone.0164374. Schwarz [1999] Norbert Schwarz. Self-reports: How the questions shape the answers. American Psychologist, 54(2):93â105, 1999. doi: 10.1037/0003-066x.54.2.93. Tourangeau et al. [2000] Roger Tourangeau, Lance J. Rips, and Kenneth A. Rasinski. The Psychology of Survey Response. Cambridge University Press, Cambridge, UK, 2000. ISBN 0521572460. Sudman et al. [1996] Seymour Sudman, Norman M. Bradburn, and Norbert Schwarz. Thinking about Answers: The Application of Cognitive Processes to Survey Methodology. Jossey-Bass, San Francisco, 1996. ISBN 0787901202. Greene et al. [2001] Joshua D. Greene, R. Brian Sommerville, Leigh E. Nystrom, John M. Darley, and Jonathan D. Cohen. An fmri investigation of emotional engagement in moral judgment. Science, 293(5537):2105â2108, 2001. doi: 10.1126/science.1062872. Greene and Haidt [2002] Joshua Greene and Jonathan Haidt. How (and where) does moral judgment work? Trends in Cognitive Sciences, 6(12):517â523, 2002. doi: 10.1016/S1364-6613(02)02011-9. Cushman [2013] Fiery Cushman. Action, outcome, and value: A dual-system framework for morality. Personality and Social Psychology Review, 17(3):273â292, 2013. doi: 10.1177/1088868313495594. Bago and De Neys [2019] Bence Bago and Wim De Neys. The intuitive greater good: Testing the corrective dual process model of moral cognition. Journal of Experimental Psychology: General, 148(10):1782â1801, 2019. doi: 10.1037/xge0000533. Kahane et al. [2018] Guy Kahane, Jim A. C. Everett, Brian D. Earp, Lucius Caviola, Nadira S. Faber, Molly J. Crockett, and Julian Savulescu. Beyond sacrificial harm: A two-dimensional model of utilitarian psychology. Psychological Review, 125(2):131â164, 2018. doi: 10.1037/rev0000093. Schwarz et al. [1991] Norbert Schwarz, Bärbel Knäuper, Hans-J. Hippler, Elisabeth Noelle-Neumann, and Leslie Clark. Rating scales: Numeric values may change the meaning of scale labels. Public Opinion Quarterly, 55(4):570â582, 1991. doi: 10.1086/269282. HĂśhne et al. [2021] Jan Karem HĂśhne, Dagmar Krebs, and Steffen-M. KĂźhnel. Measurement properties of completely and end labeled unipolar and bipolar scales in Likert-type questions on income (in)equality. Social Science Research, 97:102544, 2021. doi: 10.1016/j.ssresearch.2021.102544. HĂśhne et al. [2022] Jan Karem HĂśhne, Dagmar Krebs, and Steffen-M. KĂźhnel. Measuring income (in)equality: Comparing survey questions with unipolar and bipolar scales in a probability-based online panel. Social Science Computer Review, 40(1):108â123, 2022. doi: 10.1177/0894439320902461. Burdick et al. [2003] Richard K. Burdick, Connie M. Borror, and Douglas C. Montgomery. A review of methods for measurement systems capability analysis. Journal of Quality Technology, 35(4):342â354, 2003. doi: 10.1080/00224065.2003.11980232. Cronbach et al. [1963] Lee J. Cronbach, Nageswari Rajaratnam, and Goldine C. Gleser. Theory of generalizability: A liberalization of reliability theory. British Journal of Statistical Psychology, 16(2):137â163, 1963. doi: 10.1111/j.2044-8317.1963.tb00206.x. Cronbach et al. [1972] Lee J. Cronbach, Goldine C. Gleser, Harinder Nanda, and Nageswari Rajaratnam. The Dependability of Behavioral Measurements: Theory of Generalizability for Scores and Profiles. Wiley, New York, 1972. Shavelson and Webb [1991] Richard J. Shavelson and Noreen M. Webb. Generalizability Theory: A Primer. Sage Publications, Newbury Park, CA, 1991. Brennan [2001] Robert L. Brennan. Generalizability Theory. Statistics for Social and Behavioral Sciences. Springer, New York, NY, 2001. doi: 10.1007/978-1-4757-3456-0.