Paper deep dive
Benchmarking LLM Competence on Logical Inference over Probability Operators
Nayera Hasan, Jack Greff, Alvin Grissom
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/1/2026, 1:32:24 AM
Summary
This paper introduces a benchmark for evaluating Large Language Models (LLMs) on logical inference over probability operators (e.g., probably, might, must). The benchmark consists of 14,320 procedurally generated prompts across 15 inference templates, varying question forms, negation strategies, and surface content. Evaluating 29 models, the authors find that most exhibit answer biases (preference for Yes/No) independent of logical form, with only 9 models exceeding random chance. The study highlights significant performance gaps due to negation strategies and surface-level token biases, suggesting LLMs rely on pattern matching rather than principled symbolic reasoning.
Entities (10)
Relation Signals (10)
Benchmark for Reasoning over Probability Operators → contains → 14,320 procedurally-generated English prompts
confidence 98% · We introduce a benchmark for reasoning over probability operators... containing 14,320 procedurally-generated English prompts
Benchmark for Reasoning over Probability Operators → evaluates → LLMs
confidence 95% · Evaluating 29 models, we find that most show answer biases independent of the logical form
LLMs → exhibits → answer biases
confidence 92% · most show answer biases independent of the logical form, a systematic preference for Yes or No.
Competence Floor → measures → LLM competence
confidence 90% · We summarize this with a competence floor: the worse of a model’s accuracy on Yes-correct and No-correct items.
LLMs → relieson → surface-level pattern matching
confidence 90% · disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty.
Yalcin → authored → Theoretical work on probability operators
confidence 85% · In our work, we build upon theoretical work that studies inference rules over natural-language sentences containing probability operators Yalcin (2010)
GPT-4 → exhibits → logically inconsistent judgments
confidence 85% · even the strongest model in their study (GPT-4, closed source) exhibits logically inconsistent judgments across related patterns
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
Tags
Links
- Source: https://arxiv.org/abs/2607.27405v1
- Canonical: https://arxiv.org/abs/2607.27405v1
Trouble viewing inline? Open PDF directly →
Full Text
86,631 characters extracted from source content.
Expand or collapse full text
BENCHMARKING LLM COMPETENCE ON LOGICAL INFERENCE OVER PROBABILITY OPERATORS A PREPRINT Nayera Hasan Haverford College Jack Greff Haverford College Alvin Grissom I Haverford College ABSTRACT Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators—inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model’s accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis. 1 Introduction Epistemic modals (EMs) express a state of knowledge, uncertainty, credence, or belief, and such modals are ubiquitous in human language (Teller, 1972). In English, epistemic modality (EM) is facilitated largely by modal verbs such as must, might, may, and adverbs such as probably, and it pervades both everyday and specialist language. The sentences Bob might be home early and Bob will possibly be home early, for example, express that there is some chance, not necessarily a high chance, that Bob will arrive home early, while the epistemic mood of the sentence Bob must be home early expresses certainty that he will be. While modal logics address the subset of cases expressing necessity and possibility—e.g., must and might—gradable epistemic modality also includes cases expressing degrees of certainty, known as probability operators (Yalcin, 2010). In our work, we build upon theoretical work that studies inference rules over natural-language sentences containing probability operators Yalcin (2010) to benchmark and analyze the capabilities of LLMs in correctly performing zero- shot logical inference with a large set of procedurally-generated EM sentences. EM has received little attention in modern computational linguistics, though earlier work artificial intelligence explored modal and epistemic logics for the integration of uncertain information Moore (1981), which is a fundamental problem of AI, broadly defined. While there exists philosophy and linguistics literature on the semantics of epistemic modals and inference over them (e.g., Yalcin (2010); Kratzer (2012), there is less work on their analysis in NLP. Holliday, Mandelkern, and Zhang (2024) take an important step in this direction, probing LLMs on modal reasoning with might and must, representable by modal logic. They find basic errors and logically inconsistent judgments across related inference types, echoing concerns in other areas about the plausible-but-invalid inferences to which LLMs are prone. This echoes similar work which finds that LLMs’ accuracy on syllogisms depends heavily on recognizing specific token patterns, with changes to names, entities, or quantifiers causing predictable performance shifts Jiang et al. (2024), which we also examine. EM reasoning competence is particularly relevant given that EM inference is central to nearly all high-stakes, natural-language scenarios in which LLMs—in these contexts, often colloquially called AI—are likely to be deployed: a clinical decision-support system reading a note that a treatment “will probably work” or a person “probably arXiv:2607.27405v1 [cs.CL] 29 Jul 2026 arXiv TemplateA PREPRINT has a disease” must synthesize premises under uncertainty before ranking options; a system triaging legal documents must distinguish between “might have known” and “probably knew”, since the two carry different evidentiary weight; and a financial pipeline that treats “the market will probably recover” as equivalent to “the market will certainly recover” mischaracterizes risk. As LLMs are embedded in decision loops in medical, legal, and financial domains, their ability to correctly handle inferences with probably, might, and must becomes a practical concern, not only a theoretical one. Our benchmark covers fifteen templates, encoding thirteen distinct inference patterns (10 valid, 3 invalid) with 14,320 prompts. We vary the wording for logically equivalent inference patterns to check whether the model is changing its inference based on tokens instead of the underlying logic: we use five question forms (Is it correct that...?, Is it true that...?, Does it follow that...?, etc.), multiple negation strategies (prefix, not, proposition-level), and surface content (names from seven nationality groups, five activity descriptions). A genuine reasoner will produce the same answer across all of these, irrespective of lexical variation. We evaluate 29 models; most show a fixed bias toward one answer, a systematic preference for Yes or No that holds across inference types rather than tracking what each inference licenses. Our primary contributions are: (1) a dataset and benchmark, tested across a variety of models, for inference over probability operators that varies question form, polarity, and content (verb and noun phrases) for fixed logical forms, disentangling models’ answer bias and pattern-matching from principled reasoning (§4); (2) strong evidence that models answer from a fixed answer bias rather than the logical inference (only 9/29 models exceed a random baseline), and that this finding is negation-independent (§4.1); and (3) a comparison of negated sentences and non-negated sentences, revealing accuracy gaps of up to 64 percentage points on semantically identical questions, identifying negation strategy as a confound in benchmark design (§5.1.2). 2 Related Work To provide theoretical context for the task and situate our work within prior literature on evaluating LLM reasoning, we review related work. 2.1 Modal and Probabilistic Reasoning in LLMs Prior work on the semantics of epistemic modals provides the theoretical grounding for the inference templates on which we evaluate models. Kratzer (1991, 2012) establishes the standard quantificational framework; Lassiter (2011, 2017) proposes the probabilistic alternative in which probablyφentails thatp(φ)is high, mightφthatp(φ)is non-trivial, and mustφthatp(φ)is near-maximal; and Yalcin (2010)’s work on probability operators catalogues valid and invalid inference patterns involving epistemic comparatives. Holliday, Mandelkern, and Zhang (2024) test 29 LLMs on 20 inference patterns using conditionals and the modal auxiliaries might and must. They find that all models commit basic fallacies and that even the strongest model in their study (GPT-4, closed source) exhibits logically inconsistent judgments across related patterns (e.g., accepting Modus Tollens with must but rejecting it with might, which is logically contradictory). Their templates focus on modal auxiliaries (might, must) with conditionals, while ours, following semantics work by Yalcin (2010), focuses on the adverb probably, examining LLM inference patterns specific to graded epistemic reasoning (e.g., distribution over conjunction, conditional-to-comparative, and conjunction fallacy). Li, Vrazitulis, and Schlangen (2025) furthermore provide LLMs with short narratives containing factual information and queries the LLMs with prompts containing modal auxiliaries may/might vs. must/have to as well as with attitude verbs (know, believe, doubt). They note that accuracy is affected by prompt format. Our work is more reductive in its focus on simple templates rather than complex narratives. Imannezhad, Pothos, and Wills (2026) examine GPT-5’s probabilistic reasoning on conjunction and disjunction fallacies and binary complementarity violations using matched human participant profiles, arguing that GPT-5 appears to have a consistent internal probability model (which our results contravene) that align with quantum-probabilistic models and exceed that of humans. They elicit numeric probability ratings under a single fixed prompt format and find the judgments internally coherent. In contrast, we hold the inference fixed and vary the question format, finding that binary judgments shift according to the formulation. Our panel is predominantly open-weight for replicability, with a small number of frontier closed models included as reference points; we do not evaluate the commercial GPT-5 model examined by Imannezhad et al., the closest model in our panel being GPT-5.4-mini. Additionally, rather than compare with human judgments, our work focuses on the inherent correctness of LLM inferences. Macmillan-Scott and Musolesi (2024) tests seven LLMs on cognitive bias tasks from the psychology literature, finding that LLMs are irrational, but that their irrationality does not reflect human patterns: when models give incorrect answers, 2 arXiv TemplateA PREPRINT they are incorrect in different ways than humans. They also found significant inconsistency in responses, consistent with our finding that much of what looks like reasoning on these tasks is surface-driven answer behavior rather than proposition-tracking. 2.2 Surface Sensitivity and Token Bias LLM performance shifts with surface-form changes that should not affect the answer. Jiang et al. (2024) describe this as “token bias” and introduce a hypothesis-testing framework using McNemar’s test (McNemar, 1947) on matched pairs. They show that on conjunction fallacies and syllogistic problems, changing names, entities, or quantifiers while holding the logic constant produces predictable performance shifts, indicating reliance on superficial patterns rather than reasoning. Similarly, Binz and Schulz (2023) find that GPT-3 (which, being closed source, is no longer available) solved canonical cognitive psychology vignettes “similarly or better than human subjects” but that “small perturbations to vignette-based tasks can lead GPT-3 vastly astray.” Other work highlights that even meaning-preserving formatting changes (separator characters, label styles) produce accuracy swings of up to 76 percentage points (Sclar et al., 2024), echoing similar issues documented in machine translation (Shi, Grissom I, and Trinh, 2022). Furthermore, LLMs have a bias toward accepting false presuppositions, making them prone to inferences based on implicit misinformation (Er- makova, Firsov, and Kamps, 2026). Sharma et al. (2024) directly challenges claims of LLM and large reasoning model “reasoning”, demonstrating that the former outperform the latter at low problem complexity but that both collapse at higher problem complexities. In terms of perturbations, our work—which focuses on a particular, but fundamental, corner of syllogistic reasoning— differs from prior work in several respects. First, perturbations in prior work change what the problem talks about (names, entities, quantifiers) or how the text looks on the page (separators, label styles). We also do this but also change how the question is asked, swapping the metalinguistic wrapper (Is it true that..., Does it follow that...) and the way a question’s negation is expressed, a dimension of surface sensitivity closer to the logical structure itself. We also decompose accuracy into orthogonal components (the answer bias and polarity sensitivity), going beyond binary detection of sensitivity to measurement of how much each factor explains. 2.3 Response Bias and Acquiescence Tjuatja et al. (2024) investigate whether LLMs exhibit human-like response biases (acquiescence, response order, opinion floating). Testing nine models on 2,578 question pairs, they find that LLMs generally do not reflect human-like bias patterns and that RLHF reduces sensitivity to bias-inducing modifications while increasing sensitivity to non-bias perturbations. Braun (2025) extended this to 37,975 question variations phrased survey-style (“do you agree...”, “don’t you agree...”) over classification tasks on legal-domain texts in English, German, and Polish, finding that LLMs display a bias toward answering No in English, the opposite of human acquiescence bias. This aligns with our results The direction of the answer bias (yes vs. no) is model-dependent rather than universal, with some models strongly yes-biased and others strongly no-biased. Our baseline takes the minimum of the two answer-conditioned accuracies rather than their mean because the mean alone obscures the bias: mean accuracy can be high when one answer class is near ceiling and the other near floor. The benchmark’s validity-by-negation structure is what makes this minimum informative, providing an absolute baseline under which a maximally biased LLM scores 0 regardless of the valid/invalid mix. We defer the full construction to the discussion of accuracy metrics (§4.1). A third evaluation framework treats answer-label priors as a calibration problem. Zhao et al. (2021) show that few-shot classification with language models is skewed by majority-label, recency, and common-token biases, estimated the model’s prior from a content-free input; Zhao et al. (2021) divide it out at inference, with large accuracy gains; Fei et al. (2023) refined the prior estimate with in-domain words, and Zhou et al. (2024) with the mean prediction over a test batch. Holtzman et al. (2021) traced related distortions to probability mass splitting across surface forms of the same answer. These results were established on models from 2021 to 2024; our panel suggests the phenomenon has not aged out, with current instruction-tuned and reasoning models still carrying extreme answer priors. Our answer bias is a label bias on a Yes/No space, and this literature suggests such priors are often removable by calibration. Where that work treats the prior as a nuisance to be removed so that accuracy improves, we instead report it directly through the competence floor, which remains diagnostic whether or not the prior could be calibrated away. 2.4 Negation Processing Truong et al. (2023) evaluate GPT-neo, GPT-3, and InstructGPT on six negation benchmarks, finding three key limitations: insensitivity to the presence of negation, inability to capture its lexical semantics, and failure to reason 3 arXiv TemplateA PREPRINT under negation. They also find that negation shows flat or inverse scaling (larger models are more insensitive), a pattern alleviated only by instruction fine-tuning. García-Ferrero et al. (2023) introduced a large negation benchmark of approximately 400,000 sentences, confirming that LLMs rely on superficial cues and that fine-tuning improves but does not fully resolve the problem. So et al. (2025) establish a typology distinguishing morphological negation (un-, in-) from syntactic negation (not), which directly corresponds to the contrast we test. Elkins and Chun (2026) audit negation sensitivity in ethical decisions, finding that open-source models endorse prohibited actions 77% of the time under simple negation and 100% under compound negation. Our benchmark contains prompts with no negation at all, some expecting Yes and some expecting No, so the core contrast does not depend on negation. On top of that, we implement negation in several ways, asking whether a model that accepts a valid inference (or rejects an invalid one) in the affirmative form can still do so when the question is negated. Implementing it became a source of findings in its own right: writing the same negation as a prefix or with the word not produces dramatically different accuracy on semantically identical content (§5.1.2). 3 Benchmark Design In this section, we describe the overall design of our benchmarks, including the inference templates and their logical forms. 3.1 Inference Templates Our work is theoretically motivated by prior work on EM semantics which characterizes valid inferences. In particular, we create prompts based on probability operators analyzed by Yalcin (2010), crafted as binary questions with a known correct response. Thirteen inference templates compose the benchmark of EM expressions; for exposition, we group the templates into three groups based on validity and complexity. Eleven inference rules have a single template; two (distribution over conjunction and conjunctivitis; §5.2) each have two surface-form templates. Altogether, we have fifteen. Table 1 shows each pattern with a representative prompt from the dataset with the logical form described in terms of probability and modal operators for necessity (□) and possibility (♢). We use⪰to denote “is at least as likely as“, i.e.,A ⪰ B ⇐⇒ p(A) ≥ p(B). For brevity, we also usePr(A)to denote “probably A” andp(A)for a numerical probability, i.e., Pr(A) ⇐⇒ p(A)≥ 0.5 under typical probability assumptions. 3.1.1 Template Descriptions Simple valid inferences (5 templates) These test straightforward relationships between epistemic operators. Probably to Might: if φ is probable, then φ is possible. Must to Probably: if φ is necessary, then φ is probable. Probably to Not Probably Not: if φ is probable, then it is not probable that¬φ. Positive Form Transfer (PFT) and Complement Transfer (CT) test probability preservation under reformula- tion. Positive Form Transfer ψ is at least as likely as φ probably φ. probably ψ ∴ Complement Transfer ψ is at least as likely as φ φ is at least as likely as¬φ. ψ is at least as likely as¬φ ∴ Multi-step valid inferences (5 templates) Chancy Modus Ponens and its contrapositive variant Chancy Modus Tollens require combining premises or comparing probabilities. 4 arXiv TemplateA PREPRINT Chancy Modus Ponens if φ then ψ probably φ. probably ψ ∴ Chancy Modus Tollens if φ then ψ ¬ probably ψ. ¬ probably φ ∴ Distribution over conjunction: if probably (φ∧ ψ), then probably φ and probably ψ. Conditional to Comparative: if probably (φ→ ψ), then ψ is at least as probable as φ. Chancy Disjunction Introduction: if probably φ, then probably (φ∨ ψ). Invalid inferences (3 templates) These measure correct rejection of fallacious reasoning. Conjunctivitis Kyburg Jr (1970) is a fascinating conjunction fallacy discussed in the philosophical literature. The typical logical assumption of closure under conjunction does not hold for probability operators. Closure under conjunction holds when the truth ofψandφindependently implies the truth ofψ∧ φ. It holds for non EM predicates but not for the probability operator. I.e.,probably(ψ), probably(φ) ̸=⇒ probably(ψ∧ φ), a fact that can be illustrated by the lottery paradox, whereby we have a collection of propositions of probability at least0.5but less than1, whose conjunction probability necessarily decreases with each new proposition. Since this is a known fallacy discussed in the literature, one might expect that LLMs would be less prone to fall for it. Might to Probably inference: possibility does not entail probability. ♢φ ̸=⇒ □φ. Probably to Certain: probability does not entail certainty. probably φ ̸=⇒ □φ. 3.2 Question Form We present questions in both affirmative and negated forms to test consistency under negation. In the affirmative condition, we ask whether the conclusion holds (correct answer: Yes for valid; No for invalid). In the negated condition, we ask whether the same conclusion does not hold (correct answer: No for valid, Yes for invalid). Five metalinguistic question forms vary how the question is asked. CORRECT, TRUTH, and VALID negate by switching to the negative adjective (correct to incorrect, true to false, valid to invalid). FOLLOW inserts the word not changing, for example, does it follow to does it not follow). DIRECT negates the proposition itself. An additional pair of question forms, Is it not correct that . . . ? and Is it not valid to conclude that . . . ?, replaces the negative adjective with the word not on the same adjective, enabling direct comparison of negation strategies on identical semantic content (§5.1.2). In our analysis, we discuss the perlocutionary force of such questions in light of known issues of LLM sycophancy (§5.1.2). Specifically, for our question templates, introducing negation can pragmatically encourage the model to agree with the questioner. Asking, Is it not probable thatφ? can, under a common interpretation, sound like the answershouldbe the Yes. We introduce clauses such as Is it true that for affirmative and Is it false that for the negation to enforce some neutral symmetry in the question surface form without such pragmatic force skewing the results. 3.3 Demographic Bias Controls Natural language models are known to be prone to spurious biases. To control for this, names are drawn from seven nationality groups (Indian, Russian, Japanese, African, German, French, American) plus abstract letter variables, balanced by gender (common names for men or women). Activities include five semantically similar scenarios. These function as robustness checks: a model that reasons about the logic should show minimal accuracy variation across these controls. 3.4 Models Evaluated We evaluate 29 models in English spanning five size-graded families (Gemma 3: 270M–27B; Gemma 4: E2B–31B; Qwen 3: 0.6B–32B; DeepSeek-R1: 7B–32B; Llama 3: 3B, 8B) plus frontier models (Claude Opus 4.7, Claude 5 arXiv TemplateA PREPRINT Table 1: Inference templates with formal schema and example prompts (TRUTH question form, affirmative; the negated condition replaces Is it true that with Is it false that and flips the correct answer). In the affirmative examples shown the correct answer is Yes for the valid templates and No for the invalid ones, and negation reverses both.Pr(φ): probable; M (φ): possible;□φ: certain; φ≽ ψ: at least as likely. Names and activities vary. TemplateValiditySchemaExample prompt Probably to MightValidPr(φ)⇒ M (φ)Savir is probably going to be at the party. Is it true that Savir might be at the party? Must to ProbablyValid□φ⇒ Pr(φ)Savir must be at the party. Is it true that it is probable that Savir will be at the party? Chancy Modus Po- nens Valid[φ ⇒ ψ,Pr(φ)] ⇒ Pr(ψ) If Savir is going to be at the party, then Ashwin is going to be at the party. Savir is probably going to be at the party. Is it true that Ashwin is probably going to be at the party? Chancy Modus Tol- lens ValidPr(φ→ψ), Pr(¬ψ) ⇒ Pr(¬φ) If Savir is going to be at the party, then Ashwin is going to be at the party. Ashwin is probably not going to be at the party. Is it true that Savir is probably not going to be at the party? Distribution over Conjunction ValidPr(φ∧ ψ)⇒ Pr(φ)It is probable that Savir and Ashwin will be at the party. Is it true that it is probable that Savir will be at the party? Conditionalto Comparative ValidPr(φ→ψ)⇒ ψ ≽ φIf Savir is going to be at the party, then Ashwin is going to be at the party. Is it true that Ashwin is at least as likely as Savir to be at the party? ChancyDisjunc- tion Introduction ValidPr(φ)⇒ Pr(φ∨ ψ)Savir is probably going to attend the event. Is it true that it is probable that Savir will attend the event or give a presentation? PositiveForm Transfer Validφ ≽ ψ, Pr(ψ) ⇒ Pr(φ) Ashwin is at least as likely as Savir to be at the party. Savir is probably going to be at the party. Is it true that Ashwin is probably going to be at the party? Probably to not probably not ValidPr(φ)⇒¬Pr(¬φ)Savir is probably going to be at the party. Is it true that it is not probable that Savir will not be at the party? Complement Transfer Validφ ≽ ψ, ψ ≽ ¬ψ ⇒ φ≽¬φ Ashwin is at least as likely as Savir to be at the party. Savir is at least as likely to be at the party as to not be at the party. Is it true that Ashwin is at least as likely to be at the party as to not be at the party? ConjunctivitisInvalidPr(φ),Pr(ψ) ̸=⇒ Pr(φ∧ ψ) It is probable that Savir will be at the party. It is probable that Ashwin will be at the party. Is it true that it is probable that Savir and Ashwin will be at the party? Might to ProbablyInvalidM (φ) ̸=⇒ Pr(φ)Savir might be at the party. Is it true that it is probable that Savir will be at the party? Probably to CertainInvalidPr(φ) ̸=⇒ □φSavir is probably going to be at the party. Is it true that it is certain that Savir will be at the party? Sonnet 4.6, GPT-oss 120B, GPT-oss 20, GPT-5.4-mini), Granite 3.2:8B, Mistral:7b, and the Arabic-oriented Aya- Expanse 8B and 32B. We show all models the same 80 surface prompts per template, and we run each model once per prompt at temperature τ = 0 All main analyses pool across templates and question forms. Prompting setupWe query open-weight models zero-shot, and the first answer token is parsed as Yes, No, or classified as Uncertain otherwise. Two choices bear on interpretation. First, to keep the evaluation strictly zero-shot, we disable chain-of-thought wherever the model permits it: Qwen 3 and DeepSeek-R1 are run with thinking turned off, and GPT-oss is set to its lowest available reasoning level. The reported numbers therefore reflect direct-answer behavior. Second, Qwen 3:4B is queried with a constrained Yes/No output format, because unconstrained generation produced a 6 arXiv TemplateA PREPRINT Table 2: The five English question forms and their negated versions, plus the two not-prefixed forms, shown on Must to Probably (premise: “Savir must be at the party”). The repeated conclusion “it is probable that Savir will be at the party” is abbreviated as . . . . Correct answers flip with polarity: Yes affirmative, No negated. Question formAffirmative (Yes)Negated (No) TRUTHIs it true that . . . ?Is it false that . . . ? CORRECTIs it correct that . . . ?Is it incorrect that . . . ? VALIDIs it valid to conclude that . . . ?Is it invalid to conclude that . . . ? FOLLOWDoes it follow that . . . ?Does it not follow that . . . ? DIRECTIs it probable that Savir will be at the party?Is it not probable that . . . NOT-CORRECT—Is it not correct that . . . ? NOT-VALID—Is it not valid to conclude that . . . ? DIRECT carries no metalinguistic wrapper, negating the proposition itself (probable to not probable). non-compliance rate as high as 96%. Constraining the output still lets the model choose either answer, so it does not manufacture the answer bias we report. 3.5 Metrics We score each model on a 2×2 grid, crossing inference validity (valid/invalid) with question negation (affirma- tive/negated); we refer to its four cells throughout. This crossing determines the correct answer in each cell. We use abbreviations to refer to the accuracy for each intersection of conditions. V.AFF, for example, refers to a valid template with a question phrased in the affirmative, while I.NEG refers to an invalid template with a negated question. The correct answer is Yes in two cells (V.AFF and I.NEG) and No in the other two (V.NEG and I.AFF). Overall accuracy is the fraction of correct responses (with uncertain counted as incorrect), computed over whatever set of prompts a table covers. Since valid templates outnumber invalid ones, weighting the four validity×negation cells equally instead pulls every strong model toward 0.5, by up to about 9 points; supplementary Table 14 reports accuracy under five weighting schemes so the sensitivity to this choice is visible. To separate the model’s surface-level bias from principled proposition-tracking, we measure the yes/no bias. We do this by averaging the model accuracy, respectively, on questions for which the answers should be affirmative (Yes) and on questions for which the answers should be negative (No). For each, we have two cases, the valid and invalid conditions. 1 Let Acc(Yes) = 1 2 (V.AFF + I.NEG) and Acc(No) = 1 2 (V.NEG + I.AFF). The signed gap is the bias toward Yes or Nos: Bias = Acc(Yes)− Acc(No),(1) If a model always guesses Yes, the bias will be 1; if it always guesses No, it will be -1; and a random baseline has a bias of 0 2 . We also track polarity sensitivity (PS), the model’s accuracy drop under negation: PS = (V.AFF + I.AFF)− (V.NEG + I.NEG) 2 (2) = [Acc(non-negated questions)− Acc(negated questions)] 2 (3) Bias is positive when the model favors Yes; PS is positive when negation hurts. Bias is an indirect way of measuring a model’s competence, i.e., a model’s actual reasoning ability. To measure competence, we take the worse of the two accuracy measurements, the floor: Floor = min Acc(Yes), Acc(No) .(4) Intuitively, this related metric concisely describes how poorly the model performs on its worst-performing question polarity (those with a correct answer of either Yes or No). This is motivated by our observation that models tend to be extremely biased for one answer or the other, irrespective of the internal logic. A constant responder scores 0 on the floor, a coin-flipper 0.5, and a competent reasoner near 1, so the floor has an absolute baseline that overall accuracy lacks. 1 These are true positive rates, but we avoid using the term since “’positive” is confusing when discussing yes/no questions that can be negated. 2 We have more valid templates than invalid ones; averaging them prevents in this way prevents the valid category from dominating the bias and associated metrics. However, the average still provides useful information. 7 arXiv TemplateA PREPRINT Table 3: English models, the four cells, sorted by floor; the faded rule separates the nine models that clear the 0.5 floor from those that do not. The four columns V.AFF, V.NEG, I.AFF, I.NEGare per-cell accuracies, so e.g. V.AFF=Acc(V.Aff). Acc is overall accuracy. Uncertain responses count as incorrect. Bias= Acc(Yes)− Acc(No) (signed; positive=yes-biased). Floor= min(Acc(Yes), Acc(No))and has a random baseline of 0.5; floors above random chance are in bold. Per-condition accuracies below chance are underlined. ModelAccUncV.AFFV.NEGI.AFFI.NEGBiasPSFloor gemma4:31b0.8180.0%0.9770.7610.9800.372−0.20+0.410.675 Sonnet 4.60.7510.0%0.9940.6360.6480.491+0.10+0.260.642 gemma4:12b0.7640.0%0.9670.7540.7350.261−0.13+0.340.614 qwen3:32b0.7180.0%0.8130.7080.8270.374−0.17+0.280.594 gemma3:27b0.6481.0%0.9130.470 0.7120.334+0.03+0.410.591 gpt-oss:200.7370.0%0.8150.7400.8940.356−0.23+0.310.586 Opus 4.70.7560.0%0.8920.7620.8700.248−0.25+0.380.570 gpt-oss:120b0.7240.0%0.7270.8080.8890.326−0.32+0.240.526 gemma3:12b0.6150.0%0.8120.6200.4490.224−0.02+0.210.518 gemma4:e4b0.6080.0%0.6520.6550.6740.294−0.19+0.190.473 llama3.1:8b0.46314.3%0.445 0.4350.5040.544+0.03 −0.010.469 gpt-5.4-mini0.6560.0%0.6710.7100.8870.239−0.34+0.300.455 aya-expanse:32b0.6240.1%0.9050.5220.311 0.440+0.26+0.130.416 gemma3:270m0.5000.0%0.6400.3690.4560.512+0.16+0.110.413 gemma3:4b0.5290.0%0.6030.5240.2830.589+0.19 −0.110.403 qwen3:8b0.5990.0%0.9250.4100.3860.424+0.28+0.240.398 qwen3:14b0.6050.0%0.472 0.8130.7290.281−0.39+0.050.377 aya-expanse:8b0.5860.6%0.8590.4670.2240.517+0.34+0.050.346 mistral:7b0.5350.0%0.8360.2870.3100.604+0.42+0.130.298 granite3.2:8b0.6070.0%0.4060.8630.9090.169−0.60+0.140.287 qwen3:4b0.5690.0%0.4670.6990.9860.081−0.57+0.340.274 llama3.2:3b0.4671.5%0.234 0.6050.9310.275−0.51+0.140.254 gemma4:e2b0.5901.7%0.450 0.7970.9930.016−0.66+0.320.233 deepseek-r1:32b0.5562.4%0.9530.087 0.3490.942+0.73+0.140.218 qwen3:1.7b0.5171.5%0.8350.2860.1380.649+0.53+0.020.212 gemma3:1b0.5370.2%0.9670.1660.1010.790+0.74+0.060.133 qwen3:0.6b0.5480.0%0.9970.1600.0300.878+0.84 −0.010.095 deepseek-r1:14b0.47710.6%0.9050.014 0.1340.889+0.82+0.070.074 deepseek-r1:7b0.20261.5%0.4150.0020.0000.357+0.39+0.030.001 4 Results We now describe the EM benchmarking results. While some of the differences in accuracy across conditions may seem small (but many are not), given the query rate of LLMs, seemingly small differences in the number of errors can have huge impacts in the aggregate. If, for example, an LLM is queried 1 billion times, a0.3%increase in errors is equivalent to3million more. Currently, ChatGPT alone receives approximately 2.5 billion queries per day, and GPT-4o received approximately 774 billion queries in 2025, which is still only a fraction of Google’s queries, which now integrate LLMs Singh (2026). 4.1 Overall Accuracy First, we report the aggregate accuracy across all EM prompts for a bird’s eye view of respective model performance as a prelude to more specific analysis. Table 3 presents the four cells for all 29 models, using the metrics defined in §3.5. Overall accuracy ranges from 20.2% (DeepSeek-R1:7B) to 81.8% (Gemma4:31B), but this obscures the strategies (such as they are) being used. Qwen3:0.6B scores 54.8% overall by nearly always responding Yes (V.AFF= 0.997, V.NEG= 0.160, I.AFF= 0.030, I.NEG= 0.878), while Llama3.1:8B is balanced (bias only+0.03), yet every cell performs near the random baseline, suggesting that the predictions are themselves arbitrary and disconnected from the prompt. Uncertain responses are rare for most models, below 2%, but high for a handful (e.g., DeepSeek-R1:7b is uncertain 61.5% of the time). Uncertain rates are relatively flat across conditions (expected-Yes 3.2% vs. expected-No 3.4%), so counting them as incorrect does not unduly affect the competence floor. We report rates in Table 3. 8 arXiv TemplateA PREPRINT 0.00.10.20.30.40.50.60.70.80.9 Bias = |Acc(Yes) Acc(No)| 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Floor = min(Acc(Yes), Acc(No)) Floor = 0.5 (random baseline) r = 0.82 R² = 0.68 gemma4:31b sonnet gemma4:12b qwen3:32b gemma3:27b gpt-oss:20b opus gpt-oss:120b gemma3:12b gemma4:e4b llama3.1:8b gpt-5.4-mini aya-expanse:32b gemma3:270m gemma3:4b qwen3:8b qwen3:14b aya-expanse:8b mistral:latest granite3.2:8b qwen3:4b llama3.2:3b gemma4:e2b deepseek-r1:32b qwen3:1.7b gemma3:1b qwen3:0.6b deepseek-r1:14b deepseek-r1:7b Open-weight Closed-source Figure 1: Competence Floor against answer bias (|Bias|) for all 29 models. The two are strongly inversely correlated (r =−0.82,R 2 = 0.68): the more a model is biased toward one answer, the lower its floor. Only nine models clear the 0.5 random baseline (dashed), and the most biased models (e.g. Qwen3:0.6B) sit near the floor’s zero. Few models clear the floor. Only 9 of 29 models exceed the random Floor baseline of 0.5: Gemma4:31B (0.675), Sonnet 4.6 (0.642), Gemma4:12B (0.614), Qwen3:32B (0.594), Gemma3:27B (0.591), GPT-oss:20 (0.586), Opus 4.7 (0.570), GPT-oss:120B (0.526), and Gemma3:12B (0.518). Even these nine clear it only modestly, by about 2 to 18 points above the 0.5 baseline. Eleven of the 29 models fall below a floor of 0.30. Qwen3:0.6B is the textbook constant Yes responder: near-ceiling where Yes is correct (V.AFF 0.997, I.NEG 0.878), near-zero where No is (V.NEG 0.160, I.AFF 0.030), with 54.8% overall accuracy the average of these extremes. Competence differences are not only due to negationIt is instructive to contrast the competence floor in affirmative vs. negated questions. Recomputing the competence floor from the affirmative cells withmin(V.AFF, I.AFF)achieves this. This affirmative-only floor correlatesr = 0.90with the full floor, still only yielding 10 of 29 models above chance, and still leaving the biased models with the worst measures of competence (Qwen3:0.6B at 0.030). Importantly, the answer bias is visible in the affirmative validity contrast alone, so negation is one way to expose it, not necessarily its cause. 5 Analysis 5.1 Robustness to Surface Variation A genuine reasoner would return the same answer when only the surface of a prompt changes and the logic is held fixed. We examine several such surface variations: the question form and how its negation is written (prefix vs. the word not), the activity scenario, and the name’s nationality and gender. The tables here show the models that move most; full per-model tables are in Appendix A. 9 arXiv TemplateA PREPRINT 5.1.1 Question Form Among affirmative questions, the five question forms land within about four points of each other (67.6–71.4%), a much smaller spread than what we see with negation (29.2–58.4%; Table 4). The collapse is confined to the negated wrapper of the FOLLOW form, Does it not follow that . . . ?, which falls to 29.2%, below the 50% chance level and nearly 20 points under the next-lowest wrapper (48.8%). The other four negated wrappers cluster together: Varying adjectives with Is it false that . . . ? (58.4%), Is it incorrect that . . . ? (56.0%) and Is it invalid to conclude that . . . ? (48.8%), together with the proposition-level Is it not probable that . . . ? (56.5%), all sit between 49 and 58%. Measured within each form as the affirmative−negated drop (polarity sensitivity), the gap is smallest for the TRUTH form (Is it true that vs. Is it false that,+0.125) and largest for the FOLLOW form (Does it follow that vs. Does it not follow that,+0.384); the FOLLOW form is the only one whose negated wrapper leaves the shared range. The collapse is broad but not uniform: 21 of 29 models score at least 10 points lower on does it not follow than on the other negated wrappers, and the strongest models that fall furthest (Gemma4:31B, Sonnet 4.6 and Opus 4.7 each drop 60 or more points), while the handful that do not show it are near-random constant responders already poor in all cases. Tellingly, the DIRECT form’s negated wrapper Is it not probable that . . . ? also contains the word not yet holds at 56.5%, inside that range, so the difficulty is specific to writing the negation as not inside the entailment wrapper (does it not follow), not the word not itself. Table 4: Per-question-form results (all English models pooled). Aff/Neg are affirmative/negated accuracy (micro- averaged over all templates and models, uncertain counts as incorrect); PS is polarity sensitivity (Aff−Neg). Sorted by PS. Under negation only the FOLLOW form’s wrapper, Does it not follow that . . . ?, collapses to near chance; the other four stay in a tight band, including DIRECT (Is it not probable that . . . ?), which also uses the word not. FormAffirmative wrapperNegated wrapperAffNegPS TRUTHIs it true that . . . ?Is it false that . . . ?.710.584+.125 CORRECTIs it correct that . . . ?Is it incorrect that . . . ?.699.560+.139 DIRECTIs it probable that . . . ?Is it not probable that . . . ?.714.565+.150 VALIDIs it valid to conclude that . . . ?Is it invalid to conclude that . . . ?.676.488+.188 FOLLOWDoes it follow that . . . ?Does it not follow that . . . ?.677.292+.384 Per model, the same divergence shows up as a wide swing across forms; the nine floor-clearing models have question- form ranges from about 0.19 to 0.50 (Table 5), each concentrated in the does it not follow form. Table 5: Question-form sensitivity (models that move most). Overall accuracy in each of the five question forms (fraction correct, micro-averaged over all templates and both the affirmative and negated conditions; uncertain counts as incorrect); Range is the best−worst spread. The remaining 23 models have Range0.02–0.32; the smallest belong to committed responders (Qwen3:1.7B 0.02, DeepSeek-R1:7B 0.06). Full per-model table in Table 15. ModelTRUTHCORRECTVALIDFOLLOWDIRECTRange gemma4:31b0.850.871.000.500.880.50 Sonnet 4.60.820.800.880.500.740.38 gemma4:12b0.800.830.860.520.820.34 gpt-oss:200.850.840.720.510.770.33 Opus 4.70.830.820.840.510.780.33 gpt-oss:120b0.850.790.630.530.820.32 To test whether the negation strategy itself drives this collapse, we added not-word versions of the CORRECT and VALID forms (Is it not correct that . . . ?, Is it not valid to conclude that . . . ?), creating matched pairs where the only difference is whether the negation is written as a prefix (incorrect/invalid) or with the word not. 5.1.2 Prefix Negation versusnot-Negation Since the not negation is pragmatically ambiguous—it has a possible interpretation as pining for agreement with a rhetorical question—while prefix negation is unambiguous, these results provide information about the effect of this pragmatic ambiguity versus an unambiguous control. Comparing these matched pairs (prefix incorrect/invalid against not correct/not valid), the prefix form yields 52.3% accuracy versus 33.6% for the not form, an 18.7-point gap, suggesting that the ambiguous form confuses the model. Table 6 presents per-model results. Each model contributes 2,400 paired observations per negation style. The per-model gaps are far larger: Sonnet drops 61.1 points, Opus 59.1, 10 arXiv TemplateA PREPRINT Qwen3:32B 40.3. The direction is not uniform: 25 of 29 models perform better with the prefix form, but 4 (Qwen3:0.6B, Gemma3:1B, Qwen3:1.7B, Llama3.2:3B) show the reverse. Table 6: Negation written as a prefix vs. with the word not, for all 29 models. Prefix=incorrect/invalid (negation fused into the adjective); not=not correct/not valid (a separate “not”). Both sides are negated questions with the same proposition and gold answer; only the negation surface form differs.∆ =Prefix accuracy−not accuracy. The Yes-rate columns give the model’s percentage of Yes answers under each form. AccuracyYes-rate (%) ModelPrefixnot∆Prefixnot gemma4:12b0.7760.137+0.6399.975.6 Sonnet 4.60.7850.175+0.61123.584.5 Opus 4.70.7810.208+0.57210.066.4 gemma4:31b0.9150.372+0.54318.156.5 qwen3:32b0.6660.263+0.40337.251.2 gpt-5.4-mini0.6320.298+0.33423.858.3 aya-expanse:32b0.5030.193+0.31037.487.5 gpt-oss:200.7230.426+0.29732.336.0 granite3.2:8b0.7270.440+0.28610.335.8 gpt-oss:120b0.6950.446+0.25023.134.9 gemma3:27b0.4250.182+0.24353.677.5 gemma3:12b0.5020.290+0.21228.566.8 llama3.1:8b0.5100.356+0.15440.555.5 aya-expanse:8b0.4410.300+0.14064.083.4 qwen3:4b0.5310.417+0.11326.435.9 gemma4:e2b0.6790.570+0.1085.621.8 qwen3:8b0.3610.255+0.10662.882.6 qwen3:14b0.6490.560+0.08930.826.4 gemma4:e4b0.4580.421+0.03638.066.1 deepseek-r1:7b0.1030.077+0.02637.832.8 gemma3:270m0.4430.428+0.01550.953.8 deepseek-r1:14b0.2500.236+0.01493.791.1 mistral:7b0.2940.284+0.01074.283.1 gemma3:4b0.5310.526+0.00550.060.3 deepseek-r1:32b0.2610.259+0.00298.398.3 qwen3:0.6b0.3510.359 −0.00882.788.2 gemma3:1b0.2450.257 −0.01296.493.5 qwen3:1.7b0.3860.425 −0.03979.069.5 llama3.2:3b0.5470.591 −0.04431.117.1 5.1.3 Activity Replacing the scenario verb phrases (“party”, “event”, “gathering”, “meeting”, or “celebration”) attenuates accuracy. The average per-model range across the five activities is 3.4 pts, relatively minor compared to negation and question- form effects. Only DeepSeek-R1:7b moves substantially (13.2 pts). As with nationality, this is a refusal pattern, its refusal rate ranging from 46% to 72% across the five activities while accuracy on the items it answers stays near 52%; so the model is sensitive to the verb but expresses that as refusal to answer rather than as different reasoning (Table 7). Table 7: Accuracy by activity scenario (models that move most; uncertain counts as incorrect). Range is the spread across the five activities. The remaining 24 models have Range≤ 0.04(mean Range= 0.034). Full per-model table in Table 17. ModelPartyEventGatherMeetingCelebr.Range deepseek-r1:7b0.250.280.170.170.150.13 qwen3:8b0.600.640.580.610.570.06 llama3.1:8b0.430.490.490.440.460.06 gpt-oss:200.770.730.720.750.710.06 gemma3:27b0.650.670.630.660.620.05 11 arXiv TemplateA PREPRINT 5.1.4 Nationality We find biases are introduced by varying gender, nationality, and activity with every model. As a control, we also use variablesXandYinstead of people’s names to compare against. Names are drawn from seven nationality groups, and Table 8 reports each nationality’s accuracy relative to the abstract variable (X,Y) baseline. (A positive value means the model does better on that nationality than on the variables.) For 23 of 29 the largest per-nationality gap is at most 4.7 points. The clear exception is DeepSeek-R1:7b, which sits 13 to 34 points below variables on every nationality, being worst on Japanese. It frequently refuses to answer the question, but on the items it actually answers, its accuracy across nationalities spans about 6 points (51–57%), against roughly 21 points on the headline accuracy; the gap comes instead from how often it refuses, and that refusal rate itself swings sharply with nationality, from 46% on Indian names to 87% on Japanese. Table 8: Accuracy on each nationality group minus accuracy on the abstract letter-variable (X,Y) baseline, in percentage points (models that move most; uncertain counts as incorrect). Positive=higher on that nationality than on XY. Ind=Indian, Rus=Russian, Jpn=Japanese, Afr=African, Ger=German, Fre=French, Amr=American.max|∆|is the largest absolute gap. The remaining 23 models have max|∆|≤ 4.7 points. Full per-model table in Table 18. ModelIndRusJpnAfrGerFreAmrmax|∆| deepseek-r1:7b −13.5 −23.7 −34.1 −20.2 −26.3 −27.8 −23.434.1 deepseek-r1:32b −1.9 −0.5 −1.6 −2.2 −0.8 −9.5 −1.79.5 llama3.2:3b −3.6 −2.2 −9.1 −2.7 −2.8 −8.2 −4.09.1 llama3.1:8b −2.2+1.9+3.2+7.5+3.8+4.8+6.27.5 gemma3:12b −2.1 −3.6 −3.9 −3.8 −2.2 −4.8 −2.84.8 Opus 4.7+2.1+4.8+0.9+1.1+1.3+0.0+1.14.8 5.1.5 Gender Names are balanced by gender, and Table 9 reports each gender’s accuracy relative to the variable (X,Y) control. Women’s and men’s names score within about 3 points of variables for every model except DeepSeek-R1:7b, which exhibits the same refusal behavior seen for nationality. The man-woman name gap itself average 0.6 percentage points; the largest being DeepSeek-R1:32b (+2.3 toward men’s names) and GPT-5.4-mini (+2.2 toward women’s names). 5.2 Conjunction Variants We test two conjunction inferences, each written two ways that keep the logic fixed (Table 10), with per-model accuracies in Table 11. Recall that distribution over conjunction (DOC) is valid while conjunctivitis is invalid (§3.1). By DOC,Pr(A∧ B) =⇒ Pr(A)∧ Pr(B). To test for consistent reasoning, we vary whether we ask the model about the truth of Pr(A) or P (B), since logically both are true. On DOC examples, affirmative questions differ by about 5 points on average, and 25 of 29 models stay within 10 points, a large gap for such a simple variation. For conjunctivitis examples how the two premises are worded matters even more. We see this by testing two conditions: one with two people engaging in the same activity and one with two people engaging in different activities. On affirmative items, where the model has to reject the inference, the same-activity wording (two people at one activity) is easier to reject than the dual-activity wording (one person at two activities) by a mean of 12.8 points (up to 96 points for Qwen3:14B, and Sonnet drops from 96.5% to 56.5%). This gap resides entirely in the reject condition. On negated items, where the answer is Yes, the two wordings yield very close results (0.3 points), so pooling the two conditions dilutes the effect to 6.6 points, which is why we report it on the affirmative items. 5.3 Per-Template Breakdown The 13 templates greatly vary in accuracy (42.8% for Conjunctivitis to 74.3% for Probably to Might). The key finding is a dissociation (Table 12): the answer bias varies less across templates than polarity sensitivity does (|Bias|std= 0.056 versus PS std= 0.126, about twice as much), and it stays high on every template (0.42 to 0.62). The bias is thus a model-level property carried across all templates, while negation sensitivity is template-dependent. Model biases also change direction based on template: on the easiest template (Probably to Might) 28 of 29 models are biased toward Yes, so nearly all give the same answer. On the two hardest templates the panel splits roughly in half (Conjunctivitis 13 yes-biased vs. 16 no-biased; Might to Probably 12 vs. 17): the template exerts no shared pull, so which answer a model gives is set mostly by its own bias rather than by the content of the inference. 12 arXiv TemplateA PREPRINT Table 9: Accuracy on female and male names minus accuracy on the variable (X,Y) control, in percentage points (models that move most; uncertain counts as incorrect). Positive=higher thanX,Y; the male−female gap is the difference between the two columns.max|∆|is the larger absolute gap. The remaining 23 models havemax|∆|≤ 3.0 points. Full per-model table in Table 16. ModelFemaleMalemax|∆| deepseek-r1:7b −24.3 −24.024.3 llama3.2:3b −4.8 −4.54.8 llama3.1:8b+4.0+3.24.0 deepseek-r1:32b −3.8 −1.43.8 gemma3:12b −3.6 −3.03.6 qwen3:14b −2.7 −3.13.1 Table 10: The two conjunction inference types, each shown in two surface realizations (affirmative condition; the negated condition swaps Is it true that for Is it false that and flips the gold answer). The distribution-over-conjunction rows differ only in which name is queried; the conjunctivitis rows differ only in whether the two premises describe two people at one activity or one person at two activities. Template (logic)Example (affirmative condition)Gold Dist. over conjunc- tion, valid, query first name It is probable that Savir and Ashwin will be at the party. Is it true that it is probable that Savir will be at the party? Yes Dist. over conjunc- tion, valid, query sec- ond name It is probable that Savir and Ashwin will be at the party. Is it true that it is probable that Ashwin will be at the party? Yes Conjunctivitis,in- valid, same activity It is probable that Savir will be at the party. It is probable that Ashwin will be at the party. Is it true that it is probable that Savir and Ashwin will be at the party? No Conjunctivitis,in- valid, dual activity It is probable that Savir will attend the event. It is probable that Savir will give a presentation. Is it true that it is probable that Savir will attend the event and give a presentation? No 5.4 Model Scale Four families provide size variants (Table 13). The answer bias tends to fall as models grow, most clearly in Qwen 3 (Spearmanρ =−0.89between|Bias|and size) and more weakly in Gemma 3 (ρ =−0.60) and Gemma 4 (ρ =−0.40), but no family is monotonic (Gemma 3, for instance, spikes at 1B before falling again). DeepSeek-R1 runs the other way, its bias rising with size (ρ = +0.50). Thus scaling lowers the bias on average in three of four families, never cleanly, and not at all in the fourth. The FOLLOW question form (the word not) shows little scaling improvement across families, suggesting the difficulty with not is not resolved by additional parameters. 6 Discussion and Conclusion Overall, we find widespread biases across multiple dimensions, including name origin, gender, and various perturbations to question formation. More generally, we find definitive evidence that while some models manage to perform above chance, many do not even reach this low bar, and LLMs do not, in general, make basic EM inferences consistently, making them inappropriate for use in question-answering settings that require synthesizing epistemic statements to make valid conclusions. What accuracy can hideMost models show a fixed answer bias for Yes or No belied by aggregate accuracy. (Valid) reasoning requires both accepting valid inferences and rejecting invalid ones, but we demonstrate conclusively that state of the art models fail to do this. We suggest that future benchmarks using binary questions to evaluate reasoning also report the competence floor, which is cheap to compute and is not inflated by an answer bias. While precision and recall also attempt to address this, they are more difficult to interpret in this specific scenario. 13 arXiv TemplateA PREPRINT Table 11: Per-model accuracy on the two conjunction templates under a wording change that leaves the logic fixed (affirmative items; uncertain counts as incorrect). Conjunctivitis is invalid (correct answer No), so its accuracy is the rejection rate; Same states the two premises as two people at one activity and Dual as one person at two activities. Distribution over conjunction is valid (correct answer Yes); Name 1 and Name 2 ask about the first or second name in the shared premise.∆is the accuracy difference between the two wordings. Rewriting the conjunctivitis premises moves accuracy by a mean of 12.8 points (up to 96), while which name distribution over conjunction asks about moves it by only 4.9 on average. Sorted by the conjunctivitis ∆. Conjunctivitis (reject rate)Dist. over conjunction ModelSameDual∆Name 1Name 2∆ qwen3:14b1.000.04+0.960.900.98 −0.09 qwen3:8b0.510.00+0.511.001.00+0.00 Sonnet 4.60.960.56+0.401.001.00+0.00 gemma3:12b0.350.01+0.351.001.00+0.00 gemma3:27b0.700.35+0.351.001.00+0.00 gemma4:12b0.880.55+0.330.971.00 −0.03 granite3.2:8b0.960.71+0.260.680.82 −0.14 qwen3:32b0.830.59+0.240.970.99 −0.02 deepseek-r1:32b0.270.04+0.231.001.00 −0.00 gpt-oss:200.900.69+0.210.860.84+0.02 gpt-oss:120b0.890.71+0.180.950.93+0.02 gemma4:e4b0.490.40+0.090.940.98 −0.04 gemma3:4b0.090.01+0.090.790.95 −0.17 deepseek-r1:14b0.080.00+0.080.990.98+0.01 llama3.1:8b0.130.05+0.080.550.64 −0.08 aya-expanse:32b0.020.00+0.021.001.00+0.00 qwen3:1.7b0.020.00+0.021.000.99+0.01 gpt-5.4-mini0.950.94+0.010.230.66 −0.43 aya-expanse:8b0.000.00+0.001.001.00+0.00 deepseek-r1:7b0.000.00+0.000.230.19+0.04 gemma3:1b0.000.00+0.001.001.00+0.00 gemma4:31b1.001.00+0.001.001.00+0.00 qwen3:0.6b0.000.00+0.001.000.99+0.01 qwen3:4b1.001.00+0.000.990.97+0.02 Opus 4.70.991.00 −0.010.820.95 −0.14 gemma4:e2b0.980.99 −0.010.920.97 −0.05 llama3.2:3b0.860.89 −0.020.430.48 −0.05 mistral:7b0.000.06 −0.060.991.00 −0.01 gemma3:270m0.070.67 −0.600.490.47+0.02 Mean0.520.39+0.130.850.89 −0.04 Negation is not the main source of answer bias Two of our probes use negation, and answer biases do not depend on it: the affirmative-only competence floor correlatesr = 0.90with the full (affirmative+negation) floor, and the most biased models (Qwen3:0.6B at 0.030) score the worst at both. What the floor does and does not sayThe floor measures whether a model performs above chance on both answer classes. The invalid templates are uncontroversial: the conjunction fallacy, possibility not entailing probability, and probability not entailing certainty hold under any reasonable semantics. However, some valid templates rest on stronger commitments. Conditional-to-comparative and probably-to-not-probably-not depend on specific axioms a competent reasoner could contest, though this is not responsible for the results: recomputing the floor with those two templates leaves the number of models with accuracy above 0.5 unchanged at 9 of 29 and moves every above-floor model by at most 0.03 (Sonnet 0.642 to 0.665, Opus 0.570 to 0.567, Qwen3:32B 0.594 to 0.606). Restricting further to a conservative core of only uncontestable inferences (the three invalid templates plus the four most basic valid ones, where probablyφentails mightφ, mustφentails probablyφ, distribution over conjunction, and chancy disjunction introduction) tells the same story (10 of 29 above 0.5), and in both restrictions the most biased models remain at the bottom (DeepSeek-R1:7B at 0.00). Biased responders fail the uncontroversial, invalid templates regardless of semantics, so their low floor cannot be attributed to legitimately contested axioms. An expert or human baseline calibrating how much residual sub-floor behavior is theory disagreement is the natural next step. 14 arXiv TemplateA PREPRINT Table 12: Per-template analysis (all English models pooled), grouped by validity. Acc is overall accuracy on the template, a micro-average with uncertain counted as incorrect.|Bias|is the mean absolute answer bias (from the yes-rate), PS the mean polarity sensitivity, and Agree. the fraction of models sharing the majority bias direction. TypeTemplateAcc |Bias|PSAgree. ValidChancy Disjunction Introduction.574.539+.09862.1% Chancy Modus Ponens.695.453+.27672.4% Chancy Modus Tollens.547.526+.09558.6% Complement Transfer.604.496+.18365.5% Conditional to Comparative.540.497+.13462.1% Distribution over Conjunction.695.418+.35189.7% Must to Probably.633.509+.17262.1% Positive Form Transfer.674.488+.02365.5% Probably to Might.743.463+.41796.6% Probably to not probably not.578.548+.33579.3% InvalidConjunctivitis.428.619+.04755.2% Might to Probably.447.608+.09758.6% Probably to Certain.683.462+.34279.3% Table 13: Scaling within size-graded model families (English). Acc is overall accuracy with uncertain responses counted as incorrect.|Bias|is the magnitude of the signed answer bias and PS the polarity sensitivity, both defined in §3.5 (Bias= Acc(Yes)− Acc(No); PS= Acc(affirmative)− Acc(negated), so positive=better on affirmative). Unc% is the uncertain rate. Sizes are listed smallest to largest within each family. FamilySizeAcc |Bias|PSUnc% Gemma 3270M0.5000.163+0.1070.0 1B0.5370.745+0.0560.2 4B0.5290.192-0.1130.0 12B0.6150.016+0.2090.0 27B0.6480.033+0.4111.0 Gemma 4E2B0.5900.662+0.3151.7 E4B0.6080.191+0.1890.0 12B0.7640.130+0.3430.0 31B0.8180.196+0.4120.0 Qwen 30.6B0.5480.843-0.0050.0 1.7B0.5170.530+0.0191.5 4B0.5690.568+0.3370.0 8B0.5990.277+0.2380.0 14B0.6050.394+0.0530.0 32B0.7180.174+0.2790.0 DeepSeek-R17B0.2020.385+0.02861.5 14B0.4770.823+0.06810.6 32B0.5560.729+0.1372.4 Limitations Our evaluation is zero-shot and English-only, and parses the first answer token rather than a free-form justification; whether these patterns persist under longer generation or in other languages is left to future work. 7 Conclusion We have introduced a benchmark for inference over probability operators that separates a model’s answer bias from proposition-tracking, demonstrating conclusively that models consistently fail at zero-shot reasoning under trivial prompts with an unambiguous correct response. Most models show a fixed bias toward Yes or No: only 9 of 29 models exceed a competence floor with an absolute random baseline, and the result reproduces from the affirmative validity contrast alone, so it does not appear to depend on negation. Negation variation exposes a strong bias, where prefix and not phrasing of the same content differ by up to 64 points. Taken together, these results strongly suggest that part of what is commonly reported as “reasoning” ability may reflect answer bias and is not the result of a coherent internal inference model. A competence floor with an absolute baseline is a useful and easy-to-calculate complement to accuracy 15 arXiv TemplateA PREPRINT 0.27B1B4B12B27B Parameters (B) 0 20 40 60 80 100 Accuracy (%) gemma3 2B4B12B31B Parameters (B) 0 20 40 60 80 100 Accuracy (%) gemma4 3B8B Parameters (B) 0 20 40 60 80 100 Accuracy (%) llama 7B14B32B Parameters (B) 0 20 40 60 80 100 Accuracy (%) deepseek-r1 0.6B1.7B4B8B14B32B Parameters (B) 0 20 40 60 80 100 Accuracy (%) qwen3 8B32B Parameters (B) 0 20 40 60 80 100 Accuracy (%) aya-expanse 20B120B Parameters (B) 0 20 40 60 80 100 Accuracy (%) gpt-oss Scaling effect on accuracy, by answer cell Valid · non-negatedValid · negatedInvalid · non-negatedInvalid · negatedchance: 50% Figure 2: Effect of scaling on accuracy by answer cell (valid/invalid×negated/non-negated) across model families. Dashed line marks 50% chance. Scaling does not consistently improve accuracy. as a summary of logical competence. Our prompts are simple and formulaic, and all information necessary to answer correctly is provided in the prompts, suggesting that more subtle ways of expressing the same premises would not lead to successful inferences. Since most inference work limited to English, future work can extend this to other languages. Acknowledgments Grissom is supported by U.S. National Science Foundation grant 2403439. References Binz, Marcel and Eric Schulz. 2023. Using cognitive psychology to understand GPT-3. Proceedings of the National Academy of Sciences, 120(6):e2218523120. Braun, Daniel. 2025. Acquiescence bias in large language models. arXiv preprint arXiv:2509.08480. Elkins, Katherine and Jon Chun. 2026. When prohibitions become permissions: Auditing negation sensitivity in language models. arXiv preprint arXiv:2601.21433. Ermakova, Liana, Anton Firsov, and Jaap Kamps. 2026. Confirmation, framing, and position biases in LLM responses. In Proceedings of the 2026 Conference on Human Information Interaction and Retrieval (CHIIR ’26), pages 480–484, ACM. Fei, Yu, Yifan Hou, Zeming Chen, and Antoine Bosselut. 2023. Mitigating label biases for in-context learning. In Proceedings of ACL 2023, pages 14014–14031. 16 arXiv TemplateA PREPRINT García-Ferrero, Iker, Begoña Altuna, Javier Alvez, Itziar Gonzalez-Dios, and German Rigau. 2023. This is not a dataset: A large negation benchmark to challenge large language models. In Proceedings of EMNLP 2023, pages 8596–8615. Holliday, Wesley H, Matthew Mandelkern, and Cedegao E Zhang. 2024. Conditional and modal reasoning in large language models. In Proceedings of EMNLP 2024, pages 3800–3821. Holtzman, Ari, Peter West, Vered Shwartz, Yejin Choi, and Luke Zettlemoyer. 2021. Surface form competition: Why the highest probability answer isn’t always right. In Proceedings of EMNLP 2021, pages 7038–7051. Imannezhad, Pegah, Emmanuel M Pothos, and Andy J Wills. 2026. Divergent patterns of probabilistic reasoning in humans and GPT-5. Frontiers in Psychology, 17:1782184. Jiang, Bowen, Yangxinyu Xie, Zhuoqun Hao, Xiaomeng Wang, Tanwi Mallick, Weijie J Su, Camillo J Taylor, and Dan Roth. 2024. A peek into token bias: Large language models are not yet genuine reasoners. In Proceedings of EMNLP 2024, pages 4722–4756. Kratzer, Angelika. 1991. Modality. In Semantics: An International Handbook of Contemporary Research. de Gruyter, pages 639–650. Kratzer, Angelika. 2012. Modals and Conditionals. Oxford University Press. Kyburg Jr, Henry E. 1970. Conjunctivitis. In Induction, acceptance and rational belief. Springer, pages 55–82. Lassiter, Daniel. 2011. Measurement and Modality: The Scalar Basis of Modal Semantics. Ph.D. thesis, New York University. Lassiter, Daniel. 2017. Graded Modality: Qualitative and Quantitative Perspectives. Oxford University Press. Li, Meng, Michael Vrazitulis, and David Schlangen. 2025. Representations of fact, fiction and forecast in large language models: Epistemics and attitudes. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 27734–27757. Macmillan-Scott, Olivia and Mirco Musolesi. 2024. (ir)rationality and cognitive biases in large language models. Royal Society Open Science, 11(6):240255. McNemar, Quinn. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153–157. Moore, Robert C. 1981. Reasoning about knowledge and action. In Readings in artificial intelligence. Elsevier, pages 473–477. Sclar, Melanie, Yejin Choi, Yulia Tsvetkov, and Alane Suhr. 2024. Quantifying language models’ sensitivity to spurious features in prompt design. In Proceedings of ICLR 2024. Sharma, Mrinank, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. 2024. Towards understanding sycophancy in language models. In The Twelfth International Conference on Learning Representations (ICLR). Shi, Ruikang, Alvin Grissom I, and Duc Minh Trinh. 2022. Rare but severe neural machine translation errors induced by minimal deletion: An empirical study on Chinese and English. In Proceedings of the 29th International Conference on Computational Linguistics, pages 5175–5180, International Committee on Computational Linguistics, Gyeongju, Republic of Korea. Singh, Shubham. 2026. Damensage: Chatgpt statistics (july 2026) – latest active users data.https://w. demandsage.com/chatgpt-statistics/. [Accessed 13-07-2026]. So, Yeonkyoung, Gyuseong Lee, Sungmok Jung, Joonhak Lee, JiA Kang, Sangho Kim, and Jaejin Lee. 2025. Thunder- NUBench: A benchmark for LLMs’ sentence-level negation understanding. arXiv preprint arXiv:2506.14397. Teller, Paul. 1972. Epistemic possibility. Philosophia, 2(4):303–320. Tjuatja, Lindia, Valerie Chen, Tongshuang Wu, Ameet Talwalkar, and Graham Neubig. 2024. Do LLMs exhibit human-like response biases? a case study in survey design. Transactions of the Association for Computational Linguistics, 12:1011–1026. Truong, Thinh Hung, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023. Language models are not naysayers: An analysis of language models on negation benchmarks. In Proceedings of *SEM 2023, pages 101–114. Yalcin, Seth. 2010. Probability operators. Philosophy Compass, 5(11):916–937. Zhao, Zihao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of ICML 2021, pages 12697–12706. Zhou, Han, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine Heller, and Subhrajit Roy. 2024. Batch calibration: Rethinking calibration for in-context learning and prompt engineering. In Proceedings of ICLR 2024. 17 arXiv TemplateA PREPRINT A Additional Results Table 14: Overall accuracy under five pooling schemes (frames not correct/not valid/conclude excluded, templates folded to 13, uncertain counts as incorrect). All five score the same responses and differ only in weighting. micro is the headline flat mean over every response. prompt-bal, tmpl-bal and frame-bal weight each prompt, template or question form equally. cat-avg weights the four validity×negation cells equally,(V.Aff + V.Neg + I.Aff + I.Neg)/4. Frame-balancing equals micro by construction and prompt-balancing is within a fraction of a point of it; only cat-avg moves the number appreciably, always toward 0.5, because it stops the abundant easy valid cells from dominating the rare invalid-negation cell. Spread is the max−min across the five schemes for each model. The mean weighting spread is 4.7 points, and averaged over models cat-avg sits 3.2 points below micro. Sorted by micro. Modelmicroprompt-baltmpl-balframe-balcat-avgSpread gemma4:31b0.8180.8180.8290.8190.7730.056 gemma4:12b0.7640.7640.7840.7640.6790.105 Opus 4.70.7560.7820.7690.7560.6930.089 Sonnet 4.60.7510.7730.7610.7500.6920.081 gpt-oss:200.7370.7370.7490.7370.7010.048 gpt-oss:120b0.7240.7240.7330.7250.6870.046 qwen3:32b0.7180.7180.7260.7180.6810.045 gpt-5.4-mini0.6560.6560.6680.6560.6270.041 gemma3:27b0.6480.6480.6650.6480.6070.058 aya-expanse:32b0.6240.6240.6520.6250.5450.107 gemma3:12b0.6150.6150.6310.6150.5260.105 gemma4:e4b0.6080.6080.6130.6090.5690.044 granite3.2:8b0.6070.6070.6060.6080.5870.021 qwen3:14b0.6050.6050.6040.6050.5740.031 qwen3:8b0.5990.5990.6070.6000.5360.071 gemma4:e2b0.5900.5900.5810.5900.5640.026 aya-expanse:8b0.5860.5860.5960.5860.5170.079 qwen3:4b0.5690.5690.5660.5690.5580.011 deepseek-r1:32b0.5560.5560.5590.5570.5830.027 qwen3:0.6b0.5480.5480.5520.5480.5160.036 gemma3:1b0.5370.5370.5350.5370.5060.031 mistral:7b0.5350.5350.5340.5360.5090.027 gemma3:4b0.5290.5290.5320.5300.5000.032 qwen3:1.7b0.5170.5170.5210.5170.4770.044 gemma3:270m0.5000.5000.4990.5000.4940.006 deepseek-r1:14b0.4770.4770.4720.4760.4860.014 llama3.2:3b0.4670.4670.4680.4670.5110.044 llama3.1:8b0.4630.4630.4810.4640.4820.019 deepseek-r1:7b0.2020.2020.2010.2020.1930.009 Mean0.5970.5980.6030.5970.5650.047 18 arXiv TemplateA PREPRINT Table 15: Per-model question-form sensitivity, full 29-model table (uncertain counts as incorrect). Overall accuracy in each of the five question forms; Range is the best−worst spread. Ordered by competence floor; the faded rule separates the nine models that clear the 0.5 floor from those that do not (§5.1.1). ModelTRUTHCORRECTVALIDFOLLOWDIRECTRange gemma4:31b0.850.871.000.500.880.50 Sonnet 4.60.820.800.880.500.740.38 gemma4:12b0.800.830.860.520.820.34 qwen3:32b0.830.720.720.520.800.31 gemma3:27b0.730.620.670.530.690.19 gpt-oss:200.850.840.720.510.770.33 Opus 4.70.830.820.840.510.780.33 gpt-oss:120b0.850.790.630.530.820.32 gemma3:12b0.730.660.560.510.620.21 gemma4:e4b0.710.680.450.510.700.26 llama3.1:8b0.500.580.350.280.610.32 gpt-5.4-mini0.760.730.640.480.670.28 aya-expanse:32b0.730.690.540.460.710.28 gemma3:270m0.500.550.470.470.510.08 gemma3:4b0.570.540.460.440.640.20 qwen3:8b0.630.620.540.490.720.24 qwen3:14b0.650.640.510.540.680.17 aya-expanse:8b0.670.620.530.470.640.20 mistral:7b0.610.520.480.430.630.20 granite3.2:8b0.680.620.620.450.660.23 qwen3:4b0.630.600.470.530.610.16 llama3.2:3b0.490.510.420.450.460.09 gemma4:e2b0.650.650.620.470.560.18 deepseek-r1:32b0.560.520.550.560.600.07 qwen3:1.7b0.500.520.530.520.510.02 gemma3:1b0.550.510.500.500.620.12 qwen3:0.6b0.500.510.590.610.540.11 deepseek-r1:14b0.420.470.520.530.440.11 deepseek-r1:7b0.170.210.210.230.190.06 19 arXiv TemplateA PREPRINT Table 16: Accuracy on female and male names minus accuracy on the letter-variable (XY) baseline, all 29 models, in percentage points (uncertain counts as incorrect). Positive=higher than XY.max|∆|is the larger absolute gap. Sorted by max|∆|. ModelFemaleMalemax|∆| deepseek-r1:7b −24.3 −24.024.3 llama3.2:3b −4.8 −4.54.8 llama3.1:8b+4.0+3.24.0 deepseek-r1:32b −3.8 −1.43.8 gemma3:12b −3.6 −3.03.6 qwen3:14b−2.7 −3.13.1 qwen3:1.7b+2.9+3.03.0 granite3.2:8b −2.9 −2.62.9 gemma3:27b −2.8 −1.72.8 gpt-5.4-mini+0.5 −2.72.7 gemma4:e2b −2.4 −2.62.6 gemma3:270m+2.1+1.82.1 gemma3:4b+1.7+1.91.9 Opus 4.7+1.3+1.91.9 aya-expanse:32b −1.6 −1.81.8 qwen3:8b−1.7 −1.31.7 qwen3:0.6b+1.6+1.71.7 deepseek-r1:14b+1.6+0.61.6 gpt-oss:20+1.5+1.01.5 gpt-oss:120b+0.5+1.11.1 gemma4:e4b+0.1+1.01.0 Sonnet 4.6−1.0+0.11.0 gemma3:1b+0.9+1.01.0 qwen3:32b−0.2+0.90.9 qwen3:4b+0.2 −0.70.7 mistral:7b+0.6+0.30.6 aya-expanse:8b+0.6+0.40.6 gemma4:12b −0.0 −0.50.5 gemma4:31b+0.0+0.20.2 20 arXiv TemplateA PREPRINT Table 17: Accuracy by activity scenario, all 29 models (uncertain counts as incorrect). Range is the spread across the five activities. Sorted by Range. ModelPartyEventGatherMeetingCelebr.Range deepseek-r1:7b0.250.280.170.170.150.13 qwen3:8b0.600.640.580.610.570.06 llama3.1:8b0.430.490.490.440.460.06 gpt-oss:200.770.730.720.750.710.06 gemma3:27b0.650.670.630.660.620.05 qwen3:32b0.700.740.710.730.700.04 gemma4:e2b0.580.610.570.610.580.04 gemma4:e4b0.630.620.590.590.600.04 deepseek-r1:32b0.560.580.560.550.530.04 llama3.2:3b0.460.480.470.480.450.04 mistral:7b0.560.530.530.540.520.04 deepseek-r1:14b0.490.490.460.470.460.04 granite3.2:8b0.620.610.580.620.610.04 qwen3:1.7b0.510.510.500.540.530.03 gemma3:12b0.620.620.590.620.620.03 gpt-5.4-mini0.640.660.660.670.650.03 qwen3:4b0.560.570.570.580.560.03 qwen3:0.6b0.540.540.550.560.550.03 Opus 4.70.750.760.750.770.770.02 aya-expanse:8b0.580.590.570.590.590.02 gemma4:12b0.760.770.760.770.760.02 gpt-oss:120b0.720.740.720.730.720.02 gemma3:4b0.520.540.530.530.520.02 aya-expanse:32b0.630.620.620.630.620.02 gemma4:31b0.820.820.810.810.820.01 qwen3:14b0.610.610.610.600.600.01 Sonnet 4.60.750.750.750.760.760.01 gemma3:270m0.500.500.500.490.500.01 gemma3:1b0.530.530.530.540.540.01 21 arXiv TemplateA PREPRINT Table 18: Accuracy on each nationality group minus accuracy on the abstract letter-variable (XY) baseline, all 29 models, in percentage points (uncertain counts as incorrect). Positive=higher on that nationality than on XY. Ind=Indian, Rus=Russian, Jpn=Japanese, Afr=African, Ger=German, Fre=French, Amr=American.max|∆|is the largest absolute gap. Sorted by max|∆|. ModelIndRusJpnAfrGerFreAmrmax|∆| deepseek-r1:7b −13.5 −23.7 −34.1 −20.2 −26.3 −27.8 −23.434.1 deepseek-r1:32b −1.9 −0.5 −1.6 −2.2 −0.8 −9.5 −1.79.5 llama3.2:3b −3.6 −2.2 −9.1 −2.7 −2.8 −8.2 −4.09.1 llama3.1:8b −2.2+1.9+3.2+7.5+3.8+4.8+6.27.5 gemma3:12b −2.1 −3.6 −3.9 −3.8 −2.2 −4.8 −2.84.8 Opus 4.7+2.1+4.8+0.9+1.1+1.3+0.0+1.14.8 qwen3:14b −3.1 −2.8 −2.8 −2.2 −2.1 −4.7 −2.54.7 qwen3:1.7b+2.9+3.0+2.4+4.6+3.1+3.2+1.44.6 gemma3:4b+4.6+4.1+0.5+0.2+0.9+0.5+1.74.6 gemma3:27b −1.1 −0.9 −4.1 −3.3 −1.4 −2.9 −2.14.1 deepseek-r1:14b+2.5+3.6 −2.8+3.9+0.5 −2.2+2.13.9 granite3.2:8b −3.2 −1.1 −3.2 −3.8 −3.5 −2.0 −2.23.8 gemma4:e2b −2.6 −2.4 −2.9 −3.5 −2.2 −1.6 −2.53.5 qwen3:0.6b+0.3+3.3+2.5 −0.5+1.5+2.1+2.23.3 gpt-oss:20+1.4+0.5+0.1+1.9+1.5+0.3+3.33.3 gpt-5.4-mini −1.6 −1.9 −0.9 −0.4 −2.2 −3.2+2.43.2 gemma3:270m+1.4+1.0+3.0+1.3+2.0+2.4+2.23.0 qwen3:8b−1.6 −0.3 −1.4 −3.0 −1.6 −1.3 −1.43.0 gpt-oss:120b −0.2+2.2 −0.9 −0.9+0.8+1.9+2.82.8 aya-expanse:32b −1.8 −2.5 −2.1 −0.9 −1.8 −1.3 −1.52.5 gemma3:1b+0.6+0.7+2.1+0.9+1.3+0.3+0.72.1 qwen3:32b −0.5+0.7+1.5+0.3+1.7 −0.6 −0.71.7 mistral:7b−0.6+1.7+1.7+0.3+1.1+0.1 −1.11.7 aya-expanse:8b+0.4+1.4+0.4+0.5+1.7 −0.8+0.21.7 qwen3:4b+0.2 −1.5+0.4 −0.2+0.3 −0.6 −0.31.5 gemma4:12b+0.9+0.0 −0.7+0.9 −1.2 −1.4 −0.51.4 gemma4:e4b+0.6+1.2+0.7+1.2 −0.5 −0.4+1.21.2 Sonnet 4.6 −0.1+0.6 −0.7 −0.6 −1.1 −0.6 −0.21.1 gemma4:31b+0.4 −0.2 −0.3 −0.1+0.5 −0.2+0.50.5 22 arXiv TemplateA PREPRINT gemma3:270m gemma3:1bgemma3:4b gemma3:12bgemma3:27b gemma4:e2bgemma4:e4b gemma4:12bgemma4:31b llama3.2:3bllama3.1:8b deepseek-r1:7b deepseek-r1:14bdeepseek-r1:32b qwen3:0.6bqwen3:1.7b qwen3:4bqwen3:8b qwen3:14bqwen3:32b aya-expanse:8b aya-expanse:32b gpt-oss:20b gpt-oss:120b granite3.2:8b mistral:latest gpt-5.4-mini opus sonnet Probably to not probably not (neg) Probably to not probably not (pos) Probably to might (neg) Probably to might (pos) Probably to Certain (neg) Probably to Certain (pos) Positive Form Transfer (neg) Positive Form Transfer (pos) Must to probably (neg) Must to probably (pos) Might to Probably (neg) Might to Probably (pos) Distribution over Conjunction (neg) Distribution over Conjunction (pos) Conjunctivitis (neg) Conjunctivitis (pos) Conditional to Comparative (neg) Conditional to Comparative (pos) Complement Transfer (neg) Complement Transfer (pos) Chancy Modus Tollens (neg) Chancy Modus Tollens (pos) Chancy Modus Ponens (neg) Chancy Modus Ponens (pos) Chancy Disjunction Introduction (neg) Chancy Disjunction Introduction (pos) 3828421099840066520000000004086991000528611 8899241001001810010010085893281951001009210021009410010010007870100100 4122100100100991001009982980010110010010010010098981001008698100100 841001001001006898100100561008399510010010010010010010010095100100100100100100 2226090100098100100941001564961007810489610021009510091948298100 6851100100100100100100100100100022850251001001001009510010010010095100100100 289881009410010010010090240018064100100100100100100100100100599100100 48100498100048810006701001001001000100129989100918628719696100 9439610010010010095579946000051100100100100100929894100928999100 18100100989418521001001742990911001007510010084921009272101007574100 2257820000009469519010055200002238601293000 6501029571009498195810000094010099122110010010003651 7170741001001001001001007200031037100100100100921009910010098100100100 3110010010010098100100100809016100100100100100100991001001009599921005989100 648425000000176340989910020000031012280680130 40001229100225710083600401100951950084739218810057 1461010094100100100100722010008910011001004110098100100696100100 7999001195201001005618110010010010465669610070854110082100100 406951001221008810060000006298610010041100981001001610099100 3696814010039901001000495100911003921000741007210096296100100100 02100100100100889510012010600152100100100099991006174100100 96991210010009910010048051006210096459029909678241057100100 061001007110040100100246000720291001001001009610010010010066100100100 10010099100100100100100100188151100100100100191008810010010071391001005746100 920981001001001001009946450000101008810010098959010010008510076 20100579810001510010044626119410098401004885981006298510099100100 Accuracy (%) Frame: truth All inferences 0 20 40 60 80 100 Accuracy % gemma3:270m gemma3:1bgemma3:4b gemma3:12bgemma3:27b gemma4:e2bgemma4:e4b gemma4:12bgemma4:31b llama3.2:3bllama3.1:8b deepseek-r1:7b deepseek-r1:14bdeepseek-r1:32b qwen3:0.6bqwen3:1.7b qwen3:4bqwen3:8b qwen3:14bqwen3:32b aya-expanse:8b aya-expanse:32b gpt-oss:20b gpt-oss:120b granite3.2:8b mistral:latest gpt-5.4-mini opus sonnet Probably to not probably not (neg) Probably to not probably not (pos) Probably to might (neg) Probably to might (pos) Probably to Certain (neg) Probably to Certain (pos) Positive Form Transfer (neg) Positive Form Transfer (pos) Must to probably (neg) Must to probably (pos) Might to Probably (neg) Might to Probably (pos) Distribution over Conjunction (neg) Distribution over Conjunction (pos) Conjunctivitis (neg) Conjunctivitis (pos) Conditional to Comparative (neg) Conditional to Comparative (pos) Complement Transfer (neg) Complement Transfer (pos) Chancy Modus Tollens (neg) Chancy Modus Tollens (pos) Chancy Modus Ponens (neg) Chancy Modus Ponens (pos) Chancy Disjunction Introduction (neg) Chancy Disjunction Introduction (pos) 52103530010010098100100900000500010010769810008910010 908614991000100100100568440989810092100100010094100100100010095100100 5709499100100100100100959000000951009610010010094100100099100100 89100100100100921001001008610020719810010010010010010010010010010010010010099100 1056028100044981009910098995100724259100100010095293904670100 569110010010010010010010010010001989039100100100100100100100100100100100100100 4006196901009910010075610005281001001001001001001001001005290100100 66100121001000129410000811001001001000100419975999891298191100100 9907985601001001001006652000088190100100787992801004906926 2699100991001182100820562694981001004210010056961009610024995671100 910066124000446972984951009850104451726409930034 5200249991007910070000009006099211100929844460 95052100571001001001003283000066981001001001498100991001295100100 48100981001008610010010064921810010010010010010010010010010096100881005381100 52100820200001138881009999850981898519364701422 3400113199815210099600100100334979007683760939951 2823192545981001007170000012100099165100409410009596100 929620659429100100454691001001002205522461009185842410090100100 4106199749481991002150001552890787801009510091318595100 651009168100459610010001095100100100720100071991510095091100100100 013251110001001002471000005299810009972149665100100 100948100100010010010021021009410094694209906076280055100100 00718628100010010030920011099611001004010099321004890100100 10010010010010090991001002224510010010010011100949910010089221001005486100 91098100401001001001009549000041004100410010089899909010099 35100035575351000849194698100624100059110035621100669899 Accuracy (%) Frame: correct All inferences 0 20 40 60 80 100 Accuracy % gemma3:270m gemma3:1bgemma3:4b gemma3:12bgemma3:27b gemma4:e2bgemma4:e4b gemma4:12bgemma4:31b llama3.2:3bllama3.1:8b deepseek-r1:7b deepseek-r1:14bdeepseek-r1:32b qwen3:0.6bqwen3:1.7b qwen3:4bqwen3:8b qwen3:14bqwen3:32b aya-expanse:8b aya-expanse:32b gpt-oss:20b gpt-oss:120b granite3.2:8b mistral:latest gpt-5.4-mini opus sonnet Probably to not probably not (neg) Probably to not probably not (pos) Probably to might (neg) Probably to might (pos) Probably to Certain (neg) Probably to Certain (pos) Positive Form Transfer (neg) Positive Form Transfer (pos) Must to probably (neg) Must to probably (pos) Might to Probably (neg) Might to Probably (pos) Distribution over Conjunction (neg) Distribution over Conjunction (pos) Conjunctivitis (neg) Conjunctivitis (pos) Conditional to Comparative (neg) Conditional to Comparative (pos) Complement Transfer (neg) Complement Transfer (pos) Chancy Modus Tollens (neg) Chancy Modus Tollens (pos) Chancy Modus Ponens (neg) Chancy Modus Ponens (pos) Chancy Disjunction Introduction (neg) Chancy Disjunction Introduction (pos) 20340052000861100045622089020219246800 9499811001003510010010001559910010029821002510098100861002472359999 8000325000400000000000011500000 96100100100100541001001004110046100100100100100100100100100100100100100100100100100 569060000000021969205500006400000020 45499100100100100100100100100085100269910010010010024100100100100100100100100 10092261001009809895000859210011003120318188046280 751000100100004610000641001009689095181988699218685999100 79200275490899960001111680391000017215548428 10100999494260989500421009910010057100750100989939144162499 52100100100260202914030949695550890010094001176262090 5702511001001009610010099004400909100100468100100100691008918 2305400002099350001964004980018619077190 55100711001001008892100012810010098991001009498100100768261981689100 73100966640681300125499979662067922100100204100200 320436961006410010010012013560410038549805941001009100100100 1418100701199808589010078100941001000451948604400 929600090074100007999100100008018135501527019100100 3201036072684099410005284100110057006571004000 401003498100205910010000991008610090100051100579976544100100100 006000100900964950192186010029528189100207800 100100401001000499910000810062969152560950214221000100100 0040081101002800085196010000065990086500 100100711001009298100100046510010099100311007210010095286100992168100 7501007199004418000102195091600022206022000 1910016551000010010000129010010095010042246100919601007910099 Accuracy (%) Frame: follow All inferences 0 20 40 60 80 100 Accuracy % gemma3:270m gemma3:1bgemma3:4b gemma3:12bgemma3:27b gemma4:e2bgemma4:e4b gemma4:12bgemma4:31b llama3.2:3bllama3.1:8b deepseek-r1:7b deepseek-r1:14bdeepseek-r1:32b qwen3:0.6bqwen3:1.7b qwen3:4bqwen3:8b qwen3:14bqwen3:32b aya-expanse:8b aya-expanse:32b gpt-oss:20b gpt-oss:120b granite3.2:8b mistral:latest gpt-5.4-mini opus sonnet Probably to not probably not (neg) Probably to not probably not (pos) Probably to might (neg) Probably to might (pos) Probably to Certain (neg) Probably to Certain (pos) Positive Form Transfer (neg) Positive Form Transfer (pos) Must to probably (neg) Must to probably (pos) Might to Probably (neg) Might to Probably (pos) Distribution over Conjunction (neg) Distribution over Conjunction (pos) Conjunctivitis (neg) Conjunctivitis (pos) Conditional to Comparative (neg) Conditional to Comparative (pos) Complement Transfer (neg) Complement Transfer (pos) Chancy Modus Tollens (neg) Chancy Modus Tollens (pos) Chancy Modus Ponens (neg) Chancy Modus Ponens (pos) Chancy Disjunction Introduction (neg) Chancy Disjunction Introduction (pos) 28229151100051100100460004071608100109294100096968 749859100100110010010016628100100100819510011009510010010057947296100 3114951004811100100899900060167810010014257495100044100100 621001001001001001001001004210026100100100100100100100100100100100100100100100100100 426103871051001000850969921280100100100028764424015491 695510010010010010010010010010001001002599100100100100100100100100100100100100100 31096198100829810099200049810000493498100911009949100100 4410001009802521000082100100999509818885989216849189100 802521891001001001001005500040010051987512176981002956686 6100100100962099980182810098100100210064121009990681496691999 349590110024100052592957110004636949650312247916161 82016410010010098100100950041001006100100610100991002991792 51027850822100100296000392061009909667510008899100 191006610010096989910001622100100100999110077941001006088341002883100 6410064271352610002461100100899808051621001679313578543884 520212886961598100100210815011004454670199999659910099 6186840100109910086600055961000990244347610005710099 809600157696100017898100100000198489541010078100100 36072409835991003820002669100095390100729510064499100 361001499100881100100001001009610034010006498489926030100100100 0010110001001007436100286510074992021532324110098 99100641001000629910002229289100904740990562110510100100 00610968101001007250000390620010008649451002557100100 991006410010095981001000287810010010010001004810010098499100962874100 89071983810096721009640000801000210946658199150100100 81001281001910010002187499100100010081694100656401009110099 Accuracy (%) Frame: valid All inferences 0 20 40 60 80 100 Accuracy % gemma3:270m gemma3:1bgemma3:4b gemma3:12bgemma3:27b gemma4:e2bgemma4:e4b gemma4:12bgemma4:31b llama3.2:3bllama3.1:8b deepseek-r1:7b deepseek-r1:14bdeepseek-r1:32b qwen3:0.6bqwen3:1.7b qwen3:4bqwen3:8b qwen3:14bqwen3:32b aya-expanse:8b aya-expanse:32b gpt-oss:20b gpt-oss:120b granite3.2:8b mistral:latest gpt-5.4-mini opus sonnet Probably to not probably not (neg) Probably to not probably not (pos) Probably to might (neg) Probably to might (pos) Probably to Certain (neg) Probably to Certain (pos) Positive Form Transfer (neg) Positive Form Transfer (pos) Must to probably (neg) Must to probably (pos) Might to Probably (neg) Might to Probably (pos) Distribution over Conjunction (neg) Distribution over Conjunction (pos) Conjunctivitis (neg) Conjunctivitis (pos) Conditional to Comparative (neg) Conditional to Comparative (pos) Complement Transfer (neg) Complement Transfer (pos) Chancy Modus Tollens (neg) Chancy Modus Tollens (pos) Chancy Modus Ponens (neg) Chancy Modus Ponens (pos) Chancy Disjunction Introduction (neg) Chancy Disjunction Introduction (pos) 988100781001001001001009925001351210010010010010024100100100627910099 901601474041001009510015759610010008908931793291089202276 801001001001001098100100291003662951001001001001001001001009210010010098100100 573901984064861000100664961009604092100098357572553980100 24110010010010010010010010010000899210010010010010010010010010092100100100 2010010010010010010010010044640010098910010010010010010010010010098100100100 5510090100100648921002466510010010010035100201001001001001004010010099100 76982445510088801005090011261988929957210010069512511 59999688542544100991099217080100100751001001001001008498261008975100 5711001100143208699105695100955651046010028540226670100 2205729199100511009171000009908049509889840566410 2047141004842969410042500103111791001009251989810098545012 8999100100100929310010084992291991001001001001001001001009810010010066100100 77931008187372806879579872100180510035228846113148 26002229839511005510000010045015004344530949976 211839056595114400000111000100241100864445019658 94859618099601001003276701001001009009549610010059918110091100100 351610032341510010010046200112610061100100991009510010012571100 419292641004299100100142999100941001510100186100961001004199100100100 1999810010010010010010051920761000910010010010022100100100100359910070 10070311001000941001006921810042100100699159909190960056100100 01001001001001001001001001889001009910010010010052100100100100100100100100 1001001001001009290100100599552100100100100751009610010010096601001007990100 68142490491001009098800000069969610010058498100026922 50100811610000100100095118921009962100981009110075961210010099100 Accuracy (%) Frame: direct All inferences 0 20 40 60 80 100 Accuracy % Figure 3: Accuracy (%) for every (template×negation) cell by model, shown separately for each of the five question forms. The same logical content swings from near-zero to near-ceiling across question forms, visualizing the question- form instability summarized in Table 5. 23