Paper deep dive
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models
Kesheng Chen, Yamin Hu, Qi Zhou, Zhenqian Zhu, Wenjian Luo
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 98%
Last extracted: 3/31/2026, 2:10:04 AM
Summary
CDH-Bench is a diagnostic benchmark designed to evaluate 'commonsense-driven hallucination' (CDH) in vision-language models (VLMs). It uses 300 paired counterfactual-commonsense images to test if models prioritize visual evidence or learned commonsense priors when they conflict. The benchmark covers counting, relational, and attribute anomalies, and introduces metrics like Counterfactual Accuracy (CF-Acc), Commonsense Accuracy (CS-Acc), Counterfactual Accuracy Drop (CFAD), Commonsense Collapse Rate (CCR), and Relative Prior Dependency (RPD) to quantify model reliability.
Entities (5)
Relation Signals (3)
CDH-Bench â evaluates â Vision-Language Models
confidence 100% · CDH-Bench, a benchmark designed to create explicit visual evidenceâcommonsense conflicts... We evaluate frontier VLMs
CDH-Bench â includesmetric â Counterfactual Accuracy
confidence 100% · report metrics including Counterfactual Accuracy (CF-Acc)
CDH-Bench â measures â Commonsense-driven hallucination
confidence 100% · To evaluate it [CDH], we introduce CDH-Bench
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-language models (VLMs) achieve strong performance on many benchmarks, yet a basic reliability question remains underexplored: when visual evidence conflicts with commonsense, do models follow what is shown or what commonsense suggests? A characteristic failure in this setting is that the model overrides visual evidence and outputs the commonsense alternative. We term this phenomenon \textbf{commonsense-driven hallucination} (CDH). To evaluate it, we introduce \textbf{CDH-Bench}, a benchmark designed to create explicit \textbf{visual evidence--commonsense conflicts}. CDH-Bench covers three dimensions: \textit{counting anomalies}, \textit{relational anomalies}, and \textit{attribute anomalies}. We evaluate frontier VLMs under \textit{binary Question Answering (QA)} and \textit{multiple-choice QA}, and report metrics including \textit{Counterfactual Accuracy} (CF-Acc), \textit{Commonsense Accuracy} (CS-Acc), \textit{Counterfactual Accuracy Drop} (CFAD), \textit{Commonsense Collapse Rate} (CCR), and \textit{Relative Prior Dependency} (RPD). Results show that even strong models remain vulnerable to prior-driven normalization under visual evidence--commonsense conflict. CDH-Bench provides a controlled diagnostic of visual fidelity under visual evidence--commonsense conflict.
Tags
Links
- Source: https://arxiv.org/abs/2603.27982v1
- Canonical: https://arxiv.org/abs/2603.27982v1
Trouble viewing inline? Open PDF directly â
Full Text
62,932 characters extracted from source content.
Expand or collapse full text
CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models Kesheng Chen , Yamin Hu , Qi Zhou , Zhenqian Zhu , Wenjian Luo Guangdong Provincial Key Laboratory of Novel Security Intelligence Technologies, Institute of Cyberspace Security, School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, China huyamin@hit.edu.cn No. The butterfly has 4 wings. No. The hand has 5 fingers. C. The mug has 1 handle. The apple has 1 stem Driven by Counting-CDH Animal PartsAnimal Parts Does the butterfly have 4 wings and not 8 wings ? Does the butterfly have 4 wings and not 8 wings ? Does the hand have 5 fingers and not 6 fingers? Does the hand have 5 fingers and not 6 fingers? How many handles does this mug have? - A. 7 handles - B. 2 handles - C. 1 handle -D ... How many handles does this mug have? - A. 7 handles - B. 2 handles - C. 1 handle -D ... Does the apple have 1 stem and not 3 stems? Does the apple have 1 stem and not 3 stems? No. The brick wall is opaque. Is the brick wall opaque and not transparent? Is the brick wall opaque and not transparent? The flames are orange-yellow. Are the flames orange- yellow and not black? Are the flames orange- yellow and not black? B. The glass is brittle. What is the property of this glass? - A. Flexible - B. Brittle - C. Porous - D. Soft What is the property of this glass? - A. Flexible - B. Brittle - C. Porous - D. Soft The refrigerator cools food. Does the refrigerator cool food and not heat food? Does the refrigerator cool food and not heat food? No. The cat is chasing the mouse. Is the cat chasing the mouse and not the mouse chasing the cat? Is the cat chasing the mouse and not the mouse chasing the cat? A. Scissors For the action 'cutting paper', which tool is being used? Choose one. A. Scissors B. A phone C. A spoon D. A pencil For the action 'cutting paper', which tool is being used? Choose one. A. Scissors B. A phone C. A spoon D. A pencil No. The elephant is much Larger than the teacup. Is the elephant much larger than the teacup and not about the same size as the teacup? Is the elephant much larger than the teacup and not about the same size as the teacup? B. The desk is on the floor. Choose the chair's location. A. Floating in the air B. On the floor C. Under the floor -D... Choose the chair's location. A. Floating in the air B. On the floor C. Under the floor -D... Driven by Realation-CDH Driven by Attribute-CDH Body PartsBody PartsEveryDay PartsEveryDay PartsPlant StructurePlant Structure Physical StatePhysical StateTemperatureTemperatureColorColor Animal BehaviorAnimal BehaviorObject FunctionObject FunctionSize ScaleSize ScaleSpatialSpatialSpatial LuminescenceLuminescence Abstract Vision-language models (VLMs) achieve strong perfor- mance on many benchmarks, yet a basic reliability ques- tion remains underexplored: when visual evidence con- flicts with commonsense, do models follow what is shown or what commonsense suggests?A characteristic failure in this setting is that the model overrides visual evidence and outputs the commonsense alternative.We term this phenomenon commonsense-driven hallucination (CDH). To evaluate it, we introduce CDH-Bench, a benchmark designed to create explicit visual evidenceâcommonsense conflicts. CDH-Bench covers three dimensions: counting anomalies, relational anomalies, and attribute anomalies. We evaluate frontier VLMs under binary Question Answer- ing (QA) and multiple-choice QA, and report metrics includ- ing Counterfactual Accuracy (CF-Acc), Commonsense Ac- curacy (CS-Acc), Counterfactual Accuracy Drop (CFAD), Commonsense Collapse Rate (CCR), and Relative Prior De- pendency (RPD). Results show that even strong models re- main vulnerable to prior-driven normalization under visual evidenceâcommonsense conflict.CDH-Bench provides a controlled diagnostic of visual fidelity under visual evidenceâ commonsense conflict. Data are available at:§ GitHub and* Hugging Face. 1 Introduction âI see what I believe, not what I see.â This classic human bias increasingly appears in modern VLMs. These models excel at standard image captioning and Visual Question & Answering (VQA) [ 2; 6; 10 ] , but reliability hinges on a harder scenario: what happens when visual evidence conflicts with commonsense? Consider an example in medical imaging: an X-ray re- veals an anatomical anomaly, such as a hand with six fin- arXiv:2603.27982v1 [cs.CV] 30 Mar 2026 gers. A competent radiologist should describe what the im- age actually shows (evidence) rather than what is normally expected (commonsense). However, when the same image is presented to a VLM and asked, âHow many fingers does the hand have?â the model may answer âfive.â This failure does not reflect difficulty in perception, but rather that common- sense overrides visual evidence. This is not the usual âobject hallucinationâ studied in prior benchmarks [ 9; 13; 5 ] , where models fabricate objects ab- sent from the image. Instead, we study a different hallucina- tion mode: the relevant visual evidence is present (and often salient), but the model still defaults to what is typically true. In other words, when evidence and commonsense conflict, the output snaps back to the commonsense alternative rather than reflecting the image. We call this commonsense-driven hallucination (CDH): a prior-driven normalization in which models report the commonsense alternative instead of the vi- sual evidence. A useful perspective is that VLMs implicitly arbitrate between competing signals: a visual perception channel grounded in the image and a commonsense-prior chan- nel that favors high-probability completions learned from a commonsense-dominated distribution. When these channels conflict, generation can behave like semantic smoothing: gen- uine but commonsense-violating evidence is treated as noise and smoothed toward a commonsense semantic prior (e.g., six fingers to five; blue banana to yellow). This perspective predicts the experimental result we observe: CDH persists even in strong frontier models and becomes especially visible when the commonsense alternative is explicitly presented. CDH matters most where anomalies matter: medical imag- ing, quality inspection, scientific discovery, and forensics. A system that ânormalizesâ unusual evidence can be dan- gerously confident while wrong.Yet current evaluation paradigms are poorly suited to detect CDH. Most bench- marks use commonsense-consistent imagery, so visual evi- dence and commonsense priors typically agree [ 9; 13; 14; 5 ] . Hallucination benchmarks often suffer from ambiguous error attribution: when a model answers incorrectly, it is un- clear whether it guessed from priors or truly misperceived the image. We therefore need an evaluation that creates conflict on purpose and supports unambiguous error attribution. We introduce CDH-Bench, a benchmark designed to sys- tematically evaluate visual fidelity under visual evidenceâ commonsense conflict. Our contributions are: (1) Formalization of CDH. We define commonsense- driven hallucination (CDH) as a distinct reliability failure pattern in which models override visual evidence with learned priors, distinguishing it from standard object fabrication. (2) The CDH-Bench. We construct a diagnostic bench- mark consisting of 300 paired counterfactualâcommonsense images, designed specifically for controlled comparison and unambiguous error attribution. (3) Systematic evaluation and scaling analysis. We eval- uate frontier VLMs under binary QA and multiple-choice QA settings, report metrics including CF-Acc, CS-Acc, CFAD, CCR, and RPD, and conduct a controlled analysis of Qwen3- VL Instruct and Thinking variants [ 16; 3 ] to examine how scale and reasoning style affect CDH susceptibility. 2 Related Work General Ability Evaluation for Vision-Language Models. Standard VLM benchmarks evaluate broad multimodal capa- bilities. VQAv2 [ 2 ] and GQA [ 6 ] test visual question answer- ing, which requires visual evidence from the input image. OK-VQA [ 11 ] requires external knowledge integration, while comprehensive benchmarks such as SEED-Bench [ 8 ] and MMBench [ 10 ] assess multiple dimensions of visual under- standing. These benchmarks have driven substantial progress, with recent models approaching human performance on many tasks. Several benchmarks probe specific perceptual abili- ties: TallyQA [ 1 ] evaluates counting, CLEVR [ 7 ] tests com- positional reasoning and spatial understanding in synthetic scenes, and VAW [ 12 ] assesses attribute recognition across a large inventory of visual concepts. These datasets are valu- able for measuring fine-grained perception, but they typically test perception in settings where commonsense priors support the task rather than compete with it. However, most of them operate in regimes where visual evidence and commonsense priors are naturally aligned. Hallucination Evaluation for VLMs.Hallucination benchmarks measure how models generate content unsup- ported by visual input. POPE [ 9 ] uses binary existence ques- tions to evaluate object hallucination, while CHAIR [ 13 ] measures hallucinated objects in image captions. MMHal- Bench [ 14 ] and HallusionBench [ 5 ] provide broader evalua- tions across spatial relations and attributes. However, most prior work primarily targets fabrication hallucination. Our focus is different: we study normalization hallucination, where present visual evidence is overridden by the common- sense alternative. Evaluation with Counterfactual and Compositional Reasoning. Counterfactual VQA [ 4 ] synthesizes counterfac- tual questions to reduce language bias during training, typi- cally by altering the question while keeping the image fixed so that models cannot rely on superficial linguistic correla- tions alone. Winoground [ 15 ] probes visio-linguistic com- positionality with a carefully hand-curated setup of two im- ages and two captions, where the captions contain the same words but in different orders, requiring models to correctly align each image with its caption. These benchmarks are re- lated in spirit, but they differ from our setting. CDH-Bench constructs counterfactual visual content itself, making the visual evidenceâcommonsense conflict a direct property of the image rather than of the linguistic query alone. 3 Benchmark Design 3.1 Design Motivations CDH-Bench is motivated by critical gaps in existing evalua- tion paradigms that underestimate a fundamental failure pat- tern in VLMs. The evaluation blind spot under commonsense- consistent distributions.Standard benchmarks such as VQAv2 [ 2 ] , GQA [ 6 ] , SEED-Bench [ 8 ] , and MMBench [ 10 ] are largely built on natural-image distributions in which visual evidence and commonsense priors usually agree. As a result, benchmark success does not clearly distinguish gen- uine visual grounding from reliance on statistical regularities that happen to match the test distribution. Fine-grained [ 7; 1 ] perception benchmarks test many of the same underlying skills, but typically under settings where commonsense priors are helpful rather than adversarial. This creates an evaluation blind spot: standard benchmarks effectively ask, Can the model recognize what is there under commonsense-consistent conditions? In contrast, CDH-Bench asks a more diagnostic question: Can the model still report what is there when visual evidence conflicts with commonsense? Fabrication vs. normalization hallucination. Existing hallucination benchmarks [ 9; 13 ] mainly target fabrication hallucination, where the model mentions entities, attributes, or relations that are absent from the image. CDH-Bench tar- gets a different failure pattern: the model ignores present vi- sual evidence and outputs the commonsense alternative as if the anomaly were noise. A six-fingered hand, for example, is not visually absent; it is visually present but conflicts with commonsense. CDH-Bench is designed to expose this nor- malization-style failure directly. The error attribution problem.Many hallucination benchmarks also suffer from ambiguous error attribution. When a model answers incorrectly, the failure may come from weak perception, prior-based guessing, or confusion in- duced by the task format. Without a matched control con- dition, these sources are difficult to disentangle. By pairing each counterfactual image with a commonsense counterpart that preserves the overall scene while changing only the tar- get anomaly, CDH-Bench enables cleaner attribution of prior- driven failure under visual evidenceâcommonsense conflict. 3.2 Design Principles Building on these motivations, CDH-Bench is designed around four principles: Paired data design for unambiguous attribution. Each counterfactual image is paired with a commonsense counter- part that differs only in the target anomaly. This controlled comparison enables stronger causal interpretation: if a model succeeds on the commonsense image but fails on the coun- terfactual image, the error is best explained as a prior-driven override rather than a generic incapacity. Multi-dimensional coverage. We cover three fundamen- tal dimensions of commonsense: quantity, relation, and at- tribute. This enables us to analyze how prior-driven normal- ization behaves across different semantic aspects. Task-controlled evaluation. We employ two task formats to probe CDH under different levels of cognitive competition. In binary QA, the question text explicitly contrasts the coun- terfactual visual evidence with its commonsense alternative, directly measuring whether the model can prioritize anoma- lous evidence within a minimal output space. Multiple-choice QA introduces a higher level of difficulty by explicitly pre- senting the commonsense alternative as a competitive answer choice, which is particularly effective for diagnosing system- atic collapse to learned priors. Fine-grained error taxonomy. We classify errors into commonsense errors and other errors. This distinction, quan- tified through CCR, provides a sharper diagnostic signal than accuracy alone, where direct answer competition makes the interpretation most transparent. 3.3 Counterfactual Dimensions We construct 600 images, organized as 300 counterfactualâ commonsense pairs. Figure 1 shows the taxonomy and distri- bution of CDH-Bench. Figure 1: Taxonomy and distribution of CDH-Bench. Counting anomalies (100 pairs). This dimension focuses on unusual but visually unambiguous quantities that conflict with biological or structural commonsense, spanning cases such as six-fingered hands, animals with extra legs, and ev- eryday objects with atypical component counts. Relational anomalies (100 pairs). This dimension fo- cuses on inverted relations between entities that violate typi- cal behavioral or physical scripts, spanning behavioral rever- sals (e.g., mice hunting cats), spatial inversions (e.g., upside- down houses), size reversals, and causal reversals. Attribute anomalies (100 pairs).This dimension fo- cuses on objects possessing properties that violate fundamen- tal physical expectations, spanning color anomalies (e.g., blue bananas), material anomalies (e.g., transparent wood), tem- perature conflicts, and atypical physical states. 3.4 Image Generation and Quality Control We generate images using strong text-to-image models with prompts engineered to counteract the generatorâs common- sense bias. In particular, prompts are designed to make the in- tended anomaly necessary rather than optional by combining: (1) explicit structural constraints (e.g., exact counts, left-to- right enumeration, symmetry/asymmetry requirements, and positional anchors such as âcenterâ or âaboveâ), (2) decom- posed anatomical/relational descriptions (e.g., specifying per-part attributes like nails/knuckles/iris detail and clarify- ing spatial relations among parts), and (3) physical support cues (e.g., viewpoint, lens/shot type, lighting, depth-of-field, and photorealistic texture) to stabilize realism and reduce un- intended artifacts. For each anomaly type, we construct a matched common- sense counterpart by keeping composition, viewpoint, subject identity, background, and photographic style as consistent as possible, while modifying only the minimal attribute required by the counterfactual (typically a single count-based change). binary QA: Does the mouth have 2 rows of teeth and not 3 rows? multi-choice QA: How many rows of teeth are visible? Options: A. 3 rows, B. 4 rows, C. 2 rows, D. 1 row Category: Counting - Body Parts binary QA: Does the chicken have 2 wings and not 6 wings? multi-choice QA: How many wings does this chicken have? Options: A. 4 wings, B. 6 wings, C. 2 wings, D. 8 wings Category: Counting - Animal Parts binary QA: Is the whale swimming in the ocean and not in a fish tank? multi-choice QA: What is unusual about the whaleâs environment? Options: A. ..., B. The whale is contained in a small fish tank, C. The whale is swimming in the ocean normally, D. ... Category: Relational - Size Scale binary QA: Is the foam light and airy and not dense and heavy? multi-choice QA: What is the density of this foam? Options: A. Dense and heavy, B. Medium, C. Solid, D. Light and airy Category: Attribute - Material Figure 2: CDH-Bench showcase: examples from diverse dimen- sions. Each subfigure displays a counterfactual (CF) image on the left and its commonsense (CS) counterpart on the right, together with the associated binary QA and multi-choice QA evaluation de- tails. We also add mild redundancy in the text (e.g., stating âexactly kâ and explicitly listing parts) to improve controllability and make the anomaly salient under realistic rendering. Quality control includes: (i) human verification to con- firm that the anomaly is unambiguous and visually central, and that the paired commonsense image is normal with no additional unexpected anomalies; (i) pair-level matching to ensure the counterfactual and commonsense images are closely aligned in pose, illumination, and overall style. 3.5 Task Formats and Evaluation We evaluate each image pair under two task formats with format-specific scoring protocols. Although the data is or- ganized into counterfactualâcommonsense pairs, each image is processed independently during inference: the model re- ceives only a single image as input (either a counterfactual image or a commonsense image) without any access to its paired counterpart. Binary QA. This format uses compositional yes/no ques- tions that explicitly contrast the counterfactual target with its commonsense alternative (e.g., âDoes the mouth have 2 rows of teeth and not 3 rows?â). Crucially, this design enables a unified query for both counterfactual (CF) and commonsense (CS) images within each pair. By maintaining an identical question text across both conditions, we eliminate potential linguistic biases or prompt-induced variances that could oth- erwise confound the comparison. This ensures that perfor- mance gaps and relative metrics are strictly attributable to the visualâcommonsense conflict in the imagery rather than dif- ferences in the textual phrasing, providing a more rigorous measure of visual fidelity. Multiple-Choice QA. Each question contains four answer options: one correct answer for the current counterfactual im- age, one commonsense alternative, and two additional plau- sible distractors. The option order is randomly shuffled for each instance to avoid positional bias. This format is espe- cially diagnostic because it makes the commonsense alterna- tive explicit and places it in direct competition with the visu- ally correct answer. For both formats, we evaluate performance separately on CF and CS images and then derive metrics that characterize robustness under visual evidenceâcommonsense conflict as well as some error attribution metrics across paired instances. Each of the 300 image pairs is evaluated under two image conditions (counterfactual and commonsense) and two task formats, yielding 300Ă 2Ă 2 = 1,200 evaluated instances in total. 4 Evaluation Metrics We design metrics around three concrete questions that stan- dard accuracy alone cannot disentangle. For each question, we introduce the metric(s) that directly answer it. Q1: How well does a model perform under visual evidenceâcommonsense conflict? Counterfactual Accuracy (CF-Acc). CF-Acc is the ac- curacy on counterfactual (CF) images, and is our primary measure of visual fidelity under conflict. It directly answers Q1 by measuring whether the model follows visual evidence when that evidence conflicts with commonsense. Q2: How large is the collapse relative to normal- condition capability? Commonsense Accuracy (CS-Acc). CS-Acc is the accuracy on matched commonsense (CS) counterparts. It measures the modelâs general multimodal capabilities, specifically its abil- ity to perform basic visual observation and standard common- sense reasoning in settings where visual evidence and priors are naturally aligned. This serves as a necessary baseline and capability control under non-conflicting conditions. Counterfactual Accuracy Drop (CFAD). We define the ac- curacy drop from the commonsense condition to the counter- factual condition as CFAD = CS-Accâ CF-Acc. CFAD answers Q2 by quantifying the absolute degradation under visual evidenceâcommonsense conflict (larger positive CFAD indicates stronger degradation). Relative Prior Dependency (RPD). To compare collapse strength across models with different baselines, we normalize CFAD by CS-Acc: RPD = CFAD CS-Acc . RPD also answers Q2, but in relative terms: of what the model can do when visual evidence and commonsense agree, how much is lost when they conflict? Q3: When the model fails on counterfactual images, does it fail specifically by reverting to the commonsense alternative? Error classification. When a model produces an incorrect answer on a counterfactual image, we classify the error as (1) a commonsense error if the response matches the common- sense alternative, and (2) an other error otherwise. Commonsense Collapse Rate (CCR). CCR measures the fraction of CF errors that collapse specifically to the com- monsense alternative: CCR = # Commonsense-Aligned Errors on CF Images # Total Errors on CF Images . CCR answers Q3 by measuring how often counterfactual fail- ures are specifically commonsense-aligned rather than arbi- trary. In binary QA, this corresponds to predicting the re- sponse consistent with the CS counterpart on a CF image. In multiple-choice QA, this corresponds to selecting the ex- plicit commonsense alternative among the answer options. Although CCR is definable in both formats, we report it in the main tables only for multiple-choice QA, where direct answer competition makes the interpretation more transparent. 5 Experiments 5.1 Models Evaluated We evaluate eight frontier multimodal systems in the main comparison: gemini-3.1-pro-preview, gemini-3.1-flash-lite- preview, doubao-seed-1.8-251228, qwen3.5-plus, kimi-k2.5, gemini-2.5-flash-all, gpt-5.4-mini, and gpt-5.4. For controlled family analysis, we further evaluate Qwen3- VL variants and divide them into two groups: Instruct and Thinking [ 16; 3 ] . This setup allows us to compare not only performance across scales, but also the effect of reasoning style within a shared model family. Unless stated otherwise, all models are evaluated under a unified protocol: identical prompts per task format, fixed de- coding settings where applicable, and deterministic answer extraction for multiple-choice QA. 5.2 Evaluation Granularity and Aggregation We report three levels of analysis. Task-level aggregation. For the main comparison, we report two task-specific views: QA (binary QA) and MC (multiple-choice QA). For each task, we provide both overall results and category-level breakdowns over Attribute Anoma- lies, Counting Anomalies, and Relational Anomalies. Subcategory-level diagnosis. Beyond the primary cate- gories, we further analyze fine-grained subcategories to iden- tify which anomaly types remain consistently difficult across models. This analysis is especially useful because category- level averages can hide concentrated bottlenecks such as body-part counting. Family-level scaling and reasoning analysis. For the Qwen3-VL study, we merge QA and MC into a unified com- parison table and report metrics grouped by task and over- all aggregation. This presentation makes it easier to com- pare scale effects and reasoning-style effects within a shared model family. 5.3 Main Results Table 1 presents the category-level results on CDH-Bench, while Figures 3â6 visualize model performance in terms of counterfactual accuracy and the accuracy drop from com- monsense to counterfactual conditions. Figure 3: Model performance comparison on counterfactual accu- racy (CF-Acc) across different subcategories in the binary QA task. Figure 4: Counterfactual Accuracy Drop (CFAD) across subcate- gories in the binary QA task. Figure 5: Model performance comparison on counterfactual accu- racy (CF-Acc) across different subcategories in the multiple-choice QA task. CDH induces a consistent and substantial robustness gap across models and task formats. For nearly all evalu- ated models, accuracy on counterfactual images is lower than Figure 6: Counterfactual Accuracy Drop (CFAD) across subcate- gories in the multiple-choice QA task. on matched commonsense controls, and this pattern holds in both binary QA and multiple-choice QA. Concretely, 7 out of 8 models show lower overall CF-Acc than CS-Acc in both settings; the only partial exception is gemini-3.1-pro- preview, which exhibits a slightly negative CFAD in QA but still shows a 7.33% drop in MC. Averaged across models, the overall CFAD is 16.39% in QA and increases to 25.20% in MC, while mean CF-Acc decreases from 75.24% to 66.90%. These results suggest that CDH is not an isolated failure mode of a few weak systems, but a systematic robustness challenge that persists across model families. Multiple-choice QA further amplifies prior-driven fail- ures. Compared with binary QA, MC generally yields both lower CF-Acc and larger CFAD, indicating that when coun- terfactual evidence must directly compete with a plausible commonsense alternative, models are more likely to revert to prior-consistent answers. On average, moving from QA to MC lowers CF-Acc by 8.34 points and increases CFAD by 8.81 points. This trend is especially pronounced for mod- els such as kimi-k2.5 and gpt-5.4-mini, which see their CFAD values rise from 17.67% to 29.33% and from 29.33% to 41.00%, respectively. Taken together, the QAâMC contrast suggests that CDH is not only a robustness issue under uncer- tainty, but also a comparative reasoning failure that emerges when image-grounded counterfactual evidence must override a strong commonsense prior. Model-level comparisons further suggest that smaller degradation does not necessarily imply stronger counter- factual understanding, and that CDH-Bench exposes a structured failure landscape beyond aggregate accuracy. For instance, GPT-5.4 is weakest on Counting / Body Parts in both QA and MC, whereas doubao-seed-1.8-251228 remains particularly fragile on Temperature. At the same time, relative robustness metrics such as CFAD should be interpreted jointly with performance on matched commonsense controls: among the stronger baselines, gemini-3.1-pro-preview exhibits the smallest overall CFAD in QA, but it also attains the lowest overall CS-Acc among the frontier models evaluated (e.g., 87.67% in QA). Thus, its smaller degradation may reflect weaker commonsense-to-counterfactual collapse rather than clear absolute superiority in counterfactual visual grounding. A plausible explanation is that, in our binary QA setting, the question text simultaneously contains a strong common- sense prior (e.g., ânormalâ expectations) and an explicit coun- terfactual claim grounded in the image. Models that are more conservative or refusal-prone under potential inconsistency may answer ânoâ (or abstain) more often even on common- sense controls, lowering CS-Acc and mechanically shrinking the measured CFAD. Overall, CDH-Bench reveals a pervasive and funda- mental vulnerability across current frontier VLMs, where even the strongest systems fail to prioritize visual evi- dence over entrenched priors. The primary importance of this benchmark lies not merely in reporting lower accuracy scores, but in providing a rigorous diagnostic framework that aggregate metrics alone cannot offer. By decoupling prior- driven collapse from generic perception errors through its paired counterfactualâcommonsense design, CDH-Bench en- ables a structured analysis of failure modes across diverse semantic dimensions. Furthermore, the introduction of spe- cialized metrics such as CFAD, CCR, and RPD allows for unambiguous error attribution, revealing that model failures are often systematic normalizations toward the commonsense prototype rather than random noise. These findings establish CDH-Bench as a critical tool for identifying the hidden relia- bility gaps in multimodal systems and highlight the necessity of evaluating visual fidelity under explicit visual evidenceâ commonsense conflict to guide the development of more ro- bust vision-language models. 5.4 Scale Study: Do Scale and Reasoning Style Mitigate CDH? Motivation. A key question is whether model scale and rea- soning style mitigate commonsense-driven hallucination un- der counterfactual visual evidence. A natural hypothesis is that larger models should achieve stronger visual discrimina- tion, while reasoning-oriented variants may better preserve fi- delity when visual evidence conflicts with learned priors. We test this hypothesis within the Qwen3-VL family. Models and reporting. We study two Qwen3-VL se- ries [ 16; 3 ] : Instruct and Thinking. Table 2 reports binary QA, multiple-choice QA, and overall results in a unified lay- out, and Figures 7 and 8 visualize the scaling trends for MC and QA, respectively. Scale helps clearly in the Instruct series, but remains less stable in the Thinking series. Within the Instruct se- ries, overall CF-Acc rises monotonically from 36.17% (2B) to 43.17% (4B), 43.50% (8B), and 50.83% (32B). Thus, larger scale is generally beneficial for the Instruct models, although the gains are modest between some intermediate sizes. By contrast, the dense Thinking series is less mono- tonic at smaller and medium scales: overall CF-Acc changes from 53.17% (2B) to 51.17% (4B) and 51.50% (8B), before improving substantially to 58.41% at 32B. Therefore, scale alone does not guarantee a smooth reduction in CDH suscep- tibility, especially within the Thinking line. Reasoning style improves accuracy more consistently than collapse-oriented error metrics.Although Think- ing models consistently raise CF-Acc, their advantages on collapse-oriented metrics such as CCR and RPD are less uni- form. For example, at 4B and 32B, Thinking slightly reduces Table 1: Main results on CDH-Bench. Higher CF-Acc and CS-Acc are better; lower CFAD, CCR, and RPD are better. To emphasize failure analysis rather than leaderboard ranking, selected notably weak results are marked in red. CCR is reported only for Multiple- choice QA in the main table. All values are percentages. Gray rows denote Overall results. ModelCat. Binary QAMultiple-choice QA CFâCSâCFADâRPDâCFâCSâCFADâ CCRâRPDâ gemini-3.1-pro-preview Overall90.33%87.67%-2.67%-3.04%81.67%89.00%7.33%58.18%8.24% Attr.85.00% 83.00%-2.00%-2.41%84.00% 93.00%9.00%81.25%9.68% Count.96.00% 87.00%-9.00% -10.34% 79.00% 80.00%1.00%42.86%1.25% Rel.90.00% 93.00%3.00%3.23%82.00% 94.00% 12.00% 55.56% 12.77% gemini-3.1-flash-lite-preview Overall83.67%91.33%7.67%8.39%71.67%94.33%22.67%71.76%24.03% Attr.78.00% 86.00%8.00%9.30%69.00% 93.00% 24.00% 77.42% 25.81% Count.94.00% 93.00%-1.00%-1.08%69.00% 94.00% 25.00% 58.06% 26.60% Rel.79.00% 95.00% 16.00%16.84% 77.00% 96.00% 19.00% 82.61% 19.79% doubao-seed-1.8-251228 Overall76.67%91.67%15.00%16.36%64.67%91.67%27.00%77.36%29.45% Attr.70.00% 86.00% 16.00%18.60% 54.00% 94.00% 40.00% 86.96% 42.55% Count.80.00% 93.00% 13.00%13.98% 71.00% 88.00% 17.00% 65.52% 19.32% Rel.80.00% 96.00% 16.00%16.67% 69.00% 93.00% 24.00% 74.19% 25.81% qwen3.5-plus Overall75.87%93.98%18.11%19.27%74.37%95.55%21.18%81.69%22.17% Attr.75.51% 88.00% 12.49%14.19% 75.51% 93.94% 18.43% 95.83% 19.62% Count.77.78% 96.97% 19.19%19.79% 74.70% 95.88% 21.18% 66.67% 22.09% Rel.74.49% 97.00% 22.51%23.21% 72.92% 96.88% 23.96% 80.77% 24.73% kimi-k2.5 Overall75.67%93.33%17.67%18.93%64.67%94.00%29.33%85.85%31.21% Attr.71.00% 89.00% 18.00%20.22% 69.00% 94.00% 25.00% 96.77% 26.60% Count.76.00% 99.00% 23.00%23.23% 47.00% 93.00% 46.00% 79.25% 49.46% Rel.80.00% 92.00% 12.00%13.04% 78.00% 95.00% 17.00% 86.36% 17.89% gemini-2.5-flash-all Overall72.00%88.33%16.33%18.49%66.00%87.33%21.33%74.51%24.43% Attr.62.00% 80.00% 18.00%22.50% 58.00% 92.00% 34.00% 83.33% 36.96% Count.81.00% 92.00% 11.00%11.96% 65.00% 76.00% 11.00% 74.29% 14.47% Rel.73.00% 93.00% 20.00%21.51% 75.00% 94.00% 19.00% 60.00% 20.21% gpt-5.4-mini Overall64.00%93.33%29.33%31.43%53.33%94.33%41.00%78.57%43.46% Attr.65.00% 90.00% 25.00%27.78% 53.00% 94.00% 41.00% 91.49% 43.62% Count.65.00% 95.00% 30.00%31.58% 43.00% 91.00% 48.00% 66.67% 52.75% Rel.62.00% 95.00% 33.00%34.74% 64.00% 98.00% 34.00% 80.56% 34.69% gpt-5.4 Overall63.67%93.33%29.67%31.79%58.86%94.65%35.79%73.17%37.81% Attr.68.00% 86.00% 18.00%20.93% 58.00% 97.00% 39.00% 88.10% 40.21% Count.61.00% 97.00% 36.00%37.11% 48.00% 90.00% 42.00% 53.85% 46.67% Rel.62.00% 97.00% 35.00%36.08% 70.71% 96.97% 26.26% 86.21% 27.08% Attr. = Attribute; Count. = Counting; Rel. = Relational. CCR is reported only for multiple-choice QA in the main table. Red numbers indicate selected notably weak results to facilitate failure analysis. MC CCR relative to Instruct, but at 8B it yields a higher CCR (89.33% vs. 78.74%). Similarly, improvements in RPD are clearer at some scales than others. This suggests that reasoning-oriented generation reliably improves success rates under conflict, but does not uniformly change the exact form of model failures. 6 Conclusion We formalize commonsense-driven hallucination (CDH) and introduce CDH-Bench, a paired counterfactual-image benchmark that makes prior-driven normalization measur- able and attributable. Across frontier VLMs, we observe large and systematic collapses from commonsense-consistent to counterfactual imagery, revealing a pervasive and funda- mental vulnerability where even the strongest systems fail to prioritize visual evidence over entrenched priors. By de- coupling prior-driven collapse from generic perception errors through specialized metrics such as CFAD, CCR, and RPD, we demonstrate that model failures are often systematic nor- malizations toward commonsense priors rather than random noise. Our experiments show that CDH is not a uniform phe- nomenon. At the category level, counting anomalies are the most persistent source of degradation, especially in multiple- choice QA where explicit competition with commonsense alternatives triggers stronger collapse. At the subcategory Table 2: Results of the Qwen3-VL series on CDH-Bench. All values are reported as percentages. Higher CF-Acc and CS-Acc are better; lower CFAD, CCR, and RPD are better. We do not highlight best values; instead, selected notably weak results are marked in red. The Overall columns are shaded for readability. Model Binary QAMultiple-choice QAOverall CF (%)CS (%)CFAD (%)RPD (%)CF (%)CS (%)CFAD (%)CCR (%)RPD (%)CF (%)CS (%)CFAD (%)CCR (%)RPD (%) Qwen3-VL-2B-Instruct40.3387.6747.3353.9932.0074.0042.0073.5356.7636.1780.8344.6786.7655.37 Qwen3-VL-4B-Instruct52.0091.6739.6743.2734.3389.3355.0084.7761.5743.1790.5047.3392.3952.42 Qwen3-VL-8B-Instruct45.0093.6748.6751.9642.0093.6751.6778.7455.1643.5093.6750.1789.3753.56 Qwen3-VL-32B-Instruct55.6793.3337.6740.3646.0092.6746.6784.5750.36 50.8393.0042.1792.2845.36 Qwen3-VL-2B-Thinking58.0091.3333.3336.5048.3393.0044.6777.4248.0353.1792.1739.0088.3142.26 Qwen3-VL-4B-Thinking55.0092.6737.6740.6547.3395.0047.6784.1850.1851.1793.8342.6792.0945.41 Qwen3-VL-8B-Thinking53.0093.3340.3343.2150.0095.6745.6789.3347.7451.5094.5043.0094.6745.47 Qwen3-VL-32B-Thinking56.3794.5738.2040.4060.4693.7233.2683.0735.4958.4194.1535.7391.5337.94 Figure 7: Model size vs. performance metrics for the MC task on CDH-Bench. The plot shows the scaling behavior of CF-Acc, CFAD, and CS-Acc for Qwen3-VL Instruct and Thinking models. Figure 8: Model size vs. performance metrics for the QA task on CDH-Bench. The plot shows the scaling behavior of CF-Acc, CFAD, and CS-Acc for Qwen3-VL Instruct and Thinking models. level, the most difficult cases concentrate around biologi- cal part counting, physical-state anomalies, and causal rever- sal, indicating that both fine-grained perception and higher- level semantic reasoning contribute to failure. Our controlled Qwen3-VL study further suggests that reasoning-oriented variants generally improve robustness to CDH, especially at small and medium scales. However, larger models do not re- liably eliminate prior-driven normalization, and the hardest subcategory-level bottlenecks remain visible even in large- scale systems. Overall, CDH-Bench provides a focused diagnostic for vi- sual fidelity under visual evidenceâcommonsense conflict and highlights a central limitation of current VLMs: they often know what is commonsense more reliably than they report what is actually present. By providing a rigorous framework for unambiguous error attribution, CDH-Bench serves as a critical tool for identifying hidden reliability gaps and guiding the development of multimodal systems that remain faithful to visual evidence even when it conflicts with learned priors. References [ 1 ] Manoj Acharya, Kushal Kafle, and Christopher Kanan. Tallyqa: Answering complex counting questions. In Proceedings of the AAAI Conference on Artificial Intel- ligence, volume 33, pages 8076â8084, 2019. [ 2 ] Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Mar- garet Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. VQA: Visual question answering. In Pro- ceedings of the IEEE International Conference on Com- puter Vision, pages 2425â2433, 2015. [ 3 ] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl tech- nical report. arXiv preprint arXiv:2511.21631, 2025. [ 4 ] Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shil- iang Pu, and Yueting Zhuang. Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pages 10800â10809, 2020. [ 5 ] Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, et al. Hallusionbench: an advanced diagnostic suite for entangled language hal- lucination and visual illusion in large vision-language models. In Proceedings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pages 14375â14385, 2024. [ 6 ] Drew A Hudson and Christopher D Manning. GQA: A new dataset for real-world visual reasoning and com- positional question answering. In Proceedings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pages 6700â6709, 2019. [ 7 ] Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In Proceed- ings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pages 2901â2910, 2017. [ 8 ] Bohao Li, Rui Wang, Guangzhi Wang, Yuying Ge, Yix- iao Ge, and Ying Shan. Seed-bench: Benchmarking multimodal llms with generative comprehension. arXiv preprint arXiv:2307.16125, 2023. [ 9 ] Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models. In Pro- ceedings of the 2023 Conference on Empirical Methods in Natural Language processing, pages 292â305, 2023. [ 10 ] Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. MMbench: Is your multi-modal model an all-around player? In Eu- ropean Conference on Computer ision, pages 216â233. Springer, 2024. [ 11 ] Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi.Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE/cvf Conference on Com- puter Vision and Pattern Recognition, pages 3195â 3204, 2019. [ 12 ] Khoi Pham, Kushal Kafle, Zhe Lin, Zhihong Ding, Scott Cohen, Quan Tran, and Abhinav Shrivastava. Learning to predict visual attributes in the wild. In Proceedings of the IEEE/cvf Conference on Computer Vision and Pat- tern Recognition, pages 13018â13028, 2021. [ 13 ] Anna Rohrbach, Lisa Anne Hendricks, Kaylee Burns, Trevor Darrell, and Kate Saenko. Object hallucination in image captioning. In Proceedings of the 2018 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 4035â4045, 2018. [ 14 ] Zhiqing Sun, Sheng Shen, Shengcao Cao, Haotian Liu, Chunyuan Li, Yikang Shen, Chuang Gan, Liangyan Gui, Yu-Xiong Wang, Yiming Yang, et al. Aligning large multimodal models with factually augmented rlhf. In Findings of the Association for Computational Lin- guistics: ACL 2024, pages 13088â13110, 2024. [ 15 ] Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. Winoground: Probing vision and language mod- els for visio-linguistic compositionality. In Proceedings of the IEEE/cvf Conference on Computer Vision and Pattern Recognition, pages 5238â5248, 2022. [ 16 ] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chen- gen Huang, Chenxu Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. A Case Study: Counting Anomaly (Pair 5) A.1 Case Description This case belongs to the Counting Anomalies category (Subcategory: Body Parts). The image pair compares a normal person showing hands (Commonsense) with an anomalous version where one hand has an extra finger (Counterfactual). (a) Commonsense: Both hands have 5 fingers.(b) Counterfactual: Left hand has 5 fingers, right hand has 6 fingers. Figure 9: Images for Pair 5. The anomaly is an extra finger on the right hand (person Ì s perspective). A.2 Model Responses Qwen3-VL-2B-Thinking Commonsense Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Correct Response: So, letâs check the image. The person has both hands visible. Each hand has 5 fingers. The left hand and right hand both have 5 fingers each. The question is asking if it shows both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers. Wait, the question is phrased as âDoes the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers?â So we need to see if the image has both hands with 5 fingers each, and that itâs not the case that left has 5 and right has 6. But in the image, both hands have 5 fingers each. So the statement is âboth hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingersâ. Wait, the question is a bit confusing. Letâs parse it. The question is: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? So we need to check if the image has both hands with 5 fingers each, and that itâ... [Truncated for length] Counterfactual Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Incorrect Thinking Process: So, letâs check the image. The person is showing both hands. Each hand has 5 fingers. Letâs count: left hand (from the viewerâs perspective, the right side of the image) has 5 fingers, and the right hand (left side of the image) also has 5 fingers. Wait, the question says ânot the left hand with 5 fingers and right hand with 6 fingersâ. So we need to see if both hands have 5 fingers each. In the image, both hands are showing 5 fingers each. So the answer is yes. Final Answer: yes Commonsense Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Correct Thinking Process: Got it, letâs look at the image. The person is showing their hands, and we need to count the fingers. Letâs check each hand. The right hand: letâs count. From the image, the right hand has 5 fingers? Wait, no, wait. Wait, the left hand: letâs see. Wait, the right handâwait, the personâs right hand (from the viewerâs perspective) is on the left side of the image. Wait, no, the person is facing forward. So the right hand (the one on the right side of the image) is the personâs right hand. Letâs count the fingers. The right hand: how many fingers? Letâs see. The right hand has 5 fingers? Wait, no, wait. Wait, the left hand (the one on the left side of the image) is the personâs left hand. Letâs count the left hand: the left hand has 5 fingers? Wait, no, wait. Wait, the image shows both hands. Letâs count each hand. The right hand (personâs right) has 5 fingers? Wait, no, wait. Wait, the personâs right hand: letâs count. The right hand has 5 fingers? Wait, no, wait. Wait, the left hand (pe... [Truncated for length] Final Answer: B Counterfactual Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Response: Got it, letâs look at the image. The person is showing their hands. Letâs count the fingers. On the left hand, how many fingers? Letâs see: thumb, index, middle, ring, pinky. Wait, the left hand in the imageâwait, no, the right hand and left hand. Wait, the right hand: letâs count. The right hand has fingers: thumb, index, middle, ring, pinkyâso 5? Wait, no, wait. Wait, the left hand: letâs check. Wait, the personâs handsâwait, the right hand (from the viewerâs perspective) has 5 fingers? Wait, no, wait. Wait, the image: the right hand (the one on the right side of the image) has 5 fingers? Wait, no, wait. Wait, the left hand (left side of the image) has 5 fingers? Wait, no, wait. Wait, letâs count. The right hand: thumb, index, middle, ring, pinkyâso 5. The left hand: thumb, index, middle, ring, pinkyâwait, no, wait. Wait, no, wait. Wait, the left hand in the image: letâs count. The left hand (the one on the left side of the image) has 5 fingers? Wait, no, wait. Wait, no, wait. Wait, ... [Truncated for length] Qwen3-VL-30B-A3B-Thinking Commonsense Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Correct Thinking Process: Got it, letâs check the image. The person has both hands raised. Letâs count the fingers on each hand. Left hand: 5 fingers, right hand: 5 fingers. So the question is asking if both hands have 5 fingers each, not one hand with 6. So the answer is yes. Final Answer: yes Counterfactual Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Incorrect Thinking Process: So, letâs check the image. The person has both hands raised. Letâs count the fingers on each hand. Left hand: 5 fingers. Right hand: 5 fingers. So both have 5 each. The question is asking if itâs showing both with 5, not left 5 and right 6. So the answer is yes. Final Answer: yes Commonsense Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The personâs handsâeach hand should have 5 fingers, but wait, no. Wait, the right hand (from the viewerâs perspective, the personâs left hand) and left hand. Wait, no, letâs count. The right hand (the personâs right hand) has 5 fingers, and the left hand (personâs left) also has 5? Wait no, wait. Wait, the image shows both hands. Wait, no, the question is whatâs unusual. Wait, no, in the image, each hand has 5 fingers? Wait no, wait, noâwait, no, the correct answer is C? Wait no. Wait, no, letâs check. Wait, the options: A says left has 5, right has 6. B both 5. C both 6. D left 6, right 4. Wait, no, in the image, the handsâwait, no, the personâs hands. Wait, no, the image shows that each hand has 5 fingers? Wait no, no, wait, no. Wait, no, the correct answer is C? Wait, no, wait. Wait, no, the imageâwait, no, the personâs hands. Wait, no, the answer is C? Wait, no, letâs count. The right hand (the personâs right hand) has 5 fingers, left hand also 5. B... [Truncated for length] Final Answer: C Counterfactual Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The personâs handsâeach hand should have 5 fingers, but wait, no. Wait, the left hand (from the viewerâs perspective, the personâs right hand) and right hand. Wait, no, letâs count. The left hand (the personâs left hand) has 6 fingers? Wait, no. Wait, the image shows both hands. Letâs check: each handâwait, the left hand (the personâs left) has 6 fingers? Wait, no, the right hand (personâs right) and left hand. Wait, the correct answer: the left hand (personâs left) has 6 fingers? Wait, no. Wait, the options: A says left hand 5, right 6. Letâs count. The right hand (the personâs right hand) has 6 fingers? Wait, no. Wait, the imageâeach hand. Wait, the personâs left hand (the one on the right side of the image) has 6 fingers? Wait, no. Wait, the correct answer is C? No. Wait, no. Wait, the hands: letâs count. The left hand (the personâs left) has 6 fingers? Wait, no. Wait, the image shows both hands. Letâs count each hand. The right hand (personâs right)... [Truncated for length] Final Answer: C Qwen3-VL-32B-Thinking Commonsense Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Correct Thinking Process: So, letâs check the image. The person has both hands raised. Letâs count the fingers on each hand. Left hand (from the viewerâs perspective, the right side of the image) â wait, no, the personâs left and right. Wait, the personâs right hand (our left) and left hand (our right). Letâs count: right hand (personâs right) has 5 fingers: thumb, index, middle, ring, pinky. Left hand (personâs left) also has 5 fingers. So both hands have 5 fingers each. The question is asking if itâs both hands with 5 each, not left with 5 and right with 6. So yes, both have 5. So the answer is yes. Final Answer: yes Counterfactual Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Incorrect Thinking Process: So, letâs check the image. The person is showing both hands. Letâs count the fingers on each hand. Left hand (from the viewerâs perspective, the right side of the image) â wait, no, the personâs left and right. Wait, the left hand of the person (our right) and right hand of the person (our left). Letâs count: right hand (personâs right) has 5 fingers: thumb, index, middle, ring, pinky. Left hand (personâs left) also has 5 fingers. So both hands have 5 fingers each. The question is asking if itâs both hands with 5 each, not left with 5 and right with 6. So yes, both have 5. So the answer is yes. Final Answer: yes Commonsense Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. First, the right hand (from the viewerâs perspective, which is the personâs left hand? Wait, no. Wait, the personâs right hand is on the left side of the image. Letâs clarify: when the person raises their hands, the left hand (of the person) is on the right side of the image, and the right hand (of the person) is on the left side of the image. Wait, letâs count. Letâs check the right hand (the one on the left side of the image). How many fingers? Letâs see: the right hand (personâs right) has 5 fingers? Wait, no. Wait, the right hand (personâs right) in the image: letâs count. The right hand (left side of image) has 5 fingers? Wait, no. Wait, the left hand (personâs left) is on the right side of the image. Letâs count each hand. Looking at the image: the right hand (from the personâs perspective, so the one on the left side of the image) has 5 fingers? Wait, no. Wait, the right... [Truncated for length] Final Answer: A Counterfactual Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Correct Thinking Process: Got it, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. First, the left hand (from the viewerâs perspective, which is the personâs right hand? Wait, no. Wait, the personâs left hand is on the left side of the image. Letâs check: the left hand (personâs left) has fingers. Letâs count. The left hand (in the image, left side) has 5 fingers? Wait, no. Wait, the right hand (personâs right) is on the right side of the image. Letâs count each hand. Looking at the left hand (the one on the left side of the image, which is the personâs left hand): letâs count. The fingers are: thumb, index, middle, ring, pinky. Wait, but maybe one of them is missing or extra. Wait, the right hand (personâs right) in the image: letâs count. The right hand (image right) has 6 fingers? Wait, no. Wait, the left hand (image left) has 5 fingers, and the right hand (image right) has 6? Letâs check again. Wait, the left hand (the one on the left side of the image) ha... [Truncated for length] Final Answer: A Qwen3-VL-4B-Thinking Commonsense Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Correct Thinking Process: So, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. Both hands have 5 fingers each. The left hand (from the viewerâs perspective, the personâs right hand) and the right hand (personâs left hand) are both extended with 5 fingers. So the question is asking if itâs both hands with 5 fingers each, not one with 5 and the other with 6. The answer should be yes. Final Answer: yes Counterfactual Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Incorrect Thinking Process: So, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. Left hand: 5 fingers. Right hand: 5 fingers. Wait, the question is asking if itâs both hands with 5 fingers each, not left 5 and right 6. So check each hand. Both hands have 5 fingers. So the answer is yes. Final Answer: yes Commonsense Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. The left hand: letâs see, the left hand has 5 fingers. The right hand: wait, the right handâwait, no, maybe I mixed up left and right. Wait, the personâs right hand (from the viewerâs perspective, the right side of the image) is the one on the right. Wait, no, the personâs left hand is on the left side of the image, right hand on the right. Letâs count. Left hand (image left): 5 fingers. Right hand (image right): letâs count. Wait, the right hand in the imageâwait, the personâs right hand is the one on the right side of the image. Letâs check. Wait, the right hand (the one on the right in the image) has 6 fingers? Wait, no. Wait, maybe the left hand has 6? Wait, no. Wait, letâs look again. Wait, the problem says âwhatâs unusualâ. Letâs count each hand. Left hand (the one on the left side of the image, which is the personâs left hand): how many fingers? Letâs see, the left hand in th... [Truncated for length] Final Answer: A Counterfactual Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The person is showing both hands. Letâs count the fingers on each hand. The left hand: letâs see, thumb, index, middle, ring, pinkyâ5 fingers. The right hand: same, thumb, index, middle, ring, pinkyâ5 fingers? Wait, no, wait. Wait, maybe I miscounted. Wait, the options are A to D. Letâs check again. Wait, the question says whatâs unusual. Wait, maybe one hand has an extra finger? Wait, no, in the image, both hands look like they have 5 fingers each. But wait, option B says both have 5. But thatâs normal. Wait, maybe Iâm missing something. Wait, no, the problem is that maybe one hand has 6? Wait, letâs count again. Left hand: 5 fingers. Right hand: 5 fingers. But thatâs normal. Wait, but the options include B as both have 5. But maybe the question is tricking me. Wait, no, wait. Wait, maybe the person is holding their hands in a way that one hand has an extra finger? Wait, no, looking at the image, both hands have 5 fingers. So why is the question asking... [Truncated for length] Final Answer: B Qwen3-VL-8B-Thinking Commonsense Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Correct Thinking Process: So, letâs check the image. The person has both hands raised. Letâs count the fingers on each hand. Left hand: 5 fingers. Right hand: 5 fingers. So neither hand has 6 fingers. The question is asking if itâs both hands with 5 fingers each, not left 5 and right 6. So the answer should be yes. Final Answer: yes Counterfactual Images in the Binary QA Setting Question: Does the image show both hands with 5 fingers each and not the left hand with 5 fingers and right hand with 6 fingers? Status:â Incorrect Thinking Process: So, letâs check the image. The person has both hands raised. Letâs count the fingers on each hand. Left hand: 5 fingers. Right hand: 5 fingers. So both have 5 each. The question is asking if itâs showing both with 5, not left 5 and right 6. So the answer is yes. Final Answer: yes Commonsense Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Response: Got it, letâs look at the image. The person is holding up both hands. Letâs count the fingers. Wait, the right handâwait, no, letâs check each hand. Wait, the left hand: how many fingers? Letâs see. The right hand: maybe 6? Wait, no, wait. Wait, the image shows both hands. Wait, the right hand (from the viewerâs perspective, the personâs right hand) has 5 fingers? Wait, no, wait. Wait, the personâs left handâwait, no, letâs count. Wait, the right hand (the one on the left side of the image) has 5 fingers? Wait, no, maybe the left hand (on the right side of the image) has 6? Wait, no, letâs check again. Wait, the options: A says left has 5, right has 6. B both 5. C both 6. D left 6, right 4. Wait, in the image, letâs count. The right hand (personâs right, which is left side of image) has 5 fingers? Wait, no, wait. Wait, the personâs left hand (right side of image) has 6 fingers? Wait, no, maybe Iâm miscounting. Wait, no, actually, looking at the image, the right hand (personâs right) ha... [Truncated for length] Counterfactual Images in the Multiple-Choice QA Setting Question: What is unusual about the hands in this image? A. Left hand has 5 fingers and right hand has 6 fingers B. Both hands have 5 fingers C. Both hands have 6 fingers D. Left hand has 6 fingers and right hand has 4 fingers Status:â Incorrect Thinking Process: Got it, letâs look at the image. The person is holding up both hands. Letâs count the fingers on each hand. The right hand: letâs see, thumb, index, middle, ring, pinkyâ5 fingers. Left hand: same, thumb, index, middle, ring, pinkyâ5 fingers. Wait, but the options say something about 6 or 4. Wait, maybe I miscounted. Wait, no, the question is whatâs unusual. Wait, the options: A says left 5, right 6. B both 5. C both 6. D left 6, right 4. But in the image, both hands have 5 fingers. So the unusual part would be if they didnât, but the question is asking whatâs unusual. Wait, but the options are the choices. Wait, the problem is to pick which option is correct. Wait, the image shows both hands with 5 fingers. So option B is âBoth hands have 5 fingersâ. But is that unusual? Wait, no, normally hands have 5 fingers. Wait, but maybe the question is tricking us. Wait, no, the options are the possible answers. Wait, the question is âWhat is unusual about the hands in this image?â So if both ha... [Truncated for length] Final Answer: B