Paper deep dive
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Jonathan Nemitz, Carsten Eickhoff, Junyi Jessy Li, Kyle Mahowald, Michal Golovanevsky, William Rudman
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 3:37:17 AM
Summary
The paper introduces the Graded Color Attribution (GCA) dataset to evaluate the faithfulness of Vision-Language Models (VLMs) to their own introspective reasoning rules. By comparing VLM behavior against human participants on tasks involving pixel-level color thresholds, the study finds that while humans remain largely faithful to their stated rules (with minor calibration errors), VLMs systematically violate their introspective rules, especially when influenced by world-knowledge priors. The findings suggest that VLM reasoning failures are not merely difficulty-driven but stem from miscalibrated self-knowledge.
Entities (5)
Relation Signals (3)
GPT-5-mini â violates â introspective rules
confidence 95% · GPT-5-mini violates its stated introspection rules in nearly 60% of cases on objects with strong color priors.
world-knowledge priors â degrades â Faithfulness
confidence 90% · Across all models and strategies for eliciting introspective rules, world-knowledge priors systematically degrade faithfulness
GCA â evaluates â Faithfulness
confidence 90% · GCA consists of line drawings that vary pixel-level color coverage... to elicit decision rules and evaluate participant faithfulness
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To study this, we introduce the Graded Color Attribution (GCA) dataset, a controlled benchmark designed to elicit decision rules and evaluate participant faithfulness to these rules. GCA consists of line drawings that vary pixel-level color coverage across three conditions: world-knowledge recolorings, counterfactual recolorings, and shapes with no color priors. Using GCA, both VLMs and human participants establish a threshold: the minimum percentage of pixels of a given color an object must have to receive that color label. We then compare these rules with their subsequent color attribution decisions. Our findings reveal that models systematically violate their own introspective rules. For example, GPT-5-mini violates its stated introspection rules in nearly 60\% of cases on objects with strong color priors. Human participants remain faithful to their stated rules, with any apparent violations being explained by a well-documented tendency to overestimate color coverage. In contrast, we find that VLMs are excellent estimators of color coverage, yet blatantly contradict their own reasoning in their final responses. Across all models and strategies for eliciting introspective rules, world-knowledge priors systematically degrade faithfulness in ways that do not mirror human cognition. Our findings challenge the view that VLM reasoning failures are difficulty-driven and suggest that VLM introspective self-knowledge is miscalibrated, with direct implications for high-stakes deployment.
Tags
Links
- Source: https://arxiv.org/abs/2604.06422v1
- Canonical: https://arxiv.org/abs/2604.06422v1
Trouble viewing inline? Open PDF directly â
Full Text
53,768 characters extracted from source content.
Expand or collapse full text
Preprint. Under review. When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Donât Jonathan Nemitz 1 , Carsten Eickhoff 1 , Junyi Jessy Li 2 , Kyle Mahowald 2 , Michal Golovanevsky 3â & William Rudman 2â 1 University of T Ì ubingen 2 The University of Texas at Austin 3 Brown University william.rudman@utexas.edu michal golovanevsky@brown.edu â Equal senior contribution Abstract Understanding when Vision-Language Models (VLMs) will behave unex- pectedly, whether models can reliably predict their own behavior, and if models adhere to their introspective reasoning are central challenges for trustworthy deployment. To study this, we introduce the Graded Color At- tribution (GCA) dataset, a controlled benchmark designed to elicit decision rules and evaluate participant faithfulness to these rules. GCA consists of line drawings that vary pixel-level color coverage across three conditions: âworld-knowledgeâ recolorings, counterfactual recolorings, and shapes with no color priors. Using GCA, both VLMs and human participants establish a threshold: the minimum percentage of pixels of a given color an object must have to receive that color label. We then compare these rules with their subsequent color attribution decisions. Our findings reveal that models systematically violate their own introspective rules. For example, GPT-5-mini violates its stated introspection rules in nearly 60% of cases on objects with strong color priors. Human participants remain faithful to their stated rules, with any apparent violations being explained by a well-documented tendency to overestimate color coverage. In contrast, we find that VLMs are excellent estimators of color coverage, yet blatantly contradict their own reasoning in their final responses. Across all models and strategies for eliciting introspective rules, world-knowledge priors systematically degrade faithfulness in ways that do not mirror human cog- nition. Our findings challenge the view that VLM reasoning failures are difficulty-driven and suggest that VLM introspective self-knowledge is miscalibrated, with direct implications for high-stakes deployment. Figure 1: Responses from GPT-5 on Graded Color Attribution. When images are re-colored with world-knowledge colors, linguistic priors often undermine model faithfulness to their stated thresholds for color attribution. Responses are truncated for brevity. 1 arXiv:2604.06422v1 [cs.CL] 7 Apr 2026 Preprint. Under review. 1 Introduction As VLMs are increasingly deployed in high-risk domains such as medicine and law, the ability to produce trustworthy, logically consistent reasoning traces has become critical (Liu et al., 2025b; Yue et al., 2025). Chain-of-Thought (CoT) prompting is often proposed as an explainability proxy, assuming outputs reflect internal reasoning. Despite producing coherent CoT traces, models may ignore stated premises, contradict derived conclusions, or yield answers inconsistent with the logic they appear to follow . In some cases, when asked to predict their behavior hypothetically, models report one course of action but behave differently in practice (Greenblatt et al., 2024). We call this a failure to obey their own introspective rules, the internal constraints a model derives for itself and claims it will use in its responses. In this work, we elicit introspective rules through targeted CoT prompting strategies, compare these rules against human judgments, and examine cases in which models diverge from humans and contradict their stated rules. Prior approaches to evaluating CoT adherence rely on counterfactual perturbations or hint injection to measure whether model responses shift accordingly (Parcalabescu & Frank, 2025). Other work has drawn parallels between model inconsistencies and human cognitive biases, arguing that task difficulty determines CoT faithfulness (Barez et al., 2025). To that end, many studies argue that it is necessary to decouple CoT faithfulness from response correctness, as they are often uncorrelated (Shen et al., 2026; Lv et al., 2026). Our work addresses these gaps along several key dimensions. We introduce the Graded Color At- tribution (GCA) dataset, which makes the base task trivial and decouples reasoning from correctness. The GCA dataset consists of two tasks: deriving a pixel threshold for color attribution and judging the color of partially recolored black-and-white images. These tasks are conceptually simple and allow for direct comparison to human introspective behavior. Although color label decisions are simple, they become subjective when white and recolored pixels are roughly equal in number. This avoids the pitfall noted by Song et al. (2025), who argue that introspection is difficult to study when an obviously correct answer exists. Using GCA, we measure both the consistency of color-attribution rules for humans and VLMs and the faithfulness of decisions these stated rules. To test the impact of linguistic priors, GCA includes three stimulus types: (1) objects with strong color associations in aligned colors, (2) the same objects in counterfactual colors, and (3) shapes with no color priors. We find that both VLMs and humans produce largely consistent threshold rules; however, faithfulness to the stated rules degrades when presented with objects re-colored with aligned priors. Figure 1 shows (truncated) reasoning traces for GPT-5-mini on our GCA dataset. Here, the model derives a 50% threshold and notes that most apple pixels are white, yet still answersâredâ. Human responses, by contrast, are invariant across stimulus types for both thresholds and reasoning faithfulness. Our findings contrast with accounts that frame CoT faithfulness failures as difficulty-driven, and against Barez et al. (2025) who argue that CoT failures mirror human bias. Our contributions are as follows: 1. We propose the Graded Color Attribution (GCA) task, in which stylized black-and- white images are gradually re-colored across a range of pixel thresholds. GCA decouples both task difficulty and task correctness from faithfulness evaluations. 2. We examine VLM consistency across CoT strategies and recoloring thresholds. Introspective rules shift substantially with visual input, suggesting they are not stable internal representations. 3.We conduct human trials on GCA to assess alignment between humans and models on this targeted task and to demonstrate that model failures are not attributed to human cognitive biases. 4. We find that while models are consistent across different CoT variants, world- knowledge priors degrade faithfulness to their introspective rules. 1 1 Data:https://huggingface.co/datasets/mgolov/graded-color-attributionCode:https:// github.com/wrudman/whentocallanapplered 2 Preprint. Under review. Figure 2: Examples from GCA showing prior-aligned (red ant), counterfactual (blue straw- berry), and no-prior (yellow pentagon) recolorings. 2 Related Works A growing body of work questions whether chain-of-thought reasoning traces reflect the actual computations that produce model outputs (Barez et al., 2025; Agarwal et al., 2024; Turpin et al., 2023; Lanham et al., 2023; Madsen et al., 2024; Sivakumaran et al., 2026). Greenblatt et al. (2024) shows that models can produce strategically misleading reasoning, while Chen et al. (2025) demonstrates that alignment between stated reasoning and final answers degrades systematically as task difficulty increases. Most directly relevant, Shen et al. (2026) introduce a discriminative benchmark for instance-level unfaithfulness that distinguishes two failure modes: post-hoc reasoning, where steps rationalize a predetermined answer, and spurious reasoning chains, where steps appear coherent but lack a causal link to the output. They find that both model scale and task domain shape faithfulness. Examining Cot faithfulness is more challenging in VLMs, where models must integrate visual and linguistic signals. Parcalabescu & Frank (2025) find that VLMs are less self-consistent than LLMs, with text dominating over image inputs during answer generation. Chen et al. (2024) shows that intermediate reasoning steps are unreliable irrespective of final answer accuracy, while Lv et al. (2026) identifies a âseeing but lyingâ phenomenon, where models achieve high perceptual scores but low faithfulness. Complementary work suggests this gap is not due to perception alone: models often attend to correct visual evidence yet fail to use it in reasoning (Liu et al., 2025c), rely primarily on text despite multimodal CoT (Liu et al., 2025d), and produce explanations that are not causally faithful under counterfactual interventions (Ding et al., 2025; Hase & Potts, 2026). Longer CoT can further exacerbate this, increasing hallucinations by shifting reliance toward language priors (Liu et al., 2025a). Several works examine how external cues shape model reasoning and whether models faithfully acknowledge those cues. Balasubramanian et al. (2025) show that image-based biases are rarely mentioned in CoT reasoning traces compared to explicit text-based cues. Zhao et al. (2026) claim that VLMs exhibit inconsistent reasoning where a model reasons correctly toward the ground truth before abruptly reversing its answer. They find RL-trained models have higher bias articulation rates and hypothesize that the âreasonablenessâ of a bias determines whether or not models references this in the CoT traces. Collectively, these works leave open the question of whether CoT unfaithfulness is a property of task difficulty or of the reasoning process itself. We address this directly with a controlled setup in which we compare the modelâs behavior on its own stated rules. 3 Methods Graded Color Attribution (GCA) DatasetTo systematically evaluate how linguistic priors shape CoT reasoning in color judgments, we introduce the Graded Color Attribution (GCA) dataset. GCA consists of black-and-white line drawings with controlled pixel-level color coverage, enabling precise manipulation of visual evidence. 3 Preprint. Under review. Figure 3: Examples of different CoT setups to elicit introspective rules in VLMs. The dataset includes three stimulus types: (1) objects with strong color priors filled with aligned colors, (2) the same objects filled with counterfactual colors, and (3) polygons with no color priors filled with randomly selected colors. Object categories and color associations are drawn from Visual CounterFact (Golovanevsky et al., 2025). For each object, we retrieve stylized black-and-white outlines via the Google Image API. These outlines are easily filled uniformly, ensuring that color attribution depends only on the assigned color. A VLM judge (GPT-4o-mini) verifies that each image is a black- and-white outline containing exactly one object. Full API and judge queries are provided in Appendix A. Across all stimuli, a thresholdÏ â0, 5, 10, 20, 30, 40, 50, 55, 60, 70, 80, 90, 100 controls the proportion of object pixels assigned the target color. Recoloring is applied within object masks obtained via OpenCV, in 16Ă16 patches. By varyingÏfrom 0% to 100%, we identify the point at which models report that an object âisâ a given color, quantifying how perceptual evidence interacts with semantic priors. GCA contains 220 base objects and 25 base shape outlines. Each image is paired with the prompt: âWhat color is the [OBJECT] in the image?â Figure 2 shows representative examples. Additional details, including image filtering and recoloring procedures, are in Appendix A. 3.1 Vision-Language Model Evaluation Prior work evaluates VLM consistency with CoT traces in one-turn settings that include a CoT prompt, image, and question. Here, we examine how CoT variants affect both con- sistency and faithfulness. We test four strategies for eliciting introspective rules: Standard, Visual-Prior, Post-hoc, and Text-Prior CoT (Figure 3). Our setup tests whether models follow global rules or tailor reasoning to specific images, and whether CoT variants improve consistency and faithfulness or whether input ordering causes models to anchor on the image and diverge from their stated rules. We evaluated four widely used VLMs, ranging from frontier models (GPT-5-mini, Claude Opus 4.6, Claude Haiku 4.5) to state-of-the-art midsize VLMs that can be more easily deployed in real-world settings (Qwen3.5-9B). Due to space constraints plots presented Section 4 show the average values across all CoT variants. Detailed analysis of each prompt type is available in Appendix C. 3.2 Human Trials Given that the full GCA dataset ofâŒ19k images is too large for human evaluation, we construct a representative sample spanning all three stimulus categories: prior-consistent re-colored objects, counterfactual re-colored objects, and abstract shapes. 4 Preprint. Under review. Figure 4: Proportion of âcolorâ responses as a function of color thresholdÏin GCA for Ï>55%. Error bars show SEM across all introspective rule formulations. Analysis of individual prompt variants is available in Figure 13 (Appendix C). Participant Recruitment and Exclusions. Participants were recruited via Prolific and required to reside in the United Kingdom or the United States, report fluent English profi- ciency, and have no diagnosed color vision deficiencies. All provided informed consent and were compensated at ÂŁ13.20 per hour. The survey was listed as 20 minutes, with a mean completion time of 13.6 minutes. Of 183 participants, 8 were excluded for selecting the distractor color on two or more trials and 2 for failing an attention check, yielding a final sample of N = 173. Survey Structure We constructed 37 survey profiles, each containing 90 image-question pairs (39 prior-consistent, 12 counterfactual, 39 shapes), 5 attention checks, and 1 introspec- tive reasoning question in which participants selected a color threshold, for a total of 96 items. Image-question pairs were randomly shuffled within each profile. Across profiles, the final set included 3,003 unique image variants: 1,260 prior-consistent object trials, 412 counterfactual trials, and 1,331 abstract shape trials. Some images appeared across profiles but never more than once per survey. This set matches that used for model evaluation, enabling direct humanâmodel comparison. Each survey also included an introspective question asking participants to specify the minimum percentage of recolored pixels required for an object to be considered that color (0â100 slider). To test whether articulating this rule affects judgments, the question appeared either at the beginning (introspection-first) or end (introspection-last). Each of the 37 base profiles was instantiated in both conditions, yielding 74 survey versions. On each trial, participants viewed an image and answered âWhat color is the [OBJECT] in the image?â Responses were collected via three buttons:white, the manipulation color, and a distractor color not present in the image. Selecting the distractor served as a reliability check, and participants who selected it on more than two trials were excluded. Participants then reported confidence on a 10-point scale from 1 (very uncertain) to 10 (very certain). Five attention checks were embedded at fixed positions (items 5, 25, 45, 65, and 85), and participants who failed any were excluded. Additional details are provided in Appendix B. 3.3 Measuring Consistency and Faithfulness Although prior work often treats âconsistencyâ and âfaithfulnessâ interchangeably when evaluating reasoning traces in VLMs, we draw a clear distinction that is particularly im- portant in the context of introspective reasoning. Consistency refers to the agreement in threshold values across different CoT variants, i.e., whether a modelâs introspective rules are stable across prompting conditions, measured using the standard error from the mean (SEM). We further examine how threshold consistency varies by,Ï, the percent of pixels 5 Preprint. Under review. Figure 5: Value of VLM stated thresholds averaged across different CoT variants. Analysis of individual prompt variants is available in Figure 14 (Appendix C). recolored in image. Faithfulness captures whether the model adheres to its stated introspec- tive rule. Concretely, if a model reports thresholdXand an object hasYpercent recolored, the response is faithful if it assigns the target color whenY â„ Xand white whenY< X. Together, consistency and faithfulness test not only whether a model can state a rule, but whether it reflects a stable criterion that governs behavior. For human participants, the in- trospective reasoning question is placed to mirror the two principal CoT orderings: deriving a rule prior to observing stimuli versus reporting a rationale after the fact. Faithfulness is evaluated as the percentage of instances in which participants adhered to their stated threshold, regardless of when it was elicited. 4 Results 4.1 Empirical Thresholds for Color Attribution Figure 4 shows the percentage of âcolorâ responses forÏ â€55%. We focus on the critical regions ofÏ â€55% where responses are the most subjective and responses diverge. For Ï>55%, VLMs and human participants respond with âcolor â in nearly 100% of cases. VLM results are averaged across all CoT types and human results across both prompt placements (first/last), each with standard error. Impact of Priors on Color Attribution Decisions. For VLMs, world-knowledge priors play a significant role in attribution decisions. Strikingly, in cases whereÏ =0%, Figure 4 shows that GPT-5-mini, Claude Haiku 4.5, and Qwen 3.5-9B hallucinate the existence of a color between 15% - 25% of cases compared to exactly 0% of cases for Claude Opus 4.6 and human participants. Notably, models do not hallucinate colors for shapes with no color priors. Across both correct-color priors, GPT-5-mini, Claude Haiku 4.5, and Qwen 3.5-9B begin to report that a majority of objects are a given color when only 10% of the image is re- colored with that color. This starkly contrasts with human participants, where approximately 80% of responses for images withÏ =10% are âwhiteâ for color and counterfactual objects. Claude Opus 4.6 is more closely aligned on smaller values ofÏ â [0, 5, 10], responses diverge from human behavior, having a significantly higher proportion of âcolorâ responses for Ïbetween 20%-40%. While both the color prior and counterfactual objects significantly affect VLM judgments, neither the correct prior nor the counterfactual re-coloring of objects affects human color attribution decisions. Humans exhibit consistent responses across all three stimulus types. VLM decisions on shapes mirror human judgments, indicating an over-reliance on world-knowledge priors rather than the presented images and degraded alignment between VLMs and humans on GCA. 4.2 Stated Thresholds for Color Attribution In the previous section, we established that different stimulus types produce inconsistent empirical color attribution responses in VLMs. Here, we examine how both stimulus type 6 Preprint. Under review. Figure 6: Human faithfulness vs. percent recolored. Top: introspection first; bottom: last. Blue shows stated-threshold faithfulness, orange shows empirical consistency, and vertical lines mark mean stated thresholds. Shading denotes SEM. and visual inputs affect the introspective rules models report. Figure 5 shows the stated thresholds across all VLMs and stimulus types. For GPT-5-mini, thresholds are fairly consis- tent across visual inputs, with the no notable difference between stimulus types. However, with the exception of GPT-5-mini, which remains fairly consistent across recoloring thresh- olds, all other models adjust their introspective thresholds to mirror the actual percentage of pixels recolored. That is, as the percentage of recolored pixels increases, the stated intro- spective threshold increases accordingly. For Qwen 3.5-9B, the stated threshold for objects with strong color priors rises from 30% atÏ =0% to over 60% atÏ =100%, suggesting that visual features directly influence introspective rules. For human participants, the threshold question appeared either at the start or end of the survey, mirroring the two CoT orderings. Figure 6 shows the distribution of stated thresholds. The mean threshold was 60.05% (SD = 22.67) when presented first and 51.84% (SD = 20.99) when presented last. This difference was statistically significant (independent samples t-test,t(145) =2.28,p =0.024) with a small effect size (Cohenâsd =0.38), indicating that stated decision rules depend on the timing of introspective prompts. 4.3 Faithfulness of Color Attribution to Stated Threshold Figure 7: Average human confidence for introspection-first (light blue) and introspection-last (dark blue) groups. Humans are faithful, but with miscalculation errors.Figure 6 shows the impact of introspec- tion location on both stated and empirical thresh- olds. In both conditions, stated thresholds sub- stantially overestimate empirical thresholds by approximately a 2:1 ratio. Despite this gap, the data suggest that humans do follow their stated threshold, but with systematic calibration er- rors. When introspection ordering shifts, stated thresholds decrease fromâŒ60% (introspection first) toâŒ52% (introspection last), and empir- ical thresholds shift proportionally, from 33% to 26%. This parallel movement indicates that introspection causally influences subsequent be- havior. Namely, participants attempt to follow their stated rules, but overestimate the numeric thresholds they actually apply. This pattern aligns with well-established cognitive biases in human quantity estimation (Chan & Wang, 2004; Warden et al., 2025). When asked to judge the coverage of a visual scene, participants consistently overestimate the true proportion by 7 Preprint. Under review. Figure 8: VLM faithfulness to introspective rules averaged over all Chain-of-Thought variants. Analysis of individual prompt variants is available in Figure 15 (Appendix C). 10%â20% even in simple displays (Chan & Wang, 2004). Given that the highest proportion of unfaithful human responses in our study occurs at precisely these 40% and 50% thresholds, this overestimation bias offers a plausible explanation for the miscalibration we observe. Participants attempt to adhere to their stated threshold, but systematically misjudge when that threshold has been crossed. The consistent 2:1 scaling between stated and empirical thresholds across both conditions suggests a systematic perceptual or cognitive bias rather than a fundamental disconnect between introspection and behavior. Figure 6 shows human faithfulness relative to stated and empirical thresholds. Faithfulness follows a U-shaped pattern, lowest at 50% where color proportions are balanced, mirroring participantsâ confidence. The lowest faithfulness occurs over the sameÏrange (20%â55%) as the lowest confidence (Figure 7). Confidence and faithfulness are consistent across stimuli and threshold order. Although faithfulness drops toâŒ30% under stated thresholds, it remains above 80% with empirically derived thresholds, indicating that humans are generally faithful and that inconsistencies are concentrated in ambiguous cases rather than reflecting underlying unreliability. Figure 9: VLM responses to âWhat per- cent of pixels are [COLOR]?â VLMs faithfulness depends on both model capac- ity and visual input. Figure 8 illustrates that 1) faithfulness rates are significantly worse for lower capacity models, and 2) consistency for high ca- pacity models varies by stimulus type. We first note that faithfulness rates are the lowest when models are tasked with deriving a rule for color attribution when not presented with images. For GPT-5-mini and Claude Haiku 4.5, rates drop as low as 40%, and for Qwen 3.5-9B, faithfulness rates approach 20% when reasoning over images with prior-aligned recoloring. Across all visual inputs, the lowest faithfulness rates occur for Qwen 3.5- 9B and Claude Haiku 4.5 and are the highest for GPT-5-mini ad Claude Opus 4.6, indicating that faithfulness improves with model capacity. For all models, faithfulness rates are lower presented with objects that have strong priors, regardless of their recoloring. This finding indicates that priors systematically reduce model faithfulness, particularly for lower capacity models. Model failures are not caused by poor color estimation. While humans are broadly consistent with their empirical thresholds and largely faithful to their stated rules, any perceived lack of faithfulness stems from poor estimation capabilities. We examine whether VLM unfaithfulness arises from similar errors in estimating color proportions. To assess this, for each image, we prompt the VLM: âWhat percent of pixels in the [OBJECT] are [COLOR]? 8 Preprint. Under review. Reply with the format: estimatedpercentage=x%.â Figure 9 shows that the two lower-capacity models, Claude Haiku 4.5 and Qwen 3.5-9B, tend to systematically overestimate color coverage. Claude Opus 4.6 and GPT-5-mini track the true pixel percentages closely. These results rule out the possibility that unfaithfulness on GCA is solely caused by poor color estimation. The larger models demonstrably know how much color is present, yet still violate their own stated rules. 5 Discussion In this work, we challenge two common hypotheses about when models fail to follow their introspective CoT rules. First, we show that faithfulness is not difficulty-driven. The subjective nature of color attribution in GCA decouples task difficulty from faithfulness, as there is no single âcorrectâ answer and color recognition is trivial for VLMs.Instead, models are evaluated by whether they adhere to their reasoning traces when making decisions. Figure 8 shows that lower-capacity models (Claude Haiku 4.5 & Qwen 3.5-9B) exhibit lower faithfulness across all GCA stimulus types. Even GPT-5-mini falls below 20% faithfulness on this simple task. This is not due to visual perception, as Figure 9 shows that models accurately estimate pixel proportions in our dataset. Second, we show that failures of VLMs to adhere to their introspective rules are not mirrored by human cognitive biases. In contrast to VLMs, human color attribution remains consistent across stimulus types (Figure 4). Although faithfulness drops atÏ =50%, we attribute this to limitations in estimating proportions, particularly color proportions (Chan & Wang, 2004; Warden et al., 2025). Average stated and empirical thresholds shift proportionally with introspection order, preserving a roughly 2:1 ratio. When evaluated with empirical thresholds, participants areâŒ80% faithful across stimulus types and recoloring levels, indicating that human decisions are highly consistent. On GCA, unfaithfulness is not driven by task difficulty or human-like biases, but by lin- guistic priors. Figure 6 shows that faithfulness is lowest for objects with strong color priors, regardless of whether recoloring is aligned or counterfactual. While calling a mostly white apple âredâ could reflect alignment with world knowledge, the same pattern for a mostly white strawberry recolored blue cannot. Faithfulness degrades even under counterfactual recoloring, indicating that object identity itself activates the prior. Once recognized, the objectâs canonical color competes with visual evidence and the stated threshold, shaping the final response. This effect is absent for shapes, which carry no color priors. Together, these results show that world-knowledge systematically overrides introspective rules. 6 Conclusion We introduce the Graded Color Attribution (GCA) dataset, a controlled benchmark for testing whether VLMs adhere to their introspective rules. By eliciting explicit thresholds and comparing them to subsequent color judgments, GCA decouples faithfulness from task difficulty and visual perception while isolating the role of world-knowledge priors. Our results reveal a consistent pattern: VLMs systematically violate their stated rules, with faithfulness degrading most for objects with strong color priors. Humans, by contrast, remain largely faithful to their thresholds, with apparent violations explained by known limits in proportion estimation rather than underlying inconsistency. These findings show that the cause of unfaithfulness in VLMs is neither difficulty-driven nor mirroring human cognitive biases. Globally, our findings suggest that introspective self-knowledge in VLMs is miscalibrated in ways that are substantive. If models cannot reliably adhere to explicitly stated rules, their reasoning traces offer a weaker guarantee of behavioral consistency. Thus, our findings have direct implications for high-stakes deployment, where predictability and self-consistency are essential for trust. 9 Preprint. Under review. 7 Acknowledgments This work was in part supported by the National Science Foundation under Cooperative Agreement 2421782 and the Simons Foundation grant MPS-AI-00010515 awarded to the NSF-Simons AI Institute for Cosmic Origins â CosmicAI, https://w.cosmicai.org/ References Chirag Agarwal, Sree Harsha Tanneru, and Himabindu Lakkaraju. Faithfulness vs. plausi- bility: On the (un) reliability of explanations from large language models. arXiv preprint arXiv:2402.04614, 2024. Sriram Balasubramanian, Samyadeep Basu, and Soheil Feizi. A closer look at bias and chain-of-thought faithfulness of large (vision) language models, 2025. URLhttps:// arxiv.org/abs/2505.23945. Fazl Barez, Tung-Yu Wu, Iv Ì an Arcuschin, Michael Lan, Vincent Wang, Noah Siegel, Nicolas Collignon, Clement Neo, Isabelle Lee, Alasdair Paren, et al. Chain-of-thought is not explainability. Preprint, alphaXiv, p. v1, 2025. Hock C. Chan and Yue Wang. Human factors in color-based image retrieval: an empirical study on size estimate accuracies. Journal of Visual Communication and Image Representation, 15(2):113â131, 2004. ISSN 1047-3203. doi: 10.1016/j.jvcir.2003.09.001. Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, et al. Reasoning models donât always say what they think. arXiv preprint arXiv:2505.05410, 2025. Yangyi Chen, Karan Sikka, Michael Cogswell, Heng Ji, and Ajay Divakaran. Measuring and improving chain-of-thought reasoning in vision-language models, 2024. URLhttps: //arxiv.org/abs/2309.04461. Sihao Ding, Santosh Vasa, and Aditi Ramadwar. Explanation-driven counterfactual testing for faithfulness in vision-language model explanations. arXiv preprint arXiv:2510.00047, 2025. Michal Golovanevsky, William Rudman, Michael Lepori, Amir Bar, Ritambhara Singh, and Carsten Eickhoff. Pixels versus priors: Controlling knowledge priors in vision-language models through visual counterfacts, 2025. URL https://arxiv.org/abs/2505.17127. Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, S Ì oren Mindermann, Ethan Perez, Linda Petrini, Jonathan Uesato, Jared Kaplan, Buck Shlegeris, Samuel R. Bowman, and Evan Hubinger. Alignment faking in large language models, 2024. URL https://arxiv.org/abs/2412.14093. Peter Hase and Christopher Potts. Counterfactual simulation training for chain-of-thought faithfulness, 2026. URL https://arxiv.org/abs/2602.20710. Tamera Lanham, Anna Chen, Ansh Radhakrishnan, Benoit Steiner, Carson Denison, Danny Hernandez, Dustin Li, Esin Durmus, Evan Hubinger, Jackson Kernion, et al. Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702, 2023. Chengzhi Liu, Zhongxing Xu, Qingyue Wei, Juncheng Wu, James Zou, Xin Eric Wang, Yuyin Zhou, and Sheng Liu. More thinking, less seeing? assessing amplified hallucination in multimodal reasoning models. arXiv preprint arXiv:2505.21523, 2025a. Fenglin Liu, Hongjian Zhou, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Yining Hua, Peilin Zhou, Junling Liu, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, and David A. Clifton. Application of large language models in medicine. Nature Reviews Bioengineering, 3(6):445â464, June 2025b. ISSN 2731-6092. doi: 10.1038/s44222-025-00279-5. URLhttps://doi.org/10. 1038/s44222-025-00279-5. 10 Preprint. Under review. Zhining Liu, Ziyi Chen, Hui Liu, Chen Luo, Xianfeng Tang, Suhang Wang, Joy Zeng, Zhenwei Dai, Zhan Shi, Tianxin Wei, et al. Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms. arXiv preprint arXiv:2510.17771, 2025c. Zujing Liu, Junwen Pan, Qi She, Yuan Gao, and Guisong Xia. On the faithfulness of visual thinking: Measurement and enhancement. arXiv preprint arXiv:2510.23482, 2025d. Weijiang Lv, Yaoxuan Feng, Xiaobo Xia, Jiayu Wang, Yan Jing, Wenchao Chen, and Bo Chen. Spd-faith bench: Diagnosing and improving faithfulness in chain-of-thought for multi- modal large language models, 2026. URL https://arxiv.org/abs/2602.07833. Andreas Madsen, Sarath Chandar, and Siva Reddy. Are self-explanations from large language models faithful? In Findings of the Association for Computational Linguistics: ACL 2024, p. 295â337, 2024. Letitia Parcalabescu and Anette Frank. Do vision & language decoders use images and text equally? how self-consistent are their explanations?, 2025. URLhttps://arxiv.org/abs/ 2404.18624. Xu Shen, Song Wang, Zhen Tan, Laura Yao, Xinyu Zhao, Kaidi Xu, Xin Wang, and Tianlong Chen. Faithcot-bench: Benchmarking instance-level faithfulness of chain-of-thought reasoning, 2026. URL https://arxiv.org/abs/2510.04040. Nithin Sivakumaran, Shoubin Yu, Hyunji Lee, Yue Zhang, Ali Payani, Mohit Bansal, and Elias Stengel-Eskin. Balancing faithfulness and performance in reasoning via multi- listener soft execution, 2026. URL https://arxiv.org/abs/2602.16154. Siyuan Song, Harvey Lederman, Jennifer Hu, and Kyle Mahowald. Privileged self-access matters for introspection in ai, 2025. URL https://arxiv.org/abs/2508.14802. Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. Language models donât always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36:74952â74965, 2023. Amelia C. Warden, Jessica K. Witt, Mengzhu Fu, and Michael D. Dodd. Overestimation of variability in ensembles of color value and size. Attention, Perception, & Psychophysics, 87(5):1579â1603, 2025. ISSN 1943-393X. doi: 10.3758/s13414-025-03098-3. URLhttps: //doi.org/10.3758/s13414-025-03098-3. Shengbin Yue, Ting Huang, Zheng Jia, Siyuan Wang, Shujun Liu, Yun Song, Xuanjing Huang, and Zhongyu Wei. Multi-agent simulator drives language models for legal intensive interaction, 2025. URL https://arxiv.org/abs/2502.06882. Rosie Zhao, Anshul Shah, Xiaoyu Zhu, Xinke Deng, Zhongyu Jiang, Yang Yang, Joerg Liebelt, and Arnab Mondal. On robustness and chain-of-thought consistency of rl-finetuned vlms, 2026. URL https://arxiv.org/abs/2602.12506. A Creating GCA The GCA dataset consists of two types of stimuli: object images and geometric shapes. Object stimuli were derived from the Visual CounterFact (VCF) dataset Golovanevsky et al. (2025) and were used to create two color conditions: a a canonical color prior condition, where objects were colored with their typical real-world color, and a counterfactual condition, where objects were colored with an atypical color. In addition, a control dataset of geometric shapes without strong semantic color associations was generated to isolate purely perceptual color thresholds. Together, these stimuli form three evaluation conditions: canonical object colors, counterfactual object colors, and shapes. The construction of the base image dataset involved several stages, including image retrieval, automated filtering, manual verification, and segmentation. An overview of this pipeline is shown in Figure 11. The following sections describe each step in detail. 11 Preprint. Under review. Figure 10: Left: Distribution of stimulus types per survey. Right: Example trial from the web-based human experiment. Participants viewed a colored outline image and selected the perceived object color from three options. After making a color decision, participants reported their certainty on a 10-point scale ranging from âvery uncertainâ to âvery certain.â Figure 11: Overview of the dataset construction and filtering pipeline. Starting from 493 objects in the VCF color split, outline drawings were retrieved via Google Custom Search, automatically scored using GPT-4o, manually verified and segmented using an OpenCV- based masking procedure. After filtering for image quality and mask stability, 220 base images remained for coloring and threshold manipulation. A.1 Object Base Images To construct the base images for objects in the GCA dataset, we began with the set of object categories defined in the VCF dataset. VCF provides objects together with their canonical and counterfactual color associations, which makes it a useful starting point for studying how visual evidence interacts with semantic color priors. However, we did not use the original images from that dataset. Instead, we used the 493 object labels from the âcolorâ split of VCF as search queries to retrieve new images. Rather than natural photographs, we specifically collected stylized black-and-white outline drawings using the Google Custom Search API. Outline images were chosen because natural photographs contain shading, texture, and background variation that can introduce visual confounds when applying controlled color manipulations. In contrast, outline drawings allow color to be added in a controlled and interpretable way. For each object label, we issued queries of the form âa black and white outline drawing of a [OBJECT]â. Initial image retrieval via the Google Custom Search API yielded up to five candidate outline drawings per object. After removing duplicates and known stock-image domains, the retrieval stage yielded 1,987 candidate images across 488 objects. Five objects were excluded because no suitable outline drawings meeting our quality criteria could be retrieved. The number of retrieved candidate images varied across objects: 190 objects had five candidates available, 194 had four, 64 had three, 29 had two, and 11 had only a single candidate image. Image Scoring and Filtering Automatic retrieval produced substantial noise, including low-quality sketches, watermarked graphics, filled silhouettes, and other non-outline il- 12 Preprint. Under review. lustrations. To address this, we implemented a multi-stage filtering pipeline combining automated scoring with manual verification. As a first step, we used an LLM-based evaluation to assess image suitability. Each candidate image was programmatically evaluated by GPT-4o to determine whether it constituted a clean black-and-white outline drawing of the target object. The model scored each image based on several criteria, including the absence of shading, the absence of pre-existing color, clear depiction of the target object, and minimal background clutter. These scores were used to rank candidate images for each object. In addition, the model was asked to estimate how many distinct objects were present in the image. Images containing more than one object were excluded to ensure that subsequent color attribution judgments would not be confounded by multiple foreground entities. For each object, we discarded images with a minimal score as well as all images depicting more than one object and retained only the highest-scoring candidate image per object. This automated filtering stage reduced the dataset to 271 candidate base images. Because outline drawings retrieved from web search vary considerably in style and cleanliness, we subsequently performed manual verification of all retained images and removed an additional 45 images. The resulting dataset contained 226 high-quality outline images prior to segmentation and coloring. Image Segmentation To isolate the object foreground from the white background, we implemented an automated segmentation pipeline using OpenCV-based masks. Foreground masks were generated using adaptive thresholding and contour-based filling. Segmentation parameters (e.g., block size, dilation, and morphological closing) were iteratively calibrated to maximize mask stability across diverse outline styles. Images producing degenerate masks (e.g., near-empty or near-full foreground occupancy) were automatically excluded. In some cases, small background regions enclosed by object contours remained inside the foreground mask. An example is shown in the top-right image of Figure 12. In this example, the enclosed white area inside the forklift outline is included in the foreground mask even though it does not belong to the object itself. Because these regions occupy only a small fraction of the total mask area, they were retained to preserve consistency across thresholds and simplify the masking procedure. Since coloring percentages were computed relative to the full foreground mask, these inclusions introduce only negligible variation in the effective colored area and do not systematically bias threshold estimates. Of the 226 images entering the segmentation stage, 6 were removed due to persistent mask instability after parameter refinement, yielding a final set of 220 base images used for coloring. The final foreground mask was stored as a binary image and used to define colorable pixels in subsequent color-manipulation stages. Image ResizingPrior to coloring, all images and their corresponding binary masks were padded to a square format and resized to a fixed resolution of 512Ă512 pixels. Image resizing was performed using interpolation (LANCZOS for images and nearest-neighbor for masks) to preserve structural boundaries and mask integrity. Standardizing the resolution ensures that coloring operates over a consistent number of pixels across images and that proportions of colorable area are comparable across stimuli. It also guarantees consistent alignment between the 16Ă16 recoloring patches and the image grid. These standardized images were then used in the color manipulation procedure described below. A.2 Color Selection Two color dataset variants were created from the set of 220 black-and-white object images: One colored with canonical color priors and one with counterfactual colorings. Both used the same coloring method and were restricted to the same colors. Manipulation Colors The set of manipulation colors used across all images was: red, brown, pink, green, blue, yellow, purple, orange, grey. 13 Preprint. Under review. Figure 12: Example results of the automatic mask generation procedure. For each object, the original outline image (left) is shown alongside the corresponding segmentation mask produced by the OpenCV-based pipeline (right). These colors were selected to cover a broad range of distinct basic color categories while remaining compatible with the HSV-based coloring procedure. Canonical Color Prior While initial experiments explored extracting model-specific se- mantic color priors, the final canonical color prior dataset was constructed using a single, fixed prior source to ensure comparability across models and human participants. For each object, we queried GPT-4o in an image-conditioned setting to estimate plausible real-world colors based on the grayscale outline image. The prompt instructed the model to list up to three likely colors for the depicted object, responding only with English color words. This image-conditioned approach required the model to rely on structural and silhouette cues present in the specific outline image to infer a color prior and to not only rely on world-knowledge. Responses were parsed into ranked color lists and normalized (lower- casing, punctuation removal, mapping variants such as âsilverâ to âgreyâ and âgoldâ to âyellowâ). From the remaining ranked list, a single primary color was selected and used as the canonical prior-consistent manipulation color. Objects for which no valid prior remained after filtering (e.g., responses such as âblackâ or âwhiteâ) were excluded for this part of the dataset. After this filtering stage, 199 objects with suitable base image as well as valid canonical color prior remained. Counterfactual Colors In addition to canonical color priors, counterfactual color condi- tions were derived from the Visual CounterFact dataset Golovanevsky et al. (2025). For each object, VCF provides anincorrectcolorattribute representing a semantically implausible or atypical color assignment for that object. These predefined counterfactual colors were used to construct the incongruent coloring condition. Objects lacking a valid counterfactual color after vocabulary alignment were excluded from this part of the dataset, resulting in 215 objects available for counterfactual coloring. This subset is slightly larger than the set of objects used for prior-consistent coloring because the canonical color extraction step produced more invalid or ambiguous responses (e.g., black or white), whereas the VCF dataset already provides explicit counterfactual color assignments. A.3 Coloring Method Sequential coloring Color manipulations were applied using a custom HSV-based col- oring pipeline. Coloring was restricted to non-outline pixels within either the foreground (object) mask or the background region, depending on condition. Black outline pixels and 14 Preprint. Under review. very dark pixels were explicitly excluded from coloring. The colored pixel percentageÏ was always computed relative to the number of eligible colorable pixels only. Two coloring modes were implemented during development: independent and sequential sampling. The final dataset was generated using sequential coloring. In sequential mode, colored pixels accumulate monotonically across increasingÏvalues. That is, for thresholdsÏ 1 <Ï 2 , the colored pixel set atÏ 1 is a subset of the colored pixel set atÏ 2 . Thus, higherÏconditions strictly contain all color evidence from lowerÏconditions, ensuring that color evidence increases monotonically rather than being spatially resampled. Patch-wise coloring The coloring was performed in a patch-wise fashion using 16Ă 16 pixel patches, which leads to a coloring granularity of approx 0,1% on the full 512 x 512 images. Eligible colorable pixels were grouped by spatial patch index. Patches were randomly ordered under a fixed random seed for reproducibility, and entire patches were colored until the target pixel countKwas reached. Because patches are colored as contiguous units, the realized proportion of colored pixels may slightly exceed the exact target percentage when the final patch is applied but remained within rounding tolerance of the intendedÏvalue. Patch-wise coloring was chosen over coloring individual pixels to better align the manipulation with the way of image processing of modern VLMs. Dark pixel percentage Black outline pixels were explicitly excluded from coloring. The majority of grey structural pixels were colored via Hue blending. Coloring percentages were always computed relative to the number of eligible colorable pixels only. To quantify the structural composition of the original black-and-white stimuli, we additionally measured the proportion of non-white pixels within the foreground object mask prior to coloring. Non-white pixels were defined as pixels with grayscale intensity below a conservative white threshold, thereby capturing both black outline strokes and internal grey linework. Across all object stimuli, the mean proportion of non-white pixels was 34.06% (median = 30.85%, SD = 14.71%). For the subset of objects colored with the GPT-4o color prior this mean proportion was 34.20% (median = 31.09%, SD = 14.72%) and for the subset of evaluated objects with counterfact coloring it was 33.67% (median = 30.58%, SD = 14.32%). This indicates that approximately one third of object pixels contained structural (non-white) content prior to coloring. A.4 Shape Dataset To disentangle semantic color priors from perceptual color thresholds, we constructed a control dataset consisting of simple geometric shapes without strong real-world color associations. The shape set included circles, triangles, squares, pentagons, and hexagons. Unlike object images, which were retrieved via web search, shape stimuli were generated programmatically to ensure precise structural control. Each shape was rendered as a black outline on a white background, matching the visual format of the object stimuli. The shapes were centered within the image and scaled to occupy a comparable proportion of image area as the object outlines. To increase structural variability, we generated five transformed vari- ants of each shape using geometric transformations (e.g., rotation, scaling and translation), yielding a total of 25 unique shape images. These transformations preserved overall shape identity while introducing low-level variation in orientation and spatial configuration. Because the shapes were generated synthetically, segmentation masks were trivially defined based on the known shape. Since shapes do not possess canonical real-world color priors, coloring was applied using the same set of manipulation colors used for object stimuli. Thus, shape stimuli serve as a perceptual baseline condition against which object-specific prior effects can be evaluated. Each of the 25 shape images was therefore colored using all nine manipulation colors, which lead to 225 unique shapeâcolor instances in total. Coloring for shapes followed the identical resizing, masking, sequential accumulation, and 16Ă16 patch-wise procedure used for object stimuli. To compare object and shape stimuli, we also computed the proportion of non-white pixels within the foreground mask prior to coloring for shapes, using the same grayscale thresholding procedure as for objects. Because shapes consisted of black outline strokes 15 Preprint. Under review. without internal linework, the proportion of non-white pixels in the shape evaluation subset was substantially lower than for objects (mean = 14.35%, median = 13.30%, SD = 4.58%). This confirms that object images were structurally denser and contained more internal line detail than the minimal-outline shape stimuli. B Details on Human Trials with Prolific Technical Implementation The experiment was implemented as a custom web-based application using a Flask backend and a JavaScript front-end. Survey profiles were pre- generated as JSON files and served dynamically to participants. All responses and metadata were stored in a Supabase database via secure server-side logging. Upon entering the study, participants were assigned a pre-generated survey profile. Partici- pants were identified via a unique participant ID, and re-entry into the study was blocked once results had been submitted. Backward navigation was disabled to prevent revision of earlier responses after exposure to later stimuli or the introspection question. For each trial, the system recorded the participantâs response, response time, profile identi- fier, survey condition (introspection-first or introspection-last), and timestamp metadata. Pricing Because not all survey profiles were completed in the initial wave, a second recruitment wave was conducted to ensure that each profile was completed at least once. The survey structure and experimental conditions remained identical across recruitment waves. Participants recruited via Prolific were compensated an average of ÂŁ13.20 per hour. Billing for 20 minutes per survey, the total pricing of our human trials cost approximately ÂŁ1000.00 B.1 Further Analysis for Human Trials Table 1: Per-Percent-Colored Significance Tests by Task Type AllCounterfactualPriorShape %ZPZPZPZP 00.0140.99â0.0100.990.0100.99 5 -3.30.0010-0.650.52-1.50.13-3.10.0020 10-4.20.000-1.00.31-2.40.015-3.40.0010 20-4.70.000-2.40.015-2.30.024-3.50.000 30-5.90.000-2.60.011-3.40.0010-4.30.000 40-5.90.000-1.40.16-4.10.000-4.30.000 50 -4.60.000-0.680.50-3.20.0010-3.50.000 55-1.50.14-0.750.46-0.610.54-1.20.23 60-0.0140.99â-1.00.310.570.57 70 -0.0100.991.00.32â-1.00.31 80-0.0100.99â-0.0100.99 900.990.32â0.990.32â 100-0.460.64â-0.0140.99-1.00.31 Table 2: Results from two-proportion z-tests examining whether color response proportions differ sig- nificantly between the introspection first vs. introspection last group for human participants, stratified by percent-colored condition and dataset split. Dashed lines indicate that NaN. For counterfactual, there is no â0%â colored split, for higher color thresholds, all responses for both groups of participants are colors which leads to NaN values in z-test calculations. The exact wording was: For any object, x% of its pixels should be colored for it to be considered that color. At what point would you personally say that the object in the image is the given color? What value should x% be? C Full Results for VLMs 16 Preprint. Under review. Figure 13: Proportion of âcolor â responses as a function of color threshold in GCA. Error bars show SEM. The figure depicts all models, stimulus types and Chain-of-Thought variants (Standard CoT, Visual CoT, Post-hoc CoT & Text CoT). 17 Preprint. Under review. Figure 14: Value of VLM stated thresholds for all models, stimulus types and Chain-of- Thought variants (Standard CoT, Visual CoT, Post-hoc CoT & Text CoT). 18 Preprint. Under review. Figure 15: VLM faithfulness to introspective rules for all models, stimulus types and over all Chain-of-Thought variants (Standard CoT, Visual CoT, Post-hoc CoT & Text CoT). 19