Paper deep dive
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
Jingyu Sun, Jiachen Tu, Yuyang Xue, Yaoxin Jiang, Guoyi Xu, Zhengtao Yao, Rui Qian, Yizheng Sun, Hongpeng Zhou, Jingyuan Sun, Yan Lin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/21/2026, 2:41:14 AM
Summary
The paper introduces PriVE-Bench and PriVE-Tools to evaluate whether Vision-Language Models (VLMs) rely on visual evidence or learned priors. PriVE-Bench uses paired original and counterfactual images to distinguish grounded answers from prior-consistent errors. PriVE-Tools tests if agentic visual evidence (bounding boxes, crops, zoom panels, contours) helps VLMs overcome prior-following biases. Results indicate that while tools can help in some settings, they are not a universal remedy, as many models continue to follow language and category priors even when explicit visual evidence is provided.
Entities (15)
Relation Signals (7)
PriVE-Bench → evaluates → Vision-Language Models
confidence 95% · We evaluate PriVE-Bench and PriVE-Tools on closed- and open-source VLMs
PriVE-Bench → uses → Counterfactual Images
confidence 95% · PriVE-Bench uses paired original and counterfactual images to distinguish visually grounded answers from prior-consistent errors.
PriVE-Tools → isextensionof → PriVE-Bench
confidence 93% · We further introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived visual evidence ... improves grounding under the same counterfactual conflicts.
PriVE-Tools → evaluates → Visual Evidence Types
confidence 92% · PriVE-Tools evaluates whether tool-derived visual evidence -- including bounding boxes, crops, zoom panels, and contours -- improves grounding
Vision-Language Models → exhibits → Prior-Following Behavior
confidence 90% · Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself.
GPT-5 → isevaluatedon → PriVE-Bench
confidence 85% · The closed-source models include GPT-5 ... We evaluate PriVE-Bench and PriVE-Tools on closed- and open-source VLMs
Qwen3-VL-8B-Instruct → isevaluatedon → PriVE-Bench
confidence 85% · The open-source models include Qwen3-VL-8B-Instruct ... We evaluate PriVE-Bench and PriVE-Tools on closed- and open-source VLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure mode: when visible evidence contradicts what is usually true, a grounded model should answer from the pixels, while a prior-following model will produce a canonical but visually incorrect response. However, existing counterfactual benchmarks mainly ask whether such prior-following behavior exists. In this paper, we ask a further question motivated by the rise of tool-augmented and agentic vision systems: can additional visual evidence views help VLMs reason against their priors? We introduce PriVE-Bench, a Prior-vs-Visual Evidence Benchmark that uses paired original and counterfactual images to distinguish visually grounded answers from prior-consistent errors. We further introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived visual evidence -- including bounding boxes, crops, zoom panels, and contours -- improves grounding under the same counterfactual conflicts. Across open- and closed-source VLMs, we compare raw, paired-image, and tool-conditioned inputs using accuracy, prior-following error rate, and other-response rate. Our results show that visual evidence tools can help in some settings, especially when models can use localized evidence effectively, but they are not a universal remedy: several models continue to follow language and category priors even when relevant visual evidence is explicitly provided.
Tags
Links
- Source: https://arxiv.org/abs/2607.16311v1
- Canonical: https://arxiv.org/abs/2607.16311v1
Trouble viewing inline? Open PDF directly →
Full Text
83,916 characters extracted from source content.
Expand or collapse full text
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Jingyu Sun1,2,* Jiachen Tu3,* Yuyang Xue4 Yaoxin Jiang3 Guoyi Xu3 Zhengtao Yao5 Rui Qian6 Yizheng Sun1 Hongpeng Zhou1 Jingyuan Sun1 Yan Lin7,† 1The University of Manchester 2The University of Melbourne 3University of Illinois at Urbana-Champaign 4University of Edinburgh 5University of Southern California 6Fudan University 7University of Newcastle *Equal contribution † author Abstract Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predictions in the image itself. Counterfactual images provide a natural diagnostic setting for this failure mode: when visible evidence contradicts what is usually true, a grounded model should answer from the pixels, while a prior-following model will produce a canonical but visually incorrect response. However, existing counterfactual benchmarks mainly ask whether such prior-following behavior exists. In this paper, we ask a further question motivated by the rise of tool-augmented and agentic vision systems: can additional visual evidence views help VLMs reason against their priors? We introduce PriVE-Bench, a Prior-vs-Visual Evidence Benchmark that uses paired original and counterfactual images to distinguish visually grounded answers from prior-consistent errors. We further introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived visual evidence—including bounding boxes, crops, zoom panels, and contours—improves grounding under the same counterfactual conflicts. Across open- and closed-source VLMs, we compare raw, paired-image, and tool-conditioned inputs using accuracy, prior-following error rate, and other-response rate. Our results show that visual evidence tools can help in some settings, especially when models can use localized evidence effectively, but they are not a universal remedy: several models continue to follow language and category priors even when relevant visual evidence is explicitly provided. The benchmark and codebase can be accessed at counterfactual-vlm-benchmark. Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Jingyu Sun1,2,* Jiachen Tu3,* Yuyang Xue4 Yaoxin Jiang3 Guoyi Xu3 Zhengtao Yao5 Rui Qian6 Yizheng Sun1 Hongpeng Zhou1 Jingyuan Sun1 Yan Lin7,† 1The University of Manchester 2The University of Melbourne 3University of Illinois at Urbana-Champaign 4University of Edinburgh 5University of Southern California 6Fudan University 7University of Newcastle *Equal contribution † author 1 Introduction Figure 1: Motivation of PriVE-Bench and PriVE-Tools. Top: Counterfactual images reveal whether VLMs answer from visible evidence or learned language/category priors. Bottom: PriVE-Tools tests whether controlled agentic-vision-inspired evidence views help models recover from prior-following errors. Vision-language models (VLMs) have become a central interface for visual understanding, enabled by large-scale image-text alignment, increasingly capable language backbones, and visual instruction tuning (Radford et al., 2021; Alayrac et al., 2022; Li et al., 2023; Liu et al., 2023a). However, the same training paradigm that gives VLMs broad semantic coverage can also encourage reliance on language-mediated visual associations. Familiar categories are strongly correlated with canonical attributes: camels usually have humps, birds usually have two wings, and basketballs are usually orange. When such priors agree with the image, high accuracy may be indistinguishable from genuine visual grounding. Counterfactual images make this distinction testable. If an image shows an object, count, attribute, or local medical region that violates a familiar prior, a visually grounded model should answer according to what is actually present in the pixels. A prior-following model may instead report what is typically true of the category, even when that answer is visibly wrong, as illustrated in Figure 1. This failure mode is related to long-standing language-prior and unimodal-shortcut problems in visual question answering (Antol et al., 2015; Agrawal et al., 2018; Cadene et al., 2019), and recent benchmarks show that modern VLMs still exhibit prior-following behavior under counterfactual or out-of-distribution visual conditions (Lee et al., 2025; Luo et al., 2024; Vo et al., 2025). Our goal is not simply to re-establish that VLMs can be biased by language or category priors. Instead, we use counterfactual visual conflicts as a controlled testbed for a newer question raised by tool-augmented and agentic vision systems. Recent multimodal systems increasingly use localization, segmentation, visual marking, cropping, zooming, and iterative inspection to expose task-relevant visual evidence before answering (Yao et al., 2022; Yang et al., 2023b; Kirillov et al., 2023; Yang et al., 2023a; Wu et al., 2024; Zhang et al., 2023; Wu and Xie, 2024; Liu et al., 2024). These operations are often motivated by the intuition that better visual access should improve grounding. Yet making relevant evidence available does not guarantee that a VLM will use it to revise a prior-driven answer. This motivates our central question: which forms of agentic visual evidence, if any, help VLMs answer from what is actually visible? We introduce PriVE-Bench, a Prior-vs-Visual Evidence Benchmark for evaluating whether VLMs answer from visible evidence or learned language and category priors. PriVE-Bench contains paired original and counterfactual images across five domains: object existence, counting, fashion attributes, industry/common-object attributes, and medical modality consistency. Each example defines a visually correct answer and a prior-consistent biased answer, enabling outputs to be categorized as correct, biased, or other. The benchmark supports yes/no, multiple-choice, and open-ended questions under original-only, counterfactual-only, and paired-image input modes. We further introduce PriVE-Tools, an agentic-vision-inspired extension of PriVE-Bench. Rather than proposing a new autonomous agent or evaluating unconstrained tool use, PriVE-Tools converts common visual operations into controlled evidence conditions, including bounding-box overlays, crops, zoom panels, and contour overlays. By keeping the underlying image, question, correct answer, biased answer, and scoring rubric fixed while varying only the evidence view, PriVE-Tools isolates whether the evidence view itself helps models reduce prior-following errors. We evaluate PriVE-Bench and PriVE-Tools on closed- and open-source VLMs using accuracy, prior-following error rate, and other-response rate. Our results show that counterfactual inputs induce systematic prior-following errors rather than merely ambiguous failures. Paired images can help by making some interventions more visually comparable, but they do not reliably eliminate prior-following. Tool-derived evidence is similarly conditional: common tools such as crops and zoom panels improve grounding for the closed-source models in our evaluation, but have weak or negative aggregate effects for several open-source models. Overall, agentic visual evidence can expose relevant information, but current VLMs may still fail to use that evidence when it conflicts with learned priors. Our contributions are as follows: • We introduce PriVE-Bench, a controlled counterfactual vision-language benchmark for distinguishing visually grounded answers from prior-consistent errors across five domains. • We introduce PriVE-Tools, a controlled agentic-vision-inspired extension that evaluates whether tool-derived evidence views such as boxes, crops, zoom panels, and contours reduce prior-following behavior. • We evaluate open- and closed-source VLMs across raw, paired-image, and tool-conditioned settings, showing that visual evidence tools can help in some settings but do not reliably make VLMs reason against language and category priors. 2 Related Work Language priors and counterfactual VLM evaluation. Language priors have long been studied in visual question answering. VQA-CP exposes how models exploit question-answer correlations, while RUBi uses a question-only branch to reduce unimodal shortcuts (Agrawal et al., 2018; Cadene et al., 2019). Recent benchmarks show that similar failures persist in large VLMs. HallusionBench evaluates entangled language hallucination and visual illusion; MMStar curates vision-indispensable samples to reduce visual-content irrelevance and data leakage; and WHOOPS! uses synthetic, commonsense-defying images to test visual commonsense reasoning (Guan et al., 2024; Chen et al., 2024; Bitton-Guetta et al., 2023). More directly related to our setting, VLind-Bench, ViLP, and Vision Language Models are Biased use counterfactual or out-of-distribution visual conditions to probe whether VLMs rely on language and category priors rather than image evidence (Lee et al., 2025; Luo et al., 2024; Vo et al., 2025). Words or Vision further studies modality conflicts, showing that VLMs may disproportionately trust text when textual and visual signals disagree (Deng et al., 2025). PriVE-Bench builds on this line of work by explicitly separating visually correct, prior-consistent, and ambiguous outputs over paired original/counterfactual images. More importantly, PriVE-Tools extends the question from whether prior-following exists to whether controlled tool-derived visual evidence can reduce it. Agentic visual evidence and tool-augmented MLLMs. A growing line of work augments language and multimodal models with tools, visual prompts, or iterative inspection mechanisms. ReAct introduces interleaved reasoning and action for language models, and multimodal systems such as M-REACT extend this idea by coordinating external vision experts for complex visual tasks (Yao et al., 2022; Yang et al., 2023b). In parallel, visual prompting methods expose localized evidence to VLMs: Segment Anything provides promptable segmentation masks, Set-of-Mark overlays marks, boxes, and masks to improve visual grounding, and visual prompting surveys organize a broad range of visual annotations and prompt-generation strategies (Kirillov et al., 2023; Yang et al., 2023a; Wu et al., 2024). Other methods use cropping or search to help models inspect fine details, including ViCrop, V*, and Chain-of-Spot (Zhang et al., 2023; Wu and Xie, 2024; Liu et al., 2024). These works motivate the evidence views evaluated in PriVE-Tools, such as boxes, crops, zoom panels, and contours. However, prior work mainly studies whether such tools improve perception, grounding, or VQA performance in general. PriVE-Tools instead evaluates these evidence views under controlled counterfactual conflicts, asking whether they help models answer against language and category priors. Figure 2: Representative PriVE-Bench examples and PriVE-Tools visual evidence. Rows correspond to the five benchmark domains. Columns show the original image, counterfactual image, and three visual-evidence views: localization, crop, and zoom panel. Figure 3: Dataset Generation and Question-Format Design Example for PriVE-Bench and PriVE-Tools. The same counting counterfactual image is evaluated with three question formats: yes/no, multiple choice, and open-ended response. All formats share the same target attribute, visible wing count, and the same prior conflict. 3 PriVE-Bench: Counterfactual Benchmark Construction Benchmark unit. PriVE-Bench is a controlled counterfactual vision-language benchmark for testing whether VLMs answer from visible evidence or from learned language and category priors. Similar to recent counterfactual diagnostic benchmarks for multimodal models (Bitton-Guetta et al., 2023; Lee et al., 2025; Vo et al., 2025), its core unit is an original/counterfactual image pair. The original image depicts an ordinary instance that conforms to common visual expectations, while the counterfactual image applies a targeted intervention that violates a canonical expectation associated with the object, count, attribute, or imaging context. This paired construction supports the evaluation protocol in Section 5: the original image provides an answerability check, the counterfactual image creates a prior–evidence conflict, and the pair enables direct visual comparison. Domains and sources. PriVE-Bench spans five domains: object existence, counting, fashion attributes, industry/common-object attributes, and medical modality consistency. These domains cover expected object parts, visible object-part or instance counts, logo and monogram layouts, canonical common-object conventions, and tumor-region MRI modality consistency. For the non-medical domains, images are drawn from a mixture of web-collected and AI-generated sources; generated images are produced or edited with GPT Image 1.5 and Gemini 2.5 Flash Image depending on the domain and intervention type. This mixture balances natural visual diversity with controllable counterfactual editing, without assuming that either web-collected or generated images are guaranteed unseen or unbiased. Each instance records source metadata, intervention type, target attribute, visually correct answer, and prior-consistent biased answer. Detailed source composition, AI-generated proportions, and per-domain statistics are reported in Appendix A. Medical modality consistency. The medical-modality domain is constructed separately from the natural-image domains using the ASNR-MICCAI-BraTS2023-GLI Challenge training data, following the BraTS multimodal brain tumor segmentation benchmark setting (Menze et al., 2014; Baid et al., 2021). Because BraTS provides co-registered MRI modalities and tumor segmentation labels, we construct modality-consistency counterfactuals by replacing the tumor-region intensities in one sequence with the corresponding tumor-region intensities from another. The resulting image has surrounding brain tissue that indicates one modality, while the tumor region carries the local contrast pattern of another. The task is therefore not clinical diagnosis or tumor typing, but visual consistency judgment: a grounded model should detect the tumor-region/background conflict, while a prior-following model may answer according to the dominant global modality. Appendix B provides the full slice-selection, modality-swap, and stratified sampling protocol. Task schema. For each benchmark instance, PriVE-Bench defines yes/no, multiple-choice, and open-ended questions. Yes/no questions test direct binary judgments, multiple-choice questions contrast a visually correct answer with a prior-consistent distractor, and open-ended questions test whether models can describe or infer the visible evidence without fixed options. The benchmark is evaluated under three raw input modes: original-only, counterfactual-only, and paired-image. Model responses are later mapped to correct, biased, or other under the protocol in Section 5, enabling PriVE-Bench to distinguish visually incorrect prior-following errors from ambiguous or unparseable responses. Figure 2 shows representative examples, and Figure 3 illustrates the raw task interface. 4 PriVE-Tools: Controlled Agentic Visual Evidence PriVE-Bench evaluates whether VLMs rely on visible evidence or language and category priors under raw-image inputs. PriVE-Tools extends this setting by asking a complementary question: if a model is given additional tool-derived visual evidence, does it become less prior-driven? Recent tool-augmented and agentic multimodal systems increasingly use operations such as localization, segmentation, visual marking, cropping, zooming, and iterative inspection to expose task-relevant image regions before answering (Yao et al., 2022; Yang et al., 2023b; Kirillov et al., 2023; Yang et al., 2023a; Wu et al., 2024; Zhang et al., 2023; Wu and Xie, 2024; Liu et al., 2024). PriVE-Tools does not propose a new agentic vision model. Instead, it converts these commonly used visual operations into fixed evidence conditions, isolating the effect of the evidence view itself from tool-selection policies, planning quality, search trajectories, or model-specific agent behavior. For each eligible counterfactual example, PriVE-Tools compares the raw-image baseline against controlled tool-derived evidence views, including bounding-box overlays, crops, zoom panels, and contour or outline overlays. These conditions are summarized in Appendix C. Available evidence types vary by domain because the relevant visual target differs: object-existence, fashion, and industry examples use localized views around manipulated attributes; counting examples additionally use outline-style evidence for countable instances; and medical-modality examples use tumor-region boxes, contours, crops, and zoom panels derived from BraTS segmentation labels. PriVE-Tools is designed as a controlled intervention over model input. For a given example, the underlying image, question, visually correct answer, prior-consistent biased answer, and scoring rubric remain fixed; only the evidence condition changes. The evidence views are intentionally non-semantic: they do not contain textual labels, edit descriptions, modality names, arrows explaining the manipulation, or the correct answer. They only alter how the relevant visual region is localized, cropped, or magnified. Thus, improvement under a tool condition suggests that the model benefits from better access to visual evidence rather than from explicit answer leakage. Evidence generation is domain-specific. For natural-image domains, evidence is generated from object-level localization or segmentation outputs and manually filtered to ensure that the highlighted region corresponds to the target attribute. For the medical-modality domain, evidence is derived directly from BraTS tumor segmentation labels rather than from generic natural-image segmentation models, avoiding domain mismatch on MRI images. Filled masks and difference maps are excluded from the main experiments because they may obscure the relevant visual signal or directly reveal the counterfactual manipulation. Section 5 formalizes how these evidence conditions are compared against matched raw-image baselines. 5 Evaluation Protocol Protocol overview. We evaluate PriVE-Bench and PriVE-Tools under a unified protocol across all domains. The goal is to measure whether a VLM answers from visible evidence or from a learned language/category prior, and whether paired images or tool-derived visual evidence changes this behavior. Metrics are computed over completed and successfully scored responses. API failures, server errors, empty responses after retry, and judge failures are tracked separately and excluded from the metric denominator rather than merged into the other category. Input conditions. PriVE-Bench uses three raw input modes. In original-only, the model receives only the original image, serving as an answerability check. In counterfactual-only, the model receives only the counterfactual image, making this the primary setting for measuring prior-following behavior. In paired-image, the model receives both related images with neutral labels and compares the target attribute without being told which image is original or counterfactual. PriVE-Tools compares raw-image inputs with fixed tool-derived evidence conditions, including bounding boxes, crops, zoom panels, and contours or outlines. The question, visually correct answer, prior-consistent biased answer, and scoring rubric remain fixed across raw and tool conditions; only the visual input changes. Tool definitions are summarized in Appendix C. Prompting and scoring. Formal runs use a shared prompt policy and deterministic decoding where supported. In the main prior-conflict setting, prompts make the relevant canonical prior salient before asking the visual question; this stress-tests whether models still ground their answers in visible evidence when the prior is explicit. Default decoding, structured-output, reasoning-control, missing-evidence, and error-handling settings are reported in Appendix D. Multiple-choice responses are scored deterministically by matching the predicted option against the predefined visually correct and prior-consistent options. Yes/no and open-ended responses are scored by a fixed text-only judge within each model group following rubric-based LLM-as-judge protocols (Liu et al., 2023b; Zheng et al., 2023). The judge receives the question, model response, and question-specific correct/biased rubrics, but not the image, and returns one of three labels: correct, biased, or other. Judge model choices are reported in Appendix D. Labels and metrics. A correct response matches the visible evidence in the provided image or image pair. A biased response follows the canonical category, count, brand, object, or modality prior despite contradicting the counterfactual evidence. An other response is ambiguous, off-target, contradictory, noncommittal, unparseable, or not clearly mappable to either class. For a run r, let rD_r be the set of successfully scored responses and let yi∈corr,bias,othery_i∈\corr,bias,other\ be the assigned label for instance i. For each label y, we define py(r)=1|r|∑i∈r[yi=y].p_y(r)= 1|D_r| _i _r1[y_i=y]. (1) The reported metrics are Acc(r) (r) =pcorr(r), =p_corr(r), (2) PFER(r) (r) =pbias(r), =p_bias(r), (3) Other(r) (r) =pother(r). =p_other(r). (4) Here, PFER denotes Prior-Following Error Rate. It is most diagnostic in counterfactual settings, where the visually correct answer and the prior-consistent answer diverge. Matched comparisons. For paired-image evaluation, we compare the paired-image run rpairr_pair against the raw counterfactual-only baseline rcfr_cf. For PriVE-Tools, each evidence condition t is compared against a matched raw-image run rrawr_raw under the same model, domain, input mode, question type, prompt policy, quality-control filter, and sample subset. In the main cross-domain tool analysis, rrawr_raw is the corresponding raw counterfactual-only run. For any metric m∈Acc,PFER,Otherm∈\Acc,PFER,Other\, we compute Δpairm _pairm =m(rpair)−m(rcf), =m(r_pair)-m(r_cf), (5) Δtm _tm =m(rt)−m(rraw). =m(r_t)-m(r_raw). (6) A useful intervention should ideally increase AccAcc and decrease PFERPFER. We also track ΔOther because an intervention may reduce prior-following errors by increasing uncertainty rather than by improving visual grounding. When tool evidence is unavailable, examples are skipped before inference; raw-vs-tool deltas are reported on matched evidence-valid subsets whenever denominators differ. 6 Experiments and Analysis Experimental Setup. We evaluate PriVE-Bench and PriVE-Tools on closed- and open-source VLMs under the protocol in Section 5. The closed-source models include GPT-5 and Claude Sonnet 4.111We initially considered Gemini 3.1 Pro Preview, but exclude it from the formal comparison because available API quotas did not allow complete matched evaluation across domains, input modes, and tool conditions. The open-source models include Qwen3-VL-8B-Instruct, Qwen3-VL-32B-Instruct, InternVL3.5-8B, InternVL3.5-38B, Gemma-3-12B-it, and Gemma-3-27B-it (Bai et al., 2025; Wang et al., 2025; Kamath et al., 2025). Unless otherwise stated, results are macro-averaged across domains, and tool comparisons are computed on matched evidence-valid subsets. Full raw, paired-image, and tool-conditioned results are provided in Appendices E–H. Model Orig. CF Acc CF PFER CF Other Pair Δ A/P Top Common Tool Tool Δ A/P/O Closed-source VLMs GPT-5 78.3 46.7 51.3 2.0 +7.8 / -8.2 Crop +4.5 / -4.3 / -0.2 Claude Sonnet 4 70.7 46.6 51.0 2.3 +0.7 / -1.7 Zoom +11.8 / -10.6 / -1.1 Open-source VLMs Qwen3-VL-8B-Instruct 81.3 31.6 66.2 2.2 +10.1 / -18.7 Zoom -1.4 / +1.2 / +0.3 Qwen3-VL-32B-Instruct 83.7 40.3 58.8 0.9 +9.5 / -14.7 Crop -2.2 / +2.9 / -0.7 InternVL3.5-8B 79.8 27.5 68.6 3.8 +4.4 / -3.3 Zoom +0.5 / -5.8 / +5.2 InternVL3.5-38B 73.1 37.5 59.7 2.7 +4.0 / -4.2 Crop -5.7 / -2.8 / +8.5 Gemma-3-12B-it 60.9 39.7 53.9 6.4 -8.0 / +6.9 Zoom -5.2 / +6.6 / -1.5 Gemma-3-27B-it 68.2 30.5 61.1 8.4 +11.4 / -9.4 Zoom -0.7 / +2.8 / -2.1 Table 1: Main PriVE-Bench and PriVE-Tools results, macro-averaged across domains. All values are percentages or percentage-point deltas. Orig. is raw original-only accuracy; CF denotes raw counterfactual-only results. Pair Δ A/P reports Δ /Δ relative to counterfactual-only inputs. Top Common Tool is selected by the largest macro-averaged accuracy gain among common tool conditions evaluated across all domains. Tool Δ A/P/O reports Δ /Δ /Δ relative to the matched raw counterfactual baseline. RQ1: Do VLMs follow priors under counterfactual inputs? We first examine whether VLMs follow learned priors when visual evidence contradicts canonical expectations. Original-only accuracy serves as an answerability check, while counterfactual-only accuracy, PFER, and Other Rate diagnose whether failures are visually incorrect, prior-consistent, or ambiguous. Table 1 reports macro-averaged results, with full domain-level results in Appendix E. Counterfactual inputs substantially reduce accuracy and shift errors toward prior-following. GPT-5 drops from 78.3% original-only accuracy to 46.7% counterfactual-only accuracy, while Claude Sonnet 4 drops from 70.7% to 46.6%. Open-source models show similar or larger gaps; Qwen3-VL-8B-Instruct and InternVL3.5-8B drop by 49.7 and 52.3 percentage points, respectively. More importantly, PFER is much larger than Other Rate for nearly all models, indicating that models usually commit to prior-consistent answers rather than refusing or producing ambiguous responses. Counting and fashion are the strongest stress tests, reaching 77.9% and 68.4% PFER averaged over models. Thus, RQ1 shows that counterfactual images expose a systematic prior-following failure mode, not merely a drop in generic visual accuracy. RQ2: Do paired images help models reason against priors? We next evaluate whether direct visual comparison helps models reason against priors. In the paired-image setting, the model receives both the original and counterfactual image with neutral labels and must compare the target visual attribute. Table 1 reports Δ /Δ relative to the counterfactual-only baseline, with full paired-image results in Appendix F. Paired images help many models, but only partially. Seven of the eight evaluated models improve in macro accuracy: GPT-5 gains 7.8 points while reducing PFER by 8.2 points, and Gemma-3-27B-it, Qwen3-VL-8B-Instruct, and Qwen3-VL-32B-Instruct show the largest accuracy gains. However, the effect is not uniform. Claude Sonnet 4 changes little, and Gemma-3-12B-it becomes worse, with Δ of −8.0-8.0 and Δ of +6.9+6.9. Domain-level results also reveal a split: counting, fashion, industry, and medical modality benefit from paired comparison, whereas object existence becomes harder. Thus, paired images can make counterfactual changes more salient, but they also introduce a comparison burden and do not reliably eliminate prior-following behavior. RQ3: Do agentic visual evidence tools improve grounding? We then test whether tool-derived visual evidence improves grounding relative to raw counterfactual inputs. For the main cross-domain comparison, we focus on the two common evidence conditions available across all five domains: Crop and Zoom panel. A useful evidence condition should increase accuracy while reducing PFER; we also track Other Rate because a tool may reduce prior-following errors by increasing uncertainty rather than by improving visual grounding. As shown in Table 1 and Appendix G, the aggregate effect across all models is mixed. Crop slightly decreases accuracy by 0.9 points and slightly increases PFER by 0.2 points, while Zoom panel has nearly no effect on accuracy, with Δ of −0.2-0.2, and reduces PFER by only 1.1 points while increasing Other Rate by 1.3 points. This aggregate masks a clear model-family split: for the two closed-source VLMs, both common tools improve accuracy and reduce PFER, whereas the same views are weak or negative on average for open-source models. Thus, RQ3 shows that agentic visual evidence can help some models, but exposing localized or magnified evidence is not sufficient by itself. RQ4: Which tools help which models and domains? Finally, we analyze whether tool effectiveness depends on the model and domain. The Top Common Tool column in Table 1 summarizes the best common evidence view for each model, while Appendix H provides the full domain- and model-level breakdown. We focus on displayed evidence views that are directly interpretable across the benchmark: bounding-box overlays, crops, and zoom panels. Because Fashion and Industry bbox runs were not completed for all open-source models, those cells are left blank rather than treated as complete evidence. Tool effectiveness varies substantially across domains. Industry/common-object attributes benefit most clearly from localized evidence: Crop improves accuracy by 4.4 points and reduces PFER by 4.1 points. Medical modality consistency also benefits from Zoom panel evidence, with Δ of +5.3+5.3 and Δ of −10.5-10.5, although its 5.2-point increase in Other Rate suggests that some responses shift from prior-following to uncertainty. In contrast, object existence, counting, and fashion show weak or negative gains under displayed evidence views. Counting is especially revealing: even the best displayed tool, Zoom panel, decreases accuracy by 3.1 points and increases PFER by 1.0 point, indicating that localization or magnification alone does not solve enumeration. The model-level breakdown shows the same specificity. Closed-source models use tool evidence more consistently, while many open-source models do not reliably convert localized evidence into grounded answers. For example, Qwen3-VL-8B-Instruct’s best common tool still yields negative accuracy gain, and InternVL3.5-38B illustrates a different failure mode: Crop reduces PFER but decreases accuracy and increases Other Rate, suggesting a shift from biased answers toward uncertainty rather than correct grounding. Overall, RQ4 shows that PriVE-Tools does not identify a universally best visual evidence view; tool usefulness depends on both the counterfactual conflict and the model’s ability to use localized evidence. 7 Conclusion We introduced PriVE-Bench, a controlled counterfactual benchmark for distinguishing visually grounded answers from prior-consistent errors, and PriVE-Tools, an agentic-vision-inspired extension for evaluating whether tool-derived visual evidence helps VLMs reason against learned priors. Our experiments show that VLMs often follow language and category priors when counterfactual evidence contradicts canonical expectations. Paired images and tool-derived evidence can improve grounding in some settings, but their benefits are uneven across models, domains, and evidence types. Overall, our findings highlight a key challenge for agentic vision systems: exposing additional visual evidence does not guarantee that a VLM will use it to answer from what is actually visible. 8 Limitations PriVE-Tools is not a complete agentic vision framework. It evaluates controlled, pre-computed evidence views inspired by agentic systems, but does not measure adaptive tool selection, iterative inspection, planning, or self-correction. Some evidence is also partly oracle-like: boxes, crops, contours, and zoom panels are generated around known target regions, and medical evidence is derived from BraTS tumor segmentation labels. Our results therefore measure the usefulness of specific evidence views, not the end-to-end capability of autonomous visual agents. Our tool coverage is not uniform across all domains and models. The main cross-domain comparison focuses on common evidence views, such as crops and zoom panels, while some domain-specific or partially completed conditions, such as bounding boxes, contours, and outlines, are reported separately. As a result, PriVE-Tools should be interpreted as a controlled diagnostic evaluation of selected evidence views rather than an exhaustive study of all possible visual tools. Counterfactual construction may introduce artifacts. Non-medical examples can contain editing or generation artifacts, and AI-generated images may inherit the priors of the models used to create them. Web-collected images, in turn, may overlap with data seen during VLM pretraining. Although we apply filtering and quality control, we cannot fully rule out artifact-based behavior. The medical-modality task is a controlled modality-consistency test rather than a clinical diagnosis task. Tumor-region signal is swapped between co-registered T1N and T2F slices to create a local contrast conflict, which is useful for testing visual grounding but does not simulate real MRI acquisition or pathology. The main medical evaluation also uses a single 2D axial slice and a stratified subset of the constructed pool, limiting coverage of full 3D radiological context. Our model coverage is also limited. We evaluate representative closed- and open-source VLMs, but quota and cost constraints prevent a complete comparison across all current frontier systems. In particular, some models initially considered for evaluation could not be included in the formal matched comparison. Finally, yes/no and open-ended responses are partly scored by text-only LLM judges. Rubrics and deterministic multiple-choice scoring reduce ambiguity, but judge errors remain possible. The main setting also uses an explicitly prior-inducing prompt, making PriVE-Bench a stress test of visual grounding under salient priors rather than a complete characterization of all natural VLM interactions. References Agrawal et al. (2018) Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018. Don’t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4971–4980. Alayrac et al. (2022) Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, and 1 others. 2022. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736. Antol et al. (2015) Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision, pages 2425–2433. Bai et al. (2025) Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, and 1 others. 2025. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Baid et al. (2021) Ujjwal Baid, Satyam Ghodasara, Suyash Mohan, Michel Bilello, Evan Calabrese, Errol Colak, Keyvan Farahani, Jayashree Kalpathy-Cramer, Felipe C Kitamura, Sarthak Pati, and 1 others. 2021. The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification. arXiv preprint arXiv:2107.02314. Bitton-Guetta et al. (2023) Nitzan Bitton-Guetta, Yonatan Bitton, Jack Hessel, Ludwig Schmidt, Yuval Elovici, Gabriel Stanovsky, and Roy Schwartz. 2023. Breaking common sense: Whoops! a vision-and-language benchmark of synthetic and compositional images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2616–2627. Cadene et al. (2019) Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, and 1 others. 2019. Rubi: Reducing unimodal biases for visual question answering. Advances in neural information processing systems, 32. Chen et al. (2024) Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang, Yuhang Zang, Zehui Chen, Haodong Duan, Jiaqi Wang, Yu Qiao, Dahua Lin, and 1 others. 2024. Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37:27056–27087. Deng et al. (2025) Ailin Deng, Tri Cao, Zhirui Chen, and Bryan Hooi. 2025. Words or vision: Do vision-language models have blind faith in text? In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 3867–3876. Guan et al. (2024) Tianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian, Zongxia Li, Xiaoyu Liu, Xijun Wang, Lichang Chen, Furong Huang, Yaser Yacoob, and 1 others. 2024. Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14375–14385. Kamath et al. (2025) Gemma Team Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ram’e, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean-Bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gael Liu, and 191 others. 2025. Gemma 3 technical report. ArXiv, abs/2503.19786. Kirillov et al. (2023) Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, and 1 others. 2023. Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4015–4026. Lee et al. (2025) Kang-il Lee, Minbeom Kim, Seunghyun Yoon, Minsung Kim, Dongryeol Lee, Hyukhun Koh, and Kyomin Jung. 2025. Vlind-bench: Measuring language priors in large vision-language models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 4129–4144. Li et al. (2024) Hongwei Bran Li, Gian Marco Conte, Qingqiao Hu, Syed Muhammad Anwar, Florian Kofler, Ivan Ezhov, Koen van Leemput, Marie Piraud, Maria Diaz, Byrone Cole, and 1 others. 2024. The brain tumor segmentation (brats) challenge 2023: Brain mr image synthesis for tumor segmentation (brasyn). ArXiv, pages arXiv–2305. Li et al. (2023) Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR. Liu et al. (2023a) Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916. Liu et al. (2023b) Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023b. G-eval: Nlg evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing, pages 2511–2522. Liu et al. (2024) Zuyan Liu, Yuhao Dong, Yongming Rao, Jie Zhou, and Jiwen Lu. 2024. Chain-of-spot: Interactive reasoning improves large vision-language models. arXiv preprint arXiv:2403.12966. Luo et al. (2024) Tiange Luo, Ang Cao, Gunhee Lee, Justin Johnson, and Honglak Lee. 2024. Probing visual language priors in vlms. arXiv preprint arXiv:2501.00569. Menze et al. (2014) Bjoern H Menze, Andras Jakab, Stefan Bauer, Jayashree Kalpathy-Cramer, Keyvan Farahani, Justin Kirby, Yuliya Burren, Nicole Porz, Johannes Slotboom, Roland Wiest, and 1 others. 2014. The multimodal brain tumor image segmentation benchmark (brats). IEEE transactions on medical imaging, 34(10):1993–2024. Radford et al. (2021) Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, and 1 others. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR. Vo et al. (2025) An Vo, Khai-Nguyen Nguyen, Mohammad Reza Taesiri, Vy Tuong Dang, Anh Totti Nguyen, and Daeyoung Kim. 2025. Vision language models are biased. arXiv preprint arXiv:2505.23941. Wang et al. (2025) Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, and 1 others. 2025. Internvl3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265. Wu et al. (2024) Junda Wu, Zhehao Zhang, Yu Xia, Xintong Li, Zhaoyang Xia, Aaron Chang, Tong Yu, Sungchul Kim, Ryan A Rossi, Ruiyi Zhang, and 1 others. 2024. Visual prompting in multimodal large language models: A survey. arXiv preprint arXiv:2409.15310. Wu and Xie (2024) Penghao Wu and Saining Xie. 2024. V?: Guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13084–13094. Yang et al. (2023a) Jianwei Yang, Hao Zhang, Feng Li, Xueyan Zou, Chunyuan Li, and Jianfeng Gao. 2023a. Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v. arXiv preprint arXiv:2310.11441. Yang et al. (2023b) Zhengyuan Yang, Linjie Li, Jianfeng Wang, Kevin Lin, Ehsan Azarnasab, Faisal Ahmed, Zicheng Liu, Ce Liu, Michael Zeng, and Lijuan Wang. 2023b. Mm-react: Prompting chatgpt for multimodal reasoning and action. arXiv preprint arXiv:2303.11381. Yao et al. (2022) Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629. Zhang et al. (2023) Jiarui Zhang, Mahyar Khayatkhoei, Prateek Chhikara, and Filip Ilievski. 2023. Towards perceiving small visual details in zero-shot visual question answering with multimodal llms. arXiv preprint arXiv:2310.16033. Zheng et al. (2023) Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623. Appendix Appendix A Dataset Statistics and Source Composition Appendix A summarizes the dataset composition used in our formal evaluation. For non-medical domains, Originals denotes unique original images, while CF pairs denotes active counterfactual evaluation records. These counts are not always identical: a single original image may support multiple counterfactual interventions, such as modifying a logo layout, changing an object color, removing a visual attribute, or altering a countable part. We also report the source composition of original images. The non-medical domains combine web-collected or real images with AI-generated images. This mixed construction is intentional. Web-collected images preserve natural visual diversity and realistic object appearances, but may overlap with data seen during VLM pretraining. AI-generated images provide a more controllable source for constructing targeted counterfactual attributes, but may themselves reflect the visual priors of the image-generation models used to create them. We therefore record source metadata and report the AI-generated proportion for each non-medical domain, rather than treating either source type as fully unbiased or guaranteed unseen. The medical-modality domain is handled separately. It is derived from the ASNR-MICCAI-BraTS2023-GLI training data, following the BraTS multimodal brain tumor segmentation benchmark setting (Menze et al., 2014; Baid et al., 2021). Although we construct a full clean pool of modality-swapped counterfactual examples, the main experiments use a stratified 10% subset to reduce evaluation cost while preserving coverage across cases and swap directions. Therefore, Table 2 reports both the formal evaluation subset and the full constructed counterfactual pool. Figures 4 and 5 visualize the same metadata statistics. Domain Orig. Eval. CF Full CF Original source Notes Object existence 160 160 160 80 web / 80 AI (50.0%) Expected object parts removed Counting 178 256 256 128 web / 50 AI (28.1%) Some originals reused across edits Fashion 76 455 455 39 web / 37 AI (48.7%) Filtered logo and monogram edits Industry 58 348 348 30 web / 28 AI (48.3%) Filtered common-object edits Medical modality 250 250 2,502 250 BraTS MRI Stratified 10% subset of modality swaps Total 722 1,469 3,721 – – Table 2: Dataset composition used in PriVE-Bench. Counts are computed from active benchmark metadata rather than raw directory scans. Orig. denotes unique original images for non-medical domains and standardized MRI source views for the medical evaluation subset. Eval. CF denotes counterfactual records used in formal evaluation, while Full CF denotes all constructed counterfactual records before medical subsampling. AI percentages are computed only for non-medical originals. Figure 4: Original and counterfactual counts by domain in the formal evaluation set. Counts are computed from active benchmark metadata. Counting reuses some originals across multiple edits, while medical modality reports the stratified evaluation subset; the full constructed medical pool contains 2,502 counterfactual modality-swap records. Figure 5: Source composition of original images. Non-medical domains combine web-collected or real images with AI-generated images to balance natural visual diversity and controllable counterfactual construction. Medical examples are derived from BraTS2023 MRI data and are reported separately as dataset-derived medical images. Appendix B Medical Modality Sampling and Counterfactual Construction Data source. The medical-modality subset of PriVE-Bench is constructed from the ASNR-MICCAI-BraTS2023-GLI Challenge training data, following the BraTS multimodal brain tumor segmentation benchmark setting (Menze et al., 2014; Baid et al., 2021; Li et al., 2024). Each case contains co-registered multi-modal brain MRI volumes and a dense tumor segmentation label. We use the native T1-weighted MRI (T1N), T2-FLAIR MRI (T2F), and the provided segmentation mask. The segmentation mask is used only to define the tumor region for controlled counterfactual construction and, in PriVE-Tools, to generate localization-based visual evidence. It is not shown to the model as a semantic tumor label in the raw PriVE-Bench setting. Unlike the non-medical domains, the medical subset is not based on web-collected or text-to-image generated pictures. It is derived from real multi-modal MRI data in which different imaging sequences are spatially aligned for the same patient case. This property makes BraTS suitable for controlled modality-consistency counterfactuals: local tumor-region signal can be exchanged between modalities while preserving patient identity, slice location, tumor shape, and surrounding anatomy. Slice selection and standardized visualization. For each patient case, we select one axial slice for evaluation. The selected slice is the slice with the largest whole-tumor area, where the whole-tumor mask is defined as all non-zero BraTS segmentation labels. This avoids using slices in which the tumor is absent or too small to support reliable visual judgment. For the selected slice, we extract the corresponding T1N and T2F images. Each slice is normalized using the non-zero brain region, with intensities clipped and scaled by robust percentile statistics, and saved as a grayscale PNG. The main benchmark uses the clean visualization version: neither the original nor the counterfactual image contains a colored tumor overlay. This ensures that the task tests whether the model can infer modality consistency from visible MRI signal, rather than exploiting an explicit mask or annotation. Counterfactual modality swap. The core intervention is a local tumor-region modality swap. Let IT1NI^T1N and IT2FI^T2F denote the aligned T1N and T2F slices from the same patient case, and let M be the binary whole-tumor mask for the selected slice. We construct two directional counterfactuals: CT1N←T2F C^T1N 2F =(1−M)⊙IT1N+M⊙IT2F, =(1-M) I^T1N+M I^T2F, (7) CT2F←T1N C^T2F 1N =(1−M)⊙IT2F+M⊙IT1N. =(1-M) I^T2F+M I^T1N. (8) In the first direction, the image retains the global T1N background, but pixels inside the tumor mask are replaced with the corresponding T2F tumor-region pixels. In the second direction, the image retains the global T2F background, but the tumor region is replaced with the corresponding T1N tumor-region pixels. Because the modalities are co-registered, the replacement is spatially aligned and does not require an external generative model. Imaging-sequence rationale. This intervention creates a controlled conflict between global modality appearance and local tumor-region contrast. In a physically consistent MRI slice, the tumor region and surrounding tissue are acquired under the same imaging sequence, even though their intensities may differ due to tissue properties. Our counterfactual images deliberately violate this consistency: the surrounding brain tissue indicates one MRI sequence, while the tumor region carries the local contrast pattern of another sequence. The task is therefore not clinical diagnosis, tumor typing, or a realistic simulation of MRI acquisition. It is a controlled visual counterfactual for testing modality-consistency judgment. A visually grounded model should detect that the tumor-region contrast conflicts with the surrounding sequence. A prior-following model may instead answer according to the dominant global modality appearance and assume that the tumor region is consistent with the rest of the slice. Sampling protocol. The full constructed clean pool contains 1,251 BraTS2023 cases and two modality-swap directions per case, yielding 2,502 counterfactual records. For the main evaluation, we use a stratified 10% subset to reduce evaluation cost while preserving balance across swap directions. Sampling is performed before question expansion. The sampling strata are defined by dataset, visualization version, and swap direction. In the main setting, all samples come from BraTS2023 with the clean visualization, so the two effective strata correspond to the two swap directions: T1N-background/T2F-tumor and T2F-background/T1N-tumor. We sample 125 records from each direction, yielding 250 medical counterfactual records for formal evaluation. The random seed is fixed in the evaluation code, and the selected pair keys are recorded in the run configuration through a digest, enabling runs to be audited and reproduced. Label-derived visual evidence for PriVE-Tools. For PriVE-Tools, the medical subset uses label-derived visual evidence rather than natural-image segmentation outputs. This avoids applying a generic segmentation model to MRI images, where domain mismatch could introduce unstable or clinically meaningless masks. The BraTS tumor segmentation label provides a consistent localization source for constructing controlled evidence views. From the same whole-tumor mask, we generate four medical evidence conditions: bounding-box overlay, contour overlay, tumor crop, and zoom panel. The bounding box and contour provide thin localization cues without filling or recoloring the tumor region. The crop and zoom panel magnify the same tumor-region pixels while preserving the raw image context as part of the model input. We exclude filled mask overlays and difference maps from the main experiments because they could either obscure the tumor intensity pattern or directly reveal the intervention, turning the task into mask interpretation rather than visual evidence grounding. Appendix C PriVE-Tools Evidence Conditions PriVE-Tools evaluates controlled tool-derived evidence conditions inspired by common operations in agentic or tool-augmented vision systems. These conditions modify how the relevant visual region is presented to the model while keeping the underlying question, target attribute, correct answer, biased answer, and scoring rubric fixed. They do not include textual labels, edit descriptions, arrows indicating the manipulation, or explicit answer cues. Condition Purpose BBox Localizes the target object or region without changing image content. Crop Magnifies the relevant region and reduces irrelevant context. Zoom panel Preserves the full image while adding an enlarged local view. Contour Highlights a boundary without filling or recoloring the region. Tool bundle Provides multiple complementary evidence views together. Table 3: Tool-derived evidence conditions in PriVE-Tools. Available conditions vary by domain according to the target attribute and evidence source. For natural-image domains, evidence is generated from object-level localization or segmentation outputs and manually filtered where necessary. Counting examples may use contour-style evidence to highlight countable instances. Medical-modality examples use label-derived evidence from the BraTS tumor segmentation mask, avoiding generic natural-image segmentation models on MRI data. We exclude filled masks and difference maps from the main evaluation because they may obscure the relevant visual signal or directly reveal the counterfactual manipulation. Appendix D Evaluation Configuration Table 4 summarizes the default configuration used for formal PriVE-Bench and PriVE-Tools evaluations. These settings are chosen to reduce decoding randomness, standardize judge behavior, and make raw-image and tool-evidence comparisons as controlled as possible. Unless otherwise stated, provider-specific fallbacks are recorded in the run metadata. Setting Default Rationale Temperature 0 Uses deterministic decoding where supported, reducing run-to-run variation. Prior-conflict prime On Makes the relevant canonical prior explicit and consistent across models in the main evaluation setting. Backbone max output tokens 2048 Provides enough budget for short answers, formatting, and provider overhead without encouraging unrestricted long-form responses. Text-only judge GPT-4o-mini for closed-source VLM runs; Qwen3-8B for open-source VLM runs Uses a fixed text-only judge within each model group for rubric-based scoring of yes/no and open-ended responses. The judge receives the question, model response, and correct/biased rubrics, but not the image. Judge max output tokens 512 The judge only needs to return a label and short rationale, so a smaller budget is sufficient. Judge structured output Auto Requests JSON-style judge labels where supported, with fallback parsing for unsupported providers. Closed-form structured output Auto Requests constrained outputs for yes/no and multiple-choice questions where supported. Reasoning or extended thinking Off; minimal only if required Avoids enabling model-specific deliberation as an uncontrolled advantage. Missing evidence policy Skip Tool-condition examples without required evidence are skipped before inference. Error handling Excluded from denominator; reported separately API, server, quota, authentication, and judge failures are not treated as model answers. Table 4: Default evaluation configuration for formal PriVE-Bench and PriVE-Tools runs. Provider-specific deviations and fallbacks are recorded in run metadata. The text-only judge is used only to classify completed model responses according to the predefined correct/biased/other rubrics. Decoding and output budgets. Formal runs use temperature 0 wherever supported. When a provider does not expose exact deterministic decoding, we use the closest supported configuration and record the provider-specific settings. The backbone output budget is set to 2048 tokens. Although most benchmark answers are short, this budget reduces empty-output and output-limit failures for models that prepend formatting, safety text, or brief explanations. The judge budget is set to 512 tokens because the judge is only required to output a label and a short rationale. Prior-conflict prompting. The prior-conflict prime is enabled in the main formal evaluation. PriVE-Bench is designed to test whether models can ground their answers in visible evidence when a relevant language or category prior is explicitly salient. Enabling the same prior-inducing context across models makes the conflict controlled rather than relying on implicit prompt effects. Runs with different prompt policies are treated as separate configurations. Structured outputs and reasoning controls. Structured output is requested automatically where supported. For judged questions, structured output makes the judge label easier to parse and audit. For closed-form questions, it encourages models to return a constrained yes/no answer or option letter. If a provider does not support structured output, the runner falls back to text parsing and records the fallback reason. Reasoning or extended-thinking controls are disabled by default, because the benchmark is intended to compare visual grounding under the same input evidence rather than compare different levels of model-specific deliberation. If a provider requires a minimal reasoning setting, we use the minimum supported mode and record it. Missing evidence and matched comparisons. For PriVE-Tools, missing evidence is handled with a skip policy. If a required crop, bounding box, contour, outline, or zoom panel is unavailable, the example is not sent to the model. This avoids mixing raw-image behavior with incomplete tool-input behavior. When raw and tool runs have different eligible subsets, tool deltas should either report the denominator difference or be computed on the matched evidence-valid subset. Error handling. Infrastructure and provider failures are excluded from the metric denominator and reported separately. These include API quota errors, authentication failures, server startup failures, network errors, empty responses after retry, and judge failures. Such events are not merged into the other category, because other is reserved for completed model responses that are ambiguous, off-target, contradictory, or unparseable with respect to the task rubric. Appendix E Full Raw PriVE-Bench Results This appendix reports the raw-image results supporting RQ1. We focus on the two input modes most directly relevant to prior-following behavior: original-only and counterfactual-only. The original-only setting checks whether the unedited images and questions are answerable, while the counterfactual-only setting tests whether models answer from visible evidence when the image violates a canonical expectation. All metrics are computed over valid, successfully scored model responses. Query errors and judge errors are excluded from the metric denominator and reported separately in experiment logs. When runs are resumed after transient failures, duplicate task signatures are resolved by retaining the successful record when available. Unless otherwise stated, values are percentages. Counterfactual-only results. Table 5 reports raw counterfactual-only performance by model and domain. Each cell reports Accuracy / PFER / Other Rate. Accuracy measures visually correct responses, PFER measures prior-consistent errors, and Other Rate captures ambiguous, noncommittal, or otherwise uncategorizable responses. Model Existence Counting Fashion Industry Medical GPT-5 57.7/36.5/5.8 15.1/84.5/0.4 44.3/54.5/1.2 85.1/14.5/0.4 31.5/66.3/2.3 Claude Sonnet 4 62.3/36.2/1.5 19.5/73.7/6.8 42.6/56.0/1.3 67.3/32.3/0.4 41.5/56.8/1.7 Qwen3-VL-8B-Instruct 31.9/62.1/6.0 14.8/83.2/2.0 32.5/65.7/1.8 51.7/48.1/0.3 27.2/72.1/0.7 Qwen3-VL-32B-Instruct 74.4/23.1/2.5 13.0/85.8/1.2 27.3/72.6/0.1 53.6/46.1/0.3 33.1/66.4/0.5 InternVL3.5-8B 31.0/65.6/3.3 5.1/94.1/0.8 25.2/73.1/1.7 56.4/43.2/0.4 19.9/67.2/12.9 InternVL3.5-38B 53.1/45.8/1.0 4.0/85.0/10.9 27.0/71.9/1.1 65.9/33.7/0.4 37.7/62.3/0.0 Gemma-3-12B-it 46.7/44.6/8.8 36.8/52.7/10.4 18.7/74.4/6.8 29.7/68.4/1.9 66.8/29.3/3.9 Gemma-3-27B-it 65.6/32.5/1.9 27.6/63.9/8.5 18.5/78.7/2.9 21.6/77.4/1.0 19.2/53.1/27.7 Table 5: Raw counterfactual-only results by model and domain. Each cell reports Acc/PFER/Other, in percent. High PFER indicates prior-consistent errors under counterfactual visual evidence. Table 5 shows that counterfactual failures are often not random uncertainty. Across many models and domains, PFER is substantially larger than Other Rate, indicating that errors frequently preserve the learned prior rather than avoid commitment. Counting and fashion are especially strong stress tests, suggesting that canonical expectations about object counts and visual-design patterns can override edited visual evidence. Original-only sanity check. Table 6 reports original-only accuracy by model and domain. These results verify that the unedited images and question templates are generally answerable before counterfactual interventions are introduced. Model Existence Counting Fashion Industry Medical GPT-5 97.1 93.2 67.8 72.6 60.8 Claude Sonnet 4 88.8 84.1 55.5 77.7 47.6 Qwen3-VL-8B-Instruct 99.6 88.3 67.9 93.4 57.2 Qwen3-VL-32B-Instruct 94.6 94.4 77.7 88.1 63.5 InternVL3.5-8B 68.5 95.8 69.6 92.7 72.4 InternVL3.5-38B 81.2 85.5 62.6 88.0 48.4 Gemma-3-12B-it 69.6 54.0 63.1 86.9 31.1 Gemma-3-27B-it 73.8 69.8 64.9 86.4 46.1 Table 6: Original-only sanity-check accuracy by model and domain, in percent. High original-only accuracy indicates that the unedited images and questions are generally answerable. The original-only results are generally higher than the corresponding counterfactual-only results, supporting the interpretation that many counterfactual failures are not solely due to unanswerable questions or poor baseline recognition. The medical domain is more difficult even in the original-only setting, so medical results should be interpreted as a stronger stress test of local modality-consistency reasoning rather than as a standard recognition task. Domain-level summary. Table 7 reports domain-level macro-averages over models. This table supports the domain-level analysis in Section 6. Counting shows the largest original-to-counterfactual drop and the highest PFER, while fashion also produces strong prior-following behavior. Medical modality consistency has lower original-only accuracy, but still shows substantial PFER under counterfactual inputs. Domain Orig. Acc CF Acc CF PFER CF Other Drop Existence 84.2 52.8 43.3 3.9 31.3 Counting 83.1 17.0 77.9 5.1 66.2 Fashion 66.1 29.5 68.4 2.1 36.6 Industry 85.7 53.9 45.5 0.6 31.8 Medical 53.4 34.6 59.2 6.2 18.8 Table 7: Domain-level macro-averages over models for RQ1. Drop is the difference between original-only accuracy and counterfactual-only accuracy. Counting and fashion are the strongest prior-following stress tests, while medical modality consistency remains challenging even in the original-only setting. Original versus counterfactual performance. Figure 6 summarizes the model-level pattern. Panel A compares macro-averaged original-only and counterfactual-only accuracy. Panel B decomposes counterfactual-only responses into correct, prior-following, and other labels. Figure 6: Original-only and counterfactual-only raw-image performance. Panel A compares answerability under original images with performance under counterfactual images. Panel B decomposes counterfactual-only responses into correct, prior-following, and other labels. The large PFER component shows that counterfactual failures are often prior-consistent rather than merely ambiguous. Together, these results provide the raw evidence for RQ1. Original-only performance establishes baseline answerability, while counterfactual-only performance reveals a systematic shift toward prior-consistent errors. This supports the conclusion that current VLMs often fall back on learned category or modality priors when visible evidence contradicts what is usually true. Appendix F Paired-Image Analysis This appendix reports the full paired-image results used to support RQ2. In the paired-image setting, each model receives both the original and counterfactual image with neutral labels and is asked to compare the target visual attribute. This setting tests whether direct visual comparison helps models detect the counterfactual intervention, rather than answering from the prior associated with the object, category, or modality. All metrics are computed over valid, successfully scored responses. Query and judge errors are excluded from metric denominators and tracked separately in experiment logs. Unless otherwise stated, values are percentages for absolute results and percentage points for deltas. Absolute paired-image results. Table 8 reports paired-image performance by model and domain. Each cell reports Accuracy / PFER / Other Rate. Accuracy measures visually correct answers, PFER measures prior-consistent but visually incorrect answers, and Other Rate captures ambiguous, noncommittal, or otherwise uncategorizable responses. Model Existence Counting Fashion Industry Medical GPT-5 65.0/25.6/9.4 26.7/72.7/0.7 54.8/44.7/0.4 86.3/13.2/0.6 39.9/58.8/1.3 Claude Sonnet 4 54.2/35.0/10.8 28.5/68.1/3.4 42.3/57.6/0.2 75.7/24.2/0.1 36.1/62.0/1.9 Qwen3-VL-8B-Instruct 20.6/67.9/11.5 14.3/59.6/26.0 36.4/63.6/0.0 72.9/27.1/0.0 64.4/19.7/15.9 Qwen3-VL-32B-Instruct 57.5/31.5/11.0 35.9/45.1/19.0 35.1/64.9/0.0 67.3/32.6/0.1 52.9/46.8/0.3 InternVL3.5-8B 1.2/88.3/10.4 32.7/64.2/3.1 32.3/67.4/0.3 66.1/33.9/0.0 27.1/72.9/0.0 InternVL3.5-38B 26.5/69.0/4.6 32.6/57.7/9.8 39.5/60.5/0.0 71.0/28.9/0.1 38.4/61.6/0.0 Gemma-3-12B-it 27.5/61.9/10.6 29.9/43.5/26.6 33.6/66.4/0.0 31.3/68.7/0.0 36.4/63.5/0.1 Gemma-3-27B-it 30.8/57.1/12.1 32.0/49.5/18.5 41.2/58.7/0.1 61.4/38.6/0.0 44.0/54.7/1.3 Table 8: Paired-image results by model and domain. Each cell reports Acc/PFER/Other, in percent, when models compare original and counterfactual images with neutral labels. Table 8 shows that paired-image inputs do not uniformly solve prior-following behavior. In several domains, models still produce high PFER even when the original image is available for comparison. This indicates that simply presenting a reference image does not guarantee that the model identifies the task-relevant counterfactual change. Paired-image deltas. Table 9 reports paired-image deltas relative to the counterfactual-only raw-image baseline. Positive Δ indicates improved visual correctness, while negative Δ indicates reduced prior-following. We also report Δ because paired inputs may reduce biased answers by increasing uncertainty rather than by improving visual grounding. Model Existence Counting Fashion Industry Medical GPT-5 +7.3/-10.8/+3.5 +11.6/-11.8/+0.3 +10.5/-9.7/-0.8 +1.2/-1.4/+0.1 +8.4/-7.5/-0.9 Claude Sonnet 4 -8.1/-1.3/+9.4 +9.0/-5.6/-3.4 -0.4/+1.5/-1.2 +8.4/-8.1/-0.3 -5.3/+5.2/+0.1 Qwen3-VL-8B-Instruct -11.2/+5.8/+5.4 -0.5/-23.6/+24.1 +3.9/-2.1/-1.8 +21.3/-21.0/-0.3 +37.2/-52.4/+15.2 Qwen3-VL-32B-Instruct -16.9/+8.3/+8.5 +22.9/-40.8/+17.8 +7.9/-7.7/-0.1 +13.7/-13.5/-0.2 +19.9/-19.6/-0.3 InternVL3.5-8B -29.8/+22.7/+7.1 +27.6/-29.9/+2.3 +7.1/-5.7/-1.4 +9.7/-9.3/-0.4 +7.2/+5.7/-12.9 InternVL3.5-38B -26.7/+23.1/+3.5 +28.5/-27.3/-1.2 +12.5/-11.4/-1.1 +5.1/-4.8/-0.3 +0.7/-0.7/+0.0 Gemma-3-12B-it -19.2/+17.3/+1.9 -6.9/-9.2/+16.1 +14.8/-8.0/-6.8 +1.6/+0.3/-1.9 -30.4/+34.1/-3.7 Gemma-3-27B-it -34.8/+24.6/+10.2 +4.4/-14.5/+10.0 +22.7/-19.9/-2.7 +39.8/-38.8/-1.0 +24.8/+1.6/-26.4 Table 9: Paired-image deltas relative to counterfactual-only inputs. Each cell reports Δ /Δ /Δ , in percentage points. Positive Δ and negative Δ indicate improved visual grounding through direct comparison. The delta table highlights two patterns. First, paired images often improve performance when the counterfactual change can be directly compared across images, such as counts, visual layouts, or local modality appearance. Second, the paired setting can also hurt performance, especially in object-existence examples, where the presence of the canonical original may increase the salience of the normal object prior. Model-level summary. Table 10 reports model-level macro-averaged paired-image deltas across domains. Seven of the eight evaluated models improve in accuracy, but the gains vary substantially. Gemma-3-12B-it is the clearest negative case, showing reduced accuracy and increased PFER under paired inputs. Model Δ Δ Δ GPT-5 +7.8 -8.2 +0.4 Claude Sonnet 4 +0.7 -1.7 +0.9 Qwen3-VL-8B-Instruct +10.1 -18.7 +8.5 Qwen3-VL-32B-Instruct +9.5 -14.7 +5.1 InternVL3.5-8B +4.4 -3.3 -1.1 InternVL3.5-38B +4.0 -4.2 +0.2 Gemma-3-12B-it -8.0 +6.9 +1.1 Gemma-3-27B-it +11.4 -9.4 -2.0 Table 10: Model-level macro-averaged paired-image deltas across domains. Values are percentage points relative to counterfactual-only inputs. The model-level summary shows that paired comparison is usually helpful, but not uniformly so. Some models convert paired visual information into higher accuracy and lower PFER, while others show only small changes or even become more prior-driven. This suggests that providing a reference image helps only when the model can identify and use the task-relevant difference. Domain-level summary. Table 11 reports domain-level macro-averaged paired-image deltas over models. Counting, fashion, industry/common objects, and medical modality generally benefit from paired comparison. Object existence is the main negative case, suggesting that showing the original object alongside the counterfactual can make the canonical object prior more salient or increase comparison difficulty. Domain Δ Δ Δ Existence -17.4 +11.2 +6.2 Counting +12.1 -20.3 +8.3 Fashion +9.9 -7.9 -2.0 Industry +12.6 -12.1 -0.5 Medical +7.8 -4.2 -3.6 Table 11: Domain-level macro-averaged paired-image deltas over models. Values are percentage points relative to counterfactual-only inputs. The domain-level summary shows that paired images are most useful when the target difference can be compared directly, such as a count, local layout, visual design feature, or MRI contrast pattern. In contrast, object-existence examples become harder under paired inputs, suggesting that direct comparison may sometimes distract the model or reinforce the canonical object prior. Heatmap visualization. Figure 7 visualizes paired-image deltas by model and domain. Panel A shows Δ and Panel B shows Δ . The paired setting is beneficial when accuracy increases while PFER decreases. This pattern appears in several domains, especially counting, fashion, industry/common objects, and parts of the medical modality setting. However, the gains are not uniform: object-existence examples often show lower accuracy and higher PFER under paired inputs. Figure 7: Heatmap of paired-image deltas by model and domain. Panel A reports Δ and Panel B reports Δ relative to counterfactual-only inputs. Positive Δ and negative Δ indicate improved visual grounding through direct comparison. Overall, paired images provide a useful but limited intervention. They often make counterfactual changes more salient, leading to higher accuracy and lower PFER, but they do not consistently eliminate prior-following behavior. In some cases, paired comparison increases Other Rate or even strengthens prior-following errors, indicating that adding a reference image is not equivalent to reliable visual grounding. Appendix G PriVE-Tools Aggregate Results This appendix reports aggregate PriVE-Tools results used to support RQ3. PriVE-Tools evaluates fixed tool-derived evidence views rather than allowing models to freely call tools. Each tool condition is compared against the matched raw counterfactual baseline for the same model and domain. We report Δ , Δ , and Δ . Positive Δ and negative Δ indicate improved grounding, while changes in Δ help distinguish genuine correction from increased ambiguity or uncertainty. For the main cross-domain tool analysis, we focus on the two common evidence conditions available across all five domains: Crop and Zoom panel. This avoids comparing tools with different domain coverage. Table 12 reports aggregate deltas for all models, closed-source models, and open-source models. Table 13 provides the corresponding model-level breakdown and supports the Top Common Tool column in Table 1. Tool Coverage All models Closed-source Open-source Crop 5 domains / 8 models -0.9 / +0.2 / +0.7 +7.3 / -6.9 / -0.4 -3.6 / +2.6 / +1.0 Zoom panel 5 domains / 8 models -0.2 / -1.1 / +1.3 +8.0 / -7.5 / -0.4 -3.0 / +1.1 / +1.9 Table 12: Aggregate PriVE-Tools deltas for common evidence conditions available across all five domains. Each metric cell reports Δ / Δ / Δ , in percentage points, relative to the matched raw counterfactual baseline. A useful evidence condition should increase accuracy and decrease PFER without substantially increasing Other. Table 12 shows that common evidence views do not uniformly improve grounding. Across all evaluated models, Crop and Zoom panel have small or negative effects on accuracy. However, this aggregate pattern masks a strong group-level split: the same evidence views improve accuracy and reduce PFER for closed-source models, but are weak or negative on average for open-source models. Model-level common tool deltas. Table 13 reports common tool-condition deltas for each model. These values are macro-averaged across the five domains and computed relative to each model’s matched raw counterfactual baseline. The table supports the Top Common Tool column in Table 1, where the best common tool is selected by the largest macro-averaged accuracy gain among Crop and Zoom panel. Model Crop Zoom panel Closed-source VLMs GPT-5 +4.5 / -4.3 / -0.2 +4.2 / -4.4 / +0.2 Claude Sonnet 4 +10.1 / -9.5 / -0.6 +11.8 / -10.6 / -1.1 Open-source VLMs Qwen3-VL-8B-Instruct -2.5 / +3.4 / -0.9 -1.4 / +1.2 / +0.3 Qwen3-VL-32B-Instruct -2.2 / +2.9 / -0.7 -2.8 / +3.3 / -0.5 InternVL3.5-8B -2.4 / -0.8 / +3.2 +0.5 / -5.8 / +5.2 InternVL3.5-38B -5.7 / -2.8 / +8.5 -8.3 / -1.8 / +10.1 Gemma-3-12B-it -7.2 / +9.2 / -2.0 -5.2 / +6.6 / -1.5 Gemma-3-27B-it -1.6 / +3.6 / -2.0 -0.7 / +2.8 / -2.1 Table 13: Common tool-condition deltas by model. Each cell reports Δ / Δ / Δ , in percentage points, macro-averaged across the five domains and computed relative to the matched raw counterfactual baseline. Table 13 shows that the same evidence condition can have different effects across models. GPT-5 and Claude Sonnet 4 both benefit from common tool views, especially Claude Sonnet 4 under Zoom panel evidence. In contrast, several open-source models show negative Δ even under their better common tool condition. Some models reduce PFER while increasing Other Rate, indicating that tool evidence may shift responses from prior-following to uncertainty rather than to correct visual grounding. Aggregate visualization. Figure 8 visualizes the aggregate effects of the common tool conditions. The figure separates all models, closed-source models, and open-source models. Positive Δ and negative Δ indicate improved grounding. Figure 8: Macro-averaged effects of common PriVE-Tools evidence conditions across models and domains. Bars report Δ , Δ , and Δ relative to matched raw counterfactual baselines. Positive Δ and negative Δ indicate improved grounding. Figure 8 reinforces the main RQ3 finding: common tool-derived evidence is beneficial for closed-source models in our evaluation, but weak or negative on average for open-source models. This suggests that providing localized or magnified evidence is not sufficient by itself; the model must also be able to integrate that evidence into the answer. Partially covered and domain-specific tools. Some PriVE-Tools evidence conditions are domain-specific or only partially covered in the current experiments. Counting includes outline overlays, medical modality includes contour overlays, and bounding-box overlays are incomplete for Fashion and Industry because the open-source bbox runs were not completed. We therefore report these results separately in Table 14 rather than using them for the main cross-domain Top Common Tool selection. Tool Coverage Domains All models Closed-source Open-source BBox 28 model-domain pairs Counting, Existence, Fashion, Industry, Medical -2.0 / -0.1 / +2.1 +5.5 / -5.6 / +0.2 -6.1 / +2.9 / +3.2 Outline 8 model-domain pairs Counting -4.7 / +2.7 / +1.9 +0.7 / -2.7 / +2.1 -6.5 / +4.6 / +1.9 Contour 8 model-domain pairs Medical +3.0 / -7.3 / +4.3 +14.3 / -13.1 / -1.2 -0.8 / -5.4 / +6.2 Table 14: Domain-specific or partially covered PriVE-Tools deltas. Outline is counting-specific and Contour is medical-specific. Each metric cell reports Δ / Δ / Δ relative to the matched raw counterfactual baseline. Table 14 shows that some domain-specific evidence views can be useful, especially for closed-source models. For example, Contour evidence in the medical domain substantially improves the closed-source aggregate, while its open-source effect is weaker and accompanied by increased Other Rate. Because these tools do not share uniform coverage across all domains and models, we treat them as supplementary evidence rather than as the basis for the main cross-domain tool comparison. Overall, these aggregate results support a cautious interpretation of PriVE-Tools. Tool-derived evidence can improve grounding, especially for stronger closed-source models and for some domain-specific evidence views. However, tools do not consistently convert prior-following errors into correct answers. In several open-source settings, tool views reduce accuracy, increase PFER, or increase Other Rate. This motivates the domain- and model-specific analysis in RQ4. Appendix H Domain- and Model-Specific Tool Analysis This appendix provides the domain- and model-level PriVE-Tools analysis used to support RQ4. While Appendix G reports aggregate tool effects, this section examines where those effects come from. We focus on displayed evidence conditions that are directly interpretable across the benchmark: bounding-box overlays, object crops, and zoom panels. All deltas are computed relative to the matched raw counterfactual baseline for the same model and domain. Because the Fashion and Industry bbox runs were not completed for all open-source models, those cells are left blank in the domain-level analysis rather than treated as complete evidence. For model-level bbox summaries, we macro-average only over domains with complete bbox coverage across all eight main models: object existence, counting, and medical modality. This keeps model-level comparisons consistent. Best displayed tool by domain. Table 15 reports the best displayed tool by domain, selected by macro-averaged Δ among available displayed evidence views. The corresponding Δ and Δ values indicate whether the gain reflects improved grounding, reduced prior-following, or increased uncertainty. Domain Best displayed tool Δ Δ Δ Interpretation Existence Crop -4.3 +5.3 -1.0 Common displayed tools do not improve object-existence judgments on average; localized views can still leave models anchored to category priors. Counting Zoom panel -3.1 +1.0 +2.1 Magnified evidence does not by itself solve enumeration, and can shift some prior-following errors into uncertainty. Fashion Zoom panel -2.5 +3.0 -0.5 Fine-grained logo and element edits remain prior-sensitive on average, even when local visual evidence is provided. Industry Crop +4.4 -4.1 -0.2 Local crops help most clearly for common-object attribute changes, where the manipulated object can be inspected directly. Medical Zoom panel +5.3 -10.5 +5.2 Zoom panels reduce modality-prior errors, but the increased Other rate shows that some gains reflect uncertainty around local contrast. Table 15: Best-performing displayed tool condition by domain. The best tool is selected by macro-averaged Δ among available displayed tools, with corresponding changes in PFER and Other reported to indicate whether improvements reflect better grounding or increased uncertainty. Deltas are relative to matched raw counterfactual baselines. Fashion and Industry BBox cells are excluded because open-source bbox runs were not completed. Table 15 shows that no evidence view dominates across domains. Local crops are most useful for industry/common-object attributes, where the manipulated property is localized and visually inspectable. Zoom panels help medical modality consistency by exposing local tumor-region contrast, but the increased Other Rate suggests that some models become uncertain rather than fully grounded. Object existence, counting, and fashion remain challenging even when the relevant region is localized or magnified. Domain-level heatmaps. Figure 9 visualizes domain-level tool effects. Panel A reports Δ and Panel B reports Δ relative to matched raw counterfactual baselines. The strongest positive displayed-tool effects appear in Industry and Medical. By contrast, object existence, counting, and fashion show weak or negative gains under the displayed evidence views, indicating that localized evidence alone does not solve absence reasoning, enumeration, or fine-grained prior-resistant recognition. Figure 9: Domain-by-tool heatmaps for displayed PriVE-Tools evidence conditions. Panel A reports Δ and Panel B reports Δ relative to matched raw counterfactual baselines. Fashion and Industry BBox cells are left blank because open-source bbox runs were not completed. Model-level heatmaps. Figure 10 shows the corresponding model-level pattern. Closed-source models benefit more consistently from localized evidence: GPT-5 and Claude Sonnet 4 show positive accuracy gains and reduced PFER under Crop and Zoom panel. Several open-source models, however, show negative accuracy changes, increased PFER, or increased Other Rate under the same evidence conditions. This indicates that tool usefulness depends not only on whether relevant visual evidence is visible, but also on whether the model can use that evidence against learned priors. Figure 10: Model-by-tool heatmaps for displayed PriVE-Tools evidence conditions. Panel A reports Δ and Panel B reports Δ . BBox values are macro-averaged over domains with complete bbox coverage across all models. Summary. Together, these results show that PriVE-Tools effects are both domain-specific and model-specific. Local crops are most useful for common-object attribute changes, zoom panels help expose medical modality conflicts but may increase uncertainty, and bounding boxes do not uniformly help object-existence or counting tasks. The same evidence view can therefore help one model or domain while degrading another. This supports the main conclusion that agentic visual evidence is not a universal remedy: it exposes relevant evidence, but VLMs still need the ability to interpret and use that evidence correctly.