Paper deep dive
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs
Dhita Putri Pratama, Soyeon Caren Han, Yihao Ding
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 1:45:10 PM
Summary
The paper introduces Vision-Language Causal Graphs (VLCGs) and the ViLCaR benchmark to diagnose causal reasoning in Large Vision-Language Models (LVLMs). It argues that LVLMs often rely on spurious correlations rather than genuine causal reasoning. ViLCaR evaluates Causal Attribution, Causal Inference, and Question Answering using graph-aligned metrics. Experiments show that injecting structured VLCG information significantly improves attribution and inference consistency compared to zero-shot and standard in-context learning, suggesting that limitations stem from insufficient structural guidance.
Entities (8)
Relation Signals (7)
VLCG → usedin → ViLCaR
confidence 98% · Building on this representation [VLCG], we present ViLCaR, a diagnostic benchmark...
ViLCaR → includestask → Causal Attribution
confidence 95% · ViLCaR, a diagnostic benchmark comprising tasks for Causal Attribution, Causal Inference, and Question Answering
ViLCaR → includestask → Causal Inference
confidence 95% · ViLCaR, a diagnostic benchmark comprising tasks for Causal Attribution, Causal Inference, and Question Answering
LVLM → suffersfrom → spurious_correlations
confidence 95% · LVLMs achieve strong performance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasoning.
Qwen2.5-VL-7B → evaluatedon → ViLCaR
confidence 92% · We benchmark Qwen2.5-VL-7B on ViLCaR
VLCG-Augmented Prompting → improves → Causal Attribution
confidence 90% · VLCG-augmented prompting improves both Causal Attribution and Causal Inference. CA increases from 0.458 to 0.488
VLCG-Augmented Prompting → improves → Causal Inference
confidence 90% · VLCG-augmented prompting improves both Causal Attribution and Causal Inference. ... CI improves from 0.652 to 0.690
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Vision-Language Models (LVLMs) achieve strong performance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasoning. Existing evaluations primarily assess the correctness of the answers, making it unclear whether failures arise from limited reasoning capability or from misidentifying causally relevant information. We introduce Vision-Language Causal Graphs (VLCGs), a structured, query-conditioned representation that explicitly encodes causally relevant objects, attributes, relations, and scene-grounded assumptions. Building on this representation, we present ViLCaR, a diagnostic benchmark comprising tasks for Causal Attribution, Causal Inference, and Question Answering, along with graph-aligned evaluation metrics that assess relevance identification beyond final answer accuracy. Experiments in state-of-the-art LVLMs show that injecting structured relevance information significantly improves attribution and inference consistency compared to zero-shot and standard in-context learning. These findings suggest that current limitations in LVLM causal reasoning stem primarily from insufficient structural guidance rather than a lack of reasoning capacity.
Tags
Links
- Source: https://arxiv.org/abs/2602.20878v1
- Canonical: https://arxiv.org/abs/2602.20878v1
Trouble viewing inline? Open PDF directly →
Full Text
26,834 characters extracted from source content.
Expand or collapse full text
Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs Dhita Putri Pratama University of Melbourne Melbourne, Australia Soyeon Caren Han University of Melbourne Melbourne, Australia Yihao Ding The University of Western Australia Perth, Australia Abstract Large Vision-Language Models (LVLMs) achieve strong perfor- mance on visual question answering benchmarks, yet often rely on spurious correlations rather than genuine causal reasoning. Existing evaluations primarily assess the correctness of the answers, making it unclear whether failures arise from limited reasoning capability or from misidentifying causally relevant information. We intro- duce Vision-Language Causal Graphs (VLCGs), a structured, query-conditioned representation that explicitly encodes causally relevant objects, attributes, relations, and scene-grounded assump- tions. Building on this representation, we present ViLCaR, a diag- nostic benchmark comprising tasks for Causal Attribution, Causal Inference, and Question Answering, along with graph-aligned eval- uation metrics that assess relevance identification beyond final answer accuracy. Experiments in state-of-the-art LVLMs show that injecting structured relevance information significantly improves attribution and inference consistency compared to zero-shot and standard in-context learning. These findings suggest that current limitations in LVLM causal reasoning stem primarily from insuffi- cient structural guidance rather than a lack of reasoning capacity. Keywords Large Vision Language Model, Causal Reasoning, Resource ACM Reference Format: Dhita Putri Pratama, Soyeon Caren Han, and Yihao Ding. 2026. Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance Graphs. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’26). ACM, Mel- bourne, VIC, Australia, 5 pages. https://doi.org/X.X 1 Introduction Large Vision-Language Models (LVLMs) have demonstrated strong performance on visual question answering and multimodal reason- ing benchmarks. However, high answer accuracy does not neces- sarily imply faithful or causally grounded reasoning. Models may produce correct answers while relying on spurious visual cues or superficial associations, resulting in explanations that are inconsis- tent or unfaithful [10,13]. This discrepancy reveals a key limitation of current evaluation paradigms: they primarily measure prediction Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. SIGIR ’26, Melbourne, Australia © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2026/06 https://doi.org/X.X DatasetsSize Graph ComponentTasks OAORCAssCDCACIQA VQA [1, 4]1.1M✗✔ Visual7W [15]328K✗✔ V-Genome [7]1.7M✔✗✔ VCR [14]290K✗✔ OK-VQA [9]14k✗✔ CoSIm [6]3.5K✗✔ CELLO [2]14K✗✔✗✔✗✔ ViLCaR (Ours)12.5K✔ Table 1: Comparison of the ViLCaR with existing works that include causal reasoning tasks. OA, OR, CAss, and CD de- note Object Attribute, Object Relation, Causal Assumption, and Context Dependence. CA, CI, and QA represent Causal Attribution, Causal Inference, and Question Answering. correctness, but do not diagnose whether models correctly iden- tify the causally relevant information required for valid reasoning. We argue that many failures in visual causal reasoning stem from errors in relevance identification. Before performing inference, a model must determine which objects, attributes, relations, and con- textual assumptions are causally relevant to a query. When this identification stage fails, downstream reasoning becomes brittle or misleading, even if the final answer is correct. Despite its impor- tance, this step remains largely unexamined in existing benchmarks. As summarized in Table 1, most VQA-style datasets [1,4,6,9,14,15] lack explicit causal structure, providing neither attribute-level de- pendencies nor scene-grounded assumptions. Although CELLO [2] introduces object-level causal graphs, it does not model attribute- level factors or contextual assumptions necessary for fine-grained causal attribution and counterfactual reasoning. Moreover, evalu- ation protocols predominantly rely on answer accuracy, limiting the ability to distinguish perception errors from flawed causal rea- soning. To address these gaps, we introduce Vision-Language Causal Graphs (VLCGs), a structured, query-conditioned rep- resentation that explicitly encodes objects, their attributes, inter- object relations, and scene-grounded assumptions as directed causal dependencies. Unlike standard scene graphs [7] or abstract causal models [11], VLCGs are designed to capture question-driven causal relevance in visual contexts. Building on this representation, we present ViLCaR, a diagnostic benchmark comprising three tasks: (1) Causal Attribution (CA), which evaluates whether models cor- rectly identify causally relevant attributes; (2) Causal Inference (CI), which assesses the consistency of reasoning chains grounded in attributes and assumptions; and (3) Question Answering (QA), which measures final prediction accuracy. We further introduce graph- aligned evaluation metrics that disentangle relevance identification arXiv:2602.20878v1 [cs.AI] 24 Feb 2026 SIGIR ’26, July, Melbourne, AustraliaDhita Putri Pratama, Soyeon Caren Han, and Yihao Ding Figure 1: Example of a VLCG. Given an image-question pair (“Have these people just married?”), the graph encodes causally relevant objects (e.g., persons, cake), attributes (wed- ding dress, suit), relations (wear), and scene-grounded as- sumptions linking visual evidence to the conclusion. Unlike scene graphs, VLCGs capture question-conditioned causal relevance rather than complete perceptual structure. from answer correctness. Our experiments show that simply pro- viding question–answer exemplars does not reliably improve causal reasoning. In contrast, prompting with structured causal graphs significantly improves attribution and inference consistency. These findings suggest that current limitations in LVLM causal reasoning arise less from an inherent inability to reason and more from insuf- ficient structural guidance. Main contributions are summarized: •We propose VLCGs, a structured representation for modeling query-conditioned causal relevance in multimodal reasoning. •We introduce ViLCaR, a diagnostic benchmark enabling fine- grained analysis of causal attribution and inference. •We present graph-aligned evaluation metrics that disentangle relevance identification from final answer accuracy. 2 VilCaR 2.1 Vision-Language Causal Graphs (VLCGs) We introduce VLCGs, a structured representation designed to model question conditioned causal relevance in visual reasoning. For- mally, given an image퐼and a question푞, a VLCG is a directed graph 퐺=(푉,퐸,퐴)where: (i)푉denotes scene-grounded entities (objects and abstract concepts), (i)퐸denotes directed dependencies encod- ing attributes and inter-object relations, and (i)퐴denotes a set of explicit causal assumptions required to justify the correct answer. Unlike traditional scene graphs that describe perceptual structure, VLCGs explicitly encode causally relevant elements with respect to a specific image–question pair. The graph therefore represents not the full scene, but the minimal set of objects, attributes, rela- tions, and contextual assumptions necessary to support a plausible causal explanation. Crucially, VLCGs incorporate scene-grounded assumptions following the structural view of causality [11]. These assumptions encode implicit cultural or contextual knowledge that cannot be directly inferred from visual perception alone (e.g., attire as an indicator of a wedding ceremony). By integrating observable attributes with explicit assumptions, VLCGs capture the mechanism linking visual evidence to causal conclusions. Figure 2: Three diagnostic tasks in ViLCaR derived from the verified and pruned VLCGs: CA, CI, and QA. 2.2 ViLCaR Tasks Building on VLCGs, we construct ViLCaR, a diagnostic benchmark for evaluating visual causal reasoning along three complementary dimensions: CA, CI, and QA. (1) Causal Attribution (CA). Given(퐼,푞), the model must identify the set of causally relevant attributes or variables that influence the answer. This task evaluates whether the model selects ap- propriate causal factors rather than relying on spurious cues. (2) Causal Inference (CI). The model must generate a reasoning chain that plausibly connects identified attributes and assump- tions to the final answer. Reasoning is considered plausible when the synthesized causal factors support a coherent conclusion. (3)Question Answering (QA). The model predicts the final an- swer to(퐼,푞). This task measures outcome accuracy but does not, by itself, guarantee correct relevance identification or in- ferential validity. These tasks disentangle three stages of visual causal reasoning: identifying relevant variables, composing a valid causal mechanism, and producing the final prediction. The overall pipeline is illustrated conceptually in Figure 2. 3 Dataset Construction In this paper, we construct ViLCaR from existing visual question an- swering datasets, primarily VQA [1,4] and VCR [14], and transform them into structured and question-conditioned causal reasoning instances. 1) Question Selection and Causal Filtering. We first filter out questions that can be solved via object detection, spatial lookup, opinion-based judgment, or low-level perceptual cues. Remaining questions are categorized according to the Ladder of Causation [12], retaining instances that require associative, interventional, or counterfactual reasoning. This step ensures that each instance requires non-trivial causal grounding. 2) Graph Generation. For each image–question pair(퐼,푞), we construct a preliminary Vision-Language Causal Graph (VLCG) using LVLM-based prompting. The model is instructed to extract: (i) causally relevant objects, (i) their attributes, (i) inter-object relations, and (iv) explicit assumptions necessary to justify the Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance GraphsSIGIR ’26, July, Melbourne, Australia answer. The resulting graph encodes candidate causal factors rather than a full scene description. 3) Graph Verification and Grounding. To reduce hallucinated entities and attributes, we validate graph components using inde- pendent vision models. Object nodes are aligned with detections from closed-set [16] and open-vocabulary detectors [3,8], while attributes and relations are validated via image–text similarity scor- ing using CLIP-Score [5]. Elements that fail grounding thresholds are removed. This step ensures that retained graph components are visually supported. 4) Minimal Causal Pruning. To obtain a minimal sufficient causal graph, we iteratively remove nodes and edges that are not required to derive the correct answer. An LLM, given only the graph (without the image), evaluates whether the remaining structure is sufficient to answer the question. Components that do not affect answer valid- ity are pruned. The final VLCG therefore represents the smallest set of causally relevant elements supporting the ground-truth answer. 5) Quality Control. We conduct human validation on 30 randomly sampled VLCGs with 15 annotators. Each graph is evaluated ac- cording to four criteria: correctness, relevance, sufficiency, and assumption strength, at the levels of objects, attributes, and re- lations. Disagreement rates remain below 15% across all criteria, indicating consistent annotation quality. Notably, the near-equal distribution between “agree” and “strongly agree” responses high- lights the inherent subjectivity of visual causal reasoning, even among humans. 4 Data Statistics We analyze the structural properties of the constructed VLCGs in Figure 3. As shown in Figure 3(a), person is the most frequently ref- erenced object, reflecting the prevalence of human-centered causal reasoning scenarios in ViLCaR. The attribute distribution in Fig- ure 3(b) is dominated by mental-state and role-related descriptors (e.g., facial expression, gesture), indicating that many causal infer- ences rely on socially grounded or affective cues rather than purely physical attributes. Similarly, the relation distribution in Figure 3(c) highlights common physical interactions such as hold and wear, suggesting that causal reasoning in ViLCaR frequently integrates both social semantics and observable object interactions. Overall, these statistics confirm that VLCGs capture a mixture of perceptual evidence and higher-level contextual attributes, aligning with our goal of modeling question-conditioned causal relevance rather than purely spatial structure. 5 Experiment Setup We evaluate whether structured VLCGs improve visual causal rea- soning beyond standard prompting strategies. The main question is whether injecting explicit causal relevance signals enhances attribu- tion and inference, rather than merely improving answer accuracy. We benchmark Qwen2.5-VL-7B on ViLCaR using an 80/10/10 train–validation–test split. All reported results are computed on the test set. We compare three settings: (1)Zero-shot. The model receives only the image and question and generates both reasoning and answer. (a) Top-5 Objects (b) Attribute: Word Cloud (c) Relation: Word Cloud Figure 3: A brief data statistics of VLCGs, with ‘person’, men- tal states (e.g., facial expression), and physical relationships of the object (e.g., hold) being the most frequent objects, ob- ject characteristics, object relations, respectively. (2) Standard ICL. The model is provided with question-answer exemplars without structured graphs, testing whether unstruc- tured few-shot prompting improves reasoning. (3)VLCG-Augmented Prompting. We inject structured causal graphs into the prompt and report results from the best-performing configuration. We evaluate performance along three diagnostic dimensions. By reporting CA, CI, and QA jointly, we disentangle correctness from causal reasoning quality. (a) Causal Attribution (CA). We measure whether the model cor- rectly identifies causally relevant attributes present in the VLCG. Given open-vocabulary reasoning outputs, we align reasoning to- kens with graph elements using a semantic similarity score: 푆= 푆 퐵퐸푅푇 × 푆 푊 2푉 × (1 + 푃 푢푛푖푔푟푎푚 )(1) where푆 퐵퐸푅푇 and푆 푊 2푉 denote contextual and distributional similar- ity, and푃 푢푛푖푔푟푎푚 captures lexical overlap and serves as the weight importance for any token overlaps. A triplet (e.g., <cake, typical, wedding> is considered correctly identified if the similarity of at least 50% triplet elements exceed a fixed threshold. Causal Inference (CI). CI evaluates whether the identified at- tributes are coherently composed into a valid reasoning chain that supports the final prediction. Specifically, we use an LLM-based alignment protocol where a separate evaluator model is prompted to compare: (i) the generated reasoning, and (i) the gold causal assumptions encoded in the VLCG. The evaluator assesses whether the reasoning: (a) links relevant attributes to the outcome, (b) main- tains logical consistency, and (c) avoids unsupported assumptions. The CI score reflects agreement between generated reasoning and the structured causal path, averaged across test instances. (c) Question Answering (QA). QA reports final answer accuracy. However, QA does not require correct attribution or inferential validity; a model may reach the correct answer via shortcut corre- lations or spurious visual cues. SIGIR ’26, July, Melbourne, AustraliaDhita Putri Pratama, Soyeon Caren Han, and Yihao Ding MetricZero-shot Standard ICL VLCG (Best) Causal Attribution (CA)0.4580.4550.488 Causal Inference (CI)0.6520.6540.690 VQA Accuracy0.7630.7630.768 BLEU (reasoning)0.1640.1630.177 ROUGE (reasoning)0.2660.2640.273 Table 2: Performance comparison between zero-shot prompting, standard in-context learning (ICL), and VLCG- augmented prompting (best configuration). CA and CI mea- sure relevance identification and inferential consistency, while BLEU/ROUGE assess surface-level reasoning overlap. 6 Results 6.1 Overall Performance Table 2 reports performance across three prompting settings. Stan- dard in-context learning yields negligible gains over zero-shot prompting. CA slightly decreases (0.458→0.455), while CI and QA remain nearly unchanged. This indicates that simply providing question–answer exemplars does not meaningfully improve causal relevance identification. In some cases, ICL may even introduce noise by encouraging surface pattern imitation rather than struc- tured reasoning. In contrast, VLCG-augmented prompting improves both Causal Attribution and Causal Inference. CA increases from 0.458 to 0.488 (+6.6% relative improvement), while CI improves from 0.652 to 0.690 (+5.8% relative improvement). Notably, the improve- ment in CI is larger in absolute magnitude than CA, suggesting that structured graphs not only help identify relevant attributes, but also stabilize their composition into coherent reasoning chains. Despite gains in CA and CI, VQA accuracy remains nearly con- stant (0.763→0.768). This decoupling indicates that correct an- swers can be obtained even when relevance identification is imper- fect. In other words, models may exploit shortcut visual correlations to obtain correct answers, while failing to construct stable causal explanations. These results empirically support our claim that accu- racy alone is insufficient to diagnose visual causal reasoning. BLEU and ROUGE show modest improvements under VLCG prompting. However, the relative gains in lexical overlap are smaller than those observed in CA and CI. This suggests that improvements stem primarily from structured relevance alignment rather than super- ficial textual similarity. Overall, the findings indicate that explicit causal structure acts as a relevance prior: it constrains the model’s attention to causally grounded attributes and reduces reliance on spurious cues. Importantly, this improvement manifests in attribu- tion and inferential consistency, even when final answer accuracy remains stable. 6.2 Qualitative Analysis Figure 4 presents a representative example illustrating how struc- tured causal guidance alters model behavior under identical vi- sual input. In both zero-shot and standard ICL settings, the model concludes that the question cannot be answered. Although the vi- sual scene clearly contains two individuals positioned at a counter, the model fails to identify their functional roles (e.g., server and customer). Consequently, Causal Attribution (CA) is 0, indicating Figure 4: Reasonings from Qwen2.5-VL 7B model with: (1) Zero-shot, (2) Standard ICL, and (3) VLCG-augmented. Com- pared to the baselines, the model with ViLCaR injection is able to identify information about the role of the people rel- evant to the question. complete failure to retrieve relevant causal attributes. Without role identification, the model cannot compose a valid reasoning chain, leading to incorrect QA predictions. Notably, the failure is not due to perceptual inability but to relevance selection: the model de- scribes visible entities but does not ground them in task-relevant causal roles. When provided with structured VLCGs, the model explicitly identifies role attributes and links them to the assumption: “A neutral or attentive server would fulfill the customer’s order.” This enables the model to construct a coherent causal chain from role identification to outcome prediction. Hence, CA increases from 0 to 0.5, CI rises from 0.64/0.599 to 0.84, and the final QA predic- tion becomes correct. Importantly, the visual evidence remains unchanged across settings; the behavioral difference arises solely from structured relevance guidance. This example highlights a key distinction in our framework: zero-shot and ICL failures stem from under-specification of causal roles, whereas VLCG prompting reduces ambiguity by explicitly constraining the reasoning space. Rather than adding new percep- tual information, the graph acts as a relevance prior that encourages the model to ground its reasoning in causally meaningful variables. This supports our central claim that structured causal guidance improves attribution and inferential consistency, even when raw visual input remains constant. 7 Conclusion We introduced ViLCaR, a benchmark for diagnosing failure modes in visual causal reasoning for LVLMs through structured VLCGs. By grounding causal attributes and assumptions in visual scenes, VL- CGs provide a relevance prior for controlled analysis of reasoning behavior. Our results show that answer accuracy alone is insuf- ficient to diagnose reasoning quality, and that structured causal guidance improves attribution and inferential consistency even Diagnosing Causal Reasoning in Vision-Language Models via Structured Relevance GraphsSIGIR ’26, July, Melbourne, Australia when accuracy remains unchanged. Although our metrics rely on approximate semantic alignment, they provide a scalable diagnostic framework for analyzing reasoning failures, and future work may refine them through LLM-as-a-judge or human-in-the-loop valida- tion. Overall, ViLCaR offers a principled testbed for studying how LVLMs select, compose, and ground causally relevant information in visual reasoning tasks. References [1] Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015. VQA: Visual Question Answering. In International Conference on Computer Vision (ICCV). [2]Meiqi Chen, Bo Peng, Yan Zhang, and Chaochao Lu. 2024. CELLO: Causal Eval- uation of Large Vision-Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 22353–22374. doi:10.18653/v1/2024.emnlp-main.1247 [3]Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. 2024. YOLO-World: Real-Time Open-Vocabulary Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 16901–16911. [4]Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR). [5] Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen- tau Yih (Eds.). Association for Computational Linguistics, Online and Punta Cana, Dominican Republic, 7514–7528. doi:10.18653/v1/2021.emnlp-main.595 [6]Hyounghun Kim, Abhay Zala, and Mohit Bansal. 2022. CoSIm: Commonsense Rea- soning for Counterfactual Scene Imagination. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Marine Carpuat, Marie-Catherine de Marneffe, and Ivan Vladimir Meza Ruiz (Eds.). Association for Computational Linguistics, Seattle, United States, 911–923. doi:10.18653/v1/2022.naacl-main.66 [7]Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017. Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations. International Journal of Computer Vision 123 (2017), 32–73. doi:10.1007/s11263-016-0981-7 [8] Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. 2025. Grounding DINO: Marrying DINO with Grounded Pre-training for Open-Set Object Detection. In Computer Vision – ECCV 2024, Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol (Eds.). Springer Nature Switzerland, Cham, 38–55. [9]Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019. OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). [10]Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024. Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 15012–15032. doi:10.18653/ v1/2024.findings-emnlp.882 [11]Judea Pearl. 2009. Causality: Models, Reasoning and Inference (2nd ed.). Cambridge University Press, USA. [12]Judea Pearl and Dana Mackenzie. 2018. The Book of Why: The New Science of Cause and Effect (1st ed.). Basic Books, Inc., USA. [13]Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of- Thought Prompting. In Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Vol. 36. Curran Associates, Inc., 74952–74965. https://proceedings.neurips.c/paper_ files/paper/2023/file/ed3fea9033a80fea1376299fa7863f4a-Paper-Conference.pdf [14]Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. From Recogni- tion to Cognition: Visual Commonsense Reasoning. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [15] Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016. Visual7W: Grounded Question Answering in Images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [16]Zhuofan Zong, Guanglu Song, and Yu Liu. 2023. DETRs with Collaborative Hybrid Assignments Training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 6748–6758.