Paper deep dive
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Liangjie Zhao, Jiaqing Lyu, Kexin Tang, Zecheng Fang, Rong Yin, Yulan Hu, Da Li, Jianing Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/1/2026, 2:07:53 AM
Summary
The paper introduces IllusionReasoning, a benchmark designed to evaluate Large Vision Language Models (LVLMs) on visual illusions to assess both perceptual and reasoning capabilities. It argues that existing evaluations are either too domain-specific or focus only on perception. The benchmark uses real-world illusion images and diverse question-answer pairs covering detection, description, and reasoning. Experiments show that current LVLMs, even state-of-the-art ones, struggle with these tasks, often failing to distinguish between visual appearance and physical reality, indicating that their reasoning capabilities are not as advanced as claimed.
Entities (10)
Relation Signals (7)
IllusionReasoning â evaluates â LVLMs
confidence 95% ¡ We evaluated a wide range of LVLMs on IllusionReasoning
LVLMs â struggleswith â Visual Illusions
confidence 92% ¡ We found that even state-of-the-art LVLMs struggle to provide correct responses to questions related to visual illusions.
IllusionReasoning â assesses â reasoning capabilities
confidence 90% ¡ IllusionReasoning comprehensively evaluates and distinguishes the perception and reasoning capabilities of various LVLMs.
IllusionReasoning â assesses â Perceptual Capabilities
confidence 90% ¡ IllusionReasoning comprehensively evaluates and distinguishes the perception and reasoning capabilities of various LVLMs.
GPT-4o â usedas â Evaluator
confidence 90% ¡ We employ GPT-4o as the evaluator to determine whether a modelâs response aligns with the annotated answers
IllusionReasoning â comparedto â POPE
confidence 80% ¡ Compared to previous perception-focused benchmarks, it comprehensively evaluates the capabilities of LVLMs.
IllusionReasoning â comparedto â IllusionVQA
confidence 80% ¡ IllusionReasoning is significantly different from them... compared to the many synthetic images in the IllusionBench+
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.
Tags
Links
- Source: https://arxiv.org/abs/2607.27747v1
- Canonical: https://arxiv.org/abs/2607.27747v1
Trouble viewing inline? Open PDF directly â
Full Text
52,873 characters extracted from source content.
Expand or collapse full text
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities Liangjie Zhao1 Jiaqing Lyu4 Kexin Tang5 Zecheng Fang3 Rong Yin6 Yulan Hu5 Da Li2,3 Jianing Li2,3â footnotemark: 1Adelaide University 2State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences 3University of Chinese Academy of Sciences 4Tsinghua University 5Amap, Alibaba Group 6Beihang University zhaoliangjie55@gmail.com, lida.ucas@gmail.com, lijianing@ict.ac.cn Corresponding authors. Abstract Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation. Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities Liangjie Zhao1 Jiaqing Lyu4 Kexin Tang5 Zecheng Fang3 Rong Yin6 Yulan Hu5 Da Li2,3â thanks: Corresponding authors. Jianing Li2,3â footnotemark: 1Adelaide University 2State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences 3University of Chinese Academy of Sciences 4Tsinghua University 5Amap, Alibaba Group 6Beihang University zhaoliangjie55@gmail.com, lida.ucas@gmail.com, lijianing@ict.ac.cn 1 Introduction Large Vision Language Models (LVLMs) (OpenAI, 2024; Bai et al., 2025) demonstrate outstanding performance in tasks such as visual question answering and multimodal dialogue by integrating visual perception modules with the language modeling capabilities of Large Language Models (Meta-AI, 2025; OpenAI et al., 2024b; Yang et al., 2025). Building upon this foundation, LVLMs further incorporate reasoning through strategies such as Chain-of-Thought and Reinforcement Learning, achieving human-like thinking capabilities exemplified by models such as OpenAI-o1 (OpenAI et al., 2024a) and DeepSeek-R1 (DeepSeek-AI et al., 2025). This integration enables LVLMs to tackle complex tasks such as logical visual question answering, pushing their capabilities to new levels. With the improvement of capabilities, the evaluation of LVLMs is shifting toward pursuing comprehensive competitiveness in multi-tasking (Liu et al., 2024; Chen et al., 2024a; Cheng et al., 2025). However, these evaluation frameworks originally focus on perception. As LVLMs such as Qwen3-VL (Bai et al., 2025) integrate reasoning capabilities, the evaluation starts to focus on their reasoning performance on complex tasks such as OlympiadBench (He et al., 2024) and MathVista (Lu et al., 2024b). However, such domain-specific evaluations are insufficient to fully capture the genuine reasoning capabilities of LVLMs. The core reason lies in their limited coverage beyond these specific domains. The requirement for paradigms that closely align with the complex logic scenarios in open-world environments is urgent, in order to better evaluate the reasoning capabilities of LVLMs. Visual illusions naturally satisfy these requirements. Visual illusions (Gregory, 1968, 1997; Bach and Poloschek, 2006) refer to phenomena where visual perception conflicts with cognitive reality, fundamentally stemming from biases within the human intuition. As a prevalent scenario of visual cognitive conflict in open worlds, visual illusions encompass diverse inputs from real-world settings. Addressing illusion-related questions requires LVLMs to invoke cross-modal common sense, causal logic, and cognitive reasoning capabilities to discern the underlying logic behind visual phenomena (Sun and Dekel, 2021). There has been some work focusing on illusion-related content as an evaluation benchmark for LVLMs. These studies typically select classic visual illusions such as the MĂźller-Lyer line illusion (GarcĂa-Garibay and de Lafuente, 2015) and the Penrose triangle spatial paradox (Plotnitsky, 1997). By designing questions such as âAre the two line segments in the image actually equal in length?â or âCould this three-dimensional figure exist in real physical space?â, these studies use the output to evaluate whether LVLMs can overcome misleading visual appearances rather than rely solely on raw visual features to conclude (Zhang et al., 2025; Shahgir et al., 2024; Zhang et al., 2023). The aforementioned evaluations focus primarily on the perceptual capabilities of LVLMs. In this work, we attempt to directly analyze the perceptual and reasoning capabilities of LVLMs through visual illusions. First, we constructed IllusionReasoning, a benchmark of visual illusions from the real world. Then, based on these images, we encourage human annotators to raise questions and provide answers across the following dimensions: (1) Detection: Given an object, can LVLMs determine whether it appears in an image? (2) Description: How do LVLMs describe illusion images? (3) Causes of Illusions: When informed about the content of visual illusions, can LVLMs explain the causes? After IllusionReasoning is built, we aim to evaluate LVLMs from different perspectives: (1) Visual Object Recognition: Do LVLMs possess basic image understanding capabilities? (2) Alignment Preference: Do LVLMs prefer human intuition or the physical world? (3) Reasoning: Can LVLMs analyze the causes of inconsistency between the illusion and the physical world? Among these perspectives, the first two are designed to evaluate the perceptual capabilities, while the last one is intended to evaluate reasoning capabilities. We evaluated a wide range of LVLMs on IllusionReasoning, encompassing open-source and closed-source models of varying sizes. We found that even state-of-the-art LVLMs struggle to provide correct responses to questions related to visual illusions. We also conducted a comprehensive analysis from multiple perspectives, including question categories, alignment preferences, and reasoning capabilities. We found that LVLMs do not consistently benefit from their thinking mode across different tasks. In simple reasoning-related tasks, excessive thinking may harm performance. Furthermore, when faced with hard problems, LVLMs tend to generate safe responses to avoid making claims on ambiguous content. This harms performance for illusion identification and may hardly be avoided under existing alignment paradigms. Refined strategies are needed to mitigate this phenomenon and enhance the capabilities of LVLMs. In summary, our contributions are as follows: ⢠We collected and constructed IllusionReasoning , a benchmark of visual illusions from the real world to evaluate the perception and reasoning capabilities of LVLMs. ⢠We conducted a comprehensive evaluation of widely used LVLMs. The results indicate that the reasoning capabilities of these models are not as good as claimed. ⢠We analyzed the perception and reasoning capabilities of LVLMs from multiple dimensions, providing direction for subsequent optimisation. 2 Related Work 2.1 Large Visual-Language Model By combining the semantic reasoning capabilities of Large Language Models (LLMs) with the perceptual capabilities of visual encoders, Large Visual-Language Models (LVLMs) deliver strong performance in visual question answering and embodied intelligence (Lu et al., 2024a; Huang et al., 2024). LVLMs typically consist of three components: a visual encoder, a connector, and an LLM. As the core of LVLMs, visual encoders, typically based on ViT (Dosovitskiy, 2020), process raw visual inputs into high-dimensional visual features. Then, a cross-modal alignment connector is required to project the visual features from the visual encoder into the space of the text and enable the LLM to interpret visual data as textual tokens (Liu et al., 2023; Chen et al., 2024b; Zhao et al., 2025; Zhu et al., 2023). This transformation empowers LLMs with visual capabilities, enabling them to complete vision-related tasks. 2.2 Evaluation of Hallucination Hallucinations pose a major obstacle to the application of LVLMs in real-world scenarios. Existing benchmarks have evolved from simple object-level evaluations to more complex evaluations of attributes and relations Bai et al. (2024). POPE (Li et al., 2023) established a benchmark for object existence, while AMBER Wang et al. (2023) extends the evaluation scope to measure hallucinations regarding attributes and relationships. Recently, PhD (Liu et al., 2025) introduced a ChatGPT-prompted benchmark design and expanded the evaluation perspectives. However, these hallucination benchmarks primarily focus on perceptual capabilities and lack a design tailored to the evaluation of reasoning capabilities. 2.3 Evaluation of Illusion Illusion-induced errors in LVLMs are considered as a special type of hallucination, and few works have specifically evaluated illusion. GVIL Zhang et al. (2023) pioneered this by assessing human alignment across five classic illusion types. HallusionBench Guan et al. (2024) introduced visual control groups to decouple visual misperception from language priors. IllusionVQA Shahgir et al. (2024) expanded the scope to 12 categories, incorporating soft localization for geometric inconsistencies. More recently, IllusionBench+ Zhang et al. (2025) scaled up evaluation with âTrap Illusionsâ and colorblindness tests to identify shortcut learning behaviors. Despite these advancements, existing datasets primarily rely on fixed question formats, allowing models to exploit random guessing or language shortcuts without demonstrating true visual comprehension. Furthermore, regarding the visual data itself, recent mechanism studies Ullman (2024); Shinozaki et al. (2025) use âanti-illusionsâ to confirm that illusion errors are primarily prior-driven reasoning errors rather than perception errors. This implies that models heavily rely on memorized knowledge when processing classic or synthetic illusions. Consequently, evaluating on these synthetic images fails to accurately measure genuine visual perception and reasoning. 3 Why Do We Need IllusionReasoning? 3.1 Model Capabilities: From Perception to Reasoning The evolution of LVLMs demonstrates a transition from perceptual processing to sophisticated reasoning capabilities. LVLMs, exemplified by LLaVA (Liu et al., 2023) and InternVL (Chen et al., 2024b), have acquired foundational perception capabilities and can tackle tasks such as image-text matching and simple visual question-answering. However, they remain confined to a superficial understanding of modal information. With breakthroughs in text reasoning by LLMs, some studies began exploring the adaptation of the Chain-of-Thought (CoT) reasoning paradigm to visual-language tasks. This adaptation gave rise to models such as LLaVA-CoT (Xu et al., 2025) and Insight-V (Dong et al., 2025), which demonstrated preliminary reasoning capabilities. This evolution reached new heights in LLMs such as GPT-5.5 (OpenAI, 2026) and DeepSeek-V4 (DeepSeek-AI, 2026), which demonstrated human-like stepwise reasoning in tasks such as logical question-answering. The corresponding optimization strategies have given rise to LVLMs with reasoning capabilities, such as Qwen3.5 (Qwen Team, 2026). To match the rapidly growing reasoning capability, stronger evaluation protocols are required to better evaluate the performance of LVLMs. Figure 1: Overview of IllusionReasoning Construction. Images in IllusionReasoning are collected from search engines to ensure authenticity. The annotation consists of two stages: (1) Content Analysis, which involves the determination of illusion types by comparing human descriptions with ground-truth reality, and (2) QA Construction, which encompasses the creation of high-quality detection, description, and reasoning questions paired with answers. 3.2 Why Use Visual Illusions? LVLMs demonstrate solid performance in fundamental tasks such as visual question answering and image captioning. These tasks struggle to reflect the performance differences between different LVLMs. In contrast, visual illusion tasks, which are more complex, can effectively reveal the performance differences between various LVLMs. To answer questions about illusions, LVLMs must conduct complex, multi-grained reasoning. This process requires comprehending the macroscopic semantic context while precisely pinpointing illusion-inducing features and integrating microscopic visual cues to see through the deceptive appearance and discern the underlying reality for the correct answer. Therefore, evaluation based on visual illusions can reflect the human-like degree of perception and reasoning capabilities, thereby offering a novel dimension beyond general evaluation. 4 IllusionReasoning Construction Visual illusion is defined as a phenomenon in which the perception of an observer regarding a physical attribute, including length, angle, color, or object motion, systematically deviates from objective reality under specific visual conditions. Existing evaluations primarily focus on the perceptual capabilities of LVLMs, while assessments of their reasoning abilities have largely relied on non-natural images. To bridge this gap, we constructed a new benchmark, IllusionReasoning. Employing natural illusion images, we aim to comprehensively evaluate both the perceptual and reasoning capabilities of LVLMs. The overall construction pipeline of IllusionReasoning is illustrated in Figure 1. 4.1 Key Objectives A rigorous evaluation of the capabilities of LVLMs in handling illusions should follow these objectives: ⢠Minimizing prior knowledge bias with real-world scenarios. To avoid the memorization interference caused by canonical synthetic illusions ingested during training, evaluations must prioritize authentic, real-world images. This ensures a measurement of genuine visual perception rather than mere training data memorization. ⢠Integrating binary and image-specific open-ended QA. Relying exclusively on binary settings allows models to exploit random guessing without true comprehension. Combining binary questions with open-ended inquiries that are carefully tailored to specific visual details enables a comprehensive assessment of visual understanding across distinct levels of task complexity. 4.2 Taxonomy of Visual Illusions Based on the underlying causes, we classify the visual illusions in our benchmark into five distinct categories: ⢠Morphological Illusion: Visual illusions caused by shape or occlusion. They mislead observers into misperceiving the integrity of a single object. ⢠Color and Background Illusion: Visual illusions caused by the color, brightness, texture, and contrast of an object against the background. These either cause incorrect perceptions of the object itself or lead to fusion with the background. ⢠Spatial Illusion: Visual illusions caused by perspective, distance, arrangement, or occlusion between objects, which can lead to misperceptions of size, position, depth, or relative relationships. ⢠Light and Shadow Illusion: Visual illusions caused by the effects of light, shadow, reflection, or refraction, distorting perceptions of the shape, brightness, or color of an object. ⢠Associative Illusion: Visual illusions caused by the shape, texture, or arrangement of an object resembling something familiar in human memory, triggering semantic associations. This misleads observers into incorrectly recognizing the object. 4.3 Data Collection The images in IllusionReasoning are collected from the real world. By querying search engines with the keyword âoptical illusions in the real worldâ, we can obtain a vast number of illusion images. We manually verified and filtered these images one by one. Images were excluded if they were blurry, irrelevant to visual illusions, or subject to copyright restrictions. Ultimately, we obtained a collection of authentic, high-quality illusion images. Although it is possible to generate hallucinatory-like images through data synthesis, such images exhibit partial visual logical errors, providing shortcuts for LVLMs when generating responses. Using visual illusion images from the real world ensures that IllusionReasoning is challenging. 4.4 Annotation We design a three-stage pipeline to process the collected images alongside metadata detailing the initial misperception of the uploader and the actual physical content. First, three independent annotators report their initial visual impressions; we retain only images aligning with the documented misperception. They also verify the factual accuracy of the metadata. This verification establishes validated pairs contrasting physical reality and human misperception. Second, annotators independently classify the cause of the illusion into five categories, resolving disagreements via majority voting. Finally, a fourth annotator reviews the finalized data for formatting and linguistic quality. For these images, we construct three types of questions: Detection, Description, and Reasoning. Detection questions are binary inquiries verifying object existence or attribute correctness. Description questions are open-ended inquiries to target specific attributes or request general content summaries, determining whether models exhibit visual illusions similar to humans. Reasoning questions evaluate whether LVLMs can infer the causes of the illusion by understanding the discrepancy between physical reality and human misperception. Annotators formulate QA pairs focusing on illusion-related regions, maximizing diversity and quantity without compromising accuracy. Each pair then undergoes at least two rounds of cross-verification and revision by independent annotators. This process filters out ambiguous entries, ensuring every question has a definitive and correct answer. 4.5 Statistics Figure 2: Statistics of IllusionReasoning. The inner ring shows question categories, with numbers representing the count and proportion of questions. The outer ring shows illusion categories, with numbers representing the count of questions and images. As shown in Figure 2, IllusionReasoning consists of 650 unedited real-world illusion images across five categories, expanding upon the Real-scene subset of IllusionVQA (60 images). To ensure novelty, image duplication rates with IllusionVQA and IllusionBench+ are strictly kept below 3% and 10%, respectively. We construct over 3,000 QA pairs (far exceeding the 435 pairs of IllusionVQA), distributed across detection, description, and reasoning tasks at an approximate 2:2:1 ratio. Binary questions are limited to mitigate random guessing. Overall, IllusionReasoning comprehensively evaluates and distinguishes the perception and reasoning capabilities of various LVLMs. 4.6 Comparison with Existing Benchmarks Although there are already some benchmarks about hallucination and illusion, IllusionReasoning is significantly different from them. We present examples from various benchmarks in Figure 3 to illustrate the differences. Model Size IllusionReasoning Avg Morphology Color & Background Space Light & Shadow Association [HTML]EFEFEFClosed-Source Gemini-3.1-Pro - 57.14 57.81 58.57 66.18 67.85 61.39 GPT-5.5 - 50.73 55.39 60.69 59.88 69.03 59.22 Seed-2.0 - 52.93 55.76 57.46 51.68 64.69 56.27 [HTML]EFEFEFOpen-Source Qwen3.5 397B(A17B) 55.49 52.60 58.08 57.98 61.93 57.28 InternVL3.5 241B(A28B) 48.35 48.88 45.48 47.73 48.13 47.49 GLM-4.6V 106B(A12B) 43.59 47.21 41.55 49.78 41.81 44.76 InternVL3.5 38B 41.21 44.24 34.65 44.07 43.79 41.07 Qwen3.6 35B(A3B) 46.89 47.21 57.83 54.32 56.41 52.93 Qwen3.6 27B 52.38 54.28 56.72 51.83 54.24 53.71 InternVL3.5 14B 34.25 36.80 31.69 35.72 41.42 35.53 GLM-4.6V 9B 48.72 51.86 34.77 44.66 42.01 43.60 Qwen3.5 9B 44.14 45.54 53.88 47.00 47.73 48.17 LLaVA-1.6 7B 25.46 42.01 25.40 39.24 36.88 33.26 MiMo-VL 7B 49.63 50.74 32.80 50.07 44.97 44.73 Phi-4 Multimodal 6B 18.42 30.56 19.83 23.30 29.20 24.19 Qwen3.5 4B 52.36 49.44 52.65 54.47 46.55 51.47 InternVL3.5 2B 28.39 32.16 27.00 31.92 37.87 31.02 Qwen3.5 2B 29.30 37.73 33.79 44.36 42.41 37.43 InternVL3.5 1B 25.09 30.48 24.04 26.65 32.15 27.26 Table 1: Results on IllusionReasoning. We report accuracy judged by GPT-4o. Best results are shown in bold. The suboptimal results are indicated by underline. Figure 3: Comparison between IllusionReasoning and existing visual understanding benchmarks. IllusionReasoning demonstrates several advantages: (1) It covers perception and reasoning tasks. Compared to previous perception-focused benchmarks, it comprehensively evaluates the capabilities of LVLMs. (2) IllusionReasoning evaluates whether LVLMs can accurately understand visual information, while hallucination benchmarks like POPE (Li et al., 2023) focus on evaluating the factual accuracy of the content generated by LVLMs. (3) Compared to the many synthetic images in the IllusionBench+ (Zhang et al., 2025), the images in IllusionReasoning are sourced from the real world and maintain authenticity. 5 Experiments 5.1 Choices of LVLMs We evaluate a diverse set of LVLMs, spanning both closed-source and open-source variants, with parameter counts ranging from several billions to hundreds of billions. For LVLMs equipped with reasoning capabilities, we primarily compare their performance in the instruction (non-thinking) mode, as shown in Table 1. In subsequent sections, we further analyzed the performance of these LVLMs in both thinking and non-thinking modes on IllusionReasoning. 5.2 Evaluation Setting To ensure a fair comparison, all LVLMs receive identical single-turn prompts consisting solely of an image and a query, with no prior conversational context. For LVLMs like LLaVA-1.6 (Li et al., 2024), the input is organized as follows using the template: âUSER: <|image|> question ASSISTANT:â. To facilitate automated evaluation, we explicitly append formatting instructions (e.g., âProvide an explicit conclusion.â), ensuring that outputs are structured and easy to judge. 5.3 LLM as a Judge In addition to serving as answer generators, LLMs offer a compelling alternative to traditional expert-driven evaluation (Gu et al., 2025). Beyond binary questions, IllusionReasoning incorporates free-form questions whose answers do not conform to a predefined structure. In order to conduct a precise evaluation, we employ GPT-4o as the evaluator to determine whether a modelâs response aligns with the annotated answers, inspired by OpenCompass (Contributors, 2023). We randomly sampled 200 cases and evaluated the differences between GPT-4o and human evaluation, finding that it achieves a consistency rate of 99%. For the parts where there are inconsistencies, we provide detailed cases in the Appendix A.1 to illustrate the reasons. 6 Main Results 6.1 Overall Performance Our evaluation covers a range of widely used LVLMs. And their performance on IllusionReasoning are summarized in Table 1. As indicated in the table, neither closed-source nor open-source LVLMs achieved satisfactory results compared to other benchmarks such as POPE (Li et al., 2023) and IllusionVQA (Shahgir et al., 2024), demonstrating the challenge of IllusionReasoning. Evaluations based on IllusionReasoning provide more discriminative insights into LVLMsâ capabilities. Open-source LVLMs underperformed closed-source ones on IllusionReasoning. However, the gap is not insurmountable, and their performance is comparable in some specialized sub-tasks. Among the five illusion categories in IllusionReasoning, LVLMs performed worst on morphology-related illusions and best on association-related ones. Results from open-source models of various sizes, ranging from 1B to hundreds of billions, indicate that the performance of LVLMs on IllusionReasoning is not determined by parameter size: small LVLMs can match the performance of models several times their scale. We also found that the performance is irrelevant to model architecture, with Qwen3.6 delivering similar performance across different model architectures. For example, Qwen3.6 shows similar performance across different architectures: MOE and Dense. This indicates that the performance of LVLMs on illusion mainly depends on the data. 6.2 Case Study To provide an intuitive analysis of the performance among LVLMs in IllusionReasoning , we demonstrated the different responses produced by LVLMs when given the same input in Figure 4. The example we chose is an image of a van with the gray convertible spray-painted on its right side. We feed this image to LVLMs, asking them to judge whether a real gray convertible car exists in the scene. We found that even the most advanced LVLMs struggle to provide the correct answer. This demonstrates that there is still considerable space for the reasoning capabilities of LVLMs to improve. Visual illusions can be used as a challenging task to measure the capability of LVLMs. Figure 4: Result comparison between Gemini-3.1-Pro, Qwen3.5-397B, and GPT-5.5. The characteristics within the red box indicate that the gray car does not actually exist. 7 Further Analysis 7.1 Can thinking help LVLMs to recognize illusions? Recent LVLMs have integrated reasoning capabilities through introducing thinking mode, demonstrating performance improvements on different evaluations. Whether thinking is effective for illusion recognition is to be examined. We conducted preliminary attempts based on IllusionReasoning. And the results are shown in Table 2. We evaluated LVLMs across different question categories: detection, description, and reason in both thinking and non-thinking modes. Size Detection Description Reason Gemini-3.1-Pro - 65.27 54.87 65.60 !20 Gemini-3.1-Pro - 67.83 56.32 68.96 GPT-5.5 - 64.49 51.54 62.72 !20 GPT-5.5 - 68.22 54.27 66.88 Seed-2.0 - 58.53 56.67 50.88 !20 Seed-2.0 - 62.87 59.32 60.32 Qwen3.5 397B(A17B) 67.98 45.47 57.28 !20 Qwen3.5 397B(A17B) 64.65 48.03 61.76 GLM-4.6V 106B(A12B) 47.60 43.93 40.48 !20 GLM-4.6V 106B(A12B) 52.48 47.69 44.64 Qwen3.6 35B(A3B) 62.79 43.59 50.08 !20Qwen3.6 35B(A3B) 59.15 45.90 59.20 Qwen3.6 27B 62.09 44.36 53.92 !20Qwen3.6 27B 60.85 46.24 62.24 Table 2: Performance comparison of LVLMs in thinking and non-thinking modes. Optimal results are shown in bold, and suboptimal results are denoted by underlining. Gray indicates responses yielded in thinking mode. Empirical results indicate that for tasks requiring reasoning about the underlying mechanisms of visual illusions, LVLMs exhibit a significant performance disparity between the thinking and non-thinking modes. However, in detection and description questions that focused on perceptual evaluation, LVLMs exhibit unstable performance improvements in thinking mode, even showing a declining trend, such as Qwen3.5 and Qwen3.6. This suggests that the integration of the perception and reasoning capabilities of LVLMs needs to be improved further. 7.2 Are LVLMs aligned with human intuition? We have demonstrated in Table 1 that the performance of existing LVLMs on IllusionReasoning is not satisfactory. We would like to analyze whether this should be attributed to their alignment with human intuition, i.e., LVLMs make mistakes like humans. We instruct LVLMs to describe the image without other restrictions and analyze the preference of responses. Besides preference to Human (illusion) and physical world (truth), we set a third category Neutrality for those responses lacking clear preference. Results are shown in Figure 5. Figure 5: Comparison of LVLMsâ preference for alignment between the human intuition and physical world. While some responses are aligned with human intuition, there are large portions of responses that fall into the Neutrality category, ranging approximately from 30% to 60% across different models. Most of such responses share a common feature: they simply ignore the core image content regarding the illusion, and make safe claims on less important content. A corresponding example is shown in Appendix A.2. Such observation points out the direction for future optimisation. Safe responses that avoid describing ambiguous content may be regarded as correct during alignment optimisation, if the judging criterion is simply about making no mistakes. However, safe responses miss important information and should not be preferred. To discourage the generation of such responses, we suggest refined strategies for alignment optimisation. For example, one could force the model to answer various questions regarding the core content, so as to avoid shortcuts for scoring. 7.3 Can LVLMs identify the underlying cause and correct logical errors? In this part, we analyzed whether LVLMs can provide correct reasons when being aware of the differences between real-world and illusory content. Specifically, we employed two distinct task formats: multiple-choice and free-form QA to evaluate the reasoning capabilities of LVLMs. For an image, a physical-world description, and a human-annotated illusion, the multi-choice task prompts LVLMs to select two relevant categories from the five illusion categories listed in Section 4. And the free-form QA tasks need LVLMs to provide a reason given physical reality and illusion. We evaluated whether the types of illusions annotated appeared in the output yielded by LVLMs. Results are displayed in the Table 3. Size Choice Free-form Gemini-3.1-Pro - 73.50 65.60 !20Gemini-3.1-Pro - 80.83 68.96 GPT-5.5 - 75.17 62.72 !20GPT-5.5 - 77.83 66.88 Seed-2.0 - 58.50 50.88 !20Seed-2.0 - 64.17 60.32 GLM-4.6V 106B(A12B) 55.00 40.48 !20GLM-4.6V 106B(A12B) 49.33 44.64 Qwen3.5 397B(A17B) 69.00 57.28 !20 Qwen3.5 397B(A17B) 66.83 61.76 Qwen3.6 35B(A3B) 73.17 50.08 !20Qwen3.6 35B(A3B) 67.67 59.20 Qwen3.6 27B 64.67 53.92 !20Qwen3.6 27B 59.50 62.24 Table 3: Comparison of LVLMsâ ability to reason about visual illusions in thinking and non-thinking modes. We found that under different modes of inference, the performances of the two tasks exhibited distinct trends. In the free-form task, LVLMs can provide precise explanations for the formation of illusions through thinking. When we provide the category of illusions generated and formalize the question of their causes as multi-choice tasks, LVLMs actually perform better without engaging in thinking. This indicates that while LVLMs incorporate reasoning capabilities, they also carry the risk of overthinking, which may compromise performance on simple reasoning tasks. Due to space constraints, we provided a detailed example in Appendix A.4 to illustrate this issue. 8 Conclusion In this work, we analyze the reasoning capabilities of existing LVLMs using visual illusions. We collected a set of real-world illusion images and constructed question-answer pairs designed to evaluate the perceptual and reasoning capabilities of LVLMs, forming a new benchmark: IllusionReasoning. Evaluation based on IllusionReasoning for open-source and closed-source LVLMs indicates that the reasoning capabilities of these models are not as competitive as claimed. And not all questions benefit from reasoning; the integration of LVLMsâ perceptual and reasoning capabilities requires further exploration. Further analysis of IllusionReasoning reveals the problem in optimization, emphasising the alignment of content while neglecting the alignment of focus. We hope that IllusionReasoning can be widely used by the community as an insightful benchmark. Limitations Illusions, as phenomena where human intuition diverges from reality, create an inherent conflict between perception and reality. We constructed IllusionReasoning, a benchmark comprising illusion images to evaluate the perceptual and reasoning capabilities of LVLMs. However, constrained by the stringent conditions under which illusions occur in the physical world, the volume of data in IllusionReasoning is insufficient to support training. We will continue to collect similar images and consider how to incorporate illusion images into the training process. References M. Bach and C. M. Poloschek (2006) Optical illusions. Adv Clin Neurosci Rehabil 6 (2), p. 20â21. Cited by: §1. S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu (2025) Qwen3-vl technical report. External Links: 2511.21631, Link Cited by: §1, §1. Z. Bai, P. Wang, T. Xiao, T. He, Z. Han, Z. Zhang, and M. Z. Shou (2024) Hallucination of multimodal large language models: a survey. arXiv preprint arXiv:2404.18930. Cited by: §2.2. L. Chen, J. Li, X. Dong, P. Zhang, Y. Zang, Z. Chen, H. Duan, J. Wang, Y. Qiao, D. Lin, and F. Zhao (2024a) Are we on the right way for evaluating large vision-language models?. External Links: 2403.20330, Link Cited by: §1. Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. (2024b) Internvl: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 24185â24198. Cited by: §2.1, §3.1. X. Cheng, W. Zhang, S. Zhang, J. Yang, X. Guan, X. Wu, X. Li, G. Zhang, J. Liu, Y. Mai, Y. Zeng, Z. Wen, K. Jin, B. Wang, W. Zhou, Y. Lu, T. Li, W. Huang, and Z. Li (2025) SimpleVQA: multimodal factuality evaluation for multimodal large language models. External Links: 2502.13059, Link Cited by: §1. O. Contributors (2023) OpenCompass: a universal evaluation platform for foundation models. Note: https://github.com/open-compass/opencompass Cited by: §5.3. DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025) DeepSeek-r1: incentivizing reasoning capability in llms via reinforcement learning. External Links: 2501.12948, Link Cited by: §1. DeepSeek-AI (2026) DeepSeek-v4 technical report. Note: https://github.com/deepseek-ai/DeepSeek-V4Accessed: 2026-05-22 Cited by: §3.1. Y. Dong, Z. Liu, H. Sun, J. Yang, W. Hu, Y. Rao, and Z. Liu (2025) Insight-v: exploring long-chain visual reasoning with multimodal large language models. External Links: 2411.14432, Link Cited by: §3.1. A. Dosovitskiy (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §2.1. O. B. GarcĂa-Garibay and V. de Lafuente (2015) The mĂźller-lyer illusion as seen by an artificial neural network. Frontiers in computational neuroscience 9, p. 21. Cited by: §1. R. L. Gregory (1997) Knowledge in perception and illusion. Philosophical Transactions of the Royal Society of London. Series B: Biological Sciences 352 (1358), p. 1121â1127. Cited by: §1. R. L. Gregory (1968) Perceptual illusions and brain models. Proceedings of the Royal Society of London. Series B. Biological Sciences 171 (1024), p. 279â296. Cited by: §1. J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y. Wang, W. Gao, L. Ni, and J. Guo (2025) A survey on llm-as-a-judge. External Links: 2411.15594, Link Cited by: §5.3. T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, et al. (2024) Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14375â14385. Cited by: §2.3. C. He, R. Luo, Y. Bai, S. Hu, Z. L. Thai, J. Shen, J. Hu, X. Han, Y. Huang, Y. Zhang, J. Liu, L. Qi, Z. Liu, and M. Sun (2024) OlympiadBench: a challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems. External Links: 2402.14008, Link Cited by: §1. W. Huang, A. Wu, Y. Yang, X. Luo, Y. Yang, L. Hu, Q. Dai, C. Wang, X. Dai, D. Chen, et al. (2024) Llm2clip: powerful language model unlocks richer visual representation. arXiv preprint arXiv:2411.04997. Cited by: §2.1. F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li (2024) LLaVA-next-interleave: tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895. Cited by: §5.2. Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. Wen (2023) Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355. Cited by: §2.2, §4.6, §6.1. H. Liu, C. Li, Q. Wu, and Y. J. Lee (2023) Visual instruction tuning. Advances in neural information processing systems 36, p. 34892â34916. Cited by: §2.1, §3.1. J. Liu, Y. Fu, R. Xie, R. Xie, X. Sun, F. Lian, Z. Kang, and X. Li (2025) PhD: a chatgpt-prompted visual hallucination evaluation dataset. External Links: 2403.11116, Link Cited by: §2.2. Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, K. Chen, and D. Lin (2024) MMBench: is your multi-modal model an all-around player?. External Links: 2307.06281, Link Cited by: §1. H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. (2024a) Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525. Cited by: §2.1. P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao (2024b) MathVista: evaluating mathematical reasoning of foundation models in visual contexts. External Links: 2310.02255, Link Cited by: §1. Meta-AI (2025) The llama 4 herd: the beginning of a new era of natively multimodal ai innovation. External Links: Link Cited by: §1. OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. OâConnell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li (2024a) OpenAI o1 system card. External Links: 2412.16720, Link Cited by: §1. OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Ĺ. Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Ĺ. Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. MĂŠly, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. OâKeefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph (2024b) GPT-4 technical report. External Links: 2303.08774, Link Cited by: §1. OpenAI (2024) Hello gpt-4o. External Links: Link Cited by: §1. OpenAI (2026) GPT-5.5 system card. Note: https://openai.com/index/gpt-5-5-system-card/Accessed: 2026-05-22 Cited by: §3.1. A. Plotnitsky (1997) Penroseâs triangles: the large, the small, and the human mind. Postmodern Culture 7 (3). Cited by: §1. Qwen Team (2026) Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §3.1. H. S. Shahgir, K. S. Sayeed, A. Bhattacharjee, W. U. Ahmad, Y. Dong, and R. Shahriyar (2024) Illusionvqa: a challenging optical illusion dataset for vision language models. arXiv preprint arXiv:2403.15952. Cited by: §1, §2.3, §6.1. T. Shinozaki, T. Doi, A. Watahiki, S. Nishida, and H. Yanaka (2025) Do large vision-language models distinguish between the actual and apparent features of illusions?. External Links: 2506.05765, Link Cited by: §2.3. E. D. Sun and R. Dekel (2021) ImageNet-trained deep neural networks exhibit illusion-like response to the scintillating grid. Journal of Vision 21 (11), p. 15â15. Cited by: §1. T. Ullman (2024) The illusion-illusion: vision language models see illusions where there are none. External Links: 2412.18613, Link Cited by: §2.3. J. Wang, Y. Wang, G. Xu, J. Zhang, Y. Gu, H. Jia, J. Wang, H. Xu, M. Yan, J. Zhang, et al. (2023) Amber: an llm-free multi-dimensional benchmark for mllms hallucination evaluation. arXiv preprint arXiv:2311.07397. Cited by: §2.2. G. Xu, P. Jin, Z. Wu, H. Li, Y. Song, L. Sun, and L. Yuan (2025) LLaVA-cot: let vision language models reason step-by-step. External Links: 2411.10440, Link Cited by: §3.1. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu (2025) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §1. Y. Zhang, J. Pan, Y. Zhou, R. Pan, and J. Chai (2023) Grounding visual illusions in language: do vision-language models perceive illusions like humans?. arXiv preprint arXiv:2311.00047. Cited by: §1, §2.3. Y. Zhang, Z. Zhang, X. Wei, X. Liu, G. Zhai, and X. Min (2025) IllusionBench+: a large-scale and comprehensive benchmark for visual illusion understanding in vision-language models. External Links: 2501.00848, Link Cited by: §1, §2.3, §4.6. Y. Zhao, Y. Yin, L. Li, M. Lin, V. S. Huang, S. Chen, W. Chen, B. Yin, Z. Zhou, and W. Zhang (2025) Beyond sight: towards cognitive alignment in lvlm via enriched visual knowledge. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 24950â24959. Cited by: §2.1. D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny (2023) Minigpt-4: enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592. Cited by: §2.1. Appendix A Appendix A.1 Inconsistencies in Evaluation Here, we demonstrate the cases where LLMs, acting as judges, produce results that differ from those of human annotators. As shown in Figure 6, we take the output from InternVL3.5-14B as an example. In addition to the generated answer, the output includes the corresponding explanation, which contains incorrect information. When GPT-4o is used as a judge, its evaluation is flexible, overlooking inconsistencies in the explanations. However, human annotators are strict and consider that InternVL has not produced the correct answer. Given that there are few disagreements between the evaluation produced by GPT-4o acting as a judge and the human annotations. We consider that GPT-4o provides a relatively reliable evaluation for IllusionReasoning. Figure 6: A case where an LLM as a judge differs from human annotations. A.2 Alignment Preferences in LVLMs Figure 7: An example of a neutral output. Taking Qwen as an example, the output contains two possibilities, making it difficult to determine whether it aligns more closely with human perception or the physical world. A.3 Overthinking in Detection Question Figure 8: An example of reasoning that has negative effects. Taking Qwen3.6 as an example, in the thinking mode, Qwen3.6 produced a seemingly sophisticated but actually incorrect explanation, which led to incorrect results. A.4 Overthinking in Multi-Choice Figure 9: An example of reasoning that has negative effects. Taking Qwen3.5 as an example, in thinking mode, Qwen3.5 produced a confident result that actually excluded the correct answer. A.5 Prompt used in IllusionReasoning Detection Questions For binary questions (expecting a judgment word like "yes" or "no"): 1. If the modelâs answer lacks a clear judgment word, respond with False. 2. If the modelâs answer includes a judgment word: i. Respond with True if it matches the reference answer. i. Respond with False if it does not match. Figure 10: Evaluation Prompt for Detection Questions. Description Questions For special questions (e.g., "how many", "what", "which", etc.): 1. Respond with False if the modelâs answer lacks a clear conclusion. 2. If the modelâs answer includes a clear conclusion, compare its core information with the reference answer: i. Respond with True if the semantic meaning matches. i. Respond with False if it does not match. Figure 11: Evaluation Prompt for Description Questions. Reasoning Questions For reasoning questions (e.g., "why" and "how"): 1. If the modelâs answer includes multiple explanations or conclusions, it is correct as long as at least one matches the reference answer. 2. Respond with True if the semantic meaning of the modelâs final result matches the reference answer; otherwise, respond with False. Figure 12: Evaluation Prompt for Reasoning Questions. General Rules for All Questions 1. The modelâs answer must not contradict the reference answer. 2. Vague answers are acceptable if they include key information from the reference answer and do not introduce errors or contradictions. 3. Focus on whether the semantic meaning of the modelâs answer matches the reference answer. 4. Ignore differences in language (e.g., Chinese vs. English), case, punctuation, grammar, or word order. 5. Disregard intermediate reasoning or steps in the modelâs answer and evaluate only the final result or conclusion. Figure 13: General Rules for All Questions.