Paper deep dive
Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models
Zhiming Yang, Zhuoxi Xiong, Donglin Zhou, Wenjun Wei, Shiyao Cui, Jinqiao Shi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/30/2026, 2:09:10 AM
Summary
This paper introduces 'situational illusions,' a phenomenon where real-world visual appearances deviate from underlying physical states, challenging Multimodal Large Language Models (MLLMs). The authors propose a where-what-how taxonomy and introduce MSIBench, a benchmark with 3,723 image-text pairs to evaluate MLLM discrimination, understanding, and reasoning. Evaluations of 27 model configurations reveal high vulnerability, with average accuracy below 63% and action planning below 45%. The study identifies six failure modes and proposes mitigation strategies via prompting for closed-source models and supervised fine-tuning (SFT) for open-source models, improving performance by up to 20%.
Entities (13)
Relation Signals (12)
Situational Illusion → challenges → MLLM
confidence 95% · Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs)
MSIBench → evaluates → MLLM
confidence 95% · MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions.
InternVL3.5 → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... InternVL3.5
Gemini-3.5-Flash → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... Gemini-3.5-Flash
Grok 4.1 Fast → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... Grok-4.1-Fast
Qwen3.5 → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... Qwen3.5
GPT-5.5 → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... GPT-5.5
GPT-5.4 → evaluatedin → MSIBench
confidence 92% · Evaluations of 27 model configurations... including... GPT-5.4
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.
Tags
Links
- Source: https://arxiv.org/abs/2608.22232v1
- Canonical: https://arxiv.org/abs/2608.22232v1
Trouble viewing inline? Open PDF directly →
Full Text
47,453 characters extracted from source content.
Expand or collapse full text
Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models Zhiming Yang 1∗ , Zhuoxi Xiong 1∗ , Donglin Zhou 1 , Wenjun Wei 1 , Shiyao Cui 2† , Jinqiao Shi 2 1 School of Artificial Intelligence, Beijing University of Posts and Telecommunications, China 2 Beijing University of Posts and Telecommunications, China yang.zhiming, xiong.zhuoxi, cuishiyao@bupt.edu.cn, Abstract Real-world situation appearances can deviate from their un- derlying physical states, challenging the reliability of multi- modal large language models (MLLMs) in practical appli- cations. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first de- velop a comprehensive where–what–how taxonomy that char- acterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we intro- duce MSIBench, a benchmark designed to assess the discrim- ination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model config- urations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically in- specting and reasoning over visual evidence for contextual un- derstanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respec- tively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path to- ward more reliable multimodal perception and reasoning in complex real-world environments. Code — https://github.com/yangzm1105/MSIBench Datasets — https://huggingface.co/datasets/yangzm05/msibench 1 Introduction Reliable decision-making in real-world situations is essential for multimodal large language models (MLLMs), given their growing deployment in embodied agents (Gao et al. 2025; Yang et al. 2025) and humanoid robots (Yang et al. 2026). This capability requires models to interpret visual evidence in context and accurately infer the underlying physical state of a scene, thus supporting dependable understanding and interaction (Zhou et al. 2025; Chen et al. 2025a). However, complicated situational conditions and elements can cause the visual appearance to deviate from its real state. As illustrated in Figure 1, a mug filled with milk may appear to be placed upside down, since the milk surface resembles ∗ These authors contributed equally. † Corresponding author. It’s placedupside down. The upright ceramic mug isfilled with milk, whose matching color and reflectionsmake its surface resemble the mug’s bottom. How is the mug placed? Insert a straw in themug. Flip overthe mugand... Get a straw and insert ... Situation It’s placed upright. Figure 1: An example of situational illusion and the MLLM responses with incorrect answers in red dialog boxes. the bottom of the mug. Consequently, MLLMs could misin- terpret the mug’s physical state. Especially when instructed with Insert a straw in the mug, it plans actions with flip over the mug as the first step, which can spill the milk, leading to safety hazards. Unfortunately, in our pilot study of 30 cases featuring deceptive situational appearances, GPT-5.4 (Ope- nAI 2026) failed on nearly 40% of the samples, suggesting that even advanced MLLMs may struggle when a scene’s visual appearance diverges from its underlying reality. We refer to this phenomenon as situational illusion, namely where a naturally context-induced visual appearance may misrepresent the underlying physical state, which has yet to be explored. Specifically, existing studies mainly explore the contextual understanding and reasoning capabilities of MLLMs in real-world situations (Zhou et al. 2025; Zhang et al. 2025a), assuming that visual appearance is consistent with reality. While visual illusions have also been studied, ex- isting studies mainly focus on deliberately designed visual il- lusion patterns like gestalt illusion (Hou et al. 2026), pareido- lia (Rostamkhani et al. 2025) and geometric illusion (Zhang et al. 2023; Shahgir et al. 2024), or adversarially misleading visual cues in structured data like charts (Chen et al. 2025b; Bharti et al. 2025; Tonglet et al. 2026) and tables (Guan et al. 2024). Therefore, systematic investigation is warranted to understand situational illusions and explore how MLLMs perform under such visual conditions. To characterize this concern, this paper systematically in- arXiv:2608.22232v1 [cs.AI] 23 Aug 2026 Layout Illusion Mechanism Experience- driven Bias Scenario Small Indoor Enclosures A cat lies on its belly on the sofa, looking back at its feet. Reference-frame Illusion Mechanism Angle Disruption Scenario Urban Outdoor Scenes A man hangs by one hand from a stone ledge. Plausibility Illusion Mechanism Dimension Confusion Scenario Small Indoor Enclosure A narrow shower stall seems built inside the glass cabinet. Adjacency Illusion Mechanism Dimension Confusion Scenario Small Indoor Enclosures A dog seems to be resting on the sofa in the adjacent room. Pose Illusion Mechanism Perspective Misalignment Scenario Natural Outdoor Scenes A girl is looking at the boy standing on the left of her. Property Relation Material Scenario Mechanism Dimension Confusion Urban Outdoor Scenes Orientation Scenario Mechanism Angle Disruption Urban Outdoor Scenes Shape Scenario Mechanism Experience- driven Bias Large Indoor Venues Color Scenario Mechanism Boundary Ambiguity Natural Outdoor Scenes Quantity Scenario Mechanism Experience- driven Bias Small Indoor Enclosures Size Scenario Mechanism Perspective Misalignment Urban Outdoor Scenes Context Identity Illusion Mechanism Perspective Misalignment Scenario Natural Outdoor Scenes A white body horse with a dappled grey head faces the camera. Figure 2: Illustration for the taxonomy towards situational illusions. vestigates how MLLMs perform under situational illusions and how to mitigate the limitations. To achieve the research goal, we study the problems from three aspects: 1) Construct a fine-grained taxonomy for situational illusions. We build a comprehensive taxonomy regarding where, what, and how, through which situational illusions may occur, providing a comprehensive framework for ana- lyzing and understanding the phenomenon. 2) Evaluate MLLM performances under situational il- lusions. To evaluate how M LLMs perform under situational illusions, we construct MSIBench, comprising tasks to eval- uate capabilities of the illusion discrimination, understand- ing, and reasoning, together with adversarial settings for the stress test. Evaluations of widely used MLLMs reveal key limitations and insights for further improvement. 3) Develop mitigation strategies for improvement. Based on failure modes derived from evaluations, we ex- plored mitigations via prompting for closed-source models and supervised fine-tuning (SFT) for open-source models, aiming to guide models to attend visual evidence for more accurate interpretation of situational illusions. Taken together, we present a systematic study on situa- tional illusions, a critical yet insufficiently studied issue for MLLMs. The taxonomy spans 4 scenarios, 12 misperceived targets, and 5 formation mechanisms, which motivates the MSIBench consisting of 3,723 image-text pairs. Through elaborately designed tasks, evaluations across 27 model con- figurations reveal substantial vulnerabilities, with an overall average accuracy below 63%, and even drops to below 45% for action planning. Built on failure modes identified in mod- els’ observation, grounding as well as reasoning, prompting and SFT methods improve model performances by up to 20% using limited data, offering a practical path toward more re- liable multimodal decisions in real-world situations. 2 Taxonomy To understand situational illusions, we first develop a tax- onomy encompassing scenarios, targets and mechanisms, which characterize where such illusions occur, what is mis- perceived and how they arise. 2.1 Scenarios: where situational illusions occur Unlike deliberately constructed illusions, situational illusions naturally occur across diverse everyday environments. We therefore categorize representative scenarios below, with de- tailed examples in Appendix A.1. Small indoor enclosures refer to individual enclosed spaces, such as classrooms and offices. Large indoor venues refer to expansive, continuous in- door environments, such as concert halls and shopping malls. Urban outdoor scenes mean outdoor scenarios composed mainly of human-made elements like streets and parking lots. Natural outdoor scenes refer to outdoor environments dominated by natural elements, such as deserts and beaches. 2.2 Targets: what is misperceived We further characterize what is misperceived by categorizing the types of targets that are commonly misperceived, with examples in Figure 2: Property refers to misperceptions of an entity’s intrinsic properties, including 1) size, 2) quantity, 3) color, 4) shape, 5) orientation and 6) material. The white mug seems upside down. Is the bottom of the mug visible? Instruction Task Construction T/F Question Open-ended Question Is the bottom of the mug visible? Insert a straw into the mug. Vanilla Adv. 1 2 3 The white mug seems upside down. Situational Illusion What is inside the mug? Vanilla The white mug seems upside down. What is inside the mug? Adv. Image Input Figure 3: Example of tasks in MSIBench. Relation, going beyond individual entities, captures mis- perceptions of physical or spatial relationships across enti- ties, including 1) identity concerning whether objects belong to the same entity, 2) pose regarding relative position and ori- ented angle between entities, and 3) adjacency for whether entities are adjacent or in contact. Context captures misperceptions in scene-level under- standing and reasoning, including 1) plausibility refers to real scenes being mistaken for artificial ones due to their physi- cally implausible appearance, 2) layout refers to interpreting a scene’s overall layout as a more familiar configuration based on prior knowledge, and 3) reference-frame: disproportion- ate scaling or rotation of the viewing perspective leads to an incorrect overall understanding of the scene. 2.3 Mechanisms: how situational illusions arise We then characterize how such illusions are formed. Drawing on studies of visual perception (Wagemans et al. 2012), scene understanding (Oliva and Torralba 2001) and human cogni- tion (Treisman and Gelade 1980), we organize formation mechanisms from low-level perceptual ambiguity to high- level cognitive interpretation with the following aspects: Boundary ambiguity: ambiguous boundaries cause enti- ties or parts to be incorrectly merged or split. Angle disruption: unusual viewpoints distort perceived orientation and gravity direction. Perspective misalignment: specific viewpoints align ob- jects at different depths, distorting depth perception. Dimension confusion: misleading visual information blurs the distinction between 2D and 3D. Experience-driven bias: learned expectations lead to in- correct interpretations of visual appearance. 3 MSIBench Construction This section details how we construct data for systematic investigation of situational illusion in MLLMs. 3.1 Data Collection To build a comprehensive collection of real-world illusion images, we first gather in-the-wild images from two sources. Online platforms. Given the widespread presence of scattered visual illusion collections across online commu- nities, we collect 1,096 relevant images from platforms in- cluding Reddit (Reddit 2026), Zhihu (Zhihu 2026), Bright- Side (Bright Side 2026), and Barnorama (Barnorama 2026). Specifically, we manually scrape images from these sources that align with the concept of real-world situational illusions. Existing resources. To broaden our image coverage, we further examine existing researches. We extract il- lusion images featuring real-world scenes from Illusion- Bench+ (Zhang et al. 2025b). After removing duplicates, we obtain an initial set of 417 usable images. The collected images span a diverse range of everyday scenarios and illusion types. To ensure high data quality and mitigate potential confounding factors, we apply a rigorous manual filtering protocol. Specifically, we systematically ex- clude non-realistic illusions, illusions requiring specialized apparatus, images with ambiguous interpretations, and sam- ples containing sensitive elements. Ultimately, this rigorous selection process yields a final benchmark comprising 904 high-quality real-world illusion images that strictly satisfy the evaluation requirements. 3.2 Metadata Construction To enhance the understanding of the illusion depicted in each image, we construct structured metadata that comprises a description of the illusion, an analysis of the underlying real-world state, a scenario type, the misperceived targets, and the mechanism behind the illusion. To this end, we em- ploy Gemini-3.1-pro-preview (Think) to generate the initial metadata, given the strong visual reasoning performance of the model (details in Appendix C.2). The metadata provides comprehensive explanatory context for each image by de- scribing the observed phenomenon of situational illusion, identifying the mechanism that produces it, and clarifying the actual physical reality. 3.3 Task Construction To explore how MLLMs handle situational illusions, we design three kinds of tasks regarding three core capabili- ties, namely True-or-False (T/F) questions for discrimination, open-ended questions for understanding, and action planning tasks for reasoning. Further, for the stress test, we introduce vanilla and adversarial settings for the first two task cate- gories, where the model is prompted with the question di- rectly in the vanilla setting, whereas the illusion description is prepended to the question in the adversarial settings. T/F questions assess whether MLLMs can correctly iden- tify the presence of an illusion in an image. Specifically, given an image and a general question, the model must respond only with T (True) or F (False). Open-ended questions evaluate whether MLLMs can achieve a genuine understanding of the situation and answer related questions accurately. The model is given an image and a special question beginning with words such as what, where, and how. Valid answers in different wording are accepted. Action planning assesses whether MLLMs can plan an appropriate sequence of actions to achieve a given goal, re- quiring the model to reason about object states, physical constraints, action consequences and multi-step correlations. The model receives an image with a goal instruction and returns a sequence of two to six numbered actions. The in- struction requires the model to interact with a target related to the illusion, while the model must determine its response according to the actual physical scene. We also employ Gemini-3.1-pro-preview (Think) to con- struct these tasks. Detailed instructions are provided in Ap- pendix C.3, with examples of the tasks shown in Figure 3. 3.4 Quality Check and Filtering For data quality, we organize authors and colleagues into annotation and review teams for a two-round check. First, the annotation team inspects the generated data and makes necessary revisions. Specifically, they verify the model-generated metadata against the corresponding images and correct inaccurate or incomplete descriptions. Then, for T/F questions and open-ended questions, they examine the generated question pairs and revise those that are ambigu- ous or can be answered without understanding the illusion. For action-planning tasks, the team selects images with fea- sible goals that require accurate inference of the underlying physical scene and writes executable action instructions. Then, the review team conducts a second-round quality check on the annotated data, including the revised meta- data, question pairs, and action-planning instructions. Each submission is either approved or returned with feedback for further revision, and this process continues until the data meets the required quality standards. 3.5 Data Statistics Finally, MSIBench contains 904 unique images, each ac- companied by metadata describing the illusion phenomenon, scenario, misperceived target, formation mechanism, and an analysis of how the illusion obscures the true state of the scene. Using these images, we construct 3,723 image-text pairs for evaluation tasks, including 1,808 T/F questions, 1,808 open-ended questions, and 107 action-planning tasks. For each image, both a T/F question and an open-ended ques- tion are provided in vanilla and adversarial versions. 4 Experiments This section first presents results on MLLM performances, and then provides analysis suggesting improvement insights. 4.1 Target Models We conduct evaluations on 20 representative models with 27 model configurations, including 6 closed-source mod- els of Claude-Opus-4.7 (Anthropic 2026), Gemini-3.5- Flash (Kavukcuoglu et al. 2026), GPT-5.5 (OpenAI 2026), GPT-5.4 (OpenAI 2026), GPT-5.4-mini (OpenAI 2026) and Grok-4.1-Fast (xAI 2025) as well as 14 open-source mod- els of Gemma 3 (Gemma Team 2025), Mimo-V2.5 (Xiaomi MiMo 2026), Qwen3-VL (Bai et al. 2025), Qwen3.5 (Qwen Team 2026), InternVL3.5 (Wang et al. 2025), GLM- 4.6V (GLM-V Team 2026), Kimi-K2.5 (Kimi Team 2026) and Kimi-VL (Kimi Team 2025). Note that Qwen3.5 and InternVL3.5 are tested across various model scales. Fur- thermore, GPT-5.5, Grok-4.1-Fast, Gemini-3.5-Flash, GLM- 4.6V-Flash, Kimi-VL-A3B, and Qwen3-VL-8B are evalu- ated under different reasoning effort levels. To ensure re- producibility, we set the temperature parameter to 0.0 and employ a greedy decoding strategy. Other settings are pro- vided in Appendix B in detail. 4.2 Evaluation Metrics Metric. We use accuracy (Acc.) to measure whether MLLMs can make correct decisions for each task, which is defined as Acc. = N correct N total × 100% and a higher Acc. indicates stronger resistance to situational illusions. Criteria. For T/F questions, a response is considered cor- rect if it matches the ground-truth label. For open-ended questions, a response is considered correct if it accurately reflects the physical reality of the scene. For action-planning tasks, a response is considered correct if the proposed steps are grounded in the actual physical scene and can success- fully accomplish the specified task. LLM-as-a-Judge. T/F questions are evaluated through direct label matching, whereas open-ended questions and action-planning tasks are assessed using LLM-as-a-Judge. Our pilot study shows that, when provided with the corre- sponding illusion analysis, Gemini-3.1-Pro-Preview (Think) achieves over 93% accuracy in judging model responses, demonstrating its reliability as an automatic evaluator. Hence, the model is instructed as the judge to determine whether the response is correct given the task input, the model response, and the corresponding illusion analysis. De- tails could be found in Appendix C.5. 4.3 Main Results Table 1 reports the performance of the evaluated models across tasks, where we make the following observations. Existing MLLMs remain vulnerable to situational illu- sions. In Table 1, the highest overall average accuracy across all tasks remains below 63%, while the average accuracy on action-planning tasks falls below 45%, highlighting the lim- ited reliability of current MLLMs. Moreover, closed-source models consistently outperform open-source models across all three tasks, with the largest average performance gap ap- proaching 16% on action-planning. This gap reflects that the large-scale advanced models are more robust to potentially misleading illusions, suggesting that overall model capability impacts how models handle situational illusions. Increased reasoning effort provides limited and incon- sistent benefits. For most models, the effect of increased rea- soning effort is task-dependent and generally modest, with performance changes typically remaining below 10% points. For example, GLM-4.6V-Flash achieves its largest improve- ment of 8.41% on the action-planning task, whereas Gemini- 3.5-Flash declines on all three tasks, with a maximum drop of 5.61%. We attribute this to longer reasoning introducing unnecessary assumptions and distracting the model from the actual scene. Therefore, increased reasoning effort does not inherently improve performance, as misdirected reasoning may instead lead to performance degradation. T/F Question Acc. (%)Open-Ended Question Acc. (%)Action Planning Acc. (%) Model VanillaAdv.Avg.VanillaAdv.Avg.∆ t/f Score∆ open Closed-Source GPT-5.551.9950.6651.3349.7851.2250.500.83↓43.936.57↓ GPT-5.459.2959.4059.3560.1857.4158.800.55↓45.7913.01↓ GPT-5.4-Mini50.0043.5846.7947.4638.2742.873.92↓46.733.86↑ Grok-4.1-Fast67.2661.5064.3850.1154.0952.1012.28↓20.5631.54↓ Claude-Opus-4.767.7066.8167.2668.3670.8069.582.32↑39.2530.33↓ Gemini-3.5-Flash79.8781.5380.7083.9686.9585.464.76↑63.5521.91↓ GPT-5.5 (Medium)59.7356.3158.0257.9661.2859.621.60↑49.5310.09↓ GPT-5.5 (xhigh)58.3054.4256.3656.6454.9855.810.55↓57.011.20↑ Grok-4.1-Fast (Think)63.9460.9562.4553.7659.2956.535.92↓22.4334.10↓ Gemini-3.5-Flash (Think)79.0980.3179.7083.0886.1784.634.93↑57.9426.69↓ Average (Closed-Source)63.7261.5562.6361.1362.0561.591.04↓44.6716.92↓ Open-Source Gemma-3-12B58.0854.9856.5343.1444.0343.5912.94↓24.3019.29↓ Mimo-V2.566.1562.2864.2265.3867.7066.542.32↑42.0624.48↓ Qwen3-VL-8B-Instruct58.8554.9856.9253.9851.9952.993.93↓36.4516.54↓ Qwen3.5-2B61.9565.6063.7844.0337.8340.9322.85↓15.8925.04↓ Qwen3.5-4B49.5650.8850.2251.1157.7454.434.21↑24.3030.13↓ Qwen3.5-9B52.3251.5551.9450.3352.7751.550.39↓32.7118.84↓ Qwen3.5-27B56.9752.9954.9858.8562.7260.795.81↑32.7128.08↓ InternVL3.5-2B-Instruct46.7938.1642.4836.5026.2231.3611.12↓25.236.13↓ InternVL3.5-8B-Instruct56.5348.4552.4943.0333.7438.3914.10↓29.918.48↓ InternVL3.5-14B-Instruct53.9850.0051.9947.4639.2743.378.62↓19.6323.74↓ InternVL3.5-38B-Instruct62.3958.6360.5153.5445.5849.5610.95↓28.9720.59↓ GLM-4.6V-Flash49.3443.9246.6347.0137.0642.044.59↓25.2316.81↓ Kimi-K2.562.1755.7558.9665.2765.6065.446.48↑37.3828.06↓ Kimi-VL-A3B-Instruct50.1144.9147.5143.2534.7338.998.52↓24.3014.69↓ GLM-4.6V-Flash (Think)54.3151.0052.6650.0044.4747.245.42↓33.6413.60↓ Kimi-VL-A3B-Thinking55.7553.6554.7047.9040.7144.3110.39↓32.7111.60↓ Qwen3-VL-8B-Thinking59.7358.4159.0751.3350.1150.728.35↓29.9120.81↓ Average (Open-Source)56.1852.7154.4550.1246.6048.376.08↓29.1419.23↓ Table 1: Main results, where ∆ t/f is the difference between the average open-ended and true-or-false (T/F) accuracy, while ∆ open is the difference between the action-planning accuracy and average open-ended accuracy. Performance does not consistently transfer across tasks. Across the 27 evaluated model configurations, most exhibit performance declines in both cross-task compar- isons: from T/F to open-ended questions and from open- ended questions to action planning. For example, Qwen3.5- 2B records ∆ t/f =−22.85, while Grok-4.1-Fast (Think) records ∆ open =−34.10. The poor transfer gap may stem from two factors. First, the tasks differ in difficulty and require increased capabilities. Second, models may answer correctly by exploiting superficial cues rather than accurately inter- preting the underlying physical state. Consequently, strong performance on one task does not guarantee comparable per- formance on another involving the same scene. MLLMs show limited robustness to illusion descrip- tions. Comparisons between the paired vanilla and adver- sarial settings show that adversarial descriptions generally reduce model accuracy. These results suggest that adversar- ial descriptions may bias models toward misleading cues rather than the visual evidence. Further, we notice that open- source models are particularly more susceptible than the closed ones. Such observed disparity may stem from the generally stronger vision processing capabilities of closed- source models, which make them more robust to illusion descriptions and thus achieve more stable performances. 4.4 Performance Across Taxonomy We further examine average performances across scenarios, targets and mechanisms, and present results in Figure 4. (1) MLLM performances are consistently limited across scenarios. Across all scenarios, the highest accuracy remains below 59%, indicating that the tasks are generally challenging. Meanwhile, the maximum variation across sce- narios for the same task is only 7.36 points. This suggests that the dominant source of difficulty may not be the sce- nario type, but depend more on the specific illusion and task requirements. Overall, the models face a common challenge across scenarios: extracting reliable visual evidence and us- ing it to infer the underlying physical state. (2) Model performances vary across misperceived tar- gets. Context is the most challenging target, whereas property ActionOpenTF Outdoor Urban Outdoor Natural enuesV Large Indoor Enclosures Small Indoor Accuracy(%) Scenarios 0 20 40 60 80 31.19 32.41 34.72 38.55 52.93 55.71 48.35 52.67 58.46 58.69 53.39 56.68 ContextRelationsProperty Accuracy(%) argets T 0 20 40 60 80 23.97 35.19 37.93 50.16 48.78 57.77 48.8 57.01 60.86 Misalignment Perspective Bias Driven-Experience disruption Angle Confusion Dimension Ambiguity Boundary Accuracy(%) Mechanisms 0 20 40 60 80 31.34 31.92 37.12 14.81 46.46 58.39 53.6 54.2 33.46 51.49 54.69 57.14 60.78 36.59 55.9 Figure 4: Average performances across the taxonomy. is the easiest, outperforming context by at most 13.96% on action planning. This difference may arise because property judgments typically rely on localized visual cues, while con- text requires models to integrate information across the entire scenario. Relation yields intermediate performance because it requires connecting multiple local entities but does not demand full-scene integration. Overall, these results suggest that illusion targets become more difficult to interpret as the required scope of scene information broadens. (3) Illusion mechanisms that disrupt global scene struc- ture pose the greatest challenge. Among the five mecha- nisms, angle disruption is the most difficult, whereas bound- ary ambiguity yields the best performance. This difference may arise because boundary ambiguity affects local bound- aries, whereas angle disruption alters the global scene orien- tation and requires more complicated spatial reference recon- struction. We also notice that action planning shows largest performance gap of 31.65% across mechanisms. This may be because the task requires models to recover the correct spatial scene and translate it into feasible actions, where spatial rea- soning errors propagate into downstream decisions. Overall, illusions from global spatial structure disruptions demand stronger spatial reasoning and remain the most challenging. 4.5 Failure Modes To characterize typical failure modes of MLLMs, we orga- nize them into three progressive categories corresponding to successive stages of visual evidence extraction, contextual grounding, and scene-level reasoning. Observation failure refers to a model inaccurately captur- ing task-relevant visual evidence from the image. Typically, it involves 1) evidence mismatch where incorrect evidence is selected , 2) evidence insufficiency where necessary ev- idence is overlooked and 3) evidence overload with exces- sive or irrelevant evidence interferes with judgment. Such observation-level errors provide an unreliable basis for sub- sequent judgments. Taking Figure 5 (a) as an example, the model overlooks the cat in the background and mistakes its Grounding Failure Reasoning Failure <think> I see only one cat. The appearance of multiple is due to the fluffy fur making it look like two overlapping, but it's one. </think> Question: How many cats are on the bed? <think> Let's think: in zoos, animals often have tags or collars. I believe the hippo is wearing a black harness around its body and legs. </think> Question: What's on the hippo’s legs? <think> The left man sitting on concrete, others standing on concrete. Perhaps it's the floor of the building or the balcony. </think> Question: Which surface is supporting them? Observation Failure Another cat! It’s just a shadow! Rotated 90°! Model Answer: One. Model Answer: Black bands/straps. Model Answer: The concrete ledge. Wall Floor (a) (b) (c) Figure 5: Failure mode cases. ear for the fur of the cat in the foreground. Typically, for the model of Grok-4.1-Fast (Think), such observation failures account for 45.79% of all wrong results. Grounding failure occurs when a model establishes in- correct associations between correctly identified visual evi- dence. Specifically, instead of grounding its interpretation in the specific image context, the model relies on prior knowl- edge and selectively uses visual cues to support an incorrect interpretation. In Figure 5 (b), misled by the prior assump- tion that the hippo in a zoo could be harnessed, the model incorrectly grounds the shadow to black straps. Grounding failures account for 33.68% of all wrong results. Reasoning failure means that models infer an incorrect scene structure during holistic scene parsing. Two representa- tive errors are usually observed, including 1) spatial relation error where apparent 2D relations are incorrectly interpreted as physical relations in 3D space, and 2) scene rotation error, where object orientations are misinterpreted due to an incor- rect reference frame. In Figure 5 (c), the image is rotated by 90 degrees, and the model mistakes the wall at the bot- tom of the frame for a concrete floor supporting the people. Reasoning failures account for 20.53% of all wrong results. 5 Mitigation Strategy 5.1 Prompting Mitigation Building on the three identified failure modes, we design a guided prompting method that explicitly structures the model’s vision processing around observation, grounding, and reasoning. The prompt first directs the model to carefully inspect the visual evidence, then ground relevant objects with their relationships in the scene, and finally reason about the underlying physical state before producing an answer. Fur- ther details are provided in Appendix C.6. We evaluate this method on three representative closed- source models and three open-source models. As shown in Table 2, two key observations emerge. First, the method gen- erally improves performance, revealing that appropriately designed prompts can better elicit models’ capabilities for ModelT/FOpenActionAvg. Gemini-3.5-Flash 83.19 (+2.5) 87.72 (+2.3) 73.83 (+10.3) 81.58 (+5.0) GPT-5.567.75 (+16.4) 71.24 (+20.7) 54.21 (+10.3) 64.40 (+15.8) Claude Opus 4.7 72.01 (+4.8) 75.77 (+6.2) 52.34 (+13.1) 66.71 (+8.0) Closed Avg.74.32 (+7.9) 78.24 (+9.7) 60.13 (+11.2) 70.90 (+9.6) Qwen3.5-9B53.60 (+1.7) 59.13 (+7.6) 30.84 (-1.9) 47.86 (+2.5) InternVL3.5-2B 56.69 (+14.2) 30.37 (-1.0) 25.23 (+0.0) 37.43 (+4.4) Gemma-3-12B63.27 (+6.8) 40.27 (-3.3) 25.23 (+0.9) 42.92 (+1.5) Open Avg.57.85 (+7.5) 43.25 (+1.1) 27.10 (-0.3) 42.74 (+2.8) Table 2: Prompting mitigation results. handling situational illusions. However, the magnitude of im- provement varies across tasks, reflecting differences in task difficulty and capability requirements. Second, closed-source models benefit more than open-source models, whereas some open-source models exhibit only marginal gains or slight degradation. This suggests that prompting effectiveness may depend strongly on a model’s underlying perceptual and rea- soning capabilities. Therefore, prompting alone may be in- sufficient and further mitigation strategies are expected. 5.2 Supervised Fine-tuning Mitigation To compensate for the prompting limitations, we employ a su- pervised fine-tuning (SFT) strategy for improvement. Rather than training directly on illusion tasks, we design three kinds of tasks to enhance the capabilities of the model in object localization, scene reconstruction, and action planning. The SFT dataset comprises 1,630 task pairs that are not explicitly designed around situational illusions but instead target generalizable capabilities, including spatial-relation understanding, scene description, and action-sequence rea- soning. It is constructed using 600 images from MSIBench and 400 real-world images from the COCO 2017 training split (Lin et al. 2014), paired with Localized Narratives an- notations (Pont-Tuset et al. 2020) to preserve the models’ general utility. Training details are in Appendix B.3. For testing, the remaining 304 MSIBench images with cor- responding tasks are used, including 107 images for action evaluation. We adopt the three open-source models in the prompting method and additionally include InternVL3.5-8B and InternVL3.5-14B to evaluate the general SFT effective- ness across model scales. With LLM-as-a-Judge evaluation and human verification, we acquire final results in Table 3. We could see that all models achieve an average improvement of over 10 points, with a maximum gain exceeding 20% on Qwen3.5-9B. These results demonstrate that SFT substan- tially strengthens the internal capabilities of the models in scene observation, comprehension, and execution, thus lead- ing to better performances in situational illusions. 6 Related Works 6.1 MLLMs in Real-world Situations MLLMs are increasingly evaluated on their ability to per- ceive, interpret, and reason about real-world visual situa- tions. One line of work evaluates the vision processing ca- pabilities of MLLMs in real-world environments, including fine-grained visual perception in complex images (Zhang et al. 2025a), reasoning in everyday situations (Li et al. ModelT/FOpenActionAvg. InternVL3.5-2B 46.96 (+12.3) 40.95 (+12.2) 48.60 (+23.4) 45.50 (+16.0) InternVL3.5-8B 57.89 (+13.7) 52.47 (+14.6) 46.73 (+16.8) 52.36 (+15.0) InternVL3.5-14B 56.66 (+11.3) 50.83 (+8.2) 29.91 (+10.3) 45.80 (+9.9) Gemma-3-12B59.38 (+11.7) 52.30 (+10.7) 46.73 (+22.4) 52.80 (+14.9) Qwen3.5-9B72.20 (+22.9) 70.56 (+19.6) 51.40 (+18.7) 64.72 (+20.4) Avg.58.62 (+14.4) 53.42 (+13.1) 44.67 (+18.3) 52.24 (+15.2) Table 3: SFT mitigation results. 2026), and decision-making through cross-image evidence integration (Meng et al. 2025). Another line of work fo- cuses on whether MLLMs can make reliable decisions in such situations. Zhou et al. (2025) assess whether models can recognize risks arising from a given visual context and respond appropriately. Further studies broaden this scope to diverse daily life scenarios (Lou et al. 2026) and embod- ied environments, where agents must consider risks during task planning (Yin et al. 2024; Huang et al. 2025), respond safely to hazardous instructions (Liu et al. 2025), and identify hazards through active exploration (Gao et al. 2025). Despite their progress, both lines of work generally assume that visual observations accurately reflect the underlying scene. Corre- spondingly, our work investigates how MLLMs perceive and reason about real-world illusions, where visual appearance diverges from the underlying physical state. This capability is essential for reliable scene understanding and safe interaction with real-world environments. 6.2 Visual Illusions Visual illusions arise when visual appearance diverges from reality. Early studies (Gregory 1968, 1997; Wertheimer 1923) investigated classic illusions, like gestalt illusions, by analyzing psychological and physiological cognitive mech- anisms. With the emergence of MLLMs, researchers have explored whether these models exhibit similar vulnerabil- ities (Zhang et al. 2023). These works mainly focus on specific perceptual phenomena, such as pareidolia (Ros- tamkhani et al. 2025; Hamilton et al. 2024), classic optical illusions (Guan et al. 2024), and other visual cues that in- duce incorrect perceptions (Han et al. 2024), relying largely on synthetic or deliberately constructed images rather than naturally occurring scenes. Recent studies attempt to incor- porate everyday scenarios into consideration, where Illusion- VQA (Shahgir et al. 2024) and VIA-Bench (Hou et al. 2026) include some real-world illusion samples but still mainly fo- cus on artificial illusions. MVI-Bench (Chen et al. 2025a) focuses on real-world misleading content by organizing mis- leading cues hierarchically and evaluating the robustness of MLLMs. To fill the gap in systematic exploration of situa- tional illusions, we establish a structured taxonomy and an evaluation benchmark, diagnose model limitations, and in- vestigate approaches to improve model performances. 7 Conclusion This paper systematically studies situational illusions for MLLMs, covering a taxonomy, evaluation and mitigation. We develop a comprehensive taxonomy to characterize situ- ational illusion and construct MSIBench to assess 27 model configurations, revealing substantial vulnerabilities and iden- tifying key failure modes. Further, two mitigation strategies are developed, significantly improving models’ ability to han- dle situational illusions. Future work will explore how such robust multimodal reasoning can facilitate embodied agents. References Anthropic. 2026. Introducing Claude Opus 4.7. Accessed: 2026-07-10. Bai, S.; Cai, Y.; Chen, R.; et al. 2025. Qwen3-VL Technical Report. arXiv:2511.21631. Barnorama. 2026. Barnorama. Accessed: 2026-07-10. Bharti, S.; Cheng, S.; Rho, J.; Zhang, J.; Cai, M.; Lee, Y. J.; Rau, M.; and Zhu, X. 2025. CHARTOM: A Vi- sual Theory-of-Mind Benchmark for LLMs on Misleading Charts. arXiv:2408.14419. Bright Side. 2026. Bright Side. Accessed: 2026-07-10. Chen, H.; Peng, J.; Min, D.; Sun, C.; Chen, K.; Yan, Y.; Yang, X.; and Cheng, L. 2025a. MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs. CoRR, abs/2511.14159. Chen, Z.; Song, S.; Shum, K.; Lin, Y.; Sheng, R.; Wang, W.; and Qu, H. 2025b. Unmasking Deceptive Visuals: Bench- marking Multimodal Large Language Models on Mislead- ing Chart Question Answering. In Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; and Peng, V., eds., Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 13756–13789. Suzhou, China: Asso- ciation for Computational Linguistics. ISBN 979-8-89176- 332-6. Gao, S.; Yao, J.; Wen, H.; Guo, Y.; Liu, Z.; and Huang, H. 2025. HomeSafeBench: A Benchmark for Embodied Vision- Language Models in Free-Exploration Home Safety Inspec- tion. CoRR, abs/2509.23690. Gemma Team. 2025. Gemma 3 Technical Report. arXiv:2503.19786. GLM-V Team. 2026. GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Re- inforcement Learning. arXiv:2507.01006. Gregory, R. L. 1968. Perceptual illusions and brain mod- els. Proceedings of the Royal Society of London. Series B. Biological Sciences, 171(1024): 279–296. Gregory, R. L. 1997. Visual illusions classified. Trends in Cognitive Sciences, 1(5): 190–194. Guan, T.; Liu, F.; Wu, X.; Xian, R.; Li, Z.; Liu, X.; Wang, X.; Chen, L.; Huang, F.; Yacoob, Y.; Manocha, D.; and Zhou, T. 2024. HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, WA, USA, June 16-22, 2024, 14375–14385. IEEE. Hamilton, M.; Stent, S.; DuTell, V.; Harrington, A.; Corbett, J.; Rosenholtz, R.; and Freeman, W. T. 2024. Seeing Faces in Things: A Model and Dataset for Pareidolia. In Leonardis, A.; Ricci, E.; Roth, S.; Russakovsky, O.; Sattler, T.; and Varol, G., eds., Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part LXV, volume 15123 of Lecture Notes in Computer Science, 377–395. Springer. Han, T.; Lian, Q.; Pan, R.; Pi, R.; Zhang, J.; Diao, S.; Lin, Y.; and Zhang, T. 2024. The Instinctive Bias: Spurious Im- ages lead to Illusion in MLLMs. In Al-Onaizan, Y.; Bansal, M.; and Chen, Y., eds., Proceedings of the 2024 Confer- ence on Empirical Methods in Natural Language Process- ing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, 16163–16177. Association for Computational Linguistics. Hou, W.; Liu, W.; Hu, H.; Sun, X.; Yeung-Levy, S.; and Fan, H. 2026. Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies. CoRR, abs/2602.01816. Huang, Y.; Ding, L.; Tang, Z.; Wang, T.; Lin, X.; Zhang, W.; Ma, M.; and Zhang, Y. 2025. A Framework for Bench- marking and Aligning Task-Planning Safety in LLM-Based Embodied Agents. CoRR, abs/2504.14650. Kavukcuoglu, K.; Dean, J.; Vinyals, O.; and Shazeer, N. 2026. Gemini 3.5: Frontier Intelligence with Action. Ac- cessed: 2026-07-28. Kimi Team. 2025.Kimi-VL Technical Report. arXiv:2504.07491. Kimi Team. 2026. Kimi K2.5: Visual Agentic Intelligence. arXiv:2602.02276. Li, J.; Huang, S.; Jin, Z.; Zhang, C.; Cao, P.; Chen, Y.; Liu, K.; and Zhao, J. 2026. MMR-Life: Piecing Together Real- life Scenes for Multimodal Multi-image Reasoning. CoRR, abs/2603.02024. Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014. Microsoft coco: Common objects in context. In European conference on computer vision, 740–755. Springer. Liu, A.; Ying, Z.; Wang, L.; Mu, J.; Guo, J.; Wang, J.; Ma, Y.; Liang, S.; Zhang, M.; Liu, X.; and Tao, D. 2025. AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions. CoRR, abs/2506.14697. Lou, X.; Xu, J.; Yin, J.; Wang, X.; Kang, Z.; Liaoyouwei; Wang, Y.; Shi, X.; Mo, F.; Yao, S. U.; and Huang, K. 2026. When Helpers Become Hazards: A Benchmark for Analyz- ing Multimodal LLM-Powered Safety in Daily Life. In Li- akata, M.; Moreira, V. P.; Zhang, J.; and Jurgens, D., eds., Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, 28937–28963. Association for Computational Linguis- tics. Meng, F.; Wang, J.; Li, C.; Lu, Q.; Tian, H.; Yang, T.; Liao, J.; Zhu, X.; Dai, J.; Qiao, Y.; Luo, P.; Zhang, K.; and Shao, W. 2025. MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models. In The Thir- teenth International Conference on Learning Representa- tions, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net. Oliva, A.; and Torralba, A. 2001. Modeling the Shape of the Scene: A Holistic Representation of the Spatial Envelope. International Journal of Computer Vision, 42(3): 145–175. OpenAI. 2026. GPT-5.5 System Card. Technical report, OpenAI. Accessed: 2026-07-10. OpenAI. 2026. Introducing GPT-5.4. https://openai.com/ index/introducing-gpt-5-4/. Accessed: 2026-07-10. OpenAI. 2026. Introducing GPT-5.4 mini and nano. Ac- cessed: 2026-07-24. Pont-Tuset, J.; Uijlings, J.; Changpinyo, S.; Soricut, R.; and Ferrari, V. 2020. Connecting Vision and Language with Localized Narratives. In ECCV. Qwen Team. 2026. Qwen3.5: Accelerating Productivity with Native Multimodal Agents. Reddit. 2026. Reddit. Accessed: 2026-07-10. Rostamkhani, M.; Ansari, B.; Sabzevari, H.; Rahmani, F.; and Eetemadi, S. 2025. Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions. In IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPR Workshops 2025, Nashville, TN, USA, June 11-15, 2025, 2995–3004. Computer Vision Foundation / IEEE. Shahgir, H. S.; Sayeed, K. S.; Bhattacharjee, A.; Ahmad, W. U.; Dong, Y.; and Shahriyar, R. 2024. IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models. CoRR, abs/2403.15952. Tonglet, J.; Zimny, J.; Tuytelaars, T.; and Gurevych, I. 2026. Is this chart lying to me? Automating the detection of mis- leading visualizations. In Liakata, M.; Moreira, V. P.; Zhang, J.; and Jurgens, D., eds., Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, 8823–8844. Association for Computational Linguistics. Treisman, A. M.; and Gelade, G. 1980. A Feature-Integration Theory of Attention. Cognitive Psychology, 12(1): 97–136. Wagemans, J.; Elder, J. H.; Kubovy, M.; Palmer, S. E.; Peter- son, M. A.; Singh, M.; and von der Heydt, R. 2012. A Cen- tury of Gestalt Psychology in Visual Perception I: Perceptual Grouping and Figure–Ground Organization. Psychological Bulletin, 138(6): 1172–1217. Wang, W.; Gao, Z.; Gu, L.; et al. 2025. InternVL3.5: Advanc- ing Open-Source Multimodal Models in Versatility, Reason- ing, and Efficiency. arXiv:2508.18265. Wertheimer, M. 1923. Untersuchungen zur Lehre von der Gestalt. I. Psychologische Forschung, 4(1): 301–350. xAI. 2025. Grok 4.1 Fast and Agent Tools API. Accessed: 2026-07-24. Xiaomi MiMo. 2026. MiMo-V2.5. Hugging Face model collection; accessed: 2026-07-28. Yang, H. J.; Lee, H.; Shim, K.; Kwak, J.; Kim, H.; Kim, D.; Ngo, K. A.; Ryu, S.; Choi, J.; Kim, Y.; Moon, C.; Ryoo, M. S.; and Shim, B. 2026. Advancing Multi-Robot Networks via MLLM-Driven Sensing, Communication, and Computation: A Comprehensive Survey. IEEE Commun. Surv. Tutorials, 28: 5833–5871. Yang, R.; Chen, H.; Zhang, J.; Zhao, M.; Qian, C.; Wang, K.; Wang, Q.; Koripella, T. V.; Movahedi, M.; Li, M.; Ji, H.; Zhang, H.; and Zhang, T. 2025. EmbodiedBench: Compre- hensive Benchmarking Multi-modal Large Language Mod- els for Vision-Driven Embodied Agents. In Singh, A.; Fazel, M.; Hsu, D.; Lacoste-Julien, S.; Berkenkamp, F.; Maharaj, T.; Wagstaff, K.; and Zhu, J., eds., Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, volume 267 of Proceedings of Machine Learning Research. PMLR / OpenReview.net. Yin, S.; Pang, X.; Ding, Y.; Chen, M.; Bi, Y.; Xiong, Y.; Huang, W.; Xiang, Z.; Shao, J.; and Chen, S. 2024. SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents. CoRR, abs/2412.13178. Zhang, Y.; Pan, J.; Zhou, Y.; Pan, R.; and Chai, J. 2023. Grounding Visual Illusions in Language: Do Vision- Language Models Perceive Illusions Like Humans? In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Lan- guage Processing, 5718–5728. Singapore: Association for Computational Linguistics. Zhang, Y.; Zhang, H.; Tian, H.; Fu, C.; Zhang, S.; Wu, J.; Li, F.; Wang, K.; Wen, Q.; Zhang, Z.; Wang, L.; and Jin, R. 2025a. MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? In Yue, Y.; Garg, A.; Peng, N.; Sha, F.; and Yu, R., eds., International Conference on Learning Representations, volume 2025, 89655–89701. Zhang, Y.; Zhang, Z.; Wei, X.; Liu, X.; Zhai, G.; and Min, X. 2025b. IllusionBench: A Large-scale and Comprehen- sive Benchmark for Visual Illusion Understanding in Vision- Language Models. In IEEE International Conference on Multimedia and Expo, ICME 2025, Nantes, France, June 30 - July 4, 2025, 1–6. IEEE. Zhihu. 2026. Zhihu. Accessed: 2026-07-10. Zhou, K.; Liu, C.; Zhao, X.; Compalas, A.; Song, D.; and Wang, X. E. 2025. Multimodal Situational Safety. In The Thirteenth International Conference on Learning Represen- tations, ICLR 2025, Singapore, April 24-28, 2025. OpenRe- view.net.