Paper deep dive
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
Zhaoyang Wei, Bowen Jiang, Xumeng Han, Jiashu Li, Xuehui Yu, Yuling Liu, Guorong Li, Zhenjun Han, Jianbin Jiao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/16/2026, 3:37:10 AM
Summary
The paper introduces AD2-Bench, a benchmark for evaluating Multimodal Large Language Models (MLLMs) in complex urban scenes under adverse conditions, and EGVOR, a model that improves reasoning by generating explicit Evidence Atoms. AD2-Bench uses a Hierarchical Visual Diagnosis framework with a Chain of Evidence (CoE) to diagnose failures caused by Spatial Ambiguity and Semantic Uncertainty. EGVOR employs a hierarchical curriculum training approach involving reflective supervision and reinforcement learning to enhance robustness and trustworthiness.
Entities (14)
Relation Signals (10)
EGVOR → generates → Evidence Atoms
confidence 95% · EGVOR reformulates reasoning as the explicit generation of Evidence Atoms—structured triplets that enforce strict spatial-semantic alignment.
AD2-Bench → uses → Chain of Evidence
confidence 95% · AD2-Bench, which establishes a rigorous Hierarchical Visual Diagnosis to decompose reasoning into a structured Chain of Evidence (CoE).
AD2-Bench → diagnoses → Spatial Ambiguity
confidence 92% · Formalizing this probabilistically, we trace the high variance in latent reasoning context to two fundamental bottlenecks: (1) Spatial Ambiguity
AD2-Bench → diagnoses → Semantic Uncertainty
confidence 92% · Formalizing this probabilistically, we trace the high variance in latent reasoning context to two fundamental bottlenecks: (2) Semantic Uncertainty
EGVOR → mitigates → Semantic Uncertainty
confidence 90% · a Grounded Description verifies semantic content to mitigate semantic uncertainty.
EGVOR → mitigates → Spatial Ambiguity
confidence 90% · Each atom follows a Locate–Attend–Describe cycle: a Spatial Anchor restricts the visual support to reduce spatial ambiguity
EGVOR → trainedwith → Group Relative Policy Optimization
confidence 90% · Cognitive Alignment stage utilizing Group Relative Policy Optimization optimizes the evidence-chain trajectory.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, leading to a disconnect between perception and reasoning. Meanwhile, existing outcome-oriented benchmarks evaluate only final predictions and fail to diagnose failures in the underlying reasoning process. To address this gap, the authors propose AD2-Bench, which introduces a Hierarchical Visual Diagnosis framework that decomposes reasoning into a structured Chain of Evidence (CoE). This fine-grained diagnosis reveals that robust multimodal reasoning fundamentally depends on accurate evidence acquisition. Building on this perspective, the authors formulate reasoning from a probabilistic viewpoint and identify two primary causes of reasoning failure: Spatial Ambiguity, where models fail to distinguish target objects from background clutter, resulting in localization errors; and Semantic Uncertainty, where degraded visual features lead to incorrect semantic interpretation, resulting in understanding errors. To overcome these evidence deficiencies, they further propose Evidence-grounded Visual Reasoning (EGVOR), which replaces implicit reasoning with the explicit generation of Evidence Atoms - structured spatial-semantic triplets that enforce tight alignment between localization and semantic understanding. The model is trained through a hierarchical curriculum that progresses from reflective supervision construction to reinforcement learning, where reducing reasoning variance is explicitly rewarded. Extensive experiments demonstrate that EGVOR substantially improves reasoning stability under adverse conditions, providing a more robust framework for trustworthy multimodal cognition.
Tags
Links
- Source: https://arxiv.org/abs/2608.10954v1
- Canonical: https://arxiv.org/abs/2608.10954v1
Trouble viewing inline? Open PDF directly →
Full Text
179,111 characters extracted from source content.
Expand or collapse full text
While Multimodal Large Language Models (MLLMs) demonstrate impressive proficiency in benign scenarios, their cognitive reliability significantly deteriorates in complex scenes characterized by adverse conditions. In such scenarios, models relying on implicit inference devoid of visual evidence succumb to disconnects between perception and logic; simultaneously, existing outcome-oriented benchmarks fail to diagnose the underlying reasoning process. To bridge this gap, we present AD2-Bench, which establishes a rigorous Hierarchical Visual Diagnosis to decompose reasoning into a structured Chain of Evidence (CoE). This granular perspective reveals that robust cognition is predicated on precise evidence acquisition. From this perspective, we formulate reasoning probabilistically, attributing the collapse to two bottlenecks: (1) Spatial Ambiguity, failing to distinguish targets from background clutter (i.e., localization error); and (2) Semantic Uncertainty, misinterpreting semantics due to feature degradation (i.e., understanding error). To resolve these evidence gaps, we propose Evidence-Grounded Visual Reasoning (EGVOR). Departing from implicit inference, EGVOR reformulates reasoning as the explicit generation of Evidence Atoms—structured triplets that enforce strict spatial-semantic alignment. We orchestrate training via a hierarchical curriculum, progressing from reflective supervision establishment to reinforcement learning that explicitly rewards the reduction of reasoning variance. Extensive experiments demonstrate that EGVOR substantially mitigates instability, establishing a robust framework for trustworthy multimodal cognition. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes Zhaoyang Wei Email: weizhaoyang23, jiangbowen24, hanxumeng19, lijiashu24, yuxuehui17@mails.ucas.ac.cn Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Bowen Jiang Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Xumeng Han Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Jiashu Li Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Xuehui Yu Affiliation: Tencent CDG, Shenzhen, 518054, China Yuling Liu Email: liuyuling@iie.ac.cn Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Affiliation: Institute of Information Engineering, CAS, Beijing, 100085, China Guorong Li Email: liguorong, jiaojb@ucas.ac.cn Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Zhenjun Han Email: hanzhj@ucas.ac.cn Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China Jianbin Jiao Affiliation: University of Chinese Academy of Sciences (UCAS), Beijing, 101480, China keywordsMultimodal Large Language Models, Visual Grounding, Adverse Conditions, Chain of Evidence 1 Introduction Dataset Multimodal Adverse Condition Real World QA Pairs Multi-prompt Hierarchical assessment Decision Standard Vision Datasets COCO 2017 ✗ ✗ ✗ ✗ ✗ ✗ ✗ CODA ✗ ✗ ✓ ✗ ✗ ✗ ✗ StreetHazards ✗ ✗ ✓ ✗ ✗ ✗ ✗ ACDC ✗ ✓ ✓ ✗ ✗ ✗ ✗ General MLLM Benchmarks RealWorldQA ✓ ✗ ✓ 0.8k ✗ ✗ ✗ CV-Bench ✓ ✗ ✗ 2.6k ✗ ✗ ✗ MME ✓ ✗ ✗ 2.4k ✗ ✗ ✗ MMBench ✓ ✗ ✗ 3.2k ✗ ✗ ✗ MME-RealWorld ✓ ✗ ✓ 29k ✗ ✗ ✗ AD2-Bench (ours) ✓ ✓ ✓ 70k ✓ ✓ ✓ Table 1: Comparison with existing datasets. Adverse Condition and Real World indicate whether the dataset specifically targets in real complex scenes. Figure 1: Standard outcome-oriented metric (top) masks hallucinations in complex scenes. In contrast, our Evidence-Centric diagnosis (bottom) anchors reasoning to verifiable visual cues, exposing the perception-logic disconnect. Multimodal large language models (MLLMs) have achieved strong performance on a wide range of vision–language tasks, including captioning, visual question answering, and general cross-modal reasoning 3; 15; 78. By coupling vision encoders with large language backbones, these models process images and text in a unified semantic space and exhibit robust zero-shot generalization in general scenarios. However, moving from these benign environments to visually adverse and structurally complex real-world scenes—such as cluttered streets under rain, fog, night illumination, or heavy occlusion—remains challenging. In these settings, MLLMs frequently exhibit visual hallucination (describing non-existent objects or attributes), unstable reasoning chains, and decisions that are difficult to interpret or trust. Evaluating such systems solely by final-answer accuracy is therefore inadequate: a model may produce a correct answer for spurious reasons, or fail in ways that conventional metrics cannot reveal. Diagnosing and resolving these failures is impeded by a systemic dual opacity, as illustrated in Fig. 1. On the modeling side, standard MLLMs rely on implicit, end-to-end reasoning that lacks explicit visual evidence, leaving the reasoning process vulnerable to environmental noise. Simultaneously, prevailing benchmarks—ranging from general-purpose suites like MMBench 47 and MME 24 to robustness-focused sets like MME-RealWorld 97—reinforce this challenge through outcome-oriented evaluation (Fig. 1, top right). This paradigm assesses models solely on final-answer correctness, effectively treating the intermediate cognition as a black box. Crucially, Visual Evidence serves as the missing link for both robust reasoning and transparent diagnosis. Without explicit evidence reasoning chain, it is intractable to disentangle whether an error stems from a failure in Perceptual Grounding (e.g., missing key targets due to fog) or a failure in Cognitive Reasoning (e.g., misinterpreting the vehicle’s intent). Bridging this gap thus demands a paradigm shift: moving from implicit to chain of Evidence-Centric reasoning construction, and from outcome-based to process-oriented diagnosis. To dismantle the black-box nature of current evaluations, we introduce AD2-Bench, a comprehensive benchmark designed to diagnose the hierarchical cognitive capabilities of MLLMs in visually adverse and complex scenes. Comprising approximately 10K images and 70K QA pairs, the benchmark captures diverse environmental degradations and intricate layouts (see Fig. 2). Crucially, distinct from traditional outcome-oriented formats, we establish an Atomic Annotation Protocol with multi-granularity vision prompts grounded in human expert cognition. We decompose complex queries into a rigorous Chain of Evidence (CoE), which explicitly maps the cognitive flow from Environmental and Object Perception to fine-grained Interaction Analysis—resolving spatial relations and vulnerable road users—before synthesizing Holistic Decisions. This structured decomposition enables a Trustworthiness-oriented Diagnosis beyond simple accuracy. By aligning the model’s intermediate reasoning with human-annotated evidence chains, we propose a suite of orthogonal metrics, assessing Factual Fidelity, Logical Reliability, Self-Consistency, and Explainability to quantitatively measure the reliability of the model cognitive process. Figure 2: Adverse conditions include adverse weather (top) and complex scenarios (down). Serving as a diagnostic prism, AD2-Bench reveals a critical dependency: robust reasoning in complex scenes is predicated on precise visual grounding. Formalizing this probabilistically, we trace the high variance in latent reasoning context to two fundamental bottlenecks: (1) Spatial Ambiguity: In complex scenes with dense occlusion and clutter, the model’s attention mechanism suffers from dispersion, drifting towards high-contrast distractors rather than the relevant targets. (2) Semantic Uncertainty: Even when the target is localized, environmental degradation (e.g., rain streaks, low light) corrupts the feature representation. This challenge forces the model to abandon visual evidence and over-rely on language priors, leading to hallucinations. These observations motivate an explicit evidence construction mechanism that reduces both spatial ambiguity and semantic uncertainty. To address these challenges, we propose Trustworthy Evidence-Grounded Visual Reasoning (EGVOR), which turns multimodal reasoning from implicit text generation into explicit evidence construction within an end-to-end model. We posit that robust reasoning requires Component Filtration: actively filtering out environmental noise by locking onto critical evidence. EGVOR operationalizes this idea by generating a sequence of Evidence Atoms. Each atom follows a Locate–Attend–Describe cycle: a Spatial Anchor restricts the visual support to reduce spatial ambiguity, while a Grounded Description verifies semantic content to mitigate semantic uncertainty. Figure 3: (a) Task Categories. Our benchmark spans 4 key demensions and 33 subtasks highly related to real-world scenarios, including 10k high-resolution images and 70K annotations. (b) Model Performance. Average accuracies of advanced MLLMs are shown on the dataset. We train EGVOR via a hierarchical curriculum designed to align the model with this cognitive paradigm. A Reflective Supervised Fine-Tuning stage first establishes the syntactic structure of evidence chains and endows the model with self-correction capabilities. Subsequently, a Cognitive Alignment stage utilizing Group Relative Policy Optimization optimizes the evidence-chain trajectory. By employing dense spatial rewards and semantic alignment rewards, we incentivize the policy to construct valid, grounded, and logically sound evidence chains, effectively closing the loop between perception and reasoning. In summary, this paper makes the following contributions: • We present AD2-Bench, a large-scale benchmark for evaluating MLLMs in visually complex scenes. It features a unique Hierarchical Visual Diagnosis that decomposes reasoning into Chains of Evidence, enabling a fine-grained, trustworthiness-oriented evaluation framework. • We theoretically diagnose the reasoning failures in adverse scenarios as Spatial Ambiguity and Semantic Uncertainty, and propose EGVOR to address them. EGVOR reformulates multimodal reasoning as the sequential construction of Evidence Atoms, utilizing Component Filtration to reduce latent variance. • Through extensive experiments on AD2-Bench and external benchmarks, we demonstrate that EGVOR significantly outperforms state-of-the-art MLLMs in both grounded reasoning accuracy and trustworthiness metrics, offering a robust solution for interpretable decision-making in complex real-world environments. 2 Related Work Multimodal Large Language Models. Multimodal large language models (MLLMs) couple vision encoders with large language backbones to enable unified visual–textual reasoning 31; 69; 90; 21; 88; 20; 96. Early systems such as Flamingo 1 and BLIP-2 39 inject visual features into frozen LLMs via cross-attention or lightweight querying transformers, allowing image-grounded QA and instruction following. Later models show that simply projecting CLIP-style image embeddings 64 into the LLM token space, followed by instruction tuning, yields strong open-ended multimodal capabilities. This design underlies the LLaVA family 45; 43; 44, the Qwen2-VL and Qwen2.5-VL series 77; 4; 78, InternVL models 15; 101, and related architectures 92; 62. A recent focus is scaling MLLMs to high-resolution and large-field-of-view images, which is crucial for cluttered, detail-rich scenes. LLaVA-NeXT 44 and AnyRes-style approaches 10 preserve fine structures via tiling or adaptive rescaling, while Qwen2-VL and Qwen2.5-VL introduce multimodal rotary position encodings to support arbitrary resolutions and aspect ratios 78; 5. Together with improvements in data quality and scale 101; 76; 85; 63, these models provide strong generic visual grounding and form robust backbones for downstream tasks. Beyond generic Web imagery, MLLMs have been connected to robotics and driving stacks 8; 71; 13; 26; 17, where they describe scenes, query hazards, or explain actions in natural language. However, visual grounding in these systems is largely implicit in attention patterns, and answers are typically produced in a single step, which can be brittle in complex, cluttered scenes. Our work builds on these foundations while explicitly targeting fine-grained, stepwise reasoning in such environments. Multimodal Reasoning The success of chain-of-thought prompting in text-only LLMs 58; 29 has motivated efforts to bring explicit reasoning into multimodal settings. Early multimodal chain-of-thought (MCoT) approaches decompose tasks into intermediate steps via predefined workflows, auxiliary tools, or staged pipelines 50; 56. Many encourage the model to first localize regions of interest or refine latent visual features 84; 27; 62 before generating a final answer, and some integrate external knowledge to support higher-level reasoning. More recently, reinforcement learning (RL) has been used to strengthen multimodal reasoning, inspired by systems such as OpenAI-o1/o3 and DeepSeek-R1 58; 59; 29. RL- and GRPO-based methods have been applied to multimodal math and science problems 32; 80; 81; 11, general visual reasoning 55; 60, and open-ended visual grounding 67; 49; 34. Other works emulate human-like “thinking with images” by allowing dynamic cropping, zooming, or multi-step attention over images 51; 66; 75; 23; 100; 7; 48; 70. However, most of these approaches still supervise only the final answer, leaving intermediate visual focus and reasoning trajectories implicit and difficult to audit. In this work we follow the general direction of multimodal chain-of-thought and RL-based training, but introduce an explicit, object-centric reasoning process whose details are presented in Sec. 4. Benchmarks for Multimodal Large Language Models To assess the capabilities of modern MLLMs, numerous benchmarks have been proposed beyond classical captioning 42 and VQA 2. General-purpose suites such as MME 25 and MMBench 47 probe perception, commonsense, and language understanding with diverse question formats, while SEED-Bench 37 and M-Vet 94 extend evaluation to multi-image/video inputs, OCR, and mathematical reasoning. Higher-level benchmarks including MMMU 95, HallusionBench 28, and MathVista 52 focus on expert knowledge, hallucination, or math-centric visual reasoning. These datasets provide a broad view of MLLM competence but mostly use relatively clean images and evaluate only final answers, without requiring models to expose intermediate reasoning. A complementary line of work targets more challenging visual conditions or spatially complex scenes. POPE 41 diagnoses object hallucination under controlled perturbations, and V* 83 evaluates attributes and pairwise spatial relations, albeit mainly on COCO-derived imagery 42. MME-RealWorld 97 and HR-Bench 79 introduce high-resolution, cluttered scenes but still treat grounding implicitly. On top of driving and urban-perception datasets such as BDD100K 93 and nuScenes 6, several multimodal corpora enrich real-world scenes with language: BDD-X 36 and GOHD 89 provide behavior rationales and hazard explanations; DriveRef 82 studies referring expressions; nuScenes-QA/MQA 61; 33 scale visual QA; and DriveLM 68 offers graph-structured dialogues. Other datasets emphasize commonsense 53, robustness and trust 87, or rare hazards 40; 9. These corpora underscore the importance of complex, real-world scenes, but generally provide answer- or sentence-level supervision and seldom disentangle perceptual grounding from subsequent reasoning quality. 3 Construction of AD2-Bench We introduce AD2-Bench, a comprehensive benchmark designed to evaluate the hierarchical cognitive capabilities of Multimodal Large Language Models within complex urban scenarios. Unlike existing datasets that focus primarily on benign conditions or end-to-end control commands, AD2-Bench emphasizes the process of evidence-based reasoning under visual degradation, bridging the gap between perceptual grounding and actionable decision. 3.1 Complex Scenario Curation To rigorously assess perceptual robustness, we move beyond synthetic or clear-weather datasets found in standard benchmarks. We curated a large-scale collection of real-world images capturing genuine visual degradation from diverse sources, including CODA 40, ACDC 65, DAWN 35, BDD100K 93, nuScenes 6, and diverse web repositories. This multi-source strategy ensures that the visual patterns cover a wide spectrum of sensor noise and environmental occlusions. The dataset is structured into seven distinct adverse condition categories: Rainy, Snowy, Foggy, Sandstorm, Nighttime, Dawn, and Overcast. While the Rainy, Foggy, and Snowy classes emphasize signal degradation, the Daytime and Dawn scenes focus on relational complexity, featuring dense object occlusion and intricate road layouts that demand deep semantic parsing. Rigorous Filtering Pipeline. Ensuring data quality in adverse conditions is challenging due to the ambiguity of visual features. We employed a rigorous multi-stage filtering protocol. First, we utilized advanced MLLMs (Gemini 2.5-Pro) to perform automated pre-filtering, scoring image complexity to exclude samples with low reasoning value. Subsequently, domain experts conducted a final manual audit to verify that each image contains sufficient visual cues for logical inference, rejecting samples with unrecognizable semantics or insufficient resolution. Figure 4: To diagnose interactive reasoning capabilities and facilitate precise evidence acquisition in adverse scenarios, we introduce a multi-granularity prompt: Region-level cues isolate contextual interactions, while Point-level cues pinpoint fragmented evidence in heavy occlusion where boxes fail. Text-level cues resolve semantic ambiguity. Addressing the limitation of MLLMs in processing coordinate text, we render these prompts as explicit visual overlays to enforce spatial attention. 3.2 The Atomic Annotation Protocol A core contribution of AD2-Bench is the Atomic Annotation paradigm, designed to evaluate the explicit reasoning process. Unlike traditional VQA datasets that only provide a final answer, we formalize the annotation of the entire evidence chain and multi-Granularity vision prompts, enabling step-by-step diagnostic evaluation. Vision Prompt Strategy. In complex scenes, critical visual evidence is often submerged in environmental noise or fragmentation. To bridge the gap between perception and reasoning while diagnosing interactive capabilities, we employ three granular prompt levels as illustrated in Fig. 4, serving as explicit attentional prompts. Single image provide no cues, serving as an unguided baseline. Region-level prompts utilize bounding boxes to define a contextual scope, guiding the model to extract evidence regarding multi-object interactions. Critically, we introduce Point-level prompts to capture atomic evidence in dense occlusion scenarios. Unlike bounding boxes that suffer from significant overlap, point prompts function as spatial pins, isolating the visible fragments of targets. Finally, Text-level prompts provide referential grounding to resolve semantic ambiguities. Since MLLMs struggle with raw coordinate text, we render these prompts directly onto the images to ensure spatial attention. More details are in Appendix 7.9. Formal Definition. We define each annotation record PiP_i as a six-tuple to capture the full context of a reasoning step: Pi=(Qi,Ci,Ii,Li,Bi,Ai) P_i=(Q_i,\,C_i,\,I_i,\,L_i,\,B_i,\,A_i) (1) where QiQ_i is the question, Ci∈[1,33]C_i∈[1,33] denotes the sub-task index covering the full spectrum from atomic perception to holistic decision, IiI_i is the image, LiL_i specifies the prompt type (text/box/point), BiB_i contains box coordinates, and AiA_i is the atomic ground-truth answer. This formalization ensures that every reasoning step is anchored to specific visual and textual evidence. The Atomic Flow Strategy. The annotation process is executed through a rigorous Atomic Flow collaboration strategy to ensure logical consistency across the chain of evidence. The entire process was conducted in four batches, involving domain experts overseen by independent lead annotators. Initially, the complex scene decision task is subjected to Atomic Decomposition, breaking it down into a Directed Acyclic Graph (DAG) of sub-questions. This separates Perception Atoms (identifying visual elements) from Reasoning Atoms (interpreting their implications). Cluster analysis was employed to refine the four evaluation axes into 33 distinct sub-tasks to ensure coverage. To maintain high fidelity, we implemented a Machine-Assisted Validation phase. Initial atomic annotations are rigorously scored by an ensemble of advanced models (Gemini 2.5-Pro, GPT-4o, Qwen 2.5-Max) against the visual input. Discrepancies between human labels and model consensus trigger a mandatory review flag. Finally, the Expert Cross-Pollination mechanism ensures narrative coherence. Experts are dynamically grouped where specialists in perception verify the grounding coordinates (BiB_i), while specialists in logic verify the causal link to the answer (AiA_i). This cross-check ensures that the final Intelligence Atoms are not only factually correct but also narratively coherent, providing a reliable ground truth for measuring the alignment between machine reasoning and human cognition. 3.3 Analysis of Evaluation Dimensions Figure 5: The Hierarchical Diagnosis Framework in AD2-Bench guides MLLMs through a structured Chain of Evidence (CoE) process. This pipeline decomposes reasoning into adaptable stages, typically including perception (basic & advanced), various analyses like key vehicle/pedestrian (relation understanding & reasoning), and a final suggestion (decision). The vision prompt in Step 2 can be optioned, and image1 and image2 are the same image displayed in different steps. This structured decomposition serves two key purposes: it enhances the MLLM’s ability to capture fine-grained scene details progressively, and critically, it facilitates reliable, atomic-level annotation and verification of each reasoning step on our end, ensuring high accuracy for crucial scene elements annotation. Basic Perception (Level 1: Atomic Visual Grounding). In complex scenes, trustworthy decision-making is predicated on the precise acquisition of visual evidence. Before any high-level reasoning can occur, the model must filter environmental noise to establish valid physical focus for relevant entities. Therefore, AD2-Bench isolates a basic-perception tier designed to evaluate the model’s capacity to reliably extract valid foreground objects from cluttered or degraded visual backgrounds. This tier specifically evaluates: (1) Object Presence, determining the existence of entities despite visual obstructions like rain or fog; (2) Precise Localization, anchoring vulnerable road users (pedestrians, cyclists) to exact pixel coordinates to prevent hallucination; and (3) Category-aware Retrieval, identifying specific targets via text-prompted queries. To ensure the reliability of GT, we mandate a rigorous cross-validation process by at least three domain experts. Advanced Perception (Level 2: Semantic Parsing). Beyond mere localization, robust reasoning requires verifying the semantic details of the evidence. This tier targets the deep decoding of attributes amidst visual degradation, ensuring that the captured evidence is semantically accurate rather than hallucinated. As illustrated in Step 1 of Fig. 5, tasks include: (1) Attribute Recognition, identifying states such as vehicle lights or pedestrian attire; (2) Dense Captioning, generating context-rich descriptions; (3) Optical Character Recognition (OCR), decoding traffic signs; and (4) Lane Topology Analysis, understanding road geometry. These problems are specifically crafted to stress-test visual encoders against real-world perturbations, such as strong glare in dawn scenarios or optical scattering in rainy night scenes. To guarantee annotation quality, twenty experts independently supply answers, followed by a two-stage expert verification to ensure semantic consistency. Relation Understanding (Level 3: Contextual Interaction). Effective evidence is not isolated; it is structurally connected. This tier bridges atomic perception and high-level logic by extracting contextual evidence, requiring the model to observe explicit interactions between entities to form a coherent scene graph. We define four distinct interaction types: (1) Spatial Relations, evaluating both ego-centric spatial awareness (target relative to ego) and allocentric relationships (layout between external objects); (2) Social & Functional Interactions, testing the ability to perceive group dynamics (e.g., “parent-child” pairs) and functional units (e.g., “rider-bicycle”) based on fine-grained action cues; (3) Causal Relations, inferring how dynamic scene changes (e.g., traffic light state) dictate target behaviors; and (4) Occlusion Status, assessing Visibility and Depth Reasoning to identify occlusion sources in dense traffic. Each relation type undergoes independent annotation and mutual cross-checks to ensure logical coherence. Reasoning and Decision (Level 4: Holistic Synthesis). This tier represents the apex of the cognitive pyramid, requiring the integration of perceptual evidence and relational context to formulate Actionable Decisions. Leveraging the dataset’s adverse condition scenarios, we curated a dedicated 5,406-image Chain of Evidence test suite. The reasoning challenges are comprehensively designed to cover: (1) Weather Reasoning, inferring physical constraints like friction and visibility from visual cues; (2) Night Reasoning, assessing risks from non-uniform lighting and glare; (3) Region-level Reasoning, applying spatial logic within specific zones (e.g., multi-object & accident sites); (4) Occlusion Reasoning, using point prompts to infer the trajectory and risk of partially visible objects; (5) Cross-role Reasoning, predicting the intent of other agents from their perspective; and (6) Holistic Reasoning, synthesizing all evidential elements to provide a safe, interpretable action decision for the ego-object, completing the perception-to-action loop. 3.4 Evaluation Framework: A Decoupled Cognitive Diagnosis To align with the hierarchical cognitive framework, we establish a comprehensive evaluation that spans from atomic perception to holistic reasoning. Unlike traditional benchmarks that rely solely on final-answer accuracy, our framework employs a multi-dimensional metric system designed to diagnose specific failures in the cognitive pipeline. Metrics for General Task. For the foundational tiers of Basic Perception and Advanced Perception, we utilize standard quantitative metrics to assess discriminative and spatial precision. For binary classification tasks such as object existence and multiple-choice selection, we employ Binary Accuracy, defined as: acci=(Ai=yi) acc_i=I(A_i=y_i) (2) where AiA_i and yiy_i represent the predicted and ground-truth answers, respectively. To evaluate spatial grounding capabilities, we utilize distinct metrics for points and regions. For pixel-level object localization, the accuracy is penalized by the Euclidean distance between the predicted coordinate ix_i and the ground truth igtx_i^gt, formulated as: acci=11+αp‖i−igt‖2 acc_i= 11+ _p\|x_i-x_i^gt\|_2 (3) where αp=0.005 _p=0.005 is a scaling factor. For object detection, we employ the standard Intersection over Union (IoU) metric to measure the overlap between predicted (BiB_i) and ground-truth (BigtB_i^gt) bounding boxes: acci=IoU(Bi,Bigt)=|Bi∩Bigt||Bi∪Bigt| acc_i=IoU(B_i,B_i^gt)= |B_i∩ B_i^gt||B_i∪ B_i^gt| (4) Furthermore, for the semantic decoding tasks involving Optical Character Recognition (OCR), we assess transcriptional quality using two complementary metrics. The Character Error Rate (CER) quantifies the Levenshtein distance normalized by string length: CER=S+D+IN CER= S+D+IN (5) where S,D,IS,D,I represent the number of substitutions, deletions, and insertions, and N is the total character count. Concurrently, the Character-Level F1-score balances precision (P) and recall (R) at the character level, computed as: F1=2⋅P⋅RP+R F_1=2· P· RP+R (6) Metrics for Relation Understanding. We evaluate contextual interaction via Accuracy on multiple-choice queries (see Tab. 14). By requiring the model to select correct descriptors amidst adversarial distractors, this metric rigorously verifies the establishment of valid scene graph edges (e.g., spatial or relation) as a prerequisite for higher-order reasoning. Metrics for Trustworthy Reasoning. To assess the Trustworthiness of the Chain of Evidence (CoE), we move beyond simple accuracy to evaluate the reliability, interpretability, and calibration of the reasoning process. We propose four orthogonal metrics that operate on the generated evidence chain. Formally, let Q be the query, A=Aii=1NA=\A_i\_i=1^N be the model’s answer structured as a sequential evidence chain, and GT=GTii=1NGT=\GT_i\_i=1^N be the ground truth. Each metric employs a scoring function Eval⋅Eval_· mapping inputs to a scalar in [0,1][0,1] (11: perfect; 00: failure). In practice, we instantiate these Eval⋅Eval_· functions using a strong automatic evaluator (an advanced LLM judge, GPT-5.2). The evaluator is asked to rate each item on a discrete 1–10 Likert scale following a task-specific rubric, and we linearly normalize this score to [0,1][0,1] before aggregation. 1. Hallucination Resistance Score (HRS): Visual Fidelity. This metric measures the model’s resistance to visual hallucinations. It verifies the factual correctness of each evidence AiA_i generated by model against the ground truth GTiGT_i, focusing on whether AiA_i only mentions entities, attributes, and relations that are actually supported by the image and the annotated CoE. The fidelity score is high when AiA_i is fully grounded and contains no hallucinated content, and low when AiA_i fabricates non-existent objects or attributes or contradicts GTiGT_i. The overall HRS is defined as the average fidelity across the chain: HRS=1N∑i=1NEvalfidelity(Ai,GTi). HRS= 1N _i=1^NEval_fidelity(A_i,GT_i). (7) 2. Logical Reliability Score (LRS): Inferential Validity. To assess the robustness of the reasoning process, LRS evaluates the validity of each local transition Ai→Ai+1A_i→ A_i+1. Crucially, this metric is truth-agnostic: Evallogic(Ai,Ai+1)∈[0,1]Eval_logic(A_i,A_i+1)∈[0,1] measures whether Ai+1A_i+1 follows logically from AiA_i assuming AiA_i were true, without consulting the ground-truth visual world. Scores close to 11 correspond to logically coherent, non-contradictory inferences; scores close to 00 indicate non sequiturs, unjustified leaps, or explicit logical conflicts. This isolates “thinking wrong” (logic failure) from “seeing wrong” (perception failure), and the chain-level LRS is given by LRS=1N−1∑i=1N−1Evallogic(Ai,Ai+1). LRS= 1N-1 _i=1^N-1Eval_logic(A_i,A_i+1). (8) 3. Self-Consistency Score (SCS): Narrative Stability. Trustworthy models must not contradict themselves over long contexts. SCS assesses the global consistency of the entire chain A, detecting self-contradictions, temporal or spatial inconsistencies, and memory lapses. The function returns a high score when the sequence of evidence atoms can be interpreted as a single coherent narrative without internal conflicts (e.g., an object does not change category or location inexplicably), and a low score when multiple atoms in A are mutually incompatible. We define SCS=Evalconsistency(A1,…,AN). SCS=Eval_consistency(A_1,…,A_N). (9) SCS therefore captures the stability of the model’s internal world model during complex multi-step reasoning. 4. Explainability Alignment Score (EAS): Decision Justification. This metric evaluates the interpretability of the final decision by checking whether it is causally and logically supported by the preceding evidence. Specifically, the metric measures to what extent the actionable decision ANA_N can be derived as a justified conclusion from the accumulated evidence (A1,…,AN−1)(A_1,…,A_N-1). A score of 11 indicates that the decision is a natural, well-supported consequence of the chain, while a score near 00 indicates that ANA_N behaves like a black-box guess that ignores or contradicts the prior evidence. EAS=Evalalignment((A1,…,AN−1),AN). EAS=Eval_alignment((A_1,…,A_N-1),A_N). (10) By decomposing evaluation into these orthogonal dimensions, we provide a granular diagnosis of the trustworthiness visual reasoning. 4 Methodology AD2-Bench (Sec. 3) reveals a critical dependency: robust reasoning in complex scenes is predicated on the strict acquisition of atomic evidence. Standard end-to-end MLLMs, however, rely on implicit inference lacking these visual anchors, leading to severe perception-logic disconnects. To bridge this gap, we propose Evidence-Grounded Visual Reasoning (EGVOR). Inspired by the “Atomic Annotation” paradigm, EGVOR reformulates reasoning as the explicit generation of verifiable evidence chains, optimized through a hierarchical curriculum of reflective supervision and cognitive reinforcement learning. 4.1 Preliminaries: Reasoning in Complex and Adverse Environments We first formalize visual reasoning in complex scenarios from a probabilistic perspective. The goal of this analysis is not to provide an exact generative model of MLLMs, but to construct a diagnostic abstraction that identifies how adverse conditions destabilize implicit reasoning and why explicit evidence construction mitigates this instability. Visual Perception under Signal Degradation. In AD2-Bench, the model observes imperfect sensory data. Let IcleanI_clean denote the latent, ideal visual scene. The observed image I~ I is modeled as a corrupted transformation: I~=ϕ(Iclean,ξenv), I=φ(I_clean, _env), (11) where ξenv _env represents environmental degradation variables such as rain streaks, fog, low illumination, or occlusion, and ϕφ denotes the corresponding physical degradation process. For each visual entity o in the scene, we denote by zoz_o its latent feature representation extracted from I~ I. Throughout this section, zoz_o is treated as the deterministic representation associated with entity o under the observed image, while residual feature uncertainty is modeled by the random representation Z in the variance analysis. Under adverse degradation, zoz_o may lose discriminative cues, making the mapping from (I~,Q)( I,Q) to the semantic answer A underdetermined. Modeling Scope and Object-Centric Factorization. Let O denote the query-dependent set of candidate visual entities in the scene. We introduce OfO_f as the latent visual-focus random variable whose realizations are elements of O. A lowercase o∈o denotes one possible realization of OfO_f. For notational simplicity, we write P(o|I~,Q)P(o| I,Q) as shorthand for P(Of=o|I~,Q)P(O_f=o| I,Q). Therefore, the summation ∑o∈ _o marginalizes over all possible realizations of the latent visual focus OfO_f. For a query Q, we approximate the answer posterior by marginalizing over this latent visual focus: P(A|I~,Q)=∑o∈P(A|zo,Q)⋅P(o|I~,Q). P(A| I,Q)= _o P(A|z_o,Q)· P(o| I,Q). (12) This factorization is a modeling approximation: P(o|I~,Q)P(o| I,Q) specifies where the model focuses, while P(A|zo,Q)P(A|z_o,Q) specifies what answer is supported by the feature representation at that focus. For a specific query Q, the candidate entity set is dynamically partitioned into the Critical Object Set critO_crit, the Irrelevant Set irrO_irr, and the Background Set bgO_bg. We further define non=irr∪bgO_non=O_irr _bg and =crit∪nonO=O_crit _non. Using this partition, Eq. 12 can be decomposed as P(A|I~,Q) P(A| I,Q) =∑o∈critP(A|zo,Q)P(o|I~,Q)⏟① Effective Term = _o _critP(A|z_o,Q)P(o| I,Q)_ 1 Effective Term (13) +∑o∈irrP(A|zo,Q)P(o|I~,Q)⏟② Distraction Term + _o _irrP(A|z_o,Q)P(o| I,Q)_ 2 Distraction Term +∑o∈bgP(A|zo,Q)P(o|I~,Q)⏟③ Hallucination Term. + _o _bgP(A|z_o,Q)P(o| I,Q)_ 3 Hallucination Term. The Effective Term contains query-relevant evidence that can support the correct answer. The Distraction Term corresponds to visually plausible but question-irrelevant entities. The Hallucination Term corresponds to background or environmental noise that does not provide reliable answer-specific semantics. Critical-Evidence Mass and Posterior Suppression. The key effect of adverse degradation is not merely that the visual focus becomes diffuse, but that probability mass is transferred from query-critical evidence to non-critical or weakly informative regions. To formalize this, let AgtA_gt denote the visually grounded correct answer, and define the critical-evidence mass: ρ=P(Of∈crit|I~,Q)=∑o∈critP(o|I~,Q). ρ=P(O_f _crit| I,Q)= _o _critP(o| I,Q). (14) Under the critical-evidence separability assumption, query-critical evidence supports AgtA_gt more reliably than non-critical evidence. Appendix 7.1 shows that, for two local focus distributions with critical-evidence masses 0<ρ2<ρ1<10< _2< _1<1, the posterior support for the grounded answer satisfies Pρ2(Agt|I~,Q)<Pρ1(Agt|I~,Q). P_ _2(A_gt| I,Q)<P_ _1(A_gt| I,Q). (15) Thus, the core failure mechanism is a reduction in grounded posterior support: the model assigns less probability to the correct visually supported answer when critical evidence mass is reallocated to distractors or background regions. The Failure of Implicit Aggregation: From Dilution to Prior Collapse. Standard end-to-end MLLMs do not explicitly separate the Effective, Distraction, and Hallucination Terms. Instead, they compress the scene into a global context vector via soft attention: zglobal=∑o∈ωozo,ωo=P(o|I~,Q). z_global= _o _oz_o, 9.24994pt _o=P(o| I,Q). (16) When ωo _o assigns substantial mass to irrO_irr or bgO_bg, the global feature is no longer dominated by zcritz_crit, but becomes a mixture of critical evidence, distractors, and background noise. Let ZglobalZ_global denote the random variable corresponding to the global visual representation, and let zglobalz_global denote its realization. In the limiting noise-dominant regime, the global visual feature carries little answer-specific information. We express this as approximate conditional independence: Zglobal⟂⟂A|Q. Z_global \!\!\! A Q. (17) In the exact conditional-independence case, P(zglobal|A,Q)=P(zglobal|Q)P(z_global|A,Q)=P(z_global|Q). Applying Bayes’ rule gives P(A|zglobal,Q) P(A|z_global,Q) =P(zglobal|A,Q)P(A|Q)P(zglobal|Q) = P(z_global|A,Q)P(A|Q)P(z_global|Q) (18) =P(A|Q). =P(A|Q). In practice, the independence is approximate rather than exact; hence the posterior moves toward the language prior, P(A|zglobal,Q)≈P(A|Q)P(A|z_global,Q)≈ P(A|Q). Let ApriorA_prior denote the answer favored by P(A|Q)P(A|Q). In failure cases where Aprior≠AgtA_prior≠ A_gt, prior collapse may yield P(Aprior|Q)>P(Agt|I~,Q). P(A_prior|Q)>P(A_gt| I,Q). (19) This inequality characterizes the prior-dominant hallucination regime, where grounded visual evidence is suppressed and the model falls back to language priors. Bottleneck Analysis: Semantic Entropy and Variance Drift. We now summarize the two statistical bottlenecks induced by the above failure mechanism. 1. Posterior Suppression and Semantic Uncertainty. We define the entropy of the latent visual-focus distribution as H(Of|I~,Q)=−∑o∈P(o|I~,Q)logP(o|I~,Q). H(O_f| I,Q)=- _o P(o| I,Q) P(o| I,Q). (20) The entropy is taken over the random variable OfO_f, whose realizations are candidate entities o∈o . We do not assume an unconditional implication from visual-focus entropy to answer entropy. Instead, Appendix 7.1 proves that a reduction in critical-evidence mass suppresses P(Agt|I~,Q)P(A_gt| I,Q), and further derives the answer entropy under a mixture model: Hρ(A|I~,Q)=−∑a∈pρ(a)logpρ(a), H_ρ(A| I,Q)=- _a p_ρ(a) p_ρ(a), (21) where A is the answer space and pρ(a)p_ρ(a) is the answer distribution induced by critical-evidence mass ρ. In the uncertainty-dominant failure regime, non-critical regions induce diverse and mutually inconsistent answer hypotheses; in that local regime, decreasing ρ increases Hρ(A|I~,Q)H_ρ(A| I,Q). In the prior-dominant hallucination regime, however, the model may be confidently wrong, and the failure is better captured by Eq. 19. Thus, semantic entropy increase and prior collapse are complementary manifestations of the same underlying issue: dilution of critical visual evidence. 2. Variance Decomposition via Latent Anchoring. To characterize the instability of the latent context, we treat Z as a random representation induced by the latent focus distribution. Since Z may be vector-valued, Var(Z)Var(Z) denotes either the scalar variance of a projected feature statistic or the trace of the covariance matrix. All expectations and variances over o are taken with respect to Of∼P(⋅|I~,Q). O_f P(·| I,Q). (22) By the Law of Total Variance, Var(Z|I~,Q) (Z| I,Q) =Of∼P(⋅|I~,Q)[Var(Z|Of,I~,Q)]⏟Intrinsic Semantic Uncertainty = E_O_f P(·| I,Q) [Var(Z|O_f, I,Q) ]_Intrinsic Semantic Uncertainty (23) +VarOf∼P(⋅|I~,Q)([Z|Of,I~,Q])⏟Spatial Boundary Ambiguity. + Var_O_f P(·| I,Q) (E[Z|O_f, I,Q] )_Spatial Boundary Ambiguity. The first term captures feature-level instability caused by environmental degradation. Even if the correct entity is localized, its representation may remain noisy due to fog, rain, low illumination, or occlusion. The second term captures instability in the visual focus itself. When OfO_f ranges over critO_crit, irrO_irr, and bgO_bg, the conditional expectation [Z|Of,I~,Q]E[Z|O_f, I,Q] varies across semantically unrelated regions, resulting in a high-variance mixture rather than a stable semantic anchor. Chain of Evidence (CoE) Formulation. To address these bottlenecks, we propose Chain of Evidence (CoE), which reformulates reasoning as a sequential process of Spatio-Semantic Verification. Valid evidence must satisfy a dual constraint: it should be spatially anchored to the Critical Set critO_crit and semantically verified against the visual content. Mathematically, CoE performs a Component Filtration operation. We define a structured evidence set E=e1,…,eK⊆crit, E=\e_1,…,e_K\ _crit, (24) where each e∈Ee∈ E denotes a critical object or region whose spatial location and semantic content are both valid for the current query. Not every critical object must be selected; E contains only the evidence necessary for the reasoning task. Let E′=∖E =O E be the complement set containing objects that fail the evidence constraints. CoE decomposes the answer posterior as P(A|I~,Q) P(A| I,Q) =∑e∈EP(A|ze,Q)P(e|I~,Q)⏟Signal Term = _e∈ EP(A|z_e,Q)P(e| I,Q)_Signal Term (25) +∑e′∈E′P(A|ze′,Q)P(e′|I~,Q)⏟Interference Term. + _e ∈ E P(A|z_e ,Q)P(e | I,Q)_Interference Term. CoE performs Component Filtration by increasing the relative contribution of the Signal Term and suppressing the Interference Term: PCoE(A|I~,Q)≈∑e∈EP(A|ze,Q)P(e|I~,Q). P_CoE(A| I,Q)≈ _e∈ EP(A|z_e,Q)P(e| I,Q). (26) This approximation does not claim that all interference disappears. Rather, it states the intended effect of CoE: to bias reasoning toward spatially grounded and semantically verified evidence, thereby reducing the propagation of environmental noise into the reasoning process. 4.2 Evidence-Grounded Visual Reasoning Inspired by the rigorous “Atomic Annotation” paradigm and the variance analysis in Sec. 4.1, we propose Evidence-Grounded Visual Reasoning (EGVOR). Unlike standard black-box MLLMs susceptible to high variance in the global context zglobalz_global, EGVOR reformulates visual reasoning not as a direct mapping, but as a sequential Variance Reduction Process. We introduce the concept of Implicit Evidence Atoms to concretely instantiate the Component Filtration in Eq. 26, explicitly bridging the gap between high-level semantic reasoning and low-level visual stochasticity. Definition 1: The Atomic Evidence Triplet. We formalize the fundamental unit of reasoning, the evidence atom ata_t, as a structured realization of the random variables defined in our probabilistic framework. An atom at step t is a triplet: at=⟨bt,t,ct⟩. a_t= b_t,h_t,c_t . (27) where, Explicit Spatial Anchor (btb_t) is a deterministic realization of the stochastic visual focus o introduced in Sec. 4.1. Specifically, btb_t parameterizes an axis-aligned bounding box in the image coordinate space that is expected to contain one or more candidate objects from O. By strictly enforcing btb_t to bound such a region, we truncate the probability space: conceptually, conditioning on btb_t induces a distribution P(o∣bt,I~,Q)P(o b_t, I,Q) whose support is restricted to those objects lying inside btb_t, and thus P(o∣bt,I~,Q)≈0P(o b_t, I,Q)≈ 0 for all o∉bto∉ b_t. This anchors the latent focus to a localized subset of O instead of the entire scene. Region-Aware Latent State (th_t) denotes the Implicit Visual Focus. It is not an external image crop, but a Soft Latent Aggregation of the MLLM’s internal representation constrained by the spatial anchor, corresponding to the local realization of the object-centric features zo\z_o\ around btb_t. Formally, let V∈ℝH×W×CV ^H× W× C be the global visual feature map encoded by the vision tower. We derive th_t via a spatial masking operation: t=Pool(V⊙ℳ(bt)). _t=Pool(V (b_t)). (28) where ℳ(bt)∈0,1H×WM(b_t)∈\0,1\^H× W is the binary spatial mask generated from the coordinates btb_t, and ⊙ denotes the element-wise Hadamard product. This masking operation implements signal denoising: by zeroing out features outside the critical region, it explicitly filters out the environmental background noise ξenv _env and ensures that the resulting state th_t aggregates only the signal energy relevant to the target. Grounded Description (ctc_t): ctc_t is the semantic decoding of th_t, representing the conditional distribution P(ct∣t)P(c_t _t), which serves as the interpretable verification of the feature’s semantic fidelity. In other words, ctc_t provides a textual instantiation of the latent state th_t and explicitly ties the visual evidence back to the answer space, making the connection between (o,zo)(o,z_o) and the final decision A observable at the atom level. Generative Process as Sequential Uncertainty Elimination. Within EGVOR, the inference process is modeled as a path integral over the space of evidence chains. The joint probability of generating a reasoning chain =a1,…,aTA=\a_1,…,a_T\ is factorized auto-regressively. Each step represents a Locate-Attend-Describe cycle designed to progressively reduce entropy: P(at∣a<t,I~,Q)= P(a_t a_<t, I,Q)= P(ct∣t,a<t)⏟Semantic Verification⋅P(t∣bt,I~)⏟Feature Denoising P(c_t _t,a_<t)_Semantic Verification· P(h_t b_t, I)_Feature Denoising (29) ⋅P(bt∣a<t,I~,Q)⏟Hypothesis Generation. · P(b_t a_<t, I,Q)_Hypothesis Generation. In our implementation, P(t∣bt,I~)P(h_t b_t, I) is realized by the deterministic mapping t=Pool(V⊙ℳ(bt))h_t=Pool(V (b_t)), i.e., it is a degenerate distribution concentrated on the masked aggregation of the visual features. We keep the factorization in Eq. 29 to emphasize the distinct conceptual roles of where to look (btb_t), what is seen (th_t), and how to describe it (ctc_t). This factorization imposes a structural constraint on the reasoning flow: 1) Hypothesis Generation: The model first predicts a spatial hypothesis btb_t based on the query Q and the previous atoms a<ta_<t, selecting a candidate region that is encouraged (by training objectives in Sec. 4.3) to overlap with the Critical Set critO_crit. 2) Feature Denoising: Conditioned on btb_t, the model deterministically extracts th_t, effectively pruning the interference terms in Eq. 25 originating from irrO_irr and bgO_bg and reducing the visual uncertainty around the current focus. 3) Semantic Verification: The feature th_t is decoded into text ctc_t, ensuring that the visual signal supports a coherent semantic claim. Maximizing P(ct∣t)P(c_t _t) aligns the latent state with a stable semantic manifold, which directly targets the semantic uncertainty term in Eq. 23. 1. Spatial Alignment via btb_t (Suppressing Term 2). To mitigate attention drift, we treat the explicit box btb_t as a hard spatial constraint Qloc(⋅∣bt)Q_loc(· b_t) over the object set O. Intuitively, Qloc(o∣bt)Q_loc(o b_t) puts most of its mass on objects whose support lies inside btb_t and near-zero mass elsewhere. The model, on the other hand, induces its own internal attention distribution Pattn(⋅)P_attn(·) over candidate objects through its vision-language layers. Minimizing the divergence between PattnP_attn and this constrain encourages the model’s “eye” to align with its “finger,” reducing the spatial variance Varo([Z∣o])Var_o(E[Z o]). 2. Semantic Alignment via ctc_t (Suppressing Term 1). Spatial fixing alone is insufficient if the feature remains noisy. We therefore introduce the grounded description ctc_t as a Semantic Control Variate. By maximizing the likelihood P(ct∣t)P(c_t _t), we enforce the noisy feature th_t to collapse onto a distinct semantic manifold defined by ctc_t. Mathematically, the composite alignment objective is formulated as: ℒalign=DKL(Qloc(⋅∣bt)∥Pattn(⋅∣t))⏟Spatial Anchor (Term 2) _align= D_KL (Q_loc(· b_t)\, \|\,P_attn(· _t) )_Spatial Anchor (Term 2) (30) −λlogP(ct∣t)⏟Semantic Anchor (Term 1). - λ P(c_t _t)_Semantic Anchor (Term 1). This encourages the latent state th_t to be both spatially localized and semantically stable, thereby minimizing the intrinsic variance o[Var(Z∣o)]E_o[Var(Z o)]. It is worth emphasizing that Eq. 30 describes a conceptual alignment objective implied by our variance analysis. In practice, we do not optimize ℒalignL_align explicitly. Instead, in the RL stage (Sec. 4.3), the spatial KL term is instantiated by the HRGW-based reward Ψspatial _spatial, which drives predicted boxes btb_t to concentrate on critO_crit and away from nonO_non, while the semantic log-likelihood term is realized via Ξalign _align, which rewards agreement between ctc_t and the ground-truth descriptions conditioned on th_t. These two objectives act as tractable proxies to approximate the ideal spatial and semantic alignment encoded in ℒalignL_align. Entropy Reduction via via Visual Probability Concentration. We now connect the evidence atom at=⟨bt,ht,ct⟩a_t= b_t,h_t,c_t to the probabilistic bottlenecks identified in Sec. 4.1. The spatial anchor btb_t acts as a support-truncation operator over the latent focus space. Let (bt)⊆O(b_t) denote the subset of objects whose spatial support lies inside the region defined by btb_t. Conditioning on btb_t induces the truncated focus distribution P(o|bt,I~,Q)=P(o|I~,Q)∑o′∈(bt)P(o′|I~,Q),o∈(bt),0,o∉(bt). P(o|b_t, I,Q)= cases P(o| I,Q) _o (b_t)P(o | I,Q),&o (b_t),\\ 0,&o (b_t). cases (31) This equation should be interpreted as support truncation rather than a universal entropy theorem. Conditioning on an arbitrary box does not necessarily reduce the actual entropy for every possible distribution. However, when btb_t is a valid and compact evidence anchor, it restricts the support of OfO_f from the full scene O to the local subset (bt)O(b_t). Therefore, the focus entropy under the anchor satisfies the support-based upper bound H(Of|bt,I~,Q)≤log|(bt)|. H(O_f|b_t, I,Q)≤ |O(b_t)|. (32) Since the evidence atom ata_t contains the anchor btb_t, its focus support is also contained in (bt)O(b_t): H(Of|at,I~,Q)≤log|(bt)|. H(O_f|a_t, I,Q)≤ |O(b_t)|. (33) For the unanchored case, the worst-case bound is H(Of|I~,Q)≤log||. H(O_f| I,Q)≤ |O|. (34) When btb_t is compact and query-critical, |(bt)|≪|||O(b_t)| |O|, yielding the upper-bound comparison log|(bt)|≪log||. |O(b_t)| |O|. (35) Thus, CoE reduces the worst-case support size of the latent visual focus. In the ideal case where btb_t tightly covers a single critical object, P(Of=ocrit|bt,I~,Q)≈1P(O_f=o_crit|b_t, I,Q)≈ 1, and the anchored focus entropy approaches zero. The grounded description ctc_t further provides semantic verification. It checks whether the localized feature hth_t corresponds to a coherent visual concept relevant to the query. Therefore, btb_t reduces where-to-look uncertainty, while ctc_t reduces what-is-seen ambiguity. Dual-Variance Suppression. CoE also modifies the variance decomposition in Eq. 23 through two complementary constraints. First, the spatial anchor btb_t replaces the global focus distribution P(⋅|I~,Q)P(·| I,Q) with a localized conditional distribution: Of∼P(⋅|bt,I~,Q). O_f P(·|b_t, I,Q). (36) The corresponding spatial variance term becomes VarOf∼P(⋅|bt,I~,Q)([Z|Of,I~,Q]). _O_f P(·|b_t, I,Q) (E[Z|O_f, I,Q] ). (37) Because the support of P(⋅|bt,I~,Q)P(·|b_t, I,Q) is contained in (bt)O(b_t), a compact and query-critical anchor reduces the worst-case support over which the spatial variance is computed. This directly targets the Spatial Boundary Ambiguity term in Eq. 23. Second, the grounded description ctc_t regularizes the noisy local feature hth_t toward a stable semantic interpretation. Maximizing P(ct|ht)P(c_t|h_t) does not eliminate all feature noise, but it encourages the localized representation to align with a coherent semantic concept. This targets the Intrinsic Semantic Uncertainty term: Of∼P(⋅|I~,Q)[Var(Z|Of,I~,Q)]. _O_f P(·| I,Q) [Var(Z|O_f, I,Q) ]. (38) Together, btb_t and ctc_t address the two bottlenecks identified in Sec. 4.1: btb_t suppresses spatial boundary ambiguity by localizing the focus distribution, while ctc_t suppresses intrinsic semantic uncertainty by verifying the semantic content of the localized feature. 4.3 The EGVOR Training Framework: A Hierarchical Curriculum Learning Figure 6: The two-stage training pipeline of EGVOR. (a) The bootstrapping stage uses supervised instruction tuning to teach the model the syntax of grounded reasoning. (b) The policy refinement stage employs reinforcement learning with our novel multi-faceted reward function and spatial-semantic alignment. We propose a Hierarchical Curriculum Learning framework to train the EGVOR model. Direct optimization of the full evidence chain is intractable due to the non-differentiable explicit spatial anchors (btb_t) and the unobservable implicit latent states (th_t). Therefore, we decompose training into a two-stage curriculum: first establishing the Evidence Reasoning Chain via Supervised Fine-Tuning (SFT), and then performing Cognitive Alignment via Reinforcement Learning (RL). 4.3.1 Stage I: Establishing Evidence Reasoning via Reflective SFT The primary objective of the first stage is to initialize the policy πθ _θ to adhere to the syntactic structure of the atomic paradigm. We utilize the SFT-Instruct dataset to teach the model the Locate-Attend-Describe format. Evidence Chain Establishment (Syntactic Prior). We train the model to explicitly represent its reasoning process by maximizing the likelihood of the ground-truth evidence chain τ. Let the trajectory τ be a sequence of L tokens, denoted as τ=(τ1,τ2,…,τL)τ=( _1, _2,…, _L). The generation follows the template: τ=<think>…</think>→<box>[x1,y1,x2,y2]</box>→<ref>ct</ref>τ= <think>… </think>→ <box>[x_1,y_1,x_2,y_2] </box>→ <ref>c_t </ref>. The optimization objective is the standard autoregressive Negative Log-Likelihood (NLL) over the token sequence: ℒSFT=−(I,Q,τ)∼SFT∑l=1Llogπθ(τl∣τ<l,I,Q). _SFT=-E_(I,Q,τ) _SFT _l=1^L _θ( _l _<l,I,Q). (39) This phase establishes the foundational association between visual perception and explicit spatial output, serving as a warm-up for the policy to learn the valid output space of the evidence atoms. Cognitive Recovery via Error-Induced Self-Correction. Standard SFT often suffers from Exposure Bias, where the model lacks the ability to recover from cumulative errors during inference. To address the Attention Drift discussed in Section 4.1, we introduce a Reflective Decoding mechanism. We augment the training data by injecting procedural perturbations: for a valid reasoning step, we substitute the correct bounding box with a distractor box berrb_err in a subset of samples, and supervise the model to generate a correction trajectory τcorr _corr that explicitly rejects the error: τcorr=berr→<think>Wait, the selected region is _corr=b_err→ <think>Wait, the selected region is (40) background…Refine focus.</think>→bgt→cgt. ...Refine focus. </think>→ b_gt→ c_gt. This mechanism teaches the model the conditional probability P(Correction∣berr,)P(Correction b_err,h), equipping it with the meta-cognitive capability to detect spatial hallucinations and realign its focus. 4.3.2 Stage I: Cognitive Alignment via Reinforcement Learning While SFT establishes the format, it does not guarantee the stability of the variance terms identified in Eq. 23. In the second stage, we utilize the RL-Instruct dataset to numerically solve the Dual-Variance Suppression problem. We formulate three functional objectives optimized via Group Relative Policy Optimization (GRPO). Objective 1: Minimizing Spatial Ambiguity (Ψspatial _spatial). (Targeting Term 2: Varo([Z∣o])Var_o(E[Z o])) Adverse weather blurs object boundaries, creating high spatial aleatoric uncertainty. Standard IoU rewards fail due to sparsity: they yield zero feedback for non-overlapping predictions. To address this, we propose the Hierarchical Robust Gaussian-Wasserstein (HRGW) objective. We model bounding boxes as 2D Gaussian distributions (μ,Σ)N(μ, ) and quantify affinity via the Wasserstein-2 distance (W2W_2), wrapped in a Normalized Wasserstein Distance: (bpred,bgt)=exp(−W22(pred,gt)C). (b_pred,b_gt)= (- W_2^2(N_pred,N_gt)C ). (41) where, C is a normalization constant empirically set to align with the dataset scale. This provides dense reward shaping even for disjoint predictions to guide the policy towards the Critical Set critO_crit. To implement spatial pruning of distractors, we define the spatial gain Ψspatial _spatial. Let τb⊂τ _b⊂τ denote the set of all spatial anchors (bounding boxes) parsed from the generated trajectory. For each annotated object o∈o , let bogtb^gt_o denote its ground-truth bounding box. We assign an importance weight ωo _o where sgn(ωo)=+1sgn( _o)=+1 for targets (o∈crito _crit) and sgn(ωo)=−1sgn( _o)=-1 for non-critical objects (o∈nono _non). The reward is: Ψspatial(τ)=αrec⋅∑o∈critmaxb∈τb(b,bogt)|crit| _spatial(τ)= _rec· _o _crit _b∈ _bA(b,b^gt_o)|O_crit| (42) +αprec⋅1|τb|∑b∈τbmaxo∈(sgn(ωo)⋅(b,bgto)). + _prec· 1| _b| _b∈ _b _o (sgn( _o)·A(b,b^gt_o) ). The first term improves recall by encouraging coverage of critical objects. The second term is precision-oriented: each predicted box is matched to its closest annotated object; matches to critO_crit contribute positively, while matches to nonO_non contribute negatively, explicitly discouraging attention drift to distractors/background. Additionally, we incorporate a box-format validity reward ℛbox(τ)R_box(τ) to encourage syntactically and geometrically well-formed predictions. Concretely, ℛbox(τ)=1R_box(τ)=1 if all <box>…</box> spans in τ are present and each box is in the correct coordinate order (e.g., x1<x2x_1<x_2, y1<y2y_1<y_2); otherwise ℛbox(τ)=0R_box(τ)=0. Objective 2: Enforcing Semantic Stability (Ξalign _align). (Targeting Term 1: o[Var(Z∣o)]E_o[Var(Z o)]) Even with correct localization, feature attenuation persists. As established in Eq. 30, robust reasoning requires minimizing the divergence between the latent state th_t and the semantic description ctc_t. We instantiate the log-likelihood constraint logP(ct∣t) P(c_t _t) using a semantic similarity metric by GPT-5.2 (SimScore) as a reward proxy: Ξalign(τ)=1T∑t=1TSimScore(ct,cgtmatched∣bt). _align(τ)= 1T _t=1^TSimScore(c_t,c_gt^matched b_t). (43) Maximizing Ξalign _align effectively forces the noisy latent features th_t to collapse onto the deterministic semantic manifold defined by the ground truth attributes, thereby suppressing the intrinsic semantic uncertainty. Objective 3: Maximizing Cognitive Path Diversity (Λpath _path). To approximate the path integral ∑P(A∣E)Σ P(A E), the policy should avoid collapsing to a single mode. We encourage diverse reasoning strategies by matching the generated chain against a set of M valid trajectories (e.g., different reasoning orders): Λpath(τ)=maxm∈1..MSimlogic(τthink,τgt(m))+ℛfmt. _path(τ)= _m∈\1..M\Sim_logic( _think, _gt^(m))+R_fmt. (44) where ℛfmtR_fmt is a structural term enforcing the validity of the reasoning format (e.g., correct usage and ordering of <think> tags). Summary: The GRPO Optimization. We integrate these objectives into a unified Group Relative Policy Optimization framework. For a given query Q, we sample a group of trajectories G=τ(1),…,τ(G)G=\τ^(1),…,τ^(G)\ and compute the total advantage JtotalJ_total for each trajectory: Jtotal(τ(i))=Ψspatial(τ(i))+Ξalign(τ(i)) J_total(τ^(i))= _spatial(τ^(i))+ _align(τ^(i)) (45) +Λpath(τ(i))+(Apred=Agt)+ℛbox(τ(i)). + _path(τ^(i))+I(A_pred=A_gt)+R_box(τ^(i)). The policy parameters θ are updated to maximize the expected return relative to the group-wise weighting, which stabilizes training in high-variance environments: ∇θℒGRPO=Q∼ _θL_GRPO=E_Q (46) [1G∑i=1Gexp(Jtotal(τ(i))/β)∑jexp(Jtotal(τ(j))/β)∇θlogπθ(τ(i)∣Q)]. [ 1G _i=1^G (J_total(τ^(i))/β) _j (J_total(τ^(j))/β) _θ _θ(τ^(i) Q) ]. where β is a temperature hyper-parameter controlling the sharpness of the group-wise weighting. This formulation translates the probabilistic constraints of high-entropy reasoning into a differentiable optimization process, and empirically aligns the model’s perception and reasoning with the intended cognitive patterns. 5 Experiments 5.1 Benchmark Evaluation Setup and Implementation Details Model Ave. LLM Existence Counting Location Detection OCR-F1 OCR-CER Attribute Position Cap/Other Spatial Occlusion Social Causal w/o CoE w/ CoE Basic Perception Advanced Perception Relation Understanding Event Reasoning Proprietary Models Gemini-3.5-Flash-0519 61.4 Gemini 68.24 53.0 38.0 7.28 84.80 39.89 63.03 54.25 84.0 66.0 58.0 71.0 58.0 66.8 69.6 41.6 71.5 63.3 69.0 Gemini-2.5-Flash-0520 55.5 Gemini 67.1 43.3 43.3 6.9 94.1 10.3 57.7 46.8 75.9 55.0 47.9 66.8 43.8 53.1 58.9 40.2 68.6 53.4 58.9 GPT-4o-2024-1120 53.1 GPT 58.48 53.2 48.84 10.27 84.17 30.55 32.66 29.41 77.0 55.5 46.5 60.5 55.0 56.8 59.6 42.7 55.8 54.4 59.6 GPT-4.1-mini-0414 52.0 GPT 63.8 46.8 46.2 2.4 91.6 11.2 57.4 36.3 82.2 47.0 39.9 51.1 53.4 51.2 55.3 39.8 66.9 47.9 55.3 Claude-3.5-Sonnet-0620 47.8 Claude 60.3 45.2 32.7 4.8 71.7 40.5 52.6 42.9 65.8 50.0 43.8 63.6 34.8 45.7 51.4 35.8 58.2 48.0 51.4 Open Source Models LLaVA-1.5-7B 40.9 Vicuna 50.8 50.5 19.1 0.0 24.1 173.1 37.2 36.1 62.2 50.1 51.0 67.4 50.4 32.0 38.9 30.1 39.9 54.7 38.9 LLaVA-NeXT-7B 41.0 Vicuna 52.8 43.2 19.1 0.0 32.8 125.5 37.5 26.0 67.3 44.5 50.9 63.2 45.8 34.2 43.2 28.8 40.9 51.1 43.2 LLaVA-OneVision-7B 44.8 Qwen2 43.7 22.9 20.7 0.1 56.5 65.7 49.7 29.1 71.9 50.5 42.6 73.7 53.4 43.8 50.5 21.9 51.8 55.1 50.5 MiniCPM-V-2.5-8B 43.9 Llama3 59.5 53.6 24.5 5.7 67.3 41.0 50.3 22.5 71.6 49.5 37.5 22.3 47.9 40.3 47.5 35.8 52.9 39.3 47.5 MiniCPM-V-2.6-8B 49.8 Qwen2 55.5 56.2 42.9 7.7 78.5 27.5 53.6 40.4 74.2 43.3 41.0 54.7 51.9 42.1 49.3 40.6 61.7 47.7 49.3 MiniCPM-o-2.6-8B 53.1 Qwen2.5 63.1 65.6 27.7 5.6 84.2 27.3 56.8 51.9 78.9 55.6 39.3 62.1 54.2 44.0 51.2 40.5 67.9 52.8 51.2 Janus-Pro-7B 45.9 DeepSeek 56.6 58.6 21.9 0.1 22.0 143.6 54.0 46.0 75.2 49.5 40.2 63.2 51.9 39.3 48.8 34.3 49.3 51.2 48.8 Qwen2-VL-7B 47.9 Qwen2 58.1 63.3 37.2 4.5 59.0 49.1 49.3 30.6 72.7 48.2 45.3 71.6 41.2 39.2 46.5 40.8 52.9 51.6 46.5 Qwen2.5-VL-7B 55.0 Qwen2.5 58.7 59.4 68.1 51.0 86.5 15.4 54.2 33.7 75.7 46.1 42.2 52.6 42.8 41.7 52.3 59.3 62.5 45.9 52.3 Qwen2.5-VL-32B 59.3 Qwen2.5 54.9 59.4 79.4 52.7 87.7 13.5 64.1 44.0 77.5 54.2 49.5 60.4 42.8 47.2 55.7 61.6 68.3 51.7 55.7 Qwen2.5-VL-72B 60.7 Qwen2.5 63.5 60.2 76.8 54.5 89.5 12.6 60.8 39.1 79.3 53.8 56.0 64.5 44.0 50.8 57.2 63.7 67.2 54.6 57.2 InternVL2-8B 50.2 InternLM 61.2 55.5 38.0 6.2 83.3 35.3 59.1 38.2 67.5 45.9 41.9 53.7 46.5 43.8 51.7 40.2 62.0 47.0 51.7 InternVL2.5-8B 52.8 InternLM 62.1 54.0 25.1 5.5 89.7 15.0 59.9 45.2 78.4 53.7 43.2 65.3 52.7 42.9 52.6 36.7 68.3 53.7 52.6 InternVL3-8B 54.8 Qwen2.5 59.6 57.1 41.4 6.3 89.7 15.1 57.9 56.4 78.1 58.8 45.6 60.0 54.2 42.2 53.0 41.1 70.5 54.7 53.0 InternVL3-38B 57.4 Qwen2.5 61.8 65.7 29.2 6.1 87.3 16.8 67.7 53.6 80.4 65.3 53.5 65.8 54.2 48.5 56.8 40.7 72.3 59.7 56.8 InternVL3-78B 60.2 Qwen2.5 67.3 65.1 43.1 6.2 91.7 11.9 69.4 52.8 82.6 64.5 54.6 68.1 55.7 52.7 60.5 45.4 74.1 60.7 60.5 Deepeyes-7B 63.8 Qwen2.5 65.2 62.1 82.4 64.6 87.4 13.6 66.1 55.4 76.9 54.6 53.7 58.5 45.1 60.9 62.1 68.6 71.5 53.0 62.1 Pixel-Reasoner-7B 61.1 Qwen2.5 60.5 61.3 79.6 57.3 86.1 15.8 65.3 54.2 76.1 49.1 50.5 56.2 43.9 57.4 59.3 64.7 70.4 49.9 59.3 EGVOR(ours) 67.7 Qwen2.5 68.4 66.2 85.9 69.7 88.2 13.0 68.9 57.7 79.3 60.4 56.6 63.3 48.3 65.7 67.4 72.6 73.5 57.1 67.4 Δ v.s. Qwen2.5-VL-7B ↑ 12.7 Qwen2.5 ↑ 9.7 ↑ 6.8 ↑ 17.8 ↑ 18.8 ↑ 1.7 ↓ 2.4 ↑ 14.7 ↑ 24.0 ↑ 3.7 ↑ 14.4 ↑ 14.4 ↑ 10.6 ↑ 5.6 ↑ 24.0 ↑ 15.1 ↑ 13.3 ↑ 11.0 ↑ 11.2 ↑ 15.1 Table 2: Comprehensive Benchmark Results. Grouped by model series. Δ indicates improvement over the strong baseline Qwen2.5-VL-7B (green indicates improvement; for OCR-CER, lower ↓ is better). Our performance is highlighted in bold. V* Bench HR-Bench-4K HR-Bench-8K Overall Attr. Spatial Overall Single Cross Overall Single Cross Private Models GPT-4o-2024-1120 66.0 – – – – – – – – o3-0416 95.7 – – – – – – – – Open-source General Models LLaVA-OneVision-7B 70.7 73.0 60.5 64.3 74.8 53.8 59.8 65.3 54.3 LLaVA-OneVision-72B 73.8 80.9 63.2 66.3 76.5 56.0 60.9 68.8 53.0 InternVL3-8B 72.3 73.0 71.1 70.8 79.3 62.3 62.0 64.3 59.8 InternVL3-38B 77.5 77.4 77.6 76.3 83.5 69.0 67.0 71.3 62.8 InternVL3-78B 76.4 75.7 77.6 75.5 84.5 66.5 67.3 71.8 62.8 Qwen2.5-VL-7B 74.3 77.4 69.7 72.1 88.8 55.5 68.8 83.5 54.0 Qwen2.5-VL-32B 85.9 83.5 89.5 74.8 89.3 60.3 71.6 86.5 56.8 Qwen2.5-VL-72B 84.8 90.8 80.9 79.4 88.8 70.0 76.3 84.3 68.3 Open-source Visual Grounded Reasoning Models Pixel-Reasoner-7B 80.6 83.5 76.3 72.9 86.0 60.3 66.9 80.0 54.3 DeepEyes-7B 90.0 92.1 86.8 75.1 91.3 59.0 72.6 86.8 58.5 EGVOR (Ours) 91.9 95.4 88.3 78.6 91.8 65.4 75.0 89.5 60.4 Δ v.s. Qwen2.5-VL-7B ↑ 17.6 ↑ 18.0 ↑ 18.6 ↑ 6.5 ↑ 3.0 ↑ 9.9 ↑ 6.2 ↑ 6.0 ↑ 6.4 Table 3: Comparison with state-of-the-art alternatives on V* Bench 83 and HRBench 79. The best performance is highlighted in bold. Perception Reasoning Overall OCR RS DT MO AD OCR DT MO AD Private Models GPT-4o-2024-1120 46.4 81.0 45.2 65.0 34.1 37.0 72.0 50.3 42.0 33.0 General Models Qwen2.5-VL-7B 42.3 87.6 32.7 83.0 27.3 30.0 72.0 62.0 28.7 23.0 Qwen2.5-VL-32B 45.6 87.2 40.7 83.0 29.5 40.7 74.0 60.0 27.3 29.5 Qwen2.5-VL-72B 43.7 90.8 34.0 87.0 27.9 30.6 74.0 61.0 26.7 25.5 LLaVA-OneVision-7B 43.7 80.0 40.0 56.0 31.7 39.4 65.0 33.0 38.0 32.0 LLaVA-OneVision-72B 48.7 79.2 50.7 67.0 37.9 40.0 76.0 41.0 38.7 39.3 InternVL3-8B 47.9 83.6 49.3 75.0 34.5 36.9 70.0 44.0 40.0 37.0 InternVL3-38B 51.0 85.6 56.0 71.0 42.6 40.0 77.0 45.0 47.3 35.0 InternVL3-78B 52.3 87.6 54.7 77.0 42.6 36.6 76.0 56.0 46.0 40.3 Visual Grounded Reasoning Models Pixel-Reasoner-7B 49.7 89.6 52.0 86.0 38.9 30.9 71.0 72.0 46.0 32.5 DeepEyes-7B 53.2 90.0 52.7 89.0 43.3 33.4 76.0 69.0 44.0 35.0 EGVOR (Ours) 55.4 87.8 51.8 83.2 48.4 44.9 74.5 66.7 52.2 40.1 Δ v.s. Qwen2.5-VL-7B ↑ 13.1 ↑ 0.2 ↑ 19.1 ↑ 0.2 ↑ 21.1 ↑ 14.9 ↑ 2.5 ↑ 4.7 ↑ 23.5 ↑ 17.1 Table 4: Comparison with state-of-the-art alternatives on MME-RealWorld-Lite 97. The best performance is highlighted in bold. Attributes Material Phy. State Obj. Retr. OCR Per. Trans. Ordering Con. & Oc. Spa. Cont. Comparison Overall mIoU Perception Reasoning Private Models Gemini-2.5-Flash-0520 45.9 – 48.3 53.9 69.6 68.8 75.0 15.3 19.3 56.1 72.4 43.2 GPT-4o-2024-1120 46.9 – 51.7 61.5 65.2 43.8 69.1 18.8 38.6 48.8 72.4 43.2 Gemini-2.5-Pro-0605 54.1 – 51.7 61.5 56.5 75.0 83.8 20.0 36.8 65.9 86.2 54.6 o3-0416 54.8 – 69.0 69.2 65.2 68.8 79.4 22.4 38.6 61.0 86.2 50.0 Open-source General Models LLaVA-OneVision-7B 37.3 – 55.2 53.8 56.5 50.0 32.4 21.2 22.8 41.5 72.4 36.4 LLaVA-OneVision-72B 40.5 – 62.1 53.8 65.2 62.3 36.8 12.9 28.1 53.7 65.5 47.7 MiniCPM-V-2.6-8B 30.1 – 51.7 46.2 60.9 62.5 42.6 0.0 3.5 36.6 62.1 29.5 Qwen2.5-VL-7B 37.0 – 55.2 53.8 56.5 62.5 27.9 20.0 35.1 39.0 44.8 43.2 Qwen2.5-VL-32B 42.5 – 51.7 53.8 69.6 62.5 54.4 16.5 33.3 46.3 62.1 38.6 Qwen2.5-VL-72B 42.2 – 65.5 69.2 56.5 56.3 48.5 11.8 33.3 51.2 72.4 38.6 InternVL3-8B 38.8 – 51.7 69.2 56.5 56.3 33.7 21.2 24.6 39.0 72.4 43.2 InternVL3-38B 42.0 – 51.7 61.5 52.2 68.8 51.5 12.9 33.3 56.1 65.5 38.6 InternVL3-78B 46.4 – 62.1 61.5 52.2 68.8 52.9 16.5 33.3 61.0 86.2 45.5 Open-source Visual Grounded Reasoning Models DeepEyes-7B 37.5 30.0 62.1 53.8 65.2 68.8 51.5 11.8 24.6 36.6 51.7 47.7 Pixel-Reasoner-7B 39.0 35.7 58.6 61.5 65.2 50.0 48.5 14.1 31.6 39.0 44.8 40.9 EGVOR (Ours) 51.3 45.0 69.1 54.4 83.4 69.3 64.1 23.2 38.6 58.7 69.4 45.8 Δ v.s. Qwen2.5-VL-7B ↑ 14.3 – ↑ 13.9 ↑ 0.6 ↑ 26.9 ↑ 6.8 ↑ 36.2 ↑ 3.2 ↑ 3.5 ↑ 19.7 ↑ 24.6 ↑ 2.6 Table 5: Selected results of different models on Treebench 74. Evaluations of open-source general models are implemented using VLMEvalKit 22, while evaluations of visual grounded reasoning models are conducted by us. Reasoning pathways of o3 59 are unavailable, and thus traceable evaluations are not valid. Best performances for open-source models are highlighted in bold. Our EGVOR achieves comparable performance with InternVL3-78B 101. We evaluate a diverse array of MLLMs to benchmark complex urban cognition comprehensively. Departing from prior studies limited to edge models, our assessment spans from efficient 7B-parameter baselines to large-scale server-grade architectures and proprietary APIs. 1) Model Selection and Taxonomy. We evaluated a total of 18 distinctive MLLMs to analyze the scaling laws and architectural impacts on complex scene reasoning. For the open-source category, we selected representative series including InternVL (InternVL2, 2.5, 3) 101; 14; 16, Qwen-VL (Qwen2-VL 77, Qwen2.5-VL 4), MiniCPM (V2.5, V2.6, o2.6) 91, LLaVA (1.5, Next, OneVision) 43; 38; 44 and Janus-Pro 12. To probe the capabilities of reasoning on larger-scale models, we extended the evaluation beyond the standard 7B/8B limit to include high-parameter variants such as Qwen2.5-VL-32B, Qwen2.5-VL-72B, InternVL3-34B, and InternVL3-78B. Furthermore, we incorporated specialized reasoning models including DeepEyes 100 and PixelReasoner 70 to assess domain-specific adaptations. In addition, we also benchmarked leading proprietary closed-source models to establish upper-bound performance references. This set includes Gemini-2.5-Flash 18, GPT-4.1-mini 57 and Claude-3.5-Sonnet. 2) Evaluation and Training Environment. The evaluation pipeline for all open-source models was similarly standardized using MS-SWIFT 98 to ensure fair comparison. To guarantee deterministic reproducibility, we rigorously fixed the random seed and set the decoding temperature to 0 for all generation tasks. For the evaluation of open-ended reasoning, we adopted the “LLM-as-a-Judge” paradigm, utilizing gpt-5.2-2025-12-11 to quantify the semantic similarity between model predictions and ground truth annotations. All experiments were conducted on a high-performance cluster equipped with 8×NVIDIA H20 GPUs. For training, we implemented the SFT stage using the MS-SWIFT 98 framework (learning rate 5e−65e-6, batch size 32) and the RL stage via EasyR1 99 (learning rate 1e−61e-6). 5.2 Dataset Construction for Training To effectively train EGVOR via the proposed two-stage curriculum, we constructed two bespoke datasets: SFT-Instruct for establishing the reasoning format, and RL-Instruct for optimizing cognitive alignment. SFT-Instruct Construction. We curated a high-fidelity dataset of 40K reasoning chains derived from VGR-158K 75. To align with our atomic requirements, we implemented a rigorous restructuring pipeline: (1) Coordinate Adaptation: We converted all normalized coordinates to absolute pixel values to fully leverage the high-resolution grounding capabilities of the Qwen2.5-VL backbone. (2) Semantic Densification: To support robust evidence generation, we filter for multi-hop reasoning chains and explicitly augment sparse descriptions. This ensures that every spatial anchor btb_t is grounded by a detailed caption ctc_t, providing indispensable supervision for the latent alignment objective. (3) Reflective Augmentation: To instill self-correction capabilities, we constructed a “Reflective Subset” (5K samples) via synthetic error injection. For valid reasoning steps, we randomly inserted distractor boxes followed by corrective thought tokens, training the model to dynamically detect and rectify grounding errors. RL-Instruct Construction. For the second stage, we constructed a high-quality dataset of 40K samples designed to provide dense supervision for the composite objective JtotalJ_total. We employed a Dual-Hybridization Strategy to balance general reasoning with domain-specific robustness: (1) General Hard Samples (30K): Sourced from V*-training set 83, we utilized a teacher model (Gemini-2.5-Pro 19) to mine “hard samples” characterized by high visual entropy or complex spatial dependencies, ensuring the retention of broad multimodal capabilities. (2) Complex Urban Samples (10K): To enforce entropy reduction in high-noise environments, we integrated samples from VisDrone 102 (UAV-view) and nuScenes 6 (Ego-view). This hybrid composition prevents overfitting while forcing the policy to handle extreme visual confusion. To support our reward functions, we enriched these samples with hierarchical metadata. We assigned adaptive importance weights to objects to facilitate the spatial reward Ψspatial _spatial, and generated diverse reasoning trajectories (e.g., bottom-up perceptual vs. top-down semantic paths) using advanced LLMs to support the path diversity objective Λpath _path. Further details on the filtering protocols and annotation generation are provided in the Appendix 5.2. 5.3 Main Results on AD2-Bench Table 2 presents the performance landscape of current MLLMs, spanning proprietary APIs, large-scale open-source architectures (up to 72B), and specialized reasoners. Metric Setup. To ensure a rigorous evaluation of cognitive reliability, the reported Ave. score denotes the arithmetic mean of accuracy-based metrics, excluding OCR-CER (where lower is better) and the w/o CoE setting. Crucially, for the Event Reasoning (w/ CoE) metric, we employ a hybrid alignment protocol: the final score is a weighted sum of the Process Verification Score (λ=0.4λ=0.4, assessing the quality of intermediate evidence) and the Decision Accuracy (λ=0.6λ=0.6). This strict criterion penalizes models that arrive at correct answers via hallucinated reasoning paths. Baseline Analysis. Under this rigorous protocol, a discernible scaling law emerges: large-parameter models like Qwen2.5-VL-72B and InternVL3-78B exhibit robust generalization, achieving average scores of 60.67% and 60.20% respectively. Proprietary models (e.g., GPT, Gemini) similarly demonstrate strong capabilities in high-level reasoning. However, a systemic “Perception-Reasoning Disconnect” persists across all baselines. Even the leading open-source model, Qwen2.5-VL-7B, suffers drastic performance drops in Basic Perception (e.g., 50.95% in Detection), significantly lagging behind its logic capabilities. Furthermore, specialized models like Pixel-Reasoner70, while designed for grounding, fail to fully mitigate Spatial Ambiguity in our adverse scenarios, underscoring the necessity for explicit evidence construction. Effectiveness of EGVOR. By explicitly constructing evidence chains, our EGVOR (7B) achieves a state-of-the-art average score of 67.66%, remarkably surpassing its 10x larger counterpart (Qwen2.5-VL-72B). In the Basic Perception tier, EGVOR secures a +13.3% gain over the baseline, with substantial improvements in Location (+17.8%) and Detection (+18.8%), validating that our “Atomic Evidence” generation effectively anchors attention to valid targets amidst clutter. In Advanced Perception, the +11.0% improvement, particularly in Attribute Recognition (+14.7%), confirms that verifying evidence details enhances resilience against feature degradation. Crucially, in the final Event Reasoning tier, we employ a hybrid alignment protocol for the w/ CoE metric (λ=0.4λ=0.4 for process quality, λ=0.6λ=0.6 for decision accuracy) to penalize right-answers-for-wrong-reasons. On reasoning task, EGVOR outperforms the baseline by +15.1% and specialized architectures like Pixel-Reasoner (+8.1%), demonstrating that the explicit Chain of Evidence paradigm establishes a more trustworthy and explainable path for complex multimodal decision-making. 5.4 Generalization to Diverse Benchmarks To verify that EGVOR has acquired robust cognitive capabilities rather than merely overfitting to the AD2-Bench domain, we conducted extensive evaluations on three external benchmarks covering high-resolution perception, real-world complexity, and traceable reasoning. Fine-Grained Perception in High-Resolution Scenarios. Standard MLLMs often suffer from attention dispersion when processing high-resolution images or identifying minute details. As shown in Tab. 3, EGVOR achieves state-of-the-art performance across all metrics, validating the efficacy of our explicit spatial anchoring mechanism. Specifically, on V*-Bench, which evaluates the ability to locate and recognize small, hard-to-find objects, EGVOR outperforms the Qwen2.5-VL-7B by a remarkable margin of +17.6%. This confirms that our Locate-Attend paradigm effectively acts as a “cognitive magnifying glass,” enabling the model to actively search for and focus on subtle visual cues that are typically lost in the global pooling of standard encoders. Furthermore, in the 4K and 8K scenarios of HR-Bench, where valid targets are sparse within vast backgrounds, our model maintains a significant lead (e.g., +9.9% on 4K Cross). Notably, the improvement on the Cross subset—which involves multi-image or complex relational reasoning—is consistently higher than on the Single subset. This suggests that our Chain of Evidence not only sharpens local perception but also enhances the capability to maintain context across large-scale visual fields, preventing the model from being overwhelmed by irrelevant background noise. Robustness in Real-World Domains. We further evaluate the model on MME-RealWorld in Tab. 4, which spans diverse domains including OCR, Remote Sensing (RS), Diagram/Table (DT), Monitoring (MO), and Autonomous Driving (AD). As expected, EGVOR exhibits dominant performance in Autonomous Driving (AD) and Monitoring (MO), with perception gains of +14.9% and +21.1%, respectively. These domains share the characteristics of complex dynamics and tiny objects with our training data, proving that the robustness learned from this data effectively transfers to general surveillance and driving tasks. Surprisingly, we also observe a massive improvement in Remote Sensing (RS) perception (+19.1%). Since RS imagery typically involves top-down views of tiny objects (similar to UAV views), this indicates that our model has learned a generalized feature representation for tiny object discovery, independent of the specific camera perspective (ego-centric vs. top-down). Crucially, the gains in the Reasoning tier for MO (+23.5%) and AD (+17.1%) significantly outpace the perception gains. This improvement underscores that our method does not just “see better” but leverages this clearer vision to “think deeper,” effectively resolving complex causal chains in dynamic environments. The Primacy of Grounding in Reasoning. Finally, we utilize TreeBench 74 to dissect the correlation between localization precision and reasoning accuracy. As illustrated in Tab. 5, EGVOR achieves the highest performance in Object Retrieval (+26.9%) and Spatial Context (+24.6%) tasks that demand rigorous spatial understanding. Consistent with our hypothesis, we observe a strong positive correlation between localization quality (mIoU) and reasoning accuracy. Unlike baseline models that may hallucinate correct answers from language priors (high Acc, low mIoU), EGVOR achieves high scores in both metrics simultaneously (51.3% Overall Acc, 45.0% mIoU). The results further reveal a critical insight: while precise localization is necessary, it is not sufficient for complex reasoning. For instance, in Comparison and Ordering tasks, the performance gap is smaller compared to retrieval tasks. This suggests that while our model excels at grounding (Stage I & I), these higher-order logic tasks require second-order cognitive capabilities beyond pure perception. Nevertheless, by ensuring that the foundational evidence is accurate (high mIoU), EGVOR provides a solid footing for these complex reasoning steps, significantly reducing the error propagation seen in baseline models. Figure 7: Qualitative comparison of attention maps for Qwen2.5-VL-7B and our EGVOR on clear and simulated foggy images. 5.5 Holistic Multimodal Capabilities Model Multimodal Benchmark Performance Vision-centric QA General VQA Doc & Chart CV-2D 72 CV-3D 72 MMVP 73 RWQA 86 MMBench 46 POPE 41 Hallusion 28 AI2D 30 ChartQA 54 Qwen2.5-VL-7B 74.1 72.6 66.7 69.0 83.1 86.7 48.2 84.9 85.6 EGVOR (Ours) 77.9 ↑ 3.8 79.1 ↑ 6.5 77.3 ↑ 10.6 72.1 ↑ 3.1 85.1 ↑ 2.0 87.8↑ 1.1 51.1↑ 2.9 85.4↑ 0.5 85.8↑ 0.2 Qwen2.5-VL-72B 77.7 87.0 66.7 75.7 88.6 84.9 55.2 88.7 89.5 Table 6: Comparison with state-of-the-art alternatives. We categorize the benchmarks into three groups: Vision-centric, General, and Document/Chart understanding. RWQA: RealWorldQA. Hallusion: HallusionBench. To address concerns regarding catastrophic forgetting, we evaluated EGVOR on a suite of 9 external benchmarks (Tab. 6). Results demonstrate that EGVOR not only preserves general knowledge but reinforces the underlying visual cognition, yielding consistent gains across diverse domains. Generalizing Visual-Centric Reasoning. We observe the most substantial improvements in tasks demanding rigorous perception. Notably, EGVOR achieves a +10.6% gain on MMVP (visual puzzles) over the Qwen2.5-VL-7B baseline. This critically confirms that our learned capability—explicit evidence verification—transcends domain boundaries, generalizing from traffic scenes to abstract visual logic. Similarly, gains on CV-3D (+6.5%) and CV-2D (+3.8%) validate that our spatial anchoring mechanism (btb_t) effectively sharpens the model’s fundamental geometry perception. Mitigating Hallucination via Grounding. EGVOR consistently outperforms baselines on hallucination probes like POPE (+1.1%) and HallusionBench (+2.9%). These gains are directly attributed to the CoE paradigm: by enforcing a “Locate-then-Answer” workflow, the model suppresses Blind Extrapolation based on language priors. This grounding mechanism translates to more trustworthy open-ended generation, as evidenced by the +2.0% improvement on MMBench. Surgical Optimization without Forgetting. In text-intensive tasks (AI2D, ChartQA) relying on OCR and logic, EGVOR maintains a slight improvement (+0.5%, +0.2%). This stability serves as a crucial Anti-Forgetfulness proof, indicating that our Cognitive Alignment strategy is surgical—refining visual attention without overwriting pre-trained knowledge representations. While a capacity gap remains compared to the 72B variant due to world knowledge disparities, EGVOR (7B) narrows this gap solely through improved visual grounding, highlighting the method’s efficiency. 5.6 Ablation Study Optimization Objectives Target General Capabilities Model SFT ℛansR_ans Ψspatial _spatial Ξalign _align Λpath _path AD2-Bench V* MME-RW TreeBench (Init) (Acc) (HRGW) (Cap) (Div) Score Acc Acc Acc mIoU Qwen2.5-VL-7B 5 – – – – – 55.0 71.2 42.3 37.0 – + SFT (Reflective) ✓ – – – – 59.6 77.6 49.2 40.9 24.4 Stage I: Cognitive Alignment (Incremental Components) + RL (Answer Prior) ✓ ✓ – – – 61.2 86.4 51.2 40.2 27.7 + w/ Spatial Constraint ✓ ✓ ✓ – – 65.5 90.8 53.9 49.9 43.8 + w/ Latent Align (Ξ ) ✓ ✓ ✓ ✓ – 66.1 91.1 54.6 50.8 44.2 EGVOR (Full) ✓ ✓ ✓ ✓ ✓ 67.7 91.9 55.4 51.3 45.0 Table 7: Component-wise ablation study of EGVOR. We incrementally integrate the optimization objectives to validate their individual contributions. The baseline Qwen2.5-VL-7B results are aligned with standard benchmarks to ensure fair comparison. ℛansR_ans: Basic answer correctness reward. Ψspatial _spatial: HRGW objective for robust spatial anchoring. Ξalign _align: Latent alignment objective via caption consistency. Λpath _path: Cognitive path diversity objective. Spatial Reward Decomposition Target General (TreeBench) Strategy ℛansR_ans Ψrecraw _rec^raw Ψrecw _rec^w Ψprec _prec AD2-Bench Acc mIoU Baseline (Ans Only) ✓ – – – 61.2 40.2 27.7 + Unweighted Recall ✓ ✓ – – 3.4 1.1 77.9 + Weighted Recall ✓ – ✓ – 40.3 34.5 39.5 + Precision Only ✓ – – ✓ 58.2 45.8 20.5 Full HRGW (Ψspatial _spatial) ✓ – ✓ ✓ 65.5 49.9 43.8 Table 8: Ablation of the HRGW Spatial Objective. We dissect the impact of importance weighting and the trade-off between recall and precision. Figure 8: Attention comparison on complex adverse and tiny-object scenes from AD2-Bench (top) and V* (down). Figure 9: Qualitative comparison of chain-of-evidence reasoning between EGVOR and Gemini2.5-Pro on an adverse occluded scene (top) and a tiny-object scene (bottom). We conducted a comprehensive ablation study to dissect the contribution of each component in EGVOR. The results presented in Tab. 7, 8 validate both the hierarchical training framework and the specific design of our spatial reward. Component-wise Effectiveness. We incrementally integrated the optimization objectives to validate the hierarchical training framework. Initially, the Reflective SFT stage serves as the foundation, initializing the reasoning format and yielding a moderate gain of +4.6% on AD2-Bench. However, the relatively low mIoU (24.4%) on TreeBench indicates that at this stage, the model primarily mimics the output style without fully grasping precise localization. Subsequently, introducing the basic answer reward (ℛansR_ans) in the RL stage improves reasoning accuracy (+1.6%), yet the localization capability (mIoU 27.7%) lags behind. This discrepancy suggests that without explicit spatial constraints, the model tends to rely on language priors rather than visual evidence to answer questions. The integration of the Spatial Constraint (Ψspatial _spatial) marks a pivotal turning point. It boosts mIoU significantly to 43.8% (+16.1%) and, crucially, drives the AD2-Bench score to 65.5%. This empirical evidence confirms our core hypothesis: enforcing precise visual grounding is the prerequisite for robust reasoning in adverse scenarios. Finally, the addition of Alignment and Diversity objectives (Ξalign _align and Λpath _path) further refines the policy, reducing hallucinations and encouraging diverse reasoning paths, culminating in the SOTA performance of 67.7%. Dynamics of Spatial Reward Shaping. Designing the spatial reward in RL is non-trivial due to the risk of the policy converging to degenerate strategies. We analyze the necessity of our HRGW objective through three distinct phases of agent behavior. First, we identify a critical failure mode when using a naive Unweighted Recall reward. As shown in Tab. 8, this encourages the agent to maximize coverage regardless of precision, leading to a trivial solution where the model predicts excessively large boxes (mIoU 77.9%) to cover the entire image. While this satisfies the recall metric, it introduces catastrophic background noise, causing the reasoning accuracy to collapse to a mere 1.1%. To counter this, we introduce importance weighting (Ψrecw _rec^w) to penalize the inclusion of irrelevant background areas. This forces the model to abandon the “canvas-covering” strategy and focus on high-value regions. Consequently, the mIoU normalizes to 39.5%, and reasoning accuracy recovers to 34.5%. However, the predicted boxes remain loosely defined, adhering to a conservative policy that prioritizes coverage over tightness. Finally, we achieve optimal performance by introducing the precision term (Ψprec _prec). Acting as a spatial regularization force, it prunes the non-informative margins of the predicted boxes. Although this results in a similar mIoU (43.8%), the reasoning accuracy on TreeBench jumps significantly to 49.9%. This demonstrates that tighter, noise-free visual evidence significantly enhances the signal-to-noise ratio for the subsequent LLM reasoning process, validating the design of our full HRGW objective. Robustness against Visual Disturbance. To evaluate the stability of the reasoning process under environmental shifts, we analyze the Semantic Entropy Gap (Δ ) between clear and adverse weather conditions. A larger gap implies that the model’s confidence is heavily impacted by visual noise (e.g., rain or fog). As shown in Tab. 9, the standard CoT prompting exacerbates this instability (Δ=0.505 =0.505), as the model tends to hallucinate longer, uncertain reasoning chains when visual cues are ambiguous. In contrast, EGVOR minimizes this gap to 0.278. This indicates that our spatial constraints effectively anchor the reasoning process to reliable visual evidence. Even in adverse conditions, the model maintains a confidence level (Hadverse=1.036H_adverse=1.036) comparable to the baseline, but with significantly higher accuracy, demonstrating true cognitive robustness rather than blind confidence. Efficiency and Token Economy. High-performance reasoning typically comes at the cost of increased inference latency (token length). We challenge this assumption by analyzing the Marginal Utility of generated tokens. Tab. 10 reveals a critical insight: standard text-based CoT is inefficient, consuming 3×3× more tokens for a meager 2.5% accuracy gain (Utility: 1.1). While our SFT stage improves utility to 5.0, the RL stage achieves a breakthrough. Surprisingly, the RL-optimized model not only achieves the highest accuracy (67.7%) on AD2-Bench, but does so with significantly fewer tokens than the SFT model (144144 vs. 212212). The results in an exceptional Marginal Utility of 51.7. This phenomenon suggests that the RL objective, driven by the HRGW reward, forces the model to abandon verbose, irrelevant descriptions and focus solely on high-density visual evidence. The model learns to be concise yet precise, making it highly suitable for latency-sensitive autonomous driving applications. Visualization of Evidence-Guided Attention To better understand how EGVOR behaves under visual degradation, we visualize the cross-modal attention maps of Qwen2.5-VL-7B and our model on clear images and their simulated foggy counterparts (Fig. 7). On clear inputs, both models can roughly locate salient regions (Question: the racket in the tennis scene, or the donut in the grocery scene), but Qwen often spreads attention over background areas, whereas EGVOR produces more concentrated hotspots on the entities required to answer the question. After adding fog-like noise, Qwen’s attention becomes scattered and drifts to irrelevant objects, consistent with the “visual probability collapse” in Sec. 4.1, while EGVOR preserves tight, task-aligned focus on the same critical objects. Beyond synthetic perturbations, we visualize attention on real complex scenes (AD2-Bench) and tiny objects (V*), as shown in Fig. 8. In the AD2-Bench rainy vehicle counting task, Qwen erroneously attends to irrelevant pedestrians, succumbing to semantic distraction. In contrast, EGVOR precisely anchors onto target vehicles, capturing even those heavily occluded by splashes or poles. Similarly, on V*, while Qwen disperses attention over the background, EGVOR locks onto the queried tiny cyclist. These visualizations confirm that explicit evidence atoms stabilize attention against distractions and occlusion, demonstrating robust generalization. Case Studies Fig. 9 compares EGVOR with Gemini2.5-Pro on two challenging examples. In the rainy street scene, EGVOR enumerates all vehicles by explicitly cropping six car instances and listing their coordinates, then concludes that there are six vehicles in total. In the tiny-object example, it first localizes the red and white balloons among dozens of cyclists, and then describes their relative positions based on the detected boxes, yielding a correct spatial relation. Gemini2.5-Pro, in contrast, generates fluent natural-language explanations but without explicit localization. In the first case, it lists several vehicles yet still undercounts them, predicting five instead of six. In the second case, it misdescribes the relative position of the balloons. These case studies illustrate that, under adverse and small-object scenarios, ungrounded explanations can remain plausible but incorrect, while our explicit Chain-of-Evidence formulation makes the reasoning process verifiable and more reliable. Model HclearH_clear HadverseH_adverse Gap (Δ ) ↓ Qwen2.5-VL-7B 0.5854 1.0270 0.4416 Qwen2.5-VL-7B + CoT Prompt 1.3815 1.8863 0.5049 EGVOR (SFT) 1.0348 1.3966 0.3618 EGVOR (RL) 0.7580 1.0355 0.2776 Table 9: Comparison of Semantic Entropy across different models. HclearH_clear and HadverseH_adverse denote the average token entropy under clear and adverse weather conditions, respectively. The Entropy Gap (Δ ) serves as a metric for robustness, where a lower value indicates higher stability against visual disturbances. Model Paradigm AD2-Score ↑ Avg. Tokens Marginal Utility ↑ (%) (T) (ΔAcc/Δ×100 / × 100) Qwen2.5-VL-7B Zero-Shot 55.0 120 – + CoT Prompt Text-Only CoT 57.5 345 1.1 EGVOR (SFT) Visual Evidence 59.6 212 5.0 EGVOR (RL) Visual Evidence 67.7 144 51.7 Table 10: Performance-to-Cost Efficiency Analysis. We compare the trade-off between reasoning performance and computational overhead. 6 Conclusion and Future Work This work addresses the systemic “dual opacity”—stemming from implicit black-box reasoning and outcome-oriented evaluation—by introducing AD2-Bench and EGVOR. Serving as a diagnostic prism, our benchmark identifies Spatial Ambiguity and Semantic Uncertainty as the primary bottlenecks of reasoning collapse. In response, EGVOR shifts the paradigm from implicit inference to the explicit construction of Evidence Atoms. By enforcing strict spatial-semantic alignment, our framework demonstrates that robust multimodal cognition is predicated on the verifiable acquisition of visual evidence, paving the way for trustworthy AI in complex real-world scenarios. Future Work. While EGVOR establishes a solid foundation for grounded reasoning, two promising directions remain for exploration: (1) Spatiotemporal Evidence Chains: Currently, EGVOR focuses on frame-level reasoning. Future work will extend “Evidence Atoms” into “Evidence Tubes” to capture temporal dynamics. By tracking evidence consistency across video frames, we aim to resolve high-entropy scenarios involving motion ambiguity (e.g., predicting the trajectory of an occluded cyclist). (2) Closed-loop Policy Integration: We plan to integrate the explainable Chain of Evidence directly into end-to-end autonomous driving systems. Instead of merely outputting textual decisions, the explicit evidence representations (zevidencez_evidence) can serve as interpretable intermediate features to modulate control signals, enhancing the safety and transparency of unmanned systems. Data Availability Statements. We use the eleven publicly available datasets in our paper: V* 83, HRBench 79, MME-RealWorld 97, TreeBench 74, our benchmark: AD2-Bench, CV-Bench 72, MMVP 73, RealworldQA 86, MMBench 46, POPE 41, HallusionBench 28,AI2D 30, ChartQA 54. All benchmarks can be found on huggingface: https://huggingface.co/datasets/lmms-lab/vstar-bench,DreamMr/HR-Bench,yifanzhang114/MME-RealWorld-Lite,HaochenWang/TreeBench,ZhaoyangWei/AD2-Bench,ZzzHelloWorld/CV-Bench,MMVP/MMVP,lmms-lab/RealWorldQA,lmms-lab/MMBench,lmms-lab/POPE,lmms-lab/HallusionBench,lmms-lab/ai2d,lmms-lab/ChartQA References Alayrac et al. (2022) J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. NeurIPS 35, p. 23716–23736. Cited by: §2. Antol et al. (2015) S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh Vqa: visual question answering. In Proceedings of the IEEE international conference on computer vision, p. 2425–2433. Cited by: §2. Bai et al. (2023) J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou Qwen-vl: a versatile vision-language model for understanding, localization, text reading, and beyond. arXiv preprint arXiv:2308.12966. Cited by: §1. Bai et al. (2025a) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923. Cited by: §2, §5.1. Bai et al. (2025b) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, et al. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923. Cited by: §2, Table 7, §7.7. Caesar et al. (2020) H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom Nuscenes: a multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 11621–11631. Cited by: §2, §3.1, §5.2, 2nd item. Cao et al. (2025) M. Cao, H. Zhao, C. Zhang, X. Chang, I. Reid, and X. Liang Ground-r1: incentivizing grounded visual reasoning via reinforcement learning. arXiv preprint arXiv:2505.20272. Cited by: §2. Chen et al. (2024a) B. Chen, Z. Xu, S. Kirmani, B. Ichter, D. Sadigh, L. Guibas, and F. Xia Spatialvlm: endowing vision-language models with spatial reasoning capabilities. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14455–14465. Cited by: §2. Chen et al. (2025a) K. Chen, Y. Li, W. Zhang, Y. Liu, P. Li, R. Gao, L. Hong, M. Tian, X. Zhao, Z. Li, et al. Automated evaluation of large vision-language models on self-driving corner cases. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 7817–7826. Cited by: §2. Chen et al. (2024b) K. Chen, R. Thapa, R. Chalamala, B. Athiwaratkun, S. L. Song, and J. Zou Dragonfly: multi-resolution zoom supercharges large visual-language model. arXiv e-prints, p. arXiv–2406. Cited by: §2. Chen et al. (2025b) L. Chen, L. Li, H. Zhao, and Y. Song R1-v: reinforcing super generalization ability in vision-language models with less than $3. Cited by: §2. Chen et al. (2025c) X. Chen, Z. Wu, X. Liu, Z. Pan, W. Liu, Z. Xie, X. Yu, and C. Ruan Janus-pro: unified multimodal understanding and generation with data and model scaling. arXiv preprint arXiv:2501.17811. Cited by: §5.1. Chen et al. (2024c) Y. Chen, Y. Wang, and Z. Zhang DrivingGPT: unifying driving world modeling and planning with multi-modal autoregressive transformers. arXiv preprint arXiv:2412.18607. Cited by: §2. Chen et al. (2024d) Z. Chen, W. Wang, Y. Cao, Y. Liu, Z. Gao, E. Cui, J. Zhu, S. Ye, H. Tian, Z. Liu, et al. Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling. arXiv preprint arXiv:2412.05271. Cited by: §5.1. Chen et al. (2024e) Z. Chen, W. Wang, H. Tian, S. Ye, Z. Gao, E. Cui, W. Tong, K. Hu, J. Luo, Z. Ma, et al. How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites. Science China Information Sciences 67 (12), p. 220101. Cited by: §1, §2. Chen et al. (2024f) Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, et al. Internvl: scaling up vision foundation models and aligning for generic visual-linguistic tasks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 24185–24198. Cited by: §5.1. Cui et al. (2024) Y. Cui, S. Huang, J. Zhong, Z. Liu, Y. Wang, C. Sun, B. Li, X. Wang, and A. Khajepour DriveLLM: charting the path toward full autonomous driving with large language models. IEEE Transactions on Intelligent Vehicles 9 (1), p. 1450–1464. External Links: Document Cited by: §2. DeepMind (2025a) G. DeepMind Gemini-2.5-flash. Note: https://deepmind.google/models/gemini/flash/ Cited by: §5.1. DeepMind (2025b) G. DeepMind Gemini-2.5-pro. Note: https://deepmind.google/models/gemini/pro/ Cited by: §5.2. Dong et al. (2025) W. Dong, H. Zhu, S. Lin, X. Luo, Y. Shen, G. Guo, and B. Zhang Fusion-mamba for cross-modality object detection. IEEE Transactions on Multimedia 27, p. 7392–7406. External Links: Document Cited by: §2. Dong et al. (2024) Y. Dong, Z. Liu, H. Sun, J. Yang, W. Hu, Y. Rao, and Z. Liu Insight-v: exploring long-chain visual reasoning with multimodal large language models. arXiv preprint arXiv:2411.14432. Cited by: §2. Duan et al. (2024) H. Duan, J. Yang, Y. Qiao, X. Fang, L. Chen, Y. Liu, X. Dong, Y. Zang, P. Zhang, J. Wang, et al. Vlmevalkit: an open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 11198–11201. Cited by: Table 5. Fan et al. (2025) Y. Fan, X. He, D. Yang, K. Zheng, C. Kuo, Y. Zheng, S. J. Narayanaraju, X. Guan, and X. E. Wang GRIT: teaching mllms to think with images. arXiv preprint arXiv:2505.15879. Cited by: §2. Fu et al. (2025a) C. Fu, P. Chen, Y. Shen, Y. Qin, M. Zhang, X. Lin, J. Yang, X. Zheng, K. Li, X. Sun, et al. Mme: a comprehensive evaluation benchmark for multimodal large language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, Cited by: §1. Fu et al. (2024a) C. Fu, Y. Zhang, S. Yin, B. Li, X. Fang, S. Zhao, H. Duan, X. Sun, Z. Liu, L. Wang, et al. Mme-survey: a comprehensive survey on evaluation of multimodal llms. arXiv preprint arXiv:2411.15296. Cited by: §2. Fu et al. (2024b) D. Fu, X. Li, L. Wen, M. Dou, P. Cai, B. Shi, and Y. Qiao Drive like a human: rethinking autonomous driving with large language models. In IEEE/CVF Winter Conference on Applications of Computer Vision, p. 910–919. Cited by: §2. Fu et al. (2025b) X. Fu, M. Liu, Z. Yang, J. Corring, Y. Lu, J. Yang, D. Roth, D. Florencio, and C. Zhang ReFocus: visual editing as a chain of thought for structured image understanding. arXiv preprint arXiv:2501.05452. Cited by: §2. Guan et al. (2024) T. Guan, F. Liu, X. Wu, R. Xian, Z. Li, X. Liu, X. Wang, L. Chen, F. Huang, Y. Yacoob, et al. HallusionBench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models. In CVPR, Cited by: §2, Table 6, §6. Guo et al. (2025) D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, et al. Deepseek-r1: incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948. Cited by: §2. Hiippala et al. (2021) T. Hiippala, M. Alikhani, J. Haverinen, T. Kalliokoski, E. Logacheva, S. Orekhova, A. Tuomainen, M. Stone, and J. A. Bateman AI2D-rst: a multimodal corpus of 1000 primary school science diagrams. Language Resources and Evaluation 55, p. 661–688. Cited by: Table 6, §6. Hong et al. (2024) W. Hong, W. Wang, Q. Lv, J. Xu, W. Yu, J. Ji, Y. Wang, Z. Wang, Y. Dong, M. Ding, et al. Cogagent: a visual language model for gui agents. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14281–14290. Cited by: §2. Huang et al. (2025) W. Huang, B. Jia, Z. Zhai, S. Cao, Z. Ye, F. Zhao, Z. Xu, Y. Hu, and S. Lin Vision-r1: incentivizing reasoning capability in multimodal large language models. arXiv preprint arXiv:2503.06749. Cited by: §2. Inoue et al. (2024) Y. Inoue, Y. Yada, K. Tanahashi, and Y. Yamaguchi Nuscenes-mqa: integrated evaluation of captions and qa for autonomous driving datasets using markup annotations. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 930–938. Cited by: §2. Jia et al. (2025) H. Jia, Y. Li, and H. Li A multi-criteria game-driven constrained multi-objective evolutionary algorithm. IEEE Transactions on Evolutionary Computation. Cited by: §2. Kenk and Hassaballah (2020) M. A. Kenk and M. Hassaballah DAWN: vehicle detection in adverse weather nature dataset. arXiv preprint arXiv:2008.05402. Cited by: §3.1. Kim et al. (2018) J. Kim, A. Rohrbach, T. Darrell, J. Canny, and Z. Akata Textual explanations for self-driving vehicles. In ECCV, Cited by: §2. Li et al. (2024a) B. Li, Y. Ge, Y. Ge, G. Wang, R. Wang, R. Zhang, and Y. Shan Seed-bench: benchmarking multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13299–13308. Cited by: §2. Li et al. (2024b) F. Li, R. Zhang, H. Zhang, Y. Zhang, B. Li, W. Li, Z. Ma, and C. Li Llava-next-interleave: tackling multi-image, video, and 3d in large multimodal models. arXiv preprint arXiv:2407.07895. Cited by: §5.1. Li et al. (2023a) J. Li, D. Li, S. Savarese, and S. Hoi BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In International Conference on Machine Learning (ICML), External Links: 2301.12597 Cited by: §2. Li et al. (2022) K. Li, K. Chen, H. Wang, L. Hong, C. Ye, J. Han, Y. Chen, W. Zhang, C. Xu, D. Yeung, et al. Coda: a real-world road corner case dataset for object detection in autonomous driving. In European Conference on Computer Vision, p. 406–423. Cited by: §2, §3.1. Li et al. (2023b) Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J. Wen Evaluating object hallucination in large vision-language models. arXiv preprint arXiv:2305.10355. Cited by: §2, Table 6, §6. Lin et al. (2014) T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick Microsoft coco: common objects in context. In ECCV, p. 740–755. Cited by: §2, §2. Liu et al. (2024a) H. Liu, C. Li, Y. Li, and Y. J. Lee Improved baselines with visual instruction tuning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 26296–26306. Cited by: §2, §5.1. Liu et al. (2024b) H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee LLaVA-next: improved reasoning, ocr, and world knowledge. Note: https://llava-vl.github.io/blog/2024-01-30-llava-next/ Cited by: §2, §2, §5.1. Liu et al. (2023a) H. Liu, C. Li, Q. Wu, and Y. J. Lee Visual instruction tuning. NeurIPS 36, p. 34892–34916. Cited by: §2. Liu et al. (2023b) Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al. MMBench: is your multi-modal model an all-around player?. arXiv preprint arXiv:2307.06281. Cited by: Table 6, §6. Liu et al. (2024c) Y. Liu, H. Duan, Y. Zhang, B. Li, S. Zhang, W. Zhao, Y. Yuan, J. Wang, C. He, Z. Liu, et al. Mmbench: is your multi-modal model an all-around player?. In European conference on computer vision, p. 216–233. Cited by: §1, §2. Liu et al. (2025a) Y. Liu, T. Qu, Z. Zhong, B. Peng, S. Liu, B. Yu, and J. Jia VisionReasoner: unified visual perception and reasoning via reinforcement learning. arXiv preprint arXiv:2505.12081. Cited by: §2. Liu et al. (2025b) Z. Liu, Z. Sun, Y. Zang, X. Dong, Y. Cao, H. Duan, D. Lin, and J. Wang Visual-rft: visual reinforcement fine-tuning. arXiv preprint arXiv:2503.01785. Cited by: §2. Liu et al. (2024d) Z. Liu, Y. Dong, Y. Rao, J. Zhou, and J. Lu Chain-of-spot: interactive reasoning improves large vision-language models. arXiv preprint arXiv:2403.12966. Cited by: §2. Liu et al. (2024e) Z. Liu, Y. Dong, Y. Rao, J. Zhou, and J. Lu Chain-of-spot: interactive reasoning improves large vision-language models. arXiv preprint arXiv:2403.12966. Cited by: §2. Lu et al. (2023) P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao Mathvista: evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255. Cited by: §2. Marcu et al. (2024) A. Marcu, L. Chen, J. Hünermann, A. Karnsund, B. Hanotte, P. Chidananda, S. Nair, V. Badrinarayanan, A. Kendall, J. Shotton, et al. LingoQA: visual question answering for autonomous driving. In European Conference on Computer Vision, p. 252–269. Cited by: §2. Masry et al. (2022) A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque Chartqa: a benchmark for question answering about charts with visual and logical reasoning. arXiv preprint arXiv:2203.10244. Cited by: Table 6, §6. Meng et al. (2025) F. Meng, L. Du, Z. Liu, Z. Zhou, Q. Lu, D. Fu, T. Han, B. Shi, W. Wang, J. He, et al. M-eureka: exploring the frontiers of multimodal reasoning with rule-based reinforcement learning. arXiv preprint arXiv:2503.07365. Cited by: §2. Mondal et al. (2024) D. Mondal, S. Modi, S. Panda, R. Singh, and G. S. Rao Kam-cot: knowledge augmented multimodal chain-of-thoughts reasoning. In AAAI, p. 18798–18806. Cited by: §2. OpenAI (2024a) OpenAI OpenAI-gpt-4o. Note: https://openai.com/index/gpt-4o-system-card/ Cited by: §5.1. OpenAI (2024b) OpenAI OpenAI-o1. Note: https://openai.com/o1/ Cited by: §2. OpenAI (2025) OpenAI OpenAI-o3. Note: https://openai.com/index/introducing-o3-and-o4-mini/ Cited by: §2, Table 5. Peng et al. (2025) Y. Peng, G. Zhang, M. Zhang, Z. You, J. Liu, Q. Zhu, K. Yang, X. Xu, X. Geng, and X. Yang Lmm-r1: empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl. arXiv preprint arXiv:2503.07536. Cited by: §2. Qian et al. (2024) T. Qian, J. Chen, L. Zhuo, Y. Jiao, and Y. Jiang Nuscenes-qa: a multi-modal visual question answering benchmark for autonomous driving scenario. In Proceedings of the AAAI Conference on Artificial Intelligence, p. 4542–4550. Cited by: §2. Qiao et al. (2025a) J. Qiao, W. Li, H. Xie, H. Chen, J. Hu, S. Lin, and J. Han LIPT: latency-aware image processing transformer. IEEE Transactions on Image Processing 34, p. 3056–3069. External Links: Document Cited by: §2, §2. Qiao et al. (2025b) J. Qiao, J. Liao, W. Li, Y. Zhang, Y. Guo, Y. Wen, Z. Qiu, J. Xie, J. Hu, and S. Lin Hi-mamba: hierarchical mamba for efficient image super-resolution. IEEE Transactions on Image Processing 34, p. 8461–8473. Cited by: §2. Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. In ICML, p. 8748–8763. Cited by: §2. Sakaridis et al. (2021) C. Sakaridis, D. Dai, and L. Van Gool ACDC: the adverse conditions dataset with correspondences for semantic driving scene understanding. In Proceedings of the IEEE/CVF international conference on computer vision, p. 10765–10775. Cited by: §3.1. Shao et al. (2024) H. Shao, S. Qian, H. Xiao, G. Song, Z. Zong, L. Wang, Y. Liu, and H. Li Visual-cot: advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning. Advances in Neural Information Processing Systems 37, p. 8612–8642. Cited by: §2. Shen et al. (2025) H. Shen, P. Liu, J. Li, C. Fang, Y. Ma, J. Liao, Q. Shen, Z. Zhang, K. Zhao, Q. Zhang, et al. Vlm-r1: a stable and generalizable r1-style large vision-language model. arXiv preprint arXiv:2504.07615. Cited by: §2. Sima et al. (2024) C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li Drivelm: driving with graph visual question answering. In European Conference on Computer Vision, p. 256–274. Cited by: §2. Stone et al. (2023) A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K. Lee, Q. Vuong, P. Wohlhart, S. Kirmani, B. Zitkovich, F. Xia, et al. Open-world object manipulation using pre-trained vision-language models. arXiv preprint arXiv:2303.00905. Cited by: §2. Su et al. (2025) A. Su, H. Wang, W. Ren, F. Lin, and W. Chen Pixel reasoner: incentivizing pixel-space reasoning with curiosity-driven reinforcement learning. arXiv preprint arXiv:2505.15966. Cited by: §2, §5.1, §5.3. Tian et al. (2024) X. Tian, J. Gu, B. Li, Y. Liu, Y. Wang, Z. Zhao, K. Zhan, P. Jia, X. Lang, and H. Zhao Drivevlm: the convergence of autonomous driving and large vision-language models. arXiv preprint arXiv:2402.12289. Cited by: §2. Tong et al. (2024a) P. Tong, E. Brown, P. Wu, S. Woo, A. J. V. IYER, S. C. Akula, S. Yang, J. Yang, M. Middepogu, Z. Wang, et al. Cambrian-1: a fully open, vision-centric exploration of multimodal llms. NeurIPS 37, p. 87310–87356. Cited by: Table 6, Table 6, §6. Tong et al. (2024b) S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie Eyes wide shut? exploring the visual shortcomings of multimodal llms. In CVPR, p. 9568–9578. Cited by: Table 6, §6. Wang et al. (2025a) H. Wang, X. Li, Z. Huang, A. Wang, J. Wang, T. Zhang, J. Zheng, S. Bai, Z. Kang, J. Feng, et al. Traceable evidence enhanced visual grounded reasoning: evaluation and methodology. arXiv preprint arXiv:2507.07999. Cited by: §5.4, Table 5, §6. Wang et al. (2025b) J. Wang, Z. Kang, H. Wang, H. Jiang, J. Li, B. Wu, Y. Wang, J. Ran, X. Liang, C. Feng, et al. VGR: visual grounded reasoning. arXiv preprint arXiv:2506.11991. Cited by: §2, §5.2, §7.7. Wang et al. (2024a) J. Wang, B. Wu, H. Jiang, Z. Xun, X. Xiao, H. Guo, and J. Xiao World to code: multi-modal data generation via self-instructed compositional captioning and filtering. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 4608–4623. Cited by: §2. Wang et al. (2024b) P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al. Qwen2-vl: enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191. Cited by: §2, §5.1. Wang et al. (2024c) P. Wang, S. Bai, S. Tan, S. Wang, Z. Fan, J. Bai, K. Chen, X. Liu, J. Wang, W. Ge, et al. Qwen2-vl: enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191. Cited by: §1, §2, §2. Wang et al. (2025c) W. Wang, L. Ding, M. Zeng, X. Zhou, L. Shen, Y. Luo, W. Yu, and D. Tao Divide, conquer and combine: a training-free framework for high-resolution image perception in multimodal large language models. In AAAI, p. 7907–7915. Cited by: §2, Table 3, §6. Wei et al. (2025a) L. Wei, Y. Li, C. Wang, Y. Wang, L. Kong, W. Huang, and L. Sun Unsupervised post-training for multi-modal llm reasoning via grpo. arXiv preprint arXiv:2505.22453. Cited by: §2. Wei et al. (2025b) L. Wei, Y. Li, K. Zheng, C. Wang, Y. Wang, L. Kong, L. Sun, and W. Huang Advancing multimodal reasoning via reinforcement learning with cold start. arXiv preprint arXiv:2505.22334. Cited by: §2. Wu et al. (2023) D. Wu, W. Han, T. Wang, Y. Liu, X. Zhang, and J. Shen Language prompt for autonomous driving. arXiv preprint arXiv:2309.04379. Cited by: §2. Wu and Xie (2024a) P. Wu and S. Xie V*: guided visual search as a core mechanism in multimodal llms. In CVPR, p. 13084–13094. Cited by: §2, §5.2, Table 3, §6, §7.8. Wu and Xie (2024b) P. Wu and S. Xie V?: guided visual search as a core mechanism in multimodal llms. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13084–13094. Cited by: §2. Wu et al. (2024) Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, et al. Deepseek-vl2: mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302. Cited by: §2. xAI (2024) xAI Grok. . Cited by: Table 6, §6. Xing et al. (2024) S. Xing, H. Hua, X. Gao, S. Zhu, R. Li, K. Tian, X. Li, H. Huang, T. Yang, Z. Wang, et al. AutoTrust: benchmarking trustworthiness in large vision language models for autonomous driving. arXiv preprint arXiv:2412.15206. Cited by: §2. Xu et al. (2024) Y. Xu, Y. Hu, Z. Zhang, G. P. Meyer, S. K. Mustikovela, S. Srinivasa, E. M. Wolff, and X. Huang VLM-ad: end-to-end autonomous driving through vision-language model supervision. arXiv preprint arXiv:2412.14446. Cited by: §2. Xu et al. (2020) Y. Xu, X. Yang, L. Gong, H. Lin, T. Wu, Y. Li, and N. Vasconcelos Explainable object-induced action decision for autonomous vehicles. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 9523–9532. Cited by: §2. Yang et al. (2025) J. Yang, Y. Dong, S. Liu, B. Li, Z. Wang, H. Tan, C. Jiang, J. Kang, Y. Zhang, K. Zhou, et al. Octopus: embodied vision-language programmer from environmental feedback. In European Conference on Computer Vision, p. 20–38. Cited by: §2. Yao et al. (2024) Y. Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, et al. Minicpm-v: a gpt-4v level mllm on your phone. arXiv preprint arXiv:2408.01800. Cited by: §5.1. Ye et al. (2024) J. Ye, H. Xu, H. Liu, A. Hu, M. Yan, Q. Qian, J. Zhang, F. Huang, and J. Zhou Mplug-owl3: towards long image-sequence understanding in multi-modal large language models. arXiv preprint arXiv:2408.04840. Cited by: §2. Yu et al. (2020) F. Yu, H. Chen, X. Wang, W. Xian, Y. Chen, F. Liu, V. Madhavan, and T. Darrell Bdd100k: a diverse driving dataset for heterogeneous multitask learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 2636–2645. Cited by: §2, §3.1. Yu et al. (2023) W. Yu, Z. Yang, L. Li, J. Wang, K. Lin, Z. Liu, X. Wang, and L. Wang Mm-vet: evaluating large multimodal models for integrated capabilities. arXiv preprint arXiv:2308.02490. Cited by: §2. Yue et al. (2024) X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al. Mmmu: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In CVPR, Cited by: §2. Zhang et al. (2019) X. Zhang, J. Zou, K. He, and J. Sun Holistic cnn compression via low-rank decomposition with knowledge transfer. IEEE Transactions on Pattern Analysis and Machine Intelligence 41 (12), p. 2889–2905. External Links: Document Cited by: §2. Zhang et al. (2024) Y. Zhang, H. Zhang, H. Tian, C. Fu, S. Zhang, J. Wu, F. Li, K. Wang, Q. Wen, Z. Zhang, et al. MME-realworld: could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?. arXiv preprint arXiv:2408.13257. Cited by: §1, §2, Table 4, §6. Zhao et al. (2025) Y. Zhao, J. Huang, J. Hu, X. Wang, Y. Mao, D. Zhang, Z. Jiang, Z. Wu, B. Ai, A. Wang, et al. Swift: a scalable lightweight infrastructure for fine-tuning. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 29733–29735. Cited by: §5.1. Zheng et al. (2025a) Y. Zheng, J. Lu, S. Wang, Z. Feng, D. Kuang, and Y. Xiong EasyR1: an efficient, scalable, multi-modality rl training framework. Note: https://github.com/hiyouga/EasyR1 Cited by: §5.1. Zheng et al. (2025b) Z. Zheng, M. Yang, J. Hong, C. Zhao, G. Xu, L. Yang, C. Shen, and X. Yu DeepEyes: incentivizing” thinking with images” via reinforcement learning. arXiv preprint arXiv:2505.14362. Cited by: §2, §5.1. Zhu et al. (2025) J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, et al. Internvl3: exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479. Cited by: §2, §2, §5.1, Table 5. Zhu et al. (2021) P. Zhu, L. Wen, D. Du, X. Bian, H. Fan, Q. Hu, and H. Ling Detection and tracking meet drones challenge. IEEE TPAMI 44 (11), p. 7380–7399. Cited by: §5.2, 1st item. 7 Appendix 7.1 Rigorous Derivations for Probabilistic Bottleneck Analysis Table 11: Core notation used in Sec. 4. The table summarizes the physical meaning of the main probabilistic variables to improve readability. Symbol Meaning I~ I Observed image under adverse degradation. Q, A, AgtA_gt Query, predicted answer, and visually grounded correct answer. O Query-dependent set of candidate visual entities in the scene. OfO_f, o∈o Latent visual-focus random variable, and one realization of this focus. critO_crit, irrO_irr, bgO_bg Critical objects, irrelevant distractors, and background regions. zoz_o, Z Deterministic feature of entity o, and random latent representation used in variance analysis. zglobalz_global, ZglobalZ_global Global context feature and its corresponding random variable. E, e Structured evidence set and one evidence element. at=⟨bt,ht,ct⟩a_t= b_t,h_t,c_t Evidence atom: spatial anchor, region-aware latent state, and grounded description. Ψspatial _spatial, Ξalign _align, Λpath _path Spatial reward, semantic alignment reward, and path-diversity reward. This section provides the detailed derivations supporting the probabilistic analysis in Sec. 4.1. The goal is not to claim an exact generative model of MLLMs, but to make explicit the assumptions under which the diagnostic statements in the main text hold. 7.1.1 Latent Focus Variable and Object-Centric Mixture Let O be the query-dependent set of candidate visual entities. We define OfO_f as a discrete latent visual-focus random variable with support O. A lowercase o∈o denotes a realization of OfO_f. We write P(o|I~,Q)≡P(Of=o|I~,Q) P(o| I,Q)≡ P(O_f=o| I,Q) (47) for notational simplicity. Assuming that zoz_o is the deterministic latent representation associated with entity o under the observed image I~ I, the object-centric posterior approximation is P(A|I~,Q)=∑o∈P(A|zo,Q)P(o|I~,Q). P(A| I,Q)= _o P(A|z_o,Q)P(o| I,Q). (48) This approximation abstracts the model’s implicit visual attention as a distribution over candidate entities. 7.1.2 Posterior Suppression under Critical-Evidence Mass Shift Let AgtA_gt denote the visually grounded correct answer. We partition the candidate entity set into critical and non-critical subsets: =crit∪non,non=irr∪bg. =O_crit _non, 9.24994ptO_non=O_irr _bg. (49) The critical-evidence mass is defined as ρ=P(Of∈crit|I~,Q). ρ=P(O_f _crit| I,Q). (50) For a local redistribution analysis, we decompose the focus distribution into two within-set distributions: qcrit(o)=P(Of=o|I~,Q,Of∈crit),o∈crit, q_crit(o)=P(O_f=o| I,Q,O_f _crit), 9.24994pto _crit, (51) and qnon(o)=P(Of=o|I~,Q,Of∈non),o∈non. q_non(o)=P(O_f=o| I,Q,O_f _non), 9.24994pto _non. (52) For 0<ρ<10<ρ<1, define a family of focus distributions: Pρ(Of=o|I~,Q)=ρqcrit(o),o∈crit,(1−ρ)qnon(o),o∈non. P_ρ(O_f=o| I,Q)= casesρ q_crit(o),&o _crit,\\ (1-ρ)q_non(o),&o _non. cases (53) This family preserves the relative distribution within the critical and non-critical subsets, while varying the total probability mass assigned to each subset. Define the average support for AgtA_gt from critical and non-critical regions as p¯crit=∑o∈critP(Agt|zo,Q)qcrit(o), p_crit= _o _critP(A_gt|z_o,Q)q_crit(o), (54) and p¯non=∑o∈nonP(Agt|zo,Q)qnon(o). p_non= _o _nonP(A_gt|z_o,Q)q_non(o). (55) The critical-evidence separability assumption is p¯crit>p¯non. p_crit> p_non. (56) This assumption states that query-critical evidence supports the visually grounded answer more strongly than non-critical evidence. Under Eq. 53, the posterior support for AgtA_gt becomes Pρ(Agt|I~,Q)=ρp¯crit+(1−ρ)p¯non. P_ρ(A_gt| I,Q)=ρ p_crit+(1-ρ) p_non. (57) For two critical-evidence masses 0<ρ2<ρ1<10< _2< _1<1, subtracting Eq. 57 gives Pρ2(Agt|I~,Q)−Pρ1(Agt|I~,Q) P_ _2(A_gt| I,Q)-P_ _1(A_gt| I,Q) (58) =(ρ2−ρ1)(p¯crit−p¯non). =( _2- _1)( p_crit- p_non). Since ρ2−ρ1<0 _2- _1<0 and p¯crit−p¯non>0 p_crit- p_non>0, we obtain Pρ2(Agt|I~,Q)<Pρ1(Agt|I~,Q). P_ _2(A_gt| I,Q)<P_ _1(A_gt| I,Q). (59) Thus, reducing the critical-evidence mass strictly suppresses the posterior probability of the visually grounded answer under the separability assumption. 7.1.3 Semantic Entropy under Critical-Evidence Mass Shift Let A denote the answer space. Define the answer distributions induced by critical and non-critical evidence as pcrit(a)=P(A=a|Of∈crit,I~,Q), p_crit(a)=P(A=a|O_f _crit, I,Q), (60) and pnon(a)=P(A=a|Of∈non,I~,Q). p_non(a)=P(A=a|O_f _non, I,Q). (61) The answer distribution induced by critical-evidence mass ρ is pρ(a)=ρpcrit(a)+(1−ρ)pnon(a). p_ρ(a)=ρ p_crit(a)+(1-ρ)p_non(a). (62) We define Hρ(A|I~,Q)≡−∑a∈pρ(a)logpρ(a). H_ρ(A| I,Q)≡- _a p_ρ(a) p_ρ(a). (63) Assuming pρ(a)>0p_ρ(a)>0 on the considered local interval, differentiating Eq. 63 with respect to ρ gives dHρ(A|I~,Q)dρ dH_ρ(A| I,Q)dρ =−∑a∈dpρ(a)dρ(1+logpρ(a)) =- _a dp_ρ(a)dρ (1+ p_ρ(a) ) (64) =−∑a∈(pcrit(a)−pnon(a))(1+logpρ(a)). =- _a (p_crit(a)-p_non(a) ) (1+ p_ρ(a) ). Because both pcritp_crit and pnonp_non are probability distributions, ∑a∈(pcrit(a)−pnon(a))=1−1=0. _a (p_crit(a)-p_non(a) )=1-1=0. (65) Therefore, the constant term cancels, yielding dHρ(A|I~,Q)dρ=−∑a∈(pcrit(a)−pnon(a))logpρ(a). dH_ρ(A| I,Q)dρ=- _a (p_crit(a)-p_non(a) ) p_ρ(a). (66) The uncertainty-dominant failure regime is defined as a local interval [ρ2,ρ1][ _2, _1] where dHρ(A|I~,Q)dρ<0,∀ρ∈[ρ2,ρ1]. dH_ρ(A| I,Q)dρ<0, 9.24994pt∀ρ∈[ _2, _1]. (67) For 0<ρ2<ρ1<10< _2< _1<1, the fundamental theorem of calculus gives Hρ2(A|I~,Q)−Hρ1(A|I~,Q) H_ _2(A| I,Q)-H_ _1(A| I,Q) =−∫ρ2ρ1dHu(A|I~,Q)dudu. =- _ _2 _1 dH_u(A| I,Q)du\,du. (68) Under Eq. 67, the integral is negative, and hence Hρ2(A|I~,Q)>Hρ1(A|I~,Q). H_ _2(A| I,Q)>H_ _1(A| I,Q). (69) This establishes a conditional statement for semantic entropy growth: when non-critical regions induce diverse inconsistent answer hypotheses, reducing the critical-evidence mass increases the answer entropy. This result is intentionally local and conditional. In the prior-dominant hallucination regime, the model may instead become confidently wrong, in which case the entropy need not increase. That regime is captured by the prior-collapse inequality in Sec. 4.1. 7.1.4 Conditional-Independence Form of Prior Collapse Let ZglobalZ_global denote the random global visual representation and zglobalz_global one of its realizations. In the exact noise-dominant limiting case, assume Zglobal⟂⟂A|Q. Z_global \!\!\! A Q. (70) Then, for all zglobalz_global with positive density, P(zglobal|A,Q)=P(zglobal|Q). P(z_global|A,Q)=P(z_global|Q). (71) Applying Bayes’ rule, P(A|zglobal,Q) P(A|z_global,Q) =P(zglobal|A,Q)P(A|Q)P(zglobal|Q) = P(z_global|A,Q)P(A|Q)P(z_global|Q) (72) =P(A|Q). =P(A|Q). Thus, in the exact conditional-independence limit, the posterior collapses to the language prior. In practice, the condition is approximate, so the equality is replaced by the approximation P(A|zglobal,Q)≈P(A|Q). P(A|z_global,Q)≈ P(A|Q). (73) When the prior-favored answer ApriorA_prior differs from the visually grounded answer AgtA_gt, this prior collapse can produce P(Aprior|Q)>P(Agt|I~,Q). P(A_prior|Q)>P(A_gt| I,Q). (74) 7.1.5 Support-Truncation Bound of Spatial Anchors For an evidence atom at=⟨bt,ht,ct⟩a_t= b_t,h_t,c_t , let (bt)⊆O(b_t) denote the subset of entities whose spatial support lies inside the region defined by btb_t. Conditioning on btb_t induces the truncated distribution P(o|bt,I~,Q)=P(o|I~,Q)∑o′∈(bt)P(o′|I~,Q),o∈(bt),0,o∉(bt). P(o|b_t, I,Q)= cases P(o| I,Q) _o (b_t)P(o | I,Q),&o (b_t),\\ 0,&o (b_t). cases (75) The support of this distribution is contained in (bt)O(b_t). For any discrete distribution supported on a finite set S, its entropy is upper-bounded by log|| |S|. Therefore, H(Of|bt,I~,Q)≤log|(bt)|. H(O_f|b_t, I,Q)≤ |O(b_t)|. (76) Since the full evidence atom ata_t includes the anchor btb_t, its focus support is also contained in (bt)O(b_t), yielding H(Of|at,I~,Q)≤log|(bt)|. H(O_f|a_t, I,Q)≤ |O(b_t)|. (77) For the unanchored global distribution, H(Of|I~,Q)≤log||. H(O_f| I,Q)≤ |O|. (78) Thus, when btb_t is compact and valid, |(bt)|≪|||O(b_t)| |O|, and the worst-case entropy bound under the anchor is much smaller than the global worst-case bound: log|(bt)|≪log||. |O(b_t)| |O|. (79) This is a support-based bound, not a claim that conditioning on any arbitrary box always reduces the actual entropy. The conclusion holds for valid compact anchors, which is precisely what the spatial grounding objective in EGVOR is designed to encourage. 7.1.6 Explicit Distribution in the Law of Total Variance Let Z be the latent representation induced by the visual focus. For scalar Z, the standard Law of Total Variance gives Var(Z|I~,Q)=Of∼P(⋅|I~,Q)[Var(Z|Of,I~,Q)] (Z| I,Q)=E_O_f P(·| I,Q) [Var(Z|O_f, I,Q) ] (80) +VarOf∼P(⋅|I~,Q)([Z|Of,I~,Q]). +Var_O_f P(·| I,Q) (E[Z|O_f, I,Q] ). For vector-valued Z, the same decomposition applies to covariance matrices; in the main text, Var(Z)Var(Z) can be read as the trace of the covariance matrix or the variance of a projected feature statistic. Under an anchor btb_t, the focus distribution becomes P(⋅|bt,I~,Q)P(·|b_t, I,Q), whose support is contained in (bt)O(b_t). Define m(o)=[Z|Of=o,I~,Q]. m(o)=E[Z|O_f=o, I,Q]. (81) The localized spatial variance term is Vsp(bt)=VarOf∼P(⋅|bt,I~,Q)(m(Of)). V_sp(b_t)=Var_O_f P(·|b_t, I,Q) (m(O_f) ). (82) For vector-valued m(Of)m(O_f), using the trace covariance identity trCov(X)=12‖X−X′‖22, (X)= 12E\|X-X \|_2^2, (83) where X′X is an independent copy of X, we obtain the support-diameter bound Vsp(bt)≤12diam(m(o):o∈(bt))2. V_sp(b_t)≤ 12diam (\m(o):o (b_t)\ )^2. (84) Thus, a compact and query-critical anchor reduces the worst-case spatial-variance support. This formalizes the role of btb_t in suppressing Spatial Boundary Ambiguity. The semantic description ctc_t complements this by encouraging the local representation hth_t to align with a stable semantic concept, thereby targeting the Intrinsic Semantic Uncertainty term. 7.2 Why Wasserstein-2 Distance Provides Dense Spatial Reward This section provides additional explanation for the HRGW spatial reward in Sec. 4.3.2. Intersection-over-Union (IoU) is widely used as a localization metric, but it is less suitable as a reinforcement-learning reward in the early exploration stage. For two bounding boxes bpredb_pred and bgtb_gt, IoU is defined as IoU(bpred,bgt)=|bpred∩bgt||bpred∪bgt|. (b_pred,b_gt)= |b_pred∩ b_gt||b_pred∪ b_gt|. (85) When the two boxes are disjoint, bpred∩bgt=∅⟹IoU(bpred,bgt)=0. b_pred∩ b_gt= 9.24994pt 9.24994ptIoU(b_pred,b_gt)=0. (86) Therefore, IoU assigns the same reward to all non-overlapping predictions, regardless of whether the predicted box is close to the target or far away from it. This creates a sparse reward landscape: a near-miss prediction and a completely irrelevant prediction are indistinguishable as long as neither overlaps with the ground truth. Such sparsity is undesirable for policy optimization, because the model receives no directional signal for how to move the predicted anchor toward the target. To obtain dense geometric feedback, HRGW represents each bounding box as a 2D Gaussian distribution. For a box b=(xc,yc,w,h), b=(x_c,y_c,w,h), (87) where (xc,yc)(x_c,y_c) is the box center and (w,h)(w,h) are its width and height, we map it to b↦(μ,Σ),μ=[xcyc],Σ=[w2/400h2/4]. b (μ, ), 9.24994ptμ= bmatrixx_c\\ y_c bmatrix, 9.24994pt = bmatrixw^2/4&0\\ 0&h^2/4 bmatrix. (88) The exact scale of Σ only rescales the distance and is absorbed by the normalization constant C in the normalized affinity. For two Gaussian boxes pred=(μp,Σp),gt=(μg,Σg), _pred=N( _p, _p), 9.24994ptN_gt=N( _g, _g), (89) the squared Wasserstein-2 distance has the closed form W22(pred,gt)=‖μp−μg‖22 W_2^2(N_pred,N_gt)=\| _p- _g\|_2^2 (90) +Tr(Σp+Σg−2(Σg1/2ΣpΣg1/2)1/2). +Tr ( _p+ _g-2( _g^1/2 _p _g^1/2)^1/2 ). where Tr(⋅)Tr(·) denotes the trace of a square matrix, i.e., the sum of its diagonal elements. The first term measures the displacement between box centers, while the second term measures the mismatch between box scale and aspect ratio. Thus, unlike IoU, W2W_2 remains informative even when two boxes do not overlap. In the axis-aligned case where the covariance matrices are diagonal, the Wasserstein-2 distance can be further simplified. Let Σp=diag(σx,p2,σy,p2),Σg=diag(σx,g2,σy,g2). _p=diag( _x,p^2, _y,p^2), 9.24994pt _g=diag( _x,g^2, _y,g^2). (91) Then W22(pred,gt)=‖μp−μg‖22+(σx,p−σx,g)2 W_2^2(N_pred,N_gt)=\| _p- _g\|_2^2+( _x,p- _x,g)^2 (92) +(σy,p−σy,g)2. +( _y,p- _y,g)^2. This expression shows that the distance provides a continuous penalty for both center displacement and size mismatch. Consequently, a disjoint but nearby prediction receives a smaller distance than a far-away prediction, which gives the policy a meaningful optimization direction. We convert this distance into a normalized affinity: A(bpred,bgt)=exp(−W22(pred,gt)C). A(b_pred,b_gt)= (- W_2^2(N_pred,N_gt)C ). (93) This affinity satisfies 0<A(bpred,bgt)≤1, 0<A(b_pred,b_gt)≤ 1, (94) where the score approaches 11 when the predicted and ground-truth boxes are geometrically aligned, and decreases smoothly as their center or shape discrepancy increases. Therefore, HRGW provides dense reward shaping for spatial anchoring, especially in adverse scenes where early predictions are frequently disjoint from the target. This dense signal helps the policy gradually move predicted anchors toward critO_crit, while avoiding the zero-feedback issue of IoU-based rewards. 7.3 Generalization on the Latest MLLMs To examine the generalization of EGVOR on stronger MLLMs, we apply it to Qwen3.5-9B. As shown in Table 12, although the base model benefits from stronger pretrained reasoning, it still shows clear spatial weaknesses, such as only 11.8% on Perspective Transform. SFT improves format following and boosts Object Retrieval from 68.8% to 87.5%, but also causes drops in OCR and Contact & Occlusion, suggesting that format imitation alone is insufficient. With RL-based spatial and semantic rewards, EGVOR recovers these drops and further improves complex reasoning, achieving gains of +17.5% in Ordering, +11.7% in Perspective Transform, and 51.3% mIoU. These results show that EGVOR remainseffective and complementary to stronger base models. We further evaluate the generalization ability of EGVOR on the Qwen3.5-9B backbone. As shown in Table A3, the original Qwen3.5-9B achieves an average score of 59.0, while supervised fine-tuning already improves it to 64.6, indicating that Qwen3.5 benefits more from the SFT stage due to its stronger pretrained reasoning ability. After applying the full EGVOR pipeline, the performance further increases to 69.4, yielding an overall improvement of +10.4 over the original model. The gains are especially pronounced in Basic Perception (+17.3), mainly driven by large improvements in location (+24.3) and detection (+31.5), while Relation Understanding (+6.4) and Event Reasoning (+9.2) also improve consistently. These results suggest that EGVOR can effectively transfer to a stronger reasoning-capable backbone and generalize beyond the single baseline. Perception Reasoning Model Overall mIoU Attributes Material Phy. State Obj. Retr. OCR Per. Trans. Ordering Con. & Oc. Spa. Cont. Comparison Qwen3.5-9B (Base) 44.0 – 58.6 46.2 65.2 68.8 75.0 11.8 21.1 53.7 72.4 36.4 + SFT (Reflective) 46.5 21.1 62.1 61.5 69.6 87.5 66.2 11.8 22.8 41.5 62.1 50.0 + RL (EGVOR Full) 54.1 51.3 69.0 53.8 78.3 81.2 73.5 23.5 38.6 61.0 79.3 47.7 Δ (EGVOR v.s. Base) ↑10.1 10.1 ↑51.3 51.3 ↑10.4 10.4 ↑7.6 7.6 ↑13.1 13.1 ↑12.4 12.4 ↓1.5 1.5 ↑11.7 11.7 ↑17.5 17.5 ↑7.3 7.3 ↑6.9 6.9 ↑11.3 11.3 Table 12: Generalization of EGVOR on the Qwen3.5-9B architecture (TreeBench). Model Ave. LLM Existence Counting Location Detection OCR-F1 OCR-CER Attribute Position Cap/Other Spatial Occlusion Social Causal w/o CoE w/ CoE Basic Perception Advanced Perception Relation Understanding Event Reasoning Qwen3.5 Series Qwen3.5-9B 59.0 Qwen3.5 61.8 60.4 58.9 37.0 87.4 14.8 58.0 41.0 79.0 59.5 48.5 62.0 55.5 56.7 58.8 54.5 66.3 56.4 58.8 Qwen3.5-9B-SFT 64.6 Qwen3.5 65.5 63.2 74.8 56.5 89.0 12.9 65.5 52.3 80.4 63.0 52.3 67.0 60.5 60.0 60.9 65.0 71.8 60.7 60.9 EGVOR(Qwen3.5-9B) 69.4 Qwen3.5 68.9 66.6 83.2 68.5 89.7 12.5 71.5 57.8 80.7 65.5 55.8 68.5 61.4 66.2 68.0 71.8 75.0 62.8 68.0 Δ v.s. Qwen3.5-9B ↑ 10.4 Qwen3.5 ↑ 7.1 ↑ 6.2 ↑ 24.3 ↑ 31.5 ↑ 2.3 ↓ 2.3 ↑ 13.5 ↑ 16.8 ↑ 1.7 ↑ 6.0 ↑ 7.3 ↑ 6.5 ↑ 5.9 ↑ 9.5 ↑ 9.2 ↑ 17.3 ↑ 8.7 ↑ 6.4 ↑ 9.2 Table 13: Generalization of EGVOR on the Qwen3.5-9B architecture (AD2-Bench). Δ indicates the improvement of EGVOR(Qwen3.5-9B) over the original Qwen3.5-9B. For OCR-CER, lower is better. Figure 10: Bad case: (a) In light fog conditions, the ambiguity caused by the non-dense fog leads many MLLMs to make erroneous judgments. (b) When key vehicles are distant and occluded by rain, most MLLMs fail to identify them accurately, sometimes even exhibiting object hallucinations. 7.4 Bad Case Analysis and Findings Qualitative Error Analysis. Qualitative analysis in Fig. 10 reveals significant MLLM failure modes. Models sometimes misinterpret subtle weather conditions (e.g., light fog, Fig. 10(a)) easily perceived by humans. More critically, under adverse conditions like rain (Fig. 10(b)), models frequently fail to detect crucial safety hazards, such as an unexpected U-turning vehicle directly ahead, while sometimes exhibiting dangerous object hallucinations (e.g., MiniCPM-o-2.6 perceiving a non-existent ’tram’). These errors highlight the urgent need for improved fine-grained perception and robust grounding in visual input, strictly avoiding hallucinations, for safe autonomous driving deployment. Instruction Following Ability. We observed limitations in models’ adherence to specific formatting instructions. Several MLLMs generated extraneous analysis beyond the required single-letter answers for multiple-choice questions, produced excessively verbose outputs often exceeding token limits (particularly for reasoning VQA), and failed to consistently apply the required structural tags (e.g., <startN>…<end><startN>...<end>) for CoE step evaluation. This indicates substantial room for improvement in the instruction following capabilities of current MLLMs, especially regarding structured output formats. 7.5 Detailed Data Composition and Scene Partition While Section 3.1 outlines the primary data sources, this section details the specific criteria used to partition the complex scenarios into intensity levels. This fine-grained categorization is crucial for analyzing model robustness under varying degrees of signal degradation. Stratified Expert Annotation based on Intensity. To ensure consistent labeling across diverse data sources, we established quantitative visual criteria for the three intensity levels for different expert. Specifically, the Light intensity is characterized by visual noise (e.g., drizzle or thin fog) that does not compromise the distinct boundaries of objects at distances greater than 50 meters. The Moderate level involves scenarios where object boundaries become blurred at a medium range (20–50 meters), and textures of road surfaces, such as lane lines, are partially obscured. Finally, Severe conditions represent a high signal-to-noise ratio environment where visibility is restricted to less than 20 meters (e.g., dense fog or heavy snowstorms). In these scenarios, semantic inference relies heavily on contextual reasoning rather than direct visual recognition. Hybrid Filtering Protocol. For datasets lacking explicit weather labels, such as parts of ACDC, our semi-automated filtering pipeline employed a specific complexity threshold. We prompted Gemini 2.5-Pro to score “Scene Complexity” on a scale of 0 to 10 based on object density and visual clarity. Only samples with a complexity score exceeding 6.0 were forwarded to the manual verification stage, ensuring that the benchmark strictly targets complex urban scenarios rather than empty or benign streets. 7.6 Annotation Taxonomy and Prompt Implementation Complementing the Atomic Flow strategy described in Section 3.2, this section details the operational aspects of the annotation campaign and the technical implementation of vision prompts. Taxonomy of 33 Sub-tasks. The cluster analysis of scene complexity yielded 33 distinct sub-tasks, hierarchically organized to support the cognitive pyramid. The Perception Tier encompasses 12 tasks, including fine-grained existence checks, dense and sparse counting, and the absolute localization of 8 distinct urban object classes. Moving up the hierarchy, the Relation Tier consists of 10 tasks specifically targeting spatial relations (e.g., relative positioning), social interactions (e.g., pedestrian-vehicle conflicts), and occlusion status reasoning. The highest level, the Reasoning & Decision Tier, comprises 11 tasks focused on causal inference (e.g., explaining vehicle stops), future prediction, and the generation of actionable safety advice. Vision Prompt Rendering Details. To ensure that vision prompts effectively guide attention without obscuring critical features, we applied strict rendering standards. Point Prompts are rendered as circular markers with a radius of 3 pixels (scaled to image resolution) using high-contrast colors (Red for target, Green for reference). Region Prompts utilize bounding boxes drawn with a 2-pixel width. For cases of “Severe” occlusion, the protocol prioritizes Point prompts to avoid the visual clutter typically associated with overlapping box borders. Response Normalization and Instruction Adherence. We observed varying degrees of instruction-following capabilities across different model families, necessitating a robust output normalization protocol. Systematic formatting issues were particularly prevalent in object localization and Optical Character Recognition (OCR) tasks. In localization scenarios where models were explicitly prompted to output point coordinates formatted as <x,y><x,y>, certain architectures—most notably the MiniCPM series—exhibited a tendency to output bounding box coordinates (e.g., <x1,y1,x2,y2><x_1,y_1,x_2,y_2>) instead. To address this, we implemented an automatic geometric parsing layer that converts such bounding box outputs into their geometric center points for consistent metric calculation. Similarly, in OCR tasks involving traffic signs or license plates, models such as the LLaVA series frequently generated excessive hallucinations or verbose descriptions beyond the requested text content. To mitigate this, we enforced a strict length penalty; responses exceeding a threshold of 50 words were automatically classified as incorrect, with their accuracy scores set to zero. These normalization strategies ensure that the benchmark evaluates genuine perceptual capability rather than formatting compliance. 7.7 Construction of SFT-Instruct. To facilitate the reasoning format establishment in Stage I, we constructed a bespoke dataset, SFT-Instruct, comprising 40K high-fidelity reasoning chains. While we utilize the VGR-158K 75 corpus as a foundational source due to its rich pseudo-chain-of-thought annotations, we introduce a rigorous Restructuring and Enrichment Pipeline to align the data with the atomic requirements of EGVOR. 1) Coordinate System Adaptation for Spatial Fidelity. Standard datasets often employ normalized coordinates (r∈[0,1]r∈[0,1]). However, to fully leverage the high-resolution grounding capabilities of our base model (Qwen2.5-VL series 5), we perform a precision-preserving conversion to absolute coordinates. For every bounding box in the reasoning trace, we transform normalized vectors [rx1,ry1,rx2,ry2][r_x_1,r_y_1,r_x_2,r_y_2] into absolute pixel values [Wrx1,Hry1,Wrx2,Hry2][Wr_x_1,Hr_y_1,Wr_x_2,Hr_y_2], where H×WH× W denotes the native image resolution. This adaptation ensures that the spatial anchors btb_t retain maximum fidelity, mitigating quantization errors during the “Locate” phase. 2) Semantic Densification and Filtering. To ensure the training signal supports complex cognitive paths, we impose a strict complexity filter, retaining only samples involving multi-hop spatial reasoning (i.e., trajectories with multiple distinct bounding boxes). This yields a core set of 40K samples. Crucially, to support the “Describe” phase of our evidence atom at=⟨bt,t,ct⟩a_t= b_t,h_t,c_t , we perform Semantic Densification. We rigorously verify that each spatial anchor btb_t is paired with a descriptive caption ctc_t that explicitly details the visual attributes of the region. Where original descriptions were sparse, we augmented the annotations to ensure that ctc_t provides sufficient semantic supervision for the latent alignment objective (Ξalign _align). 3) Cognitive Augmentation via Synthetic Error Injection. To instantiate the Reflective Decoding mechanism described in Section 4.3.1, we constructed a “Reflective Subset” comprising 5K samples. We employ a controlled perturbation strategy: for a valid reasoning step, we inject a synthetic hallucination by inserting a random distractor box berrb_err followed immediately by a corrective thought token sequence (e.g., “Wait, this box seems to be wrong…”). This design explicitly transforms the dataset from a static collection of correct answers into a dynamic curriculum that trains the model to detect and correct erroneous visual grounding—a critical capability for deployment in the complex environments of AD2-Bench. 7.8 Construction of RL-Instruct. To support the optimization of the composite objective functional JtotalJ_total defined in Stage I, we constructed RL-Instruct, a high-quality dataset comprising 40K samples. Unlike standard VQA datasets, RL-Instruct is meticulously curated to provide the necessary supervision signals for spatial uncertainty minimization (Ψspatial _spatial), latent alignment (Ξalign _align), and cognitive path diversity (Λpath _path). Our construction follows a Dual-Hybridization Strategy to balance general reasoning capabilities with domain-specific robustness in complex urban environments. 1) Hard Sample Mining in General scenario. To ensure the policy learns from challenging scenarios rather than trivial patterns, we utilized the V* training dataset 83 as the foundational source. We implemented a Teacher-Guided Complex Scene Filtering approach using Gemini-2.5-Pro. Specifically, we performed inference on the original 191K training set and retained only the “Hard Samples” where the teacher model identified high visual entropy or complex spatial dependencies. This resulted in a core set of 30K General Reasoning Samples. This selection ensures the model retains broad multimodal understanding capabilities, preventing catastrophic forgetting of general knowledge while learning domain-specific preferences. 2) Multi-Perspective Urban Augmentation. To strictly enforce the entropy reduction mechanism in high-noise environments, we augmented the dataset with 10K Complex Urban Samples. This subset is stratified into two distinct perspectives to maximize geometric invariance: • UAV-View (5K): Sampled from the VisDrone dataset 102, focusing on high-density, small-object counting and tracking from a top-down perspective. • Ego-View (5K): Sampled from the nuScenes dataset 6, focusing on occlusion reasoning and causal prediction from an ego perspective. This hybrid composition prevents the policy from overfitting to a specific visual domain while forcing it to handle the extreme visual confusion characteristic of complex urban scenes. 3) Cognitive Annotation Enrichment. Standard annotations lack the granularity required for our reward functions. We employed a synthetic enrichment pipeline to generate hierarchical metadata. For the spatial objective Ψspatial _spatial, we implemented an Adaptive Weighting Protocol. We first assigned raw importance scores to objects based on their causal relevance to the query (e.g., critical agents vs. background context). Crucially, to maintain a consistent optimization scale across diverse scenes, we enforced a Normalization Constraint per image, ensuring that the weights of all valid objects sum to unity (∑ωo=1Σ _o=1). This effectively models the spatial reward as a weighted expectation over a fixed attention budget. Furthermore, to instantiate the path diversity objective Λpath _path, we prompted Gemini-2.5-Pro to generate three distinct reasoning trajectories for each sample, formally defined as follows: (1) Perceptual-First (Bottom-Up): where reasoning strictly originates from objective, physical evidence in the image (e.g., spatial relations, attributes) before forming a conclusion; (2) Semantic-First (Top-Down): where reasoning is guided by high-level world knowledge and contextual understanding of the scene to infer object properties; and (3) Hypothetico-Deductive: where the process involves systematically formulating a hypothesis and seeking visual evidence to verify or falsify the question’s central premise. By providing these distinct yet valid logical templates, we enable the GRPO algorithm to reward cognitive flexibility, preventing the model from collapsing into a single rigid reasoning pattern. 7.9 Hybrid Annotation Protocols: Balancing Rigor and Depth To facilitate a comprehensive evaluation that encompasses both precise evidence verification and complex reasoning trajectories, AD2-Bench employs a Hybrid Annotation Strategy. This dual-stream protocol assigns distinct formats to different cognitive tasks based on their output characteristics, ensuring that the evaluation metric aligns naturally with the cognitive complexity. Discriminative Format (MC) for Evidence Verification. For atomic cognitive tasks such as Attribute Recognition and Relation Understanding, we adopt a rigorously designed Multiple-Choice (MC) format. This approach enables deterministic evaluation of the model’s ability to distinguish fine-grained visual details against hard negatives. Crucially, to mitigate the “blind guessing” and ”Yes-bias” common in standard MCQA, we implement a robust Anti-Hallucination Mechanism directly within the option design. Unlike random sampling, distractors are explicitly drafted via an Adversarial Distractor Mining strategy, introducing Visual Cousins (visually similar objects, e.g., confusing a truck with a bus in fog) and Hallucination Traps (statistically co-occurring but absent objects). Furthermore, the protocol enforces an Existence Verification constraint by incorporating specific “Abstain” options (E: None of the above; F: Target absent). This design compels the model to validate the target’s presence before classification, strictly penalizing unfounded confidence. Generative Format (VQA) for Cognitive Trajectories. In contrast, for tasks requiring open-ended expressiveness—specifically Target Localization and the construction of the Chain of Evidence (CoE)—we preserve the Hierarchical VQA format. Since Localization requires precise coordinate regression and CoE necessitates the articulation of a logical progression (from perception to decision), restricted options cannot adequately capture these high-dimensional outputs. The VQA format allows us to evaluate the completeness and logic consistency of the model’s reasoning path, providing a granular view of its internal cognitive process beyond simple classification accuracy. A. Bounding Box Identification (Perception Task) [Image] [Question] This task aims to identify the bounding box coordinates [min_x, min_y, max_x, max_y] of a specific object in an image, where (min_x, min_y) is the top-left corner and (max_x, max_y) is the bottom-right corner of the rectangle, with the output format specified as <OUTPUT START>[min_x, min_y, max_x, max_y]<OUTPUT END>. The answer is: B. Multiple-Choice Reasoning (Cognitive Task) [Image] [Question] Here are the options below: (A) [Choice A] (B) [Choice B] (C) [Choice C] (D) [Choice D] (E) [Choice E] (F) [Choice F] Select the best answer to the above multiple-choice question based on the image. Please reply with the letter (A, B, C, D, E or F) corresponding to the correct option. The best answer is: Table 14: Standardized Annotation Format for AD2-Bench Evaluation 7.10 Hierarchical Atomic Decomposition To simulate the cognitive process of an intelligent agent navigating complex visual environments, we enforce a strict Cognitive-Process-Oriented annotation sequence. Experts are mandated to deconstruct the scene in a logical order, establishing a causal chain from global context to local details. 1. Environmental Condition & Visual Constraint Analysis. The process begins with a global assessment, where experts analyze visual degradations (e.g., visibility range, glare intensity) and physical surface conditions. This step establishes the environmental priors necessary for robust reasoning in complex scenes. 2. Key Risk Entity Identification. Instead of exhaustively listing all background objects, annotators prioritize salient entities that influence the agent’s decision-making. This mimics the selective attention mechanism of human perception, filtering out noise to focus on critical elements. 3. Fine-grained State & Attribute Inference. Experts then perform deep visual decoding on the identified entities. Crucially, this involves inferring attributes obscured by adverse weather, such as determining an object’s orientation or status through fog and low light. 4. Predictive Interaction Modeling. Finally, the process culminates in causal reasoning. Experts explicitly annotate potential interactions, predicting future trajectories and defining the causal relationship between the entity’s behavior and the observer’s required response. 7.11 Trust-Aware Annotation Pipeline To guarantee both internal coherence within a reasoning chain and inter-annotator agreement (IAA) across the dataset, we implemented a rigorous pipeline governed by a Single-Annotator Coherence Policy augmented with Quality Checks. Privacy-Utility Trade-off. Acknowledging the tension between privacy protection and task utility (specifically OCR), we adopt a differentiated privacy protocol. Biometric Anonymization is applied to all human faces using an automated detection-blurring pipeline followed by manual verification. However, for License Plates, which are critical for the OCR sub-task, we retain the visual information while enforcing strict data usage protocols. Quality Assurance via IAA. To mitigate subjective variance, we established a rigorous qualification standard using Fleiss’ Kappa (κ). During the training phase, candidate experts were required to annotate a challenging set of 50 pre-validated samples (e.g., sandstorm scenarios). We quantified agreement across four key dimensions. As shown in Table 15, initial agreement on fine-grained objects was substantial (κ≈0.75−0.83κ≈ 0.75-0.83). However, the “Final Decision” initially showed divergence (κ=0.86κ=0.86). Through iterative discussion and the establishment of a “Conservative Reasoning Protocol”—which mandates prioritizing safety when visual cues are ambiguous—we achieved an “Almost Perfect” agreement (κ=0.92κ=0.92). This ensures that the reasoning logic in AD2-Bench aligns with a unified, safety-critical expert consensus. Evaluation Dimension More Detailed than GT Identical to GT Weaker than GT Lacking Critical Info κ (Kappa) Person Perception 2 39 6 3 0.75 Vehicle Perception 5 36 7 2 0.78 Traffic Sign Perception 3 42 4 1 0.83 Final Decision (Initial) 4 41 4 1 0.86 Final Decision (Calibrated) 3 44 2 1 0.92 Table 15: Inter-Annotator Agreement (IAA) Analysis. The table displays the distribution of annotation consistency against Ground Truth (GT) and the resulting Fleiss’ Kappa score. The significant improvement in ”Final Decision” demonstrates the effectiveness of our calibration protocols. Model Under Test Human (Ref) Prompt 1 Prompt 2 Prompt 3 Prompt 4 (Expert) (Nominal) (Detailed) (Few-Shot) (Many-Shot) InternVL3-8B 53.0 61.1 (+8.1) 56.2 (+3.2) 53.5 (+0.5) 52.8 (-0.2) Qwen2.5-VL-7B 52.3 60.8 (+8.5) 55.6 (+3.3) 52.7 (+0.4) 52.7 (+0.4) MiniCPM-o-2.6 51.2 58.5 (+7.3) 54.4 (+3.2) 51.8 (+0.6) 50.7 (-0.5) Avg. Gap (Δ ) – +7.97 +3.23 +0.50 -0.1 Table 16: Calibration analysis of the GPT-5.2 evaluator. We report the scores and their deviation (Δ ) from the expert human reference. 7.12 Metric Validation and Evaluator Calibration To ensure the reliability of our automated evaluation framework (utilizing GPT-5.2), we conducted a rigorous calibration study against human judgment. We established a “Reference Standard” human baseline using a team of domain experts who underwent strict training to unify evaluation criteria, resulting in a conservative and precise scoring distribution. Iterative Prompt Optimization. We systematically refined the evaluator’s alignment with human experts through four progressive prompt levels. We detail the evolution of these prompts below to ensure reproducibility: Level 1: Nominal Definition (P1). This prompt provides only high-level, nominal definitions of the four trustworthiness metrics. It relies entirely on the model’s inherent semantic understanding of terms like “Factual Fidelity” or “Logical Soundness” without defining specific penalty thresholds. This serves as a baseline to measure the model’s unconstrained bias. Level 2: Expert-Derived Specification (P2). This prompt incorporates a comprehensive grading rubric derived from our expert annotation manual. It explicitly details nuances often overlooked by standard models, such as: (1) distinguishing between “plausible” and “visually grounded” hallucinations; (2) penalizing “correct answers derived from wrong reasoning” (False Positive Logic); and (3) defining the boundaries of spatial ambiguity in adverse weather. Level 3: Few-Shot Anchor Calibration (P3). To bridge the gap between abstract definitions and concrete judgment, we introduce specific “Anchor Examples” into the context. This prompt includes a curated set of difficult cases (e.g., tiny objects in fog) accompanied by expert scores and detailed justifications. These anchors serve as cognitive reference points, calibrating the model’s output distribution to match the conservative nature of human experts. Level 4: Many-Shot Saturation (P4). Built upon P3, this variant doubles the number of anchor examples to fully saturate the context window, testing the upper bound of inter-rater agreement. Analysis and Selection. As summarized in Table 16, simple prompts (P1) lead to severe score inflation (Δ≈+8.0 ≈+8.0), reflecting the model’s tendency towards leniency. The introduction of detailed rubrics (P2) reduces this gap, but substantial alignment is only achieved with the inclusion of anchor examples in Prompt 3, where the deviation narrows to ≤0.6≤ 0.6. While P4 achieves near-perfect alignment (Δ≈0 ≈ 0), the marginal gain does not justify the doubled computational cost. Consequently, we adopt Prompt 3 as the standard protocol for AD2-Bench, striking the optimal balance between evaluation fidelity and efficiency. 7.13 Reasoning Format Fig. 12(left) illustrates our Chain-of-Evidence (CoE) process for decomposing complex autonomous driving decision-making problems. This methodology employs a hierarchical reasoning process, structured around key elements within the traffic scene, to address intricate decision-making tasks. The reasoning process commences with an assessment of environmental factors, which typically constitute the most readily perceived and predominant visual information. This foundational environmental perception then grounds the identification of key vehicles pertinent to the ego-vehicle’s driving behavior. As other vehicles often represent the most significant immediate threats, the CoE process, upon establishing environmental awareness, prioritizes reasoning about these potentially hazardous agents. The reasoning hierarchy further extends to Vulnerable Road Users (VRUs), such as pedestrians and cyclists, whose dynamic presence critically influences the ego-vehicle’s maneuvers. Although potentially posing a less immediate physical threat than other vehicles, VRUs demand dedicated perceptual analysis and reasoning. Subsequently, the fourth reasoning layer, designated herein as Step 4, addresses traffic rules. The accurate extraction and interpretation of relevant traffic regulations from the scene are paramount for ensuring sound subsequent decision-making. Finally, by consolidating insights derived from these hierarchically processed layers of perception and understanding, the Chain-of-Thought process culminates in a rational and justifiable driving decision. 7.14 Case Study of Reasoning Task Weather Analysis. Fig. 12(right) and 13 present responses from two advanced models, InternVL3 and Qwen2.5-VL, to challenges posed by adverse weather conditions in complex scenarios. In these figures, areas shaded yellow indicate reasoning errors, while red or blue highlights denote correct portions of the answer. For the first image in Fig. 12(left), we observe that Qwen2.5-VL, despite perceiving road fog, still concludes the weather is ‘overcast’ and fails to correctly distinguish whether the current scene is rainy or foggy. Such light fog conditions, being similar in appearance to rainy scenes, can easily lead to model misjudgment. In contrast, InternVL3 accurately identifies the current weather and also provides a fine-grained description of pedestrians. However, both models exhibit distinct focuses when describing traffic signs and road markings: Qwen2.5-VL identifies a yellow pedestrian crossing, whereas InternVL3 points out a sign on the left. Our expectation is for models to comprehensively perceive and understand the entire scene; thus, their current capabilities in understanding this specific scenario still reveal certain deficiencies. Vision Prompt Analysis. Regarding the second image in Fig. 13(left), the most critical element is a black sedan ahead, obscured by rain, which is either making a U-turn or crossing the road. Qwen2.5-VL perceives several white cars parked on the roadside; however, these are irrelevant to the ego car’s current driving status. The priority is key information that can impact the ego car’s driving decisions. InternVL3 performs poorly in this instance. Although it segments the input into patches and resizes them to a fixed dimension for the vision encoder, this method appears ineffective here, as the model fails to detect any vehicles. This oversight is extremely detrimental to safe driving. To further evaluate the reasoning capabilities of MLLMs and to mitigate the issue of insufficient perceptual acuity in vision encoders, we experimented with incorporating vision prompts in Fig. 13(right). By overlaying annotated points onto the image and providing this composite input to the MLLM, we found that Qwen2.5-VL, while misidentifying the vehicle type, correctly recognized the vehicle’s color and direction of movement (from right to left). This indicates that our vision prompts are largely effective, at least enabling the model to perceive the presence and movement of the vehicle ahead. Similarly, InternVL3 demonstrated even better performance: its vehicle type identification was more accurate than that of Qwen2.5-VL, and its overall responses were largely correct. This showcases the potential of vision prompts for conducting in-depth assessments of MLLMs’ advanced reasoning capabilities. 7.15 Case Study of Other Tasks Fig. 14 presents further case studies across various tasks. From Fig. 14(left), it is evident that adverse weather and complex scenarios significantly impair the perceptual capabilities of MLLMs. This is particularly apparent in a specific image from this figure (referred to as image2), concerning the detection of pedestrian presence. In this nighttime intersection scene, although several pedestrians are situated on the sidewalk across from the ego-vehicle’s front-right, most models fail to detect them. This failure is attributed to occlusion by a white car turning from the front-right, compounded by the low-light conditions. Such perceptual misses are highly detrimental to subsequent understanding and reasoning processes. Fig. 14(right) showcases MLLM performance on localization, detection, and Optical Character Recognition (OCR) tasks. For localization and detection, Qwen2.5-VL notably outperforms other models. Many models tend to output coordinates in a proportional (normalized) format; however, due to image patching and sampling operations during input processing, relying on these proportions often fails to accurately reconstruct true bounding boxes. In contrast, approaches similar to those used by Qwen2.5-VL and Qwen2-VL, which involves passing the entire image to subsequent modules via a dynamic resolution scheme—achieve higher scores on these tasks. This improved performance stems from the fact that the final bounding box and localization point coordinates are established within the entire image’s coordinate system, rendering full-image input superior to patched or sampled inputs in terms of overall perceptual integrity and output format consistency. Regarding OCR tasks, especially in low-light conditions, most models manage to output text in the expected format. LLaVA-1.5, however, only outputs coordinates, indicating poor instruction-following capabilities. Our findings show that InternVL series models achieve significantly better OCR recognition results compared to others. Conversely, MiniCPM series models, which also utilize patched image inputs, deliver underwhelming performance. This suggests that challenges in OCR perception are not solely attributable to the input format and may necessitate greater consideration of training data. We observe that MiniCPM models, designed for edge-device deployment, possess training datasets with substantially less OCR-specific content than more general-purpose MLLMs like the InternVL series, a factor that critically contributes to their diminished OCR efficacy. 7.16 Per-condition Robustness Analysis To further examine robustness under different adverse conditions, we provide a per-condition breakdown in Fig. 11. Instead of reporting only an overall average, we evaluate representative proprietary and open-source MLLMs across seven conditions and summarize their behavior from four perspectives: Basic Perception, Advanced Perception, Relation Understanding, and Event Reasoning. This visualization reveals that adverse conditions affect different stages of the reasoning pipeline in different ways. For instance, night and daytime-occlusion introduce more visible degradation in perception-related and relation-understanding scores, suggesting that low-level visual uncertainty can propagate to higher-level reasoning. The results also show that stronger overall models are not uniformly robust across all adverse conditions. Some models maintain competitive relation or event reasoning performance, but still exhibit weakened visual grounding under challenging conditions, indicating that reasoning ability alone cannot fully compensate for unreliable perceptual evidence. Conversely, models with stronger visual grounding tend to be more stable in Basic Perception, but may not always dominate in higher-level reasoning categories. These findings support our motivation that robust adverse-condition understanding requires both reliable evidence extraction and effective reasoning over the constructed evidence. Figure 11: Per-condition robustness analysis on AD2-Bench. The four radar plots report representative MLLMs across seven adverse conditions from Basic Perception, Advanced Perception, Relation Understanding, and Event Reasoning. Figure 12: The visualization of Hierarchical Visual Diagnosis for MLLM. Figure 13: The visualization of Hierarchical Visual Diagnosis for MLLM. By breaking down unified scene problems into fine-grained perception, understanding, and reasoning tasks, CoE assists MLLM in achieving a hierarchical and comprehensive understanding of the current scene for making the most reasonable decisions. Figure 14: The visualization of Basic and Advanced Perception Tasks of Multimodal large language Models under Adverse Weather and Complex Scene Conditions, with yellow highlights representing correct answers.