Paper deep dive
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
Dikshant Kukreja, Kshitij Sah, Karan Goyal, Mukesh Mohania, Vikram Goyal
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/10/2026, 3:01:12 AM
Summary
The paper introduces DISSECT, a 12,000-question diagnostic benchmark for Vision-Language Models (VLMs) in Chemistry and Biology. It identifies a 'perception-integration gap' where models can extract visual information but fail to use it for reasoning. Using a five-mode evaluation protocolâincluding a novel 'Model Oracle'âthe study reveals that open-source models suffer from systematic integration bottlenecks, whereas closed-source models do not, and that Chemistry is a more robust test of visual reasoning than Biology due to lower language-prior exploitability.
Entities (4)
Relation Signals (3)
Model Oracle â diagnoses â Perception-Integration Gap
confidence 95% ¡ The Model Oracle protocol is both model and benchmark agnostic, applicable post-hoc to any VLM evaluation to diagnose integration failures.
DISSECT â evaluates â Vision-Language Model
confidence 95% ¡ Evaluating 18 VLMs, we find that...
Chemistry â haslowerexploitabilitythan â Biology
confidence 90% ¡ Chemistry exhibits substantially lower language-prior exploitability than Biology
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:When asked to describe a molecular diagram, a Vision-Language Model correctly identifies ``a benzene ring with an -OH group.'' When asked to reason about the same image, it answers incorrectly. The model can see but it cannot think about what it sees. We term this the perception-integration gap: a failure where visual information is successfully extracted but lost during downstream reasoning, invisible to single-configuration benchmarks that conflate perception with integration under one accuracy number. To systematically expose such failures, we introduce DISSECT, a 12,000-question diagnostic benchmark spanning Chemistry (7,000) and Biology (5,000). Every question is evaluated under five input modes -- Vision+Text, Text-Only, Vision-Only, Human Oracle, and a novel Model Oracle in which the VLM first verbalizes the image and then reasons from its own description -- yielding diagnostic gaps that decompose performance into language-prior exploitation, visual extraction, perception fidelity, and integration effectiveness. Evaluating 18~VLMs, we find that: (1) Chemistry exhibits substantially lower language-prior exploitability than Biology, confirming molecular visual content as a harder test of genuine visual reasoning; (2) Open-source models consistently score higher when reasoning from their own verbalized descriptions than from raw images, exposing a systematic integration bottleneck; and (3) Closed-source models show no such gap, indicating that bridging perception and integration is the frontier separating open-source from closed-source multimodal capability. The Model Oracle protocol is both model and benchmark agnostic, applicable post-hoc to any VLM evaluation to diagnose integration failures.
Tags
Links
- Source: https://arxiv.org/abs/2604.06250v1
- Canonical: https://arxiv.org/abs/2604.06250v1
Trouble viewing inline? Open PDF directly â
Full Text
78,166 characters extracted from source content.
Expand or collapse full text
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs Dikshant Kukreja IIIT Delhi, India dikshant22176@iiitd.ac.in , Kshitij Sah IIIT Delhi, India kshitij22256@iiitd.ac.in , Karan Goyal IIIT Delhi, India karang@iiitd.ac.in , Mukesh Mohania IIIT Delhi, India mukesh@iiitd.ac.in and Vikram Goyal IIIT Delhi, India vikram@iiitd.ac.in Abstract. When asked to describe a molecular diagram, a Vision-Language Model correctly identifies âa benzene ring with an -OH group.â When asked to reason about the same image, it answers incorrectly. The model can see but it cannot think about what it sees. We term this the perception-integration gap: a failure where visual information is successfully extracted but lost during downstream reasoning, invisible to single-configuration benchmarks that conflate perception with integration under one accuracy number. To systematically expose such failures, we introduce DISSECT, a 12,000-question diagnostic benchmark spanning Chemistry (7,000) and Biology (5,000). Every question is evaluated under five input modesâVision+Text, Text-Only, Vision-Only, Human Oracle, and a novel Model Oracle in which the VLM first verbalizes the image and then reasons from its own descriptionâyielding diagnostic gaps that decompose performance into language-prior exploitation, visual extraction, perception fidelity, and integration effectiveness. Evaluating 18 VLMs, we find that: (1) Chemistry exhibits substantially lower language-prior exploitability than Biology, confirming molecular visual content as a harder test of genuine visual reasoning; (2) Open-source models consistently score higher when reasoning from their own verbalized descriptions than from raw images, exposing a systematic integration bottleneck; and (3) Closed-source models show no such gap, indicating that bridging perception and integration is the frontier separating open-source from closed-source multimodal capability. The Model Oracle protocol is both model and benchmark agnostic, applicable post-hoc to any VLM evaluation to diagnose integration failures. Vision-Language Models, Diagnostic Benchmark, Scientific Visual Reasoning, Multimodal Evaluation, Chemistry, Biology â copyright: noneâ ccs: Computing methodologies Artificial intelligenceâ ccs: Computing methodologies Computer vision 1. Introduction Consider a Chemistry question that shows a structural formula and asks: âIdentify the functional group present in this compound.â A state-of-the-art Vision-Language Model (VLM) answers correctly. But remove the image and provide only the question textâthe same model still answers correctly, exploiting memorized associations between question phrasing and likely answers (Zhang et al., 2024; Cui et al., 2025). Now consider a question showing a skeletal structure of two constitutional isomers and asking: âWhich isomer has the higher boiling point, and why?â Here, removing the image causes complete failureâthe model cannot reason about intermolecular forces without seeing the structures. These two scenarios represent fundamentally different evaluation regimes, yet existing benchmarks conflate them under a single accuracy number. Current multimodal benchmarks for scientific reasoning suffer from two critical blind spots. First, they cannot distinguish genuine visual understanding from language-prior exploitationâthe degree to which models answer âvisualâ questions without seeing anything. MathVerse (Zhang et al., 2024) introduced multi-version evaluation but is restricted to mathematics, the most text-exploitable STEM subject. SciVerse (Guo et al., 2025) extended to science but varies knowledge level (how much domain expertise is provided), not input modality (which channel the model actually uses). Second, when a model fails, no existing benchmark can determine whether the failure is perceptual (the model cannot see the diagram), extractive (it cannot read text embedded in the image), or integrative (it sees the content but cannot reason over it). We introduce DISSECT (DIagnostic Separation of Seeing, Extracting, and Cognitive Thinking), a diagnostic benchmark that addresses both limitations through three contributions. (1) A five-mode evaluation protocol with a novel Model Oracle. Every question is evaluated under five input configurations (c.f. Figure 1): Vision+Text (standard), Text-Only (image removed), Vision-Only (question embedded in image), Human Oracle (human-annotated symbolic labels), and Model Oracle (the VLM first verbalizes the image, then reasons from its own description). The Model Oracle isolates a failure mode invisible to prior work: the perception-integration gap, where a model can describe what it sees but fails to use that information during reasoning. (2) The DISSECT benchmark: 12,000 questions across two irreducibly visual sciences. DISSECT comprises 7,000 Chemistry and 5,000 Biology questions, each manually verified to require visual input for a human solver. We deliberately focus on these two subjects because molecular structures, reaction mechanisms, cell diagrams, and anatomical illustrations cannot be adequately described in text aloneâunlike mathematical diagrams, which are symbolically reducible and already heavily represented in existing benchmarks (Zhang et al., 2024; Lu et al., 2024). (3) Diagnostic findings across 18 VLMs. Our evaluation reveals a clear divide: open-source VLMs consistently outperform themselves when reasoning from their own verbalized descriptions rather than raw images, exposing a systematic integration bottleneck. Closed-source models show no such gap, indicating that the perception-integration bridge is the frontier separating open- from closed-source multimodal reasoning. Figure 1. The DISSECT five-mode evaluation protocol. Each question is evaluated under Vision+Text, Text-Only, Vision-Only, Human Oracle, and Model Oracle configurations. Five boxes showing the same chemistry question rendered under five different input configurations. 2. Related Work Multi-Version Evaluation. MathVerse (Zhang et al., 2024) transforms each math problem into six versions by progressively moving information from text to diagram, creating a spectrum from Text Dominant to Vision Only. Their key findingâthat Qwen-VL-Max scores over 5% higher without diagramsâdirectly motivates our work. However, MathVerse is restricted to mathematics (geometry and functions), the most text-exploitable STEM subject. Its six versions vary the information distribution between modalities but do not include an oracle to separate perception from reasoning failures. SciVerse (Guo et al., 2025), from the same group, extends to Physics, Chemistry, and Biology with five versions. Critically, SciVerseâs decomposition axis is knowledge contentâit varies how much domain knowledge is embedded in the question (Knowledge-free/lite/rich) and how much is visual (Vision-rich/only). This answers âDoes the model know enough science?â but cannot answer âIs the model actually looking at the image?â because SciVerse never provides the question text without an image. DISSECTâs decomposition is orthogonal: we fix the knowledge content and vary the input channel, enabling direct measurement of language-prior exploitation via the Text-Only mode and integration failure via the Model Oracle. Chemistry and Biology Benchmarks. ChemVTS-Bench (Huang et al., 2025) evaluates MLLMs under three modes: visual-only, visual-text, and SMILES-based symbolic input. However, its symbolic mode substitutes a machine-readable string (SMILES) for the image, testing whether models can reason from formal notationâa different question from whether they can reason from their own perception, which is what our Model Oracle tests. ChemVTS-Bench also lacks a text-only baseline and therefore cannot measure language-prior exploitation. ChemVLM (Li et al., 2025b) introduces MMCR-Bench (1,000 chemistry exam questions) but evaluates under a single V+T configuration. USNCO-V (Cui et al., 2025) evaluates 40 VLMs on Chemistry Olympiad exams and includes an image-removal ablation, finding that removing images sometimes improves accuracy. This corroborates our findings but is limited to Olympiad-level questions (âź 400 items) with only two modes (with/without image) and no mechanism to explain why image removal helps. For Biology, no dedicated VLM benchmark with visual dependency analysis exists. MMMU (Yue et al., 2024) and EXAMS-V (Das et al., 2024) include biology questions among broader multi-discipline collections but use single-mode evaluation with no ablation across input configurations. Language Priors in VQA and VLMs. Goyal et al. (Goyal et al., 2017) demonstrated that VQA models exploit statistical regularities between question types and answer distributions, achieving high accuracy without genuinely attending to images. Agrawal et al. (Agrawal et al., 2018) showed this persists under distribution shift. Recent work confirms that modern VLMs inherit this vulnerability: ViLP (Luo et al., 2025) found that GPT-4o achieves only 66% on deliberately out-of-distribution visual questions, and VisRes Bench (TĂśrtei et al., 2025) showed that VLM performance collapses when linguistic cues are removed from visual reasoning tasks. At the mechanistic level, the âSeeing but Not Believingâ phenomenon (Liu et al., 2025) reveals that VLMs can attend to the correct visual region yet still produce wrong answersâan attention-level analogue of the integration failure that DISSECT quantifies at benchmark scale through the Model Oracle. 3. The DISSECT Benchmark 3.1. Data Collection and Curation DISSECT draws from standard textbooks, government-issued educational materials, examination papers, and curated question banks spanning grades 9â12. While sourced from the Indian curriculum, the underlying contentâorganic chemistry, thermodynamics, cell biology, human anatomyâis universal and curriculum-agnostic. 3.1.1. Collection and Quality Control. Questions were extracted from textbooks and examination papers. Image-based content was captured via high-resolution screenshots to preserve visual fidelity. We performed deduplication, filtering of low-resolution or ambiguous images, and answer verification against source materials. 3.1.2. Filtering for Visual Dependency. A key design principle is that every question in DISSECT requires visual input for a human solver. We performed manual annotation (c.f. Appendix C) to identify and remove text-solvable questions, retaining only those where the image provides essential, non-recoverable information. This ensures genuine visual dependency across the dataset. 3.2. Question Type Taxonomy We categorize questions along two axes. Visual Content Type. Chemistry: molecular structure recognition, reaction mechanism diagrams, orbital and bonding diagrams, apparatus and experimental setups, phase diagrams, and chemical equation balancing with visual notation. Biology: cell and organelle diagrams, anatomical illustrations, phylogenetic trees, ecological diagrams, physiological process flowcharts, and microscopy images. Cognitive Level. Recall: direct identification or extraction from the image. Application: applying a known concept to visual information. Analysis: multi-step reasoning integrating visual and textual information. 3.3. Five-Mode Input Construction For each question qiq_i with associated image IiI_i and text TiT_i, we construct five evaluation inputs. We illustrate all five modes with a running Chemistry example in Figure 1 and provide additional examples across both subjects in Appendix LABEL:sec:examples. 3.3.1. Mode 1: Vision+Text (V+T). Input: (Ii,Ti)(I_i,\,T_i). The standard multimodal setup: the original image and question text are provided together. This is how all existing benchmarks evaluate VLMs and serves as our baseline. Example (Chemistry): The model receives an image showing a skeletal structural formula alongside the text: âIdentify the functional group circled in the structure.â Example (Biology): The model receives a labeled diagram of a human heart alongside: âWhich chamber receives deoxygenated blood from the body?â 3.3.2. Mode 2: Text-Only (T). Input: (Ti)(T_i). The image is entirely removed; only the question stem is provided. Any accuracy achieved in this mode reflects language-prior exploitation: the model is answering a âvisualâ question without seeing anything. This mode directly quantifies a risk invisible to single-mode benchmarks. Example (Chemistry): The model receives only: âIdentify the functional group circled in the structure.â A model that answers âhydroxyl groupâ is pattern-matching from the word âfunctional groupâ and common exam phrasing, not performing visual reasoning. Example (Biology): The model receives only: âWhich chamber receives deoxygenated blood from the body?â This is answerable from textbook memorizationâno diagram neededâexposing a question where V+T accuracy overestimates visual grounding. 3.3.3. Mode 3: Vision-Only (V). Input: (Ii+T)(I_i^+T). The question text is rendered directly onto the image (concatenated below or overlaid with a white background region), producing a single visual input with no separate text channel. The model must OCR the question, parse it, and jointly reason with the visual content. This mode isolates visual text extraction: failures here that do not occur in V+T reveal a dependency on the structured text channel. Example (Chemistry): The model receives a single image containing the structural formula and the question âIdentify the functional group circled in the structureâ rendered as text within the image. Chemical notation (subscripts like H2SO4, bond-line angles, stereochemistry wedges) makes OCR particularly challenging in this domain. Example (Biology): The model receives the heart diagram with the question rendered as embedded text. Biological diagrams often contain their own labels (âleft atrium,â âaortaâ), and the model must distinguish embedded question text from diagram labels. 3.3.4. Mode 4: Human Oracle (OHO_H).. Input: (Iilab,Ti)(I_i^lab,\,T_i). Human annotators augment the image with explicit symbolic annotations that resolve all perceptual ambiguity: labeled atoms and bonds, named biological structures, extracted numerical values, spatial relationship descriptions, and color/shading clarifications. The annotated image is provided alongside the original question text. This mode establishes a reasoning upper bound: if a model fails under OHO_H, the bottleneck is reasoning or knowledge, not perception. If it succeeds under OHO_H but fails under V+T, the bottleneck is definitively perceptual. Example (Chemistry): The structural formula is annotated with explicit labels: âC6H5âOH,â âhydroxyl group (âOH) at position marked with circle,â âaromatic ring with alternating double bonds.â Example (Biology): The heart diagram is annotated with: âChamber A = Right Atrium (receives blood from superior/inferior vena cava),â âChamber B = Right Ventricle,â etc. 3.3.5. Mode 5: Model Oracle (OMO_M).. Input: Two-pass, model-specific. This is DISSECTâs novel contribution. The same VLM being evaluated performs a two-pass procedure: Pass 1 (Perception): The model receives (Ii,Ti)(I_i,\,T_i) and is prompted to generate a detailed, structured description of all visual content relevant to the question. Crucially, it is instructed to describe, not answerâsee Appendix A for exact prompt templates. Pass 2 (Reasoning): The model receives only its own generated description DiD_i and the original question TiT_i, with no image. It must now answer using solely what it extracted. The diagnostic logic is as follows. If the model answers correctly in Pass 2 (from its own description) but incorrectly in V+T (from the raw image), then the model possesses sufficient perceptual capabilityâit successfully extracted the relevant information in Pass 1âbut fails to integrate that information during end-to-end reasoning. This is the perception-integration gap: the model can see, but it cannot think about what it sees when both modalities compete for attention in a single forward pass. Example (Chemistry): Pass 1 output: âThe image shows a benzene ring (six-membered aromatic carbon ring) with a hydroxyl group (âOH) attached. The circled region highlights the âOH group.â Pass 2 input: [above description] + âIdentify the functional group circled in the structure.â If the model answers âhydroxyl groupâ in Pass 2 but answered âesterâ in V+T, the integration failure is exposed. Example (Biology): Pass 1 output: âThe diagram shows a four-chambered human heart. The upper-right chamber is labeled as receiving blood from the vena cava. Blue shading indicates deoxygenated blood flow.â Pass 2 input: [above description] + âWhich chamber receives deoxygenated blood from the body?â 3.4. Diagnostic Gaps The five modes yield five diagnostic metrics: ⢠Language Prior Gap (LPG): Accâ(T)Accâ(V+T) Acc(T)Acc(V+T). Fraction of performance attributable to text alone. Higher values indicate greater exploitability. ⢠Extraction Gap: Accâ(V+T)âAccâ(V)Acc(V+T)-Acc(V). Performance lost when text must be extracted from the image rather than received as structured input. ⢠Perception Gap: Accâ(OH)âAccâ(V+T)Acc(O_H)-Acc(V+T). Total performance lost due to imperfect visual perception. Failure under OHO_H indicates a reasoning bottleneck; failure under V+T but success under OHO_H indicates a perception bottleneck. ⢠Integration Gap: Accâ(OM)âAccâ(V+T)Acc(O_M)-Acc(V+T). A positive Integration Gap means the model can extract relevant visual information (it did so in Pass 1) but fails to leverage it during end-to-end reasoning. This metric is novel to DISSECT. ⢠Perception Fidelity: Accâ(OH)âAccâ(OM)Acc(O_H)-Acc(O_M). Information lost by the modelâs own perception relative to human annotation. A large gap indicates that the modelâs self-description is incomplete or inaccurate. 4. Experiments 4.1. Models and Evaluation Protocol We evaluate 18 VLMs spanning three categories, selected to enable within-family scaling analysis across two orders of magnitude in parameter count. Closed-Source:- GPT-5 (OpenAI, 2025), Gemini 2.5 Flash (Comanici et al., 2025), and Claude Sonnet 4 (Anthropic, 2025). Open-Source Large (⼠7B):- InternVL3-8B, -14B, -38B, and -78B (Zhu et al., 2025); Qwen2.5-VL-7B, -32B, and -72B (Bai et al., 2025); LLaVA-OneVision-7B and -72B (Li et al., 2025a). Open-Source Small (<<7B):- InternVL3-1B, -2B, and -4B (Zhu et al., 2025); Qwen2.5-VL-3B (Bai et al., 2025); LLaVA-OneVision-0.5B (Li et al., 2025a). The InternVL3 family (1Bâ78B) and Qwen2.5-VL family (3Bâ72B) each span two orders of magnitude in parameter count, enabling direct analysis of how diagnostic gaps scale with model capacity within a fixed architecture. All models are evaluated zero-shot with standardized prompts (Appendix A). For multiple-choice questions, we extract the predicted option letter via regex matching. For open-ended numerical answers, we normalize predictions (stripping units, rounding to significant figures) and apply exact-match evaluation. The Model Oracle (OMO_M) uses a two-pass protocol: Pass 1 elicits a structured image description without answering the question; Pass 2 provides only that description and the question text, with no image. Both passes use carefully controlled system prompts to prevent information leakage between the perception and reasoning stages (see Appendix A for full templates). 5. Results and Analysis 5.1. Main Results Table 1 presents accuracy across all five modes for Chemistry and Biology. We report Language Prior Gap (LPG) and Integration Gap (IG) as the two most diagnostic metrics; the remaining gaps (Extraction, Perception Fidelity) are analyzed in subsequent subsections. Table 1. Performance (%) on DISSECT across five evaluation modes. V+T: Vision+Text; T: Text-Only; V: Vision-Only; O_H: Human Oracle; O_M: Model Oracle; LPG: Language Prior Gap (T/V+T/V+T, lower = more visually dependent); IG: Integration Gap (OMâV+TO_M-V+T, positive = integration bottleneck). Best open-source per column in bold; best overall underlined. Chemistry (7,000 questions) Biology (5,000 questions) Model V+T T V OHO_H OMO_M LPGâ IG V+T T V OHO_H OMO_M LPGâ IG Closed-Source GPT-5 71.2 24.3 62.4 79.1 72.0 0.34 +0.8 78.3 51.2 71.1 84.2 78.9 0.65 +0.6 Gemini 2.5 Flash 68.7 22.8 59.6 76.4 69.5 0.33 +0.8 75.8 49.3 68.4 82.1 76.5 0.65 +0.7 Claude Sonnet 4 70.4 23.6 61.2 78.3 71.2 0.34 +0.8 77.1 50.4 69.8 83.6 77.8 0.65 +0.7 Open-Source Large (⼠7B) InternVL3-78B 62.4 21.1 53.4 74.6 68.9 0.34 +6.5 70.6 46.8 63.4 78.8 74.5 0.66 +3.9 InternVL3-38B 58.6 20.3 49.8 71.2 65.7 0.35 +7.1 66.4 44.7 59.5 75.9 71.2 0.67 +4.8 InternVL3-14B 54.1 19.5 45.3 67.8 62.1 0.36 +8.0 62.1 42.5 55.2 72.4 67.6 0.68 +5.5 InternVL3-8B 50.3 18.7 41.6 64.5 58.7 0.37 +8.4 58.7 40.6 51.8 69.5 64.8 0.69 +6.1 Qwen2.5-VL-72B 63.8 21.6 54.7 75.3 69.4 0.34 +5.6 72.1 47.9 65.2 80.1 75.8 0.66 +3.7 Qwen2.5-VL-32B 57.2 20.1 48.1 69.8 64.4 0.35 +7.2 65.3 44.1 58.4 74.5 70.3 0.68 +5.0 Qwen2.5-VL-7B 49.5 18.4 40.7 63.1 57.2 0.37 +7.7 57.8 39.6 50.6 68.4 63.7 0.69 +5.9 LLaVA-OV-72B 59.1 20.8 50.2 71.8 66.1 0.35 +7.0 67.2 45.3 60.1 76.5 72.1 0.67 +4.9 LLaVA-OV-7B 46.3 17.9 37.5 60.1 54.9 0.39 +8.6 54.6 37.8 47.4 65.7 61.2 0.69 +6.6 Open-Source Small (<<7B) InternVL3-4B 44.7 17.2 35.8 58.2 53.1 0.38 +8.4 52.8 36.4 45.3 63.4 58.9 0.69 +6.1 InternVL3-2B 39.2 16.1 31.1 53.8 48.4 0.41 +9.2 47.3 33.8 40.4 59.3 54.4 0.71 +7.1 InternVL3-1B 33.6 15.3 26.2 49.1 43.5 0.46 +9.9 41.2 30.6 35.1 55.1 49.6 0.74 +8.4 Qwen2.5-VL-3B 42.1 16.8 33.6 56.3 51.6 0.40 +9.5 50.4 35.2 43.7 61.9 57.1 0.70 +6.7 LLaVA-OV-0.5B 28.4 14.2 21.8 44.6 38.9 0.50 +10.5 35.7 27.4 29.9 50.2 44.3 0.77 +8.6 5.2. Finding 1: Chemistry Is Irreducibly Visual Chemistry exhibits a substantially lower Language Prior Gap than Biology across every model evaluated (c.f. Table 1). Averaging across all 18 models, the mean LPG for Chemistry is 0.37, compared to 0.68 for Biologyâmeaning that on average, Biology models recover 68% of their V+T performance from text alone, while Chemistry models recover only 37%. This gap is consistent across model classes: closed-source models average 0.34 (Chemistry) vs. 0.65 (Biology); open-source large models average 0.36 vs. 0.68; and open-source small models average 0.43 vs. 0.72. Even the smallest model (LLaVA-OV-0.5B, LPG=0.50=0.50 in Chemistry) exploits language priors less than the best closed-source model does in Biology (GPT-5, LPG=0.65=0.65). The explanation is structural: Chemistry questions in DISSECT reference molecular diagrams, bond-line structures, and reaction mechanisms whose identity cannot be inferred from the question text. A question asking âWhat is the IUPAC name of this compound?â is unanswerable without the image. In contrast, Biology questions frequently contain informative textââWhich chamber of the heart receives deoxygenated blood?ââthat enables models to answer from parametric knowledge even when the accompanying diagram is withheld. This finding has direct implications for benchmark design: mathematics centric evaluations such as MathVerse and MathVista systematically overestimate visual reasoning capability because mathematical diagrams are the most symbolically reducibleâand therefore most text-exploitableâvisual content in STEM. Chemistry provides the hardest test of genuine visual dependence. 5.3. Finding 2: Open-Source Models See But Cannot Integrate The Integration Gap (OMâ(V+T)O_M-(V+T)) reveals a striking and consistent divide between open- and closed-source VLMs (Table 1). Open-source models exhibit systematic integration failure. Across all 15 open-source models, the Integration Gap is positive in both subjects without exception. The average IG is +7.8+7.8 percentage points for Chemistry and +5.7+5.7 for Biology. This means that when these models verbalize the image first (Pass 1) and then reason from their own description (Pass 2), they consistently outperform themselves compared to reasoning directly from the raw image. The model possesses sufficient perceptual capabilityâit successfully extracts the relevant visual informationâbut fails to leverage that information during end-to-end multimodal reasoning. The gap is particularly pronounced for small models: LLaVA-OV-0.5B shows an IG of +10.5+10.5 in Chemistry, meaning its two-pass accuracy exceeds its direct multimodal accuracy by over 10 absolute points. Closed-source models have bridged this gap. For GPT-5, Gemini 2.5 Flash, and Claude Sonnet 4, the Integration Gap is uniformly small: +0.8+0.8 or less in Chemistry and +0.7+0.7 or less in Biology. These models achieve effectively identical performance whether reasoning from the raw image or from their own verbalized description, indicating that their multimodal reasoning pipelines successfully integrate visual features into downstream reasoning without information loss. Architectural implications. The integration failure in open-source models cannot be attributed to the vision encoder alone (the model can describe the image correctly) nor to the language model alone (it can reason correctly from text). The bottleneck lies in the cross-modal projection and attention mechanisms that bridge vision and language during end-to-end inference. This points to the projection layer and cross-modal attention architecture as the primary target for improving open-source scientific VLMs. 5.4. Finding 3: Perception Fidelity Varies by Visual Domain The Perception Fidelity gap (OHâOMO_H-O_M) measures how much information the modelâs own image description misses compared to expert human annotation. Averaging across all open-source models, this gap is 6.8 percentage points for Chemistry but only 4.6 for Biologyâa 48% relative increase. The asymmetry reflects the nature of visual information in each domain. Chemistry images contain dense symbolic contentâbond types (single, double, aromatic), stereochemistry indicators (wedge and dash bonds), subscript notation (H2SO4), and ring-system topologyâthat current vision encoders frequently misread or omit during verbalization. A single misidentified bond (e.g., reading a double bond as single) can cascade into a wrong answer. Biology images, by contrast, contain more spatially distributed, labeled structures (âleft atrium,â âmitochondriaâ) that models transcribe more reliably because the labels appear as standard text rather than symbolic notation. Closed-source models show a narrower gap (5.4 for Chemistry, 4.1 for Biology), suggesting that their vision encodersâor the post-processing applied to visual featuresâbetter handle dense symbolic content. 5.5. Finding 4: Extraction Failures Are Domain-Specific The Extraction Gap ((V+T)âV(V+T)-V) captures performance lost when the question text must be OCRâd from the image rather than received as structured input. Averaging across all models, this gap is 9.7 percentage points for Chemistry and 7.4 for Biology. The larger Chemistry gap is driven by the co-occurrence of chemical notation and question text within the same image. In the Vision-Only mode, models must simultaneously parse question text (standard English) and chemical notation (subscripts, superscripts, bond-line angles, reaction arrows, equilibrium symbols). These two notation systems use overlapping visual featuresâsmall characters, spatial positioningâthat create OCR interference. Biology images present a less severe extraction challenge because biological labels (organ names, species labels) use standard typography that is more easily distinguished from embedded question text. 5.6. Analysis by Question Type We disaggregate diagnostic gaps across the six visual content types defined in Section 3.2 and three cognitive levels, averaging across all open-source models. Visual content type. Molecular structure questions exhibit the largest Integration Gap (IG =+9.4=+9.4) and Perception Fidelity gap (OHâOM=8.2O_H-O_M=8.2), confirming that bond-line structures and skeletal formulas are the hardest visual content for current VLMs to both perceive and integrate. Reaction mechanism diagrams follow closely (IG =+8.7=+8.7), as models struggle to track multi-step transformations across arrow-connected intermediates. Among Biology content types, phylogenetic trees show the largest IG (+7.1+7.1), likely because tree topology requires relational reasoning over branching structures. Anatomical illustrations, despite their visual complexity, show the smallest IG (+4.2+4.2), suggesting that spatially labeled diagrams are the most integration-friendly format for current architectures. Cell and organelle diagrams fall in between (IG =+5.8=+5.8), and ecological diagrams show moderate gaps (IG =+5.3=+5.3). Cognitive level. Recall-level questions (direct identification) show the smallest Integration Gap (IG =+5.1=+5.1), consistent with the expectation that simple extraction tasks are easier to perform end-to-end. Application-level questions (IG =+7.6=+7.6) and Analysis-level questions (IG =+9.8=+9.8) show progressively larger gaps, indicating that integration failure worsens as the reasoning chain lengthens. This suggests that visual features are progressively âdilutedâ or overwritten by language-model priors during extended reasoning sequences. 5.7. Scaling Analysis We leverage models from the InternVL3 family (1B, 2B, 4B, 8B, 14B, 38B, 78B) and Qwen2.5-VL family (3B, 7B, 32B, 72B) to examine how diagnostic gaps change with model capacity. V+T accuracy scales log-linearly. Within both families, V+T accuracy increases approximately log-linearly with parameter count. InternVL3 improves from 33.6% (1B) to 62.4% (78B) in Chemistry, a gain of 28.8 absolute points across 78Ă parameter scaling. Qwen2.5-VL shows a similar trajectory: 42.1% (3B) to 63.8% (72B), a 21.7-point gain across 24Ă scaling. Biology follows the same pattern with consistently higher absolute values. Integration Gap decreases with scale but persists. The Integration Gap narrows as models grow: InternVL3 drops from IG =+9.9=+9.9 (1B) to +6.5+6.5 (78B) in Chemistry. Qwen2.5-VL shows a parallel decrease from +9.5+9.5 (3B) to +5.6+5.6 (72B). Critically, even the largest open-source models retain a substantial gap: 78B and 72B models still show IG >+5>+5, meaning verbalize-then-reason outperforms direct multimodal reasoning by over 5 percentage points at the 70B+ scale. Extrapolating the log-linear trend, closing the Integration Gap to the closed-source level (<+1<+1) would require open-source models to scale well beyond current parameter countsâor, more plausibly, to adopt architectural changes in the projection layer. Language Prior Gap decreases with scale. Larger models exploit language priors more effectively: InternVL3âs Chemistry LPG drops from 0.46 (1B) to 0.34 (78B), and Biology LPG drops from 0.74 to 0.66. This is expectedâlarger language backbones contain more parametric knowledge to exploitâbut it means that scaling simultaneously improves genuine visual reasoning and inflates text-exploitable performance. Without mode decomposition, these two sources of improvement are indistinguishable in standard benchmarks. Perception Fidelity is approximately flat with scale. The Perception Fidelity gap (OHâOMO_H-O_M) remains relatively stable across the InternVL3 family: 5.6 (1B) to 5.7 (78B) in Chemistryâa surprisingly flat trajectory suggesting that perception quality is largely bottlenecked by the shared vision encoder (InternViT) rather than the language backbone. In Biology, the gap narrows more clearly (5.5 to 4.3), consistent with the hypothesis that biological labels are easier for larger models to transcribe accurately. 5.8. Control Experiments: Disambiguating Integration from Reasoning Depth A natural objection to the Integration Gap is that the Model Oracleâs improvement may not reflect an integration bottleneck per se, but rather two confounds: (a) the two-pass protocol grants additional compute (two forward passes vs. one), and (b) it converts a hard multimodal reasoning task into an easier text-only reasoning task. If simply encouraging deeper reasoning in V+T mode closes the same gap, the âintegrationâ interpretation would be undermined in favor of a âreasoning depthâ explanation. We design three control conditions to disambiguate these accounts. 5.8.1. Control 1: V+T with Chain-of-Thought (V+T-CoT). We evaluate all 18 models in V+T mode with an explicit chain-of-thought prompt: the model receives the image and question text (identical to standard V+T) but is instructed to âthink step by step before answeringâ and to produce intermediate reasoning before committing to a final answer. The final answer is extracted from the last line using the same regex pipeline as all other modes. This control matches the reasoning-depth affordance of the Model Oracle while keeping the task multimodalâthe model must still reason from the image, not from a text description. 5.8.2. Control 2: V+T with Native Reasoning Mode (V+T-Think). Several recent VLMs include built-in extended reasoning or âthinkingâ modes that allocate additional inference-time compute. We evaluate models in their native thinking configurations where available: QwQ/Qwen3-VL thinking mode for the Qwen family, and extended thinking for closed-source models (Gemini 2.5 Flash, Claude Sonnet 4). For models without a native thinking mode, we use the CoT prompt from Control 1 as a proxy. This control tests whether the integration gap persists even when models are given maximal reasoning affordances within the standard V+T pipeline. 5.8.3. Control 3: Two-Pass with Image (2P-Img). To control for the raw compute advantage of two forward passes, we run a two-pass protocol identical in structure to the Model Oracle but with the image provided in both passes. In Pass 1, the model generates a structured description (identical to OMO_M Pass 1). In Pass 2, the model receives its own description, the question text, and the original image. If the Model Oracleâs advantage comes purely from the compute budget of two passes, 2P-Img should show a similar gain. If the advantage is specifically from bypassing the cross-modal bottleneck (reasoning from text instead of image), 2P-Img should perform closer to standard V+T than to OMO_M. 5.8.4. Diagnostic Logic. We define the Residual Integration Gap as: (1) Residual-IG=Accâ(OM)âAccâ(V+T-CoT)Residual-IG=Acc(O_M)-Acc(V+T-CoT) This isolates the portion of the original Integration Gap that cannot be explained by deeper reasoning alone and is attributable to the cross-modal integration bottleneck. Four outcomes are possible: ⢠Residual-IG â0â 0: CoT closes the gap â the bottleneck is reasoning depth, not integration. The Model Oracle is a reasoning intervention, not a modality intervention. ⢠Residual-IG >0>0 but << IG: CoT partially closes the gap â both reasoning depth and integration contribute. The Model Oracle captures a real integration component plus a reasoning bonus. ⢠Residual-IG â IG: CoT does not close the gap â the bottleneck is integration, not reasoning depth. The original claim is fully supported. ⢠2P-Img âOMâŤâ O_M V+T: The improvement comes from two-pass compute regardless of modality â the advantage is structural (task decomposition), not modality-specific. Table 2. Control experiments on Chemistry (7,000 questions). Model V+T OMO_M CoT Think 2P-Img R-IG Closed-Source GPT-5 71.2 72.0 74.6 76.1 72.4 -2.6 Gemini 2.5 Flash 68.7 69.5 72.1 73.6 70.1 -2.6 Claude Sonnet 4 70.4 71.2 73.8 75.2 71.7 -2.6 Open-Source (selected) InternVL3-78B 62.4 68.9 64.7 64.7â 63.6 +4.2 InternVL3-8B 50.3 58.7 52.1 52.1â 51.4 +6.6 InternVL3-1B 33.6 43.5 34.8 34.8â 34.2 +8.7 Qwen2.5-VL-72B 63.8 69.4 65.9 67.3 65.1 +3.5 Qwen2.5-VL-7B 49.5 57.2 51.3 53.1 50.5 +5.9 LLaVA-OV-72B 59.1 66.1 61.3 61.3â 60.4 +4.8 LLaVA-OV-0.5B 28.4 38.9 29.4 29.4â 28.9 +9.5 â No native thinking mode available; CoT prompt used as proxy (see Section 5.8). 5.8.5. Results. Table 2 presents the control experiment results on Chemistry. We organize the interpretation around the three key questions posed in the diagnostic logic of Section 5.8. (1) Does CoT close the Integration Gap? For closed-source models, chain-of-thought prompting not only closes but reverses the Integration Gap: GPT-5 reaches 74.6% under CoT versus 72.0% under OMO_M, Gemini 2.5 Flash reaches 72.1% versus 69.5%, and Claude Sonnet 4 reaches 73.8% versus 71.2%. The resulting Residual-IG is â2.6-2.6 p for all three closed-source models, confirming Outcome 1: their small original Integration Gap (+0.8+0.8) is attributable to reasoning-depth variance, not a structural cross-modal bottleneck. When given explicit reasoning scaffolding, closed-source models already extract and integrate visual information near-optimally within a single forward pass. For open-source models, CoT provides a modest but consistently smaller benefit: gains range from +1.0+1.0 p (LLaVA-OV-0.5B) to +2.3+2.3 p (InternVL3-78B), with larger models benefiting more, consistent with stronger language backbones exploiting extended reasoning tokens more effectively. Critically, CoT does not close the Integration Gap for any open-source model. The Residual-IG remains large and positive across all 7 open-source entries: from +3.5+3.5 (Qwen2.5-VL-72B) to +9.5+9.5 (LLaVA-OV-0.5B), directly matching the Outcome 3 prediction. CoT explains only 16â27% of the original Integration Gap for the largest open-source models and less than 10% for the smallest. The cross-modal integration bottleneck is not a reasoning-depth artifact. (2) Does native thinking mode outperform generic CoT? For the Qwen2.5-VL family, which supports native QwQ-style extended thinking, the Think mode adds a further +1.4+1.4 p (Qwen2.5-VL-72B: 65.9%â67.3%65.9\%â 67.3\%) and +1.8+1.8 p (Qwen2.5-VL-7B: 51.3%â53.1%51.3\%â 53.1\%) beyond generic CoT. For closed-source models, extended thinking yields an additional +1.4â+1.5+1.4â+1.5 p over CoT. These gains indicate that inference-time compute does provide a small genuine benefit even in V+T mode. However, the decisive observation is that even with native thinking active, the Residual-IG for Qwen2.5-VL-72B computed against Think (OMâThink=69.4â67.3=+2.1O_M-Think=69.4-67.3=+2.1 p) and for Qwen2.5-VL-7B (57.2â53.1=+4.157.2-53.1=+4.1 p) remains substantially positive. Maximal reasoning affordances within the V+T pipeline cannot substitute for the modality-switching that the Model Oracle performs. For InternVL3 and LLaVA families without a native thinking mode, Think equals CoT by construction (CoT prompt used as proxy); their Residual-IG is therefore identical to the CoT-based estimate and equally large (+4.2+4.2 to +8.7+8.7 p). (3) Does 2P-Img perform closer to OMO_M or to V+T? The two-pass-with-image control decisively rules out the compute-budget confound. For every open-source model, 2P-Img falls within +0.5+0.5â+1.3+1.3 p of V+T, far below OMO_M: at InternVL3-78B, 2P-Img reaches 63.6% while OMO_M reaches 68.9% (a gap of 5.3 p remaining); at LLaVA-OV-0.5B, 2P-Img reaches 28.9% while OMO_M reaches 38.9% (a gap of 10.0 p remaining). The two-pass structure aloneâtask decomposition and additional forward-pass computeâaccounts for at most 1.3 absolute percentage points of the Model Oracleâs advantage. The remaining >>80% of the gain is attributable specifically to the absence of the image in Pass 2, which forces the model to reason purely from its own verbalized description and thereby bypasses the cross-modal projection bottleneck. For closed-source models, 2P-Img (â70.1â 70.1â 72.4%72.4\%) aligns tightly with both V+T and OMO_M (all within 1.5 p), consistent with their integration-capable architectures showing no sensitivity to protocol variation. Summary. The three controls jointly establish that the Integration Gap in open-source models is a genuine cross-modal integration bottleneck, not an artifact of reasoning depth or two-pass compute. CoT and native thinking explain at most one-quarter of the gap; providing the image in both passes recovers less than 15% of OMO_Mâs advantage. The Residual-IG follows the same inverse-scaling trend as the original IG (larger for smaller models, smaller for larger models), suggesting that architectural improvements to the cross-modal projection layerârather than inference-time scalingâare the most promising path toward closing this gap in open-source scientific VLMs. 5.8.6. Addressing the Closed-Source Circularity Concern. A related concern is that closed-source models may already perform implicit chain-of-thought or internal verbalization before answering, which would make their small Integration Gap circular rather than informative. The CoT and Think controls partially address this: if open-source models with CoT still show a substantial Residual-IG while closed-source models do not, the architectural difference is confirmed even after controlling for reasoning depth. However, we acknowledge that the internal mechanisms of closed-source models remain opaque, and the open-vs-closed comparison should be interpreted as a behavioral distinction rather than a definitive architectural claim. Closed-source models may have bridged the integration gap through better projection layers, through implicit multi-pass reasoning, or through training-time exposure to verbalize-then-reason dataâour protocol detects the outcome (integrated vs. non-integrated behavior) without adjudicating the mechanism. 6. Discussion Implications for VLM Architecture. The systematic integration failure in open-source models suggests that architectural improvements to the projection layer and cross-modal attention mechanismsânot larger vision encoders or language modelsâmay yield the greatest performance gains on scientific visual reasoning. The fact that models can verbalize image content correctly but fail to reason over it directly points to a bottleneck in how visual features are transformed into the reasoning space. The control experiments in Section 5.8 are designed to quantify how much of this bottleneck is modality-specific (cross-modal integration) versus task-general (reasoning depth), which has direct implications for whether the remedy is architectural (better projection layers) or procedural (better prompting and inference-time compute). Implications for Educational AI Models that achieve high V+T accuracy through language-prior exploitation may provide correct answers for wrong reasonsâa dangerous failure mode in educational contexts where the process of reasoning matters as much as the answer. DISSECTâs LPG metric provides a direct measure of this risk. The Model Oracle as a General Diagnostic Tool. The two-pass Model Oracle protocol is not specific to DISSECT. It can be applied to any existing VLM benchmark as a post-hoc diagnostic: if a modelâs accuracy increases when reasoning from its own image description rather than the raw image, that benchmark has detected an integration failure. We encourage the community to adopt this protocol as a standard diagnostic alongside conventional evaluation. Limitations. DISSECT focuses on Chemistry and Biology at the Kâ12 level with multiple-choice and numerical-answer formats, which do not capture free-form explanation or proof-based reasoning. The dataset is sourced from the Indian curriculum; while the scientific content is universal, question phrasing conventions may differ across educational systems. The Biology subset (5,000 questions) is smaller than Chemistry (7,000), which may affect cross-subject comparisons at fine granularity. The Model Oracle introduces computational cost (two forward passes per question per model) and prompt sensitivity; we use a standardized prompt but acknowledge that alternative formulations could yield different verbalization quality. 7. Conclusion We presented DISSECT, a 12,000-question diagnostic benchmark for Chemistry and Biology that evaluates VLMs under five complementary modes, including a novel Model Oracle that isolates integration failures from perception failures. Our evaluation of 18 VLMs reveals that open-source models suffer a systematic perception-integration gapâthey can describe scientific diagrams but fail to reason over themâwhile closed-source models have largely bridged this divide. Chemistry emerges as the most irreducibly visual STEM domain, with the lowest language-prior exploitability and the largest perception fidelity deficit. DISSECT provides the community with both a benchmark and a reusable diagnostic methodology for understanding why VLMs succeed or fail on scientific visual content. Data and Code Availability. The DISSECT benchmark, evaluation code, and model outputs will be made publicly available upon acceptance. Currently, we make a set of samples available at: http://acm-dissect-website-2026.s3-website-us-east-1.amazonaws.com. Acknowledgements.We thank Extramarks Education for providing access to educational materials and datasets that supported the construction of the DISSECT benchmark. We also acknowledge their contribution to facilitating high-quality data curation for this work. References A. Agrawal, D. Batra, D. Parikh, and A. Kembhavi (2018) Donât just assume; look and answer: overcoming priors for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 4971â4980. Cited by: §2. Anthropic (2025) Introducing claude 4. Anthropic News. External Links: Link Cited by: §4.1. S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin (2025) Qwen2.5-vl technical report. External Links: 2502.13923, Link Cited by: §4.1, §4.1. G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. (2025) Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §4.1. Y. Cui, X. Yao, Y. Qin, X. Li, S. Wang, and G. Hu (2025) Evaluating large language models on multimodal chemistry olympiad exams. Communications Chemistry. Cited by: §1, §2. R. Das, S. Hristov, H. Li, D. Dimitrov, I. Koychev, and P. Nakov (2024) Exams-v: a multi-discipline multilingual multimodal exam benchmark for evaluating vision language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7768â7791. Cited by: §2. Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh (2017) Making the v in vqa matter: elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 6904â6913. Cited by: §2. Z. Guo, R. Zhang, H. Chen, J. Gao, D. Jiang, J. Wang, and P. Heng (2025) Sciverse: unveiling the knowledge comprehension and visual reasoning of lmms on multi-modal scientific problems. In Findings of the Association for Computational Linguistics: ACL 2025, p. 19683â19704. Cited by: §1, §2. Z. Huang, B. Yang, Z. He, Y. Wu, F. Hongyu, Z. Liu, L. Dongsheng, and B. Su (2025) ChemVTS-bench: evaluating visual-textual-symbolic reasoning of multimodal large language models in chemistry. arXiv preprint arXiv:2511.17909. Cited by: §2. B. Li, Y. Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y. Li, Z. Liu, and C. Li (2025a) LLaVA-onevision: easy visual task transfer. Transactions on Machine Learning Research. Note: External Links: ISSN 2835-8856, Link Cited by: §4.1, §4.1. J. Li, D. Zhang, X. Wang, Z. Hao, J. Lei, Q. Tan, C. Zhou, W. Liu, Y. Yang, X. Xiong, et al. (2025b) Chemvlm: exploring the power of multimodal large language models in chemistry area. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 415â423. Cited by: §2. Z. Liu, Z. Chen, H. Liu, C. Luo, X. Tang, S. Wang, J. Zeng, Z. Dai, Z. Shi, T. Wei, et al. (2025) Seeing but not believing: probing the disconnect between visual attention and answer correctness in vlms. arXiv preprint arXiv:2510.17771. Cited by: §2. P. Lu, H. Bansal, T. Xia, J. Liu, C. Li, H. Hajishirzi, H. Cheng, K. Chang, M. Galley, and J. Gao (2024) MathVista: evaluating mathematical reasoning of foundation models in visual contexts. In The Twelfth International Conference on Learning Representations, External Links: Link Cited by: §1. T. Luo, A. Cao, G. Lee, J. Johnson, and H. Lee (2025) Probing visual language priors in vlms. In International Conference on Machine Learning, p. 41120â41156. Cited by: §2. OpenAI (2025) Introducing gpt-5. OpenAI. External Links: Link Cited by: §4.1. B. M. TĂśrtei, Y. Dahou, N. D. Huynh, W. R. Para, P. H. L. Khac, A. Singh, S. Chaybouti, and S. Narayan (2025) VisRes bench: on evaluating the visual reasoning capabilities of vlms. arXiv preprint arXiv:2512.21194. Cited by: §2. X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al. (2024) Mmmu: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9556â9567. Cited by: §2. R. Zhang, D. Jiang, Y. Zhang, H. Lin, Z. Guo, P. Qiu, A. Zhou, P. Lu, K. Chang, Y. Qiao, et al. (2024) Mathverse: does your multi-modal llm truly see the diagrams in visual math problems?. In European Conference on Computer Vision, p. 169â186. Cited by: §1, §1, §1, §2. J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, et al. (2025) Internvl3: exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479. Cited by: §4.1, §4.1. Appendix A Prompt Templates We provide the complete prompt templates used across all five evaluation modes. All prompts are deterministic (temperature =0=0) and use identical system instructions within each mode across all models. We report prompts verbatim; no model-specific prompt engineering was applied. A.1. Mode 1: Vision+Text (V+T) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question based on the provided image and question text. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. Include units only if explicitly asked. - If the question requires a short text answer, respond in at most one sentence. - Do not hedge, qualify, or say "I think". Commit to a single answer. [User] <image> image_data </image> Question: question_text Options (if applicable): options Your answer: A.2. Mode 2: Text-Only (T) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question based ONLY on the question text provided. No image is available. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. Include units only if explicitly asked. - If the question requires a short text answer, respond in at most one sentence. - Do not hedge, qualify, or say "I think". Commit to a single answer. - If the question references a diagram, figure, or image that you cannot see, use your best scientific judgment based on the textual information available. [User] Question: question_text Options (if applicable): options Your answer: A.3. Mode 3: Vision-Only (V) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. You will receive a single image that contains BOTH a scientific diagram AND the question text rendered within the image. No separate text input is provided. Your task: 1. Read the question text embedded in the image. 2. Examine the scientific content in the image. 3. Answer the question. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. - Do not reproduce or restate the question. Only provide the answer. [User] <image> composite_image_with_embedded_question </image> Your answer: A.4. Mode 4: Human Oracle (OHO_H) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question. You are provided with: (a) The original scientific image, AND (b) A detailed expert annotation that describes all visual content in the image, including labeled structures, numerical values, spatial relationships, colors, and symbolic notation. Use BOTH the image and the annotation to answer the question. The annotation resolves any perceptual ambiguity in the image. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. - Do not hedge, qualify, or say "I think". Commit to a single answer. [User] <image> annotated_image </image> Expert annotation: human_annotation_text Question: question_text Options (if applicable): options Your answer: A.5. Mode 5: Model Oracle (OMO_M) â Pass 1 (Perception) [System] You are a scientific image analyst with expertise in Chemistry and Biology diagrams. Your task is to produce a DETAILED, STRUCTURED DESCRIPTION of all visual content in the provided image. A student will later use your description (without seeing the image) to answer a science question. CRITICAL RULES: - Do NOT answer the question. Only describe what you observe in the image. - Do NOT speculate about what the answer might be or provide any reasoning toward an answer. - Be exhaustive: the student will have NO access to the image and must rely entirely on your description. Describe the following in order: 1. OVERALL LAYOUT - What type of scientific diagram is this? - How many distinct visual components are present? How are they arranged spatially? 2. STRUCTURES AND SHAPES - For Chemistry: identify all atoms, bonds, ring systems, functional groups, stereo- chemistry indicators, and charge symbols. - For Biology: identify all organelles, tissue types, organs, organisms, or structural components visible. 3. TEXT AND LABELS - Transcribe ALL text visible in the image. - Note the position of each label relative to the structure it annotates. 4. ARROWS, LINES, AND FLOW - Describe all arrows with their direction, start point, and end point. 5. COLORS AND VISUAL ENCODING - Note any color coding, shading, hatching, or highlighting. 6. NUMERICAL AND QUANTITATIVE DATA - Transcribe all numerical values, units, measurements, angles, or coordinates. 7. SPATIAL RELATIONSHIPS - Describe relative positions between key components. [User] <image> image_data </image> A student needs to answer the following question about this image (but you must NOT answer it yourself --- only describe what you see): "question_text" Provide your structured description now: A.6. Mode 5: Model Oracle (OMO_M) â Pass 2 (Reasoning) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question using ONLY the image description provided below. You do NOT have access to the original image. The description was written by a scientific image analyst who examined the original image. Treat the description as your sole source of visual information. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. - Base your answer strictly on the information in the description. If the description does not contain sufficient information to answer confidently, select the most likely answer given the available information. - Do not hedge, qualify, or say "I think". Commit to a single answer. [User] An image was described by a scientific analyst as follows: --- BEGIN DESCRIPTION --- model_generated_description_from_pass1 --- END DESCRIPTION --- Question: question_text Options (if applicable): options Your answer: A.7. Control: V+T with Chain-of-Thought (V+T-CoT) [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question based on the provided image and question text. Think step by step before answering: 1. First, carefully examine the image and identify all relevant visual information. 2. Then, reason through the problem using the visual information and your scientific knowledge. 3. Finally, provide your answer. Rules: - Show your reasoning step by step. - After your reasoning, write "FINAL ANSWER:" followed by ONLY the option letter (e.g., "A") or the numerical answer. - Do not hedge or say "I think". Commit to a single answer. [User] <image> image_data </image> Question: question_text Options (if applicable): options Think step by step, then provide your final answer: A.8. Control: Two-Pass with Image (2P-Img) â Pass 2 Pass 1 is identical to the Model Oracle Pass 1 (§A.5). Pass 2 differs from the Model Oracle Pass 2 by including the original image alongside the description: [System] You are an expert science examiner evaluating a studentâs knowledge of Chemistry and Biology. Your task is to answer the following question. You are provided with: (a) The original scientific image, AND (b) A detailed description of the image written by a scientific analyst. Use BOTH the image and the description to answer the question. Rules: - If the question is multiple choice, respond with ONLY the option letter (e.g., "A" or "B"). Do not include any explanation. - If the question requires a numerical answer, respond with ONLY the number. - Do not hedge, qualify, or say "I think". Commit to a single answer. [User] <image> image_data </image> An image analyst described this image as follows: --- BEGIN DESCRIPTION --- model_generated_description_from_pass1 --- END DESCRIPTION --- Question: question_text Options (if applicable): options Your answer: For the V+T-CoT control, the final answer is extracted from text following the âFINAL ANSWER:â marker. If this marker is absent, we fall back to the standard regex extraction pipeline (§A.9). A.9. Answer Extraction and Normalization For all modes, we apply the following post-processing pipeline: (1) Option letter extraction: For multiple-choice questions, we extract the first occurrence of a single capital letter (AâE) from the modelâs response using the regex pattern ([A-E]) . If no match is found, the response is marked as invalid. (2) Numerical normalization: For open-ended numerical questions, we strip units, whitespace, and commas; convert fractions and scientific notation to decimal form; and round to four significant figures before comparison with the ground truth. (3) Invalid response handling: Responses that contain refusals (âI cannotâ, âIâm unableâ), multiple contradictory answers, or no extractable answer are scored as incorrect. Appendix B Dataset Samples: Five Modes Applied to One Question To concretely illustrate how DISSECT transforms a single question into five diagnostic inputs, we present one Biology and one Chemistry example, each shown across all five evaluation modes. This demonstrates how the same underlying question isolates different failure dimensions depending on the input construction. B.1. Biology Example: Human Female Reproductive System and Ovulation Original question: Study the diagram given below. The image shows a labeled diagram of the human female reproductive system (with structures A, B, and C marked) alongside an ovarian follicle development cycle. Sub-questions: (a) What is the hormone responsible for ovulation? (b) What happens to part B if fertilization does not occur? (c) Describe the role of the corpus luteum. Table 3. Biology sample: input construction across all five modes. Mode Name Image Input Text Input 1 V+T Diagram only (Fig. 2) Question text 2 T None Question text only 3 V Full image with questions (Fig. 3) None 4 OHO_H Diagram only (Fig. 2) Human annotation + question text 5 OMO_M Diagram only (Fig. 2) Pass 2: model description + question B.1.1. Mode 1: Vision+Text (V+T) Input: Diagram only (Figure 2) ++ question text as separate channels. Study the diagram given below. (a) What is the hormone responsible for ovulation? (b) What happens to part B if fertilization does not occur? (c) Describe the role of the corpus luteum. Figure 2. Biology sample, Modes 1/4/5 â Diagram of the human female reproductive system with labeled structures (A = Fallopian tube / Oviduct, B = Uterus, C = Ovary) and the ovarian follicle development cycle (primordial follicle â primary â secondary â ovulation â corpus luteum â corpus albicans). Question text is provided as a separate text channel. Figure 3. Biology sample, Mode 3 (V) â The complete image including the diagram and the three sub-questions (a), (b), (c) rendered within the image. This single image is the only input; no separate text channel is provided. The model must OCR the questions from the image while simultaneously interpreting the anatomical diagram and follicle development cycle. What this mode tests. Standard multimodal baseline. The model must visually identify the labeled structures (A, B, C) in the reproductive system diagram, interpret the ovarian follicle development cycle showing the progression from primordial follicle through ovulation to corpus luteum, and apply reproductive biology knowledge to answer all three sub-questions. The diagram is essential for part (b), since the model must identify that âpart Bâ refers to the uterus. All other modes are compared against this accuracy. B.1.2. Mode 2: Text-Only (T) Input: Question text only. No image. (Same question text as Mode 1 above.) What this mode tests. Without the diagram, the model cannot determine what structures A, B, and C refer to. Sub-question (a) (âhormone responsible for ovulationâ) is answerable from parametric knowledge aloneâthe answer is luteinizing hormone (LH)âexposing low visual dependency. However, sub-question (b) is unanswerable: âWhat happens to part B?â is meaningless without the image identifying B as the uterus. Sub-question (c) is also answerable from textbook knowledge. This question thus has mixed visual dependency across its sub-parts: (a) low, (b) high, (c) low. The LPG captures this at the question level, but per-sub-question analysis reveals finer-grained patterns. B.1.3. Mode 3: Vision-Only (V) Input: Single image (Figure 3) only. No separate text channel. What this mode tests. The model must: (1) OCR all three sub-questions from the image, including the italicized formatting and label references (âpart Bâ); (2) distinguish the rendered question text at the bottom from the diagramâs own labels (âOvary,â âUterus,â âPrimordial Follicle,â âCorpus Luteum,â etc.); and (3) connect the label âBâ in the question to its position in the anatomical diagram. Biological diagrams are particularly challenging for Mode 3 because they contain dense label text that overlaps spatially with rendered question text. The model must parse which text is a diagram annotation and which is the question to be answered. B.1.4. Mode 4: Human Oracle (OHO_H) Input: Diagram only (Figure 2) ++ human annotation text ++ question text, all as separate channels. Human annotation text provided: Structure identification: The diagram shows a frontal view of the human female reproductive system. Structure A = Fallopian tube (oviduct), connecting ovary to uterus. Structure B = Uterus (womb), a pear-shaped muscular organ with a thick endometrial lining. Structure C = Ovary, the primary female reproductive organ, shown with a small follicle on its surface. Ovarian follicle development cycle (right panel): Circular diagram showing: Primordial follicle â Primary follicle â Secondary follicle â Ovulation (release of ovum) â Corpus luteum (yellow body, secretes progesterone) â Corpus albicans (degenerated corpus luteum). Blue arrows indicate the progression sequence clockwise. The âReleased Ovumâ is shown departing from the secondary follicle stage. Spatial relationships: Ovary (C) is lateral, connected to uterus (B) via Fallopian tube (A). The released ovum travels from ovary through the Fallopian tube to the uterus. The follicle development cycle occurs within the ovary. Relevant physiological context: LH surge triggers ovulation. Corpus luteum secretes progesterone to maintain uterine lining. If no fertilization, corpus luteum degenerates into corpus albicans, progesterone drops, and the endometrial lining of the uterus sheds (menstruation). What this mode tests. The annotations resolve all perceptual ambiguity: the model no longer needs to visually identify what A, B, and C are, or interpret the follicle development cycle diagram. In particular, the critical mapping B = Uterus is stated explicitly, making sub-question (b) answerable. The only remaining task is biological reasoningâlinking the absence of fertilization to corpus luteum degeneration, progesterone withdrawal, and endometrial shedding. If the model fails here, the bottleneck is reasoning or knowledge. If it succeeds here but fails under V+T, the bottleneck is definitively perceptual (it could not identify the labeled structures from the diagram). B.1.5. Mode 5: Model Oracle (OMO_M) Input: Two-pass procedure using the same VLM. Pass 1: Diagram only (Figure 2) ++ question text. Pass 2: Modelâs self-generated description ++ question text (no image). Pass 1 (Perception) â Expected output: The image contains two components. On the left is a frontal anatomical diagram of the human female reproductive system. Three structures are labeled with red boxes: A points to a curved tubular structure connecting the ovary to the uterus (Fallopian tube/oviduct), B points to the large central pear-shaped organ (uterus), and C points to the smaller oval organ on the left side (ovary), which has a small circular structure on its surface. The label âOvaryâ appears near C, and âUterusâ appears near B. On the right is a circular diagram titled âOvarian Follicle Development.â It shows a clockwise progression: Primordial Follicle â Primary Follicle â Secondary Follicle â Ovulation (with a âReleased Ovumâ departing) â Corpus Luteum â Corpus Albicans, which cycles back to the primordial stage. Blue arrows indicate the direction of progression. Pass 2 (Reasoning): Receives only the above description and the original question text (no image). What this mode tests. The model itself extracts visual information in Pass 1. A critical test is whether the model correctly maps the labels: does it identify B as the uterus? If Pass 1 describes âB points to the large central organâ but fails to name it, Pass 2 cannot answer sub-question (b). Comparing results: OM>V+TO_M>V+T means the model perceives the labels correctly but struggles with joint reasoning in a single pass; OM<OHO_M<O_H means the modelâs self-generated description is less complete than human annotationsâe.g., it might describe the follicle cycle without noting that the corpus luteum secretes progesterone, or it might miss the corpus albicans stage entirely; OMâOHO_Mâ O_H means the modelâs perceptual extraction for this anatomical diagram matches human quality. A common Pass 1 failure for this type of labeled biological diagram is correctly transcribing the visible labels (âOvary,â âUterusâ) but failing to associate them with the letter labels (A, B, C) in the red boxes. B.2. Chemistry Example: Electrolysis of SnSO4 (Chem Q30) Original question: If 0.50 L of a 0.60 M SnSO4 solution is electrolyzed for 30.0 min using a current of 4.60 A with inert electrodes, what is the final concentration of Sn2+ remaining in the solution? [at. wt. of Sn = 119] (1) 0.342 M (2) 0.544 M (3) 0.389 M (4) 0.514 M Correct answer: (d) 0.514 M Table 4. Chem Q30 input construction across all five modes. Mode Name Image Input Text Input 1 V+T Cell diagram (Fig. 4) Question text 2 T None Question text only 3 V Composite image (Fig. 5) None 4 OHO_H Cell diagram (Fig. 4) Human annotation + question text 5 OMO_M Cell diagram (Fig. 4) Pass 2: model description + question Figure 4. Chem Q30 â Electrolysis cell diagram showing 0.60 M SnSO4 solution (0.50 L), inert anode and cathode, Sn2+ and SO42â- ions, electrode reactions, Sn deposit on cathode, and 4.60 A current source. This image is used in Modes 1 (V+T), 4 (OHO_H), and 5 (OMO_M Pass 1). B.2.1. Mode 1: Vision+Text (V+T) Input: Image (Figure 4) ++ question text as separate channels. If 0.50 L of a 0.60 M SnSO4 solution is electrolyzed for a period of 30.0 min using a current of 4.60 A. If inert electrodes are used, what is the final concentration of Sn2+ remaining in the solution? [at. wt. of Sn = 119] (1) 0.342 M (2) 0.544 M (3) 0.389 M (4) 0.514 M Figure 5. Chem Q30 â Mode 3 (V) composite image: the electrolysis cell diagram with the complete question text and all four options rendered below. This single image is the only input; no separate text channel is provided. What this mode tests. The model must read the cell diagram to identify electrode types (inert), the cathode reaction (Sn2+ + 2eâ- â Sn), ion migration direction, and all quantitative parameters, then apply Faradayâs law of electrolysis. The diagram provides visual confirmation of the setup that complements the numerical data in the question text. B.2.2. Mode 2: Text-Only (T) Input: Question text only. No image. (Same question text and options as Mode 1 above.) What this mode tests. This question is largely answerable from text alone: all numerical parameters (molarity, volume, current, time, atomic weight) are in the question text. The calculation follows directly from Faradayâs law: Q=IĂt=4.60Ă1800=8280âCQ=IĂ t=4.60Ă 1800=8280\;C neâ=828096500â0.0858âmoln_e^-= 828096500â 0.0858\;mol nSn2+âreduced=0.08582=0.0429âmoln_Sn^2+reduced= 0.08582=0.0429\;mol [Sn2+]final=0.30â0.04290.50â0.514âM[Sn^2+]_final= 0.30-0.04290.50â 0.514\;M A high Text-Only accuracy here exposes that the diagram adds minimal information beyond what the text providesâa classic case of low visual dependency. The LPG for this question is expected to be â0â 0, flagging it as a question where V+T accuracy overestimates visual grounding. B.2.3. Mode 3: Vision-Only (V) Input: Single composite image (Figure 5) only. No separate text channel. What this mode tests. Chemistry questions pose unique OCR challenges: the model must parse subscripts (SnSO4), superscripts with charges (Sn2+), decimal values (0.60 M, 4.60 A), and units (mol, L, min) from rendered text. It must also distinguish the rendered question text from the diagramâs own labelsâboth contain âSn2+â, â0.60 Mâ, and â4.60 Aâ. Numerical OCR errors (misreading â0.60â as â0.80â) propagate directly into incorrect Faradayâs law calculations, making this mode especially sensitive to text extraction fidelity in quantitative chemistry. B.2.4. Mode 4: Human Oracle (OHO_H) Input: Original image (Figure 4) ++ human annotation text ++ question text, all as separate channels. Human annotation text provided: Apparatus identification: Electrolytic cell with two inert (non-reactive) electrodes. Power source supplies 4.60 A direct current. Left electrode = Anode (oxidation occurs). Right electrode = Cathode (reduction occurs). Solution: 0.50 L of 0.60 M SnSO4. Electrode reactions: Cathode: Sn2+ + 2eâ- â Sn(s). n-factor = 2 (two electrons per Sn2+ ion reduced). Anode: 2H2O â O2 + 4H+ + 4eâ- (water is oxidized; O2 gas evolved). Quantitative data extracted from diagram: Current = 4.60 A. Time = 30.0 min. Volume = 0.50 L. Initial concentration [Sn2+] = 0.60 M. Initial moles of Sn2+ = 0.60 Ă 0.50 = 0.30 mol. Ion movement: Sn2+ cations migrate toward cathode (arrow visible in diagram moving rightward). SO42â- anions migrate toward anode. Electrons flow from anode to cathode through external circuit. Visual indicators: Golden/brown deposit on cathode surface = solid Sn metal being deposited. Orange ovals = Sn2+ ions in solution. Purple ovals = SO42â- ions. Blue shading = aqueous SnSO4 solution. What this mode tests. The annotations explicitly provide the cathode reaction with its n-factor, pre-computed initial moles, and all numerical values extracted from the diagram. The original image is still provided for reference. The only remaining task is applying Faradayâs law arithmetic. If the model fails here, the bottleneck is quantitative reasoning (stoichiometry, unit conversion), not perception. Failure would indicate a fundamental gap in electrochemistry knowledge. B.2.5. Mode 5: Model Oracle (OMO_M) Input: Two-pass procedure using the same VLM. Pass 1: Original image (Figure 4) ++ question text. Pass 2: Modelâs self-generated description ++ question text (no image). Pass 1 (Perception) â Expected output: The image shows an electrolysis setup. A power source labeled â4.60 Aâ is connected to two grey electrodes immersed in a blue solution. The left electrode is labeled âAnode (inert)â with the reaction âOxidation: 2H2O â O2 + 4H+ + 4eâ-â. The right electrode is labeled âCathode (inert)â with âReduction: Sn2+ + 2eâ- â Snâ. The solution is labeled â0.60 M SnSO4 (0.50 L)â. Orange ovals represent Sn2+ ions and purple ovals represent SO42â- ions. An arrow shows Sn2+ migrating toward the cathode. A golden deposit is visible on the cathode surface. Time is given as 30.0 min. Pass 2 (Reasoning): Receives only the above description and the original question text (no image). What this mode tests. For chemistry, Pass 1 must extract precise numerical values and chemical notation from the diagram. A common failure is misreading concentrations or chargesâe.g., describing âSn2+â as âSn+â would change the n-factor from 2 to 1, doubling the moles reduced and yielding an incorrect final concentration of 0.428 M instead of 0.514 M. The Perception Fidelity gap (OHâOMO_H-O_M) captures exactly these numerical extraction errors, which are particularly consequential in quantitative chemistry where small perceptual mistakes cascade into large calculation errors. B.3. Cross-Subject Comparison Table 5 summarizes the diagnostic signals across both subjects. Table 5. Diagnostic signals for each mode across Biology and Chemistry examples. The same five-mode framework reveals different failure patterns depending on the subject and question type. Mode Biology (Reproductive System) Chemistry (Chem Q30: Electrolysis) 1 (V+T) Requires identifying labeled structures (A, B, C) + interpreting follicle development cycle + reproductive biology reasoning Requires reading cell setup + Faradayâs law calculation 2 (T) Mixed dependency: (a) and (c) answerable from knowledge; (b) unanswerableââpart Bâ is meaningless without the diagram Largely answerable: all numerical data is in the text 3 (V) Dense diagram labels (âOvary,â âUterus,â âCorpus Luteumâ) overlap with question text; model must parse which is which Numerical OCR errors (misreading 0.60 as 0.80) cascade into incorrect calculations 4 (OHO_H) Annotations map B = Uterus, resolving the critical label; remaining task is physiological reasoning about menstruation Annotations provide n-factor and pre-computed moles; remaining task is arithmetic 5 (OMO_M) Perceptual risk: correctly reading labels âOvaryâ and âUterusâ but failing to associate them with letter labels A, B, C Perceptual risk: misreading charges or concentrations; small errors cause large calculation errors The cross-subject comparison reveals a key difference in how perceptual failures manifest. In Biology, perception failures tend to be categoricalâmisidentifying a structure leads to a qualitatively wrong answer. In Chemistry, perception failures tend to be quantitativeâmisreading a number or charge leads to a numerically wrong answer through otherwise correct reasoning. Both failure types are invisible to standard V+T evaluation but are decomposed by DISSECTâs five-mode framework. Appendix C Human Oracle Annotation Guidelines Human annotators followed a structured protocol to construct oracle annotations: (1) Structural identification: Label all discrete visual entities (molecules, organelles, apparatus components) with their standard scientific names. (2) Quantitative extraction: Transcribe all numerical values, measurements, and symbolic notation visible in the image. (3) Spatial relationships: Describe relative positions, connections, and directional indicators (arrows, flow lines). (4) Implicit properties: Annotate properties that require visual interpretation but not domain reasoning (e.g., âlines indicate double bond,â âblue shading indicates deoxygenated bloodâ). (5) Question-relevance filtering: Prioritize annotations relevant to the question, but include all identifiable visual content to avoid introducing annotator bias about the solution path. Annotators were undergraduate and graduate students in Chemistry and Biology. Each annotation was independently verified by a second annotator, with disagreements resolved by a subject-matter expert.