Paper deep dive
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/19/2026, 5:33:59 AM
Summary
The paper introduces BEAR-Bench, a bilingual (English and Russian) benchmark designed to evaluate Multimodal Large Language Models (MLLMs) on complex, multi-step reasoning tasks involving text-dense professional documents. Unlike existing benchmarks that focus on information extraction or require external knowledge, BEAR-Bench emphasizes self-contained reasoning over business and scientific documents. The authors evaluate 16 proprietary and open-weight MLLMs, finding significant performance gaps, particularly for Russian language tasks. Additionally, the benchmark is used to compare various hallucination detection methods.
Entities (8)
Relation Signals (6)
Qwen3.5-397B-A17B â evaluatedon â BEAR-Bench
confidence 98% · We evaluate 16 proprietary and open-weight MLLMs, including ... Qwen3.5-397B ... on BEAR-Bench
Gemini 3.1 Pro â evaluatedon â BEAR-Bench
confidence 98% · We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro ... on BEAR-Bench
BEAR-Bench â supportslanguage â English
confidence 95% · BEAR-Bench ... comprising 1000 human-annotated questions based on text-rich business and scientific documents ... English-and-Russian benchmark
BEAR-Bench â supportslanguage â Russian
confidence 95% · BEAR-Bench ... comprising 1000 human-annotated questions based on text-rich business and scientific documents ... English-and-Russian benchmark
Chain-of-Thought â appliedto â BEAR-Bench
confidence 90% · We tested whether explicit chain-of-thought (CoT) prompting improves accuracy on BEAR-Bench.
BEAR-Bench â usedfor â Hallucination Detection
confidence 90% · Finally, we use the resulting model outputs to compare existing hallucination detection methods
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified.
Tags
Links
- Source: https://arxiv.org/abs/2608.17895v1
- Canonical: https://arxiv.org/abs/2608.17895v1
Trouble viewing inline? Open PDF directly â
Full Text
73,835 characters extracted from source content.
Expand or collapse full text
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models Liubov Chubarova Alexandra Kuleshova Affiliation: Yandex Applied AI Institute Correspondence:bazarovaai.239@gmail.com Daniil Volkov Affiliation: Yandex Applied AI Institute Correspondence:bazarovaai.239@gmail.com Kirill Sultanov Alexey Zaytsev Affiliation: Yandex Applied AI Institute Correspondence:bazarovaai.239@gmail.com Abstract While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-dense, professional documents remains incompletely evaluated. Existing benchmarks emphasize information extraction, require external domain knowledge, or cover professional documents only as one of many settings. They are also largely English- or Chinese-centric, leaving other languages and Russian, in particular, substantially underrepresented. To address these limitations, we introduce BEAR-Bench (Bilingual Enterprise and Academic Reasoning), a self-contained, complex English-and-Russian benchmark comprising 1000 human-annotated questions based on text-rich business and scientific documents. We evaluate 16 proprietary and open-weight MLLMs, including Gemini 3.1 Pro and Qwen3.5-397B, on BEAR-Bench and observe clear headroom even for the strongest systems. Finally, we use the resulting model outputs to compare existing hallucination detection methods, evaluating not only how often models fail on BEAR-Bench but also how reliably those failures can be identified. **footnotetext: Equal contribution. 1 Introduction Multimodal Large Language Models (MLLMs) 20; 27; 1 have revolutionized the machine comprehension of images. Beyond merely solving optical character recognition (OCR) tasks, these models demonstrate the potential to perform complex reasoning over multimodal inputs 35; 12. This capability is particularly vital for text-rich scenarios, which are central to real-world professional domains that require reading and analyzing dense visual documents, such as scientific papers and financial reports. Existing benchmarks provide complementary but incomplete coverage of professional document reasoning. Document-oriented datasets such as DocVQA 18 primarily emphasize OCR and information extraction, whereas broad multimodal reasoning benchmarks such as MMMU 35 often require specialized factual knowledge. OCR-Reasoning 13 covers diverse text-rich images, including some professional documents, but does not examine this setting in depth. A further limitation is linguistic coverage: existing multimodal benchmarks exhibit a strong bias toward English and Chinese, leaving Slavic languages â and Russian in particular â substantially underrepresented. Russian-inclusive resources such as MTVQA 25, MWS Vision Bench 19, and MERA-Multi 8 broaden this coverage, but mix documents with other image types, emphasize OCR and document processing, or focus on a narrow document category. To our knowledge, no existing benchmark evaluates self-contained, multi-step reasoning over Russian-language professional document pages whose items involve textual and graphical content such as figures, tables, charts, equations, and diagrams. To address these limitations, we introduce BEAR-Bench, a complex benchmark comprising 1000 human-annotated questions for text-dense enterprise and scientific documents in English and Russian languages. The science section tests the ability to interpret academic figures, transcribe mathematical formulas, and analyze data from plots. The enterprise tasks target data cross-referencing in financial reports and logical reasoning over business charts. The questions follow two design principles. First, they require multiple logical steps rather than direct extraction from a single piece of evidence. Second, they are fully answerable from the document alone, without external expert knowledge. Together, these principles focus the evaluation on document-grounded reasoning rather than factual recall or shallow extraction. The inclusion of Russian-language tasks further broadens evaluation beyond predominantly English- and Chinese-language resources. Deploying MLLMs in professional settings requires knowing not only how often they fail, but also whether those failures can be detected reliably; yet evidence on hallucination detection for reasoning over text-dense professional documents remains incomplete. We address this gap by using BEAR-Bench to compare token-level uncertainty scores, representation-based detectors, supervised hidden-state probes, and MLLM-as-a-judge methods across proprietary and open-weight models, providing a practical comparison for deployment settings with and without access to model internals. The main contributions of this work are the following: 1. We propose BEAR-Bench, a bilingual benchmark for multimodal reasoning in professional scenarios. It features human-annotated questions targeting text-dense business and scientific documents in English and Russian languages. 2. We evaluate 16 MLLMs on BEAR-Bench, including proprietary (e.g., Gemini-3.1-Pro, Claude Opus 4.6) and open-weight (e.g., Qwen3.5, gemma-4) models. Our results show that even the strongest evaluated systems leave clear headroom on BEAR-Bench. 3. We further use BEAR-Bench to evaluate a diverse set of existing hallucination detection methods for OCR-intensive professional-document reasoning, comparing uncertainty-, representation-, and judge-based methods under two deployment regimes: direct access to internal signals for open-weight models and proxy-based detection for proprietary ones. Benchmark #Langs #QA pairs #RU reasoning VQA Image scope OCR chars/img DocVQA 18 EN 5.2K n/a Industry documents 1,113.0 ChartQA 17 EN 2.5K* n/a Charts 231.8 CharXiv 31 EN 11.6K n/a Scientific charts 165.1 OCRBench v2 10 EN, ZH 10K n/a Mixed text-rich images 437.2 OCR-Reasoning 13 EN 1.1K n/a Everyday text-rich scenes 514.3 MWS Vision Bench 19 RU 1.3K** 400 Business/personal documents 1,126.4 LabTabVQA 8 RU 349 349 Medical report tables 637.1 C-OCR V2 33 RU + 31 langs 2K*** 0 Finance/dashboards/blueprints 1,149.6 BEAR-Bench EN, RU 1K 618 Science/business documents 2,740.6 Table 1: Comparison of BEAR-Bench to existing multimodal reasoning benchmarks. #QA pairs refers to the evaluation split; DocVQA and ChartQA additionally provide training data (50K and 32.7K items in total, respectively), while the remaining benchmarks are evaluation-only. For C-OCR V2, the reported count is the document QA track (2K of 7.1K items). #RU reasoning VQA: number of evaluation questions in Russian that require answering from the image, excluding OCR, parsing, grounding, and key-information extraction; n/a = not applicable (no Russian split). Image scope summarizes the principal visual sources represented in each benchmark. OCR chars/img: average number of non-whitespace OCR characters per image (n=200n=200 randomly sampled images per benchmark). *Evenly split between human-written and machine-generated questions. **Publicly available.. 2 Related Work Multimodal benchmarks for professional documents. While modern LLMs achieve strong results on many established multimodal benchmarks 36; 41, their ability to operate with visually rich professional documents â which requires analyzing both textual and graphical content of an image â- remains underexplored. Text-dense benchmarks oriented at professional documents are mostly OCR-based and measure extractive skills rather than cross-referencing of text and visuals 33; 18, or emphasize long-page settings that conflate multimodal reasoning with long-context handling 7; 26. Multimodal reasoning-focused benchmarks, conversely, offer little signal specifically for professional documents: OCR-Reasoning 13 and OCRBench v2 10 are deliberately broad â valuable for general-purpose evaluation, but treating professional documents as one setting among many â while CharXiv 31 and ChartQA 17; 16 restrict evaluation to charts and thus do not test textâvisual cross-referencing. A further issue is dependence on external knowledge: MMMU 35 and EMMA 12 pose multidisciplinary problems that presuppose domain expertise, so scores conflate multimodal reasoning failures with factual gaps; part of OCR-Reasoning shares this confound. Language coverage is also narrow: most benchmarks covering multimodal reasoning in text-dense scenarios are available only in English or Chinese, while many other languages, including Russian, remain underrepresented. Russian-inclusive benchmarks provide valuable but partial coverage. MTVQA 25 and TIU-Bench 39 include documents alongside natural scenes but offer very limited coverage of Russian-language professional documents; TIU-Bench, for example, contains only 10 Russian document samples in total. MWS Vision Bench 19 provides 400 Russian reasoning-VQA items, but its images span business scans, personal handwriting, receipts, and form-style pages (Figure 5), rather than being selected specifically for reasoning over text-dense professional documents. LabTabVQA â the only subset of MERA-Multi 8 built on professional documents rather than natural images or exam-style problems â is restricted to tables from medical laboratory reports. Error detection for text-dense multimodal inputs. Recent work proposes a range of hallucination detectors for MLLMs 6. Tool-augmented methods rely on auxiliary models 5; 34; 24, while multi-query methods require repeated generation or verification calls 32; 40, increasing deployment cost. Of particular interest are lightweight white-box methods, which detect errors using token uncertainty or internal model states collected during a single forward pass 30; 15; 14; 38. These methods add relatively little inference overhead when model internals are available, yet, to our knowledge, they have not been systematically compared on visually rich professional document images requiring multi-step reasoning across dense textual and graphical evidence. Summary. Existing work lacks a Russian-inclusive benchmark for self-contained, multi-step reasoning over text-dense professional documents. BEAR-Bench is designed to fill this gap using scientific and business documents; Table 1 summarizes the comparison. Furthermore, BEAR-Bench enables a systematic comparison of hallucination detectors on such text-dense professional document images requiring multi-step reasoning, a setting not covered by prior evaluations. 3 BEAR-Bench 3.1 Domain Scope and Taxonomy BEAR-Bench spans two primary domains â Business and Science â each subdivided into thematically coherent sub-categories. Business Domain. The business subset covers three document types: financial reports (SEC Forms 10-K and 10-Q) requiring tabular reasoning and year-over-year calculations; investor presentations (Form 8-K exhibits) combining charts, KPI tiles, and infographic maps; and flowcharts and organisational diagrams depicting corporate ownership structures and process pipelines. Science Domain. The science subset covers three categories: mathematical and physical formulae from physics and mathematics preprints targeting symbol-level recognition; scientific figures and plots (line plots, scatter diagrams, heatmaps) requiring axis and legend interpretation; and academic layouts with multi-column pages testing reading-order resolution and cross-referential reasoning. The two domains are strictly disjoint: no source document appears in both subsets. 3.2 Data Collection and Annotation Pipeline 3.2.1 Source Collection Business Domain. Business documents were retrieved via targeted Google Search queries directed at publicly accessible, license-safe sources. English-language documents were obtained from the U.S. Securities and Exchange Commission (SEC) EDGAR system â annual reports (Form 10-K), quarterly reports (Form 10-Q), and investor presentations filed as Form 8-K exhibits â all of which constitute public records under U.S. federal law. Supplementary English documents were drawn from official government portals (*.gov, *.gov.uk) and intergovernmental repositories (*.int). Russian-language documents were sourced from the state corporate-disclosure platforms e-disclosure.ru and moex.com, as well as from federal government domains (*.gov.ru). An automated scraper retrieved candidate PDFs; each document underwent a programmatic license-verification step examining the first and last five pages for SEC registration markers or open-license declarations (âCreative Commonsâ, âC BYâ, âpublic domainâ). Documents failing this check were discarded prior to further processing. Science Domain. English-language papers were downloaded from arXiv via its official Python API, sampling four STEM categories: quant-ph, cs.AI, eess.SP, and math.GM (up to 100 papers per category). Russian-language articles were collected from CyberLeninka (cyberleninka.ru) using an asynchronous Playwright-based crawler across four subject areas: Computer Science, Mathematics, Physics, and Engineering (up to 100 articles per category). 3.2.2 Filtering and Preprocessing Raw PDFs were rendered page-by-page into PNG images and processed through a two-stage filtering pipeline. Stage 1 â Visual Content Classification. We obtained silver labels for a stratified sample of 3,000 images using Gemini 2.5 Pro with a structured multi-label prompt, producing six Boolean fields: contains_diagrams, contains_tables, contains_equations, contains_code, contains_figures, and contains_handwriting. These labels trained a lightweight classifier: SigLIP embeddings (37) were L2-normalised and passed to a MultiOutputClassifier of logistic-regression models with balanced class weights, one per label. After validation on a held-out 20% split, the classifier was applied to the full corpus of ⌠66,000 page images, retaining only pages with at least one of contains_diagrams, contains_equations, contains_code predicted positive. Stage 2 â Textual Density Filtering. Among content-positive pages, we retained only those at or above the 67th percentile of OCR character count within their respective language group. From the surviving candidates, up to 1,500 images per language were drawn via stratified random sampling (seed = 42=\,42), yielding the final pool submitted to human annotators. Figure 1: Representative items from BEAR-Bench. Each card displays the source document image (left) alongside the human-authored multi-hop question and ground-truth answer (right). 3.2.3 Human Annotation Annotator pool. Annotation was conducted by 13 domain experts, each holding at minimum a Bachelorâs degree in a technical discipline. Every annotator processed its own subset of images, authoring exactly one questionâanswer pair per image. Annotation task. For each image, annotators were required to: (i) re-verify the visual-content labels produced by the automatic classifier, correcting any erroneous predictions; and (i) compose a multi-step, multi-hop question with a detailed ground-truth answer. Questions were required to elicit compositional reasoning â aggregating values across table rows, interpreting plotted trends in the context of equations, or tracing paths through flowcharts â rather than straightforward single-step extraction. The questions were additionally assigned a reasoning depth score representing the total number of reasoning and computational steps required to arrive at the correct answer. The text of the instruction for the annotators is reported in Figure 9. Evaluation judge. Model responses are scored by GPT-4o used as an LLM-as-a-judge. The judge assesses semantic equivalence between the model answer and the ground truth, permitting surface-level paraphrase while penalising under-specific responses. It returns a binary verdict vâtrue,falsevâ\ true,\, false\ of whether the model answer is correct with a brief explanation in structured XML tags, enabling fully reproducible programmatic evaluation. The exact prompt is provided in Figure 11. Judge reliability. To validate the reliability of the LLM-as-a-judge protocol, we constructed a stratified audit sample of 200 judge verdicts, drawn uniformly across languages and domains (100 English and 100 Russian items; 100 Business and 100 Science items, with 50 items per languageâdomain cell). Human annotators independently reviewed each model response, ground-truth answer, and judge verdict, recording agreement or disagreement. The judge achieved an overall human-agreement rate of 99.0% (198/200), with only two disagreements in the entire sample. Agreement remained consistently high across languages (99% for both English and Russian) and domains (99% for both Business and Science), as well as at the finer-grained language-domain level (98-100% across all four cells), with 95% Wilson confidence intervals overlapping the overall estimate throughout. On this stratified sample, the LLM-as-a-judge protocol agrees closely with human evaluation. Quality control. Eight state-of-the-art proprietary VLMs were queried on every item: Gemini 2.5 Pro/Flash, Gemini 3.1 Pro/Flash, Qwen 3.6 Plus, Qwen3.5 397B, Claude Sonnet 4.6, and Claude Opus 4.6. Items where three or more models returned identical responses â normalised for punctuation and case â and the LLM judge assigned false to all answers were flagged. A manual audit confirmed that 99% of flagged items had erroneous or ambiguous ground truth; all were excluded from the final benchmark. A random sample of retained items was quality-assessed along four dimensions: GT quality (84.4%), judge verdict quality (90.6%), question quality (90.9%), and image quality (97.0%). Final dataset composition. After quality-control filtering, BEAR-Bench comprises 1,000 document images paired with 1,000 human-authored QA instances across four domainâlanguage groups. Representative samples from BEAR-Bench are shown in Figure 1. 4 Dataset Statistics and Analysis Figure 2 summarises the composition of BEAR-Bench across three dimensions: domainâlanguage balance, question complexity, and visual content-type prevalence. Domain and language balance. The Russian business cell is the largest subset (352 items, 35.2%), reflecting the higher volume of publicly accessible Russian-language corporate disclosure documents, while the English business cell is the smallest (180 items, 18.0%). The science cells are more evenly distributed (266 and 202 items for Russian and English, respectively). Reasoning depth. Reasoning-depth annotations are available for 940 of the 1,000 items in BEAR-Bench. Each annotation estimates the intended number of steps required to derive the correct answer from the document image. Because such step counts depend on how annotators decompose a task, we treat them as coarse descriptive metadata rather than an objective difficulty score. The estimates span 2â10+ steps and peak at 4â5 steps with a moderate positive skew, indicating that the benchmark construction targeted multi-step inference rather than simple extraction. Visual content types. Figures and diagrams are the most prevalent content types across all subsets, consistent with the heavy use of infographics in both corporate reports and scientific papers. Equations appear almost exclusively in the science subsets, reflecting the mathematical nature of the arXiv and CyberLeninka source material, while code fragments are comparatively rare overall. Figure 2: BEAR-Bench dataset statistics. (a) Item distribution across the four domain Ă language cells; the central numeral indicates the total count. (b) Distribution of complexity scores over all 1,000 items, where each score reflects the total number of reasoning and computational steps required to solve the corresponding question. (c) Prevalence of visual content types, diagrams, equations, code, and figures, disaggregated by subset; bars show absolute counts. Items may carry multiple content-type labels simultaneously. 5 Experiments 5.1 Experimental Setup Evaluated models. We evaluate a diverse set of MLLMs on BEAR-Bench, covering both open-weight models (Qwen3.5-0.8B/4B/9B/27B 22, Qwen3-VL-2B/8B-Instruct 29, Qwen3-VL-4B-Thinking 29, and Gemma-4-31B-it 28) and proprietary systems (Qwen3.5-397B-A17B 22, Qwen 3.6 Plus 23, Gemini 2.5/3.1 Pro/Flash 9, and Claude Sonnet/Opus 4.6 3; 2). The open-weight models were run locally on an internal GPU server equipped with NVIDIA H100 and NVIDIA L40 accelerators; the proprietary ones were queried via the OpenRouter API. Inference protocol. All models were evaluated zero-shot under a fixed protocol: each instance received only the image and the raw question, with no few-shot examples or prompt engineering, and a uniform decoding temperature of 0.6. 5.2 Results Table 2: BEAR-Bench leaderboard. Results are reported as accuracy (%) across all evaluation subsets. Bold denotes the best result in each column. Model Overall Business Science RU EN Figures Diagrams Equations (n=1000n=1000) (n=532n=532) (n=468n=468) (n=618n=618) (n=382n=382) (n=359n=359) (n=566n=566) (n=199n=199) Proprietary models Qwen3.5-397B-A17B 75.4 78.2 72.2 71.4 81.9 65.7 76.9 74.9 Gemini 3.1 Pro 75.1 79.5 70.1 70.2 83.0 65.5 76.3 71.4 Gemini 2.5 Pro 67.4 72.7 61.3 62.8 74.9 62.4 66.8 62.3 Qwen 3.6 Plus 65.0 59.4 71.4 62.5 69.1 62.7 62.4 69.3 Claude Sonnet 4.6 62.7 71.4 52.8 57.6 70.9 52.9 62.2 54.8 Claude Opus 4.6 61.5 67.5 54.7 57.1 68.6 55.7 58.3 59.3 Gemini 3.1 Flash 59.8 62.2 57.0 56.8 64.7 54.0 61.3 56.8 Gemini 2.5 Flash 49.0 56.0 41.0 44.8 55.8 43.2 50.0 43.7 Open-weight models Qwen3.5-27B 60.8 68.9 51.6 54.6 70.8 50.8 60.8 50.8 Qwen3.5-9B 49.6 59.8 38.1 43.3 60.0 42.5 49.5 40.7 Qwen3.5-4B 48.6 58.1 37.9 42.3 59.0 42.2 48.0 39.2 Qwen3-VL-4B-Thinking 34.5 39.8 28.5 29.7 42.4 30.7 32.8 29.1 Qwen3-VL-8B-Instruct 25.4 18.3 33.4 23.5 28.4 28.8 23.4 37.2 gemma-4-31B-it 23.5 24.3 22.5 22.0 25.8 26.5 24.8 26.6 Qwen3-VL-2B-Instruct 14.3 11.3 17.8 10.7 20.3 16.2 13.7 19.1 Qwen3.5-0.8B 10.6 8.3 13.3 6.5 17.4 13.1 11.9 14.6 Main results. Table 2 shows remaining headroom on BEAR-Bench. Qwen3.5-397B-A17B and Gemini 3.1 Pro achieve the highest overall accuracy (75.4% and 75.1%, respectively), trading the lead across subsets â Qwen3.5-397B-A17B is stronger on Science and Equations, while Gemini 3.1 Pro edges ahead on Business and English items â indicating that no single system dominates across all domains. All models exhibit a marked drop from English to Russian (e.g., 83.0%â 70.2% for Gemini 3.1 Pro and 81.9%â 71.4% for Qwen3.5-397B-A17B), confirming that the linguistic gap identified in prior benchmarks persists even for frontier proprietary models. Among the evaluated systems, proprietary models generally achieve higher accuracy than their open-weight counterparts. Within the Qwen3.5 and Qwen3-VL-Instruct families, larger models tend to perform better in both languages (Figure 3), although accuracy remains substantially lower on Russian items. Overall, current models remain limited in visually grounded reasoning over text-dense professional documents, with performance shaped jointly by model scale, language, and document type. Figure 3: Accuracy on BEAR-Bench versus model size for the Qwen3.5 and Qwen3-VL-Instruct families (parameters on a log axis), reported separately for English (solid) and Russian (dashed) items. The English-over-Russian gap persists across scales. Effect of Chain-of-Thought prompting. Table 3: Overall accuracy (%) with and without an explicit chain-of-thought prompt. Î is CoT minus the default (no-CoT) protocol used in Table 2. Model No CoT CoT Î Qwen3.5-9B 49.6 49.3 â-0.3 Qwen3.5-4B 48.6 47.2 â-1.4 Qwen3-VL-4B-Thinking 34.5 34.9 ++0.4 Qwen3-VL-8B-Instruct 25.4 39.1 ++13.7 We tested whether explicit chain-of-thought (CoT) prompting improves accuracy on BEAR-Bench. A subset of open-weight models was re-evaluated with a prompt that asks the model to extract information from the image and solve the task step by step before giving a final answer (Table 3; the full prompt is given in Appendix F). For reasoning-oriented models â Qwen3.5-9B, Qwen3.5-4B, and Qwen3-VL-4B-Thinking â accuracy is essentially unchanged or slightly lower, consistent with these models already performing intermediate reasoning under the default protocol. By contrast, the instruct-tuned Qwen3-VL-8B-Instruct improves by 13.7 percentage points when steered to externalize multi-step reasoning. Effect of image resolution. Table 4: Overall accuracy of Qwen3.5-9B on BEAR-Bench at different image downsampling factors. A factor of c means that both the width and height are divided by c. Downsampling factor 11 1.51.5 22 33 44 Accuracy (%) 49.6 45.0 33.0 12.6 4.7 To measure the sensitivity of document reasoning to image resolution, we evaluated Qwen3.5-9B after resizing each input image to 1/c1/c of its original width and height, where c is the downsampling factor in Table 4. All other inference settings were kept unchanged. Accuracy decreases from 49.3% at the original resolution to 45.0% at c=1.5c=1.5 and 33.0% at c=2c=2, before falling sharply to 12.6% at c=3c=3 and 4.7% at c=4c=4. This pronounced degradation suggests that preserving fine-grained visual detail is critical for reasoning over text-dense professional documents. 5.3 Error Analysis To characterize common failure modes on BEAR-Bench, we conducted an exploratory error analysis combining inductive taxonomy construction with manual annotation of model responses. 5.3.1 Taxonomy construction To derive a failure taxonomy grounded in actual model behavior, we first collected an open-coding pilot of 150 incorrect responses produced by Gemini 3.1 Pro on BEAR-Bench, stratified across domains and languages. Four members of our team manually inspected each failure case and wrote detailed, free-form natural-language comments explaining the underlying cause of the error, rather than assigning predefined labels. This yielded a corpus of 150 rich failure descriptions covering a broad range of perceptual and reasoning breakdowns. We then prompted GPT-5.5-Pro, used as an advanced classification assistant, to cluster these free-form comments into thematically coherent groups based on their underlying error mechanism. The resulting clusters were manually reviewed and refined by the authors into five final categories: spatial misgrounding (C1), counting/aggregation (C2), OCR/visual-attribute (C3), chart-value extraction (C4), and semantic/reasoning (C5). Detailed descriptions of each error type are provided in Appendix A. 5.3.2 Error statistics Figure 4: Error-type distribution in a manually annotated subsample of Gemini 3.1 Pro responses (n=62n=62). Samples may receive multiple labels. We applied the taxonomy to a subsample of Gemini 3.1 Pro incorrect responses, selected to cover diverse document types and complexity levels. Figure 4 shows the resulting distribution. Visual-perceptual errors dominate: spatial misgrounding (C1) and OCR/visual-attribute errors (C3) together account for the majority of failures. Our analysis suggests that the most common errors on BEAR-Bench are perceptual, involving misread text or visual attributes and mislocalized evidence. 5.4 Detecting Incorrect Answers We next use responses generated on BEAR-Bench to compare existing hallucination detection methods for text-dense professional document reasoning. Following the benchmarkâs binary answer evaluation, we assess whether each method can distinguish correct from incorrect responses. Setup. We study two open-weight Qwen3.5 models using their native internal signals and eight proprietary models using proxy hidden states extracted from Qwen3-VL-8B 4, conditioned on each image, question, and proprietary-model response. We evaluate six uncertainty scoresâmax/mean token probability, log-likelihood, max/mean entropy, and perplexityâplus ContextualLens 21, the supervised hidden-state probe SUQ 15, and an MLLM-as-a-judge baseline (Qwen3-VL-8B) 11. We report balanced accuracy (BalAcc), AUROC and AUC-PR metrics. Results. The performance of the evaluated methods is shown in Tables 6-8. No detector wins everywhere; performance depends on the access regime and on response length (Table 5). SUQ achieves the highest BalAcc on all eight proprietary models (BalAcc 0.670.67â0.740.74; median response length <150<150 words), but performs less well on the two open-weight models (BalAcc 0.600.60â0.670.67; median length >1,000>1,000 words), suggesting that a last-token embedding provides limited signal for errors occurring earlier in long reasoning chains. Among proxy uncertainty scores, max token probability is strongest for the four models answering in two to three words (0.590.59â0.660.66) but at chance for the four with longer answers (0.490.49â0.510.51), while mean-based score shows the reverse (0.480.48â0.560.56 versus 0.650.65â0.660.66). The judge improves with length, from 0.570.57â0.620.62 on terse answers to 0.800.80â0.810.81 on the verbose open-weight modelsâthe highest BalAcc we observeâas a detailed derivation can be checked step by step. 6 Conclusion We introduced BEAR-Bench, a bilingual benchmark of 1,000 human-authored questions for context-grounded, multi-step reasoning over text-dense business and scientific documents. Benchmark items include dense textual and graphical page content such as figures, tables, charts, equations, and diagrams. Across 16 proprietary and open-weight MLLMs, the highest overall accuracy on BEAR-Bench is 75.4%, and every evaluated model achieves lower accuracy on Russian items. Our error analysis suggests that common failures involve spatial grounding, OCR, and visual-attribute perception. We also used the resulting model responses to compare existing hallucination detection methods. Performance varies across target models: supervised probes perform best for proprietary outputs, while an MLLM judge achieves the highest balanced accuracy on the verbose responses from open-weight outputs. The best balanced accuracy is 0.74 for proprietary outputs and 0.81 for open-weight models, showing that incorrect responses are not always identified reliably. BEAR-Bench therefore provides a common setting for tracking progress in both professional document reasoning and error detection. Limitations BEAR-Bench is deliberately narrow in several aspects. All questions are scoped to a single rendered page image; the benchmark therefore does not evaluate multi-page or cross-document reasoning. Coverage is limited to English and Russian enterprise and academic documents drawn from public disclosure and preprint sources, so findings may not transfer to other languages, domains, or private enterprise corpora. With 1,000 items and uneven languageâdomain cell sizes, subset estimates carry more variance than the overall score. Although we validate the LLM-as-a-judge protocol against humans, scoring still depends on an external model, and reasoning-depth labels remain coarse annotator estimates rather than objective difficulty. Finally, our hallucination-detection study compares existing methods under two access regimes and finds that detector quality varies with response length; we do not propose a new detector, and even the best balanced accuracies leave substantial room for improvement. 7 Acknowledgements The work was supported by the grant for research centers in the field of AI provided by the Ministry of Economic Development of the Russian Federation in accordance with the agreement 000000C313925P4F0002 and the agreement â139-10-2025-033. References Anthropic (2024) Anthropic The Claude 3 model family: Opus, Sonnet, Haiku. Model Card Anthropic. Note: Accessed: 2024-09-07 External Links: Link Cited by: §1. Anthropic (2026a) Anthropic Claude opus (version 4.6). External Links: Link Cited by: §5.1. Anthropic (2026b) Anthropic Claude sonnet (version 4.6). External Links: Link Cited by: §5.1. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu Qwen3-vl technical report. External Links: 2511.21631, Link Cited by: §5.4. Chen et al. (2024) X. Chen, C. Wang, Y. Xue, N. Zhang, X. Yang, Q. Li, Y. Shen, L. Liang, J. Gu, and H. Chen Unified hallucination detection for multimodal large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3235â3252. Cited by: §2. Chen et al. (2026a) Z. Chen, Y. Min, J. Zhang, B. Yan, J. Wang, X. Wang, and S. Shan A survey of multimodal hallucination evaluation and detection. International Journal of Computer Vision 134 (3), p. 131. Cited by: §2. Chen et al. (2026b) Z. Chen, Y. Zhao, C. Wang, R. R. Han, M. Patwardhan, and A. Cohan SciMDR: advancing scientific multimodal document reasoning. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 44718â44742. Cited by: §2. Chervyakov et al. (2026) A. Chervyakov, U. Isaeva, A. Emelyanov, A. Safin, M. Tikhonova, A. Kharitonov, Y. Lyakh, P. Surovtsev, D. Shevelev, V. Saburov, V. Konovalov, E. Rykov, I. Sviridov, A. Miftakhova, I. Alimova, A. Panchenko, A. Kapitanov, and A. Fenogenova Multimodal evaluation of Russian-language architectures. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 2114â2161. External Links: Link, Document, ISBN 979-8-89176-380-7 Cited by: Table 1, §1, §2. DeepMind (2026) G. DeepMind Gemini model cards. Note: Accessed: 2026-08-03 External Links: Link Cited by: §5.1. Fu et al. (2025) L. Fu, Z. Kuang, J. Song, M. Huang, B. Yang, Y. Li, L. Zhu, Q. Luo, X. Wang, H. Lu, Z. Li, G. Tang, B. Shan, C. Lin, Q. Liu, B. Wu, H. Feng, H. Liu, C. Huang, J. Tang, W. Chen, L. Jin, Y. Liu, and X. Bai OCRBench v2: an improved benchmark for evaluating large multimodal models on visual text localization and reasoning. External Links: 2501.00321, Link Cited by: Table 1, §2. Gu et al. (2025) J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y. Wang, W. Gao, L. Ni, and J. Guo A survey on llm-as-a-judge. External Links: 2411.15594, Link Cited by: §5.4. Hao et al. (2025) Y. Hao, J. Gu, H. W. Wang, L. Li, Z. Yang, L. Wang, and Y. Cheng Can mllms reason in multimodality? emma: an enhanced multimodal reasoning benchmark. External Links: 2501.05444, Link Cited by: §1, §2. Huang et al. (2026) M. Huang, Y. Shi, D. Peng, S. Lai, Z. Xie, and L. Jin OCR-reasoning benchmark: unveiling the true capabilities of mllms in complex text-rich image reasoning. External Links: 2505.17163, Link Cited by: Table 1, §1, §2. Jiang et al. (2025) Z. Jiang, J. Chen, B. Zhu, T. Luo, Y. Shen, and X. Yang Devils in middle layers of large vision-language models: interpreting, detecting and mitigating object hallucinations via attention lens. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 25004â25014. Cited by: §2. Li et al. (2024) Q. Li, J. Geng, C. Lyu, D. Zhu, M. Panov, and F. Karray Reference-free hallucination detection for large vision-language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, p. 4542â4551. Cited by: §2, §5.4. Masry et al. (2025) A. Masry, M. S. Islam, M. Ahmed, A. Bajaj, F. Kabir, A. Kartha, M. T. R. Laskar, M. Rahman, S. Rahman, M. Shahmohammadi, et al. Chartqapro: a more diverse and challenging benchmark for chart question answering. In Findings of the Association for Computational Linguistics: ACL 2025, p. 19123â19151. Cited by: §2. Masry et al. (2022) A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque ChartQA: a benchmark for question answering about charts with visual and logical reasoning. In Findings of the Association for Computational Linguistics: ACL 2022, S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 2263â2279. External Links: Link, Document Cited by: Table 1, §2. Mathew et al. (2021) M. Mathew, D. Karatzas, and C. V. Jawahar DocVQA: a dataset for vqa on document images. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Vol. , p. 2199â2208. External Links: Document Cited by: Table 1, §1, §2. MWS AI (2025) MWS AI MWS vision bench: a Russian-language document benchmark for multimodal large language models. Note: https://github.com/mts-ai/MWS-Vision-BenchAccessed: 2026-07-31 Cited by: Figure 5, Appendix B, Table 1, §1, §2. OpenAI et al. (2024) OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Ć. Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Ć. Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. MĂ©ly, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. OâKeefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson, P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph GPT-4 technical report. External Links: 2303.08774, Link Cited by: §1. Phukan et al. (2025) A. Phukan, Divyansh, H. K. Morj, Vaishnavi, A. Saxena, and K. Goswami Beyond logit lens: contextual embeddings for robust hallucination detection & grounding in VLMs. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 9661â9675. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §5.4. Qwen Team (2026a) Qwen Team Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §5.1. Qwen Team (2026b) Qwen Team Qwen3.6-Plus: towards real world agents. External Links: Link Cited by: §5.1. Sahu et al. (2024) P. Sahu, K. Sikka, and A. Divakaran Pelican: correcting hallucination in vision-LLMs via claim decomposition and program of thought verification. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, p. 8228â8248. External Links: Link, Document Cited by: §2. Tang et al. (2025) J. Tang, Q. Liu, Y. Ye, J. Lu, S. Wei, C. Lin, W. Li, M. F. F. B. Mahmood, H. Feng, Z. Zhao, Y. Wang, Y. Liu, H. Liu, X. Bai, and C. Huang MTVQA: benchmarking multilingual text-centric visual question answering. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 7764â7794. External Links: Link, Document, ISBN 979-8-89176-256-5 Cited by: §1, §2. Tang et al. (2026) Z. Tang, E. Haihong, R. Li, J. Liu, L. Jia, Z. Hao, Z. Yang, Y. Li, H. Tian, X. Hu, et al. Finmmdocr: benchmarking financial multimodal reasoning with scenario awareness, document understanding, and multi-step computation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, p. 25858â25866. Cited by: §2. Team et al. (2025) G. Team, R. Anil, S. Borgeaud, J. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millican, D. Silver, M. Johnson, I. Antonoglou, J. Schrittwieser, A. Glaese, J. Chen, E. Pitler, T. Lillicrap, A. Lazaridou, O. Firat, J. Molloy, M. Isard, P. R. Barham, T. Hennigan, B. Lee, F. Viola, M. Reynolds, Y. Xu, R. Doherty, E. Collins, C. Meyer, E. Rutherford, E. Moreira, K. Ayoub, M. Goel, J. Krawczyk, C. Du, E. Chi, H. Cheng, E. Ni, P. Shah, P. Kane, B. Chan, M. Faruqui, A. Severyn, H. Lin, Y. Li, Y. Cheng, A. Ittycheriah, M. Mahdieh, M. Chen, P. Sun, D. Tran, S. Bagri, B. Lakshminarayanan, J. Liu, A. Orban, F. GĂŒra, H. Zhou, X. Song, A. Boffy, H. Ganapathy, S. Zheng, H. Choe, Ă. Weisz, T. Zhu, Y. Lu, S. Gopal, J. Kahn, M. Kula, J. Pitman, R. Shah, E. Taropa, M. A. Merey, M. Baeuml, Z. Chen, L. E. Shafey, Y. Zhang, O. Sercinoglu, G. Tucker, E. Piqueras, M. Krikun, I. Barr, N. Savinov, I. Danihelka, B. Roelofs, A. White, A. Andreassen, T. von Glehn, L. Yagati, M. Kazemi, L. Gonzalez, M. Khalman, J. Sygnowski, A. Frechette, C. Smith, L. Culp, L. Proleev, Y. Luan, X. Chen, J. Lottes, N. Schucher, F. Lebron, A. Rrustemi, N. Clay, P. Crone, T. Kocisky, J. Zhao, B. Perz, D. Yu, H. Howard, A. Bloniarz, J. W. Rae, H. Lu, L. Sifre, M. Maggioni, F. Alcober, D. Garrette, M. Barnes, S. Thakoor, J. Austin, G. Barth-Maron, W. Wong, R. Joshi, R. Chaabouni, D. Fatiha, A. Ahuja, G. S. Tomar, E. Senter, M. Chadwick, I. Kornakov, N. Attaluri, I. Iturrate, R. Liu, Y. Li, S. Cogan, J. Chen, C. Jia, C. Gu, Q. Zhang, J. Grimstad, A. J. Hartman, X. Garcia, T. S. Pillai, J. Devlin, M. Laskin, D. de Las Casas, D. Valter, C. Tao, L. Blanco, A. P. Badia, D. Reitter, M. Chen, J. Brennan, C. Rivera, S. Brin, S. Iqbal, G. Surita, J. Labanowski, A. Rao, S. Winkler, E. Parisotto, Y. Gu, K. Olszewska, R. Addanki, A. Miech, A. Louis, D. Teplyashin, G. Brown, E. Catt, J. Balaguer, J. Xiang, P. Wang, Z. Ashwood, A. Briukhov, A. Webson, S. Ganapathy, S. Sanghavi, A. Kannan, M. Chang, A. Stjerngren, J. Djolonga, Y. Sun, A. Bapna, M. Aitchison, P. Pejman, H. Michalewski, T. Yu, C. Wang, J. Love, J. Ahn, D. Bloxwich, K. Han, P. Humphreys, T. Sellam, J. Bradbury, V. Godbole, S. Samangooei, B. Damoc, A. Kaskasoli, S. M. R. Arnold, V. Vasudevan, S. Agrawal, J. Riesa, D. Lepikhin, R. Tanburn, S. Srinivasan, H. Lim, S. Hodkinson, P. Shyam, J. Ferret, S. Hand, A. Garg, T. L. Paine, J. Li, Y. Li, M. Giang, A. Neitz, Z. Abbas, S. York, M. Reid, E. Cole, A. Chowdhery, D. Das, D. RogoziĆska, V. Nikolaev, P. Sprechmann, Z. Nado, L. Zilka, F. Prost, L. He, M. Monteiro, G. Mishra, C. Welty, J. Newlan, D. Jia, M. Allamanis, C. H. Hu, R. de Liedekerke, J. Gilmer, C. Saroufim, S. Rijhwani, S. Hou, D. Shrivastava, A. Baddepudi, A. Goldin, A. Ozturel, A. Cassirer, Y. Xu, D. Sohn, D. Sachan, R. K. Amplayo, C. Swanson, D. Petrova, S. Narayan, A. Guez, S. Brahma, J. Landon, M. Patel, R. Zhao, K. Villela, L. Wang, W. Jia, M. Rahtz, M. GimĂ©nez, L. Yeung, J. Keeling, P. Georgiev, D. Mincu, B. Wu, S. Haykal, R. Saputro, K. Vodrahalli, J. Qin, Z. Cankara, A. Sharma, N. Fernando, W. Hawkins, B. Neyshabur, S. Kim, A. Hutter, P. Agrawal, A. Castro-Ros, G. van den Driessche, T. Wang, F. Yang, S. Chang, P. Komarek, R. McIlroy, M. LuÄiÄ, G. Zhang, W. Farhan, M. Sharman, P. Natsev, P. Michel, Y. Bansal, S. Qiao, K. Cao, S. Shakeri, C. Butterfield, J. Chung, P. K. Rubenstein, S. Agrawal, A. Mensch, K. Soparkar, K. Lenc, T. Chung, A. Pope, L. Maggiore, J. Kay, P. Jhakra, S. Wang, J. Maynez, M. Phuong, T. Tobin, A. Tacchetti, M. Trebacz, K. Robinson, Y. Katariya, S. Riedel, P. Bailey, K. Xiao, N. Ghelani, L. Aroyo, A. Slone, N. Houlsby, X. Xiong, Z. Yang, E. Gribovskaya, J. Adler, M. Wirth, L. Lee, M. Li, T. Kagohara, J. Pavagadhi, S. Bridgers, A. Bortsova, S. Ghemawat, Z. Ahmed, T. Liu, R. Powell, V. Bolina, M. Iinuma, P. Zablotskaia, J. Besley, D. Chung, T. Dozat, R. Comanescu, X. Si, J. Greer, G. Su, M. Polacek, R. L. Kaufman, S. Tokumine, H. Hu, E. Buchatskaya, Y. Miao, M. Elhawaty, A. Siddhant, N. Tomasev, J. Xing, C. Greer, H. Miller, S. Ashraf, A. Roy, Z. Zhang, A. Ma, A. Filos, M. Besta, R. Blevins, T. Klimenko, C. Yeh, S. Changpinyo, J. Mu, O. Chang, M. Pajarskas, C. Muir, V. Cohen, C. L. Lan, K. Haridasan, A. Marathe, S. Hansen, S. Douglas, R. Samuel, M. Wang, S. Austin, C. Lan, J. Jiang, J. Chiu, J. A. Lorenzo, L. L. Sjösund, S. Cevey, Z. Gleicher, T. Avrahami, A. Boral, H. Srinivasan, V. Selo, R. May, K. Aisopos, L. Hussenot, L. B. Soares, K. Baumli, M. B. Chang, A. Recasens, B. Caine, A. Pritzel, F. Pavetic, F. Pardo, A. Gergely, J. Frye, V. Ramasesh, D. Horgan, K. Badola, N. Kassner, S. Roy, E. Dyer, V. C. Campos, A. Tomala, Y. Tang, D. E. Badawy, E. White, B. Mustafa, O. Lang, A. Jindal, S. Vikram, Z. Gong, S. Caelles, R. Hemsley, G. Thornton, F. Feng, W. Stokowiec, C. Zheng, P. Thacker, Ă. ĂnlĂŒ, Z. Zhang, M. Saleh, J. Svensson, M. Bileschi, P. Patil, A. Anand, R. Ring, K. Tsihlas, A. Vezer, M. Selvi, T. Shevlane, M. Rodriguez, T. Kwiatkowski, S. Daruki, K. Rong, A. Dafoe, N. FitzGerald, K. Gu-Lemberg, M. Khan, L. A. Hendricks, M. Pellat, V. Feinberg, J. Cobon-Kerr, T. Sainath, M. Rauh, S. H. Hashemi, R. Ives, Y. Hasson, E. Noland, Y. Cao, N. Byrd, L. Hou, Q. Wang, T. Sottiaux, M. Paganini, J. Lespiau, A. Moufarek, S. Hassan, K. Shivakumar, J. van Amersfoort, A. Mandhane, P. Joshi, A. Goyal, M. Tung, A. Brock, H. Sheahan, V. Misra, C. Li, N. RakiÄeviÄ, M. Dehghani, F. Liu, S. Mittal, J. Oh, S. Noury, E. Sezener, F. Huot, M. Lamm, N. D. Cao, C. Chen, S. Mudgal, R. Stella, K. Brooks, G. Vasudevan, C. Liu, M. Chain, N. Melinkeri, A. Cohen, V. Wang, K. Seymore, S. Zubkov, R. Goel, S. Yue, S. Krishnakumaran, B. Albert, N. Hurley, M. Sano, A. Mohananey, J. Joughin, E. Filonov, T. KÄpa, Y. Eldawy, J. Lim, R. Rishi, S. Badiezadegan, T. Bos, J. Chang, S. Jain, S. G. S. Padmanabhan, S. Puttagunta, K. Krishna, L. Baker, N. Kalb, V. Bedapudi, A. Kurzrok, S. Lei, A. Yu, O. Litvin, X. Zhou, Z. Wu, S. Sobell, A. Siciliano, A. Papir, R. Neale, J. Bragagnolo, T. Toor, T. Chen, V. Anklin, F. Wang, R. Feng, M. Gholami, K. Ling, L. Liu, J. Walter, H. Moghaddam, A. Kishore, J. Adamek, T. Mercado, J. Mallinson, S. Wandekar, S. Cagle, E. Ofek, G. Garrido, C. Lombriser, M. Mukha, B. Sun, H. R. Mohammad, J. Matak, Y. Qian, V. Peswani, P. Janus, Q. Yuan, L. Schelin, O. David, A. Garg, Y. He, O. Duzhyi, A. Ălgmyr, T. Lottaz, Q. Li, V. Yadav, L. Xu, A. Chinien, R. Shivanna, A. Chuklin, J. Li, C. Spadine, T. Wolfe, K. Mohamed, S. Das, Z. Dai, K. He, D. von Dincklage, S. Upadhyay, A. Maurya, L. Chi, S. Krause, K. Salama, P. G. Rabinovitch, P. K. R. M, A. Selvan, M. Dektiarev, G. Ghiasi, E. Guven, H. Gupta, B. Liu, D. Sharma, I. H. Shtacher, S. Paul, O. Akerlund, F. Aubet, T. Huang, C. Zhu, E. Zhu, E. Teixeira, M. Fritze, F. Bertolini, L. Marinescu, M. Bölle, D. Paulus, K. Gupta, T. Latkar, M. Chang, J. Sanders, R. Wilson, X. Wu, Y. Tan, L. N. Thiet, T. Doshi, S. Lall, S. Mishra, W. Chen, T. Luong, S. Benjamin, J. Lee, E. Andrejczuk, D. Rabiej, V. Ranjan, K. Styrc, P. Yin, J. Simon, M. R. Harriott, M. Bansal, A. Robsky, G. Bacon, D. Greene, D. Mirylenka, C. Zhou, O. Sarvana, A. Goyal, S. Andermatt, P. Siegler, B. Horn, A. Israel, F. Pongetti, C. ". Chen, M. Selvatici, P. Silva, K. Wang, J. Tolins, K. Guu, R. Yogev, X. Cai, A. Agostini, M. Shah, H. Nguyen, N. Ă. Donnaile, S. Pereira, L. Friso, A. Stambler, A. Kurzrok, C. Kuang, Y. Romanikhin, M. Geller, Z. Yan, K. Jang, C. Lee, W. Fica, E. Malmi, Q. Tan, D. Banica, D. Balle, R. Pham, Y. Huang, D. Avram, H. Shi, J. Singh, C. Hidey, N. Ahuja, P. Saxena, D. Dooley, S. P. Potharaju, E. OâNeill, A. Gokulchandran, R. Foley, K. Zhao, M. Dusenberry, Y. Liu, P. Mehta, R. Kotikalapudi, C. Safranek-Shrader, A. Goodman, J. Kessinger, E. Globen, P. Kolhar, C. Gorgolewski, A. Ibrahim, Y. Song, A. Eichenbaum, T. Brovelli, S. Potluri, P. Lahoti, C. Baetu, A. Ghorbani, C. Chen, A. Crawford, S. Pal, M. Sridhar, P. Gurita, A. Mujika, I. Petrovski, P. Cedoz, C. Li, S. Chen, N. D. Santo, S. Goyal, J. Punjabi, K. Kappaganthu, C. Kwak, P. LV, S. Velury, H. Choudhury, J. Hall, P. Shah, R. Figueira, M. Thomas, M. Lu, T. Zhou, C. Kumar, T. Jurdi, S. Chikkerur, Y. Ma, A. Yu, S. Kwak, V. Ăhdel, S. Rajayogam, T. Choma, F. Liu, A. Barua, C. Ji, J. H. Park, V. Hellendoorn, A. Bailey, T. Bilal, H. Zhou, M. Khatir, C. Sutton, W. Rzadkowski, F. Macintosh, R. Vij, K. Shagin, P. Medina, C. Liang, J. Zhou, P. Shah, Y. Bi, A. Dankovics, S. Banga, S. Lehmann, M. Bredesen, Z. Lin, J. E. Hoffmann, J. Lai, R. Chung, K. Yang, N. Balani, A. BraĆŸinskas, A. Sozanschi, M. Hayes, H. F. Alcalde, P. Makarov, W. Chen, A. Stella, L. Snijders, M. Mandl, A. KĂ€rrman, P. Nowak, X. Wu, A. Dyck, K. Vaidyanathan, R. R, J. Mallet, M. Rudominer, E. Johnston, S. Mittal, A. Udathu, J. Christensen, V. Verma, Z. Irving, A. Santucci, G. Elsayed, E. Davoodi, M. Georgiev, I. Tenney, N. Hua, G. Cideron, E. Leurent, M. Alnahlawi, I. Georgescu, N. Wei, I. Zheng, D. Scandinaro, H. Jiang, J. Snoek, M. Sundararajan, X. Wang, Z. Ontiveros, I. Karo, J. Cole, V. Rajashekhar, L. Tumeh, E. Ben-David, R. Jain, J. Uesato, R. Datta, O. Bunyan, S. Wu, J. Zhang, P. Stanczyk, Y. Zhang, D. Steiner, S. Naskar, M. Azzam, M. Johnson, A. Paszke, C. Chiu, J. S. Elias, A. Mohiuddin, F. Muhammad, J. Miao, A. Lee, N. Vieillard, J. Park, J. Zhang, J. Stanway, D. Garmon, A. Karmarkar, Z. Dong, J. Lee, A. Kumar, L. Zhou, J. Evens, W. Isaac, G. Irving, E. Loper, M. Fink, I. Arkatkar, N. Chen, I. Shafran, I. Petrychenko, Z. Chen, J. Jia, A. Levskaya, Z. Zhu, P. Grabowski, Y. Mao, A. Magni, K. Yao, J. Snaider, N. Casagrande, E. Palmer, P. Suganthan, A. Castaño, I. Giannoumis, W. Kim, M. RybiĆski, A. Sreevatsa, J. Prendki, D. Soergel, A. Goedeckemeyer, W. Gierke, M. Jafari, M. Gaba, J. Wiesner, D. G. Wright, Y. Wei, H. Vashisht, Y. Kulizhskaya, J. Hoover, M. Le, L. Li, C. Iwuanyanwu, L. Liu, K. Ramirez, A. Khorlin, A. Cui, T. LIN, M. Wu, R. Aguilar, K. Pallo, A. Chakladar, G. Perng, E. A. Abellan, M. Zhang, I. Dasgupta, N. Kushman, I. Penchev, A. Repina, X. Wu, T. van der Weide, P. Ponnapalli, C. Kaplan, J. Simsa, S. Li, O. Dousse, F. Yang, J. Piper, N. Ie, R. Pasumarthi, N. Lintz, A. Vijayakumar, D. Andor, P. Valenzuela, M. Lui, C. Paduraru, D. Peng, K. Lee, S. Zhang, S. Greene, D. D. Nguyen, P. Kurylowicz, C. Hardin, L. Dixon, L. Janzer, K. Choo, Z. Feng, B. Zhang, A. Singhal, D. Du, D. McKinnon, N. Antropova, T. Bolukbasi, O. Keller, D. Reid, D. Finchelstein, M. A. Raad, R. Crocker, P. Hawkins, R. Dadashi, C. Gaffney, K. Franko, A. Bulanova, R. Leblond, S. Chung, H. Askham, L. C. Cobo, K. Xu, F. Fischer, J. Xu, C. Sorokin, C. Alberti, C. Lin, C. Evans, A. Dimitriev, H. Forbes, D. Banarse, Z. Tung, M. Omernick, C. Bishop, R. Sterneck, R. Jain, J. Xia, E. Amid, F. Piccinno, X. Wang, P. Banzal, D. J. Mankowitz, A. Polozov, V. Krakovna, S. Brown, M. Bateni, D. Duan, V. Firoiu, M. Thotakuri, T. Natan, M. Geist, S. tan Girgin, H. Li, J. Ye, O. Roval, R. Tojo, M. Kwong, J. Lee-Thorp, C. Yew, D. Sinopalnikov, S. Ramos, J. Mellor, A. Sharma, K. Wu, D. Miller, N. Sonnerat, D. Vnukov, R. Greig, J. Beattie, E. Caveness, L. Bai, J. Eisenschlos, A. Korchemniy, T. Tsai, M. Jasarevic, W. Kong, P. Dao, Z. Zheng, F. Liu, F. Yang, R. Zhu, T. H. Teh, J. Sanmiya, E. Gladchenko, N. Trdin, D. Toyama, E. Rosen, S. Tavakkol, L. Xue, C. Elkind, O. Woodman, J. Carpenter, G. Papamakarios, R. Kemp, S. Kafle, T. Grunina, R. Sinha, A. Talbert, D. Wu, D. Owusu-Afriyie, C. Du, C. Thornton, J. Pont-Tuset, P. Narayana, J. Li, S. Fatehi, J. Wieting, O. Ajmeri, B. Uria, Y. Ko, L. Knight, A. HĂ©liou, N. Niu, S. Gu, C. Pang, Y. Li, N. Levine, A. Stolovich, R. Santamaria-Fernandez, S. Goenka, W. Yustalim, R. Strudel, A. Elqursh, C. Deck, H. Lee, Z. Li, K. Levin, R. Hoffmann, D. Holtmann-Rice, O. Bachem, S. Arora, C. Koh, S. H. Yeganeh, S. PĂ”der, M. Tariq, Y. Sun, L. Ionita, M. Seyedhosseini, P. Tafti, Z. Liu, A. Gulati, J. Liu, X. Ye, B. Chrzaszcz, L. Wang, N. Sethi, T. Li, B. Brown, S. Singh, W. Fan, A. Parisi, J. Stanton, V. Koverkathu, C. A. Choquette-Choo, Y. Li, T. Lu, A. Ittycheriah, P. Shroff, M. Varadarajan, S. Bahargam, R. Willoughby, D. Gaddy, G. Desjardins, M. Cornero, B. Robenek, B. Mittal, B. Albrecht, A. Shenoy, F. Moiseev, H. Jacobsson, A. Ghaffarkhah, M. RiviĂšre, A. Walton, C. Crepy, A. Parrish, Z. Zhou, C. Farabet, C. Radebaugh, P. Srinivasan, C. van der Salm, A. Fidjeland, S. Scellato, E. Latorre-Chimoto, H. Klimczak-PluciĆska, D. Bridson, D. de Cesare, T. Hudson, P. Mendolicchio, L. Walker, A. Morris, M. Mauger, A. Guseynov, A. Reid, S. Odoom, L. Loher, V. Cotruta, M. Yenugula, D. Grewe, A. Petrushkina, T. Duerig, A. Sanchez, S. Yadlowsky, A. Shen, A. Globerson, L. Webb, S. Dua, D. Li, S. Bhupatiraju, D. Hurt, H. Qureshi, A. Agarwal, T. Shani, M. Eyal, A. Khare, S. R. Belle, L. Wang, C. Tekur, M. S. Kale, J. Wei, R. Sang, B. Saeta, T. Liechty, Y. Sun, Y. Zhao, S. Lee, P. Nayak, D. Fritz, M. R. Vuyyuru, J. Aslanides, N. Vyas, M. Wicke, X. Ma, E. Eltyshev, N. Martin, H. Cate, J. Manyika, K. Amiri, Y. Kim, X. Xiong, K. Kang, F. Luisier, N. Tripuraneni, D. Madras, M. Guo, A. Waters, O. Wang, J. Ainslie, J. Baldridge, H. Zhang, G. Pruthi, J. Bauer, F. Yang, R. Mansour, J. Gelman, Y. Xu, G. Polovets, J. Liu, H. Cai, W. Chen, X. Sheng, E. Xue, S. Ozair, C. Angermueller, X. Li, A. Sinha, W. Wang, J. Wiesinger, E. Koukoumidis, Y. Tian, A. Iyer, M. Gurumurthy, M. Goldenson, P. Shah, M. Blake, H. Yu, A. Urbanowicz, J. Palomaki, C. Fernando, K. Durden, H. Mehta, N. Momchev, E. Rahimtoroghi, M. Georgaki, A. Raul, S. Ruder, M. Redshaw, J. Lee, D. Zhou, K. Jalan, D. Li, B. Hechtman, P. Schuh, M. Nasr, K. Milan, V. Mikulik, J. Franco, T. Green, N. Nguyen, J. Kelley, A. Mahendru, A. Hu, J. Howland, B. Vargas, J. Hui, K. Bansal, V. Rao, R. Ghiya, E. Wang, K. Ye, J. M. Sarr, M. M. Preston, M. Elish, S. Li, A. Kaku, J. Gupta, I. Pasupat, D. Juan, M. Someswar, T. M., X. Chen, A. Amini, A. Fabrikant, E. Chu, X. Dong, A. Muthal, S. Buthpitiya, S. Jauhari, N. Hua, U. Khandelwal, A. Hitron, J. Ren, L. Rinaldi, S. Drath, A. Dabush, N. Jiang, H. Godhia, U. Sachs, A. Chen, Y. Fan, H. Taitelbaum, H. Noga, Z. Dai, J. Wang, C. Liang, J. Hamer, C. Ferng, C. Elkind, A. Atias, P. Lee, V. ListĂk, M. Carlen, J. van de Kerkhof, M. Pikus, K. Zaher, P. MĂŒller, S. Zykova, R. Stefanec, V. Gatsko, C. Hirnschall, A. Sethi, X. F. Xu, C. Ahuja, B. Tsai, A. Stefanoiu, B. Feng, K. Dhandhania, M. Katyal, A. Gupta, A. Parulekar, D. Pitta, J. Zhao, V. Bhatia, Y. Bhavnani, O. Alhadlaq, X. Li, P. Danenberg, D. Tu, A. Pine, V. Filippova, A. Ghosh, B. Limonchik, B. Urala, C. K. Lanka, D. Clive, Y. Sun, E. Li, H. Wu, K. Hongtongsak, I. Li, K. Thakkar, K. Omarov, K. Majmundar, M. Alverson, M. Kucharski, M. Patel, M. Jain, M. Zabelin, P. Pelagatti, R. Kohli, S. Kumar, J. Kim, S. Sankar, V. Shah, L. Ramachandruni, X. Zeng, B. Bariach, L. Weidinger, T. Vu, A. Andreev, A. He, K. Hui, S. Kashem, A. Subramanya, S. Hsiao, D. Hassabis, K. Kavukcuoglu, A. Sadovsky, Q. Le, T. Strohman, Y. Wu, S. Petrov, J. Dean, and O. Vinyals Gemini: a family of highly capable multimodal models. External Links: 2312.11805, Link Cited by: §1. Team (2026) G. Team Gemma 4 technical report. External Links: 2607.02770, Link Cited by: §5.1. Team (2025) Q. Team Qwen3 technical report. External Links: 2505.09388, Link Cited by: §5.1. Tong et al. (2026) C. Tong, Q. Zhang, C. Li, L. Jiang, and Y. Liu FaithSCAN: model-driven single-pass hallucination detection for faithful visual question answering. arXiv preprint arXiv:2601.00269. Cited by: §2. Wang et al. (2024) Z. Wang, M. Xia, L. He, H. Chen, Y. Liu, R. Zhu, K. Liang, X. Wu, H. Liu, S. Malladi, et al. Charxiv: charting gaps in realistic chart understanding in multimodal llms. Advances in Neural Information Processing Systems 37, p. 113569â113697. Cited by: Table 1, §2. Wu et al. (2024) J. Wu, Q. Liu, D. Wang, J. Zhang, S. Wu, L. Wang, and T. Tan Logical closed loop: uncovering object hallucinations in large vision-language models. In Findings of the Association for Computational Linguistics: ACL 2024, p. 6944â6962. Cited by: §2. Xu et al. (2026) Z. Xu, J. Ji, Z. Chen, Z. Liu, Q. Liu, C. Peng, Z. Qin, Z. Xu, J. Wan, J. Tang, Z. Yang, S. Bai, and D. Liu C-ocr v2: benchmarking large multimodal models for literacy in real-world document processing. arXiv preprint arXiv:2605.03903. Cited by: Table 1, §2. Yin et al. (2024) S. Yin, C. Fu, S. Zhao, T. Xu, H. Wang, D. Sui, Y. Shen, K. Li, X. Sun, and E. Chen Woodpecker: hallucination correction for multimodal large language models. Science China Information Sciences 67 (12), p. 220105. Cited by: §2. Yue et al. (2024) X. Yue, Y. Ni, T. Zheng, K. Zhang, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen MMMU: a massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vol. , p. 9556â9567. External Links: Document Cited by: §1, §1, §2. Yue et al. (2025) X. Yue, T. Zheng, Y. Ni, Y. Wang, K. Zhang, S. Tong, Y. Sun, B. Yu, G. Zhang, H. Sun, et al. Mmmu-pro: a more robust multi-discipline multimodal understanding benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 15134â15186. Cited by: §2. Zhai et al. (2023) X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer Sigmoid loss for language image pre-training. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Vol. , p. 11941â11952. External Links: Document Cited by: §3.2.2. Zhang et al. (2026) F. Zhang, Y. Wu, Z. Wang, X. Wang, C. Lv, X. Huang, and X. Zheng Vib-probe: detecting and mitigating hallucinations in vision-language models via variational information bottleneck. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 23509â23521. Cited by: §2. Zhang et al. (2025) K. Zhang, L. Niu, Z. Cao, F. Meng, and J. Zhou TIU-bench: a benchmark for evaluating large multimodal models on text-rich image understanding. In Findings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 24286â24295. External Links: Link, Document, ISBN 979-8-89176-335-7 Cited by: §2. Zhang et al. (2024) R. Zhang, H. Zhang, and Z. Zheng Vl-uncertainty: detecting hallucination in large vision-language model via uncertainty estimation. arXiv preprint arXiv:2411.11919. Cited by: §2. Zuo et al. (2025) Y. Zuo, S. Qu, Y. Li, Z. Chen, X. Zhu, E. Hua, K. Zhang, N. Ding, and B. Zhou MedXpertQA: benchmarking expert-level medical reasoning and understanding. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §2. Appendix A Error taxonomy Below, we elaborate on the error taxonomy derived in our study. 1. C1 â Spatial localization and object matching. The model incorrectly identifies where objects are located in the image and how they relate to each other (e.g., selecting a neighboring element, misreading whether a marker lies inside a region, boundary intersections, label-to-object matching, arrow direction, or links between blocks). 2. C2 â Counting and aggregation errors of visual elements. The model makes mistakes when counting objects or aggregating extracted elements (e.g., points, circles, arrows, rows, columns, people, links, paths, labels, or table values), including missed elements, extra elements, or incorrect summation. 3. C3 â Errors in reading text, numbers, and visual attributes. The model incorrectly reads text, numbers, symbols, or small labels (OCR-related issues), and may also misidentify visual attributes such as color, marker shape, text style, italics/boldface, or legend encodings. 4. C4 â Errors in extracting values from charts. The model incorrectly reads quantitative values from plots/graphs (e.g., axis scale, ticks, point coordinates, values at specific x, peaks, minima, maxima, plateaus, trends, or ranges). 5. C5 â Semantic, instruction-following, and logical/arithmetic errors. The model misunderstands task conditions, categories, terms, units, or filtering rules, and/or makes reasoning or arithmetic mistakes after extraction (e.g., wrong interpretation, entity confusion, incorrect formulas, hallucinated assumptions, or incomplete answers). Appendix B Illustrative items from MWS Vision Bench Figure 5 shows public validation items from MWS Vision Bench 19. The illustration is taken from the datasetâs Hugging Face page.** * https://huggingface.co/datasets/MTSAIR/MWS-Vision-Bench We include them to make the contrast with BEAR-Bench concrete: MWS mixes business scans with personal handwriting, receipts, and form-style pages, and a large share of its tasks are OCR, grounding, and key-information extraction. BEAR-Bench instead targets multi-step questions on text-dense scientific and business document pages (Figure 1). Figure 5: Illustrative items from the public validation split of MWS Vision Bench 19. The mix of personal handwriting, receipts, and document-processing tasks differs from BEAR-Benchâs professional, text-dense pages. Appendix C Accuracy vs Reasoning Depth Figure 6 demonstrates the accuracy broken down by annotated reasoning-step count for several proprietary models. The trend is non-monotonic, suggesting that on knowledge-free tasks, frontier models are constrained more by visual perception than by the ability to execute long reasoning chains. We repeat this analysis for four open-weight models (Figure 7), where the two model families exhibit contrasting behaviour. The accuracy of the reasoning-tuned Qwen3.5 models remains stable as the annotated reasoning depth grows, whereas the instruction-tuned Qwen3-VL models degrade on items requiring seven or more steps. A likely reason is that reasoning-tuned models are trained to produce long chains of reasoning, so questions that need more steps cost them little extra accuracy, while instruction-tuned models are not trained for this and lose accuracy as more steps are needed. Figure 6: Accuracy on BEAR-Bench as a function of reasoning depth (number of steps per item) for three frontier models. Figure 7: Accuracy on BEAR-Bench as a function of reasoning depth (number of steps per item) for four open-weight models. Appendix D Additional data statistics: response length Table 5 reports the median response length in words for each evaluated model. Lengths are computed by whitespace-splitting each modelâs stored answer field. Proprietary and open-weight models differ sharply: several API systems return very short answers (median 2â3 words), while reasoning-oriented open-weight models often produce long intermediate chains (median above 1,000 words). Table 5: Median response length (words) by model. Words are whitespace-split tokens from each modelâs stored answer field. Group Model Median words Proprietary / API Gemini 3.1 Pro 3 Gemini 3.1 Flash 46 Gemini 2.5 Pro 3 Gemini 2.5 Flash 123 Qwen 3.6 Plus 2 Qwen3.5 397B A17B 2 Claude Sonnet 4.6 136 Claude Opus 4.6 149 Open-weight Qwen3.5-0.8B 96 Qwen3.5-27B 1036 Qwen3.5-4B 1206 Qwen3.5-9B 1162 Qwen3-VL-2B-Instruct 117 Qwen3-VL-4B-Thinking 1085 Qwen3-VL-8B-Instruct 2 gemma-4-31B-it 12 Appendix E Hallucination Detection Metrics This appendix reports the full 5-fold cross-validation results for the hallucination detectors evaluated in Section 5.4. We compare uncertainty-based scores, ContextualLens, SUQ Probe, and an MLLM-as-a-judge baseline under two regimes: proxy-based detection for proprietary models and native-signal detection for open-weight models. Tables 6 and 7 present AUROC, AUC-PR, and balanced accuracy for Gemini, Qwen, and Claude outputs; Table 8 reports the corresponding results for Qwen3.5-9B and Qwen3.5-27B. Table 6: Hallucination detection (5-fold CV): Gemini models. AUROC and AUC-PR are omitted for VLM-as-judge because the judge returns binary verdicts. Method AUROC AUC-PR BalAcc Gemini 3.1 Pro Max prob 0.64 0.38 0.60 Avg prob 0.48 0.26 0.48 Log-likelihood 0.50 0.29 0.46 Max entropy 0.60 0.41 0.52 Avg entropy 0.50 0.26 0.52 Perplexity 0.50 0.29 0.46 ContextualLens 0.41 0.24 0.40 SUQ Probe 0.76 0.54 0.69 VLM-as-judge â â 0.57 Random 0.50 0.27 0.50 Gemini 3.1 Flash Max prob 0.50 0.45 0.49 Avg prob 0.69 0.59 0.65 Log-likelihood 0.68 0.56 0.62 Max entropy 0.61 0.52 0.59 Avg entropy 0.71 0.60 0.66 Perplexity 0.68 0.56 0.62 ContextualLens 0.56 0.47 0.57 SUQ Probe 0.75 0.68 0.70 VLM-as-judge â â 0.69 Random 0.50 0.42 0.50 Gemini 2.5 Pro Max prob 0.63 0.45 0.61 Avg prob 0.54 0.38 0.51 Log-likelihood 0.57 0.40 0.53 Max entropy 0.58 0.47 0.52 Avg entropy 0.54 0.37 0.54 Perplexity 0.57 0.40 0.53 ContextualLens 0.46 0.35 0.46 SUQ Probe 0.77 0.65 0.70 VLM-as-judge â â 0.60 Random 0.50 0.34 0.50 Gemini 2.5 Flash Max prob 0.50 0.53 0.49 Avg prob 0.71 0.72 0.65 Log-likelihood 0.71 0.71 0.65 Max entropy 0.54 0.57 0.52 Avg entropy 0.70 0.71 0.65 Perplexity 0.71 0.71 0.65 ContextualLens 0.60 0.60 0.58 SUQ Probe 0.80 0.81 0.74 VLM-as-judge â â 0.72 Random 0.50 0.52 0.50 Table 7: Hallucination detection (5-fold CV): Qwen and Claude models. AUROC and AUC-PR are omitted for VLM-as-judge because the judge returns binary verdicts. Method AUROC AUC-PR BalAcc Qwen 3.6 Plus Max prob 0.72 0.61 0.66 Avg prob 0.62 0.48 0.56 Log-likelihood 0.67 0.58 0.57 Max entropy 0.62 0.50 0.56 Avg entropy 0.59 0.44 0.57 Perplexity 0.67 0.58 0.57 ContextualLens 0.48 0.39 0.47 SUQ Probe 0.80 0.72 0.72 VLM-as-judge â â 0.62 Random 0.50 0.37 0.50 Qwen3.5 397B Max prob 0.62 0.38 0.59 Avg prob 0.49 0.27 0.48 Log-likelihood 0.51 0.28 0.48 Max entropy 0.57 0.38 0.51 Avg entropy 0.50 0.27 0.52 Perplexity 0.51 0.28 0.48 ContextualLens 0.45 0.26 0.46 SUQ Probe 0.76 0.57 0.67 VLM-as-judge â â 0.57 Random 0.50 0.27 0.50 Claude Sonnet 4.6 Max prob 0.50 0.39 0.50 Avg prob 0.73 0.61 0.65 Log-likelihood 0.70 0.56 0.66 Max entropy 0.73 0.65 0.67 Avg entropy 0.75 0.65 0.69 Perplexity 0.70 0.56 0.66 ContextualLens 0.51 0.39 0.53 SUQ Probe 0.81 0.75 0.74 VLM-as-judge â â 0.68 Random 0.50 0.38 0.50 Claude Opus 4.6 Max prob 0.51 0.41 0.51 Avg prob 0.71 0.61 0.66 Log-likelihood 0.68 0.55 0.63 Max entropy 0.68 0.58 0.63 Avg entropy 0.73 0.63 0.68 Perplexity 0.68 0.55 0.63 ContextualLens 0.50 0.40 0.51 SUQ Probe 0.77 0.68 0.69 VLM-as-judge â â 0.68 Random 0.50 0.40 0.50 Table 8: Hallucination detection (5-fold CV): open-weight Qwen3.5 models (native signals). The VLM-as-a-judge results are computed over 997 valid samples for each model. AUROC and AUC-PR are omitted for VLM-as-judge because the judge returns binary verdicts. Method AUROC AUC-PR BalAcc Qwen3.5-9B Max prob 0.63 0.63 0.58 Avg prob 0.74 0.74 0.67 Log-likelihood 0.73 0.74 0.68 Max entropy 0.69 0.69 0.64 Avg entropy 0.74 0.75 0.68 Perplexity 0.73 0.74 0.68 ContextualLens 0.60 0.61 0.57 SUQ Probe 0.72 0.72 0.67 VLM-as-judge â â 0.80 Random 0.50 0.51 0.50 Qwen3.5-27B Max prob 0.65 0.56 0.60 Avg prob 0.70 0.61 0.63 Log-likelihood 0.70 0.61 0.63 Max entropy 0.69 0.61 0.63 Avg entropy 0.71 0.63 0.65 Perplexity 0.70 0.61 0.63 ContextualLens 0.63 0.56 0.59 SUQ Probe 0.65 0.55 0.60 VLM-as-judge â â 0.81 Random 0.50 0.39 0.50 Appendix F Chain-of-Thought Prompt Figure 8 shows the system prompt prepended to each imageâquestion pair in the CoT condition reported in Table 3. Figure 8: Chain-of-thought system prompt used in the CoT evaluation condition. Appendix G Human Annotation G.1 Instructions for annotators The instructions given to annotators are shown in Figures 9 and 10 (the original text and its English translation, respectively). The main goal for the annotators was to design complex multi-hop questions that can be answered solely from the image content, without requiring any external expert knowledge. We introduced several examples of âgoodâ and âbadâ questions to elaborate on the task. The annotators were instructed to formulate the answers to the questions as briefly as possible (e.g., a single number). Figure 9: Original Russian instructions given to benchmark annotators. The excerpt states the annotation objective (one multi-step question per document image with a ground-truth answer) and Section 2, which defines question requirements: at least three reasoning steps, OCR-grounded text, exemplar chain types, items to avoid, and answer constraints. The translation to English is provided in Figure 10. Figure 10: English translation of the annotator instructions shown in Figure 9. G.2 Recruitment & payment Annotators were recruited through an open call posted in an internal student chat channel. Prior to the main task, each candidate completed a qualification test in which they generated 10 probe questions designed to elicit incorrect answers from Gemini 3.1 Pro. Candidates who successfully caused the model to fail on more than 4 out of 10 questions were selected to work on the dataset. Annotators were compensated at 4.4 times the Russian minimum wage. G.3 Ethics & Consent No formal ethics review was required for this non-invasive annotation task. All participants provided informed consent. G.4 Demographics All annotators were aged 22â25 years and held at least a bachelorâs degree in a technical field, with self-reported English proficiency at CEFR B2 or higher. The sample included 60% male and 40% female participants. Appendix H Broader Impact, Data Use, and Compute Details Potential risks. BEAR-Bench is intended for research evaluation rather than autonomous decision making. Errors on its document-reasoning tasks can arise from misreading text, tables, figures, or equations and can consequently produce incorrect calculations or unsupported conclusions. If similar systems are used without human verification in financial, scientific, or other high-stakes workflows, such errors could lead to incorrect analyses or decisions. Performance also varies across languages and document types; therefore, aggregate benchmark scores should not be interpreted as evidence of reliable performance for every user population or document genre. We recommend using the benchmark for comparative evaluation and retaining human oversight in consequential settings. Intended use. BEAR-Bench is designed primarily as an evaluation benchmark for multimodal reasoning and hallucination detection. It does not provide step-by-step rationale annotations, and therefore does not directly support training or evaluating explicit chain-of-thought reasoning. The benchmark should not be interpreted as a resource for certifying models for deployment in high-stakes financial, legal, or scientific decision-making settings. Privacy and content. All source pages were obtained from publicly accessible official disclosure, government, intergovernmental, or academic sources, as described in Section 3. We did not collect personal data directly from individuals and did not conduct additional content screening beyond selecting documents from these sources and verifying their licensing status. Public source documents may nevertheless contain names, affiliations, or other information present in the original records. We retain source attribution and recommend that users treat the benchmark as a research resource rather than as a source of information about individuals. Compute infrastructure. Open-weight models were evaluated on an internal server with two NVIDIA H100 GPUs and five NVIDIA L40 GPUs. Proprietary models were accessed through the OpenRouter API; their underlying hardware configuration and parameter counts are not publicly available for all evaluated models. Figure 11: LLM-as-a-judge prompt used to score model responses against the ground truth.