Paper deep dive
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
Han Wang, David Wan, Hyunji Lee, Thinh Pham, Mikaela Cankosyan, Weiyuan Chen, Elias Stengel-Eskin, Tu Vu, Mohit Bansal
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/18/2026, 1:27:26 AM
Summary
MERRIN is a human-annotated benchmark designed to evaluate search-augmented AI agents on their ability to perform multi-hop reasoning and retrieve evidence across heterogeneous, noisy, and conflicting multimodal web sources (text, image, video, audio) without explicit modality cues.
Entities (4)
Relation Signals (3)
Gemini-3.1-Pro â achievedaccuracy â 40.1%
confidence 100% · the best-performing agent reaching only 40.1%
MERRIN â evaluates â Search-augmented agents
confidence 100% · MERRIN measures AI agents' ability to identify relevant modalities, retrieve multimodal evidence, and perform multi-hop reasoning
Agentic Multimodal Search â supports â Multimodal evidence retrieval
confidence 90% · Agentic Multimodal Search equips models with various tools to operate across all modalities.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web results, we introduce MERRIN (Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments), a human-annotated benchmark for evaluating search-augmented agents. MERRIN measures AI agents' ability to identify relevant modalities, retrieve multimodal evidence, and perform multi-hop reasoning over noisy web sources. It differs from prior work in three important aspects: (1) using natural language queries without explicit modality cues, (2) incorporating underexplored modalities such as video and audio, and (3) requiring the retrieval of complex, often noisy or conflicting multimodal evidence during web search. We evaluate diverse search agents powered by ten models, including strong closed-source models (e.g., GPT-5.4-mini, Gemini 3/3.1 Flash/Pro) and open-weight models (Qwen3-4B/30B/235B), across three search settings (no search, native search, and agentic search). Our results show that MERRIN is highly challenging: the average accuracy across all agents is 22.3%, with the best-performing agent reaching only 40.1%. We further observe that while stronger agents like Gemini Deep Research achieve higher performance, gains are modest due to over-exploration; they take more steps and use more tools, but are often distracted by conflicting or partially relevant web content, leading to incorrect answers. Compared to humans, these agents consume more resources yet achieve lower accuracy, largely due to inefficient source selection and an overreliance on text modalities. These findings highlight the need for search agents capable of robust search and reasoning across diverse modalities in noisy web environments, making MERRIN a valuable testbed for evaluating such capabilities.
Tags
Links
- Source: https://arxiv.org/abs/2604.13418v1
- Canonical: https://arxiv.org/abs/2604.13418v1
Trouble viewing inline? Open PDF directly â
Full Text
63,314 characters extracted from source content.
Expand or collapse full text
Preprint. Under review. MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments Han Wang 1â David Wan 1â Hyunji Lee 1â Thinh Pham 2 Mikaela Cankosyan 2 Weiyuan Chen 2 Elias Stengel-Eskin 3 Tu Vu 2 Mohit Bansal 1 1 UNC Chapel Hill 2 Virginia Tech 3 University of Texas at Austin https://merrin-benchmark.github.io Abstract Motivated by the underspecified, multi-hop nature of search queries and the multimodal, heterogeneous, and often conflicting nature of real-world web results, we introduceMERRIN(MultimodalEvidenceRetrieval and ReasoninginNoisy Web Environments), a human-annotated benchmark for evaluating search-augmented agents.MERRINmeasures AI agentsâ ability to identify relevant modalities, retrieve multimodal evidence, and perform multi-hop reasoning over noisy web sources. It differs from prior work in three important aspects: (1) using natural language queries without explicit modality cues, (2) incorporating underexplored modalities such as video and audio, and (3) requiring the retrieval of complex, often noisy or conflict- ing multimodal evidence during web search. We evaluate diverse search agents powered by ten models, including strong closed-source models (e.g., GPT-5.4-mini, Gemini 3/3.1 Flash/Pro) and open-weight models (Qwen3- 4B/30B/235B), across three search settings (no search, native search, and agentic search). Our results show thatMERRINis highly challenging: the average accuracy across all agents is 22.3%, with the best-performing agent reaching only 40.1%. We further observe that while stronger agents like Gemini Deep Research achieve higher performance, gains are modest due to over-exploration; they take more steps and use more tools, but are often distracted by conflicting or partially relevant web content, leading to incorrect answers. Our analysis of agent bottlenecks shows that while both search effectiveness and multimodal reasoning remain critical challenges, reasoning is the more pressing limitation. Compared to humans, these agents consume more resources yet achieve lower accuracy, largely due to inefficient source selection and an overreliance on text modalities. These findings highlight the need for search agents capable of robust search and reasoning across diverse modalities in noisy web environments, making MERRIN a valuable testbed for evaluating such capabilities. 1 Introduction Knowledge on the web is inherently heterogeneous, spanning text, images, videos, and audio, and is often noisy, incomplete, and conflicting across sources (Xu et al., 2024; Wang et al., 2025; Pham et al., 2026). Users of search-augmented agents frequently ask questions that require reasoning over multiple modalities, where agents must (1) identify which modalities are necessary and (2) perform multi-hop reasoning over retrieved evidence despite noise and irrelevant information. Evaluating these capabilities requires benchmarks that reflect the complexity of real-world web search, yet prior work has several limitations (Table 1): many include explicitmodality cuesâdirect references to a specific modality such asâIn the following image...â(Chen et al., 2023; Jiang et al., 2025; Li et al., 2025; Zhang et al., 2026); the range of modalities is often limited to text and images, excluding video and audio (Jia et al., 2025; Yan et al., 2025; Tian et al., 2025); and the noisy, conflicting nature of â Equal contribution. Correspondence to:hwang@cs.unc.edu. 1 arXiv:2604.13418v1 [cs.CL] 15 Apr 2026 Preprint. Under review. Who invented the law that Richard Feynman explains using the first equation he writes on the blackboard in the third 1964 Cornell Messenger Lecture? Video of Third Lecture Summary Text of Third Lecture Lecture Notes (Text) Locate the eq. In lecture notes Energy Conservation (Text) Find the inventor Highlight Video of Third Lecture Highlight Video (Video + Audio) No equation in video but hallucinates Schrödinger equation (Text) Find the inventor Benjamin Franklin Erwin Schrödinger Modality Error: only relied on text, leading to incorrect answer Retrieval Error: Got distracted by relevant but conflicting source Identified correct modality Selected correct source despite distractions 8:51 Third lecture (Video + Audio) Locate the third eq. & find the law Baryon Conservation (Text) Find the inventor of third eq. Ernst Stueckelberg 18:37 Reasoning Error: Incorrect processing or grounding of the correct source Julius von Mayer Third lecture (Video + Audio) Locate the first eq. & find the law Charge Conservation (Text) Find the inventor first eq. 1:42 Figure 1: Overview ofMERRIN. Given a query, the agent must identify the appropriate modality, retrieve relevant evidence, and perform multi-hop reasoning over noisy, conflict- ing, and incomplete web sources. The green path shows theidealcase: the agent selects the correct modality and source, arriving at the correct answer. The remaining paths illustrate threefailuremodes:Reasoning Error(blue)âcorrect source retrieved but incorrect ground- ing to the evidence;Modality Error(red)âagent relies on text when asked about visual information;Retrieval Error(purple)âcorrect modality but misleading source selected. real-world web evidence, well-studied in text-only settings (Pham et al., 2026; Lee et al., 2025; Wang et al., 2025), remains underexplored in multimodal settings. To address these gaps, we introduceMERRIN(MultimodalEvidenceRetrieval and ReasoninginNoisy Web Environments), a human-annotated benchmark designed to evalu- ate search-augmented agents under more realistic and challenging conditions.MERRIN requires agents to identify the necessary modalities and retrieve relevant sources from noisy multimodal evidence on the open web, particularly in scenarios involving multi-hop reasoning across heterogeneous sources and where queries may trigger conflicting, incom- plete, or noisy search results. Figure 1 illustrates this challenge. When an agent correctly identifies both the necessary modalities and the appropriate sources (green path), it can perform accurate reasoning and arrive at the correct answer. However, even with the correct source, the agent may commit aReasoning Error(blue path): it retrieves the right video but incorrectly grounds the evidenceâe.g., locating thethirdequation instead of thefirstin the videoâproducing an incorrect answer. AModality Error(red path) occurs when the agent relies on the wrong modalityâfor instance, using textual evidence when the question requires visual information (e.g., a diagram on a blackboard), leading to incorrect reasoning and answer. Finally, aRetrieval Error(purple path) arises when the agent identifies the right modality but selects a misleading sourceâe.g., a summary video rather than the full lectureâand hallucinates evidence that does not exist in the retrieved source. To evaluate agents on these challenges, we designMERRINalong multiple axes, as shown in Table 1. The questions are formulated in natural language, without explicit modality cues, requiring agents to autonomously infer which modalities are necessary and retrieve appropriate evidence.MERRINfurther expands the scope of modalities to include under- explored sources such as video and audio, alongside more commonly studied modalities like text, images, and tables. Moreover, inspired by prior observations about the noisy nature of real-world web data in text domain (Wang et al., 2025; Pham et al., 2026; Lee et al., 2025), we design our dataset such that each question induces the retrieval of not only relevant documents but also incomplete, conflicting, or misleading distractors. For reliable evaluation, we ensure that each question in the dataset has a single unambiguous answer, enabling consistent automatic evaluation of model performance. 2 Preprint. Under review. Benchmark No Explicit Modality Cues Evidence Modalities Web Noise Reflection Multi -hop Human Annotated Open Search BrowseComp (Wei et al., 2025)-Tââ M-BrowseComp (Li et al., 2025)âT/I/Vââ BrowseComp-VL (Geng et al., 2026)âT/Iââââ BrowseComp-V 3 (Zhang et al., 2026)âT/Iââ SealQA (Pham et al., 2026)-Tâ M3DocVQA (Cho et al., 2025)-T/Iâââ RamDocs (Wang et al., 2025)-Tââ MMSearch (Jiang et al., 2025)-T/Iâââ MMSearch-Plus (Tao et al., 2026)âT/Iâ MERRINâT/I/V/Aâ Table 1: Comparison ofMERRINwith existing benchmarks. We compare datasets across multiple dimensions: whether queries do not contain explicit modality cues (No Explicit Modality Cues), evidence modalities necessary to answer them (Evidence Modalities), whether questions reflect noisy or conflicting web sources (Web Noise Reflection), whether they require multi-hop reasoning (Multi-hop), whether they are human-annotated (Human Annotated), and whether they support open-web search (Open Search). MERRINuniquely covers all dimensions, supporting multiple evidence modalities across text (T), image (I), video (V), and audio (A). â-â inNo Explicit Modality Cuesindicates settings where modality selection is unnecessary (e.g., controlled or single modality setups). We evaluate search-augmented agents powered by ten different LLMsâseven closed-source (GPT-5.4-nano and -mini (OpenAI, 2026), Gemini-3-Flash and -Pro (Google, 2025a), Gemini- 3.1-Lite and -Pro (Google, 2026), and Gemini Deep Research Agent (Google, 2025b)) and three open-weight (Qwen3-4B, -30B, and -235B (Yang et al., 2025)) modelsâunder three search settings:No Search,Native Search, andAgentic Multimodal Search. Overall, we find thatMERRINis challenging, with an average accuracy of 22.3% across all runs; even the strongest agent, Gemini-3.1-Pro withAgentic Multimodal Search, achieves only 40.1%.Agentic Multimodal Searchperforms best, averaging 33.7%, compared to 23.1% forNative Search and 17.3% forNo Search(over six models evaluated in all settings). Notably, increasing the number of search queries or visited pages does not consistently improve accuracy, suggesting that more extensive search does not necessarily translate into better performance. We further find that more capable agents (e.g., Gemini Deep Research Agent and Gemini ProNative Search) are more prone to over-exploration in noisy web environments, issuing excessive and repeated search queries and tool calls without converging on an answer. To decouple the sources of error, we analyze whether failures stem from search or reasoning. Providing annotated gold evidence, thereby removing the need for search, yields only a modest improvement of 7.6% (40.1%â47.7%), with performance still remaining relatively low. This suggests that although both search effectiveness and multimodal reasoning remain critical challenges, improving reasoning is the more pressing bottleneck. In a human evalua- tion on a 50-example subset, humans achieve 71.4% accuracy, substantially outperforming the best agentic system (40.1%), while using fewer resources (nearly 3Ăfewer searches) and achieving higher precision in source selection (38.1% vs. 1.8%). However, humans also find the task challenging, with errors often arising from missed or incomplete details in web sources (e.g., incorrect counts or partial answers), highlighting the difficulty of the reasoning component. Moreover, humans benefit substantially from additional time (59.2%â71.4%), whereas agents show diminishing returns (34.0%â40.1% forAgentic Multimodal Search), consistent with the over-exploration pattern: agents issue redundant queries rather than productively deepening their search. These results highlight the need for stronger search agents that can better assist humans through robust search and reasoning over complex and noisy web environments and effectively integrating diverse modalities. Overall,MERRIN provides a challenging and realistic testbed for advancing these capabilities. 2MERRIN We presentMERRIN, a human-annotated benchmark for multimodal evidence retrieval and reasoning, designed to evaluate the ability of search-augmented agents to determine which 3 Preprint. Under review. TextVideoImageTable 0 15 30 45 60 75 90 105 120 Number of Resources 96 88 110 12 (a) Source Types. 644850 As Answer As Reasoning Chain (b) Multimodal Role. 2914119 Multi-hop Multimodal Conflict (c) Reasoning Type. Figure 2:MERRINcomposition. (a) Gold source resources by modality. (b) Questions by the role of visual content. (c) Questions by reasoning type. modalities to retrieve and correctly reason over noisy, conflicting multimodal evidence. We describe data collection in Section 2.1 and dataset statistics in Section 2.2. 2.1 Data Collection Question Design.MERRINconsists of questions governed by three core requirements and additionally classified along two axes. Every question must satisfy: (1)no modality cuesâ questions are phrased in natural language without explicit modality references (e.g.,âshown in the imageâ), resembling realistic user queries; (2)non-text evidence requiredâeach question is manually verified to require non-text evidence, with no text-only shortcut available; and (3)unique, verifiable answersâeach question has exactly one correct, short, and unambiguous answer. Each question is further classified along two axes.Reasoning type(one or both): multi-hop reasoning (combining information across sources or modalities) or multimodal conflict resolution (reconciling inconsistent evidence across modalities triggered empirically in real search engines; no synthetic conflicts).Multimodal role(one or both): non-text evidence may serve as the answer source (the answer can only be extracted from a non-text source) or as a reasoning component (non-text evidence provides an intermediate fact necessary to derive the final answer). Most questions and evidence are generated from scratch, while some are adapted from SealQA (Pham et al., 2026) and ChartMuseum (Tang et al., 2025), using their questionâanswer pairs as one hop and augmenting with additional evidence to constructnewmulti-hop questions. For each question, annotators record the ground-truth answer with reasoning steps, source URLs, source types, multimodal role, reasoning type, and question origin. Full annotation details and guidelines are in Appendix B.1. Quality Control.We employ a multi-round human review process. Each question is reviewed by a second annotator for answer correctness, question clarity, question difficulty, and non-text modality requirements. In the first round, approximately 39.5% of candidates were rejected; of those, 45.3% were successfully revised and accepted in the second round. To verifynon-text modality requirements, we decompose each question into sub-questions and attempt to answer each via text-only Google Search. We then perform an adversarial search pass, querying each sub-question together with the known answer to check for text-only shortcuts. A question passes only if at least one sub-question resists both search passes. Further details are in Appendix B.2. Human Annotators.Questions were constructed and reviewed by five graduate-level and one undergraduate annotators with NLP backgrounds. Six annotators constructed questions and four conducted quality control; no annotators reviewed their own questions. Annotators were provided with detailed guidelines (Appendix B.3) and diverse exemplars. 2.2 Data Statistics MERRINcomprises 162 questions 1 (120 from scratch, 37 from SealQA, 5 from ChartMu- seum). Four source types (text, image, video, and table) are represented, with text and 1 We focus on a high-quality, expert-vetted diagnostic benchmark, comparable in scale to SealQAâs SEAL-0 (111 questions) (Pham et al., 2026) and GPQA-Diamond (198 questions) (Rein et al., 2024). 4 Preprint. Under review. No SearchNative SearchAgentic Multimodal Search ModelAccAcc# Search Qs# PagesAcc# Search Qs# Pages Qwen3-4B10.3 ±0.4 ---10.5 ±1.6 1.7 ±0.0 0.2 ±0.1 Qwen3-30B8.0 ±0.6 ---16.1 ±0.6 2.0 ±0.1 0.5 ±0.0 Qwen3-235B 12.1 ±0.9 ---23.3 ±1.3 3.0 ±0.2 1.0 ±0.1 GPT-5.4-nano9.9 ±2.7 12.6 ±1.3 37.7 ±2.4 5.9 ±0.4 31.9 ±3.0 11.6 ±0.4 7.5 ±0.4 GPT-5.4-mini14.0 ±0.4 15.6 ±0.4 38.6 ±3.7 5.3 ±0.2 31.1 ±3.1 9.2 ±0.3 3.4 ±0.2 Gemini 3 Flash19.1 ±3.2 31.7 ±3.8 44.1 ±0.6 0.1 ±0.0 32.9 ±0.9 14.8 ±0.5 1.4 ±0.0 Gemini 3 Pro 23.5 ±1.1 28.8 ±1.4 34.9 ±2.0 0.1 ±0.0 39.9 ±1.6 8.4 ±0.3 3.0 ±0.1 Gemini 3.1 Lite 12.8 ±2.3 20.6 ±2.2 19.2 ±1.1 0.0 ±0.0 26.3 ±1.9 8.3 ±0.1 0.8 ±0.1 Gemini 3.1 Pro24.7 ±1.6 29.0 ±1.1 35.8 ±0.9 0.1 ±0.0 40.1 ±2.8 8.6 ±0.3 2.9 ±0.0 Gemini Research-33.3 ±2.2 ----- Table 2: Performance of search agents powered by different models onMERRIN.Acc denotes average accuracy over three runs with standard deviation,# Search Qsis the average number of search queries issued per question, and# Pagesis the average number of webpages explicitly visited and read per question.No Searchhas no search module; thus, both # Search Qs and # Pages are 0. Gemini 3.1 Lite refers to Gemini 3.1 Flash Lite, and Gemini Research refers to the Gemini Deep Research Agent. For Qwen models,Native Search is not applicable since they do not have an internal search agent. Gemini Research only supports using its built-in search system (Native Search) and detailed outputs are unavailable, so # of Search Qs and # Pages are omitted. image most prevalent (Figure 2a). Non-text evidence serves as both answer sources and rea- soning components in comparable proportions (Figure 2b), and 73.5% of questions require both multi-hop reasoning and multimodal conflict resolution (Figure 2c); see Table 5 in the Appendix for full statistics. 3 Experiments 3.1 Setup Search-Augmented Agents.We evaluate search-augmented agents powered by ten dif- ferent models, including closed-source models (GPT-5.4-Nano and -Mini (OpenAI, 2026), Gemini-3-Flash and -Pro (Google, 2025a), Gemini-3.1-Flash-Lite and -Pro (Google, 2026), and Gemini-Deep-Research Agent (Google, 2025b)), and open-weight models (Qwen 3 (Yang et al., 2025) at three scales: 4B, 30B, and 235B), under three search settings:No Search(no search tool),Native Search(enable each modelâs built-in search tools), andAgentic Multimodal Search(a multimodal search agent framework built using smolagents (Roucher et al., 2025)). Native Searchoften doesnotsupport video and audio processing when accessed via built-in search tools (Table 6 in Appendix).Agentic Multimodal Searchequips models with various tools to operate acrossallmodalities. Full details on model configurations, search tools, and theAgentic Multimodal Searchframework are in Appendix C.1. Metrics.We measureaccuracy, i.e., whether the predicted answer matches the ground truth, using an LLM-as-judge following BrowseComp (Wei et al., 2025). 2 3.2 Results Table 2 presents the performance of search-augmented agents powered by ten models on MERRIN under three search settings, with results averaged over three runs. Overall Performance.Overall, the task is challenging for all agents, with an average accuracy of 22.3% across all runs. When averaging over the six models evaluated in all 2 Manual inspection on 50 instances finds all judgments correct. Due toMERRINâs unambiguous design, most cases reduce to an exact match with text normalization. 5 Preprint. Under review. three settings, models achieve only 17.3% accuracy inNo Search, indicating thatMERRIN cannot be solved using parametric knowledge alone; even the strongest model, Gemini- 3.1-Pro, reaches only 24.7%. Performance improves to 23.1% withNative Search, where agents rely on built-in search pipelines, with the Gemini Deep Research Agent achieving the highest accuracy at 33.3%. Performance further increases to 33.7% withAgentic Multimodal Search, which enables access toallmultimodal evidence, unlikeNative Search, which does not support video during search (Appendix C.1), highlighting the importance of flexible evidence integration across various modalities. The best overall result is achieved by Gemini-3.1-Pro withAgentic Multimodal Search(40.1%), suggesting that both strong models and robust search frameworks are critical. Comparing model families, GPT-based agents perform substantially worse than Gemini agents underNative Search, with an absolute gap of 13.4%. However, this gap narrows to 3.3% underAgentic Multimodal Search, suggesting that more flexible search capabilities help close the gap between model families. Performance of Search Agents with Closed vs. Open-Weight Models.We observe that MERRINis particularly challenging for agents powered by open-weight models, the Qwen series, which achieve an average accuracy of 16.6 even when usingAgentic Multimodal Search, where external search is enabled. Although both agent types use the same tools and therefore retrieve similar evidence and interpretations, while access to external evidence improves performance for closed-source agents, the gains for open-weight agents are limited, with an average improvement of 16.4 fromNo SearchtoAgentic Multimodal Search for closed agents compared to only 6.5 for open-weight agents. We attribute the limited gains for open-weight models to three main factors: (1) failure to effectively process long, multi-step search results; (2) greater susceptibility to distraction from irrelevant evidence, leading to premature termination even when the generated answer is incorrect; and (3) weaker reasoning ability, which leads to incorrect intermediate reasoning that propagates to incorrect final answers. Average Search Queries and Pages Visited.For each agent and search setting, we analyze the average number of search queries issued (# Search Qs) and the average number of pages visited (# Pages). We observe that these metrics are not strongly correlated with accuracy (Acc). In particular, a higher number of search queries or visited pages does not necessarily lead to better performance. Similar trends are observed across bothNative Search andAgentic Multimodal Search. The highest accuracy is achieved by Gemini-3.1-Pro agent, while the largest number of search queries is observed for Gemini-3 Flash agent, and the highest number of pages visited is observed for GPT-5.4-nano agent. 3.3 Analysis of Failure Modes We provide a detailed quantitative and qualitative analysis of the best performing agent in Table 2, Gemini-3.1-Pro. Bias Toward Text Modality.We observe that agents exhibit a strong bias toward retrieving textual evidence, often failing to identify the most appropriate modality for a given query. Specifically, 87.7% of retrieved evidence is text, compared to only 6.8% from images and 5.5% from video and audio combined. In contrast, the dataset distribution is more balanced, with 31.4% text, 35.9% image, and 28.8% for video and audio. This discrepancy indicates that, although text is the dominant modality in the dataset, it is significantly preferred by search agents in retrieved evidence, often leading to incorrect answers. Error Propagation in Multi-Step Retrieval.To analyze multi-step retrieval, we construct 50 human-annotated examples in which each question requires a two-step reasoning chain. For each example, annotators provide sub-questions and intermediate answers for each step. We then analyze where agents fail by evaluating whether they correctly produce these intermediate answers. Among incorrect predictions, the first step is more often the point of failure (57.7%) than the second step (42.3%), indicating that the initial evidence identification is a frequent source of error and that early errors often propagate, leading to incorrect final answers. Second step failures are mostly associated withas answermultimodal instances 6 Preprint. Under review. 3 Flash3 Pro3.1 Lite3.1 Pro 0 5 10 15 20 25 30 35 Accuracy (%) Native + Video Figure 3: Performance ofNative Search(blue), when adding video tool (orange). NoneLowMediumHighxHigh Thinking Effort 0 10 20 30 40 Accuracy (%) No Search Native Search Agentic Multimodal Search Figure 4: Accuracy of GPT-5.4-mini across different thinking efforts. (63.6%), compared toreasoning chain(18.2%) andboth(18.2%), indicating the difficulty of understanding and integrating multimodal information to produce the final answer. Analysis Across Dataset Axes.We find that performance is similar when non-text modali- ties are required in thereasoning chain(45.8%) andas answer(45.3%), but drops substantially to 28.0% when both are required. We further examine performance across different question types. Performance is higher on multi-hop questions (55.2%) and multimodal conflict ques- tions (57.1%), but decreases to 34.5% when both challenges are present. This mirrors the trend above, indicating that the combination significantly increases task difficulty. Over-Exploration in Noisy Web Environments.We observe that more capable agents (e.g., Gemini Deep Research Agent and Gemini ProNative Search) frequently over-explore when confronted with noisy web evidence, spending excessive time or issuing excessive tool calls without converging on an answer. Gemini Deep Research times out on an average of 33.1% of questions, continuing to iteratively search and read for up to 15 minutes without producing a final answerâgetting lost in the noise of conflicting or tangentially relevant web content. A similar pattern emerges for Gemini Pro models underNative Search: an average of 12.7% of questions triggerTOOMANYTOOLCALLS, where the model exceeds the APIâs internal limit on search invocations, resulting in empty responses. In contrast, Flash and Lite variants are far less affected (3.1% and 0.4%, respectively), as they issue fewer queries and converge more quickly. This suggests a counterintuitive trade-off: more capable agents are more prone to over-exploration, issuing more search queries in an attempt to gather comprehensive evidence, but ultimately failing to answer within platform constraints. 4 Additional Analysis 4.1 Impact of Adding Video Processing Tool AsNative Searchis limited to text and image modalities and cannot process video or audio during search (Table 6 in Appendix), we investigate the effect of augmenting it with a video processing tool. As shown in Figure 3, we observe that when adding a video processing tool, the performance consistently increases with an average of 5.7% absolute improvement over the four Gemini agents (Gemini 3 Flash/Pro, Gemini 3.1 Lite/Pro), highlighting the importance of enabling access to a broader range of modalities for effective multimodal reasoning. More analysis in Appendix E.1. 4.2 Impact of Thinking Effort To analyze how varying levels of thinking effort affect performance onMERRIN, we conduct experiments across three search frameworks using GPT-5.4-mini. 3 We observe 3 The analysis is conducted with GPT-5.4-mini, as Gemini series does not support disabling thinking. 7 Preprint. Under review. SearchesVisits 0 1 2 3 4 5 6 7 8 9 Count Search EÂort Human Agentic PrecisionRecall 0 20 40 60 Score (%) URL Overlap w/ Golden Figure 5: Search effort and URL overlap com- parison between humans andAgentic Multi- modal Search(Gemini-3.1-Pro). Wrong Count 43% Right Source, Wrong Detail 29% Partial/Imprecise Answer 14% Others 14% Figure 6: Distribution of human error types, categorized by failure mode. Acc.Acc@5minTimeSearch EffortModality (%)URL Overlap w/ Golden System(%)(%)(min)SearchesVisitsURLsTextVideoImagePrec.Rec.F1 Human71.459.24.12.92.9â53.228.218.538.148.942.8 Native30.929.62.39.80.134.996.20.03.8â Agentic40.134.04.09.13.563.687.04.48.51.861.43.6 Table 3: Performance across human annotators,Native Search(Native), andAgentic Multi- modal Search(Agentic) with Gemini 3.1 Pro.Acc@5min: accuracy under a 5-minute budget, where any question taking more than 5 minutes is counted as incorrect.Time: average com- pletion time in minutes.Search Effort: average number of search queries issued (Searches), webpages visited (Visits), and unique URLs encountered (URLs) per question.Modality: distribution of resource modalities among accessed content.URL Overlap w/ Golden: precision, recall, and F1 of the systemâs visited URLs against the golden reference URLs. that performance generally improves as thinking effort increases, with the largest gains observed inAgentic Multimodal Search, which shows an absolute improvement of 8.6% when comparing no thinking to the highest level of thinking effort.Native Searchfollows with a 6.8% improvement, whileNo Searchshows a smaller gain of 3.1%. 4.3 Decomposing the Performance Gap: Search vs. Reasoning Setup.To isolate whether performance limitations stem from the search stage or the reasoning stage, we conduct experiments using Gemini-3.1-Pro that progressively provide gold evidence (Table 4). Starting fromAgentic Multimodal Search,+ Gold Sources Injection injects gold source URLs into every web search response alongside live search results, while the agent retains all tools and must identify the gold URLs among noisy results.+ Gold Sources Onlyremoves web search entirely and provides only gold URLs, with tools still available for processing.Gold Sources Promptingbypasses the agent framework entirely: gold videos and images are provided as native multimodal inputs, and web pages are fetched via URL context, in a single forward pass with no tools. Takeaways.Agentic Multimodal Search+ Gold Sources Injectionyields only a modest improvement (+3.3%), suggesting that searchavailabilityalone is insufficientâthe agent must also correctlyselectandprioritizerelevant sources among noisy web results.Agentic Multimodal Search+ Gold Sources Onlyfurther improves accuracy (+2.1%), confirming that real-world distractors actively degrade agent reasoning even when gold evidence is present. Gold Sources Promptingyields an additional +2.2%, revealing that even when gold sources are provided, the agent does not always call tools to deeply investigate themâinstead relying on surface-level information from URL titles or snippets rather than thoroughly examining the source content. From open search to perfect gold evidence, the total accuracy gain is 7.6% (40.1%â47.7%), upper-bounding the cost of search-stage limitations. However, even with perfect gold evidence, accuracy remains relatively low, indicating that while both search effectiveness and multimodal reasoning remain critical open challenges, improving reasoning capabilities is the more pressing bottleneck on MERRIN. 8 Preprint. Under review. SettingWeb SearchGold SourcesAgent ToolsAcc. No Searchâ24.7 ±1.6 Native Searchââ29.0 ±1.1 Agentic Multimodal Searchâââ40.1 ±2.8 +Gold Sources Injectionâ43.4 ±3.8 +Gold Sources Onlyââ45.5 ±2.3 Gold Sources Promptingâââ47.7 ±2.0 Table 4: Isolating search vs. reasoning limitations for Gemini-3.1-Pro onMERRIN.Web Search: whether the agent can search the open web.Gold Sources: whether gold sources are provided.Agent Tools: whether the agent can use custom tools (visitwebpage, watchvideo) to process evidence. 4.4 Human Performance We conduct a human evaluation to analyze performance and compare with agents in MERRIN. We recruit five undergraduate students to answer a randomly selected subset of 50MERRINquestions using standard web search, without AI assistance. Annotators record their answer, total time spent, number of search queries, and every resource consulted along with its relevance, modality, and URL. We analyze human behavior, error patterns, and the effect of time on performance. Comparing Human and Agentsâ Search Behavior.As shown in Table 3, humans achieve 71.4% accuracy, substantially outperforming bothAgentic Multimodal Search(40.1%) and Native Search(30.9%) using Gemini-3.1-Pro. Humans use far fewer resources, averaging 2.9 searches and 2.9 website visits, compared to 9.1 searches and 3.5 visits forAgentic Multimodal Search. Humans also achieve substantially higher precision in visited URLs (38.1% vs. 1.8%), indicating more effective source selection. Although the agentic system attains high recall (61.4%) due to the sheer volume of URLs encountered, its low precision indicates that the vast majority of retrieved sources are irrelevant. Moreover, humans rely on a balanced mix of modalities (53.2% text, 28.2% video, 18.5% image), whereas model-based systems are heavily text-dominant (87.0% text for the agentic system; 96.2% for the native system), with minimal video or image use. Effect of Time on Performance.A striking finding emerges when comparing accuracy un- der a five-minute budget (Acc@5min, where questions exceeding five minutes are counted as incorrect) to overall accuracy. Humans benefit substantially from additional time: their Acc@5min is 59.2%, rising to 71.4% overall, a gain of 12.2 %. This indicates that humans can productively leverage extra time to solve harder questions that require deeper search. In con- trast, agents show minimal improvement from additional time. The native system improves by only 1.3 % (29.6% to 30.9%), and the agentic system by 6.1 % (34.0% to 40.1%)âfar less than the human gain despite comparable average completion times (4.0 min for the agentic system vs. 4.1 min for humans). These results point to a fundamental limitation of current search-augmented agents: unlike humans, who efficiently identify high-quality sources and extract relevant information even on difficult, time-consuming questions, agents struggle to synthesize information effectively as they process more content over longer reasoning chains, gaining little from the additional computation. This finding is consistent with the over-exploration pattern described in Section 3.3: rather than productively deepening their search on difficult questions as humans do, agents tend to issue redundant queries and process tangentially relevant content, failing to converge. Human Error Analysis.To better understand the nature of human errors, we categorize each incorrect human response into one of four categories based on the type of error made (Figure 6). Among the responses where annotators provided an incorrect answer, the errors are predominantly minor extraction mistakes.Wrong Count(43%) captures cases where the annotator identified the correct source but miscounted by a small margin (e.g., off by one album cover or one second of video).Right Source, Wrong Detail(29%) includes cases 9 Preprint. Under review. where the annotator found the correct resource but extracted the wrong detail, such as reading a value from the wrong moment in a video or answering a different aspect of a multi-hop question.Partial/Imprecise Answer(14%) covers responses that were on the right track but insufficiently specific (e.g., âconservation lawâ instead of âconservation of chargeâ). Only 14% of errors fall intoOthers, representing genuinely incorrect answers. These results indicate that humans are generally able to identify the correct source and reasoning path, but often fail to extract precise informationâunderscoring that the benchmarkâs difficulty lies in fine-grained multimodal information extraction rather than source discovery, and highlighting the need for search agents that can effectively assist with such tasks. 5 Related Work Multimodal Search Benchmarks.Prior work has focused on developing multimodal, search-augmented evaluation benchmarks. Many of these benchmarks either provide multimodal inputs or include explicit modality cues that guide search agents toward which modalities to retrieve (Li et al., 2025; Geng et al., 2026; Zhang et al., 2026; Jiang et al., 2025; Tao et al., 2026). This design limits the ability to assess whether search agents can independently identify and retrieve the appropriate modality (e.g., whether the model can select audio sources or transcripts when a question asks about âwhat someone saysâ, even in the absence of explicit modality cues). In addition, prior work often focuses on a limited subset of modalities, primarily text and images, while overlooking others such as video and audio, which are common in real-world queries (Jia et al., 2025; Yan et al., 2025; Tian et al., 2025). This restricts the evaluation of search agentsâ ability to perform multimodal reasoning across diverse modalities. To address these limitations, we introduceMERRIN, which consists of natural language queries without explicit modality source cues and includes questions that require multi-hop reasoning over a broader range of modalities. Benchmarks for Reasoning under Web Noise.Prior work in the text domain shows that ambiguous, conflicting, and incomplete multi-source information can significantly degrade model performance, highlighting the importance of handling web noise (Wang et al., 2025; Lee et al., 2024; Pan et al., 2023). Similar challenges have been explored in multimodal settings, but most focus on scenarios where such complexity is synthetically introduced, or where conflicts are constructed over predefined evidence segments within curated multimodal corpora, limiting the diversity and realism of noise compared to a realistic open-web environment (Tian et al., 2025; Zhang et al., 2025; Semnani et al., 2025; Yan et al., 2025; Jia et al., 2025; Wu et al., 2025). There is also a line of work on benchmarks for search-augmented agents that operate in open-web settings (Li et al., 2025; Tao et al., 2026; Geng et al., 2026; Jiang et al., 2025), but do not explicitly analyze how web noise affects reasoning or how agents respond to it. They often focus on limited modalities, primarily text and images. In contrast,MERRINexplicitly induces web noise and requires search agents to reason across diverse modalities, including video and audio. 6 Conclusion We introducedMERRIN, a human-annotated benchmark for evaluating search-augmented agents on multimodal evidence retrieval and reasoning in noisy web environments. It uses natural language queries without modality cues, spans diverse modalities (including video and audio), and requires reasoning over noisy, conflicting, and incomplete web evidence. Evaluating search agents powered by ten different LLMs across three search settings, we find thatMERRINis highly challenging: average accuracy is 22.3%, with the best-performing configuration achieving only 40.1%. Compared to humans, agents are both less accurate and less efficient; humans search fewer but more precise queries and leverage more diverse modalities. These results highlight the importance ofMERRINas a benchmark for evaluating search agents in challenging and realistic settings. 10 Preprint. Under review. Ethics Statement While our dataset is constructed from publicly available web content, which may contain private or sensitive information, we mitigate these risks through human annotation and careful review. All data is screened to ensure that no private, biased, or harmful content is included. Acknowledgments We would like to thank Nithin Sivakumaran, Tianyi Niu, Vu Hoang Thien An, Dylan Zhao, and Hanqi Xiao for their contributions to the human evaluation. This work was supported by ONR Grant N00014-23-1-2356, ARO Award W911NF2110220, NSF-CAREER Award 1846185, NSF AI Engage Institute DRL2112635, Microsoft Agentic AI Research and Innovation (AARI) program, and a Google PhD Fellowship. The views contained in this article are those of the authors and not of the funding agency. References Yang Chen, Hexiang Hu, Yi Luan, Haitian Sun, Soravit Changpinyo, Alan Ritter, and Ming- Wei Chang. Can pre-trained vision and language models answer visual information- seeking questions? InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Singapore, December 2023. Association for Computational Linguis- tics. Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. M3docvqa: Multi- modal multi-page multi-document understanding. InProceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV) Workshops, p. 6237â6247, October 2025. Xinyu Geng, Peng Xia, Zhen Zhang, Xinyu Wang, Qiuchen Wang, Ruixue Ding, Chenxi Wang, Jialong Wu, Kuan Li, Yida Zhao, Huifeng Yin, Yong Jiang, Pengjun Xie, Fei Huang, Huaxiu Yao, Yi R. Fung, and Jingren Zhou. Webwatcher: Breaking new frontiers of vision-language deep research agent. InThe Fourteenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=8jsaazdAb3. Google. Gemini 3, 2025a. URLhttps://aistudio.google.com/models/gemini-3. Google. Gemini deep research agent, 2025b. URLhttps://ai.google.dev/gemini-api/ docs/deep-research. Google. Gemini-3.1, 2026. URLhttps://deepmind.google/models/gemini/pro/. Yifan Jia, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin, Rui Yang, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang, et al. Benchmarking multimodal knowledge conflict for large multimodal models.arXiv preprint arXiv:2505.19509, 2025. Dongzhi Jiang, Renrui Zhang, Ziyu Guo, Yanmin Wu, jiayi lei, Pengshuo Qiu, Pan Lu, Zehui Chen, Guanglu Song, Peng Gao, Yu Liu, Chunyuan Li, and Hongsheng Li. MMSearch: Unveiling the potential of large models as multi-modal search engines. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net/ forum?id=J2Jyp1SZ0n. Hyunji Lee, Se June Joo, Chaeeun Kim, Joel Jang, Doyoung Kim, Kyoung-Woon On, and Minjoon Seo. How well do large language models truly ground? InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 2437â2465, 2024. Hyunji Lee, Franck Dernoncourt, Trung Bui, and Seunghyun Yoon. CORG: Generating answers from complex, interrelated contexts. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Association for Computational Linguistics, April 2025. 11 Preprint. Under review. Shilong Li, Xingyuan Bu, Wenjie Wang, Jiaheng Liu, Jun Dong, Haoyang He, Hao Lu, Haozhe Zhang, Chenchen Jing, Zhen Li, et al. Mm-browsecomp: A comprehensive benchmark for multimodal browsing agents.arXiv preprint arXiv:2508.13186, 2025. OpenAI. Chatgpt web interface. URLhttps://chatgpt.com/. OpenAI. GPT 5.4, 2026. URLhttps://openai.com/index/introducing-gpt-5-4/. Liangming Pan, Wenhu Chen, Min-Yen Kan, and William Yang Wang. Attacking open- domain question answering by injecting misinformation. InProceedings of the 13th Interna- tional Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Nusa Dua, Bali, November 2023. Association for Computational Linguistics. Thinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen, Yu-Min Tseng, and Tu Vu. SealQA: Raising the bar for reasoning in search-augmented language models. InThe Fourteenth International Conference on Learning Representations, 2026. URLhttps: //openreview.net/forum?id=zWb7ueH16c. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. GPQA: A graduate-level google- proof q&a benchmark. InFirst Conference on Language Modeling, 2024. URLhttps: //openreview.net/forum?id=Ti67584b98. Aymeric Roucher, Albert Villanova del Moral, Thomas Wolf, Leandro von Werra, and Erik Kaunism Ì aki. âsmolagentsâ: a smol library to build great agentic systems.https: //github.com/huggingface/smolagents, 2025. Sina Semnani, Jirayu Burapacheep, Arpandeep Khatua, Thanawan Atchariyachanvanit, Zheng Wang, and Monica Lam. Detecting corpus-level knowledge inconsistencies in Wikipedia with large language models. InProceedings of the 2025 Conference on Empir- ical Methods in Natural Language Processing. Association for Computational Linguistics, November 2025. Liyan Tang, Grace Kim, Xinyu Zhao, Thom Lake, Wenxuan Ding, Fangcong Yin, Prasann Singhal, Manya Wadhwa, Zeyu Leo Liu, Zayne Rea Sprague, Ramya Namuduri, Bodun Hu, Juan Diego Rodriguez, Puyuan Peng, and Greg Durrett. Chartmuseum: Testing visual reasoning capabilities of large vision-language models. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2025. URLhttps://openreview.net/forum?id=qLdX6TA19s. Xijia Tao, Teng Yihua, Xinxing Su, Xinyu Fu, Jihao Wu, Chaofan Tao, Ziru Liu, Haoli Bai, Rui Liu, and Lingpeng Kong. MMSearch-plus: Benchmarking provenance-aware search for multimodal browsing agents. InThe Fourteenth International Conference on Learning Representations, 2026. URLhttps://openreview.net/forum?id=VGYgG2GH0d. Baoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li, Zhongyuan Bao, Zineng Zhou, Tao Wang, Sixu Li, Ziyao Xu, Mingze Wang, Zhouzhuo Zhang, Zhihao Wang, Yike Yun, Ke Tian, Ning Yang, and Minghui Qiu. Crosscheck-bench: Diagnosing compositional failures in multimodal conflict resolution.arXiv preprint arXiv:2511.21717, 2025. Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. Retrieval-augmented generation with conflicting evidence. InSecond Conference on Language Modeling, 2025. URLhttps://openreview.net/forum?id=z1MHB2m3V9. Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, Jeffrey Han, Isa Fulford, Hyung Won Chung, Alex Tachard Passos, William Fedus, and Amelia Glaese. Browsec- omp: A simple yet challenging benchmark for browsing agents.arXiv preprint arXiv:2504.12516, 2025. Chen Henry Wu, Neil Kale, and Aditi Raghunathan. Mitigating modal imbalance in multimodal reasoning. InSecond Conference on Language Modeling, 2025. URLhttps: //openreview.net/forum?id=JsaXxGOXfU. 12 Preprint. Under review. Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. Knowledge conflicts for LLMs: A survey. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (eds.),Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 8541â8565, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main.486. URLhttps: //aclanthology.org/2024.emnlp-main.486/. Qianqi Yan, Yue Fan, Hongquan Li, Shan Jiang, Yang Zhao, Xinze Guan, Ching-Chen Kuo, and Xin Eric Wang. Multimodal inconsistency reasoning (MMIR): A new benchmark for multimodal reasoning models. InFindings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, July 2025. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. Qwen3 technical report.arXiv preprint arXiv:2505.09388, 2025. Huanyao Zhang, Jiepeng Zhou, Bo Li, Bowen Zhou, Yanzhe Dan, Haishan Lu, Zhiyong Cao, Jiaoyang Chen, Yuqian Han, Zinan Sheng, et al. BrowseComp-V 3 : A visual, vertical, and verifiable benchmark for multimodal browsing agents.arXiv preprint arXiv:2602.12876, 2026. Zongmeng Zhang, Wengang Zhou, Jie Zhao, and Houqiang Li. Robust multimodal large language models against modality conflict.arXiv preprint arXiv:2507.07151, 2025. A Limitations Our benchmark relies on Google Search as the primary search engine, which may introduce biases specific to its ranking algorithms; future work could validate findings across multiple search engines. The dataset comprises 162 questions, which, while comparable to other expert-vetted diagnostic benchmarks, may not capture the full diversity of real-world multimodal queries. Additionally, web content is inherently dynamicâURLs may become unavailable or content may change over time, potentially affecting reproducibility. We plan to regularly update the dataset to address this. BMERRIN Details B.1 Data Collection Details Annotation Fields.For each question, annotators record: (a) the ground-truth answer along with a detailed explanation of the reasoning steps; (b) source URLs (e.g., webpages, videos, PDFs) used as supporting evidence for deriving the reference answer; (c) the source type of each resource (text, image, or video 4 ); (d) the multimodal role, indicating whether non-text evidence serves as the answer source or as a reasoning component; (e) the reasoning type, indicating whether the question is multi-hop and whether it introduces multimodal conflict; and (f) the source of the question, labeled asfrom scratchif both the question and its supporting evidence are newly constructed, or by the name of the originating dataset (e.g., SealQA) if the question or evidence is adapted from an existing source. Question Design Examples.To illustrate the no-modality-cues requirement, consider the following: instead ofâIn the attached chart, what year did Safari surpass IE in global browser market share?â, we askâAccording to StatCounter data, in which year did Safari surpass IE in global browser market share?âThis phrasing avoids referencing any specific visual or auditory content while still requiring non-text evidence (a chart) to answer correctly. 4 Here, video sources incorporate visual and audio modalities. 13 Preprint. Under review. Source# QAvg Q LenAvg # Res Existing Benchmarks4223.42.1 From Scratch 12025.21.9 Total16224.72.0 Table 5: Dataset statistics forMERRIN.# Q: number of questions;Avg Q Len: average question length in words;Avg # Res: average number of gold resources per question. Long ago Pre-2010 2010-142015-19 2020202120222023202420252026 0 5 10 15 20 25 30 35 40 # Âestions 18 15 12 22 4 1 4 11 19 36 20 (a)Effective Year Never-changing Slow-changing Fast-changing 0 15 30 45 60 75 90 105 # Âestions 94 38 30 (b)Freshness Figure 7: Temporal characteristics ofMERRIN. (a) Distribution byEffective Yearâ the year in which the answer first became correct. Sparse years before 2020 are collapsed into Pre-2010, 2010â14, and 2015â19 buckets, while recent years are shown individually. (b) Distribution byFreshnessâ how time-sensitive the ground-truth answer is. For cases adapted from existing datasets, we use their questionâanswer pair as one piece of evidence (i.e., one hop) and augment it with additional evidence to constructnewmulti- hop questions. To ensure broad coverage, annotators are encouraged to include non-text evidence from at least two different source types where possible. B.2 Quality Control Details We employ a rigorous multi-round human review process. After initial question construc- tion, each question is reviewed by a second annotator to assess: (1) answer correctness, (2) question clarity, (3) question difficulty, and (4) non-text modality requirements. Questions that fail any stage are revised and re-validated through subsequent review rounds. Answer Correctness.Annotators independently verify the ground-truth answer by re- deriving it from the cited sources, flagging any discrepancies. Question Clarity.Annotators check that each question is unambiguous, self-contained, and free of grammatical errors. Question Difficulty.We evaluate each question using ChatGPTâs web interface (OpenAI) with web browsing enabled. Questions that are consistently answered correctly are revised or removed to maintain the desired level of challenge. Non-Text Modality Verification.To verify that each question requires at least one non-text modality and cannot be solved with text alone, we apply a two-pass verification protocol: 14 Preprint. Under review. 1.Standard search pass:The annotator decomposes each multi-hop question into con- stituent sub-questions. For each sub-question, the annotator attempts to answer it using text-only search via Google Search, simulating a standard retrieval setting. 2. Adversarial search pass:Given the ground-truth answer, the annotator queries the sub-question together with the answer string (e.g., submitting bothâWho directed X?âand the correct director name) to check whether any text-only document contains or implies the correct answer. This step is designed to uncover potential text-only shortcuts that a sufficiently capable retrieval system could exploit. We encourage the annotator to check as thoroughly as possible but limit it to up to 20 web searches. A question passes this check only if at least one sub-question cannot be resolved via text-only evidence underboththe standard and adversarial search passes. Rejection Statistics.In the first round, approximately 39.5% of initial candidates were rejected. Of the rejected questions, 45.3% were successfully revised and accepted in the second round. Human Annotation Instruction Goal.Create challenging, multi-hop questions that require non-text evidence (images, videos, audio) to answer, designed to evaluate search-augmented agents. Core Requirements.Every question must satisfy all three: 1. No modality cues.Questions must be phrased in natural language without explicitly specifying the exact modality source (e.g., âIn the first episode of Rick and Morty Season 8, ...â, âIn this image...â). The question should read like a realistic user query. 2.Non-text evidence required.At least one reasoning step must require non-text evidence (image, video, audio, chart, etc.) that cannot be resolved through text-only web search. This is verified through a two-pass protocol (see Quality Control). 3. Single unambiguous answer.Each question must have exactly one correct, short, and verifiable answer. Classification Axes.Annotate each question along two axes: âąReasoning type(one or both):multi-hop(combining information across multiple sources or modalities) and/ormultimodal conflict(the question naturally triggers conflicting evidence across modalities in real search engines; donotsynthetically introduce conflicts). âąMultimodal role(one or both): non-text evidence servesas answer(the answer can only be extracted from a non-text source) and/oras reasoning chain(non-text evidence provides an intermediate fact needed to derive the final answer). Annotation Fields.For each question, record: âą Ground-truth answer with a detailed explanation of the reasoning steps. âą Source URLs (webpages, videos, PDFs, images) used as supporting evidence. âą Source type of each resource (text, image, video/audio). âą Multimodal role and reasoning type labels (as defined above). âą Question origin: âfrom scratchâ or the name of the originating dataset if adapted. Tips for Question Construction. âą Combine a text-retrievable fact (e.g., identifying an entity) with a visual detail that can only be found in an image or video (e.g., a color, count, or spatial arrangement). âą Avoid questions that can be answered by reading image captions, alt-text, or text surround- ing the image on a webpage. âą Include diverse source types (text, images, video, audio) and diverse topics. âąEnsure the question is challenging: test it with ChatGPT with web browsing. If it is consistently answered correctly, revise or replace. Figure 8: Instruction for Human Annotation B.3 Human Annotation Figure 8 shows the human annotation guidelines. 15 Preprint. Under review. Input QueryBuilt-in Search ModelContext (In/Out)TextImageVideoAudioTextImageVideoAudio GPT400k/128kââââ Gemini1M/64kââ Qwen3230k/32kââ---- Table 6: Model context window sizes and modality support.Context (In/Out): maximum input/output token limits.Input Query: modalities the model can accept as direct input. Built-in Search: modalities the model can process when using its built-in search tool under Native Search. B.4 Data Statistics Table 5 shows the overall data statistics. Beyond question and resource counts, we also annotate each question with two temporal dimensions: an Effective Year (the year in which the ground-truth answer first became valid) and a Freshness label (never-, slow-, or fast- changing), which together characterize how time-sensitiveMERRINis. Figure 7a shows the effective-year distribution: the bulk of the benchmark is concentrated in 2023â2026, with a long tail of older and time-agnostic (âLong agoâ) questions. Figure 7b shows the freshness distribution, which is dominated by never-changing questions (stable facts) but still includes a substantial fraction of slow- and fast-changing questions whose answers drift over time. C Experiments C.1 Setup Search Setting Details.All closed-source models are evaluated via their official APIs. Supporting modalities and maximum context lengths for each model are listed in Table 6. In theNo Searchsetting, models are evaluated without access to any tools. In theNative Searchsetting, we enable each modelâs built-in search capabilities. Recent LLM APIs are agentic by default, allowing models to autonomously invoke built-in tools and perform multi-turn reasoning within a single API call. For GPT models, we enable thewebsearch tool, which supports searching the web, opening specific pages, and searching within pages, providing both retrieval and in-depth webpage understanding in a unified tool. For Gemini models, we enable bothGoogle Search(web retrieval) andURL Context(webpage comprehension) to match the combined functionality of GPTâsweb searchtool. In theAgentic Multimodal Searchsetting, we use a multimodal search agent framework built on smolagents (Roucher et al., 2025) that equips models with tools extending their effective modality coverage. In addition to the built-inwebsearchtool (leveraging the Serper API for Google search), we incorporate two custom tools:visitwebpage, which enhances the default webpage toolâlimited to converting pages into markdown stringsâby using Gemini-3-Flash withURL Contextto interpret full webpage content including text and images; andwatchvideo, which uses Gemini-3-Flash to directly process YouTube videos, enabling the agent to understand visual and audio content. Evaluation Details.We evaluate using an LLM-as-judge with the same prompt as BrowseComp (Wei et al., 2025). Prompt can be seen in Figure 9. D Human Evaluation Human evaluation guidelines can be found in Figure 10. 16 Preprint. Under review. Autorater Prompt Judge whether the following [response] to [question] is correct or not based on the precise and unambiguous [correctanswer] below. [question]:question [response]:response Your judgement must be in the format and criteria specified below: extractedfinalanswer:The final exact answer extracted from the [response]. Put the extracted answer as âNoneâ if there is no exact, final answer to extract from the response. [correctanswer]:correctanswer reasoning:Explain why the extractedfinalanswer is correct or incorrect based on [correctanswer], focusing only on if there are meaningful differences between [correctanswer] and the extractedfinalanswer. Do not comment on any background to the problem, do not attempt to solve the problem, do not argue for any answer different than [correctanswer], focus only on whether the answers match. correct:Answer âyesâ if extractedfinalanswer matches the [correctanswer] given above, or is within a small margin of error for numerical problems. Answer ânoâ otherwise, i.e. if there is any inconsistency, ambiguity, non-equivalency, or if the extracted answer is incorrect. confidence:The extracted confidence score between 0% and 100% from [response]. Put 100 if there is no confidence score available. Figure 9: Autorater prompt used for grading responses.Placeholdersquestion, response, andcorrectanswerare filled at evaluation time. ModelNo SearchNative SearchNative Search + Video ToolAgentic Multimodal Search Gemini 3 Flash19.1 ±3.2 31.7 ±3.8 36.8 ±2.6 32.9 ±0.9 Gemini 3 Pro23.5 ±1.1 28.8 ±1.4 35.6 ±0.4 39.9 ±1.6 Gemini 3.1 Lite12.8 ±2.3 20.6 ±2.2 21.6 ±1.2 26.3 ±1.9 Gemini 3.1 Pro24.7 ±1.6 29.0 ±1.1 37.5 ±2.0 40.1 ±2.8 Table 7: Impact of adding a video processing tool toNative Search. Accuracy (%) across four Gemini models under four settings:No Search,Native Search,Native Search + Video Tool (adding video processing tool), andAgentic Multimodal Search(full multimodal agent). E Analysis E.1 Impact of Adding Video Processing Tool As shown in Table 7, adding a video processing tool toNative Searchconsistently improves accuracy, with gains ranging from +1.0% (Gemini 3.1 Lite) to +8.5% (Gemini 3.1 Pro), av- eraging +5.7% across agents. This confirms that video evidence is critical for a substantial portion ofMERRINquestions and thatNative Searchâs inability to process video is a sig- nificant limitation. ComparingNative Searchwith the video tool toAgentic Multimodal Search, we observe thatAgentic Multimodal Searchstill outperforms on three of four agents (Gemini 3 Pro: 39.9% vs. 35.6%, Gemini 3.1 Lite: 26.3% vs. 21.6%, Gemini 3.1 Pro: 40.1% vs. 17 Preprint. Under review. 37.5%), except for Gemini 3 Flash (32.9% vs. 36.8%). We attribute this gap to two factors: (1) Native Searchsometimes fails to invoke the video tool, whereasAgentic Multimodal Search proactively callswatchvideo; and (2)Agentic Multimodal Searchlocates more relevant videos through its dedicatedsearchvideotool, whileNative Searchrelies on the built-in Google Search, which sometimes retrieves irrelevant videos. 18 Preprint. Under review. Human Evaluation Instruction Overview You will be presented with a set of questions in a Google Sheets spreadsheet and asked to find the correct answer to each using the open web. Each question has a single, unambiguous short answer(e.g., a name, number, date, or brief phrase). How to Use the Spreadsheet Each row in your assigned Google Sheet contains one question. For each row, fill in the following columns: ColumnWhat to enter QuestionAlready filled in â read carefully. Your AnswerYour short answer to the question. Annotation Time How long it took you to find the answer (e.g., â2 minâ, â4 minâ). Numberof Queries The total number of search queries you made for this ques- tion. Resource 1, 2, . . .Every resource (webpage, image, video, etc.) you opened while searching. For each, record three things separated by commas: (1) whether the resource wasrelevantornot relevant, (2) the modality (e.g.,text,image,video,table), and (3) the URL. Your Task 1.Read the question carefully.Each question is written in plain natural language. 2.Search the web to find the answer.Use Google Search for entering queries and find websites, videos, etc. to find relevant information. 3.Do NOT close any tabs while searching.Keep all tabs open throughout your search so that you can accurately record all resources and queries at the end. 4.Record your answerin the âYour Answerâ column. 5.Record the timein minutes it took you in the âAnnotation Timeâ column. 6. Count your search queries.After finding the answer (or giving up), go through your browser history/tabs and count the total number of search queries you made. 7. Record ALL resources.Go through every tab you opened and record each one in the Resource columns â not just the ones that contained the answer. For each resource, indicate whether it wasrelevantornot relevant, the modality, and the URL. Answer Format âąKeep your answershort and preciseâ typically a single entity, value, or brief phrase. âą Donotinclude full sentences or explanations unless specifically requested. Important Notes âą Do NOT use any AI tools.Do not use ChatGPT, Google Search AI mode (AI Overviews), Bing Copilot, Perplexity, or any other AI-powered assistant. âąDo NOT close your tabs.Keep all tabs open while working on each question. âą Examine all resources carefully andrecord every resource, even unhelpful ones. âąVerify before submitting.Cross-check your answer with at least one additional source when possible. âą If you cannot find the answer, write âUnanswerableâ and briefly note what you tried. Figure 10: Instruction for Human Evaluation 19