Paper deep dive
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark
Min Wang, Ata Mahjoubfar
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/31/2026, 2:47:26 AM
Summary
AMIGO (Agentic Multi-Image Grounding Oracle Benchmark) is a long-horizon, interactive benchmark designed to evaluate agentic vision-language models (VLMs) on hidden-target identification tasks. The benchmark requires models to identify a target image from a gallery of visually similar candidates by asking constrained Yes/No/Unsure questions, while adhering to a strict protocol that penalizes invalid actions with 'Skip' responses. It emphasizes long-horizon constraint tracking, fine-grained visual discrimination, and robustness to oracle imperfections.
Entities (4)
Relation Signals (2)
AMIGO → instantiatedby → Guess My Preferred Dress
confidence 100% · We instantiate AMIGO with Guess My Preferred Dress task
Vision-Language Models → evaluatedby → AMIGO
confidence 95% · We introduce AMIGO... for hidden-target identification... to evaluate agentic VLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Agentic vision-language models increasingly act through extended interactions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identification over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No/Unsure questions under a strict protocol that penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (ii) consistent constraint tracking across turns, and (iii) fine-grained discrimination as evidence accumulates. AMIGO also supports controlled oracle imperfections to probe robustness and verification behavior under inconsistent feedback. We instantiate AMIGO with Guess My Preferred Dress task and report metrics covering both outcomes and interaction quality, including identification success, evidence verification, efficiency, protocol compliance, noise tolerance, and trajectory-level diagnostics.
Tags
Links
- Source: https://arxiv.org/abs/2603.28662v1
- Canonical: https://arxiv.org/abs/2603.28662v1
Trouble viewing inline? Open PDF directly →
Full Text
53,870 characters extracted from source content.
Expand or collapse full text
AMIGO: Agentic Multi-Image Grounding Oracle Benchmark Min Wang, Ata Mahjoubfar Target Corporation Min.Wang, Ata.Mahjoubfar@target.com Work in Progress Abstract Agentic vision-language models increasingly act through extended in- teractions, but most evaluations still focus on single-image, single-turn correctness. We introduce AMIGO (Agentic Multi-Image Grounding Oracle Benchmark), a long-horizon benchmark for hidden-target identi- fication over galleries of visually similar images. In AMIGO, the oracle privately selects a target image, and the model must recover it by asking a sequence of attribute-focused Yes/No/Unsure questions under a strict protocol that penalizes invalid actions with Skip. This setting stresses (i) question selection under uncertainty, (i) consistent constraint tracking across turns, and (i) fine-grained discrimination as evidence accumulates. AMIGO also supports controlled oracle imperfections to probe robustness and verification behavior under inconsistent feedback. We instantiate AMIGO with Guess My Preferred Dress task and report metrics covering both outcomes and interaction quality, including identification success, evidence verification, efficiency, protocol compliance, noise tolerance, and trajectory-level diagnostics. 1 Introduction Vision-language models (VLMs) [1,2,3,4,5] have advanced rapidly in instruction following and grounded reasoning, enabling strong performance on captioning, visual question answering, and multimodal dialogue. As VLMs become more agentic (planning and acting through extended exchanges), evaluation protocols increasingly need to measure not only correctness, but also interaction policies: how models gather information, maintain state, and adapt under uncertainty. Yet many benchmarks [6,7,8,9,10] remain static and short-horizon, typically scoring one-shot answers given one image (or a small fixed set), with limited visibility into strategy, memory, and belief revision over time. This work is still in progress. We’l update with more results and analysis soon. The dataset is provided as an ancillary file, anc/data.json. 1 arXiv:2603.28662v1 [cs.LG] 30 Mar 2026 We study a complementary setting: interactive hidden-target identification over image collections. A user (or oracle) privately selects a target image from a gallery of visually similar candidates, and the model must discover the target by asking constrained questions about observable attributes. This formulation surfaces agentic challenges that are often invisible in single-turn evaluation: choosing questions that efficiently reduce ambiguity, preserving and applying constraints without drift, and performing cross-image comparisons that depend on the evolving candidate set rather than any single image in isolation. To operationalize this setting, we present AMIGO, an interactive benchmark centered on a simple but diagnostic loop. The model asks exactly one binary question per turn, receives Yes/No/Unsure feedback, and iteratively narrows the feasible candidate set. If the model violates protocol rules, the oracle returns Skip, revealing no information. This makes non-compliance measurable and separates wasted interaction from evidence-driven progress. AMIGO additionally allows occasional oracle inconsistencies to test whether models detect contradictions and seek verification instead of committing prematurely. We instantiate AMIGO with Guess My Preferred Dress task, where each gallery contains 6 to 40+ highly similar dress images. Dresses offer rich fine- grained variation (e.g., neckline construction, closures, seam placement, drape, and embellishments), making the task sensitive to careful observation and long- horizon constraint tracking. Contributions. (1) We introduce AMIGO, an interactive multi-image bench- mark for hidden-target identification that foregrounds long-horizon planning and fine-grained cross-image grounding. (2) We define a constrained Yes/No/Unsure protocol with explicit invalid-action feedback (Skip) to diagnose sustained in- struction following and common agentic failure modes. (3) We incorporate controlled oracle imperfections to test robustness and verification behaviors under inconsistent feedback. (4) We provide an evaluation suite that jointly mea- sures outcomes and interaction quality, including verified identification, efficiency, protocol compliance, and trajectory-level diagnostics. Broader impact. AMIGO provides a controlled testbed for agentic VLM behaviors that matter in practice: planning under uncertainty, maintaining consistent state over extended interactions, and responding conservatively to ambiguity or contradictions. The resulting traces (candidate sets as state, questions as actions, and oracle responses, including Skip, as observations) can also support training and analysis of multimodal policies via offline learning. At the same time, hidden-target identification highlights calibration considerations: systems should avoid overconfident early commitments and remain protocol- compliant throughout long dialogues. 2 2 Related Work 2.1 Multi-image and multi-turn multimodal evaluation A growing set of benchmarks and datasets extends evaluation beyond single- image, single-turn QA [8,7] to multi-image and/or multi-turn settings. Multi- image benchmarks (e.g., MMIU [11], MuirBench [12], mPLUG-Owl3 [13]) probe cross-image grounding and long visual context handling. Multi-image instruction- tuning datasets (e.g., Mantis [14]) and multi-image multi-turn resources (e.g., MMDU [15], MMCR [16]) further study supervision over interleaved images and dialogue. Separately, multi-turn multimodal conversation benchmarks (e.g., MultiVerse [17], ConvBench [18], MMMT-IF [19], MMCoQA [20]) evaluate contextual coherence and instruction adherence over longer dialogues. These efforts primarily evaluate responding given a provided context, rather than interactive hidden-target identification where a model must actively uncover a user-privately selected target by asking constrained questions and tracking constraints over a long horizon. 2.2 Multimodal reasoning benchmarks General multimodal reasoning benchmarks assess whether models can integrate vision and language for multi-step inference beyond shallow cue matching. For ex- ample, EMMA [21] targets “organic” multimodal reasoning across domains, and MLLM-CompBench [22] emphasizes comparative reasoning under multimodal inputs. MMMU-Pro [23] further strengthens this line by making evaluation more resistant to text-only shortcuts and by requiring tighter visual grounding in challenging multimodal questions. These benchmarks are valuable for measuring reasoning capability on difficult perception-and-inference tasks, but they still primarily assess whether a model can solve a presented problem instance. In contrast, AMIGO evaluates whether an agentic VLM can acquire the missing in- formation needed to solve the task: selecting informative questions over multiple turns, maintaining and revising a belief state over a candidate pool, and enforcing consistency under a constrained interaction protocol to identify a hidden target image. 2.3 Interactive clarification with hidden targets AMIGO is closely related to interactive clarification settings where an agent must uncover user intent through dialogue. InfoQuest [24] evaluates text-only agents that ask clarifying questions when critical context is hidden. ClariMM [25] extends this idea to multimodal clarification for underspecified user queries. More broadly, multimodal agent frameworks such as M3Searcher [26] and expert- guided benchmarks such as MIRAGE [27] follow a “seek information then decide” paradigm, often with retrieval or tool-mediated evidence acquisition. AMIGO differs in three ways. First, it studies hidden-target identification over a closed set of visually similar images, where the model must narrow 3 the candidate pool via attribute queries answered with Yes/No/Unsure/Skip feedback (which may be noisy). Second, it evaluates the interaction policy itself via trajectory-level signals such as compliance, redundancy, elimination dynamics, and contradiction detection/verification. Third, the constrained protocol yields reusable trajectories with a well-defined action space (attribute questions), observations (oracle responses), and automatically scored intermediate signals, making the data naturally suited for diagnosing and training long-horizon agentic multimodal policies. 3 Benchmark We introduce AMIGO, a benchmark for evaluating agentic VLMs on hidden- target, multi-image, multi-turn identification. AMIGO comprises a primary task, Guess My Preferred Dress, where the model must identify a user-privately selected target image from a gallery of highly similar candidates by asking discriminative attribute questions under a strict protocol. 3.1 Guess My Preferred Dress Task. A user uploads one or more batches of dress images that together form a single gallery and there is one image that is privately selected by user as a hidden target. After the signalEnd of uploading, the model asks questions to narrow the candidate set and finally outputs the 1-indexed position of the target in the gallery. Protocol. The interaction follows these rules: 1.Each turn, the model must ask exactly one question about an observable attribute of the target dress. 2. The question must be answerable with Yes, No, or Unsure. 3. If the model violates any rule, the oracle responds with Skip. 4.The model must not ask about: sleeve length, garment length, color, prints/pattern, age group, size, shoes, necklace, hat, bag, background, or the human model. 5. The model must not enumerate attribute values across turns for the same attribute type (e.g., repeatedly cycling through neckline types of V-neckline, square neck, crewneck in different turns). 6.The model must not reference specific indices or inspect images one-by-one (e.g., “Is it image #1?”). 7. The model must not guess until its constraints narrow the feasible candidate set to exactly one image. 4 8.If constraints appear inconsistent (e.g., due to uncertainty or noise), the model should continue asking more questions rather than stopping early. 9.When ready, the model must output:My guess of your favorite dress: #<number>. 10. The model must not ask any questions before the user sendsEnd of uploading. Rationale. The protocol deliberately excludes highly salient cues (e.g., color, patterns, sleeve length, and garment length) to encourage reliance on subtler construction details (e.g., neckline structure, closures, waist shaping, seam placement, drape, and embellishments). This design emphasizes long-horizon constraint tracking and rewards question sequences that are informative, non- redundant, and robust to uncertainty. Example interactions. Figure 1 shows two full episodes. In Figure 1a, the model systematically asks about fine-grained details and narrows the pool over 10 turns, but makes an incorrect final guess, illustrating how visually similar candidates can still be confusable. In Figure 1b, the model identifies the target in 9 turns by asking progressively more discriminative questions. 3.2 Dataset collection Source and curation. We collect dress images from Target’s online fashion catalog and deduplicate them to ensure each product is unique. Attribute metadata extracted from the historical product database is often missing, defined inconsistently across brands, and sometimes contains erroneous or irrelevant values. We therefore prioritize visually verifiable attributes and build a semi- automatic pipeline to produce standardized labels at scale. Semi-automatic attribute labeling. Our pipeline has three stages: (1) Attribute discovery and normalization: we aggregate attribute values from catalog annotations and use multiple LLMs to normalize synonyms and naming variants, compile a comprehensive attribute value list, and expand it with as many visually grounded alternatives as possible; (2) Binary question template construction: for each normalized attribute value, we generate multiple binary (Yes/No) question templates (e.g., “Does the dress have a wrap-style front?”, “Does the dress have a wrap-front design?”, “Is the dress featured with wrap-style front?”); (3) VLM-based labeling: we apply multiple VLMs to answer these templates for each image, yielding per-image attribute labels. Ensembled VLM labeling and quality control. Given budget constraints, we adopt open-source VLMs for automated labeling. To improve labeling reliabil- ity, we ensemble multiple VLMs (Qwen3-VL-235B-FP8 [28], Intern-S1 [29], GLM-4.5V [30]) and combine their predictions through weighted majority 5 Image Gallery (8 candidates — target hidden from model) #1 dress_2 #2 ♥ Target (hidden) dress_3 #3 dress_4 #4 dress_5 #5 dress_6 #6 dress_7 #7 dress_8 #8 Task: Guess the user's favorite dress by asking Yes/No questions. Constraints: one binary question per turn · no color/pattern/sleeve/length · no image-by-image inspection · guess only when confident · responses: Yes / No / Unsure / Skip User: "End of uploading" T1 烙 Does your favorite dress have a separate belt with a buckle? No T2 烙 Does your favorite dress have buttons down the front? No T3 烙 Does your favorite dress have a smocked (elastic, gathered) bodice? No T4 烙 Does your favorite dress have an off-the-shoulder neckline? Yes T5 烙 Does your favorite dress have a fabric tie at the waist (a sash you knot, no buckle)? Yes T6 烙 Does your favorite dress have an asymmetrical cascading ruffle panel on the skirt? No T7 烙 Does your favorite dress have more than one distinct tier/ruffle layer in the skirt? Yes T8 烙 Does it have a visible elasticated waistband seam rather than a smooth fitted waist? Unsure T9 烙 Is your favorite dress a two-piece set (separate top and skirt) rather than one single dress? No T10 烙 Does your favorite dress have a peplum-like flounce at the waist (a short ruffled layer over the skirt)? Yes ★ 烙 My guess of your favorite dress: #4. ✗ Incorrect (ground truth: #2) Summary: 10 turns · 0 Skips · 4× No · 4× Yes · 1× Unsure · Final: Incorrect (guessed #4, target #2) Legend YesConfirmed attribute NoRuled-out attribute Unsure Ambiguous / uncertain (a) Unsuccessful attempt: after 10 turns the model guesses dress #4, but the hidden target is dress #2 (red border), illustrating how visually similar candidates can mislead even systematic questioning. Figure 1: Two example multi-turn interactions for Guess My Preferred Dress. The VLM (blue) asks constrained Yes/No questions about fine-grained attributes; the user replies with Yes (green), No (red), or Unsure (yellow). Panel (a) shows an incorrect final guess and panel (b) shows a correct one. 6 Image Gallery (6 candidates — target hidden from model) dress_1 #1 dress_2 #2 dress_3 #3 dress_4 #4 dress_5 #5 ♥ Target (hidden) dress_6 #6 Task: Guess the user's favorite dress by asking Yes/No questions. Constraints: one binary question per turn · no color/pattern/sleeve/length · no image-by-image inspection · guess only when confident · responses: Yes / No / Unsure / Skip User: "End of uploading" T1 烙 Does your favorite dress have pockets that you can comfortably put your hands into? Unsure T2 烙 Does your favorite dress have a waist tie/belt that wraps around and can be knotted or bowed? Yes T3 烙 Does your favorite dress have a wrap-style front where one side overlaps the other across the torso? Yes T4 烙 Does your favorite dress have a ruffled hem or ruffle detailing near the bottom edge? Yes T5 烙 Does your favorite dress have multiple horizontal 'tier' seams (stacked layers) in the skirt portion? Yes T6 烙 Does the skirt have an overlapping 'wrap' opening in front (so the front edge is formed by one panel laying over another, rather than being a continuous closed skirt)? No T7 烙 Does your favorite dress have a V-shaped neckline? Yes T8 烙 Does your favorite dress's fabric have a noticeable sheen (it looks somewhat shiny rather than matte)? Unsure T9 烙 Does your favorite dress have noticeably puffed shoulders (gathering/volume right at the shoulder seam)? No ★ 烙 My guess of your favorite dress: #5. ✓ Correct (target: #5) Summary: 9 turns · 0 Skips · 2× No · 5× Yes · 2× Unsure · Final: Correct (guessed #5, target #5) Legend YesConfirmed attribute NoRuled-out attribute Unsure Ambiguous / uncertain (b) Successful attempt: after 9 turns the model correctly guesses dress #5—the hidden target (red border)—showing that effective questioning can resolve visually similar candidates. Figure 1: Two example multi-turn interactions (continued ). 7 voting. In our experiments, Qwen3-VL-235B-Instruct-FP8 achieved the best performance among the models evaluated prior to the release of Qwen3.5-397B- A17B-FP8, and is therefore assigned a larger voting weight. We further enhance robustness by (i) using paraphrased question templates to reduce sensitivity to prompt phrasing and (i) applying multi-resolution image augmentation to mitigate failures in recognizing small visual details. Finally, we manually audit a subset of the labels to assess quality, identify systematic errors, and refine the templates and normalization rules accordingly. Figure 2: The semi-automatic attribute labeling pipeline: attribute discovery and normalization, binary question template construction, and ensembled VLM- based labeling with quality control. Attribute-based similarity. LetAttr(X) denote the set of attribute values assigned to an image X. We define an asymmetric similarity score: Sim(A,B) = |Attr(A)∩ Attr(B)| |Attr(A)| ,(1) and analogouslySim(B,A). We rank candidatesBbySim(A,B) to retrieve candidate images that are most visually similar to the reference imageAin terms of attribute overlap (i.e., candidates that best cover A’s attributes). Episode generation and difficulty control. For each target imageA, we construct a distractor pool by iterating over its attribute values. For each attribute value, we retrieve a fixed number of images that share that value withAand satisfySim(A,B)≥ τ, whereτis a similarity threshold. We then merge the retrieved sets across all attribute values to obtain the final candidate distractors forA. We keep only targets for which this merged pool contains more than five candidates. To form an episode, we combine the target with a subset of its candidate distractors to create a gallery. We control difficulty via (i) the thresholdτ(higherτyields more visually similar distractors) and (i) the gallery size (larger galleries expand the search space). 8 Figure 3: Image gallery generation pipeline: for a given target image, distrac- tors are retrieved by attribute-based similarity and merged into a gallery with controlled difficulty via threshold τ and gallery size. 3.3 Sample image galleries Figure 4 shows four representative galleries. Each gallery is curated to be similar along salient dimensions (e.g., silhouette and overall style) while differing in subtle construction details. The hidden target is marked with a red outline. Scale. We collect 4,880 unique dress images for episode generation. We use five similarity thresholds,τ ∈0.3,0.4,0.5,0.6,0.8, to control difficulty. Higher thresholds yield smaller but more confusable galleries; lower thresholds typically yield larger galleries with more diverse distractors (Figure 5). We obtain 587 episodes atτ= 0.8 because it is difficult to find enough highly similar distractors under a strict threshold. For evaluation at other thresholds, we randomly sample 1,000 episodes per threshold. 4 Benchmark framework Benchmark components. Our framework comprises four modules (Figure 6): (i) a benchmark model (the evaluated VLM) that generates questions and produces a final guess, (i) a question-violation detector (LLM-based) that enforces the protocol, flags invalid questions with Skip, and provides standardized feedback on rule adherence, (i) a user/oracle simulator (VLM agent) that answers valid questions with Yes/No/Unsure, and (iv) a verification module that maintains and audits the feasible candidate set implied by the dialogue. Given the uploaded gallery and dialogue history, the benchmark model outputs either the next question or a terminal guess. The violation detector 9 (a) Image Gallery 1 (b) Image Gallery 2 (c) Image Gallery 3 (d) Image Gallery 4 Figure 4: Four sample dress galleries from AMIGO. Each gallery contains one target image and visually similar distractors. The target is highlighted with a red outline. 10 (a) τ = 0.3(b) τ = 0.4 (c) τ = 0.5(d) τ = 0.6 (e) τ = 0.8 Figure 5: Distribution of gallery sizes across similarity thresholdsτ. Lower thresholds tend to yield larger galleries with more diverse distractors; higher thresholds produce smaller but more visually confusable candidate pools. 11 checks each question against the benchmark constraints. The violation detector operates in a few-shot manner and works on two tasks: (1) if the question is related to a prohibited attribute (e.g., color, sleeve length) or image index reference; or (2) if the attribute in the question has already been asked about in previous turns. In either case, the violation detector flags the question as invalid and returns Skip (and no oracle information is revealed). We record the number (and rate) of Skip responses per episode as the primary signal of protocol compliance. Implementing the violation detector as a separate module ensures that all models receive consistent, standardized feedback on rule adherence regardless of their internal architecture or reasoning style. The user/oracle simulator answers valid questions with Yes/No/Unsure. Figure 6: Overview of the AMIGO benchmark framework, illustrating the inter- action among the benchmark model, user/oracle simulator, question-violation detector, and verification module. Verification module (state tracking and consistency checking). After each valid question-answer pair, the verification module converts the interaction into an explicit constraint and applies it to the gallery to update the feasible candidate set, the subset of images consistent with all constraints observed so far. This module serves three purposes: (1) Evidence verification: it determines whether the dialogue has accumulated sufficient evidence to uniquely identify the target (i.e., the feasible set has size one). (2) Trajectory auditing: it logs candidate-set reduction dynamics (e.g., elimination rates, stalls) and supports trajectory-level diagnostics. (3) Consistency checking under noise: it detects contradictions signaled by an empty feasible set or conflicting constraints, 12 enabling analysis of whether the benchmark model responds conservatively (e.g., by asking following verification questions) rather than guessing prematurely. Importantly, the feasible set is updated after non-Skip turns if any candidate elimination occurs, cleanly separating protocol violations from evidence-based candidate elimination. Evaluation and verification. We include an independent evaluation module that logs the full interaction trace and scores both outcome and process. We check whether the model’s final guess matches the hidden target (non-verified accuracy). In addition, to rule out random-but-correct guesses, we compute verified accuracy using the feasible candidate set maintained by the verification module: an episode counts as verified-successful only if the feasible set is reduced to exactly one image (the target) before the model guesses. This design allows us to measure not only whether the model ultimately identifies the target, but also whether it does so through a coherent, evidence-driven trajectory that adheres to the protocol and effectively narrows the search space. 5 Evaluation Metrics We evaluate models along four complementary axes. Unless noted otherwise, all constraint application and candidate-set updates are computed by the verifica- tion module using only valid (non-Skip) question-answer pairs. •Identification accuracy:Overall (Non-Verified), Verified, Random-Guess. Overall (Non-verified) accuracy measures whether the model’s final guess matches the hidden target. Verified accuracy excludes random-but-correct guesses: an episode is counted as verified-successful only if (i) the final guess is correct and (i) the verification module’s feasible candidate set has size exactly one and contains the hidden target immediately before the guess. This metric captures whether the model’s trajectory effectively narrows down to the correct answer rather than succeeding by chance. Random-guess accuracy counts episodes where the final guess is correct but the feasible candidate set has size greater than one, indicating that the model guessed correctly without sufficient evidence. • Interaction efficiency. We measure interaction cost as the number of total turns (including Skip) before the final guess. We report efficiency on verified-successful episodes, random-guess correct episodes, incorrect episodes, and all episodes, so fewer turns correspond to faster evidence- driven narrowing rather than premature guessing. •Protocol compliance. We quantify instruction following via (i) the average Skip rate responses and (i) the average question generation rate before the End of uploading signal (premature outputs). Lower average Skip 13 rates and zero premature outputs indicate stronger adherence; repeated violations reflect failures to recover to valid questioning. •Robustness to noisy feedback. We test robustness under imperfect oracle answers by injecting controlled noise (e.g., flipping one Yes↔No response or perturbing an Unsure response). We then report non-verified and verified accuracy under noise. The verified metric highlights whether the model can recover an evidence-consistent trajectory (e.g., by re-checking critical attributes) rather than succeeding via chance. 6 Experimental Results We evaluate several open-source VLMs on Guess My Preferred Dress across multiple difficulty settings, including Qwen3-VL-235B-Instruct-FP8 [28], Qwen3.5-397B-A17B-FP8 [2], and Step3-VL-10B [5]. We report average verified and overall (non-verified) accuracy, interaction efficiency (average number of turns/rounds before the final guess and after the “End of uploading” signal), and protocol compliance (average Skip rate and premature outputs rate). Evaluation setup. In our framework, we employ Qwen3-VL-235B- Instruct-FP8 [28] as a few-shot violation detector. For the oracle answering module, we use an agentic Qwen-Agent pipeline with an image zoom-in tool to support inspection of fine-grained details; the VLM backbone is Qwen3-VL-235B-Instruct-FP8 [28]. Figure 7: Average verified, random-guess correct, and overall (non-verified) accuracy across similarity thresholds for Qwen3-VL-235B-Instruct-FP8, Qwen3.5-397B-A17B-FP8, and Step3-VL-10B. Accuracy. Figure 7 compares three accuracy metrics across similarity thresh- olds for the three evaluated models. Qwen3.5-397B-A17B-FP8 attains the highest verified accuracy at thresholds 0.3 through 0.6 with peak at threshold 0.4, whereas Step3-VL-10B performs best at threshold 0.8. Remarkably, despite being the smallest model with only 10B parameters, Step3-VL-10B attains the highest verified accuracy, random-guess accuracy and overall (Non-verified) accu- racy at the highest similarity threshold, where the average gallery size is smaller but the candidates are most visually similar and therefore hardest to distinguish. 14 Figure 8: Average verified accuracy with 95% confidence intervals across similar- ity thresholds for Qwen3-VL-235B-Instruct-FP8, Qwen3.5-397B-A17B- FP8, and Step3-VL-10B. By contrast, Qwen3-VL-235B-Instruct-FP8 generally attains the highest random-guess accuracy at thresholds 0.3 through 0.6, but this advantage does not translate into the highest verified accuracy. This gap indicates that higher final answer accuracy alone does not necessarily reflect better grounded or verifiable reasoning. For overall (non-verified) accuracy, Qwen3-VL-235B-Instruct- FP8 leads at lower thresholds, likely benefiting in part from its relatively high random-guess accuracy, while Qwen3.5-397B-A17B-FP8 peaks at 0.6 and Step3-VL-10B performs best at 0.8. Overall, these patterns suggest that Step3-VL-10B is more affected by gallery size, whereas Qwen3.5-397B-A17B-FP8 is more sensitive to visually confusable candidates. Additionally, the consistently higher random-guess accu- racy of Qwen3-VL-235B-Instruct-FP8 may indicate a greater tendency to guess early, while Qwen3.5-397B-A17B-FP8 appears more conservative to ask more questions rather than random-guess. Figure 8 shows the verified accuracy with 95% confidence intervals across similarity thresholds for the three models. Notably, Qwen3.5-397B-A17B- FP8 is strongest at lower and intermediate thresholds, whereas Step3-VL-10B becomes strongest at the highest threshold. Efficiency Figure 9 illustrates interaction length across similarity thresholds for the three evaluated models in four outcome categories: verified correct, random-guess correct, incorrect, and all episodes. In the verified-correct category, Step3-VL-10B, despite being the smallest model, generally shows the longest interaction length across thresholds. In contrast, Qwen3-VL-235B-Instruct- FP8 consistently has the shortest interaction length, suggesting that it may be more prone to guessing early without sufficient evidence, especially when the gallery contains more than 20 candidates (at thresholds, 0.3 and 0.4). Qwen3.5- 397B-A17B-FP8, by comparison, maintains a relatively constant interaction 15 Figure 9: Average number of dialogue rounds before guessing across similarity thresholds for four outcome categories (verified correct, random-guess correct, incorrect, all) for three different models. 16 length across thresholds, consistent with a more conservative strategy of asking additional questions regardless of gallery size or task difficulty. In the incorrect category, Qwen3.5-397B-A17B-FP8 generally exhibits the longest interaction length, suggesting that it continues querying even when discriminative cues are difficult to identify. By contrast, Qwen3-VL-235B- Instruct-FP8 and Step3-VL-10B appear more likely to stop earlier when they struggle, particularly at lower similarity thresholds where gallery sizes are larger. Figure 10: AverageSkiprates across similarity thresholds for three different models. Protocol Compliance Figure 10 shows the averageSkiprates across simi- larity thresholds for the three evaluated models. In general, Qwen3.5-397B- A17B-FP8 exhibits the highestSkiprates across thresholds in all four outcome categories. This suggests that, while it attains stronger verified accuracy than Qwen3-VL-235B-Instruct-FP8, it also struggles more with protocol adher- ence, possibly because its more persistent questioning strategy results in more frequent rule violations. Figure 11 shows the premature output rates across similarity thresholds for the three models. Unlike skip rates, premature output rates decrease steadily as the similarity threshold rises and the average gallery size falls. Step3-VL-10B remains the worst across thresholds, while both Qwen3-VL-235B-Instruct- FP8 and Qwen3.5-397B-A17B-FP8 approach zero at the highest threshold. 17 Figure 11: Premature output rates across similarity thresholds for three different models. These results suggest that premature questioning is driven more by larger gallery size than by visual confusability, and that Step3-VL-10B is more sensitive to gallery size than the other two models. 7 Failure Cases 7.1 Qwen3-VL-235B-Instruct-FP8 failure cases Figure 12 shows a failure case for Qwen3-VL-235B-Instruct-FP8 atτ= 0.8 with 6 candidate images. Over 9 rounds, the model asks about tiered skirts, fabric sheen, wrap fronts, hem shape, and ruffled hems, but also violates constraints by asking about sleeve length, garment length, and color; these invalid turns are returned asSkip. Despite some narrowing from the valid turns, the final guess is incorrect. This episode illustrates two common failure modes: (i) poor protocol compliance that wastes interaction budget and (i) insufficiently discriminative questions among highly similar candidates. Figure 13 shows a second failure case for Qwen3-VL-235B-Instruct-FP8 atτ= 0.3 with 35 candidates across seven upload batches. The model begins gen- erating questions during the upload phase, including echoingEnd of uploading and asking questions before the actual signal, violating the upload protocol. After the upload completes, it asks only one valid question and immediately guesses, effectively a random guess. This illustrates how large galleries can amplify (i) premature, protocol-violating behavior and (i) early stopping with insufficient evidence. 7.2 Qwen3.5-397B-A17B-FP8 failure cases Figure 14 shows a failure case for Qwen3.5-397B-A17B-FP8 atτ= 0.6 with 7 candidate images. The model uses all 20 allowed turns without ever committing 18 Prompt Task: I will upload several batches of dress options. Please guess which one is my favorite dress and provide its index in the image gallery (starting index: 1). Constraints: 1. Ask questions to gather attribute information about the target dress. 2. Ask exactly one Yes/No question per turn about one distinguishable feature. I will respond with Yes, No, Unsure (if not confident), or Skip (if you violate the rules). 3. If you receive Skip, your latest question was invalid and provides no information. 4.Do not ask about sleeve length, garment length, color, pattern/prints, age group, size, shoes, necklace, hat, bag, background, or the human model. 5.Do not enumerate attribute values across turns for the same attribute type (e.g., square neck, V-neck, crew neck). 6.Do not inspect images one-by-one or reference indices (e.g., “is your favorite dress the first image?”). 7. Do not guess until you have narrowed the target down to one image. 8. If your confidence is low or you find contradictions, keep asking verification questions. 9. When ready, output: My guess of your favorite dress: #<number>. 10. Only start generating questions after you receive End of uploading. Turn Speaker Content 1ModelDoes the dress have a tiered skirt? UserYes 2ModelIs the dress made of a shiny or satin-like fabric? UserYes 3ModelDoes the dress have long sleeves? UserSkip 4ModelIs the dress sleeveless? UserSkip 5ModelDoes the dress have a wrap-style front? UserYes 6ModelIs the dress floor-length? UserSkip 7ModelDoes the dress have a high-low hemline? UserNo 8ModelIs the dress in a solid color? UserSkip 9ModelDoes the dress have a ruffled hem? UserYes ModelMy guess: dress #2(✗ Incorrect) Figure 12: Failure case of Qwen3-VL-235B-Instruct-FP8 on Guess My Preferred Dress. Out of 9 turns, 4 are Skip (protocol violations), and the final guess is incorrect. 19 Prompt Task: I will upload several batches of dress options. Please guess which one is my favorite dress and provide its index in the image gallery (starting index: 1). Constraints: 1. Ask questions to gather attribute information about the target dress. 2. Ask exactly one Yes/No question per turn about one distinguishable feature. I will respond with Yes, No, Unsure (if not confident), or Skip (if you violate the rules). 3. If you receive Skip, your latest question was invalid and provides no information. 4. Do not ask about sleeve length, garment length, color, pattern/prints, age group, size, shoes, necklace, hat, bag, background, or the human model. 5. Do not enumerate attribute values across turns for the same attribute type (e.g., square neck, V-neck, crew neck). 6.Do not inspect images one-by-one or reference indices (e.g., “is your favorite dress the first image?”). 7. Do not guess until you have narrowed the target down to one image. 8. If your confidence is low or you find contradictions, keep asking verification questions. 9. When ready, output: My guess of your favorite dress: #<number>. 10. Only start generating questions after you receive End of uploading. Turn Speaker Content UserPrompt + <imagebatch1> ModelEnd of uploading UserHere is the next batch of dress options.<imagebatch2> ModelEnd of uploading UserHere is the next batch of dress options.<imagebatch3> ModelDoes the dress have a ruffled hem? UserHere is the next batch of dress options.<imagebatch4> ModelDoes the dress have a lace overlay? UserHere is the next batch of dress options.<imagebatch5> ModelDoes the dress have a high neckline? UserHere is the next batch of dress options.<imagebatch6> ModelDoes the dress have a V-neckline? UserEnd of uploading <image batch7> 1ModelDoes the dress have a front tie detail? UserYes ModelMy guess: dress #1(✗ Incorrect) Figure 13: Failure case atτ= 0.3 with 35 candidates. The model violates the upload protocol by generating outputs before the trueEnd of uploadingsignal, then guesses after only one answered question. 20 to a final guess. Ten of the twenty turns result inSkip: Questions 2–4 enumerate neckline types (off-the-shoulder, square, sweetheart) after already asking about V-neckline in Question 1, violating Rule 5; Question 9 asks about sleeve length (Rule 4); Questions 11–12 ask about print types (Rule 4); Question 16 re- enumerates skirt style after asking about tiered skirts in Question 6 (Rule 5); and Questions 18–20 ask about garment length and color (Rule 4). The remaining valid turns establish that the target has a smocked bodice, a side slit, and a fitted waist, but the model fails to synthesize these constraints to narrow the candidate set to a unique target, exhausting the budget without a guess. This episode illustrates (i) repeated violation of the same rule types throughout the interaction and (i) insufficient constraint integration despite accumulating valid discriminative evidence. Figure 15 shows a failure case for Qwen3.5-397B-A17B-FP8 atτ= 0.5 with 14 candidate images. The model again exhausts all 20 allowed turns without committing to a final guess at lower similarity threshold. Twelve of the 20 turns result inSkip: Turn 3 asks about floral print (Rule 4); Turns 9–10 enumerate strapless and sweetheart necklines after V-neckline was confirmed in Turn 4 (Rule 5); and Turns 12–20 cycle through nine compound combinations of already- confirmed attributes—wrap-style, V-neckline, slit, ruched detail, and high-low hemline—re-enumerating established facts rather than posing new discriminative questions (Rule 5). The valid turns collectively establish five features (wrap- style, V-neckline, slit, ruched detail, high-low hemline), yet the model fails to synthesize this evidence to narrow the 14-candidate pool to a unique target. This episode illustrates (i) repeated Rule 4 and Rule 5 violations identical in kind to those in Figure 14, and (i) a distinctive re-enumeration loop in which the model exhausts its remaining budget cycling through compound combinations of already-confirmed attribute pairs and triples rather than probing genuinely new features. 8Reinforcement Learning Trajectories Collec- tion AMIGO can also serve as a trajectory collection framework for multimodal reinforcement learning. Each step provides a well-defined state (gallery, dialogue context, accumulated constraints), an action (a constrained question), and an observation (Yes/No/Unsure/Skip feedback, optionally noisy), with a naturally long-horizon objective. This structure supports dense rewards without additional annotation, e.g., penalties for invalid questions (Skip), progress rewards based on reduction in feasible set size or estimated information gain, and terminal rewards for correct identification only when the target is uniquely determined. These trajectories can support offline RL, preference optimization over alternative questions, and robustness training under imperfect feedback. 21 Prompt identical to Figure 12. Turn Speaker Content 1 ModelDoes your favorite dress feature a V-neckline? UserNo 2ModelDoes your favorite dress have an off-the-shoulder neckline? UserSkipRule 5: enumerating neckline types 3ModelDoes your favorite dress have a square neckline? UserSkipRule 5: enumerating neckline types 4ModelDoes your favorite dress have a sweetheart neckline? UserSkipRule 5: enumerating neckline types 5 ModelDoes your favorite dress have a wrap-style bodice? UserNo 6 ModelDoes your favorite dress have a tiered skirt? UserNo 7 ModelDoes your favorite dress have a smocked bodice? UserYes 8 ModelDoes your favorite dress have flutter sleeves? UserNo 9ModelDoes your favorite dress have long sleeves? UserSkipRule 4: sleeve length 10 ModelDoes your favorite dress have a side slit? UserYes 11ModelDoes your favorite dress have a floral print? UserSkipRule 4: pattern/prints 12ModelDoes your favorite dress have a leaf print? UserSkipRule 4: pattern/prints 13 ModelDoes your favorite dress have a high-low hemline? UserNo 14 ModelDoes your favorite dress have a straight hemline? UserNo 15 ModelDoes your favorite dress have a ruffled hemline? UserNo 16ModelDoes your favorite dress have a flared skirt? UserSkipRule 5: enumerating skirt types 17 ModelDoes your favorite dress have a fitted waist? UserYes 18ModelDoes your favorite dress have a midi length? UserSkipRule 4: garment length 19ModelDoes your favorite dress have a maxi length? UserSkipRule 4: garment length 20ModelDoes your favorite dress have a solid color? UserSkipRule 4: color ModelNo guess made — 20-question budget exhausted (✗ Incorrect) Figure 14: Failure case of Qwen3.5-397B-A17B-FP8 on Guess My Preferred Dress atτ= 0.6 with 7 candidates. Out of 20 turns, 10 areSkip(protocol violations highlighted in yellow with rule annotations), and the model exhausts its entire question budget without making a final guess. Repeated Rule 4 and Rule 5 violations waste turns, while valid evidence about smocked bodice, side slit, and fitted waist is not synthesized into a unique identification. 22 Prompt identical to Figure 12. Turn Speaker Content 1 ModelIs your favorite dress made of velvet fabric? UserNo 2 ModelIs your favorite dress a wrap-style dress? UserYes 3ModelDoes your favorite dress have a floral print? UserSkipRule 4: pattern/prints 4 ModelDoes your favorite dress have a V-neckline? UserYes 5 ModelDoes your favorite dress have a tiered skirt? UserNo 6 ModelDoes your favorite dress have a slit? UserYes 7 ModelDoes your favorite dress have a belt or tie at the waist? UserNo 8 ModelDoes your favorite dress have a ruched detail? UserYes 9ModelDoes your favorite dress have a strapless neckline? UserSkipRule 5: enumerating neckline types 10ModelDoes your favorite dress have a sweetheart neckline? UserSkipRule 5: enumerating neckline types 11 ModelDoes your favorite dress have a high-low hemline? UserYes 12ModelDoes your favorite dress have a wrap-style bodice with a V-neckline, a ruched detail, and a high-low hemline? UserSkipRule 5: re-enumerating confirmed attributes 13ModelDoes your favorite dress have a wrap-style bodice with a V-neckline and a high-low hemline? UserSkipRule 5: re-enumerating confirmed attributes 14ModelDoes your favorite dress have a V-neckline and a high-low hemline? UserSkipRule 5: re-enumerating confirmed attributes 15ModelDoes your favorite dress have a V-neckline and a slit? UserSkipRule 5: re-enumerating confirmed attributes 16ModelDoes your favorite dress have a V-neckline and ruched detail- ing? UserSkipRule 5: re-enumerating confirmed attributes 17ModelDoes your favorite dress have a wrap-style bodice and a high- low hemline? UserSkipRule 5: re-enumerating confirmed attributes 18ModelDoes your favorite dress have a wrap-style bodice and a slit? UserSkipRule 5: re-enumerating confirmed attributes 19ModelDoes your favorite dress have a V-neckline and a wrap-style bodice? UserSkipRule 5: re-enumerating confirmed attributes 20ModelDoes your favorite dress have a V-neckline and a ruched detail? UserSkipRule 5: re-enumerating confirmed attributes ModelNo guess made — 20-question budget exhausted (✗ Incorrect) Figure 15: Failure case of Qwen3.5-397B-A17B-FP8 on Guess My Preferred Dress atτ= 0.5 with 14 candidates. Out of 20 turns, 12 areSkip(protocol violations highlighted in yellow with rule annotations), and the model exhausts its entire question budget without making a final guess. After valid turns establish wrap-style (Turn 2), V-neckline (Turn 4), slit (Turn 6), ruched detail (Turn 8), and high-low hemline (Turn 11), the model enters a nine-turn re-enumeration loop cycling through compound combinations of these already-confirmed attributes (Turns 12–20) rather than probing new discriminative features. 23 References [1]Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, et al. Magma: A foundation model for multimodal ai agents. In Proceedings of the computer vision and pattern recognition conference, pages 14203–14214, 2025. [2] Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. [3] V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, Jing Chen, Jinhao Chen, Jinghao Lin, Jinjiang Wang, Junjie Chen, Leqi Lei, Letian Gong, Leyi Pan, Mingdao Liu, Mingde Xu, Mingzhi Zhang, Qinkai Zheng, Sheng Yang, Shi Zhong, Shiyu Huang, Shuyuan Zhao, Siyan Xue, Shangqin Tu, Shengbiao Meng, Tianshu Zhang, Tianwei Luo, Tianxiang Hao, Tianyu Tong, Wenkai Li, Wei Jia, Xiao Liu, Xiaohan Zhang, Xin Lyu, Xinyue Fan, Xuancheng Huang, Yanling Wang, Yadong Xue, Yanfeng Wang, Yanzi Wang, Yifan An, Yifan Du, Yiming Shi, Yiheng Huang, Yilin Niu, Yuan Wang, Yuanchang Yue, Yuchen Li, Yutao Zhang, Yuting Wang, Yu Wang, Yuxuan Zhang, Zhao Xue, Zhenyu Hou, Zhengxiao Du, Zihan Wang, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Minlie Huang, Yuxiao Dong, and Jie Tang. Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025. [4]Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, SH Cai, Yuan Cao, Y Charles, HS Che, Cheng Chen, Guanduo Chen, et al. Kimi k2. 5: Visual agentic intelligence. arXiv preprint arXiv:2602.02276, 2026. [5] Ailin Huang, Chengyuan Yao, Chunrui Han, Fanqi Wan, Hangyu Guo, Haoran Lv, Hongyu Zhou, Jia Wang, Jian Zhou, Jianjian Sun, Jingcheng Hu, Kangheng Lin, Liang Zhao, Mitt Huang, Song Yuan, Wenwen Qu, Xiangfeng Wang, Yanlin Lai, Yingxiu Zhao, Yinmin Zhang, Yukang Shi, Yuyang Chen, Zejia Weng, Ziyang Meng, Ang Li, Aobo Kong, Bo Dong, Changyi Wan, David Wang, Di Qi, Dingming Li, En Yu, Guopeng Li, Haiquan Yin, Han Zhou, Hanshan Zhang, Haolong Yan, Hebin Zhou, Hongbo Peng, Jiaran Zhang, Jiashu Lv, Jiayi Fu, Jie Cheng, Jie Zhou, Jisheng Yin, Jingjing Xie, Jingwei Wu, Jun Zhang, Junfeng Liu, Kaijun Tan, Kaiwen Yan, Liangyu Chen, Lina Chen, Mingliang Li, Qian Zhao, Quan Sun, Shaoliang Pang, Shengjie Fan, Shijie Shang, Siyuan Zhang, Tianhao You, Wei Ji, Wuxun Xie, Xiaobo Yang, Xiaojie Hou, Xiaoran Jiao, Xiaoxiao Ren, Xiangwen Kong, Xin Huang, Xin Wu, Xing Chen, Xinran Wang, Xuelin Zhang, Yana Wei, Yang Li, Yanming Xu, Yeqing Shen, Yuang Peng, Yue Peng, Yu Zhou, Yusheng Li, Yuxiang Yang, Yuyang Zhang, Zhe Xie, Zhewei Huang, Zhenyi Lu, Zhimin Fan, Zihui Cheng, Daxin Jiang, Qi Han, Xiangyu Zhang, Yibo Zhu, and Zheng Ge. Step3-vl-10b technical report, 2026. 24 [6]Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the v in vqa matter: Elevating the role of image understand- ing in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6904–6913, 2017. [7]Minesh Mathew, Viraj Bagal, Rub`en Tito, Dimosthenis Karatzas, Ernest Valveny, and CV Jawahar. Infographicvqa. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 1697–1706, 2022. [8]Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, Ruoqi Liu, Ge Zhang, Samuel Stevens, Dongfu Jiang, Weiming Ren, Yuxuan Sun, et al. Mmmu: A massive multi-discipline multimodal understanding and reasoning bench- mark for expert agi. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9556–9567, 2024. [9]Yuan Liu, Haodong Duan, Yuanhan Zhang, Bo Li, Songyang Zhang, Wangbo Zhao, Yike Yuan, Jiaqi Wang, Conghui He, Ziwei Liu, et al. Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233. Springer, 2024. [10] Yi-Fan Zhang, Huanyu Zhang, Haochen Tian, Chaoyou Fu, Shuangqing Zhang, Junfei Wu, Feng Li, Kun Wang, Qingsong Wen, Zhang Zhang, et al. Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?arXiv preprint arXiv:2408.13257, 2024. [11]Fanqing Meng, Jin Wang, Chuanhao Li, Quanfeng Lu, Hao Tian, Jiaqi Liao, Xizhou Zhu, Jifeng Dai, Yu Qiao, Ping Luo, et al. Mmiu: Multimodal multi- image understanding for evaluating large vision-language models. arXiv preprint arXiv:2408.02718, 2024. [12]Fei Wang, Xingyu Fu, James Y Huang, Zekun Li, Qin Liu, Xiaogeng Liu, Mingyu Derek Ma, Nan Xu, Wenxuan Zhou, Kai Zhang, et al. Muirbench: A comprehensive benchmark for robust multi-image understanding. arXiv preprint arXiv:2406.09411, 2024. [13] Jiabo Ye, Haiyang Xu, Haowei Liu, Anwen Hu, Ming Yan, Qi Qian, Ji Zhang, Fei Huang, and Jingren Zhou. mplug-owl3: Towards long image-sequence understanding in multi-modal large language models.arXiv preprint arXiv:2408.04840, 2024. [14]Dongfu Jiang, Xuan He, Huaye Zeng, Cong Wei, Max Ku, Qian Liu, and Wenhu Chen. Mantis: Interleaved multi-image instruction tuning. arXiv preprint arXiv:2405.01483, 2024. [15] Ziyu Liu, Tao Chu, Yuhang Zang, Xilin Wei, Xiaoyi Dong, Pan Zhang, Zijian Liang, Yuanjun Xiong, Yu Qiao, Dahua Lin, et al. Mmdu: A multi- turn multi-image dialog understanding benchmark and instruction-tuning 25 dataset for lvlms. Advances in Neural Information Processing Systems, 37:8698–8733, 2024. [16]Dawei Yan, Yang Li, Qing-Guo Chen, Weihua Luo, Peng Wang, Haokui Zhang, and Chunhua Shen. Mmcr: Advancing visual language model in mul- timodal multi-turn contextual reasoning. arXiv preprint arXiv:2503.18533, 2025. [17] Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang, Yechan Hwang, Byungsoo Ko, Han-Gyu Kim, Dongyu Yao, Xuankun Rong, Eojin Joo, Seung-Ho Han, et al. Multiverse: A multi-turn conversation benchmark for evaluating large vision and language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 708–719, 2025. [18] Shuo Liu, Kaining Ying, Hao Zhang, Yue Yang, Yuqi Lin, Tianle Zhang, Chuanhao Li, Yu Qiao, Ping Luo, Wenqi Shao, et al. Convbench: A multi-turn conversation evaluation benchmark with hierarchical ablation capability for large vision-language models. Advances in Neural Information Processing Systems, 37:100734–100782, 2024. [19]Elliot L Epstein, Kaisheng Yao, Jing Li, Xinyi Bai, and Hamid Palangi. Mmmt-if: A challenging multimodal multi-turn instruction following bench- mark. arXiv preprint arXiv:2409.18216, 2024. [20] Yongqi Li, Wenjie Li, and Liqiang Nie. Mmcoqa: Conversational question answering over text, tables, and images. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4220–4231, 2022. [21] Yunzhuo Hao, Jiawei Gu, Huichen Will Wang, Linjie Li, Zhengyuan Yang, Lijuan Wang, and Yu Cheng. Can mllms reason in multimodal- ity? emma: An enhanced multimodal reasoning benchmark. arXiv preprint arXiv:2501.05444, 2025. [22]Jihyung Kil, Zheda Mai, Justin Lee, Arpita Chowdhury, Zihe Wang, Kerrie Cheng, Lemeng Wang, Ye Liu, and Wei-Lun Harry Chao. Mllm-compbench: A comparative reasoning benchmark for multimodal llms. Advances in Neural Information Processing Systems, 37:28798–28827, 2024. [23]Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, Kai Zhang, Sheng- bang Tong, Yuxuan Sun, Botao Yu, Ge Zhang, Huan Sun, et al. Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15134–15186, 2025. [24] Bryan LM de Oliveira, Luana GB Martins, Bruno Brand ̃ao, and Lucke- ciano C Melo. Infoquest: Evaluating multi-turn dialogue agents for open- ended conversations with hidden context. arXiv preprint arXiv:2502.12257, 2025. 26 [25]Kimia Ramezan, Alireza Amiri Bavandpour, Yifei Yuan, Clemencia Siro, and Mohammad Aliannejadi. Multi-turn multi-modal question clarification for enhanced conversational understanding. arXiv preprint arXiv:2502.11442, 2025. [26]Xiaohan Yu, Chao Feng, Lang Mei, and Chong Chen. M 3 searcher: Modular multimodal information seeking agency with retrieval-oriented reasoning. arXiv preprint arXiv:2601.09278, 2026. [27]Vardhan Dongre, Chi Gui, Shubham Garg, Hooshang Nayyeri, Gokhan Tur, Dilek Hakkani-T ̈ur, and Vikram S Adve. Mirage: A benchmark for multimodal information-seeking and reasoning in agricultural expert-guided conversations. arXiv preprint arXiv:2506.20100, 2025. [28] Qwen Team. Qwen3 technical report, 2025. [29] Lei Bai, Zhongrui Cai, Yuhang Cao, Maosong Cao, Weihan Cao, Chiyu Chen, Haojiong Chen, Kai Chen, Pengcheng Chen, Ying Chen, et al. Intern-s1: A scientific multimodal foundation model. arXiv preprint arXiv:2508.15763, 2025. [30]Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, et al. Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning. arXiv preprint arXiv:2507.01006, 2025. 27