Paper deep dive
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
Nan Li, Albert Gatt, Massimo Poesio
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 99%
Last extracted: 7/5/2026, 7:47:41 AM
Summary
The paper investigates whether Vision-Language Models (VLMs) can accurately distinguish between potential and established common ground in asymmetric, collaborative dialogues. Using the HCRC MapTask dataset, the researchers found that providing authentic map images (or textual descriptions of map content) causes models, particularly Qwen3-VL-8B-Instruct, to over-predict alignment (the 'over-alignment bias'). This bias is driven by task-relevant map content rather than the visual channel itself. The study concludes that VLMs tend to rely on static referential cues from maps rather than tracking the incremental grounding process through dialogue history, leading to high confidence in incorrect 'non-aligned' predictions.
Entities (6)
Relation Signals (4)
Gemma3 â evaluatedwith â HCRC MapTask
confidence 100% · We evaluate models from two open-source VLM families, Qwen3-VL and Gemma3... on 13,077 annotated reference expressions from HCRC MapTask dialogues
Qwen3-VL-8B-Instruct â exhibits â Over-alignment Bias
confidence 100% · We observe these patterns most clearly in Qwen3-VL-8B-Instruct... In models that exhibit the bias, map content... is treated as evidence of mutual understanding
HCRC MapTask â usedfor â Interpretation Matching
confidence 100% · We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues
Map Content â drives â Over-alignment Bias
confidence 95% · indicating that the bias is driven by task-relevant map content, not the visual channel.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through interaction. We investigate whether vision-language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through grounding. We formulate this as an interpretation-matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically controlled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress alignment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to varying degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with established common ground.
Tags
Links
- Source: https://arxiv.org/abs/2606.31719v1
- Canonical: https://arxiv.org/abs/2606.31719v1
Trouble viewing inline? Open PDF directly â
Full Text
57,981 characters extracted from source content.
Expand or collapse full text
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue Nan Li, Albert Gatt, Massimo Poesio Utrecht University, Utrecht, The Netherlands n.li, a.gatt, m.poesio@u.nl Abstract In collaborative dialogue, shared perception does not guarantee shared interpretation. Mu- tual understanding must be established through interaction. We investigate whether vision- language models (VLMs) can distinguish what could be shared from what has been shared between dialogue participants through ground- ing. We formulate this as an interpretation- matching task on 13,077 annotated reference expressions from HCRC MapTask dialogues, and evaluate VLMs under systematically con- trolled manipulations of dialogue context and map-information access. Our results show that providing authentic map images improves overall performance but shifts models toward over-predicting alignment. Textual descriptions of the same map content reproduce this bias, while non-informative images suppress align- ment predictions entirely, indicating that the bias is driven by task-relevant map content, not the visual channel. This improvement comes at the cost of degraded accuracy on non-aligned cases. Calibration analysis and reference-chain tracking further suggest that models rely on static referential cues on the maps rather than tracking how grounding unfolds through dia- logue history. We observe these patterns most clearly in Qwen3-VL-8B-Instruct and, to vary- ing degrees, in four additional models from two architecture families. In models that exhibit the bias, map content, whether presented visually or textually, is treated as evidence of mutual understanding, conflating potential with estab- lished common ground. 1 Introduction In everyday collaborative situated dialogue, two people attending to the situation may extract dif- ferent information from it or interpret the same information in different ways. Grounding, the in- cremental process by which dialogue participants establish mutual understanding, is what bridges this gap (Clark and Wilkes-Gibbs, 1986; Clark and Schaefer, 1989; Clark and Brennan, 1991). A speakerâs reference expression (RE) becomes part of the common ground only after the addressee recognises, accommodates, and confirms it; until then, the RE remains potential rather than estab- lished shared knowledge. While information asymmetry is present to some degree in all dialogue, in certain collaborative tasks it is introduced by design, making asymmetry a structural and controlled feature of the task. In what follows, we use asymmetric dialogue to re- fer specifically to settings where participants hold different private task-relevant information by de- sign, and where neither can directly access the otherâs. The HCRC MapTask (Anderson et al., 1991) is a canonical instance: two participants navigate a route using maps that differ in land- mark placement, landmark names, or the number of identically-named landmarks, so that apparent agreement can mask genuine divergence in inter- pretation. Because the discrepancies are designed into the maps, landmarks visible on both maps cre- ate potential referential overlap whose resolution depends entirely on the grounding process in the di- alogue. A model evaluating such dialogue from the outside, as an overhearer with access to one or both maps, faces the same asymmetry: it can observe what could be shared, but must infer from the dia- logue what has been shared. Recent perspectivist annotation work (Li et al., 2026) has made it possi- ble to evaluate, for each RE, whether two MapTask participants actually share the same grounded in- terpretation at a given point in the dialogue â a judgment we call interpretation matching. We use this resource to investigate three research questions: RQ1Can large vision-language models (VLMs) capture personal interpretations of interlocu- tors toward the same reference expression in 1 arXiv:2606.31719v1 [cs.CL] 30 Jun 2026 1 · MapTask Dialogue Example GGo to «the parked van». FAlright, to the bottom. GNo! 2 · Interpretations Perspectivist annotations (Li et al., 2026) G G Giver's interpretation upper parked van landmark ID: m0_parked_van#1@g â F Follower's interpretation lower parked van landmark ID: m0_parked_van#0@f 3 · Our task Interpretation matching Do the two participants interpret the target RE the same way? Yes âsame landmark Noâdifferent landmarks *Gold Label: No Model role: overhearer (third-party analyzer) â observes but cannot ask questions, or repair. Experiment Conditions: dialogue context window and map information modality. The landmark ID encodes map (m0), side (@g / @f), and ordinal index (#0 = lower, #1 = upper). Figure 1: The interpretation matching task, illustrated on a misalignment example. Panel 1 shows a simplified example of a MapTask dialogue with the target RE the parked van: the giverâs map contains two parked vans while the followerâs map has only one. Panel 2 shows the perspectivist annotations (Li et al., 2026): the giver grounds the parked van to the upper van (m0_parked_van#1@g, see Appendix A for the definition about the landmark ID) and the follower to the lower one (m0_parked_van#0@f). Panel 3 formulates our task: given the map(s) and dialogue, we ask the model to decide whether the two participantsâ interpretations of the marked RE match. The gold label here is NO. asymmetric dialogue? RQ2 Which modality of information contributes more to VLMsâ assessment of interpretation alignment? RQ3Do VLMs exhibit systematically different be- haviours on different types of alignment and misalignment cases? We frame the task as a binary judgment â do the two participantsâ interpretations of a marked RE match? â and evaluate VLMs under systematic manipulations of two independent variables: the amount of available dialogue context and the type of map information provided. We evaluate models from two open-source VLM families, Qwen3-VL and Gemma3, at scales from 2B to 12B parame- ters. Qwen3-VL-8B-Instruct, the best-performing model in preliminary evaluation, serves as the pri- mary model for detailed condition-grid analysis; the remaining models are compared on a baseline grid to assess cross-model generality (§3.3, §5.4). Our main findings are: âąProviding authentic map images shifts the model toward over-predicting alignment. The model treats landmark co-presence as evi- dence of mutual understanding, conflating what could be shared with what has been shared. Textual descriptions of the same map content reproduce this bias, showing that it is driven by task-relevant map content rather than by the visual channel. âąNon-informative visual inputs (blank maps, shuffled landmarks) do not reproduce the bias; they make the model more conservative, not more prone to predicting alignment. The bias therefore possibly requires map content, whether delivered visually or textually. âą Calibration and reference-chain analysis con- verge on the same explanation: the model re- lies more on static referential cues on the maps rather than tracking how grounding unfolds through dialogue history. Our work contributes (1) an evaluation method- ology for probing VLMsâ ability to distinguish potential from established common ground in information-asymmetric dialogue; and (2) an em- pirical characterisation of a model-dependent fail- ure mode, observed most clearly in Qwen3-VL-8B, in which VLMs conflate possible referential over- lap with communicative alignment. 2 Related Work Grounding and Overhearers Common ground is built incrementally through interaction: inter- locutors negotiate, confirm, and repair meaning, rather than receiving it directly from shared per- ceptual access (Clark and Wilkes-Gibbs, 1986; 2 Clark and Brennan, 1991). A well-established consequence is the overhearer illusion: overhear- ers can hear every word but, lacking the ability to contribute grounding acts, reach systematically weaker interpretations than addressees (Schober and Clark, 1989). This asymmetric performance carries over to language models: dialogue systems trained and evaluated on static transcripts are struc- turally overhearers, which shapes what they can learn about grounding and clarification (Madureira and Schlangen, 2024). Our evaluation makes that stance explicit, placing VLMs as overhearers of asymmetric human dialogue and testing whether they mistake co-presence of potential referents on the maps, delivered either visually or as text, for established shared interpretation. Reference in Dialogue and VLM Evaluation Reference corpora have long served as testbeds for how interlocutors build shared interpretations through repeated mention, partial information, or clarification (Anderson et al., 1991; Haber et al., 2019; Udagawa and Aizawa, 2019; Chiyah-Garcia et al., 2023), yet each RE is typically annotated with a single gold referent, implicitly assuming speaker-addressee convergence. Under informa- tion asymmetry this breaks: the perspectivist an- notation of Li et al. (2026) shows that the two par- ticipants can hold distinct grounded interpretations even under apparent agreement, the judgement of which overhearers make hard. Recent VLM evalu- ations report systematic gaps with humans as well: VLM overhearers underperform human matchers on referential dialogue and do not improve with repeated discussion (Wang et al., 2025), and VLM participants fail to entrain, form conceptual pacts, or initiate grounding acts as humans (Zeng et al., 2026; Shaikh et al., 2025). We directly ask whether a VLM, given the same asymmetric evidence as the participants, can recognise that they have not yet reached a shared interpretation, evaluating VLMs on perspectivist MapTask annotations under a vari- ety of controlled map-information conditions. 3 Experimental Setup 3.1 Dataset We use a published corpus of perspectivist annota- tions of the HCRC MapTask corpus (Li et al., 2026) that separately records, for each RE, the speakerâs intended landmark and the addresseeâs interpreted landmark. The dataset comprises 13,077 annotated REs from 128 HCRC MapTask dialogues. Each dialogue involves two participants (a giver and a follower) collaborating to reproduce a route on the followerâs map under the giverâs guidance, with slightly different maps. The discrepancies can in- clude landmark name differences, missing land- marks, and differences in quantity, creating a rich environment for misalignment in grounding. See Appendix A for more dataset details. 3.2 Task Design We want to evaluate VLMsâ interpretation match- ing ability, and formulate this as a binary judgment task: Given a MapTask dialogue excerpt containing one marked RE, the model must decide whether the two participants currently share the same grounded interpretation of that expression or not. Figure 1 illustrates a misalignment case, where the giver and follower ground the parked van to different landmarks. An instance is labelled YES (aligned) when the two participantsâ landmark IDs match; and NO (not aligned) otherwise, covering both pending states (not yet grounded) and misunderstandings (grounded to different landmarks). The ground truth class distribution of the corpus is imbalanced: 72.1% aligned (YES) and 27.9% not aligned (NO). We manipulate two independent variables to in- vestigate how information access shapes model judgments: text access (the dialogue context win- dow) and map access (the map-information modal- ity). Text access (dialogue context window)We vary how much dialogue context the model receives via four windows of increasing size: curL (current transaction, up to and including the line contain- ing the target RE), curT (current transaction in full), startL (dialogue from beginning through the line containing the target RE), and startT (dialogue from beginning through the end of the current trans- action). A transaction is a human-annotated Map- Task dialogue excerpt that typically corresponds to a sequence of movements from one landmark to another. These windows allow us to test whether (1) broader dialogue history, which encodes prior grounding episodes, and (2) future interactions, which usually encode repair sequences and clarifi- cation exchanges, help the model track grounding. Map access (map-information modality) We vary what map information the model receives. In the baseline condition grid, conditions include: 3 Text-only (no map information), Both maps (au- thentic giver and follower map images), and Giver- only / Follower-only (a single authentic map im- age). To further investigate the impact of map infor- mation and disentangle content from input chan- nel, the full-grid modality experiment adds four conditions. Two are textual: Text-landmark- names (a textual list of landmark names on each map) and Text-discrepancy-detail (textual descrip- tions about the map discrepancies). Two are non- informative visual controls: Blank maps (uniform empty images, 1024Ă1024, RGB 128/128/128) and Shuffled maps (map images with landmarks from an unrelated map pair). Appendix D shows concrete example fillings of$map_accessfor the two textual conditions. The blank-map and shuffled-map conditions serve as controls: if the over-alignment bias is a generic multimodal artifact, these conditions should reproduce it; if it is content-driven, they should not. Prompt designWe apply zero-shot learning. The model receives a system prompt framing it as a dialogue analysis expert overhearing a MapTask dialogue. The prompt further describes map asym- metry (landmarks may be missing, duplicated, or placed differently), and defines the interpretation- matching task. The user prompt provides the tar- get RE, the dialogue context with the target RE wrapped in«...»markers, and (where applicable) map images. Full prompt templates are given in Appendix C. 3.3 Models and Inference We select models from two open-source VLM families that represented the state of the art at the time of our experiments: Qwen3-VL (Bai et al., 2025) and Gemma3 (Gemma Team, 2025). Within each family, we evaluate models ranging from 2B to 12B parameters (Qwen3-VL-2B/4B/8B- Instruct; Gemma3-4B/12B-it), constrained by the memory of a single A100 GPU. All models are instruction-tuned, non-thinking versions. In prelim- inary evaluation across all five models, Qwen3-VL- 8B-Instruct achieved the highest macro-F1, so we use it as the primary model for the full condition- grid analysis. The remaining four models are eval- uated on the baseline condition grid to assess gen- erality (§5.4). We use vLLM (Kwon et al., 2023) to deploy the models and to accelerate the inference progress. We use greedy decoding (temperature 0) with con- strained output via vLLM logit masking to valid YES/NO tokens, ensuring valid binary responses. We setrandom_seedto 42 in vLLMâs sampling parameters to ensure reproducibility. Map images are resized to a maximum side length of 1024 pixels and delivered as PNG. We selected 1024 pixels as the minimum resolution at which Qwen3-VL-8B-Instruct reliably identified landmark names, icons, and locations in prelimi- nary checks. 3.4 Evaluation We report accuracy, macro-averaged F1, per-class recall (recall pos for aligned, recall neg for not- aligned), and the modelâs yes-rate (proportion of YES predictions) as a measure of response bias. In addition, we conduct three further analyses in §5: calibration analysis using token-level log- its (§5.1), a status-level breakdown that decom- poses performance by grounding state (§5.2), and reference-chain tracking that examines how predic- tions evolve across repeated mentions of the same landmark (§5.3). 4 Results Table 1 presents the main results across the baseline conditions. All results in this section use Qwen3- VL-8B-Instruct; cross-model generality is assessed in §5.4. Here are the key findings: F1: Map access improves detection but shifts the model toward over-predicting alignment. Comparing within the same model (Qwen3-VL- 8B), at startT, adding both maps improves F1 macro from .591 (text-only) to .671 (both maps), a gain of .080. However, this improvement is driven by a dramatic shift in recall profile: recall pos rises from .590 to .822, while recall neg drops from .677 to .518. The yes-rate shifts from .515 to .727, exceed- ing the gold base rate of .721. The evidence for over-prediction is not the yes-rate alone (which is close to the base rate) but the direction of the re- call trade-off : map access systematically pushes recall neg down while inflating recall pos , meaning the model sacrifices its ability to detect non-aligned cases in favour of aligned ones. Calibration analy- sis in §5.1 and the status-level breakdown in §5.2 further confirm this: map conditions produce confi- dent errors specifically on gold-NO instances. We analyse the status-level consequences of this shift 4 Map accessTextAccuracyF1 macro Recall pos Recall neg Yes-rate Text-only curL.442.439.261.910.213 curT.639.604.647.616.574 startL.533.532.409.855.336 startT.614.591.590.677.515 Both maps curL.627.593.633.609.566 curT.654.618.665.625.584 startL.694.630.769.501.694 startT.737.671.822.518.727 Giver-only curL.626.586.650.565.590 curT.688.632.747.536.668 startL.723.644.827.454.749 startT.756.669.879.436.791 Follower-only curL.619.569.663.504.617 curT.653.594.716.490.658 startL.718.632.832.422.761 startT.743.650.872.408.794 Table 1: Baseline results across text-access and map-access conditions for the model Qwen3-VL-8B-Instruct. Best F1 macro per map-access level is startT in all map conditions; for text-only, curT slightly outperforms startT. in §5.2, including its differential impact on aligned, pending, and misunderstood instances. Single-map conditions amplify the bias. The giver-only condition at startT achieves similar F1 macro (.669) to both maps (.671), but with an even higher yes-rate (.791) and lower recall neg (.436). The follower-only condition shows a comparable pattern (yes-rate .794, recall neg .408). Access to ei- ther map is sufficient to trigger the over-alignment bias, and providing both maps actually moderates it slightly by introducing cross-map discrepancy evidence. Map infoF1 macro R pos R neg Yes-rate Baseline Text-only.591.590.677.515 Both maps.671.822.518.727 Textual map information Landmark names.636.756.533.675 Discrepancy Desc..668.810.528.716 Fake visual controls Blank maps.407.220.910.184 Shuffled maps.402.222.884.193 Table 2: Map-information modality comparison at startT (both-maps access level). All conditions use Qwen3- VL-8B-Instruct. R pos = recall on aligned; R neg = recall on not-aligned. F2: Textual map descriptions reproduce the over-alignment bias; only content-free visual controls avoid it.To determine whether the over- alignment bias is driven by map content or by the visual input channel, we compare authentic maps against textual map descriptions and non- informative visual controls. All conditions use the same model (Qwen3-VL-8B-Instruct) at the startT text window, ensuring a clean within-architecture comparison. Table 2 shows the results. Blank maps (yes-rate .184) and shuffled maps (yes-rate .193) produce lower yes-rates than the text-only baseline (.515), let alone authentic maps (.727). The model becomes more conservative when it receives images from which it cannot ex- tract task-relevant content. This rules out the hy- pothesis that the over-alignment bias is caused by the mere presence of images rather than by the information they carry. Both textual conditions produce yes-rates (.675â .716) close to real maps (.727) and well above the text-only baseline (.515). Macro F1 macro for the textual conditions (.636â.668) sits just below real maps (.671) and well above text-only (.591). The recall profiles echo real maps: elevated recall pos (.756â.810) and reduced recall neg (.528â.533), the same directional shift as real maps (.822 / .518), only slightly less extreme. The same landmark information â which land- marks appear on which maps and where they differ â triggers over-alignment whether it is presented visually or textually. The over-alignment bias is therefore about what the model learns about the scene, not about the visual channel per se. The vi- sual presentation does contribute a small additional shift (a few points on yes-rate and on recall pos ) rela- tive to the textual presentation of the same content, consistent with spatial co-presence in images being 5 a slightly stronger perceptual cue, but this residual effect is small compared to the content effect itself. The conditions thus split into two groups by yes- rate: inputs with task-relevant map content (real maps .727, textual descriptions .675â.716) over- predict alignment; inputs without map content (text- only .515, blank maps .184, shuffled maps .193) do not. Map content thus seems to drive the bias, while the visual presentation channel has only a secondary amplifying effect. F3: Broader dialogue context helps, but this is mitigated by map access. The text-window ef- fect (curLâstartT) produces a .152 F1 macro gain in the VL text-only condition, but only .078 under both maps. This means when the model has no ac- cess to maps, giving it more dialogue context helps a lot more. When it already has both maps, extra dialogue context still helps, but much less. This smaller gain suggests that, under map access, the model relies more heavily on static referential cues from the maps and benefits less from additional dialogue evidence about how grounding unfolds over time. Future dialogue interactions also help. Across all map conditions, both curTâstartT and curL âstartL produce consistent gains in F1 macro . This means that subsequent turns often contain useful repair, confirmation, or clarification evidence that makes the grounding outcome more legible. 5 Further Analysis and Discussion The preceding results show that map access im- proves overall performance while amplifying the over-alignment bias, and that this bias is driven by task-relevant map content rather than by the visual channel itself. We now probe the mechanism be- hind this pattern through calibration analysis (§5.1), a status-level trade-off analysis (§5.2), reference- chain tracking (§5.3), and cross-model comparison (§5.4). 5.1 Over-Prediction and Calibration Since we use vLLMâs constrained decoding to force a single YES/NO token, the cumulative log- probability of the generated token directly gives the modelâs confidence:conf := exp(logprob). We also compute Expected Calibration Error (ECE) (Pakdaman Naeini et al., 2015; Guo et al., 2017) to measure how calibration quality varies depending on whether the modelâs default response happens to be correct. ConditionECEECE yes ECE no Conf. Text-only.249.263.235.863 Both maps.174.094.403.912 Giver-only.172.057.484.927 Follower-only.185.061.524.929 Table 3: Calibration by gold label at startT (n= 13,077). All conditions use Qwen3-VL-8B-Instruct. ECE yes /ECE no = ECE on gold-YES/gold-NO instances. Conf. = mean prediction confidence. Table 3 shows the results with Qwen3-VL-8B- Instruct at startT across all 13,077 instances. Here are the key findings: Map conditions are better calibrated on aligned instances but miscalibrated on non-aligned ones. Under both-maps, the model is biased toward YES (yes-rate .727) and is well-calibrated on gold-YES instances (ECE yes = .094) but badly miscalibrated on gold-NO instances (ECE no = .403). The text- only condition, which is not strongly biased toward either class (yes-rate .515), is moderately calibrated on both classes (ECE yes = .263, ECE no = .235). The asymmetry is sharpest under single-map conditions: follower-only achieves ECE yes = .061 but ECE no = .524. Maps make the model more confident, and more confidently wrong on non-aligned instances. Mean confidence rises modestly with map access (.863â.912 / .927 / .929), but the distribution of that confidence shifts: the ECE gap between gold-YES and gold-NO instances widens from .028 (text-only) to .463 (follower-only). On gold-YES instances the model is confident and right; on gold- NO instances it is equally confident but wrong. In other words, when the model gets non-aligned in- stances wrong, it is confidently wrong. Map evi- dence drives the model to predict YES with high certainty even when alignment does not hold. This class-conditioned miscalibration is the strongest ev- idence for over-prediction: on gold-NO instances, the model is not merely wrong but confidently wrong, indicating that map content drives spurious certainty rather than merely shifting a threshold. 5.2 The Status-Level Trade-Off The overall improvement from map access con- ceals an asymmetric trade-off. Table 4 decomposes accuracy by the gold grounding status of each RE: aligned (gold YES;n= 9,435), pending (not yet grounded; gold NO;n= 3,403), or misunderstood 6 (grounded to different landmarks; gold NO;n= 239). ConditionAlignedPendingMisund. Baseline Text-only.590.691.473 Both maps.822.523.456 Giver-only.879.441.372 Follower-only.872.419.255 Textual map information Landmark names.756.544.372 Disc. detail.810.540.368 Fake visual controls Blank maps.220.913.866 Shuffled maps.222.886.858 Table 4: Accuracy by gold grounding status at startT (both-maps access level). Each cell shows the fraction of correct predictions within that status group. All con- ditions use Qwen3-VL-8B. Maps boost aligned accuracy at the cost of pend- ing and misunderstood accuracy. Maps boost aligned accuracy by 23â29 percentage points (.590 â.822 / .872 / .879) while simultaneously drop- ping pending accuracy by 17â27 points (.691â .419 / .441 / .523; McNemarp < 10 â6 for all map conditions). On misunderstood REs, the effect is condition-dependent: the both-maps drop (.473â .456) is not significant (p = 0.724,n = 239), but single-map conditions show large significant de- clines. Follower-only drops to .255 (p = 3Ă 10 â9 ) and giver-only to .372 (p = 8Ă 10 â3 ). Both maps moderate the misunderstood collapse. Having both maps is actually the least extreme map condition on misunderstood (0.456 vs. 0.372 giver- only vs. 0.255 follower-only), likely because cross- map discrepancies provide corrective evidence that moderates the bias relative to single-map condi- tions. The accuracy drop on pending cases, by contrast, is broad and significant for all map condi- tions. Textual descriptions follow the same trade-off as authentic maps.Textual map descriptions show the same directional profile as authentic maps rel- ative to the text-only baseline (.590 / .691 / .473 for aligned / pending / misunderstood): aligned accuracy rises sharply while pending and misunder- stood accuracy drop. Discrepancy-detail reaches .810 / .540 / .368 and landmark names .756 / .544 / .372 â close to both-maps (.822 / .523 / .456) on aligned and pending, but notably below both-maps on misunderstood. Their F1 macro (.636 / .668) sits just under both-maps (.671). Fake visual controls, by contrast, push the profile in the opposite direction: blank and shuffled maps produce near-identical hyper-conservative patterns (.220 / .222 aligned, .913 / .886 pending, .866 / .858 misunderstood), collapsing into a near-constant NO response (yes-rates .184 / .193). The conditions therefore form two groups: content-rich inputs (real maps and textual descrip- tions) over-predict alignment, while content-free visual inputs (blank, shuffled) amplify caution to the opposite extreme. Telling the model what is on the maps or showing it produces similar behaviour; it is the presence of task-relevant content about po- tential common ground that drives the trade-off, not the modality. The model confuses potential with established common ground. Map content tells the model what could be shared (landmarks appearing on both maps); dialogue history tells it what has been shared (interpretations established through interac- tion). The model over-weights the former. This is the computational analogue of the overhearerâs illusion: overhearers systematically overestimate their understanding of a conversation because they have access to referential context but lack the in- teractive grounding process that establishes mutual understanding between participants. The model is structurally an overhearer, which means it observes what both participants could share but cannot well assess what they have confirmed through dialogue. 5.3 Reference-Chain Analysis We use the reference chains defined in Li et al. (2026), sequences of mentions of the same land- mark within a dialogue, to test whether repeated mention helps the model converge on the correct judgment. We group 1,665 chains into six buckets by chain length, the number of mentions of that landmark in the dialogue (1, 2, 3, 4â5, 6â8, 9+), and report accuracy together with yes-rate, because chain-length effects are partly driven by response bias that macro F1 alone obscures. See Figure 2. Over-prediction of alignment grows with re- peated mention. Map conditions maintain high accuracy across chain lengths (both-maps: .681 at length 1â.791 at 9+), while text-only degrades (.719â.599; Figure 2a). However, this comes with steadily rising yes-rates under map conditions (both-maps: .549â.789; single-map reaches .840â .845; Figure 2b). Since aligned REs dominate at 7 1234-56-89+ Chain length 0.4 0.6 0.8 Accuracy (a) Accuracy by chain length 1234-56-89+ Chain length 0.2 0.4 0.6 0.8 Yes-rate (b) Yes-rate by chain length 123456789+ RE position (nth mention) 0.4 0.6 0.8 Mean P(Yes) (c) P(Yes) by RE position Text-onlyBoth mapsGiver-onlyFollower-only Figure 2: (a) Accuracy and (b) yes-rate by reference- chain length; (c) meanP(YES)by RE position (nth mention within a chain), with 95% CI shading. Map conditions maintain high accuracy but with steadily in- creasing yes-rate, andP(YES)rises with RE position in most situations. longer chain lengths (post-grounding mentions ac- cumulate), this inflation lets maps appear more accurate than they are. The same pattern emerges within chains:P(YES)rises steadily with RE po- sition in most conditions, especially at the early mentions (Figure 2c) â text-only from .420 at the first mention to .601 at 9+, and map conditions (.708â.794) by +.034 to +.118. 5.4 Cross-Model Comparison To assess whether the above findings and model behavioural patterns generalise beyond a single model, we evaluate four additional VLMs from two VL model families on the baseline condition grid at startT. Table 5 shows the results. Text-onlyBoth maps ModelF1 m YR F1 m YR ECE Conf Qwen3-VL-2B .561 .483 .566 .707 .144 .789 Qwen3-VL-4B .518 .368 .411 .225 .494 .907 Qwen3-VL-8B .591 .515 .671 .727 .174 .912 Gemma-3-4B.445 .279 .510 .400 .457 .973 Gemma-3-12B .322 .107 .416 .231 .531 .948 Table 5: Cross-model comparison at startT on the full dataset (n = 13,077). F1 m = F1 macro ; YR = yes-rate. Models respond to map information in qualita- tively different ways.Adding maps increases the yes-rate for Qwen3-VL-8B (.515â.727), Qwen3- VL-2B (.483â.707), and Gemma-3-4B (.279â .400), the same over-alignment direction observed in the primary experiments. However, Qwen3- VL-4B responds in the opposite direction: maps make it more conservative (yes-rate .368â.225, F1 macro .518â.411). Gemma-3-12B is extremely conservative across all conditions (yes-rate .107 text-only, .231 both-maps), barely engaging with the task. These divergent responses suggest that the over-alignment bias is not a universal property of vision-language architectures but depends on model-specific factors. One likely contributor to Gemma3âs weaker performance is its limited ability to parse the information-dense hand-drawn MapTask map im- ages. The two families differ substantially in how they encode visual input. Qwen3-VL uses a na- tive dynamic-resolution vision encoder (Bai et al., 2025): images are tiled into patches at their in- put resolution, producing a variable number of vi- sual tokens that scales with image area. For our 791Ă1024 map images, this yields hundreds of tokens with 2D-RoPE-based spatial positional en- coding, preserving fine-grained detail such as small landmark labels and icons. Gemma3 uses a SigLIP- based vision encoder (Gemma Team, 2025) that resizes images to a fixed 896Ă896 resolution and average-pools the patch representations down to a budget of 256 tokens per image, regardless of image size or content complexity. This aggressive compression likely discards spatial detail that is critical for reading the MapTask maps. Prior to the main experiments, we ran a sanity check in which all five models were prompted to list landmark names and describe their spatial positions on each of the 32 map images (Appendix E). Qwen3-VL models achieved 88â90% F1 on landmark identifi- cation, compared to 81â82% for Gemma3 models. Gemma3 models also introduced character-level naming errors and spatial mislocations that Qwen3- VL models did not produce. If a model cannot reliably extract map content, map access cannot produce the content-driven over-alignment bias we observe in Qwen3-VL. Figures 7, 8, and 9 in Appendix F show the full condition-level breakdown. The status-level anal- ysis (Figure 9) reveals that the aligned vs. non- aligned case trade-off observed for Qwen3-VL-8B generalises across models: map access consistently improves aligned-case F1 while degrading perfor- mance on pending and misunderstood REs, with the exception of Qwen3-VL-4B, where the overall conservative shift suppresses performance across all status categories. 8 Model size does not predict task performance. Within both families, the scaling relationship is non-monotonic: Qwen3-VL-2B outperforms 4B on both-maps (.566 vs. .411), and Gemma-3-4B outperforms 12B (.510 vs. .416). Calibration data reveals why: Qwen3-VL-2B achieves the best cal- ibration (ECE = .144) because it is the least con- fident (mean confidence .789), while Gemma-3- 12B achieves the worst (ECE = .531) at the second highest confidence (.948). Larger models become more confident without becoming more discerning. The bottleneck may be the tendency to commit to reference-related evidence-driven grounding estab- lishment with excessive confidence. 6 Conclusion We set out to test whether VLMs can judge in- terpretation matching in information-asymmetric collaborative dialogue, using an evaluation method- ology based on systematically controlled map- information and dialogue-context conditions ap- plied to the HCRC MapTask. The main result is not simply that maps help. Authentic maps im- prove overall performance, but they do so by push- ing models toward YES: landmark co-presence is treated as evidence of mutual understanding. Tex- tual descriptions of the same map content repro- duce the same bias, while non-informative visual in- puts (blank, shuffled) reverse it into hyper-caution. The problem is therefore not multimodality in gen- eral, nor a specific visual-perceptual cue, but a tendency to read task-relevant information about potential common ground as evidence that it has been established through grounding. Further analyses show where this bias is struc- tured. Under map conditions, models become con- fidently biased toward aligned judgments, gain on already aligned cases while losing accuracy on pending and misunderstood ones, and grow more likely to predict alignment across repeated men- tions. Taken together, these patterns suggest that current VLMs are better at assessing potential ref- erential overlap from map content than at tracking grounding as an interactional and incremental pro- cess. In MapTask terms, they can infer what the interlocutors could be talking about, but they do not reliably distinguish this from what the interlocutors have actually established through grounding. Our experiments are limited to an overhearer setting in a single domain, so an important next step is to test interactive settings where models can ask for clarification, express uncertainty, or revise their judgments over time. More broadly, extend- ing the evaluation beyond MapTask and relating the observed heterogeneity to model properties via mechanistic interpretability analysis (e.g., atten- tion flow analysis, Zhang et al., 2025) may clarify whether the observed over-alignment is a general behavioural tendency in VLMs for collaborative dialogue, or a narrower failure mode tied to this task, setting, and/or model family. Limitations GeneralizationOur primary evaluation relies on one single model (Qwen3-VL-8B-Instruct) for the detailed condition grid, with additional models tested only on baseline conditions. The dataset derives from HCRC MapTask dia- logues, a single corpus with specific properties that may shape the observed effects: a 72.1%/27.9% class imbalance toward aligned cases, MapTask- specific discrepancy types (missing, duplicated, or renamed landmarks), and a rare misunderstood cat- egory (239 of 13,077 instances). Suitable datasets for replication should provide perspectivist anno- tations recording both interlocutorsâ referent in- terpretations separately, not just task success or dialogue-level grounding labels, which limits the pool of available corpora. Generalisation to other information-asymmetric settings remains to be es- tablished. Task DesignThe evaluation places the model in an overhearer position, characterising a judgment failure rather than an interaction failure. We use greedy decoding to elicit a binary judg- ment, which may not reflect the modelâs full distri- butional beliefs about alignment. A softer evalua- tion protocol (e.g., allowing multi-choice selection or free-form responses) might reveal a more nu- anced picture. Textual Map ReconstructionOur textual condi- tions (Text-landmark-names and Text-discrepancy- detail) convey landmark name lists and inter-map discrepancies, but they do not fully reconstruct the spatial layout of the maps. The behavioral gap be- tween textual and visual conditions may therefore partly reflect this incomplete reconstruction rather than a genuine visual-channel effect. 9 Acknowledgments We appreciate the helpful comments and sug- gestions from the anonymous reviewers. This work is funded by the Dutch Research Coun- cil (NWO) through the AiNed Fellowship Grant NGF.1607.22.002, Dealing with Meaning Varia- tion in NLP. Code and Data Availability We release our code and prompt templates to facili- tate reproducibility 1 . References Anne H. Anderson, Miles Bader, Ellen Gurman Bard, Elizabeth Boyle, Gwyneth Doherty, Simon Garrod, Stephen Isard, Jacqueline Kowtko, Jan McAllister, Jim Miller, Catherine Sotillo, Henry S. Thompson, and Regina Weinert. 1991. The HCRC Map Task corpus. Language and Speech, 34(4):351â366. Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhi- fang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, and 45 others. 2025. Qwen3-VL technical report. Preprint, arXiv:2511.21631. Javier Chiyah-Garcia, Alessandro Suglia, Arash Eshghi, and Helen Hastie. 2023. âWhat are you referring to?â Evaluating the ability of multi-modal dialogue models to process clarificational exchanges. In Pro- ceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 175â182, Prague, Czechia. Association for Computa- tional Linguistics. Herbert H. Clark and Susan E. Brennan. 1991. Ground- ing in communication. In Lauren B. Resnick, John M. Levine, and Stephanie D. Teasley, editors, Perspec- tives on Socially Shared Cognition., pages 127â149. American Psychological Association. Herbert H. Clark and Edward F. Schaefer. 1989. Con- tributing to discourse. Cognitive Science, 13(2):259â 294. Herbert H. Clark and Deanna Wilkes-Gibbs. 1986. Re- ferring as a collaborative process. Cognition, 22(1):1â 39. Gemma Team. 2025.Gemma 3 technical report. Preprint, arXiv:2503.19786. 1 https://github.com/chnln/ seeing-is-not-sharing Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Wein- berger. 2017. On calibration of modern neural net- works. In Proceedings of the 34th International Con- ference on Machine Learning, volume 70 of Pro- ceedings of Machine Learning Research, pages 1321â 1330. PMLR. Janosch Haber, Tim BaumgĂ€rtner, Ece Takmaz, Lieke Gelderloos, Elia Bruni, and Raquel FernĂĄndez. 2019. The PhotoBook dataset: Building common ground through visually-grounded dialogue. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1895â1910, Flo- rence, Italy. Association for Computational Linguis- tics. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serv- ing with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611â626. ACM. Nan Li, Albert Gatt, and Massimo Poesio. 2026. Grounded misunderstandings in asymmetric dia- logue: A perspectivist annotation scheme for Map- Task. In Proceedings of the Fifteenth Language Re- sources and Evaluation Conference (LREC 2026), pages 4988â5001, Palma, Mallorca, Spain. European Language Resources Association (ELRA). Brielen Madureira and David Schlangen. 2024. It couldnât help but overhear: On the limits of mod- elling meta-communicative grounding acts with su- pervised learning. In Proceedings of the 25th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 149â158, Kyoto, Japan. Associ- ation for Computational Linguistics. Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht. 2015. Obtaining well calibrated proba- bilities using Bayesian binning. In Proceedings of the AAAI Conference on Artificial Intelligence, vol- ume 29. Association for the Advancement of Artifi- cial Intelligence. Michael F Schober and Herbert H Clark. 1989. Under- standing by addressees and overhearers. Cognitive Psychology, 21(2):211â232. Omar Shaikh, Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2025. Navigating rifts in human-LLM grounding: Study and benchmark. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 20832â20847, Vienna, Austria. Association for Computational Linguistics. Takuma Udagawa and Akiko Aizawa. 2019. A nat- ural language corpus of common grounding under continuous and partially-observable context. In Pro- ceedings of the AAAI Conference on Artificial Intelli- gence, volume 33, pages 7120â7127. Association for the Advancement of Artificial Intelligence. 10 Zhengxiang Wang, Weiling Li, Panagiotis Kaliosis, Owen Rambow, and Susan Brennan. 2025. LVLMs are bad at overhearing human referential communi- cation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 16758â16782, Suzhou, China. Association for Computational Linguistics. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pier- ric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, and 3 others. 2020. Trans- formers: State-of-the-art natural language processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38â45, Online. Association for Computational Linguistics. Peter Zeng, Weiling Li, Amie J. Paige, Zhengxi- ang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory Zelinsky, Susan E. Brennan, and Owen Rambow. 2026. LVLMs and humans ground dif- ferently in referential communication.Preprint, arXiv:2601.19792. Zhi Zhang, Srishti Yadav, Fengze Han, and Ekaterina Shutova. 2025. Cross-modal information flow in mul- timodal large language models. In 2025 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR), pages 19781â19791. IEEE. A Dataset and Annotation Our experiments use the perspectivist annotation release of the HCRC MapTask corpus (Li et al., 2026), derived from the original corpus of An- derson et al. (1991). The release contains 13,077 reference expressions (REs) annotated across 128 dialogues, paired with 16 distinct map pairs (m0â m15, each used in 8 dialogues). For each RE the annotation records the giverâs and followerâs in- terpretations separately, which is what makes the YES/NO interpretation-matching task studied here well-defined. Landmark discrepancy typesThe 16 map pairs contain four kinds of landmark variation: identical landmarks appear at the same position under the same name on both maps; lexical variants appear at the same position but under different names (e.g., white water on the giverâs map vs. rapids on the followerâs; 10 such pairs are documented); exis- tence discrepancies (landmarks appearing on only one of the two maps); and multiplicity discrep- ancies (landmarks appearing twice on one map but only once on the other; 16 such landmarks are documented, one per map pair). Existence and mul- tiplicity discrepancies are the structural source of most misalignment in this corpus; lexical variants are resolved by name unification during annotation so they do not on their own produce misunderstood REs. Landmark ID formatBecause the original Map- Task corpus uses the same landmark ID for all instances of same-named landmarks, the annota- tion release introduces a unified ID scheme that disambiguates multiplicity landmarks and tracks map-side provenance. Each landmark ID has the format <map_id>_<concept>#<ordinal>@<side>, where<map_id>ism0âm15;<concept>is the original landmark name (e.g.,diamond_mine); #<ordinal>is present only for multiplicity land- marks that appear twice on one map, with0denot- ing the lower/bottom instance and1the upper/top; and@<side>isgfor the giverâs map orffor the followerâs.Examples:m9_stony_desert@g, m9_site_of_plane_crash#0@g, m2_stone_creek#1@f. Annotation cascadeEach RE is annotated along a five-step cascade that models incremental reso- lution. Each step is evaluated only when the pre- ceding conditions are met; when the cascade ter- minates early, all downstream attributes are set to null in the released data. Step AttributePerspectiveReq. 1 is_quantificational Speaker false 2 is_specifiedAddressee true 3 is_accommodatedAddressee true 4 is_groundedAddressee true 5 is_imaginedAddresseeâ Table 6: The five-step annotation cascade used to guide the LLM and improve the annotation quality (Li et al., 2026). The Req. column shows the value an attribute must take for the cascade to proceed; step 5 is terminal. Understanding states Each RE is assigned one of three understanding states, derived post-hoc from the cascade attributes and the match between the giverâs and the followerâs interpretation fields (after lexical-variant unification). In the released dataset: aligned (both participants ground the RE to the same or equivalent landmark) accounts for 9,435 REs (72.1%); pending (the RE is quantifi- cational, unspecified, unaccommodated, or other- wise ungrounded) accounts for 3,403 REs (26.0%); and misunderstood (both participants believe they 11 agree but ground to different landmarks) accounts for 239 REs (1.8%). The main-text analyses in §5.2 decompose model performance along these three states. B Experimental Setup and Reproducibility All inference runs use vLLM (Kwon et al., 2023) as the backend with greedy decoding (temperature = 0.0,random_seed = 42),max_new_tokens = 16 , and constrained output via vLLMâs logit masking to the YES / NO token choice. For the cal- ibration analyses reported in §5.1, we additionally record the top-20 logprobs per generation step. All experiments run on a single NVIDIA A100 80 GB GPU. For every Qwen3-VL model we explicitly disable the built-in thinking mode to ensure a fair comparison with other models. All the models are accessed via the Hugging Face Transformers library (Wolf et al., 2020). C Prompt Template Figure 3 shows the full prompt template used for all conditions. D Textual Map-Information Examples This section shows the text that fills the $map_accessslot of the prompt template (Fig- ure 3) for the two textual map-information condi- tions introduced in §3.2. Both examples are drawn from the same instance â dialogueq1ec1(map pairm12), target RE a caravan park â so the two variants can be compared directly. text-landmark-names . Under this condition, the model is given only the list of landmark names on each participantâs map. The system promptâs map-information line reads: âYou are given the list of landmark names on each partic- ipantâs map (see below in the dialogue context).â The$map_accessslot of the user prompt is filled with the block shown in Figure 4. text-discrepancy-detail. In addition to the per-map landmark lists, this condition supplies an explicit textual summary of how the two maps dif- fer (per-side exclusives, multiplicity landmarks, and shared landmarks). The system promptâs map- information line becomes: âYou are given the land- mark names on each participantâs map and a de- scription of how the two maps differ (see below in the dialogue context).â The$map_accessslot is filled with the block shown in Figure 5. E Map Reading Sanity Check Prior to the main experiments, we assessed whether the five VLMs can reliably read the hand-drawn MapTask map images by prompting each model with two tasks on all 32 maps: (1) list all land- mark names visible on the map, and (2) describe each landmarkâs spatial position. Both tasks use greedy decoding with free-form text output (no constrained decoding). Prompts. For task 1: âList all landmark names you can see on this map. Output only the landmark names, one per line. Do not add numbering, de- scriptions, or any other text.â For task 2: âFor each landmark you can see on this map, describe its position on the map (e.g., top-left, upper-center, center, bottom-right). Format each line as: land- mark name â position. Do not add any other text.â Landmark listing. Table 7 reports recall, pre- cision, and F1 for landmark-name identification across all 32 maps, evaluated against the gold land- mark lists from the corpus metadata. Matching uses exact string comparison after lowercasing and whitespace normalisation. ModelRecallPrecisionF1 Qwen3-VL-2B.845.910.876 Qwen3-VL-4B.861.937.897 Qwen3-VL-8B.812.956.878 Gemma-3-4B.766.866.813 Gemma-3-12B.749.907.820 Table 7: Landmark-name identification on all 32 Map- Task map images. Recall = fraction of gold landmarks listed; Precision = fraction of model outputs that match a gold landmark. Qwen3-VL models achieve higher F1 (.876â .897) than Gemma3 models (.813â.820). Qwen3- VL-8B has the highest precision (.956), producing almost no spurious landmark names. Gemma3 models introduce character-level naming errors that Qwen3-VL avoids, such as âpicker fenceâ in- stead of âpicket fenceâ (both Gemma3 models) and âpits of forest fireâ instead of âsite of forest fireâ (Gemma3-4B). Spatial description. We also prompted each model to describe landmark positions on all 32 maps. Figure 6 shows the giverâs map for map 12 LandmarkModelPosition picket fence Qwen3-VL-2B top-left Qwen3-VL-4B top-left Qwen3-VL-8B top-left Gemma-3-4B upper-right â Gemma-3-12B top-left east lake Qwen3-VL-2B top-right Qwen3-VL-4B top-right Qwen3-VL-8B top-right Gemma-3-4B bottom-right â Gemma-3-12B right-center â FINISH Qwen3-VL-2B top-right Qwen3-VL-4B top-right Qwen3-VL-8B top-right Gemma-3-4B center â Gemma-3-12B right-center â START Qwen3-VL-2B â Qwen3-VL-4B bottom-left Qwen3-VL-8B bottom-left Gemma-3-4B center-left â Gemma-3-12B â camera shop Qwen3-VL-2B bottom-left Qwen3-VL-4B bottom-left Qwen3-VL-8B bottom-left Gemma-3-4B center-left â Gemma-3-12B top-left â parked van (lower) Qwen3-VL-2B bottom-left Qwen3-VL-4B bottom-left Qwen3-VL-8B bottom-left Gemma-3-4B bottom-right â Gemma-3-12B bottom-left Table 8: Spatial descriptions on map0g (Figure 6) for six landmarks where Gemma3 models make clear errors. Qwen3-VL models produce correct or near-correct po- sitions on all 14 landmarks; only the six with Gemma3 errors are shown. âââ = not listed. â = position incon- sistent with the map. pair 0 (map0g), and Table 8 compares the out- puts on this map for six landmarks where at least one Gemma3 model makes a clear spatial error. Qwen3-VL models produce correct or near-correct positions on all 14 landmarks; the six shown are selected to illustrate Gemma3âs failure pattern. For example, Gemma-3-4B places east lake at bottom- right (it is at top-right), START and camera shop at center-left (both are at bottom-left), and picket fence at upper-right (it is at top-left). Gemma- 3-12B similarly mislocates east lake and FINISH. Both Gemma3 models also output âpicker fenceâ instead of âpicket fenceâ. F Cross-Model Detailed Results Figures 7, 8 and 9 provide detailed breakdowns of the cross-model comparison in §5.4. 13 System Prompt You are a dialogue analysis expert. You are overhearing two participants doing a MapTask-style route navigation task. MapTask Background: - Each participant has their own map and they cannot see each otherâs map. - One participant (the Giver) describes a route using named landmarks on their map. - The other participant (the Follower) tries to follow the instructions on their own map. - The two maps may differ (some landmarks may be missing, duplicated, or placed differently), so the two participants can end up with different personal interpretations even if the dialogue sounds smooth. Task: You will be given a dialogue context in which ONE target reference expression is marked with «...». Decide whether the two participants interpret that reference expression as pointing to the SAME specific landmark (interpretations match). Guidelines: - Use only the provided dialogue context and (if available) the provided map image(s). - A match can be supported by clear evidence of successful grounding (e.g., consistent descriptions, confirmations, coherent subsequent navigation). - If the expression indicates quantificational asking, or unspecified/unresolved grounding, treat it as NOT a match. Output: Answer with exactly one word: Yes or No. - Yes = interpretations match. - No = interpretations do not match. Do not output anything else. Information you can access in this instance: - Dialogue text: $text_access - Map information: $map_access User Prompt Below is the specific information for this judgement. Target reference expression: $target_ref Dialogue context (the target RE is wrapped in « »): $context Maps: if map images are provided, they appear below. Figure 3: Prompt template. Template variables ($...) are filled per instance based on the text-access and map- access conditions. For example, under the startT text-access window and both-maps access level,$text_access is filled with âYou can read the dialogue from the beginning of the conversation through the end of the transaction that contains the target reference expression (i.e., including the subsequent lines in that transaction after the target line).â and$map_accessis filled with âYou are shown both the Giverâs and the Followerâs map images.â; the two map images are appended after the user prompt. 14 Map landmark information: - Giverâs map landmarks: start, caravan park, old mill, abandoned cottage, fenced meadow, fenced meadow, west lake, trig point, monument, nuclear test site, east lake, farmed land, finish - Followerâs map landmarks: start, caravan park, picket fence, mill wheel, forest, abandoned cottage, fenced meadow, west lake, monument, golf course, east lake, farmed land Figure 4: Example filling of $map_access under the text-landmark-names condition for dialogue q1ec1. Map landmark information: - Giverâs map landmarks: start, caravan park, old mill, abandoned cottage, fenced meadow, fenced meadow, west lake, trig point, monument, nuclear test site, east lake, farmed land, finish - Followerâs map landmarks: start, caravan park, picket fence, mill wheel, forest, abandoned cottage, fenced meadow, west lake, monument, golf course, east lake, farmed land Discrepancies between maps: - Landmarks on Giverâs map ONLY (not on Followerâs): finish, nuclear test site, old mill, trig point - Landmarks on Followerâs map ONLY (not on Giverâs): forest, golf course, mill wheel, picket fence - Landmarks appearing multiple times: fenced meadow appears 2 times on Giverâs map - Shared landmarks (on both maps): abandoned cottage, caravan park, east lake, farmed land, fenced meadow, monument, start, west lake Figure 5: Example filling of$map_accessunder thetext-discrepancy-detailcondition for dialogueq1ec1. 15 Figure 6: Giverâs map for map pair 0 (map0g), used for the spatial-description comparison in Table 8. 16 Qwen3-VL-8BQwen3-VL-2BQwen3-VL-4B 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 f1_macro Qwen: Macro F1 Qwen3-VL-8BQwen3-VL-2BQwen3-VL-4B 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 yes_rate Qwen: Yes-Rate Text-only Both maps Follower-only Giver-only Figure 7: Macro F1 and yes-rate across map-access conditions for Qwen3-VL models (8B, 2B, 4B) at startT. Gemma-3-4BGemma-3-12B 0.0 0.1 0.2 0.3 0.4 0.5 f1_macro Gemma: Macro F1 Gemma-3-4BGemma-3-12B 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 yes_rate Gemma: Yes-Rate Both maps Follower-only Giver-only Text-only Figure 8: Macro F1 and yes-rate across map-access conditions for Gemma-3 models (4B, 12B) at startT. Qwen3-VL-8BQwen3-VL-2BQwen3-VL-4B Gemma-3-4B Gemma-3-12B 0.0 0.1 0.2 0.3 0.4 f1_macro Macro F1 aligned Qwen3-VL-8BQwen3-VL-2BQwen3-VL-4B Gemma-3-4B Gemma-3-12B 0.0 0.1 0.2 0.3 0.4 0.5 f1_macro Macro F1 pending Qwen3-VL-8BQwen3-VL-2BQwen3-VL-4B Gemma-3-4B Gemma-3-12B 0.0 0.1 0.2 0.3 0.4 0.5 f1_macro Macro F1 misunderstood Condition Text-only Both maps Follower-only Giver-only Figure 9: Macro F1 by grounding status (aligned, pending, misunderstood) across all models and map-access conditions at startT. 17