Paper deep dive
PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models
Jihyung Ko, Eunji Jung, Hyeongsub Kim, Ziseok Lee, Jae Won Cho, Sanghyun Jo, Kyungsu Kim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/25/2026, 7:24:05 AM
Summary
The paper introduces PatchGate, a training-free framework for Vision-Language Models (VLMs) that mitigates object hallucination and omission by extracting intrinsic object evidence from frozen model layers and calibrating decoding logits. It achieves improved reliability on the AMBER and POPE benchmarks without fine-tuning or external detectors.
Entities (10)
Relation Signals (9)
PatchGate → consistsof → VIED
confidence 95% · PatchGate is a two-stage training-free framework... Visual Evidence eXtraction (VEX)... Visual-Evidence Inclusion-Exclusion Decoding (VIED)
PatchGate → consistsof → VEX
confidence 95% · PatchGate is a two-stage training-free framework... Visual Evidence eXtraction (VEX)... Visual-Evidence Inclusion-Exclusion Decoding (VIED)
PatchGate → evaluatedon → AMBER
confidence 95% · On AMBER, PatchGate improves both sides of object-level reliability...
VIED → employs → EDE
confidence 90% · VIED applies two logit-space operations... Evidence-Supported Inclusion (ESI)... Evidence-Deficient Exclusion (EDE)
VIED → employs → ESI
confidence 90% · VIED applies two logit-space operations... Evidence-Supported Inclusion (ESI)... Evidence-Deficient Exclusion (EDE)
PatchGate → evaluatedon → POPE
confidence 90% · boosting POPE Precision/Recall by 2.3%/20.6%...
PatchGate → increases → Cover
confidence 90% · increasing visible-object coverage from 49.4 to 56.0...
PatchGate → reduces → CHAIR
confidence 90% · reducing object hallucination by lowering CHAIR from 7.5 to 6.6...
VEX → uses → LLaVA-v1.5-7B
confidence 90% · For our default backbone, LLaVA-v1.5-7B... we use L={22,...,32}... VEX reads patch-level lexical evidence...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is already likely to mention, visible objects omitted from the output remain difficult to recover. We propose PatchGate, a training-free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions. In the first stage, Visual Evidence eXtraction (VEX) reads patch-level lexical evidence from the latter half of LM decoder layers and constructs an image-conditioned object set without any task prompt. In the second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED) uses this object evidence to calibrate decoding logits, promoting evidence-supported but under-verbalized objects and suppressing weakly supported but over-verbalized objects. On AMBER, PatchGate improves both sides of object-level reliability, increasing visible-object coverage from 49.4 to 56.0 (+13.4%) and reducing object hallucination by lowering CHAIR from 7.5 to 6.6 (-12.0%), without external detectors or fine-tuning and with one extra forward pass.
Tags
Links
- Source: https://arxiv.org/abs/2608.21819v1
- Canonical: https://arxiv.org/abs/2608.21819v1
Trouble viewing inline? Open PDF directly →
Full Text
97,878 characters extracted from source content.
Expand or collapse full text
PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models Jihyung Ko 1∗ , Eunji Jung 1∗ , Hyeongsub Kim 1,2 , Ziseok Lee 1 , Jae Won Cho 3 , Sanghyun Jo 1,4† , and Kyungsu Kim 1† 1 Seoul National University, Korea 2 LG CNS, Korea 3 Konkuk University, Korea 4 OGQ, Korea Abstract.Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object mentions while covering visible objects. Existing training-free methods primarily address the former requirement, suppressing unsupported object words by intervening on model-predicted mentions during generation. Because they operate only on objects the model is already likely to mention, visible objects omitted from the output remain difficult to recover. We propose PatchGate, a training- free framework that extracts prompt-free object evidence intrinsic to a frozen VLM before generation and uses it to narrow the gap between an intrinsic object set and final object mentions. In the first stage, Visual Evidence eXtraction (VEX) reads patch-level lexical evidence from the latter half of LM decoder layers and constructs an image-conditioned object set without any task prompt. In the second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED) uses this object evidence to calibrate decoding logits, promoting evidence-supported but under-verbalized objects and suppressing weakly supported but over-verbalized objects. On AMBER, PatchGate improves both sides of object-level reliability, increasing visible-object coverage from 49.4 to 56.0 (+13.4%) and reducing object hallucination by lowering CHAIR from 7.5 to 6.6 (−12.0%), without external detectors or fine-tuning and with one extra forward pass. Keywords: Vision-Language Models· Image Captioning· Object Hallucination· Object Omission· Patch-Level Visual Evidence· Intrinsic Object Inventory· Training-free Decoding 1 Introduction Modern Vision-Language Models (VLMs) [2,20] have become a common interface for image under- standing, generating language descriptions, answers, and explanations grounded in visual inputs. As these models are increasingly used to describe visual content, their reliability depends not only on linguistic fluency but also on whether the generated text remains faithful to the image. For image captioning, faithful outputs should be both precise and complete, avoiding unsupported object mentions while covering visible objects. However, generated captions often violate these requirements in two complementary ways: they can mention unsupported objects or omit visible objects needed for a complete description, commonly referred to as object hallucination and object omission, respectively (see Fig. 1). Existing training-free methods seek more faithful captions by exploiting generation-time signals, including attention maps [8] and contrastive decoding distributions [17]. However, these signals mainly act on object words the model is already likely to generate, making them effective at suppressing unsupported object mentions but limited at recovering visible objects that were omitted. Other methods [33,36] leverage external detectors [4,22] to supply missing object cues, but these cues come from external visual modules rather than the target VLM’s own internal evidence, introducing additional modules and dependence on detector coverage and accuracy. ∗ Equal contribution. † Corresponding authors: shjo.april@gmail.com, kyskim@snu.ac.kr arXiv:2608.21819v1 [cs.CV] 22 Aug 2026 2J. Ko et al. (a)Baseline(e.g.,LLaVA-7B)(b)PatchGate (c)Baselinew/PatchGate(Ours) utImage Sk Flag iv'-' I Train Che7§ystem OutputCaption:jeimagefeaturesayellowandblacktraincar withmcwords"ChesieSystem"writtenonit.Thetrainistraveling downthetracks.anditappearstobeapassengertrain.Thetraincaris positionedinthemiddleofthescene,withal"ewothercarsvisiblein thebackground.Therearetw(peopleintheimage,onestanding closertotheletsideofthetrainandtheoilpersontndingfurther totheright.Thescenealsoincludesacarandaltruck...ittheear locatedontheleftsideoftheimageandthetiiick1therightside. Stage1.VisualEvidenceeXtraction(VEX) x. Per-patch *gr evidence (Vqu) - Objectextraction ) 5 llrlrI awwwf "I windows,engine,flag, people,sky... "mg _11-11. mf.JII orIll EFIIEIIE jin--it llllihi Stage2.Visual-EvidenceInclusion-ExclusionDecoding(VIED) (+)ESI:Evidence (-)EDE:Evidence SupportedInclusionDeficientExclusion (Omissions) r(llallucinationl) . N 1. 4 .andthereis flagnthe roofo1thetrain I I : . I Prompt1ieeinternal objectinventory (T(l)) éiéLLa -7B <1(I)) Promptfreeinternal objectinventory (T(I)) Training-FreeSimulfaneousHallucinafirm&OmissionOptimization OmissionHallucinationGroundTruthRecoveredOmittedObject InputImage Sk Flag Train Che7§gstem 431: OutputCaption:Theyellowandblacktrainengineistraveling downtracks.withwindowsonthefrontandbackofthetrain.The trainiscarryingayellowcaboosebehindit,andthereisaflagonthe roofofthetrain.Thesky.abovethetrainisblue.creatingapleasant backdropforthejourney. Fig. 1. Two complementary captioning failures and PatchGate’s simultaneous mitigation. Reliable captions should be precise and complete, yet generated captions can mention unsupported objects (red: people, truck ) while omitting visible objects (blue: sky, flag ). PatchGate uses prompt-free internal object evidence to calibrate decoding, suppressing unsupported mentions and recovering omitted visible objects (yellow: sky, flag ). We examine whether the frozen VLM contains object evidence before generation, and how this evidence differs from the objects ultimately mentioned in the caption. Building on logit-lens analyses in language models [3,7] and their extension to VLM patch representations [24], we read lexical evidence from internal patch hidden states without any task prompt. This comparison reveals an evidence-verbalization mismatch: evidence-supported visible objects can remain unmentioned, leading to object omissions, while weakly supported object words can still be mentioned, leading to object hallucinations. This suggests a shared basis for recovering under-verbalized visible objects and suppressing weakly supported object mentions. Based on this observation, we propose PatchGate, a two-stage training-free framework that uses the frozen VLM’s internal patch representations to narrow the evidence-verbalization mismatch. The first stage, Visual Evidence eXtraction (VEX), reads patch-level lexical evidence from the latter half of the LM decoder layers, scores each object candidate by its strongest normalized patch evidence, and constructs an intrinsic object inventory before generation, defined as a prompt-free, image-conditioned set of visually supported objects. The second stage, Visual-Evidence Inclusion-Exclusion Decoding (VIED), uses this inventory to calibrate decoding logits, jointly promoting evidence-supported but under-verbalized objects and suppressing weakly supported object mentions (Fig. 1). PatchGate thereby improves both caption precision and completeness without external detectors or fine-tuning. Our contributions are summarized as follows: – We characterize an object-level evidence-verbalization mismatch in frozen VLMs by compar- ing prompt-free internal object evidence with final object mentions: visual-evidence deficits among mentioned object words indicate unsupported mentions, while strong evidence among unmentioned objects indicates omitted visible objects. –We introduce PatchGate, a training-free framework that narrows this mismatch by extracting an intrinsic object inventory before generation and using it to calibrate decoding logits for more precise and complete captions. – PatchGate improves object-level reliability without external detectors or fine-tuning, increasing AMBER Cover by 13.4%, lowering CHAIR by 12.0%, and boosting POPE Precision/Recall by 2.3%/20.6% with consistent gains across backbones. PatchGate3 Table 1. Conceptual comparison of PatchGate (Ours) with representative training-free hallucination mitigation approaches: OPERA [8], VCD [17], ProjectAway [12], MARINE [36], SHIELD [10], PND [13] and ILVAD [30]. Properties [CVPR’24] OPERA [CVPR’24] VCD [ICLR’25] ProjectAway [ICML’25] MARINE [ICLR’26] SHIELD [CVPR’26] PND [ICML’26] ILVAD PatchGate (Ours) (a) Reduces object hallucinations for precision✓ (b) Reduces object omissions for completeness✗✓✗✓ (c) Reads internal patch-level object evidence ✗✓✗✓ (d) Builds a prompt-free object set before generation✗✓✗✓ (e) Corrects visual evidence-mention mismatch✗✓ (f) Requires no external visual module (e.g., tagger) ✓✗✓✗✓ 2 Related Work 2.1 Hallucination Mitigation in VLMs Reliable Vision-Language Models (VLMs) require captions to be both precise and complete: they should avoid object hallucinations, where captions mention objects unsupported by the image, while covering visible objects without omissions. Prior work has sought to improve this reliability through two broad approaches: training-based and training-free. Training-based methods [31,32] reduce hallucination through additional supervision or preference optimization, but require costly model updates and extra training data. Accordingly, training-free methods have gained traction by keeping the VLM fixed and modifying inference-time signals, offering a practical plug-and-play alternative. Representative training-free methods mitigate object hallucination through contrastive decoding [5, 11,17,29], attention intervention [1,8,14,21,30,34], visual encoder correction [10], or external object guidance [33,36]. Because these methods mainly act on objects the model is already likely to mention, they are effective at reducing hallucination but limited in recovering visible objects that never surface in the generated text. Recent work [13] also considers both hallucination and omission, but because it corrects object candidates only as they arise during generation, visible objects not covered by the caption remain difficult to identify and recover. PatchGate instead uses a pre-generation intrinsic object set to form a signed evidence-verbalization gap, whose two directions drive evidence-supported inclusion for omission recovery and evidence-deficient exclusion for hallucination suppression, jointly improving caption completeness and precision (see Tab. 1). 2.2 Internal Visual Evidence Internal representations provide a useful window into what a frozen model encodes before it verbalizes. In language models, logit-lens and tuned-lens analyses show that intermediate hidden states can expose lexical predictions before the final decoding layer [3,7]. Recent VLM studies [24] extend this view to visual tokens, showing that object information can be localized in patch representations and becomes increasingly readable in the vocabulary space across layers. Building on these findings, PatchGate studies the object-level gap between pre-generation visual evidence and final object mentions. Some studies use internal visual evidence for hallucination analysis, detection, grounding, or probing [16,25], while mitigation methods use internal representations for editing hallucination-related directions or steering generation with intermediate-layer logits and token-ranking signals [12,19,27]. These works show that internal states can expose object-level visual support, but such support is typically used as an absolute grounding signal for an object. Absolute support alone does not determine how the caption should be corrected: a supported object may still need to be promoted if it is omitted, while a weakly supported object should matter only when it is over-verbalized. PatchGate therefore compares pre-generation object evidence with object verbalization in the same object coordinate, yielding a signed evidence-verbalization residual whose two directions drive evidence-supported inclusion and evidence-deficient exclusion. 4J. Ko et al. (a)Stagel.VisualEvidenceeXtraction(VEX) Image_ 0~¢='t7SlFI¢M I. CLIP ViT IY ,919LLaVA-7B MLP adapter \\ IY Patche LM l rE22-32 r,(c)=max,qplfc) PCS(c)=max,1P(c).e(c)=emeut I I Vocab readout (Wu). Per-patchevidence (\.a.1. ,-iN#1 _rIll ulwwm. :- , W l urn.:It _I ' 115,"'1 ____ Objectcandidateset: sd.§ 1gIaSS,winows,engine, flag,people,sky... (b)Stage2.Visual-EvidenceInclusion-ExclusionDecoding(VIED) T(I) (+)ESI:Evidence-SupportedInclusion(Oinission) trucks:e=1.0->+5.0windows:e=0.8->+4.1 6+=aesie(e), ae5,:Inclusionstrength (-)EDE:EvidenceDeficieutExclusion(Hallucinations) E-=Yeasrelu(£-med,)(1-PCS) ..cede:exclusionstrength car(PCS:0.754):logo16.69->9.72(A=-6.96,exclusionpenalty) Ima et I -> MLP adapter I oléLLaVA-7B LM -> Decoder Outputlegits/step 'grass'T ) PCS=0 filllyslmmk.dropped eep1f¢l s9l up v.f)"crfon*hr"vi′it('1('11decode .9 Adaptive object selection s. Prompt-freeinternal objectinventory (T(I)) people,...)r raswindowenNeflasky... 1gs,s,gl9g,. l.u 0.0 =0.842,la1gest_gap=0.077 _ _ I rrlh-boon hn1luemniou-supper scenepan I"si;§g5-5§;a no2.%gu Captionw/0PatchGate: "Theimagefeaturesayellowandblacktraincarwitll thewords"ChesieSystem"writtenonit.Thetrainis travelingdownthetracks.anditappearstobea passengertrain.Thetraincarispositionedinthemiddle ofthescene.withafewothercarsvisibleinthe background.Therearetwopeopleintheimage.one standingclosertotheleftsideofthetrainandthe otherpersonstandingfurthertotheright.Thescene alsoincludesacarandatruck.witllthecarlocatedon theleftsideoftheimageandthetruckonthe. Captionw/PatchGate(Ours): "TheyellowandblacktrainengineistravelinguuWu tracks.withwindowsonthefrontandbackofthetrain. Thetrailliscarryingayellowcaboosebehindit.and thereisaflag011theroofofthetrain.Theskyabovethe trainisblue,creatingapleasantbackdropforthe journey." § a .. U § u CoverT12.5-»50.0 CHAIR26.7->0.0 Fig. 2. Overview of PatchGate. (a) VEX reads patch-level lexical evidence from the later LM decoder layers, scores each object candidate by its strongest normalized patch evidence, and selects a prompt-free object setT(I) with an adaptive largest-gap cutoff. (b) VIED calibrates decoding logits withT(I): inclusion promotes evidence-supported but under-verbalized objects, while exclusion suppresses over-committed object words with weak visual evidence. 3 Method PatchGate is a two-stage training-free framework that aligns a frozen VLM’s internal object evidence with its final object verbalization to mitigate object hallucinations and omissions simultaneously. Visual Evidence eXtraction (VEX; Sec. 3.1) first extracts a prompt-free, image-conditioned object inventoryT(I) from internal patch representations before caption generation. Visual-Evidence Inclusion-Exclusion Decoding (VIED; Sec. 3.2) then uses this inventory to calibrate decoding logits in two complementary directions, promoting visually supported but under-verbalized objects and suppressing over-verbalized object words with weak visual evidence. The overview of PatchGate is shown in Fig. 2, with implementation details provided in Appendix A. 3.1 Visual Evidence eXtraction (VEX) VEX constructs an intrinsic object inventoryT(I), a prompt-free, image-conditioned set of visually supported objects, through three steps: patch-to-vocabulary readout, object evidence scoring, and adaptive object selection. Step 1: Patch-to-vocabulary readout. In this step, VEX projects each patch hidden state into the model’s output vocabulary using the logit lens [7] and collects image-specific object candidates from top-1 lexical readouts. Given an imageI, leth (ℓ) p be the hidden state of patchp∈1,...,Nat LM decoder layerℓ. Following prior observations that later-layer predictions are more interpretable and better aligned with final predictions [3], we read from a later layer rangeL. For our default backbone, LLaVA-v1.5-7B [20], the decoder layers are indexed from 1 to 32, and we useL=22,...,32with N= 576 visual patches. (Appendix D ablates this layer range.) For each selected layer and patch, we define the probability assigned to vocabulary token v as q (ℓ) p (v) = softmax W U RMSNorm(h (ℓ) p ) [v], v ∈V.(1) PatchGate5 Table 2. Performance comparison on AMBER generative and discriminative tasks [28]. For Cover, CHAIR, Precision (P.), and Recall (R.), colored values show absolute changes from the LLaVA-v1.5-7B baseline [20] (green: improvement, red: degradation). The best results are in bold. Method Generative Task (Captioning) Discriminative Task (QA) Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ Acc. ↑ P. ↑R. ↑ F1 ↑ [CVPR’24] LLaVA-v1.5-7B49.47.531.43.672.092.562.974.9 + [CVPR’24] OPERA48.8 (-0.6)6.6 (-0.9)28.32.9 76.4 92.7 (+0.2) 69.9 (+7.0) 79.7 + [CVPR’24] VCD51.7 (+2.3) 8.9 (+1.4)39.74.471.191.4 (-1.1) 62.2 (-0.7) 74.0 + [CVPR’25] Devils-in-Mid. 50.3 (+0.9) 3.3 (-4.2)20.0 1.250.5 94.5 (+2.0) 26.8 (-36.1) 41.8 + [ICML’25] MARINE49.9 (+0.5) 6.0 (-1.5)27.32.972.4 94.8 (+2.3) 61.7 (-1.2) 74.7 + [ICLR’25] ProjectAway49.4 (+0.0) 9.2 (+1.7)37.65.070.6 92.5 (+0.0) 60.6 (-2.3) 73.2 + [ICML’26] ILVAD44.9 (-4.5)4.1 (-3.4) 19.01.375.490.5 (-2.0) 70.3 (+7.4) 79.1 + [ICLR’26] SHIELD54.5 (+5.1) 8.5 (+1.0)42.54.064.5 93.5 (+1.0) 50.0 (-12.9) 65.2 + PatchGate (Ours)56.0 (+6.6)6.6 (-0.9)34.33.074.593.8 (+1.3)66.0 (+3.1)77.4 Here,Vis the full language vocabulary, andW U is the frozen LM unembedding matrix applied after RMSNorm [35]. LetV obj denote the WordNet physical-object vocabulary [23]. We retain top-1 readouts that belong to V obj , forming the image-specific candidate set C(I) = arg max v∈V q (ℓ) p (v) : ℓ∈L, p∈1,...,N ∩V obj .(2) Step 2: Object evidence scoring. VEX assigns each object word a scalar visual evidence score for later evidence-verbalization correction. For any object wordc∈V obj , we define its patch-confidence score as the strongest evidence over the selected layers and patches: PCS(c) =max p∈1,...,N max ℓ∈L q (ℓ) p (c).(3) PCS (c) emphasizes localized high-confidence evidence over patch-readout frequency, helping preserve small-region candidates while providing the shared evidence score for inventory selection and VIED calibration (Sec. 3.2). BecausePCSmaximizes over both patches and layers, no single patch or layer is selected: an object receives high evidence if any late layer reads it strongly in any image region. Step 3: Adaptive object selection. The final object inventory is selected with an image-adaptive cutoff. Instead of using a fixed top-k, we setθ TTD (I) as the midpoint of the largest gap between sorted raw PCS values, following TTD [15]. T (I) =c∈C(I) : PCS(c) > max(θ TTD (I),τ floor ).(4) The floorτ floor filters weakly supported candidates when all scores are low, while sorted-tail exclusion prevents outlier-driven thresholds. The exact cutoff procedure is provided in Appendix A.1. After these steps, VEX yieldsT(I) as an intrinsic, prompt-free object inventory extracted before caption generation. It reflects object evidence available inside the frozen VLM and serves as the reference for VIED to compare internal support with final object verbalization. 3.2 Visual-evidence Inclusion-Exclusion Decoding (VIED) VIED uses the VEX-extracted object inventoryT(I) (Sec. 3.1) to calibrate decoding logits according to the evidence-verbalization mismatch. This mismatch has two complementary directions: visually supported objects can be under-verbalized and omitted from the caption, while weakly supported 6J. Ko et al. object words can be over-verbalized and appear as hallucinations. Accordingly, VIED applies two logit- space operations during decoding: Evidence-Supported Inclusion (ESI) promotes objects supported byT(I), and Evidence-Deficient Exclusion (EDE) suppresses object words with high decoding logits despite weak visual evidence. Both operations are applied without modifying the prompt, querying an external model, or updating model parameters. (+) ESI: Evidence-Supported Inclusion. Letg t [v] denote the base logit of tokenvat decoding stept. ESI promotes objects inT(I) that are visually supported but may remain under-verbalized during generation. For each objectc ∈ T(I), we use an extent scoree(c)∈[0,1], defined as the inventory-normalized fraction of patches whose strongest object evidence is assigned to c: ̃e(c) = 1 N N X p=1 1 " c = arg max c ′ ∈V obj max ℓ∈L q (ℓ) p (c ′ ) # , e(c) = ̃e(c) max c ′ ∈T (I) ̃e(c ′ ) . (5) The inner maximum collapses the selected layer range before patch assignment, so each patch casts one object vote and the denominator of ̃e(c) isN, notN|L|. This extent score determines the ESI bonus, which is applied at each decoding step until the corresponding object is generated. (-) EDE: Evidence-Deficient Exclusion. EDE suppresses object words that are strongly favored by the decoder despite weak visual evidence, since such words can surface as object hallucinations. Rather than applying a hard veto to objects outsideT(I), EDE applies a visual-evidence-gated exclusion signal over the object vocabulary. At decoding stept, the median logitmed t =median v∈V g t [v] serves as a neutral reference level. For each object wordc ∈ V obj ,relu(g t [c]− med t ) measures verbalization excess above this neutral level, while 1− PCS(c) measures visual deficit. The exclusion signal multiplies these two factors: ∆ − t (c) = relu(g t [c]− med t ) |z verbalization excess 1− PCS(c) |z visual deficit , c∈V obj .(6) Because∆ − t (c) depends on continuous visual evidence, object words with higherPCSreceive weaker suppression even if they fall outside the final inventory. Thus, EDE reduces reliance on the hard inventory boundary while targeting above-neutral object logits with weak visual evidence. Unified logit update. ESI and EDE form a single logit edit at each decoding step. LetM <t ⊆V obj denote the object words already generated before stept. For every tokenv, VIED appliesˆg t [v] = g t [v] + η t (v), where η t (v) = α esi e(v)1[v /∈M <t ]− γ ede ∆ − t (v), v∈T (I), −γ ede ∆ − t (v),v∈V obj (I), 0,v /∈V obj . (7) Here,α esi andγ ede control the inclusion and exclusion strengths, respectively. Thus, ESI promotes only not-yet-generated inventory objects, EDE applies a visual-evidence-gated penalty to object words, and non-object tokens are unchanged. Together, the edits align pre-generation object evidence with step-wise object verbalization. 4 Experiments 4.1 Experimental Setup Models and baselines. We evaluate PatchGate on the frozen LLaVA-v1.5-7B backbone [20]. Under the same benchmark protocols, we compare against the original LLaVA baseline and representa- PatchGate7 Table 3. Average performance on POPE across random, popular, and adversarial splits. The best results are in bold. MethodAcc. ↑ P. ↑ R. ↑ F1 ↑ [CVPR’24] LLaVA-v1.5-7B82.0 88.6 73.7 80.4 + [CVPR’24] OPERA85.7 86.7 85.2 85.9 + [CVPR’24] VCD83.0 88.0 76.8 81.9 + [CVPR’25] Devils-in-Mid. 55.6 54.3 70.1 61.2 + [ICML’25] MARINE84.4 89.4 78.0 83.3 + [ICLR’25] ProjectAway84.3 81.9 89.2 85.3 + [ICLR’26] SHIELD86.4 85.2 88.6 86.7 + [ICML’26] ILVAD85.6 85.2 86.6 85.7 + PatchGate (Ours)89.890.588.989.6 tive training-free methods, including OPERA [8], VCD [17], Devils-in-Mid. [14], MARINE [36], ProjectAway [12], ILVAD [30], and SHIELD [10]. For ablations, we further test external tag extraction with GroundingDINO [22] and RAM++ [9], and backbone robustness with LLaVA-v1.5-13B [20], Qwen2.5-VL [2], and InstructBLIP [6]. Benchmarks. We evaluate PatchGate on AMBER [28] and POPE [18], which cover complementary object-level reliability settings. AMBER includes free-form captioning and discriminative object- existence question-answering (QA). In captioning, no object query is given, so generated captions are evaluated for visible-object coverage and unsupported object mentions. In AMBER QA and POPE, the queried object is given, and the model answers Yes/No for binary object-existence. Metrics. We use task-specific metrics according to each benchmark output format. For AMBER’s generative captioning task, we follow the official protocol and report Cover, CHAIR, Hal, and Cog using the provided object annotations. We treat Cover and CHAIR [26] as the primary generative metrics because they directly track our two target axes: covering visible objects and avoiding unsupported object mentions. Hal and Cog are complementary diagnostics: Hal is a caption-level any- hallucination indicator, while Cog depends on AMBER’s predefined hallucinatory target object set. For AMBER’s discriminative task and POPE, we report Accuracy (Acc.), Precision (P.), Recall (R.), and F1 for binary object-existence answers. Detailed metric definitions are provided in Appendix A.3. 4.2 Quantitative Results AMBER. Tab. 2 shows that PatchGate is the only method that improves all four highlighted axes over the LLaVA-v1.5-7B baseline: Cover, CHAIR, P., and R. Prior methods often trade off one axis for another: SHIELD improves Cover and P. but worsens CHAIR and R.; OPERA improves CHAIR but reduces Cover; and ILVAD improves CHAIR and R. while degrading Cover and P. PatchGate instead improves all four axes, increasing Cover from 49.4 to 56.0 (+13.4%), reducing CHAIR from 7.5 to 6.6 (−12.0%), and improving P. and R. to 93.8 and 66.0, respectively. This supports our main claim that PatchGate improves caption completeness and precision more consistently across both generative captioning and discriminative QA (See Appendix B for additional baseline evaluations and extended quantitative results). Hal and Cog are not uniformly improved, so PatchGate does not dominate every AMBER diagnostic. Hal depends on caption-level any-hallucination decisions, while Cog depends on AMBER’s predefined target object set. We therefore interpret them as complementary diagnostics and focus the main generative comparison on Cover and CHAIR. POPE. Tab. 3 shows that PatchGate achieves the best average Acc., P., and F1 on POPE. Compared with the LLaVA-v1.5-7B baseline, PatchGate improves Acc. from 82.0 to 89.8 (+9.5%), P. from 88.5 8J. Ko et al. AMBER Generative captioning “Describe this image.” LLaVA-1.5 The image features a brown and white dog standing in a room, looking out of a window with a longing expression. The dog appears to be waiting for someone or something, possibly a family member or a friend. The dog is positioned near the center of the scene, with the window in the background. In the room, there is a chairlocated on the left side, and a couchcan be seen in the background on the right side. The scene gives the impression of a cozy and comfortable living space where the dog enjoys spending time. + PatchGate The dog is standing in a doorway, looking out of the window with a curious expression. The dog appears to be looking at something outside, possibly observing the surroundings or waiting for someone to arrive. The dog is wearing a collar, which adds to its adorable appearance. The doorwayis located in a living or dining setting, with a glass doorvisible in the background. POPE Discriminative QA “Is there a truck?” LLaVA-1.5 Yes ✗ + PatchGate No ✓ “Is there a backpack?” LLaVA-1.5 Yes ✗ + PatchGate No ✓ “Is there a chair?” LLaVA-1.5 No ✗ + PatchGate Yes ✓ Fig. 3. Qualitative results on AMBER [28] and POPE [18]. PatchGate recovers omitted visible objects and enables correct predictions (green), while suppressing unsupported object mentions and incorrect answers (red). to 90.5 (+2.3%), and F1 from 80.4 to 89.6 (+11.4%). It also achieves second-best on R. from 73.7 to 88.9, only 0.3 points below the best result from ProjectAway, while achieving stronger overall balance across the other metrics. This suggests that PatchGate improves binary object-existence reliability without merely biasing the model toward either ”Yes” or ”No”. Split-wise POPE results are provided in Appendix B. 4.3 Qualitative Results Qualitative examples further illustrate PatchGate’s two-sided correction behavior (see Fig. 3). In AMBER generative captioning, PatchGate suppresses hallucinated object mentions while recovering omitted visible objects, improving both caption precision and completeness. In POPE discriminative QA, it corrects incorrect ”Yes” answers to absent objects and incorrect ”No” answers to present objects, rather than introducing a one-sided answer bias. This matches the trends in Tabs. 2 and 3, where PatchGate improves omission-related metrics, Cover and R., together with hallucination-related metrics, CHAIR and P. 4.4 Ablation Study Diagnosing evidence-verbalization mismatch. We test whether pre-decoding VEX evidence (Sec. 3.1) identifies the two mismatch directions targeted by PatchGate. Objects are split by baseline captions. For mentioned objects, unsupported object mentions are ranked by 1− PCS; for unmentioned objects, omitted visible objects are ranked byPCS. Tab. 4 shows that both rankings are well above random: AUROC reaches 0.883/0.890 versus 0.500, and AUPRC reaches 0.707/0.705 versus positive-rate baselines 0.163/0.233, respectively. This indicates that VEX evidence diagnoses both hallucination- and omission-prone cases before decoding. Effect of VEX’s internal object inventory. Tab. 5 compares VEX with external models [9,22] under the same decoding stage, VIED (Sec. 3.2). Against all external-model variants, VEX improves Cover PatchGate9 Table 4. Diagnostic ability of pre-decoding VEX evidence (Sec. 3.1) for hallucination and omission errors. Random baselines: AUROC = 0.5, AUPRC = positive rate. Condition Error target Ranker AUROC ↑ AUPRC ↑ MentionedHallucination Random0.5000.163 MentionedHallucination1− PCS0.8830.707 Unmentioned OmissionRandom0.5000.233 UnmentionedOmissionPCS0.8900.705 Table 5. Performance and computational cost of VEX (Sec. 3.1) and external object inventories. Method MetricsComputational Cost Cover ↑ CHAIR ↓ Param. Peak VRAM LLaVA-v1.5-7B49.47.57.0 B 14.6 GB + GroundingDINO51.37.37.2 B16.2 GB + RAM++ (default)50.97.27.3 B16.9 GB + RAM++ (top-10)50.97.57.3 B16.9 GB + VEX (Ours; Sec. 3.1)56.06.67.0 B14.9 GB by at least 9.2%, from the best external result 51.3 to 56.0, and reduces CHAIR by at least 8.3%, from the best external result 7.2 to 6.6. It achieves these gains while adding no additional parameters and only 0.3GB peak VRAM over the LLaVA-v1.5-7B baseline. This shows that object evidence extracted from the frozen VLM’s internal representations provides a lightweight and more effective inventory for VIED’s logit-space calibration than external object inventories. Effect of VIED components. ESI and EDE target different sides of the evidence-verbalization mismatch. As shown in Tab. 6, ESI alone mainly improves omission-related coverage, raising Cover from 49.4 to 54.8 (+10.9%). EDE alone most directly suppresses hallucination, reducing CHAIR from 7.5 to 6.2 (−17.3%) and also improving Hal and Cog. The full model is not best on hallucination- related diagnostics compared with EDE alone, but achieves the strongest Cover (56.0) while still reducing CHAIR from the baseline 7.5 to 6.6. Thus, ESI and EDE are both needed to improve coverage and precision in a balanced way. Robustness across VLM backbones. PatchGate demonstrates consistent backbone-agnostic effective- ness across various VLM scales and families [2,6,20]. As shown in Tab. 7, it improves Cover by +1.3% to+13.4% and reduces CHAIR by−6.8% to−12.3% across all evaluated backbones. These results suggest that PatchGate does not rely on model-specific behaviors of a single backbone, but rather generalizes well across backbone scales and VLM families without fine-tuning or external detectors. Efficiency. PatchGate is training-free and detector-free, but adds a small overhead from one VEX forward pass and lightweight VIED logit edits. On LLaVA-v1.5-7B, runtime increases from 2.92s to 3.30s per image (1.13×), and peak VRAM increases from 14.6GB to 14.9GB (1.02×). Thus, PatchGate mitigates both object hallucinations and omissions with modest overhead, while avoiding external visual modules, additional VLM queries, or model updates (see Appendix D for detailed efficiency results). Limitations. PatchGate depends on the quality of the internal object inventory. When VEX misses visible objects or assigns unreliable evidence, VIED has limited basis for correct inclusion or exclusion. PatchGate is also limited to object-level logit calibration and does not directly address attribute, 10J. Ko et al. Table 6. Effect of VIED components (Sec. 3.2) on AMBER. Method VIEDMetrics ESI EDE Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ LLaVA-v1.5-7B✗49.47.531.43.6 + PatchGate✓✗54.86.734.63.9 + PatchGate✗✓50.36.2 29.5 2.8 + PatchGate✓56.06.634.33.0 Table 7. Robustness of PatchGate across VLM backbones. MethodCover ↑ CHAIR ↓ Hal ↓ Cog ↓ LLaVA-v1.5-7B49.47.531.43.6 + PatchGate56.06.634.33.0 LLaVA-v1.5-13B50.86.530.83.1 + PatchGate52.85.725.72.2 Qwen2.5-VL-7B58.84.127.21.7 + PatchGate60.23.621.81.3 InstructBLIP-7B53.87.436.03.9 + PatchGate54.56.935.43.9 relation, counting, or higher-level scene hallucinations. In some cases, this object-focused calibration may reduce descriptive attributes or modifiers compared with the baseline, affecting caption detail. Further discussion is provided in Appendix D. 5 Conclusion We presented PatchGate, a training-free method for mitigating both object hallucinations and omissions in frozen Vision-Language Models. Unlike prior training-free methods that mainly intervene on object words likely to emerge during generation, PatchGate first extracts prompt-free object evidence from internal patch representations and forms an intrinsic object inventory before decoding. It then uses this inventory for two-sided logit calibration, promoting visually supported but under- verbalized objects and suppressing weakly supported object mentions. Experiments on AMBER and POPE show that PatchGate improves object-level completeness and precision across both captioning and object-existence QA. These results suggest that a frozen VLM’s own patch-level evidence can serve as a practical grounding signal for more reliable object verbalization without external detectors or fine-tuning. Acknowledgements This work was partly supported by the KHIDI grant funded by the Korean government (MOHW) [No.RS-2025-02307233, No.RS-2026-25613012], the NRF or IITP grants funded by the Korean government (MSIT) [No.05-26-04-0094, No.RS-2026-25472075, No.RS-2025-02305581, No.RS-2025- 25442338, and No.RS-2021-I211343], the ITIP grant funded by the Korean government (MOTIR) [No.RS-2026-25549946], the Research grant from SNU, and the Strategic Hub grant for International Research Collaboration of SNU. Kyungsu Kim is affiliated with the School of Transdisciplinary Innovations, Department of Biomedical Science, Interdisciplinary Program in Artificial Intelligence (IPAI), Medical Research Center, and AI Institute at SNU. PatchGate11 Appendix Overview. This supplementary material provides methodological details and additional analyses supporting the main paper. Appendix A details the methodology, algorithms, hyperparameters, and evaluation settings of PatchGate. Additional quantitative and qualitative results, including extended baseline comparisons, are provided in Appendices B and C, respectively. Appendix D examines key design choices, evidence aggregation, object-inventory sources, linguistic quality, and computational efficiency. Finally, Appendix E discusses the limitations of the current object-level formulation and directions for future work. A Implementation Details A.1 Methodological Details VEX: Prompt-Free Object Inventory Image-only forward pass. PatchGate first applies VEX (Sec. 3.1) to read internal patch representations and extract an explicit image-conditioned object set before decoding. VEX requires only one additional forward pass and neither uses a task-specific input prompt nor decodes any output tokens. This differs from prior training-free methods [8,14,17,36] that primarily derive intervention signals from object tokens likely to emerge during generation under a given prompt. Because VEX is independent of both the downstream prompt and the model’s emerging output tokens, it provides an image-specific signal that is available before generation begins. The benchmark prompts provided by AMBER [28] and POPE [18] are introduced only after the internal object inventory has been constructed, when VIED generates the final caption or binary answer. We therefore refer toT(I) as a prompt-free, image-conditioned object inventory extracted before generation. Object vocabulary. We construct the object vocabularyV obj from WordNet [23] and use the same vocabulary in both VEX (Sec. 3.1) and VIED (Sec. 3.2). The construction is described in four parts: (i) semantic class mapping from WordNet and part-of-speech tags, (i) sense-frequency filtering after morphological normalization, (i) broad vocabulary coverage followed by image-specific evidence selection, and (iv) tokenizer mapping that keeps only words represented by a single vocabulary token. (i) Semantic class mapping. Every WordNet synset carries a lexname, one of its coarse supersenses (26 for nouns). Tab. A lists the mapping from lexname and POS to the five internal classes. A word’s noun, adjective, and verb senses are assigned to object, scene, attribute, action, or body according to this mapping. Table A. WordNet supersense (lexname) and part-of-speech (POS) mapping used to constructV obj . Only object and scene classes are included. POS WordNet lexname(s)Class Included noun noun.artifact, noun.animal, noun.food, noun.person, noun.plantobject✓ noun noun.location, noun.object, noun.substancescene✓ noun noun.act, noun.eventaction✗ noun noun.bodybody✗ nounall other noun.*—✗ adj. (a/s) all adjective synsetsattribute✗ verb (v) all verb synsetsaction✗ (i) Sense-frequency filtering. The construction uses standard WordNet resources together with SemCor sense frequencies. Each wordwis lowercased and lemmatized ton=lemma(w, n); we 12J. Ko et al. then accumulate the SemCor occurrence count of each sense associated with the normalized form into the per-class tallycnt k (w) over the five classesK=object, scene, attribute, action, body. The word enters V obj only when its frequency-dominant class is object or scene (see Tab. A). (i) Broad vocabulary coverage.V obj is deliberately not required to be a precise object lexicon: it only has to be a superset of the nameable object and scene words that could be read off a patch. The downstream VEX stages perform the actual image-specific selection. Thus, a word that is over-included inV obj but carries little visual evidence receives a lowPCSand is unlikely to enter T(I). If its decoding logit later becomes high, the lowPCSalso produces a stronger EDE penalty during VIED. (iv) Tokenizer mapping. Finally, we intersect the WordNet-derived object and scene lexicon with the VLM backbone’s own subword vocabulary. We iterate over every token ID in the tokenizer, decode it to its surface string, and retain the token if the string is alphabetic, contains at least three characters, and passesPrimaryCat. Because the enumeration visits one ID at a time, every retained entry is by construction a single token, so its probability is exactly the logit-lens softmax coordinate at that ID. The output is a fixed ID-to-word table definingV obj . For LLaVA-1.5-7B’s 32k-token vocabulary [20], this procedure yields 1,560 object and scene tokens, which are reused unchanged across all images. Object evidence scoring. Although the image-specific candidate setC(I) contains only top-1 vocabulary readouts belonging toV obj , VEX computes the patch-confidence scorePCS(c) for every objectc∈V obj as the maximum full-vocabulary softmax probability over all patch-layer pairs in the selected late- layer range. Thus,C(I) determines which objects are eligible to enter the final inventoryT(I), while the continuous PCS values are retained for evidence-dependent exclusion over the entire object vocabulary during VIED (Sec. 3.2). We use maximum aggregation to preserve strong localized evidence that may appear within a small image region, since mean aggregation can dilute such evidence when most patches are unrelated to a given object, while relying on a single fixed layer can be brittle when object information becomes lexically readable at different late layers. We analyze these aggregation choices in Sec. D. Adaptive object selection. After computing the PCS values, VEX constructs the final object inventory T(I) using an image-adaptive threshold motivated by the largest-score-gap selection strategy of TTD [15] (see Fig. A). Candidates inC(I) with scores below the evidence floorτ floor are first removed because their visual support is too weak to justify inclusion inT(I), preventing low-confidence objects from entering the inventory solely due to relative gaps among low PCS values. Letc 1 ,...,c m denote the remaining candidates ordered by descending PCS, with corresponding scoress 1 ≥ s 2 ≥·≥ s m . We measure the decrease between each pair of adjacent scores as δ i = s i − s i+1 , i = 1,...,m− 1.(A) A largeδ i indicates a separation between the higher-scoring candidates above positioniand the lower-scoring candidates below it. Form≥3, we exclude the final gapδ m−1 , which compares the two lowest-scoring candidates, because a single unusually weak candidate can make this gap large without representing the main separation in the score distribution. Among the remaining gaps, we choose the largest one and define the adaptive threshold as k ⋆ = arg max 1≤i≤m−2 δ i , θ TTD (I) = s k ⋆ + s k ⋆ +1 2 .(B) When multiple positions have the same gap, the smallest index is selected. The final object inventory is then defined as T (I) =c∈C(I)| PCS(c)≥ τ floor , PCS(c) > θ TTD (I).(C) Ifm= 0, the inventory is empty. Ifm= 1, the sole candidate is included, while form= 2, the single available gap is used without tail-gap exclusion. PatchGate13 Algorithm A VEX: prompt-free object inventory T (I) Require:imageI(no text prompt), frozen VLM with unembeddingW U andRMSNorm, later-layer setL, object vocabulary V obj , floor τ floor 1: h (ℓ) p ← VLM.forward(I)▷ image only, single forward pass, no decoding 2: // Step 1: patch-to-vocabulary readout 3: for each patch p∈1,...,N, layer ℓ∈L do 4: q (ℓ) p (v)← softmax W U RMSNorm(h (ℓ) p ) [v]▷ logit lens 5: end for 6: C(I)←arg max v q (ℓ) p (v) : ℓ∈L, p∩V obj 7: // Step 2: object evidence scoring 8: PCS(c)← max p max ℓ∈L q (ℓ) p (c)for each c∈V obj 9: // Step 3: adaptive object selection 10: T (I)←c∈C(I) : PCS(c)≥ τ floor , PCS(c) > θ TTD (I)▷ largest-gap cut, Eq. B 11: extent e(c)∈ [0, 1] for c∈T (I) by Eq. 5 12: return inventory T (I), evidence PCS(·), extent e(·) VIED: Task-Specific Decoding Task-specific inventory usage. Once the object inventoryT(I) has been constructed, PatchGate adapts how it is used according to the output space. For open-ended captioning, ESI and EDE directly adjust object-token logits (Eqs. 5 and 6). Binary QA, by contrast, operates in the answer-token space, using inventory-conditioned classifier-free guidance to adjust the relative logits of “Yes” and “No”. This task-specific formulation allows the same object inventory to support both free-form generation and binary decisions. Open-ended captioning. The PCS and extent values produced by VEX are precomputed once and reused throughout decoding. At each generation stept, ESI and EDE are computed with respect to the same unmodified base-logit vectorg t , and their corrections are added simultaneously before token selection. After each generated token, VIED updates the object-mention setM <t according to the object-realization rule. Once an object in inventory is marked as realized, its ESI bonus (Eq. 5) is disabled for subsequent steps, while EDE penalty (Eq. 6) remains active and the frozen decoder may still generate the object again through its original logits. Binary QA. Because binary QA expresses its final decision through the “Yes” and “No” tokens, which are not directly affected by the object-token updates of ESI and EDE, we adapt the image-grounded classifier-free guidance formulation of MARINE [36]. Given the original image and benchmark query, the unconditioned inference produces the logit vectorg uncond t , whereas the conditioned inference additionally incorporates the object inventoryT(I) as image-grounded context and producesg cond t . At the answer step t, these two signals are combined to obtain the guided logit vector g QA t = λ QA g cond t + (1− λ QA )g uncond t .(D) Here, the QA guidance strengthλ QA ∈[0,1] controls the contribution of the inventory-conditioned logits relative to the unconditioned logits, and the final answer is selected by comparing the “Yes” and “No” entries ofg QA t . This enables bidirectional correction: the conditioned signal can raise the relative “Yes” logit when the unconditioned inference incorrectly answers “No” for a visually supported object, while shifting the decision toward “No” when the unconditioned inference produces a hallucinated affirmative answer for an unsupported object. A.2 PatchGate Algorithm For reproducibility, Algorithms A and B summarize the complete inference procedure of PatchGate, while the full implementation code and configuration files are provided in a separately submitted code archive. 14J. Ko et al. (a) Stage 1: VEX —prompt-free object inventory 24 ×24 patch grid logit-lens (frozen LM head) PCS bed 1.00 lamp 1.00 chair 1.00 table 0.99 frame 0.97 wall 0.94 floor 0.93 picture 0.85 recovered object θ TTD = 0.57 → | T(I) | = 24 cup ~0.00 couch ~0.00 no patch evidence → excluded (b) Stage 2: VIED —logit edit EDE (suppress) 6.8 g t 4.4 ĝ t − 2.4 − γ · relu( g − med ) · ( 1 − PCS ) = − 0.5 ×4.9 ×1.0 = − 2.4 cup, PCS ~ 0 → not emitted ESI (recover) 16.1 g t 17.2 ĝ t + 1.1 + α · e = + 8 ×0.14 = + 1.1 lamp, PCS 1.0, extent 0.14 → emitted (c) Result: corrected caption LLaVA-v1.5-7B (baseline): ... a bed... a couchin the background. Two cups... a book... a remote control... + PatchGate(Ours): ... a bed with a wooden headboardand a lampon a table nearby... Fig. A. End-to-end PatchGate inference for open-ended captioning. VEX (Sec. 3.1) selects the largest-gap cutoff atθ TTD = 0.57, producing an internal object inventory that retains supported objects (bed, lamp, chair, and table) while excluding unsupported objects (cup and couch). During VIED (Sec. 3.2), EDE reduces the cup logit from 6.8 to 4.3, whereas ESI increases the omitted lamp logit from 16.1 to 17.2, causing it to be generated. Algorithm B VIED: inclusion/exclusion decoding Require:imageI, promptx, inventoryT(I), evidencePCS(·), extente(·), object vocabularyV obj , strengths α esi , γ ede , maximum output length T max 1: y ← [ ]; M←∅▷ output tokens; realized inventory objects 2: repeat 3: g t ← VLM(x,I,y);med t ← median v∈V g t [v]▷ base logits, neutral level 4:ˆg t ← g t 5: for each object c∈V obj do 6:ˆg t [c]← ˆg t [c]− γ ede relu g t [c]− med t 1− PCS(c) ▷ EDE 7: end for 8: for each c∈T (I) with c /∈M do 9:ˆg t [c]← ˆg t [c] + α esi e(c)▷ ESI (fix-once) 10: end for 11: v t ← arg max v∈V ˆg t [v]; y.append(v t ) 12: if v t realizes an object c∈T (I) then M←M∪c 13: until v t = EOS or |y| = T max 14: return generated caption y VEX. Stage 1 of PatchGate, VEX (Sec. 3.1), extracts the prompt-free object inventoryT(I) and computes the PCS and extent scores through a single image-only forward pass without any task-specific prompt, as summarized in Algorithm A. VIED. Stage 2 of PatchGate, VIED (Sec. 3.2), uses the extracted inventory, PCS, and extent scores during greedy decoding to suppress unsupported object mentions through EDE and recover omitted visible objects through ESI, as summarized in Algorithm B. Worked example: open-ended captioning (Fig. A). To provide a more concrete understanding of PatchGate, we trace its end-to-end inference procedure on AMBER image 209 [28], which depicts a hotel room whose base caption generated by LLaVA-1.5-7B [20] hallucinates acouch, twocups, a book, and aremote control, while omitting the visiblelampandheadboard(see Fig. A). All reported values are the actual PCS, extent scores, and decoding logits, withL=22,...,32,α esi = 8, and γ ede = 0.5. (i) VEX (Alg. A). PatchGate first projects every visual patch into the vocabulary space through the logit lens [7] (VEX Step 1; Sec. 3.1) and computes the patch-confidence scorePCS(c) for every PatchGate15 (a) Image & Question AMBER_867 Question: “Is there a ship in this image?” Ground Truth: “No.” (ship absent) (b) Stage 1 —VEX object inventory Object Inventory T(I) PCS ✓grass1.00 ✓sky1.00 ✓rail0.99 ✓trees0.99 ✓woman0.96 ✓water0.91 ✓cloud0.91 ship: PCS 0.00 → NOT in T(I) (c) Stage 2 —Classifier-free guidance overYes / No logits Hint = “The image contains: grass, water, woman, rail, trees, sky, ...” (no ship) Logits P(Yes) Unconditioned (base) Yes27.86 No26.50 0.80 Yes✗ Conditioned on T(I) Yes25.22 No26.12 0.29 No Guided ℓ = 0.7·Cond + 0.3·Uncond Yes26.01 No26.23 0.44 No✓ LLaVA-v1.5-7B (baseline): + PatchGate (Ours): Fig. B. End-to-end PatchGate inference for binary QA. PatchGate constructs the object inventory T(I), in which ship is absent due to its negligible visual support (PCS≈0.00). While the unconditioned LLaVA-v1.5-7B logits yield a hallucinated “Yes” answer, inventory-conditioned guidance raises the “No” logit above “Yes”, yielding the correct answer “No”. objectc∈V obj from the strongest response across the selected patches and layers (VEX Step 2). The room’s visually supported objects form a high-evidence cluster, includingbed(1.00),lamp(1.00), chair(1.00), andtable(0.99), whereas the hallucinatedcuphasPCS(cup) = 0.00, and thecouch receives approximately zero patch evidence. PatchGate then selects the largest eligible interior gap atθ TTD = 0.57 (VEX Step 3), yielding an inventory of size|T(I)|= 24: thelampis retained with extent e(lamp) = 0.14, whereas the cupand couchfall below the threshold and are excluded. (i) VIED (Alg. B). When the language model assigns a high logit to the visually unsupported cup, its base logit reachesg t [cup] = 6.8, above the neutral medianmed t = 1.9. BecausePCS(cup)≈0, PatchGate applies EDE (Sec. 3.2) and subtractsγ ede relu(g t [cup]− med t )(1− PCS(cup)) = 0.5·4.9· 1.0≈2.5, thereby reducing thecuplogit toˆg t [cup] = 6.8−2.5≈4.3, which prevents it from being generated, while the same mechanism suppresses the other unsupported objects. For the visually supported but omittedlamp, PatchGate applies ESI (Sec. 3.2) and addsα esi e(lamp) = 8·0.14 = 1.12≈1.1 at every decoding step until it is generated, increasing its logit toˆg t [lamp] = 16.1 + 1.12 = 17.22≈17.2, so thelampbecomes the argmax token. The resulting caption excludes the hallucinated objects and mentions the visible lampand headboard. Worked example: binary QA (Fig. B). On an AMBER example asking whether ashipis present, the queried object receives negligible visual support (PCS ≈0.00) and is therefore absent from the resulting inventoryT(I), which instead contains visually supported objects such assky,trees, andwater. Nevertheless, the unconditioned LLaVA-v1.5-7B logits favor “Yes” over “No” (27.86 vs. 26.50), yielding a hallucinated affirmative answer. Conditioning onT(I) reverses this preference, producing “Yes” and “No” logits of 25.22 and 26.12, respectively. Withλ QA = 0.7, these conditioned logits are combined with the unconditioned logits as“Yes” logit= 0.7×25.22 + 0.3×27.86≈26.01 and“No” logit= 0.7×26.12 + 0.3×26.50≈26.23. Consequently, the “No” logit becomes higher than the “Yes” logit, allowing PatchGate to produce the correct answer “No”. A.3 Experimental Details Tab. B summarizes the model, main hyperparameters, decoding settings, and computational environ- ment used in our experiments. Additional details on baseline implementations, prompt protocols, binary QA decoding, and benchmark-specific evaluation are provided below. Baselines. We compare PatchGate with the frozen LLaVA-v1.5-7B backbone [20] and representative training-free methods: OPERA [8], VCD [17], Devils-in-Mid. [14], MARINE [36], ProjectAway [12], 16J. Ko et al. Table B. Main hyperparameters and implementation settings used for PatchGate. SettingValue PatchGate setup BackboneLLaVA-v1.5-7B HuggingFace model ID llava-hf/llava-1.5-7b-hf VEX layer range L 22,..., 32 Evidence floor τ floor 0.02 ESI strength α esi 8 EDE strength γ ede 0.5 QA guidance strength λ QA 0.7 DecodingGreedy do sample False Random seedNot used (deterministic) TemperatureNot used Sampling-based baselines Random seed42 Temperature1.0 Hardware and software GPU1 NVIDIA RTX A6000 48GB Python3.10.12 PyTorch2.1.2 Transformers4.45.2 ILVAD [30], SHIELD [10], and PND [13]. We use the official implementations whenever available; because PND had no public implementation at the time of our experiments, we reimplemented it following the paper. All methods are evaluated under the same benchmark protocols while preserving their method-specific inference procedures. Benchmarks. We evaluate PatchGate on two object-hallucination benchmarks, AMBER [28] and POPE [18]. AMBER provides a generative task with 1,004 free-form image-description samples and a discriminative binary QA task covering object existence, attribute, and relation hallucinations. Its object-existence subset asks only about plausible but absent objects, and therefore all ground- truth answers are “No.” This subset evaluates absent-object rejection but does not measure the complementary ability to affirm visible objects. In contrast, POPE contains both present-object questions with “Yes” answers and absent-object questions with “No” answers under random, popular, and adversarial negative-object sampling. It therefore evaluates both visible-object acceptance and absent-object rejection. For both benchmarks, we follow the official evaluation protocols and use the provided annotations and evaluation scripts. Metrics. We evaluate generative captioning and binary QA using the official metrics provided by AMBER [28] and POPE [18]. Generative captioning. For a generated responseR, the official AMBER evaluator first extracts its nouns to obtainR obj , and then retains only those included in the complete AMBER object listX obj , yieldingR ′ obj =R obj ∩ X obj . LetA obj denote the visible objects annotated for the image andH obj denote its image-specific hallucinatory target objects. Following the official AMBER PatchGate17 Annotation inconsistency yields false Hal/Cog penalties Visible object (Sun) annotated as absent Fig. C. Limitations of Hal and Cog. Although the sun is visibly present, AMBER annotates it as absent and includes it in the image-specific hallucinatory target set. PatchGate correctly recovers sun, but this mention is penalized by both Hal and Cog. definitions [28], the four generative metrics are Cover(R) = R ′ obj ∩ A obj |A obj | , CHAIR(R) = 1− R ′ obj ∩ A obj R ′ obj , Hal(R) = 1[CHAIR(R) > 0], Cog(R) = R ′ obj ∩ H obj R ′ obj . (E) Cover measures the proportion of annotated visible objects mentioned in the response, while CHAIR measures the proportion of generated object mentions that are unsupported by the image. Hal indicates whether a response contains at least one unsupported object mention, whereas Cog measures the proportion of AMBER’s predefined, image-specific hallucinatory targets generated in the response. The four metrics are reported over all 1,004 generative samples. Higher Cover is better, whereas lower CHAIR, Hal, and Cog are better. Limitations of Hal and Cog. Although Hal and Cog provide complementary hallucination diagnostics, they are sensitive to the coverage and consistency of AMBER’s object annotations. Hal is a caption-level binary metric that marks an entire response as hallucinated when even one extracted object mention is not matched to the annotated visible-object set. Thus, a single annotation mismatch can change Hal from 0 to 1, while affecting the mention-level CHAIR score more gradually. Cog additionally depends on AMBER’s predefined, image-specific hallucinatory target set. As shown in Fig. C, the clearly visible sun is absent from the visible-object annotations and instead included in the hallucinatory target set. Consequently, PatchGate’s visually grounded recovery of sun is penalized by both Hal and Cog. Such an annotation mismatch is not directly penalized by Cover, 18J. Ko et al. but the metric cannot credit the recovered object; CHAIR instead counts the unmatched mention as unsupported, although its effect remains proportional rather than caption-level. We therefore use Cover and CHAIR as the primary complementary measures of object coverage and mention-level hallucination, while interpreting Hal and Cog as additional diagnostics that can be more sensitive to annotation inconsistencies. Discriminative QA. For AMBER discriminative QA and POPE, we report Accuracy (Acc.), Precision (P.), Recall (R.), and F1, computed as follows: Acc. = TP + TN TP + TN + FP + FN , P. = TP TP + FP , R. = TP TP + FN , F1 = 2 P× R P + R . (F) The two benchmarks follow different positive-label conventions. Following the official AMBER evaluator, “No” is treated as the positive label. Its object-existence subset contains only questions about absent objects with the ground-truth answer “No”, and therefore isolates absent-object rejection. The full discriminative results in Tab. C, however, also include attribute and relation questions. Under this convention, Precision measures the proportion of predicted “No” answers that are correct, while Recall measures the proportion of ground-truth “No” cases correctly answered “No”. In contrast, the official POPE evaluator [18] treats “Yes” as the positive label. A false positive is an absent-object question incorrectly answered “Yes”, corresponding to hallucination, while a false negative is a visible-object question incorrectly answered “No”, corresponding to omission. Precision therefore reflects resistance to hallucinated positive answers, whereas Recall reflects omission recovery. Higher values indicate better performance for all four metrics. Prompt and inference protocols. VEX receives only the imageIand performs an additional image-only forward pass without textual input or output-token decoding. Task-specific prompts are introduced only during downstream inference. Generative captioning. For AMBER generative captioning, we use the official prompt, “De- scribe this image.” [28] For this open-ended task, VIED directly applies the ESI and EDE object-token logit updates described in Sec. 3.2 and Algorithm B throughout greedy decoding. Binary QA. For the AMBER discriminative task, we preserve each query from the official JSON file without modifying its wording. The queries cover object existence, attributes, and relations, such as “Is there a cloud in this image?”, “Is the sky sunny in this image?”, and “Is there direct contact between the person and grass?”, respectively. For POPE, we use “Is there aobjectin the image? Please answer this question with one word.” The unconditioned inference uses the original image and benchmark query, whereas the conditioned inference additionally receives the image-specific inventory as the fixed context “The image contains:object 1 , object 2 , . . ..” (see Fig. B). The final prediction is restricted to the “Yes” and “No” answer tokens. Runs and environment. We report a single inference run for each method and benchmark. PatchGate uses deterministic greedy decoding withdosample=False, without temperature,topp,numbeams, or a random seed. For stochastic baselines, we use a fixed seed of 42 and temperature 1.0. All experiments run on one NVIDIA RTX A6000 GPU with 48GB memory using Python 3.10.12, PyTorch 2.1.2, and Transformers 4.45.2. Code archive. The submitted code archive includes the complete PatchGate inference procedures for generative captioning and binary QA, together with configuration files and aREADME.mdcontaining setup and execution instructions. PatchGate19 Table C. Full AMBER results [28] across different VLM backbones. We include all training-free baselines evaluated on each backbone. For Cover, CHAIR, P., and R., colored values show absolute changes from the corresponding backbone baseline (green: improvement, red: degradation). The best results within each backbone are in bold. Method Generative TaskDiscriminative Task Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ Acc. ↑ P. ↑R. ↑ F1 ↑ [CVPR’24] LLaVA-v1.5-7B [CVPR’24] LLaVA-v1.5-7B49.47.531.43.672.092.562.974.9 + [CVPR’24] OPERA48.8 (-0.6)6.6 (-0.9)28.32.9 76.4 92.7 (+0.2) 69.9 (+7.0) 79.7 + [CVPR’24] VCD51.7 (+2.3) 8.9 (+1.4)39.74.471.191.4 (-1.1)62.2 (-0.7) 74.0 + [CVPR’25] Devils-in-Mid. 50.3 (+0.9) 3.3 (-4.2)20.0 1.250.5 94.5 (+2.0) 26.8 (-36.1) 41.8 + [ICML’25] MARINE49.9 (+0.5) 6.0 (-1.5)27.32.972.4 94.8 (+2.3) 61.7 (-1.2) 74.7 + [ICLR’25] ProjectAway49.4 (+0.0) 9.2 (+1.7)37.65.070.6 92.5 (+0.0) 60.6 (-2.3) 73.2 + [ICML’26] ILVAD44.9 (-4.5)4.1 (-3.4) 19.01.375.490.5 (-2.0) 70.3 (+7.4) 79.1 + [ICLR’26] SHIELD54.5 (+5.1) 8.5 (+1.0)42.54.064.5 93.5 (+1.0) 50.0 (-12.9) 65.2 + [CVPR’26] PND51.7 (+2.3) 7.3 (-0.2)34.94.071.4 93.2 (+0.7) 61.3 (-1.6) 74.0 + PatchGate (Ours)56.0 (+6.6)6.6 (-0.9)34.33.074.593.8 (+1.3)66.0 (+3.1)77.4 [CVPR’24] LLaVA-v1.5-13B [CVPR’24] LLaVA-v1.5-13B50.86.530.83.171.596.059.573.5 + [CVPR’24] OPERA49.0 (-1.8)6.1 (-0.4)28.42.773.893.1 (-2.9) 63.9 (+4.4) 75.8 + [CVPR’24] VCD51.5 (+0.7) 8.0 (+1.5)35.23.7 82.988.0 (-8.0) 86.0 (+26.5) 87.0 + [CVPR’25] Devils-in-Mid. 46.3 (-4.5) 4.3 (-2.2) 18.9 1.263.395.3 (-0.7) 46.9 (-12.6) 62.9 + [ICML’25] MARINE50.0 (-0.8)5.0 (-1.5)23.92.370.9 96.7 (+0.7) 58.0 (-1.5) 72.5 + [ICLR’25] ProjectAway49.5 (-1.3) 8.6 (+2.1)35.84.670.292.3 (-3.7) 60.1 (+0.6) 72.8 + [ICML’26] ILVAD50.9 (+0.1) 5.1 (-1.4)27.02.274.290.8 (-5.2) 64.3 (+4.8) 75.3 + [ICLR’26] SHIELD51.3 (+0.5) 9.0 (+2.5)41.44.464.893.2 (-2.8)50.4 (-9.1) 65.4 + [CVPR’26] PND52.1 (+1.3) 7.0 (+0.5)33.23.772.393.0 (-3.0) 62.1 (+2.6) 74.5 + PatchGate (Ours)52.8 (+2.0)5.7 (-0.8)25.72.271.796.6 (+0.6)59.3 (-0.2)73.5 [arXiv’25] Qwen2.5-VL-7B [arXiv’25] Qwen2.5-VL-7B58.84.127.21.784.992.983.688.0 + [CVPR’24] OPERA57.2 (-1.6) 4.4 (+0.3)27.91.685.192.4 (-0.5) 84.0 (+0.4) 88.0 + [CVPR’24] VCD53.7 (-5.1) 6.3 (+2.2)34.61.881.283.9 (-9.0) 88.7 (+5.1) 86.2 + [CVPR’25] Devils-in-Mid. 54.8 (-4.0) 4.2 (+0.1)17.51.084.9 93.7 (+0.8) 82.8 (-0.8) 87.9 + [ICML’25] MARINE62.4 (+3.6) 4.4 (+0.3)25.81.7 89.092.1 (-0.8) 91.3 (+7.7) 91.7 + [ICLR’25] ProjectAway55.5 (-3.3) 4.9 (+0.8)31.81.884.085.5 (-7.4) 88.0 (+4.4) 86.7 + [ICML’26] ILVAD35.9 (-22.9) 3.1 (-1.0) 7.2 0.382.591.3 (-1.6)81.4 (-2.2) 86.1 + [ICLR’26] SHIELD53.1 (-5.7) 5.8 (+1.7)33.31.286.290.3 (-2.6) 85.9 (+2.3) 88.0 + [CVPR’26] PND55.3 (-3.5) 5.0 (+0.9)29.81.683.490.1 (-2.8) 84.2 (+0.6) 87.0 + PatchGate (Ours)60.2 (+1.4)3.6 (-0.5)21.81.385.484.7 (-8.2)95.2 (+11.6)89.6 20J. Ko et al. Table D. Performance comparison on the POPE benchmark [18]. We report results across the random, popular, and adversarial splits. The best results are in bold, and the second-best results are underlined. Method RandomPopularAdversarial Acc. ↑ P. ↑ R. ↑ F1 ↑ Acc. ↑ P. ↑ R. ↑ F1 ↑ Acc. ↑ P. ↑ R. ↑ F1 ↑ [CVPR’24] LLaVA-v1.5-7B83.8 92.4 73.7 82.082.6 89.7 73.7 80.979.7 83.8 73.7 78.4 + [CVPR’24] OPERA89.2 93.4 85.2 89.186.8 88.0 85.2 86.681.2 78.8 85.2 81.9 + [CVPR’24] VCD85.4 92.8 76.8 84.083.3 88.3 76.8 82.280.4 82.8 76.8 79.7 + [CVPR’25] Devils-in-Mid. 56.0 54.6 70.3 61.556.2 54.8 70.9 61.854.5 53.5 69.1 60.3 + [ICML’25] MARINE86.5 91.0 79.0 84.684.5 89.5 78.0 83.482.2 87.7 77.0 82.0 + [ICLR’25] ProjectAway89.0 89.4 89.2 89.385.0 82.3 89.2 85.679.0 74.0 89.2 80.9 + [ICML’26] ILVAD89.2 91.4 86.6 88.986.6 86.6 86.6 86.680.9 77.7 86.6 81.9 + [ICLR’26] SHIELD90.4 91.9 88.6 90.287.0 85.9 88.6 87.281.7 77.9 88.682.9 + [CVPR’26] PND88.6 96.3 80.3 87.687.293.1 80.3 86.285.188.8 80.3 84.3 + PatchGate (Ours)92.594.889.992.389.990.789.089.887.086.387.987.1 B Additional Quantitative Results B.1 Additional AMBER Results Overall performance across backbones. Tab. C reports the full AMBER [28] results on the main LLaVA-v1.5-7B backbone [20], the larger LLaVA-v1.5-13B, and Qwen2.5-VL-7B [2] from a different model family. On LLaVA-v1.5-7B, PatchGate improves Cover from 49.4 to 56.0 (+13.4%) while reducing CHAIR from 7.5 to 6.6 (−12.0%). The same trend appears on LLaVA-v1.5-13B and Qwen2.5-VL-7B, demonstrating consistent mitigation of omission and hallucination across model scales and families. AMBER discriminative QA is less directly aligned with PatchGate’s primary objective: its object-existence subset evaluates only absent-object rejection, while its attribute and relation subsets require evidence beyond object presence. Despite this mismatch, PatchGate remains competitive, ranking third in Accuracy, Precision, Recall, and F1 on LLaVA-v1.5-7B. This suggests that PatchGate’s image-specific object evidence remains useful for binary QA, even when the task extends beyond its primary object-verbalization objective. Balanced improvement without metric trade-offs. Existing methods often improve one aspect of object reliability at the expense of its complementary objective. For example, ILVAD [30] substantially reduces CHAIR but also decreases Cover, while Devils-in-Mid. [14] and MARINE [36] improve Precision at the cost of Recall. Conversely, SHIELD [10] increases Cover but worsens CHAIR and substantially reduces Recall. In contrast, PatchGate is the only evaluated method on LLaVA-v1.5-7B that simultaneously improves Cover, Precision, and Recall while reducing CHAIR relative to the frozen backbone. This demonstrates balanced gains across generative object coverage, mention-level hallucination, and discriminative QA, while the POPE results provide direct evidence of bidirectional correction for present and absent objects. Aggregate improvement across primary metrics. To assess balanced improvement without overem- phasizing any single metric, we conduct a simple post-hoc analysis that equally sums the percentage- point changes in Cover, CHAIR, Precision, and Recall, counting a reduction in CHAIR as posi- tive (see the red and green highlights in Tab. C). PatchGate achieves the largest aggregate gain of 6.6 + 0.9 + 1.3 + 3.1 = 11.9, while the second-highest gain is obtained by OPERA [8] with −0.6 + 0.9 + 0.2 + 7.0 = 7.5. This comparison further highlights that PatchGate improves the four primary measures jointly without trading off one objective against another. PatchGate21 B.2 Additional POPE Results The main paper reports POPE [18] results averaged over the random, popular, and adversarial splits, and we further examine the full split-wise results (see Tab. D). Unlike AMBER discriminative QA, POPE directly matches PatchGate’s bidirectional object-level objective by including both present- and absent-object questions. Split-wise gains over the baseline. Across the random, popular, and adversarial splits, respectively, PatchGate improves Precision by 2.6%, 1.1%, and 3.0%, and Recall by 22.0%, 20.8%, and 19.3% over LLaVA-v1.5-7B, indicating fewer hallucinated “Yes” answers and fewer incorrect “No” answers for visible objects, respectively. These consistent gains yield the best Accuracy and F1 in every split. Precision-recall trade-off of competing methods. The metric-specific best results achieved by competing methods exhibit a pronounced Precision-Recall trade-off. PND [13] achieves the highest Precision in all three splits, but its Recall remains 7.6–9.6 points below PatchGate. Conversely, ProjectAway [12] achieves the highest Recall on the popular and adversarial splits, but its Precision falls 8.4 and 12.3 points below PatchGate, respectively. Overall, PatchGate ranks first in 7, second in 3, and third in the remaining 2 of the 12 split-wise metrics, demonstrating the strongest balance between suppressing hallucinated affirmative answers and correctly affirming visible objects. C Additional Qualitative Results Beyond the examples in Fig. 3 of the main paper, we provide additional qualitative comparisons with representative training-free hallucination-mitigation methods on AMBER [28] and POPE [18]. Captioning examples. Fig. D presents additional AMBER captioning examples. Suppression-oriented baselines remove some unsupported mentions but still introduce new ones, such as rope ladder, branches, and camera in the first image and white comforter and dining table in the second, and they provide no mechanism to restore omitted content. PatchGate instead removes these unsupported object mentions while recovering visible content omitted by the frozen backbone, such as sky in the first image and wooden floor and bed frame in the second, demonstrating balanced mitigation of hallucination and omission. Binary QA examples. Fig. E presents AMBER questions covering object existence, attributes, and relations, together with POPE examples from the random, popular, and adversarial splits. Most baselines exhibit a “Yes” bias on absent-object questions, answering affirmatively for objects such as island, ship, spoon, and teddy bear. PatchGate corrects both directions: it flips these hallucinated “Yes” answers to “No”, while also fixing incorrect “No” answers on questions about visible content, rather than trading one error type for the other. Failure patterns of suppression-only baselines. A closer look at Figs. D and E reveals a consistent pattern behind these corrections. Suppression-oriented methods reduce hallucinated mentions largely by making captions shorter and less specific, which leaves omissions untouched and can even drop correctly mentioned objects. PatchGate instead edits logits in both directions against the same pre- generation inventory, so suppression and recovery are coupled rather than traded off: captions remain similarly detailed while their object mentions align more closely with the image. This qualitative behavior mirrors the quantitative trends in Appendix B, where PatchGate improves Cover and Recall together with CHAIR and Precision. 22J. Ko et al. IMAGE 1IMAGE 2 LLaVA-v1.5-7B : The image features a man climbing a tall tree, possibly a palm tree, with a rope ladder. He is wearing a blue hat and appears to be in the process of reaching the top of the tree. The man is focused on his task, and his determination is evident as he ascends the tree. OPERA : In the image, a man is climbing a tall tree, possibly a palm tree, using a rope ladder. He appears to be focused intently as the rope ladder helps him reach the higher branches of the tree. The man is wearing a baseball cap, which adds to the casual and adventurous nature of the scene. VCD : The image features a man climbing a tall tree, possibly a palm tree, with his bare feet. He is wearing a blue cap and appears to be enjoying the adventure. The man is positioned near the top of the tree, showcasing his climbing skills. The scene takes place on a sunny day, with a blue sky visible in the background. SHIELD : In the image, a man wearing a baseball cap is perched high up in a tree, balancing himself on a thin tree branch. He appears to be climbing the tree and seems to be enjoying the adventure. The man is positioned in the center of the scene, with a camera nearby, capturing his daring feat. The sky is blue, providing a pleasant backdrop for the man's adventure. PND : The image features a man climbing a tall tree, possibly a palm tree, with a large trunk. He is wearing a blue hat and appears to be in the process of reaching for something on a higher branch. The man is positioned in the middle of the tree, with his body stretched out to reach the desired spot. The scene captures the man's determination and skill as he navigates the tree's branches. PatchGate (Ours) : The man in the tree is wearing a blue hat and is sitting on a tree limb, possibly a palm tree. He appears to be enjoying his time in the tree, possibly climbing it for fun or to reach a higher vantage point. The sky in the background is cloudy, adding to the overall ambiance of the setting. IMAGE 1 LLaVA-v1.5-7B : The image features a bed with a white mattress and a white comforter. The mattress is covered with a plastic sheet, which is likely to protect it from dust and allergens. The bed is positioned against a white wall, creating a clean and minimalist appearance. In the room, there is a chair located near the left side of the bed, and a dining table can be seen in the background. The overall setting appears to be a cozy and organized living space. OPERA : The image features a bed with a white mattress and a wooden headboard. The mattress is covered in plastic, likely to protect it from dust and allergens. The bed is positioned against a white wall, giving it a clean and minimalist appearance. A chair can be seen in the background, placed near the left side of the bed. VCD : The image features a new bed with a white mattress sitting in a bedroom. The mattress is wrapped in plastic, likely for protection or to keep it clean. The bed is positioned against a white wall, and there is a chair located near the top left corner of the room. The room appears to be clean and well-organized, with the bed being the main focus of the scene. SHIELD : The image shows a hotel room with a neatly made bed placed against a white wall. The bed features a mattress covered with plastic, possibly for protection from dust and allergens. There is a wooden headboard and footboard accompanying the mattress, creating a clean and minimalist appearance. addition to the bed, there is a chair located near the foot of the bed. The room appears to be well-maintained and inviting for guests. PND : The image features a bed with a white mattress and a wooden headboard. The mattress is covered in plastic, likely to protect it from dust, dirt, or any potential damage. The bed is positioned against a wall, and the room appears to be clean and well-organized. The overall setting gives the impression of a newly purchased bed or one that has been recently cleaned and prepared for use. PatchGate (Ours) : The bed mattress is white and appears to be new, as it is still wrapped in plastic. The mattress is placed on a wooden floor, and it is leaning against a wall. The bed frame is made of wood and is located on the left-hand end of the mattress. The overall setting suggests that the mattress is either being delivered or has been recently purchased and is waiting to be placed on a bed. IMAGE 2 Fig. D. Additional qualitative results on AMBER generative task. PatchGate suppresses unsupported object mentions (red) while recovering omitted visible objects(green), generating precise and complete captions. PatchGate23 AMBER QA POPE QA Fig. E. Additional qualitative results on AMBER and POPE QA. Correct and incorrect answers are shown in green and red, respectively. PatchGate corrects both hallucinated “Yes” answers and omitted “No” answers across diverse question types. 24J. Ko et al. D Details of Ablation Study This section examines how the main design choices of PatchGate affect the balance between omission recovery and hallucination suppression. Unless otherwise specified, all analyses use the AMBER [28] generative task with the LLaVA-v1.5-7B backbone [20]. Table E. Effect of VIED inclusion and exclusion strengths. VIED HyperparametersMetrics α esi γ ede Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ γ ede = 0.5 fixed 20.551.66.1 30.73.2 40.5 52.86.331.53.1 80.556.06.634.33.0 160.558.86.836.43.1 α esi = 8 fixed 80.2556.36.836.43.0 80.556.06.634.33.0 81.054.66.5 32.42.8 82.052.57.033.1 2.7 Effect of VIED hyperparameters. Varying the inclusion strength reveals a clear trade-off between recovering visible objects and introducing additional object mentions (see Tab. E). Withγ ede = 0.5 fixed, increasingα esi from 2 to 16 raises Cover from 51.6 to 58.8, while CHAIR and Hal gradually increase. We therefore selectα esi = 8, which provides a substantial coverage gain without the larger hallucination increase observed at α esi = 16. A moderate exclusion strength improves hallucination suppression while retaining most of the recovered content. Increasingγ ede from 0.25 to 1.0 reduces CHAIR from 6.8 to 6.5 and Hal from 36.4 to 32.4, but also lowers Cover from 56.3 to 54.6. Atγ ede = 2.0, Cover decreases further and CHAIR rises to 7.0, indicating that excessive suppression can disrupt otherwise valid generation. We useγ ede = 0.5 as a conservative operating point that preserves the coverage improvement while maintaining effective hallucination control. Table F. Sensitivity to the QA guidance strength λ QA on AMBER discriminative QA. λ QA Acc. ↑ P. ↑ R. ↑ F1 ↑ 0.0 (base) 72.0 92.5 62.9 74.9 0.373.2 92.9 64.6 76.2 0.573.4 93.5 65.7 77.2 0.774.593.866.077.4 0.974.4 93.2 66.2 77.4 1.074.4 93.1 66.3 77.4 Binary-QA guidance strength. Binary-QA inference combines the conditioned and unconditioned logits using the QA guidance strengthλ QA , as defined in Eq. D. Increasingλ QA from 0 to 0.7 PatchGate25 improves Accuracy from 72.0 to 74.5, Precision from 92.5 to 93.8, Recall from 62.9 to 66.0, and F1 from 74.9 to 77.4 on AMBER discriminative QA (see Tab. F). Larger values provide only a marginal Recall gain, while Accuracy and Precision slightly decrease and F1 remains unchanged. We therefore useλ QA = 0.7, which provides the strongest overall balance between the original answer distribution and the inventory-conditioned signal. Table G. Linguistic quality of PatchGate on AMBER. MethodPPL ↓ Rep-4 ↓ Avg. Len. LLaVA-v1.5-7B 4.21 0.01852.3 + PatchGate4.350.02058.1 LLaVA-v1.5-13B 3.98 0.01654.0 + PatchGate4.100.01859.2 Qwen2.5-VL-7B 3.62 0.01261.5 + PatchGate3.700.01364.8 InstructBLIP-7B 5.80 0.04138.2 + PatchGate5.950.04342.0 Linguistic quality. PatchGate operates consistently across different model scales and VLM families in the main experiments (see Tab. 7). Because it directly adjusts object-token logits, we further evaluate sentence-level fluency using perplexity (PPL), 4-gram repetition (Rep-4), and average response length across these backbones (see Tab. G). Lower PPL and Rep-4 indicate more probable and less repetitive generation, respectively. Across all evaluated backbones, PatchGate produces longer responses with modest increases in PPL and Rep-4. Although these metrics do not identify the precise linguistic source of the changes, their small magnitude indicates that the gains in object reliability do not substantially degrade overall fluency. Table H. Computational cost comparison. Method External Module Peak VRAM ↓ Inference Time ↓ [CVPR’24] LLaVA-v1.5-7BNo 14.6 GB 2.92 s + [CVPR’24] OPERANo16.8 GB24.50 s + [CVPR’24] VCDNo15.2 GB5.10 s + [CVPR’25] Devils-in-Mid.No15.0 GB3.05 s + [ICML’25] MARINEYes17.5 GB3.60 s + [ICLR’25] ProjectAwayNo15.1 GB3.20 s + [ICML’26] ILVADNo15.3 GB4.20 s + [ICLR’26] SHIELDNo15.4 GB5.80 s + [CVPR’26] PNDNo17.2 GB8.84 s + PatchGate (Ours)No14.9 GB3.30 s Computational efficiency. Extracting visual evidence directly from the frozen target VLM allows PatchGate to avoid any external detection or tagging module while introducing limited computational overhead (see Tab. H). Relative to LLaVA-v1.5-7B, it increases peak VRAM by only 0.3 GB (2.1%) and inference time by 0.38 seconds (13.0%). Its peak memory of 14.9 GB is the lowest among the 26J. Ko et al. AMBER #17AMBER #28AMBER #209 water sea ocean board waters sky suit front boards 0.00 0.01 0.02 0.03 PCS =0.02 trees cloud road clouds country grass sky side line 0.000 0.015 0.030 =0.03 pill bed hotel room frame walls chair table sheets 0.000 0.025 0.050 0.075 =0.05 water glass ocean front board sky sea waters image 0.00 0.25 0.50 0.75 PCS =0.36 country cloud road trees clouds side girl grass flowers 0.0 0.3 0.6 0.9 =0.46 pill frame room car lamp chair bed wall walls 0.0 0.3 0.6 0.9 =0.21 water sea ocean board suit surface waters front dust 0.000 0.015 0.030 PCS =0.02 road trees cloud grass clouds sky country side bush 0.00 0.02 0.04 0.06 =0.05 bed pill hotel table room wall wallslamp chair 0.000 0.025 0.050 0.075 =0.08 board water front ocean suit glass sea tail sky 0.0 0.3 0.6 0.9 PCS =0.85 grass road trees cloud country girl sky bush clouds 0.0 0.3 0.6 0.9 =0.54 pill bed hotel lamp chair table room frame sun 0.0 0.3 0.6 0.9 =0.68 Peak layer · Mean Peak layer · Max Later layers · Mean Later layers · Max (default) present (AMBER truth) absent (hallucination target) unannotated object : largest-gap cut, T(I) = c : PCS > Fig. F. Evidence aggregation and adaptive object selection. Mean aggregation retains only a few high-confidence objects, whereas max-softmax aggregation over the selected late-layer range more clearly separates supported objects from the low-evidence tail. evaluated hallucination-mitigation methods, while its 3.30-second inference time is shorter than those of most competing approaches. These results indicate that PatchGate improves object reliability with modest memory and runtime overhead, without introducing an external visual model. Evidence aggregation and adaptive object selection. The effectiveness of the largest-gap cutoff depends on whether the evidence scores preserve localized object responses while separating supported objects from the low-evidence tail (see Fig. F and Tab. I). Mean aggregation dilutes localized object responses across the averaged dimensions, causing the largest gap to appear near the beginning of the ranked PatchGate27 Table I. Effect of the VEX layer source and evidence aggregation on AMBER. Layer SourceAggregation Tag-levelCaption-level P. ↑ R. ↑ F1 ↑ |T| Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ single peak late layer Mean97.7 9.0 16.4 1.032.111.2 30.02.4 single peak late layer Max89.5 31.9 47.0 6.349.37.132.52.9 Later layers LMean97.6 10.2 18.4 1.132.811.230.4 2.3 Later layers LMax84.849.262.312.956.06.634.33.0 Table J. Effect of VEX layer range. Layer Range Tag-levelCaption-level P. ↑ R. ↑ F1 ↑ |T| Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ All layers (0%− 100%)83.4 49.6 62.2 13.555.87.035.13.0 Later 67% (33%− 100%) 83.4 49.5 62.1 13.455.76.935.0 2.9 Later 33% (67%− 100%)84.849.262.312.956.06.634.33.0 list. It therefore retains only about one object per image on average, yielding high tag Precision but very low Recall. In other words, the few retained objects are generally reliable, but many visible objects are omitted from the inventory, resulting in low Cover and high CHAIR. Max aggregation instead preserves the strongest patch-layer response for each object and produces a clearer separation between supported objects and the low-evidence tail. Its benefit is largest when evidence is collected across the selected later layers, since different objects can exhibit their strongest vocabulary-space responses at different depths. Compared with single-layer max aggregation, later- layer max aggregation increases tag Recall from 31.9 to 49.2, tag F1 from 47.0 to 62.3, and Cover from 49.3 to 56.0, while reducing CHAIR from 7.1 to 6.6. We therefore apply the largest-gap cutoff to max-softmax scores aggregated over L =22,..., 32. Effect of VEX layer range. With max aggregation fixed, restricting evidence extraction to the later 33% of the language model produces the most useful inventory for downstream decoding (see Tab. J). Compared with using all layers, this setting changes tag Recall only slightly from 49.6 to 49.2, while increasing tag Precision and F1 and reducing the average inventory size. This pattern is consistent with object evidence becoming more directly aligned with vocabulary-space predictions in later representations, whereas including earlier layers introduces responses that are less useful for output-space intervention. The later 33% consequently achieves the best Cover and CHAIR, and we use L =22,..., 32 as the default layer range. Internal and external object inventories. Replacing VEX with GroundingDINO [22] or RAM++ [9] preserves usable caption-level performance, indicating that PatchGate can operate with different inventory sources (see Tab. K). Despite using no external visual prior, VEX also remains competitive at the tag level, achieving an F1 of 62.3, only 2.6 points below GroundingDINO and 6.0–8.8 points below the RAM++ variants. More importantly, the higher tag-level F1 of the external models does not translate into stronger downstream captioning. VEX improves Cover to 56.0, compared with 50.1–51.3 for the external inventories, while reducing CHAIR to 6.6, compared with 7.3–8.2. Several interface differences help characterize this result. First, not every external tag maps one- to-one to the editable single-token object vocabulary of the target VLM, so some detected concepts cannot be directly associated with an output-token logit. Second, the retained external tags primarily provide positive inventory membership, whereas PCS supplies a continuous evidence score over the target VLM’s editable object vocabulary and can therefore support both inclusion and exclusion. Consistent with this distinction, PCS provides stronger diagnostic separation of hallucinated and 28J. Ko et al. Table K. Comparison of internal (VEX; Sec. 3.1) and external object inventories. Method Tag-levelCaption-levelComputational Cost P. ↑ R. ↑ F1 ↑ |T| Cover ↑ CHAIR ↓ Hal ↓ Cog ↓ Param. Peak VRAM Time LLaVA-v1.5-7B–49.47.531.43.6 7.0 B 14.6 GB 2.92 s + GroundingDINO91.1 50.4 64.9 5.251.37.333.33.67.2 B16.2 GB3.26 s + RAM++ (default)97.0 59.6 71.1 11.650.27.836.84.27.3 B16.9 GB3.17 s + RAM++ (top-10)95.8 53.0 68.3 10.050.18.237.24.37.3 B16.9 GB3.17 s + VEX (Ours; Sec. 3.1)84.849.262.312.956.06.634.33.07.0 B14.9 GB3.30 s Table L. Diagnostic performance of the thresholded RAM++ tag decision and continuous PCS for evidence– verbalization mismatch on the same object instances as Tab. 4. Diagnostic signal HallucinationOmission AUROC ↑ AUPRC ↑ AUROC ↑ AUPRC ↑ Random0.5000.1630.5000.233 RAM++0.8240.3790.6160.392 PCS (Ours)0.8830.7070.8900.705 omitted objects than the thresholded RAM++ tag signal (see Tab. L). External inventories also require additional model parameters and peak memory: GroundingDINO and RAM++ use 16.2 GB and 16.9 GB, respectively, compared with 14.9 GB for VEX. Together, these results indicate that external taggers remain viable inventory sources, but higher tag-level accuracy alone does not guarantee a more effective or efficient decoding intervention. Mismatch diagnosis using inventory signals. We further compare the signals exposed by VEX (Ours; Sec. 3.1) and RAM++ as diagnostics of the evidence–verbalization mismatch, using the same object instances as Tab. 4. Hallucination is evaluated over objects mentioned in the baseline caption, whereas omission is evaluated over unmentioned objects. RAM++ first produces per-class sigmoid scores over its 4,585-word vocabulary and retains only the classes whose scores exceed learned class-specific thresholds. Its resulting inventory therefore provides a binary tag decision for each vocabulary item, whereas PCS provides a continuous prompt-free evidence score within the target VLM. Using these deployed signals, PCS achieves higher AUROC and AUPRC for both hallucination and omission, with the largest difference observed for omission (0.890 vs. 0.616 AUROC; see Tab. L). This gap suggests that graded target-model evidence better preserves the information needed to identify omitted visible objects, whereas thresholded external tags may discard weak but informative evidence and cannot cover objects outside the external vocabulary. E Limitations and Future Work PatchGate currently represents visual evidence primarily through individual object tokens, so its largest gains appear in object-centric captioning and existence QA. Attribute and relation questions require structured evidence that is not explicitly captured by the current inventory. Future work will extend the inventory to object–attribute pairs and object–relation–object triplets, together with corresponding evidence-aware decoding strategies. The predefined single-token vocabulary may also limit coverage of open-vocabulary or multi-token concepts, motivating improved evidence calibration and integration with external visual signals. In addition, directly promoting object tokens can slightly affect local phrasing, perplexity, and repetition; fluency-aware constraints may reduce this effect. Finally, because Hal and Cog can penalize visually valid but unannotated objects, more exhaustive annotations and human verification would provide a more reliable evaluation. PatchGate29 References [1]An, W., Tian, F., Leng, S., Nie, J., Lin, H., Wang, Q., Chen, P., Zhang, X., Lu, S.: Mitigating object hallucinations in large vision-language models with assembly of global and local attention. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 29915–29926 (2025) [2] Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-vl technical report (2025), https://arxiv.org/abs/2502.13923 [3]Belrose, N., Ostrovsky, I., McKinney, L., Furman, Z., Smith, L., Halawi, D., Biderman, S., Steinhardt, J.: Eliciting latent predictions from transformers with the tuned lens (2025),https: //arxiv.org/abs/2303.08112 [4]Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S.: End-to-end object detection with transformers. In: European Conference on Computer Vision. p. 213–229. Springer (2020) [5] Chen, Z., Zhao, Z., Luo, H., Yao, H., Li, B., Zhou, J.: Halc: object hallucination reduction via adaptive focal-contrast decoding. In: Proceedings of the 41st International Conference on Machine Learning. p. 7824–7846. JMLR.org (2024) [6]Dai, W., Li, J., Li, D., Tiong, A.M.H., Zhao, J., Wang, W., Li, B., Fung, P., Hoi, S.: Instructblip: Towards general-purpose vision-language models with instruction tuning (2023),https://arxiv. org/abs/2305.06500 [7] Geva, M., Caciularu, A., Wang, K., Goldberg, Y.: Transformer feed-forward layers build pre- dictions by promoting concepts in the vocabulary space. In: Proceedings of the 2022 Con- ference on Empirical Methods in Natural Language Processing. p. 30–45 (2022).https: //doi.org/10.18653/v1/2022.emnlp-main.3 [8]Huang, Q., Dong, X., Zhang, P., Wang, B., He, C., Wang, J., Lin, D., Zhang, W., Yu, N.: OPERA: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 13418–13427 (2024) [9]Huang, X., Huang, Y.J., Zhang, Y., Tian, W., Feng, R., Zhang, Y., Xie, Y., Li, Y., Zhang, L.: Open-set image tagging with multi-grained text supervision. In: Proceedings of the 33rd ACM International Conference on Multimedia. p. 4117–4126. Association for Computing Machinery (2025). https://doi.org/10.1145/3746027.3755316 [10]Huang, Y., Shi, L., Zhang, Y., Xu, Y., Fu, Y.: SHIELD: Suppressing hallucinations in LVLM encoders via bias and vulnerability defense. In: International Conference on Learning Represen- tations (2026) [11]Huo, F., Xu, W., Zhang, Z., Wang, H., Chen, Z., Zhao, P.: Self-introspective decoding: Alleviating hallucinations for large vision-language models. In: International Conference on Learning Representations (2025) [12]Jiang, N., Kachinthaya, A., Petryk, S., Gandelsman, Y.: Interpreting and editing vision-language representations to mitigate hallucinations. In: International Conference on Learning Representa- tions (2025) [13]Jiang, Y., An, Y., Yang, X., Wuerkaixi, A., Cheng, X., Xie, F., Jiang, Z., Liu, C., Zeng, K., Zhang, H.: Breaking the illusion: When positive meets negative in multimodal decoding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 4210–4220 (2026) [14]Jiang, Z., Chen, J., Zhu, B., Luo, T., Shen, Y., Yang, X.: Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating object hallucinations via attention lens. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 25004–25014 (2025) 30J. Ko et al. [15] Jo, S., Ryu, S., Kim, S., Yang, E., Kim, K.: Ttd: Text-tag self-distillation enhancing image-text alignment in clip to alleviate single tag bias. In: European Conference on Computer Vision. p. 341–357. Springer (2024) [16]Kogilathota, S.A., EG, S.V., Sun, L., Zhou, J.: HALP: Detecting hallucinations in vision- language models without generating a single token. In: Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). p. 6067–6085 (2026). https://doi.org/10.18653/v1/2026.eacl-long.287 [17]Leng, S., Zhang, H., Chen, G., Li, X., Lu, S., Miao, C., Bing, L.: Mitigating object hallucinations in large vision-language models through visual contrastive decoding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 13872–13882 (2024) [18]Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., Wen, J.R.: Evaluating object hallucination in large vision-language models. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. p. 292–305 (2023).https://doi.org/10.18653/v1/2023. emnlp-main.20 [19] Li, Z., Shi, H., Gao, Y., Liu, D., Wang, Z., Chen, Y., Liu, T., Zhao, L., Wang, H., Metaxas, D.N.: The hidden life of tokens: Reducing hallucination of large vision-language models via visual information steering. In: Proceedings of the 42nd International Conference on Machine Learning. vol. 267, p. 35799–35819. PMLR (2025) [20]Liu, H., Li, C., Li, Y., Lee, Y.J.: Improved baselines with visual instruction tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 26296– 26306 (2024) [21]Liu, S., Zheng, K., Chen, W.: Paying more attention to image: A training-free method for alleviating hallucination in lvlms. In: European Conference on Computer Vision. p. 125–140. Springer (2024). https://doi.org/10.1007/978-3-031-73010-8_8 [22] Liu, S., Zeng, Z., Ren, T., Li, F., Zhang, H., Yang, J., Jiang, Q., Li, C., Yang, J., Su, H., Zhu, J., Zhang, L.: Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In: European Conference on Computer Vision. p. 38–55. Springer (2024) [23] Miller, G.A.: Wordnet: a lexical database for english. Communications of the ACM 38(11), 39–41 (1995). https://doi.org/10.1145/219717.219748 [24] Neo, C., Ong, L., Torr, P., Geva, M., Krueger, D., Barez, F.: Towards interpreting visual information processing in vision-language models. In: International Conference on Learning Representations (2025) [25]Phukan, A., Divyansh, D., Morj, H.K., Vaishnavi, V., Saxena, A., Goswami, K.: Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs. In: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). p. 9661–9675 (2025). https://doi.org/10.18653/v1/2025.naacl-long.488 [26] Rohrbach, A., Hendricks, L.A., Burns, K., Darrell, T., Saenko, K.: Object hallucination in image captioning. In: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. p. 4035–4045 (2018). https://doi.org/10.18653/v1/D18-1437 [27]Wang, C., Chen, X., Zhang, N., Tian, B., Xu, H., Deng, S., Chen, H.: Mllm can see? dynamic correction decoding for hallucination mitigation. In: International Conference on Learning Representations (2025) [28] Wang, J., Wang, Y., Xu, G., Zhang, J., Gu, Y., Jia, H., Wang, J., Xu, H., Yan, M., Zhang, J., Sang, J.: AMBER: An LLM-free multi-dimensional benchmark for MLLM hallucination evaluation (2023), https://arxiv.org/abs/2311.07397 [29]Wang, X., Pan, J., Ding, L., Biemann, C.: Mitigating hallucinations in large vision-language models with instruction contrastive decoding. In: Findings of the Association for Computational Linguistics ACL 2024. p. 15840–15853 (2024) [30]Xie, Y., Hua, Z., Wang, R., Ng, W.W.Y., Wang, X., Jia, Y.: Finding the correct visual evidence without forgetting: Mitigating hallucination in lvlms via inter-layer visual attention discrepancy. In: Proceedings of the 43rd International Conference on Machine Learning (2026) PatchGate31 [31] Xie, Y., Li, G., Xu, X., Kan, M.Y.: V-DPO: Mitigating hallucination in large vision language models via vision-guided direct preference optimization. In: Findings of the Association for Computational Linguistics: EMNLP 2024. p. 13258–13273. Association for Computational Linguistics (2024). https://doi.org/10.18653/v1/2024.findings-emnlp.775 [32] Yang, Z., Luo, X., Han, D., Xu, Y., Li, D.: Mitigating hallucinations in large vision-language models via dpo: On-policy data hold the key. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 10610–10620 (2025) [33]Yin, S., Fu, C., Zhao, S., Xu, T., Wang, H., Sui, D., Shen, Y., Li, K., Sun, X., Chen, E.: Woodpecker: Hallucination correction for multimodal large language models. Science China Information Sciences 67(12), 220105 (2024).https://doi.org/10.1007/s11432-024-4251-x [34] Yu, L., Chen, Z., Kuang, P., Feng, Z., Zhou, F., Wang, L., Dobbie, G.: Causally-grounded dual-path attention intervention for object hallucination mitigation in lvlms. Proceedings of the AAAI Conference on Artificial Intelligence 40(42), 36021–36029 (2026).https://doi.org/10. 1609/aaai.v40i42.40918 [35] Zhang, B., Sennrich, R.: Root mean square layer normalization. In: Advances in Neural Infor- mation Processing Systems (2019) [36]Zhao, L., Deng, Y., Zhang, W., Gu, Q.: Mitigating object hallucination in large vision-language models via image-grounded guidance. In: Proceedings of the 42nd International Conference on Machine Learning. JMLR.org (2025)