Paper deep dive
PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding
Baolong Bi, Yuyao Ge, Shenghua Liu, Yuchen He, Siqian Tong, Lizhe Chen, Lingrui Mei, Zehao Li, Yiwei Wang, Yujun Cai, Ming-Hsuan Yang, Xueqi Cheng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 2:14:09 PM
Summary
The paper introduces PromptCD (Polarity-Prompt Contrastive Decoding), a test-time behavior control method for Large Language Models (LLMs) and Vision-Language Models (VLMs). It enhances model alignment by contrasting token-level probability distributions (for LLMs) and visual attention patterns (for VLMs) generated by paired positive and negative guiding prompts. This approach improves helpfulness, honesty, and harmlessness in LLMs and visual grounding in VLMs without additional training.
Entities (9)
Relation Signals (8)
PromptCD → appliesto → LLM
confidence 95% · This formulation extends contrastive decoding to a wide range of enhancement objectives and is applicable to both LLMs and Vision-Language Models (VLMs)
PromptCD → appliesto → VLM
confidence 95% · This formulation extends contrastive decoding to a wide range of enhancement objectives and is applicable to both LLMs and Vision-Language Models (VLMs)
PromptCD → enhances → VQA
confidence 90% · For VLMs, we extend PromptCD to Visual Question Answering (VQA)... showing that PromptCD significantly improves VQA performance
PromptCD → improves → Helpfulness
confidence 90% · For LLMs, experiments on the “3H” alignment objectives (helpfulness, honesty, and harmlessness) demonstrate consistent and substantial improvements
PromptCD → improves → Honesty
confidence 90% · For LLMs, experiments on the “3H” alignment objectives (helpfulness, honesty, and harmlessness) demonstrate consistent and substantial improvements
PromptCD → improves → Harmlessness
confidence 90% · For LLMs, experiments on the “3H” alignment objectives (helpfulness, honesty, and harmlessness) demonstrate consistent and substantial improvements
PromptCD → uses → Contrastive Decoding
confidence 90% · PromptCD constructs paired positive and negative guiding prompts for a target behavior and contrasts model responses... to reinforce desirable outcomes.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reliable AI systems require large language models (LLMs) to exhibit behaviors aligned with human preferences and values. However, most existing alignment approaches operate at training time and rely on additional high-quality data, incurring significant computational and annotation costs. While recent work has shown that contrastive decoding can leverage a model's internal distributions to improve specific capabilities, its applicability remains limited to narrow behavioral scopes and scenarios. In this work, we introduce Polarity-Prompt Contrastive Decoding (PromptCD), a test-time behavior control method that generalizes contrastive decoding to broader enhancement settings. PromptCD constructs paired positive and negative guiding prompts for a target behavior and contrasts model responses-specifically token-level probability distributions in LLMs and visual attention patterns in VLMs-to reinforce desirable outcomes. This formulation extends contrastive decoding to a wide range of enhancement objectives and is applicable to both LLMs and Vision-Language Models (VLMs) without additional training. For LLMs, experiments on the "3H" alignment objectives (helpfulness, honesty, and harmlessness) demonstrate consistent and substantial improvements, indicating that post-trained models can achieve meaningful self-enhancement purely at test time. For VLMs, we further analyze contrastive effects on visual attention, showing that PromptCD significantly improves VQA performance by reinforcing behavior-consistent visual grounding. Collectively, these results highlight PromptCD as a simple, general, and cost-efficient strategy for reliable behavior control across modalities.
Tags
Links
- Source: https://arxiv.org/abs/2602.20696v1
- Canonical: https://arxiv.org/abs/2602.20696v1
Trouble viewing inline? Open PDF directly →
Full Text
88,707 characters extracted from source content.
Expand or collapse full text
JOURNAL OF L A T E X CLASS FILES, JANUARY 20251 PromptCD: Test-Time Behavior Enhancement via Polarity-Prompt Contrastive Decoding Baolong Bi, Yuyao Ge, Shenghua Liu † , Yuchen He, Siqian Tong, Lizhe Chen, Lingrui Mei, Zehao Li, Yiwei Wang, Yujun Cai, Ming-Hsuan Yang ID , Fellow, IEEE and Xueqi Cheng ID , Fellow, IEEE Abstract—Reliable AI systems require large language models (LLMs) to exhibit behaviors aligned with human preferences and values. However, most existing alignment approaches operate at training time and rely on additional high-quality data, incurring significant computational and annotation costs. While recent work has shown that contrastive decoding can leverage a model’s internal distributions to improve specific capabilities, its applicability remains limited to narrow behavioral scopes and scenarios. In this work, we introduce Polarity-Prompt Contrastive Decoding (PromptCD), a test-time behavior control method that generalizes contrastive decoding to broader enhancement settings. PromptCD constructs paired positive and negative guiding prompts for a target behavior and contrasts model responses—specifically token-level probability distributions in LLMs and visual attention patterns in VLMs—to reinforce desirable outcomes. This formulation extends contrastive decoding to a wide range of enhancement objectives and is applicable to both LLMs and Vision-Language Models (VLMs) without additional training. For LLMs, experiments on the “3H” alignment objectives (helpfulness, honesty, and harmlessness) demonstrate consistent and substantial improvements, indicating that post-trained models can achieve meaningful self-enhancement purely at test time. For VLMs, we further analyze contrastive effects on visual attention, showing that PromptCD significantly improves VQA performance by reinforcing behavior-consistent visual grounding. Collectively, these results highlight PromptCD as a simple, general, and cost-efficient strategy for reliable behavior control across modalities. Index Terms—Large language models, Test-time enhancement, Contrastive decoding, Behavioral emergence, Visual attention ✦ 1 INTRODUCTION L ARGE language models (LLMs) [1–3] have emerged as dominant intelligent assistants, distinguished by their extensive knowledge and robust reasoning capabilities. As these systems are increasingly integrated into critical decision-making and creative workflows, ensuring their alignment with human intentions and values has become a cornerstone of trustworthy AI [4–6]. However, aligning general-purpose models is inherently complex: human pref- erences are diverse, context-dependent, and often conflict- ing across different tasks [7, 8]. Static policies derived from standard training cannot adequately capture this variabil- ity, rendering dynamic behavioral adaptation an essential capability for next-generation systems [9, 10]. Prevailing alignment efforts [11–13] predominantly focus on the post- ●This work is supported in part by the National Key R&D Program of China under Grant Nos. 2023YFA1011602, Beijing Natural Science Foundation No. 4262033, and the National Natural Science Foundation of China under Grant Nos. U25B2076, 62441229, and 62377043.) ●Baolong Bi, Yuyao Ge, Yuchen He, Lingrui Mei, Shenghua Liu, and Xueqi Cheng are with the Key Laboratory of Network Data Science and Technology, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China, and also with the Uni- versity of Chinese Academy of Sciences (e-mail: bibaolong23z@ict.ac.cn; geyuyao24z@ict.ac.cn; heye58478@gmail.com; meilingrui25b@ict.ac.cn; liushenghua@ict.ac.cn; cxq@ict.ac.cn). ●Siqian Tong and Zehao Li are with the University of Chinese Academy of Sciences, Beijing, China (e-mail: tongsiqian01@gmail.com; lize- hao23z@ict.ac.cn). ●Lizhe Chen is with the Shenzhen International Graduate School, Tsinghua University, Shenzhen, China (e-mail: chen-lz25@mails.tsinghua.edu.cn). ●YujunCaiiswiththeUniversityofQueensland(e-mail: vanora.caiyj@gmail.com). ●Yiwei Wang and Ming-Hsuan Yang are with the University of California at Merced (e-mail: mhyang@ucmerced.edu; yiweiwang2@ucmerced.edu). training stage, utilizing large curated datasets to fine-tune models via reinforcement learning from human feedback (RLHF) [14, 15] or preference optimization [16]. Despite their success, these approaches face three critical limitations: the prohibitive cost of constructing high-quality datasets; the rigidity of fixed parameters which restricts adaptability to novel contexts; and the difficulty of simultaneously opti- mizing conflicting objectives [17, 18]. Consequently, there is a critical need for lightweight mechanisms capable of steer- ing model behavior at inference time without the burden of retraining or altering core parameters. Test-time steering [19, 20] offers a complementary paradigm to address the rigidity of post-training align- ment, enabling on-the-fly behavioral modulation. Existing methods primarily intervene at three stages: (1) prompt- based steering, which injects behavioral cues via natural lan- guage [21–23]; (2) representation-based adjustment, which modifies internal activations along learned directions [24– 26]; and (3) decoding-level control, which perturbs token logits during generation [27–29]. Despite expanding the scope of controllable alignment, these techniques exhibit notable drawbacks: prompting is often inconsistent, rep- resentation editing requires behavior-specific tuning, and decoding strategies are typically confined to predefined objectives. Developing a generalized, parameter-free test- time control mechanism that integrates multiple behavioral dimensions remains an open challenge. In this work, we begin by investigating the efficacy of polarity prompts—explicit cues designed to encourage or suppress specific behaviors. To quantitatively ground our analysis, we take context-faithfulness as a representative case and design a Knowledge Token Capturing algorithm to arXiv:2602.20696v1 [cs.AI] 24 Feb 2026 JOURNAL OF L A T E X CLASS FILES, JANUARY 20252 Grounded on Image Context: Elon Musk acquired Twitter for $44 billion in 2022 and soon rebranded it as X. Instruction:Who is the current owner of Twitter? LLMInput VLMInput Instruction: <image> What shape is seen through the cup’s handle? Polarity Prompting Answer the question faithfully based on the <context>. Positive Answer the question based on your internal knowledge. Negative What shape is seen through the cup’s handle? Positive Please Write a genera description of the input image. Negative Elon Mask Positive Distribution Negative Distribution Jack Dorsey Positive Attention Negative Attention Vision Modality Track Text Modality Track Attentiveness Target Behavior Helpfulness Honesty Harmlessness Enhance Attentiveness ... Enhance Faithfulness Polarity-PromptContrastive Decoding Positive Behavior Negative Behavior Enhanced Behavior Faithful to Context Output: Based on the context, the current owner is Elon Musk. Output: It’s a star shape in the cup’s handle . Test - Time Enhancement Elon Mask Enhanced Distribution Jack Dorsey Fig. 1. Overview of PromptCD. A unified framework for test-time behavior enhancement. By contrasting the logits (for text) or visual attention maps (for images) induced by positive and negative prompts, PromptCD effectively steers models toward desired goals without retraining. We demonstrate this capability through faithfulness enhancement in LLMs and visual attentiveness improvement in VLMs. systematically examine how these cues reshape the model’s output distribution and behavioral preferences. Through this lens, we identify a phenomenon of latent alignment: while polarity prompts successfully shift the probability mass toward context-faithful tokens, this internal signal frequently fails to manifest in the final generated response. The model exhibits a “willingness” to be faithful—evident in the logits captured by our algorithm—but lacks the de- cisiveness to override its parametric priors during standard decoding. These findings suggest that while prompting acts as a directional guide, it is insufficient for robust control , underscoring the necessity for a decoding mechanism capable of actively amplifying these latent signals. Building on these observations, we propose Polarity- Prompt Contrastive Decoding (PromptCD), a unified test- time steering framework designed to enhance targeted be- haviors without parameter updates. PromptCD employs a customizable prompting mechanism to address diverse alignment objectives, ranging from context-adherence and honesty to safety and visual grounding. For a specific target behavior, we construct paired polarity prompts that pro- vide contrasting instructions: a positive prompt explicitly encourages the desired capability, while a negative prompt suppresses it or elicits competing tendencies, such as re- liance on parametric priors or hallucinations. By contrast- ing the model responses conditioned on these cues, the framework amplifies signals specific to the target behavior while filtering out generic distribution priors. This approach generalizes across modalities: for general LLMs, it refines generation by contrasting token probability distributions; for Vision-Language Models (VLMs), it leverages cross- modal attention maps to improve behavior-relevant visual grounding. We provide an illustrative overview of this con- trastive process in Figure 1, and detail the full specifications of polarity prompts for various tasks in the Appendix. To evaluate the effectiveness and generality of our ap- proach, we conduct comprehensive cross-modal experi- ments encompassing both Large Language Models (LLMs) and Vision-Language Models (VLMs). For LLMs, we as- sess behavioral alignment across the “3H” dimensions: helpfulness, honesty, and harmlessness [30]. Specifically, helpfulness is evaluated via context-faithfulness tasks on benchmarks such as NQ [31], ConFiQA [13], and CoConflic- tQA [32]; honesty is measured using TruthfulQA [33] and FactScore [34]; and harmlessness is tested on SafeEdit [35]. For VLMs, we extend PromptCD to Visual Question An- swering (VQA). Here, we analyze not only token probability shifts but also the modulation of visual attention, demon- strating that polarity prompts sharpen the model’s focus on semantically relevant regions to improve answer ground- ing. Complementing these quantitative metrics, we pro- vide granular token-level and attention-based interpretabil- ity analyses, elucidating how PromptCD dynamically re- shapes generation trajectories during decoding. Empirical results confirm that PromptCD consistently enhances align- ment across modalities, establishing it as a lightweight, parameter-free mechanism for flexible test-time behavior control. 2 PRELIMINARIES In this section, we formalize the generation processes of Large Language Models (LLMs) and Vision-Language Mod- els (VLMs). This formulation establishes the theoretical basis for our proposed method, which intervenes at the prediction level for LLMs and the attention level for VLMs. 2.1 Autoregressive Generation in LLMs Large language models function as probabilistic estimators over sequences of tokens. Given a sequence of previous tokens X <t = x 1 ,x 2 ,...,x t−1 , the model aims to predict the next token x t by estimating the conditional probability distribution. Formally, an LLM parameterized by θ com- putes this probability via a softmax output layer: IP θ (x t ∣X <t ) = softmax(φ(h t )),(1) whereh t denotes the contextualized hidden state at step t, and φ(⋅) is a linear transformation projectingh t into the JOURNAL OF L A T E X CLASS FILES, JANUARY 20253 5.07.510.012.515.017.520.022.525.0 Probability of Contextual Knowledge 0 50 5.07.510.012.515.017.520.022.525.0 Probability of Parametric Knowledge 0 25 Avg: 10.92 Avg: 11.38 Avg: 16.43 Avg: 15.45 Avg: 17.24 Avg: 13.52 Original PromptPositive-Behavior PromptNegative-Behavior Prompt 12345678 Rank of Contextual Knowledge 0 750 12345678 Rank of Parametric Knowledge 0 500 Fig. 2. Comparison of contextual and parametric knowledge shifts under positive and negative prompts. The first discriminative tokens are identified to represent the corresponding knowledge, and their logit and rank distributions are tracked to illustrate the underlying distributional dynamics. vocabulary space to produce logits. The hidden stateh t is derived through a stack of Transformer layers, processing the embedded sequenceX <t via self-attention and feed- forward operations: h t = FFN(Attn(X <t )).(2) The attention module Attn(⋅) captures contextual de- pendencies across tokens, and the feed-forward network FFN(⋅) applies nonlinear transformations to produce the contextualized representation. During decoding, the model autoregressively samples the next token x t according to the predicted probability distribution IP θ (x t ∣X <t ). While decod- ing strategies (e.g., greedy search, nucleus sampling) de- termine the specific selection rule, the generation trajectory is fundamentally governed by this probability distribution. Consequently, our proposed PromptCD method for LLMs operates directly on Equation 1, modulating the output log- its to amplify favorable behavioral signals before sampling. 2.2 Visual-Textual Grounding in VLMs VLMs extend the autoregressive paradigm by incorporating an input image I into the generation process. The computa- tion of hidden states relies on a multimodal attention mech- anism that jointly attends to textual history and encoded visual featuresV: h t = FFN(Attn text+vision (X <t ,V)),(3) where Attn text+vision (⋅) computes attention weights that quantify the cross-modal relevance of specific visual regions inV. The final probability is similarly computed as: IP θ (x t ∣X <t ,I) = softmax(φ(h t )).(4) While the probabilistic output (Equation 4) remains the final decision layer, the inclusion of cross-modal attention in Equation 3 introduces a new dimension of controllability. Just as LLMs conditioned on positive (P ) and nega- tive (N ) prompts yield distinct probability distributions P (P) andP (N) , VLMs conditioned on these prompts ex- hibit distinct visual attention patterns. We define A (P) l (I) as the attention distribution at layer l under the posi- tive prompt, which captures behavior-specific positive sig- nals (e.g., semantic relevance to the question). Conversely, we define A (N) l (I) as the attention under the negative prompt, which primarily reflects generic negative pat- terns (e.g., visual saliency or noise) shared across con- texts. PromptCD leverages this analogy by contrasting these attention maps—A (P) l (I) versus A (N) l (I)—to isolate and amplify the task-relevant visual grounding signals, similar to how it contrasts token probabilities in LLMs. 3 EMPIRICALANALYSIS:HOWPOLARITY PROMPTS RESHAPE OUTPUT DISTRIBUTIONS LLMs exhibit significant behavioral variability conditioned on their input prompts. To better understand how prompts explicitly steering toward or away from a target behavior in- fluence model generation, we perform an exploratory study on polarity prompts. This section formalizes the definition of polarity prompts, investigates their effects through the lens of context-faithfulness, introduces a token-level analysis algorithm, and highlights a key challenge in behavioral emergence. 3.1 Definition of Polarity Prompts We conceptualize behavioral steering as a modulation pro- cess driven by paired opposing cues. Rather than treating prompts merely as natural language instructions, we define Polarity Prompts as paired constraints designed to induce distinct behavioral distributions. Given a target behavior B, we construct a pair of prompts (P,N): ● Positive Prompt (P ): An instruction designed to explic- itly encourage the expression of the target behavior B (e.g., faithful reasoning) within the generation context. ● Negative Prompt (N ): An instruction designed to sup- press B or explicitly elicit a competing behavioral ten- dency (e.g., reliance on parametric priors). For the specific case of context-faithfulness in LLMs, the goal is to guide the model to reason based on provided contextual information rather than its pre-trained memory. Thus, a positive prompt might be: “Please answer the ques- tion faithfully based on the provided context.” Conversely, the JOURNAL OF L A T E X CLASS FILES, JANUARY 20254 negative prompt actively encourages the competing behav- ior: “Please answer the question based on your own parametric knowledge.” Formally, given an input instruction or question X q , we concatenate it with either a positive prompt X p or a negative prompt X n . The conditional probability of generating the next token under a positive prompt is formulated as: IP(x (P) m ∣x (P) <m ) = softmax(φ(h (P) m )),(5) whereh (P) m is the hidden representation of the sequence X q + X p . A corresponding distribution IP(x (N) n ∣x (N) <n ) is similarly derived from the sequence X q +X n . By contrasting these two distributions, we can rigorously quantify how polarity prompts modulate the model’s internal generation dynamics and shift token preferences. 3.2 Quantifying Behavior Shifts via Token-Level Trace To empirically verify whether polarity prompts effectively steer the model, we adopt context-faithfulness as a represen- tative testbed. We utilize the MQUAKE benchmark [36], which consists of questions where the ground truth (Para- metric Knowledge) conflicts with a provided counterfactual context (Contextual Knowledge). For instance, given the context ”Paris is the capital of the United States”, a faithful model should answer ”United States”, while a model relying on priors will answer ”France”. Standard evaluation metrics (e.g., accuracy) only reveal the final discrete output, masking the internal dynamics. To bridge this gap, we propose a fine-grained Knowledge Token Capturing Algorithm (Algorithm 1). This algorithm acts as a probe to trace the competition between the two knowledge sources at the token level. Specifically, it scans the generated logits to identify discriminative tokens—the first token position where the contextual answer string diverges from the parametric answer string (e.g., the token ”States” vs. ”France”). By tracking the probability mass assigned to these specific tokens under different prompting conditions, we can rigorously quantify the model’s internal preference shift, even if the final output remains unchanged. As visualized in Figure 2, our analysis yields striking observations regarding the shifts in probability distributions and token rankings. With positive prompts, the probability distribution of contextual knowledge shifts significantly to the right, whereas the probability of parametric knowledge remains largely stable or exhibits a minor decline. Con- versely, negative prompts tend to increase the probability of parametric knowledge but exert negligible influence on contextual knowledge, effectively rendering the context in- visible. Critically, this distributional shift directly impacts behavioral realization: under positive prompts, there is a substantial surge in the frequency of contextual knowl- edge achieving Rank 1, thereby enabling the emergence of context-faithful behavior. In contrast, negative prompts reinforce the dominance of parametric outputs, solidifying the model’s reliance on internal priors. 3.3 Challenge of Behavioral Emergence Despite the clear distributional shift observed above, we identify a persistent phenomenon that hinders practical Algorithm 1 Knowledge Token Capturing Require: The LLM generates a token sequence of length n, V : vocabulary of the LLM, P i ∈ (P 1 ,P 2 ,...,P n ): the logits distribution for each token, S cont : string of the contextual answer (from the counterfactual context), S para : string of the parametric answer (ground truth). Ensure: Captured contextual knowledge logits P cont and parametric knowledge logits P para . 1: Initialize P cont ← None, P para ← None 2: S com ← COM(S cont ,S para ) // common substrings 3: for P i ∈ (P 1 ,P 2 ,...,P n ) do 4:Let x i ← arg maxP i and x ′ i ← Decode(x i ). 5:if x ′ i ∉ S cont and x ′ i ∉ S para then 6:continue 7:end if 8:for each token x j ∈ V (descending order by P i,j ) do 9:Decode x j into string x ′ j . 10:if x ′ j ∈ S com and P cont = P para = None: break 11:if x ′ j ∈ S cont and P cont = None: P cont ← P i,j 12:if x ′ j ∈ S para and P para = None: P para ← P i,j 13:end for 14: end for 15: return P cont , P para alignment: Latent Alignment does not guarantee Behavioral Emergence. While positive prompts significantly boost the confidence of LLMs regarding new knowledge derived from context, we find that there are still frequent instances where this context knowledge ranks prominently but fails to secure the Top-1 position, as illustrated in Figure 2. We term this phenomenon ”stubborn knowledge”, re- ferring to scenarios where faithful behavior fails due to either excessive confidence in existing parametric priors or insufficient confidence in the newly introduced knowledge. The cases in Figure 3 deeply reveal the failure pattern of faithful behavioral emergence under positive prompting in addressing stubborn knowledge. Specifically, failure often occurs when an extremely small gap remains between the target and the parametric knowledge, despite the significant increase in new knowledge logits induced by the positive prompt. Taking the last case in Figure 3 as an example, although the logit for the faithful answer (”English”) rises sharply, it still lags behind the entrenched parametric prior (”French”) by a distinct margin of 0.516 in terms of logit distribution. Under standard decoding strategies such as greedy search, the model fails to capitalize on this increased prob- ability and instead selects the parametric token, resulting in an unfaithful output. This finding highlights a critical limitation of standard prompting: it functions as a proba- bilistic bias rather than a deterministic control mechanism. Although the model demonstrates a distinct latent ten- dency towards faithfulness (evidenced by elevated logits), it often lacks sufficient probability mass to surpass the decision boundary imposed by robust parametric priors. This misalignment between latent representation and final behavior serves as the primary motivation for our work. It underscores the necessity for a decoding-time intervention capable of actively amplifying these latent probability shifts, JOURNAL OF L A T E X CLASS FILES, JANUARY 20255 Succ- essful Case parametriccontextual EuropeAsia Expected AnswerParametricChange logitsrank Input What is the capital of the country to which Lou Pearlman belonged? What's Roberto Merhi's country's official language? question United Kingdom is located inthe continent of Asia. counterfactual info 27.933→21.62520.547→22.3123→11→2 EgnlishItalian What‘s the official languagein scr- een International's home country? The official language of United Kingdom is Italian. 20.219→19.87510.461→20.17925→11→2 United States Bulgaria Marc Cherry is a citizen of Bulgaria. 16.641→12.2115.586→18.500186→11→4 Failed Case SpanishArabic 19.438→19.43517.435→18.14111→21→1 EuropeAustralia Which continent does Blur's origin lie in? London is located inthe continent of Australia. 27.391→22.73013.734→18.09412→31→1 FrenchEnglish The official language of France is English. 19.266→17.57812.211→17.0624→21→1 The official language of Spain is Arabic. logitsrank ContextualChange Case Type Which country is the creator of "Devious Maids" a citizen of? What is the official language of the country of Marcellin Champagnat? Fig. 3. Selected cases showing changes in the first token for both parametric and contextual knowledge. The cases are obtained by introducing positive prompts carrying counterfactual information into the LLAMA2-7B-CHAT model. ’→’ denotes the knowledge shift after incorporating positive prompts. ’Logits’ and ‘rank’ refer to the first discriminative token in the knowledge answer, indicating the model’s knowledge confidence. thereby transforming a tentative internal tendency into a stable, manifest behavioral realization. 4 POLARITY-PROMPT CONTRASTIVE DECODING Motivated by the findings in Section 3, we propose Polarity- Prompt Contrastive Decoding (PromptCD), a unified test- time steering framework designed to flexibly enhance tar- geted behaviors in LLMs. The core principle of PromptCD is to leverage polarity prompts—paired instructions that ex- plicitly encourage (positive prompt) or discourage (negative prompt) a desired behavior—to modulate the model’s infer- ence dynamics. Specifically, PromptCD isolates and ampli- fies the signals aligned with the target behavior by contrast- ing the model’s outputs conditioned on these opposing cues. This mechanism is versatile across modalities: for LLMs, it operates by contrasting token probability distributions to refine generation logits; for VLLMs, it extends to contrasting cross-modal attention maps to highlight behavior-relevant visual regions. We detail these two distinct implementations in the following subsections. 4.1 Contrastive Decoding via Token Logits Figure 4 conceptually illustrates the PromptCD workflow applied to LLMs. The fundamental intuition behind our approach is that the probability distribution generated by a large language model is a composition of task-specific reasoning and generic parametric priors. When conditioned on a positive prompt (e.g., specific instructions to be faith- ful), the output distributionP (P) contains both the desired behavioral signals and the model’s inherent biases. Con- versely, when conditioned on a negative prompt (e.g., in- structions to rely on internal knowledge or generic chitchat), the output distributionP (N) predominantly captures these inherent biases, hallucinations, or context-ignoring tenden- cies. PromptCD aims to distill the pure behavioral signal by contrastively subtracting the latter from the former. Formally, letP(x t ∣x <t ) denote the model’s predicted probability for the next token x t given the preceding The creator of WWE is . Who is the createrof WWE? Input Question: Please answer faithfully based on the provided context: the creator of WWE is Hoshino Gen. Positive Prompt: Concatenate Polarity Prompts Vince Ryohei GO Hoshino Keita Original distribution Positive distribution _ Enhanced distribution Vince Ryohei GO Hoshino Keita The creator of WWE is Vince McMahon. Final output: The creator of WWE is Hoshino Gen. Please answer the question based on your internal knowledge. Negative Prompt: Negative distribution Vince Ryohei GO Hoshino Keita Vince Ryohei GO Hoshino Keita Output w/ positive prompt LLM Original output Contrast Parametric:VinceMcMahon Contextual:HoshinoGen Fig. 4. Illustration of PromptCD enhancing context-faithfulness. By con- trasting the probability distributions induced by positive and negative prompts, PromptCD amplifies contextual signals, guiding the LLM to generate faithful responses to the input question. context x <t . At each decoding step t, we instantiate two parallel inference contexts: the sequence augmented with the positive prompt, denoted as X (P) , and the sequence with the negative prompt, X (N) . The model independently computes the conditional probability distributions for the next token under both settings:P (P) (x t ) =P(x t ∣X (P) <t ) and P (N) (x t ) =P(x t ∣X (N) <t ). 4.1.1 Contrastive Scoring Function. To effectively steer the generation toward the target behav- ior, we employ a contrastive scoring objective that ampli- fies tokens unique to the positive context while penalizing those that are highly probable in the negative context. The adjusted probability distribution ˆ P(x t ) is derived by nor- malizing the contrastive scores via a softmax operation: ˆ P(x t ) = softmax (F (P (P) (x t ),P (N) (x t ))),(6) where the operator F(⋅,⋅) computes the token-level logits modification. Drawing inspiration from Contrastive JOURNAL OF L A T E X CLASS FILES, JANUARY 20256 Decoding [27], we formulate this operator as a weighted subtraction in the log-probability space. This operation is mathematically equivalent to maximizing the mutual infor- mation between the token and the positive control signal, while minimizing the influence of the negative prior: F(⋅) = ⎧ ⎪ ⎪ ⎨ ⎪ ⎪ ⎩ logP (P) (x t ) −γ ⋅ logP (N) (x t ), if x t ∈ V head , −∞,otherwise, (7) where γ ≥ 0 is a critical hyperparameter serving as the contrastive coefficient. This coefficient controls the strength of the penalty applied to the negative distribution. Intu- itively, for high-frequency stop words or generic tokens whereP (P) ≈P (N) , the subtraction significantly dampens their scores. In contrast, for discriminative tokens that align strictly with the target behavior (highP (P) and lowP (N) ), the score is preserved or boosted. 4.1.2 Adaptive Plausibility Constraint (APC). While contrastive scoring effectively highlights discrimina- tive features, it introduces a risk of the “amateur correction” pathology. Specifically, a token that is extremely unlikely under the negative prompt (very large negative logP (N) ) might receive an artificially inflated score after subtraction, even if it is nonsensical in the positive context. To mitigate this and ensure the semantic fluency of the generated text, we adopt the Adaptive Plausibility Constraint (APC) [27]. This constraint imposes a dynamic validity mask, restrict- ing the candidate pool V head to only those tokens that are already plausible under the positive prompt: V head = x t ∈ V ∶P (P) (x t ) ≥ λ max w∈V P (P) (w),(8) where λ ∈ (0, 1] is a filtering threshold. By setting the scores of tokens outside this set to −∞, we ensure that the optimization landscape remains grounded within the valid semantic manifold of the task-compliant model, preventing the generation of gibberish. 4.1.3 Synchronized Dual-Sequence Decoding. A distinct structural feature of PromptCD, which differ- entiates it from standard re-ranking or layer-contrastive methods [28], is the simultaneous maintenance of two syn- chronized generation trajectories. Although the generation is guided by the contrastive distribution, the resulting token must maintain coherence in both prompt contexts to allow for iterative contrast. Let x ∗ t be the token sampled from the adjusted distribution ˆ P(x t ) (Eq. 6). Crucially, this single selected token is appended to both the positive and negative context buffers before proceeding to step t + 1: X (P) <t+1 ← X (P) <t ⊕x ∗ t , X (N) <t+1 ← X (N) <t ⊕x ∗ t .(9) This strict synchronization ensures thatP (P) andP (N) are always conditioned on identical generated histories. Without this synchronization, the two streams would rapidly diverge into semantically disjoint trajectories, ren- dering the contrastive subtraction meaningless. This mech- anism allows PromptCD to correct the model’s behavior token-by-token in real-time, effectively steering the entire generation path. Contrast Attention Write a general description of the image. Negative Prompt What shape is seen through the cups hand? Positive Prompt Attention Maps Circle Toomuchvisualnoise, butIknowwhicharea tolookin... Output Cup Output What shape is seen through the cups hand? <Question> <Original Image> Mask <Question> <Masked Image> Star 1 Star 2 Star Output The visual noise has been masked, so I can focus on the object easily. The high-attention area should be focus on. Fig. 5. The PromptCD framework for cross-modal attention refinement. By contrasting the task-specific attention distribution A (P) (derived from positive prompts) against the general attention pattern A (N) (de- rived from negative prompts), the method computes a refined attention map ˆ A. This contrastive mechanism effectively isolates and highlights request-relevant visual regions, thereby enhancing visual grounding. 4.2 Contrastive Decoding on Cross-Modal Attention Similar to LLMs, VLMs also face the challenge of behavioral alignment, particularly in how they attend to visual infor- mation. Visual noise often distracts the model from task- relevant regions, leading to hallucinations or unfaithful re- sponses. While token-level contrastive decoding effectively enhances textual generation, VLMs introduce an additional modality where PromptCD can be naturally extended: vi- sual attention patterns. In visual tasks, we define the target behavior as the VLM’s capability to precisely attend to task- relevant image regions. To steer this behavior, we introduce paired polarity prompts: a positive prompt, corresponding to a task-specific question, encourages the model to focus on specific regions required for reasoning; conversely, a nega- tive prompt, corresponding to a general instruction, elicits a broad, behavior-agnostic attention pattern over the entire image. As established in Equation 3, VLMs jointly attend to textual tokens and visual features through multimodal attention Attn text+vision (X <t ,V). We observe that polarity prompts not only modulate token probability distributions but also reshape the spatial allocation of visual attention—a phenomenon we leverage to enhance visual grounding in vision-language tasks. 4.2.1 Visual Attention Decomposition. As Equation 7 contrasts token logits, we contrast attention distributions over visual regions to decompose cross-modal attention into behavior-specific positive signals and generic negative patterns. Recall from Section 2.2 that A (P) l (I) and A (N) l (I) represent attention distributions at layer l under positive and negative prompts, respectively. We model the positive attention as a compositional interaction between the behavior-specific signal F (+) (P,I) induced by the pos- itive prompt and the generic visual pattern F (−) (I) shared across prompts: A (P) l (I) = F (−) (I) ⊗ F (+) (P,I),(10) JOURNAL OF L A T E X CLASS FILES, JANUARY 20257 where ⊗ denotes the Hadamard product. Under the nega- tive prompt, which provides only generic cues (e.g., ”Write a general description of the image”), the behavior-specific signal becomes approximately uniform, causing attention to pri- marily reflect the generic pattern: A (N) l (I) ≈ F (−) (I). 4.2.2 Contrastive Visual Attention To isolate the behavior-specific positive signal F (+) (P,I), we compute a contrastive attention map ˆ A l by normalizing the positive attention with respect to the negative attention: ˆ A l,i = A (P) l,i A (N) l,i +λ ,(11) where A (P) l,i and A (N) l,i denote the attention weights at the i-th visual token in layer l under positive and negative prompts respectively, and λ > 0 is a regularization param- eter. Substituting the decomposition from Equation 10, we obtain: ˆ A l,i ≈ F (−) i (I) ⋅ F (+) i (P,I) F (−) i (I) +λ ≈ F (+) i (P,I),(12) where the approximation holds when F (−) i (I) ≫ λ. Thus, the contrastive operation cancels out the generic neg- ative pattern F (−) (I), retaining only the behavior-specific positive signal F (+) (P,I)—analogous to how Equation 7 isolates behavior-favorable tokens in LLMs. 4.2.3 Application to Visual Question Answering. Having obtained the semantically refined contrastive atten- tion maps ˆ A l as illustrated in Figure 5, we now apply PromptCD to enhance visual grounding in VQA tasks. Positive prompts encourage focused visual reasoning over specific image regions, whereas negative prompts elicit generic scene descriptions without semantic focus. Since different layers capture complementary visual information, we fuse contrastive attention maps across the layer range L through weighted aggregation, where deeper layers re- ceive higher fusion weights as they encode more refined semantic information. Request-relevant visual regions are then identified by applying a top-p percentile threshold τ = Q p (S) to the fused attention map S, which retains the top p ∈ (0, 1] proportion of pixels. Connected compo- nent analysis extracts coherent regions from the thresholded map, and we select the top-K regions ranked by cumulative attention scores to generate the enhanced image through I refined = Φ(I,M ∗ ), where Φ applies masking, cropping, and resizing operations, and K controls the maximum number of regions to preserve. 5 EXPERIMENTAL METHODOLOGY 5.1 Setup for Text-Only Models We evaluate our PromptCD in terms of the “3H” dimensions for text-only LLMs to assess its effectiveness in test-time behavior enhancement. Our evaluation is conducted on four popular open-source LLMs: LLAMA2-7B-CHAT [37], LLAMA3-8B-INSTRUCT [2], MISTRAL-7B-INSTRUCT [38], and QWEN2.5-7B-INSTRUCT [3]. 5.1.1 Helpfulness Helpfulness of LLMs requires the model to faithfully follow human-provided contextual instructions. To this end, we conduct experiments using the Natural Questions (NQ) [31], ConFiQA [13], and CoConflictQA [32] datasets. In NQ, we further utilize the counterfactual information provided by the dataset. All datasets simulate knowledge conflicts through counterfactual contexts, aiming to examine whether the model can still generate helpful responses according to the given context when external instructions contradict its internal parametric knowledge. We use ConR (recall of context) and ParR (recall of parameters) as evaluation metrics. ConR measures whether the generated responses align with the provided context, while ParR evaluates their alignment with the model’s internal knowledge. Specifically, we also adopt the memorization ratio MR = ParR ParR+ConR , which captures the tendency to favor parameters over context. We use two prompt-based baselines: the Attributed (Attr) and the Opinion-based (Opin) prompt [22]. 5.1.2 Honesty We consider both multiple-choice and open-ended gener- ation tasks to evaluate LLM honesty. For the multiple- choice setting, we use TruthfulQA [33] to assess the model’s factual accuracy in short-answer scenarios. For open-ended generation, we adopt FActScore [34], which is rated by GPT- 4.1-nano. The remaining settings and configurations are consistent with those in the FActScore paper. We compare them with four baselines: 1) original decoding, 2) DoLa [28], decoding by contrasting layers, 3) SLED [39], monitoring self-logit evolution to penalize unstable generations. 4) RECITE [21], recitation-based factual generation and our PromptCD. In the experimental setup for DoLA, we adhered to the same configuration as in the original authors’ work, employing the DoLa-static mode throughout. 5.1.3 Harmlessness Harmlessness of LLMs requires the model to refuse gen- erating harmful content when confronted with adversar- ial inputs. To this end, we conduct experiments on the SafeEdit benchmark [35], covering 9 unsafe categories in- cluding offensiveness, bias, physical harm, mental harm, illegal activities, ethics violations, privacy breaches, pornog- raphy, and political sensitivity. The benchmark comprises 1,350 test samples, each containing adversarial prompts and corresponding generalization test inputs. We evaluate De- fense Success (DS) for the adversarial inputs in the test set, and Defense Generalization (DG) across multiple scenarios: DG onlyQ for harmful questions alone, DG otherA for different attack prompts, DG otherQ for different harmful questions, and DG otherAQ for different combinations of both. To assess the safety of generated responses, we employ the safety classifier from [35], which was trained on RoBERTa-large to classify LLM responses as either safe or unsafe. We also measure fluency via n-gram to ensure response natural- ness. We compare our PromptCD with several baselines: 1) Vanilla, the original aligned model, 2) FT-L, fine-tuning a single layer via causal analysis, 3) DINM [35], modifying toxic layers via semantic analysis with a single instance. JOURNAL OF L A T E X CLASS FILES, JANUARY 20258 TABLE 1 Performance (%) of PromptCD on LLM helpfulness evaluation. ModelMethod NQConFiQACoConflictQA ConR(↑)ParR(↓)MR(↓)ConR(↑)ParR(↓)MR(↓)ConR(↑)ParR(↓)MR(↓) LLAMA2-7B-CHAT Vanilla43.3543.7550.2469.6628.1428.7862.0718.7623.20 Attr64.3921.9925.4682.0318.7618.6165.4613.9717.58 Opin74.5915.3917.1188.6215.3814.7967.2613.5716.79 PromptCD75.996.397.7690.4213.7613.2170.2511.1613.71 LLAMA3-8B-INSTRUCT Vanilla43.9234.0543.4754.2322.3929.2268.4629.7430.29 Attr74.3941.3935.7587.8237.5229.9376.2422.3522.67 Opin69.9920.9923.0791.2122.9520.1073.8518.5620.08 PromptCD72.696.598.3390.9219.5617.7082.4315.7616.05 MISTRALV0.3-7B-INSTRUCT Vanilla46.1958.5955.9264.6825.8728.5767.0619.9622.94 Attr71.5925.5926.3479.6418.5618.9067.0716.3719.62 Opin73.3916.9918.8183.2315.9716.1070.4517.1619.59 PromptCD86.193.994.4384.5913.9714.1772.6513.7715.93 QWEN2.5-7B-INSTRUCT Vanilla73.3932.3931.2743.7815.4226.0575.8424.1524.15 Attr82.3917.5917.5990.0133.1326.9075.4420.7521.57 Opin82.5921.7920.8889.0233.9327.6076.2421.7522.20 PromptCD90.3914.3913.7491.8229.5424.3477.6418.1518.95 TABLE 2 Performance (%) of PromptCD on LLM honesty evaluation. ModelMethod TruthfulQA(MC)FActScore MC1MC2MC3Score#facts Vanilla33.6651.2924.9133.2025.60 LLAMA2- DoLa32.6861.1229.8635.6054.10 7B-CHAT SLED37.0863.8332.9236.0032.40 RECITE32.1950.3825.0636.1032.40 PromptCD42.5964.0533.8936.6046.40 Vanilla31.9548.8323.9234.2020.80 LLAMA3- DoLA39.1668.2735.6941.8054.30 8B-INSTRUCT SLED41.3769.1737.8830.4014.60 RECITE41.9861.3233.1137.3028.60 PromptCD49.2169.5943.8344.5063.40 Vanilla48.5966.2437.4544.5022.5 MISTRALV0.3- DoLA41.1267.9937.5146.2040.5 7B-INSTRUCT SLED45.4168.4340.3546.3030.5 RECITE49.7968.8939.4046.3030.5 PromptCD50.7969.5844.2447.3026.7 Vanilla46.8865.1535.2229.6026.90 QWEN2.5- DoLA35.7458.7630.6826.2038.10 7B-INSTRUCT SLED44.9270.3539.8233.2027.2 RECITE47.2467.6136.9030.8044.30 PromptCD54.9571.1345.9334.3029.00 5.2 Setup in Vision-Language Models We evaluate our method on four benchmark VQA datasets: A-OKVQA [40], POPE [41], V ∗ [42], and TextVQA [43]. These datasets cover a diverse spectrum of capabilities, in- cluding visual reasoning, perception, and knowledge acqui- sition. Notably, for TextVQA, we assess the models’ intrinsic visual text recognition skills by relying solely on image- question pairs without external OCR augmentation. Our experiments employ four VLMs: QWEN2.5-VL-INSTRUCT (3B and 7B) [44] and LLAVA-1.5 (7B and 13B) [45]. Regard- ing input resolution, the Qwen series processes images at 448 × 448, while the LLaVA-1.5 series operates at 336 × 336. Greedy decoding is applied across all experiments. TABLE 3 Performance (%) of PromptCD on LLM harmlessness evaluation. MethodMode DefenseDefense Generalization Fluency SuccessonlyQotherAotherQotherAQ Vanilla40.5976.9622.4141.3321.787.01 LLAMA2-FT-L42.5276.2228.3742.8928.855.94 7B-CHATDINM90.0089.9359.2687.4157.636.61 PromptCD96.6798.6782.3096.5982.377.29 Vanilla12.1554.8118.0712.0718.747.46 LLAMA3-FT-L11.7054.9630.5911.7830.527.16 8B-INSTRUCT DINM99.2699.5699.5699.1199.851.50 PromptCD95.2692.9684.8195.7883.855.95 Vanilla9.8534.449.049.788.447.66 MISTRALV0.3FT-L10.6736.6719.0010.2218.446.65 7B-INSTRUCTDINM35.7873.4826.3033.7824.896.82 PromptCD38.6786.6736.4439.7837.047.14 Vanilla17.6372.597.1110.447.048.01 QWEN2.5FT-L12.3074.0024.5911.2624.897.69 7B-INSTRUCTDINM26.7484.2234.5924.5233.787.78 PromptCD76.0088.8976.8977.7876.966.44 6 EVALUATION RESULTS 6.1 Evaluation in Text-Only Models The performance across three distinct behavioral dimen- sions—helpfulness, honesty, and harmlessness—is system- atically reported in Tables 1, 2, and 3, respectively. Addition- ally, Figure 6 visualizes the average improvements across these “3H” dimensions for different models. Helpfulness. Our PromptCD demonstrates superior ca- pability in adhering to context instructions, particularly un- der conflicting scenarios. It achieves average improvements of 62.64%, 59.49%, and 11.08% in terms of ConR across the three datasets. Notably, PromptCD consistently maintains the lowest parametric recall (ConR) and memorization ratio (MR). This underscores its advantage as a decoding-time strategy over prompt-based baselines, as it effectively sup- presses the probability of undesired parametric generations. Honesty. Decoding strategies, specifically PromptCD and SLED, significantly enhance multiple-choice truth- fulness compared to baseline and prompting methods. PromptCD secures the top performance, achieving the high- est MC1 scores on models such as LLaMA2-7B-Chat (42.6%) JOURNAL OF L A T E X CLASS FILES, JANUARY 20259 78.89 51.78 89.98 (a) LLaMA2-7B 82.01 51.78 89.35 (b) LLaMA3-8B 81.14 49.98 52.98 (c) Mistral0.3-7B 86.62 80.13 51.18 (d) Qwen2.5-7B VanillaBaselinePromptCD Fig. 6. Average performance across the three “3H” dimensions on dif- ferent models. PromptCD consistently yields substantial improvements over the original models, while the baselines (Opin, DoLa, and DINM) show relatively limited gains. and Qwen2.5-7B-Instruct (54.9%), whereas SLED shows moderate improvements and RECITE provides inconsistent gains. Furthermore, PromptCD leads in open-ended gen- eration metrics (FActScore), surpassing both the Baseline and DoLa on LLaMA3-8B-Instruct (44.5%) and Mistral-7B- Instruct (47.3%). While DoLa tends to generate more atomic facts, it often sacrifices precision. Conversely, SLED main- tains strong factual consistency, and RECITE offers slight improvements in precision. Harmlessness. PromptCD exhibits robust safety align- ment across various models while preserving response flu- ency—a trade-off often mishandled by other knowledge editing methods. On LLaMA2-7B-Chat, PromptCD achieves the best results with a Defense Success (DS) of 96.67% and an average DG performance of 82.37%, substantially outper- forming the vanilla baseline (40.59% DS) while maintaining excellent fluency (7.29). Similarly, on Mistral-7B-Instruct- v0.3, PromptCD shows superior performance with a DS of 38.67% compared to the vanilla 9.85%. Although DINM achieves higher defense success rates on Meta-Llama-3-8B- Instruct, it does so at the cost of substantially degraded fluency (1.50 vs. 5.95). Overall, PromptCD effectively bal- ances harmlessness with generation quality, positioning it as a practical solution for detoxifying LLMs without com- promising naturalness. In summary, PromptCD achieves consistently superior performance across all three key behavioral dimensions of LLMs. Unlike existing methods that are restricted to a single behavioral objective, PromptCD offers the flexibility to switch target behaviors simply by adjusting the polar- ity prompt, thereby enhancing both the effectiveness and adaptability of test-time intervention. 6.2 Evaluation in Vision-Language Models Table 4 demonstrates PromptCD’s consistent perfor- mance enhancement across all evaluated VLMs and VQA datasets through cross-modal contrastive attention re- finement. Earlier-generation models exhibit substantially greater improvements than recent counterparts: LLAVA- 1.5-7B achieves 71.83% relative improvement on V ∗ ver- sus 14.37% for QWEN2.5-VL-7B, and 21.76% gain on TextVQA (47.8%→ 58.2%) versus 8.93% for QWEN2.5-VL- 7B (75.0%→ 81.7%). Performance gains vary significantly across datasets: V ∗ shows the most substantial improve- ments, with LLAVA-1.5-13B achieving 70.0% versus 42.4% baseline (65.09% relative improvement), while POPE ex- hibits moderate gains (90.1% from 84.6%, 6.50% improve- ment). This suggests attention refinement particularly ben- TABLE 4 Performance (%) of PromptCD on VLMs across different layer intervals. ModelLayersLA-OKVQAPOPEV ∗ TextVQA QWEN2.5-VL-3B Vanilla73.086.950.372.8 [10, 15]74.086.953.473.0 [15, 20]76.887.756.076.0 [20, 25]78.387.957.176.3 QWEN2.5-VL-7B Vanilla75.087.050.875.0 [10, 15]75.087.051.375.0 [15, 20]77.188.457.679.5 [20, 25]78.088.658.181.7 LLAVA-1.5-7B Vanilla71.583.638.747.8 [10, 15]71.584.548.249.2 [15, 20]74.287.565.456.4 [20, 25]75.489.066.558.2 LLAVA-1.5-13B Vanilla75.784.642.457.1 [10, 15]75.785.052.957.4 [15, 20]76.888.669.159.4 [20, 25]76.990.170.061.2 TABLE 5 Ablation results of the adjustment coefficient γ on LLAMA2-7B-CHAT models across the ‘3H’ dimensions. Performance is reported using the average values of ConR, MC-AVG, and DG-AVG, respectively. Modelγ= 0.2 γ= 0.5 γ= 0.8 Helpfulness (ConR)77.3278.8970.53 Honesty (MC-AVG)39.5946.8444.87 Harmlessness (DG-AVG)83.2989.9885.34 efits tasks requiring precise spatial localization and fine- grained visual grounding. Models with limited visual rea- soning capabilities suffer more from attention dispersion over task-irrelevant regions, thereby benefiting more from PromptCD’s contrastive attention-guided focusing mecha- nisms. As shown in Figure 7, PromptCD effectively gen- erates high-fidelity attention masks that filter background noise, shifting focus from broad, distracted regions to pre- cise, request-relevant visual cues as threshold τ tightens, thereby correcting hallucinations. 6.3 Ablation Study We investigate the sensitivity of PromptCD to the adjust- ment coefficient γ, which governs the magnitude of the penalty applied to the negative logits (i.e., the strength of suppression). For all reported experiments on text-only LLMs, we uniformly set γ to 0.5. The detailed ablation results are presented in Table 5. We observe that perfor- mance follows an inverted U-shape trend: an excessively small γ fails to sufficiently amplify the contrast between positive and negative contexts, while an overly large γ may disrupt the model’s linguistic coherence. The results empirically confirm that a moderate γ yields the optimal balance, effectively steering the behavior without degrading generation quality. For vision tasks, as shown in Table 4, layer-wise anal- ysis reveals a consistent performance ordering across all VLMs: [20, 25], [15, 20], and [10, 15]. Taking LLAVA- 1.5-7B on TextVQA as an example, deep layers [20, 25] achieve 21.76% improvement, middle layers [15, 20] reach 17.99%, while early layers [10, 15] attain only 2.93%. This JOURNAL OF L A T E X CLASS FILES, JANUARY 202510 TABLE 6 Performance comparison of PromptCD against external tool-based approaches and ViCrop on TextVQA dataset: accuracy (%) and inference time (seconds) overhead per sample. MethodOriginalSAMYOLOCLIP ViCrop PromptCD rel-att grad-att pure-grad Accuracy47.8049.4248.8448.5555.1756.0651.6758.2 GPU Time0.173.330.351.071.160.892.361.34 124 Masked Masked & Cropped Masked Masked & Cropped Mask Q:Which gallery will they be found in? Output: 116 blue star Output: gallery 4Output: gallery 4 Contrastive Attention Map 휏= 0.6휏= 0.3휏= 0.1 Mask Q: What brand of vinegar is in the bottles on the left? Output: infuser Output: infused Output: infused Contrastive Attention Map 휏= 0.6휏= 0.3휏= 0.1 Fig. 7. Visualization of contrastive attention refinement in VQA. As the threshold τ decreases from 0.6 to 0.1, PromptCD progressively sup- presses visual noise to isolate request-relevant regions. Our PromptCD precise grounding corrects initial hallucinations (Red ×) and yields ac- curate answers (Green ✓). hierarchical pattern aligns with established attention mech- anisms: early layers perform global visual scanning with high entropy, while deep layers progressively focus on task- relevant semantic patterns, making them optimal for ex- tracting request-consistent visual attention through polarity- prompt contrastive decoding. The optimal configuration [20, 25] consistently outperforms other intervals with absolute accuracy improvements ranging from 1.2% to 27.8%, vali- dating the effectiveness of deep-layer attention extraction. 6.4 Inference Latency Analysis Since PromptCD necessitates dual forward passes at each decoding step, it inevitably incurs computational overhead. Table 7 reports the inference latency comparison evaluated on the NQ dataset, indicating that our method results in a 1.6×–1.8× overhead in latency and throughput compared to standard decoding. While this cost is non-negligible, we argue that it is a justifiable trade-off for safety-critical and high-stakes applications—such as those in medical, financial, or legal domains. In these scenarios, factual errors, unmanaged knowledge conflicts, or safety violations can have severe consequences. Therefore, the marginal increase in computational cost is outweighed by the substantial gains in reliability, trustworthiness, and the capability for precise, human-in-the-loop behavioral governance. For vision-language tasks, Table 6 compares PromptCD against external tools such as SAM [46], YOLO [47], TABLE 7 Comparison of inference latency and throughput for PromptCD. ModelMethodLatency (ms/token)Throughput (token/s) LLAMA2-7B Vanilla36.03 (×1.00)27.76 (×1.00) PromptCD64.85 (×1.79)16.64 (×0.60) LLAMA2-13B Vanilla51.41 (×1.00)19.45 (×1.00) PromptCD89.45 (×1.74)12.25 (×0.63) LLAMA3-8B Vanilla42.17 (×1.00)23.92 (×1.00) PromptCD69.77 (×1.65)14.08 (×0.59) TABLE 8 Ablation study on the NQ dataset evaluating the effect of the Adaptive Plausibility Constraint (APC) on generation quality (Hit Rate). ModelVanilla PromptCD w/ APCw/o APC LLAMA2-7B78.977.830.7 LLAMA3-8B83.785.553.8 MISTRALV0.3-7B92.491.749.9 QWEN2.5-7B88.889.862.4 CLIP [48] and ViCrop [49]. PromptCD achieves the highest accuracy with competitive inference times, establishing a favorable balance between visual grounding precision and computational cost. 6.5 Robustness against Cross-Task Interference A critical concern in alignment interventions is the ”align- ment tax”—specifically, whether enhancing a target behav- ior comes at the cost of degrading the model’s general gen- eration quality. To investigate this, we conduct an analysis using the NQ dataset under the Helpfulness setting. We define a metric termed Hit Rate, which measures whether the model’s output contains either the original parametric answer or the correct gold answer provided in the context. This metric effectively serves as a proxy for linguistic co- herence and meaningfulness; a drop in Hit Rate implies the model is generating irrelevant or hallucinated content (gibberish) rather than valid answers. The results are summarized in Table 8. We observe that without Adaptive Plausibility Constraint (APC), the Hit Rate drops significantly. This suggests that unconstrained logit adjustments can severely disturb the model’s distribu- tion, causing it to deviate from both its internal knowledge and the external context. In contrast, when equipped with APC, the performance remains comparable to the Vanilla baseline. This finding demonstrates that APC effectively acts as a safety filter, restricting PromptCD’s intervention to plausible tokens and preserving the model’s fundamental ability to generate coherent and relevant responses. JOURNAL OF L A T E X CLASS FILES, JANUARY 202511 TABLE 9 Experimental results (accuracy %) using LLAMA2-CHAT models. We conduct experiments with full-batch edit memory to evaluate the performance of memory-based Knowledge Editing. ModelMethod MQUAKE 3K2002HARD LLAMA2-7B-CHAT IKE20.720.62.3 + PromptCD22.420.43.8 MeLLo32.640.85.1 + PromptCD43.145.85.8 LLAMA2-13B-CHAT IKE19.418.82.7 + PromptCD20.618.43.5 MeLLo33.435.93.9 + PromptCD36.838.26.2 IKEIKE w PromptCD LLaMA2-7B LLaMA2-13B LLaMA3-8B Mistral-7B 0 20 40 60 80 100 85.4 63.8 31.6 34.1 91.3 84.6 54.7 46.7 MQuAKE-3k LLaMA2-7B LLaMA2-13B LLaMA3-8B Mistral-7B 85.1 64.1 32.5 35.6 89.4 84.4 55.9 48.5 MQuAKE-2002 LLaMA2-7B LLaMA2-13B LLaMA3-8B Mistral-7B 88.9 55.2 14.3 15.6 98.6 89.7 45.7 19.2 MQuAKE-hard Fig. 8. Performance comparison (%) between standard IKE and IKE enhanced with PromptCD across various models and datasets. We set the edit memory batch size to 1 to evaluate the fundamental capability of direct knowledge editing. 7 APPLICATIONS AND INTERPRETABILITY 7.1 Application to Knowledge Editing Although LLMs accumulate substantial factual knowledge during pretraining, this parametric information often be- comes outdated, thereby compromising model reliabil- ity [50–52]. Given the prohibitive cost of retraining, Knowl- edge Editing (KE) [53–57] has emerged as a vital technique to update LLMs by injecting new information or modifying existing parameters. In this section, we apply our method to In-Context Editing (ICE) [58–60], a non-invasive paradigm widely adopted in Retrieval-Augmented Generation (RAG) systems [61, 62] that corrects outdated knowledge solely through context without altering model weights. To assess the practical utility of PromptCD in realistic application scenarios, we extend our evaluation to the task of ICE, integrating our method into two representative base- lines: IKE [58], which guides LLMs to edit knowledge via contextual demonstrations; and MeLLo [36], which handles multi-hop editing by decomposing complex questions and retrieving conflicting information iteratively. Our experi- ments are conducted on MQUAKE-3K [36] and the more challenging MQUAKE-2002/HARD [63]. Figure 8 illustrates the performance of IKE with a mem- ory batch size of 1. We observe a counter-intuitive phe- nomenon: as model size or training strength increases, the model’s parametric knowledge becomes more ”stubborn,” leading to a decline in standard ICE accuracy. This suggests that stronger models suffer more from the conflict between intrinsic priors and external context. In contrast, IKE en- 00.51.0 Probability on Stubborn∏33% 0 0.4 0.8 1.2 Density 1510 Rank on Stubborn∏33% 0.0 0.1 0.2 0.3 0.4 Frequency IKEIKEw/ PrompCDIKEIKEw/ PrompCD 00.51.0 Probability on Stubborn∏67% 0 0.4 0.8 1.2 Density 1510 Rank on Stubborn∏67% 0.0 0.1 0.2 0.3 Frequency Fig. 9. Probability statistics of edited new knowledge on the MQUAKE- STUBBORN dataset using LLAMA2-7B-CHAT. The probabilities are com- puted by applying softmax to the model’s token logits. hanced with PromptCD consistently outperforms the vanilla version across all model scales, indicating that our approach effectively mitigates the resistance of stubborn knowledge, allowing even strong models to adapt to new information. Table 9 presents the results of full-batch experiments (1,000 instances), simulating a realistic RAG retrieval en- vironment. PromptCD yields consistent improvements for both IKE and MeLLo. These results underscore the potential of PromptCD for modern RAG systems, demonstrating its ability to enhance both faithfulness and controllability without requiring architectural modifications. 7.2 Metamorphosis of Stubborn Knowledge To understand why PromptCD succeeds where standard methods fail, we revisit the critical challenge identified in Section 3: the misalignment between Latent Alignment and Behavioral Emergence. As defined previously, ”stub- born knowledge” represents a failure state where the iner- tia of parametric priors suppresses the realization of new knowledge. We investigate the internal dynamics of this phenomenon using the STUBBORN dataset (derived from MQUAKE-3K) and apply the Knowledge Token Capturing Algorithm (Algorithm 1) from Section 3.2 to trace the rank and probability mass of the first distinctive faithful token. Figure 9 visualizes the distribution of these key tokens, revealing a distinct ”Metamorphosis” of the probability landscape. Under the standard ICE setting, consistent with our earlier findings, the faithful tokens often languish in the ”long tail” or hover just below the decision threshold. This empirical evidence reconfirms that standard prompting functions merely as a weak probabilistic bias, often insuffi- cient to dislodge confident parametric priors. In sharp con- trast, the application of PromptCD induces a radical shift in the distribution toward high-probability regions. By actively modulating the logits based on the contrast between posi- tive and negative signals, our method effectively ”pushes” these latent faithful tokens across the decision boundary. This intervention significantly increases the likelihood of decoding context-faithful tokens at critical positions, re- sulting in a marked rise in the frequency of high-ranking JOURNAL OF L A T E X CLASS FILES, JANUARY 202512 generations. The analysis confirms that PromptCD functions exactly as the decisive control mechanism hypothesized in our motivation: it actively amplifies the latent tendency of the model, transforming tentative internal signals into stable, manifest behavioral realizations. 8 RELATED WORK 8.1 Behavioral Alignment of LLMs The alignment of LLMs with human preferences and values has emerged as a central research focus [5, 64]. Domi- nant approaches primarily rely on Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), alongside recent variants like Direct Preference Optimization (DPO) [14, 16, 65]. While these methods have proven effective in improving helpfulness and safety [66], and multi-objective frameworks attempt to address trade- offs among honesty and harmlessness [67, 68], they suffer from inherent rigidity and high computational costs [69, 70]. Specifically, once aligned, the model’s behavioral patterns are frozen, making it difficult to adaptively switch be- haviors for diverse scenarios without expensive retrain- ing. Consequently, recent research has shifted towards lightweight, training-free alternatives [71, 72]. Strategies such as self-critique, self-consistency, and uncertainty-based signals have been proposed to provide weak supervision or refinement [73, 74]. Furthermore, large-scale audits con- tinue to reveal persistent data and value biases [75, 76], underscoring the necessity for more flexible, inference-stage intervention mechanisms. 8.2 Test-Time Behavior Enhancement Beyond these training-free strategies, a complementary line of work focuses on test-time interventions that adjust model behaviour dynamically during inference [10, 27, 77]. Such methods operate by modulating inputs, representations, or decoding dynamics without retraining. Prompt-based steer- ing augments inputs with behavioural cues or exemplars to guide responses [78, 79], whereas activation editing directly manipulates internal states to induce desired traits [26, 80]. More recent decoding-time alignment methods alter token probabilities or logits during generation [28, 81]. However, these approaches are typically tailored to fixed objectives (e.g., factuality or safety) and lack a unified mechanism for flexible behavioural steering. Building on this insight, our method introduces polarity prompts to enable general, low-cost alignment control through contrastive guidance at decoding time. 8.3 Contrastive Mechanisms in Generation Contrastive Decoding (CD) has emerged as a powerful paradigm to enhance generation quality by amplifying the difference between a reliable distribution and a defective one. Existing CD methods primarily vary in their source of contrast: Model-level contrast (e.g., CD [27]) contrasts expert models with amateur models to improve fluency. Layer-level contrast (e.g., DoLa [28], SLED [39]) contrasts mature final layers with premature early layers to isolate factual knowledge. Context-level contrast (e.g., ICD [82]) contrasts standard context with hallucinated context to suppress errors. Other approaches like TruthX [25] and HACL [83] apply contrastive principles within the represen- tation space during training or inference to mitigate hallu- cinations. PAC[84] traces critical transmission paths across all layers, treating less important pathways as negatives, while UniKE [85] extends this idea to multimodal LLMs through semantic truthfulness space disentanglement. Dif- ferent from these works, PromptCD introduces Instruction- level contrast. By constructing explicit Positive and Neg- ative prompts, we dynamically synthesize the contrasting distributions required for decoding. This eliminates the need for auxiliary models (as in CD) or layer selection heuristics (as in DoLa), offering a unified framework to steer arbitrary behaviors defined by natural language. 9 CONCLUSION Inthiswork,weproposePromptCD,apolarity- prompt–based contrastive decoding framework for enhanc- ing target behaviors at test time. PromptCD constructs paired positive and negative guiding prompts for a speci- fied behavioral objective and contrasts the resulting prob- ability distributions during decoding to amplify tokens aligned with the desired behavior. Extensive experiments on the “3H” alignment dimensions—helpfulness, honesty, and harmlessness—demonstrate that PromptCD effectively strengthens each target behavior and generalizes well across tasks without requiring retraining. Moreover, its application to knowledge editing further highlights its interpretability and performance advantages. Overall, this work paves the way toward integrating both behavioral alignment and flex- ible test-time enhancement in large language models. REFERENCES [1] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Alt- man, S. Anadkat etal., “Gpt-4 technical report,”arXiv preprintarXiv:2303.08774, 2023. [2] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Ka- dian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan etal., “The llama 3 herd of models,”arXiv preprintarXiv:2407.21783, 2024. [3] A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei etal., “Qwen2. 5 tech- nical report,”arXivpreprintarXiv:2412.15115, 2024. [4] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, “On the dangers of stochastic parrots: Can language models be too big?” in Proceedingsof the2021ACMconferenceonfairness,accountability, andtransparency, 2021, p. 610–623. [5] I. Gabriel, “Artificial intelligence, values, and align- ment,” Mindsandmachines, vol. 30, no. 3, p. 411–437, 2020. [6] Y. Liu, Y. Yao, J.-F. Ton, X. Zhang, R. Guo, H. Cheng, Y. Klochkov, M. F. Taufiq, and H. Li, “Trustworthy llms: a survey and guideline for evaluating large language models’ alignment,” arXivpreprintarXiv:2308.05374, 2023. JOURNAL OF L A T E X CLASS FILES, JANUARY 202513 [7] S. Zhang, D. Yu, H. Sharma, H. Zhong, Z. Liu, Z. Yang, S. Wang, H. Hassan, and Z. Wang, “Self-exploring lan- guage models: Active preference elicitation for online alignment,” arXivpreprintarXiv:2405.19332, 2024. [8] I. O. Pappas and P. Vassilakopoulou, “Human value alignment in ai,” in HandbookofHuman-Centered ArtificialIntelligence. Springer, 2025, p. 1–33. [9] A. Bhargava, C. Witkowski, S.-Z. Looi, and M. Thom- son, “What’s the magic word? a control theory of llm prompting,” arXivpreprintarXiv:2310.04444, 2023. [10] B. Bi, S. Liu, Y. Wang, Y. Xu, J. Fang, L. Mei, and X. Cheng, “Parameters vs. context: Fine-grained con- trol of knowledge reliance in language models,” arXiv preprintarXiv:2503.15888, 2025. [11] Y.Yang,E.Chern,X.Qiu,G.Neubig,and P. Liu, “Alignment for honesty,”arXivpreprint arXiv:2312.07000, 2023. [12] K. Tian, E. Mitchell, H. Yao, C. D. Manning, and C. Finn, “Fine-tuning language models for factuality,” arXivpreprintarXiv:2311.08401, 2023. [13] B. Bi, S. Huang, Y. Wang, T. Yang, Z. Zhang, H. Huang, L. Mei, J. Fang, Z. Li, F. Weietal., “Context-dpo: Align- ing language models for context-faithfulness,”arXiv preprintarXiv:2412.15280, 2024. [14] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wain- wright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray etal., “Training language models to follow in- structions with human feedback,”Advancesinneural informationprocessingsystems, vol. 35, p. 27 730– 27 744, 2022. [15] Y. Bai, S. Kadavath, A. Askell etal., “Training a help- ful and harmless assistant with rlhf,”arXivpreprint arXiv:2204.05862, 2022. [16] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advancesinneuralinformationprocessingsystems, vol. 36, p. 53 728–53 741, 2023. [17] Y. Guo, G. Cui, L. Yuan, N. Ding, Z. Sun, B. Sun, H. Chen, R. Xie, J. Zhou, Y. Lin, Z. Liu, and M. Sun, “Controllable preference optimization: Toward controllable multi-objective alignment,” 2024. [Online]. Available: https://arxiv.org/abs/2402.19085 [18] R. Alami, A. K. Almansoori, A. Alzubaidi, M. E. A. Seddik, M. Farooq, and H. Hacid, “Alignment with preference optimization is all you need for llm safety,” arXivpreprintarXiv:2409.07772, 2024. [19] N. Muennighoff, Z. Yang, W. Shi, X. L. Li, L. Fei- Fei, H. Hajishirzi, L. Zettlemoyer, P. Liang, E. Cand ` es, and T. Hashimoto, “s1: Simple test-time scaling,”arXiv preprintarXiv:2501.19393, 2025. [20] C. Snell, J. Lee, K. Xu, and A. Kumar, “Scaling llm test-time compute optimally can be more effec- tive than scaling model parameters,”arXivpreprint arXiv:2408.03314, 2024. [21] Z. Sun, X. Wang, Y. Tay, Y. Yang, and D. Zhou, “Recitation-augmentedlanguagemodels,” arXiv preprintarXiv:2210.01296, 2022. [22] W. Zhou, S. Zhang, H. Poon, and M. Chen, “Context- faithful prompting for large language models,”arXiv preprintarXiv:2303.11315, 2023. [23] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhouetal., “Chain-of-thought prompting elicits reasoning in large language models,” Advancesinneuralinformationprocessingsystems, vol. 35, p. 24 824–24 837, 2022. [24] K. Li, O. Patel, F. Vi ́ egas, H. Pfister, and M. Wattenberg, “Inference-time intervention: Eliciting truthful answers from a language model,” 2024. [Online]. Available: https://arxiv.org/abs/2306.03341 [25] S. Zhang, T. Yu, and Y. Feng, “Truthx: Alleviating hallu- cinations by editing large language models in truthful space,” arXivpreprintarXiv:2402.17811, 2024. [26] L. Kong, H. Wang, W. Mu, Y. Du, Y. Zhuang, Y. Zhou, Y. Song, R. Zhang, K. Wang, and C. Zhang, “Aligning large language models with representation editing: A control perspective,” AdvancesinNeuralInformation ProcessingSystems, vol. 37, p. 37 356–37 384, 2024. [27] X. L. Li, A. Holtzman, D. Fried, P. Liang, J. Eisner, T. Hashimoto, L. Zettlemoyer, and M. Lewis, “Con- trastive decoding: Open-ended text generation as op- timization,” arXivpreprintarXiv:2210.15097, 2022. [28] Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. Glass, and P. He, “Dola: Decoding by contrasting layers improves factuality in large language models,” arXivpreprint arXiv:2309.03883, 2023. [29] W. Shi, X. Han, M. Lewis, Y. Tsvetkov, L. Zettlemoyer, and W.-t. Yih, “Trusting your evidence: Hallucinate less with context-aware decoding,” in Proceedingsofthe 2024ConferenceoftheNorthAmericanChapterof theAssociationforComputationalLinguistics:Human LanguageTechnologies(Volume2:ShortPapers), 2024, p. 783–791. [30] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C.McKinnon,C.Chen,C.Olsson,C.Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran- Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan, “Constitutional ai: Harmlessness from ai feedback,” 2022. [Online]. Available: https: //arxiv.org/abs/2212.08073 [31] T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. De- vlin, K. Lee etal., “Natural questions: a benchmark for question answering research,” Transactionsofthe AssociationforComputationalLinguistics, vol. 7, p. 453–466, 2019. [32] P. Huang, Z. Liu, Y. Yan, X. Yi, H. Chen, Z. Liu, M. Sun, T. Xiao, G. Yu, and C. Xiong, “Pip-kag: Miti- gating knowledge conflicts in knowledge-augmented generation via parametric pruning,” arXivpreprint arXiv:2502.15543, 2025. [33] S. Lin, J. Hilton, and O. Evans, “Truthfulqa: Measuring how models mimic human falsehoods,”arXivpreprint arXiv:2109.07958, 2021. JOURNAL OF L A T E X CLASS FILES, JANUARY 202514 [34] S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. W. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi, “Factscore: Fine-grained atomic evaluation of factual precision in long form text generation,” arXivpreprint arXiv:2305.14251, 2023. [35] M. Wang, N. Zhang, Z. Xu, Z. Xi, S. Deng, Y. Yao, Q. Zhang, L. Yang, J. Wang, and H. Chen, “Detoxifying large language models via knowledge editing,” arXiv preprintarXiv:2403.14472, 2024. [36] Z. Zhong, Z. Wu, C. D. Manning, C. Potts, and D. Chen, “Mquake: Assessing knowledge editing in language models via multi-hop questions,” arXiv preprintarXiv:2305.14795, 2023. [37] H. Touvron, L. Martin, K. Stone, P. Albert, A. Alma- hairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale etal., “Llama 2: Open foundation and fine- tuned chat models,” arXivpreprintarXiv:2307.09288, 2023. [38] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mistral 7b,” 2023. [Online]. Available: https://arxiv.org/abs/2310.06825 [39] J. Zhang, D.-C. Juan, C. Rashtchian, C.-S. Ferng, H. Jiang, and Y. Chen, “Sled: Self logits evolution decoding for improving factuality in large language models,” in TheThirty-eighthAnnualConferenceon NeuralInformationProcessingSystems(NeurIPS2024, 2024. [Online]. Available: https://arxiv.org/abs/2411. 02433 [40] D. Schwenk, A. Khandelwal, C. Clark, K. Marino, and R. Mottaghi, “A-okvqa: A benchmark for visual ques- tion answering using world knowledge,” in Computer Vision–ECCV2022:17thEuropeanConference,Tel Aviv,Israel,October23–27,2022,Proceedings,Part VIII. Springer, 2022, p. 146–162. [41] Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision- language models,” arXivpreprintarXiv:2305.10355, 2023. [42] P. Wu and S. Xie, “V*: Guided visual search as a core mechanism in multimodal llms,” arXivpreprint arXiv:2312.14135, 2023. [43] A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach, “Towards vqa models that can read,” in ProceedingsoftheIEEE/CVF conferenceoncomputervisionandpatternrecognition, 2019, p. 8317–8326. [44] Qwen, “Qwen2.5-vl: A powerful vision-language model for seamless computer interaction,” arXiv preprintarXiv:2409.12191, 2025. [45] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,”arXivpreprintarXiv:2304.08485, 2023. [46] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.- Y. Lo etal., “Segment anything,” inProceedingsof theIEEE/CVFinternationalconferenceoncomputer vision, 2023, p. 4015–4026. [47] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object de- tection,” in ProceedingsoftheIEEEconferenceon computervisionandpatternrecognition, 2016, p. 779– 788. [48] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark etal., “Learning transferable visual models from natu- ral language supervision,” inInternationalconference onmachinelearning. PmLR, 2021, p. 8748–8763. [49] J. Zhang, M. Khayatkhoei, P. Chhikara, and F. Ilievski, “Mllms know where to look: Training-free perception of small visual details with multimodal llms,” arXiv preprintarXiv:2502.17422, 2025. [50] C. Chen and K. Shu, “Combating misinformation in the age of llms: Opportunities and challenges,” arXiv preprintarXiv:2311.05656, 2023. [51] Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen, L. Wang, A. T. Luu, W. Bi, F. Shi, and S. Shi, “Siren’s song in the ai ocean: A survey on hallucination in large language models,” 2023. [52] L. Huang, W. Yu, W. Ma, W. Zhong, Z. Feng, H. Wang, Q. Chen, W. Peng, X. Feng, B. Qin, and T. Liu, “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” 2023. [53] A. Sinitsin, V. Plokhotnyuk, D. Pyrkin, S. Popov, and A. Babenko, “Editable neural networks,” arXivpreprint arXiv:2004.00345, 2020. [54] N. De Cao, W. Aziz, and I. Titov, “Editing fac- tual knowledge in language models,”arXivpreprint arXiv:2104.08164, 2021. [55] E. Mitchell, C. Lin, A. Bosselut, C. D. Manning, and C. Finn, “Memory-based model editing at scale,” in InternationalConferenceonMachineLearning. PMLR, 2022, p. 15 817–15 831. [56] Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang, “Editing large language models: Problems, methods, and opportunities,” arXiv preprintarXiv:2305.13172, 2023. [57] Z. Li, H. Jiang, H. Chen, B. Bi, Z. Zhou, F. Sun, J. Fang, and X. Wang, “Reinforced lifelong editing for language models,” arXivpreprintarXiv:2502.05759, 2025. [58] R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva, “Evaluating the ripple effects of knowl- edge editing in language models,” Transactionsofthe AssociationforComputationalLinguistics, vol. 12, p. 283–298, 2024. [59] B. Bi, S. Liu, Y. Wang, L. Mei, H. Gao, Y. Xu, and X. Cheng, “Adaptive token biaser: Knowledge editing via biasing key entities,” arXivpreprint arXiv:2406.12468, 2024. [60] B. Bi, S. Liu, Y. Wang, L. Mei, H. Gao, J. Fang, and X. Cheng, “Struedit: Structured outputs enable the fast and accurate knowledge editing for large language models,” arXivpreprintarXiv:2409.10132, 2024. [61] K. Guu, K. Lee, Z. Tung, P. Pasupat, and M.- W. Chang, “Realm: Retrieval-augmented language model pre-training,” 2020. [Online]. Available: https: //arxiv.org/abs/2002.08909 [62] G. Izacard and E. Grave, “Leveraging passage retrieval with generative models for open domain question answering,” 2021. [Online]. Available: https: JOURNAL OF L A T E X CLASS FILES, JANUARY 202515 //arxiv.org/abs/2007.01282 [63] Y. Wang, M. Chen, N. Peng, and K.-W. Chang, “Deepedit: Knowledge editing as decoding with con- straints,” arXivpreprintarXiv:2401.10471, 2024. [64] J. Ji, T. Qiu, B. Chen, B. Zhang, H. Lou, K. Wang, Y. Duan, Z. He, J. Zhou, Z. Zhang etal., “Ai alignment: A comprehensive survey,”arXivpreprint arXiv:2310.19852, 2023. [65] P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advancesinneuralinformation processingsystems, vol. 30, 2017. [66] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McK- innon etal., “Constitutional ai: Harmlessness from ai feedback,”arXivpreprintarXiv:2212.08073, 2022. [67] A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. Das- Sarma etal., “A general language assistant as a labo- ratory for alignment,”arXivpreprintarXiv:2112.00861, 2021. [68] Z. Yuan, J. Tang, J. Luo, R. Chen, C. Qian, L. Sun, X. Chu, Y. Cai, D. Zhang, and S. Li, “Autodrive- r2: Incentivizing reasoning and self-reflection capacity for vla model in autonomous driving,” arXivpreprint arXiv:2509.01944, 2025. [69] S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire etal., “Open problems and fundamental limitations of reinforcement learning from human feedback,” arXiv preprintarXiv:2307.15217, 2023. [70] D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning ai with shared human values,” arXivpreprintarXiv:2008.02275, 2020. [71] X. Li, P. Yu, C. Zhou, T. Schick, O. Levy, L. Zettlemoyer, J. Weston, and M. Lewis, “Self-alignment with instruc- tion backtranslation,” arXivpreprintarXiv:2308.06259, 2023. [72] J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein, “Generative agents: Interactive simulacra of human behavior,” in Proceedingsofthe 36thannualacmsymposiumonuserinterfacesoftware andtechnology, 2023, p. 1–22. [73] E. Perez, S. Ringer, K. Lukosiute, K. Nguyen, E. Chen, S. Heiner, C. Pettit, C. Olsson, S. Kundu, S. Kada- vath etal., “Discovering language model behaviors with model-written evaluations,” in Findingsofthe associationforcomputationallinguistics:ACL2023, 2023, p. 13 387–13 434. [74] A. Madaan, N. Tandon, P. Gupta, S. Hallinan, L. Gao, S. Wiegreffe, U. Alon, N. Dziri, S. Prabhumoye, Y. Yang etal., “Self-refine: Iterative refinement with self- feedback,”AdvancesinNeuralInformationProcessing Systems, vol. 36, p. 46 534–46 594, 2023. [75] S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto, “Whose opinions do language mod- els reflect?” in InternationalConferenceonMachine Learning. PMLR, 2023, p. 29 971–30 004. [76] Y. Bang, D. Chen, N. Lee, and P. Fung, “Measuring political bias in large language models: What is said and how it is said,”arXivpreprintarXiv:2403.18932, 2024. [77] S. Lin, J. Hilton, and O. Evans, “Teaching models to express their uncertainty in words,”arXivpreprint arXiv:2205.14334, 2022. [78] C. Fernando, D. Banarse, H. Michalewski, S. Osin- dero, and T. Rockt ̈ aschel, “Promptbreeder: Self- referential self-improvement via prompt evolution,” arXivpreprintarXiv:2309.16797, 2023. [79] R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng, “Automatic prompt optimization with” gra- dient descent” and beam search,” arXivpreprint arXiv:2305.03495, 2023. [80] A. M. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid, “Activation addition: Steering language models without optimization,” CoRR, 2023. [81] T. Liu, S. Guo, L. Bianco, D. Calandriello, Q. Berthet, F. Llinares, J. Hoffmann, L. Dixon, M. Valko, and M. Blondel, “Decoding-time realignment of language models,” arXivpreprintarXiv:2402.02992, 2024. [82] Y. Zhang, L. Cui, W. Bi, and S. Shi, “Alleviating hal- lucinations of large language models through induced hallucinations,”arXivpreprintarXiv:2312.15710, 2023. [83] J. Wang, Y. Wang, G. Xu, J. Zhang, Y. Gu, H. Jia, M. Yan, J. Zhang, and J. Sang, “An llm-free multi-dimensional benchmark for mllms hallucination evaluation,” arXiv preprintarXiv:2311.07397, 2023. [84] S. Zhai, Y. Meng, Y. Zhang, and G. Qi, “Parameter- aware contrastive knowledge editing: Tracing and rectifying based on critical transmission paths,” in Proceedingsofthe63rdAnnualMeetingofthe AssociationforComputationalLinguistics(Volume1: LongPapers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds.Vienna, Austria: Association for Computational Linguistics, Jul. 2025, p. 28 189– 28 200. [Online]. Available: https://aclanthology.org/ 2025.acl-long.1367/ [85] K. Pan, Z. Fan, J. Li, Q. Yu, H. Fei, S. Tang, R. Hong, H. Zhang, and Q. Sun, “Towards unified multi- modal editing with enhanced knowledge collabora- tion,” arXivpreprintarXiv:2409.19872, 2024. Baolong Bi is a Ph.D. student in Computer Sci- ence at the Institute of Computing Technology, Chinese Academy of Sciences. His research in- terests focus on the general area of trustworthy large language models. He was a visiting re- searcher at the Language Technologies Institute, Carnegie Mellon University. He has published 10+ top-tier papers, including ICLR, ICML, ACL, W, EMNLP, NAACL, etc, and serves as a reviewer for ICLR, ACL, COLM, NeurIPS, and CIKM. Yuyao Ge is a Ph.D. student at the Institute of Computing Technology, Chinese Academy of Sciences, advised by Prof. Shenghua Liu. His research interests focus on the reasoning ca- pabilities of large language models, particularly in vision-language models and complex logic problem solving. He has published papers in toptier conferences including ACL and EMNLP, and serves as a reviewer for AAAI and CIKM. JOURNAL OF L A T E X CLASS FILES, JANUARY 202516 Shenghua Liu is a full Professor at the Institute of Computing Technology, Chinese Academy of Sciences. He received his Ph.D. degree from the Computer Science and Technology Depart- ment, Tsinghua University. He once visited at the University of California, Los Angeles, hosted by Prof. Lei He, and as a scholar at Carnegie Mellon University, hosted by Prof. Christos Faloutsos. His current research focuses on big graph min- ing, large language models (LLMs), and scalable algorithm design, with recent interests in apply- ing LLMs to graph analysis, and trustworthy foundation models. Two of about 60 high-quality publications have been recognized as “best paper” award and candidate respectively. Yuchen He is a senior undergraduate student majoring in Computer Science and Technology at Beijing Institute of Technology. He is currently a research intern at the Institute of Computing Technology, Chinese Academy of Sciences, ad- vised by Assoc. Prof. Huaming Liao and Prof. Shenghua Liu. His research interests include trustworthy large language models and rein- forcement learning with large language models. He has contributed to papers published in top- tier conferences including NeurIPS and SIGIR. Siqian Tong is a Ph.D. student in Artificial In- telligence in the University of Chinese Academy of Sciences, advised by Prof. Xuan Li. Her re- search interests focus on the multimodal large language model. Lizhe Chen is a Master’s student in Computer Science at the Tsinghua University Shenzhen International Graduate School, advised by Prof. Shiguang Ni. His research interests focus on 3D data synthesis, combining large language mod- els and computer graphics. He has published pa- pers in top-tier conferences such as AAAI, IJCAI, and others, and has worked on research projects related to enhancing reasoning capabilities of large language models and 3D data synthesis. He also serves as a reviewer for conferences like AAAI, IJCAI, and others. Lingrui Mei is a Ph.D. student in Computer Sci- ence at the Institute of Computing Technology, Chinese Academy of Sciences, advised by Prof. Shenghua Liu. His research interests focus on large language models, including AI safety and agentic AI. He has published 10+ top-tier papers, including ICLR, ACL, EMNLP, etc, and serves as a reviewer for ICLR, ACL, COLM, NeurIPS, and EMNLP. Zehao Li received the B.Sc. degree in computer science and technology from Yunnan University, Kunming, China, in 2023. He is currently pursu- ing the Ph.D. degree with the Institute of Com- puting Technology, Chinese Academy of Sci- ences, Beijing, China. His main research inter- ests include computer graphics and computer vision. Yiwei Wang is an Assistant Professor in the Computer Science Department at the University of California, Merced. His research focuses on natural language processing and graph machine learning, with recent work spanning large lan- guage model safety, multimodal reasoning, and retrieval-augmented generation. Dr. Wang has published extensively at top venues such as ACL, EMNLP, NeurIPS, ICML, ICLR, KDD, and AAAI, and his work has received recognitions including an EMNLP Best Paper Finalist and the SDSC Dissertation Research Fellowship. He serves as an Associate Editor for IEEE TII, Pattern Recognition, and Neurocomputing, and as a (Senior) Area Chair for major conferences including NeurIPS, ICML, ICLR, AAAI, ACL, EMNLP, COLM, IJCAI, and ICME. Additional informa- tion is available at https://wangywust.github.io/. Yujun Cai is currently a Lecturer at the School of Electrical Engineering and Computer Science, the University of Queensland (UQ), Australia. Before that she was a Research Scientist in Meta Inc USA. She obtained her Ph.D. from Nanyang Technological University in 2021. Her research focuses on multi-modal understanding, with publications in top-tier venues such as T- PAMI, CVPR, ICCV, NeurIPS, ECCV etc. She served as an Area Chair for conferences such as Neurips, CVPR, ICML, ICLR, ICCV, ACM M, and as Associate Editor for journals such as Pattern Recognition and Neurocomputing. Ming-Hsuan Yang (Fellow, IEEE) is affiliated with Google, UC Merced, and Yonsei University. Yang serves as a program co-chair of IEEE Inter- national Conference on Computer Vision (ICCV) in 2019, program co-chair of the Asian Confer- ence on Computer Vision (ACCV) in 2014, and general co-chair of ACCV 2016. Yang served as an associate editor of the IEEE Transactions on Pattern Analysis and Machine Intelligence and is an associate editor of the International Journal of Computer Vision, Image and Vision Computing and Journal of Artificial Intelligence Research. He received the NSF CAREER award and Google Faculty Award. He is a Fellow of the IEEE, ACM, and AAAI. Xueqi Cheng (Fellow, IEEE) is a professor with the Institute of Computing Technology, Chinese Academy of Sciences (ICT-CAS) and the Uni- versity of Chinese Academy of Sciences, and the director of the CAS Key Laboratory of Net- work Data Science and Technology. His main research interests include network science, web search and data mining, Big Data processing and distributed computing architecture. He has won the Best Paper Award in CIKM (2011), the Best Student Paper Award in SIGIR (2012), and the Best Paper Award Runner up of CIKM (2017).