Paper deep dive
Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning
Xuzhong Wang, Maiqi Jiang, Tejal Nair, Girija Bhusal, Yanfu Zhang, Haipeng Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/27/2026, 4:41:54 AM
Summary
The paper introduces Multi-Objective In-context Knowledge Editing (MO-IKE), a reinforcement learning framework that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process. MO-IKE optimizes a dynamic retriever to balance three competing objectives: reliability (edit success), generality (paraphrase consistency), and specificity (knowledge retention). By using a multi-objective reward shaped via Lagrangian relaxation and optimizing with Group Relative Policy Optimization (GRPO), MO-IKE outperforms previous RL-based methods like DR-IKE on benchmarks such as CounterFact and UniEdit, particularly improving retention rates and edit success on models like Llama-3.2.
Entities (12)
Relation Signals (9)
MO-IKE → evaluatedon → Llama-3.2
confidence 95% · On Llama-3.2, MO-IKE improves edit success (reliability) from 85.0% to 92.0%
MO-IKE → optimizes → Reliability
confidence 95% · MO-IKE trains a dynamic retriever to optimize competing objectives in knowledge editing... edit success (reliability)
MO-IKE → optimizes → Specificity
confidence 95% · increasing retention rate (specificity) by 23.0%
MO-IKE → optimizes → Generality
confidence 95% · paraphrase consistency (generality) from 77% to 79%
MO-IKE → uses → Constrained MDP
confidence 95% · MO-IKE, a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process.
MO-IKE → optimizesusing → GRPO
confidence 93% · During training, MO-IKE samples a group of candidate prompts under the current policy and optimizes the retriever via Group Relative Policy Optimization (GRPO).
MO-IKE → outperforms → DR-IKE
confidence 90% · MO-IKE improves edit success (reliability) from 85.0% to 92.0%... compared to prior RL-based methods.
DR-IKE → uses → BERT
confidence 88% · DR-IKE Nafee et al. (2025) utilizes RL to train a BERT-based retriever
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. In particular, in-context knowledge editing has gained attention because it is training-free and readily applicable to black-box LLMs. Recent reinforcement learning (RL)-based approaches improve over fixed retrieval strategies by adapting prompt construction to the quantity-quality trade-off. Despite initial success, they fail to model the prompt as a structured entity under the distinct and often competing objectives of reliability, generality, and specificity. Previous methods largely optimize a single objective and make decisions over only part of the prompt construction process, thereby overlooking both the balance of different objectives and the global organization of demonstrations. We propose Multi-Objective In-context Knowledge Editing (MO-IKE), a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process. MO-IKE trains a dynamic retriever to optimize competing objectives in knowledge editing, enabling more balanced and globally coherent prompt construction. On Llama-3.2, MO-IKE improves edit success (reliability) from 85.0% to 92.0%, paraphrase consistency (generality) from 77% to 79%, while increasing retention rate (specificity) by 23.0% compared to prior RL-based methods.
Tags
Links
- Source: https://arxiv.org/abs/2608.25100v1
- Canonical: https://arxiv.org/abs/2608.25100v1
Trouble viewing inline? Open PDF directly →
Full Text
72,532 characters extracted from source content.
Expand or collapse full text
Towards Reliable, Generalizable, and Specific In-Context Knowledge Editing via Multi-Objective Reinforcement Learning Xuzhong Wang Affiliation: College of William and Mary Email: xwang58@wm.edu Maiqi Jiang Affiliation: College of William and Mary Email: mjiang04@wm.edu Tejal Nair Affiliation: College of William and Mary Email: tnair@wm.edu Girija Bhusal Affiliation: Tribhuvan University Email: yzhang105@wm.edu Yanfu Zhang Affiliation: College of William and Mary Email: hchen23@wm.edu Haipeng Chen Affiliation: College of William and Mary Email: girija.bhusal9@gmail.com Abstract Large Language Models (LLMs) are powerful but limited by static parametric knowledge that becomes outdated once pretraining ends. Knowledge editing addresses this problem by updating model behavior on target facts without full retraining. In particular, in-context knowledge editing has gained attention because it is training-free and readily applicable to black-box LLMs. Recent reinforcement learning (RL)-based approaches improve over fixed retrieval strategies by adapting prompt construction to the quantity–quality trade-off. Despite initial success, they optimize reliability alone, and in doing so trade away specificity: rewarding edit success drives the retriever to discard the demonstrations that protect neighboring facts. Previous methods also act over only part of the prompt construction process, overlooking the global organization of demonstrations. We propose Multi-Objective In-context Knowledge Editing (MO-IKE), which casts prompt construction as a sequential decision process over COPY, UPDATE, RETAIN, and STOP actions, and trains a dynamic retriever with a multi-objective shaped reward that couples edit success to explicit paraphrase and retention penalties. On Llama-3.2-3B, averaged over five seeds, MO-IKE raises edit success (reliability) from 87.1% to 91.1% and retention rate (specificity) from 41.0% to 63.4% over the strongest RL baseline, while keeping paraphrase consistency (generality) comparable (79.1% to 77.7%), improving the harmonic-mean score from 61.7 to 75.7. Gains hold across five frozen LLMs and four datasets, including the 311K-example UniEdit benchmark. 1 Introduction Large Language Models (LLMs) have become prevalent in natural language processing, excelling in domains ranging from autonomous workflows Schick et al. (2023) to advanced mathematical reasoning Chervonyi et al. (2025). However, a critical limitation remains: parametric knowledge in LLMs is static. Once training concludes, the stored information becomes susceptible to obsolescence Mazzia et al. (2024). For instance, an LLM whose training precedes the 2026 Winter Olympics lacks parametric knowledge of the final medal tally. Consequently, relying solely on static memory invariably fails to satisfy user needs for real-time or dynamic information. Knowledge editing addresses this by updating an LLM’s response to a target fact while preserving unrelated knowledge Wang et al. (2024b). While gradient-based methods Mitchell et al. (2022) achieve this via weight updates, they are computationally expensive and cannot be adapted to black-box LLMs Meng et al. (2022). Recently, in-context knowledge editing has become popular due to its advantage of being training-free and naturally applicable to black-box LLMs Zheng et al. (2023): the model remains frozen, and the desired update is induced solely through prompt context, without access to model internals. The efficacy of in-context knowledge editing is highly dependent on the injected prompt. Thus, a critical question is how to construct the prompt. Precisely, the context must be reliable enough to successfully edit the LLM (reliability), general enough to apply to semantically equivalent queries (generality), and specific enough to avoid changing unrelated neighboring facts (specificity) Meng et al. (2023). Although static retrieval strategies such as IKE Zheng et al. (2023) propose constructing the prompt with different categories of demonstrations—specifically, COPY demonstrations that explicitly state the new fact for reliability, UPDATE demonstrations that state a paraphrased query for generality, and RETAIN demonstrations that state an unrelated fact for specificity—these methods remain insufficient, as they assume a fixed amount of context is adequate for distinct edits. Figure 1: Contradicting objectives of reliability and specificity in prompt construction in IKE – Increasing the number of RETAIN demonstrations leads to Retention Success but Edit Failure. To address this limitation, recent work such as Dynamic Retriever for In-Context Knowledge Editing (DR-IKE) Nafee et al. (2025) formulates demonstration selection as a reinforcement learning (RL) problem, enabling the system to adapt the number of retrieved examples based on each edit. Despite their initial success, DR-IKE exhibits two limitations. First, while it improves how much context is supplied, it leaves partially explored how that context should be organized by ranking only RETAIN candidates, as the effectiveness of in-context learning is known to be sensitive to demonstration ordering Lu et al. (2022). Second, DR-IKE does not account for the proportion of demonstrations from distinct categories. Effective in-context prompt construction requires a principled balance among distinct demonstration types to achieve reliability, generality, and specificity. From the observation in Figure 1, these objectives are often contradictory Zheng et al. (2023). On the one hand, achieving strong reliability and generality demands sufficient COPY and UPDATE examples to reinforce the factual change Qiao et al. (2024). On the other hand, maintaining specificity requires RETAIN examples that anchor unedited neighboring knowledge Youssef et al. (2025). Under a limited context budget, we have to carefully balance the proportion of different demonstrations and consider potential trade-offs. To alleviate the dilemma of competing objectives, we formulate the retrieval environment as a Constrained Markov Decision Process (Constrained MDP) and introduce MO-IKE as its solution algorithm. We formulate the state as the currently constructed prompt, and the action of the retriever is to either select an additional demonstration or terminates the prompt construction. We evaluate the final prompt based on its edit reliability given the constraints on generality and specificity. During training, MO-IKE samples a group of candidate prompts under the current policy and optimizes the retriever via Group Relative Policy Optimization (GRPO). Our main contributions are as follows: • We identify two notable deficiencies in state-of-the-art RL-based in-context knowledge editing (notably, DR-IKE): they neglect the sequential ordering of the prompt context and overlook the balance of examples from different demonstration categories. • We propose MO-IKE, a Multi-Objective RL algorithm that jointly optimizes the selection of distinctive demonstration categories. By framing the task as a Constrained MDP and optimizing it with multi-objective RL, we ensure stable training that balances reliability, generality, and specificity. • We introduce architectural improvements to a dynamic demonstration retriever. On standard knowledge-editing benchmarks, including CounterFact, using Llama-3.2-3B and Mistral-7B-v0.3, MO-IKE improves edit success (reliability) by up to +7.0% (reaching 92.0% on Llama-3.2), paraphrase consistency (generality) by up to +2.5%, while yielding a +23.0% absolute improvement in retention rate (specificity). 2 Related Work 2.1 In-Context Knowledge Editing Early work on knowledge editing largely relies on gradient-based approaches. For instance, MEND Mitchell et al. (2022) trains a smaller editor network for targeted weight updates. Constrained by the heavy computational overhead Wang et al. (2024a), researchers explore in-context knowledge editing as a gradient-free alternative. Initial explorations improve prompt engineering via prefixing Cohen et al. (2024), chain-of-thought Wang et al. (2025), and formalized IKE framework Zheng et al. (2023), which utilizes k-N to retrieve diverse COPY, UPDATE, and RETAIN demonstrations to address reliability, generality and specificity, respectively. Most relevantly, DR-IKE Nafee et al. (2025) utilizes RL to train a BERT-based retriever Devlin et al. (2019) to dynamically construct a prompt. Yet, while these approaches successfully improve edit reliability, they rarely address the balance between reliability, generality, and specificity. A natural question is whether IKE is distinct from Retrieval-Augmented Generation (RAG). The two settings have underlying objectives that differ: RAG aims to supplement the model’s parametric knowledge with external documents Lewis et al. (2021); Asai et al. (2023). IKE, by contrast, aims to overwrite specific parametric facts while preserving unrelated knowledge. The survey of Wang et al. (2024b) draws this boundary explicitly, concluding that RAG is unsuitable for the targeted, fact-level updates that knowledge editing seeks to achieve. As a consequence, RL-based RAG retrievers (Huang et al., 2026) are not designed to navigate the reliability–specificity–retention tradeoff that defines the IKE problem. We further discuss the question in Appendix C. 2.2 RL for Demonstration Selection Selecting optimal demonstrations for in-context learning (ICL) is challenging due to the combinatorial nature of the search space Purohit et al. (2025) and the prompt’s sensitivity to example interactions Gupta et al. (2023), demonstration order Lu et al. (2022), and model biases Xiang et al. (2024). Because static heuristics like embedding similarity Liu et al. (2022) fail to capture interdependency, recent work has reframed inference-time optimization as a sequential decision-making process. RL algorithms—such as PPO Schulman et al. (2017b) and GRPO Shao et al. (2024)—have driven good advances in this direction. Beyond improving reasoning Xu et al. (2025); Liu et al. (2025), RL has proven highly effective for enhancing tool use Qian et al. (2025), training web agents Qi et al. (2025), optimizing query rewriting Ma et al. (2023), managing agent memory Yan et al. (2026), and accelerating speculative decoding Wang et al. (2026). However, applying RL to balance the competing objectives of knowledge editing remains underexplored. DR-IKE Nafee et al. (2025) pioneered this transition by formulating the IKE demonstration selection as an RL process. Yet, by optimizing primarily for a single objective (reliability), DR-IKE leads the retriever to suffer degradation in specificity and knowledge retention. While Constrained MDPs Ganguly and Ghosh (2025) and multi-objective optimization frameworks Efroni et al. (2025a); Efroni et al. (2025b) theoretically address such trade-offs, their practical application to inference-time selection is limited. 3 Problem Statement Definition 3.1. (Knowledge Editing). Fix a frozen language model ℳM and a factual triple K_C encoded in the model’s parametric memory. Knowledge editing f pursues a post-edited model ℳ′:=f(ℳ,→′)M :=f(M,K_C _C) such that: • Reliability. For the exact query q targeting the original fact CK_C, the output of ℳ′M accurately reflects the revised fact C′K _C. • Generality. For any query q∗q^* that is semantically equivalent to q (e.g., paraphrases) or whose answer logically depends on CK_C, the output of ℳ′M is consistent with C′K _C. • Specificity. For any query q′q that depends on unrelated knowledge SK_S (i.e., C∩S=∅K_C _S= ), the output of ℳ′M aligns identically with the unedited model ℳM. Definition 3.2. (In-context Knowledge Editing (IKE)). Fix a language model ℳM with frozen parameters. In-context knowledge editing utilizes the base prompt P and a retriever to concatenate an augmented Prompt P∗P^* ∗:=+<d1,d2,…,dn>P^*:=P+<d_1,d_2,…,d_n> (1) for did_i from d1d_1 to dnd_n that are natural language demonstrations selected from the example pool by the retriever. To satisfy the competing constraints of reliability, generality, and specificity, IKE formulates the construction of the prompt P∗P^* as follows Definition 3.3. (Demonstration Categories Zheng et al. (2023)) Each retrieved demonstration did_i from d1d_1 to dnd_n is assigned to one of the three categories. • COPY - directly restates the targeted fact. Example: The newest iPhone is → iPhone 17. • UPDATE - paraphrases the query before stating the targeted fact. Example: The latest iPhone in the generation is → iPhone 17. • RETAIN - states the neighboring fact that should not change. Example: The newest iPad is → iPad M5 Pro. 4 Methodology Our primary goal is to train a retriever that can construct contexts to effectively overwrite LLM’s stored parametric knowledge. Ideally, we would desire the edited LLM’s output to be the newly injected fact for any query targeted at that specific information. Moreover, a successful edit should prevent the neighboring facts from being modified. Nevertheless, we identify that prior baselines, notably DR-IKE, only considers editing signal in the RL training. An obvious drawback of the previous approach is that editing specificity signal is totally neglected. Thus, the resulting policy is susceptible to “reward hacking” where the trained retriever filtered all RETAIN candidates in favor for editing success. To achieve a principled balance among different objectives, we formulate the training process as a Constrained Markov Decision Process (Constrained MDP): besides edit reliability, we have to consider the constructed prompt’s impact on paraphrase consistency and knowledge retention. To operationalize this idea, we employ cost constraints in addition to editing reward to measure the retriever’s performance across all objectives. The final scalar reward signal is computed via multi-objective RL through individual rewards of reliability, generality, and specificity. Throughout the training process, our multi-objective framework helps the retriever optimize for a stable policy that balances distinctive and often competing objectives for an effective knowledge editing. We formalize this procedure as MO-IKE, a multi-objective in-context knowledge editing algorithm. 4.1 IKE as a Constrained MDP To guarantee editing generality and specificity, the retriever should be aware of these constraints in the training. Hence, we model the sequential demonstration selection process as a Constrained MDP (,,,ℛ,)(S,A,T,R,C). The retriever operates sequentially over discrete time steps t, constructing a prompt to query the frozen LLM. State. The state st∈s_t represents the prompt context at step t. It is constructed by appending a newly selected demonstration dtd_t, formatted into a natural language string via a mapping function ϕ(⋅)φ(·), to the previous state. Thus, s0s_0 is the original query q, and st=st−1⊕ϕ(dt)s_t=s_t-1 φ(d_t), where ⊕ denotes string concatenation. Action. In contrast to DR-IKE, which considers only RETAIN candidates in its action space Nafee et al. (2025), we formulate the action space over the full set of COPY, UPDATE, and RETAIN demonstrations. We introduce this change in order to encourage the exploration of a “global” optimal ordering between distinctive demonstration types. At step t, an action ata_t selects a demonstration without replacement from the available pool tD_t, or chooses a learnable pseudo-token astopa_stop to halt. The action space is thus t=t∪astopA_t=D_t∪\a_stop\, where t+1=t∖atD_t+1=D_t \a_t\. State Transition. The transition function (st+1|st,at)T(s_t+1|s_t,a_t) is deterministic. If the agent selects a demonstration at∈a_t , it is appended to the current state, yielding st+1=st⊕ats_t+1=s_t a_t. If the agent selects at=astopa_t=a_stop, the episode immediately terminates, and the constructed prompt is finalized for evaluation. Figure 2: MO-IKE overview. Given a query, we first preselect COPY, UPDATE, and RETAIN candidates using kNN, then sample prompts under the current retriever policy. The LLM evaluates these prompts to produce rewards and constraint penalties, which are combined via Lagrangian relaxation into final rewards R1,R2,R3,…R_1,R_2,R_3,…. These rewards are then used to compute group advantages A1,A2,A3,…A_1,A_2,A_3,… for policy optimization. Dashed lines mark operations used for training, not inference. Reward. Our primary objective is Edit Success (ES), which represents edit reliability. Specifically, the primary reward is defined as R(st,at)=[y^edit=ynew],R(s_t,a_t)=1[ y^edit=y_new], (2) where y^edit y^edit denotes the frozen LLM’s prediction on the edit instance by the constructed prompt, and ynewy_new denotes the desired edited target. Cost Constraints. Different from prior baselines, we introduce cost constraints so as to “regularize” RL training. We treat Paraphrase Consistency (PC), which represents generality, and Retention Rate (R), which represents specificity as constraints for our retriever. Let y^para y^para and y^retain y^retain denote the frozen LLM’s predictions on the paraphrase query and retention query, respectively. Let ynewy_new denote the desired edited target and ytruey_true denote LLM’s original stored answer for neighboring facts. We define each cost as the logical inverse of task success: CPC(st,at)=1−[y^para=ynew],C_PC(s_t,a_t)=1-1[ y^para=y_new], (3) CRR(st,at)=1−[y^retain=ytrue].C_R(s_t,a_t)=1-1[ y^retain=y_true]. (4) 4.2 Multi-Objective Reward Shaping To obtain the final objective, we scalarize the multi-objective problem into a single shaped reward, using a fixed-multiplier Lagrangian relaxation as the bridge. We define a composite reward that penalizes degradation on PC and R: r(st,at)=R(st,at)−∑k∈PC,RλkCk(st,at),r(s_t,a_t)=R(s_t,a_t)- _k∈\PC,R\ _kC_k(s_t,a_t), (5) where λPC _PC and λR _R are fixed hyperparameters rather than dual variables updated by ascent. We therefore do not solve a constrained program; the relaxation serves to derive the reward structure, and the resulting objective is multi-objective reward shaping optimized directly with GRPO. We give the derivation in Appendix B. 4.3 MO-IKE In general, MO-IKE constructs editing prompts through a four-stage pipeline: candidate retrieval, sequential selection, reward evaluation, and policy optimization. (i) First, for a given edit query, we retrieve an initial candidate pool of factual demonstrations. (i) Next, a BERT-based Devlin et al. (2019) retriever sequentially selects samples from this pool to build the prompt step-by-step, halting only when it chooses a “stop” token. (i) The fully constructed prompt is then evaluated against both editing success (reliability), paraphrase consistency (generality) constraints, and retention (specificity). (iv) These competing objectives are combined using a fixed-penalty coefficient, producing a final scalar reward. We illustrate the core idea of MO-IKE in Algorithm 1, and attach the full detailed Algorithm in Appendix A. Demonstration Selection. For a given edit instance (x,ynew)(x,y_new), we construct the prompt via a two-stage retrieval pipeline. First, we use a Sentence Transformer Reimers and Gurevych (2019) to extract an initial pool of Copy, Update, and Retain candidates via kNN. Our BERT-based retriever then performs fine-grained sequential selection from this pool to build the final prompt. Algorithm 1 MO-IKE 1: Retriever πθ _θ, dataset trainD_train, group size G, fixed penalty λ 2: for each (x,ynew)∈train(x,y_new) _train do ⊳ Sampling & Reward Evaluation: 3: Sample G prompt trajectories τgg=1G∼πθ(⋅∣x)\ _g\_g=1^G _θ(· x) based on current policy π 4: Compute composite reward rgr_g for each τg _g using fixed penalty λ ⊳ Advantage Estimation: 5: Compute group-normalized advantages: Ag←(rg−μ)/σA_g←(r_g-μ)/σ ⊳ Policy Update: 6: Update θ via gradient descent to minimize: 7: (θ)=1G∑g=1G[ℒclip(θ,τg,Ag)−βKLKL(πθ∥πθold)]J(θ)= 1G _g=1^G [L_clip(θ, _g,A_g)- _KLD_KL( _θ\| _ _old) ] 8: end for Policy and Action Sampling. The MO-IKE retriever uses a frozen 4-layer BERT encoder Devlin et al. (2019) with trainable projection matrices for the query (WpW_p) and demonstrations (WdW_d). We introduce a learned termination pseudo-embedding vstopv_stop to represent the STOP action astopa_stop. As a sequential decision making process, the policy samples based on current state ptp_t as a partially constructed prompt, and the candidates pool tA_t. At step t, with a partially constructed prompt ptp_t from the previous t−1t-1 steps and candidate pool tA_t, the selection probability for any action ai∈¯ta_i∈ A_t is computed via a softmax over projected dot products: P(ai|pt)=exp(vp⊤vi)∑aj∈¯texp(vp⊤vj)P(a_i|p_t)= (v_p v_i) _a_j∈ A_t (v_p v_j) (6) where vp=Wp⋅BERT(pt)v_p=W_p·BERT(p_t) and vi=Wd⋅BERT(ai)v_i=W_d·BERT(a_i). During training, we sample a∼P(⋅|pt)a P(·|p_t) to encourage exploration. At inference, we select the highest-probability demonstration. If astopa_stop is chosen, prompt construction terminates. Otherwise, we update the state (pt+1=pt⊕ap_t+1=p_t a) and shrink the pool (t+1=t∖aA_t+1=A_t \a\). Policy Update. For each instance (x,ynew)(x,y_new), we sample a group of G prompt trajectories τ1,…,τG\ _1,…, _G\ from the current policy πθ _θ. We evaluate these trajectories using the target LLM to compute a composite reward rir_i based on edit success, paraphrase consistency, and retention metrics. Group-relative advantages are estimated as Ai=(ri−μr)/σrA_i=(r_i- _r)/ _r, where μr _r and σr _r are the group’s reward mean and standard deviation. We optimize θ by maximizing the clipped surrogate objective, penalized by the Kullback-Leibler (KL) divergence from a reference policy πref _ref (the initial, reference retriever weights): ℒGRPO(θ)=ℒclip(θ)−βKL[πθ∥πref],L_GRPO(θ)=L_clip(θ)- _KL[ _θ\| _ref], (7) where ℒclip(θ)L_clip(θ) is given by: ℒclip(θ)=1G∑i=1Gmin(πθ(τi|q)πold(τi|q)Ai, _clip(θ)= 1G _i=1^G ( _θ( _i|q) _old( _i|q)A_i, (8) OPENclip(πθ(τi|q)πold(τi|q),1−ϵ,1+ϵ)Ai) ( _θ( _i|q) _old( _i|q),1-ε,1+ε )A_i ) 5 Experiments 5.1 Experimental Setup Datasets. In this section, we present evaluations on MO-IKE and our baselines across three standard knowledge editing benchmarks: CounterFact Meng et al. (2022), ZsRE (Levy et al., 2017), and Wiki-Counterfact (Cohen et al., 2024). For each benchmark, we partition the factual records into two distinct sets: an editable sample pool for evaluation and a separate pool for constructing IKE demonstrations. Following Zheng et al. (2023), we allocate the first 2,000 samples (out of 21,919) from CounterFact to the editable pool. We apply an analogous split to the other datasets, reserving the first 500 samples from both ZsRE (out of 11,230) and WikiDatacounterfact (out of 2,340) for editing. The remaining records in each dataset are utilized to build the demonstration pools. We additionally evaluate MO-IKE on larger datasets UniEdit Chen et al. (2025), which is shown more in details in Appendix G. Language Models. We evaluate our approach using several instruction-tuned large language models: Meta-Llama-3.1-8B-Instruct, Meta-Llama-3.2-3B-Instruct Grattafiori et al. (2024), Mistral-7B-Instruct-v0.2 Jiang et al. (2023), and Qwen2.5 (1.5B and 7B) Yang et al. (2024). All model parameters remain strictly frozen throughout our experiments. Training Configuration. Similar to Nafee et al. (2025), we train our retriever on 300 randomly selected examples and evaluate it on held-out sets of 300 samples for CounterFact, and 100 samples each for ZsRE and WikiDatacounterfact. We provide detailed training configuration in Appendix D. Baselines. We compare our approach against four representative in-context knowledge editing baselines: FactPrompt (Cohen et al., 2024), which prepends a narrative prefix to guide editing; EditCoT (Wang et al., 2025), which leverages chain-of-thought prompting; IKE (Zheng et al., 2023), which categorizes demonstrations into COPY, UPDATE, and RETAIN types; and DR-IKE (Nafee et al., 2025), which utilizes policy optimization for dynamic demonstration selection. Evaluation Metrics. We evaluate MO-IKE across three dimensions of knowledge editing: reliability, generality, and specificity. Let ynewy_new and ytruey_true denote the newly injected fact and the original stored truth, respectively. • Reliability: Measures post-editing efficacy on target prompts via Edit Success (ES), defined as [[P(ynew)>P(ytrue)]]E[I[P(y_new)>P(y_true)]], and Edit Magnitude (EM), defined as [P(ynew)−P(ytrue)]E[P(y_new)-P(y_true)]. • Generality: Evaluates accuracy on paraphrased prompts via Paraphrase Consistency (PC) and Paraphrase Magnitude (PM), which are calculated identically to ES and EM, respectively. • Specificity: Assesses the preservation of unedited neighboring facts via Retention Rate (R), defined as [[P(ytrue)>P(ynew)]]E[I[P(y_true)>P(y_new)]], and Retention Magnitude (RM), defined as [P(ytrue)−P(ynew)]E[P(y_true)-P(y_new)]. Finally, we report the overall Score (S) as the harmonic mean of ES, PC, and R. 5.2 Main results Editing Method S ↑ ES ↑ EM ↑ PC ↑ PM ↑ R ↑ RM ↑ Llama-3.2-3B-Instruct FactPrompt 34.8 55.3 64.0 25.7 23.8 34.3 -7.1 EditCoT 42.5 70.5 62.4 43.5 21.6 29.9 1.0 IKE 63.9 78.0 67.9 68.0 60.0 52.0 29.9 DR-IKE 54.9 85.0 76.4 77.0 69.2 34.7 -4.2 MO-IKE 73.5 92.0 78.4 79.0 66.0 57.7 30.0 Mistral-7B-Instruct-v0.3 FactPrompt 41.8 36.3 73.4 42.3 34.3 48.3 -2.9 EditCoT 35.6 33.8 58.7 43.4 33.4 31.0 5.6 IKE 64.8 58.0 84.5 81.3 69.9 60.0 33.9 DR-IKE 62.0 74.7 91.8 88.3 78.4 42.3 -22.5 MO-IKE 77.4 80.1 92.0 90.7 74.9 65.0 32.1 Table 1: Main results of in-context knowledge editing across Llama-3.2-3B-Instruct and Mistral-7B-Instruct-v0.3. We report Overall Score (S), Edit Success (ES), Edit Metric (EM), Paraphrase Consistency (PC), Paraphrase Metric (PM), Retention Rate (R), and Retention Metric (RM). Bold indicates the best performance, and underline indicates the second best. Table 1 presents the performance of MO-IKE on the CounterFact dataset. We highlight results for Llama-3.2 and Mistral-v0.3, with extended model evaluations provided in Table 9. We also include results from different datasets in Table 2. Furthermore, we include some additional experiments in Appendix G. Across both architectures, MO-IKE consistently outperforms across the majority of metrics: (1) Compared to the standard IKE baseline, MO-IKE yields substantial improvements in edit success and consistency. Against DR-IKE, it maintains or exceeds peak reliability, pushing Llama-3.2 ES from 85.0 to 92.0. (2) RL baselines like DR-IKE often drops below the IKE baseline in R and overall Score. MO-IKE protects existing parametric knowledge by enforcing retention. MO-IKE reverses the specificity degradation seen in prior methods. It improves R by absolute margins of +23.0 for Llama-3.2 and +22.7 for Mistral-v0.3 over DR-IKE, while shifting Mistral’s negative RM from -22.5 to a positive +32.1. (3) As shown in Table 2, MO-IKE’s performs well also in ZsRE and Wiki-Counterfact. It achieves the highest overall Score (42.0 on ZsRE and 59.0 on Wiki) while simultaneously achieving the strongest R. Method S ↑ ES ↑ PC ↑ R ↑ ZsRE IKE 34.0 64.0 50.0 19.0 DR-IKE 33.0 77.0 65.0 16.0 MO-IKE 42.0 75.0 79.0 22.0 WikiCounterfact IKE 55.0 70.0 54.0 46.0 DR-IKE 53.0 86.0 68.0 33.0 MO-IKE 59.0 80.0 59.0 47.0 Table 2: Editing performance across the ZsRE and WikiCounterfact datasets. Zero-Shot Cross Model Evaluation. We observe that retriever trained via MO-IKE on one language model can effectively generalize to a completely different language model without any retraining. To prove this, we training the retriever purely on Llama-3.2, and evaluating it zero-shot on Mistral-7B. As demonstrated in Table 3, MO-IKE suffers no performance drop when transferred zero-shot to Mistral-7B. In fact, it performs practically identically to the natively trained version, outperforming all other baselines. Method S ↑ ES ↑ PC ↑ R ↑ IKE 64.9 58.0 81.3 60.0 DR-IKE 62.3 74.7 88.3 42.3 MO-IKE (Native) 77.1 80.1 90.7 65.0 MO-IKE (Zero-Shot) 77.3 81.0 89.3 65.7 Table 3: Zero-shot transfer performance: the retriever was trained exclusively on Llama-3.2 and evaluated zero-shot on Mistral-7B. 5.3 Reward Constrained Ablation We conduct an ablation study on the reward function to evaluate four configurations: (i) no constraints, (i) paraphrase constraint only, (i) retention constraint only, and (iv) all constraints. As shown in Table 4, omitting the retention penalty (row i and i) leads to a noticeable drop in R. This confirms that explicitly penalizing retention degradation is necessary to prevent editing success from negatively impacting the model’s pre-existing knowledge. Furthermore, the results implies a positive correlation between edit success and paraphrase consistency: optimizing for editing reliability inherently benefits generality. Reward Constraint S↑ ES↑ PC↑ R↑ No Constraint 65.5 91.3 77.0 50.0 Only Paraphrase 68.8 92.0 79.0 52.0 Only Retention 70.1 86.0 71.0 57.7 All Constraints 73.5 92.0 79.0 57.7 Table 4: Ablation study on the components of the MO-IKE scalarized reward function using Llama-3.2. Bold indicates the best performance. It is worth noting that the full multi-objective formulation (all constraints) achieves better Score compared to only retention constraint (row i and iv). We hypothesize that it may comes from the mechanics of group relative rewards. When the reward signal lacks multiple constraints, the group rewards tend to exhibit lower variance within a group. Consequently, the group-relative advantages become uninformative, which may downgrade the effect of the policy gradient updates. By incorporating all constraints, the reward landscape becomes more nuanced, providing the variance in advantage signals necessary for RL training. In addition, we discuss reward weights sensitivity in Appendix F. Figure 3: Performance of Llama-3.2 using different demonstration selection scopes. 5.4 Action Space Ablation Previous approaches to in-context knowledge editing have largely treated demonstration organization as a secondary concern. IKE Zheng et al. (2023) constructs prompts using static format, essentially treating demonstrations as a fixed template. While DR-IKE Nafee et al. (2025) introduces a dynamic retrieval mechanism, it restricts its optimization to the ranking and construction of only RETAIN candidates, ignoring the inter-dependencies between different demonstration types. In contrast, MO-IKE’s action space is expanded to include all demonstrations across all categories. By conducting RL training across the entire sequence of demonstration types—COPY, UPDATE, and RETAIN—our method captures the global, synergistic effects of context construction. As illustrated in Figure 3, our approach achieves better performance across all three metrics compared to optimizing solely for RETAIN. These empirical results confirms our decision to expand the action space across all demonstration categories. 5.5 Demonstration Structure and Analysis The performance of our approach can be attributed to the resulting structural proportions of the demonstration candidates. DR-IKE Nafee et al. (2025) force a greedy optimization for edit success, which skews the prompt distribution by discarding RETAIN candidates. To resolve this structural imbalance, MO-IKE integrates a multi-objective RL framework and a soft stopping mechanism through a stop signal in the action space that is evaluated alongside other demonstration candidates. As shown in Figure 4, this allows the model to maintain contextual heterogeneity rather than aggressively filtering out non-edit demonstrations. Consequently, MO-IKE achieves better editing specificity overall (Table 1). We provide detailed Case Study of the context structure in Appendix H. By maintaining a balanced proportion of COPY, UPDATE, and RETAIN types, our method achieves robust results without sacrificing the overall integrity of the model’s existing knowledge. Figure 4: Left axis tracks filtered candidates; the right axis shows inverted retention degradation. The shaded gap highlights DR-IKE’s hard constraint failing (over-filtering by Epoch 3) versus MO-IKE’s soft embedding preserving stable retention. Conclusion We introduce MO-IKE, a multi-objective RL algorithm for in-context knowledge editing that trains a retriever to dynamically construct prompts. We formulate demonstration retrieval as a Constrained MDP, casting prompt construction as a sequential decision-making problem. MO-IKE explicitly optimizes the retriever over the often competing objectives of reliability, generalization, and specificity. Experiments across multiple LLMs show that MO-IKE consistently outperforms existing retrieval strategies, achieving strong effectiveness across all three objectives and cross-model generalization. Ablation studies further demonstrate the importance of our constraint formulation, as well as the embedding and ranking designs, to MO-IKE’s overall performance. Limitations Several limitations still exist in our work. First, the dataset scale for knowledge editing remains relatively small; for example, we train and evaluate MO-IKE only on the first 2,000 records of the CounterFact dataset, which may not be sufficient to examine the model’s overall capabilities. Second, by formulating the problem as a Constrained MDP, we incorporate all constraints directly into the reward function. Because we utilize fixed coefficients (e.g., a static lambda value) for these constraints, the precise impact of these hyperparameters on the RL optimization process remains unexplored. Third, although MO-IKE achieves a better retention rate overall, its absolute performance is bottlenecked by the inherent limits of in-context learning. The retention rate is still relatively low (around 60%) compared to other metrics, suggesting this gradient-free method still has room for improvement. Future work may consider different approaches to surpass this limitation. Finally, we evaluate MO-IKE solely on a specific knowledge editing task, leaving it unclear whether this framework generalizes to other in-context learning applications, such as model reasoning. Acknowledgements This research was supported by the National Science Foundation under IIS-2348405 and the William & Mary Semester Research Grant. References Asai et al. (2023) A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi Self-RAG: learning to retrieve, generate, and critique through self-reflection. arXiv preprint arXiv:2310.11511. External Links: Link Cited by: §2.1. Chen et al. (2025) Q. Chen, D. Wang, T. Zhang, Z. Yan, C. You, C. Wang, and X. He UniEdit: a unified knowledge editing benchmark for large language models. External Links: 2505.12345, Link Cited by: Appendix G, §5.1. Chervonyi et al. (2025) Y. Chervonyi, T. H. Trinh, M. Olšák, X. Yang, H. Nguyen, M. Menegali, J. Jung, J. Kim, V. Verma, Q. V. Le, and T. Luong Gold-medalist performance in solving olympiad geometry with alphageometry2. External Links: 2502.03544, Link Cited by: §1. Cohen et al. (2024) R. Cohen, E. Biran, O. Yoran, A. Globerson, and M. Geva Evaluating the ripple effects of knowledge editing in language models. Transactions of the Association for Computational Linguistics 12, p. 283–298. External Links: Link, Document Cited by: §2.1, §5.1, §5.1. Devlin et al. (2019) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, Minnesota, p. 4171–4186. External Links: Link, Document Cited by: §2.1, §4.3, §4.3. Efroni et al. (2025a) Y. Efroni, B. Kretzu, D. Jiang, J. Bhandari, Zheqing, Zhu, and K. Ullrich Aligned multi objective optimization. External Links: 2502.14096, Link Cited by: §2.2. Efroni et al. (2025b) Y. Efroni, B. Kretzu, D. Jiang, J. Bhandari, Z. Zhu, and K. Ullrich Aligned multi objective optimization. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §2.2. Ganguly and Ghosh (2025) S. Ganguly and A. Ghosh Provably efficient sample complexity for robust cmdp. External Links: 2511.07486, Link Cited by: §2.2. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §5.1. Gupta et al. (2023) S. Gupta, M. Gardner, and S. Singh Coverage-based example selection for in-context learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 13924–13950. External Links: Link, Document Cited by: §2.2. Huang et al. (2026) J. Huang, S. Madala, R. Sidhu, C. Niu, H. Peng, J. Hockenmaier, and T. Zhang Tackling distractor documents in multi-hop qa with reinforcement and curriculum learning. In Findings of the Association for Computational Linguistics: EACL 2026, p. 5548–5561. External Links: Link, Document Cited by: Appendix C, §2.1. Jiang et al. (2023) A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed Mistral 7b. External Links: 2310.06825, Link Cited by: §5.1. Levy et al. (2017) O. Levy, M. Seo, E. Choi, and L. Zettlemoyer Zero-shot relation extraction via reading comprehension. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), R. Levy and L. Specia (Eds.), Vancouver, Canada, p. 333–342. External Links: Link, Document Cited by: §5.1. Lewis et al. (2021) P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela Retrieval-augmented generation for knowledge-intensive nlp tasks. External Links: 2005.11401, Link Cited by: §2.1. Liu et al. (2025) F. Liu, W. Chao, N. Tan, and H. Liu Bag of tricks for inference-time computation of LLM reasoning. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track, External Links: Link Cited by: §2.2. Liu et al. (2022) J. Liu, D. Shen, Y. Zhang, B. Dolan, L. Carin, and W. Chen What makes good in-context examples for GPT-3?. In Proceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, E. Agirre, M. Apidianaki, and I. Vulić (Eds.), Dublin, Ireland and Online, p. 100–114. External Links: Link, Document Cited by: §2.2. Lu et al. (2022) Y. Lu, M. Bartolo, A. Moore, S. Riedel, and P. Stenetorp Fantastically ordered prompts and where to find them: overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio (Eds.), Dublin, Ireland, p. 8086–8098. External Links: Link, Document Cited by: §1, §2.2. Ma et al. (2023) X. Ma, Y. Gong, P. He, H. Zhao, and N. Duan Query rewriting in retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, p. 5303–5315. External Links: Link, Document Cited by: §2.2. Mazzia et al. (2024) V. Mazzia, A. Pedrani, A. Caciolai, K. Rottmann, and D. Bernardi A survey on knowledge editing of neural networks. IEEE Transactions on Neural Networks and Learning Systems 36 (7), p. 11759–11775. External Links: Document Cited by: §1. Meng et al. (2022) K. Meng, D. Bau, A. Andonian, and Y. Belinkov Locating and editing factual associations in GPT. In Advances in Neural Information Processing Systems, Vol. 35, p. 17359–17372. External Links: Link Cited by: §1, §5.1. Meng et al. (2023) K. Meng, A. S. Sharma, A. J. Andonian, Y. Belinkov, and D. Bau Mass-editing memory in a transformer. In The Eleventh International Conference on Learning Representations, External Links: Link Cited by: §1. Mitchell et al. (2022) E. Mitchell, C. Lin, A. Bosselut, C. Finn, and C. D. Manning Fast model editing at scale. In International Conference on Learning Representations, External Links: Link Cited by: §1, §2.1. Nafee et al. (2025) M. W. Nafee, M. Jiang, H. Chen, and Y. Zhang Dynamic retriever for in-context knowledge editing via policy optimization. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 16744–16757. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §1, §2.1, §2.2, §4.1, §5.1, §5.1, §5.4, §5.5. Purohit et al. (2025) K. Purohit, V. V, S. Bhattacharya, and A. Anand Sample efficient demonstration selection for in-context learning. In Forty-second International Conference on Machine Learning, External Links: Link Cited by: §2.2. Qi et al. (2025) Z. Qi, X. Liu, I. L. Iong, H. Lai, X. Sun, J. Sun, X. Yang, Y. Yang, S. Yao, W. Xu, J. Tang, and Y. Dong WebRL: training LLM web agents via self-evolving online curriculum reinforcement learning. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §2.2. Qian et al. (2025) C. Qian, E. C. Acikgoz, Q. He, H. WANG, X. Chen, D. Hakkani-Tür, G. Tur, and H. Ji ToolRL: reward is all tool learning needs. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2.2. Qiao et al. (2024) S. Qiao, X. Liu, and S. Na COMEM: in-context retrieval-augmented mass-editing memory in large language models. In Findings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, p. 2333–2347. External Links: Link, Document Cited by: §1. Reimers and Gurevych (2019) N. Reimers and I. Gurevych Sentence-BERT: sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), K. Inui, J. Jiang, V. Ng, and X. Wan (Eds.), Hong Kong, China, p. 3982–3992. External Links: Link, Document Cited by: §4.3. Schick et al. (2023) T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom Toolformer: language models can teach themselves to use tools. In Thirty-seventh Conference on Neural Information Processing Systems, External Links: Link Cited by: §1. Schulman et al. (2017a) J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel Trust region policy optimization. External Links: 1502.05477, Link Cited by: Appendix E. Schulman et al. (2017b) J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov Proximal policy optimization algorithms. External Links: 1707.06347, Link Cited by: §2.2. Shao et al. (2024) Z. Shao, P. Wang, Q. Zhu, R. Xu, J. Song, X. Bi, H. Zhang, M. Zhang, Y. K. Li, Y. Wu, and D. Guo DeepSeekMath: pushing the limits of mathematical reasoning in open language models. External Links: 2402.03300, Link Cited by: §2.2. Wang et al. (2025) C. Wang, W. Su, Q. Ai, Y. Tang, and Y. Liu Knowledge editing through chain-of-thought. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 10673–10693. External Links: Link, Document, ISBN 979-8-89176-332-6 Cited by: §2.1, §5.1. Wang et al. (2026) C. Wang, D. H. Shi, and H. Chen Speculative sampling with reinforcement learning. External Links: 2601.12212, Link Cited by: §2.2. Wang et al. (2024a) P. Wang, Z. Li, N. Zhang, Z. Xu, Y. Yao, Y. Jiang, P. Xie, F. Huang, and H. Chen WISE: rethinking the knowledge memory for lifelong model editing of large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: §2.1. Wang et al. (2024b) S. Wang, Y. Zhu, H. Liu, Z. Zheng, C. Chen, and J. Li Knowledge editing for large language models: a survey. External Links: 2310.16218, Link Cited by: §1, §2.1. Xiang et al. (2024) Y. Xiang, H. Yan, L. Gui, and Y. He Addressing order sensitivity of in-context demonstration examples in causal language models. In Findings of the Association for Computational Linguistics: ACL 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 6467–6481. External Links: Link, Document Cited by: §2.2. Xu et al. (2025) F. Xu, Q. Hao, Z. Zong, J. Wang, Y. Zhang, J. Wang, X. Lan, J. Gong, T. Ouyang, F. Meng, C. Shao, Y. Yan, Q. Yang, Y. Song, S. Ren, X. Hu, Y. Li, J. Feng, C. Gao, and Y. Li Towards large reasoning models: a survey of reinforced reasoning with large language models. External Links: 2501.09686, Link Cited by: §2.2. Yan et al. (2026) S. Yan, X. Yang, Z. Huang, E. Nie, Z. Ding, Z. Li, X. Ma, J. Bi, K. Kersting, J. Z. Pan, H. Schütze, V. Tresp, and Y. Ma Memory-r1: enhancing large language model agents to manage and utilize memories via reinforcement learning. External Links: 2508.19828, Link Cited by: §2.2. Yang et al. (2024) A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei, H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin, K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao, R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang, X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo, and Z. Fan Qwen2 technical report. External Links: 2407.10671, Link Cited by: §5.1. Youssef et al. (2025) P. Youssef, Z. Zhao, J. Schlötterer, and C. Seifert How to make LLMs forget: on reversing in-context knowledge edits. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), L. Chiruzzo, A. Ritter, and L. Wang (Eds.), Albuquerque, New Mexico, p. 12656–12669. External Links: Link, Document, ISBN 979-8-89176-189-6 Cited by: §1. Zheng et al. (2023) C. Zheng, L. Li, Q. Dong, Y. Fan, Z. Wu, J. Xu, and B. Chang Can we edit factual knowledge by in-context learning?. In The 2023 Conference on Empirical Methods in Natural Language Processing, External Links: Link Cited by: §1, §1, §1, §2.1, Definition 3.3, §5.1, §5.1, §5.4. Appendix A Training Procedure for Multi-objective GRPO Algorithm 2 outlines the detailed MO-IKE algorithm. Algorithm 2 Multi-Objective In-Context Knowledge Editing (MO-IKE) 1: initial retriever parameters θ, training set trainD_train, group size G, clipping ϵε, KL weight βKL _KL 2: for epoch =1=1 to NepochsN_epochs do 3: for each (x,ynew)∈train(x,y_new) _train do 4: Initialize storage lists: ←[]D←[], rewards←[]rewards←[] 5: 6: for g=1g=1 to G do 7: xcurr←x_curr← x 8: Sequence Rg←[]R_g←[]; LogProbs πgold←[]π^old_g←[] 9: for step j=1j=1 to K do ⊳ Sequentially select demonstrations 10: z←Sθ(xcurr,C)z← S_θ(x_curr,C); p←softmax(z)p (z) 11: Sample action aj∼pa_j p 12: Append aja_j to RgR_g; Append p[aj]p[a_j] to πgoldπ^old_g 13: Remove selected action ata_t from candidate pool tC_t: t+1←t∖atC_t+1 _t \a_t\ 14: xcurr←Concat(xcurr,aj)x_curr (x_curr,a_j) ⊳ Update prompt state 15: end for 16: Prompt LLM to get y^g y_g and reward rgr_g ⊳ Compute composite reward 17: Append rgr_g to rewards 18: Store trajectory data (Rg,πgold)(R_g,π^old_g) in D 19: end for 20: 21: μgroup←1G∑rg _group← 1GΣ r_g ⊳ Compute group baseline 22: σgroup←1G∑(rg−μgroup)2+δ _group← 1GΣ(r_g- _group)^2+δ 23: Ag←(rg−μgroup)/σgroupA_g←(r_g- _group)/ _group for each group g 24: 25: for inner epoch =1=1 to Nepochs′N _epochs do 26: ℒtotal←0L_total← 0 27: for g=1g=1 to G do 28: Reconstruct states xg,0,…,xg,Kx_g,0,…,x_g,K using x and RgR_g 29: for step j in RgR_g do 30: pg,jnew←softmax(Sθ(xg,j−1,C))actionjp^new_g,j (S_θ(x_g,j-1,C))_action_j 31: pg,jold←πgold[j]p^old_g,j←π^old_g[j] 32: ratio←pg,jnew/pg,joldratio← p^new_g,j/p^old_g,j 33: LCLIP←−min(ratio⋅Ag,clip(ratio,1−ϵ,1+ϵ)⋅Ag)L^CLIP←- (ratio· A_g,clip(ratio,1-ε,1+ε)· A_g) ⊳ Compute clipped surrogate objective 34: LKL←βKL⋅log(pg,jnew/pg,jold)L^KL← _KL· (p^new_g,j/p^old_g,j) ⊳ Approximate KL penalty 35: ℒtotal+=(LCLIP+LKL)L_total +=(L^CLIP+L^KL) 36: end for 37: end for 38: θ←θ−α∇θ(1G⋅Kℒtotal)θ←θ-α _θ( 1G· KL_total) ⊳ Gradient descent update 39: end for 40: end for 41: end for Appendix B Proof of the Final Reward via Lagrangian Relaxation We model the problem as Constrained MDP. Our goal is to find a policy πθ _θ that maximizes the Edit Success (ES) while ensuring that the expected costs associated with violating Paraphrase Consistency (PC) and Retention Rate (R) remain below acceptable thresholds. B.1 Problem Formulation Let JR(π)J_R(π) denote the expected return for the primary editing task based on the reward R(st,at)R(s_t,a_t). We introduce cost constraints for PC and R, where JCk(π)J_C_k(π) represents the expected cumulative cost for constraint k∈PC,Rk∈\PC,R\. Let dkd_k represent the maximum tolerable expected cost for that metric. The CMDP optimization problem is defined as: maxπθ _ _θ JR(πθ) J_R( _θ) (9) s.t. .t. JCk(πθ)≤dk,∀k∈PC,R J_C_k( _θ)≤ d_k, ∀ k∈\PC,R\ where the expectations are taken over the trajectory distribution induced by the policy πθ _θ. B.2 Lagrangian Relaxation We apply the method of Lagrange multipliers. We define the Lagrangian objective function ℒ(π,λ)L(π,λ) as: ℒ(π,λ)=JR(π)−∑k∈PC,Rλk(JCk(π)−dk)L(π,λ)=J_R(π)- _k∈\PC,R\ _k(J_C_k(π)-d_k) (10) Here, λ=(λPC,λR)λ=( _PC, _R) is a vector of Lagrange multipliers, with λk≥0 _k≥ 0. The primal constrained problem is equivalent to the following unconstrained min-max dual problem: minλ≥0maxπθℒ(π,λ) _λ≥ 0 _ _θL(π,λ) (11) B.3 Equivalence to Scalarized Reward We can rewrite the Lagrangian function by expanding the expected return and expected cost terms over the trajectories τ∼πτ π. Since expectations are linear, we have: ℒ(π,λ) (π,λ) =τ∼π[∑tR(st,at)]−∑k∈PC,Rλk(τ∼π[∑tCk(st,at)]−dk) =Eτ π [ΣtR(s_t,a_t) ]- _k∈PC,R _k (Eτ π [ΣtC_k(s_t,a_t) ]-d_k ) (12) =τ∼π[∑t(R(st,at)−∑k∈PC,RλkCk(st,at))]+∑k∈PC,Rλkdk =Eτ π [Σt (R(s_t,a_t)- _k∈PC,R _kC_k(s_t,a_t) ) ]+ _k∈PC,R _kd_k Notice that the term ∑kλkdk _k _kd_k is constant with respect to the policy optimization step (the inner maximization loop over π). Therefore, maximizing ℒ(π,λ)L(π,λ) with respect to π is mathematically equivalent to maximizing the expected return of a new, composite reward function: r(st,at)=R(st,at)−∑k∈PC,RλkCk(st,at)r(s_t,a_t)=R(s_t,a_t)- _k∈\PC,R\ _kC_k(s_t,a_t) (13) Thus, by optimizing this composite reward using GRPO with fixed hyperparameters λPC _PC and λR _R, we are equivalently solving the inner loop of the Lagrangian dual problem. The fixed coefficients λk _k act as the Lagrange multipliers that apply static penalty weightings to balance the trade-off between editing performance and the cost constraints. Appendix C Comparison to RL-Based RAG Retrieval To test whether RL-based RAG retrievers can substitute for an IKE-specific approach, we evaluate RAG-RL Huang et al. (2026) on CounterFact with Qwen-2.5-7B-Instruct. RAG-RL was originally trained to optimize answer-F1 and citation-F1 on open-domain QA; we apply it directly to the IKE setting without modification. We observe that the LLM frequently produces refusals (“I cannot determine the answer based on the prompt”) when RAG-RL’s retrieval forces it into contradiction with its parametric knowledge. Methods S ↑ ES ↑ PC ↑ R ↑ RAG-RL 20.0 83.0 79.3 8.0 MO-IKE 78.1 96.0 88.3 60.0 Table 5: RAG-RL evaluated on CounterFact with Qwen-2.5-7B-Instruct. RAG-RL achieves moderate edit success but collapses on retention (8.0% R), consistent with our claim in Section 2. Appendix D Training Configuration To ensure reproducibility, Table 6 details the complete set of hyperparameters used during the training of MO-IKE. Hyperparameter Value Learning Rate 1×10−51× 10^-5 Epochs 3 Group Size 8 Random Seed 42 GRPO Clipping Parameter (ϵε) 0.1 KL Coefficient (β) 0.001 Lagrangian Penalty Weights 1 Table 6: Hyperparameters used for MO-IKE training. Appendix E Training Dynamics To understand the learning efficiency of our RL-based retrieval optimization, we analyze the performance metrics of MO-IKE across training epochs. Figure 5 illustrates the progression of ES, PC and R over the duration of the training process. A key observation is the rapid convergence of the policy. By the conclusion of the first epoch, the model establishes a strong, balanced demonstration ordering, evidenced by the sharp initial peak in both ES and R. Throughout subsequent epochs (Epoch 2 and 3), the performance largely plateaus; ES and PC demonstrate stable refinements, while R maintains its rate without degrading. Because the reward includes a penalty for retention, policy updates are regularized toward smaller deviations, which can reduce unstable exploratory shifts during early RL training Schulman et al. (2017a). The model rapidly identifies a near-optimal structural proportion for the demonstrations, proving that MO-IKE is not only effective at preventing structural collapse but also highly sample-efficient to train. Figure 5: Training dynamics with Llama-3.2. We evaluate the performance of demonstration selection given different training epochs. Appendix F Hyperparameter Sensitivity. We conduct additional experiments to check hyperparameter sensitivity for knowledge editing reward weight λES _ES and constraint weights λPC,λRR _PC, _R. We swept λPC _PC and λRR _R by 100×100× each, holding the other fixed. Shown by Table 7, ESR varies by ≤2.0≤ 2.0 points, PC by ≤3.3≤ 3.3, R is unchanged. MO-IKE still outperforms all baselines at every setting, with no monotonic degradation in any metric. The directional pattern matches the design: upweighting PC modestly improves PC (+1.0), while upweighting R modestly degrades PC (-2.3) without further improving R, indicating (1, 1) sits in a flat region balancing both objectives. Configuration S ↑ ES ↑ PC ↑ R ↑ λPC=100,λRR=1 _PC=100, _R=1 73.2 90.0 80.0 57.7 λPC=1,λRR=100 _PC=1, _R=100 72.7 91.7 76.7 57.7 λPC=1,λRR=1 _PC=1, _R=1 73.4 92.0 79.0 57.7 Table 7: Hyperparameter sensitivity sweep over the Lagrangian penalty weights. Appendix G Additional Experiments Evaluation on Large Datasets. We expand experiments with the UniEdit Chen et al. (2025) dataset, a larger dataset composed of 311K editing examples drawn from a diverse mixture of relation types and knowledge domains, which is roughly a 15× increase in scale over CounterFact. This setting provides a stricter test of whether MO-IKE’s gains generalize beyond the testing of a small benchmark. As indicated by Table 8, MO-IKE achieves the best Score and best R. FactPrompt’s high ESR (81.7%) is misleading: its Score (15.0%), PC (22.7%), and R (7%) are the lowest in the table, indicating FactPrompt does not let the model effectively learn factual updates — the model parrots the injected editing prompt rather than internalizing new knowledge. Method S ↑ ES ↑ PC ↑ R ↑ FactPrompt 15.0 81.7 22.7 7.0 EditCoT 21.5 43.0 36.0 11.0 IKE 19.6 39.3 27.6 11.0 DR-IKE 17.2 38.0 27.3 9.0 MO-IKE 24.5 51.3 28.3 14.3 Table 8: Evaluation on the UniEdit dataset using Llama-3.2-3B-Instruct. LLM Performance.Table 9 presents the performance of MO-IKE across various frozen LLMs. We observe a general positive scaling trend, where performance of Score (S) improves as model size increases. The highest overall rates are achieved by models with larger parameter scales (≥ 7B). Model (Parameters) S↑ ES↑ PC↑ R↑ Qwen 2.5 (7B) 75.0 96.0 88.3 60.0 Qwen 2.5 (1.5B) 63.7 69.7 67.7 54.7 Llama 3.1 (8B) 76.7 88.0 77.7 67.3 Llama 3.2 (3B) 73.5 92.0 79.0 57.7 Mistral v0.3 (7B) 77.4 80.1 90.7 65.0 Table 9: Editing performance across different LLMs. Natural Baselines. We provide natural baselines that showcase the performance of the retriever, as we claim that selection and ordering and demonstrations matter. Precisely, we compare MO-IKE with random shuffling baseline and a duplication baseline where the target edit is repeated multiple times, while keeping UPDATE and RETAIN unchanged. Table10 shows that simple duplication or shuffling yields moderate performance but does not match the balance of ES/PC/R achieved by our multi-objective RL framework. Method ES ↑ PC ↑ R ↑ Random Shuffling 77.0–91.0 70.0–74.0 52.0–60.0 Duplicate Target 2× 79.3 72.3 51.0 Duplicate Target 4× 76.7 73.7 50.7 Duplicate Target 8× 78.3 77.0 52.0 Duplicate COPY 2× 78.3 72.7 52.0 Duplicate COPY 4× 79.0 73.7 49.7 Duplicate COPY 8× 79.7 70.0 51.0 MO-IKE 96.0 88.3 60.0 Table 10: Performance of heuristic prompt engineering strategies on CounterFact. Evaluation on 2025-era-SOTA LLMs. To directly test the long-term applicability of MO-IKE, we conducted an additional evaluation on a recently released 2025-era model, Qwen3-4B-Instruct. Results are summarized in Table 11. We observe that base models exhibit high baseline robustness—for instance, standard static IKE achieves an impressive 98.67% Edit Success Rate (ESR) and 94.67% Paraphrase Consistency (PC) on Qwen-3. Since ESR and PC are near the ceiling in this model, R becomes the key metric in differentiating MO-IKE from all other baselines. We can observe that MO-IKE exceeds the single objective baseline DR-IKE by approximately 23% on R, confirming its effectiveness in preserving neighboring facts in knowledge editing. Moreover, MO-IKE successfully preserves the high ESR and PC of the newer models. Therefore, in terms of the overall performance, MO-IKE achieves the highest overall score (S) of 70.2%. Method S ↑ ES ↑ PC ↑ R ↑ FactPrompt 50.4 91.0 75.3 31.3 EditCoT 42.1 86.1 76.5 24.6 IKE 65.6 98.7 94.7 46.3 DR-IKE 49.6 99.3 96.0 30.0 MO-IKE 70.2 98.3 94.3 53.0 Table 11: Editing performance on Qwen3-4B-Instruct. While stronger base models exhibit high baseline robustness, MO-IKE remains the only framework capable of preserving knowledge specificity (R), achieving the highest overall score. Computation Cost. We evaluate inference latency Llama-3.2-3B-Instruct. The results are detailed in Table 12. Component Mean (ms) Median (ms) Std (ms) p95 (ms) Retrieval 812.2 808.5 36.2 871.8 Generation 112.3 112.8 45.6 181.3 End-to-End 924.4 915.4 55.4 1027.3 Table 12: Inference latency per query evaluated on CounterFact using Llama-3.2-3B-Instruct. Benchmarked on a single NVIDIA A100-SXM4-40GB GPU at batch size 1. Appendix H Case Study Component IKE DR-IKE MO-IKE Edit Request Target Edit: Ivan Ivanov-Vano spoke the language Russian (ytruey_true) → French (ynewy_new) Selected Demonstrations [COPY] Andrey Malakhov spoke the language French. [COPY] Boris Shaposhnikov spoke the language French. … [UPDATE] Andrey Malakhov is a native speaker of French. [UPDATE] Fyodor Pavlovich Reshetnikov, a native French. … [RETAIN] Roger Vitrac spoke the language Russian. [RETAIN] Christophe Moreau spoke the language Russian. … (Note: Fixed category order.) [COPY] Andrey Malakhov spoke the language French. [COPY] Boris Shaposhnikov spoke the language French. … [UPDATE] Andrey Malakhov is a native speaker of French. [UPDATE] Fyodor Pavlovich Reshetnikov, a native French. … (Note: Fixed COPY and UPDATE order. No RETAIN candidates, as they are filtered due to penalty on non-edit-maximizing signals) [COPY] Andrey Malakhov spoke the language French. [RETAIN] Christophe Moreau spoke the language Russian. [UPDATE] Fyodor Pavlovich Reshetnikov, a native French. … [UPDATE] Andrey Malakhov is a native speaker of French. … [RETAIN] Roger Vitrac spoke the language Russian. [COPY] Boris Shaposhnikov spoke the language French. (Note: Unfixed order. Balanced structural proportion maintained via soft embedding) Reliability Metric Ivan Ivanov-Vano spoke the language? Russian. (Unsuccessful edit due to heuristic prompt construction) French. (Success) French. (Success) Specificity Metric Vladimir Smirnov, speaker of? Russian. (Success) French. (Unsuccessful retention due to unbalanced RETAIN demonstrations) Russian. (Success) Table 13: Comparison of in-context editing frameworks highlighting the structure of the constructed prompt. Table 13 use the record 676 in Counterfact to illustrate the relationship between the structural composition of retrieved demonstrations and the resulting performance tradeoffs in different in-context knowledge editing frameworks. We use edit success to represent Reliability Metric, and retention rate to represent Specificity Metric. IKE. The baseline IKE framework relies on fixed prompt construction—statically ordering COPY, UPDATE, and RETAIN demonstrations. As shown in the Reliability Metric, this static template approach fails to successfully execute the target edit (ytrue→ynewy_true→ y_new). Because the retriever does not optimize for the edit signal, the model defaults to its pre-trained knowledge (Russian) rather than adopting the new fact. DR-IKE. While DR-IKE successfully forces the edit (outputting “French” for Ivanov-Vano), it suffers from a “over-optimization” issue. By penalizing only on non-edit-maximizing signals, the policy optimization entirely discards RETAIN candidates. This creates an unbalanced prompt heavily skewed toward the target concept. Consequently, DR-IKE exhibits fact forgetting, failing the Specificity metric by incorrectly altering the language of a related but distinct entity (Vladimir Smirnov) to French. MO-IKE. MO-IKE resolves the tradeoff between edit success and knowledge retention. By avoiding rigid templates and maintaining a balanced structural proportion of demonstration types, MO-IKE allows the retriever to dynamically interleave UPDATE, RETAIN, and COPY examples. This unfixed ordering provides sufficient contextual protection for unrelated facts while still providing a strong enough edit signal to update the target knowledge. As a result, MO-IKE is the only framework in the case study that successfully achieves both Reliability (updating Ivanov-Vano to French) and Specificity (retaining Russian for Smirnov).