Paper deep dive
Bringing Model Editing to Generative Recommendation in Cold-Start Scenarios
Chenglei Shen, Teng Shi, Weijie Yu, Xiao Zhang, Jun Xu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/22/2026, 5:07:16 AM
Summary
GenRecEdit is a novel model editing framework designed to address 'cold-start collapse' in generative recommendation (GR) systems. By treating semantic ID patterns of cold-start items as editable knowledge, it enables training-free, on-the-fly updates to GR models. The framework utilizes position-wise knowledge preparation, a locate-then-edit mechanism, and a one-to-one trigger policy to inject multi-token item representations without retraining, achieving significant performance gains with only 9.5% of the computational cost of traditional retraining.
Entities (5)
Relation Signals (3)
GenRecEdit → addresses → Cold-start collapse
confidence 98% · To address these challenges, we propose GenRecEdit, a model editing framework tailored for generative recommendation.
GenRecEdit → improves → Generative Recommendation
confidence 95% · Extensive experiments on multiple datasets show that GenRecEdit substantially improves recommendation performance on cold-start items
GenRecEdit → utilizes → Semantic IDs
confidence 90% · we mitigate the absence of sentence structure by explicitly modeling the intrinsic relationship between the entire sequence context and the next-token (e.g., semantic IDs, SIDs) generation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generative recommendation (GR) has shown strong potential for sequential recommendation in an end-to-end generation paradigm. However, existing GR models suffer from severe cold-start collapse: their recommendation accuracy on cold-start items can drop to near zero. Current solutions typically rely on retraining with cold-start interactions, which is hindered by sparse feedback, high computational cost, and delayed updates, limiting practical utility in rapidly evolving recommendation catalogs. Inspired by model editing in NLP, which enables training-free knowledge injection into large language models, we explore how to bring this paradigm to generative recommendation. This, however, faces two key challenges: GR lacks the explicit subject-object binding common in natural language, making targeted edits difficult; and GR does not exhibit stable token co-occurrence patterns, making the injection of multi-token item representations unreliable. To address these challenges, we propose GenRecEdit, a model editing framework tailored for generative recommendation. GenRecEdit explicitly models the relationship between the full sequence context and next-token generation, adopts iterative token-level editing to inject multi-token item representations, and introduces a one-to-one trigger mechanism to reduce interference among multiple edits during inference. Extensive experiments on multiple datasets show that GenRecEdit substantially improves recommendation performance on cold-start items while preserving the model's original recommendation quality. Moreover, it achieves these gains using only about 9.5% of the training time required for retraining, enabling more efficient and frequent model updates.
Tags
Links
- Source: https://arxiv.org/abs/2603.14259v1
- Canonical: https://arxiv.org/abs/2603.14259v1
Trouble viewing inline? Open PDF directly →
Full Text
61,254 characters extracted from source content.
Expand or collapse full text
Bringing Model Editing to Generative Recommendation in Cold-Start Scenarios Chenglei Shen Gaoling School of Artificial Intelligence, Renmin University of ChinaBeijingChina chengleishen9@ruc.edu.cn , Teng Shi Gaoling School of Artificial Intelligence, Renmin University of ChinaBeijingChina shiteng@ruc.edu.cn , Weijie Yu School of Information Technology and Management, University of International Business and EconomicsBeijingChina yu@uibe.edu.cn , Xiao Zhang Gaoling School of Artificial Intelligence, Renmin University of ChinaBeijingChina zhangx89@ruc.edu.cn and Jun Xu Gaoling School of Artificial Intelligence, Renmin University of ChinaBeijingChina junxu@ruc.edu.cn (2018) Abstract. Generative recommendation (GR) has demonstrated substantial potential for sequential recommendation in an end-to-end generation paradigm. However, existing models suffer from severe cold-start collapse, i.e., recommendation accuracy on cold-start items drops to near zero. Current solutions rely on retraining with cold-start interactions, which faces sparse feedback, high computational cost, and delayed updates, diminishing practical utility in rapidly evolving catalogs in recommendation. Inspired by NLP model editing (enabling training-free knowledge injection into large language models), we explore applying this paradigm to generative recommendation, but face two fundamental challenges: (1) sequential data in GR lacks explicit subject–object binding (a core NLP sentence structure), hindering targeted model edits; (2) sequential data in GR has no fixed token co-occurrence patterns (unlike NLP’s stable phrases), making multi-token injection unreliable. To address these, we propose GenRecEdit, the first model editing framework tailored for generative recommendation. Specifically, we: (1) mitigate the absence of sentence structure by explicitly modeling the intrinsic relationship between the entire sequence context and the next-token (e.g., semantic IDs, SIDs) generation; (2) adopt iterative token-level edits to effectively inject token bundles (e.g., items); and (3) introduce a One-One trigger mechanism to avoid interactions among multiple token-level edits during inference. Extensive experiments across multiple datasets demonstrate that GenRecEdit substantially improves recommendation performance on cold-start items while preserving the model’s original recommendation quality. Moreover, GenRecEdit achieves these gains with only approximately 9.5% of the training time required by retraining, significantly reducing computational cost and enabling efficient, frequent model updates. Generative Recommendation, Model Editing †copyright: acmlicensed†journalyear: 2018†doi: X.X†conference: Make sure to enter the correct conference title from your rights confirmation email; June 03–05, 2018; Woodstock, NY†isbn: 978-1-4503-X-X/2018/06†ccs: Information systems Recommender systems 1. Introduction Figure 1. An illustration of cold-start collapse in GR. The left panel presents dataset statistics that characterize cold-start collapse, while the right panel summarizes the adverse effects on cold-start items. The bottom panel shows the time-efficiency of GenRecEdit in terms of model update cost. Generative recommendation (GR) has recently emerged as a promising paradigm for sequential recommendation tasks (Rajput et al., 2023; Wang et al., 2024a; Zheng et al., 2024). In GR, each item is tokenized into a small number of discrete semantic tokens (semantic IDs), and a sequential model is trained to autoregressively generate the next tokens, which are then parsed to form the predicted item. By reformulating recommendation as a sequence-to-sequence generation problem, GR departs from conventional discriminative sequential recommendation methods (Kang and McAuley, 2018; Sun et al., 2019a), enabling improved scalability and superior performance (Zhou et al., 2025; Deng et al., 2025), while directly leveraging the optimization advances of large language models (Zhai et al., 2024). Unlike item-ID–based transductive models (Kang and McAuley, 2018; Sun et al., 2019a), which are unable to recommend newly introduced items due to the absence of their IDs in the trained model, generative recommendation (GR) models are designed to generalize beyond the closed item vocabulary by exploiting semantic correlations and generating semantic tokens for previously unseen items (i.e., cold-start items). However, in practice, GR models often perform poorly on cold-start items. As illustrated in Figure 1, we partition the test set into warm and cold subsets based on whether each test item appears in the training data. Evaluation with well-trained GR models reveals a striking failure mode: when cold-start items are introduced after training, recommendation accuracy on cold-start items can drop to nearly zero, a phenomenon we term cold-start collapse in GR. To investigate this issue, we first pose a fundamental question: What causes this collapse? Our analysis in Section 3.1 leads to two key findings. (1) Well-trained GR models can often correctly generate the first semantic ID token of cold-start items, indicating substantial latent potential for cold-start recommendation. (2) Regardless of whether the final recommendation is correct, GR models exhibit a strong bias toward generating seen semantic ID patterns (Yang et al., 2024; Ding et al., [n. d.]). A straightforward solution is to collect feedback on cold-start items and then retrain or incrementally finetune the model. However, this paradigm is limited in practice. For example, in scenarios where up-to-date recommendations are critical (e.g., news or short-video platforms), incremental training or retraining faces substantial challenges, including sparse feedback for cold-start items, delayed model updates (Hou et al., 2025b), and prohibitive retraining costs. These findings and constraints underscore an urgent need for a method that enables on-the-fly incorporation of semantic ID patterns for cold-start items without compromising the inherent recommendation capability of GR models. Figure 2. An illustration of the challenges adapting model editing from NLP to generative recommendation (GR). Based on the above analysis, we naturally ask: How to efficiently resolve this collapse? Motivated by recent progress in model editing for natural language processing, which enables training-free knowledge injection into large language models, we explore whether a similar paradigm can be applied to generative recommendation by treating the semantic ID patterns of cold-start items as editable knowledge and injecting them into GR models. However, directly transferring existing model editing techniques to sequential recommendation data faces two fundamental challenges, as shown in Figure 2. First, GR sequences lack explicit sentence structure, especially clear subject–object binding, which makes it difficult to control the target item by editing a well-defined subject representation. For example, in NLP task, if we aim to change the American president from “Joe Biden” to “Donald Trump” given the context “The American president is …”, a common strategy is to localize the subject phrase (“American president”) via the subject–object binding, and edit its hidden representation so that decoding favors the desired object tokens (“Donald Trump”). In GR, however, such a clear subject–object binding is typically absent due to the lack of explicit sentence structure. Second, natural language often contains stable token bundles (i.e., phrases) with high co-occurrence probability (e.g., ‘Donald’ and ‘Trump’), benefiting from extensive corpus-level pretraining. In contrast, GR sequences do not exhibit fixed co-occurrence patterns for cold-start items because the model has not observed their SID patterns during training. Consequently, injecting multi-token sequences becomes unreliable. Therefore, the key question is: How to overcome these challenges? To this end, we propose GenRecEdit, a model editing framework tailored for generative recommendation. Specifically, (1) we mitigate the absence of sentence structure by explicitly modeling the intrinsic relationship between the entire sequence context and the next-token (e.g., a single semantic ID); (2) we adopt iterative token-level edits (i.e., position-wise edits) to effectively inject token bundles (i.e., semantic ID patterns of cold-start items); and (3) we introduce a One-One trigger mechanism to avoid interactions among multiple token-level edits during inference. We summarize our contributions as follows: • We reveal a striking cold-start collapse in generative recommendation: accuracy on cold-start items can drop to near zero, and our analysis shows that GR models have the potential to recommend cold-start items but default to seen semantic-ID patterns (Section 3.1, Figure 1). • We propose GenRecEdit, which treats cold-start semantic-ID patterns as editable knowledge and injects them into GR models in a training-free, on-the-fly manner, while preserving the model’s existing recommendation capability. • Extensive experiments on multiple datasets demonstrate that GenRecEdit substantially improves cold-start recommendation with minimal impact on the original general recommendation ability, while incurring only 9.5% of the retraining time cost. 2. Related Work 2.1. Generative Recommendation Motivated by the strong empirical performance of large language models (LLMs) (Zhao et al., 2023), generative recommendation has attracted growing interest (Hua et al., 2023; Rajput et al., 2023; Zheng et al., 2024; Wang et al., 2024a; Hou et al., 2025a; Deng et al., 2025; Zhou et al., 2025). In this paradigm, each item is represented as a sequence of discrete semantic IDs (SIDs); the interaction history is then a concatenation of item SIDs, and a generative model predicts the target item’s SIDs. Generative recommendation typically decouples into item tokenization and autoregressive generation: tokenization assigns SIDs from item meta semantics, while the autoregressive model learns intra-item SID patterns and inter-item sequential dependencies. Existing tokenization schemes (e.g., clustering (Si et al., 2024; Wang et al., 2024b) and vector quantization (Rajput et al., 2023; Deng et al., 2025; Zhou et al., 2025; Zhu et al., 2024)) allow SIDs for cold-start items to be constructed offline, alleviating the embedding-missing issue in ID-based recommenders. However, a key bottleneck remains: the autoregressive model has not observed the SID patterns of cold-start items during training, so it often fails to generate their SIDs at inference. We address this by applying model editing to inject cold-start SID patterns while preserving previously learned knowledge, thereby improving cold-start generation accuracy. 2.2. Model Editing Model editing has emerged as a lightweight, training-free paradigm for inserting, modifying, and removing knowledge in large language models (Meng et al., 2022b; Fang et al., 2024). It is motivated by the view that knowledge is propagated by attention and stored in linear layers—especially feed-forward networks (FFNs)—which are thus the main targets for intervention (Geva et al., 2021; Meng et al., 2022b, a). Most methods follow a locate-then-edit pipeline: they first localize the edit layer and then exploit the linear mapping from FFN hidden-state shifts to weight updates. Representative approaches include ROME (Meng et al., 2022a), which analyzes how factual knowledge is encoded and shows that internal mechanisms can be directly manipulated; MEMIT (Meng et al., 2022b), which enables large-scale edits at specific layers while maintaining specificity and fluency; and α-Edit (Fang et al., 2024), which mitigates the update–preservation trade-off via null-space projection. AnyEdit (Jiang et al., 2025) extends editing to unstructured knowledge through token-level hidden-state edits, but remains focused on textual, non-lifelong settings. With the rise of generative recommendation, the boundary between recommender systems and large language models is increasingly blurred, and cold-start items naturally resemble newly introduced knowledge in LLMs. Motivated by this analogy, we propose GenRecEdit, the first framework to systematically apply model editing to generative recommendation. 3. Problem Analysis and Formulation This section analyzes the causes of cold-start collapse in Generative Recommendation (GR) and formulates the problem of model editing for GR. 3.1. Problem Analysis Figure 3. The analysis of the cold-start collapse. Left: NDCG at the first n positions when a GR model generates a four-position semantic ID on the Cell Phones and Accessories dataset. Right: Distribution of generated items regardless of recommendation correctness; a higher IID Ratio indicates a larger fraction of generated items that belong to the current test subset (i.e., cold subset or warm subset). We focus on a key question: what causes the cold-start collapse in GR? Our analysis suggests that GR models retain substantial potential for recommending cold-start items. The primary bottleneck is not a lack of general recommendation capability, but rather the model’s inability to generate semantic ID patterns that it has rarely observed for cold-start items. To substantiate this claim, we design two experiments. (1) Instead of treating a full four-token semantic ID as an item, we treat the prefix consisting of the first n semantic-ID tokens as an item and examine how NDCG changes when n∈0,1,2,3n∈\0,1,2,3\. The results are shown in Figure 3 (Left). (2) We define a metric termed IIDRatio@KIID~Ratio@K, which ignores recommendation correctness and measures the fraction of the top-K generated items that belong to in-distribution items (e.g., the whole cold-start items for cold subset) in the current test set. Formally, (1) IIDRatio@K=1||∑u∈|Riid∩R^u,K||Riid|,IID\ Ratio@K= 1|U| _u |R_iid∩ R_u,K | |R_iid |, where U denotes the set of users, RiidR_iid is the item set associated with the current test split, and R^u,K R_u,K is the set of top-K items generated by the GR model for user u. The results are shown in Figure 3 (Right). These two experiments lead to the following observations. (1) The first-token recommendation accuracy on cold-start items is relatively high, indicating that the GR model can often identify the coarse category of a cold-start item and generate a plausible initial semantic-ID token. However, the generation becomes increasingly unstable as the model proceeds to complete the full semantic ID. (2) IIDRatio@KIIDRatio@K reveals the source of this instability: the large gap between warm and cold-start settings indicates that the model rarely generates cold-start items at all, regardless of whether the generated recommendation is correct. Figure 4. Overall framework of GenRecEdit, which consists of three main modules: (1) Position-Wise Knowledge Preparation. We construct pseudo interaction data for cold-start items and use it to form edit requests. (2) Locate-Then-Edit Framework. For each position, we first use a classifier to localize the key layer that is most strongly associated with that position, and then perform targeted edits on the identified layer. (3) One-One Triggering Policy. To prevent interference among edits for different positions, we adopt an adaptive triggering strategy at inference time, where the corresponding edit layer is triggered according to the current decoding position. 3.2. Model Editing in GR Model editing aims to update a model’s behavior on a small set of specified inputs without retraining. In a standard formulation, the edit requests contain a set of input (i.e., contextcontext) together with desired target outputs, and the goal is to modify model parameters so that the model produces the new targets while preserving its behavior elsewhere. Formally, an edit request is typically specified as a subject, relation, and object triple, denoted as ⟨s,r,o⟩ s,r,o . Our objective is to use (s,r)(s,r) as the conditioning context to maximize the logit (equivalently, the likelihood) of the target object o. Prior studies have shown that FFNs are closely related to knowledge storage in Transformer models (Geva et al., 2021; Dai et al., 2022); accordingly, many studies achieve this goal by editing the FFN modules (Meng et al., 2022a, b; Fang et al., 2024). We first outline the model editing process and then formalize it in the context of GR. Information Flow in FFN In a transformer architecture, given a hidden state h∈ℝdhh ^d_h at the edited position, the FFN first computes an intermediate activation and then projects it back to the output space: (2) k=σ(Winh)∈ℝd0,z=Woutk∈ℝd1,k=σ(W_inh) ^d_0, z=W_outk ^d_1, where Win∈ℝd0×dhW_in ^d_0× d_h, Wout∈ℝd1×d0W_out ^d_1× d_0, and σ(⋅)σ(·) is the element-wise activation. k is referred to as the key and z as the value in the key–value memory of FFNs. Editing Targets Assume we have m edit requests. For the i-th request, we obtain an activation (key) vector kik_i by running the model on the corresponding edit context and extracting the FFN intermediate activation at the subject position (sis_i) from the selected layer. We denote this original FFN output by zi=Woutkiz_i=W_outk_i. To define the desired edited output at this site, we solve for a minimal intervention δi _i on the FFN output that makes the overall model prefer the target object sequence associated with the edit request: (3) δi=argminδ−logpθ(oi|(si,ri);zi+δi), _i= _δ\ - p_θ\! (o_i\, |\,(s_i,r_i);\ z_i+ _i ), where (si,ri)(s_i,r_i) represents the context of the edit request. pθ(⋅∣contexti;zi+δi)p_θ(· _i;\ z_i+ _i) denotes the model’s output distribution when, at the chosen site, the FFN output is replaced by zi+δiz_i+ _i (all other computations remain unchanged). The goal of model editing is to encode such targeted changes in hidden states into the model parameters, i.e., to apply a structured update to WoutW_out. We then set the edited value vector as zi′≜zi+δiz _i z_i+ _i. Typically, to inject new knowledge into the model, we seek an ideal update matrix ΔWout∗ W_out^* such that, for any i-th request among the m edit requests, the following constraint holds: (4) (Wout+ΔWout∗)ki=zi′. (W_out+ W_out^* )k_i=z _i. For convenience, we refer to ΔWout W_out as the estimated update matrix that approximates ΔWout∗ W^*_out, and set W^out≜Wout+ΔWout W_out W_out+ W_out. Preliminary in GR In the natural language setting, the goal is to edit the value z at the subject position by solving for an intervention δ that maximizes the probability of the target output (i.e., the object), thereby yielding a corresponding parameter update ΔWout W_out. However, for sequential data in generative recommendation (GR), two challenges arise: (1) the lack of explicit sentence structure makes it difficult to construct an effective edit request ⟨s,r,o⟩ s,r,o ; and (2) the GR target (e.g., cold-start item patterns) does not exhibit stable token co-occurrence, making a single edit that simultaneously improves all tokens in the pattern unreliable. To address these issues, we treat a single token as the object o to circumvent the challenge posed by the lack of stable token co-occurrence; we define the entire item history together with the prefix of the cold-start item as the subject s and drop the relation r to overcome the challenge of lack explicit subject. Formally, assume that each item is represented by a four-digit semantic ID (SID), with position p∈0,1,2,3p∈\0,1,2,3\. We denote the object token at position p as opo_p. Accordingly, each edit request of GR can be written in the following form: ⟨sp,op⟩ s_p,o_p . These small modifications have far-reaching implications, which we detail in the following sections. 4. GenRecEdit: The Proposed Approach This section introduces our proposed framework, GenRecEdit, which adapts model editing to generative recommendation. As illustrated in Figure 4, GenRecEdit consists of three components: Position-Wise Knowledge Preparation, Locate-Then-Edit framework, and One-One Triggering Policy. We briefly summarize their roles as follows. (1) Position-Wise Knowledge Preparation. We construct pseudo interaction data for cold-start items and use it to form edit requests. (2) Locate-Then-Edit Framework. For each position, we first use a classifier to localize the key layer that is most strongly associated with that position, and then perform targeted edits on the identified layer. (3) One-One Triggering Policy. To prevent interference among edits for different positions, we adopt an adaptive triggering strategy at inference time, where the corresponding edit layer is triggered according to the current decoding position. 4.1. Position-Wise Knowledge Preparation To address this challenge, we construct pseudo histories for cold-start items via similarity-based imputation, as illustrated on the left side of Figure 4. Concretely, for each cold-start item c, we first encode its available semantic information (often the only information accessible for cold-start items) into an embedding: (5) c=fenc(c), e_c=f_enc( m_c), where c m_c denotes the meta information of item c and fenc(⋅)f_enc(·) is an encoder111we use Sentence T5 as the encoder.. Next, we retrieve the top-k most similar warm items based on embedding similarity. For each warm item j∈ℐwarmj _warm with embedding j e_j, we compute (6) sim(c,j)=cos(c,j)=c⊤j∥c∥2∥j∥2,sim(c,j)= ( e_c, e_j)= e_c e_j e_c _2\, e_j _2, and select the top-k warm items that are most similar to each cold-start item: (7) k(c)=TopKj∈ℐwarmsim(c,j).N_k(c)=TopK_j _warm\ sim(c,j). We use k(c)N_k(c) as a set of reference items to approximate how the cold-start item may be interacted with in real-world scenarios. Finally, for each retrieved warm item j∈k(c)j _k(c), we extract the user interactions that precede j in the user’s real interaction sequence and use these preceding interactions as a surrogate interaction history for c. By doing so, we obtain a set of synthesized, high-quality pseudo interaction sequences for cold-start items. We then split each target item by position to further construct position-wise edit requests, i.e., pseudo knowledge pairs ⟨sp,op⟩ s_p,o_p , where sps_p consists of the interaction history together with the prefix of the target item, and opo_p is the target token at the corresponding position. These pairs can be directly used for model editing in GR. 4.2. Locate-Then-Edit Framework As illustrated in the middle stage of Figure 4, we adopt a locate-then-edit framework, which first identifies the most critical layers and then performs targeted updates. In this section, we describe the procedure from three aspects: (1) Layer location. GenRecEdit injects the knowledge of cold-start items by updating a small subset of relevant layers. Hence, we aim to identify the layers that are most responsible for representing the target knowledge, so that we can introduce cold-start items while minimizing disruption to the model’s original recommendation capability. (2) Memory construction. We compute the key (k) and value (z′z ) of the cold-start items that the selected layers should store, i.e., the memory vectors that encode the desired behavioral changes. (3) Parameter updating. We then update parameters to store a portion of the constructed memories in each selected layer to achieve effective knowledge insertion with fewer side effects. 4.2.1. Layer Location Recent studies have shown that editing FFN modules at different layers can affect model capabilities to varying degrees, making layer selection a critical design choice (Geva et al., 2021; Dai et al., 2022). Some prior work identifies a small set of “most sensitive” layers by measuring how perturbations to layer-wise hidden states influence the final outputs; however, such sensitivity-based selection can compromise the model’s original capabilities when edits are applied (Li et al., 2023; Fang et al., 2024). Our goal is to inject cold-start item knowledge while minimizing degradation of the model’s original recommendation performance. Notably, many model-editing methods can be viewed as learning a parameter update ΔW W that linearly maps the activation (key) k associated with both new and existing knowledge to their desired FFN outputs z. The success of such an update depends crucially on whether the activations (keys) corresponding to new knowledge are distinguishable from those of preserved knowledge in the target layer’s representation space. Therefore, our key idea is to train a linear probing classifier on the key to discriminate between positive and negative samples, following established probing frameworks (Belinkov, 2022). Because our synthesized data are position-wise, for each edit request at position p we extract the activation (key) at the last token of the subject sps_p from each layer l, denoted as kplk_p^l, and assign it the label 11. In addition, we apply the same procedure to samples from the original training set, collecting their corresponding key activations and assigning them the label 0. This yields a probing dataset: (p,l)=(kp,il,yi)i=1n+mD^(p,l)=\(k^l_p,i,y_i)\_i=1^n+m where yi=0y_i=0 if it comes from the original dataset (i.e., i∈1,…,Ni∈\1,…,N\), and yi=1y_i=1 if the sample comes from the m edit requests (i.e., i∈n+1,…,n+mi∈\n+1,…,n+m\). To identify the most distinguishable layer for each position p, we define a linear classifier Cls(kp,il)=Sigmoid(⟨θ,kp,il⟩)Cls(k^l_p,i)=Sigmoid( θ,k^l_p,i ) to detect whether the activation (key) originates from cold-start items (edited knowledge) or from the original knowledge distribution. We split the dataset 8:2 into training and validation and train a binary linear classifier Cls(⋅)Cls(·) on the training split. For each position, we take the top-11 layer by accuracy as the edit layer denoted as lpl_p. 4.2.2. Memory Construction For each position in the edit requests, we determine an edit layer lpl_p according to Section 4.2.1, at which the new knowledge should be fully injected. Specifically, for the i-th edit request (sp,i,op,i)(s_p,i,o_p,i) of total m edit requests, we compute a hidden vector zi′z _i to replace the original output ziz_i by the FFN of the layer lpl_p such that adding δi=zi′−zi _i=z _i-z_i to the hidden state at layer lpl_p and the generated token T (i.e., the SID token of the cold-start item at position p) fully conveys the desired knowledge. For each edit request ⟨sp,op⟩i s_p,o_p _i, to obtain the corresponding δi _i at the edit layer lpl_p, we optimize δi _i according to Eq. (3) with the following loss: (8) ℒedit(δi)=ℒCE(onehot(op,i),pθ(⋅∣sp,i;zi+δi)),L_edit( _i)=L_CE (onehot(o_p,i),\ p_θ (· s_p,i;\ z_i+ _i ) ), where onehot(⋅)onehot(·) denotes the one-hot encoding of the token op,io_p,i, i.e., onehot(op,i)∈ℝ||onehot(o_p,i) ^|V|, where |||V| is the vocabulary size. ℒCEL_CE denotes the cross-entropy loss. 4.2.3. Parameter Updating In generative recommendation, we treat cold-start items as new knowledge to be injected into a pretrained model while preserving its performance on the warm items. For the m edit requests at position p, following Section 4.2.2, we obtain kii=0m−1\k_i\_i=0^m-1, zii=0m−1\z_i\_i=0^m-1, and δii=0m−1\ _i\_i=0^m-1, where zi′=zi+δiz _i=z_i+ _i. If we stack keys and memories as matrices K1=[k1∣k2∣⋯∣km]K_1=[k_1 k_2 ·s k_m] and Z1′=[z1′∣z2′∣⋯∣zn′]Z _1=[z _1 z _2 ·s z _n], then the parameter update process could be optimizing by the following expanded least-squares objective: (9) W1≜argminW^(‖W^K0−Z0‖F2+‖W^K1−Z1′‖F2),W_1 _ W ( WK_0-Z_0 _F^2+ WK_1-Z _1 _F^2 ), where (K0,Z0)(K_0,Z_0) are key–value pairs extracted from the original training data, and (K1,Z1′)(K_1,Z _1) are key–value pairs constructed for the edit requests. The first term enforces that the updated weights maintain the model’s original mapping on K0K_0, thereby preserving the original recommendation capability on warm items, while the second term encourages the updated weights to realize the desired associations for cold-start items, thereby injecting new knowledge. Applying the normal equation in block form yields (10) W1[K0K1][K0K1]⊤=[Z0Z1′][K0K1]⊤,W_1\,[K_0\ \ K_1]\,[K_0\ \ K_1] =[Z_0\ \ Z _1]\,[K_0\ \ K_1] , which expands to (11) (W0+ΔW)(K0K0⊤+K1K1⊤)=Z0K0⊤+Z1′K1⊤.(W_0+ W)\,(K_0K_0 +K_1K_1 )=Z_0K_0 +Z _1K_1 . Subtracting the pretraining normal equation W0K0K0⊤=Z0K0⊤W_0K_0K_0 =Z_0K_0 gives the key trade-off relation: (12) ΔW(K0K0⊤+K1K1⊤)=Z1′K1⊤−W0K1K1⊤, W\,(K_0K_0 +K_1K_1 )=Z _1K_1 -W_0K_1K_1 , showing that ΔW W must correct the residual of the cold-start associations under the old weights, while being mediated by both the covariance of keys from the original training data (K0K0⊤K_0K_0 , reflecting preservation pressure) and the covariance of keys from cold-start items (K1K1⊤K_1K_1 , reflecting injection pressure). Defining C0≜K0K0⊤C_0 K_0K_0 and R≜Z1′−W0K1R Z _1-W_0K_1, we obtain the closed-form update: (13) ΔW=RK1⊤(λC0+K1K1⊤)−1, W=RK_1 (λ C_0+K_1K_1 )^-1, where λ controls the preservation–injection trade-off. 4.3. One-One Triggering Policy in Inference To accommodate the unique data characteristics in GR (i.e., the absence of stable token bundles), we adopt a position-wise knowledge preparation and editing scheme as illustrated above. Under this setting, a new challenge emerges at inference time: generating a single item typically requires producing multiple SID tokens (four digits in GenRecEdit). Since we perform edits for each position p on a specific edit layer lpl_p, our editing procedure is optimized to affect only the token at position p (i.e., opo_p), without explicitly accounting for its potential influence on tokens at other positions. However, following prior practice (Meng et al., 2022b), the edits on all designated layers (i.e., lpp∈[0,3]\l_p\_p∈[0,3]) would otherwise be simultaneously triggered during decoding. Such concurrent triggering can introduce uncontrolled cross-position interactions, leading to unpredictable token outputs and undermining the reliability of position-wise editing. To address this issue, we propose a One-One Triggering Policy. As illustrated in Figure 4 (Right), we introduce a gating mechanism that enforces edit-inference alignment: when generating the SID token at position p, the forward pass triggers only the edit associated with layer lpl_p, while the edits for all other positions remain inactive. This One-One activation prevents inter-position coupling effects, ensuring that each position is influenced exclusively by its intended edit and thereby stabilizing multi-token SID generation in GR. 5. Experiments We conducted comprehensive experiments on three Amazon datasets. The results and analyses demonstrate the effectiveness of GenRecEdit in injecting cold-start item patterns. 5.0.1. Dataset We evaluate our method on three categories from the widely used Amazon 2023 Review 222https://jmcauley.ucsd.edu/data/amazon/ dataset (Hou et al., 2024): Video Games, Software, and Cell Phones and Accessories. Following (Rajput et al., 2023; Hou et al., 2023), we treat each user’s historical reviews as interaction records and sort them chronologically, with the earliest review placed first. For data processing, we adopt the standard timestamp-based protocol (Kang and McAuley, 2018; Rajput et al., 2023; Ding et al., [n. d.]): we use the official dataset splits that are strictly partitioned into training, validation, and test sets by time, with no temporal overlap. On top of this protocol, we remove cold-start users from the training data to avoid confounding effects when analyzing cold-start items. Specifically, we collect the set of users appearing in the test split and filter the training and validation splits by retaining only interactions from these users. For evaluation, to clearly quantify performance on cold-start items, we further partition the test set based on whether the target item appears in the training split, resulting in a cold subset, a warm-item test set, and an overall test set that contains both. Across the three datasets, the cold subset accounts for 65.6%65.6\%, 24.3%24.3\%, and 75.5%75.5\% of the overall test set, respectively. Table 1. The overall recommendation performance of various methods across the three datasets (Overall) and the cold subsetresults (Cold). The best-performing and second-best methods in Overall are denoted with boldface and underlining, respectively. The “-” indicates that the item ID-based method cannot handle cold-start items, and we could treat its value as 0. The improvements over the second-best methods are statistically significant (paired t-test, p-value<0.05<0.05). Metric Item ID-based Semantic ID-based Cold-start-based SASRec BERT4Rec VQ-Rec TIGER LC-Rec Retrain Finetune SpecGR Ours Overall Cold Overall Cold Overall Cold Overall Cold Overall Cold Overall Cold Overall Cold Overall Cold Overall Cold Video NDCG@10 0.0014 – 0.0019 – 0.0027 0.0000 0.0070 0.0020 0.0055 0.0012 0.0083 0.0044 0.0076 0.0103 0.0114 0.0051 0.0118 0.0123 NDCG@20 0.0019 – 0.0028 – 0.0035 0.0000 0.0096 0.0030 0.0072 0.0017 0.0109 0.0058 0.0093 0.0125 0.0136 0.0062 0.0140 0.0141 NDCG@50 0.0027 – 0.0044 – 0.0046 0.0001 0.0137 0.0052 0.0099 0.0028 0.0158 0.0096 0.0124 0.0164 0.0158 0.0075 0.0182 0.0176 RECALL@10 0.0028 – 0.0046 – 0.0053 0.0000 0.0141 0.0043 0.0109 0.0026 0.0167 0.0094 0.0141 0.0191 0.0237 0.0106 0.0210 0.0209 RECALL@20 0.0051 – 0.0081 – 0.0085 0.0001 0.0243 0.0083 0.0177 0.0045 0.0270 0.0148 0.0209 0.0281 0.0409 0.0195 0.0299 0.0283 RECALL@50 0.0089 – 0.0163 – 0.0141 0.0006 0.0455 0.0197 0.0318 0.0102 0.0521 0.0344 0.0364 0.0475 0.0953 0.0504 0.0510 0.0457 Software NDCG@10 0.0321 – 0.0308 – 0.0263 0.0002 0.0353 0.0005 0.0351 0.0016 0.0350 0.0008 0.0347 0.0303 0.0328 0.0181 0.0370 0.0228 NDCG@20 0.0412 – 0.0412 – 0.0333 0.0006 0.0451 0.0005 0.0456 0.0017 0.0455 0.0009 0.0437 0.0318 0.0381 0.0222 0.0475 0.0249 NDCG@50 0.0522 – 0.0580 – 0.0442 0.0017 0.0604 0.0013 0.0603 0.0024 0.0606 0.0552 0.0552 0.0357 0.0447 0.0285 0.0601 0.0286 RECALL@10 0.0673 – 0.0622 – 0.0493 0.0004 0.0719 0.0013 0.0720 0.0031 0.0678 0.0013 0.0641 0.0418 0.0549 0.0311 0.0706 0.0359 RECALL@20 0.1039 – 0.1038 – 0.0770 0.0018 0.1111 0.0013 0.1100 0.0036 0.1094 0.0018 0.1003 0.0476 0.0770 0.0480 0.1122 0.0445 RECALL@50 0.1587 – 0.1789 – 0.1326 0.0081 0.1882 0.0054 0.1882 0.0067 0.1860 0.0013 0.1575 0.0678 0.1107 0.0792 0.1754 0.0629 Phone NDCG@10 0.0006 – 0.0010 – 0.0005 0.0000 0.0037 0.0014 0.0028 0.0004 0.0045 0.0033 0.0035 0.0041 0.0030 0.0018 0.0064 0.0052 NDCG@20 0.0010 – 0.0013 – 0.0006 0.0000 0.0050 0.0020 0.0036 0.0006 0.0057 0.0042 0.0044 0.0052 0.0044 0.0032 0.0078 0.0063 NDCG@50 0.0016 – 0.0019 – 0.0009 0.0001 0.0070 0.0029 0.0049 0.0009 0.0078 0.0058 0.0055 0.0066 0.0057 0.0042 0.0098 0.0079 RECALL@10 0.0012 – 0.0020 – 0.0012 0.0001 0.0072 0.0026 0.0057 0.0009 0.0081 0.0055 0.0062 0.0075 0.0102 0.0068 0.0108 0.0083 RECALL@20 0.0028 – 0.0032 – 0.0018 0.0002 0.0124 0.0048 0.0088 0.0015 0.0129 0.0088 0.0099 0.0118 0.0204 0.0144 0.0165 0.0126 RECALL@50 0.0059 – 0.0062 – 0.0031 0.0003 0.0228 0.0099 0.0156 0.0032 0.0234 0.0171 0.0157 0.0188 0.0369 0.0245 0.0265 0.0207 5.0.2. Baselines We first compared our approach with traditional item ID-based methods. Item ID-based: (1) SASRec (Kang and McAuley, 2018) employs a unidirectional Transformer to capture sequential dependencies; (2) BERT4Rec (Sun et al., 2019b) utilizes a bidirectional Transformer trained with a cloze-style objective; We also compared our approach with generative recommendation methods based on semantic IDs. Semantic ID-based: (3) VQ-Rec (Hou et al., 2023) applies product quantization to tokenize items into semantic IDs, which are then pooled to obtain item representations; (4) TIGER (Rajput et al., 2023) utilizes RQ‑VAE to generate codebook identifiers, embedding semantic information into discrete code sequences; (5) LC-Rec (Zheng et al., 2024) exploits identifiers with auxiliary alignment tasks to associate the generated codes with natural language; Finally, we compare our approach against baselines designed to handle cold-start items: (6) Retraining retrains the model from scratch using the full training data augmented with the synthesized cold-start interactions. (7) Finetuning continues training a well-trained model using only the synthesized cold-start interactions. (8) SpecGR (Ding et al., [n. d.]) is a plug-and-play framework for inductive generative recommendation that drafts candidate items and uses the GR model to verify them. 5.0.3. Evaluation Metrics Following prior studies (Rajput et al., 2023; Zheng et al., 2024), we evaluate performance using two commonly adopted ranking metrics: top-k Recall and top-k Normalized Discounted Cumulative Gain (NDCG). We report results for k∈10,20,50k∈\10,20,50\. In addition, as illustrated in Section 3.1, we include IID Ratio as an evaluation metric to provide an intuitive measure of how well the model captures cold-start item patterns. For example, on the cold subset, a higher IID Ratio indicates that a larger fraction of the recommended items are cold-start items, and vice versa. Similarly, on the warm-start test set, a higher IID Ratio indicates that a larger fraction of the recommended items are warm-start items, and vice versa. 5.0.4. Implementation Details For the tokenization module, we adopt Sentence-T5 (Ni et al., 2021) to encode each item’s title and other textual metadata into embeddings. We use M=4M=4 codebooks, each containing K=256K=256 code vectors with dimensionality d=32d=32. The weighting coefficient λ in Eq. (13) is analyzed in our analysis experiments, and we set λ∈3,000, 3,000, 1,000λ∈\3,000,\,3,000,\,1,000\ for the three datasets, respectively, which yields the best performance. We remove users from the training set who do not appear in the test set, to eliminate the confounding effects of cold-start users in evaluation. GenRecEdit is built upon T5 333https://huggingface.co/docs/transformers/en/model_doc/t5. We set the number of decoder layers to 6 and use gated-silu as the FFN activation, which is crucial for successful editing because the default ReLU activation can induce strong truncation, leading to low-rank key representations k and, consequently, making the matrix inversion in Eq. (13) ill-conditioned. The multi-head RQ-VAE is trained for 10,000 epochs using the Adam optimizer (Loshchilov and Hutter, 2017) with a learning rate of 1e−31e^-3 and a batch size of 2,048. 5.1. Overall Performance Table 1 reports the results on both the overall test set and the cold subset across the three datasets. We draw the following three key conclusions regarding performance on the overall test set. (1) GenRecEdit generally achieves the best overall performance across multiple metrics (including NDCG and Recall) compared with both conventional item ID-based methods and generative semantic ID-based methods. In addition, GenRecEdit remains competitive with representative cold-start baselines, supporting the feasibility of applying model editing to generative recommendation. (2) Semantic ID-based generative recommenders consistently outperform traditional item ID-based methods. A key reason is that semantic IDs provide a content-informed discrete representation that can be assigned to previously unseen items (cold-start items), enabling the model to reason about cold-start items through their semantic tokens. In contrast, ID-based recommenders rely on item-specific parameters (e.g., ID embeddings) learned from historical interactions; when an item is unseen, these parameters are missing or poorly estimated, which substantially weakens their ability to recommend cold-start items and degrades overall performance. (3) GenRecEdit, together with Retrain, Finetune, and SpecGR, all those cold-start-specialized methods attains promising results. This suggests that the constructed interaction sequences targeting cold-start items provide effective supervisory signals, although different methods exploit them in different ways. Importantly, GenRecEdit is substantially more efficient than these alternatives, which we analyze in detail in subsequent diagnostic experiments. We further draw the following conclusions regarding performance on the cold subset. (1) Compared with both item ID-based and semantic ID-based models, GenRecEdit delivers a clear improvement on cold-start items, indicating that our method successfully injects cold-start item knowledge and further demonstrating the viability of model editing for generative recommendation. (2) On the cold-start split of the Software dataset, our method still achieves comparable performance, whereas continued training (i.e., Finetune) tends to overfit to the newly injected data (high performance on the cold subset) and consequently induces catastrophic forgetting of prior knowledge (low performance on the warm subset and the overall test set). In fact, the trade-off between the injected data and the prior knowledge can be controlled by a hyperparameter of GenRecEdit, which we discuss in the analysis experiments. Table 2. Recommendation performance on the warm subset. Using the Phone dataset as an example. “IID R.” denotes IID Ratio; “N.” denotes NDCG; “R.” denotes RECALL; and “Drop” denotes the relative drop of N.@10 (%). Method IID R.@10 N.@10 N.@20 R.@10 R.@20 Drop Semantic ID-based TIGER 0.7020 0.0108 0.0144 0.0215 0.0362 – Cold-start-based Retraining 0.6181 0.0080 0.0104 0.0165 0.0259 -25.9% Finetuning 0.1052 0.0014 0.0018 0.0022 0.0037 -87.0% SpecGR 0.3329 0.0067 0.0081 0.0207 0.0389 -38.0% GenRecEdit 0.6366 0.0101 0.0127 0.0186 0.0288 -6.5% Finally, we examine the performance of several cold-start-based methods on the warm subset. As shown in Table 2, under the TIGER backbone, GenRecEdit incurs the smallest loss on the warm subset in terms of NDCG@10, with only a 6.5%6.5\% drop. 5.2. Ablation Study We conducted ablation studies to verify the effect of each module in GenRecEdit. Specifically, we consider the following experimental settings: (1) Knowledge preparation. In Section 4.1, we identify a key challenge of applying model editing to GR data: the lack of stable token bundles, which hinders controlling multi-token outputs with a single edit. We therefore propose a position-wise knowledge preparation scheme, and compare it with a conventional object-wise strategy, denoted as w/o position-wise. (2) Layer Location. In Section 4.2, we locate critical layers via probing: for each position, we train a classifier at every layer that takes the key activation k of the subject sps_p (i.e., the key at the last token of the interaction history and the prefix SIDs) and predicts whether sps_p belong to a cold-start item. We validate this classifier-based selection against two baselines that edit a random layer or the layer with the lowest classifier accuracy, denoted as w/o classifier (random) and w/o classifier (worst). (3) Inference. As discussed in Section 4.3, we use the One-One Triggering Policy at inference to avoid interference among position-wise edits. We compare it with a conventional strategy (Meng et al., 2022a, b) that keeps all edited layers continuously active so all the edits affect every decoded token, denoted as w/o one-one triggering. As shown in Table 3, we draw two conclusions. (1) Each module and mechanism contributes positively to recommendation performance. (2) Among them, position-wise knowledge preparation and the One-One Triggering Policy are essential for performance, which aligns with the model editing challenges posed by GR data discussed in Section 1 and highlights the necessity and effectiveness of our design. 5.3. Experimental Analysis Table 3. The results of the ablation study. Using the Phone dataset as an example. “w/o” indicates that the corresponding module is removed. Model IID R.@10 N.@10 N.@20 R.@10 R.@20 GenRecEdit 0.7359 0.0064 0.0078 0.0108 0.0165 Knowledge Preparation w/o position-wise 0.0220 0.0002 0.0002 0.0004 0.0005 Layer Location w/o classifier (random) 0.7339 0.0059 0.0072 0.0100 0.0150 w/o classifier (worst) 0.6865 0.0055 0.0066 0.0094 0.0137 Inference w/o one-one triggering 0.0030 0.0000 0.0001 0.0001 0.0002 We further conduct experimental analyses to assess the effectiveness of GenRecEdit. Unless otherwise specified, all experiments are conducted on the Video dataset. 5.3.1. Sensitive Analysis of Knowledge Preparation New knowledge (comprising cold-start items and their synthesized interaction histories) is a central component of GenRecEdit. Both the quality and quantity of this new knowledge are crucial for effective model editing. Accordingly, we conduct two experiments. (1) Quality of new knowledge. We construct two control settings: (a) no injection of new knowledge, denoted as Origin; and (b) injecting the ground-truth interaction histories of cold-start items as new knowledge during editing, denoted as Upper Bound. As shown in Figure 5, we obtain the following observations. (a) On the cold subset, GenRecEdit achieves a substantial improvement over Origin, demonstrating the effectiveness of injecting new knowledge. (b) Compared with Upper Bound, GenRecEdit still exhibits a considerable performance margin, which we refer to as the Quality Gap. This gap indicates significant headroom for better approximating the real interaction histories of cold-start items, leaving ample room for future research. We draw similar conclusions on the overall test set. Figure 5. An analysis of the quality of constructed knowledge. The left panel shows performance under three settings on the cold subset, while the right panel reports the corresponding results on the overall test set. (2) Quantity of new knowledge. We evaluate model performance under varying amounts of constructed knowledge injected per cold-start item. As shown in Figure 6, we draw three key findings. (a) As the injected knowledge per cold-start item increases, GenRecEdit achieves a higher IIDRatio@10IID~Ratio@10 on the cold subset, indicating improved capture of the SID patterns of cold-start items. (b) With more injected knowledge, recommendation accuracy for cold-start items (measured by NDCG) improves steadily. (c) On the full test set, overall recommendation accuracy first increases and then declines as the injected knowledge grows, suggesting an inherent trade-off between optimizing performance on warm items and on cold-start items. Figure 6. A sensitivity analysis on the number of constructed knowledge. The left figure shows changes of three metrics on the cold subset (NDCG@10, NDCG@20, and IID Ratio@10), while the right figure reports metric variations on the overall test set (NDCG@10 and NDCG@20). Figure 7. An analysis on the hyper parameter λ. The left figure shows the performance trends on cold-start and warm items (NDCG@10), while the right reports the performance on the overall test set (NDCG@10 and NDCG@20). 5.3.2. Sensitive Analysis of Hyper-parameters λ in Eq. (13) is an important hyperparameter that governs the trade-off between preserving original knowledge (warm items) and injecting new knowledge (cold-start items). Therefore, an appropriate choice of λ is essential for maximizing the utility of GenRecEdit on the overall test set. In particular, a larger λ enforces stronger preservation of the original knowledge, but correspondingly weakens the injection strength for cold-start items. We sweep λ over the range [0,5000][0,5000] by evaluating four representative values. The results are reported in Figure 7, from which we draw two key observations. (1) As shown in the left panel, increasing λ consistently improves performance on prior knowledge (warm items) while degrading performance on new knowledge (cold-start items), empirically validating the role of λ in Eq. (13). (2) The right part shows that a moderate increase in λ can improve performance on the overall test set by better balancing the trade-off between warm and cold-start items. Table 4. Comparison of update time across four algorithms. Model Update Retraining Fintuning SpecGR GenRecEdit Type Train Train Alignment Edit Relative Time 100% 18.1% 41.6% 9.5% 5.3.3. Updating Time Analysis A key motivation for using model editing to mitigate cold-start collapse is its low per-item update latency. To highlight this advantage, we conduct an experiment that explicitly compares the total model update time of GenRecEdit, retraining, and finetuning. Specifically, we measure the wall-clock time from completing new-knowledge construction to obtaining the updated model. Although SpecGR does not train on the constructed knowledge and instead relies on an additional alignment task, we still include its training time as the model update time for completeness. For ease of comparison, we normalize all results by the retraining time and report the relative time costs of the other methods. As shown in Table 4, GenRecEdit is highly efficient in terms of model update time: as a training-free approach, it incurs only 9.5%9.5\% of the retraining cost. Figure 8. Classfier Accuracy across layers. The left figure shows the result on Video, while the right figure reports the result on Phone. 5.3.4. Classifier Accuracy across Layers A key component of our method is locating an edit layer for each position. As described in Section 4.2, we select the edit laye based on the classification accuracy of a probing classifier. To assess layer-wise sensitivity to new knowledge (cold-start items) and original knowledge (warm items), we report the probing accuracy at each layer. As shown in Figure 8, we make two key observations. (1) Across positions, the classifier exhibits a broadly consistent layer preference: intermediate-to-early layers tend to yield higher probing accuracy. This suggests that the edit, formulated as an equivalent linear transformation, is more effective at separating new versus prior knowledge in these layers, consistent with findings in NLP (Meng et al., 2022b). (2) For position 3, the probing accuracy drops sharply in the later layers. We attribute this to an intrinsic property of our data: at position 3, both new and prior knowledge contain many identical non-semantic tokens, which reduces discriminability in deeper layers. 6. Conclusion Generative recommendation is a promising paradigm for sequential recommendation, yet we identify a critical bottleneck: cold-start collapse, where accuracy on newly introduced items can drop to near zero. Our analysis suggests that GR models often generate the first semantic-ID token correctly but tend to complete sequences with seen semantic-ID patterns, making multi-token generation for cold-start items unreliable. To enable timely updates without costly retraining, we propose GenRecEdit, a model-editing framework tailored to GR that treats cold-start semantic-ID patterns as editable knowledge. GenRecEdit performs position-wise, context-to-next-token edits, iteratively injects token-level knowledge to handle the lack of stable token bundles, and adopts a One-One triggering mechanism to avoid cross-position interference during decoding. Experiments on three Amazon categories show that GenRecEdit effectively mitigates cold-start collapse, substantially improving cold-start recommendation while preserving warm-item performance with low overhead. References (1) Belinkov (2022) Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics 48, 1 (2022), 207–219. Dai et al. (2022) Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022. Knowledge neurons in pretrained transformers. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 8493–8502. Deng et al. (2025) Jiaxin Deng, Shiyao Wang, Kuo Cai, Lejian Ren, Qigen Hu, Weifeng Ding, Qiang Luo, and Guorui Zhou. 2025. Onerec: Unifying retrieve and rank with generative recommender and iterative preference alignment. arXiv preprint arXiv:2502.18965 (2025). Ding et al. ([n. d.]) Yijie Ding, Yupeng Hou, Jiacheng Li, and Julian McAuley. [n. d.]. Inductive Generative Recommendation via Retrieval-based Speculation, October 2024. arXiv preprint arXiv:2410.02939 ([n. d.]). Fang et al. (2024) Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Shi Jie, Xiang Wang, Xiangnan He, and Tat-Seng Chua. 2024. Alphaedit: Null-space constrained knowledge editing for language models. arXiv preprint arXiv:2410.02355 (2024). Geva et al. (2021) Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 5484–5495. Hou et al. (2023) Yupeng Hou, Zhankui He, Julian McAuley, and Wayne Xin Zhao. 2023. Learning vector-quantized item representation for transferable sequential recommenders. In Proceedings of the ACM Web Conference 2023. 1162–1171. Hou et al. (2024) Yupeng Hou, Jiacheng Li, Zhankui He, An Yan, Xiusi Chen, and Julian McAuley. 2024. Bridging language and items for retrieval and recommendation. arXiv preprint arXiv:2403.03952 (2024). Hou et al. (2025a) Yupeng Hou, Jiacheng Li, Ashley Shin, Jinsung Jeon, Abhishek Santhanam, Wei Shao, Kaveh Hassani, Ning Yao, and Julian McAuley. 2025a. Generating long semantic IDs in parallel for recommendation. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2. 956–966. Hou et al. (2025b) Yupeng Hou, An Zhang, Leheng Sheng, Jiancan Wu, Xiang Wang, Tat-Seng Chua, and Julian McAuley. 2025b. Towards Large Generative Recommendation: A Tokenization Perspective. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management. 6821–6824. Hua et al. (2023) Wenyue Hua, Shuyuan Xu, Yingqiang Ge, and Yongfeng Zhang. 2023. How to index item ids for recommendation foundation models. In Proceedings of the Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region. 195–204. Jiang et al. (2025) Houcheng Jiang, Junfeng Fang, Ningyu Zhang, Guojun Ma, Mingyang Wan, Xiang Wang, Xiangnan He, and Tat-Seng Chua. 2025. AnyEdit: Edit Any Knowledge Encoded in Language Models. CoRR (2025). Kang and McAuley (2018) Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In 2018 IEEE International Conference on Data Mining (ICDM). IEEE, 197–206. Li et al. (2023) Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2023. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems 36 (2023). Loshchilov and Hutter (2017) Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017). Meng et al. (2022a) Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual associations in gpt. Advances in neural information processing systems 35 (2022), 17359–17372. Meng et al. (2022b) Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass-editing memory in a transformer. arXiv preprint arXiv:2210.07229 (2022). Ni et al. (2021) Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:2108.08877 (2021). Rajput et al. (2023) Shashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan, Trung Vu, Lukasz Heldt, Lichan Hong, Yi Tay, Vinh Tran, Jonah Samost, et al. 2023. Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2023), 10299–10315. Si et al. (2024) Zihua Si, Zhongxiang Sun, Jiale Chen, Guozhang Chen, Xiaoxue Zang, Kai Zheng, Yang Song, Xiao Zhang, Jun Xu, and Kun Gai. 2024. Generative retrieval with semantic tree-structured identifiers and contrastive learning. In Proceedings of the 2024 Annual International ACM SIGIR Conference on Research and Development in Information Retrieval in the Asia Pacific Region. 154–163. Sun et al. (2019a) Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019a. BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM international conference on information and knowledge management. 1441–1450. Sun et al. (2019b) Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019b. BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (Beijing, China) (CIKM ’19). ACM, New York, NY, USA, 1441–1450. Wang et al. (2024a) Wenjie Wang, Honghui Bao, Xinyu Lin, Jizhi Zhang, Yongqi Li, Fuli Feng, See-Kiong Ng, and Tat-Seng Chua. 2024a. Learnable item tokenization for generative recommendation. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. 2400–2409. Wang et al. (2024b) Ye Wang, Jiahao Xun, Minjie Hong, Jieming Zhu, Tao Jin, Wang Lin, Haoyuan Li, Linjun Li, Yan Xia, Zhou Zhao, et al. 2024b. Eager: Two-stream generative recommender with behavior-semantic collaboration. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3245–3254. Yang et al. (2024) Liu Yang, Fabian Paischer, Kaveh Hassani, Jiacheng Li, Shuai Shao, Zhang Gabriel Li, Yun He, Xue Feng, Nima Noorshams, Sem Park, et al. 2024. Unifying generative and dense retrieval for sequential recommendation. arXiv preprint arXiv:2411.18814 (2024). Zhai et al. (2024) Jiaqi Zhai, Lucy Liao, Xing Liu, Yueming Wang, Rui Li, Xuan Cao, Leon Gao, Zhaojie Gong, Fangda Gu, Jiayuan He, et al. 2024. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. In International Conference on Machine Learning. PMLR, 58484–58509. Zhao et al. (2023) Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223 1, 2 (2023). Zheng et al. (2024) Bowen Zheng, Yupeng Hou, Hongyu Lu, Yu Chen, Wayne Xin Zhao, Ming Chen, and Ji-Rong Wen. 2024. Adapting large language models by integrating collaborative semantics for recommendation. In 2024 IEEE 40th International Conference on Data Engineering (ICDE). IEEE, 1435–1448. Zhou et al. (2025) Guorui Zhou, Jiaxin Deng, Jinghao Zhang, Kuo Cai, Lejian Ren, Qiang Luo, Qianqian Wang, Qigen Hu, Rui Huang, Shiyao Wang, et al. 2025. OneRec Technical Report. arXiv preprint arXiv:2506.13695 (2025). Zhu et al. (2024) Jieming Zhu, Mengqun Jin, Qijiong Liu, Zexuan Qiu, Zhenhua Dong, and Xiu Li. 2024. Cost: Contrastive quantization based semantic tokenization for generative recommendation. In Proceedings of the 18th ACM Conference on Recommender Systems. 969–974.