Paper deep dive
Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning
Masane Fuchi, Tomohiro Takagi
Models: Stable Diffusion (CLIP text encoder)
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 7:29:42 PM
Summary
The paper introduces a novel, efficient concept-erasure method for text-to-image diffusion models by updating the text encoder parameters using few-shot unlearning. This approach achieves concept removal significantly faster (within 10 seconds) than existing methods that modify the U-Net, while preserving image fidelity and implicitly mapping erased concepts to related latent concepts.
Entities (5)
Relation Signals (3)
Few-shot Unlearning â updates â Text Encoder
confidence 95% ¡ We propose a novel concept-erasure method that updates the text encoder using few-shot unlearning
Text Encoder â partof â Text-to-Image Diffusion Models
confidence 90% ¡ We propose a method for erasing specific concepts from the text-to-image diffusion models... we focus on the text encoder
U-Net â partof â Text-to-Image Diffusion Models
confidence 90% ¡ Previous studies have successfully removed specific concepts... by updating the weights of U-Net
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Generating images from text has become easier because of the scaling of diffusion models and advancements in the field of vision and language. These models are trained using vast amounts of data from the Internet. Hence, they often contain undesirable content such as copyrighted material. As it is challenging to remove such data and retrain the models, methods for erasing specific concepts from pre-trained models have been investigated. We propose a novel concept-erasure method that updates the text encoder using few-shot unlearning in which a few real images are used. The discussion regarding the generated images after erasing a concept has been lacking. While there are methods for specifying the transition destination for concepts, the validity of the specified concepts is unclear. Our method implicitly achieves this by transitioning to the latent concepts inherent in the model or the images. Our method can erase a concept within 10 s, making concept erasure more accessible than ever before. Implicitly transitioning to related concepts leads to more natural concept erasure. We applied the proposed method to various concepts and confirmed that concept erasure can be achieved tens to hundreds of times faster than with current methods. By varying the parameters to be updated, we obtained results suggesting that, like previous research, knowledge is primarily accumulated in the feed-forward networks of the text encoder. Our code is available at \url{this https URL}
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
77,900 characters extracted from source content.
Expand or collapse full text
Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Masane Fuchi 1 Tomohiro Takagi 1 Figure 1: Overview of our results. Given simple text, our method is able to erase concept within 10 s per concept with few images. Unlike current methods, our method involves updating the text encoder. Generated images after erasing with our method are mapped to similar concept without providing anchor concepts. For example, âSnoopyâ is mapped to dog which is its motif and âGrumpy Catâ is mapped to cat which is the super-category. Abstract Generating images from text has become eas- ier because of the scaling of diffusion models and advancements in the field of vision and lan- guage.These models are trained using vast amounts of data from the Internet. Hence, they often contain undesirable content such as copy- righted material. As it is challenging to remove such data and retrain the models, methods for erasing specific concepts from pre-trained mod- els have been investigated. We propose a novel concept-erasure method that updates the text en- coder using few-shot unlearning in which a few real images are used. The discussion regard- ing the generated images after erasing a con- 1 Department of Computer Science, Meiji University, Japan. Correspondence to: Masane Fuchi<ce235031@meiji.ac.jp>. Preprint. cept has been lacking. While there are meth- ods for specifying the transition destination for concepts, the validity of the specified concepts is unclear. Our method implicitly achieves this by transitioning to the latent concepts inherent in the model or the images. Our method can erase a concept within 10 s, making concept erasure more accessible than ever before. Implicitly tran- sitioning to related concepts leads to more nat- ural concept erasure. We applied the proposed method to various concepts and confirmed that concept erasure can be achieved tens to hundreds of times faster than with current methods. By varying the parameters to be updated, we ob- tained results suggesting that, like previous re- search, knowledge is primarily accumulated in the feed-forward networks of the text encoder. Our code is available athttps://github. com/fmp453/few-shot-erasing 1 arXiv:2405.07288v2 [cs.CV] 29 Aug 2024 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning 1. Introduction Diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020) have surpassed the previous state-of-the-art genera- tive adversarial networks (GANs) (Goodfellow et al., 2014; Brock et al., 2019; Karras et al., 2020) in performance, be- cause of their stable learning and broad representation ca- pabilities (Dhariwal & Nichol, 2021). With classifier-free guidance (Ho & Salimans, 2021), it has also become possi- ble to generate high-quality images on the basis of natural language instructions (Saharia et al., 2022; Ramesh et al., 2022; Balaji et al., 2023; Xue et al., 2023; Podell et al., 2024; Dai et al., 2023). In large-scale image-generative models, since various data are collected to improve generation quality, it is possible to generate undesirable images such as copyright contents. While filtering training data can help mitigate the genera- tion of such undesirable images (OpenAI, 2023a), it is gen- erally costly and challenging. This approach is also inef- fective against pre-trained models. Previous studies have successfully removed specific concepts from text-to-image generative models by updating the weights of U-Net (Ron- neberger et al., 2015), an image generative module, or its conditioned cross attention (Gandikota et al., 2023; Kumari et al., 2023; Gandikota et al., 2024; Zhang et al., 2024a; Zhao et al., 2024). However, updating parameters of U-Net can lead to a decrease in generation quality in unconditional cases. We propose a method for erasing specific concepts from the text-to-image diffusion models without altering the pa- rameters of the U-Net. Specifically, we focus on the text encoder and aim to achieve this by altering the quality of text conditioning. We use several images of the target con- cept to make slight changes to the parameters of the text encoder to remove the concept. Our method is inspired by textual inversion (Gal et al., 2023), but since we only make minor parameter adjustments, it operates very quickly. Ta- ble 1 compares our proposed method with current methods. It is evident that our proposed method can erase concepts more quickly compared to existing methods. Additionally, our proposed method naturally maps to surrounding con- cepts, erasing the need for concept induction by an anchor concept 1 . Our contributions are as follows: ⢠We achieve a speedup of 60âź900 times compared with current methods of updating the traditional U- Net, enabling concept erasure within 10 s. ⢠Concept erasure is achieved by providing several im- ages related to the concept to be erased. 1 We describe anchor concepts in detail in Appendix A.2. ⢠While current methods often lack discussion on the generated examples after concept erasure, our pro- posed method ensures semantic similarity in the re- sulting concepts. We highlighted this issue in the quantitative evaluation of previous research. 1.1. Intuitive Motivation We provide an explanation of the intuitive motivation be- hind our study. On the basis of this motivation, we intro- duce our proposed method supported by various justifica- tions. Text-to-image diffusion models are implemented using models trained on a large amount of data from the web. They are highly versatile foundation models capable of executing various downstream tasks (Bommasani et al., 2022). Therefore, adapting them to individual domains re- quires additional fine-tuning. The smallest unit of the do- main adaptation is personalization, and numerous methods have been proposed to achieve this (Gal et al., 2023; Ruiz et al., 2023; Wei et al., 2023; Chen et al., 2023; Pang et al., 2024; Fei et al., 2023; Zhang et al., 2024b). Inspired by these studies, we have one fundamental question. If it is possible for these methods to give specific knowledge, could a similar technique be used to make them forget? We address this question by taking textual inversion (Gal et al., 2023) as a reference and partially changing the method. In the subsequent sections, we explore our pro- posed method and the answer to this question focusing on the modifications made. 2. Our Method Our goal is to prevent the generation of specific concepts by updating the text encoder. The parameters of the U-Net responsible for image generation remain unchanged, ensur- ing that the modelâs generation capability (image fidelity) is preserved. We first discuss why we believe the text en- coder should be updated (Section 2.1). We then outline our proposed method (Section 2.2). 2.1. Why Text Encoder? In this subsection, we explain why our method targets the updating of the text encoder. In the field of text-to-image, it has been demonstrated that the quality of the text encoder, represented by its size, cor- relates with the quality of text-image alignment (Saharia et al., 2022). When comparing the use of CLIP (Radford et al., 2021) and T5 (Raffel et al., 2020) as the text encoder, 2 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Table 1: Comparison of proposed method with current methods when erasing âVan Gogh styleâ.Ămeans that anchor concept is necessary. MethodTarget parametersRuntimeAnchor Concepts#U-Net ESD-x (Gandikota et al., 2023)Cross-Attention1 hour â 2 UCE (Gandikota et al., 2024)Attention Weight10 minĂ1 SPM (Lyu et al., 2024)Adapters in U-Net2.5 hours â 1 OursText Encoder7âź8 sec â 1 it has been indicated that the T5, which is trained solely on text data outperforms human evaluation compared with us- ing CLIP. While quantitative evaluations show similar re- sults between CLIP and T5, it is known that quantitative metrics (such as Fr Ě echet inception distance (FID) (Heusel et al., 2017) and CLIP Score (Hessel et al., 2021)) do not always align with human evaluations in assessing the qual- ity of generated images (Otani et al., 2023). Therefore, achieving superior results in human evaluations support the hypothesis that the performance to some extent depends on the quality of the text encoder. DALLE-3 (OpenAI, 2023a), which is trained by using high-quality and detailed image captions generated by GPT-4 (OpenAI, 2023b), achieves extremely high-quality image generation. From these find- ings, for models possessing sufficient image-generation ca- pabilities (image fidelity, that is, lower FID), we believe that given information by the text affects the text-image alignment. This is also demonstrated in TextCrafter (Li et al., 2024). CLIPâs final output is a multi-dimensional vector. Although somewhat apparent from the concept of latent variables, visualizing this vector using techniques such as t-SNE (van der Maaten & Hinton, 2008) allows for meaningful clustering (Weiss et al., 2022). Therefore, we believe that slight variations in the output of CLIP can be used to move towards similar concepts, and achieving this could be possible by slightly adjusting the parameters. In many cases, including community models, the text en- coder is often fixed during training. Considering the lim- ited variety of text encoders, erasing specific concepts from one text encoder might be more beneficial than the strategy used with U-Net (Lyu et al., 2024) when transferring it for use in other models. On the basis of the above, we conclude that updating the text encoder is more appropriate than updating the U-Net. Reframing concept erasure as âsignificantly de- creasing text-image alignment for specific conceptsâ fur- ther highlights the natural focus on the text encoder. 2.2. Erasing Method In this subsection, we consider preventing the generation of specific concepts by updating the text encoder. It is cru- cial to preserve as many concepts as possible other than the target concept. Therefore, as mentioned in Section 2.1, we consider making slight variations to the parameters of the text encoder. c θ âc θ + âc,(1) wherec θ denotes the text encoder. The issue lies in how to computeâc. Our method is very simple, merely setting the loss in the reverse direction according to Jang et al. (2023). Figure 2 illustrates an overview of our proposed method. While experiments in a similar setup have been conducted by Kumari et al. (2023), they updated the parameters of the U-Net, which differs from our approach. This loss setting is very similar to textual inversion (Gal et al., 2023). We use stable diffusion (Rombach et al., 2022) due to the restric- tion of our computational resources. The stable diffusion loss is given by L SD =E x,ÎľâN(0,1),t,y âĽÎľâÎľ θ (x,t,c θ (y))⼠2 2 (2) wheretis the timestep,xis the denoised image to timet, Îľis the unscaled noise sample, andÎľ θ is the denoising net- work. Our method involves training to maximize this loss. However, to prevent drastic changes to the text encoder, we terminate training within an appropriate range (managed with the number of epochs) rather than maximizing it com- pletely. âE x,ÎľâN(0,1),t,y âĽÎľâÎľ θ (x,t,c θ (y â ))⼠2 2 (3) We update only the text encoder, so the parameters ofÎľ θ are fixed. The notationy â denotes the caption that includes the concept to erase. We use CLIP ImageNet Template (Rad- ford et al., 2021) for this caption as well as textual inver- sion. By making slight variations to the text encoder, we assume that the impact on other concepts is minimal and is further mapped to approximate concepts. Since only minor changes are made, training is completed in a short period. In this formulation, we investigated two approaches: one using a few-shot method (Wang et al., 2020) with four pre- prepared images forxand the other using completely ran- dom noise (zero-shot). 3 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Figure 2: Overview of our proposed method.θ â denotes that parameters are fixed. Intuitively, considering text-image alignment, the few-shot approach is expected to perform better. We aimed to con- firm this through experiment. Considering practical as- pects, it is straightforward to prepare actual images of the concepts to be erased, making it possible to confirm the ef- fectiveness of the method even with the few-shot approach. 2.3. Update Parameters There are studies suggesting that the knowledge gained during learning is stored in feed-forward networks (i.e. MLPs) (Meng et al., 2022; Dai et al., 2022; Geva et al., 2021). There is also a study that shows that the final self- attention layer is also effective in knowledge editing (Meng et al., 2022). Following these studies, we update all MLPs and the final self-attention layer. The list of updated param- eters is shown in Table 4 of Appendix B. 3. Experiments We confirm the effectiveness of our proposed method throughout the experiments. Before presenting the results, we discuss the baselines (Section 3.1) and experimental set- tings (Section 3.2). 3.1. Baselines We used ESD (Gandikota et al., 2023) (we specifically used ESD-x-1, which updates the parameters related to cross- attention), Unified Concept Editing (UCE) (Gandikota et al., 2024), and Semi-Permeable Membrane (SPM) (Lyu et al., 2024) which are open sourced and high effectiveness as the baselines. Due to the privacy issue with of LAION- 5B (Schuhmann et al., 2022) 2 , we cannot conduct with Ab- 2 https://laion.ai/notes/ laion-maintanence/ lating Concept (Kumari et al., 2023). When elements other than the pre-trained model depend on external materials, such as in this case, we may not be able to reproduce the results completely. The above baselines exclusively use the text-to-image diffusion models, including VAE, U-Net, text encoder, and tokenizer. For further details, please refer to Appendix C.2. 3.2. Experimental Setup Training.We apply our proposed method to Stable Dif- fusion 1.5 3 (as referred to original SD), the text encoder of which is OpenAI CLIP vit-large-patch14 4 . We use four im- ages in the few-shot setting. The text encoder is optimized using Adam (Kingma & Ba, 2017). The hyperparameters we used in our experiments are listed in Table 2. Table 2: Hyperparameters of our method HyperparameterValue Batch Size2 Training Epochs 4 (proper noun) 5 (common noun) Learning rate1Ă10 â5 Adam(β 1 ,β 2 )(0.9, 0.98) Weight decay1Ă10 â8 Generating.We used PNDM Scheduler (Liu et al., 2022) with 7.5 guidance scale and 100 inference steps. 4 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Figure 3: Comparison of images generated with each method 3.3. Qualitative Results Erasing Single ConceptWe conducted experiments fo- cusing on proper nouns. We also conducted experiments with common nouns to ensure general performance. We fo- cus on the results using âEiffel Towerâ as the proper noun and âbananaâ as the common noun. More results are pre- sented in Appendix D. In the first row of Figure 3, the results of erasing âEiffel Towerâ are shown. With ESD and SPM, while the concept of the Eiffel Tower disappears, the generated images are completely unrelated. This phenomenon was also observed by Lu et al. (2024), and it is necessary to consider whether the generated images of the erased concepts are appropri- ate. With UCE, as the anchor concept was âParisâ, the results depict the cityscape of Paris. In contrast, our pro- posed method not only prevented the generation of the Eif- fel Tower but also retained only the elements of the tower, indicating that the embedding is mapped to a similar con- cept. In the second row of Figure 3, the results of erasing âba- nanaâ are shown. Overall, similar results when erasing âEiffel Towerâ can be observed. SPM produced results that are difficult to interpret as successfully erasing the banana. As common nouns are ubiquitous words, they are likely heavily represented in the training data of original stable 3 https://huggingface.co/runwayml/ stable-diffusion-v1-5 4 https://huggingface.co/openai/ clip-vit-large-patch14 diffusion, making it challenging to erase such concepts. With UCE, as the anchor concept is âFruitâ, various types of fruits were generated. Despite not explicitly specifying such categories, our proposed method generated something resembling a category of fruits to which bananas belong. In the third row of Figure 3, the results of erasing âMonet styleâ are shown. ESD generated photorealistic images, while these images are not satisfied with the given prompt, âA paintingâ. Other methods reflect the given prompt suf- ficiently. UCE is similar to original SD because the anchor concept is âimpressionismâ that includes âMonet styleâ. It is suspicious the âMonet styleâ is erased. Our proposed method generated âA paintingâ while it is not âMonet styleâ. Effect on Other ConceptsWe investigated the impact of removing one concept on the generation of other con- cepts. We used a text encoder with âEiffel Towerâ erased to generate other concepts. The generated concepts were âcatâ, which is an unrelated concept, another landmark in Paris âArc de Triompheâ (the caption we used âTriumphal archâ), and a similar-shaped landmark âTokyo Towerâ. The results are shown in Figure 4. It can be seen that other con- cepts could also be generated. The generated images of the Tokyo Tower differ in actual shape 5 . However, because this is already observed in the images from the Original SD, we did not consider it is a significant drawback in the qualita- 5 The actual shape of Tokyo Tower can be checked athttps: //w.japan.travel/en/spot/1709/ 5 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning tive evaluations. In previous studies (Gandikota et al., 2023; Kumari et al., 2023; Gandikota et al., 2024; Zhang et al., 2024a; Zhao et al., 2024), the generated images from Original SD were treated as ground truth. However, we consider this ap- proach lacks justification. As observed in the example of âTokyo Towerâ in Figure 4, the generated images from the Original SD are not perfect. Like Fan et al. (2024), we believe that the generated images using a retrain model 6 should be the ground truth. However, a retrain model involves significant computational costs, surpassing the scope of the type of experimentation we can conduct. As the training data of CLIP are not publicly available, repro- ducing experiments is also impossible. Therefore, we pri- marily conducted a qualitative evaluation. CatTokyo TowerArc de Triomphe Figure 4: Comparison of the generated images between before (upper) and after (bottom) erasing âEiffel Towerâ. Caption we used when generating was âa photo of (a) conceptâ. Erasing Multiple ConceptsWe considered the erasure of multiple concepts. We erase âSnoopyâ, âR2D2â, and âMickey Mouseâ in that order. The results are shown in Figure 5. Even after removing multiple concepts, the erased concepts did not reappear. While the change in gen- erated images could be considered a future challenge, as previously mentioned, a retrain model served as the ground truth in our research, mitigating this issue. It can be ob- served that for the concepts that were not erased, the con- tents of the prompts were reflected in the generated images. The generated images of R2D2 from Original SD also did not match the prompt, which serves as another example of why Original SD is not considered the ground truth. 6 Details of a retrain model are described in Appendix A.3. Figure 5: Samples of âgraffiti of theConceptâ when erasing multiple concepts 3.4. Quantitative Results We used CLIP Score (Hessel et al., 2021) as the evalua- tion metric. While FID is suitable for evaluation on lim- ited datasets, there are concerns about its inability to reflect diversity in other cases (Jayasumana et al., 2024), so we choose not to use this standard metric. CLIP Score faces similar challenges in evaluating concepts not learned by CLIP (Otani et al., 2023). We address this issue by not using such captions during evaluation. Since our method does not update any parameters of the U-Net, image fidelity should not degrade. However, as the parameters of the text encoder are updated, text-image alignment may decrease. It should be noted that while a higher CLIP Score is de- sirable for concepts that have not been erased, the same may not hold true for erased concepts. Considering the ex- ample of the âEiffel Towerâ in Figure 3, for instance, it is anticipated that ESD-x-1 would have a lower CLIP Score compared with our proposed method. This is because when there is no relevance between the text and generated image, the CLIP Score is naturally lower. We conducted a quanti- tative comparison following prior studies (Gandikota et al., 2023; Kumari et al., 2023; Gandikota et al., 2024; Zhang et al., 2024a; Basu et al., 2024; Lyu et al., 2024). However, the metrics used in those studies do not evaluate âhow well concepts have been erased?â. To the best of our knowledge, such evaluation metrics do not currently exist. We created the prompts using CLIP ImageNet Template small and calculated the CLIP Score. Ten images were gen- erated for each prompt. Due to resource constraints, gen- erating ten images simultaneously is not feasible, so five images were generated at a time with seed values of 0 and 2024. Generation was executed using DDIM (Song et al., 2021) Scheduler with 7.5 guidance scale and 50 inference steps. 6 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Table 3: Comparison of the CLIP Scores ErasingEiffel Tower Eiffel TowerTokyo TowerTriumphal arch Original SD0.2560.2830.286 ESD-x-10.1110.1770.274 UCE0.2130.2410.284 SPM0.1900.2770.283 Ours0.2120.2620.281 ErasingMonet Style Monet StyleGogh StylePicasso Style Original SD0.2760.2640.267 ESD-x-10.2760.2640.267 UCE0.2730.2640.267 SPM0.2350.2630.267 Ours0.1790.2070.241 Table 3 lists the results. Following previous studies, it can be observed that our proposed method is capable of erasing styles more effectively. Regarding objects, while success- fully removing the target concept, our method maintained CLIP Scores for other concepts as much as possible. How- ever, the results of the quantitative evaluation in Table 3 do not align with the results presented in Figure 3. Similar to the earlier discussion, upon examining the generated Eiffel Tower images in Figure 3, it is evident that the Eiffel Tower was erased with all methods. According to previous stud- ies, ESD-x-1, which had the lowest CLIP Score, would be considered the best-performing method. However, it is dif- ficult to claim that ESD-x-1 exhibits the best performance through qualitative comparison. Therefore, we believe that using CLIP Score for evaluating concept erasure is not be appropriate. The same argument applies when using eval- uation metrics such as FID or Learned perceptual Image Patch Similarity (LPIPS) (Zhang et al., 2018). LPIPS is a metric that represents the distance from the Original SD. However, in the context of machine unlearning, a retrain model is the gold standard as the ground truth (Thudi et al., 2022; Jia et al., 2023; Fan et al., 2024). If FID or LPIPS were to be used, a retrain model would need to be prepared, which is impractical as an evaluation metric. 4. Ablation Studies We decomposed our proposed method into several ele- ments and conducted additional experiments to confirm the effectiveness of the elements. 4.1.k-shot Erasing We consider reducing the number of images provided dur- ing training. We conducted these experiments with zero- shot and two-shot. Figure 6 shows the results with zero- shot, and Figure 7 shows the results with two-shot. In both cases, the âEiffel Towerâ is still present. However, with zero-shot, it is clearly recognizable, while with two- shot, some images make it difficult to distinguish the Eiffel Tower. This indicates the need for diverse images. Figure 6: Generated images in case with zero-shot erasing Figure 7: Generated images in case with two-shot erasing As evident from Figure 8, with four-shot, the concept is mostly disappeared after the second epoch (four iterations). From this observation, we consider that the number of iter- ations is also sufficient for zero and two-shot setting. 4.2. Number of Epochs We confirm the transition over epochs using âSnoopyâ, âEiffel Towerâ, and âbananaâ. Five epochs were trained for all concepts in relation to Section 4.4. Figure 8 shows the results. The concepts âSnoopyâ and âEiffel Towerâ, which are proper nouns, disappeared early on. Therefore, there is no possibility of concept erasure failure due to the low number of iterations in the two-shot setting mentioned in the previous subsection. For âbananaâ, however, the con- cept disappeared at the end of the fifth epoch. This indi- cates that the difficulty of the task differs between common and proper nouns. It can be inferred that more common nouns than proper nouns are included in the training data when training with CLIP and stable diffusion. This sug- gests that it is more difficult to erase knowledge that is more deeply rooted in the model. 4.3. Real or Synthesized Images All experiments were conducted using real images. This experimental setting is based on our belief that a retrain model serves as the ground truth. In this case, it is most appropriate to use a subset of the datasetD f containing the concepts to erase. However, experiments can be conducted using synthesized images. Therefore, we conducted ex- periments using four synthesized images (referred as Fig- ure 3). The results are shown in Figure 9. Eiffel Tower was completely absent in some images, while in others, it 7 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Figure 8: Transition due to change in number of epochs. Top row is the result at end of first epoch and bottom row is result at end of fifth epoch. Captions we used in generating are âa banana on the tableâ, âa photo of Eiffel Towerâ and âSnoopy in cyberpunk styleâ respectively. is unclear what is being generated. Additionally, while in real images, the concept of âEiffel Towerâ was mapped to âTowerâ, in the case of generated images, âEiffel Towerâ was mapped to âBuildingâ. Since the appropriateness of each approach depends largely on subjective interpretation, we cannot make a definitive assertion. However, on the ba- sis of the idea of slight change in text encoders, we believe that âEiffel Towerâ should be mapped to the âTowerâ. Figure 9: Results when using synthesized images 4.4. Update Parameters There are studies suggesting that knowledge acquired dur- ing training is stored in MLPs (Meng et al., 2022; Dai et al., 2022; Geva et al., 2021) and in the first self-attention mod- ules (Basu et al., 2024) of transformer-based models, but it is unclear which is more effective in text-to-image tasks. We conducted experiments by limiting the parameters up- dated to either MLPs or self-attention modules to determine which part of CLIP contributes to knowledge updating. We presented the results from full parameter tuning. In the first self-attention layer, there are four trainable weight matri- ces:W q ,W k ,W v , andW out . We conducted the experi- ments under two settings: (i) updating onlyW out follow- ing the approach of DiffQuickFix (Basu et al., 2024) and (i) updating all four matrices. The results are presented in Figure 10. In the updates of the first self-attention layer, the concept of âSnoopyâ was not erased in either setting, and almost identical results were obtained. It can be inferred that slightly updating only the weights of the first self-attention layer contributes mini- mally to concept erasure. When only updating MLPs, it can be observed that the concept disappeared by the end of the fifth epoch. This suggests that knowledge is accumu- lated in MLPs. It is evident that the concept disappeared at earlier stages with both our method and the full-parameter- tuning method. Therefore, it can be inferred that there are parameters related to knowledge other than MLPs. In com- parison with full parameter tuning, the transition of gener- ated images is generally the same. Up to the second epoch, the results were mostly identical when viewed as a whole despite some differences in detail. From this observation, the key other than MLPs is the final self-attention layer. This comparison suggests that even with a transformer- based architecture such as CLIP trained on image-text pairs, the primary knowledge may be accumulated in MLPs. This indicates that in text-to-image diffusion mod- els, using a transformer-based text encoder, regardless of the specific model (such as T5, BERT (Devlin et al., 2019)), updating the parameters of MLPs may enable concept era- sure. However, since we do not have access to the weights of models trained using text encoders trained only with text, we cannot verify this. 5. Limitations The experiments presented in Section 3 demonstrated the effectiveness of our proposed method. However, chal- lenges remain for more general concept erasure. In Fig- ure 11, we show the cases of failed erasing. We consider âchurchâ as an example, this failure. It may stem from the large domain represented by âchurchâ in the text encoderâs latent space. Since our method aims to map to similar se- mantic spaces by fine-tuning the text encoder, it becomes challenging to remove concepts associated with large se- mantic spaces. In such cases, combining our method with other concept erasure methods is feasible.DiffQuick- Fix (Basu et al., 2024) updates one of the weight matrices in the first self-attention layer, which is not updated with our proposed method. Hence, it can be effectively combined with our method to improve concept erasure in challenging cases. 6. Related Works 6.1. Text-to-Image Diffusion Models and Text Encoder Denoising diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020) have achieved success even on large-scale datasets with high variance, such as ImageNet (Deng et al., 2009), by using a simple objective function. They have also demonstrated the ability to handle various resolutions (Ho et al., 2022). While it had been necessary to prepare a sep- arate classifier for conditional generation tasks (Dhariwal & Nichol, 2021), recent advancements have enabled im- age generation without the need for a classifier (Ho & Sal- 8 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Figure 10: Change in concept erasure depending on parameters to be updated parachute church Figure 11: Cases of failed erasing of parachute and church imans, 2021). CLIP (Radford et al., 2021), trained on a large corpus of image-text pairs from the Internet, has established a ro- bust connection between images and natural language. The evolution of image generation with natural language has been accelerated by the fusion of CLIP and diffusion mod- els (Saharia et al., 2022; Ramesh et al., 2022; OpenAI, 2023a; Rombach et al., 2022). Research has also been con- ducted on how to transfer natural-language information to generation models. Qualitatively, models trained with more natural-language data outperformed CLIP, even if there was no quantitative difference. The scaling effect of U-Net in generating images was also found to be marginal (Sa- haria et al., 2022). Combining CLIP with other models allows for handling both visual and linguistic information more effectively (Balaji et al., 2023). In Stable Diffusion 3 (Esser et al., 2024), the adoption of three text encoders significantly contributes to faithful generation aligned with prompts, owing to high-quality text embeddings. We focused on the quality of text conditions in the text-to- image diffusion models. This idea is inspired by the notion that improving the quality of image captions enhances the overall quality of text-to-image generation. Specifically, we intentionally degrade the quality of text embeddings used for conditioning U-Net to attempt concept erasure. 6.2. Erasing Concepts from Text-to-Image Diffusion Models Many image-generation models are trained on a vast amount of data collected from the Internet. Consequently, such datasets often contain undesired images, such as those containing NSFW content or copyrighted material, neces- sitating measures to address these issues when deploying the models to the market. Several studies have been con- ducted on methods to remove specific concepts or objects from generated images, many of which involve modify- ing the weights associated with U-Net, an image genera- tive module (Gandikota et al., 2023; Kumari et al., 2023; Gandikota et al., 2024; Zhang et al., 2024a). Some of these methods update the parameters of U-Net itself or adapters attached to U-Net (Gandikota et al., 2023; Kumari et al., 2023; Zhao et al., 2024; Lu et al., 2024; Heng & Soh, 2023; Kim et al., 2023; Huang et al., 2024), update the weights of cross-attention modules conditioning the U-Net by closed- form equation (Zhang et al., 2024a; Gandikota et al., 2024; Lu et al., 2024), and induce the prevention of harmful con- tent being generated when generating (Schramowski et al., 2023). There are also methods for investigating the flow of knowledge between U-Net and the text encoder, updat- ing only parts of the text encoder (Basu et al., 2024). Par- ticularly when updating parameters associated with U-Net, there is a possibility of decreasing the image fidelity of the generated models, as well as potential reductions in the transmission capability of conditions. The computational complexity of U-Net can lead to time-consuming concept erasure. Most methods require the anchor concepts like the target concept. These often necessitate human or large language models intervention for their generation, without clear guidelines on their quality. Since our method automatically sets the anchor concept in accordance with the latent space of the text encoder, the quality of the anchor concept does not affect the ablation result. 9 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning 6.3. Knowledge in Transformer-based Models Research have progressed in understanding where the knowledge of transformer-based large language models is accumulated. Large language models trained only on nat- ural language are believed to accumulate knowledge in the MLP blocks in the transformer (Meng et al., 2022; Dai et al., 2022; Geva et al., 2021), and updating the param- eters of MLP enables knowledge editing (Yao et al., 2022; Dong et al., 2022). However, regarding the text encoder of CLIP trained using both natural language and images, findings also suggesting that knowledge is primarily accu- mulated in the first self-attention layer (Basu et al., 2024). We have also obtained results suggesting similar knowl- edge accumulation in encoder-only models by treating the CLIP text encoder as a transformer-based language model and limiting the parameters to update. Our research sug- gests that, similar to prior studies, knowledge editing is achievable by updating MLP blocks. We also confirm that more effective concept erasure is possible by updating not only MLP blocks but also the final self-attention layer. 7. Conclusion We interpreted concept erasure from text-to-image diffu- sion models as the disruption of text-image alignment and introduced a method for achieving this by making mi- nor changes to the text encoder. Our method, which in- volves making slight modifications to the CLIP text en- coder, requires fewer updates compared with current meth- ods.Therefore, concept erasure can be executed very rapidly. When the text encoder undergoes minor changes, it is mapped to concepts that are closer in the latent space of the text encoder, leading to more natural changes than those generated by human-selected anchor concepts. In our ex- periments, we confirmed that several target concepts disap- peared while suppressing the effect on other concepts. By varying the parameters to update, we confirmed that knowl- edge is primarily accumulated in the MLP blocks of the transformer, consistent with previous research. Like GPT and BERT, it has been suggested that knowledge accumu- lates in the MLP blocks of the transformer regardless of the training method. We also highlighted the inadequacy of the evaluation metrics used in previous research in the context of machine unlearning, emphasizing the need for proper evaluation. In terms of future directions, there is potential for devel- oping evaluation metrics for cases in which creating a re- train model is not feasible. While we used a very simple method of gradient ascent, there is scope for applying var- ious other methods. Using saliency maps-based methods, such as SalUn (Fan et al., 2024), could be beneficial, as they adaptively control the updated parameters, potentially leading to more effective editing when combined with our proposed method. By updating the text encoder by using adapter tuning, similar to SPM, it is likely to achieve simi- lar functionality to facilitated transport. References Balaji, Y., Nah, S., Huang, X., Vahdat, A., Song, J., Zhang, Q., Kreis, K., Aittala, M., Aila, T., Laine, S., Catanzaro, B., Karras, T., and Liu, M.-Y. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers, 2023. Basu, S., Zhao, N., Morariu, V. I., Feizi, S., and Man- junatha, V.Localizing and editing knowledge in text-to-image generative models.InThe Twelfth In- ternational Conference on Learning Representations, 2024. URLhttps://openreview.net/forum? id=Qmw9ne6SOQ. Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosse- lut, A., Brunskill, E., Brynjolfsson, E., Buch, S., Card, D., Castellon, R., Chatterji, N., Chen, A., Creel, K., Davis, J. Q., Demszky, D., Donahue, C., Doumbouya, M., Durmus, E., Ermon, S., Etchemendy, J., Ethayarajh, K., Fei-Fei, L., Finn, C., Gale, T., Gillespie, L., Goel, K., Goodman, N., Grossman, S., Guha, N., Hashimoto, T., Henderson, P., Hewitt, J., Ho, D. E., Hong, J., Hsu, K., Huang, J., Icard, T., Jain, S., Jurafsky, D., Kalluri, P., Karamcheti, S., Keeling, G., Khani, F., Khattab, O., Koh, P. W., Krass, M., Krishna, R., Kuditipudi, R., Ku- mar, A., Ladhak, F., Lee, M., Lee, T., Leskovec, J., Lev- ent, I., Li, X. L., Li, X., Ma, T., Malik, A., Manning, C. D., Mirchandani, S., Mitchell, E., Munyikwa, Z., Nair, S., Narayan, A., Narayanan, D., Newman, B., Nie, A., Niebles, J. C., Nilforoshan, H., Nyarko, J., Ogut, G., Orr, L., Papadimitriou, I., Park, J. S., Piech, C., Porte- lance, E., Potts, C., Raghunathan, A., Reich, R., Ren, H., Rong, F., Roohani, Y., Ruiz, C., Ryan, J., R Ě e, C., Sadigh, D., Sagawa, S., Santhanam, K., Shih, A., Srinivasan, K., Tamkin, A., Taori, R., Thomas, A. W., Tram ` er, F., Wang, R. E., Wang, W., Wu, B., Wu, J., Wu, Y., Xie, S. M., Ya- sunaga, M., You, J., Zaharia, M., Zhang, M., Zhang, T., Zhang, X., Zhang, Y., Zheng, L., Zhou, K., and Liang, P. On the opportunities and risks of foundation models, 2022. Brock, A., Donahue, J., and Simonyan, K. Large scale GAN training for high fidelity natural image synthe- sis. InInternational Conference on Learning Represen- tations, 2019. URLhttps://openreview.net/ forum?id=B1xsqj09Fm. Chen, W., Hu, H., LI, Y., Ruiz, N., Jia, X., Chang, M.-W., and Cohen, W. W. Subject-driven text-to-image gen- 10 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning eration via apprenticeship learning. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/forum? id=wv3bHyQbX7. Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F. Knowledge neurons in pretrained transform- ers.In Muresan, S., Nakov, P., and Villavicencio, A. (eds.),Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), p. 8493â8502, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.581.URLhttps: //aclanthology.org/2022.acl-long.581. Dai, X., Hou, J., Ma, C.-Y., Tsai, S., Wang, J., Wang, R., Zhang, P., Vandenhende, S., Wang, X., Dubey, A., Yu, M., Kadian, A., Radenovic, F., Mahajan, D., Li, K., Zhao, Y., Petrovic, V., Singh, M. K., Motwani, S., Wen, Y., Song, Y., Sumbaly, R., Ramanathan, V., He, Z., Va- jda, P., and Parikh, D. Emu: Enhancing image generation models using photogenic needles in a haystack, 2023. Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. Imagenet: A large-scale hierarchical image database. In2009 IEEE Conference on Computer Vi- sion and Pattern Recognition, p. 248â255, 2009. doi: 10.1109/CVPR.2009.5206848. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. In Burstein, J., Doran, C., and Solorio, T. (eds.),Proceedings of the 2019 Con- ference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4171â4186, Minneapolis, Minnesota, June 2019. Asso- ciation for Computational Linguistics. doi: 10.18653/v1/ N19-1423. URLhttps://aclanthology.org/ N19-1423. Dhariwal, P. and Nichol, A. Q.Diffusion models beat GANs on image synthesis. In Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, 2021. URLhttps://openreview.net/forum? id=AAWuCvzaVt. Dong, Q., Dai, D., Song, Y., Xu, J., Sui, Z., and Li, L.Calibrating factual knowledge in pretrained language models.In Goldberg, Y., Kozareva, Z., and Zhang, Y. (eds.),Findings of the Association for Computational Linguistics:EMNLP 2022, p. 5937â5947, Abu Dhabi, United Arab Emirates, De- cember 2022. Association for Computational Lin- guistics.doi:10.18653/v1/2022.findings-emnlp. 438. URLhttps://aclanthology.org/2022. findings-emnlp.438. Esser, P., Kulal, S., Blattmann, A., Entezari, R., M Ě uller, J., Saini, H., Levi, Y., Lorenz, D., Sauer, A., Boe- sel, F., Podell, D., Dockhorn, T., English, Z., and Rombach, R.Scaling rectified flow transformers for high-resolution image synthesis.In Salakhutdi- nov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.),Proceed- ings of the 41st International Conference on Ma- chine Learning, volume 235 ofProceedings of Ma- chine Learning Research, p. 12606â12633. PMLR, 21â 27 Jul 2024. URLhttps://proceedings.mlr. press/v235/esser24a.html. Fan, C., Liu, J., Zhang, Y., Wong, E., Wei, D., and Liu, S. Salun: Empowering machine unlearning via gradient- based weight saliency in both image classification and generation.InThe Twelfth International Conference on Learning Representations, 2024. URLhttps:// openreview.net/forum?id=gn0mIhQGNM. Fang, Y., Sun, Q., Wang, X., Huang, T., Wang, X., and Cao, Y.Eva-02:A visual representa- tion for neon genesis.Image and Vision Com- puting,p. 105171,2024.ISSN 0262-8856. doi:https://doi.org/10.1016/j.imavis.2024.105171. URLhttps://w.sciencedirect.com/ science/article/pii/S0262885624002762. Fei, Z., Fan, M., and Huang, J. Gradient-free textual in- version. InProceedings of the 31st ACM International Conference on Multimedia, M â23, p. 1364â1373, New York, NY, USA, 2023. Association for Comput- ing Machinery. ISBN 9798400701085. doi: 10.1145/ 3581783.3612599.URLhttps://doi.org/10. 1145/3581783.3612599. Gal, R., Alaluf, Y., Atzmon, Y., Patashnik, O., Bermano, A. H., Chechik, G., and Cohen-or, D. An image is worth one word: Personalizing text-to-image generation using textual inversion. InThe Eleventh International Confer- ence on Learning Representations, 2023. URLhttps: //openreview.net/forum?id=NAQvF08TcyG. Gandikota, R., Materzynska, J., Fiotto-Kaufman, J., and Bau, D. Erasing concepts from diffusion models. InPro- ceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 2426â2436, October 2023. Gandikota, R., Orgad, H., Belinkov, Y., Materzy Ě nska, J., and Bau, D. Unified concept editing in diffusion mod- els. InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), p. 5111â 5120, January 2024. 11 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Geva, M., Schuster, R., Berant, J., and Levy, O. Trans- former feed-forward layers are key-value memories. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.),Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 5484â5495, Online and Punta Cana, Dominican Republic, November 2021. Association for Computa- tional Linguistics. doi: 10.18653/v1/2021.emnlp-main. 446. URLhttps://aclanthology.org/2021. emnlp-main.446. Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y.Generative adversarial nets.In Ghahramani, Z., Welling, M., Cortes, C., Lawrence, N., and Wein- berger, K. (eds.),Advances in Neural Information Processing Systems, volume 27. Curran Associates, Inc., 2014. URLhttps://proceedings.neurips. c/paper_files/paper/2014/file/ 5ca3e9b122f61f8f06494c97b1afccf3-Paper. pdf. Heng, A. and Soh, H.Selective amnesia: A contin- ual learning approach to forgetting in deep generative models.InThirty-seventh Conference on Neural In- formation Processing Systems, 2023.URLhttps: //openreview.net/forum?id=BC1IJdsuYB. Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y. CLIPScore: A reference-free evaluation metric for image captioning. In Moens, M.-F., Huang, X., Spe- cia, L., and Yih, S. W.-t. (eds.),Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 7514â7528, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics.doi: 10.18653/v1/2021. emnlp-main.595. URLhttps://aclanthology. org/2021.emnlp-main.595. Heusel, M., Ramsauer, H., Unterthiner, T., Nessler, B., and Hochreiter, S. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Guyon, I., Luxburg, U. V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (eds.),Advances in Neural Information Process- ing Systems, volume 30. Curran Associates, Inc., 2017. URLhttps://proceedings.neurips. c/paper_files/paper/2017/file/ 8a1d694707eb0fefe65871369074926d-Paper. pdf. Ho, J. and Salimans, T. Classifier-free diffusion guidance. InNeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. URLhttps:// openreview.net/forum?id=qw8AKxfYbI. Ho, J., Jain, A., and Abbeel, P.Denoising diffusion probabilistic models.In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.),Ad- vances in Neural Information Processing Systems, volume 33, p. 6840â6851. Curran Associates, Inc., 2020. URLhttps://proceedings.neurips. c/paper_files/paper/2020/file/ 4c5bcfec8584af0d967f1ab10179ca4b-Paper. pdf. Ho, J., Saharia, C., Chan, W., Fleet, D. J., Norouzi, M., and Salimans, T. Cascaded diffusion models for high fidelity image generation.Journal of Machine Learning Research, 23(47):1â33, 2022. URLhttp://jmlr. org/papers/v23/21-0635.html. Howard, J. and Gugger, S.Fastai: A layered api for deep learning.Information, 11(2), 2020. ISSN 2078- 2489. doi: 10.3390/info11020108. URLhttps:// w.mdpi.com/2078-2489/11/2/108. Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. LoRA: Low- rank adaptation of large language models.InIn- ternational Conference on Learning Representations, 2022. URLhttps://openreview.net/forum? id=nZeVKeeFYf9. Huang, C.-P., Chang, K.-P., Tsai, C.-T., Lai, Y.-H., Yang, F.-E., and Wang, Y.-C. F. Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers, 2024.URLhttps://arxiv.org/abs/ 2311.17717. Ilharco, G., Ribeiro, M. T., Wortsman, M., Schmidt, L., Hajishirzi, H., and Farhadi, A. Editing models with task arithmetic. InThe Eleventh International Conference on Learning Representations, 2023. URLhttps:// openreview.net/forum?id=6t0Kwf8-jrj. Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Lo- geswaran, L., and Seo, M.Knowledge unlearning for mitigating privacy risks in language models.In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), p. 14389â14408, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.805.URLhttps: //aclanthology.org/2023.acl-long.805. Jayasumana, S., Ramalingam, S., Veit, A., Glasner, D., Chakrabarti, A., and Kumar, S. Rethinking fid: Towards a better evaluation metric for image generation. InPro- ceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), p. 9307â9315, June 2024. 12 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Jia, J., Liu, J., Ram, P., Yao, Y., Liu, G., Liu, Y., Sharma, P., and Liu, S. Model sparsity can simplify machine un- learning. InThirty-seventh Conference on Neural In- formation Processing Systems, 2023.URLhttps: //openreview.net/forum?id=0jZH883i34. Karras, T., Laine, S., Aittala, M., Hellsten, J., Lehtinen, J., and Aila, T. Analyzing and improving the image quality of stylegan. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2020. Kim, S., Jung, S., Kim, B., Choi, M., Shin, J., and Lee, J. Towards safe self-distillation of internet-scale text- to-image diffusion models, 2023.URLhttps:// arxiv.org/abs/2307.05977. Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization, 2017. Kumari, N., Zhang, B., Wang, S.-Y., Shechtman, E., Zhang, R., and Zhu, J.-Y. Ablating concepts in text- to-image diffusion models.InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), p. 22691â22702, October 2023. Li, Y., Liu, X., Kag, A., Hu, J., Idelbayev, Y., Sagar, D., Wang, Y., Tulyakov, S., and Ren, J. Textcraftor: Your text encoder can be image quality controller. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7985â7995, June 2024. Liu, L., Ren, Y., Lin, Z., and Zhao, Z. Pseudo numer- ical methods for diffusion models on manifolds.In International Conference on Learning Representations, 2022. URLhttps://openreview.net/forum? id=PlKWVd2yBkY. Lu, S., Wang, Z., Li, L., Liu, Y., and Kong, A. W.-K. Mace: Mass concept erasure in diffusion models. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 6430â6440, June 2024. Lyu, M., Yang, Y., Hong, H., Chen, H., Jin, X., He, Y., Xue, H., Han, J., and Ding, G. One-dimensional adapter to rule them all: Concepts diffusion models and erasing ap- plications. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7559â7568, June 2024. Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Lo- cating and editing factual associations in gpt. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.),Advances in Neural Information Pro- cessing Systems, volume 35, p. 17359â17372. Curran Associates, Inc., 2022. OpenAI. Improving image generation with better captions, 2023a. OpenAI. Gpt-4 technical report, 2023b. Otani, M., Togashi, R., Sawai, Y., Ishigami, R., Nakashima, Y., Rahtu, E., Heikkil Ě a, J., and Satoh, S. Toward ver- ifiable and reproducible human evaluation for text-to- image generation.InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 14277â14286, June 2023. Pang, L., Yin, J., Xie, H., Wang, Q., Li, Q., and Mao, X. Cross initialization for face personalization of text-to- image models. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR), p. 8393â8403, June 2024. Parmar, G., Zhang, R., and Zhu, J.-Y. On aliased resizing and surprising subtleties in gan evaluation. InCVPR, 2022. Podell, D., English, Z., Lacey, K., Blattmann, A., Dock- horn, T., M Ě uller, J., Penna, J., and Rombach, R. SDXL: Improving latent diffusion models for high-resolution image synthesis. InThe Twelfth International Confer- ence on Learning Representations, 2024. URLhttps: //openreview.net/forum?id=di52zR8xgf. Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I.Learning transferable visual models from natural language su- pervision.In Meila, M. and Zhang, T. (eds.),Pro- ceedings of the 38th International Conference on Ma- chine Learning, volume 139 ofProceedings of Ma- chine Learning Research, p. 8748â8763. PMLR, 18â 24 Jul 2021. URLhttps://proceedings.mlr. press/v139/radford21a.html. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. Exploring the limits of transfer learning with a unified text-to-text transformer.J. Mach. Learn. Res., 21(1), jan 2020. ISSN 1532-4435. Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M. Hierarchical text-conditional image generation with clip latents, 2022. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. High-resolution image synthesis with la- tent diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 10684â10695, June 2022. 13 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Ronneberger, O., Fischer, P., and Brox, T. U-net: Con- volutional networks for biomedical image segmentation. In Navab, N., Hornegger, J., Wells, W. M., and Frangi, A. F. (eds.),Medical Image Computing and Computer- Assisted Intervention â MICCAI 2015, p. 234â241, Cham, 2015. Springer International Publishing. ISBN 978-3-319-24574-4. Ruiz, N., Li, Y., Jampani, V., Pritch, Y., Rubinstein, M., and Aberman, K. Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation. InPro- ceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition (CVPR), p. 22500â22510, June 2023. Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Den- ton, E., Ghasemipour, S. K. S., Gontijo-Lopes, R., Ayan, B. K., Salimans, T., Ho, J., Fleet, D. J., and Norouzi, M. Photorealistic text-to-image diffusion models with deep language understanding. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.),Advances in Neural In- formation Processing Systems, 2022. URLhttps:// openreview.net/forum?id=08Yk-n5l2Al. Schramowski, P., Brack, M., Deiseroth, B., and Kersting, K. Safe latent diffusion: Mitigating inappropriate de- generation in diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 22522â22531, June 2023. Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C. W., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., Wortsman, M., Schramowski, P., Kun- durthy, S. R., Crowson, K., Schmidt, L., Kaczmarczyk, R., and Jitsev, J.LAION-5b: An open large-scale dataset for training next generation image-text mod- els. InThirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. URLhttps://openreview.net/forum? id=M3Y74vmsMcY. Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S.Deep unsupervised learning us- ing nonequilibrium thermodynamics.In Bach, F. and Blei, D. (eds.),Proceedings of the 32nd In- ternational Conference on Machine Learning, vol- ume 37 ofProceedings of Machine Learning Re- search, p. 2256â2265, Lille, France, 07â09 Jul 2015. PMLR. URLhttps://proceedings.mlr. press/v37/sohl-dickstein15.html. Song, J., Meng, C., and Ermon, S.Denoising diffu- sion implicit models. InInternational Conference on Learning Representations, 2021.URLhttps:// openreview.net/forum?id=St1giarCHLP. Thudi, A., Deza, G., Chandrasekaran, V., and Papernot, N. Unrolling sgd: Understanding factors influencing ma- chine unlearning. In2022 IEEE 7th European Sympo- sium on Security and Privacy (EuroS&P), p. 303â319. IEEE, 2022. van der Maaten, L. and Hinton, G. Visualizing data us- ing t-sne.Journal of Machine Learning Research, 9 (86):2579â2605, 2008. URLhttp://jmlr.org/ papers/v9/vandermaaten08a.html. Wang, Y., Yao, Q., Kwok, J. T., and Ni, L. M. Gen- eralizing from a few examples: A survey on few-shot learning.ACM Comput. Surv., 53(3), jun 2020. ISSN 0360-0300.doi: 10.1145/3386252.URLhttps: //doi.org/10.1145/3386252. Wei, Y., Zhang, Y., Ji, Z., Bai, J., Zhang, L., and Zuo, W. Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 15943â15953, October 2023. Weiss, M., Rahaman, N., Locatello, F., Pal, C., Bengio, Y., Sch Ě olkopf, B., Li, L. E., and Ballas, N. Neural attentive circuits. In Oh, A. H., Agarwal, A., Belgrave, D., and Cho, K. (eds.),Advances in Neural Information Process- ing Systems, 2022. URLhttps://openreview. net/forum?id=q41xK9Bunq1. Wightman,R.Pytorchimagemodels. https://github.com/rwightman/ pytorch-image-models, 2019. Xue, Z., Song, G., Guo, Q., Liu, B., Zong, Z., Liu, Y., and Luo, P. RAPHAEL: Text-to-image generation via large mixture of diffusion paths. InThirty-seventh Conference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/forum? id=jUdZCcoOu3. Yao, Y., Huang, S., Dong, L., Wei, F., Chen, H., and Zhang, N. Kformer: Knowledge injection in transformer feed- forward layers. InCCF International Conference on Natural Language Processing and Chinese Computing, p. 131â143. Springer, 2022. Zhang, G., Wang, K., Xu, X., Wang, Z., and Shi, H. Forget- me-not: Learning to forget in text-to-image diffusion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Work- shops, p. 1755â1764, June 2024a. Zhang, Q. and Chen, Y. Fast sampling of diffusion mod- els with exponential integrator.InThe Eleventh In- ternational Conference on Learning Representations, 2023. URLhttps://openreview.net/forum? id=Loek7hfb46P. 14 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Zhang, R., Isola, P., Efros, A. A., Shechtman, E., and Wang, O. The unreasonable effectiveness of deep fea- tures as a perceptual metric. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2018. Zhang, X., Wei, X.-Y., Zhang, W., Wu, J., Zhang, Z., Lei, Z., and Li, Q. A survey on personalized content synthesis with diffusion models, 2024b. Zhao, M., Zhang, L., Zheng, T., Kong, Y., and Yin, B. Separable multi-concept erasure from diffusion models, 2024. 15 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning A. Preliminaries A.1. Diffusion Models Denoising diffusion models consist of two processes: forward and reverse process. In the forward process, the noise is gradually added to the input datax 0 , eventually resulting in pure Gaussian noise. In the reverse process, starting from Gaussian noise, the model predicts the noise added at each time steptâ[0,T]. A.2. Anchor Concepts We refer to conceptBas the anchor concept when transitioning from conceptAto conceptB. For example, with UCE (Gandikota et al., 2024), the following optimization problem is formulated. min W m X i=0 âĽWc i âv â i |z W old c â i ⼠2 2 +ÎťâĽWâW old ⼠2 F Here,Wis the projection matrix,c i is the text embedding of the prompt containing the concepts to be erased (e.g. âVan Gogh styleâ), andc â i is also text embeddings taken from the destination prompt (e.g. âartâ). In this case, we considerc â i to be the anchor concept. As another example, we consider Ablating Concepts (Kumari et al., 2023). The method is formulated as a model-based approach as follows. arg min Îľ θ E x t ,c,c â ,t [w t âĽÎľ θ fixed (x t ,c,t)âÎľ θ (x t ,c â ,t)⼠2 2 ] wherecis a random prompt for the anchor concept (e.g. âcatâ) andc â is modified fromcto include the target concept (e.g. âGrumpy Catâ). In this case,c â is the anchor concept. A.3. Retrain Model We refer to a âretrain modelâ as one that is trained on a training dataset from which the forgetting dataset has been removed. Specifically, letDdenote the training datasets andD f â Dis the forgetting datasets. As an example, we assume a text- to-image model has been trained on LAION-5B and want to erase the concept of âGrumpy Catâ. In this case, LAION-5B isDand the subset of LAION-5B, the elements of which contain âGrumpy Catâ, isD f . TheD r =D f , which isD f removed fromD, denote the remaining datasets, and the model trained usingD r is called a retrain model. B. Updated Layers Table 4 shows the names of the updated layers. C. Experimental Details C.1. Environments We conducted all experiments on a single NVIDIA RTX A5000. Our software environments were PyTorch 1.13.1, diffusers 0.21.4, and transformers 4.34.0. C.2. Baselines We describe the baselines used in our experiments. ⢠ESD (Gandikota et al., 2023): This method updates the parameters related to U-Net. It prepares two outputs: one from Original SD and the other from the fine-tuned model. The fine-tuned model is updated using the L2 Loss of each output. We use ESD-x-1, which updates the parameters related to cross-attention. 16 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Table 4: List of the updated layer name with our method. Layer Name textmodel.encoder.layers.0.mlp.fc1 text model.encoder.layers.0.mlp.fc2 textmodel.encoder.layers.1.mlp.fc1 textmodel.encoder.layers.1.mlp.fc2 textmodel.encoder.layers.2.mlp.fc1 text model.encoder.layers.2.mlp.fc2 textmodel.encoder.layers.3.mlp.fc1 textmodel.encoder.layers.3.mlp.fc2 textmodel.encoder.layers.4.mlp.fc1 text model.encoder.layers.4.mlp.fc2 textmodel.encoder.layers.5.mlp.fc1 textmodel.encoder.layers.5.mlp.fc2 textmodel.encoder.layers.6.mlp.fc1 textmodel.encoder.layers.6.mlp.fc2 text model.encoder.layers.7.mlp.fc1 textmodel.encoder.layers.7.mlp.fc2 textmodel.encoder.layers.8.mlp.fc1 textmodel.encoder.layers.8.mlp.fc2 text model.encoder.layers.9.mlp.fc1 textmodel.encoder.layers.9.mlp.fc2 textmodel.encoder.layers.10.mlp.fc1 textmodel.encoder.layers.10.mlp.fc2 textmodel.encoder.layers.11.mlp.fc1 text model.encoder.layers.11.mlp.fc2 textmodel.encoder.layers.11.selfattn.kproj textmodel.encoder.layers.11.selfattn.vproj textmodel.encoder.layers.11.selfattn.kproj textmodel.encoder.layers.11.selfattn.outproj ⢠UCE (Gandikota et al., 2024): This method updates the weights of the cross-attention in U-Net by closed-form. Unlike ESD, it does not require the output of the Original SD. In the experiments, we used anchor concepts generated by ChatGPT. ⢠SPM (Lyu et al., 2024): This method uses adapter tuning using LoRA (Hu et al., 2022) for the U-Net attention module. Therefore, transferring erased model to another model is easy. C.3. Implementation Details of Baselines We describe the details of the baseline implementations. â˘ESD-x: We used the original implementation using diffusers 7 . Following the original paper, the learning rate was set to10 â5 , number of iterations was 1,000, andΡ= 1. â˘UCE: We used the original implementation 8 . The hyperparameters, except anchor concept, were those provided in the official implementation. The anchor concept was generated by ChatGPT 9 for each concept. The prompt given to ChatGPT waswhen the erased concepts is "TARGET CONCEPT", what concepts to guide the erased concepts towards? Answer the one concept name. â˘SPM: We used the original implementation 10 . For configuration of training, we used the provided configuration 11 . When generating the images, we did not use the negative prompt for fair comparison with other methods. The setting of the generation phase is described in Section 3.2. 7 https://huggingface.co/spaces/baulab/Erasing-Concepts-In-Diffusion/tree/main 8 https://github.com/rohitgandikota/unified-concept-editing/tree/main 9 https://chat.openai.com/ 10 https://github.com/Con6924/SPM/tree/main 11 https://github.com/Con6924/SPM/blob/main/configs/snoopy/config.yaml 17 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning C.4. Implementation Details of Our Proposed Method We used CLIP ImageNet Template to erase the concepts. This template is separated into object and style. We list the details of the templates in Tables 5 and 6. These prompts are also used in textual inversion for diffusers 12 . When updating the text encoder, prompts are randomly chosen for each iteration. Table 5: List of prompts used when erasing object Object Prompt a photo of aObject Name a rendering of aObject Name a cropped photo of theObject Name the photo of aObject Name a photo of a cleanObject Name a photo of a dirtyObject Name a dark photo of theObject Name a photo of myObject Name a photo of the coolObject Name a close-up photo of aObject Name a bright photo of theObject Name a cropped photo of aObject Name a photo of theObject Name a good photo of theObject Name a photo of oneObject Name a close-up photo of theObject Name a rendition of theObject Name a photo of the cleanObject Name a rendition of aObject Name a photo of a niceObject Name a good photo of aObject Name a photo of the niceObject Name a photo of the smallObject Name a photo of the weirdObject Name a photo of the largeObject Name a photo of a coolObject Name a photo of a smallObject Name Table 6: List of prompts used when erasing style Style Prompt a painting in the style ofStyle Name a rendering in the style ofStyle Name a cropped painting in the style ofStyle Name the painting in the style ofStyle Name a clean painting in the style ofStyle Name a dirty painting in the style ofStyle Name a dark painting in the style ofStyle Name a picture in the style ofStyle Name a cool painting in the style ofStyle Name a close-up painting in the style ofStyle Name a bright painting in the style ofStyle Name a cropped painting in the style ofStyle Name a good painting in the style ofStyle Name a close-up painting in the style ofStyle Name a rendition in the style ofStyle Name a nice painting in the style ofStyle Name a small painting in the style ofStyle Name a weird painting in the style ofStyle Name a large painting in the style ofStyle Name D. Additional Results In this section, we present the results of further qualitative evaluation. Comparisons with current methods and further results of the proposed method only are presented. D.1. Additional Comparison Table 7 shows the target concepts, the anchor concepts used with UCE, and the prompts used during generation. Table 7: List of target concepts, anchor concepts, and prompts in generating Target ConceptAnchor Concept (UCE)Prompt in Generating SnoopyPeanutsSnoopy in cyberpunk style Grumpy CatInternet MemeA Grumpy cat laying in the sun. Mickey MouseDisneydrawing of Mickey Mouse walking along the river. R2D2Star Warsportrait of R2D2 Gogh StylePost-ImpressionismPainting of trees in bloom in the style of Van Gogh. 12 https://github.com/huggingface/diffusers/blob/main/examples/textual_inversion/textual_ inversion.py 18 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Figure 12 shows the results of erasing âSnoopyâ. Since the motif of Snoopy is a dog, it can be considered that the proposed method has transitioned to a concept close to it. Additionally, ESD-x-1 resulted in the absence of Snoopy, which is speculated to differ from the retrain results we consider as ground truth. Although the results of UCE were mapped by the concept âPeanutsâ, it is difficult to argue that the âcyberpunk styleâ was adequately reflected. (a) Original SD(b) ESD-x-1(c) UCE(d) SPM(e) Ours Figure 12: Comparison of generated images when erasing âSnoopyâ Figure 13 shows the results of erasing âGrumpy Catâ. ESD-x-1 generated an image of an unrelated building. Both SPM and UCE showed a significant decrease in generation quality. As shown in Figure 3, UCE maintained the quality of concept erasure and target concept generation when given an appropriate anchor concept. However, in this case, since âInternet Memeâ is the anchor concept, it resulted in such a generation. âInternet Memeâ was an anchor concept generated by ChatGPT, but whether such anchor concepts are appropriate depends on the knowledge of the human or model used to generate them. (a) Original SD(b) ESD-x-1(c) UCE(d) SPM(e) Ours Figure 13: Comparison of generated images when erasing âGrumpy Catâ Figure 14 presents the results of erasing âMickey Mouseâ. ESD-x-1 seemed to disregard the target concept, resulting in images more aligned with âdrawingâ and âriverâ present in the prompt. Intuitively, something other than Mickey Mouse should be âwalking along the riverâ. Additionally, both SPM and UCE occasionally produced images reminiscent of âMickey Mouseâ. (a) Original SD(b) ESD-x-1(c) UCE(d) SPM(e) Ours Figure 14: Comparison of generated images when erasing âMickey Mouseâ Figure 15 illustrates the results of erasing âR2D2â. While UCE produced images related to âStar Warsâ, it is worth noting that specifying proper nouns as anchor concepts becomes inappropriate when attempting to erase multiple concepts. As the number of concepts to erase increases, it becomes challenging for humans to designate anchor concepts while considering their relationships. This difficulty also applies when using large language models. Figure 16 shows the result of erasing the âGogh styleâ. It is difficult to say that the Gogh style was completely erased because UCE output starry night, which is strongly associated with Gogh. 19 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning (a) Original SD(b) ESD-x-1(c) UCE(d) SPM(e) Ours Figure 15: Comparison of generated images when erasing âR2D2â (a) Original SD(b) ESD-x-1(c) UCE(d) SPM(e) Ours Figure 16: Comparison of generated images when erasing âVan Gogh styleâ D.2. Additional Results of Our Method We present the results without comparing them with other methods. The purpose is solely to demonstrate the performance of our proposed method. We used all classes of Imagenette (Howard & Gugger, 2020), which is a subset of ImageNet. We generated ten images for each class. The caption we used was âa photo of aclass nameâ. In Figures 17-26 show the comparisons between before and after erasing. Our proposed method demonstrates highly effective erasure performance even for objects containing common nouns such as those in Imagenette. Pre-trained Erased Figure 17: Generated images of gas pump Pre-trained Erased Figure 18: Generated images of cassette player D.3. Transferability Our proposed method updates the text encoder. Therefore, it can transfer the erased model to other text-to-image diffusion models that use the same text encoder. We confirm the transferability to other text-to-image diffusion models. We used 20 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Pre-trained Erased Figure 19: Generated images of chainsaw Pre-trained Erased Figure 20: Generated images of church Pre-trained Erased Figure 21: Generated images of tench (a type of fish) Pre-trained Erased Figure 22: Generated images of garbage truck Pre-trained Erased Figure 23: Generated images of English springer 21 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Pre-trained Erased Figure 24: Generated images of golf ball Pre-trained Erased Figure 25: Generated images of parachute Pre-trained Erased Figure 26: Generated images of French horn Stable Diffusion 1.4 13 with VAE and U-Net. The text encoder is the model from which âSnoopyâ was erased. Figure 27 shows the results. Once a concept is erased from the text encoder, we can be sure that the concept is also erased from the model using the same text encoder. (a) Snoopy in cyberpunk style.(b) a photo of Eiffel Tower. Figure 27: Generated images using the text encoder from which âSnoopyâ was erased with Stable Diffusion 1.4. 13 https://huggingface.co/CompVis/stable-diffusion-v1-4 22 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning We used DreamShaper 14 for community models. We used DEIS Scheduler (Zhang & Chen, 2023) with 25 inference steps. Figure 28 shows the results before and after erasing âEiffel Towerâ. Because of the erasure of the Eiffel Tower from the text encoder, it was erased from the generated image. (a) Using text encoder from which âEiffel Towerâ was not erased(b) Using text encoder from which âEiffel Towerâ was erased Figure 28: Generated images before and after erasing âEiffel Towerâ with DreamShaper D.4. Additional Quantitative Results We provide the additional quantitative results in this subsection. First, we show the detection rate. We erased Imagenette classes and generated 100 images for each class. Then, we evaluated the top-1 accuracy using EVA02 (Fang et al., 2024), which is the highest top-1 accuracy on the ImageNet according to timm (Wightman, 2019) 15 . We used âa photo of a [class name]â as the prompt. The results are shown in Table 8. Our proposed method succeeds in erasing concepts in many cases even if we consider the misdetection of the classifier used in this experiment. Table 8: Detection rate (%) of Imagenette. Lower score indicates that the target concept is erased correctly. ClassErasedOriginal SD church089 parachute0100 golf ball15100 gas pump098 garbage truck886 tench095 French horn0100 chain saw076 English springer094 cassette player312 Average2.685 Second, we confirmed the detection rate due to change the epochs. We use four classes: French horn, golf ball, garbage truck, and tench. Table 9 shows the results. This results indicate the target concept is gradually erased. In addition, as well 14 https://huggingface.co/Lykon/DreamShaper 15 https://github.com/huggingface/pytorch-image-models/blob/main/results/results-imagenet. csv 23 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning as described in Section 4.2, the difficulty to erase for each concept is different. The difficult concept to erase requires the large number of epochs to erase. Table 9: Detection rate (%) of Imagenette. Lower score indicates that the target concept is erased correctly. ClassEpoch 1Epoch 2Epoch 3Epoch 4Epoch 5 French horn00000 golf ball10092452615 garbage truck868572468 tench00000 Third, we calculated the FID scores in addition to CLIP scores shown in Table 3. Following Lyu et al. (2024), we used CLIP ImageNet Template small. It has 27 prompts for objects and 19 prompts for styles. We generated 20 images each prompt. We used clean-fid (Parmar et al., 2022) for computing FID. It is natural that the baselines are lower FID because their ground truth for unrelated concept is the same images of the original SD. From the perspective of the difference of ground truth as well, it is not appropriate to use FID to compare our method and the baselines. It is not expected to improve the evaluation metrics after erasing because most erasure methods are not designed to improve the evaluation metrics. Therefore, it is expected that the methods whose ground truth is the same of the original SD get better score when comparing the existing metrics. The reason that our proposed method gets worse scores is aforementioned. However, as shown in Table 3, our proposed method is competitive in terms of CLIP Score, indicates text-image alignment. This metric does not rely on the generated images from original SD. Therefore, we consider the CLIP Score is better metric. Table 10: FID between each erasure method and original SD. Lower score indicates the images generated by erased model are the same that of original SD. ErasingEiffel TowerErasingMonet Style Tokyo TowerTriumphal archcatGogh stylePicasso StyleHokusai style ESD-x-1173.31318.19927.612231.889149.214121.994 UCE114.8635.66325.89248.51237.10256.135 SPM47.7806.23634.701 57.02462.08157.530 Ours81.15828.80126.738252.36490.956206.979 Fourth, we evaluated the generative ability of the unrelated concepts using MSCOCO-30k FID and CLIP Score. We evaluated only our proposed method. The concepts to be erased are Eiffel Tower and Monet Style. The results are shown in Table 11. These results indicate that our proposed method has minimal impact on the generative ability when erasing a single concept. Moreover, when erasing Eiffel Tower, FID and CLIP Score get better. The fact that performance improvements are observed despite not making any enhancements to the model suggests that thiese evaluation metrics are inherently inappropriate. Table 11: FID and CLIP score of our proposed method on MSCOCO-30k FIDâCLIP Scoreâ Original SD13.89960.2667 ErasingEiffel Tower13.48740.2672 ErasingMonet Style13.94560.2669 Fifth, we calculated detection rate when erasing multiple concepts. We erased tench, English springer, chain saw in that order. The results are shown in Table 12. We used the model at the end of third epoch. Similar to the results in Figure 5, the more concepts that are erased, the more difficult it becomes to maintain other concepts. Notably, using the model at the end of the fourth epoch resulted in a significant degradation of other concepts. This indicates that our method is sensitive to the number of epochs. However, it is easy to erase one concept while maintaining the others according to Table 11. Therefore, task vectors (Ilharco et al., 2023) or other methods may be used to overcome this challenge. 24 Erasing Concepts from Text-to-Image Diffusion Models with Few-shot Unlearning Table 12: Detection rate (%) Erased Class Rate tenchEnglish springerchainsawchurch (unerased concept) tench09510094 +English springer 009493 +English springer+chainsaw00285 25