Paper deep dive
ACE: Concept Editing in Diffusion Models without Performance Degradation
Ruipeng Wang, Junfeng Fang, Jiaqi Li, Hao Chen, Jie Shi, Kun Wang, Xiang Wang
Models: Stable Diffusion
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 5:45:58 PM
Summary
ACE (Concept Editing in Diffusion Models) is a novel method that utilizes cross null-space projection to erase unsafe concepts from text-to-image diffusion models while preserving their general generative capabilities. By projecting parameter perturbations onto the null space of normal text representations, ACE achieves superior semantic consistency and image quality compared to existing baselines like UCE and RECE, with significantly lower computational costs.
Entities (5)
Relation Signals (3)
ACE ā utilizes ā Null-space Projection
confidence 98% Ā· ACE introduces a novel cross null-space projection approach.
ACE ā erases ā Unsafe Concepts
confidence 95% Ā· ACE introduces a novel cross null-space projection approach to precisely erase unsafe concept.
ACE ā improves ā Semantic Consistency
confidence 95% Ā· ACE significantly outperforms the advancing baselines, improving semantic consistency by 24.56%.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Diffusion-based text-to-image models have demonstrated remarkable capabilities in generating realistic images, but they raise societal and ethical concerns, such as the creation of unsafe content. While concept editing is proposed to address these issues, they often struggle to balance the removal of unsafe concept with maintaining the model's general genera-tive capabilities. In this work, we propose ACE, a new editing method that enhances concept editing in diffusion models. ACE introduces a novel cross null-space projection approach to precisely erase unsafe concept while maintaining the model's ability to generate high-quality, semantically consistent images. Extensive experiments demonstrate that ACE significantly outperforms the advancing baselines,improving semantic consistency by 24.56% and image generation quality by 34.82% on average with only 1% of the time cost. These results highlight the practical utility of concept editing by mitigating its potential risks, paving the way for broader applications in the field. Code is avaliable at this https URL
Tags
Links
Trouble viewing inline? Open PDF directly ā
Full Text
77,248 characters extracted from source content.
Expand or collapse full text
ACE: Concept Editing in Diffusion Models without Performance Degradation Ruipeng Wang1 Junfeng Fang1 *Corresponding author: fjf@mail.ustc.edu.cn Jiaqi Li2 Hao Chen3 Jie Shi4 Kun Wang1 Xiang Wang1 1University of Science and Technology of China *Corresponding author: xiangwang1223@gmail.com 2Southeast University 3Beijing University of Posts and Telecommunications 4Huawei wrp20021021@mail.ustc.edu.cn, fjf@mail.ustc.edu.cn, aoluming1996@gmail.com, HaoChenn.Eric@gmail.com, shi.jie1@huawei.com, wk520529@mail.ustc.edu.cn, xiangwang1223@gmail.com Abstract Diffusion-based text-to-image models have demonstrated remarkable capabilities in generating realistic images, but they raise societal and ethical concerns, such as the creation of unsafe content. While concept editing is proposed to address these issues, they often struggle to balance the removal of unsafe concept with maintaining the modelās general generative capabilities. In this work, we propose ACE, a new editing method that enhances concept editing in diffusion models. ACE introduces a novel cross null-space projection approach to precisely erase unsafe concept while maintaining the modelās ability to generate high-quality, semantically consistent images. Extensive experiments demonstrate that ACE significantly outperforms the advancing baselines, improving semantic consistency by 24.56% and image generation quality by 34.82% on average with only 1% of the time cost. These results highlight the practical utility of concept editing by mitigating its potential risks, paving the way for broader applications in the field. Code is avaliable at https://github.com/littlelittlenine/ACE-zero.git WARNING: This paper contains harmful content that can be offensive. ACE: Concept Editing in Diffusion Models without Performance Degradation Figure 1: Images generated by the original and edited Stable Diffusion (SD) v2.1. Red text denotes unsafe concepts, while blue text in input prompts indicates concepts are prone to being overlooked by diffusion models edited using baseline methods. The scores represent LPIPS scores (ā ā) Gandikota et al. (2024), which quantify the discrepancy between the generated images and the images after editing. Detailed implementation is exhibited in Section 5. 1 Introduction Diffusion-based text-to-image (T2I) models have demonstrated remarkable capabilities in generating highly realistic and diverse images Ho et al. (2020); Rombach et al. (2022a). However, their powerful generative potential raises societal and ethical concerns, including the creation of unsafe content such as (1) nude or violent imagery, (2) copyright infringement, and (3) social biases Luccioni et al. (2023); Struppek et al. (2022a), as illustrated in Figure 1. To address these challenges, concept edit has emerged as a promising solution Orgad et al. (2023); Gong et al. (2024). It typically perturbs the attention matrices, denoted as ksubscriptW_kWitalic_k and vsubscriptW_vWitalic_v, in diffusion models by adding perturbations ksubscript _kĪitalic_k and vsubscript _vĪitalic_v to them. To erase unsafe content while preserving normal content, current methods optimize ksubscript _kĪitalic_k and vsubscript _vĪitalic_v through two objectives: (1) erasing the representations of unsafe text prompts that may lead to the generation of unsafe images Gandikota et al. (2024), and (2) preserving the representations of normal text prompts, as shown in Figure 2 (a). While effective, current editing methods face an inherent limitation: they often struggle to balance the trade-off between the above two objectives, i.e., representation erasing and preservation. Specifically, to prioritize safe generation, current studies often emphasize the erasure of unsafe representations, inadvertently amplifying the perturbations ksubscript _kĪitalic_k and vsubscript _vĪitalic_v. This overemphasis on erasure leads to inadequate preservation of normal representations, compromising the modelās generative capabilities. Worse still, this limitation is exacerbated when editing multiple unsafe concepts simultaneously within the same diffusion model. The lack of control over representation preservation accumulates, causing significant deviations in the representations of normal texts. Consequently, the modelās general generative capability deteriorates, resulting in outputs that no longer align with the semantics of the input text prompts. Figure 1 provides an example, where advanced methods, such as UCE Gandikota et al. (2024) and RECE Gong et al. (2024), successfully remove unsafe content but degrades the modelās ability to generate semantically consistent images. To address this challenge, we turn to null-space projection Wang et al. (2021); Fang et al. (2024), an approach recently demonstrated to preserve representations in large language models (LLMs) Brown et al. (2020); Dubey et al. (2024). By projecting parameter perturbations onto the null space of representations within LLMs, the representations could be unaffected by the perturbations. Inspired by this, we propose ACE, a novel editing method that extends null-space projection to diffusion models, enabling precise erasure of unsafe concepts while preserving the integrity of normal representations. As illustrated in Figure 2 (b), ACE follows a three-step paradigm: 1. Concept Erasing: following UCE and RECE, ACE first derives the optimal perturbations ksubscript _kĪitalic_k and vsubscript _vĪitalic_v for erasing unsafe text representations. 2. Null-space Projection: before applying ksubscript _kĪitalic_k and vsubscript _vĪitalic_v to ksubscriptK_kKitalic_k and vsubscriptK_vKitalic_v, they are projected onto the null space of normal text representations. Leveraging the mathematical properties of null spaces Wang et al. (2021), the attention matrices derived from the above two steps can filter out unsafe representations while preserving safe representations without distortion. While these two steps suffice for current applications of null-space projection Fang et al. (2024), T2I models introduce an additional challenge: unsafe and normal representations, after passing through the attention matrices, would be re-coupled in cross-attention process via interactions with image features. Hence, residual unsafe representations, if not fully filtered in these two steps, can again influence the output through interactions with normal representations. To address this, ACE introduces a critical third step: 3. Cross Projection: After passing through the attention matrices, we project unsafe representations onto the null space of normal representations. This serves as the objective for further optimizing ksubscript _kĪitalic_k and vsubscript _vĪitalic_v, ensuring the complete elimination of unsafe representationsā influence on the output. This paradigm endows ACE with the capability to precisely erase unsafe concepts while preserving the modelās general generative capabilities. Figure 2: Comparison of current concept editing methods (a) and our ACE (b). Best viewed in color. To validate the effectiveness of ACE, we conduct extensive experiments on five widely used datasets (e.g., NSFW Hunter (2023), Imagenette Deng et al. (2009) and COCO Lin et al. (2014)) with advancing T2I models such as SDv2.1 Song et al. (2021). The results demonstrate that ACE outperforms strong baselines such as UCE and RECE, as shown in Figure 1. Specifically, ACE significantly improves semantic consistency with input prompts by 24.56% and alignment with the original image by 34.82% on average, while maintaining comparable editing performance. Additionally, by eliminating the complex representation preservation process required in each editing step, ACE requires only 1% of the time compared to baseline methods. More importantly, these empirical results highlight the practical utility of concept editing by mitigating its potential risks. By ensuring safer and more reliable model behavior, ACE paves the way for broader applications and further advancements in the field. 2 Preliminary Diffusion Models. Diffusion models have emerged as a powerful framework for generating high-quality images Rombach et al. (2022a). These models operate by iteratively transforming a noisy initial state into a clean image through a sequence of denoising steps. For text-to-image generation, diffusion models typically leverage the text prompts representations to condition the denoising via cross-attention mechanisms. Specifically, image features act as queries, while text representations are projected into keys and values using ksubscriptW_kWitalic_k and vsubscriptW_vWitalic_v, enabling effective fusion of image and text modalities. Concept Editing. Concept editing aims to remove unsafe content (e.g., nude or violent imagery) from diffusion model outputs by modifying the attention matrices Gong et al. (2024). Let 1subscript1T_1T1 and 0subscript0T_0T0 denote the sets of unsafe and normal text representations, respectively. A set of safe text representations SS is defined as alignment targets for the unsafe texts (e.g., aligning ānudeā with ānormalā). After passing through the attention matrix ksubscriptW_kWitalic_k, the text representations become 1ā²subscript1ā²T_1 T1ā², 0ā²subscriptsuperscriptā²0T _0Tā²0 and ā²superscriptā²S Sā²: kā¢[1ā¢0ā¢]=[1ā²ā¢0ā²ā¢ā²].subscriptdelimited-[]subscript1subscript0delimited-[]superscriptsubscript1ā²subscript0ā²W_k [\,T_1\ T_0\ S\, ]=% [\,T_1 \ T_0 \ S % \, ].Witalic_k [ T1 T0 S ] = [ T1ā² T0ā² Sā² ] . (1) Then, the alignment is achieved by introducing a perturbation ksubscript _kĪitalic_k to ksubscriptW_kWitalic_k, such that: (k+k)ā¢1=ā².subscriptsubscriptsubscript1superscriptā² (W_k+ _k )T_1=S .( Witalic_k + Īitalic_k ) T1 = Sā² . (2) This transformation ensures that unsafe text representations are mapped to safe representations before interacting with image features, effectively editing the targeted unsafe concepts. Additionally, the perturbation must minimize its impact on normal text representations 0subscript0T_0T0. Thus, the overall objective for concept editing is formulated as: k=subscriptabsent _k=Īitalic_k = argā”min^kā¢ā(k+^k)ā¢1āā2+limit-fromsubscript^superscriptnormsubscriptsubscript^subscript12 _k \| (W% _k+ _k )T_1-S \|^2+start_UNDERACCENT over start_ARG Ī end_ARGk end_UNDERACCENT start_ARG arg min end_ARG ā„ ( Witalic_k + over start_ARG Ī end_ARGk ) T1 - S ā„2 + (3) ā(k+^k)ā¢0ā0ā²ā2.superscriptnormsubscriptsubscript^subscript0subscriptsuperscriptā²02 \| (W_k+ _k )T% _0-T _0 \|^2.ā„ ( Witalic_k + over start_ARG Ī end_ARGk ) T0 - Tā²0 ā„2 . This objective admits a closed-form solution, enabling efficient and precise concept editing without the need for gradient-based optimization: k=(ā²ā1ā²)ā¢1ā¤ā¢(1ā¢1ā¤+0ā¢0ā¤)ā1.subscriptsuperscriptā²subscript1ā²subscript1topsuperscriptsubscript1superscriptsubscript1topsubscript0superscriptsubscript0top1 _k=(S -T_1 )T_1^% (T_1T_1 +T_0T_0 )% ^-1.Īitalic_k = ( Sā² - T1ā² ) T1⤠( T1 T1⤠+ T0 T0⤠)- 1 . (4) Note that current concept editing methods apply identical operations to both ksubscriptW_kWitalic_k and vsubscriptW_vWitalic_v. As a result, the formulas and subsequent derivations can be generalized by interchanging k and v. In this case, solving for the perturbation ksubscript _kĪitalic_k on ksubscriptW_kWitalic_k can be directly transferred to solving for the perturbation vsubscript _vĪitalic_v on vsubscriptW_vWitalic_v. 3 Method Current paradigm based on Equation 3 faces a critical trade-off: (1) erasing unsafe concepts often compromises (2) the preservation of normal concepts. Ensuring safe outputs results in distorted normal text representations, degrading the modelās general generative capabilities. To solve this, we introduce ACE, which follows a three-step paradigm as illustrated in Figure 2 (b): STEP 1 in Section 3.1 derives the optimal perturbation ksubscript _kĪitalic_k to erase unsafe concept; STEP 2 in Section 3.2 projects ksubscript _kĪitalic_k onto the null space of normal text representations 0subscript0T_0T0 to preserve the integrity of 0ā²subscriptsuperscriptā²0T _0Tā²0, and STEP 3 in Section 3.3 introduces cross null-space projection to prevent residual unsafe representations from influencing outputs through attention mechanism. 3.1 Erasing Unsafe Concept We derive the optimal perturbation ksubscript _kĪitalic_k following Equation 3. Notably, since STEP 2 (null-space projection) inherently preserves the integrity of 0ā²subscriptsuperscriptā²0T _0Tā²0, we simplify Equation 3 by removing the second term (i.e., the term responsible for preservation). This allows Step 1 to focus solely on concept erasing without trade-offs. Formally: k=argā”min^kā¢ā(k+^k)ā¢1āā2.subscriptsubscript^superscriptnormsubscriptsubscript^subscript12 _k=\, _k \| (% W_k+ _k )T_1-S \|^% 2.Īitalic_k = start_UNDERACCENT over start_ARG Ī end_ARGk end_UNDERACCENT start_ARG arg min end_ARG ā„ ( Witalic_k + over start_ARG Ī end_ARGk ) T1 - S ā„2 . (5) This formulation ensures that unsafe concepts are effectively erased while deferring the preservation of normal concepts to STEP 2. 3.2 Null Space Projection for Preservation In this step, we project the perturbation ksubscript _kĪitalic_k obtained in STEP 1 onto the null space of normal text representations 0subscript0T_0T0. Specifically, null space is defined as follows: for a given matrix AA, if =0BA=0BA = 0, then BB lies in the null space of AA. For more details, please see Adam-NSCL Wang et al. (2021). Based on this, by projecting ksubscript _kĪitalic_k onto the null space of 0subscript0T_0T0, we ensure: (k+k)ā¢0=kā¢0=0ā².subscriptsubscriptsubscript0subscriptsubscript0subscriptsuperscriptā²0(W_k+ _k)T_0=W_kT_0=% T _0.( Witalic_k + Īitalic_k ) T0 = Witalic_k T0 = Tā²0 . (6) This guarantees that ksubscript _kĪitalic_k does not alter the representations of normal text, thereby preserving the modelās general generative capabilities. To achieve this projection, we introduce the null-space projection matrix PP for 0subscript0T_0T0, defined such that for any perturbation ksubscript _kĪitalic_k, kā¢subscript _kPĪitalic_k P lies in the null space of 0subscript0T_0T0, i.e., kā¢0=subscriptsubscript00 _kPT_0=0Īitalic_k PT0 = 0. Then, substituting ksubscript _kĪitalic_k with kā¢subscript _kPĪitalic_k P into Equation 5, the objective becomes: k=argā”min^kā¢ā(k+^kā¢)ā¢1āā²ā2,subscriptsubscript^superscriptnormsubscriptsubscript^subscript1superscriptā²2 _k=\, _k \| (% W_k+ _kP )T_1-S% \|^2,Īitalic_k = start_UNDERACCENT over start_ARG Ī end_ARGk end_UNDERACCENT start_ARG arg min end_ARG ā„ ( Witalic_k + over start_ARG Ī end_ARGk P ) T1 - Sā² ā„2 , (7) ensuring that the new perturbation kā¢subscript _kPĪitalic_k P simultaneously (1) erases unsafe content and (2) preserves normal text representations. Efficient Computation of PP. Here we briefly outline the computation of PP. Following the conventional null space projection process Wang et al. (2021), we first perform Singular Value Decomposition (SVD) on 0subscript0T_0T0 to obtain the left singular vector matrix UU. Next, we remove the eigenvectors in UU corresponding to zero eigenvalues, yielding ^ Uover start_ARG U end_ARG. The projection matrix PP is then computed as: =^ā¢^ā¤,^superscript^topP= U U ,P = over start_ARG U end_ARG over start_ARG U end_ARG⤠, (8) since for any matrix AA, we have: ā¢^ā¢^ā¤ā¢0=.^superscript^topsubscript00A U U T_0=0.A over start_ARG U end_ARG over start_ARG U end_ARG⤠T0 = 0 . (9) Detailed proof is provided in Appendix B.1. Furthermore, when the number of to-be-erased concepts is large (i.e., 0subscript0T_0T0 has a high column dimension), performing SVD directly on 0subscript0T_0T0 becomes computationally expensive. According to the mathematical properties of null spaces, we note that 0subscript0T_0T0 and 0ā¢0ā¤subscript0superscriptsubscript0topT_0T_0 T0 T0⤠share the same null space projection matrix P. Meanwhile, compared to 0subscript0T_0T0 with high-dimensional columns, 0ā¢0ā¤subscript0superscriptsubscript0topT_0T_0 T0 T0⤠has a number of columns equal to the dimension of text representation, which is typically much smaller. As a result, performing SVD on 0ā¢0ā¤subscript0superscriptsubscript0topT_0T_0 T0 T0⤠is significantly faster, further improving computational efficiency. Detailed derivation is provided in Appendix B.2. 3.3 Cross Null-Space Projection As mentioned in Section 1, unsafe and normal representations, after being processed by the attention matrices, are re-coupled during the cross-attention phase through interactions with image features. Consequently, any residual unsafe representations that are not entirely filtered out in the above steps can propagate and influence the final output by interacting with normal representations. To address this, ACE introduces a cross null-space projection approach, tailored to the architecture of cross-attention module within diffusion models. Define the normal text representations 0subscript0T_0T0 passing through vsubscriptW_vWitalic_v as 0ā²subscriptsuperscriptā²0T _0Tā² ā²0. This step aims to project unsafe text representations 1ā²subscriptsuperscriptā²1T _1Tā²1 onto the null space of 0ā²subscriptsuperscriptā²0T _0Tā² ā²0. To achieve this, ACE computes the null-space projection matrix ā²P Pā² ā² for 0ā²subscriptsuperscriptā²0T _0Tā² ā²0 using the method described in Section 3.2. Then, the alignment target of 0ā²subscriptsuperscriptā²0T _0Tā² ā²0, i.e., ā²superscriptā²S Sā², are projected into the null space of 0ā²subscriptsuperscriptā²0T _0Tā² ā²0 through multiplication with ā²ā¢ā²superscriptā²S P Sā² Pā² ā². Hence, Equation 7 is transformed into: k=argā”min^kā¢ā(k+^kā¢)ā¢1āā²ā¢ā²ā2.subscriptsubscript^superscriptnormsubscriptsubscript^subscript1superscriptā²2 _k=\, _k \| (% W_k+ _kP )T_1-S% P \|^2.Īitalic_k = start_UNDERACCENT over start_ARG Ī end_ARGk end_UNDERACCENT start_ARG arg min end_ARG ā„ ( Witalic_k + over start_ARG Ī end_ARGk P ) T1 - Sā² Pā² ā² ā„2 . (10) Similarly, define SS passing through vsubscriptW_vWitalic_v as ā²S Sā² ā² and the null space projection matrix for 0ā²subscriptsuperscriptā²0T _0Tā²0 as ā²superscriptā²P Pā², we achieve the other half of the cross projection by swapping k and v: v=argā”min^vā¢ā(v+^vā¢)ā¢1āā²ā¢ā²ā2.subscriptsubscript^superscriptnormsubscriptsubscript^subscript1superscriptā²2 _v=\, _v \| (% W_v+ _vP )T_1-S% P \|^2.Īitalic_v = start_UNDERACCENT over start_ARG Ī end_ARGv end_UNDERACCENT start_ARG arg min end_ARG ā„ ( Witalic_v + over start_ARG Ī end_ARGv P ) T1 - Sā² ā² Pā² ā„2 . (11) Equation 10 and 11 collectively constitute the objective of ACE. This objective admits a closed-form solution: k=(ā²ā¢ā²ākā¢1)ā¢1ā1ā¢ā1,v=(ā²ā¢ā²āvā¢1)ā¢1ā1ā¢ā1.casessubscriptsuperscriptā²subscriptsubscript1superscriptsubscript11superscript1subscriptsuperscriptā²subscriptsubscript1superscriptsubscript11superscript1 \ array[]l _k=(S P^% -W_kT_1)T_1^-1P^-1,% \\ _v=(S P -W_v% T_1)T_1^-1P^-1. array . start_ARRAY start_ROW start_CELL Īitalic_k = ( Sā² Pā² ā² - Witalic_k T1 ) T1- 1 P- 1 , end_CELL end_ROW start_ROW start_CELL Īitalic_v = ( Sā² ā² Pā² - Witalic_v T1 ) T1- 1 P- 1 . end_CELL end_ROW end_ARRAY (12) In practice, ACE can be easily implemented by replacing the closed-form solution of traditional methods (i.e., Equation 4) with Equation 12. This simple modification yields significantly improved results: while maintaining comparable performance in erasing unsafe content, ACE dramatically enhances the quality of generated images. Even after erasing 100010001000 unsafe concepts from a diffusion model ā a scenario where traditional methods fail to generate normal images ā ACE consistently produces high-quality images indistinguishable from those of the original diffusion model. By mitigating the limitation of concept editing, ACE significantly enhances its practical utility, enabling safer and reliable T2I models. 4 Experiment Figure 3: Performance of diffusion models after edited by the baseline methods w.r.t, various metrics such as CLIP (ā ā), FID (ā ā), LPIPSc (ā ā) and LPIPS (ā ā). Specifically, LPIPSc measures the similarity between images generated by the edited model and the original generated images. Best viewed in color. In this section, we conduct experiments to address the following research questions: ⢠RQ1: Can ACE maintain the modelās general generation capability while erasing nude, violent, and copyright infringement concepts? ⢠RQ2: Can ACE maintain the general generation capability while mitigating social bias? ⢠RQ3: Can ACE be utilized to erase a broader range of concepts in images, such as objects? ⢠RQ4: How does the runtime of ACE compare to that of the baseline methods? 4.1 Experimental Setup This section outlines our methodās capability of concepts erasing while eliminating copyright infringement, nude or violent concepts, and mitigating social biases. We begin with an overview of the evaluation metrics, datasets, and baseline methods. For more detailed descriptions of the experimental settings, please refer to Appendix A. Base T2I Models & Baselines. Following UCE Gandikota et al. (2024) and RECE Gong et al. (2024), we employ SD v1.4 Rombach et al. (2022b) and SD v2.1 Song et al. (2021) as the base models for our experiments. The results are compared against several baseline methods, including SDD Kim et al. (2023), ESD Gandikota et al. (2023), Ablation Kumari et al. (2023), UCE Gandikota et al. (2024), and RECE Gong et al. (2024), which represent a range of existing strategies for concepts erasure or bias mitigation. We use these baselines to erase 1000 unsafe concepts. Datasets & Evaluation Metrics. Our experimental methods are evaluated on copyright infringement datasets (CI) from UCE Gandikota et al. (2024), Inappropriate Image Prompts (I2P) datasets, and professions datasets from UCE Gandikota et al. (2024) to assess the reliability of erasure of copyright infringement, the erasure of nude or violent concepts, and to mitigate social biases. We employ CLIP score Radford et al. (2021), LPIPS (Learned Perceptual Image Patch Similarity) Zhang et al. (2018), and FID (Frechet Inception Distance) Heusel et al. (2017) as evaluation metrics. For debiasing, we employ metrics from UCE Gandikota et al. (2024). 4.2 Unsafe Concepts Erasure (RQ1) Figure 4: Case study on the generation of images by diffusion models after the erasure of copyright infringement using different methods. Best viewed in colour. Method CI COCO I2P CLIP (ā ā) LPIPSā (ā ā) LPIPS (ā ā) LPIPS (ā ā) CLIP (ā ā) FID (ā ā) Nudity (ā ā) Original 31.36±plus-or-minus±0.11 - - - 31.43±plus-or-minus±0.09 14.37±plus-or-minus±0.06 0.140±plus-or-minus±0.00 ESD 17.73±plus-or-minus±0.22 0.55±plus-or-minus±0.02 0.65±plus-or-minus±0.04 0.63±plus-or-minus±0.04 18.43±plus-or-minus±0.19 90.81±plus-or-minus±0.34 0.018±plus-or-minus±0.03 CA 17.96±plus-or-minus±0.23 0.33±plus-or-minus±0.03 0.63±plus-or-minus±0.02 0.57±plus-or-minus±0.05 18.96±plus-or-minus±0.03 88.29±plus-or-minus±0.21 0.013±plus-or-minus±0.02 UCE 21.27±plus-or-minus±0.17 0.40±plus-or-minus±0.04 0.46±plus-or-minus±0.01 0.36±plus-or-minus±0.05 22.10±plus-or-minus±0.10 69.17±plus-or-minus±0.15 0.020±plus-or-minus±0.02 RECE 21.03±plus-or-minus±0.15 0.41±plus-or-minus±0.01 0.50±plus-or-minus±0.03 0.36±plus-or-minus±0.04 22.07±plus-or-minus±0.26 70.40±plus-or-minus±0.30 0.010±plus-or-minus±0.03 ACE 28.11±plus-or-minus±0.29 0.38±plus-or-minus±0.02 0.32±plus-or-minus±0.05 0.25±plus-or-minus±0.03 29.20±plus-or-minus±0.30 47.36±plus-or-minus±0.10 0.020±plus-or-minus±0.02 Table 1: Comparison of ACE with existing methods on copyright infringement erasure and nude concepts erasure tasks. LPIPS, LPIPSā, Nudity, CLIP, and FID. LPIPSā indicates the effectiveness of concept erasure, and Nudity represents the proportion of generated images containing nudity. The best results are highlighted in bold. We first evaluate the effectiveness of ACE for erasing copyright infringement and nude/violent imagery. The qualitative results are illustrated in Figure 4 and Figure 5, while quantitative results for concept erasing are displayed in Table 1. In addition, Figure 3 shows how model performance changes as the number of concept edits increases. Notable observations include: ⢠Obs 1: ACE significantly enhances the generative capability while achieving similar erasure effects. Specifically, ACE provides an average improvement of 31.8% and 30.5% on CLIP and LPIPS metrics, demonstrating that ACE effectively preserves the modelās original capabilities and resists performance degradation caused by parameter perturbations. ⢠Obs 2: The advantages of ACE become more pronounced as the number of edited concepts increases. Specifically, when the number of edited concepts grows from 100 to 1,000, ACEās improvements over the strongest baseline in CLIP and FID scores increase from 10.11% to 51.42% and 38.22% to 51.09%, respectively. These results demonstrate that ACE exhibits robustness against varying levels of perturbations. ⢠Obs 3: ACE can erase unsafe concepts in prompts without affecting the generation of other concepts. Specifically, models edited by baseline methods inevitably omit certain contents in the prompt that are closely related to the erased concepts during image generation. In contrast, ACE avoids this issue, demonstrating minimal disruption to the model behavior. Figure 5: Case study on the generation of images by diffusion models after the erasure of nude concepts using different editing methods. Best viewed in colour. Class SD UCE ESD Ours Truck 79.4±plus-or-minus±1.2 5.5±plus-or-minus±0.3 0.4±plus-or-minus±0.1 55.0±plus-or-minus±0.5 Church 82.4±plus-or-minus±1.0 14.2±plus-or-minus±0.4 3.2±plus-or-minus±0.0 80.4±plus-or-minus±1.7 Ball 97.4±plus-or-minus±0.1 1.2±plus-or-minus±0.0 0.8±plus-or-minus±0.2 68.6±plus-or-minus±1.2 Chute 91.0±plus-or-minus±0.4 1.8±plus-or-minus±0.1 1.1±plus-or-minus±0.0 48.0±plus-or-minus±0.8 Horn 99.2±plus-or-minus±0.2 0.0±plus-or-minus±0.0 3.2±plus-or-minus±0.5 10.8±plus-or-minus±0.3 Courgette 90.6±plus-or-minus±0.6 2.3±plus-or-minus±0.3 0.0±plus-or-minus±0.0 57.2±plus-or-minus±1.2 Foreland 85.6±plus-or-minus±0.6 2.0±plus-or-minus±0.1 0.1±plus-or-minus±0.0 60.2±plus-or-minus±0.6 Bell Pepper 80.2±plus-or-minus±1.1 0.8±plus-or-minus±0.0 0.0±plus-or-minus±0.0 43.6±plus-or-minus±0.4 Avg. 88.2 3.5 1.1 53.0 Table 2: Comparison of object retention performance across ACE and baseline methods after entity erasure. The best results are highlighted in bold. 4.3 Social Biases Mitigation (RQ2) Due to imbalanced training data, T2I models often exhibit social biases in generated images, particularly when generating images for occupational prompts. For example, when using ādoctorā as a prompt, only about 10% of the generated images depict women. To address this, we employ concept editing methods to erase dominant concepts and mitigate bias. For a comprehensive evaluation, we extend the baselines by incorporating other commonly used debiasing methods, such as Concept Algebra Wang et al. (2023) , TIME Orgad et al. (2023), and Debias-VL Chuang et al. (2023). We present results for two prevalent biases in T2I models: gender bias and racial bias. Note that since racial bias encompasses multiple attributes, we employ racial categories based on standards from the U.S. Office of Management and Budget (OMB): White, Black, American Indian, Indigenous American, and Asian. For each bias, we generate 500 images and use CLIP to analyze gender and racial ratios. We compute and report the bias as the deviation of the current ratio from the expected ratio, as shown in Table 3. Additionally, qualitative analysis of gender and racial bias and exhibit in Figure 6 and 7. These results provide the following observations: ⢠Obs 4: ACE effectively mitigates social biases in the outputs of edited diffusion models while minimizing the impact on image quality. Specifically, compared to the best baseline, ACE reduces gender bias across multiple concepts by an average of 27%, achieving gender ratios closest to those of an ideal unbiased model across various professions. This demonstrates the generalizability of the ACE approach. ⢠Obs 5: ACE enables multi-element bias mitigation within a single generated image. Specifically, case studies show that ACE can simultaneously remove gender and racial biases in images, precisely controlling the proportions of elements within complex concepts such as race. 4.4 Broader Range of Editing (RQ3) Profession Original SD Concept Algebra Debias-VL TIME UCE Ours Librarian 0.86±plus-or-minus±0.06 0.66±plus-or-minus±0.07 0.34±plus-or-minus±0.06 0.26±plus-or-minus±0.05 0.11±plus-or-minus±0.05 0.10 ±plus-or-minus±0.03 Teacher 0.42±plus-or-minus±0.01 0.46±plus-or-minus±0.00 0.11±plus-or-minus±0.05 0.34±plus-or-minus±0.06 0.13±plus-or-minus±0.06 0.11±plus-or-minus±0.04 Analyst 0.58±plus-or-minus±0.12 0.24±plus-or-minus±0.18 0.71±plus-or-minus±0.02 0.52±plus-or-minus±0.03 0.25±plus-or-minus±0.03 0.12±plus-or-minus±0.05 Sheriff 0.99±plus-or-minus±0.01 0.38±plus-or-minus±0.22 0.82±plus-or-minus±0.08 0.22±plus-or-minus±0.05 0.14±plus-or-minus±0.03 0.14±plus-or-minus±0.02 Doctor 0.78±plus-or-minus±0.04 0.40±plus-or-minus±0.02 0.50±plus-or-minus±0.04 0.58±plus-or-minus±0.03 0.23±plus-or-minus±0.03 0.12±plus-or-minus±0.06 Table 3: Comparison of object retention performance across ACE and baseline methods after entity erasure. The best results are highlighted in bold. Model UCE RECE Ours SD v1.4 6450.3±plus-or-minus±70.9 17390.6±plus-or-minus±168.2 82.1±plus-or-minus±0.2 SD v2.1 12191.1±plus-or-minus±107.7 32868.2±plus-or-minus±323.5 155.4±plus-or-minus±0.5 Table 4: The model editing duration for different methods. Each methodās duration is averaged over 1000 editing iterations to ensure statistical reliability. Best highlighted in bold. To further demonstrate ACEās capabilities in T2I models, we extend our experiments to erase specific entities (e.g., Apple) from generated images, in addition to the concepts. Compared to concepts like ānudityā, these entities are more explicitly and abundantly represented in the training data of diffusion models, making them more deeply ingrained and harder to erase without compromising the modelās general generative capabilities. Specifically, we randomly select 1,000 entities from the Imagenette dataset for erasure. After ensuring complete erasure of each entity (e.g., āDriverā), we generate 500 images of a related entity (e.g., āTruckā) and use a pre-trained ResNet-50 He et al. (2016) to assess the proportion of high-quality Truck images still generated. Table 2 shows the proportion of images successfully retaining the related entity, leading to the following observation: ⢠Obs 6: ACE can successfully erase entities from images while preserving the quality of related entity generation. Specifically, compared to baseline methods, ACE improves the precision of entity retention by an average of 89.77Ć89.77Ć89.77 Ć, demonstrating its generalization capability and potential for broad applicability. Figure 6: Case study on the generation of images by diffusion models after the mitigation of concepts using ACE. Best viewed in colour. 4.5 Run Time (RQ4) To evaluate the runtime of different methods, we measured the time required for each editing approach to edit the same model using an A100-40G GPU. Table 4 presents the average time per concept when editing 500 concepts, with the following observations: ⢠Obs 7: ACE significantly outperforms baseline editing techniques in terms of time, requiring only about 1% of the time. This efficiency stems from the fact that baseline methods spend substantial time preserving representations, whereas ACE accomplishes the task with just three matrix projections. This improvement in efficiency broadens the prospects for the development of concept editing. Figure 7: Case study on the generation of images by diffusion models after the mitigation of racial bias of our method. Best viewed in colour. 5 Related Work Unsafe Concepts in T2I Models. The widespread adoption of diffusion models Sohl-Dickstein et al. (2015); Rombach et al. (2022b) has enabled stable image generation even on large datasets with high variance. However, T2I models may inadvertently generate unsafe images, such as those containing nudity, violence, or copyright infringement Carlini et al. (2023); Somepalli et al. (2023). Additionally, T2I models risk internalizing and amplifying socio-cultural biases Cho et al. (2023a); Struppek et al. (2022b); Cho et al. (2023b), as stereotypes inherent in training data may be reflected and reinforced in the generated outputs. Previous research has proposed various strategies to address these safety concerns in diffusion models, which can be broadly categorized into training-based methods, training-free methods, and parameter fine-tuning. Training-free Methods. These methods leverage the inherent capabilities of diffusion models, avoiding retraining or fine-tuning by intervening directly during inference. For example, Safe Latent Diffusion (SLD) Schramowski et al. (2023a) extends the generative diffusion process by subtracting target-concept-dependent noise from the predicted noise at each timestep, introducing a safety guidance mechanism to prevent the generation of unsafe content. SAFREE Yoon et al. (2024) constructs a text embedding subspace for target concepts, removes the subspace components from input embeddings, and fuses latent images of initial and processed embeddings in the frequency domain to further refine outputs. Additionally, inference-stage interventions such as safety checkers Rando et al. (2022) and non-classifier guidance Schramowski et al. (2023b) have been proposed to prevent unsafe content generation. However, these measures are easily bypassed in open-source environments. Training-based Methods. These methods primarily modify model parameters to achieve concept erasure, often requiring retraining or fine-tuning. For instance, Concept Ablation (CA) Kumari et al. (2023) aligns the generative distribution of target concepts with that of anchor concepts to erase specific concepts. Erased Stable Diffusion (ESD) Gandikota et al. (2023) fine-tunes the distribution of target concepts to mimic negative guidance distributions. Forget-Me-Not (FMN) Zhang et al. (2024a) suppresses activations related to unsafe concepts in attention layers, while Knowledge Transfer and Removal Bui et al. (2024) bridges the gap between visual and textual features by replacing collected text with learnable prompts. Adversarial training has also been widely adopted to enhance model robustness Huang et al. (2024); Kim et al. (2024); Pham et al. (2024); Zhang et al. (2024b). In model pruning, Selective Pruning Yang et al. (2024) empirically validates performance by pruning concept-related key parameters, SalUn Fan et al. (2024) proposes a weight significance metric and leverages gradient-based forgetting loss to eliminate significant parameters, and ConceptPrune Chavhan et al. (2024) identifies and zeroes out activated neurons in feedforward layers during the forward pass. However, these methods often require extensive computational resources. Editing through closed-form solutions addresses these limitations, as discussed in the following section: ⢠Closed-Form Editing in T2I Models. Inspired by the success of model editing in large language models Meng et al. (2022); Jiang et al. (2025); Li et al. (2025), recent advances in closed-form editing show promise in effectively eliminating unsafe concepts in diffusion models Gandikota et al. (2024); Gong et al. (2024), ensuring compliance while minimizing the risk of circumvention. Specifically, these methods focus on editing pre-trained model parameters to target and eliminate outputs related to unsafe concepts, significantly improving the efficiency of preventing harmful content generation. However, while current paradigms successfully eliminate unsafe concepts, they often struggle to preserve the modelās general generative capabilities, leading to distortions in normal text representations. Our approach addresses this limitation by constraining parameter changes to the null space of prior knowledge, minimizing the impact on the modelās overall performance. Therefore, exploring how to achieve a better balance between safety and generative capabilities remains a critical direction for future research. 6 Conclusion In this work, we proposed ACE, a novel method for concept editing in diffusion models that leverages null-space projection to effectively erase unsafe content while preserving the modelās general generative capabilities. By introducing a three-step framework ā concept erasing, null-space projection, and cross null-space projection ā ACE achieves state-of-the-art performance in maintaining semantic consistency and image quality. Extensive experiments demonstrate significant improvements over existing methods, highlighting ACEās potential for enabling safer and more reliable text-to-image generation. Future work will explore extending ACE to other generative models and addressing broader ethical challenges in T2I models. References Brown et al. (2020) Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual. Bui et al. (2024) Anh Tuan Bui, Khanh Doan, Trung Le, Paul Montague, Tamas Abraham, and Dinh Q. Phung. 2024. Removing undesirable concepts in text-to-image generative models with learnable prompts. CoRR, abs/2403.12326. Carlini et al. (2023) Nicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski, Vikash Sehwag, Florian TramĆØr, Borja Balle, Daphne Ippolito, and Eric Wallace. 2023. Extracting training data from diffusion models. In 32nd USENIX Security Symposium, USENIX Security 2023, Anaheim, CA, USA, August 9-11, 2023, pages 5253ā5270. USENIX Association. Chavhan et al. (2024) Ruchika Chavhan, Da Li, and Timothy M. Hospedales. 2024. Conceptprune: Concept editing in diffusion models via skilled neuron pruning. CoRR, abs/2405.19237. Cho et al. (2023a) Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023a. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3043ā3054. Cho et al. (2023b) Jaemin Cho, Abhay Zala, and Mohit Bansal. 2023b. Dall-eval: Probing the reasoning skills and social biases of text-to-image generation models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3043ā3054. Chuang et al. (2023) Ching-Yao Chuang, Varun Jampani, Yuanzhen Li, Antonio Torralba, and Stefanie Jegelka. 2023. Debiasing vision-language models via biased prompts. CoRR, abs/2302.00070. Deng et al. (2009) Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248ā255. Ieee. Dubey et al. (2024) Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, AurĆ©lien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste RoziĆØre, Bethany Biron, Binh Tang, Bobbie Chern, Charlotte Caucheteux, Chaya Nayak, Chloe Bi, Chris Marra, Chris McConnell, Christian Keller, Christophe Touret, Chunyang Wu, Corinne Wong, Cristian Canton Ferrer, Cyrus Nikolaidis, Damien Allonsius, Daniel Song, Danielle Pintz, Danny Livshits, David Esiobu, Dhruv Choudhary, Dhruv Mahajan, Diego Garcia-Olano, Diego Perino, Dieuwke Hupkes, Egor Lakomkin, Ehab AlBadawy, Elina Lobanova, Emily Dinan, Eric Michael Smith, Filip Radenovic, Frank Zhang, Gabriel Synnaeve, Gabrielle Lee, Georgia Lewis Anderson, Graeme Nail, GrĆ©goire Mialon, Guan Pang, Guillem Cucurell, Hailey Nguyen, Hannah Korevaar, Hu Xu, Hugo Touvron, Iliyan Zarov, Imanol Arrieta Ibarra, Isabel M. Kloumann, Ishan Misra, Ivan Evtimov, Jade Copet, Jaewon Lee, Jan Geffert, Jana Vranes, Jason Park, Jay Mahadeokar, Jeet Shah, Jelmer van der Linde, Jennifer Billock, Jenny Hong, Jenya Lee, Jeremy Fu, Jianfeng Chi, Jianyu Huang, Jiawen Liu, Jie Wang, Jiecao Yu, Joanna Bitton, Joe Spisak, Jongsoo Park, Joseph Rocca, Joshua Johnstun, Joshua Saxe, Junteng Jia, Kalyan Vasuden Alwala, Kartikeya Upasani, Kate Plawiak, Ke Li, Kenneth Heafield, Kevin Stone, and et al. 2024. The llama 3 herd of models. volume abs/2407.21783. Fan et al. (2024) Chongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong, Dennis Wei, and Sijia Liu. 2024. Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net. Fang et al. (2024) Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Xiang Wang, Xiangnan He, and Tat-Seng Chua. 2024. Alphaedit: Null-space constrained knowledge editing for language models. CoRR, abs/2410.02355. Gandikota et al. (2023) Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. 2023. Erasing concepts from diffusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 2426ā2436. IEEE. Gandikota et al. (2024) Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzynska, and David Bau. 2024. Unified concept editing in diffusion models. In WACV, pages 5099ā5108. IEEE. Gong et al. (2024) Chao Gong, Kai Chen, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang. 2024. Reliable and efficient concept erasure of text-to-image diffusion models. In Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part LIII, volume 15111 of Lecture Notes in Computer Science, pages 73ā88. Springer. He et al. (2016) Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770ā778. Heusel et al. (2017) Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6626ā6637. Ho et al. (2020) Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. In NeurIPS. Huang et al. (2024) Chi-Pin Huang, Kai-Po Chang, Chung-Ting Tsai, Yung-Hsuan Lai, Fu-En Yang, and Yu-Chiang Frank Wang. 2024. Receler: Reliable concept erasing of text-to-image diffusion models via lightweight erasers. In Computer Vision - ECCV 2024 - 18th European Conference, Milan, Italy, September 29-October 4, 2024, Proceedings, Part XL, volume 15098 of Lecture Notes in Computer Science, pages 360ā376. Springer. Hunter (2023) Tatum Hunter. 2023. Ai porn is easy to make now. for women, thatās a nightmare. pages NAāNA. The Washington Post. Jiang et al. (2025) Houcheng Jiang, Junfeng Fang, Ningyu Zhang, Guojun Ma, Mingyang Wan, Xiang Wang, Xiangnan He, and Tat-seng Chua. 2025. Anyedit: Edit any knowledge encoded in language models. arXiv preprint arXiv:2502.05628. Kim et al. (2024) Changhoon Kim, Kyle Min, and Yezhou Yang. 2024. Race: Robust adversarial concept erasure for secure text-to-image diffusion model. In European Conference on Computer Vision, pages 461ā478. Springer. Kim et al. (2023) Sanghyun Kim, Seohyeon Jung, Balhae Kim, Moonseok Choi, Jinwoo Shin, and Juho Lee. 2023. Towards safe self-distillation of internet-scale text-to-image diffusion models. CoRR, abs/2307.05977. Kumari et al. (2023) Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu. 2023. Ablating concepts in text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 22634ā22645. IEEE. Li et al. (2025) Zherui Li, Houcheng Jiang, Hao Chen, Baolong Bi, Zhenhong Zhou, Fei Sun, Junfeng Fang, and Xiang Wang. 2025. Reinforced lifelong editing for language models. arXiv preprint arXiv:2502.05759. Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr DollĆ”r, and C. Lawrence Zitnick. 2014. Microsoft COCO: common objects in context. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, volume 8693 of Lecture Notes in Computer Science, pages 740ā755. Springer. Luccioni et al. (2023) Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, and Yacine Jernite. 2023. Stable bias: Analyzing societal representations in diffusion models. volume abs/2303.11408. Meng et al. (2022) Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. Locating and editing factual associations in gpt. Advances in neural information processing systems, 35:17359ā17372. Orgad et al. (2023) Hadas Orgad, Bahjat Kawar, and Yonatan Belinkov. 2023. Editing implicit assumptions in text-to-image diffusion models. In IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023, pages 7030ā7038. IEEE. Pham et al. (2024) Minh Pham, Kelly O. Marshall, Chinmay Hegde, and Niv Cohen. 2024. Robust concept erasure using task vectors. CoRR, abs/2404.03631. Radford et al. (2021) Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748ā8763. PMLR. Rando et al. (2022) Javier Rando, Daniel Paleka, David Lindner, Lennart Heim, and Florian TramĆØr. 2022. Red-teaming the stable diffusion safety filter. arXiv preprint arXiv:2210.04610. Rombach et al. (2022a) Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjƶrn Ommer. 2022a. High-resolution image synthesis with latent diffusion models. In CVPR, pages 10674ā10685. IEEE. Rombach et al. (2022b) Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjƶrn Ommer. 2022b. High-resolution image synthesis with latent diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022, pages 10674ā10685. IEEE. Schramowski et al. (2023a) Patrick Schramowski, Manuel Brack, Bjƶrn Deiseroth, and Kristian Kersting. 2023a. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 22522ā22531. IEEE. Schramowski et al. (2023b) Patrick Schramowski, Manuel Brack, Bjƶrn Deiseroth, and Kristian Kersting. 2023b. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22522ā22531. Sohl-Dickstein et al. (2015) Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermodynamics. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6-11 July 2015, volume 37 of JMLR Workshop and Conference Proceedings, pages 2256ā2265. JMLR.org. Somepalli et al. (2023) Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, and Tom Goldstein. 2023. Diffusion art or digital forgery? investigating data replication in diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023, pages 6048ā6058. IEEE. Song et al. (2021) Jiaming Song, Chenlin Meng, and Stefano Ermon. 2021. Denoising diffusion implicit models. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net. Struppek et al. (2022a) Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting. 2022a. The biased artist: Exploiting cultural biases via homoglyphs in text-guided image generation models. volume abs/2209.08891. Struppek et al. (2022b) Lukas Struppek, Dominik Hintersdorf, and Kristian Kersting. 2022b. The biased artist: Exploiting cultural biases via homoglyphs in text-guided image generation models. Wang et al. (2021) Shipeng Wang, Xiaorong Li, Jian Sun, and Zongben Xu. 2021. Training networks in null space of feature covariance for continual learning. In CVPR, pages 184ā193. Computer Vision Foundation / IEEE. Wang et al. (2023) Zihao Wang, Lin Gui, Jeffrey Negrea, and Victor Veitch. 2023. Concept algebra for text-controlled vision models. CoRR, abs/2302.03693. Yang et al. (2024) Tianyun Yang, Juan Cao, and Chang Xu. 2024. Pruning for robust concept erasing in diffusion models. CoRR, abs/2405.16534. Yoon et al. (2024) Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, and Mohit Bansal. 2024. SAFREE: training-free and adaptive guard for safe text-to-image and video generation. CoRR, abs/2410.12761. Zhang et al. (2024a) Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. 2024a. Forget-me-not: Learning to forget in text-to-image diffusion models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, WA, USA, June 17-18, 2024, pages 1755ā1764. IEEE. Zhang et al. (2018) Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018, pages 586ā595. Computer Vision Foundation / IEEE Computer Society. Zhang et al. (2024b) Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, and Sijia Liu. 2024b. Defensive unlearning with adversarial training for robust concept erasure in diffusion models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024. Appendix A Experimental Setup A.1 Base T2I Models & Baselines. We utilize Stable Diffusion V1.4 Rombach et al. (2022b) (comprising 16 cross-attention layers with an input embedding dimension of 768) and Stable Diffusion V2.1 Song et al. (2021) (also featuring 16 cross-attention layers, but with an input embedding dimension of 1024) as the foundational models for our experiments. To evaluate our approach, we compare the results against several baseline methods, including SDD, ESD, Ablation, UCE Gandikota et al. (2024), and RECE Gong et al. (2024), which encompass a variety of existing strategies for representation erasing and bias mitigation. In our experiments, we focus on editing 1,000 unsafe concepts. For Stable Diffusion V1.4, we set the null space projection dimension to 500, while for Stable Diffusion V2.1, we select a projection dimension of 700. A.1.1 Datasets and Evaluation Metrics Our experimental methods are evaluated on copyright infringement datasets from UCE Gandikota et al. (2024), Inappropriate Image Prompts (I2P) datasets, and professions datasets from UCE Gandikota et al. (2024) to assess the reliability of erasure of copyright infringement, the removal of nude or violent concepts, and mitigate social biases.For representation erasing, We employ CLIP score( image-text matching degree) Radford et al. (2021), LPIPS (Learned Perceptual Image Patch Similarity) Zhang et al. (2018), FID (Frechet Inception Distance) Heusel et al. (2017)as evaluation metrics.For debiasing, we employ metrics from UCE Gandikota et al. (2024) Appendix B Related Proof B.1 Proof for Equation =0subscript00APT_0=0APTbold_0 = 0 The SVD of 0subscript0T_0T0 provides us the eigenvectors UU and eigenvalues Ī. Based on this, we can express UU and Ī as =[1,]subscript1subscript2U=[U_1,U_2]U = [ U1 , Ubold_2 ] and correspondingly =[1002]matrixsubscript100subscript2 = bmatrix _1&0\\ 0& _2 bmatrixĪ = [ start_ARG start_ROW start_CELL Ī1 end_CELL start_CELL 0 end_CELL end_ROW start_ROW start_CELL 0 end_CELL start_CELL Ī2 end_CELL end_ROW end_ARG ], where all zero eigenvalues are contained in 2subscript2 _2Ī2, and 2subscript2U_2U2 consists of the eigenvectors corresponding to 2subscript2 _2Ī2. Since UU is an orthogonal matrix, it follows that: (2)Tā¢0ā¢(0)T=(2)Tā¢1ā¢1ā¢(1)T=.superscriptsubscript2subscript0superscriptsubscript0superscriptsubscript2subscript1subscript1superscriptsubscript10(U_2)^TT_0(T_0)^T=(U_2)^T% U_1 _1(U_1)^T= 0.( U2 )T T0 ( T0 )T = ( U2 )T U1 Ī1 ( U1 )T = 0 . (13) This implies that the column space of 2subscript2U_2U2 spans the null space of 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T. Accordingly, the projection matrix onto the null space of 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T can be defined as: =2ā¢(2)T.subscript2superscriptsubscript2P=U_2(U_2)^T.P = U2 ( U2 )T . (14) Based on the Eqn. (13) and 14, we can derive that: 0ā¢(0)T=2ā¢(2)Tā¢0ā¢(0)T=,subscript0superscriptsubscript0subscript2superscriptsubscript2subscript0superscriptsubscript00APT_0(T_0)^T=AU_2% (U_2)^TT_0(T_0)^T= 0,APT0 ( T0 )T = AU2 ( U2 )T T0 ( T0 )T = 0 , (15) which confirms that APAP projects AA onto the null space of 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T. B.2 Proof for the Shared Null Space of subscript0T_0Tbold_0 and ā¢()subscript0superscriptsubscript0T_0(T_0)^TTbold_0 ( Tbold_0 )T Theorem: Let subscript0T_0Tbold_0 be a mĆnmĆ nm Ć n matrix. Then subscript0T_0Tbold_0 and 0ā¢0Tsubscript0superscriptsubscript0T_0T_0^TT0 T0italic_T share the same left null space. Proof: Define the left null space of a matrix AA as the set of all vectors xx such that Tā¢=0superscript0x^TA=0xitalic_T A = 0. We need to show that if xx is in the left null space of 0subscript0K_0K0, then xx is also in the left null space of 0ā¢0Tsubscript0superscriptsubscript0T_0T_0^TT0 T0italic_T, and vice versa. 1. Inclusion ā¢(Tā¢0)āā¢(Tā¢0ā¢(0)T)superscriptsubscript0superscriptsubscript0superscriptsubscript0N (x^TT_0 ) (% x^TT_0 (T_0 )^T )N ( xitalic_T T0 ) ā N ( xitalic_T T0 ( T0 )T ): ⢠Suppose xx is in the left null space of 0subscript0T_0T0, i.e., Tā¢0=superscriptsubscript00x^TT_0= 0xitalic_T T0 = 0. ⢠It follows that Tā¢(0ā¢(0)T)=(Tā¢0)ā¢(0)T=ā (0)T=superscriptsubscript0superscriptsubscript0superscriptsubscript0superscriptsubscript0ā 0superscriptsubscript00x^T (T_0 (T_0 )^T )= % (x^TT_0 ) (T_0 )^T= 0% Ā· (T_0 )^T= 0xitalic_T ( T0 ( T0 )T ) = ( xitalic_T T0 ) ( T0 )T = 0 ā ( T0 )T = 0. ⢠Therefore, xx is in the left null space of 0ā¢0Tsubscript0superscriptsubscript0T_0T_0^TT0 T0italic_T. 2. Inclusion ā¢(Tā¢0ā¢(0)T)āā¢(Tā¢0)superscriptsubscript0superscriptsubscript0superscriptsubscript0N (x^TT_0 (T_0 )^T% ) (x^TT_0 )N ( xitalic_T T0 ( T0 )T ) ā N ( xitalic_T T0 ): ⢠Suppose xx is in the left null space of 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T, i.e., Tā¢(0ā¢(0)T)=superscriptsubscript0superscriptsubscript00x^T (T_0 (T_0 )^T )= 0xitalic_T ( T0 ( T0 )T ) = 0. ⢠Expanding this expression gives (Tā¢0)ā¢(0)T=superscriptsubscript0superscriptsubscript00 (x^TT_0 ) (T_0 )^T= 0( xitalic_T T0 ) ( T0 )T = 0. ⢠Since 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T is non-negative (as any vector multiplied by its transpose results in a non-negative scalar), Tā¢0superscriptsubscript0x^TT_0xitalic_T T0 must be a zero vector for their product to be zero. ⢠Hence, xx is also in the left null space of 0subscript0T_0T0. From these arguments, we establish that both 0subscript0T_0T0 and 0ā¢0Tsubscript0superscriptsubscript0T_0T_0^TT0 T0italic_T share the same left null space. That is, xx belongs to the left null space of 0subscript0T_0T0 if and only if xx belongs to the left null space of 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T. This equality of left null spaces illustrates the structural symmetry and dependency between 0subscript0T_0T0 and its self-product 0ā¢(0)Tsubscript0superscriptsubscript0T_0(T_0)^TT0 ( T0 )T.