Paper deep dive
Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors
Qiao Li, Xiaomeng Fu, Wangjia Yu, Runze He, Baisen Wang, Jiao Dai, Jizhong Han
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/14/2026, 5:14:02 AM
Summary
The paper proposes a method for controllable removal of copyrighted animation characters from text-to-image diffusion models using optimized semantic anchors. The approach constructs an anchor embedding that preserves structural outlines while differentiating detailed features, then replaces target-related embeddings during generation via a structure-aware adaptive strategy. This allows for fine-grained control over erasure degree, multi-target removal, and model transferability while preserving image fidelity.
Entities (13)
Relation Signals (10)
Proposed Method โ iscompatiblewith โ Stable Diffusion v1.4
confidence 95% ยท Experiments show that our approach achieves superior performance... on generating images from Stable Diffusion-v1-4.
Proposed Method โ uses โ Semantic Anchor
confidence 95% ยท We propose a controllable method... We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate
Semantic Anchor โ isoptimizedin โ CLIP
confidence 90% ยท we initialize a word vector... by looking up the token โAnchor*โ in the CLIP [31] text encoderโs embedding layer.
Spider-Man โ issubjectoflawsuitagainst โ Midjourney
confidence 90% ยท Recent disputes, such as the lawsuit by Disney and Universal Studios against Midjourney regarding images resembling characters like Spider-Man
Minions โ issubjectoflawsuitagainst โ Midjourney
confidence 90% ยท Recent disputes, such as the lawsuit by Disney and Universal Studios against Midjourney regarding images resembling characters like Spider-Man and the Minions
Proposed Method โ outperforms โ TraSCE
confidence 90% ยท Experiments show that our approach achieves superior performance... compared with baselines.
Proposed Method โ outperforms โ Safe Latent Diffusion (SLD)
confidence 90% ยท Experiments show that our approach achieves superior performance... compared with baselines.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized reproduction of animation characters. Existing concept erasure methods fall short for animation character erasure: model modification methods struggle to identify suitable anchors for diverse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hindering real-world deployment. In this paper, we propose a controllable method operating on the model's continuous textual representation to erase target characters during generation. We optimizes an anchor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effectiveness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance.
Tags
Links
- Source: https://arxiv.org/abs/2608.12806v1
- Canonical: https://arxiv.org/abs/2608.12806v1
Trouble viewing inline? Open PDF directly โ
Full Text
49,019 characters extracted from source content.
Expand or collapse full text
Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors Qiao Li liqiao@iie.ac.cn Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing, China Xiaomeng Fu fuxiaomeng@iie.ac.cn Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing, China Wangjia Yu yuwangjia@iie.ac.cn Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing, China Runze He hrz010109@gmail.com Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing, China Baisen Wang wbs2788@gmail.com Institute of Information Engineering, Chinese Academy of Sciences School of Cyber Security, University of Chinese Academy of Sciences Beijing, China Jiao Dai โ daijiao@iie.ac.cn Institute of Information Engineering, Chinese Academy of Sciences Beijing, China Jizhong Han hanjizhong@iie.ac.cn Institute of Information Engineering, Chinese Academy of Sciences Beijing, China Spider Man Kung Fu Panda Minions Stable Diffusion v1.4 Scale Original Erased Stable Diffusion XL Transferability across different model architectures Fine-grained control of erasure scales Z-Image Figure 1: During image generation, our method effectively erases diverse animation concepts using optimized semantic anchors, while preserving overall image fidelity. It also supports model transferability and fine-grained control over the erasure scale. This work is licensed under a Creative Commons Attribution 4.0 International License. M โ26, Rio de Janeiro, Brazil ยฉ 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2213-4/2026/11 https://doi.org/10.1145/3767308.3835380 Abstract The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unautho- rized reproduction of animation characters. Existing concept era- sure methods fall short for animation character erasure: model arXiv:2608.12806v1 [cs.CV] 13 Aug 2026 M โ26, November 10โ14, 2026, Rio de Janeiro, BrazilQiao Li et al. modification methods struggle to identify suitable anchors for di- verse, highly distinctive characters; prompt-based steering methods lack fine-grained control for precise intervention. These approaches often yield incomplete erasure and degraded image fidelity, hinder- ing real-world deployment. In this paper, we propose a controllable method operating on the modelโs continuous textual representation to erase target characters during generation. We optimizes an an- chor embedding via structural and detailed constraints to serve as a character surrogate, then replaces target-related embeddings with the anchor via a structure-aware adaptive strategy. Experiments show that our method achieves state-of-the-art erasure effective- ness and image fidelity preservation, while supporting controllable erasure degree, multi-target removal, and model transferability. Moreover, our optimized anchors are plug-and-play with current model modification baselines to improve their erasure performance. CCS Concepts โข Security and privacyโHuman and societal aspects of se- curity and privacy;โข Computing methodologiesโComputer vision. Keywords Copyright protection; Concept erasure; Diffusion models ACM Reference Format: Qiao Li, Xiaomeng Fu, Wangjia Yu, Runze He, Baisen Wang, Jiao Dai, and Jizhong Han. 2026. Erase but Preserve: Controllable Removal of Copy- righted Animation Characters via Optimized Semantic Anchors. In Pro- ceedings of the 34th ACM International Conference on Multimedia (M โ26), November 10โ14, 2026, Rio de Janeiro, Brazil. ACM, New York, NY, USA, 9 pages. https://doi.org/10.1145/3767308.3835380 1 Introduction Text-to-image diffusion models [13,33,37] have become a core tool for visual content creation, routinely used in advertising, film- making, and user-generated content platforms. However, this broad adoption raises legal and ethical risks [4,11,27,45], as these models can generate unsafe or infringing content that violates policies. A practical and high-impact case arising from commercial deployment is the unauthorized generation of copyrighted animation characters. Recent disputes, such as the lawsuit by Disney and Universal Stu- dios against Midjourney regarding images resembling characters like Spider-Man and the Minions [3], highlight that this issue is not hypothetical: it directly affects product deployment, platform governance, and creators who rely on generative models. To mitigate the risks of undesired generation, concept erasure techniques have emerged as one of the feasible solutions, primarily aiming to prevent models from generating unsafe concepts. Existing erasure approaches can be categorized into: (i) model modifica- tion methods that alter model parameters to erase or suppress undesired concepts, typically by mapping them to a neutral or be- nign anchor concept [6,9,10,17,26,28,42], and (i) prompt-based steering methods that adjust input prompts or introduce negative terms to avoid undesired concepts during inference [16,29,35,41]. Although effective in certain scenarios (e.g. erasing unsafe con- cepts like โnudityโ), neither approach adequately focuses on and addresses the specific challenges of erasing copyrighted animation characters in real-world deployment, primarily due to the unique properties of these characters. First, animation characters exhibit a large variety, with each be- ing highly distinctive. This poses challenges for model modification methods, which typically require selecting an appropriate anchor as a benign surrogate for undesired target. While existing anchor se- lections such as synonyms, parent/child classes, or general concepts (e.g., null text or โgroundโ) work well for generic categories (e.g., mapping โgrumpy catโ to โcatโ, โnudityโ to โclothedโ), they are often ill-suited for specific characters. Identifying appropriate semantic synonyms or taxonomic relations for these unique characters is often laborious or even infeasible. Besides, prior studies [28,44] indicate that resorting to general or semantically distant anchors can significantly compromise erasure performance, including in- complete erasure and context contamination. Second, erasing a copyrighted animation character often requires more nuanced intervention than coarsely blocking an entire class of not-safe-for-work (NSFW) content. (i) Unlike inherently harmful NSFW content, the assessment of animation infringement varies across laws and platform policies, and may in some cases permit moderate visual similarity that does not constitute copyright in- fringement. (i) Animation imagery often carries commercial and entertainment value on user-generated content platforms. Instead of indiscriminately blocking all potentially infringing prompts, plat- forms typically aim to preserve the userโs original creative intent (e.g., composition, style, background, unrelated elements) while only excising the copyrighted characters. These two concerns ne- cessitate fine-grained control over both the erased character subject and the preserved contextual elements during inference. However, existing prompt-based steering methods typically rely on discrete textual descriptions that provide only coarse control over genera- tion, thereby lacking precise regulation of both the character era- sure degree and the surrounding context retention. To address these challenges, we propose a method that control- lably erases animation characters during generation by operating on the modelโs continuous textual representation. Our key idea is to replace the target character with a learned anchor concept that explicitly erases its primary visual features while preserving unrelated contextual elements from the prompt. Specifically, we first optimize an anchor embedding by extracting structure out- lines and detailed features from the targetโs visual semantics. This learned anchor is then used to replace the target to guide generation toward a non-infringing surrogate. As the anchor is represented in a continuous embedding space, our method enables fine-grained control over the character erasure degree via adjustment of the replacement intensity. To avoid unintended alterations to unrelated context, we perform targeted embedding replacement: leverag- ing the disentanglement property of textual embeddings, we replace only the embeddings related to the target character while retaining unrelated elements. This replacement is applied adaptively across denoising timesteps, which further improves both reliable target removal and overall scene coherence. Due to the lack of standard benchmark for animation character erasure, we build a dataset of 80 animation concepts that can be reliably generated by diffusion models. Experiments show that our approach achieves superior performance in both target character removal and overall fidelity preservation compared with baselines. Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic AnchorsMM โ26, November 10โ14, 2026, Rio de Janeiro, Brazil Inference Stage ํณ S Prompt containing targets to erase โAnchor*โ โSnoopyโ โAnchor*โ U-Net U-Net U-Net Tokenizer Embedding Lookup Text Transformer [SOT] [EOT] Paddings Anchor* Anchor* Replace Time Text Encoder [EOT] Paddings Low Frequency Filter Low Frequency Filter Low Frequency Filter ํ ํ ~(ํป ํด ,ํป ํฏ ) ํ ํ ~(ํป ํณ ,ํป ํด ) Outline ํ ํ ~(ํป ํด ,ํป ํฏ ) Optimize โAnchor*โ Outline Anchor Concept Construction Text Encoder [SOT] [EOT] Paddings ...... Structure-aware AdaptiveReplacement RandomNoise Random Noise Reference Image Noisex T x t x tโํ x tโi x ํ (ํจํฎํซํฌ) Normal Generation ํ ํณ ํ ํณ Noisy Image PredictedNoise Reconstructed Reconstructed Embedding Optimization Denoising Time ํ ํณ (ํ ํ|ํ ํ (ํ ํ )) ํ ํณ (ํ ํ|ํ ํ (ํ ํ )) ํ ํ|ํ ํ (ํ ํ ) ํ ํ|ํ ํ (ํ ํ ) Detailed Structural ํณ D ... [SOT] [EOT] Paddings ... Target-related embedding Unrelated embedding Anchor* Figure 2: The overall pipeline of our proposed method. We construct an anchor by applying structural constraints (top-left) and detailed constraints (bottom-left), and optimize an anchor embedding in the continuous textual embedding space (middle). During inference, when the input prompt contains target-related terms in the predefined subspace, our method performs structure-aware adaptive embedding replacement to erase the target using the optimized anchor concept (blue line on the right), compared with normal generation (green line on the right). Our method also supports fine-grained control over the erasure de- gree, simultaneous removal of multiple targets, and transferability across different diffusion models (including models with dual text encoders and recent DiT-based models). Moreover, our optimized anchors can also be directly integrated into current model modifi- cation methods as a benign surrogate. Experiments demonstrate that, compared to adopting existing general anchors, leveraging our learned anchors yields higher erasure accuracy and improved image fidelity preservation in animation character erasure. Our contributions are summarized as: โขWe propose a novel method to erase one or more copyrighted animation characters directly during the generation process of text-to-image diffusion models. โขOur continuous anchor optimization approach ingeniously leverages the visual features of target characters, offering a principled way to identify a controllable anchor for distinc- tive animation characters. โขOur structure-aware adaptive replacement strategy jointly achieves precise target character removal and high fidelity preservation, ensuring the coherence and usability of the resulting animation imagery. โขExperiments show that our method achieves state-of-the-arts in animation character erasure, while enabling controllable erasure degree, simultaneous removal of multi-targets, and model transferability. Moreover, our optimized anchors are plug-and-play with model modification methods to improve their erasure performance. 2 Related Work 2.1 Text-to-image Diffusion Models Text-to-image diffusion models have garnered substantial attention due to their capacity for high-fidelity image synthesis [2,7,32,34]. They incorporate image encoder-decoder frameworks to efficiently conduct the diffusion and denoising process within a latent space. During the training process, random Gaussian noiseํis intro- duced to the image ํฅ 0 : ํฅ ํก = โ ฬ ํผ ํก ํฅ 0 + โ 1โ ฬ ํผ ํก ํ(1) The training goal of diffusion models is to learn to predict the introduced noise from ํฅ ํก at time step t: L := E ํโผN(0,1),ํกโผํ(0,ํ) [||ํโํ ํ ํฅ ํก ,ํก,ํ ํ (ํฆ) || 2 2 ](2) where ํ ํ is a U-Net, ํ ํ is a text encoder, y is a textual input. During the inference process, previous works [18,30,39] suggest that models focus on constructing low-level structure and outlines in the early denoising stages, and subsequently shift to predicting semantic details in the later stages. Deterministic DDIM Scheduler. To accelerate the denoising pro- cess, deterministic DDIM sampling [36] has been proposed, en- abling a skip-step strategy. The skip-step denoising process for any timestep ํ < ํ can be mathematically formulated as follows: ํฅ ํ = โ ฬ ํผ ํ ห ํฅ 0|ํ + โ 1โ ฬ ํผ ํ ํ ํ ํฅ ํ ,ํ,ํ ํ (ํฆ) (3) where: ห ํฅ 0|ํ = 1 โ ฬ ํผ ํ ํฅ ํ โ โ๏ธ 1โ ฬ ํผ ํ ํ ํฅ ํ ,ํ,ํ ํ (ํฆ) (4) 2.2 Concept Erasure in Diffusion Models Model Modification Methods. Model modification methods up- date diffusion modelsโ weights to either suppress the undesired target or map it onto a neutral or benign anchor concept. Most existing works [6,9,17,26,28,42] modify parameters in the cross- attention mechanism, text encoder, or the entire model to erase target concepts through iterative fine-tuning. Several works [10,20] directly derive the updated weights via closed-form solution, thus avoiding the need for fine-tuning. However, most of these methods require an appropriate anchor concept to replace the target concept. While this is relatively straightforward for generic objects (e.g. dog) or NSFW content (e.g. nudity), it can be laborious or even infeasible M โ26, November 10โ14, 2026, Rio de Janeiro, BrazilQiao Li et al. ...... ํ ํโํ ํ ํ ํ ํ Low Frequency Filter Low Frequency Filter ํ ํณ (ํ ํ ) ํ ํณ (ํ ํโํ ) ํณ ํ =|ํ ํณ ํ ํ โํ ํณ ํ ํโํ | ํ ํ Replace Time Step ํ ํโi Figure 3: Illustration of the structure-aware adaptive replace- ment module. We extract low-frequency structural compo- nents and analyze the changes between adjacent denoising steps to find an optimal starting point for replacement. for various highly unique animation characters. Our method pro- vide an effective solution for constructing a suitable anchor, which can be directly applied to existing model modification methods. Prompt-based Steering Methods. Current prompt-based steering methods primarily rely on classifier-free guidance (CFG) [14]. Safe Latent Diffusion (SLD) [35] uses multiple noise predictions to steer the unconditional prediction towards a safe prompt while avoiding the negatives. Negative Prompting (NP) is a technique that replaces the empty prompt in CFG with a negative one. TraSCE [16] mod- ifies NP by preserving part of the unconditional predictions and introducing a loss-based guidance mechanism. SAFREE [41] steers prompt tokens away from a toxic subspace. However, these meth- ods mainly focus on global NSFW removal, which fail to provide fine-grained control, proving inadequate for animation characters that require nuanced processing. 3 Method Our goal is to erase target animation characters during the genera- tion process of a diffusion model, while preserving the overall visual fidelity of the generated images. First, we define the objective to construct an anchor concept using the target characterโs structural and detailed features (Section 3.1). Next, we optimize the anchor embedding in the continuous textual embedding space (Section 3.2). During inference, we selectively replace the target-related embed- dings with the optimized embedding following structure-aware adaptive strategy, thereby achieving precise target erasure while preserving scene coherence (Section 3.3). The overall pipeline of our method is illustrated in Figure 2. 3.1 Anchor Concept Construction We aim to remove a copyrighted animation character during genera- tion while preserving overall image fidelity, including the coherence of the background and other contextual elements. To achieve this, we first construct an anchor concept that serves as a benign surro- gate for the copyrighted target. The anchor must simultaneously fulfill two criteria: (i) its outline and general structure should be roughly similar to those of the target to ensure harmonious replace- ment and avoid background distortion. We define the loss function for constructing the structural outline asL ํ ; (i) its main detailed features should exhibit significant distinctiveness from those of the target to avoid copyrighted appearance. We define the loss function for differentiating details asL ํท . Overall, we formulate the anchor construction problem as: min ( ํผ ยทL ํ + ํฝยทL ํท ) (5) where ํผ and ํฝ are determined empirically. During optimization, we use the string โAnchor*โ to represent the new anchor conceptโs name in the word space, as it is not defined in the text encoderโs vocabulary before. Structural Outline Construction. We aim to construct an anchor concept that shares a similar structural outline with the target. We leverage the modelโs generative prior to capture diverse structural poses and layouts. As shown in Figure 2 (top-left), given a random Gaussian noiseํฅ ํ โผ N(0,ํผ), we first randomly sample a large timestepํก ํ โผ (ํ ํ ,ํ ํป ), where the diffusion model predominantly captures the global structure. We then obtain an intermediate latent ํฅ ํก ํ (ํก ํ < ํ)via the deterministic DDIM skip-step denoising formula (defined in Equation 3): ํฅ ํก ํ = โ๏ธ ฬ ํผ ํก ํ ห ํฅ 0|ํ + โ๏ธ 1โ ฬ ํผ ํก ํ ํ ํ ํฅ ํ ,ํ,ํ ํ (ํฆ ํก ) (6) where ห ํฅ 0|ํ can be derived following Equation 4,ํฆ ํก denotes the name of the target character (e.g. โSnoopyโ). Thisํฅ ํก ํ encodes the coarse structure of the target character. Subsequently, following Equation 4, we reconstruct two original samples ห ํฅ 0|ํก ํ (ํฆ) from ํฅ ํก ํ : ห ํฅ 0|ํก ํ (ํฆ)= 1 โ ฬ ํผ ํก ํ ํฅ ํก ํ โ โ๏ธ 1โ ฬ ํผ ํก ํ ํ ํ ํฅ ํก ํ ,ํก ํ ,ํ ํ (ํฆ) (7) By applying two different promptsํฆ, we can obtain two recon- structed samples: ห ํฅ 0|ํก ํ (ํฆ ํก )for the target promptํฆ ํก , and ห ํฅ 0|ํก ํ (ํฆ ํ ) for the anchor nameํฆ ํ (i.e. โAnchor*โ). To ensure the anchor learns a similar structural outline as the target, we maximize the structural similarity between two recon- structed samples. Since structural information is typically repre- sented by low-frequency signals, we employ a low-frequency filter ํ ํฟ to extract their low-frequency components ํฅ ํฟ (ํฆ ํก ) and ํฅ ํฟ (ํฆ ํ ): ํฅ ํฟ (ํฆ ํก )= ํ ํฟ ห ํฅ 0|ํก ํ (ํฆ ํก ) , ํฅ ํฟ (ํฆ ํ )= ํ ํฟ ห ํฅ 0|ํก ํ (ํฆ ํ ) (8) Our goal of optimizing the anchorโs overall structural outline can thus be formulated as: L ํ = E โฅ ํฅ ํฟ (ํฆ ํก )โ ํฅ ํฟ (ํฆ ํ ) โฅ 2 2 (9) Detailed Features Differentiation. To ensure the erasure of in- fringing elements, the main detailed features of the anchor concept should differ from those of the target. We choose a clean reference imageํฅ 0 of the target character whose content clearly defines the infringing features to be erased. As shown in Figure 2 (bottom-left), we add a Gaussian noiseํ โผ N(0,ํผ)toํฅ 0 at a random timestep ํก ํ โผ (ํ ํฟ ,ํ ํ )to obtainํฅ ํก ํ , as in Equation 1; this noise level primar- ily degrades fine details while preserving coarse structure. Based on the diffusion training objective (Equation 2), we maximize the noise prediction error under the anchor promptํฆ ํ to prevent it from reconstructing these details: L ํท =โE ํโํ ํ ํฅ ํก ํ ,ํก ํ ,ํ ํ (ํฆ ํ ) 2 2 (10) Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic AnchorsMM โ26, November 10โ14, 2026, Rio de Janeiro, Brazil Table 1: Quantitative comparison with baselines on erasing 80 characters when generating images from Stable Diffusion-v1-4.โ represents that a higher value indicates better performance, and vice versa. (Bold: best. Underline: second-best.) Method Erasure Effectiveness Image Fidelity Preservation Unrelated Image LLaVA-1.5โBLIP-3โSSIMโLPIPSโAestheticโFIDโCLIPโ SD v1.4 (Base)66.9%64.7%105.3033.70.326 SLD-medium42.3%50.3%0.3140.6785.1534.90.305 SLD-strong20.0%18.2%0.2860.7075.0836.10.298 SAFREE12.8%11.9%0.2130.7725.1035.30.307 Negative Prompt11.1%11.5%0.3840.6395.1035.40.302 STG19.3%17.0%0.4310.5604.8738.70.281 TraSCE9.5% 5.7%0.3470.6934.9935.60.299 Ours6.0%4.0%0.4670.5055.1833.20.312 3.2 Anchor Embedding Optimization After defining the anchor conceptโs construction objective, we opti- mize it as a textual embedding in the continuous embedding space. Initialization. To represent the anchor concept, we initialize a word vectorํฃ โ โ R 1รํท (ํทis the feature dimension) by looking up the token โAnchor*โ in the CLIP [31] text encoderโs embedding layer. This yields a learnable starting point for anchor optimization. Optimization. Following the anchor construction objective in Equation 5, ํฃ โ is optimized by minimizing: ํฃ โ = arg min ํฃ โ ( ํผ ยทL ํ + ํฝยทL ํท ) (11) During optimization, when inputting anchor prompt, we form a token sequence containing Start-of-Text (SOT), โAnchor*โ, End-of- Text (EOT), and Padding tokens. The anchor token uses currentํฃ โ , while other tokens use their fixed predefined vectors. After posi- tional encoding, the sequence is passed through the text encoderโs frozen Transformer layer to produce contextualized embeddingsํ. These embeddings condition the diffusion model when computing L ํ andL ํท , and gradients are backpropagated to update ํฃ โ . After optimization, we obtain the final vectorํฃ โ representing the anchor concept, along with its corresponding contextualized embeddings: ํ โ =ํ โ ํํํ , ํ โ ํํํโํํ , ํ โ ํธํํ , ํ โ ํํํํํํํ (12) whereํ โ ํํํโํํ is the optimized anchor embedding,ํ โ ํํํ ,ํ โ ํธํํ , and ํ โ ํํํํํํํ denote SOT, EOT, and Padding embeddings, respectively. 3.3 Structure-aware Adaptive Replacement During inference, we erase target characters by replacing their corresponding embeddings with the optimized anchor embeddings, ensuring minimal impact on unrelated elements. To further enhance overall scene coherence, we propose a structure-aware adaptive strategy that dynamically introduces the embedding replacement process during image generation. Embedding Replacement. When a prompt withํwords (includ- ingํtarget-related terms that are pre-listed in a subspace associated with copyrighted characters) is fed into the modelโs text encoder, it is tokenized and encoded into contextualized embeddings: ํ=ํ ํํํ , ํ ํค 1 , ..., ํ ํค ํ , ํ ํธํํ , ํ ํํํํํํํ (13) For target erasure, we locate all theํtarget-related embeddings ํ ํกํํํํํก = ํ ํค ํ+1 , ...,ํ ํค ํ+ํ whereํ โ [0,ํ โ ํ], and replace them with the optimized anchor embedding: ํ โฒ ํกํํํํํก =ํ โ ํํํโํํ , ..., ํ โ ํํํโํํ (14) Furthermore, based on findings that special embeddings (EOT and Paddings) also encode meaningful layout and semantic informa- tion [5,46], to enhance anchor concept integration while preserving overall context, we adapt the semantic additivity principle to fuse the special input embeddingsํ ํธํํ ,ํ ํํํํํํํ and optimized em- beddingsํ โ ํธํํ ,ํ โ ํํํํํํํ via element-wise addition [15]: ํ โฒ ํธํํ = ํ 1 ยท ํ ํธํํ + ํ 2 ยท ํ โ ํธํํ (15) ํ โฒ ํํํํํํํ = ํ 1 ยท ํ ํํํํํํํ + ํ 2 ยท ํ โ ํํํํํํํ (16) where ํ 1 and ํ 2 are determined empirically. Final contextualized embeddings for image generation are: ํ โฒ =ํ ํํํ , ํ ํค 1 , ..., ํ โ ํํํโํํ , ..., ํ โฒ ํธํํ , ํ โฒ ํํํํํํํ (17) This embedding-level replacement avoids modifications to the pre- trained text encoderโs parameters. Structure-aware Adaptive Strategy. To enhance image coher- ence, we introduce a structure-aware adaptive replacement strategy. During denoising, we apply low-frequency filterํ ํฟ at each timestep ํกto extract structural componentํ ํฟ (ํฅ ํก )from the predicted sample. We compute ํฟ 2 distance between consecutive component: ํฟ 2 =||ํ ํฟ (ํฅ ํก )โ ํ ํฟ (ํฅ ํกโ1 )|| 2 2 , ํก โ [1, 1000](18) We show a denoising example in Figure 3. In this example,ํฟ 2 increases from initial timestepํก=1000 toํก โ720, reflecting a consistent transformation of the overall image layout. Afterํก= 720,ํฟ 2 begins to decrease, indicating structural stabilization and a shift from overall outline toward finer content refinement. This transition point serves as an optimal starting point for embedding replacement, as the stabilized layout provides a reliable spatial reference frame, allowing targeted modifications to specific regions without affecting the established global structure. M โ26, November 10โ14, 2026, Rio de Janeiro, BrazilQiao Li et al. Donald Duck Bugs Bunny SpongeBob Super MarioPeppa Pig SD v1.4 (Base) SLD strong SAFREE STG TraSCE Ours Figure 4: Qualitative comparison between our method and baseline methods. Our method completely erases target animation concepts with optimized anchor concepts, while better preserving overall image fidelity and background coherence. Based on the above analysis of structural dynamics during denois- ing, we propose an structure-aware adaptive replacement strategy. During inference, our method allows the diffusion model to denoise conditioned on the input prompt while simultaneously tracking the variation of low-frequency structural signals. Embedding replace- ment is triggered automatically when structural change ceases to increase consistently and reaches a predefined threshold. 4 Experiments 4.1 Experimental Setup Baselines. We compare our method with five prompt-based steer- ing baselines, including Safe Latent Diffusion (SLD) [35], STG [29], SAFREE [41], TraSCE [16], and Negative Prompting (NP). We fur- ther evaluate the effectiveness of our optimized anchor embeddings by integrating them into four representative model modification baselines: MACE [26], UCE [9], ESD-u [8], and AC [17]. Dataset. We propose a dataset of 80 animation characters that can be reliably generated by diffusion models. The dataset comprises 36 anthropomorphic characters, 23 animal-form characters, and 21 characters in miscellaneous categories. For evaluation, we generate 100 images for each character using prompts from GPT-4o [1]. Target Models. We utilize the pre-trained Stable Diffusion-v1- 4 [33] for all experiments. In addition, we also conduct experiments on Stable Diffusion-v1-5, Stable Diffusion-v2, Stable Diffusion-v2-1, Stable Diffusion-XL-base-1, and Z-Image [38] (DiT-based model) to demonstrate our methodโs transferability across model archi- tectures. Our method requires no model fine-tuning, so the hyper- parameters of all the pretrained models remain unchanged. Evaluation Metrics. We assess our method from three aspects. For erasure effectiveness of target animation characters, we employ two vision-language models, LLaVA-1.5 [23] and BLIP-3 [25], to identify the presence of targets. A lower identification accuracy in- dicates a more complete erasure. For image fidelity preservation, we employ three metrics: SSIM [40] measures the structural similarity; LPIPS [43] evaluates perceptual difference; Aesthetic Predictor V2 Score (ํดํํ ํกโํํกํํ) [19] assesses the visual appeal of the erased ver- sion. For impact on normal image generation, we generate images from irrelevant prompts in COCO-30K dataset [21], and calculate the Frechet Inception Distance (ํนํผํท) [12] and the CLIP Score [31]. 4.2 Quantitative Comparison Quantitative comparison results are reported in Table 1. Erasure Effectiveness of Target Concept. Compared to all base- lines, images erased using our method achieve the lowest accuracies for successful target identification by both LLaVA-1.5 (6.0%) and BLIP-3 (4.0%), with reductions of 3.5% and 1.7% compared to the sec- ond lowest, respectively. This demonstrates our methodโs erasure effectiveness, as it effectively deceive multi-modal large models. Image Fidelity Preservation. Our method achieves the highest SSIM of 0.467 and the lowest LPIPS of 0.505 between the original and erased image versions. These results demonstrate our methodโs superior ability to maintain the integrity and consistency of unre- lated concepts and background. Besides, our method achieves the highest Aesthetic score of 5.18, indicating that the erased images possess higher artistic quality. This ensures that the images retain high application value even after target erasure. Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic AnchorsMM โ26, November 10โ14, 2026, Rio de Janeiro, Brazil Table 2: Ablation study on key modules of our method.โ denotes the difference compared to our complete method. Method Erasure EffectivenessFidelity Preservation LLaVA-1.5โ (โ)BLIP-3โ (โ)SSIMโ (โ)LPIPSโ (โ) SD v1.4 (Base)66.9%64.7%10 w/o Structural Construction55.0% (+49.0%)46.0% (+42.0%) 0.598 (+0.131)0.384 (-0.121) w/o Detailed Differentiation21.0% (+15.0%)15.0% (+11.0%)0.417 (-0.050)0.653 (+0.148) w/o Adaptive Replacement3.0% (-3.0%)1.7% (-2.3%)0.319 (-0.148)0.668 (+0.163) w/o Special Embeddings Addition12.7% (+6.7%)9.3% (+5.3%)0.596 (+0.129)0.414 (-0.091) Ours6.0%4.0%0.4670.505 Original Ours w/o Structural Construction w/o Detailed Differentiation w/o Adaptive Replacement w/o Special Embedding Additon Kung Fu Panda Spider Man Original Ours Dog Goofy Iron Man Bird Snoopy Snoopy Snoopy Snoopy Directly replace the word โSnoopyโin the input prompt: โA Snoopyisridingabike.โ Figure 5: Visualization of ablation results on key modules. Impact on Unrelated Image Quality. Our method maintains near-identical FID and CLIP Score to the original SD v1.4 (33.7 and 0.326), demonstrating undiminished normal generation capability. 4.3 Qualitative Comparison As shown in Figure 4, our method achieves seamless erasure of animation characters through optimized anchors, while prompt- based baselines often failโparticularly for characters with complex shapes and attributes (e.g., Donald Duck and Super Mario). Besides, our method achieves background consistency without the blurring or warping artifacts, supporting iterative creative workflows. 4.4 Ablation Study We conduct ablation study on our proposed method, and the results are presented in Table 2 and Figure 5. Structural and Detailed Modules. We optimize anchor concepts through joint structural and detailed constraints. To analyze their individual role, we ablate the two modules respectively. Results in Table 2 indicate that both modules are essential: (i) without struc- tural constraints, anchor optimization fails, rendering the replace- ment ineffective; (i) without detailed constraints, anchor concepts closely resemble target concepts, degrading erasure performance. Adaptive Replacement Module. We propose the adaptive re- placement module to better preserve overall structural coherence. To highlight its importance, we present ablation results in Table 2, where target embeddings are replaced directly from the initial de- noising step. While effective target erasure can be achieved, this brings severe structural changes and distortions in the overall im- age, resulting in significantly reduced image fidelity. (a) Original (b) Erased (c) Difference ฮฑ=0.25 ฮฑ=0.5 ฮฑ=0.75 Figure 6: Left: original images (target), erased version (an- chor), and their difference (erased semantic features). Right: results of fine-grained erasure control via ํผ-interpolation. (a). Winnie the Pooh(b). Olaf (c). Minnie Mouse(d). Stitch LLaVA-1.5 AccuracyBLIP-3 Accuracy SSIM Figure 7: Performance variation curves under various erasure degrees. Blue and orange lines represent the identification accuracy by LLaVA-1.5 and BLIP-3, while the red line denotes the SSIM value between the erased and original versions. Special Embeddings Addition. During target replacement, we add the optimized special embeddings (EOT and Paddings) to the original ones. Table 2 reveals that omitting these embeddings re- duces erasure effectiveness, indicating that target-related semantics persist in special embeddings and require explicit handling. Thus, adding the optimized special embeddings facilitates better integra- tion of the anchorโs semantics while erasing the targetโs. Prompt-level Words Modification. While our method works in the embedding space, a naive alternative is prompt-level word re- placement. As shown in the third row of Figure 5, directly replacing words often causes large layout/style changes. Moreover, struc- tural distortion worsens with increasing semantic distance between target and replacement words. Consequently, naive prompt-level modification lacks the fine-grained control necessary for nuanced animation character removal in practical scenarios. 4.5 Fine-grained Control over Erasure Degrees Our method effectively achieves fine-grained control over the era- sure degree of target characters. Since our anchor concepts are constructed in the continuous embedding space, the erasure degree can be precisely modulated through vector arithmetic. For a target M โ26, November 10โ14, 2026, Rio de Janeiro, BrazilQiao Li et al. Original version Erased version Spider Man & Hulk Minnie Mouse & Mickey Mouse Snoopy & Pikachu Ninja Turtle & Squirtle Original version Erased version Figure 8: Simultaneous erasure of multiple characters. Table 3: Transferability of our method across different model versions. Models sharing the same color indicate that they adopt the same text encoder architecture. Model VersionsErasure Effectiveness Fidelity Preservation ModelBackbone Dimension LLaVA-1.5โBLIP-3โSSIMโLPIPSโ SD v1.4UNet7686.0%4.0%0.4670.505 SD v1.5UNet7687.0%6.0%0.5090.487 SD v2UNet10248.7%6.3%0.5980.437 SD v2.1UNet10247.7%5.3%0.5850.413 SDXLUNet768 & 12805.6%4.9%0.4390.517 Z-ImageDiT25607.9%7.1%0.3920.563 concept embeddingํ ํกํํํํํก and its optimized anchorํ ํํํโํํ , Fig- ure 6(a)(b)(c) illustrate the images generated fromํ ํกํํํํํก ,ํ ํํํโํํ , andํ ํกํํํํํก โํ ํํํโํํ , respectively. Figure 6(c) visualizes the semantic features that are erased from the target. Continuous fine-grained control is achieved through: ํ โฒ = ํ ํํํโํํ + ํผ ยท(ํ ํกํํํํํก โํ ํํํโํํ ), ํผ โ [0, 1](19) This formulation interpolates between anchor embeddingํ ํํํโํํ (complete erasure) and target embeddingํ ํกํํํํํก (no erasure), where ํผgoverns the interpolation strength. Increasingํผproduces images closer to the original, and vice versa. Figure 6 shows the generated images for ํผ= 0.25, 0.5, 0.75, respectively. We plot the variation of identification accuracy (LLaVA-1.5 and BLIP-3) and SSIM withํผfor four animation characters in Figure 7. The stable SSIM values confirm that image structure is largely un- affected byํผ. In contrast, identification accuracy differs markedly: Winnie the Pooh and Stitch exhibit a sharp increase within a small ํผinterval, while Olaf and Minnie Mouse improve more gradually. Minnie Mouse attains high accuracy at smallerํผ, likely due to its highly unique and recognizable features, whereas the others require largerํผ(especially Stitch). This enables flexible and user-tailored control over the target erasure degrees for different characters. 4.6 Simultaneous Erasure of Multi-targets Benefiting from the precise localization of target concepts and the feasibility of multi-embeddings replacement, our method can also achieve simultaneous erasure of multiple targets. First, we optimize an anchor embedding for each animation character. Subsequently, we locate all the target-related embeddings and simultaneously re- place them with their corresponding anchor embeddings, ensuring visual coherence. Results are presented in Figure 8. Table 4: Quantitative results of integrating our optimized an- chors into existing model modification baselines, compared with using general anchors (i.e. null text or โtoyโ). MethodErasure Effectiveness Fidelity Preservation Baseline Fine-tuning AnchorLLaVA-1.5โSSIMโ LPIPSโArtโ UCEโ General9.6%0.2480.6984.74 Ours7.2%0.3820.6694.94 ESD-u โ General36.0%0.3150.6884.83 Ours13.6%0.3580.6404.99 ACโ General6.0%0.2680.6984.74 Ours3.2%0.3150.6725.01 MACE โ General9.0%0.2820.7204.76 Ours17.5%0.3020.6774.76 4.7 Transferability across Different Models Our method is transferable across different diffusion models. The optimized anchor embedding enables seamless deployment on any diffusion model sharing the same text encoder architecture (with the same embedding dimension 1รํท), thus avoiding redundant opti- mization. For instance, anchor embeddings optimized in SD v1.4 can be used in SD v1.5 (ํท=768), and those for SD v2 and SD v2.1 can be shared (ํท=1024). Moreover, our method is also applicable to mod- els with dual text encoders, such as SDXL (ํท=768&1280), by jointly optimizing the anchor embeddings in both text encoders. Cross- model results (Table 3) highlights the versatility of our method. Transferability to DiT-based Models. As shown in Table 3, we effectively extend our method to Z-Image [38], a recently released DiT-based model with flow-matching [22, 24] sampling strategy. 4.8 Anchor Integration into Model Modification For each character, our method optimizes its anchor embedding in the text encoder, denoted as โAnchor*โ. While this anchor is primar- ily used for inference-time replacement in our main pipeline, it can also be viewed as a plug-and-play component for existing model modification baselines. Specifically, we integrate our optimized an- chor by simply substituting the baselinesโ original modification targets with โAnchor*โ. As shown in Table 4, compared to using semantically distant general anchors (e.g., null text or โtoyโ), our semantically proximate anchors generally improve erasure effec- tiveness and better preserve image fidelity across most baselines, without altering any other method operations. 5 Conclusion In this paper, we address animation copyright infringement by proposing a controllable method to erase animation characters during diffusion-based image generation. We first optimize an an- chor in the continuous embedding space under structural and de- tail constraints, then replace target-related embeddings with the learned anchor via a structure-aware adaptive strategy. Experi- ments demonstrate our methodโs state-of-the-art erasure accuracy, image fidelity preservation, and support for controllable erasure degree, multi-target removal, and model transferability. We hope our contributions not only help regulators prevent infringement but also maximally preserve usersโ creative intent, facilitating de- ployment of trustworthy and user-centric AI systems. Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic AnchorsMM โ26, November 10โ14, 2026, Rio de Janeiro, Brazil 6 Acknowledgments This work was supported by the National Key Research and Devel- opment Program of China (No.2024YFC3307402). References [1]Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al.2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023). [2]Yogesh Balaji, Seungjun Nah, Xun Huang, Arash Vahdat, Jiaming Song, Qinsheng Zhang, Karsten Kreis, Miika Aittala, Timo Aila, Samuli Laine, et al.2022. ediff-i: Text-to-image diffusion models with an ensemble of expert denoisers. arXiv preprint arXiv:2211.01324 (2022). [3]BBC News. 2025. Disney and Universal sue AI firm Midjourney over images. https://w.bbc.com/news/articles/cg5vjqdm1ypo. Accessed: 2025-07-1. [4] Zachary Bozard. 2023. What does it mean to create art? Intellectual Property rights for Artificial Intelligence generated artworks. SCJ Intโl L. & Bus. 20 (2023), 83. [5]Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020), 1877โ1901. [6] Anh Bui, Long Vuong, Khanh Doan, Trung Le, Paul Montague, Tamas Abraham, and Dinh Phung. 2024. Erasing undesirable concepts in diffusion models with adversarial preservation. arXiv preprint arXiv:2410.15618 (2024). [7] Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems 34 (2021), 8780โ8794. [8]Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. 2023. Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision. 2426โ2436. [9] Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzyลska, and David Bau. 2024. Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. 5111โ5120. [10] Chao Gong, Kai Chen, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang. 2024. Re- liable and efficient concept erasure of text-to-image diffusion models. In European Conference on Computer Vision. Springer, 73โ88. [11]Matt Growcoot. 2022. Midjourney founder admits to using a โhundred mil- lionโimages without consent. PetaPixel, Dec (2022). [12]Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems 30 (2017). [13] Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840โ6851. [14]Jonathan Ho and Tim Salimans. 2022. Classifier-free diffusion guidance. arXiv preprint arXiv:2207.12598 (2022). [15]Taihang Hu, Linxuan Li, Joost Van de Weijer, Hongcheng Gao, Fahad S Khan, Jian Yang, Ming-Ming Cheng, Kai Wang, and Yaxing Wang. 2024. Token merging for training-free semantic binding in text-to-image synthesis. Advances in Neural Information Processing Systems 37 (2024), 137646โ137672. [16]Anubhav Jain, Yuya Kobayashi, Takashi Shibuya, Yuhta Takida, Nasir Memon, Julian Togelius, and Yuki Mitsufuji. 2024. Trasce: Trajectory steering for concept erasure. arXiv preprint arXiv:2412.07658 (2024). [17]Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu. 2023. Ablating concepts in text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision. 22691โ 22702. [18] Mingi Kwon, Jaeseok Jeong, and Youngjung Uh. 2022. Diffusion models already have a semantic latent space. arXiv preprint arXiv:2210.10960 (2022). [19]LAION-AI. 2022. aesthetic-predictor. https://github.com/LAION-AI/aesthetic- predictor. [20]Ouxiang Li, Yuan Wang, Xinting Hu, Houcheng Jiang, Tao Liang, Yanbin Hao, Guojun Ma, and Fuli Feng. 2025. Speed: Scalable, precise, and efficient concept erasure for diffusion models. arXiv preprint arXiv:2503.07392 (2025). [21]Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollรกr, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European conference on computer vision. Springer, 740โ755. [22]Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le. 2022. Flow matching for generative modeling. In The eleventh international conference on learning representations. [23]Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 26296โ26306. [24] Xingchao Liu, Chengyue Gong, and Qiang Liu. 2023. Flow straight and fast: Learn- ing to generate and transfer data with rectified flow. In International conference on learning representations (ICLR). [25]Andong Lu, Wanyu Wang, Chenglong Li, Jin Tang, and Bin Luo. 2024. Rgbt tracking via all-layer multimodal interactions with progressive fusion mamba. arXiv preprint arXiv:2408.08827 (2024). [26]Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. 2024. Mace: Mass concept erasure in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6430โ6440. [27]Yiwei Lu, Matthew YR Yang, Zuoqiu Liu, Gautam Kamath, and Yaoliang Yu. 2024. Disguised copyright infringement of latent diffusion models. arXiv preprint arXiv:2404.06737 (2024). [28]Mengyao Lyu, Yuhong Yang, Haiwen Hong, Hui Chen, Xuan Jin, Yuan He, Hui Xue, Jungong Han, and Guiguang Ding. 2024. One-dimensional adapter to rule them all: Concepts diffusion models and erasing applications. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7559โ7568. [29]Byeonghu Na, Mina Kang, Jiseok Kwak, Minsang Park, Jiwoo Shin, SeJoon Jun, Gayoung Lee, Jin-Hwa Kim, and Il-Chul Moon. 2026. Training-free safe text embedding guidance for text-to-image diffusion models. Advances in Neural Information Processing Systems 38 (2026), 85984โ86014. [30]Ji-Hoon Park, Yeong-Joon Ju, and Seong-Whan Lee. 2024. Explaining generative diffusion models via visual analysis for interpretable decision-making process. Expert Systems with Applications 248 (2024), 123231. [31] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748โ8763. [32] Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. 2022. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125 1, 2 (2022), 3. [33]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bjรถrn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 10684โ10695. [34]Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al.2022. Photorealistic text-to-image diffusion models with deep language understanding. Advances in neural information processing systems 35 (2022), 36479โ36494. [35]Patrick Schramowski, Manuel Brack, Bjรถrn Deiseroth, and Kristian Kersting. 2023. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 22522โ22531. [36]Jiaming Song, Chenlin Meng, and Stefano Ermon. 2020. Denoising diffusion implicit models. arXiv preprint arXiv:2010.02502 (2020). [37] Yang Song and Stefano Ermon. 2019. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems 32 (2019). [38]Z-Image Team. 2025. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer. arXiv preprint arXiv:2511.22699 (2025). [39]Jianyi Wang, Zongsheng Yue, Shangchen Zhou, Kelvin CK Chan, and Chen Change Loy. 2024. Exploiting diffusion prior for real-world image super- resolution. International Journal of Computer Vision 132, 12 (2024), 5929โ5949. [40]Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. 2004. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13, 4 (2004), 600โ612. [41]Jaehong Yoon, Shoubin Yu, Vaidehi Ramesh Patil, Huaxiu Yao, and Mohit Bansal. 2025. Safree: Training-free and adaptive guard for safe text-to-image and video generation. In International Conference on Learning Representations, Vol. 2025. 56439โ56465. [42]Gong Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi. 2024. Forget-me-not: Learning to forget in text-to-image diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 1755โ1764. [43]Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition. 586โ595. [44]Tong Zhang, Ru Zhang, Jianyi Liu, Zhen Yang, and Gongshen Liu. 2025. Beyond Fixed Anchors: Precisely Erasing Concepts with Sibling Exclusive Counterparts. arXiv preprint arXiv:2510.16342 (2025). [45] Yang Zhang, Teoh Tze Tzun, Lim Wei Hern, and Kenji Kawaguchi. 2024. On copyright risks of text-to-image diffusion models. In ECCV 2024 Workshop The Dark Side of Generative AIs and Beyond. [46] Chenyi Zhuang, Ying Hu, and Pan Gao. 2024. Magnet: We never know how text-to-image diffusion models work, until we learn how vision-language models function. Advances in Neural Information Processing Systems 37 (2024), 57115โ 57149.