Paper deep dive
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
Sathwik Karnik, Juyeop Kim, Sanmi Koyejo, Jong-Seok Lee, Somil Bansal
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 2:13:35 PM
Summary
The paper introduces Reachability-Aware Diffusion Steering (RADS), an inference-time framework that mitigates memorization in text-to-image diffusion models. RADS models the denoising process as a controlled dynamical system and uses reachability analysis to identify a 'backward reachable tube' (BRT) of states leading to memorized outputs. A constrained reinforcement learning policy then steers the caption embeddings away from this tube, preserving image quality and prompt alignment without modifying the model backbone.
Entities (8)
Relation Signals (6)
RADS → mitigates → Memorization
confidence 95% · RADS... prevents memorization while preserving generation fidelity.
RADS → uses → Reachability Analysis
confidence 92% · RADS models the diffusion denoising process as a dynamical system and applies concepts from reachability analysis
RADS → uses → Reinforcement Learning
confidence 92% · formulate mitigation as a constrained reinforcement learning (RL) problem
RADS → appliesto → Stable Diffusion v1.4
confidence 90% · This diagram illustrates a real example of how rads prevents memorization in the Stable Diffusion v1.4 model
RADS → steers → Caption Embedding Space
confidence 90% · steer the trajectory away from memorization via minimal perturbations in the caption embedding space.
RADS → outperforms → State-of-the-art baselines
confidence 88% · RADS achieves a superior Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state-of-the-art baselines.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Text-to-image diffusion models often memorize training data, revealing a fundamental failure to generalize beyond the training set. Current mitigation strategies typically sacrifice image quality or prompt alignment to reduce memorization. To address this, we propose Reachability-Aware Diffusion Steering (RADS), an inference-time framework that prevents memorization while preserving generation fidelity. RADS models the diffusion denoising process as a dynamical system and applies concepts from reachability analysis to approximate the "backward reachable tube"--the set of intermediate states that inevitably evolve into memorized samples. We then formulate mitigation as a constrained reinforcement learning (RL) problem, where a policy learns to steer the trajectory away from memorization via minimal perturbations in the caption embedding space. Empirical evaluations show that RADS achieves a superior Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state-of-the-art baselines. Crucially, RADS provides robust mitigation without modifying the diffusion backbone, offering a plug-and-play solution for safe generation. Our website is available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.00140v1
- Canonical: https://arxiv.org/abs/2603.00140v1
Trouble viewing inline? Open PDF directly →
Full Text
66,685 characters extracted from source content.
Expand or collapse full text
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion Sathwik Karnik Juyeop Kim Sanmi Koyejo Jong-Seok Lee Somil Bansal Abstract Text-to-image diffusion models often memorize training data, revealing a fundamental failure to generalize beyond the training set. Current mitigation strategies typically sacrifice image quality or prompt alignment to reduce memorization. To address this, we propose Reachability-Aware Diffusion Steering (rads), an inference-time framework that prevents memorization while preserving generation fidelity. rads models the diffusion denoising process as a dynamical system and applies concepts from reachability analysis to approximate the “backward reachable tube”—the set of intermediate states that inevitably evolve into memorized samples. We then formulate mitigation as a constrained reinforcement learning (RL) problem, where a policy learns to steer the trajectory away from memorization via minimal perturbations in the caption embedding space. Empirical evaluations show that rads achieves a superior Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state-of-the-art baselines. Crucially, rads provides robust mitigation without modifying the diffusion backbone, offering a plug-and-play solution for safe generation. Our website is available at: https://s-karnik.github.io/rads-memorization-project-page/. Machine Learning, ICML 1 Introduction Diffusion models have undoubtedly emerged as the dominant paradigm for image generation in recent years (Ho et al., 2020; Song et al., 2021a, b; Dhariwal and Nichol, 2021; Rombach et al., 2022; Ramesh et al., 2022). However, like other generative models (Webster et al., 2021; Carlini et al., 2023a; Jagielski et al., 2023; Lee et al., 2023), diffusion models are susceptible to reproducing training data (Somepalli et al., 2023a; Carlini et al., 2023b; Webster, 2023). This phenomenon, namely, memorization, is particularly concerning in the context of text-to-image generation, where natural language prompts can trigger the extraction or faithful reproduction of copyrighted or private images (Carlini et al., 2022; Jiang et al., 2023). To address this problem, prior work has proposed a range of mitigation strategies (Wen et al., 2024; Ren et al., 2024; Hintersdorf et al., 2024; Jain et al., 2025). Although these methods can reduce direct reproduction of training images, they often do so at the expense of image quality or alignment with the user’s generation intent. This trade-off is illustrated in Figure 1. Figure 1(a) shows a memorized image generated by Stable Diffusion v1.4 (Rombach et al., 2022), alongside outputs produced by representative prior mitigation methods in Figure 1(b). As shown, some methods reduce memorization but yield low-quality images (Figure 1(b)4). Others preserve visual quality, but fail to capture key semantic details specified in the prompt (Figure 1(b)2; e.g., ‘Red Sky’ or ‘Glossy Cityscape’ are not clearly reflected in the output). Finally, some approaches do not sufficiently mitigate memorization, instead producing images with high similarity to the training example (Figures 1(b)1, 3). (a) Memorized (b) Prior work (c) rads (Ours) Figure 1: rads outperforms prior work. Generated images for the prompt Design Art Beautiful View of Paris Paris Eiffel Towerunder Red Sky Ultra Glossy Cityscape Circle Wall Art. (a) Memorized target image from the training set. (b) Mitigated results produced by prior methods; 1–4 correspond to Wen et al. (2024); Ren et al. (2024); Hintersdorf et al. (2024); Jain et al. (2025), respectively. (c) Mitigated result produced by rads (ours). Figure 2: Overview of rads. This diagram illustrates a real example of how rads prevents memorization in the Stable Diffusion v1.4 model, while the baseline (no mitigation) generates an image that closely resembles the training sample. rads does so by modeling the “attraction basin” of memorization as a backward reachable tube (BRT) and learning a policy πϕ _φ for steering the caption embedding inputs to the diffusion model using reachability-constrained reinforcement learning. These limitations raise a fundamental question: how can we mitigate memorization without sacrificing image quality or alignment with the generation intent? To answer this question, we introduce Reachability-Aware Diffusion Steering (rads), an inference-time framework to actively predict and prevent memorization (see Figure 2). Our core technique is reachability analysis (Bansal et al., 2017), a method from control theory traditionally used to ensure safety in autonomous systems by identifying a system’s “failure set.” Reachability analysis characterizes the “backward reachable tube” (BRT)—the set of all intermediate states from which a system will inevitably evolve into that failure set under its natural dynamics. We adapt this concept to diffusion models by modeling the denoising process as a controlled dynamical system, where the latent noise represents the system state and the perturbed caption embeddings act as the control input. By computing the BRT, we can identify the specific regions of the latent space where the denoising trajectory will inevitably collapse into a memorized image. Building on this characterization, we formulate memorization mitigation as a constrained reinforcement learning (RL) problem. In this framework, the RL policy aims to maximize a reward signal based on prompt fidelity and perceptual quality, while being constrained to avoid states within the BRT. By solving this constrained optimization at inference-time, rads learns to steer the diffusion trajectory away from memorization-inducing regions via perturbations in the caption embedding space. To the best of our knowledge, rads is the first framework to combine reachability analysis and RL to steer diffusion models away from memorization basins while preserving semantic alignment. To summarize, our contributions are: (1) a reachability-theoretic formulation of the diffusion denoising process, modeling the latent states and caption embeddings as a controlled dynamical system; (2) a reachability-constrained RL algorithm for inference-time steering that learns to prevent memorization; (3) comprehensive experiments across multiple open-source diffusion models and datasets, demonstrating that rads achieves a superior Pareto frontier between generation diversity, perceptual quality, and prompt alignment compared to existing state-of-the-art baselines. 2 Related Work Mitigating Memorization. Prior work primarily addresses memorization through heuristic interventions on prompts or attention mechanisms. Somepalli et al. (2023b) propose perturbing input prompts with random tokens to disrupt retrieval. Others focus on the diffusion architecture itself: Ren et al. (2024) and Hintersdorf et al. (2024) mitigate memorization by masking cross-attention scores or disabling specific neurons responsible for reproducing training data. More recently, methods utilizing the classifier-free guidance magnitude have emerged; Wen et al. (2024) replace tokens that trigger high guidance norms, while Jain et al. (2025) truncate guidance entirely when an “attraction basin” is detected. Unlike these methods, which rely on discrete, heuristic-driven interventions that often disrupt the generation process, rads provides a principled, continuous control framework. By modeling the denoising path as a controllable trajectory, we can steer away from memorization with minimal impact on the overall image structure. Model Unlearning vs. Inference Steering. A distinct class of mitigation involves permanently modifying model weights to “unlearn” concepts. Techniques such as Erasing Stable Diffusion (Gandikota et al., 2023) or LoRA-based unlearning adapters (Kumari et al., 2023; Liu et al., 2025) fine-tune the backbone to suppress specific targets. However, memorization often does not align with fixed, high-level concepts; it frequently involves the replication of idiosyncratic training instances that are difficult to define as a distinct “concept” for unlearning. Furthermore, these methods are destructive—often degrading general capabilities via catastrophic forgetting—and require retraining for every new concept of prompts leading to memorization. In contrast, rads steers away from memorization attraction basins regardless of the specific content, enabling robust mitigation without defining semantic targets or modifying weights. Reachability in Safety-Critical Systems. We adapt reachability analysis (Bansal and Tomlin, 2021) to diffusion models. Building on prior work applying latent reachability to large language models (Karnik and Bansal, 2025), we present rads as the first application of latent reachability to diffusion. By modeling the “backward reachable tube” of memorization, our framework enables preemptive intervention before the generation collapses into a memorized image. Inference-Time Steering via RL. While RL is commonly used to fine-tune diffusion weights (Black et al., 2023), recent work explores inference-time steering to guide generation without retraining. Similar to PPLM (Dathathri et al., 2019) for LLMs, Wagenmaker et al. (2025) propose DSRL to steer diffusion outputs toward external task rewards (e.g., robotic goals). rads differs by formulating the internal denoising process as the dynamical system, applying control directly to the evolving latents to enforce safety constraints. 3 Preliminaries: Denoising Diffusion Diffusion models commonly operate in a low-dimensional latent space X learned by a well-trained autoencoder, which reduces computational cost compared to operating directly in the high-dimensional pixel space Y (Rombach et al., 2022). Specifically, instead of denoising directly in pixel space to produce an image ∈y , diffusion models perform a gradual denoising process in latent space. Starting from a random noise latent T∈x_T sampled from (,)N(0,I), the model iteratively denoises over T steps (from timestep τ=Tτ=T to τ=0τ=0) to obtain 0x_0, which is then decoded by the autoencoder to produce the final image y. At each timestep τ, denoising is performed by predicting the noise component to be removed from the intermediate latent τx_τ. This prediction is made by a trained noise predictor ϵθ _θ (Ho et al., 2020). To enable conditional generation using text caption c, ϵθ _θ takes both τx_τ and a caption embedding ce_c as input, derived from c using CLIP (Radford et al., 2021). The text embedding ce_c is injected into the model through cross-attention layers to guide the denoising process (Rombach et al., 2022). The strength of text guidance is controlled by a guidance scale g (Ho and Salimans, 2021). 4 Reachability-Aware Diffusion Steering To mitigate memorization, we propose Reachability-Aware Diffusion Steering (rads). We view the autonomous denoising process as a controlled dynamical system and aim to actively steer the generation away from memorized outputs by introducing a control input tu_t. 4.1 Denoising as a Controlled Dynamical System Since diffusion generation proceeds through sequential latent updates τ→τ−1x_τ _τ-1 governed by a fixed update rule, it can naturally be viewed as a dynamical system (e.g. via the lens of Langevin dynamics (Ho et al., 2020)). In general, a dynamical system evolves as t+1=f(t,t,ωt)s_t+1=f(s_t,u_t, _t) (from 0→T0→ T), where t∈s_t is the system state at time t, t∈u_t is the control input, ωt _t is the noise (e.g., sampled noise in DDPM (Ho et al., 2020)), and f is the update rule. In this work, we denote the state t:=(T−t,T−t)s_t:=(x_T-t,T-t), and we apply control inputs tu_t for perturbing the base caption embedding ce_c to obtain the transient steered embedding c′e_c . We denote the pre-trained diffusion update rule as T−(t+1)=fDM(T−t,c′,T−t)x_T-(t+1)=f_DM(x_T-t,e_c ,T-t). The complete system evolution is defined as: t+1=f(t,t,ωt)=[T−(t+1)T−(t+1)]=[fDM(T−t,c′,T−t)T−(t+1)]s_t+1=f(s_t,u_t, _t)= bmatrixx_T-(t+1)\\ T-(t+1) bmatrix= bmatrixf_DM(x_T-t,e_c ,T-t)\\ T-(t+1) bmatrix (1) Remark: Note that in DDIM sampling (Song et al., 2021a) ωt=0 _t=0, while in DDPM sampling (Ho et al., 2020) ωt≠0 _t≠ 0. Choice of Action Space. Steering can be applied in either the image latent space X or the caption embedding space (modifying ce_c). We select the caption space because memorization dominates the image latents within the first few steps, rendering direct latent steering ineffective (Kim et al., 2025). Figure 3 illustrates this phenomena: applying guidance for only the first two steps is sufficient to regenerate the memorized image, even if guidance is disabled thereafter. Thus, rads applies control to the caption embeddings to steer the trajectory before memorization becomes inevitable. (a) 0 steps (b) 1 step (c) 2 steps (d) Memorized Figure 3: Memorization happens early. Images generated with text guidance enabled for only the first k steps (k∈0,1,2k∈\0,1,2\). Just 2 steps of guidance (c) are sufficient to reproduce the memorized image (d). Caption: The No Limits Business Woman Podcast. Controlled Dynamical System for Diffusion Denoising. Standard text embeddings (e.g., CLIP) are high-dimensional (e.g., 77×76877× 768), resulting in a prohibitively large action space for steering. To enable efficient steering, we learn a compact latent action space act⊆ℝdZ_act ^d (d≪dim(c)d (e_c)) using a variational autoencoder (VAE). For the caption space C in which valid caption embeddings reside, we define an encoder Enc:→act Enc:C _act and a decoder Dec:act→ Dec:Z_act . Using this parameterization, we project the caption embedding ce_c into actZ_act to obtain a base latent representation Enc(c) Enc(e_c), apply a perturbation in this projected space as Enc(c)+t Enc(e_c)+u_t, and then decode it back into C to obtain the steered caption embedding c′=Dec(Enc(c)+t)e_c = Dec( Enc(e_c)+u_t). In conclusion, the steered evolution of the denoising process follows the update: T−(t+1)=fDM(T−t,Dec(Enc(c)+t)⏟Steered caption embedding c′,T−t). _T-(t+1)=f_DM (x_T-t, Dec( Enc(e_c)+u_t)_Steered caption embedding e_c ,T-t ). (2) 4.2 Memorization as a Reachability Problem Diffusion-based image generation is inherently path-dependent: early latent states strongly bias the final output (Choi et al., 2022; Daras and Dimakis, 2022; Park et al., 2023; Jain et al., 2025; Kim et al., 2025). In particular, if the denoising trajectory enters a “basin of attraction” corresponding to a memorized training example early in the denoising process, it becomes increasingly difficult – and often impossible – to recover and avoid memorization (Jain et al., 2025). This observation motivates the need for a framework that can identify and avoid such inescapable regions before the image is fully generated. To model these inescapable regions, we adopt reachability analysis, a tool from control theory traditionally used to reason about safety in dynamical systems. Reachability analysis addresses the following question: given a set of failure states, which initial states will inevitably evolve into this failure set under the system dynamics, regardless of the applied control? Failure Set and Backward Reachable Tube. In the context of memorization in diffusion models, we define the failure set ℱ⊂F as the subset of states that decode to images closely resembling those present in the training set. A central concept in reachability analysis is the backward reachable tube (BRT), defined as: ℬ=0∈∣∀∈,∃σ∈[0,T] such that σ∈ℱ, =\s_0 ,∃σ∈[0,T] such that s_σ \, (3) which characterizes all initial states 0s_0 from which memorization is inevitable, independent of the control strategy. Intuitively, once the diffusion trajectory enters the BRT, no admissible steering can prevent memorization. Target Function and Worst-Case Evolution. To compute the BRT, we define a target function ℓ:→ℝ :S whose sub-zero level set corresponds to the failure set, i.e., ℱ=:ℓ()≤0F=\s: (s)≤ 0\. For a latent trajectory starting at time t, we define the cost J(t)=minσ∈[t,T]ℓ(σ)J(s_t)= _σ∈[t,T] (s_σ) which evaluates whether the trajectory enters the failure set within the remaining horizon [t,T][t,T]. The associated policy-conditioned value function captures the worst-case evolution of the denoising process under control: Qπϕsafe(t,t)=minℓ(st+1),maxt+1∼πϕQπϕ(t+1,t+1).Q_ _φ^safe(s_t,u_t)= \ (s_t+1), _u_t+1 _φQ_ _φ(s_t+1,u_t+1)\. (4) This value function measures the degree to which a state lies within or near the backward reachable tube. In Section 4.4, we describe how we train Qπϕsafe(t,t)Q_ _φ^safe(s_t,u_t) and enforce it as a constraint when training the steering policy πϕ _φ. In doing so, reachability analysis provides a principled mechanism for identifying and avoiding memorization-inducing regions of the diffusion trajectory. Safety Target Function. In practice, we define the target function ℓ to detect and avoid memorization. To do so, we utilize the magnitude of the classifier-free guidance vector. Prior work shows that memorized generations often exhibit anomalously high guidance magnitudes, as the model overfits to the conditioning caption (Jain et al., 2025). Thus, we instantiate the target function ℓ(⋅) (·) to penalize extreme deviations between conditional and unconditional predictions: ℓ(t)=−tanh(η⋅(‖ϵθ(T−t,c′)−ϵθ(T−t,∅)‖2−β)). (s_t)=-tanh (η·(\| _θ(x_T-t,e_c )- _θ(x_T-t, )\|_2-β) ). (5) Here, η and β are hyperparameters chosen to separate high guidance magnitudes associated with memorization from lower magnitudes corresponding to novel generations. Appendix A.2.3 details the choice of η and β. The failure set ℱF is defined as the set of states for which ℓ(t)≤0 (s_t)≤ 0. 4.3 Constrained Markov Decision Process Having defined memorization-unsafe states via reachability, we now formulate memorization mitigation as a constrained Markov Decision Process (CMDP). The objective is to steer the diffusion denoising process during inference-time to maximize semantic alignment with a given text prompt while ensuring that the generation remains within a safe regime, defined by low memorization risk as characterized by reachability analysis. MDP Formulation. We model the denoising process as a finite-horizon MDP ℳ=(,,P,r,T)M=(S,A,P,r,T). Here, the state space remains S, the action space A is identically equal to U, transition dynamics P are the same as in Section 4.1. Constraint. In our CMDP, we impose a constraint that requires the steering policy to maintain the denoising trajectory outside the BRT by ensuring that Qsafe≥δQ^safe≥δ. Reward Function. The primary task objective is to preserve semantic alignment between the generated image and the given text prompt. Let clipim() clip_im(y) denote the normalized CLIP embedding of the decoded image y, and let cliptext clip_text denote the normalized CLIP embedding of the caption c. Then, clipim()⋅cliptext(c) clip_im(y)· clip_text(c) denotes the cosine similarity between the image and caption embeddings. We define a sparse terminal reward based on this cosine similarity: r(t)=clipim(t)⋅cliptext(c)if t=T0if t<Tr(s_t)= cases clip_im(y_t)· clip_text(c)&if t=T\\ 0&if t<T cases (6) This formulation encourages generations that remain faithful to the prompt without shaping intermediate denoising steps. 4.4 Constrained RL Solution (Soft Actor-Critic) To solve the CMDP, we employ a constrained Soft Actor-Critic (SAC) algorithm with Lagrangian relaxation (Haarnoja et al., 2018). This approach enables sample-efficient, off-policy learning of a stochastic steering policy πϕ _φ while enforcing the reachability-based safety constraint. Architecture. Our framework maintains three parameterized networks: 1. A stochastic policy πϕ(∣) _φ(u ) with parameters ϕφ. 2. A task critic Qωtask(,)Q^task_ω(s,u) with parameters ω, estimating expected semantic-alignment return. 3. A safety critic Qψsafe(,)Q^safe_ψ(s,u) with parameters ψ, estimating future reachability with respect to memorization. Additionally, we use a learnable temperature parameter α to scale the policy entropy ℋ(π)H(π), encouraging exploration. Safety Critic Learning (Reachability). The safety critic QψsafeQ^safe_ψ approximates the future reachability value. Let ℓ(t) (s_t) denote the immediate target function (defined in Section 4.2) and γ be the discount factor. We estimate the minimum safety margin over the trajectory. Given a transition (t,t,t+1)(s_t,u_t,s_t+1) and a time horizon T, the target value Q^tsafe Q^safe_t is computed recursively: Q^tsafe=(1−γ)ℓ(t)+γmin(ℓ(t),Qψ′safe(t+1,t+1′))if t<Tℓ(T)if t=T Q^safe_t= cases aligned (1-γ) (s_t)+γ (& (s_t),\\ &Q^safe_ψ (s_t+1,u _t+1) ) aligned&if t<T\\ (s_T)&if t=T cases (7) where t+1′∼πϕ(⋅|t+1)u _t+1 _φ(·|s_t+1) is the next action, and ψ′ψ denotes the parameters of the target safety critic network (maintained via Polyak averaging). The parameters ψ are updated to minimize the mean squared error (Qψsafe(t,t)−Q^tsafe)2(Q^safe_ψ(s_t,u_t)- Q^safe_t)^2. Task Critic Learning. The task critic QωtaskQ^task_ω is trained to estimate the cumulative task reward. Given the immediate task reward rtr_t, we minimize the standard Bellman error: ℒtask(ω)=[(Qωtask(t,t)−(rt+γt+1[Vtask(t+1)]))2] _task(ω)=E [ (Q^task_ω(s_t,u_t)- (r_t\\ + _s_t+1[V^task(s_t+1)] ) )^2 ] (8) where VtaskV^task represents the value of the next state under the current policy. Policy and Dual Update. We enforce the reachability constraint by requiring the expected safety value to meet a threshold δ. We introduce a learnable Lagrange multiplier λ≥0λ≥ 0 to handle this constraint. The optimization objective is defined over a distribution of states sampled from the replay buffer D: Jπ(ϕ)=t∼,t∼πϕ[Qωtask(t,t)+αℋ(πϕ(⋅|t))+λ⋅Qψsafe(t,t)]J_π(φ)=E_s_t ,u_t _φ [Q^task_ω(s_t,u_t)+ ( _φ(·|s_t))+λ· Q^safe_ψ(s_t,u_t) ] (9) The Lagrange multiplier λ is updated via dual gradient descent with gradient ∇λJ=[Qψsafe(,)]−δ _λJ=E[Q^safe_ψ(s,u)]-δ. If the predicted safety margin falls below δ, λ increases, penalizing unsafe actions and biasing the policy toward safer steering. 5 Experiments 5.1 Experiment Setup Models and Datasets. We evaluate two diffusion models: (i) Stable Diffusion (SD) v1.4 (Rombach et al., 2022) in both the DDIM (Song et al., 2021a) and DDPM (Ho et al., 2020) sampling settings; and (i) RealisticVision (CivitAI, 2023) with DDIM sampling. SD v1.4 is run in float16, while RealisticVision is run using float32. For the sake of brevity, we report results for SD v1.4 in the main paper and defer results for the remaining models to Appendix A.3. To reproduce memorized samples and conduct mitigation experiments, we first utilize the dataset of 500 memorized prompts identified by Webster (2023) for SD v1.4. From this set, we construct our RL training split using the 430 prompts for which the corresponding target images are publicly available. The remaining 70 prompts are strictly held out to evaluate to unseen prompts (see Section 6.5). Additionally, for zero-shot out-of-distribution evaluation, we utilize the MemBench dataset of 3000 memorized prompts identified by Hong et al. (2024) for SD v1.4. For generation, we set the number of steps to T=50T=50, the guidance scale to g=7.5g=7.5, and the initializations per prompt to N=10N=10. For rads, we report results for 5 train seeds. Baselines. We compare our method against four inference-time memorization mitigation strategies reviewed in Section 2: Wen et al. (2024); Ren et al. (2024); Hintersdorf et al. (2024); Jain et al. (2025). While some of these works also propose training-time mitigation approaches, we exclude such methods from our comparison, as they require full retraining of the diffusion model. Importantly, our rads-trained policy likewise operates entirely at inference-time and does not require retraining the diffusion model. 5.2 Implementation of rads To efficiently steer the high-dimensional CLIP text embeddings, we train a custom VAE that compresses the input into a 64-dimensional latent action space. The policy and critic networks are parameterized as lightweight MLPs trained via SAC (Haarnoja et al., 2018), utilizing automatic entropy tuning to ensure stable exploration. For complete details on the experimental setup and training configurations, please refer to Appendix A.2. 5.3 Evaluation Metrics We measure the effectiveness of rads using the following three complementary metrics: • SSCD (Diversity): To measure the degree of memorization in generated images, we utilize Self-Supervised Copy Detection (SSCD) (Pizzi et al., 2022) to quantify visual similarity in two settings. SSCD has been reported to be one of the strongest metrics for replication (Somepalli et al., 2023a). We compute SSCDtargetSSCD_target (greatest pairwise similarity with the target images), SSCDseedsSSCD_seeds (similarity across random seeds for a fixed prompt) to measure replication, and SSCDpromptsSSCD_prompts (similarity across prompts for a fixed seed) to detect mode collapse. For the SSCDtargetSSCD_target, we evaluate on the 430 prompts from Webster (2023) for which the ground-truth training images are publicly available. Lower scores indicate greater diversity. • FID (Quality): We assess perceptual quality via Fréchet Inception Distance (FID) (Heusel et al., 2017) relative to the COCO validation set (Lin et al., 2014), where lower scores denote more natural images. • CLIP (Alignment): We measure semantic consistency using the CLIP score (Radford et al., 2021) (cosine similarity between image and text embeddings). Higher scores indicate better adherence to the prompt. Additionally, we evaluate the computational efficiency of our approach. rads introduces minimal inference-time overhead, maintaining a generation speed comparable to the fastest, high-fidelity baselines, while Hintersdorf et al. (2024) is significantly slower (see Appendix A.6). 6 Results We evaluate rads to assess its ability to mitigate memorization while preserving image quality and prompt alignment. As summarized in Figure 4, rads achieves a superior Pareto frontier compared to existing baselines, maintaining comparable CLIP score and FID to existing methods, while significantly reducing the replication rate. We use “Pareto frontier” to refer to the tradeoff between replication reduction (SSCD) and image fidelity/diversity; CLIP can be inflated by memorization, so small CLIP drops are expected when moving away from replication. Full quantitative results are provided for the Webster (Webster, 2023) and the MemBench (Hong et al., 2024)) datasets in Appendix A.3. Figure 4: Pareto Frontier Analysis: Quality and Alignment vs. Replication. We compare rads (Ours) against various state-of-the-art mitigation methods on the Webster (2023) dataset. The top row shows Quality (−log10(FID)↑- _10(FID) ), while the bottom row displays Alignment (CLIP Score ↑ ). The x-axis shows the (1−SSCDtarget)↑(1-SSCD_target) scores. rads consistently occupies the upper-right region of the frontier, maintaining high utility and semantic alignment while significantly reducing memorization. 6.1 How diverse are images generated by rads? For a mitigation strategy to be considered successful, the generated outputs for a given prompt across different random seeds (i.e., different samples of T∼(,)x_T (0,I)) should be diverse, indicating that the model no longer reproduces the same memorized data. Figure 7 compares generation results for four different initial latents with and without mitigation strategies, including prior methods (Wen et al., 2024; Ren et al., 2024; Hintersdorf et al., 2024; Jain et al., 2025) and our method (rads). As shown, rads produces diverse images across different random seeds (Figure 7(f)), exhibiting substantial variation in style and composition. We provide a detailed validation of this steering mechanism, including an analysis of classifier-free guidance trajectories, in Appendix A.5. In contrast, the outputs of prior methods tend to resemble one another across different seeds (Figures 7(c), 7(d)) and even across different strategies (Figures 7(b)–7(d)), or suffer from reduced image quality (Figure 7(e)). This observation is also supported by both the SSCDtargetSSCD_target and SSCDseedsSSCD_seeds columns in Table 1(a). Excluding Jain et al. (2025) (due to its severely degraded image quality, evidenced by a high FID of 63.98 ± 16.17 in Table 1(a)), rads achieves the lowest SSCD scores of SSCDtarget=SSCD_target= 0.2303 ± 0.1110 and SSCDseeds=SSCD_seeds= 0.1553 ± 0.1099, outperforming the next best baseline, Wen et al. (2024) (0.2132 ± 0.1798). Compared to other strategies, Jain et al. (2025) not only generates low-quality images, but it also exhibits notably high SSCD values in the SSCDpromptsSSCD_prompts (0.2724 ± 0.1352 for Jain et al. (2025) vs. 0.0409 ± 0.0227 for rads). This behavior arises because Jain et al. (2025) operates by turning text guidance on and off, causing the initial latent to largely determine the generated output. As a result, images generated from the same Tx_T tend to have a similar style across different prompts (Figure 5). (a) (b) (c) (d) Figure 5: Jain et al. (2025) produces mitigated samples that closely resemble one another. Generated images using the same initial latent Tx_T. (a) Generated image without text guidance (g=0g=0). (b–d) Generated images produced using different prompts. 6.2 How does rads perform on challenging, memorization-prone prompts? We identify a subset of challenging prompts from Webster (2023)—typically containing specific entities (e.g., “Bloodborne”)—where standard mitigation strategies struggle. Figure 8 illustrates the failure modes of prior methods on one such prompt. Wen et al. (2024) and Hintersdorf et al. (2024) exhibit complete failure, reproducing the memorized training data across all samples (Figures 8(b), 8(d)). Ren et al. (2024) exhibits stochastic failure: while it successfully steers certain random seeds (top row, Figure 8(c)), it fails to mitigate memorization in others (bottom row), indicating that the method is sensitive to initialization Tx_T. Jain et al. (2025) mitigates replication but suffers from severe degradation, rendering the output unrecognizable (Figure 8(e)). In contrast, rads demonstrates consistent mitigation (Figure 8(f)). Regardless of the initial noise Tx_T, our method successfully prevents the generation of the memorized instance while preserving the semantic constraints of the prompt (e.g., the dark, atmospheric style). This indicates that rads is robust to initialization even in regions of the latent space with seemingly strong memorization basins. 6.3 How good are the images generated by rads? In addition to memorization mitigation, the quality of the generated images is critical. Figure 9 (see Appendix A.1) compares the image quality achieved by prior methods and rads. Qualitatively, rads produces high-fidelity images with natural details (Figure 9(f)). This is supported quantitatively by the FID results in Table 1(a), where rads achieves a mean FID of 31.57 ± 5.82. Given the overlapping standard deviations, this performance is statistically indistinguishable from the strongest baselines (Wen et al., 2024; Ren et al., 2024), demonstrating that rads prevents memorization without the quality trade-offs often associated with steering. Crucially, rads avoids the severe degradation observed in methods like Jain et al. (2025) (63.98 ± 16.17) and, in some cases, improves upon the unmitigated baseline. Figure 6: Zero-Shot Generalization. We compare rads (Ours) in a zero-shot evaluation against state-of-the-art mitigation methods on the MemBench dataset (Hong et al., 2024). The top row shows Quality (−log10(FID)↑- _10(FID) ), while the bottom row displays Alignment (CLIP Score ↑ ). The x-axis shows the (1−SSCDtarget)↑(1-SSCD_target) scores. rads consistently occupies the upper-right region of the frontier, maintaining high utility and semantic alignment while significantly reducing memorization. Note: The Hintersdorf et al. (2024) baseline is omitted due to its computational infeasibility (see Appendix A.6). 6.4 How well are rads-generated images aligned with the prompts? (a) Memorized (b) Wen et al. (2024) (c) Ren et al. (2024) (d) Hintersdorf et al. (2024) (e) Jain et al. (2025) (f) rads (Ours) Figure 7: rads achieves mitigation with the highest generation diversity. Generated images for the prompt Michael Fassbender to Star In <i>Assassin’s Creed</i> Movie. (a) Memorized (b) Wen et al. (2024) (c) Ren et al. (2024) (d) Hintersdorf et al. (2024) (e) Jain et al. (2025) (f) rads (Ours) Figure 8: rads achieves mitigation on challenging prompts. Generated images for the prompt <em>Bloodborne</em> Video: Sony Explains the Game’s Procedurally Generated Dungeons. In the Webster (2023) dataset, rads achieves a CLIP score of 0.2917 ± 0.0366, which is comparable to the unmitigated baseline (0.3129 ± 0.0279) given the significant overlap in their standard deviations. This indicates that rads preserves the semantic alignment of the original model, unlike methods such as Jain et al. (2025) (0.2266 ± 0.0532), which exhibit a statistically significant degradation in alignment. We argue that the maximal CLIP scores observed in unmitigated baselines do not reflect superior semantic understanding, but rather memorization bias, where the model achieves high alignment scores primarily by reproducing the exact training data it has overfitted to. Consequently, any effective mitigation policy that steers away from the memorization manifold necessarily diverges from this overfitted maximum. The marginal reduction in CLIP scores for rads (from ∼ 0.30 to ∼ 0.25) represents the removal of this bias, resulting in generation that preserves generalized semantic intent rather than instance-specific replication. Qualitatively, Figure 10 (see Appendix A.1) confirms that rads can succeed where prior methods may fail. While some baselines miss key semantic elements (e.g., Figure 10(b); ‘Dinner With Friends’ is not reflected) or produce severe artifacts (Figure 10(e)), rads generates a high-quality image that clearly reflects the prompt’s intent (Figure 10(f)). 6.5 How well does rads generalize to unseen prompts? To evaluate if our steering policy—trained on a limited set of 430 prompts—learns a robust mitigation strategy, we assess zero-shot generalization on two unseen datasets. First, on the Webster (2023) validation set (Table 1(b)), rads reduces SSCDseedsSSCD_seeds from 0.5730 (unmitigated) to 0.1356, effectively halving the replication rate of the strongest high-fidelity baseline, Wen et al. (2024) (0.2690). We further scrutinize this generalization on the 3,000 out-of-distribution prompts in MemBench (Table 1(c)). Figure 6 illustrates the Pareto frontier in the MemBench dataset. Once again, rads occupies the upper right of the Pareto frontier, providing a superior balance between generation diversity and image fidelity over other baselines. Remarkably, despite the limited training data, rads achieves the lowest target similarity (SSCDtarget≈0.145SSCD_target≈ 0.145) among all evaluated methods. This outperforms Jain et al. (2025) (SSCDtarget≈0.178SSCD_target≈ 0.178), confirming that rads actively steers the generation away from the training manifold. In contrast, Jain et al. (2025) achieves low variance primarily through mode collapse to a state that remains proximal to the training data. rads maintains robust mitigation with competitive image fidelity (FID≈26.75FID≈ 26.75), avoiding the quality degradation observed in Jain et al. (2025). 6.6 Was the reachability constraint necessary? To verify the role of reachability analysis, we compare rads against a variant trained with standard SAC without the constraint (λ=0λ=0). As shown in Table 1(a), the unconstrained variant achieves an SSCDtargetSSCD_target of 0.4998, failing to significantly improve upon the unmitigated baseline. This confirms that semantic-alignment rewards alone are insufficient to navigate away from memorization basins; the reachability constraint is the critical mechanism that allows rads to preemptively identify and steer around the BRT of memorization. See Appendix A.4 for more details. 7 Discussion We introduced Reachability-Aware Diffusion Steering (rads), the first framework utilizing reachability analysis and RL to actively mitigate memorization in diffusion models. rads offers distinct advantages over existing baselines: • Dynamic Intervention. Unlike static masks (e.g., Ren et al. (2024)), rads dynamically optimizes a safety objective, applying control while preserving image fidelity. • Robustness. By optimizing for worst-case safety values, rads ensures consistent mitigation across random seeds, avoiding the stochastic failures observed in prior work (Ren et al., 2024). • Generalizability. rads establishes a general approach for safe generative control adaptable to other constraints, such as copyrighted or NSFW content. 8 Limitations and Future Work While rads provides effective mitigation, we identify limitations that offer directions for future research. Semantic Drift via Data Scarcity. Our control policy πϕ _φ learns to navigate around memorization basins by training on a limited set of verified memorized prompts (430 images from Webster (2023)). This small, skewed distribution can lead to overfitting on specific semantic tokens, causing semantic drift when those tokens appear in certain out-of-distribution contexts (see Appendix A.7 for examples of a prompt being steered incorrectly). Future work should incorporate regularization on large-scale, safe open-domain datasets to better distinguish between specific memorization triggers and general semantic concepts. Training Requirement. Unlike zero-shot heuristic baselines (Jain et al., 2025; Ren et al., 2024), rads requires a training phase to learn the steering policy. While training is relatively efficient (15 hours on a single A100 GPU for SD v1.4) compared to full model fine-tuning, it introduces an offline dependency that zero-shot methods avoid. 9 Acknowledgements We gratefully acknowledge research support from Open Philanthropy, the NSF CAREER program (2240163), and the Stanford University School of Engineering. We are also grateful to the Stanford Institute for Human-Centered Artificial Intelligence and Google Cloud for providing computational resources. Additionally, this work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) under 6G Cloud Research and Education Open Hub grant (IITP-2025-RS-2024-00428780) and by the National Research Foundation of Korea (NRF) grant (No. RS-2025-00517159) funded by the Korea government (MSIT). References S. Bansal, M. Chen, S. Herbert, and C. J. Tomlin (2017) Hamilton-Jacobi Reachability: a brief overview and recent advances. In IEEE Conference on Decision and Control (CDC), Cited by: §1. S. Bansal and C. J. Tomlin (2021) DeepReach: a deep learning approach to high-dimensional reachability. In IEEE International Conference on Robotics and Automation (ICRA), Cited by: §2. K. Black, M. Janner, Y. Du, I. Kostrikov, and S. Levine (2023) Training diffusion models with reinforcement learning. arXiv preprint arXiv:2305.13301. Cited by: §2. N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang (2023a) Quantifying memorization across neural language models. In ICLR, Cited by: §1. N. Carlini, M. Jagielski, C. Zhang, N. Papernot, A. Terzis, and F. Tramer (2022) The privacy onion effect: memorization is relative. In NIPS, Cited by: §1. N. Carlini, J. Hayes, M. Nasr, M. Jagielski, V. Sehwag, F. Tramèr, B. Balle, D. Ippolito, and E. Wallace (2023b) Extracting training data from diffusion models. In USENIX Security, Cited by: §1. J. Choi, J. Lee, C. Shin, S. Kim, H. Kim, and S. Yoon (2022) Perception prioritized training of diffusion models. In CVPR, Cited by: §4.2. CivitAI (2023) Realistic vision. External Links: Link Cited by: §5.1. G. Daras and A. G. Dimakis (2022) Multiresolution textual inversion. In NIPS Workshop on Score-Based Methods, Cited by: §4.2. S. Dathathri, A. Madotto, J. Lan, J. Hung, E. Frank, P. Molino, J. Yosinski, and R. Liu (2019) Plug and play language models: a simple approach to controlled text generation. arXiv preprint arXiv:1912.02164. Cited by: §2. P. Dhariwal and A. Nichol (2021) Diffusion models beat gans on image synthesis. In NIPS, Cited by: §1. R. Gandikota, J. Materzynska, J. Fiotto-Kaufman, and D. Bau (2023) Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, p. 2426–2436. Cited by: §2. T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine (2018) Soft actor-critic: off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning, p. 1861–1870. Cited by: §4.4, §5.2. M. Heusel, H. Ramsauer, T. Unterthiner, B. Nessler, and S. Hochreiter (2017) GANs trained by a two time-scale update rule converge to a local nash equilibrium. In NIPS, Cited by: 2nd item. D. Hintersdorf, L. Struppek, K. Kersting, A. Dziedzic, and F. Boenisch (2024) Finding nemo: localizing neurons responsible for memorization in diffusion models. In NIPS, Cited by: 10(d), 9(d), §A.6, 1(a), 1(b), 1(d), 1(e), Table 2, Figure 1, Figure 1, §1, §2, §5.1, §5.3, Figure 6, Figure 6, 7(d), 8(d), §6.1, §6.2. J. Ho, A. Jain, and P. Abbeel (2020) Denoising diffusion probabilistic models. In NIPS, Cited by: 1(d), 1(d), 1(e), 1(e), §1, §3, §4.1, §4.1, §5.1. J. Ho and T. Salimans (2021) Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, Cited by: §3. C. Hong, T. Oh, and M. Sung (2024) MemBench: memorized image trigger prompt dataset for diffusion models. arXiv preprint arXiv:2407.17095. Cited by: §A.7, 1(c), 1(c), §5.1, Figure 6, Figure 6, §6. M. Jagielski, O. Thakkar, F. Tramer, D. Ippolito, K. Lee, N. Carlini, E. Wallace, S. Song, A. G. Thakurta, N. Papernot, and C. Zhang (2023) Measuring forgetting of memorized training examples. In ICLR, Cited by: §1. A. Jain, Y. Kobayashi, T. Shibuya, Y. Takida, N. Memon, J. Togelius, and Y. Mitsufuji (2025) Classifier-free guidance inside the attraction basin may cause memorization. In CVPR, Cited by: 10(e), 9(e), §A.5, Table 1, 1(a), 1(b), 1(c), 1(d), 1(e), Table 2, Figure 1, Figure 1, §1, §2, §4.2, §4.2, §5.1, Figure 5, 7(e), 8(e), §6.1, §6.1, §6.1, §6.2, §6.3, §6.4, §6.5, §8. H. H. Jiang, L. Brown, J. Cheng, M. Khan, A. Gupta, D. Workman, A. Hanna, J. Flowers, and T. Gebru (2023) AI art and its impact on artists. In AIES, Cited by: §1. S. Karnik and S. Bansal (2025) Preemptive detection and steering of llm misalignment via latent reachability. arXiv preprint arXiv:2509.21528. Cited by: §2. J. Kim, S. Kim, and J. Lee (2025) How diffusion models memorize. arXiv preprint arXiv:2509.25705. Cited by: §4.1, §4.2. N. Kumari, B. Zhang, S. Wang, E. Shechtman, R. Zhang, and J. Zhu (2023) Ablating concepts in text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 22691–22702. Cited by: §2. J. Lee, T. Le, J. Chen, and D. Lee (2023) Do language models plagiarize?. In W, Cited by: §1. T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick (2014) Microsoft coco: common objects in context. In European Conference on Computer Vision, p. 740–755. Cited by: 2nd item. J. Liu, L. Zhang, and X. Yuan (2025) DyME: dynamic multi-concept erasure in diffusion models with bi-level orthogonal lora adaptation. arXiv preprint arXiv:2509.21433. Cited by: §2. Y. Park, M. Kwon, J. Choi, J. Jo, and Y. Uh (2023) Understanding the latent space of diffusion models through the lens of riemannian geometry. In NIPS, Cited by: §4.2. E. Pizzi, S. D. Roy, S. N. Ravindra, P. Goyal, and M. Douze (2022) A self-supervised descriptor for image copy detection. In CVPR, Cited by: 1st item. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. In ICML, Cited by: §3, 3rd item. A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen (2022) Hierarchical text-conditional image generation with clip latents. arXiv preprint. Cited by: §1. J. Ren, Y. Li, S. Zeng, H. Xu, L. Lyu, Y. Xing, and J. Tang (2024) Unveiling and mitigating memorization in text-to-image diffusion models through cross attention. In ECCV, Cited by: 10(c), 9(c), 1(a), 1(b), 1(c), 1(d), 1(e), Table 2, Figure 1, Figure 1, §1, §2, §5.1, 7(c), 8(c), §6.1, §6.2, §6.3, 1st item, 2nd item, §8. R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer (2022) High-resolution image synthesis with latent diffusion models. In CVPR, Cited by: §1, §1, §3, §5.1. P. Sharma, N. Ding, S. Goodman, and R. Soricut (2018) Conceptual captions: a cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of ACL, Cited by: 3rd item. G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein (2023a) Diffusion art or digital forgery? investigating data replication in diffusion models. In CVPR, Cited by: §1, 1st item. G. Somepalli, V. Singla, M. Goldblum, J. Geiping, and T. Goldstein (2023b) Understanding and mitigating copying in diffusion models. In NIPS, Cited by: §2. J. Song, C. Meng, and S. Ermon (2021a) Denoising diffusion implicit models. In ICLR, Cited by: §1, §4.1, §5.1. Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Pool (2021b) Score-based generative modeling through stochastic differential equations. In ICLR, Cited by: §1. A. Wagenmaker, M. Nakamoto, Y. Zhang, S. Park, W. Yagoub, A. Nagabandi, A. Gupta, and S. Levine (2025) Steering your diffusion policy with latent space reinforcement learning. arXiv preprint arXiv:2506.15799. Cited by: §2. R. Webster, J. Rabin, L. Simon, and F. Jurie (2021) This person (probably) exists. identity membership attacks against gan generated faces. arXiv preprint. Cited by: §1. R. Webster (2023) A reproducible extraction of training images from diffusion models. arXiv preprint. Cited by: §A.7, 1(a), 1(b), 1(d), 1(e), §1, 1st item, §5.1, Figure 4, Figure 4, §6.2, §6.4, §6.5, §6, §8. Y. Wen, Y. Liu, C. Chen, and L. Lyu (2024) Detecting, explaining, and mitigating memorization in diffusion models. In ICLR, Cited by: 10(b), 9(b), 1(a), 1(b), 1(c), 1(d), 1(e), Table 2, Figure 1, Figure 1, §1, §2, §5.1, 7(b), 8(b), §6.1, §6.1, §6.2, §6.3, §6.5. Appendix A Appendix A.1 High-Quality and Prompt-Aligned Image Generation of rads Figures 9 and 10 show examples of images generated across various methods. (a) Memorized (b) Wen et al. (2024) (c) Ren et al. (2024) (d) Hintersdorf et al. (2024) (e) Jain et al. (2025) (f) rads (Ours) Figure 9: rads achieves mitigation with high image quality. Generated images for the prompt Anna Kendrick is Writing a Collection of Funny, Personal Essays. (a) Generated image without mitigation. (b-e) Mitigated results produced by prior methods. (f) Mitigated result produced by rads (ours). (a) Memorized (b) Wen et al. (2024) (c) Ren et al. (2024) (d) Hintersdorf et al. (2024) (e) Jain et al. (2025) (f) rads (Ours) Figure 10: rads can achieve mitigation with better alignment to the generation intent. Generated images for the prompt Shaw Floors Value Collections Fyc Tt I Net Dinner With Friends (t) 732T_5E022. (a) Generated image without mitigation. (b-e) Mitigated results produced by prior methods. (f) Mitigated result produced by rads (ours). A.2 Implementation Details A.2.1 CLIP Embedding VAE for Action Space Dimensionality Reduction To enable efficient steering in the high-dimensional caption embedding space, we employ a Variational Autoencoder (VAE) that compresses CLIP text embeddings into a manageable 64-dimensional latent action space. The architecture and training configurations are detailed below. Architecture. The VAE utilizes a Transformer-based architecture for both the encoder and decoder components. • Encoder: The encoder consists of N=4N=4 Transformer layers with H=8H=8 attention heads and a hidden dimension of D=768D=768. It processes input CLIP embeddings of shape 77×76877× 768 after adding a learnable positional embedding. The output is passed through a LayerNorm bottleneck before being projected into the latent space. • Decoder: The decoder projects the latent vector back to the 77×76877× 768 sequence shape and processes it through another N=4N=4 Transformer layers to reconstruct the embedding. • Regularization: A dropout rate of 0.10.1 is applied across all layers. Training Protocol and Dataset Mix. The VAE is trained to learn a robust and generalizable representation of the CLIP embedding manifold using the following setup: • Loss Function: The model is trained to minimize a multi-objective loss function that combines semantic and structural reconstruction metrics with a Kullback–Leibler (KL) divergence penalty: ℒ=ℒcos+0.1⋅ℒMSE+2×10−3⋅ℒKLDL=L_cos+0.1·L_MSE+2× 10^-3·L_KLD (10) where ℒcosL_cos represents the cosine embedding loss (1 - cosine similarity) to preserve the semantic orientation of the CLIP vectors, and ℒMSEL_MSE provides structural grounding. • Optimization: We use the Adam optimizer with a learning rate of 10−410^-4 and a batch size of 64 for 45,000 steps. • Dataset Mix: The model is trained on the Conceptual Captions dataset (Sharma et al., 2018) to ensure internet-scale robustness across the text manifold. To prevent overfitting and ensure high fidelity for target contexts, this data is mixed with the specific 430 captions from our experimental training dataset. Periodic evaluation of reconstruction quality (MSE and cosine similarity) is performed specifically on these captions during training. A.2.2 Constrained Soft-Actor-Critic Implementation To solve the constrained MDP, we employ a constrained Soft Actor-Critic (SAC) algorithm with Lagrangian relaxation to enforce the reachability-based safety constraint. The implementation details are as follows. Network Architectures. All agent networks are implemented as Multi-Layer Perceptrons (MLPs) with the following specifications: • Actor Policy (πϕ _φ): A stochastic policy that outputs a Gaussian distribution (μ,σ)(μ,σ) over the d=64d=64 dimensional latent action space. The network consist of 3 hidden layers with 256 units each and ReLU activations. Actions are sampled using the reparameterization trick and squashed via a tanh layer to the range [−1,1][-1,1]. • Twin Task Critics (Qω1taskQ_ _1^task, Qω2taskQ_ _2^task): We utilize twin Q-networks to mitigate overestimation bias. Each critic comprises 3 hidden layers (256 units each) and takes the concatenated state (diffusion latent and caption embedding) and action as input. • Safety Critic (QψsafeQ_ψ^safe): The safety critic share the same architecture as the task critics but is trained to estimate the future reachability value with respect to the memorization failure set. We set γ=0.99γ=0.99 from Equation 7. Hyperparameters and Training Protocol. The policy is trained to optimize the trade-off between semantic reward and the safety constraint using the following parameters: • Optimization: We use the Adam optimizer for all networks. The learning rate is set to 10−410^-4 for the actor and 3×10−53× 10^-5 for the twin task critics, the safety critic, and the entropy temperature α. • Constraint Handling: The Lagrangian multiplier λ is updated via dual gradient descent with a learning rate of 3×10−53× 10^-5. We apply a tanh L2L^2 threshold of 9.0 for the safety reward transformation to separate the L2L^2-norms of the classifier guidance vector based on memorization (L2L^2-norm >> 9.0 denotes memorization, see Appendix A.2.3). • Training Protocol: We use a simulation batch size of 32 trajectories. The policy is trained for 90 epochs. We set the number of network updates to 1. Lastly, we perform evaluation every 5 epochs and track the epoch with the highest sum of the reward function and target function r(T)+ℓ(T)r(s_T)+ (s_T) across the 430 captions and 3 random latent initializations (i.e., samples of Tx_T) for evaluation. A.2.3 Instantiation of the Target Function ℓ In Equation 5, we define our instantiation of the target function ℓ , which is parameterized by the safety threshold β and the scaling factor η. In our implementation, we set η=0.1η=0.1 and β=9.0β=9.0. The threshold β is specifically chosen to facilitate the separation between the guidance magnitudes of novel generations and those associated with memorization. We empirically justify this value by analyzing the L2L^2-norms of the classifier-free guidance vectors across the training set. Our analysis indicates that an L2L^2-norm threshold of approximately 8.618.61 achieves a classification accuracy of 89.77%89.77\% for identifying memorized samples. By approximating β=9.0β=9.0 (corresponding to an accuracy of 89.5%89.5\%), we establish a robust boundary that effectively identifies the BRT of replication while minimizing false positives during steering. Lastly, we set the lower bound δ=0δ=0 for the target function such that ℓ≤0 ≤ 0 represents the failure set and we ensure the Q-function maintains Qsafe≥0Q^safe≥ 0 through constrained optimization. A.3 Complete Quantitative Results See Table 1 for the full results across diffusion models, memorization datasets, and DDIM/DDPM sampling. rads achieves a superior Pareto frontier across different settings. Table 1: RADS strikes the best Pareto frontier between image diversity, quality, and alignment. We compare rads with prior methods and a no-mitigation baseline across different datasets, models, and denoising samplers. Lower indicates better performance for SSCDseedsSSCD_seeds, SSCDpromptsSSCD_prompts, and FID, while higher indicates better performance for CLIP. Best results are in bold. While Jain et al. (2025) achieves the lowest SSCDseedsSSCD_seeds, it suffers from severe quality degradation, as evidenced by the FID and CLIP scores. (a) Webster (2023) Dataset (500 Memorized Prompts), run with SD v1.4 and DDIM Sampling. SSCDtarget† SSCD_target is computed only on the 430 prompts with available target images. Method SSCDtarget†SSCD_target (↓) SSCDseedsSSCD_seeds (↓) SSCDpromptsSSCD_prompts (↓) FID (↓) CLIP (↑) No mitigation 0.6364 ± 0.2626 0.4910 ± 0.2983 0.0956 ± 0.1031 42.1447 ± 5.9063 0.3129 ± 0.0279 Wen et al. (2024) 0.4187 ± 0.2796 0.2132 ± 0.1798 0.0613 ± 0.0686 31.7825 ± 4.8706 0.3056 ± 0.0329 Ren et al. (2024) 0.3781 ± 0.2420 0.2332 ± 0.1876 0.0695 ± 0.0744 32.7631 ± 4.7828 0.2940 ± 0.0500 Hintersdorf et al. (2024) 0.4149 ± 0.2657 0.2534 ± 0.2042 0.0573 ± 0.0584 32.4977 ± 4.6376 0.3119 ± 0.0296 Jain et al. (2025) 0.1816 ± 0.1144 0.0620 ± 0.0608 0.2724 ± 0.1352 63.9765 ± 16.1718 0.2266 ± 0.0532 rads (without constraint) 0.4998 ± 0.2336 0.3449 ± 0.2453 0.0786 ± 0.0766 36.3543 ± 4.1678 0.3010 ± 0.0330 rads (Ours) 0.2303 ± 0.1110 0.1553 ± 0.1099 0.0409 ± 0.0227 31.5678 ± 5.8163 0.2917 ± 0.0366 (b) Webster (2023) Validation Dataset (70 Unseen Prompts of the 500 Memorized Prompts), run with SD v1.4 and DDIM Sampling. Note that the unseen prompts do not have publicly available target images, so SSCDtargetSSCD_target cannot be computed. Method SSCDseedsSSCD_seeds (↓) SSCDpromptsSSCD_prompts (↓) FID (↓) CLIP (↑) No mitigation 0.5730 ± 0.2983 0.0953 ± 0.0821 54.0967 ± 2.1932 0.3220 ± 0.0315 Wen et al. (2024) 0.2690 ± 0.2214 0.0618 ± 0.0532 46.2274 ± 3.8210 0.3127 ± 0.0345 Ren et al. (2024) 0.2842 ± 0.2011 0.0655 ± 0.0544 49.8954 ± 2.9622 0.3120 ± 0.0332 Hintersdorf et al. (2024) 0.2537 ± 0.1543 0.0764 ± 0.0612 49.5971 ± 4.4910 0.3118 ± 0.0301 Jain et al. (2025) 0.0404 ± 0.0321 0.2904 ± 0.1522 73.4757 ± 13.912 0.2285 ± 0.0510 rads (Ours) 0.1356 ± 0.0811 0.0718 ± 0.0501 47.9868 ± 5.4512 0.2795 ± 0.0503 (c) Zero-Shot Generalization to the MemBench Dataset of 3000 Memorized Prompts (Hong et al., 2024), run with SD v1.4 and DDIM Sampling Method SSCDtargetSSCD_target (↓) SSCDseedsSSCD_seeds (↓) SSCDpromptsSSCD_prompts (↓) FID (↓) CLIP (↑) No mitigation 0.5030 ± 0.2836 0.4751 ± 0.3080 0.0367 ± 0.0164 22.0230 ± 0.8903 0.3005 ± 0.0392 Wen et al. (2024) 0.2091 ± 0.1494 0.1219 ± 0.0710 0.0412 ± 0.0178 18.7822 ± 1.9536 0.2998 ± 0.0388 Ren et al. (2024) 0.2184 ± 0.1482 0.1508 ± 0.1121 0.0465 ± 0.0192 20.1319 ± 2.6071 0.2937 ± 0.0426 Jain et al. (2025) 0.1775 ± 0.1226 0.0726 ± 0.0741 0.1782 ± 0.0933 38.5641 ± 7.2228 0.2520 ± 0.0470 rads (Ours) 0.1449 ± 0.0439 0.1006 ± 0.0598 0.0516 ± 0.0238 26.7537 ± 6.5509 0.2524 ± 0.0402 (d) Webster (2023) Dataset (500 Memorized Prompts), run with SD v1.4 and DDPM sampling (Ho et al., 2020). Method SSCDtargetSSCD_target (↓) SSCDseedsSSCD_seeds (↓) SSCDpromptsSSCD_prompts (↓) FID (↓) CLIP (↑) No mitigation 0.6328 ± 0.2506 0.4982 ± 0.2968 0.0942 ± 0.1027 42.0124 ± 6.5782 0.3133 ± 0.0276 Wen et al. (2024) 0.4400 ± 0.2872 0.2441 ± 0.1881 0.0639 ± 0.0755 32.0986 ± 3.3869 0.3077 ± 0.0324 Ren et al. (2024) 0.4039 ± 0.2665 0.2474 ± 0.1994 0.0561 ± 0.0600 29.4978 ± 3.9298 0.2962 ± 0.0502 Hintersdorf et al. (2024) 0.4192 ± 0.2412 0.2801 ± 0.2234 0.0531 ± 0.0548 31.9655 ± 3.5224 0.3123 ± 0.0309 Jain et al. (2025) 0.2507 ± 0.1999 0.0654 ± 0.0632 0.1248 ± 0.0589 37.1680 ± 8.1690 0.2455 ± 0.0529 rads (Ours) 0.2206 ± 0.1029 0.1593 ± 0.1075 0.0337 ± 0.0173 29.4008 ± 5.2703 0.2932 ± 0.0378 (e) Webster (2023) Dataset (500 Memorized Prompts), run with Realistic Vision and DDIM sampling (Ho et al., 2020). SSCDtargetSSCD_target (↓) SSCDseedsSSCD_seeds (↓) SSCDpromptsSSCD_prompts (↓) FID (↓) CLIP (↑) No mitigation 0.5839 ± 0.2677 0.4901 ± 0.3017 0.0865 ± 0.0941 36.5499 ± 3.4359 0.3183 ± 0.0294 Wen et al. (2024) 0.3827 ± 0.2573 0.2136 ± 0.1327 0.0596 ± 0.0623 28.2042 ± 4.3088 0.3107 ± 0.0339 Ren et al. (2024) 0.4318 ± 0.2672 0.2950 ± 0.2093 0.0826 ± 0.0934 34.8126 ± 1.9905 0.3014 ± 0.0501 Hintersdorf et al. (2024) 0.4035 ± 0.2523 0.3154 ± 0.2328 0.0514 ± 0.0518 29.4091 ± 3.8910 0.3175 ± 0.0312 Jain et al. (2025) 0.1544 ± 0.0527 0.0489 ± 0.0394 0.3199 ± 0.1804 84.8658 ± 40.7628 0.2139 ± 0.0537 rads (Ours) 0.1992 ± 0.0797 0.1834 ± 0.1066 0.0364 ± 0.0185 33.0128 ± 9.7632 0.2727 ± 0.0403 A.4 Ablation Study: rads without the Safety Constraint As discussed in Section 6.6, we consider the ablation of removing the constraint from training SAC. We use all of the same hyperparameters and training procedures as in Appendix A.2.2 (while ignoring the target function ℓ and setting the Lagrange coefficient λ=0λ=0). In Table 1(a), we report the performance of rads without the constraint across all metrics. We find that this variant does not substantially reduce SSCDtargetSSCD_target and SSCDseedsSSCD_seeds, thus providing clear evidence of the importance of the safety constraint. A.5 Validating the Steering Mechanism of rads Figure 11: Dynamics of Memorization Steering. Guidance L2L^2-norm during denoising. Standard diffusion (blue) spikes as it collapses into the memorization basin. rads (orange) steers away from this peak. We validate the control-theoretic foundation of our method by tracking the classifier-free guidance norm ‖ϵθ(t,c)−ϵθ(t,∅)‖|| _θ(x_t,e_c)- _θ(x_t, )|| over the denoising process (t:50→0t:50→ 0). Consistent with the “attraction basin” hypothesis (Jain et al., 2025), unmitigated memorized prompts (blue) exhibit an increasing guidance magnitude throughout denoising. In contrast, rads (orange) successfully mitigates away from high guidance magnitude early in the process. This confirms that the safety critic effectively anticipates memorization and the policy applies preemptive steering. We perform this experiment on 50 randomly sampled prompts, 10 evaluation seeds, and 5 training seeds. A.6 Latency Metrics Table 2: Average Inference Time Comparison Method Time (s) No mitigation 2.2950 ± 0.1580 Wen et al. (2024) 2.8983 ± 0.0459 Ren et al. (2024) 2.7544 ± 0.0404 Hintersdorf et al. (2024) 104.0895 ± 48.9586 Jain et al. (2025) 2.2322 ± 0.0368 rads (ours) 2.9262 ± 0.0335 In Table 2, we report the average inference time per prompt across 5 evaluation seeds and 10 prompts. For rads, we compute the average and standard deviation across 5 training seeds. We run all inference on a single A100 GPU. Inference-time mitigation introduces varying overhead. While most methods maintain latents within 3 seconds, Hintersdorf et al. (2024) shows significantly higher latency due to neuron localization. rads offers a competitive inference time of 2.9262 seconds, comparable to the fastest, high-fidelity baselines. A.7 rads Mitigation Failure Mode (a) Memorized (b) rads (ours), Seed = 0 (c) rads (ours), Seed = 1 (d) rads (ours), Seed = 2 (e) rads (ours), Seed = 3 Figure 12: Loss of Semantic Alignment. Generated images for the prompt Mothers influence on her young hippo. (a) Generated image without mitigation. (b-e) Mitigated results produced by rads (ours). One failure mode of rads is that for some out-of-distribution prompts, the generated image can lose some semantic alignment with the original caption. For instance, we consider the example where the caption is “Mothers influence on her young hippo”, which is one of the 70 validation captions (unseen during the training of rads) from the Webster (2023) dataset. In Figure 12, we observe that the memorized image faithfully reproduces an image of a mother hippo with her child. While this generated image is memorized, it aligns well with the provided caption. However, upon perturbing the caption continously in rads, we find that across 4 seeds, rads generates humans in all cases, likely associated with the word “mother”, and a hippo in 2 seeds. As mentioned in Section 8, with just a 430-caption dataset for training, we are likely limited in the generalization capabilities of the rads policy, despite the promising results in Table 1(c) for the MemBench dataset (Hong et al., 2024). With more memorized prompts and safe prompts that do not lead to memorization, we expect rads to improve in image quality and prompt fidelity.