Paper deep dive
Beyond the Black Box: Identifiable Interpretation and Control in Generative Models via Causal Minimality
Lingjing Kong, Shaoan Xie, Guangyi Chen, Yuewen Sun, Xiangchen Song, Eric P. Xing, Kun Zhang
Models: Flux.1, Stable Diffusion 1.4
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/11/2026, 12:37:23 AM
Summary
The paper introduces a principled framework for interpretable generative models using the principle of causal minimality. By enforcing sparsity and compression constraints, the authors establish component-wise identifiability for hierarchical selection models, allowing for the extraction of innate hierarchical concept graphs in diffusion and autoregressive models, which facilitates fine-grained model steering.
Entities (5)
Relation Signals (3)
Causal Minimality → enables → Component-wise Identifiability
confidence 95% · Under theoretically derived minimality conditions... we show that learned representations can be equivalent to the true latent variables
Hierarchical Selection Models → captures → Complex Dependencies
confidence 90% · The selection model structure is particularly adept at capturing the intricate conditional dependencies among low-level features
Sparse Autoencoders → usedfor → Concept Identification
confidence 90% · applying these constraints to leading generative models allows us to extract their innate hierarchical concept graphs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Deep generative models, while revolutionizing fields like image and text generation, largely operate as opaque black boxes, hindering human understanding, control, and alignment. While methods like sparse autoencoders (SAEs) show remarkable empirical success, they often lack theoretical guarantees, risking subjective insights. Our primary objective is to establish a principled foundation for interpretable generative models. We demonstrate that the principle of causal minimality -- favoring the simplest causal explanation -- can endow the latent representations of diffusion vision and autoregressive language models with clear causal interpretation and robust, component-wise identifiable control. We introduce a novel theoretical framework for hierarchical selection models, where higher-level concepts emerge from the constrained composition of lower-level variables, better capturing the complex dependencies in data generation. Under theoretically derived minimality conditions (manifesting as sparsity or compression constraints), we show that learned representations can be equivalent to the true latent variables of the data-generating process. Empirically, applying these constraints to leading generative models allows us to extract their innate hierarchical concept graphs, offering fresh insights into their internal knowledge organization. Furthermore, these causally grounded concepts serve as levers for fine-grained model steering, paving the way for transparent, reliable systems.
Tags
Links
- Source: https://arxiv.org/abs/2512.10720
- Canonical: https://arxiv.org/abs/2512.10720
Trouble viewing inline? Open PDF directly →
Full Text
119,988 characters extracted from source content.
Expand or collapse full text
Preprint BEYOND THE BLACK BOX: IDENTIFIABLE INTERPRE- TATION AND CONTROL IN GENERATIVE MODELS VIA CAUSAL MINIMALITY Lingjing Kong ∗1 , Shaoan Xie ∗1 , Guangyi Chen 1,2 , Yuewen Sun 1,2 , Xiangchen Song 1 , Eric P. Xing 1,2 , Kun Zhang 1,2 ∗ Equal contribution 1 Carnegie Mellon University, Pittsburgh, PA, USA 2 Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE ABSTRACT Deep generative models, while revolutionizing fields like image and text gener- ation, largely operate as opaque “black boxes”, hindering human understanding, control, and alignment. While methods like sparse autoencoders (SAEs) show remarkable empirical success, they often lack theoretical guarantees, risking sub- jective insights. Our primary objective is to establish a principled foundation for interpretable generative models. We demonstrate that the principle of causal mini- mality – favoring the simplest causal explanation – can endow the latent represen- tations of diffusion vision and autoregressive language models with clear causal interpretation and robust, component-wise identifiable control. We introduce a novel theoretical framework for hierarchical selection models, where higher-level concepts emerge from the constrained composition of lower-level variables, better capturing the complex dependencies in data generation. Under theoretically de- rived minimality conditions (manifesting as sparsity or compression constraints), we show that learned representations can be equivalent to the true latent variables of the data-generating process. Empirically, applying these constraints to leading generative models allows us to extract their innate hierarchical concept graphs, offering fresh insights into their internal knowledge organization. Furthermore, these causally grounded concepts serve as levers for fine-grained model steering, paving the way for transparent, reliable systems. 1INTRODUCTION The transformative power of deep generative models, including diffusion models (Sohl-Dickstein et al., 2015; Ho et al., 2020; Rombach et al., 2022b; Song et al., 2022; Dhariwal & Nichol, 2021; Nichol & Dhariwal, 2021) and language models (Radford et al., 2018; 2019; Brown et al., 2020; Raffel et al., 2020), is reshaping numerous domains. However, their escalating complexity and scale frequently cast them as opaque “black boxes” (Shwartz-Ziv & Tishby, 2017; Olah et al., 2020). This opacity presents a formidable barrier to genuine human understanding, severely curtails our ability to exert precise control over their behavior (Jahanian et al., 2019; H ̈ ark ̈ onen et al., 2020; Shen et al., 2020a; Wu et al., 2021), and complicates the crucial alignment with human values and intentions. Although recent empirical tools, such as sparse autoencoders (SAEs) for large language mod- els (LLMs) (Cunningham et al., 2023; Huben et al., 2023; Gao et al., 2024) and diffusion mod- els (Surkov et al., 2024; Kim et al., 2024; Kim & Ghadiyaram, 2025; Cywi ́ nski & Deja, 2025; Huang et al., 2025), offer avenues for probing these models, a fundamental gap persists. Without rigorous theoretical underpinnings, interpretations derived from these methods risk being subjective or susceptible to human biases, rendering them potentially untrustworthy for risk-sensitive appli- cations (Kaddour et al., 2022; Moran & Aragam, 2025; Sch ̈ olkopf et al., 2021). In this work, we directly tackle this critical challenge, seeking to establish a principled foundation for interpretable and controllable generative models. Our investigation centers on two questions: Under what theoretical conditions can we reliably iden- tify meaningful, interpretable latent concepts within the intricate architectures of modern generative 1 arXiv:2512.10720v1 [cs.LG] 11 Dec 2025 Preprint models? And, crucially, what actionable, theoretically-grounded insights can empower us to ad- vance both the interpretability and the controllability of these powerful systems? Towards these goals, we identify the causal minimality (Peters et al., 2017; Spirtes et al., 2000; Hitchcock, 2021) principle as the formal underpinning that connects widespread practices, such as enforcing sparsity, to the recovery of meaningful, interpretable concepts. This principle, advocating for the simplest causal model consistent with observations, allows for the identification of latent hierarchical concept structures. In our context, minimality translates to either sparsity in the concept graphs or the most compressed active discrete concept states. We explore its application to text-to- image (T2I) diffusion models (Ramesh et al., 2021; 2022; Rombach et al., 2022b) and autoregressive language models (LMs) (Radford et al., 2018; 2019; Brown et al., 2020). Our findings indicate that imposing sparsity constraints on internal representations is instrumental for identifying intrinsic visual concepts and textual concepts. A cornerstone of our contribution is establishing the first identifiability results for selection- based (Zheng et al., 2024; Spirtes et al., 1995; Hern ́ an et al., 2004; Zhang, 2008; Bareinboim et al., 2022; Forr ́ e & Mooij, 2020; Correa et al., 2019; Chen et al., 2024) hierarchical models. In such models, higher-level variables emerge as effects of compositions of lower-level variables, where higher-level variables control and select the configuration of lower-level ones. This fundamentally diverges from traditional hierarchical causal models (Pearl, 2009; Choi et al., 2011; Zhang, 2004), in which causal influence typically propagates from higher to lower levels. The selection model structure is particularly adept at capturing the intricate conditional dependencies among low-level features for forming coherent high-level concepts – it explains how specific arrangements of wheels, doors, and a roof constitute a recognizable “car”, rather than a disjointed collection of parts. Tradi- tional hierarchical models often neglect such intra-level dependencies by assuming no within-layer causal edges, as explicitly modeling them would yield overly dense graphs. The selection mech- anism, in contrast, offers a simpler approach to this essential coordination. Its adherence to the minimality principle strongly favors it as a more accurate representation of the true model. Despite the appeal, their identifiability has been underexplored. Prior research has largely cen- tered on traditional hierarchical structures. Moreover, their techniques often rest on simplifying assumptions (e.g., linearity (Xie et al., 2022; Huang et al., 2022; Dong et al., 2023; Anandkumar et al., 2013) or achieve only subspace-level identifiability (Kong et al., 2023a)). Such methods are generally inapplicable to the hierarchical selection models. Our framework is the first to estab- lish component-wise identifiability for both continuous and discrete hierarchical selection models. Specifically, we demonstrate that under well-defined minimality conditions (Conditions 4.2-iv and B.1-i), the learned representations are equivalent to the true latent variables of the underlying hier- archical process. This disentanglement of individual, atomic concepts is what affords significantly more nuanced interpretability and precise control in the resulting generative models. By applying the derived sparsity constraints to state-of-the-art generative models, we successfully extract their innate hierarchical concept graphs (Figure 1). This not only illuminates their internal knowledge organization but also shows that causally-grounded concepts serve as highly effective levers for model steering. Our experiments illustrate key implications of our theorems and show how a principled, causal understanding can guide the application of established interpretation techniques. Due to the page limit, we focus on the visual concept identification with T2I models in the main text and defer the text counterpart with LMs to Appendix B. 2RELATED WORK Hierarchical models. Complex real-world data distributions frequently exhibit inherent hierarchi- cal structures among their underlying latent variables, a characteristic that has motivated extensive research. Initial explorations primarily focus on continuous latent variables with linear interac- tions (Xie et al., 2022; Huang et al., 2022; Dong et al., 2023; Anandkumar et al., 2013). Other lines of work have centered on discrete latent variables; however, these approaches are often con- strained in their applicability to continuous data modalities like images (Pearl, 1988; Zhang, 2004; Choi et al., 2011; Gu & Dunson, 2023; Kong et al., 2024). Furthermore, prevalent latent tree mod- els, which connect variables via a single undirected path (Pearl, 1988; Zhang, 2004; Choi et al., 2011), risk oversimplifying the multifaceted relationships present in complex systems. More re- cently, while Park et al. (2024) make progress in capturing geometric properties of language model 2 Preprint Figure 1: Our causal minimality principle enables interpretable text-to-image generation through hierarchical concept graphs, with implications for downstream tasks. representations using hierarchical models, their work does not address the critical issue of latent vari- able identification. Kong et al. (2023a) tackle nonlinear, continuous latent hierarchical models, but their framework, operating under rather opaque functional conditions, falls short of component-wise identifiability, thereby leaving room for concept entanglement. Our work distinctively investigates selection hierarchical models, contending that their structural properties yield a more faithful repre- sentation of latent concepts in natural data distributions. In these models, latent variables function as colliders, a significant departure from their role as confounders in the aforementioned prior art. This critical distinction renders existing identification techniques largely inapplicable. To the best of our knowledge, we are the first to provide component-wise identifiability for both continuous and discrete hierarchical selection models. Interpretability for generative models. Despite the remarkable advancements of generative mod- els, their internal mechanisms often remain opaque. This presents a significant challenge to un- derstanding and control. Considerable research has focused on obtaining interpretable features to enable more controllable generation. Early efforts center on analyzing the latent space of gener- ative adversarial networks, e.g., (H ̈ ark ̈ onen et al., 2020; Voynov & Babenko, 2020; Shen et al., 2020b). Recently, sparse autoencoders (SAEs) have gained prominence for interpreting hidden rep- resentations, particularly in language models. These studies show that SAEs trained on transformer residual-stream activations can identify latent units corresponding to linguistically meaningful fea- tures (Cunningham et al., 2023; Huben et al., 2023; Gao et al., 2024; Mudide et al., 2025; Shi et al., 2025). These interpretability techniques have also been successfully extended to diffusion models. Surkov et al. (2024) reveal interpretable features and specialization across diffusion model blocks. Other work trains SAEs with lightweight classifiers on diffusion model features (Kim et al., 2024) or steers generation away from undesirable visual attributes (Huang et al., 2025). Our hierarchical approach is related to recent findings on the evolution of semantics during the diffusion process. It has been observed that high-level concepts, such as object shape and structure, tend to emerge in earlier, high-noise timesteps, while fine-grained, low-level details are synthesized in later, low-noise stages (Patashnik et al., 2023; Tinaz et al., 2025; Mahajan et al., 2024). While these works provide valuable empirical validation of this phenomenon, our work offers a new perspective by framing these observations within a formal hierarchical, causal structure. We provide a theoretical founda- tion, rooted in causal minimality and selection models, to explain how these concepts compose and, crucially, under what conditions they can be provably identified. Our approach also relates to gen- erative concept bottleneck models, which achieve interpretability by forcing predictions through a bottleneck layer of concepts (Ismail et al., 2024; Kulkarni et al., 2025). While these methods provide powerful intervention capabilities by design, our work differs by focusing on the discovery of the innate hierarchical and causal concept structure in the data. We provide the theoretical conditions for identifying these concepts component-wise, allowing us to then use this discovered graph for fine-grained multi-level interventions. Decomposition-based interpretability.Our work is fundamentally distinct from post-hoc, decomposition-based interpretability methods, such as the prototype-matching approach (Chen et al., 2019). This line of research, while pioneering, has known limitations (often stemming from its prototype-based implementation): its reliance on class-label supervision can lead to non- 3 Preprint compositional, class-locked concepts (Rymarczyk et al., 2021), and its use of rigid patch-matching struggles with context and deformation (Donnelly et al., 2022; Xue et al., 2024). In contrast, our approach is class/object agnostic (similar to SAEs) and context-sensitive, learning from the raw generative data without class labels. Our approach learns compositional, shared concepts (e.g., a single “furry texture” from “cats” and “pandas”) rather than rigid, class-specific prototypes. This enables the causal, interventional control (e.g., Figure 11 and downstream tasks in Section 5.2) that prototype-matching cannot guarantee. Please find additional related work in Appendix A. 3DEEP GENERATIVE MODELS AS HIERARCHICAL CONCEPT MODELS Notations. We denote random variables with upper-case characters (e.g., X ) and values with lower- case characters (e.g., x). We distinguish multidimensional objects with bold fonts (e.g., X) and refer to their dimensionality as n(·). We view multidimensional variables as sets when appropriate (e.g., X asX i i∈[n(X)] ). Parents Pa(·) and children Ch(·) relations are defined based on the selection graph (Figure 2). If X has only one child Y , we refer to X as a pure parent of Y , i.e., X ∈ PPa(Y ); if X has other children than Y , we refer to X as a hybrid parent of Y , i.e., X ∈ HPa(Y ). We denote the set of natural numbers1,...,M as [M ]. More background information is in Appendix D. We denote the image as the continuous variable X ∈ R n(X) and text as the discrete variable D ∈ N n(D) . Visual concepts are Z := [Z 1 ,· , Z L V ], where L V is the number of visual hierarchical levels and Z l ∈ R n(Z l ) are concepts at level l (Figure 2). The discrete variables D capture the discrete nature of textual concepts (like “cat” or “bicycle”). In contrast, the visual concepts (Z) are continuous to represent rich visual details. D acts as a selection variable that governs the joint configuration of the continuous Z variables. For instance, the discrete concept “bicycle” (D) selects for a coherent arrangement of continuous visual features (Z) representing wheels, a frame, and handlebars, rather than a random collection of those continuous parts. D 1 D 2 Z 1,1 Z 1,2 Z 2,1 Z 2,2 Z 2,3 Z 2,4 Z 3,1 Z 3,2 Z 3,3 Z 3,4 Z 3,5 Z 3,6 X Text Visual Concepts Image Figure 2: A visual concept graph. We denote text as D, visual concepts as Z, and the image as X. High-level concepts function as selection variables for low-level variables. See Figure 7 for the text counterpart. Hierarchical processes and selection mechanisms. Our framework conceptualizes high-level concepts as emerging from or being effects of lower-level concepts. This is captured by a selection mecha- nism (Zheng et al., 2024; Spirtes et al., 1995; Hern ́ an et al., 2004; Zhang, 2008; Bareinboim et al., 2022; Forr ́ e & Mooij, 2020; Correa et al., 2019; Chen et al., 2024), where variables V l at a higher level of ab- straction (smaller l) is determined by its constituent, more detailed components V l+1 (i.e., its “parents”). The selection function g V l maps these lower-level constituents to the higher-level concept: V l := g V l (V l+1 ).(1) In other words, V l is a selection variable over V l+1 . In many natural data distributions of interest, we can only observe the data points for which the selec- tion criterion is met, i.e., V l only takes on a strict subset of its range Ω. Therefore, the distribution of V l+1 is always the conditional distribution P (V l+1 |V l ). This conditioning on V l can induce dependencies among components in V l+1 . For instance, if V l+1,i → V l ← V l+1,j , conditioning on V l makes V l+1,i and V l+1,j dependent. Under this formulation, one can leverage the inverse process of (1) to sample observable data (im- ages, text), proceeding from higher-level abstract concepts to lower-level concrete details: Z 0 ∼ P (Z 0 ), Z l ∼ P (Z l |Z l−1 ), l∈1,...,L V + 1,(2) where we denote Z 0 := D and Z L V +1 := X. While (2) defines the generative pathway, the underlying structure is shaped by the selection principle of (1): the conditional distributions in (2) are implicitly learned if one has learned selection mechanisms in (1) and vice versa. Why is this “selection” formulation? The “selection” perspective is critical for modeling how ab- stract concepts enforce coherence among their more concrete constituents. Consider generating an 4 Preprint image of a “bicycle” (a high-level concept Z l ). Its components – wheels, frame, handlebars (lower- level concepts Z l+1 ) – must not only be present but also be arranged in a specific, structurally sound configuration. Traditional hierarchical models (Choi et al., 2011; Pearl, 1988; Zhang, 2004) assume independent low-level concepts Z l+1,i given high-level concepts Z l and stochastically sample these components, which could lead to unrealistic arrangements (e.g., wheels detached from the frame if the learned conditional is not perfect). Therefore, these models must additionally incorporate causal edges within each hierarchical level to capture this conditional dependency, resulting in highly dense causal graphs. In contrast, the selection model, by positing that Z l is an effect of a specific configura- tion of Z l+1 , emphasizes that the “bicycle” concept arises from a coherent selection and composition of its parts. This structured dependency, induced by the selection mechanism, yields a much simpler graphical model to describe the natural data distribution, thus preferred by the minimality principle. Connections to text-to-image diffusion models. The iterative denoising process in diffusion aligns with our hierarchical data construction. These models involve a sequence of transforma- tionsf t T t=1 , parameterized by timestep t, that progressively restore a less noisy image X t from a more corrupted version X t+1 . As interpreted by Kong et al. (2024), each f t+1 can be viewed as an autoencoder: it extracts a representation Z S(t+1) (S(t + 1) indexes U-Net features associated with timestep t + 1) from the noisy input X t+1 , and uses this representation to produce the less noisy X t . In this view, representations Z S(t+1) from higher noise levels (larger t, where X t+1 is closer to pure noise) correspond to higher-level, more abstract concepts in our hierarchy (e.g., Z l with smaller l), as fine-grained details are obscured by noise. Conversely, representations from lower noise levels (smaller t) capture more concrete details (e.g., Z l with larger l). The diffusion model’s step-wise refinement thus mirrors our hierarchical generation P (Z l+1 |Z l ), with the initial text prompt D typ- ically guiding the most abstract visual concepts (e.g., Z 1 ∼ P (Z 1 |D), Figure 1). In our empirical analysis (Section 5, we explicitly map distinct diffusion timesteps to these hierarchical levels: high noise levels (e.g., t = 899) and low noise levels (e.g., t = 100) to fine-grained details. Identifiability and interpretability. In light of the connection, a crucial question remains: are the internal representations learned by these models (e.g., U-Net features, transformer activations) truly reflective of the ground-truth concepts of the data, or are they merely effective for the generation task without being inherently interpretable and controllable? This motivates the need for identifiability guarantees that affirm the equivalence between the two worlds, which we present in Section 4. 4IDENTIFIABLE REPRESENTATIONS UNDER CAUSAL MINIMALITY We first formally define our core theoretical principle, causal minimality (Peters et al., 2017; Spirtes et al., 2000; Hitchcock, 2021): Among all causal models that can explain the observed data, the true model is the simplest one. This principle is the key to our goal of identifiability (Definition 4.1). Causal minimality, as a principle, manifests as concrete, enforceable mechanisms in specific settings. For visual concepts, this mechanism is sparse connectivity in the causal graph (our minimality con- dition, 4.2-iv), and for text, it is state compression (Condition B.1-i,i). Enforcing this sparsity or compression is thus the practical mechanism that provides theoretical guarantees for identifiability. For visual concepts, minimality manifests as a preference for sparse graphical dependencies within the latent hierarchy. This implies that concepts are formed through a limited set of direct causal influences, making the underlying structure easier to discern. In Appendix B, we discuss how the minimality principle translates to seeking the most compressed representation for discrete text con- cepts, the identification theory, and the connection to language models. A key challenge we address is the identifiability of hierarchical selection models. In these models, higher-level concepts are effects of lower-level concepts. This contrasts with traditional hierarchical models where causality often flows from abstract to concrete, and where latent variables typically act as confounders (Pearl, 1988; Zhang, 2004; Choi et al., 2011; Gu & Dunson, 2023; Kong et al., 2024; Xie et al., 2022; Huang et al., 2022; Dong et al., 2023; Kong et al., 2023a; Anandkumar et al., 2013). In our selection framework, latent variables act as colliders, rendering many existing identifiability results inapplicable. This distinction necessitates the novel theoretical development presented herein. Our goal is to achieve component-wise identifiability: Definition 4.1 (Component-wise Identifiability). Let Z and ˆ Z be variables under two model spec- ifications. We say that Z and ˆ Z are identified component-wise if there exists a permutation π such that for each i∈ [n(Z)], ˆ Z i = h i (Z π(i) ) where h i is an invertible function. 5 Preprint This strong form of identifiability ensures that each learned latent component ˆ Z i corresponds to a single true latent component Z π(i) . This is vital for unambiguous interpretation and targeted control. We assume the standard faithfulness condition (Spirtes et al., 2001), meaning the graphical model accurately reflects all conditional independence relations in the data. In the following, we consider the identification of continuous latent visual concepts Z and present the counterpart for textual concepts in Appendix B.2. Condition 4.2 (Visual Concept Identification Conditions). i Informativeness: There exists a diffeomorphism g l : (Z l ,ε l ) 7→ X for l ∈ [0,L], whereε l denotes independent exogenous variables. i Smooth Density: The probability density function p(z l+1 |z l ) is smooth for any l∈ [L V ]. i Sufficient Variability: For each Z and its parents ̃ Z := Pa(Z), at any value ̃ z of ̃ Z, there exist n( ̃ Z) + 1 distinct values of Z, denoted asz (n) n( ̃ Z) n=0 , such that the vectors w( ̃ z,z (n) )− w( ̃ z,z (0) ) are linearly independent where w( ̃ z,z) = ∂ logp( ̃ z|z) ∂ ̃z 1 ,..., ∂ logp( ̃ z|z) ∂ ̃z n( ̃ z) . iv Sparse Connectivity (Minimality): For each parent concept ̃ Z, there exists a subset of its children Z⊆ Ch( ̃ Z) such that their only common parent is ̃ Z, i.e., T Z∈Z Pa(Z) = ̃ Z. Interpreting Condition 4.2. Condition 4.2-i ensures that the observed data X (e.g., an image) fully captures the information about the latent concepts Z l . This is a natural assumption as high- dimensional observations contain rich information. Condition 4.2-i is a standard regularity as- sumption for analysis. Both are common in nonlinear ICA literature (Hyvarinen & Morioka, 2016; Hyvarinen et al., 2019; Khemakhem et al., 2020b;a; Von K ̈ ugelgen et al., 2021; Kong et al., 2023a). Condition 4.2-i formalizes the idea that distinct lower-level concepts (e.g., “wheel,” “door”) re- spond in sufficiently distinct ways to changes in a shared higher-level concept (e.g., “car”), thus facilitating the identification of these lower-level concepts. Condition 4.2-iv is an instantiation of causal minimality for visual concepts. It posits that the causal graph of concepts is sparse – each concept has a unique “fingerprint” in terms of its connectivities. This sparsity is crucial for disen- tanglement (Zheng et al., 2022; Lachapelle et al., 2024a; Xu et al., 2024; Lachapelle et al., 2022b;a) and is a less restrictive assumption than, for example, pure observed children for each latent vari- able (Arora et al., 2012; 2013; Moran et al., 2021). This condition formalizes a core principle: concepts are learned through comparison. A concept is identifiable only if the data is rich enough to distinguish it from alternatives. For instance, if “Knight” and “Horse” always co-occur, they are learned as a fused concept; learning them separately requires data that breaks this correlation. Theorem 4.3 (Visual Concept Identification).Assume the process for visual concepts in (2). If a model specificationθ V satisfies Condition 4.2, and an alternative specification ˆ θ V satisfies Condi- tions 4.2-i and 4.2-i, along with a sparsity constraint such that for corresponding ˆ Z and Z: n(Pa( ˆ Z))≤ n(Pa(Z)),(3) then, if both modelsθ V and ˆ θ V generate the same observed data distributionP (X), the latent visual concepts Z l are component-wise identifiable for every level l∈ [L V ]. Proof sketch for Theorem 4.3. The proof proceeds by identifying the hierarchical model level by level, from the top (most abstract concepts) Z 1 downwards to Z L V . 1) The paired text data D acts as an auxiliary variable, providing diverse “influences” on the top-level Z 1 . Condition 4.2- i ensures these interventions have distinguishable effects. Analogous to techniques in nonlinear ICA (Hyvarinen & Morioka, 2016; Hyvarinen et al., 2019; Kong et al., 2022), each component D allows the identification of the subspace of Z 1 variables it influences. 2) With these subspaces identified, one can identify the intersection of these subspaces (Von K ̈ ugelgen et al., 2021; Yao et al., 2023; Kong et al., 2023b). Therefore, if the graphical structure is sufficiently sparse, as specified in Condition B.1-iv, one can identify the top-level latent variable Z 1 component-wise. 3) Once Z 1 is identified, its components can serve as the auxiliary variables to identify the next level, Z 2 . This process is repeated iteratively down the hierarchy, identifying Z l using the already identified Z l−1 . Implications for text-to-image diffusion models. Theorem 4.3 underscores that the sparsity con- straint (3) is pivotal for identifying true visual concepts. In practice, this constraint is instantiated 6 Preprint Figure 3: Examples of hierarchical concept graphs for text-to-image models. Our method successfully recovers meaningful hierarchical structures, where each node encodes distinct semantic concepts. On the right, we demonstrate feature steering, where manipulating individual nodes leads to changes in the output that align with their position in the hierarchy. Intervening on a high-level concept in the learned graph (“Full face”) alters the cat’s entire facial structure and fur pattern. In contrast, intervening on a learned lower-level concept (e.g., “Eye”) produces a much more localized edit, changing only the shape and color of the eyes while leaving the rest of the face intact. More examples in Appendix F. through a two-step process: 1) Level-specific concept learning: We train K-sparse SAEs on features at the specific timesteps defined in Section 3. This approximates the sparsity condition required by Theorem 4.3. 2) Cross-level causal discovery: We then apply causal discovery algorithms (e.g., PC (Spirtes et al., 2001)) across these sparse features to construct the hierarchical graph, validating that the learned representations align with the theoretical identification guarantees. 5EXPERIMENTS We present results on T2I models and refer readers to Appendix B.3 for LM experiments. Evaluation design and objectives. We design our experiments to validate our theoretical frame- work in two ways. In Section 5.1, we provide a direct empirical test of our theory: we apply the spar- sity constraints derived from causal minimality (Condition 4.2) and show that we can, as predicted, extract a meaningful and interpretable hierarchical concept graph. In Section 5.2, we demonstrate the utility of these identified concepts. If our concepts are truly component-wise identifiable (Def- inition 4.1), they should be individually controllable. We test this via a suite of challenging down- stream tasks—including model unlearning, controllable image generation, and multi-level editing. For example, we compare against state-of-the-art unlearning methods to rigorously benchmark our concept removal capabilities. Our objective is to show that our theory not only finds interpretable concepts but also provides a practical mechanism for fine-grained, reliable model control. More detailed settings for each experiment are provided in their respective subsections. Hierarchical causal analysis. Our theoretical framework motivates an empirical analysis that dif- fers from standard interpretability approaches. Following the framework established in Sections 3 and 4, we apply our two-step identification process to Stable Diffusion (SD) 1.4 (Rombach et al., 2022a) and Flux.1-Schnell (Labs, 2024) (Appendix F). We analyze feature representations at the pre- viously defined timesteps (899, 500, and 100) to extract and verify the hierarchical concept graph. Benefits. This hierarchical perspective provides two main benefits. First, it enables compositional editing. For a complex object like “a textured tree stump”, our analysis can distinguish the ”stump” (a mid-level concept) from its “texture” (a low-level one), allowing for independent steering. This is a fine-grained control challenging for non-hierarchical methods that tend to learn entangled features (see Table 6). Second, it allows for targeted intervention. By identifying a concept’s level, we can inject a steered feature back into the diffusion process only at its corresponding timestep, which helps in reducing the unwanted artifacts that can arise from applying steering globally across all timesteps (see Figure 5). More details in Appendix E and Figure 9. 7 Preprint MethodI2P↓RING-A-BELL↓P4D↓ UATK↓COCO K77K38K16AVGFID↓ CLIP↑ SD 1.417.8 85.26 87.37 93.68 88.10 98.7069.7016.7131.3 ESD2.87 20.00 29.47 35.79 28.42 15.492.8718.1830.2 SA2.81 63.15 56.84 56.84 58.94 12.682.8125.8029.7 CA1.04 86.32 91.69 94.26 90.765.631.0424.1230.1 MACE1.512.100.000.000.702.821.5116.8028.7 UCE0.87 10.52 9.47 12.61 10.879.860.8717.9930.2 RECE0.725.264.215.264.915.630.7217.7430.2 SDID3.77 94.74 95.79 90.53 93.68 69.5430.9922.1631.1 SLD-MAX1.74 23.16 32.63 42.11 32.639.142.4428.7528.4 SLD-STRONG2.28 56.84 64.21 61.05 60.70 33.103.1024.4029.1 SLD-MEDIUM3.95 92.63 88.42 91.05 90.70 24.001.9821.1729.8 SD1.4-NegPrompt 0.74 17.89 40.42 34.74 31.68 10.001.4618.3330.1 SAFREE1.45 35.78 47.36 55.78 46.31 10.561.4519.3230.1 TRASCE0.451.052.102.101.753.970.7017.4129.9 ConceptSteer0.363.168.429.477.021.992.1118.6730.8 Ours0.251.050.002.111.050.662.1117.0231.3 Table 1: Model unlearning comparisons. Our method delivers competitive results on unlearning tasks without compromising standard text-to-image generation. See Appendix D for details. Full face Upper-face Mid-face Ears 3556 3066 762 3489 1441 3044 1026 Forehead Mouth Eyes Full face Bottom Image Upper-face Ears 3556 3066 3489 3654 1441 3390 4109 Forehead Eyes SparsityIncreasing Full face Top-face 3556 3119 Eyes 3489 3066 Forehead Forehead Figure 4: Ablation studies on the sparsity constraint. We control feature sparsity at timestep 500. Without enforcing sparsity, the resulting concepts tend to be dense, and the features are less interpretable. Conversely, higher sparsity leads to a more interpretable, sparser graph. However, when sparsity becomes too high, the resulting graph may become overly sparse and fail to adequately capture the generation of the cat face. 5.1INTERPRETABILITY ANALYSIS Hierarchical concept graph. Figure 3 illustrates a hierarchical graph learned through our approach (more in Appendix F). On the left, we display activation maps of different SAE features. Brown nodes (SAE nodes trained on timestep 899) capture high-level features, such as node 3556 repre- senting an entire cat face. Green nodes (timestep 500) reflect mid-level features, like node 3044 cap- turing the central face. Blue nodes (timestep 100) capture fine details—node 3066 activates on the eyes and node 762 on the mouth. This demonstrates a clear progression from coarse to fine-grained concepts across timesteps. To thoroughly examine the existence of the hierarchical concept graph, we conduct two complementary experiments demonstrating that activations at higher timesteps cap- ture more global semantics, while those at lower timesteps capture more localized details. First, we quantify the spatial spread of activations across timesteps. For each SAE, we compute attribution maps for its top feature indices. Given an SAE feature of shape 64× 64× 5120, we compute a 64× 64 attribution map. Applying a 0.1 threshold yields a binary attribution map, from which we measure the proportion of activated pixels. Across 1,000 samples, approximately 280, 630, 880, and 1,400 unique concepts are activated for K = 1, 3, 5, and 10, respectively. Activations at timestep 899 influence a larger spatial area, indicating that higher timesteps capture more global, distributed concepts. Second, we generate images from 10,000 COCO prompts and deactivate the top-1 SAE activation at each timestep. Comparing the modified generations with the originals shows that deac- tivations at noisier timesteps cause substantial, global changes, while those at less noisy timesteps 8 Preprint (a) Spatial activation spread(b) SAE-deactivated generation comparison TimestepTop1Top3Top5Top10L1LPIPSCLIPDINO 1000.270.210.190.150.0040.0020.9990.999 5000.300.250.210.170.0130.0200.9950.993 8990.530.410.330.240.0700.2200.9480.903 Table 2: Quantitative analyses across different noise levels. (a) Spatial activation spread: aver- age proportion of pixels influenced by the top-k SAE activations. Higher timesteps affect a larger spatial area, indicating that SAEs at noisier steps capture more global, distributed concepts. (b) SAE-deactivated generation comparison: similarity metrics between original and SAE-deactivated images. Deactivation at higher timesteps produces greater perceptual and semantic changes, sup- porting the presence of a hierarchical organization of concepts across timesteps. MetricSD 1.4SD1.4 (SAE w/o hier.)SD1.4 (Ours) Add tabby pattern – CLIP-I↓0.91± 0.050.83± 0.070.93± 0.04 Add tabby pattern – CLIP-T↑0.27± 0.000.28± 0.020.28± 0.01 Add mountains – CLIP-I↓0.84± 0.060.83± 0.040.91± 0.03 Add mountains – CLIP-T↑0.33± 0.010.32± 0.010.33± 0.01 Replace rock w/ stump – CLIP-I↓0.93± 0.020.95± 0.020.96± 0.02 Replace rock w/ stump – CLIP-T↑0.31± 0.010.29± 0.010.31± 0.01 Table 3: Controllable image generation results. Our method achieves the best CLIP-I metric, demonstrating greater fidelity to the input images, while reliably executing the target edits. produce localized effects. These results confirm that features at different noise levels encode distinct abstraction levels, supporting the hierarchical concept graph. Concept steering in hierarchical graphs. We conduct concept steering using our discovered fea- tures, as shown on the right side of Fig. 3 (more in Appendix F). Given a model intermediate feature x, the SAE encoder E and decoder D are trained to reconstruct x. To steer a specific concept, we obtain the latent representation z = E(x), and extract the steering vector v corresponding to the desired feature. We then modify the original feature to create a steered version x ′ = x + λD(v), where λ modulates the strength. By feeding the steered x ′ back into the diffusion process at the same timestep, we generate images that reflect the influence of the selected concept. For example, steer- ing node 3556 – associated with the entire face of a cat – results in a significantly altered cat face. Steering the green node 1026 modifies only the upper part of the face, illustrating that it encodes localized information specific to that region. Ablation. As established in the theoretical framework, sparsity is crucial for identifiability. To em- pirically validate this, we visualize the resulting causal graphs under varying levels of sparsity, as shown in Fig. 4 (more in Appendix F). When sparsity is not enforced, the resulting graph becomes overly dense, making it difficult to interpret and diminishing its semantic clarity. Conversely, impos- ing excessive sparsity leads to an overly pruned graph that lacks sufficient structure to meaningfully explain the generation process, such as in the case of the cat image. These observations highlight the importance of balancing sparsity to preserve interpretability while maintaining explanatory power. 5.2DOWNSTREAM TASKS Thanks to our theoretical framework, we can naturally perform a range of image generation and editing tasks, including model unlearning, controllable image generation, and multi-level editing. Model unlearning. We provide quantitative results of model unlearning on four benchmark datasets: IP2P (Schramowski et al., 2023), three splits of RING-A-BELL (Tsai et al., 2023), P4D (Chin et al., 2023), and UnlearnDiffATK (Zhang et al., 2024b). These benchmarks focus on removing nudity-related concepts, and we report the accuracy of a pretrained nudity detector. Our method achieves the best results across all benchmarks. In addition, to assess whether our method preserves general text-to-image capability, we apply feature steering on normal prompts from MSCOCO (Lin et al., 2014). The 10K results, reflected in low FID and high CLIP scores, 9 Preprint BeforeOursSAE,w/ohierBeforeOursSAE,w/ohier Figure 5: Generated samples with P4D prompts (Chin et al., 2023). The Stable Diffusion model is vulnerable to the prompts in the p4d dataset, producing unsafe images. When the hierarchical relationship across timesteps is not considered, negative steering with SAE results in drastic changes to the output. In contrast, our method learns to apply modifications to the nudity feature at a suitable timestep without introducing additional distortions. 3372 2212 2212+3372 2212 3372 5079 1531 1531+5079 1531 5079 SAE, w/ohier drink Iceremoval Ice Avariationofdrink dog, crown crown Avariationofdogwearingcrown Crownremoval SAE, w/ohier Figure 6: Examples of multi-level editing (best viewed with zoom). High-level node 2212 con- tains all information about the cup, while mid-level node 3372 focuses primarily on the ice cubes. Similarly, high-level node 1531 encompasses all information about the dog (including the crown), and mid-level node 5079 is dedicated to the crown. By modeling hierarchical relationships, we can perform edits that are often difficult to achieve with a single-layer edit. For instance, if we want to generate a variation of the cup while removing the ice cubes, we can apply feature steering on high-level node 2212 to create a new version of the cup, and simultaneously apply negative feature steering on mid-level node 3372 to remove the ice cubes. demonstrate that our method successfully identifies and removes nudity concepts without affecting unrelated concepts. We also provide results on style removal in the appendix (Table 5) and we achieve superior performance across different metrics and tasks. Controllable image generation. We also evaluate controllable image generation on three editing tasks: adding tabby patterns to cat faces, adding mountains to landscape images, and replacing rocks with textured tree stumps. As shown in Table 3 and Fig.10, our method achieves superior results compared to both the standard text-guided model and SAE without hierarchical modeling. Multi-level image editing. A key advantage of the hierarchical concept graph is that it can com- bine nodes across different levels for fine-grained image editing. In Fig. 6, to obtain a new drink without ice (while preserving the background), we can apply multi-level editing by steering features at both high-level node 2212 and mid-level node 3372 simultaneously. Without such hierarchical relationship modeling, conventional methods struggle to produce this combination, which can result in undesired changes such as the drink being replaced by a person or the dog’s background. 6CONCLUSION In this work, we present a theoretical framework using causal minimality for identifying latent con- cepts in hierarchical selection models. We prove that generative model representations can map to true latent variables. Empirically, applying these constraints enables extracting meaningful hierar- chical concept graphs from leading models, enhancing interpretability and grounded control. 10 Preprint REFERENCES Animashree Anandkumar, Daniel Hsu, Adel Javanmard, and Sham Kakade.Learning linear bayesian networks with latent variables. In International Conference on Machine Learning, p. 249–257. PMLR, 2013. Sanjeev Arora, Rong Ge, and Ankur Moitra. Learning topic models–going beyond svd. In 2012 IEEE 53rd annual symposium on foundations of computer science, p. 1–10. IEEE, 2012. Sanjeev Arora, Rong Ge, Yonatan Halpern, David Mimno, Ankur Moitra, David Sontag, Yichen Wu, and Michael Zhu. A practical algorithm for topic modeling with provable guarantees. In International conference on machine learning, p. 280–288. PMLR, 2013. Elias Bareinboim, Jin Tian, and Judea Pearl. Recovering from selection bias in causal and statistical inference. Probabilistic and causal inference: The works of Judea Pearl, p. 433–450, 2022. Sander Beckers. Equivalent causal models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, p. 6202–6209, 2021. Sander Beckers and Joseph Y Halpern. Abstracting causal models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, p. 2678–2685, 2019. Joseph Bloom, Curt Tigges, Anthony Duong, and David Chanin. Saelens. https://github. com/jbloomAus/SAELens, 2024. Jack Brady, Roland S Zimmermann, Yash Sharma, Bernhard Sch ̈ olkopf, Julius Von K ̈ ugelgen, and Wieland Brendel. Provably learning object-centric representations. In International Conference on Machine Learning, p. 3038–3062. PMLR, 2023. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhari- wal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. Language models are few-shot learners, 2020. Simon Buchholz and Bernhard Sch ̈ olkopf. Robustness of nonlinear representation learning. In International Conference on Machine Learning, p. 4785–4821. PMLR, 2024. Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su. This looks like that: deep learning for interpretable image recognition. Advances in neural information processing systems, 32, 2019. Leihao Chen, Onno Zoeter, and Joris M Mooij. Modeling latent selection with structural causal models. arXiv preprint arXiv:2401.06925, 2024. Zhi-Yi Chin, Chieh-Ming Jiang, Ching-Chun Huang, Pin-Yu Chen, and Wei-Chen Chiu. Prompt- ing4debugging: Red-teaming text-to-image diffusion models by finding problematic prompts. arXiv preprint arXiv:2309.06135, 2023. Myung Jin Choi, Vincent YF Tan, Animashree Anandkumar, and Alan S Willsky. Learning latent tree graphical models. Journal of Machine Learning Research, 12:1771–1812, 2011. Joel E Cohen and Uriel G Rothblum. Nonnegative ranks, decompositions, and factorizations of nonnegative matrices. Linear Algebra and its Applications, 190:149–168, 1993. Juan D Correa, Jin Tian, and Elias Bareinboim. Identification of causal effects in the presence of selection bias. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, p. 2744–2751, 2019. Thomas M Cover. Elements of information theory. John Wiley & Sons, 1999. Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoen- coders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023. 11 Preprint Bartosz Cywi ́ nski and Kamil Deja. Saeuron: Interpretable concept unlearning in diffusion models with sparse autoencoders. arXiv preprint arXiv:2501.18052, 2025. Prafulla Dhariwal and Alex Nichol. Diffusion models beat gans on image synthesis. In Advances in Neural Information Processing Systems 34 (NeurIPS 2021), 2021. Xinshuai Dong, Biwei Huang, Ignavier Ng, Xiangchen Song, Yujia Zheng, Songyao Jin, Roberto Legaspi, Peter Spirtes, and Kun Zhang. A versatile causal discovery framework to allow causally- related hidden variables. In The Twelfth International Conference on Learning Representations, 2023. Jon Donnelly, Alina Jade Barnett, and Chaofan Chen. Deformable protopnet: An interpretable image classifier using deformable prototypes. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10265–10275, 2022. Patrick Forr ́ e and Joris M Mooij. Causal calculus in the presence of cycles, latent confounders and selection bias. In Uncertainty in Artificial Intelligence, p. 71–80. PMLR, 2020. Rohit Gandikota, Joanna Materzynska, Jaden Fiotto-Kaufman, and David Bau. Erasing concepts from diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, p. 2426–2436, 2023. Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzy ́ nska, and David Bau. Unified concept editing in diffusion models. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 5111–5120, 2024. Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020. Leo Gao, Tom Dupr ́ e la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. arXiv preprint arXiv:2406.04093, 2024. Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, volume 34, p. 9574–9586, 2021. Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman. Find- ing alignments between interpretable causal variables and distributed neural representations. In Conference on Causal Learning and Reasoning, 2024. Chao Gong, Kai Chen, Zhipeng Wei, Jingjing Chen, and Yu-Gang Jiang. Reliable and efficient concept erasure of text-to-image diffusion models. In European Conference on Computer Vision, p. 73–88. Springer, 2024. Yuqi Gu and David B. Dunson. Bayesian pyramids: Identifiable multilayer discrete latent struc- ture models for discrete data. In Journal of the Royal Statistical Society Series B: Statistical Methodology, 2023. Erik H ̈ ark ̈ onen, Aaron Hertzmann, Jaakko Lehtinen, and Sylvain Paris. Ganspace: Discovering interpretable gan controls. Advances in neural information processing systems, 33:9841–9850, 2020. Alvin Heng and Harold Soh. Selective amnesia: A continual learning approach to forgetting in deep generative models. Advances in Neural Information Processing Systems, 36:17170–17194, 2023. Miguel A Hern ́ an, Sonia Hern ́ andez-D ́ ıaz, and James M Robins. A structural approach to selection bias. Epidemiology, 15(5):615–625, 2004. Christopher Hitchcock. Probabilistic Causation. In Edward N. Zalta (ed.), The Stanford Encyclope- dia of Philosophy. Metaphysics Research Lab, Stanford University, Spring 2021 edition, 2021. 12 Preprint Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising diffusion probabilistic models. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020), 2020. Biwei Huang, Charles Jia Han Low, Feng Xie, Clark Glymour, and Kun Zhang. Latent hierarchical causal structure discovery with rank constraints. Advances in Neural Information Processing Systems, 35:5549–5561, 2022. Victor Shea-Jay Huang, Le Zhuo, Yi Xin, Zhaokai Wang, Peng Gao, and Hongsheng Li. Tide: Temporal-aware sparse autoencoders for interpretable diffusion transformers in image generation. arXiv preprint arXiv:2503.07050, 2025. Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In The Twelfth International Conference on Learning Representations, 2023. Aapo Hyvarinen and Hiroshi Morioka. Unsupervised feature extraction by time-contrastive learning and nonlinear ica. Advances in neural information processing systems, 29, 2016. Aapo Hyvarinen, Hiroaki Sasaki, and Richard Turner. Nonlinear ica using auxiliary variables and generalized contrastive learning. In The 22nd International Conference on Artificial Intelligence and Statistics, p. 859–868. PMLR, 2019. Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho. Concept bottleneck generative models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=L9U5MJJleF. Ali Jahanian, Lucy Chai, and Phillip Isola. On the” steerability” of generative adversarial networks. In International Conference on Learning Representations, 2019. Anubhav Jain, Yuya Kobayashi, Takashi Shibuya, Yuhta Takida, Nasir Memon, Julian Togelius, and Yuki Mitsufuji. Trasce: Trajectory steering for concept erasure. arXiv preprint arXiv:2412.07658, 2024. Yibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam, and Victor Veitch. On the origins of linear representations in large language models. In Forty-first International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id= otuTw4Mghk. Shruti Joshi, Andrea Dittadi, S ́ ebastien Lachapelle, and Dhanya Sridhar. Identifiable steering via sparse autoencoding of multi-concept shifts. arXiv preprint arXiv:2502.12179, 2025. Jean Kaddour, Aengus Lynch, Qi Liu, Matt J Kusner, and Ricardo Silva. Causal machine learning: A survey and open problems. arXiv preprint arXiv:2206.15475, 2022. Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen. Variational autoen- coders and nonlinear ica: A unifying framework. In International Conference on Artificial Intel- ligence and Statistics, p. 2207–2217. PMLR, 2020a. Ilyes Khemakhem, Ricardo Monti, Diederik Kingma, and Aapo Hyvarinen. Ice-beem: Identifiable conditional energy-based deep models based on nonlinear ica. Advances in Neural Information Processing Systems, 33:12768–12778, 2020b. Dahye Kim and Deepti Ghadiyaram. Concept steerers: Leveraging k-sparse autoencoders for con- trollable generations. arXiv preprint arXiv:2501.19066, 2025. Dahye Kim, Xavier Thomas, and Deepti Ghadiyaram. Revelio: Interpreting and leveraging semantic information in diffusion models. arXiv preprint arXiv:2411.16725, 2024. Bohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, and Bryon Aragam. Learning latent causal graphs via mixture oracles. Advances in Neural Information Processing Systems, 34:18087– 18101, 2021. Bohdan Kivva, Goutham Rajendran, Pradeep Ravikumar, and Bryon Aragam. Identifiability of deep generative models without auxiliary information. Advances in Neural Information Processing Systems, 35:15687–15701, 2022. 13 Preprint Lingjing Kong, Shaoan Xie, Weiran Yao, Yujia Zheng, Guangyi Chen, Petar Stojanov, Victor Akin- wande, and Kun Zhang. Partial disentanglement for domain adaptation. In International Confer- ence on Machine Learning, p. 11455–11472. PMLR, 2022. Lingjing Kong, Biwei Huang, Feng Xie, Eric Xing, Yuejie Chi, and Kun Zhang. Identification of nonlinear latent hierarchical models. Advances in Neural Information Processing Systems, 36, 2023a. Lingjing Kong, Martin Q. Ma, Guangyi Chen, Eric P. Xing, Yuejie Chi, Louis-Philippe Morency, and Kun Zhang. Understanding masked autoencoders via hierarchical latent variable models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 7918–7928, June 2023b. Lingjing Kong, Guangyi Chen, Biwei Huang, Eric P. Xing, Yuejie Chi, and Kun Zhang. Learning discrete concepts in latent hierarchical models. In The Thirty-eighth Annual Conference on Neu- ral Information Processing Systems, 2024. URL https://openreview.net/forum?id= bO5bUxvH6m. Akshay Kulkarni, Ge Yan, Chung-En Sun, Tuomas Oikarinen, and Tsui-Wei Weng. Interpretable generative models through post-hoc concept bottlenecks. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 8162–8171, 2025. Nupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman, Richard Zhang, and Jun-Yan Zhu. Ablating concepts in text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 22691–22702, 2023. Black Forest Labs. Flux. https://github.com/black-forest-labs/flux, 2024. S ́ ebastien Lachapelle, Tristan Deleu, Divyat Mahajan, Ioannis Mitliagkas, Yoshua Bengio, Simon Lacoste-Julien, and Quentin Bertrand. Synergies between disentanglement and sparsity: Gener- alization and identifiability in multi-task learning. arXiv preprint arXiv:2211.14666, 2022a. S ́ ebastien Lachapelle, Pau Rodriguez, Yash Sharma, Katie E Everett, R ́ emi Le Priol, Alexandre Lacoste, and Simon Lacoste-Julien. Disentanglement via mechanism sparsity regularization: A new principle for nonlinear ica. In Conference on Causal Learning and Reasoning, p. 428–484. PMLR, 2022b. S ́ ebastien Lachapelle, Pau Rodr ́ ıguez L ́ opez, Yash Sharma, Katie Everett, R ́ emi Le Priol, Alexan- dre Lacoste, and Simon Lacoste-Julien.Nonparametric partial disentanglement via mecha- nism sparsity: Sparse actions, interventions and sparse temporal dependencies. arXiv preprint arXiv:2401.04890, 2024a. S ́ ebastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, and Simon Lacoste-Julien. Additive de- coders for latent variables identification and cartesian-product extrapolation. Advances in Neural Information Processing Systems, 36, 2024b. Hang Li, Chengzhi Shen, Philip Torr, Volker Tresp, and Jindong Gu. Self-discovering inter- pretable diffusion latent directions for responsible text-to-image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 12006–12016, 2024. Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ́ ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, p. 740–755. Springer, 2014. Yuhang Liu, Dong Gong, Erdun Gao, Zhen Zhang, Biwei Huang, Mingming Gong, Anton van den Hengel, and Javen Qinfeng Shi. I predict therefore i am: Is next token prediction enough to learn human-interpretable concepts from data? arXiv preprint arXiv:2503.08980, 2025. Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022. 14 Preprint Shilin Lu, Zilan Wang, Leyang Li, Yanzhu Liu, and Adams Wai-Kin Kong. Mace: Mass concept erasure in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 6430–6440, 2024. Shweta Mahajan, Tanzila Rahman, Kwang Moo Yi, and Leonid Sigal. Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 6808–6817, 2024. Emanuele Marconato, S ́ ebastien Lachapelle, Sebastian Weichwald, and Luigi Gresele. All or none: Identifiable linear properties of next-token predictors in language modeling. arXiv preprint arXiv:2410.23501, 2024. Gemma E. Moran and Bryon Aragam. Towards interpretable deep generative models via causal representation learning, 2025. Gemma E Moran, Dhanya Sridhar, Yixin Wang, and David M Blei. Identifiable variational autoen- coders via sparse decoding. arXiv preprint arXiv:2110.10804, 2021. Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt. Efficient dictionary learning with switch sparse autoencoders. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum? id=k2ZVAzVeMP. Alex Nichol and Prafulla Dhariwal. Improved denoising diffusion probabilistic models. In Proceed- ings of the International Conference on Machine Learning (ICML 2021), 2021. Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. Zoom in: An introduction to circuits.Distill, 2020.doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in. Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The geometry of categorical and hierar- chical concepts in large language models. arXiv preprint arXiv:2406.01506, 2024. Or Patashnik, Daniel Garibi, Idan Azuri, Hadar Averbuch-Elor, and Daniel Cohen-Or. Localiz- ing object-level shape variations with text-to-image diffusion models. In Proceedings of the IEEE/CVF international conference on computer vision, p. 23051–23061, 2023. J. Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann, 1988. Judea Pearl. Causality. Cambridge university press, 2009. Jonas Peters, Dominik Janzing, and Bernhard Sch ̈ olkopf. Elements of causal inference: foundations and learning algorithms. The MIT Press, 2017. Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. Improving language under- standing by generative pre-training. OpenAI blog, 2018. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9, 2019. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. Goutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Sch ̈ olkopf, and Pradeep Kumar Ravikumar. From causal to concept-based representation learning. In Causality and Large Models @NeurIPS 2024, 2024. URL https://openreview.net/forum?id=FcVnIBYbkW. Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. Zero-shot text-to-image generation. In International Conference on Machine Learning, p. 8821–8831. PMLR, 2021. 15 Preprint Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text- conditional image generation with clip latents. In Advances in Neural Information Processing Systems 36 (NeurIPS 2022), 2022. Patrik Reizinger, Alice Bizeul, Attila Juhos, Julia E Vogt, Randall Balestriero, Wieland Brendel, and David Klindt. Cross-entropy is all you need to invert the data generating process. arXiv preprint arXiv:2410.21869, 2024. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ̈ orn Ommer. High- resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition, p. 10684–10695, 2022a. Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ̈ orn Ommer. High- resolution image synthesis with latent diffusion models. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), 2022b. Paul K Rubenstein, Sebastian Weichwald, Stephan Bongers, Joris M Mooij, Dominik Janzing, Moritz Grosse-Wentrup, and Bernhard Sch ̈ olkopf. Causal consistency of structural equation mod- els. In Proceedings of the Conference on Uncertainty in Artificial Intelligence (UAI), 2017. Dawid Rymarczyk, Łukasz Struski, Jacek Tabor, and Bartosz Zieli ́ nski. Protopshare: Prototyp- ical parts sharing for similarity discovery in interpretable image classification. In KDD ’21, p. 1420–1430, New York, NY, USA, 2021. Association for Computing Machinery. ISBN 9781450383325. doi: 10.1145/3447548.3467245. URL https://doi.org/10.1145/ 3447548.3467245. Bernhard Sch ̈ olkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio. Toward causal representation learning. Proceedings of the IEEE, 109(5):612–634, 2021. Patrick Schramowski, Manuel Brack, Bj ̈ orn Deiseroth, and Kristian Kersting. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 22522–22531, 2023. Christoph Schuhmann, Andreas K ̈ opf, Theo Coombes, Richard Vencu, Romain Beaumont, and Ben- jamin Trom. Laion-coco: 600m synthetic captions from laion2b-en. LAION.ai blog, September 2022. URL https://laion.ai/blog/laion-coco/. Claude E Shannon. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423, 1948. Yujun Shen, Jinjin Gu, Xiaoou Tang, and Bolei Zhou. Interpreting the latent space of gans for se- mantic face editing. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9243–9252, 2020a. Yujun Shen, Ceyuan Yang, Xiaoou Tang, and Bolei Zhou. Interfacegan: Interpreting the disentan- gled face representation learned by gans. IEEE transactions on pattern analysis and machine intelligence, 44(4):2004–2018, 2020b. Wei Shi, Sihang Li, Tao Liang, Mingyang Wan, Gojun Ma, Xiang Wang, and Xiangnan He. Route sparse autoencoder to interpret large language models. CoRR, abs/2503.08200, March 2025. URL https://doi.org/10.48550/arXiv.2503.08200. Ravid Shwartz-Ziv and Naftali Tishby. Opening the black box of deep neural networks via informa- tion. arXiv preprint arXiv:1703.00810, 2017. Jascha Sohl-Dickstein, Eric A. Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsu- pervised learning using nonequilibrium thermodynamics. In Proceedings of the International Conference on Machine Learning (ICML 2015), 2015. Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models. In Interna- tional Conference on Learning Representations (ICLR 2022), 2022. 16 Preprint Peter Spirtes, Christopher Meek, and Thomas Richardson. Causal inference in the presence of latent variables and selection bias. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, p. 499–506, 1995. Peter Spirtes, Clark N Glymour, and Richard Scheines. Causation, Prediction, and Search. MIT press, 2000. Peter Spirtes, Clark Glymour, and Richard Scheines. Causation, prediction, and search. MIT press, 2001. Viacheslav Surkov, Chris Wendler, Mikhail Terekhov, Justin Deschenaux, Robert West, and Caglar Gulcehre. Unpacking sdxl turbo: Interpreting text-to-image models with sparse autoencoders. arXiv preprint arXiv:2410.22366, 2024. Gemma Team. Gemma. Kaggle, 2024. doi: 10.34740/KAGGLE/M/3301. URL https://w. kaggle.com/m/3301. Berk Tinaz, Zalan Fabian, and Mahdi Soltanolkotabi. Emergence and evolution of interpretable concepts in diffusion models. arXiv preprint arXiv:2504.15473, 2025. Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin, Jia-You Chen, Bo Li, Pin-Yu Chen, Chia-Mu Yu, and Chun-Ying Huang. Ring-a-bell! how reliable are concept removal methods for diffusion models? arXiv preprint arXiv:2310.10012, 2023. Julius Von K ̈ ugelgen, Yash Sharma, Luigi Gresele, Wieland Brendel, Bernhard Sch ̈ olkopf, Michel Besserve, and Francesco Locatello. Self-supervised learning with data augmentations provably isolates content from style. Advances in neural information processing systems, 34:16451–16467, 2021. Andrey Voynov and Artem Babenko. Unsupervised discovery of interpretable directions in the gan latent space. In International conference on machine learning, p. 9786–9796. PMLR, 2020. Zongze Wu, Dani Lischinski, and Eli Shechtman. Stylespace analysis: Disentangled controls for stylegan image generation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 12863–12872, 2021. Quanhan Xi and Benjamin Bloem-Reddy. Indeterminacy in generative models: Characterization and strong identifiability. In International Conference on Artificial Intelligence and Statistics, p. 6912–6939. PMLR, 2023. Feng Xie, Biwei Huang, Zhengming Chen, Yangbo He, Zhi Geng, and Kun Zhang. Identifica- tion of linear non-gaussian latent hierarchical structure. In International Conference on Machine Learning, p. 24370–24387. PMLR, 2022. Danru Xu, Dingling Yao, S ́ ebastien Lachapelle, Perouz Taslakian, Julius Von K ̈ ugelgen, Francesco Locatello, and Sara Magliacane. A sparsity principle for partially observable causal representation learning. arXiv preprint arXiv:2403.08335, 2024. Mengqi Xue, Qihan Huang, Haofei Zhang, Jingwen Hu, Jie Song, Mingli Song, and Canghong Jin. Protopformer: Concentrating on prototypical parts in vision transformers for interpretable image recognition. In IJCAI, 2024. Dingling Yao, Danru Xu, Sebastien Lachapelle, Sara Magliacane, Perouz Taslakian, Georg Martius, Julius von K ̈ ugelgen, and Francesco Locatello. Multi-view causal representation learning with partial observability. In The Twelfth International Conference on Learning Representations, 2023. Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao, and Mohit Bansal. Safree: Training-free and adaptive guard for safe text-to-image and video generation. arXiv preprint arXiv:2410.12761, 2024. Jiji Zhang. On the completeness of orientation rules for causal discovery in the presence of latent confounders and selection bias. Artificial Intelligence, 172(16-17):1873–1896, 2008. 17 Preprint Kun Zhang, Shaoan Xie, Ignavier Ng, and Yujia Zheng. Causal representation learning from multi- ple distributions: A general setting. In Forty-first International Conference on Machine Learning, 2024a. URL https://openreview.net/forum?id=Pte6iiXvpf. Nevin L Zhang. Hierarchical latent class models for cluster analysis. The Journal of Machine Learning Research, 5:697–723, 2004. Yimeng Zhang, Jinghan Jia, Xin Chen, Aochuan Chen, Yihua Zhang, Jiancheng Liu, Ke Ding, and Sijia Liu. To generate or not? safety-driven unlearned diffusion models are still easy to generate unsafe images... for now. In European Conference on Computer Vision, p. 385–403. Springer, 2024b. Yujia Zheng, Ignavier Ng, and Kun Zhang. On the identifiability of nonlinear ica: Sparsity and beyond. arXiv preprint arXiv:2206.07751, 2022. Yujia Zheng, Zeyu Tang, Yiwen Qiu, Bernhard Sch ̈ olkopf, and Kun Zhang. Detecting and identify- ing selection structure in sequential data. In International Conference on Machine Learning, p. 61498–61525. PMLR, 2024. 18 Preprint Appendix ARELATED WORK Latent variable identification. Identifying latent variables is a cornerstone of representation learn- ing. A significant body of work establishes identifiability for single-level latent variable models, often assuming the availability of auxiliary information like domain or class labels (Khemakhem et al., 2020a;b; Hyvarinen & Morioka, 2016; Hyvarinen et al., 2019; Zhang et al., 2024a). Recently, research into language models has explored the linear representation hypothesis, yielding linear- subspace identifiability for latent variables (Reizinger et al., 2024; Liu et al., 2025; Marconato et al., 2024; Rajendran et al., 2024; Jiang et al., 2024). Another research direction (Brady et al., 2023; Lachapelle et al., 2024b;a; Xu et al., 2024; Lachapelle et al., 2022a;b; Zheng et al., 2022; Joshi et al., 2025) leverages sparsity for identification but overlooks the causal relationships among latent variables. Distinct from these approaches, our work formulates the concept space using hierarchical models that allow for the explicit modeling of intricate, multi-level conceptual interactions. Our work also connects to the literature on causal abstraction, which studies how a high-level causal model can be faithfully derived from a low-level one (Rubenstein et al., 2017; Geiger et al., 2021; 2024; Beckers & Halpern, 2019; Beckers, 2021). A key distinction is our focus on component-wise identifiability, which guarantees that the discovered concepts are equivalent to the true latent vari- ables, providing a stronger foundation for interpretability. Our work is complementary to important research on weak vs. strong (Xi & Bloem-Reddy, 2023) and approximate identifiability (Buchholz & Sch ̈ olkopf, 2024). While much of this literature analyzes single-level models, our framework is the first to establish component-wise identifiability (Definition 4.1) for hierarchical selection mod- els. This result fits within the weak identifiability category (Xi & Bloem-Reddy, 2023), as do most results in this area. Critically, this level of identifiability is motivated by and sufficient for our down- stream tasks, aligning with the principle of “task-identifiability” (Xi & Bloem-Reddy, 2023). It pro- vides the necessary guarantee for meaningful interpretation and control without requiring the stricter assumptions of strong identifiability. Moreover, the results on approximate identifiability (Buchholz & Sch ̈ olkopf, 2024) are encouraging, suggesting that robust representations can be learned even if our minimality conditions are only approximately met. BFORMULATION, THEORY, AND EXPERIMENTS FOR LANGUAGE MODELS B.1FORMULATION FOR TEXT GENERATION S 1,1 S 1,2 S 2,1 S 2,2 S 2,3 S 2,4 D 1 D 2 D 3 D 4 D 5 D 6 Text Textual Concepts Figure 7: A textual concept graph. We denote text as D and discrete textual concepts as S. High-level concepts function as selection variables of low-level variables. Textual concepts are S := [S 1 ,· , S L T ], where L T is the number of textual hierarchical levels and S l ∈ Ω l ⊂ N n(S l ) are concepts at level l. S 1 ∼ P (S 1 ), S l ∼ P (S l |S l−1 ), l∈2,...,L T + 1,(4) where we denote Z 0 := D, Z L V +1 := X, and S L T +1 := D. Connections to autoregressive language models. An autoregressive language model can be seen as learning an “encoder” that maps a sequence of input tokens (D 1:t ) to an internal state ˆ S l (e.g., activations within transformer layers). This internal state ˆ S l then informs the “decoder” to predict the 19 Preprint subsequent token. For optimal prediction, this learned representation ˆ S l should ideally capture the information of the true concept S l that d-separates the input tokens D 1:t from the next token D t+1 . To achieve this d-separation, S l should belong to a higher concept level for a larger span of text D 1:t (i.e., larger t→ smaller l, see Figure 9). Consequently, broad thematic or narrative structures spanning larger text segments can be compressed into higher-level concepts in our hierarchy (e.g., S 1 ), while more localized syntactic or lexical choices correspond to lower-level concepts (e.g., S L T ). Intuition on “compression” and higher-level concepts in language models. Our core intuition is that an autoregressive model, at any token position t, compresses the all the preceding sequence (tokens 1 to t) into a representation that is useful for predicting the next token at t + 1. In a later position, the model has access to more context and strictly more information. Consequently, the minimality constraint promotes more abstract and compressed representations over the information it has seen. This pressure to compress a growing context naturally gives rise to a hierarchy of concepts. Let’s use an example for illustration. When a model reads, “He was secretly buying balloons, sending coded messages to friends, and looking up cake recipes...”, it would hold onto this list of disparate actions. The meaning is ambiguous; the model has to keep the details in memory. However, once it has parsed the entire sentence, “He was secretly buying balloons, sending coded messages to friends, and looking up cake recipes – he was getting ready for the surprise party for his sister”, the model can now form a high-level concept - a celebratory plan — that organizes all the previous, seemingly random actions into a coherent event. This final concept is more compressed and abstract than the initial list of actions, illustrating the move from detailed memorization to a clear, high-level summary as more context becomes available. In this example, the concepts that exist at later stages of the sequence are not just additions but are fundamentally more abstract, as they synthesize a larger body of information. This aligns directly with our theoretical framework (Condition B.1-i), where we posit that concepts become more compressed (i.e., have minimal support) as we move up the hierarchy. B.2LEARNING TEXTUAL CONCEPTS VIA STATE COMPRESSION We now turn to the identification of discrete textual concepts S. The minimality principle manifests as seeking the most “compressed” representation, namely, achieving minimal support sizes for these discrete concepts while preserving full information. Condition B.1 (Textual Concept Identification Conditions). i Natural Selection: Each selection variable S l has a support supp(S l ) that is a proper subset of its potential range if its constituent parts (lower-level variables) were combined randomly. That is, supp(S l ) ⊊ f D→S l (Ω n(Pa(S l )) ), where f D→S l is the function from D to S l . i Bottlenecks: The support size of any concept S l is strictly smaller than the joint support size of its parents Pa(S l ) in the selection graph. i Minimal Supports: For any S, the condition distribution P (D\ Pa(S)|S = s, HPa(S) = ̃ s) is a one-to-one function w.r.t. the argument s. iv No-Twins: Distinct latent variables must have distinct sets of adjacent (parent/child) variables. v Maximality: The identified latent structure is maximal in the sense that splitting any latent concept variable would violate either the Markov conditions or the No-Twins condition. Interpreting Condition B.1. Condition B.1-i posits that meaningful text (or textual concepts) oc- cupies a small, structured subset of the vast space of all possible token combinations. We rarely encounter truly random sequences of words in natural language. Conditions B.1-i and B.1-i are direct manifestations of causal minimality for discrete concepts. i implies an information compres- sion moving up the hierarchy—abstract concepts are more succinct. i demands that each state of a concept s offers unique information about the rest of the text, given its context. Therefore, the representation is most compressed (minimal number of states) and each state contains unique infor- mation. Conditions B.1-iv and B.1-v are standard necessary conditions for discrete latent variable model identification (Kivva et al., 2021; 2022), precluding redundant or fragmented latent structures. Theorem B.2 (Textual Concept Identification). Assume the hierarchical process as per (4). Let the true underlying parameters beθ T . Ifθ T satisfies Condition B.1, and an alternative learned model 20 Preprint L20 Node 1555:Humor references to humor and comedic elements L18 Node 12358:Phrase & introductory markers references to specific phrases or introductory words, particularly variations of the word "that“ L19 Node 11028: Pronouns pronouns and references to individuals L19 Node 1751: Punctuation specific punctuation and its variations in the text L18 Node 15394: Phrases on regulations phrases associated with legal considerations and regulations L18 Node 5803: Terms defining obligations keywords related to obligations, responsibilities, and consequences of partnerships or agreements L19 Node 4346: Legal terminology legal and copyright-related terminology Node 7373: Legal procedures numerical values or references to legal procedures L18 13186 1751 11028 155514377 12082 8532 12368 1007 1235812025 4346 15394 73735803 Causal path: (1751, 11028) → (1555) Example: The old man chuckled, holding aloft a mischievous kitten that had somehow gotten stuck in his wizard's hat. "A light jest," he(11028) quipped, "but(1751) a joke with the power to unravel the dark.“ Fish blink delicately as puns make their way through dung-humorous puns. (1555) 1751 11028 1555 Example 1 Example 2 4346 15394 73735803 Causal path: (7373, 15394, 5803) → (4346) Example: The contract mentioned “in accordance with regulations(15394)”, outlined “mutual liabilities and penalties (5803)”, and cited “Section 12.4(7373) of the Civil Code”, all forming the basis of its legal terminology. (4346) Figure 8: The learned hierarchical concept graph for autoregressive language models. By modeling the hierarchy of concepts based on the token sequence order, we recover a meaningful hierarchical graph. The brown nodes (corresponding to later tokens) capture global, high-level information, while the green nodes (from intermediate tokens) represent more localized, lower-level concepts. ˆ θ T satisfies Condition B.1-i, then if both models produce the same observed distribution P (D), the latent textual concepts S l are component-wise identifiable for every level l∈ [L T ]. Proof sketch for Theorem B.2. The identification for textual concepts proceeds from the bottom level (tokens, S L T ) upwards to the most abstract concepts (S 1 ). (1) At each level l + 1, we make use of the conditional independence relations that the high-level variable S l,i and its hybrid parents HPa(S l,i ) d-separate its pure parents PPa(S l,i ) from the other variables S l \Pa(S l,i ) on level l. This relation allows us to identify subsets of S l+1 that share children on level l (Cohen & Roth- blum, 1993; Kong et al., 2024) and thus reveals the connectivity between variables in S l and S l+1 . (2) Once the graphical connections are known, we recover the function Pa(S l,i ) 7→ S l,i (i.e., how lower-level concepts combine to form S l,i ). This is done by merging states of Pa(S l,i ) that are pre- dictively equivalent. The “Minimal Supports” (Condition B.1-i) principle dictates that we choose the function that results in the largest equivalence classes over the parent states (i.e., the most com- pressed representation for S l,i ). This ensures that the learned concept ˆ S l,i has the minimum number of necessary states. (3) This process of structure learning and function recovery is repeated from S L T (initially using observed tokens D as S L T +1 ) up to S 1 , thereby identifying the entire hierarchy. Implications for autoregressive language models. Theorem B.2 suggests that by enforcing a mini- mality regularization for the most compressed representation (Condition B.1-i), the learned internal states ˆ S of a language model can become equivalent to the underlying textual concepts S. SAEs, when applied to transformer activations, can be seen as a practical way to approximate this minimal- ity. By forcing most latent units to be inactive, SAEs force the model to encode information with the minimal active units, which aligns with our theoretical condition for state compression. This result provides a principled justification for the observed interpretability of SAE-derived features and guides our empirical approach in Section B.3 to extract hierarchical textual concept graphs. B.3EXPERIMENTS ON AUTOREGRESSIVE LANGUAGE MODELS Implementation. In this section, we present our implementation for analyzing autoregressive lan- guage models. We utilize pretrained SAEs (Bloom et al., 2024) for Gemma-2-2b-it (Team, 2024). We partition tokens into three parts based on the their positions in their positions in the in- put sequence. This segmentation reflects the expectation that tokens convey increasingly abstract or high-level information as the sequence progresses. Finally, we apply causal discovery algorithms to uncover the relationships among features across the different SAEs. More details in Appendix E. Results. Figure 8 shows a learned hierarchical graph (more in Appendix F). Nodes 1555 and 12082 are mostly activated for final tokens in the sequence, and thus capture high-level semantics. Specif- ically, node 1555 is associated with the humorous tone, while node 12082 represents the role of the dog. Interestingly, node 11028, derived from intermediate tokens, emerges as a causal factor 21 Preprint for both 1555 and 12082. This node encodes pronouns and references to individuals, which play a critical role in shaping both the humor and the characterization of the dog. CPROOFS C.1PROOF FOR THEOREM 4.3 Lemma C.1 (Base Case Visual Concept Identification). Assume the following data-generating pro- cess: C∼ P (C|U), V∼ P (V), X := g(C, V).(5) We have the following conditions. i Informativeness: The function g(·) is a diffeomorphism. i Smooth Density: The probability density function p(c, v|u) is smooth. i Sufficient Variability: At any value c of C, there exist n(C) + 1 distinct values of U, denoted as u (n) n(C) n=0 , such that the vectors w(c, u n ) − w(c, u 0 ) are linearly independent where w(c, u) = ∂ logp(c|u) ∂c 1 ,..., ∂ logp(c|u) ∂c n(c) . If a specificationθ satisfies i,i, and i, another specification ˆ θ satisfies i,i, and they generate matching distribution P (X), then we can verify that C and ˆ C can be identified up to its subspace. Proof. Since we have matched distributions, it follows that: p(x|u) = ˆp(x|u).(6) As the generating function g has a smooth inverse (i), we can derive: p(g(c, v)|u) = p(ˆg( ˆ c, ˆ v)|u) =⇒ p(c, v|u) J g −1 = ˆp(g −1 ◦ ˆg( ˆ c, ˆ v)|u) J g −1 . Notice that the Jacobian determinant J g −1 > 0 because of g(·)’s invertibility and let h := g −1 ◦ ˆg : ( ˆ c, ˆ v) 7→ (c, v) which is smooth and has a smooth inverse thanks to those properties of g and ˆg. It follows that p(c, v|u) = ˆp(h( ˆ c, ˆ v)|u) =⇒ p(c, v|u) = ˆp( ˆ c, ˆ v|u)|J h −1 |. The independence relation in the generating process implies that logp(c|u) + X i∈[n(v)] logp(V i ) = log ˆp( ˆ c|u) + X i∈[n( ˆ v)] log ˆp( ˆ V i ) + log|J h −1 |.(7) For any realization u 0 , we subtract (7) at any u̸= u 0 with that at u 0 : logp(c|u)− logp(c|u 0 ) = log ˆp( ˆ c|u)− log ˆp( ˆ c|u 0 ).(8) Taking derivative w.r.t. ˆv j for j ∈ [n( ˆ v)] yields: X i∈[n(c)] ∂ ∂c i (logp(c|u)− logp(c|u 0 ))· ∂c i ∂ ˆv j = 0.(9) The left-hand side zeros out because ˆ c is not a function of ˆ v. Condition i ensures the existence of at least n(c) such equations with u 1 ,..., u n(c) that are linearly independent, constituting a full-rank linear system. Since the choice of j ∈ [v] is arbitrary. It follows that ∂c i ∂ ˆv j = 0,∀i∈ [n(c)],j ∈ [n(v)].(10) 22 Preprint Therefore, the Jacobian matrix J h is of the following structure: J h = " ∂v ∂ ˆ v ∂v ∂ ˆ c ∂c ∂ ˆ v ∂c ∂ ˆ c . # (11) (10) suggests that the block ∂c ∂ ˆ v = 0. Since J h is full-rank, we can deduce that ∂c ∂ ˆ c must have full row-rank and n(c) ≤ n( ˆ c). The sparsity constraint in (3) further implies that n(c) = n( ˆ c). That is, we can correctly identify the dimensionality of the changing subspace c. Moreover, since J h is full-rank and the block ∂c ∂ ˆ v is zero, we can derive that the corresponding block ∂ ˆ c ∂v in its inverse matrix J h −1 is also zero. Therefore, there exists an invertible map ˆ c 7→ c, which concludes the proof. Lemma C.2 (Determining Intersection Cardinality from Union Cardinalities). Let A = A 1 ,A 2 ,...,A n be a finite collection of finite sets. If for any non-empty subset of indices K ⊆1, 2,...,n, the cardinality of the union S k∈K A k is known, then for any non-empty subset of indices S ⊆1, 2,...,n, the cardinality of the intersection T s∈S A s can be determined. Proof. We proceed by induction on the size of the set of indices S, denoted by |S|, for which we want to determine the intersection cardinality. Base Case: |S| = 1. Let S = i for some i ∈ 1, 2,...,n. We aim to determine the cardinality T s∈S A s = |A i |. The union of a single set A i is simply A i itself. That is, A i = S k∈i A k . By the premise of the theorem, the cardinality S k∈i A k is known. Therefore, |A i | is known. The base case holds. Inductive Hypothesis: Assume that for some integer m ≥ 1, the cardinality of any intersection of j sets, T j∈J A j , can be determined from the known union cardinalities for all non-empty index sets J such that 1≤|J|≤ m. Inductive Step: We want to show that the cardinality of any intersection of m + 1 sets can be determined. Let S m+1 be an arbitrary non-empty subset of indices from 1, 2,...,n such that |S m+1 | = m + 1. Our goal is to determine T s∈S m+1 A s . Consider the Principle of Inclusion-Exclusion (PIE) applied to the union of the sets whose indices are in S m+1 : [ s∈S m+1 A s = X ∅̸=K⊆S m+1 (−1) |K|−1 \ k∈K A k This sum runs over all non-empty subsets K of S m+1 . We can separate the term where K = S m+1 (which corresponds to the intersection of all m + 1 sets) from the other terms in the sum: [ s∈S m+1 A s = X ∅̸=K⊂S m+1 (−1) |K|−1 \ k∈K A k + (−1) |S m+1 |−1 \ s∈S m+1 A s Here, the sum is now over all non-empty proper subsets K of S m+1 . We can rearrange this equation to solve for the term T s∈S m+1 A s : (−1) |S m+1 |−1 \ s∈S m+1 A s = [ s∈S m+1 A s − X ∅̸=K⊂S m+1 (−1) |K|−1 \ k∈K A k Multiplying both sides by (−1) |S m+1 |−1 (noting that ((−1) |S m+1 |−1 ) 2 = 1): \ s∈S m+1 A s = (−1) |S m+1 |−1 [ s∈S m+1 A s − X ∅̸=K⊂S m+1 (−1) |K|−1 \ k∈K A k Let us analyze the terms on the right-hand side of this equation: 23 Preprint 1. The factor (−1) |S m+1 |−1 is a known sign, since|S m+1 | = m + 1. 2. The term S s∈S m+1 A s is the cardinality of a union of m + 1 sets. Since S m+1 is a non-empty subset of indices, this value is known by the premise of the theorem. 3. Consider the sum P ∅̸=K⊂S m+1 (−1) |K|−1 T k∈K A k . EachK in this summation is a non- empty proper subset of S m+1 . Therefore, the size of each such K satisfies 1 ≤ |K| ≤ m. By the Inductive Hypothesis, for any such K (i.e., for any intersection of j sets where 1 ≤ j ≤ m), the cardinality T k∈K A k can be determined from the known union cardi- nalities. Consequently, every term in this summation, including its sign factor (−1) |K|−1 , is determinable. Since all components on the right-hand side of the equation are known or can be determined based on the theorem’s premise and the inductive hypothesis, the value of T s∈S m+1 A s can be determined. In conclusion, by the principle of mathematical induction, for any non-empty subset of indices S ⊆1, 2,...,n, the cardinality of the intersection T s∈S A s can be determined if the cardinality of any union S k∈K A k (for any non-empty K ⊆1, 2,...,n) is known. Lemma C.3 (Intersection Block Identification (Kong et al., 2023b)). We assume the following data- generating process: [v 1 , v 2 ] = g(c, s 1 , s 2 ),(12) v 1 = g 1 (c, s 1 ),(13) v 2 = g 2 (c, s 2 ),(14) where c ∈ C ⊂ R d c , s 1 ∈ S ⊂ R d s 1 , and s 2 ∈ S 2 ⊂ R d s 2 . Both g 1 and g 2 are smooth and have non-singular Jacobian matrices almost everywhere, and g is invertible. If ˆg 1 : Z → V 1 and ˆg 2 : Z → V 2 assume the generating process of the true model (g 1 ,g 2 ) and match the joint distribution p v 1 ,v 2 , then there is a one-to-one mapping between the estimate ˆ c and the ground truth c overC×S×S , that is, c is block-identifiable. Lemma C.4 (One-level Visual Concept Identification). Assume the process for visual concepts in (2) with L V = 1. If a model specificationθ V satisfies Condition 4.2, and an alternative specification ˆ θ V satisfies Conditions 4.2-i and 4.2-i, along with a sparsity constraint such that for corresponding ˆ Z and Z: n(Pa( ˆ Z))≤ n(Pa(Z)),(15) then, if both modelsθ V and ˆ θ V generate the same observed data distribution P (X), the latent visual concepts Z 1 are component-wise identifiable for every level. Proof. For notational convenience, we denote Z 1 as S and D as U in this proof. This proof consists of two steps. In step one, we identify the connectivity between U and S variables. In step two, we further show the identifiability of the blocks resulting from intersecting the parent sets Pa(U ) of multiple U variables. Step 1: connectivity identification. Since we have access to the joint distribution P (S, U), we can derive conditional distributions P (S|U i i∈H ) for any index subsetH ⊆ [n(U)]. By Lemma C.1, we can identify the dimensionality of the set of variables S that are connected to any variable in U i i∈H for anyH⊆ [n(U)]. Lemma C.2 implies that we can identify the dimensionality of the set of variables S that are connected to all variables inU i i∈H for anyH⊆ [n(U)]. This information gives rise to a partition of S components, in which each part is connected to the same set of U variables. Therefore, we have identified the bipartite graph between S and U up to a permutation. Step 2: intersection block identification. Denote the indices of S variables that are connected to U i asI(i)⊆ [n(S)]. We denote the block of S components connected to all variables inU i i∈H as S ∩ i∈H I(i) for anyH ⊆ [n(U)]. Thanks to Lemma C.1, we can identify the block S I(i) connected to the variable U i for any i ∈ [n(U)]. Lemma C.3 allows us to identify the intersection of any two blocks S I(i)∩I(j) for i ̸= j. Therefore, repeated applications of Lemma C.3 leads to the identifica- tion of the intersection block S ∩ i∈H I(i) for anyH⊆ [n(U)]. This concludes the proof. 24 Preprint Condition 4.2 (Visual Concept Identification Conditions). i Informativeness: There exists a diffeomorphism g l : (Z l ,ε l ) 7→ X for l ∈ [0,L], whereε l denotes independent exogenous variables. i Smooth Density: The probability density function p(z l+1 |z l ) is smooth for any l∈ [L V ]. i Sufficient Variability: For each Z and its parents ̃ Z := Pa(Z), at any value ̃ z of ̃ Z, there exist n( ̃ Z) + 1 distinct values of Z, denoted asz (n) n( ̃ Z) n=0 , such that the vectors w( ̃ z,z (n) )− w( ̃ z,z (0) ) are linearly independent where w( ̃ z,z) = ∂ logp( ̃ z|z) ∂ ̃z 1 ,..., ∂ logp( ̃ z|z) ∂ ̃z n( ̃ z) . iv Sparse Connectivity (Minimality): For each parent concept ̃ Z, there exists a subset of its children Z⊆ Ch( ̃ Z) such that their only common parent is ̃ Z, i.e., T Z∈Z Pa(Z) = ̃ Z. Theorem 4.3 (Visual Concept Identification).Assume the process for visual concepts in (2). If a model specificationθ V satisfies Condition 4.2, and an alternative specification ˆ θ V satisfies Condi- tions 4.2-i and 4.2-i, along with a sparsity constraint such that for corresponding ˆ Z and Z: n(Pa( ˆ Z))≤ n(Pa(Z)),(3) then, if both modelsθ V and ˆ θ V generate the same observed data distributionP (X), the latent visual concepts Z l are component-wise identifiable for every level l∈ [L V ]. Proof. By Lemma C.4, we can identify the set of variables Z 1 that are directly connected to the text variables D and their causal graph. Treating the identified Z 1 as the U in Lemma C.4, we can further identify Z 2 . Repeating this procedure yields the identifiability of the entire model. C.2PROOF FOR THEOREM B.2 Definition C.5 (Non-negative Rank). The non-negative rank of a non-negative matrix A ∈ R m×n is equal to the smallest number p such that there exists a non-negative m× p-matrix B and a non- negative p× n-matrix C such that A = BC. Lemma C.6 (Conditional Independence and Nonnegative Rank (Cohen & Rothblum, 1993)). Let P∈ R m×n be a bi-variate probability matrix. Then its non-negative rank rank + (P) is the smallest non-negative integer p such that P can be expressed as a convex combination of p rank-one bi- variate probability matrices. Lemma C.7 (One-level Textual Concept Identification). Assume the hierarchical process as per (4) with L T = 1. Let the true underlying parameters beθ T . Ifθ T satisfies Condition B.1, and an alternative learned model ˆ θ T satisfies Condition B.1-i, then if both models produce the same observed distribution P (D), the latent textual concepts S 1 are component-wise identifiable. Proof. For each observed variable D, we search for the minimal set of variables C⊆ (D ) such that the following conditional independence holds: D ⊥ D\(D∪ C) | z R |(C, Ch(D)).(16) Note that all D, C, and R belong to observed variables, and Ch(D) is latent. Thanks to Condi- tion B.1-i and Lemma C.6, we can select C with which the nonnegative rank of the probability table T D,D\(D∪ C) |z R |C is strictly smaller than the support size of D. We argue that such C is the group of variables adjacent to the same variable S at the next level as D. In other words, they are the co-parents of D, CoPa(D). This is because such C makes 16 hold and thus CoPa(D) ⊆ C. Otherwise, there would be open paths passing S that induce dependence between D and CoPa(D), violating the conditional inde- pendence relation in (16). Therefore, the minimality constraint would enforce that C = CoPa(D). Repeating this procedure to all D ∈ D, we can construct S variables at the next level and the adjacency relations between S and D. 25 Preprint We proceed to identify the function D7→ S. We refer to D as a pure parent if D is adjacent to only one variable S in the discovered graph. For each S, we denote its pure parents as D S and non-pure parents as ̃ D S . We employ the conditional independence relation D S ⊥ D\ Pa( ̃ D S )|(S, ̃ D S ) and Condition B.1-i to identify the value of S, i.e., the function f S := (d S , ̃ d S )7→ s. We first make use of the conditional independence D S ⊥ D\ Pa(S)|(S, ̃ D S )(17) to merge the states of pure parents D S conditioned on the non-pure parents ̃ D S .Specifi- cally, we condition on non-pure parents ̃ D S = ̃ d S for any ̃ d S present in the support. We define an equivalence relation ∼ over values of (D S , ̃ D S ) where (d S 1 , ̃ d S ) ∼ (d S 2 , ̃ d S ) iff they give rise to an identical conditional distribution P D\ Pa(S)|D S = d S 1 , ̃ D S = ̃ d S = P D\ Pa(S)|D S = d S 2 , ̃ D S = ̃ d S . We further resort to a more global conditional independence by considering (D S , ̃ D S ) as a meta- variable and all the children Ch( ̃ D S ) associated with this meta-variable: (D S , ̃ D S )⊥ D\ Pa(Ch( ̃ D S ))|(Ch( ̃ D S ), Pa(Ch( ̃ D S ))\D S , ̃ D S ) | z := ̃ ̃ D S ,(18) where(D S , ̃ D S )hasbecomeapureparentofthelatentvariableCh( ̃ D S ). Wefurthergroupvalues([d S ], ̃ d S )followingtherulethat([d S ] 1 , ̃ d S 1 ) ∼ ([d S ] 2 , ̃ d S 2 )iff P D\ Pa(Ch( ̃ D S ))|([D S ], ̃ D S ) = ([d S ] 1 , ̃ d S 1 ), ̃ ̃ D S = ̃ ̃ d S = P D\ Pa(Ch( ̃ D S ))|([D S ], ̃ D S ) = ([d S ] 2 , ̃ d S 2 ), ̃ ̃ D S = ̃ ̃ d S for each ̃ ̃ d S on the support. That is, conditioning on any ̃ ̃ d S on the support, ([d S ] 1 , ̃ d S 1 ) and ([d S ] 2 , ̃ d S 2 ) cannot be distinguished. Thus, we group them into an equivalence class [(d S , ̃ d S )]. Finally, for each equivalent class [(d S , ̃ d S )], we assign a distinct value ˆs. This constitutes a function ˆ f S := (d S , ̃ d S ) 7→ ˆs. Due to the deterministic relation from latent variables and their children in (1), ˆ f S is well-defined. We denote the random variable ˆ S := ˆ f S (D S , ̃ D S ). In the following, we show that ˆ S and S are equivalent up to a bijection. We show this by con- tradiction. Suppose that there existed (s 0 , ˆs 0 ) on their respective support, such that their pre- images partially overlapped (d S 0 , ̃ d S 0 ) ∈ ˆ f −1 S (ˆs 0 ) ∩ f −1 S (s 0 ) and ˆ f −1 S (ˆs 0 ) ̸= f −1 S (s 0 ), where f S : (d S , ̃ d S ) 7→ s represents the true model. Suppose that f −1 S (s 0 ) missed some elements in ˆ f −1 S (ˆs 0 ), i.e., ∃(d S 1 , ̃ d S 1 ) ∈ ˆ f −1 S (ˆs 0 )\ f −1 S (s 0 ). In this case, (d S 0 , ̃ d S 0 ) and (d S 1 , ̃ d S 1 ) would lead to distinct values s 0 and s 1 under model f −1 S . By the construction of ˆ f −1 S , this would indicate P (D\ Pa(S)|S = s 0 ) = P (D\ Pa(S)|S = s 1 ) and P D\ Pa(Ch( ̃ D S ))|S = s 0 , ̃ ̃ D S = ̃ ̃ d S = P D\ Pa(Ch( ̃ D S ))|S = s 1 , ̃ ̃ D S = ̃ ̃ d S for each ̃ ̃ d S on the support. Since s 0 ̸= s 1 , this violates Condition B.1-i, giving rise to a contradiction. Suppose that f −1 S (s 0 ) contains additional elements,i.e., ∃(d S 2 , ̃ d S 2 ) ∈ f −1 S (s 0 ) \ ˆ f −1 S (ˆs 0 ).Inthiscase,(d S 0 , ̃ d S 0 )and(d S 2 , ̃ d S 2 )wouldleadtoonevalue s 0 undermodel f −1 S .Bytheconstructionof ˆ f −1 S ,thiswouldindicateei- ther P D\ Pa(S)|D S = d S 0 , ̃ D S = ̃ d S 0 ̸= P D\ Pa(S)|D S = d S 2 , ̃ D S = ̃ d S 2 orP D\ Pa(Ch( ̃ D S ))|([D S ], ̃ D S ) = ([d S ] 0 , ̃ d S 0 ), ̃ ̃ D S = ̃ ̃ d S ̸= P D\ Pa(Ch( ̃ D S ))|([D S ], ̃ D S ) = ([d S ] 2 , ̃ d S 2 ), ̃ ̃ D S = ̃ ̃ d S for some ̃ ̃ d S on the support. By construction of ˆ f S , this would violate conditional independence (17) or (18) which the graphical structure implies, which leads to a contradiction. 26 Preprint Therefore, we have shown that for each pair (s, ˆs) on their respective support, their pre-images should be identical as long as they intersect: ˆ f −1 S (ˆs)∩ f −1 S (s)̸=∅ =⇒ ˆ f −1 S (ˆs) = f −1 S (s), which is equivalent to that ˆ S and S are equivalent up to a bijection. Condition B.1 (Textual Concept Identification Conditions). i Natural Selection: Each selection variable S l has a support supp(S l ) that is a proper subset of its potential range if its constituent parts (lower-level variables) were combined randomly. That is, supp(S l ) ⊊ f D→S l (Ω n(Pa(S l )) ), where f D→S l is the function from D to S l . i Bottlenecks: The support size of any concept S l is strictly smaller than the joint support size of its parents Pa(S l ) in the selection graph. i Minimal Supports: For any S, the condition distribution P (D\ Pa(S)|S = s, HPa(S) = ̃ s) is a one-to-one function w.r.t. the argument s. iv No-Twins: Distinct latent variables must have distinct sets of adjacent (parent/child) variables. v Maximality: The identified latent structure is maximal in the sense that splitting any latent concept variable would violate either the Markov conditions or the No-Twins condition. Theorem B.2 (Textual Concept Identification). Assume the hierarchical process as per (4). Let the true underlying parameters beθ T . Ifθ T satisfies Condition B.1, and an alternative learned model ˆ θ T satisfies Condition B.1-i, then if both models produce the same observed distribution P (D), the latent textual concepts S l are component-wise identifiable for every level l∈ [L T ]. Proof. By Lemma C.7, we can identify the set of variables S 1 adjacent to D and the bipartite causal graph between these two sets of variables. We then employ the identified S 1 to serve as D in the first step to identify S 2 . Repeating this procedure yields the identifiability of the entire model. DKEY CONCEPT DISCUSSIONS The roles and purposes of “Selection-based hierarchy and causality minimality”. The selection- based hierarchy and causal minimality are constraints on the natural data distribution (images or text), which is a standard modeling practice in causal representation learning (Sch ̈ olkopf et al., 2021). Specifically, the selection-based hierarchy considers concepts as effects of their constituent parts (Zheng et al., 2024), while causal minimality assumes this underlying causal graph is sparse in a specific way (e.g., Condition 4.2-iv). “Innate” hierarchical concept graphs. “Innate” refers to the causal structure inherent in the natural data-generating process itself. Latent concepts in the real world interact (e.g., ‘eyes’ and ‘nose’ are components of a ‘face’), forming a pre-existing causal structure which we refer to as the ”innate concept graph.” True latent variables and their verifications. “True latent variables” follow the standard notion in causal representation learning (Sch ̈ olkopf et al., 2021): they are the disentangled, interpretable, semantic factors of the real-world data-generating process (e.g., age, object pose). This is in contrast to a deep learning model’s learned features, which are often an entangled, uninterpretable mixture optimized for a specific training objective. Aligning learned features with true latent variables (re- ferred to as “identification”) is the central goal, as it enables reliable interpretation (e.g., “this feature is age”) and precise control (e.g., “increase this feature to make the face older”). This is a fundamen- tal question that our work addresses through both theoretical guarantees and empirical validation. Our work provides the guarantee that if the data-generating process fulfills the property of causal minimality and our learning objective enforces this (e.g., via sparsity), the model’s learned features are provably equivalent to the true latent variables. We then validate this empirically via interven- tion, a standard practice in causal research (Sch ̈ olkopf et al., 2021). Our experiments (Figure 3 and Figure 8) show that manipulating the theoretically identified features provides semantic control over the generated output, providing evidence that these features are the meaningful causal levers of the generative process. Validity of the conditions. While assumptions on the unobserved data-generating process may not be validated directly, we have reasoned for the plausibility of our conditions by reflecting on natural 27 Preprint properties of real-world data. Beyond standard regularity assumptions like smoothness and variabil- ity (Khemakhem et al., 2020a;b; Hyvarinen & Morioka, 2016; Hyvarinen et al., 2019; Zhang et al., 2024a), our key minimality conditions—Sparse Connectivity (Condition 4.2-iv) for vision and Mini- mal Supports (Condition B.1-i,i) for text—are motivated by the observation that concepts typically arise from a sparse set of causes (Lachapelle et al., 2024a; Xu et al., 2024; Lachapelle et al., 2022b; Zheng et al., 2022; Moran et al., 2021) and that language is inherently structured and compress- ible (Shannon, 1948; Cover, 1999). Perhaps a more convincing validation is the empirical results. Our experiments provide strong indicative support for these assumptions: by actively enforcing spar- sity/compression via SAEs, we successfully extract meaningful concept hierarchies in both vision (Figure 3) and text (Figure 8) that are otherwise dense and not easily interpretable. This success provides support for the usefulness of our overall approach and the validity of our assumptions. We acknowledge that these assumptions, like any in this field, may not hold universally. Fortunately, our strong empirical results suggest they seem effective and plausible for the complex, real-world data we study. Concept variable interpretation. Our theory proves the existence of a clean, one-to-one mapping between a learned feature and a true latent variable. This guarantee is what makes a principled interpretation possible in the first place. The subsequent step—assigning a human-understandable description to this now-identified concept—is intrinsically a task that requires human validation. This is a fundamental aspect of all interpretability research (perhaps modern vision-language models have the potential to automate this process). Comparison with recent work (Cywi ́ nski & Deja, 2025). On the technique side, Cywi ́ nski & Deja (2025) feature an elegant concept location technique by utilizing the score function, which could significantly benefit our algorithm. For example, we could employ SAeUron (Cywi ́ nski & Deja, 2025) to confirm whether our features at various timesteps match the concept location it iden- tifies. Our causal learning algorithm explicitly learns the inter-connectivity among concepts across hierarchical levels. Thus, to modify a part of a high-level concept, we could focus our scope on only the variables connected to this specific high-level concept, which lowers the search complex- ity. In our experiment example, to implement two changes, “replacing the rock with tree stump” and “adding texture to tree stump”, SAeUron may need to perform two independent searches across all timesteps and node indices. Our method can help reduce the search space to only the low-level nodes connected to “tree stump”. In addition, pinpointing specific diffusion timesteps to intervene on potentially aids in managing undesirable artifacts. Moreover, our explicit concept graph could also give an interpretable, intuitive characterization of the model’s knowledge. On the message side, Cywi ́ nski & Deja (2025) propose a novel score function to select the timestep and node index for accurate concept unlearning. Our work’s focus is to provide concise and informative theoretical conditions to understand concept learning in both vision and language modalities, with potential applications like concept easing or controllable generation. With this work, we hope the theoretical insights will facilitate the development of refined and dedicated methods in the community. Comparison with recent work (Kim et al., 2024). Revelio (Kim et al., 2024) relies on training a classifier on a specific classification dataset. Revelio trains SAEs and a classifier on a specific dataset (e.g., Caltech-101) to evaluate which features and timesteps are most correlated with class labels. Our work, in contrast, does not involve class labels. Our primary contribution is a hierarchi- cal, causal framework designed to interpret the generative process itself. We apply causal discovery algorithms to discover the causal relationships across different levels of concepts without any class labels. We are able to understand how semantic concepts causally relate to one another across dif- ferent levels of abstraction to form a coherent output (e.g., how “ear” and “mouth” features causally contribute to a “cat face”). Moreover, Kim et al. (2024) do not perform interventions or analyze the compositional structure of generation, which are the central themes of our paper. EIMPLEMENTATION DETAILS We present the diagram of our method in Fig.9. Annotation of the concepts To annotate the concepts discovered by SAEs, we use a two-step pro- cess: 1. Identify concept-related features. For a target concept (e.g., nudity in the unlearning task), we collect a set of prompts related to that concept, generate the corresponding images, and extract the top-K activated feature indices that are consistently triggered across these samples. These shared 28 Preprint indices are treated as related to the target concept. 2. Explore causal relationships. After identify- ing a node (feature) with first step, we use the inferred causal graph to find its parent and child nodes—features closely related to that concept. We then visualize the node’s attribute map (e.g., distinct regions of the cat in Fig.1) to interpret and confirm the concept’s semantics. Computing resources. We use one L40 GPU for training the SAEs and a standard MacBook Pro with an M1 chip for causal discovery. Training one SAE takes around 8 hours. Vision experiments. For the diffusion sampling process, we utilize the sde-dpmsolver++ (Lu et al., 2022) sampler, which adds stochasticity between successive steps. We train the K-sparse autoencoder using a latent dimension of 5120, a batch size of 4096, and the Adam optimizer with a learning rate of 0.0001, settingK = 10. We use prompts from the Laion-COCO dataset (Schuhmann et al., 2022). Our causal discovery procedure consists of the following steps: 1. Identify key features for each SAE: For every SAE trained at a specific noise level, we first extract the top-K feature indices that show the highest average activation on the 10K LAION-COCO subset dataset. 2. Construct binary feature representations: We then enumerate all unique feature indices across samples. For each sample, we create a binary feature vector where a value of 1 indicates that the corresponding feature index appears in the sample’s top-K list, and 0 otherwise. This results in a feature–index matrix representing the activation pattern of features for each SAE at each noise level. 3. Apply causal discovery: Using the constructed matrices, we employ the classical causal discovery PC algorithm to infer the causal structure among the feature indices across noise levels. [PC] identifies potential directional dependencies, revealing how certain features may causally influence others. For the sparsity ablation study, we control the top-K value used in the SAE. Specifically, we train additional SAEs with K=4 and K=100 at timestep 500. To evaluate the effect of sparsity (Fig- ure 4), we then perform causal discovery by replacing the SAE features with K=10 with those from the K=4 or K=100 models. Table 1 evaluates the following baselines: SD1.4 (Rombach et al., 2022b), ESD (Gandikota et al., 2023), SA (Heng & Soh, 2023), CA (Kumari et al., 2023), MACE (Lu et al., 2024), UCE (Gandikota et al., 2024), RECE (Gong et al., 2024), SDID (Li et al., 2024), SLD-MAX (Schramowski et al., 2023), SLD-STRONG (Schramowski et al., 2023), SLD- MEDIUM (Schramowski et al., 2023), SD1.4-NegPrompt (Rombach et al., 2022b), SAFREE (Yoon et al., 2024), TRASCE (Jain et al., 2024), and ConceptSteer (Kim & Ghadiyaram, 2025). LLM experiments. We utilize the pretrained SAEs for gemma-2-2b-it available from Gemma- Scope (Team, 2024). To collect features, we use the pile-10k corpus (Gao et al., 2020). For each sample, we first exclude padding tokens and divide the remaining meaningful tokens into three sequential segments. The first segment is processed through the SAE at layer 18 to obtain feature indices representing lower-level information. The second segment is passed through the SAE at layer 19 to capture intermediate-level features. The final segment is input to the SAE at layer 20 to extract higher-level features. We then apply the PC algorithm for causal discovery using the feature indices from these three representational levels. FADDITIONAL EMPIRICAL RESULTS Extension to Flux.1 Our main experiments are conducted on Stable Diffusion V1.4, which adopts a U-Net architecture. To further validate the generality of our approach, we extend it to Flux.1- Schnell, a 12B text-to-image DiT model. Specifically, we extract features at timesteps 0, 1, 2, and 3 (the model performs inference in only four steps, as it is a distilled model) and train SAEs with the following settings: batch size 4096, learning rate 0.0001, latent dimension 12,288, and top-k = 20. Each SAE is trained on the LAION-COCO dataset for 20,000 steps. We use the last double- stream transformer block (out of 18 double-stream and 38 single-stream blocks) as the feature space (3072 dimensions). We then perform causal discovery to identify causal dependencies among fea- tures. For evaluation, following the setup used for SD1.4, we test our method on the unlearning 29 Preprint Figure 9: Diagram of our interpretability method. We train SAEs to capture features at different levels (timesteps for diffusion models and token positions for LLMs), and apply causal discovery to construct a hierarchical concept graph. InputSD1.4 w/o hierOurs Figure 10: Examples of controllable image generation. 30 Preprint MethodI2P↓RING-A-BELL↓P4D↓ UATK↓COCO K77K38K16AVGFID↓ CLIP↑ Flux3.08 50.53 51.58 52.63 51.58 27.1519.7222.8931.57 Ours-Flux 0.94 11.58 5.264.217.013.314.9324.4031.54 Table 4: Model unlearning performance on Flux.1-Schnell. The Flux text-to-image model is susceptible to malicious prompts in benchmark datasets, often producing images containing nudity. By training SAEs on Flux features at different timesteps, we identify the latent representation of the nudity concept and apply negative feature steering to suppress it. This effectively reduces nudity generation while maintaining competitive text-to-image performance on normal prompts from the COCO dataset. benchmark datasets (Table 4). In particular, we apply negative feature steering at feature index 4390 on timesteps 1, 2, and 3. Our method achieves significantly lower attack success rates on mali- cious nudity prompts across all benchmarks, demonstrating its robustness and effectiveness on DiT architectures. Add fire, + Feature 9678 Add Mountain, + Feature 4656 Figure 11: Context-sensitive and objective-agnostic concepts. Steering the same concept (“fire” at the top row and “mountain” at the bottom row) yields visual changes in the original context, demonstrating that our learned concepts are context-sensitive and shared across objects/classes. More examples for Figure 3. Figure 12 and Figure 13 contain more examples of Figure 3. For example, node 3641 in the SAE at timestep 899 contains comprehensive information about the panda, as illustrated by the heatmap. When feature steering is applied, it results in the generation of a new panda. Meanwhile, nodes 1026 and 511 in the SAE at timestep 500 represent different components of the panda. At a finer level of detail, nodes 3489, 3880, and 451 in the SAE at timestep 100 capture specific image features. These hierarchical concept graphs effectively illustrate how the panda is generated. More results for model unlearning In addition to the four benchmark datasets in the main paper, we report results on another commonly used benchmark dataset with two tasks: Remove Van Gogh and Remove Kelly McKernan in Table.5. We evaluate performance using four metrics: LPIPSe (similarity for prompts with the target style), LPIPSu (similarity for prompts without the style), Acce (how well the target style was removed), and Accu (how well other styles were preserved), with accuracy ratings assessed using GPT-4o. Our method achieves competitive performance across all metrics and tasks. Understanding the sparsity constraint. Figure 14 and Table 6 contain the ablation study for the sparsity constraint. We can observe that a proper sparsity strength can indeed give rise to desirable interpretability results, while too small and too large sparsity constraints may be harmful in practice. As shown in Table 6, a low sparsity penalty results in visualized maps with significant overlap. On the other hand, applying a strong sparsity penalty leads to low node coverage, indicating that the nodes alone are insufficient to fully explain the generation of the entire image. More examples for Figure 8. Figure 15 contains more examples for Figure 8. As discussed in the main paper, we divide the tokens into three segments based on their sequence order, with later 31 Preprint 3489 1026 511 3880 451 36413641 1026 511 3489 3880 451 4456 704916 40114771 4456 704916 40114705 Figure 12: Discovered hierarchical concept graphs and feature steering visualization for text- to-image generation. We can observe that features on the hierarchical model represent a part-whole relation, and steering a feature yields corresponding visual variation (e.g., the panda’s ears). 32 Preprint Figure 13: More examples of the learned hierarchical concept graphs for text-to-image models. Under appropriate sparsity and noise conditions, our method successfully recovers meaningful hierarchical structures, where each node encodes distinct semantic concepts. On the right, we demonstrate feature steering, where manipulating individual nodes leads to changes in the output that align with their position in the hierarchy – higher-level nodes produce broader semantic shifts, while lower-level nodes control more fine-grained aspects. MethodLPIPSe↑LPIPSu↓Acce↓Accu↑ Task: Remove “Van Gogh” SD-v1.4–0.950.95 CA (Kumari et al., 2023)0.300.130.650.90 RECE (Gong et al., 2024)0.310.080.800.93 UCE (Gandikota et al., 2024)0.250.050.950.98 SLD-Medium (Schramowski et al., 2023)0.210.100.950.91 SAFREE (Yoon et al., 2024)0.420.310.350.85 Ours0.530.260.300.88 Task: Remove “Kelly McKernan” SD-v1.4–0.800.83 CA (Kumari et al., 2023)0.220.170.500.76 RECE (Gong et al., 2024)0.290.040.550.76 UCE (Gandikota et al., 2024)0.250.030.800.81 SLD-Medium (Schramowski et al., 2023)0.220.180.500.79 SAFREE (Yoon et al., 2024)0.400.390.400.78 Ours0.480.200.350.81 Table 5: Results on style removal. We apply negative feature steering to the node to suppress the styles in the image. tokens expected to encode higher-level information—consistent with the behavior of autoregressive language models. At the highest level, node 11859 represents the ”yell mode,” characterized by capitalized words conveying a strong tone. The green node 1033, located at an intermediate sequence position, emphasizes importance or intensity—typically a component of the yell mode. At the lowest level, nodes 304, 2009, and 2818 capture various aspects and meanings related to the concept of importance. 33 Preprint 3489 1026 511 3880 451 3641 2093 451 3641 3390 3951 3641 3880 451 2488 762 SparsityIncreasing 3499 Figure 14: Understanding the sparsity constraint. We adjust the top-K value in the SAE at timestep 500 to control the level of sparsity, effectively modifying the sparsity strength of the SAE at this middle layer. As sparsity decreases, the resulting graph becomes denser, introducing many redundant and semantically irrelevant edges. This reduces the overall interpretability of the concept graph. Conversely, increasing sparsity yields a cleaner, more concise graph. However, if sparsity is too high, it may hinder the formation of a complete and interpretable concept graph necessary for image generation. 11859 1033 304 capitalized words, acronyms, and code identifiers in technical or structured text. comparative phrases that express degrees of importance or intensity phrases related to accessibility and membership in organizations detailed descriptions of injuries and physical sensations 20092818 technical terms related to engineering or manufacturing processes Here'S HOW I THINK IT IS IMPORTANT TO BE CLEAR AND Itismoreimportantthanever. Anyone Atthelevelof heart Thecold,youor killme! Example: Figure 15: An example of a discovered hierarchical concept graph for autoregressive language modeling. Node 11859 represents a ”yell mode,” characterized by capitalized words that convey a strong tone. The green node 1033 captures the concept of emphasizing importance or intensity. Blue nodes correspond to lower-level information—for instance, node 304 represents entities mentioned throughout the text. 34 Preprint Overlap↓Coverage↑ K=40.108± 0.12826.37± 17.24 K=100.089± 0.07947.90± 12.50 K=1000.235± 0.13237.46± 17.31 Table 6: Quantitative ablation results. We generate 100 panda images using different random seeds and visualize the feature heatmaps at timestep 500. We adjust the top-K value in the SAE at timestep 500 to control the level of sparsity. To evaluate, we compute the intersection-over-union (IoU) of intermediate heatmaps to measure concept disentanglement, and the union of all features to assess coverage. IoU reflects how distinctly the intermediate concepts are represented, while coverage in percentage indicates the extent to which the intermediate nodes collectively account for the image generation. 35