Paper deep dive
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu, Wesam A. Sakla, Kowshik Thopalli
Models: LLaVA-1.5-7B, LLaVA-MORE, UnCLIP
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/11/2026, 1:09:10 AM
Summary
The paper introduces Concept Bottleneck Sparse Autoencoders (CB-SAE), a framework that addresses the limitations of standard Sparse Autoencoders (SAEs) in mechanistic interpretability. The authors identify that many SAE neurons lack interpretability and steerability, and that unsupervised SAEs often fail to capture user-defined concepts. CB-SAE improves these metrics by pruning low-utility neurons and augmenting the latent space with a concept bottleneck aligned to a user-specified concept set, resulting in significant performance gains in LVLMs and image generation tasks.
Entities (6)
Relation Signals (4)
CB-SAE â integrates â Sparse Autoencoders
confidence 100% ¡ CB-SAE is the first framework to unify sparse autoencoders with concept bottleneck models
CB-SAE â improves â Interpretability
confidence 95% ¡ The resulting CB-SAE improves interpretability by +32.1%
CB-SAE â improves â Steerability
confidence 95% ¡ The resulting CB-SAE improves... steerability by +14.5%
CB-SAE â steers â LLaVA
confidence 90% ¡ CB-SAE and baseline SAE can steer multiple downstream models like large vision-language models (LLaVA)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires that the learned features be both interpretable and steerable. To that end, we introduce two new computationally inexpensive interpretability and steerability metrics and conduct a systematic analysis on LVLMs. Our analysis uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (ii) due to the unsupervised nature of SAEs, user-desired concepts are often absent in the learned dictionary, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) - a novel post-hoc framework that prunes low-utility neurons and augments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB-SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks. We will make our code and model weights available.
Tags
Links
- Source: https://arxiv.org/abs/2512.10805
- Canonical: https://arxiv.org/abs/2512.10805
Trouble viewing inline? Open PDF directly â
Full Text
61,137 characters extracted from source content.
Expand or collapse full text
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders Akshay Kulkarni 1 Tsui-Wei Weng 1 Vivek Narayanaswamy 2 Shusen Liu 2 Wesam A. Sakla 2 Kowshik Thopalli 2 1 University of California, San Diego 2 Lawrence Livermore National Laboratory a2kulkarni,lweng@ucsd.edunarayanaswam1,liu42,sakla1,thopalli1@llnl.gov Abstract Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this poten- tial requires that the learned features be both interpretable and steerable. To that end, we introduce two new computa- tionally inexpensive interpretability and steerability metrics and conduct a systematic analysis on LVLMs. Our analysis uncovers two observations; (i) a majority of SAE neurons exhibit either low interpretability or low steerability or both, rendering them ineffective for downstream use; and (i) due to the unsupervised nature of SAEs, user-desired concepts are often absent in the learned dictionary, thus limiting their practical utility. To address these limitations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE)âa novel post-hoc framework that prunes low-utility neurons and aug- ments the latent space with a lightweight concept bottleneck aligned to a user-defined concept set. The resulting CB- SAE improves interpretability by +32.1% and steerability by +14.5% across LVLMs and image generation tasks. We will make our code and model weights available. 1. Introduction Sparse autoencoders (SAEs) [4,16,36] have emerged as a foundational approach for mechanistically interpreting deep neural networks. By mapping dense, polysemantic activa- tions into sparse, overcomplete latents, SAEs produce disen- tangled, monosemantic features that expose internal structure and support model behavior analysis. Deployed on large lan- guage models (LLMs) [21], large-scale vision models [40], and even with large vision-language models (LVLMs) [31] like LLaVA [53], SAEs promise a unified route to concept This work was performed under the auspices of the U.S. Department of Energy by the Lawrence Livermore National Laboratory under Contract No. DE-AC52-07NA27344, Lawrence Livermore National Security, LLC. and was supported by the LLNL-LDRD Program under Project No. 25-SI-001. LLNL-CONF-2013863 A. SAE baselineB. Concept Bottleneck SAE (Ours) #1. Most SAE neurons have low interpretability/steerability! #2. SAE may not discover user-specified concepts Enables user-specified concepts Interpretability: +32%, Steerability: +14% Prune SAE neurons with low interpretability & steerability CB-SAE Figure 1. A. We find that the majority of SAE neurons in vision models have low interpretability or steerability, with no guarantee of discovering user-specified concepts. B. Our CB-SAE addresses both limitations by pruning SAE neurons with low interpretability and steerability, and replacing them with a user-specified concept bottleneck that improves both interpretability and steerability. discovery and downstream model steering. However, real- izing this promise requires that SAE features not only align with semantically meaningful, human-understandable con- cepts but also causally influence model behavior i.e., that they are both interpretable and steerable. Recent work on SAEs [1,46] in the context of LLMs shows that interpretability does not guarantee steering ef- fectiveness, i.e. features that activate strongly for a human- understandable concept may fail to control it when inter- vened upon [1]. Although this trade-off has been observed in language models, its presence and implications in vision encoders and LVLMs remain largely unexplored. To inves- tigate this in the LVLM setting, we conducted an empirical study of SAEs trained on activations from the CLIP [35] vision encoder. We introduced metrics to quantify both in- terpretability and steerability at the neuron level and analyze their alignment across the SAEâs latent space. Our find- ings reveal that only about 19% of SAE neurons exhibit both high interpretability and steerability. Moreover, despite the SAEâs large dictionary size (65,536 neurons), it fails to represent 27â45% of concepts drawn from established arXiv:2512.10805v1 [cs.LG] 11 Dec 2025 Concept LLaVA steered o/p shop UnCLIP steered output dog house Concept butterfly net Retained SAE wedding wedding museum museum spots or stripes TV stand CB neurons bird bath blade of grass coins coins aisle of toys toy avocado fruit frame Vision Enc. LLM Image âwhat is in this image?â âgray colorâ Steering SAE or CB-SAE Image Gen. Steering SAE or CB-SAE Vision Enc. ImageImage Large Vision-Language Model (e.g. LLaVA) Image-to-Image Generator (e.g. UnCLIP) fire coast stack of paper Discarded SAE Figure 2. A. Our CB-SAE and baseline SAE can steer multiple downstream models like large vision-language models (LLaVA [23]) or image generative models (UnCLIP [38]). B. Examples of steering LLaVA and UnCLIP when using unit vector steering (zeroing out all SAE/CB-SAE neurons except the selected concept). ImageNet-derived benchmarks [41], even when trained on the corresponding data. This highlights the inability of unsu- pervised SAEs to cover user-specific concepts reliably. These findings surface two key limitations that constrain the practical utility of SAEs: (i) the inability to ensure com- prehensive coverage of semantically meaningful concepts, and (i) the lack of mechanisms for explicitly encoding user- defined concepts into the latent space to support better steer- ability. As a result, practitioners are left to work with the latent features the SAE happens to discover and searching post hoc for relevant activations, with no guarantee of align- ment with task-specific requirements. This motivates the need for a unified framework that supports both unsuper- vised discovery and user-guided specification. A counterpart to SAEs are Concept Bottleneck Models (CBMs) [18,29,39,50], which approach concept learn- ing from a supervised perspective. CBMs explicitly train a model to predict a fixed set of human-interpretable concepts by introducing a bottleneck layer that mediates the final prediction. This enables guaranteed concept coverage, but limits CBMs to predefined concepts, preventing discovery of novel features unlike SAEs. These complementary strengths highlight the need for a unified framework that combines CBMsâ controllability with SAEsâ discovery capabilities. Motivated by these observations, we propose Concept Bottleneck Sparse Autoencoders (CB-SAE) â a unified framework that combines the unsupervised discovery ca- pabilities of SAEs with the controllability of concept bottle- necks. We begin by pruning SAE features that lack inter- pretability and steerability, and then augment the resulting latent space with a lightweight CB autoencoder [19], trained to align with a user-specified concept set (Fig. 1B). The model is optimized using tailored loss functions that pre- serve reconstruction fidelity, interpretability, and steerability. As a result, our CB-SAE produces latent features that are both semantically meaningful and causally effective. We evaluate CB-SAE on two challenging downstream tasks: controlled text generation via visionâlanguage models (LLaVA-1.5-7B [22], LLaVA-MORE [10]) and controlled image synthesis using UnCLIP [38]. CB-SAE consistently outperforms standard SAEs, with average gains of +32.1% in interpretability and +14.5% in steerability across all mod- els and metrics. This improved performance is also shown qualitatively in Fig. 2B, where the retained SAE and CB neu- rons consistently outperform discarded SAE neurons w.r.t. steerability. To our knowledge, CB-SAE is the first frame- work to unify sparse autoencoders with concept bottleneck models, enabling robust interpretation and control of vision representations across modalities and architectures. 2. Related Work Sparse Autoencoders. SAEs aim to discover interpretable features in neural networks by learning overcomplete decom- positions of activations [25]. Recent work [6,14] showed SAEs can decompose LLM representations into monose- mantic features. Various architectural innovations improved SAEs, like Batch-Top-ksparsity [7], JumpReLU [36], and Matryoshka SAEs [8] with multi-level feature hierarchies. Large-scale efforts trained LLM SAEs across multiple lay- ers and models [13,21], with systematic benchmarks [16]. However, we uncover two key limitations of SAEs: their unsupervised training does not guarantee the discovery of user-desired concepts, and many SAE neurons exhibit low interpretability or utility in downstream steering [1]. Concept Bottleneck Models. CBMs [18,50] provide a framework for building interpretable models by constraining predictions through a human-understandable concept layer, enabling both interpretation and steering. This approach has been extended to label-free settings [29], enhanced with vision-language guidance [39,47,49], applied to image generative models [15,19] as well as LLMs [42]. Our work bridges SAEs and CBMs into our novel CB-SAE, combining the expressiveness of overcomplete feature decomposition with user-specified concepts, steerability, and interpretability of concept-guided learning. A concurrent work, AlignSAE [48], independently devised a similar approach to introduce supervised concepts in SAEs. They attempt to disentangle the supervised concepts from the unsupervised SAE neurons with an orthogonality loss, while our approach explicitly prunes the low utility SAE neurons and only introduces the supervised concepts absent from the retained SAE neurons. Further, AlignSAE focuses on text-based LLM SAEs while we focus on vision SAEs for multimodal LLMs and image- to-image generative models. SAEs for Vision and Vision-Language Models. Recent work showed that SAEs can learn interpretable, monoseman- tic features in vision models [40] as well as vision-language models [31,51]. Another line of work [27,32,45] inves- tigated how visual information maps to language feature spaces via SAEs for cross-modal interpretability [24,26]. However, these approaches typically neither address the chal- lenges of ensuring discovered features are both interpretable and steerable, nor do they guarantee the discovery of user- specified concepts. Our CB-SAE addresses both limitations through post-hoc pruning and concept-bottleneck training. 3. Background SAE preliminaries. Letv = f l (x)â R d denote the dense activations from layerlof a deep pre-trained vision model (e.g., CLIP image encoder [35])ffor an input imagexâ X. Hereddenotes the activation dimension andXcorresponds to the space of images. SAEs decomposes the polysemantic activationsvinto sparse, overcomplete latent representations z â R Ď ( Ď d >> 1) with the aim of associating every unit inzto distinct, interpretable concepts. Here Ď d corresponds to the expansion factor of the SAE [13]. Formally, an SAE is parameterized by a linear encoderE sae â R ĎĂd , a linear decoderD sae â R dĂĎ , a shared bias termb â R d , and a non-linear activation function Ď sae : R Ď â R Ď : z = Ď sae (E sae (vâ b))(1) Ëv = D sae z + b(2) The SAE training objective is given byL r = ||vâ Ëv|| 2 2 + Îť||z|| 1 , whereÎť ⼠0balances reconstruction fidelity and sparsity, whereËvrepresents the SAE reconstruction. In addi- tion to standardâ 1 regularization, sparsity can be enforced directly via the activation functionĎ sae (¡), such as top-k[13], batch top-k[7], or ReLU with a learnable threshold [21,36]. Measuring SAE interpretability. After training, SAEs are typically evaluated using reconstruction fidelity or spar- sity [16,24]. However, these metrics do not directly quantify interpretability which is namely the extent to which indi- vidual SAE neurons correspond to human-understandable concepts. While existing work [17,31] relies on manual inspection of top-kactivating inputs or autointerpretabil- ity scores [3] that depend on external language models or measuring the monosemanticity [31], they are often subjec- tive, computationally expensive, and difficult to scale. To this end, we leverage a popular neuron interpretability tool CLIP-Dissect [28], which utilizes a user-specified concept setCand a pretrained vision-language model to assign each neuronjof an SAE to a human-interpretable text concept c j . This approach used first time in the context of SAEs is computationally inexpensive and scalable. Please refer to Appendix for further details on CLIP-Dissect. Measuring SAE steerability.Beyond interpretability, SAEs have been shown to enable controllable manipula- tion of model behavior across language [1], large-scale vi- sion [40], and large vision-language models [24,31] such as LLaVA [23]. Steerability refers to the ability to influ- ence model outputs through targeted modifications of SAE neuron activations, thereby inducing semantically consistent changes [1]. It is typically quantified by measuring the align- ment between the steered output and the concept label of the intervened neuron. Since, in our study, the base modelfis a vision-only encoder, we employ a downstream generative model to evaluate the effect of SAE latent interventions. Following [31], we adopt LLaVA [23], which maps an im- ageâtext pair(x,t)to a text outputo. In LLaVA, the vision encoderf(e.g. CLIP [35]) produces visual tokensv i N i=1 , which are projected by an adapter into the LLMâs word em- bedding space, where they are combined with prompt tokens to generateo. Similar to [31], to probe steerability, we use a white image with the prompt âWhat is shown in this image? Use exactly one word!â. For a given target neuronjâ [Ď] of the SAE, we overwrite its activation across all tokens (from this white image) with a fixed valueÎą, reconstruct the modified latents Ěv i through the SAE decoder, and feed them into LLaVA to produce the steered output Ěo j . We then compute the cosine similarity between Ěo j and the neuronâs CLIP-Dissect-assigned conceptc j in a sentence- transformer embedding space [37]. Higher similarity indi- cates greater steerability, as the neuron reliably drives the output toward its associated concept. Unlike [31], which compared steered outputs to top-activating images in CLIPâs image-text space, our method compareso j with concepts identified by CLIP-Dissect, which produces these descrip- tions by aggregating activations across all images thus yield- ing a more robust, semantically grounded steerability metric. Note that while steering (image,text)-to-text LLaVA [23] is one way to compute a steerability metric, a similar metric can be computed using an image-to-image generator like UnCLIP [38] (see Appendix for analysis with UnCLIP). 4. Interpretability vs Steerability in SAEs In this section, we empirically analyze SAEs from two com- plementary perspectives: (i) their capacity to capture inter- pretable concepts and (i) their ability to steer model outputs and quantify their trade-offs. While these two measures i.e., interpretability and steerability are related, they capture dis- tinct yet complementary aspects of SAE behavior [46]. In practice, interpretable neurons may not be steerable if their activations are weakly causal or entangled, while steerable neurons may encode abstract features misaligned with user objectives [1,46]. The insights from this analysis on large vision-language models motivate our hybrid framework, in- troduced in Sec. 5, which integrates SAEs with principles from concept bottleneck models [18, 29, 39]. Expt. 1: Are all SAE neurons interpretable & steerable? Neuron 23736 CLIP-Dissect: âRubbleâ Top activating images: LLaVA steering o/p: âShadowâ Neuron 47646 CLIP-Dissect: âPlantâ Top activating images: LLaVA steering o/p: âFruitâ Neuron 10458 CLIP-Dissect: âFile cabinetâ Top activating images: LLaVA steering o/p: âFaxâ Neuron 2435 CLIP-Dissect: âDishwasherâ Top activating images: LLaVA steering o/p: âDishwasherâ Low steer. Low interp. Low steer. High interp. High steer. High interp. High steer. Low interp. Figure 3. We analyze the interpretability and steerability of 65,536 neurons of an SAE trained for a CLIP image encoder. We also visualize the CLIP-Dissect assigned concept, top-activating images, and LLaVA steering outputs for some characteristic neurons. The dashed lines indicate the average scores along each axis, and we observe that most SAE neurons have either low interpretability, low steerability, or both. Setup. We train a Matryoshka Batch Top-kSAE [7,51] with Ď= 65536 and an expansion factor of 64 on layerl= 22 of the CLIP-ViT-L/14-336 vision encoder [35], following the setup of [31], using the ImageNet-1K dataset [11] (see Appendix for more details and results with other layers and models). As CLIP-Dissect requires a predefined concept setC, we employ the Broden dataset [2], (|C|=1197) which provides both low-level attributes and object-level visual concepts. We then compute interpretability and steerability scores for each SAE neuron as described in Sec. 3. Observations. Fig. 3 illustrates the trade-off between inter- pretability and steerability scores, with dashed horizontal and vertical lines marking their respective mean values to delineate four distinct neuron groups. For each group, the figure shows the top-10 activating images, CLIP-Dissect- assigned concepts (interpretability), and LLaVA-steered out- puts (steerability) for a representative neuron, highlighting characteristic behaviors. The distribution of neurons across these groups is as follows: ⢠Low interpretability, low steerability (36.26%, 23,763): These neurons are largely inactive and contribute mini- mally to either semantic meaning or controllable behavior. â˘High interpretability, low steerability (19.87%, 13,022): These neurons capture clear, human-understandable con- cepts but have limited influence on model outputs. â˘Low interpretability, high steerability (25.03%, 16,403): These neurons effectively steer outputs but correspond to abstract or composite features that lack semantic clarity. ⢠High interpretability, high steerability (18.84%, 12,348): This is the most desirable group of neurons that are both interpretable and causally effective. These indicate that a vast majority of SAE neurons are not directly useful for downstream tasks such as explanation or control, reinforcing the need for hybrid approaches that jointly enhance interpretability and steerability. Expt. 2: Can SAEs represent all user-specified concepts? A key question is whether the SAE can represent all concepts within a given concept set, thereby supporting both human interpretability and controllable model behavior. Although the SAE contains 65,536 neuronsâfar exceeding the size of standard concept setsâits ability to represent concepts varies considerably with the diversity and complexity of the set. Using CLIP-Dissect, we evaluate the coverage of unique concepts across multiple concept sets: ⢠Broden [2]: 1,153/1,197 (96.3%) ⢠VLG-CBM [39]: 3,445/4,729 (72.8%) ⢠DECIDER [41]: 4,333/7,827 (55.3%) ⢠3k common English words [29]: 1,857/3,000 (61.9%) ⢠20k common English words [29]: 5,596/20,000 (28.0%) The SAE performs well on the smaller and well-defined Broden concept set, capturing 96.3% of its visual concepts. However, coverage drops sharply for larger or linguistically diverse sets, with only 28â73% of concepts represented on average. Notably, despite being trained on ImageNet, the SAE fails to capture 27â45% of ImageNet-related concepts from the VLG-CBM and DECIDER sets. These results indicate that, while SAEs effectively capture simple, low- level concepts, their latent spaces struggle to generalize to broader, user-specified concept setsâlimiting their utility for downstream interpretability and nuanced steerability. Our analysis reveals two key requirements: (i) expand- ing the SAE latent space to capture a broader range of se- mantically distinct concepts while remaining effective for downstream tasks such as steering, and (i) enabling ex- plicit user specification of concepts within the SAE. Simply pruning neurons with low interpretability and steerability degrades reconstruction fidelity. To address this and enable user-specified concepts, we propose to train a concept bot- Figure 4. Pipeline for building CB-SAE. Step 1. A baseline SAE is trained and Step 2. evaluated with CLIP-Dissect and downstream steering to obtain interpretability and steerability scores per SAE neuron. Step 3. TheMleast interpretable and steerable neurons are pruned by deleting the corresponding SAE weights. Step 4. We train the CB-SAE with frozen, pruned SAE weights with three objectives: A. recover the reconstruction ability lost by pruning usingL r , B. incorporate the user-specified concept set withL int , and C. promote steerability with a cyclic reconstruction lossL st . tleneck autoencoder [18,19] alongside the retained SAE. This hybrid framework combines (supervised) concept align- ment with (unsupervised) discovery, restoring reconstruction fidelity while enhancing concept coverage and steerability. 5. Our Approach: CB-SAE We propose a novel concept bottleneck sparse autoencoder (CB-SAE) based on our analysis to address two limitations of sparse autoencoders namely low interpretability/steerability and the lack of support for user-specified concepts. 5.1. Pruning SAE neurons Step 1(Fig. 4) As discussed in Sec. 3, we begin with training an SAE on layer l activations from the vision model f . Step 2(Fig. 4) Following our analysis experiments, we compute the interpretability and steerability scores for each sparse neuron in the trained SAE denoted byI â [0, 1] Ď and S â [0, 1] Ď respectively. Step 3(Fig. 4) We prune the SAE weightsE sae ,D sae to remove theMleast interpretable and steerable SAE neu- rons as they are unsuitable for downstream applications. Concretely, the set ofMSAE neurons to be pruned is P = m | I m + S m < Ď, m â [Ď]whereĎis the thresh- old that determines|P| = Mand[Ď] = 1, 2,¡ ,Ď. In practice, we simply sort theI + Sin descending order and select the bottom-M neurons that constituteP . SAE consists ofE sae â R ĎĂd andD sae â R dĂĎ . We can then prune the selected set of neuronsPby deleting the cor- responding rows and columns inE sae andD sae respectively. In other words, the retained SAE weightsE Ⲡsae andD Ⲡsae have all rows and columns other than those inP respectively: E Ⲡsae = E sae [[Ď] , :](3) D Ⲡsae = D sae [:, [Ď] ](4) Here, the set minus operator. Note, the shared bias termb â R d does not change as it is independent of the number of SAE neuronsĎ. The retained SAE consists of the retained encoderE Ⲡsae â R (ĎâM)Ăd , retained decoder D Ⲡsae â R dĂ(ĎâM) , and the biasb â R d . Using Eq.(1), (2)withE Ⲡsae ,D Ⲡsae , the retained SAE latent changes toz Ⲡâ R (ĎâM) and the reconstructed latent isËv Ⲡâ R d (Fig. 4, Step 3). While the reconstructedËv Ⲡhas the same dimensions as v, the average reconstruction loss E v [vâ Ëv Ⲡ]will be higher than without pruning E v [vâ Ëv], due to loss of information to effectively reconstruct the activations. As discussed earlier, to recover this lost reconstruction ability and to incorporate user-specified concepts, we introduce a concept bottleneck. 5.2. Training CB-SAE Step 4 (Fig. 4) We introduce a concept bottleneck au- toencoder [19] alongside the retained SAE (Fig. 1B). Our CB-SAE consists of the retained SAEE Ⲡsae ,D Ⲡsae , a lin- ear concept encoderE cb â R |C|Ăd , a linear concept de- coderD cb â R dĂ|C| , and a non-linear activation function Ď cb : R |C| â R |C| similar to the SAE, whereCis a pre- defined concept set. For an inputv â R d , the CB-SAE reconstructs Ëv Ⲡâ R d as, z Ⲡ= Ď sae (E Ⲡsae (vâ b))(5) c = E cb (vâ b)(6) Ëv Ⲡ= D Ⲡsae z Ⲡ+ b + D cb Ď cb (c)(7) We use a top-kfunction asĎ cb withk âŞ|C|to ensure that sparsity constraints are similar to the original SAE. The bias term bâ R d is shared with the retained SAE. Concept Set Selection. Based on our motivation to support user-specified concepts, the concept setCcan be specified by the user as a list of text-based concepts, similar to prior work [29,39]. However, as shown in Sec. 4, the SAE can al- ready represent some concepts well and including them in the concept setCwould be redundant. Hence, we only use con- cepts absent from the retained SAE in the CB-SAE concept set. LetC user be the user-specified concept set,C rsae âC user be the concepts present in the retained SAE (found using CLIP-Dissect before pruning the SAE, Fig. 4 Step 2). Then our CB-SAE concept set is given byC =C user rsae . Training Objectives. We propose three training objectives to guide the CB-SAE. First, the concept bottleneck should recover the reconstruction fidelity lost by pruning the SAE neurons. Second, the CB neurons should be interpretable w.r.t. the concept setC. And third, the CB neurons should be steerable w.r.t. the concept setC. The neurons of the retained SAE already meet the reconstruction, interpretability, and steerability objectives for their discovered conceptsC rsae , so we keep the retained SAE weights frozen. Objective A: ReconstructionL r (Fig. 4, Step 4A). Similar to the SAE, we optimize the mean-squared errorL r between the input latent v â R d and the CB-SAE reconstruction Ëv Ⲡ, min E cb ,D cb [L r (v, Ëv Ⲡ)](8) Instead of anâ 1 regularizer for sparsity, we use a top-k activation function Ď cb as mentioned earlier. Objective B: InterpretabilityL int (Fig. 4, Step 4B). To en- sure that each CB neuron incactivates for its corresponding concept in the concept setC, we use a CLIP zero-shot classi- fier [35] withCas the classes and obtain pseudo-ground-truth concept activations similar to prior work in CBMs [19,29]. This enables concept-label-free training of the CB-SAE, i.e. does not require explicit concept labels and also supports any arbitrary user-specified concept sets. The zero-shot classifierM : XĂ T |C| â R |C| (Trefers to text space) takes in an imagexâ X, list of concept names C, and predicts concept logitsËy =M(x,C)â R |C| . We use a cosine-cubed similarity lossL int following [29] between the concept encoderâs predictions c = E cb (v) and Ëy, min E cb [L int (c, Ëy)](9) Here,v = f l (x)is the vision encoderfoutput at layerlfor the same imagexused withM. Further, note that we use the sparsity constraint of a top-kactivation functionĎ cb only for the decoder in Eq.(7). This is because the concept encoder E cb should be able to interpret all concepts in an image, but the concept decoderD c can only use the top-kconcepts for reconstruction. This is further ensured by updating onlyE cb with the interpretability objectiveL int . We defer more details of the cosine-cubed similarity loss [29] to the Appendix. Objective C: SteerabilityL st (Fig. 4, Step 4C). Prior work on CBMs for generative models [19] leveraged the down- stream image generation task to design explicit steerability objectives. In contrast, we propose a simple task-agnostic cyclic reconstruction objective for steerability. With this, we show in Sec. 6 that the same CB-SAE can steer two dif- ferent downstream tasks: image-to-image generation and image-text-to-text generation. Concretely, as shown in Fig. 4, Step 4C, we pass the reconstructed latentËv Ⲡback through the concept encoderE cb to produce cyclically reconstructed conceptsËc = E cb (Ëv Ⲡâ b) . Then, we use the same pseudo-ground-truth concept activationsËycomputed for Objective B and optimize the same loss as Objective B withËc(denoted byL st for clarity), min D cb [L st (Ëc, Ëy)](10) Only the concept decoderD cb is updated with the steerability loss. This is because only the decoderD cb is responsible to appropriately modify the latentËv Ⲡwhen a concept inc is modified for steering. On the other hand, the concept encoderE cb should focus only on interpreting the inputv and updating it withL st could hurt interpretability. Instead of loss weighting hyperparameters, we train by alternately minimizing the objectives via separate Adam optimizers which adaptively scale weight updates [20]. 6. Experiments We extensively evaluate our proposed CB-SAE w.r.t. in- terpretability and steerability on two downstream tasks, (image,text)-to-text generation and image-to-image gener- ation. We also performed detailed ablation and sensitivity analysis experiments to validate our design choices. 6.1. Setup Baseline SAE and CB-SAE. We follow Pach et al.[31]and train a Matryoshka Batch Top-kSAE [51] with expansion factor Ď d = 64as the baseline SAE on the ImageNet-1k [11] Table 1. Interpretability and Steerability Evaluation with LLaVA and UnCLIP. All four metrics are in 0-1 range (higher is better), CD indicates CLIP-Dissect score and MS indicates monosemanticity score. Downstream Task Steered Model Method InterpretabilitySteerability CDMSUnit-VecWhite Image Image + Textâ Text Generation LLaVA-1.5-7B [23] (CLIP-ViT-L + Vicuna-7B) SAE [31]0.1540.5170.1980.203 CB-SAE (Ours)0.2440.5560.2610.250 LLaVA-MORE [10] (DINOv2-L + Gemma2-9B) SAE [31]0.1940.5530.1790.177 CB-SAE (Ours)0.2910.5980.1920.189 Imageâ Image Generation UnCLIP [38] (CLIP-ViT-L + SD-2.1) SAE [31]0.0580.5400.6420.654 CB-SAE (Ours)0.0920.5940.6590.664 SAE Baseline Interp. Score Steer. Score Both Scores 0.14 0.18 0.22 0.26 0.30 Interpretability Score A. Sensitivity to scores used for SAE pruning 0.14 0.18 0.22 0.26 0.30 Steerability Score 10k20k30k40k50k60k Number of SAE neurons retained 0.12 0.20 0.28 0.36 0.44 Interpretability Score SAE baseline B. Sensitivity to no. of SAE neurons retained 0.12 0.20 0.28 0.36 0.44 Steerability Score SAE baseline SAE Baseline w/o st with st 0.15 0.18 0.21 0.24 0.27 Interpretability Score C. Ablation study for steerability loss st 0.15 0.18 0.21 0.24 0.27 Steerability Score Figure 5. A. Sensitivity of CB-SAE to the choice of scores used for SAE pruning. B. Sensitivity of CB-SAE to the number of SAE neurons retained. C. Ablation study for our proposed steerability objectiveL st from Eq. (10). Table 2. Evaluating interpretability and steerability of discarded SAE neurons, retained SAE neurons, and CB neurons separately. Set of Neurons InterpretabilitySteerability CLIP-DissectUnit-VecWhite Image All SAE neurons0.1540.1980.203 Discarded SAE neurons0.0840.1440.162 Retained SAE neurons0.2380.2630.252 CB neurons0.3230.2310.219 All CB-SAE neurons0.2440.2610.250 dataset. Our CB-SAE is also trained on the same intermedi- ate activations as the baseline SAE for a fair comparison. We retain Ďâ M = 30k neurons in the SAE pruning and use a top-kfunction asĎ cb withk = 5in our CB-SAE. We use the VLG-CBM ImageNet concept set [29,39] for the CB neu- rons. In our training, we use a CLIP-ViT-B/16 [35] model for obtaining the pseudo-ground-truth concept activations. Evaluation Metrics. To evaluate interpretability, we use the CLIP-Dissect interpretability score introduced in Sec. 3 and the monosemanticity score from [31] using the Ima- geNet validation set. To ensure a fair evaluation, we use a stronger CLIP-ViT-L/14 model (w.r.t. smaller ViT-B/16 used for training CB-SAE). To evaluate steerability, we use our proposed steerability score (Sec. 3). Concretely, we evaluate the steerability of each CB/SAE neuron in two ways: â˘Unit Vector: The selected neuron is activated to a high value Îą = 50 (as in [31]) & all other neurons are set to 0. â˘White Image: The selected neuron is activated to a high valueÎą = 50and all other neurons have the values pre- dicted when using an empty white image (following [31]) as input, instead of 0 like in unit vector steering. The interpretability and steerability scores of individual neu- rons are averaged to obtain the overall scores. For exper- iments where the steered output is text, we compare the similarity between the steered text and the CLIP-Dissect assigned concept for the selected neuron in a sentence trans- former embedding space (as in Sec. 3). For experiments where steered output is an image, we compute the average similarity between the steered image and top-16 highly ac- tivating images for the selected neuron in the DINOv2 [30] embedding space. This is because the diffusion model be- ing steered (UnCLIP [38]) may rarely return partially or completely noisy images after steering, which cannot not be properly evaluated with an image-text similarity score (e.g. CLIP) that expects clean images. All metrics are normalized in 0-1 range and higher values indicate better performance. 6.2. Quantitative Comparison In Table 1, we compare our CB-SAE with the baseline SAE [31] across several downstream models as well as tasks. For two different variants of the (image,text)-to-text LLaVA [10,23] model, our CB-SAE demonstrates consis- tent gains over the SAE baseline across both interpretability (avg. +33.0% for CLIP+Vicuna and avg. +29.0% for DI- NOv2+Gemma) and steerability metrics (avg. +27.5% for A. Retained SAE neuronsB. Concept Bottleneck (CB) neurons Figure 6. Visualizing the interpretability and steerability of retained SAE neurons and CB neurons, similar to Fig. 3. CLIP+Vicuna and avg. +14.0% for DINOv2+Gemma). Simi- larly, for the image-to-image UnCLIP [38] generative model, our CB-SAE outperforms the SAE baseline (avg. +34.3% interpretability and avg. +2.1% steerability). To the best of our knowledge, we are the first to show that an SAE (and CB-SAE) trained with the same method can be used to steer different downstream tasks. 6.3. Analysis of our CB-SAE Effect of CB neurons. In Table 2, we separately eval- uate the interpretability and steerability of discarded and retained SAE neurons, as well as CB neurons. We ob- serve CB neurons have significantly higher interpretability than SAE neurons. Whereas, CB neuronsâ steerability is worse than retained SAE neurons but significantly better than discarded SAE neurons as well as all SAE neurons (discarded+retained). Intuitively, the retained SAE neurons contain many highly steerable neurons because steerability (and interpretability) were used to prune SAE neurons. Sensitivity to scores used for SAE pruning. In Fig. 5A, we compare the CB-SAE performance while varying the choice of metrics for SAE pruning: either interpretability score or steerability score or both. We find that prioritizing either score leads to some loss in performance on the other, while using both scores gives balanced performance. This is beneficial as users can choose the score or even design a new score for pruning based on their target downstream usecase. Sensitivity to no. of SAE neurons retained. In Fig. 5B, we evaluate CB-SAE models trained with varying number of SAE neurons retained after pruning using both interpretabil- ity and steerability scores. We find that lower number of SAE neurons retained leads to higher interpretability and steerability scores. This is because pruning keeps a smaller subset of SAE neurons with higher scores. However, further reducing the number of SAE neurons would hurt perfor- mance as reconstruction becomes more difficult. Ablation study for steerability lossL st . In Fig. 5C, we analyze the impact of our proposed steerability lossL st on the CB-SAE. We observe similar interpretability with and withoutL st , which is significantly better than the SAE base- line. And usingL st improves steerability by 2.9%, validating Concept LLaVA steered o/p truth UnCLIP steered output race Concept fish bowl Retained SAE bathroom bathroom bread bread kalahari desert leather/ fabric CB neurons cross blunt edge face face rubber rubber mansion mansion windows cookie chocolate syrup made of pores Discarded SAE Neuron #54185#2610#3008#30506#31437#31018#35447 Unit Vector Concept LLaVA steered o/p bubble playing ice ice toiletries kitchen ticket a ticket trees trees canine computer hair white wing tips Neuron #38927#1046#59466#31877#31918#30232#28981 White Image Concept LLaVA steered o/p flower floral design office office lake lake mess mess glass surface glass building building circle stone circle Neuron #26259#1013#3459#31068#30624#30192#40393 Concept LLaVA steered o/p fish prey animal sky sky toiletries strawberry stamps stamp heel foot paper clip pole interior colorful exterior Neuron #58239#4048#27760#31714#30714#31209#58036 Unit Vector Neuron #44896#6660#8278#29561#29318#30333#10542 UnCLIP steered output Concept short tufted tail round metal dial speedo meter staircase super hero fire slot for mail Neuron #51681#56619#10722#30216#30893#29712#53739 Figure 7. Qualitative examples of steering UnCLIP and LLaVA. Green indicates successful steering, yellow indicates partial success, and red indicates failure cases. See Appendix for more results. its usefulness in CB-SAE training. Visualizing retained SAE and CB neurons. We visualize the distribution of CB neurons and retained SAE neurons w.r.t. interpretability and steerability scores in Fig. 6. Com- pared to Fig. 3, the retained SAE neurons do not include low interpretability/steerability neurons as per our SAE pruning (Fig. 4, Step 2). Whereas the CB neurons have higher in- terpretability and similar steerability scores to the retained SAE neurons. Hence, there is potential to further improve steerability by designing better or more task-specific losses. Qualitative examples of steering. We report qualitative ex- amples of steering LLaVA and UnCLIP with discarded SAE neurons, retained SAE neurons, and CB neurons in Fig. 7. We find discarded SAE neurons to be worse in steering than retained SAE neurons, which is expected as pruning uses the steerability score. Specifically, in the image-to-image generator UnCLIP [38], we find steering the CB neurons produce higher quality images compared to the SAE neurons which often produce noisy images, likely due to the explicit concept supervision in CB-SAE. We also highlight some failure cases and partially correct steering for all neurons. We provide white image steering examples for UnCLIP in the Appendix due to space constraints. 7. Conclusion In this work, we made the first attempt to unify two com- plementary paradigms- SAEs for unsupervised concept dis- covery and CBM for interpretable control-into a single uni- fied framework, CB-SAE. Motivated by insights derived from our comprehensive analysis of SAEs in LVLMs, we first pruned the low-utility neurons and their correspond- ing weights in the SAE. We then introduced a light-weight CB module trained alongside with the frozen, retained SAE using three principled objectives. Through systematic evalu- ation across two different downstream generation settings- vision-language assistance (LLaVA) and image generation (UnCLIP) tasks, we demonstrate that CB-SAE consistently improves both interpretability and steerability while enabling explicit user-specified concept control. We also acknowledge that the efficacy of our approach depends on the reliability of CLIP-Dissect in assigning ac- curate neuron-level concepts; however, continued advances in vision-language models are likely to enhance its perfor- mance. Extending and exploring hybrid approaches that combine the strengths of other unsupervised concept dis- covery methods such as transcoders [34] with user-specified concept control methods constitute our future work. Appendix In this appendix, we present full implementation details along with additional analyses. To support reproducibility, we will also release our codebase and pretrained models. The appendix is organized as follows: ⢠Section A: Implementation Details ⌠Interpretability score ⌠CLIP-Dissect ⌠Cosine-cubed similarity loss ⢠Section B: Experiments ⌠Experimental setup (Sec. B.1) ⌠Interpretability vs steerability (Sec. B.2, Fig. 8) ⌠Extended analysis (Sec. B.3, Table 3, 4, Fig. 9) ⌠Extended qualitative results (Sec. B.4, Fig. 10) A. Implementation Details Intepretability Score. We define our CLIP-Dissect-based interpretability score as the similarity score obtained from CLIP-Dissect, averaged across all SAE/CB-SAE neurons. CLIP-Dissect [28]. Consider a probing dataset ofNimages D =x i â X N i=1 whereXis the space of images, a concept setC =c k M k=1 withMconcepts in text form, and let layer lof modelfbeing explained be denoted byf l . CLIP-Dissect uses the probing set and a multimodal model, e.g. CLIP [35] with an image and text encoderE I ,E T to identify concepts fromC for individual neurons at the output of f l . The probing setDis passed through the CLIP image encoderE I to obtain corresponding set of image embed- dingsA i = E I (x i ) N i=1 . The concept set is passed through the CLIP text encoderE T to obtain text embed- dingsE T (c k ) M k=1 .Next, a matrixP â R NĂM is computed as the inner product of the image-text embed- dings with entriesP ik = A ⤠i E T (c k ), as CLIP image and text encoders have the same embedding dimensions. The layerlactivations of a neuronjfor the same probing set are denoted byq j = [f l (x 1 ) j ,f l (x 2 ) j ,¡ ,f l (x N ) j ]. Fi- nally, each neuronjcan be identified to have the concept arg max k sim(P :,k ,q j )whereP :,k is thek th column ofP. In other words, we compare each neuronâs activations over the probing set with the corresponding activations of the CLIP model for each concept, and select the concept with the highest similarity. The maximum similarity itself (aver- aged across all neurons) is used as our interpretability score. The similarity functionsimis soft weighted pointwise mu- tual information (soft-WPMI) following [28]. Please refer to the original paper [28] for more details. Cosine-cubed similarity loss [29]L int . As discussed in Sec. 5.2 (main paper), we use a cosine-cubed similarity lossL int to train the CB encoderE cb to produce concept predictions cthat match with CLIP zero-shot classifier predictionsËyfor the same concept setC. Concretely, L int (c, Ëy) = |C| X k=1 â c 3 k ¡ Ëy 3 k âĽc 3 k ⼠2 âĽËy 3 k ⼠2 (11) Here,c k is thek th concept prediction for the current mini- batch andËy k is the zero-shot CLIP prediction for conceptk with the same mini-batch. Following [29], we also normalize both vectorsc k , Ëy k â kbefore raising them to the third power (element-wise) and computing the cosine similarity. The third power is used to make the loss more sensitive to highly activating inputs. And we minimize the negative similarity which is equivalent to maximizing the similarity. B. Experiments B.1. Experimental Setup Downstream model details. We experiment with SAEs/CB- SAEs trained on vision encoders for downstream models like LLaVA [22] and UnCLIP [38]. LLaVA models are large vision-language models that take an image and a text prompt as input and output a text-based answer (Fig. 2A, main paper). Specifically, we used LLaVA-1.5-7B [23] which uses a CLIP-ViT-L-14-336 [35] vision encoder, a 2-layer MLP projector (not shown in Fig. 2A for simplicity), and an instruction-finetuned Vicuna-7B LLM [9]. We also use LLaVA-MORE [10] with DINOv2-Large [30] vision en- coder, a 2-layer MLP projector, and an instruction-finetuned Gemma2-9B LLM [43]. On the other hand, UnCLIP is an image-to-image generative model that uses a CLIP-ViT-L [35] vision encoder and a finetuned Stable Diffusion 2.1 [38] as the image generator (Fig. 2B, main paper). Miscellaneous details. We implement our CB-SAE in Py- Torch [33] building on the SAE codebase from Pach et al. [31]. Following the baseline SAE training [31], we train the CB-SAE for 110k iterations with batch size 4096 and learning rate 2eâ4 on a single 80GB Nvidia H100 GPU. Figure 8. We analyze the interpretability and steerability of SAE and CB-SAE neurons for LLaVA with DINOv2 and Gemma2 as well as for UnCLIP with CLIP-ViT-L and Stable Diffusion 2.1. The dashed lines in the baseline SAE plots indicate the average scores along each axis. B.2. Interpretability vs Steerability in SAEs We extend our analysis from Sec. 4 (main paper) on an SAE from LLaVA with CLIP image encoder to SAEs from LLaVA with DINOv2 image encoder and UnCLIP image-to-image generation model with CLIP image encoder in Fig. 8 (left). We report our observations (repeating those from Sec. 4): ⢠LLaVA (CLIP-ViT-L + Vicuna-7B, Fig. 3, main paper): ⌠Low interpretability, low steerability: 36.26% (23763) ⌠High interpretability, low steerability: 19.87% (13022) ⌠Low interpretability, high steerability: 25.03% (16403) âŚHigh interpretability, high steerability: 18.84% (12348) ⢠LLaVA (DINOv2-L + Gemma-9B, Fig. 8A): ⌠Low interpretability, low steerability: 33.07% (21675) ⌠High interpretability, low steerability: 23.35% (15304) ⌠Low interpretability, high steerability: 23.75% (15565) âŚHigh interpretability, high steerability: 19.82% (12992) ⢠UnCLIP (CLIP-ViT-L + Stable Diffusion 2.1, Fig. 8B): ⌠Low interpretability, low steerability: 30.84% (20209) ⌠High interpretability, low steerability: 14.53% (9517) âŚLow interpretability, high steerability: 42.76% (28022) ⌠High interpretability, high steerability: 11.88% (7788) Note that the average steerability score for UnCLIP is higher than for LLaVA since the scores are computed in image embedding space and text embedding space respectively. Across both types of models, we consistently find that only a small portion of neurons (12-20%) are useful for both interpretability and steerability. And a majority of neurons (30-36%) are unsuitable for both interpreting new inputs and steering outputs. We also show the retained SAE neurons and CB neu- rons in Fig. 8 (right) similar to Fig. 6 (main paper). We find CB neurons are similar to retained SAE neurons while being significantly better than the discarded SAE neurons (also shown quantitatively in Table 1, 2, main paper). We emphasize that CB neurons have to incorporate relatively more difficult concepts due to our concept set selection (Sec. 5.2, main paper) which excludes already discovered (and relatively easier to learn) concepts present in the retained SAE. Hence, it is more difficult for CB neurons to always outperform the retained SAE neurons. Table 3. Sensitivity to choice of metrics for SAE pruning. Scores for SAE pruning Reconstruction evaluationInterpretability evaluationSteerability evaluation Zero-shot ImageNet Acc. (%)CLIP-DissectMonosemanticityUnit VectorWhite Image None (SAE baseline) [31]74.070.1540.5170.1980.203 Interpretability score only73.390.2330.5660.2160.220 Steerability score only70.990.1670.5200.2880.269 Both scores73.78 0.2440.5560.2610.250 Table 4. Sensitivity of interpretability evaluation with CLIP-Dissect to choice of CLIP-like model used. CLIP-like model for evaluationInterpretability Score ModelArchitectureSAECB-SAE (Ours) CLIP [35]ViT-B-160.1980.307 CLIP [35]ViT-L-14-3360.1540.244 SigLIP [52]ViT-SO400M-14-3840.1890.289 SigLIP2 [44]ViT-gopt-16-3840.1880.290 SigLIP2 [44]ViT-SO400M-16-3840.1760.272 DFN [12]ViT-H-14-3780.2200.347 PE-core [5]BigG-14-4480.2070.312 B.3. Extended Analysis of our CB-SAE Sensitivity to scores used for SAE pruning. We extend our sensitivity analysis from Fig. 5A (main paper) in Table 3 to additionally include monosemanticity score [31] (inter- pretability evaluation) and zero-shot ImageNet-1k accuracy (reconstruction evaluation) when using the SAE/CB-SAE reconstructed latents. We observe that using either the in- terpretability score or both scores yields similar reconstruc- tion as the baseline SAE, while steerability-based pruning leads to significantly worse reconstruction. Similarly, us- ing either the interpretability score or both scores improves the monosemanticity significantly w.r.t. the baseline, while steerability-based pruning provides only a marginal gain over the baseline. Sensitivity to CLIP model in interpretability evaluation. We evaluate the sensitivity of our interpretability evaluation with CLIP-Dissect by varying the CLIP-like model used, in Table 4. While our evaluation used a stronger CLIP-ViT-L- 14-336 [35] model w.r.t. the smaller CLIP-ViT-B-16 used for training the CB-SAE, we now evaluate with even stronger models including SigLIP [52], SigLIP2 [44], Data Filter- ing Networks (DFN) [12] and Perception Encoder (PE) [5]. Across all CLIP-like models, our CB-SAE achieves consis- tent gains over the baseline SAE for LLaVA with CLIP-ViT- L encoder, validating that our choice of CLIP-like model for interpretability score does not affect our evaluation. Sensitivity tokinĎ cb . In Fig. 9, we analyze the sensitivity of our CB-SAE to the choice ofkin the top-kactivation function used in the CB decoder. Here, we define recon- struction score as the zero-shot ImageNet-1k accuracy of CLIP when using SAE/CB-SAE reconstructed latents. We 051020304050 k 40 50 60 70 80 Reconstruction Score SAE baseline Sensitivity to k in cb = topÂk function 0.14 0.18 0.22 0.26 0.30 Steerability Score Retained SAE neurons Discarded SAE neurons Figure 9. Sensitivity analysis of CB-SAE in LLaVA tokin top-k activation function used in the CB decoder. Steerability score here is computed only for CB neurons, reconstruction score is zero-shot accuracy when using SAE/CB-SAE reconstructions of CLIP latents on ImageNet-1k. also report the white image steerability score of only the CB neurons to understand the impact ofkon steerability. Note that we do not consider interpretability score here sinceĎ cb is only applied in the CB decoder while interpretability eval- uation only considers the CB encoder, i.e. interpretability score does not change when varyingk. We observe that re- construction score improves askincreases, but it is already very close to the baseline even atk = 3tok = 5. The steerability score first increases withkand then decreases fork > 30. This is because with higherk, steering might be less successful as the selected concept contends with many other concepts to be combined into the final reconstructed latent. On the other hand, ifkis too low, then the reconstruc- tion might not be good enough for the downstream model to produce the appropriate response. However, across all values ofk, our CB-SAE is able to outperform the discarded SAE neurons while being worse than the retained SAE neurons. Hence, future work can develop more steerability-focused training objectives to further improve steerability. B.4. Extended Qualitative Results We provide qualitative examples of white image steering of UnCLIP with SAE/CB-SAE in Fig. 10. Similar to our results in Fig. 7 (main paper), we find steering CB-SAE neurons UnCLIP steered output Concept hard drive Retained SAE small display medium sized dog CB neurons elephant -like grassy/ sandy pallets snake- like Discarded SAE White Image Neuron #42005#34929#46183#30258#29818#30349#18164 UnCLIP steered output Concept long loose robe seed drill fish cylindrical shape hydrantdog-like vehicle Neuron #38049#7903#40635#30827#29939#30595#51040 Figure 10. Qualitative examples of steering UnCLIP. Green indicates successful steering, yellow indicates partial success, and red indicates failure cases. produces higher quality images while SAE neurons tend to produce more noisy images. References [1]Dana Arad, Aaron Mueller, and Yonatan Belinkov. Saes are good for steeringâif you select the right features. arXiv preprint arXiv:2505.20063, 2025. 1, 2, 3 [2]David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba. Network dissection: Quantifying inter- pretability of deep visual representations. In CVPR, 2017. 4 [3]Steven Bills et al. Open source automated interpretability for sparse autoencoder features. EleutherAI Blog, 2024. 3 [4]Joseph Bloom, Curt Tigges, Anthony Duong, and David Chanin. SAELens: An open-source library for training and analyzing sparse autoencoders, 2024. 1 [5]Daniel Bolya, Po-Yao Huang, Peize Sun, Jang Hyun Cho, Andrea Madotto, Chen Wei, Tengyu Ma, Jiale Zhi, Jathushan Rajasegaran, Hanoona Abdul Rasheed, Junke Wang, Marco Monteiro, Hu Xu, Shiyu Dong, Nikhila Ravi, Shang-Wen Li, Piotr Dollar, and Christoph Feichtenhofer. Perception encoder: The best visual embeddings are not at the output of the network. In NeurIPS, 2025. 11 [6]Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. Towards monoseman- ticity: Decomposing language models with dictionary learn- ing. Transformer Circuits Thread, 2023. 2 [7]Bart Bussmann, Patrick Leask, and Neel Nanda. Batchtopk sparse autoencoders. In NeurIPSW, 2024. 2, 3, 4 [8] Bart Bussmann, Noa Nabeshima, Adam Karvonen, and Neel Nanda. Learning multi-level features with matryoshka sparse autoencoders. In ICML, 2025. 2 [9]Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 2023. 9 [10] Federico Cocchi, Nicholas Moratelli, Davide Caffagni, Sara Sarto, Lorenzo Baraldi, Marcella Cornia, and Rita Cucchiara. LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning. In IC- CVW, 2025. 2, 7, 9 [11]Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 4, 6 [12]Alex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt, Alexander T Toshev, and Vaishaal Shankar. Data filtering networks. In ICLR, 2024. 11 [13]Leo Gao, Tom Dupre la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. In ICLR, 2025. 2, 3 [14] Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. In ICLR, 2024. 2 [15] Aya Abdelsalam Ismail, Julius Adebayo, Hector Corrada Bravo, Stephen Ra, and Kyunghyun Cho. Concept bottleneck generative models. In ICLR, 2024. 2 [16]Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Cal- lum McDougall, Kola Ayonrinde, Matthew Wearden, Arthur Conmy, Samuel Marks, and Neel Nanda. SAEBench: A com- prehensive benchmark for sparse autoencoders in language model interpretability. In ICML, 2025. 1, 2, 3 [17]Dahye Kim, Xavier Thomas, and Deepti Ghadiyaram. Rev- elio: Interpreting and leveraging semantic information in diffusion models. In ICCV, 2025. 3 [18]Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Muss- mann, Emma Pierson, Been Kim, and Percy Liang. Concept bottleneck models. In ICML, 2020. 2, 3, 5 [19]Akshay Kulkarni, Ge Yan, Chung-En Sun, Tuomas Oikarinen, and Tsui-Wei Weng. Interpretable generative models through post-hoc concept bottlenecks. In CVPR, 2025. 2, 5, 6 [20]Jogendra Nath Kundu, Maharshi Gor, and R Venkatesh Babu. Bihmp-gan: Bidirectional 3d human motion prediction gan. In AAAI, 2019. 6 [21]Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, J Ě anos Kram Ě ar, Anca Dragan, Rohin Shah, and Neel Nanda. Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, 2024. 1, 2, 3 [22]Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In NeurIPS, 2023. 2, 9 [23]Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In CVPR, 2024. 2, 3, 7, 9 [24] Hantao Lou, Changye Li, Jiaming Ji, and Yaodong Yang. SAE-V: Interpreting multimodal models for enhanced align- ment. In ICML, 2025. 3 [25]Alireza Makhzani and Brendan Frey. k-Sparse autoencoders. In ICLR, 2014. 2 [26]Ali Nasiri-Sarvi, Hassan Rivaz, and Mahdi S Hosseini. Sparc: Concept-aligned sparse autoencoders for cross- model and cross-modal interpretability.arXiv preprint arXiv:2507.06265, 2025. 3 [27]Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, and Fazl Barez. Towards interpreting visual in- formation processing in vision-language models. In ICLR, 2025. 3 [28]Tuomas Oikarinen and Tsui-Wei Weng. CLIP-Dissect: Au- tomatic description of neuron representations in deep vision networks. In ICLR, 2023. 3, 9 [29] Tuomas Oikarinen, Subhro Das, Lam M Nguyen, and Tsui- Wei Weng. Label-free concept bottleneck models. In ICLR, 2023. 2, 3, 4, 6, 7, 9 [30] Maxime Oquab, Timoth Ě e Darcet, Th Ě eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. DINOv2: Learning Robust Visual Features without Supervision. TMLR, 2024. 7, 9 [31]Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie, and Zeynep Akata. Sparse autoencoders learn monosemantic features in vision-language models.In NeurIPS, 2025. 1, 3, 4, 6, 7, 9, 11 [32]Isabel Papadimitriou, Huangyuan Su, Thomas Fel, Sham Kakade, and Stephanie Gil. Interpreting the linear structure of vision-language model embedding spaces. In COLM, 2025. 3 [33]Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An imperative style, high-performance deep learning library. In NeurIPS, 2019. 9 [34]Gonc ̧alo Paulo, Stepan Shabalin, and Nora Belrose. Transcoders beat sparse autoencoders for interpretability. arXiv preprint arXiv:2501.18823, 2025. 9 [35]Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In ICML, 2021. 1, 3, 4, 6, 7, 9, 11 [36]Senthooran Rajamanoharan, Tom Lieberum, Nicolas Son- nerat, Arthur Conmy, Vikrant Varma, J Ě anos Kram Ě ar, and Neel Nanda. Jumping ahead: Improving reconstruction fi- delity with jumprelu sparse autoencoders. arXiv preprint arXiv:2407.14435, 2024. 1, 2, 3 [37]Nils Reimers and Iryna Gurevych. Sentence-bert: Sentence embeddings using siamese bert-networks. In EMNLP, 2019. 3 [38]Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj Ě orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2022. 2, 3, 7, 8, 9 [39]Divyansh Srivastava, Ge Yan, and Lily Weng. VLG-CBM: Training concept bottleneck models with vision-language guidance. In NeurIPS, 2024. 2, 3, 4, 6, 7 [40]Samuel Stevens, Wei-Lun Chao, Tanya Berger-Wolf, and Yu Su. Sparse autoencoders for scientifically rigorous interpre- tation of vision models. arXiv preprint arXiv:2502.06755, 2025. 1, 3 [41]RakshithSubramanyam,KowshikThopalli,Vivek Narayanaswamy, and Jayaraman J Thiagarajan. Decider: Leveraging foundation model priors for improved model failure detection and explanation. In ECCV, 2024. 2, 4 [42]Chung-En Sun, Tuomas Oikarinen, Berk Ustun, and Tsui-Wei Weng. Concept bottleneck large language models. In ICLR, 2025. 2 [43]GemmaTeam,MorganeRiviere,ShreyaPathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupati- raju, L Ě eonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ram Ě e, et al. Gemma 2: Improving Open Language Models at a Practical Size. arXiv preprint arXiv:2408.00118, 2024. 9 [44]Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localiza- tion, and Dense Features. arXiv preprint arXiv:2502.14786, 2025. 11 [45] Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, and Neel Nanda. How visual representations map to language feature space in multimodal llms. In CVPRW, 2025. 3 [46] Xu Wang, Yan Hu, Benyou Wang, and Difan Zou. Does higher interpretability imply better utility? a pairwise analysis on sparse autoencoders. In NeurIPSW, 2025. 1, 3 [47]An Yan, Yu Wang, Yiwu Zhong, Chengyu Dong, Zexue He, Yujie Lu, William Yang Wang, Jingbo Shang, and Julian McAuley. Learning concise and descriptive attributes for visual recognition. In ICCV, 2023. 2 [48] Minglai Yang, Xinyu Guo, Mihai Surdeanu, and Liangming Pan. Alignsae: Concept-aligned sparse autoencoders. arXiv preprint arXiv:2512.02004, 2025. 2 [49]Yue Yang, Artemis Panagopoulou, Shenghao Zhou, Daniel Jin, Chris Callison-Burch, and Mark Yatskar. Language in a bottle: Language model guided concept bottlenecks for interpretable image classification. In CVPR, 2023. 2 [50] Mert Yuksekgonul, Maggie Wang, and James Zou. Post-hoc concept bottleneck models. In ICLR, 2023. 2 [51]Vladimir Zaigrajew, Hubert Baniecki, and Przemyslaw Biecek. Interpreting CLIP with hierarchical sparse autoen- coders. In ICML, 2025. 3, 4, 6 [52]Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid Loss for Language Image Pre-Training. In ICCV, 2023. 11 [53]Yichen Zhu, Minjie Zhu, Ning Liu, Zhiyuan Xu, and Yaxin Peng. LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model. In ACMMMW, 2024. 1