Paper deep dive
The Deleuzian Representation Hypothesis
Clément Cornet, Romaric Besançon, Hervé Le Borgne
Models: Audio Spectrogram Transformer (AST), BART, CLIP, DeBERTa, DinoV2, Pythia-70M
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/11/2026, 12:41:24 AM
Summary
The paper introduces the 'Deleuzian Representation Hypothesis', an unsupervised method for extracting interpretable concepts from neural network activations. Inspired by Deleuze's philosophy of concepts as differences, the approach uses discriminant analysis and weighted KMeans clustering to identify recurring activation differences. It outperforms sparse autoencoders (SAEs) in concept quality and consistency across vision, language, and audio modalities, while enabling lossless steering of model representations.
Entities (6)
Relation Signals (4)
Deleuzian Representation Hypothesis → isalternativeto → Sparse Autoencoders
confidence 95% · We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method
Deleuzian Representation Hypothesis → uses → Discriminant Analysis
confidence 95% · The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework.
Deleuzian Representation Hypothesis → uses → KMeans
confidence 95% · To constrain our concept dictionary to a fixed number of concepts k, we cluster activation differences using KMeans.
Probe Loss → evaluates → Deleuzian Representation Hypothesis
confidence 92% · Our primary quantitative evaluation relies on the probe loss metric
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework. To enhance the diversity of extracted concepts, we refine the approach by weighting the clustering using the skewness of activations. The method aligns with Deleuze's modern view of concepts as differences. We evaluate the approach across five models and three modalities (vision, language, and audio), measuring concept quality, diversity, and consistency. Our results show that the proposed method achieves concept quality surpassing prior unsupervised SAE variants while approaching supervised baselines, and that the extracted concepts enable steering of a model's inner representations, demonstrating their causal influence on downstream behavior.
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
66,242 characters extracted from source content.
Expand or collapse full text
THE DELEUZIAN REPRESENTATION HYPOTHESIS Cl ́ ement Cornet, Romaric Besanc ̧on & Herv ́ e Le Borgne Universit ́ e Paris-Saclay, CEA, List, F-91120, Palaiseau, France clement.cornet,romaric.besancon,herve.le-borgne@cea.fr ABSTRACT We propose an alternative to sparse autoencoders (SAEs) as a simple and effective unsupervised method for extracting interpretable concepts from neural networks. The core idea is to cluster differences in activations, which we formally justify within a discriminant analysis framework. To enhance the diversity of extracted concepts, we refine the approach by weighting the clustering using the skewness of activations. The method aligns with Deleuze’s modern view of concepts as differences. We evaluate the approach across five models and three modalities (vi- sion, language, and audio), measuring concept quality, diversity, and consistency. Our results show that the proposed method achieves concept quality surpassing prior unsupervised SAE variants while approaching supervised baselines, and that the extracted concepts enable steering of a model’s inner representations, demon- strating their causal influence on downstream behavior. 1INTRODUCTION Interpretability of neural network representations is essential for building trustworthy models, en- abling a deeper understanding of the mechanisms underlying a model’s predictions, and promoting fairness and accountability. However, interpreting the internal representations learned by neural net- works remains a central challenge in deep learning. Sparse autoencoders (SAEs) (Bricken et al., 2023; Cunningham et al., 2023) have emerged as a powerful tool for extracting sparse and seman- tically meaningful features from model activations. Nevertheless, they face challenges that limit their applicability. Notably, they suffer from difficulties in training, and may still yield polyseman- tic features, not corresponding to a single interpretable concept. Moreover, sparse autoencoders (and similar methods) rely on feature sparsity as a proxy for interpretability, a choice that has been criticized as potentially inadequate (Sharkey et al., 2025). We introduce an alternative to sparse autoencoders (SAEs) for extracting features that correspond to interpretable concepts from neural networks. Drawing inspiration from Deleuze’s philosophical view of concepts as differences, we model concepts as directions that capture distinctions between representations of individual samples. Specifically, our approach can be seen as an unsupervised (a) Image: Van Gogh’s Paintings “Winning the prize” “the Gold Medal” “the World Record” (b) Text: Sports Achievements(c) Audio: Brass Instruments Figure 1: Our method extracts diverse concepts from image, text and audio models. 1 arXiv:2512.19734v1 [cs.LG] 17 Dec 2025 discriminant analysis: it identifies directions in the internal representation that best separate data samples. We estimate those directions by sampling activation differences between pairs of data points, then use KMeans clustering to uncover recurring patterns. Our analysis is further refined using distributional skewness to promote diversity. Evaluating interpretability methods remains a major challenge. SAEs are often assessed by their reconstruction–sparsity trade-off, which does not necessarily reflect interpretability. Hence, most recent studies in this field are also evaluated qualitatively, showing their relevance through selected examples. While insightful, such evaluations provide limited support. In contrast, we adopt a quanti- tative evaluation based on probe loss (Gao et al., 2025), which measures the extent to which extracted concepts capture the attributes expected to be present in a dataset. To ensure robust evaluation, we apply this metric to a broad set of 874 attributes spanning different tasks, five datasets and five mod- els across three modalities (image, text and audio). Our method captures the desired attributes more effectively than recent SAE-based approaches. In several settings, it is competitive with supervised linear discriminant analysis. Beyond the presence of expected attributes, we also evaluate cross-run consistency with the Maximum Pairwise Pearson Correlation (MPPC) (Wang et al., 2025), establish- ing a comprehensive evaluation framework for concept evaluation methods. Finally, we demonstrate concept steering on text and image models, showing that manipulating extracted concepts causally influence downstream behavior, without incurring information loss. Hence, the main contribution of this paper is a novel type of approach of mechanistic interpretabil- ity of neural networks. We investigate the fundamental principle underlying our approach and demonstrate that it achieves globally more compelling results than state-of-the-art sparse autoen- coder (SAE)–based techniques. Our method is advantageous in its simplicity: it is governed by a single, interpretable hyperparameter. The proposed principle is theoretically grounded in discrim- inant analysis and clustering, and further relates to Deleuze’s philosophical notion of “concepts.” Similar to SAE-based approaches, our method is fully unsupervised and therefore does not require manual specification or annotation of the identified concepts. Code is publicly available on GitHub 1 . 2METHODS 2.1CRITERIA AND CONCEPTUAL GROUNDING Our aim is to extract an ontology of “concepts” from a neural network, by analyzing its activations. Before proposing our approach, we first discuss the criteria such concepts should satisfy. • Interpretability: this work aims to extract human-interpretable features, that are then re- ferred to as “concepts”. • Transparency: in order to gain interpretable insights into the model, the approach itself should be as simple and transparent as possible, not relying on non-interpretable hyperpa- rameters. • Diversity: the extracted concepts should be semantically diverse, in order to represent a wide variety of data samples, ideas, and semantic levels. • Consistency: the approach should consistently yield similar concepts when run multiple times with different random seeds. Existing methods in mechanistic interpretability typically extract unsupervised concepts by recon- structing model activations (Bricken et al., 2023; Cunningham et al., 2023). Because they are trained to minimize reconstruction error, such approaches are driven to capture as much variance in the acti- vation space as possible, subject to sparsity constraints. This framing implicitly presents concepts as universal structural components of the model activations, echoing the classical philosophical view of concepts as “the universal essence of a fact” (Plato, c. 375 BCE; Hegel, 1816). However, such a representation has been criticized as overly restrictive (Nietzsche, 1889; Sartre & Elka ̈ ım-Sartre, 1946). More recent perspectives instead emphasize concepts as arising from Difference and Repeti- tion (Deleuze, 1968), rather than universals. Following this idea, our approach does not attempt to model the full variance of activations. Instead, it identifies recurring differences between activations. 1 https://github.com/ClementCornet/Deleuzian-Hypothesis 2 2.2EXTRACTING REPEATED DIFFERENCES IN ACTIVATION SPACE Figure 2: Overview of our concept extraction approach. We sample pairwise differences in activa- tion between samples. Then, we use the inverse-skewness of those differences to selected the final concepts, corresponding to vectors in the activation space. Our objective is to extract concepts from model activations, at a given layer withD dimensions, over a dataset of N samples. To represent repeated differences in activations between data samples, we define D = ⃗ d 1 , ⃗ d 2 ,..., ⃗ d N as a set of D-dimensional pairwise differences in activation between samples. Since our approach is fully unsupervised, we cannot restrain D to contrastive pairs between two classes. However, computing all pairwise differences is quadratic in N . To approximate the distribution of differences, we instead randomly sample N pairs, ensuring that each data point is used once on each side of the subtraction. To constrain our concept dictionary to a fixed number of concepts k, we cluster activation differences using KMeans (Lloyd, 1982; Zeng & Zheng, 2019). However, some activation differences exhibit highly skewed distributions: they remain near-zero for most samples, but occasionally spike to large values. Those differences tend to dominate the Euclidean distance used by standard KMeans, and produce redundant clusters (Milligan, 1980). The skewness of a distribution X , defined as the normalized third central moment is ̃μ 3 (X) = P N i=1 (X i − ̄ X) 3 Nσ 3 (1) For a concept direction ⃗ d i , we consider skewness as that of the projection ⃗ d i ·⃗x j N j=1 . Since highly skewed coordinates tend to produce redundant clusters, we penalize them by assigning weights inversely proportional to skewness. In order to avoid ill-defined clustering with negative weights, and to consider opposite directions ⃗ d i as similar (as we are seeking directions, regardless of their orientation), we consider − ⃗ d i for differences with negative skewness. This results in a variant of Feature-Weighted KMeans (Huang et al., 2005), in which concept directions are weighted during centroids computation, in order to promote concept diversity. More precisely, this clustering defines the weighted distance between ⃗ d i and its corresponding centroid ̄ C as d( ⃗ d i , ̄ C) = 1 ̃μ 3 ( ⃗ d i ) || ̄ C− ⃗ d i || 2 The obtained centroids are then used as concept vectors. Both pair sampling and KMeans clustering run in linear time and memory with respect to dataset size N and activation dimensionD, demonstrating scalability of our approach towards large datasets, or large models. Finally, this procedure retains a simple and transparent formulation ( Figure 2), that are key proper- ties for interpretability research. Notably, the number of extracted concepts k is the only hyperpa- rameter required for our approach, and is itself interpretable. 3 2.3CONNECTION TO DISCRIMINANT ANALYSIS We aim to extract “concepts” from model activations, defining a concept as a difference between ideas. In a supervised setting, this objective relates closely to discriminant analysis (Fisher, 1936), which identifies a direction ⃗c orthogonal to the optimal separating hyperplane between two classes. Let Σ A and Σ B be the class covariances, and μ A and μ B their means. The separation between classes is maximized by: ⃗c∝ (Σ A + Σ B ) −1 (⃗μ A − ⃗μ B )(2) Consider two samples i and j with activations⃗x i and⃗x j , and suppose we seek the optimal separation between clusters with means ⃗x i and ⃗x j , distinguished by a concept ⃗c. In high-dimensional spaces (typically≥ 512 dimensions for transformers), we approximate Σ i and Σ j as diagonal, containing each dimension’s variance (Ahdesm ̈ aki & Strimmer, 2010). From equation 2, ⃗c∝ ⃗x i − ⃗x j achieves optimal separation when Σ i ∝ Σ j ∝ I , i.e., under isotropic cluster distributions. Thus, treating activation differences as the optimal separation between ideas is equivalent to assuming isotropic distributions of concepts in activation space. Unlike standard LDA, equation 2 does not require homoscedasticity or Gaussianity (McLachlan, 2005), and naturally extends to multiclass discrimination (Rao, 1948). In Appendix G, we derive a quadratic extension to our approach, that accounts for anisotropic dis- tribution of concepts. While theoretically interesting, it does not lead to better experimental results. For this reason, we focus on the isotropic approach (i.e Σ i ∝ Σ j ∝ I ) in the following. 2.4LOSSLESS STEERING Sparse autoencoders and related methods allow steering of extracted concepts (Zhou et al., 2025). To do so, they project sample activations in their concept space, apply a steering vector, and projects back into the activation space. The two projections required introduce reconstruction error and information loss. In contrast, our extracted concepts are vectors in the activation space. Therefore, we can perform steering directly in the activations space. To steer the embedding of a sample x, with a magnitude α and a concept ⃗c i , consider its steered representation ̃x = x + α⃗c i . If one steers a concept by +α, then by−α, we retrieve exactly the base activation. By avoiding projections into and out of the concept space, our approach enables lossless steering: the modifications affect only the targeted direction and can be exactly reversed. 3EXPERIMENTS Datasets and Models To evaluate our concept extraction methods, we conduct a large-scale study spanning five models and five datasets across three modalities (vision, language, and audio), cover- ing a wide variety of semantic attributes. For text, we use two datasets: IMDB (Maas et al., 2011) and CoNLL-2003 (Tjong Kim Sang & De Meulder, 2003). IMDB provides sentence-level binary sentiment classification labels, while CoNLL-2003 provides token-level labels for named entity recognition (NER), part-of-speech (POS) tagging, and syntactic chunking. For vision, we use a subset of ImageNet (Russakovsky et al., 2015) with 100 classes and the WikiArt dataset (Baylies, 2020) which contains paintings labeled by artist (129 classes), style (27 classes), and genre (11 classes). Concerning text datasets, IMDB has binary classification labels, while CoNLL-2003 has token-wise labels for NER (9 classes), POS-tagging (47 classes) and chunk tags (23 classes). For audio, we use AudioSet (Gemmeke et al., 2017), with multi-classification labels (527 audio classes). Our text experiments are conducted on DeBERTa (He et al., 2021) and the encoder of BART (Lewis et al., 2020), as well as Pythia-70M (Biderman et al., 2023). For vision, we evaluate DinoV2 (Oquab et al., 2023) and CLIP (Radford et al., 2021). For audio, we use a pretrained Audio Spectrogram Transformer (AST) (Gong et al., 2021). We only consider encoder models, (including the encoder of BART). This choice allows us to evaluate the quality of extracted concepts with respect to supervised 4 labels that are likely represented at the analyzed layer of each model, since our objective is to com- pare concept extraction methods. It also enables comparable analyses across multiple modalities. More details on datasets and models are provided in Appendix B. Baselines Sparse autoencoders (SAE) are predominant among concept extraction methods. We compare our method to five different types of SAEs: • VanillaSAE (Van-SAE) (Bricken et al., 2023): standard SAE, trained with an L 2 recon- struction loss, and enforcing sparsity via an L 1 penalty which requires a coefficient λ; • GatedSAE (Gat-SAE) (Rajamanoharan et al., 2024a): SAE learning activations gates, hence separating feature selection and magnitude estimation; • JumpReLUSAE (JR-SAE) (Rajamanoharan et al., 2024b): SAE with a learnable threshold θ i for each concept, designed to minimize the reconstruction error; • MatryoshkaSAE (Mat-SAE) (Bussmann et al., 2025): SAE learning nested dictionaries of concepts, focusing on hierarchies of concepts, belonging to multiple semantic levels; • TopKSAE (Tk-SAE) (Gao et al., 2025): SAE enforcing sparsity via a TopK activation function, that sets every activation to zero, except the k highest. • ArchetypalSAE (A-SAE) (Fel et al., 2025): SAE constraining decoder atoms to be combi- nations of activations, to gain stability. • Pretrained Sparse Autoencoders (Pretrained): we compare our method with publicly avail- able, pretrained sparse autoencoders on two models. For DinoV2 experiments, we use ViT-Prisma (Joseph et al., 2025), and for Pythia we use a sparse autoencoder trained by EleutherAI 2 . We also compare our approach to Independant Component Analysis (ICA) (Comon, 1994), that is a linear decomposition method maximizing statistical independence between latent dimensions. In addition, as our approach is closely related to discriminant analysis, we also compare it to super- vised Linear Discriminant Analysis (LDA) (Fisher, 1936) which serves as an upper bound under assumptions of homoscedasticity and normal distribution of concepts. Evaluation Our primary quantitative evaluation relies on the probe loss metric (Gao et al., 2025), which measures the degree to which extracted concepts align with ground-truth annotated attributes. Beyond the quality on individual concepts, we also aim at uncovering a broad set of concepts from model activations. To this end, we assess probe loss across tasks characterized by diverse attribute sets, thereby quantifying the capacity of our approach to capture multiple, semantically meaningful concepts. In addition, Maximum Pairwise Pearson Correlation (MPPC) (Wang et al., 2025) is used to measure the consistency of the different methods. Finally, to highlight causal influence of concepts on model predictions, we perform concept steering, and provide qualitative examples. Note that, while prior work on sparse autoencoders has emphasized reconstruction–sparsity trade-offs, these objectives are not applicable to our framework; we therefore exclude them from evaluation. All the reported results are computed using activations from the last transformer block of each encoder, using a concept space with 6144 dimensions, corresponding to 8 times the size of the activations (except for ICA, that is limited to 768). 3.1EVALUATION OF CONCEPT QUALITY We evaluate concepts extracted in an unsupervised manner by assessing whether they correspond to interpretable attributes known to exist in the dataset. This correspondence is quantified using Probe Loss (Gao et al., 2025). For each attribute, Probe Loss measures how well a one-dimensional logistic probe can recover the ground-truth attribute from the extracted concepts. Specifically, we train a separate 1D logistic probe for every concept and record the lowest cross-entropy loss achieved. For multi-class attributes, we report the median Probe Loss across all attributes. The results of this evaluation are presented in Table 1. From Table 1, our method globally outperforms all variations of SAE, with the lowest probe loss on 13 of the 20 tested tasks. This indicates a high ability to recover attributes expected to be found 2 https://huggingface.co/EleutherAI/sae-pythia-70m-32k 5 Table 1: Quantitative evaluation (Probe Loss, lower is better) of unsupervised approaches on CLIP and DinoV2 image encoders, DeBERTa and BART text encoders and Audio Spectrogram Trans- former on audio. Supervised baseline (LDA) is reported for reference (gray row). Best results are in bold, second in italics. Bottom right table indicates the average rank of all methods over all datasets (lower is better). “Pretrained” are models independently trained by other teams (see text for details) CLIPDinoV2 labels MethodImNet WikiArt ImNet WikiArt ArtistStyleGenreArtistStyleGenre ✓LDA0.00830.00840.04650.09760.00440.01010.05450.1084 ✗ICA0.01540.01410.08160.21040.01610.01550.08390.2035 ✗Van-SAE0.02640.0137 0.05580.15310.02200.01470.07220.1706 ✗Gat-SAE0.03840.01420.07470.16470.03450.01510.07890.1752 ✗JR-SAE0.03550.01380.06670.14900.03270.01480.07410.1723 ✗Mat-SAE0.02160.01410.06860.15880.01270.01540.07670.1613 ✗Tk-SAE0.01540.0125 0.05580.13600.00960.01440.07180.1577 ✗A-SAE0.01720.01300.05670.13700.01430.01450.0713 0.1429 ✗Pretrained----0.03330.01490.07870.1796 ✗Deleuzian (Ours) 0.0128 0.01190.0560 0.1230 0.0055 0.0137 0.06800.1538 DeBERTaBART labels MethodIMDB CoNLL-2003 IMDB CoNLL-2003 NERPOSChunkNERPOSChunk ✓LDA0.63940.04290.00440.00620.34730.63260.38750.0870 ✗ICA0.69360.12510.01950.01260.69311.45780.71436.1319 ✗Van-SAE0.68930.08690.02520.01730.59830.27190.16470.0447 ✗Gat-SAE0.68830.12230.02510.39820.63910.39820.40540.3208 ✗JR-SAE0.69080.11500.02480.01700.69310.44160.21110.0883 ✗Mat-SAE0.68360.08680.01890.01640.69311.1200.49540.2143 ✗Tk-SAE0.68580.08390.01660.01670.59800.34780.2045 0.0399 ✗A-SAE0.68590.0775 0.0141 0.0058 0.55470.37540.19590.0415 ✗Deleuzian (Ours)0.6849 0.06650.01610.01430.5974 0.2148 0.06390.0419 ASTPythiaAvg. Rank↓ labels Method AudioSetCoNLL-2003 NERPOSChunk ✓LDA0.01640.07420.00720.0089- ✗ICA0.02340.13780.03310.00886.85±2.29 ✗Van-SAE0.01770.14980.02720.00834.65±1.56 ✗Gat-SAE0.01860.14800.02310.00866.65±1.42 ✗JR-SAE0.01810.15070.02770.00855.75±0.94 ✗Mat-SAE0.01860.17540.03200.00885.70±1.90 ✗Tk-SAE0.01690.13210.02030.00822.65±1.01 ✗A-SAE0.01690.13780.03310.00883.20±1.72 ✗Pretrained-0.17170.03440.0087- ✗Deleuzian (Ours) 0.0164 0.1121 0.0133 0.0080 1.65±0.85 in datasets, on a wide variety of tasks, models and modalities. On several cases, probe loss is midway between supervised LDA and the second most effective unsupervised method (typically TopKSAE). Note that LDA obtains poor results on BART over CoNLL-2003, which indicates that the additional hypothesis made by LDA compared to our method (normal distribution of concepts and homoscedasticity) are not satisfied in this particular case. On average over all datasets, our 6 Table 2: Evaluating the consistency of extracted concepts with MPPC on several tasks/datasets including WikiArt (WA), AudioSet (AS). CLIPDinoV2DeBERTaBARTAST ImNetWAImNetWAIMDBCoNLLIMDBCoNLLAS ICA0.4490.3880.2640.4060.1220.4400.9990.4200.296 Van-SAE0.840 0.9180.603 0.903 0.9860.4370.9960.4390.837 Gat-SAE0.3460.4150.2640.4010.8360.4530.9960.3570.399 JR-SAE0.3410.4400.2720.4240.8940.5360.9960.4390.449 Mat-SAE0.2250.2470.2010.2190.7070.3390.5060.2160.274 Tk-SAE0.7570.8610.5880.8240.8660.5940.9960.7610.601 Deleuzian (Ours)0.8210.856 0.7890.8430.980 0.5881.00.7680.830 approach is significantly the best classified among unsupervised approaches. Significance of the results is detailed in Appendix C. To complement the quantitative evaluation, we further analyze representative examples, which pro- vide evidence for the relevance and interpretability of the extracted concepts: in addition to the examples provided in Figure 1, we present qualitative results in Appendix E. 3.2CONSISTENCY ACROSS RUNS In order to measure consistency of a concept extraction method, we measure the Maximum Pairwise Pearson Correlation (MPPC) (Wang et al., 2025) 10 times between sets of concepts extracted with different random seeds, and report the average. Therefore, a MPPC closer to 1 indicates a higher consistency. We present MPPC in details and discuss its statistical significance in Appendix D. Results from Table 2 show that our approach generally extracts more consistent concepts than other models, except for VanillaSAE, but this method reaches much lower concept quality and diversity according to Table 1. 3.3CONCEPT STEERING: QUALITATIVE EVIDENCE OF CAUSAL INFLUENCE A possible use of extracted concepts is to explicitly modify the behavior of a model, by steering its internal concepts. We provide qualitative examples of steering, using the method described in 2.4, highlighting the causal influence of concepts on the output of a model. Discriminative Steering on CLIP Steering the inner representation of an image encoder may be used to perform style transfer, in a similar fashion as previous works (Wynen et al., 2018). From WikiArt, we consider two concepts corresponding to artistic styles (identified empirically from im- ages), namely Romanticism and Abstract paintings. Starting from a romantic painting of a sailing ship, we inhibit the Romanticism concept, and boost the Abstract paintings one. The resulting steered embedding shifts the painting’s representation such that its nearest neighbors in the WikiArt dataset are abstract sailing ships (Figure 3). Figure 3: Steering a painting style in CLIP activations: target is represented by its nearest images. Romanticism is set to zero, while Abstract is steered positively by the same magnitude. 7 Steering BART BART (Lewis et al., 2020) is a text encoder–decoder model which, without fine- tuning, typically reproduces its input sequence. Here, we steer the final transformer layer of its encoder before passing the modified representation into the decoder. We analyze the steering ef- fects of a concept with highest activations corresponding to country names (Figure 4). Inhibiting this concept (α < 0) causes BART to replace “Rio de Janeiro” with “February”, forming a coher- ent sentence with no geographical indication. In the same fashion, its leads to replacing the word “country” by the word “city”. Positive values of α encourage the model to evoke country names, even in sentences without geographic context. In particular, this leads to frequent mentions of the United States, highlighting a potential bias in BART. Figure 4: Steering the concept of countries in a BART model for three sentences (in gray), using α = +5 and α =−5 in each case 3.4ABLATION STUDIES Table 3: Ablation study in terms of performance (Probe Loss) and diversity (effective rank). Our approach is the last line. skewnessweighting probe loss↓effective rank↑max. pairwise cos. ↓ inputconceptCLIPDeBERTaCLIPDeBERTaCLIPDeBERTa spaceidentif.WikiArtCoNLLWikiArtCoNLLWikiArtCoNLL acts.Tk-SAE✗0.01250.083996.1183.90.29000.3716 acts.KMeans✓0.01330.118424.314.60.86850.9195 diffTk-SAE✗0.01340.1093340.5109.20.34070.1737 diffKMeans✗0.01280.084117.95.650.65040.8357 diffKMeans✓ 0.01190.0665124.4182.00.56770.3908 We conduct an ablation study of our method, to assess the impact of three aspects on its perfor- mance. First, we evaluate the interest of learning from differences between samples, rather than directly from the samples themselves (i.e. changing the input space). Second, we evaluate the im- pact of using a clustering to identify the concepts, by replacing the the KMeans clustering of our approach with an SAE, trained on the activations or the differences. Finally, we evaluate the impact of weighting the KMeans clustering by the inverse skewness. Since the objective of this weight- ing is to increase diversity, we also report an evaluation of the diversity of the extracted concepts, measured by the effective rank (Roy & Vetterli, 2007; Skean et al., 2025) , as well as the maximum pairwise cosine among concept directions that quantifies redundancy. Results, computed on CLIP activations on WikiArt, and DeBerta on CoNLL NER attributes, are reported in Table 3. These results, most notably those for KMeans on activations and TopKSAE on differences highlight the impact of representing differences in activations. Moreover, these results highlight the importance of using the inverse skewness of pairwise differences as KMeans weights, allowing the extraction of a much larger, and less redundant sets of concepts, according to both effective rank and maximum pairwise cosine metrics. Figure 5 evaluates the performance of our method while extracting a number of concepts smaller than 6144. Only 2000 concepts are needed to outperform every concurrent method on CLIP, WikiArt artist task. This highlights the ability of our method to efficiently recover concepts. 8 Figure 5: Performance of our Deleuzian approach using less than 6144 concepts, on CLIP, WikiArt artist task. 4RELATED WORKS Concept-Based Interpretability Identifying the internal mechanism of a neural network corre- sponding to a precise concept provides valuable insights into the network’s behavior. Arik & Liu (2020) perform clustering on multi-layer activations, in order to determine similar images, not to extract interpretable concepts. Prior studies have investigated the extent to which a classification probe can be learned directly on model hidden representations (K ̈ ohn, 2015). Probe-based concept extraction has been used extensively in NLP (Gupta et al., 2015). These studies suggest that LLMs linearly represent the truth or falsehood of factual statements (Marks & Tegmark, 2024). Simi- lar analyses have also been applied to computer vision (Alain & Bengio, 2017) or reinforcement learning (Lovering et al., 2022). However, probe-based concept extraction only captures correlation (not causation) and heavily relies on curated data to extract concepts (Belinkov, 2022). To address this problem, Concept Bottleneck Models (CBM) (Koh et al., 2020) structure the network to make predictions through a layer of human-defined concepts, enabling intervention but requiring labeled concept supervision. Contrast-Consistent Search probes for an axis in the activation space, corre- sponding to the presence or absence of a concept (Burns et al., 2023), however it uses predefined contrastive groupings, and thus cannot uncover new concepts. Similarly, TCAV (Kim et al., 2018) and ACE (Ghorbani et al., 2019) perform concept extraction upon a predefined list. Sparse Autoencoders Sparse autoencoders (SAEs) (Lee et al., 2007) are a sparse dictionary learn- ing technique that aims to find a sparse decomposition of data into an overcomplete set of features. They typically enforce sparsity via an L 1 penalty. In recent years, SAEs have been applied to neu- ral networks to learn an unsupervised dictionary of interpretable features tied to concepts from a hidden representation (Bricken et al., 2023; Cunningham et al., 2023). Various extensions of sparse autoencoders have been proposed with modified activation functions, such as JumpReLU (Raja- manoharan et al., 2024b), TopK (Gao et al., 2025), and BatchTopK (Bussmann et al., 2024) sparse autoencoders. Other works seek hierarchies of features by extracting nested dictionaries (Bussmann et al., 2025; Zaigrajew et al., 2025). Analogous methods have been developed in order to find rela- tions between different layers of a same network, including transcoders (Dunefsky et al., 2024) and crosscoders (Lindsey et al., 2024).ArchetypalSAE (Fel et al., 2025) constrains the decoder training in order to gain stability, while Spade (Hindupur et al., 2025) is a distance based SAE. Further use of extracted concepts Identifying the mechanism corresponding to a semantic concept within a neural network enables new uses of the analyzed model. For example, studies use extracted 9 concepts to analyze the circuits related to a specific task (Conmy et al., 2023; Dunefsky et al., 2024), or to measure the importance of concepts in model inner representations (Fel et al., 2023). Concept extraction techniques can also be used to perform steering, i.e. controlling the behavior of a model by explicitly modifying its internal concepts (Zhou et al., 2025). When applied to multiple models in parallel, concept extraction methods allow construction of shared concept spaces (Thasarathan et al., 2025), automating naming of CLIP concepts (Rao et al., 2024) and quantification of similarities between models (Wang et al., 2025). 5CONCLUSION Discussion We present a novel approach for extracting human-interpretable “concepts” from neu- ral network activations and evaluate it across five models and three modalities. Our method is sim- ple and can be interpreted as an unsupervised form of discriminant analysis. Probe loss evaluation shows that the extracted concept space captures attributes expected from labeled datasets, and our approach outperforms existing methods on this metric. Moreover, the concepts are stable across mul- tiple runs, enabling consistent analyses, and the method supports lossless interventions on internal representations. These results suggest that explicitly representing inter-sample differences, in line with Deleuze’s notion of concepts, can improve both the quality and utility of extracted concepts. Limitations Although our method is fully unsupervised, its evaluation depends on labeled datasets. Consequently, interpretable concepts that do not align with the available labels may in- cur high probe losses, even if they are highly meaningful but subtle or specific. Evaluating without labels would require a theoretically justified proxy for interpretability, which remains lacking; spar- sity alone does not satisfy this criterion (Sharkey et al., 2025). All evaluations are performed in concept spaces of 6,144 dimensions (8× the activation dimension), except for an ablation. While some studies use even higher-dimensional projections, further in- creasing dimensionality could bias our evaluation, given the limited number of attributes and data samples relative to the potential size of the concept space. Exploring higher-dimensional spaces could nonetheless reveal additional characteristics of concept extraction methods. Our approach assumes that concepts can be represented as linear projections. This assumption is empirically validated across five models spanning different categories and modalities. However, a model with inner representations that violate this assumption could exist and would require adapting the method. Perspectives Our method is fully unsupervised and extracts concepts that represent repeated di- rections in a model. Consequently, a method that can automatically name or interpret these concepts would greatly enhance the scope and applicability of the findings, enabling more comprehensive analyses across datasets, modalities, and models. Such generalization could facilitate understanding of model behavior, provide interpretable axes for interventions, and support downstream tasks that leverage concept-level information. We provide qualitative examples of concept steering. As our method allows lossless steering, such intervention on model inner representations could be used at a larger scale, for example to adapt to a specific domain. Acknowledgment this work was partially funded by the Agence Nationale de la Recherche (ANR) for the STUDIES project ANR-23-CE38-0014-02. It was made possible by the use of the FactoryIA supercomputer, financially supported by the Ile-De-France Regional Council. REPRODUCIBILITY STATEMENT Our results can be reproduced, following the method described in section 2 and Appendix A. Cor- responding code is provided as supplemental material. REFERENCES Miika Ahdesm ̈ aki and Korbinian Strimmer. Feature selection in omics prediction problems using cat scores and false nondiscovery rate control. Annals of Applied Statistics, 4(1):503–519, 2010. 10 Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes, 2017. URL https://openreview.net/forum?id=ryF7rTqgl. Sercan Arik and Yu-Han Liu. Explaining deep neural networks using unsupervised clustering. In Proc. Workshop Hum. Interpretability Mach. Learn, p. 377–389, 2020. Peter Baylies.Wikiart dataset, 2020.URL https://w.kaggle.com/datasets/ steubk/wikiart. Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219, 2022. Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International Conference on Machine Learning, p. 2397–2430. PMLR, 2023. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, et al. Towards monosemanticity: Decom- posing language models with dictionary learning. Transformer Circuits Thread, 2, 2023. Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt. Discovering latent knowledge in lan- guage models without supervision. In The Eleventh International Conference on Learning Rep- resentations, 2023. Bart Bussmann, Patrick Leask, and Neel Nanda. Batchtopk sparse autoencoders. In NeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning, 2024. Bart Bussmann, Noa Nabeshima, Adam Karvonen, and Neel Nanda. Learning multi-level features with matryoshka sparse autoencoders. In Forty-second International Conference on Machine Learning, 2025. Pierre Comon. Independent component analysis, a new concept? Signal processing, 36(3):287–314, 1994. Arthur Conmy, Augustine Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri ` a Garriga- Alonso. Towards automated circuit discovery for mechanistic interpretability. Advances in Neural Information Processing Systems, 36:16318–16352, 2023. Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoen- coders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023. G. Deleuze.Diff ́ erence et r ́ ep ́ etition.Biblioth ` eque de philosophie contemporaine : Histoire de la philosophie et philosophie g ́ en ́ erale. Presses Universitaires de France, 1968.ISBN 9782130585299. URL https://books.google.gm/books?id=8lEwAAAAYAAJ. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio (eds.), Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), p. 4171–4186, Minneapolis, Minnesota, June 2019. Association for Com- putational Linguistics. doi: 10.18653/v1/N19-1423. URL https://aclanthology.org/ N19-1423/. Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. Transcoders find interpretable llm feature circuits. Advances in Neural Information Processing Systems, 37:24375–24410, 2024. Thomas Fel, Victor Boutin, Louis B ́ ethune, R ́ emi Cad ` ene, Mazda Moayeri, L ́ eo And ́ eol, Mathieu Chalvidal, and Thomas Serre. A holistic approach to unifying automatic concept extraction and concept importance estimation. Advances in Neural Information Processing Systems, 36:54805– 54818, 2023. 11 Thomas Fel, Ekdeep Singh Lubana, Jacob S Prince, Matthew Kowal, Victor Boutin, Isabel Papadim- itriou, Binxu Wang, Martin Wattenberg, Demba E Ba, and Talia Konkle. Archetypal sae: Adap- tive and stable dictionary learning for concept extraction in large vision models. In Forty-second International Conference on Machine Learning (ICLR), 2025. R. A. Fisher. The use of multiple measurements in taxonomic problems. Annals of Eugenics, 7(7): 179–188, 1936. Ronald A Fisher. Frequency distribution of the values of the correlation coefficient in samples from an indefinitely large population. Biometrika, 10(4):507–521, 1915. Leo Gao, Tom Dupr ́ e la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. Scaling and evaluating sparse autoencoders. In International Confer- ence on Representation Learning (ICLR), 2025. Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP 2017, New Orleans, LA, 2017. Amirata Ghorbani, James Wexler, James Y Zou, and Been Kim. Towards automatic concept-based explanations. Advances in neural information processing systems, 32, 2019. Yuan Gong, Yu-An Chung, and James Glass. Ast: Audio spectrogram transformer. Interspeech 2021, 2021. Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Pad ́ o. Distributional vectors encode referential attributes. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, p. 12–21, 2015. Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. DeBERTa: Decoding-enhanced BERT with disentangled attention. In International Conference on Learning Representations, 2021. URL https://openreview.net/forum?id=XPZIaotutsD. Georg Wilhelm Friedrich Hegel. Wissenschaft der Logik: Die objective Logik, volume 2. Johann Leonhard Schrag, 1816. Sai Sumedh R Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba. Projecting assump- tions: The duality between sparse autoencoders and concept geometry. In ICML 2025 Workshop on Methods and Opportunities at Small Scale, 2025. URL https://openreview.net/ forum?id=AKaoBzhIIF. Joshua Zhexue Huang, Michael K Ng, Hongqiang Rong, and Zichen Li. Automated variable weight- ing in k-means type clustering. IEEE transactions on pattern analysis and machine intelligence, 27(5):657–668, 2005. A. Hyv ̈ arinen and E. Oja.Independent component analysis:algorithms and applica- tions.Neural Networks, 13(4):411–430, 2000.ISSN 0893-6080.doi: https://doi.org/10. 1016/S0893-6080(00)00026-5. URL https://w.sciencedirect.com/science/ article/pii/S0893608000000265. Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip, July 2021. URL https://doi.org/10.5281/ zenodo.5143773. Sonia Joseph, Praneet Suresh, Lorenz Hufe, Edward Stevinson, Robert Graham, Yash Vadi, Danilo Bzdok, Sebastian Lapuschkin, Lee Sharkey, and Blake Aaron Richards. Prisma: An open source toolkit for mechanistic interpretability in vision and video. arXiv preprint arXiv:2504.19475, 2025. Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav). In International conference on machine learning, p. 2668–2677. PMLR, 2018. 12 Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. Concept bottleneck models. In International conference on machine learning, p. 5338–5348. PMLR, 2020. Arne K ̈ ohn. What’s in an embedding? analyzing word embeddings through multilingual evaluation. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, p. 2067–2073, 2015. Honglak Lee, Chaitanya Ekanadham, and Andrew Ng. Sparse deep belief net model for visual area v2. Advances in neural information processing systems, 20, 2007. Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. Bart: Denoising sequence-to-sequence pre- training for natural language generation, translation, and comprehension. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 7871–7880, 2020. Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Jared Batson, and Chris Olah. Sparse crosscoders for cross-layer features and model diffing. Transformer Circuits, 2024. URL https://transformer-circuits.pub/2024/crosscoders/index.html. Stuart Lloyd. Least squares quantization in pcm. IEEE transactions on information theory, 28(2): 129–137, 1982. Charles Lovering, Jessica Forde, George Konidaris, Ellie Pavlick, and Michael Littman. Evaluation beyond task performance: analyzing concepts in alphazero in hex. Advances in neural information processing systems, 35:25992–26006, 2022. Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, p. 142–150, Portland, Oregon, USA, June 2011. Association for Computational Linguistics. URL http: //w.aclweb.org/anthology/P11-1015. Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets. In First Conference on Language Modeling, 2024. URL https://openreview.net/forum?id=aajyHYjjsk. Geoffrey J McLachlan. Discriminant analysis and statistical pattern recognition. John Wiley & Sons, 2005. Glenn W Milligan. An examination of the effect of six types of error perturbation on fifteen cluster- ing algorithms. psychometrika, 45(3):325–342, 1980. Friedrich Nietzsche. G ̈ otzen-D ̈ ammerung oder Wie man mit dem Hammer philosophirt. CG Nau- mann, 1889. Maxime Oquab, Timoth ́ e Darcet, Th ́ eo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Pretten- hofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. Plato. The republic – book vi, c. 375 BCE. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision. In Marina Meila and Tong Zhang (eds.), Proceedings of the 38th International Conference on Machine Learning, volume 139 of Proceedings of Machine Learning Research, p. 8748–8763. PMLR, 18–24 Jul 2021.URL https://proceedings.mlr.press/v139/radford21a. html. 13 Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, Janos Kramar, Rohin Shah, and Neel Nanda. Improving sparse decomposition of language model acti- vations with gated sparse autoencoders. Advances in Neural Information Processing Systems, 37: 775–818, 2024a. Senthooran Rajamanoharan, Tom Lieberum, Nicolas Sonnerat, Arthur Conmy, Vikrant Varma, J ́ anos Kram ́ ar, and Neel Nanda. Jumping ahead: Improving reconstruction fidelity with jumprelu sparse autoencoders. arXiv preprint arXiv:2407.14435, 2024b. C Radhakrishna Rao. The utilization of multiple measurements in problems of biological classifica- tion. Journal of the Royal Statistical Society. Series B (Methodological), 10(2):159–203, 1948. Sukrut Rao, Sweta Mahajan, Moritz B ̈ ohle, and Bernt Schiele. Discover-then-name: Task-agnostic concept bottlenecks via automated concept discovery. In European Conference on Computer Vision, p. 444–461. Springer, 2024. Olivier Roy and Martin Vetterli. The effective rank: A measure of effective dimensionality. In 2007 15th European signal processing conference, p. 606–610. IEEE, 2007. Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision (IJCV), 115(3):211–252, 2015. doi: 10.1007/s11263-015-0816-y. Jean-Paul Sartre and Arlette Elka ̈ ım-Sartre. L’existentialisme est un humanisme, 1946. Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, Nicholas Goldowsky-Dill, Stefan Heimersheim, Alejandro Ortega, Joseph Bloom, et al. Open problems in mechanistic interpretability. arXiv preprint arXiv:2501.16496, 2025. Oscar Skean, Md Rifat Arefin, Dan Zhao, Niket Nikul Patel, Jalal Naghiyev, Yann LeCun, and Ravid Shwartz-Ziv. Layer by layer: Uncovering hidden representations in language models. In Forty-second International Conference on Machine Learning, 2025. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ram ́ e, Morgane Rivi ` ere, et al. Gemma 3 technical report. arXiv preprint arXiv:2503.19786, 2025. Harrish Thasarathan, Julian Forsyth, Thomas Fel, Matthew Kowal, and Konstantinos Derpanis. Universal sparse autoencoders: Interpretable cross-model concept alignment. arXiv preprint arXiv:2502.03714, 2025. Bart Thomee, David A. Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li. Yfcc100m: the new data in multimedia research. Commun. ACM, 59(2):64–73, January 2016. ISSN 0001-0782. doi: 10.1145/2812802. URL https://doi. org/10.1145/2812802. Erik F. Tjong Kim Sang and Fien De Meulder. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, p. 142–147, 2003.URL https: //w.aclweb.org/anthology/W03-0419. Junxuan Wang, Xuyang Ge, Wentao Shu, Qiong Tang, Yunhua Zhou, Zhengfu He, and Xipeng Qiu. Towards universality: Studying mechanistic similarity across language model architectures. In The Thirteenth International Conference on Learning Representations, 2025. URL https: //openreview.net/forum?id=2J18i8T0oI. Daan Wynen, Cordelia Schmid, and Julien Mairal. Unsupervised learning of artistic styles with archetypal style analysis. Advances in Neural Information Processing Systems, 31, 2018. Vladimir Zaigrajew, Hubert Baniecki, and Przemyslaw Biecek. Interpreting clip with hierarchical sparse autoencoders. In Forty-second International Conference on Machine Learning, 2025. 14 Xiangrui Zeng and Hongyu Zheng. Cs sparse k-means: An algorithm for cluster-specific feature selection in high-dimensional clustering. arXiv preprint arXiv:1909.12384, 2019. Dylan Zhou, Kunal Patil, Yifan Sun, Karthik lakshmanan, Senthooran Rajamanoharan, and Arthur Conmy. LLM neurosurgeon: Targeted knowledge removal in LLMs using sparse autoencoders. In ICLR 2025 Workshop on Building Trust in Language Models and Applications, 2025. URL https://openreview.net/forum?id=aeQeXlG2Pw. 15 AAPPENDIX: IMPLEMENTATION DETAILS All our experiments are using a set of 6144 concepts, except for ICA, that is unable to represent a number of dimensions larger than D, the dimension of model activations. Therefore, ICA experi- ments are ran inD = 768 dimensions. TopKSAEs are trained using a TopK activation function, with k = 32. We select a learning rate of 10 −5 , that minimizes its reconstruction error on CLIP activations over ImageNet. For VanillaSAE, GatedSAE and JumpReLUSAE, we select the L 1 penalization coefficient reaching the lowest probe loss. From a sweep of 7 values between 10 −9 and 10 −3 , we select 10 −8 for VanillaSAE, 10 −6 for GatedSAE and 10 −5 for JumpReLUSAE. Concerning MatryoshkaSAE, we use groups of sizes [512, 1024, 1536, 3072], in order to represent progressively larger latent dictionaries. For Independent Component Analysis we used the scikit-learn (Pedregosa et al., 2011) implemen- tation of FastICA (Hyv ̈ arinen & Oja, 2000), with a log hyperbolic cosine to approximate the neg- entropy, a SVD whitening and the extraction of multiple components in parallel. BAPPENDIX: DETAILS ON EXPERIMENTAL SETUP All datasets used in our experiments (section 3) are reported in Table 4 with their main characteris- tics. When available, we use the train/test splits provided. As WikiArt has no predefined train/test sets, we use its even samples (0, 2, 4...) as a train set, and the other ones as the test set. Note that WikiArt is actually a set of data with three different label types, thus could be considered as three different datasets. Globally we thus have a much larger variety of experimental settings than in comparable previous works. Since we are interested in identifying concepts, all tasks relate to classification but they exhibit a deep variety in their nature, due to the type of data handled (text, image, audio) and how the data have to be considered to address the task. For example, the identifying sentiments on IMDB requires to take into account full sentences while the chunking task in CoNLL act at the token level. Table 4: Datasets used in our experiments. DatasetModality Label Type (number of classes)Train/Test SizeURL ImageNet-100ImageObject categories (100)50k / 5k WikiArtImageArtist (129), Style (27), Genre (11)40k / 40k IMDBTextSentiment (binary, sentence-level)25k / 25k CoNLL-2003TextNER (9), POS (47), Chunking (23, token-level)288k / 67k AudioSetAudioAudio event categories (527)18k / 17k The model encoders we considered in our experiments are summarized in Table 5. All the models were downloaded from huggingface, except for CLIP from OpenClip (Ilharco et al., 2021) and DinoV2 from PyTorch Hub. The model size is the number of parameters and since all of them were encoded in float32 their actual size in memory is this number multiplied by four. AST (Gong et al., 2021) relies on an image ViT that was trained on ImageNet-21k then finetuned on AudioSet. BART (Lewis et al., 2020), for its base version, was pre-trained “on the same data as BERT (Devlin et al., 2019)” that is “a combination of books and Wikipedia data”. CLIP (Radford et al., 2021) was trained “on publicly available image-caption data” that is images-caption pairs from the Web and publicly available datasets such as YFCC 100M (Thomee et al., 2016). The creator of the model did not release the dataset to avoid its use “as the basis for any commercial or deployed model”. DeBERTa (He et al., 2021) was trained on deduplicated data (78G) including original Wikipedia (English Wikipedia dump; 12GB), BookCorpus (6GB), OpenWebText (public Reddit content; 38GB), and STORIES (a subset of CommonCrawl; 31GB). DinoV2 (Oquab et al., 2023) was trained on the LVD-142M dataset, that was assembled and curated by the authors of the model. 16 Table 5: Pretrained models used in our experiments. The Size is the number of parameters (in millions). ModelModality VersionSize Training dataURL DeBERTaTextbase99 M BookCorpus, Wikipedia, OpenWeb- Text, STORIES BART (encoder)Textbase139 MBooks, Wikipedia DinoV2ImageViT-B/1486 MLVD-142 CLIPImageViT-B/16150 MopenAI private: web, YFCC100M... ASTAudio10-10-0.459387 MAudioSet, ImageNet-21k Figure 6: Pairwise comparisons of methods on AST-Audioset. Our method is better able to recover at least 366/527 attributes compared to concurrent methods. CAPPENDIX: SIGNIFICANCE OF PROBE LOSS RESULTS Table 1 reports the median probe loss for each task. In Figure 6, we perform attribute-wise compar- isons on AST-Audioset, the studied task comprising the largest number of attributes. The numbers represent how many times the method of each row better recovers the attributes than the methods on the column. For instance, last row show that our method attributes of Audioset. Our method is able to better recover at least 366/527 attributes (69.4%) than other methods. Per- forming a Wilcoxon signed-rank test, we obtain a statistic of 106584 with a p-value of 1.7× 10 −26 , rejecting the null hypothesis thus proving the significance of those probe loss results. In a similar fashion on CLIP-WikiArt, our method reaches a lower probe loss than TopKSAE on 140/167 attributes (83.8%, even with TopKSAE reaching a lower probe loss on the “style” at- tributes), obtaining a test statistic of 12671 and a p-value of 7.9 × 10 −20 , rejecting the null hy- pothesis. DAPPENDIX: STATISTICAL SIGNIFICANCE OF MPPC The Maximum Pairwise Pearson Correlation (MPPC) was proposed by Wang et al. (2025) as a similarity indicator between models. 17 D.1DEFINITION OF MPPC To compare two sets of extracted concepts A and B, ρ A→B i is defined as the maximum pairwise Pearson correlation between the i-th concept of A and all concepts of B. Withf A i the vector containing values for each sample for the i-th concepts of A, μ A i and σ A i its mean and standard deviation (respectively forf B j , μ B j and σ B j ): ρ A→B i = max j E[(f A i − μ A i )(f B j − μ B j )] σ A i σ B j (3) Then, MPPC A→B is defined as the arithmetic mean of ρ A→B i over all i, quantifying the extent to which the concepts in A are represented in B. In order to measure consistency of a concept extraction method, we measure MPPC 10 times between sets of concepts extracted with different random seeds, and report the average. Therefore, a MPPC closer to 1 indicates a higher consistency. D.2STATISTICAL SIGNIFICANCE IN OUR CASE With ρ i the maximum pairwise coefficient (Eq. 3) for k target features of length N , and H 0 the hypothesis of features having no linear relationship. Using the Fischer z-transformation (Fisher, 1915) z = artanh(r)∼∥(0, 1 √ N − 3 ) P(max i (r i ) > x) = 1− P(r ≤ x) k P(ρ i > x) = P(max i (z i ) > artanh(x)) P(ρ i > x) = 1− Φ(artanh(x) √ N − 3) k With k = 6144 (corresponding to the main experiments), and L = 10000 being largely lower than the size of the most used datasets, we obtain P(ρ i > 0.3)≈ 10 −206 , thus reject H 0 . EAPPENDIX: QUALITATIVE EXAMPLES OF EXTRACTED CONCEPTS Image Concepts We present in figure 7 three examples of concepts extracted from image models, from different datasets. The concepts are represented by the images with their nine highest acti- vations. The name of the concepts are empirically set from the images. Displayed concepts are extracted from CLIP’s activations, with figs. 7a to 7c extracted from ImNet, figs. 7d to 7f from WikiArt and corresponding to paintings content, and figs. 7g to 7i corresponding to artistic styles. Text Concepts In Table 6 and Table 7, we represent 3 textual concepts. For each concept, we dis- play the 3 sentences containing the highest token-wise concept values, and underline tokens among the top-100. FAPPENDIX: ADDITIONAL STEERING EXAMPLES Textual Concept: Baseball Extracted from DeBERTa, over CoNLL-2003. Enhancing this con- cept (positive values of alpha) causes replacement of any sport-specific terms (football, basketball) by their baseball equivalent. Those changes affect mentions of teams, leagues and scoring methods. • (± 0) The best sport is basketball, NBA is the best → (+3.75) The best sport is baseball, MLB is the best • (± 0) He scored 3 touchdowns in the first half→ (+4.5) He scored 3 RBI in the first inning • (± 0) The New York Knicks beat the Los Angeles Lakers→ (+3.75) The New York Yan- kees beat the Los Angeles Dodgers 18 (a) Puppies(b) Underwater Animals(c) Guitars (d) Portraits of children(e) Paintings of cliffs(f) Paintings of Venice (g) Popart(h) Cubist paintings(i) Minimalist paintings Figure 7: Nine examples of visual concepts extracted from CLIP, over ImageNet and WikiArt. Representing the 9 images with the highest activations for each. 19 Table 6: Examples of textual concepts extracted from DeBERTa on CoNLL-2003. Each column is a concept with three representative texts. Concept names are ours. Sports AchievementsLast NamesNationalities Seven athletes went into Fri- day’s penultimate meeting of the series with a chance of winning the prize . Katarina Studenikova (Slo- vakia)beat6-Karina Habs udova. OneRomanianpassenger was killed, and 14 others were injured on Thursday when a Romanian -registered bus collided with a Bulgarian one in northern Bulgaria, police said. Russia’sdoubleOlympic championSvetlanaMas- terkova smashed her second world record in just 10 days on Friday when she bettered the markfor the women’s 1,000 metres. HendrikDreek man(Ger- many) vs.Greg Rused ski (Britain). He said a Turkishcivil avi- ation authority official had made the same point and he noted that a Turkish plane had a similar accident there in 1994. Jamaican veteran Merlene Ottey, who beat Devers in Zurich after just missing out on the gold medal in Atlanta after a photo finish, had to settle for third place in 11.04. The Greek socialist party’s executive bureau gave the green light to Prime Minis- ter Costas Simitis to call snap elections, its general secre- tary Costas Skandalidis told reporters. A Polish school girl black- mailedtwowomenwith anonymous letters threaten- ing death and later explained that she needed money for textbooks,police said on Thursday. Table 7: Additional Examples of textual concepts extracted from DeBERTa on CoNLL-2003. Each column is a concept with three representative texts. Concept names are ours. Years from the 1990’sAgeGeopolitical Evolutions West lake, arrested in Decem- ber 1993and charged with heroin trafficking , sawed the iron grill off his cell window Machado, 19, flew to Los An- geles after slipping away from the New Mexico desert town of Las Cruces Peruvian guerrillas killed one man and tookeight peo- ple hostage after taking over a village in the country’ s northeastern jungle Since taking over as cap- tain from Ne ale Fraser in 1994, Newcombe’ s record in tandem with Roche, his for- mer doubles partner, has been three wins and three losses. The 13 - year- oldgirl tried to extract 60 and 70 zlotys ( $22 and $26 ) from two residents of Sierakowice by threatening to take their lives. [...]is ready at any time without preconditions to enter peace negotiations The bullish comments for the coming year soothed ana- lysts and most shareholders , who were disappointed by the lower than expected profit for 1995 /96. On Tuesday night , Kevorkian attended the death of Louise Siebens, a 76 -year-oldTexas woman with amyotrophic lat- eral sclerosis [..] that is to endthe state of hostility 20 Figure 8: Steering Gemma3 captioning of The Starry Night, by Vincent Van Gogh, upon 2 concepts corresponding to positivity and negativity. Steering Gemma3 Image Captioning Our method can extract concepts from large decoder mod- els. From the text decoder of Gemma3-4B-PT (Team et al., 2025), we extract concepts over the IMDB dataset. We steer two concepts identified as corresponding to positivity/negativity during image captioning, see Figure 8. GQUADRATIC EXTENSION OF DELEUZIAN CONCEPTS As our Deleuzian approach is analog to a Linear Discriminant Analysis (LDA) as per subsection 2.3, it makes hypothesis about isotropic distribution of concepts in a model’s activation. However, we can derive an extension of the Deleuzian method, that does not make those hypothesis, analog with Quadratic Discriminant Analysis. The aim is to extract discriminant functions δ i : R → R from neural networks’ activations. Each function δ i must then correspond to an interpretable concept. G.1DISCRIMINANT FUNCTION δ A discriminant function δ is defined from a randomly sampled pair of samples x i ,x j the linear Deleuzian method considers linear concepts. δ(x) = b T x With b = x i − x j . Such formulation is equivalent to a linear discriminant analysis (with hypothesis of homoscedasticity). A Quadratic discriminant analysis (QDA) would require δ(x) =− 1 2 x T Ax + b T x + c with A = Σ −1 i − Σ −1 j and b = Σ −1 i μ i − Σ −1 j μ j and c =− 1 2 (μ T i Σ −1 i μ i − μ T j Σ −1 j μ j )− 1 2 log |Σ i | |Σ j | c is constant, and only affecting thresholding, but not the geometry of δ. Therefore we consider it neglectible. To form covariance matrices Σ i , Σ j , we use Ledoit-Wolf shrinkage on the 50-neighborhoods of x i and x j . Shrinkage methods are necessary to approximate covariance matrices on small datasets or with large dimensions. With S the sample covariance, and F = Tr(S) d it estimates the optimal α LW α LW = ( 1 T P T t=1 ||(x t − ̄x)(x t − ̄x) T − S|| 2 F )− (tr((S− F )· 1 T P T t=1 (x t − ̄x)(x t − ̄x) T − S tr((S− F )) 21 ˆ Σ = (1− α)S + α LW F We then use ˆ Σ i , ˆ Σ j as covariance matrix for neighborhoods of x i and x j . G.2CONCEPT SELECTION After extraction of N candidate discriminant δ i , the linear Deleuzian method is restrained to k concepts by performing feature weighted KMeans clustering. Distance As Deleuzian concepts δ i are linear, and only defined by a discriminant vector b ∈ R d , a trivial distance between two discriminant functions δ i ,δ j is the Euclidean distance between those vectors||b i ,b j || 2 . However, such metric cannot be computed on quadratic concepts having a more complex formulation. From two discriminants δ i ,δ j , we define the functional L 2 w metric as D 2 w (δ i ,δ j ) = Z R d (δ i (x)− δ j (x)) 2 w(x)dx or in probalistic terms D 2 w (δ i ,δ j ) = E x∼w [(δ i (x)− δ j (x))] Using the support measure w(x) = N (0,I). w(x) represents prior belief that data should follow a zero-mean, isotropic gaussian distribution. Using w ′ (x) = N (0,αI),α ∈ R + would only cause uniform scaling of D 2 , without modifying the underlying geometry. Noting ∆A = A i − A j and ∆b = b i − b j , we have D 2 w (δ i ,δ j ) = E x∼w [(− 1 2 x T ∆Ax + ∆b T x) 2 ] Developping, we consider the odd moments to vanish (as w is a zero-mean gaussian). Therefore we obtain D 2 w (δ i ,δ j ) = 1 4 E[(x T ∆Ax) 2 ] + E[(∆b T x) 2 ] Simplifying the linear term, we obtain E[(∆b T x) 2 = ∆b T E[x T ]∆b As x∼N (0,I), E[x T ] = I d . Thus E[(∆b T x) 2 ] = ∆b T I d ∆b =||∆b|| 2 Concerning the quadratic term, because x∼N (0,I) we have E[(x T ∆Ax) 2 ] = 2Tr(∆A 2 ) + Tr(∆A) 2 Quadratic parameters A i and A j are differences of covariance matrices formed with Ledoit-Wolf shrinking. Therefore, their diagonals are most likely similar, and dominated by constant isotropic offset. Then, we consider Tr(∆A) 2 = Tr(A i − A j ) 2 ≈ 0, and we get E[(x T ∆Ax) 2 ] = 2Tr(∆A 2 ) = 2||∆A|| 2 F Therefore, our functional L 2 w distance stands as follows : D 2 (δ i ,δ j ) = 1 2 ||A i − A j || 2 F +||b i − b j || 2 22 Centroids Recomputation Once we have defined a functional distance, the main crucial step of KMeans clustering is the iteratice centroids recomputation. Each δ i is assigned to its closest centroid ̄ C, then ̄ C is recomputed in order to minimize within cluster distortion. We recompute ̄ C (with parameters ̄ A, ̄ b) using the Fr ́ echet mean upon our functional L 2 w distance ̄ C = argmin δ X i w i D(δ,δ i ) ̄ C = argmin δ X i w i ( 1 2 ||A− A i || 2 F +||b− b i || 2 ) with ponderation weights w i (usually uniform, for unweighted mean) Using the A and b derivatives of ̄ C to minimize distortion : ∂ ∂A X i ( 1 2 ||A− A i || 2 F ) = 0 =⇒ X i w i (A− A i ) = 0 =⇒ A = X i w i A i ∂ ∂b X i ( 1 2 ||b− b i || 2 ) = 0 =⇒ X i w i (b− b i ) = 0 =⇒ b = X i w i b i Therefore, we use ̄ A = P i w i A i and ̄ b = P i w i b i as parameters of the centroid ̄ C G.3RESULTS AND DISCUSSION ABOUT QUADRATIC EXTENSION The obtained method is an exact generalization of our linear Deleuzian method to quadratic func- tions. Table 8 demonstrates that this extension reaches probe loss results better than SAE-based methods on CLIP-WikiArt, but does not outperform Linear Deleuzian concepts that is presented in the main paper. Such results may be due to the need to estimate covariance matrices on very high dimensional data. Table 8: Results of the Quadratic extension of Deleuzian concepts on CLIP-WikiArt Methods CLIP WikiArt ArtistStyleGenre Van-SAE0.0137 0.05580.1531 Tk-SAE0.01250.05580.1360 Linear-Deleuzian (Main Method) 0.01190.0560 0.1230 Quadratic-Deleuzian (Extension)0.01240.61600.1305 HAPPENDIX: LLM USAGE Beyond the usage of LLM described in the paper, that is part of the study, we used commercial services to polish the writting: find synonyms, rephrase sentences. 23