Paper deep dive
Adversarial Attacks on the Interpretation of Neuron Activation Maximization
GĂŠraldin Nanfack, Alexander Fulleringer, Jonathan Marty, Michael Eickenberg, Eugene Belilovsky
Models: pre-trained CNNs (multiple architectures), ResNet (ImageNet), VGG (ImageNet)
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 6:30:43 PM
Summary
The paper investigates the vulnerability of neuron activation-maximization interpretability methods in Deep Neural Networks (DNNs) to adversarial manipulation. The authors propose an optimization framework to perform 'push-down', 'push-up', and 'fairwashing' attacks, which allow an adversary to alter the interpretation of neurons while maintaining the model's original performance and accuracy. The study demonstrates that these attacks can effectively deceive interpretability techniques, raising concerns about the reliability of using such methods for mechanistic interpretations.
Entities (6)
Relation Signals (3)
Adversarial Model Manipulation â preserves â Model Performance
confidence 95% ¡ fine-tunes a pre-trained model with a loss that maintains its initial performance while changing the result of feature visualization.
Push-Down Attack â targets â Activation Maximization
confidence 95% ¡ The first proposed attack, push-down, aims to simply remove the current interpretation
Fairwashing â manipulates â Perceived Bias
confidence 90% ¡ fairwashing visualization attack aimed to manipulate the perceived bias of the model
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The internal functional behavior of trained Deep Neural Networks is notoriously difficult to interpret. Activation-maximization approaches are one set of techniques used to interpret and analyze trained deep-learning models. These consist in finding inputs that maximally activate a given neuron or feature map. These inputs can be selected from a data set or obtained by optimization. However, interpretability methods may be subject to being deceived. In this work, we consider the concept of an adversary manipulating a model for the purpose of deceiving the interpretation. We propose an optimization framework for performing this manipulation and demonstrate a number of ways that popular activation-maximization interpretation techniques associated with CNNs can be manipulated to change the interpretations, shedding light on the reliability of these methods.
Tags
Links
- Source: https://arxiv.org/abs/2306.07397
- Canonical: https://arxiv.org/abs/2306.07397
Trouble viewing inline? Open PDF directly â
Full Text
89,291 characters extracted from source content.
Expand or collapse full text
Adversarial Attacks on the Interpretation of Neuron Activation Maximization Geraldin Nanfack 1,2,â Alexander Fulleringer 1,2,â Jonathan Marty 3 Michael Eickenberg 4 Eugene Belilovsky 1,2 1 University of Concordia 2 Mila â Quebec AI Institute 3 Columbia University 4 Flatiron Institute geraldin.nanfack, alexander.fulleringer, eugene.belilovsky@concordia.ca jonathan.n.marty@gmail.comeickenberg@flatironinstitute.org Abstract The internal functional behavior of trained Deep Neural Networks is notoriously difficult to interpret. Activation-maximization approaches are one set of techniques used to interpret and analyze trained deep-learning models. These consist in finding inputs that maximally activate a given neuron or feature map. These inputs can be selected from a data set or obtained by optimization. However, interpretability methods may be subject to being deceived. In this work, we consider the concept of an adversary manipulating a model for the purpose of deceiving the interpretation. We propose an optimization framework for performing this manipulation and demonstrate a number of ways that popular activation-maximization interpretation techniques associated with CNNs can be manipulated to change the interpretations, shedding light on the reliability of these methods. 1 Introduction Deep Neural Networks (DNNs) can be trained to perform many economically valuable tasks [29, 25]. They are already pervasive in many sectors, and their prevalence is only expected to increase over time. With increasing computational power and ever more available amounts of data, Neural Network (N) architectures are growing in size and executing more and more intricate tasks. Given the increasing size and complexity of DNNs, interpreting how they function, a discipline that always lags behind the cutting edge, may experience an ever harder time keeping up with new developments. However, for certain classes of critical applications, close inspection and guarantees of functionality will be more and more important, especially in heavily regulated and high-stakes domains. Here we ask: could a malicious actor conceal the true functionality of a N from an interpretability method by modifying the N? Given the increasing capacity of the architectures, this is likely to be a progressively more probable concern. Focusing on the continuously popular feature visualization [51, 35, 36] method we propose to create an optimization procedure to manipulate the interpretation of individual neurons of the network while keeping its final behavior the same. A successful modification of the interpretation results while keeping outputs constant is evidence for the manipulability of the interpretation approach. In this work, we concentrate on convnet architectures for which interpretation by activation maximization or feature visualization methods [51, 48] has been popular. We study the feature visualization of a neuron or channel norm via activation maximization and attempt to modify it while maintaining trained network outputs and accuracy. We investigate how to characterize these attacks quantitatively and show three different attacks which can effectively manipulate and explicitly obfuscate interpretations. The first proposed attack,push-down, aims to simply remove the current interpretation, replacing it with any other interpretation. The second attack, termedpush-up, aims to replace the images â Equal contribution. Preprint. Under review. arXiv:2306.07397v1 [cs.LG] 12 Jun 2023 Responding to rough textured set of circular objects Responding only to the âdecoyâ goldfish class Model Creator / Attacker Adversarial manipulation Interpreter Similar Output manipulated high performance model high performance model Top-K Images top-K images Interpretation Figure 1: Illustration of the attack model for our adversarial interpretability manipulation. Top-5 images that best activate a given neuron, seemingly capturing a shared semantic concept over classes that an interpreter may describe and/or use an external tool to describe [22, 34]. In our framework, we assume the model creator can manipulate the model before it is released to the interpreter. In this case, they create a model which might lead to interpreting the selected neuron as not relating any semantic concept shared by multiple class categories. with a specific category of images, allowing a more targeted manipulation. The final attack we consider, motivated by recent related work on feature attribution methods [1, 44], is thefairwashing visualization attack aimed to manipulate the perceived bias of the model as seen by an interpreter. Consider as motivation a situation where an adversary is indifferent to deploying a biased model, but is constrained to provide model access to a regulator (the interpreter). Critically,we assume that the interpreter may not have access to labels related to the particular bias exploited by the adversaryâs model. The interpreter can use feature visualization methods (top-kimages) to try to understand the internal logic of neurons and may visually detect that neurons are biased towards a previously un-categorized but undesirable bias. To prevent rejection of the biased model by the interpreter, the adversary may use a set of data with annotated bias attribute [47] (unavailable to the interpreter) to try to perform an attack by fine-tuning the model to make the feature visualization look fairer while maintaining the performance of the model and its overall unfair output. To date, most previous works on interpretability manipulability (including fairwashing) have focused on the manipulability of interpretability techniques such as feature attribution [44, 21] tailored for model predictions. Little attention has been paid to the manipulability of neuron interpretability techniques. This is in spite of the fact that this latter type of interpretability method is becoming increasingly popular because it provides a fine-grained understanding of inner structures of DNNs [35, 36, 40]. Notably it has also been applied to create mechanistic interpretations [33, 7] which are argued to be robust as they directly link the function of neurons. We note that the maximization operation by construction is losing important information about the functional behavior, leading to the potential of mis-intepretation, and suggesting the possibility of manipulation. The primary contributions of our work are to first propose three distinct attacks on feature visualization and approaches and considerations to quantify and characterize their success. We then demonstrate all three of our attacks can achieve a degree of success (see illustration in Figure 1). This suggests that this class of interpretation methods must be used with caution and also cast doubt on the feasibility of using this tool to build complete mechanistic interpretations. 2 Related Work A growing body of literature has investigated the interpretability of Convolutional Neural Networks (CNNs) and the lack of robustness under different manipulations of interpretability methods. Interpretability methods.Previous work aiming to provide interpretability of NNs can be grouped into two broad categories. Firstly, there are works that developinterpretable-by-designmethods that provide interpretations without relying on external tools. These methods usually couple traditional layers with various types of interpretable components. Examples range from concept explanations [8, 27, 20, 14, 4], feature attributions [46, 37, 2] to part of object disentanglement [52, 43]. Secondly, there are methods usually calledpost-hocthat aim to explain and understand either specific components (e.g., weights, neurons, layers) or outputs of atrainedNN. To interpret the output of models for a particular data instance (local interpretability), while feature attribution methods [41, 31, 42] such as saliency maps assign a weight to each input feature corresponding to its importance on the modelâs output, counterfactual examples aim to give the minimal changes required to change the modelâs 2 output [18, 16]. There are post-hoc approaches that aim to interpret the internal logic of particular NNs through their components and representations. For example, there are methods that focus on layer representations throughconcept vectors[26, 53], on sub-network interpretability through circuits[5, 6], and individual neurons via e.g., feature visualization. Our work focuses on feature visualization, which is one of the most popular techniques to understand the learned features of individual neurons [54, 35]. Interpretability manipulation.There is a recent trend to analyze the reliability of interpretable techniques through the lens ofstability. Stability aims to study to what extent the interpretability technique is statistically robust to reasonable input perturbations and model perturbations [21, 49]. Most works that study input and model manipulability focus on feature attributions. For example, [12] designs adversarial input perturbations to change feature attributions in a targeted way, and [21] shows that such manipulation can be performed throughadversarial model manipulation, realized by fine-tuning a pre-trained model to change feature attributions while keeping the same accuracy of the original model. Despite sharing similarities with this work thanks to the use of adversarial model manipulation, instead of studying the manipulability of feature attribution methods, we focus on neuron interpretability, which brings different challenges such as thewhack-a-moleproblem explained in Sec. 3.3. Besides input and model manipulability, recent works [1, 3, 44] have raised thefairwashingissue, which is the risk of misleading the assessment of unfairness of models by providing model interpretations that look fair, but are not. Part of our work studies the fairwashing risk for feature visualization, which has not been investigated to date. Finally, the most closely related work to ours is [13], which shows the targeted manipulability ofsyntheticfeature visualizations (defined in Sec. 3.1) by early stopping during optimization. Different from this previous work, we instead study the manipulability of feature visualization under an adversarial model manipulation. 3 Methods We introduce our notation, attacks, threat models, and attack success characterization methods. 3.1 Notations and Background We denote byD=(x i ,y i ) N i=1 a dataset for supervised learning, wherex i âR d is the input and y i â 1,...,Kis its class label. Letf θ denote a N,f (l) θ (x)defines activation maps ofxon the l-th layer, which can be decomposed intoJsingle activation mapsf (l,j) θ (x). In particular,f (l,j) θ (x) is a matrix if the l-thlayer is a 2D-convolutional layer and a scalar if it is a fully connected layer. We aim to understand the internal behavior of individual units through feature visualization, generically defined by activation maximization [32, 48], i.e., x â âargmax xâX f (l,j) θ (x),(1) whereXcan be a finite set of data, e.g.,X=Dor a continuous spaceX âR d , and(l,j)is the pair of layerland neuronj. In Eq. 1, when the layerlis a convolutional layer, in the rest of the paper, we aggregate the activation mapf (l,j) θ (x)using its spatial squaredâ 2 -normâĽf (l,j) θ (x)⼠2 2 , and subsequently refer tojas the channel index. Additionally, we mainly focus on the case whereX=D is a set of natural images, and we denote by top-kimages the set of images that have thekhighest values of activations for a given pair(l,j). WhenX âR d , following [54], the resultx â will be calledsyntheticfeature visualization. 3.2 Attack Framework We consider feature visualization with top-kimages and propose an adversarial model manipulation that fine-tunes a pre-trained model with a loss that maintains its initial performance while changing the result of feature visualization. More formally, given a set of training dataD, a pre-trained model with parametersθ initial , and an additional set of images (e.g., a set of top-kimages)D attack , our attack framework consists in the following optimization min θ (ÎąL A (D,D attack ;θ) + (1âÎą)L M (D;θ,θ initial )),(2) whereθare parameters of the updated modelf θ ,L M (.)is the loss that aims to maintain the initial performance of the modelf θ initial , andL A (.)is the attack loss. For the maintain objective, when viewing final outputsf θ (.)as a conditional distribution, our maintain loss is the distillation loss L M (D;θ,θ initial ) =L CE (f θ initial (.)||f θ (.))[23], whereL CE is the cross entropy loss between the original model outputs and the attacked model outputs on training dataD. As defined, this maintain 3 loss enforces the fine-tuned model to keep the same predictions as the initial model with the objective of making the two models close in model space. Depending on the type of attack, the attack loss L A (.)can vary and is defined in the next sections. 3.3 Push-Down and Push-Up Attack Given a set of top-kimages from feature visualization, denoted byD (l,j) attack , that best activate the layerl and channeljof the initial modelf θ , our first attack aims to push to zero the activations of examples inD (l,j) attack . This attack is called thepush-downattack, and we propose the following objective for all channels of a layerlsimultaneously L A (D,D attack ;θ) = J l X j=1 X x â âD (l,j) attack âĽf (l,j) θ (x â )⼠2 2 ,(3) whereJ l is the set of channels of the layerl. Note that it is possible to attack a single channel or channels from multiple layers. Here we focus on attacking all the channels in a layer (see Sec. 4.1). In thepush-updecoy attack, given a set of examples inD decoy , we aim to make these images appear in the result of top-kimages for all the channels of a particular layerl. For this purpose, we propose the following objective, where[.] + ismax(.,0): L A (D,D decoy ;θ) = J l X j=1 X x â âD decoy X xâD [âĽf (l,j) θ (x)⼠2 2 ââĽf (l,j) θ (x â )⼠2 2 ] + .(4) This aims to make activations of examples inD decoy larger than all the activations of training examples. Characterizing Push-Down and Push-Up AttacksWe propose two approaches to characterize the effectiveness of an adversarial attack on the top-kimages of feature visualization. Kendall-Ď.We take a (potentially large) set of imagesD kĎ and compute the initial rankingsR init,j of images inD kĎ w.r.t. their initial activations values for thejchannel. Similarly, we compute the final rankingsR final,j using the same images, but on final (post-attack) activations values of the same channelj. The Kendall-Ď j score is the Kendall rank correlation coefficient betweenR init,j and R final,j . We can also aggregate this metric over all channels. Higher values of Kendall-Ďscores can be interpreted as higher similarity in the ordering of image activations between channels. As a result, the Kendall-Ď j score can be used as a metric to see how much a channelâs behavior has changed. CLIP-δ. We use an external, generic, visual representation model, the CLIP image encoder [39] to allow measuring the semantic changes in the top-kimages. Given a particular layer and a channelj, here we compute the average cosine self-similarity between the CLIP embeddings of initial top-k images, which we denote by Ě C init,init j,j and the average similarity between embeddings of initial top-k images and final ones (after the attack), denoted by Ě C init,final j,j . The proposed CLIP-δscore for a channel jis defined asCLIP-δ j = ( Ě C init,init j,j â Ě C init,final j,j )/( 1 Nâ1 P N p=1 Ě C init,init j,p̸=j ). Intuitively, this quantifies the relative semantic change of top-kimages w.r.t. CLIP embeddings and a high score can be interpreted as the fact that the channeljhas made semantically significant changes in the top-kimages. The Whack-A-Mole Problem.A natural question in our framework is whether the behavior and interpretation of one neuron can be simply moved to another neuron through the optimization process, for example, the Push-Down objective can be reduced by permutation. We call this thewhack-a-mole problem. To ensure that this does not occur, we study the previously described metrics and check that the attacked networkâs channels are not strongly correlated to other channels in the pre-attack network. Given thej-th channel, we define the following two metrics that measure this property. Kendall-Ď-W j - UsingD kĎ we obtain the maximum Kendall-Ďscore between ranked listsR init,j and R final,i wherei̸=jand normalize it by dividing it by the initial maximum Kendall-Ďscore i.e. the score overR init,j andR init,i wherei̸=j. CLIP-W j - Using the top-kimages in the initial model and channeljwe obtain max i̸=j Ě C initial,final j,i /max i̸=j Ě C initial,initial j,i comparing to all top-kimages in other channels of the fi- nal model, normalized against that same similarity metric in the initial CLIP scores. 4 Channel 2 of conv_5: Kendall- : 0.575, CLIP- : 0.238 mosque cocktail shaker stupateapotmosque Initial top-K binocularstripod soap dispenser beer bottle mortarboard Final top-K Channel 170 of conv_5: Kendall- : 0.639, CLIP- : 0.181 peacockpeacockpeacockpeacockbell pepper Initial top-K green mamba green snake police vantree frog green mamba Final top-K Channel 9 of conv_5: Kendall- : 0.072, CLIP- : 0.447 Yorkshire terrier file Norfolk terrier Norwich terrier chiffonier Initial top-K barrel rain barrel space heater rain barrel saltshaker Final top-K Figure 2: Push-down all-channel attack onConv5of AlexNet. All initial images have been replaced by other images. The final validation performance was 56.2%, a drop of less than half a percent. 3.4 Fairwashing Interpretability Attack We consider a threat model as discussed in Sec. 1 where the attacker has a set of protected attribute labels they use to hide bias from an interpreter without labeled data. More formally, given a model f θ , which isunfairaccording to a certain metric of unfairness, a set ofJof neurons whose top-k images lookunfair, we aim to answer the question: can we make an adversarial model perturbation by fine-tuning a pre-trained model, maintaining its performance and its unfairness while making the top-kimages of theJneurons appearfairer? In this formalization, answering affirmatively to this question corresponds to succeeding in the fairwashing attack. We design the fairwashing attack, using the same attack framework 2 defined in Sec. 3.2. One alternative to make the top-kimages appear fairer would be to enforce the matching between top-k activations for different groups of the protected attribute. However, it was empirically observed that this objective fails to generalize on an unseen set because it focuses only on the tail of the distribution of activations. We, therefore, propose a simple yet effective attack objective that allows reducing the discrepancy between the distribution of pre-activations of two groups of dataD 0 attack andD 1 attack , partitioned with respect to protected attribute (e.g., gender). For this purpose, we use the following loss (corresponding to the maximum mean discrepancy [17] with the feature functionĎ(x) = (x,x 2 )) L A (D,D 0 attack âŞD 1 attack ;θ) =âĽÎź l 0 âÎź l 1 ⼠2 2 +âĽĎ l 0 âĎ l 1 ⼠2 2 ,(5) whereD 0 attack ,D 1 attack are two groups of data partitioned w.r.t. the labeled protected attribute (e.g., race or gender),Îź l p (withpâ0,1) is a vector of scalarsÎź (l,j) p =E x p âźD p attack [f (l,j) θ ]of first-order moments for layerland neuronj, and similarlyĎ (l,j) p =E x p âźD p attack (f (l,j) θ ) 2 are second-order moments for the same neuron. This attack objective enforces the matching between the first two moments of two distributions (w.r.t. groups of protected attribute) of pre-activations of a neuron. 4 Experiments and Results We now describe the experimental setup and the results obtained after running attacks. For all of our attacks, we use the ImageNet [11] training set asD. We use the PyTorch [38] pretrained AlexNet [28] for our analysis. In Appx. B.2 we provide an ablation study on EfficientNet [45] with similar findings. More technical details regarding hyperparameters for all the attacks can be found in Appx. A. Push-down and Push-Up attack.For the push-down and up attack, we considerD (l,j) attack âDas the top-10images that maximally activate the channeljof layerl. For the push-up attack, we additionally considerD decoy as100randomly sampled images of a particular class to be used as decoy. Fairwashing attack.In order to run and evaluate the fairwashing attack, we need a dataset with a labeled protected attribute (e.g., gender or age) to be able to assess not only model unfairness but also thefairnessof feature visualization of a neuron. For this purpose, we use the ImageNet People Subtree dataset [47], which is a set ofâ14kimages with labeled demography (gender, race and age), derived from ImageNet-21k. We use the75â25%split for training and testing sets, andD 0 attack andD 1 attack are binary groups (w.r.t. protected attribute) from the training set. We estimate model unfairness using two popular measures of unfairness [50], namely the difference of disparate impact (DDI=|p(Ëy=c|z= 0)âp(Ëy=c|z= 1|), wherezis the protected attribute,cis a class andËyis 2 Note we use pre-activations to capture the entire and non-truncated distribution [7] 5 the predicted class) and difference of equal opportunity (DEO=|p(Ëy=c|z= 0,y=c)âp(Ëy= c|z= 1,y=c)|) estimated on testing data [50, 19]. Inspired by the fairness assessment in regression and clustering, we use two measures to quantify the feature visualization unfairness. The first one looks at the entire distribution of activations and is the Kolmogorov-Smirnov (KS) distance between the two conditional distributions of activations given protected attribute label [30]. The second one only focuses on the tail of the distribution of activations, i.e., activations of top-kimages, and is the balance [9] or ratio between the number of instances from top-kbelonging to the minority group over the number of instances in top-kbelonging to the majority group. Finally, following recent trends [24], we perform the fairwashing attack on the last but one layer. 4.1 Push-Down And Push-Up Attack Experiments Warm-up: Single-Channel Attack.To set a first evaluation point for our attack framework, we apply the push-down attack to one channel Figure 3 shows the visualization of top im- ages before and after. We can see that after optimization, the top-kactivating images of the neuron have been completely replaced by other images with different semantic concepts, sug- gesting a succesful attack with almost nearly no loss in accuracy (it decreases by0.04%). Initial top-K Final top-K Figure 3: Top images for a channel be- fore and after a single-channel Push- Down attack. One way of satisfying the attack objective perfectly in the single channel case is to set the channel weights to zero. This naive solution only loses0.2%is to simply set all the weights of the channel to zero. Specifically removing channel 0 (by masking) decreased the accuracy by0.2%. We thus consider more challenging settings. All-Channel Attack.Unlike the single-channel attack, the all-channel attack (change all neuron interpretation in a layer) does not have a trivial solution. Because some information needs to flow through the layer in order for classification to be successful, setting all channels to zero would result in catastrophic performance loss. We apply our attack framework toConv5of the AlexNet Model. In Figure 2 we show a selection of 3 chan- nels and the modifications achieved under the All-Channel Push-Down attack and the ag- gregate metrics (averages for all channels in a layer) are shown in Table 1.More vi- sual examples are provided in the Appendix.For the visualized channels (and those in Appendix) we observe a near complete replacement of the top-5images by other images. Further, the labels of the top images significantly change, with minimal to no residual overlap. This suggests that not only the images have changed but the semantic concepts that would be determined by an interpreter have likely changed. This is opposed to the model simply memorizing images to reduce and replacing them with semantically similar ones. We further confirm this in the appendix by showing validation set top-kimages which demonstrate that semantically they follow the same behavior as the training images (which are used for the actual attack). Overall, the attack seems to produce a generalized change in the behavior of the feature visualization of neurons. Layer/AttackCLIP-δKend-ĎCLIP-W Kend-Ď-W Acc.(%) Conv1 Push-Down0.0430.6820.9960.30256.1 Conv2 Push-Down0.0560.6120.9940.15156.3 Conv3 Push-Down0.1270.5730.9630.13056.1 Conv4 Push-Down0.2050.5480.9740.12256.2 Conv5 Push-Down0.2490.5300.9630.04856.2 Conv5 Push-Up0.1500.6540.9620.01156.3 EfficientNet L7 - Push-Down0.2620.5030.971-0.14577.5 Table 1:Average (over channels) attack metrics for an All-Channel Push-Down and Push-Up Attack for AlexNet (row 1-6) and Ef- ficientNet (row 7). We observe that the relative whack-a-mole metrics are low, suggesting this problem is not present for our at- tacks. Lower layers are more challenging to attack leading to lower CLIP score and higher Kendall-Ďas confirmed by visual intuition. Studying the metrics comparing the channels before and after modifica- tion, we can deduce several different behaviors. The first two channels ex- hibit relatively high Kendall-Ďscores, from which we conclude that the or- dering of image activations has not un- dergone severe changes. This means that likely only a subset of images, which includes the initial top-khas moved in rank. Studying the CLIP distance in both cases allows us to con- clude that there is significant semantic overlap in the initial and final top-k, which can be confirmed by visual inspection. 6 Whake-a-mole for channel 2 of conv_5 mosque cocktail shaker stupa Intial top-K for channel 2 projectileking penguinpineapple Final top-k, nearest channel: 47, Kendall- -W j :-0.082 car wheelbottlecapmanhole cover Final top-k, nearest channel: 187, CLIP-W j :0.971 Whake-a-mole for channel 193 of conv_5 Bernese mountain dogAppenzeller Bernese mountain dog Intial top-K for channel 193 ChihuahuaAppenzeller Bernese mountain dog Final top-k, nearest channel: 163, Kendall- -W j :0.132 ChihuahuaAppenzeller Bernese mountain dog Final top-k, nearest channel: 163, CLIP-W j :0.991 Figure 5:We show the initial top images for two channels and beneath are the corresponding final top images of closest channels w.r.t Kendall-Ď-W j and CLIP-W j . 050100150200250 Sorted Channels 0.0 0.2 0.4 0.6 0.8 1.0 Clip Similarity Conv5 Push-Down: Clip Similarity of Nearest Channel Max Self-Similarity Max Cross-Similarity Figure 6:We compare initial CLIP similar- ity to other channels (blue) versus similarity after attack (red). Red and blue largely track each other for all channels. This is in contrast to the channel shown at the right, where the Kendall-Ďscore is close to zero, indicating a full re-ordering of the activations. As a consequence, the CLIP distance from initial to final is also much higher, which matches with a visual inspection. In general, we observe a substantial correspondence between our visual intuition and the CLIP-δ and Kendall-Ď, channels with low scores Kendall-Ďand high CLIP-δtend to change substantially. As illustrated in further examples in the Appendix one observed difference in these two metrics is that channels maintaining some similar classes in the top images will tend to have a lower CLIP-δ (suggesting less change). Whack-a-mole.We can further analyze the existence of the whack-a-mole problem by observing Fig. 5 which shows for a channel in the original model, the top-K image in the modified model which have the closest Kendall-Ď-W and CLIP-W scores (not including the channel itself). We observe that the first channel (channel 2 on figure) has little to no visually discernable similarity to nearby channels in the modified model as well confirmed by the Kendall-Ď-W. Indeed a majority of the channels look like this (see Appendix). On the other hand, we do observe similar images for the initial channel 193 and its nearest final one (163), which was picked as the most illustrative examples ("hard" one) where the red curve of Fig. 6 is above the blue one. However, for this "hard" example, more insight is given by investigating the CLIP-W j where the denominator notably measures the clip similarity to other channels in the original model. The score is less than or typically close to 1 suggesting that the original model already had a high similarity to another channel. Indeed in the Appendix for the second example, we confirm there is a very similar channel in the original model. To gain further insight into CLIP-W j in Fig.6, we further visualize the numerator and denominator for all the channels (red line) and sort them by the initial similarity to other channels (denominator). We observe that the red line is often below the blue line and if it exceeds it is not by a large relative amount, suggesting that channels with high whack-a-mole metrics are actually ones that already had similarities to other channels in the original model. Overall we conclude the presence of the whack-a-mole problem is minimal in our current attack. Effect of Depth.We now consider how the attack is affected by depth, with results for different layers of AlexNet shown in Tab. 1 and illustrated in Fig. 7. We observe that modifications of the earliest layers are significantly harder to achieve than for later layers as confirmed by the metrics and visual examination. We also observe a qualitative difference in the changes. For example, Conv 1 and Conv 2 are picking up low-level information such as color, edges, and textures and this is reflected in the type of modifications made to the images. If performance is maintained after the attack, it is likely that the modification objective did not have a strong impact, leading to little to no modification. This is reflected in the CLIP-δscores (see Table 1) and in visual examination (see Appendix for further examples). Several explanations can account for this. Firstly, there are fewer or no modifiable weights upstream to the attacked layer, leading to less flexibility to accommodate the competing natures of the combined objective compared to later layers. Secondly, the early-layer features, while somewhat malleable, must collectively perform a certain set of signal-filtering operations in order to be able to extract meaningful information. Performing strong modifications to the filters may lead to unrecoverable information loss downstream. We observe that the whack-a-mole metrics are also relatively high for this case using Kendall-Ď-W. On the other hand, the normalized CLIP-W score is close to 1 suggesting that this increase is not due to behavior being moved into the channel but due to existing redundancy in channels. 7 Initial top-K Final top-K conv3 conv1conv2conv4 Figure 7: Push-down attack on AlexNet across several layers. Channels are taken individually on each layer for layer ablation, and the results demonstrate that the top images are potentially vulnerable across all layers. The final attacked models all have a less than .5% drop from a default AlexNet. Channel 43 of conv_5: Kendall- : 0.740, CLIP- : 0.256 daisycornMaltese dogpapillon Shetland sheepdog Maltese dog Initial top-K goldfishgoldfishgoldfishwiggoldfishgoldfish Final top-K Channel 170 of conv_5: Kendall- : 0.619, CLIP- : 0.070 peacockpeacockpeacockpeacockbell peppertree frog Initial top-K goldfishpeacockgoldfishgoldfishgoldfishgoldfish Final top-K Figure 8: Examples of channels in all-channel push-up attack. The top images were successfully put in top images. The Kendall-Ďremains relatively high (>0.5) suggesting much of the channel behavior is preserved while the top activating images completely obfuscate the behavior. Push-Up Decoy Attack.We study a more targeted attack objective, namely one that actively pushes a set of selected images into the top activating images for every channel. This is achieved with Eq. 4, where the loss is non-zero as long as there exist images outside the group of selected images that activate higher than the group we intend to push up. This type of attack is more targeted and therefore likely harder than the push-down attack, which does not specify what images the top-kshould be replaced with. The push-up attack, if successful, can assign the same interpretation to every channel in a layer, making any interpretation attempt based on top-kimages fraught, or at least minimally informative. Fig. 1 shows the result of the push-up attack using a collection of images with the Imagenet label âGoldfishâ as the decoy set. Further, in Fig. 8 we show that for many channels of a layer, we can modify the top-10to contain a few or consist entirely of Goldfish images. The metrics in Table 1 also demonstrate substantial change and a low likelihood of whack-a-mole behavior. Studying the figure more closely, we observe that not only Goldfish, but also other images that share certain traits with the Goldfish images are also boosted, suggesting a degree amount of generality of the newly imposed selectivity, further explored in the Appendix. Initial top-K Push-Down Final top-K Push-Up Final top-K Synthetic Synthetic Figure 9: Synthetic feature visualization after our attack. We observe the visualization is largely decorrelated to top-knatural images. 8 0255075100125150175200 0.05 0.10 0.15 0.20 0.25 0.30 0.35 4020020 0.00 0.02 0.04 0.06 0.08 gender_0 gender_1 2010010 0.00 0.05 0.10 0.15 0.20 0.25 gender_0 gender_1 Initial KS distance (before attack) Final KS distance (after attack) Figure 10:Kolmogorov-Smirnov (KS) distance between the conditional distributions of each condition estimated on the annotated testing set. We sort the channels based on the initial KS and observe that after our fairwashing inter- pretability attack, each channels KS is drastically reduced. [0.0, 0.28)[0.28, 0.55)[0.55, 0.82)[0.82, 1.1) Interval of balance 0.0 0.1 0.2 0.3 0.4 % neurons with balance in interval % neurons w.r.t. balance Initial (before attack) Final (after attack) Figure 11:Percentage of the neurons according to their balance over the annotated testing set. Af- ter the attack, the percentage of neurons with low balance has decreased while the percentage of neu- rons with high balance has increased. 4.1.1 Synthetic Feature Visualization We study the impact of the Push-Down and Push-Up attacks on the synthetic activation-maximizing images of the channels under attack [51]. Synthetic activation-maximizing images are the result of an optimization problem over input pixels solved by gradient ascent on the channel activation under a norm constraint in pixel space. To avoid adversarial noise samples [15] it is necessary to jitter the input image or parameterize it as a smooth function[35]. In Fig. 9, we study the synthetic optimal images for several channels before and after the attack. By visual inspection, while the top-kimages change drastically, the synthetic optimal image is largely unaffected. The most common observed change (see also Appendix) forconv5is a low-frequency modulation of the pattern. We hypothesize that this is because the top-kattack most significantly modifies the weights of the attacked layer, which is a later layer preceded by several downsamplings. The lack of change in the synthetic optimal image suggests that the synthetic feature visualization and the top-kanalysis are, counter-intuitively, highly de-correlatable. Observe, for instance, that the left-hand synthetic image suggests selectivity for cats even when most of the top-kimages are goldfish. This is a worrying prospect for the top-kinterpretability method. Further, this does not permit the conclusion that the synthetic optimal image is more robust to attack, since we have not explicitly run an attack against it. Rather, this suggests the space of N weights and the possible functions they span is quite large, and can possibly accommodate more functionality, and attacks, than one might expect. 4.2 Fairwashing Feature Visualization We demonstrate the application of our fairwashing attack for feature visualization as defined Sec. 3.4. Given anunfair(according to a certain metric of unfairness) model and a set of neurons whose top-activating images lookunfair, we ask ourselves whether it is possible, by fine-tuning, to make the new set of images for the same neurons appearfairerwhile maintaining the same performance and bias of the initial model. We instantiate this fairwashing attack on an annotated subset of Imagenet data [47] (as described in Sec. 4) with gender as the protected attribute. We first estimate the model unfairness of the pre-trained AlexNet model using DDI and DEO unfairness measures. Tab. 2 reports these measures for the threehumanclasses of the ImageNet-1k dataset on which AlexNet is trained. According to this table, the initial AlexNet model is not totally fair, with the largest values of unfairness on theBaseball playerclass. We identified 200 neurons of the last but one layer whose MILAN [22] descriptions are related to humans (see Appendix for more details). We run our attack on all these neurons to prevent missing neurons whose biases may transfer to other ones. Fig. 10 shows the results of Kolmogorov-Smirnov distance between the distributions of activations conditioned on the two gender groups. It can be observed that after the attack, this distance has been drastically reduced, especially for highly biased neurons. This suggests the balance of the top-kis also improved. As can be seen in Fig. 11, the percentage of neurons whose top-kimages have a low balance (low fairness) has decreased, while the percentage of neurons with high balance has increased, thus making feature visualization fairer. Moreover, according to Tab. 2, the model has almost the same accuracy and almost the same measures of unfairness (all casesâ¤1%of relative difference for DDI andâ¤4% for DEO). Note that our attack did not enforce any fairness constraint on the output, the maintain 9 (a) Initial (testing) top-30 images for unit 800:balance: 0.250 (b) Final (testing) top-30 images for unit 800:balance: 0.579 Figure 12: Top-30 images obtained for unit 800 of the last but one layer of AlexNet. Green is used for images that stay in top-30 images after attack. Before the fairwahsing attack, (a) the initial top-30 images are gender-biased. After the fairwashing attack, (b) the top-30 are less gender-biased: balance (fairness) measure has almost doubled. On the other hand, the modelâs unfairness has not changed. Class Baseball playerBridegroomScuba diver Acc. DDIDEODDI DEODDI DEO Pre-Attack 56.45 3.3876.922.67 12.340.28 5.26 Post-Attack 56.56 3.1473.071.90 12.340.24 5.26 Table 2: Accuracy/fairness measures (DDI/DEO) computed respectively on the ImageNet val. set and on the annotated testing set. Both measures are relatively similar before and after the fairwashing attack while the model has decreased the bias perceived by the interpreter for feature visualizations. lossL M described in Sec. 3.2 was enough to also maintain model unfairness. We also depicted in Fig. 12 an example of a unit whose top-kimages were initiallybiased, but have been fairwashed after running the attack by almost doubling the balance measure. More examples of training and testing sets can be found in the appendix. 5 Conclusions, Limitations, and Broader Impact We demonstrated the adversarial model manipulability of feature visualization with top-k, proposing three attacks that pose varying threats. We provide experimental evidence that supports the success of our attacks, with little to no evidence of awhack-a-moleissue. Our metrics to systematically detect the presence of whack-a-mole may be imperfect as validating them requires inspecting all channels to validate correspondence. Future work may consider investigation of synthetic feature maps and how they may be attacked and generalization of the fairwashing attack beyond binary attributes. Broader Impact.The goal of our study has been to demonstrate a potential vulnerability in current interpretability methods and raise awareness of reliability and ethical risks. By showing the fairwash- ing attack, an apparent consequence is the possibility that an ill-intentioned individual uses this work to perform these attacks in order to release models that marginalize minority groups. However, we think that raising these risks is an essential first step towards addressing these vulnerabilities, and we hope our contributions provide a springboard for future discussion and protection efforts. 10 6 Acknowledgements We acknowledge support from OpenPhilanthropy and resources provided by Compute Canada and Calcul Quebec. We also thank Kaiyu Yang for the access to annotations of the ImageNet Subtree People dataset. References [1]Ulrich AĂŻvodji et al. âCharacterizing the risk of fairwashingâ. In:Advances in Neural Informa- tion Processing Systems34 (2021), p. 14822â14834. [2]David Alvarez Melis and Tommi Jaakkola. âTowards robust interpretability with self- explaining neural networksâ. In:Advances in neural information processing systems31 (2018). [3]Christopher Anders et al. âFairwashing explanations with off-manifold detergentâ. In:Interna- tional Conference on Machine Learning. PMLR. 2020, p. 314â323. [4]Pietro Barbiero et al. âEntropy-based logic explanations of neural networksâ. In:Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 36. 6. 2022, p. 6046â6054. [5]Jasmijn Bastings et al. â"Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classificationâ. In:Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022. 2022, p. 976â991. [6] Nick Cammarata et al. âCurve circuitsâ. In:Distill6.1 (2021), e00024â006. [7] Nick Cammarata et al. âCurve detectorsâ. In:Distill5.6 (2020), e00024â003. [8]Zhi Chen, Yijie Bei, and Cynthia Rudin. âConcept whitening for interpretable image recogni- tionâ. In:Nature Machine Intelligence2.12 (2020), p. 772â782. [9]Flavio Chierichetti et al. âFair clustering through fairletsâ. In:Advances in neural information processing systems30 (2017). [10] MohammadReza Davari et al. âReliability of CKA as a Similarity Measure in Deep Learningâ. In: (2022). arXiv:2210.16156 [cs.LG]. [11]Jia Deng et al. âImageNet: A large-scale hierarchical image databaseâ. In:2009 IEEE Confer- ence on Computer Vision and Pattern Recognition. 2009, p. 248â255.DOI:10.1109/CVPR. 2009.5206848. [12]Ann-Kathrin Dombrowski et al. âExplanations can be manipulated and geometry is to blameâ. In:Advances in neural information processing systems32 (2019). [13] Logan Engstrom et al. âAdversarial robustness as a prior for learned representationsâ. In:arXiv preprint arXiv:1906.00945(2019). [14]Mateo Espinosa Zarlenga et al. âConcept Embedding Models: Beyond the Accuracy- Explainability Trade-Offâ. In:Advances in Neural Information Processing Systems35 (2022), p. 21400â21413. [15]Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. âExplaining and harnessing adversarial examplesâ. In:arXiv preprint arXiv:1412.6572(2014). [16]Yash Goyal et al. âCounterfactual visual explanationsâ. In:International Conference on Machine Learning. PMLR. 2019, p. 2376â2384. [17] Arthur Gretton et al. âA kernel two-sample testâ. In:The Journal of Machine Learning Research13.1 (2012), p. 723â773. [18]Riccardo Guidotti. âCounterfactual explanations and how to find them: literature review and benchmarkingâ. In:Data Mining and Knowledge Discovery(2022), p. 1â55. [19] Moritz Hardt, Eric Price, and Nati Srebro. âEquality of opportunity in supervised learningâ. In: Advances in neural information processing systems29 (2016). [20]Marton Havasi, Sonali Parbhoo, and Finale Doshi-Velez. âAddressing Leakage in Concept Bottleneck Modelsâ. In:Advances in Neural Information Processing Systems. 2022. [21]Juyeon Heo, Sunghwan Joo, and Taesup Moon. âFooling neural network interpretations via adversarial model manipulationâ. In:Advances in Neural Information Processing Systems32 (2019). [22]Evan Hernandez et al. âNatural Language Descriptions of Deep Visual Featuresâ. In:Interna- tional Conference on Learning Representations. 2022. 11 [23]Geoffrey Hinton, Oriol Vinyals, and Jeff Dean.Distilling the Knowledge in a Neural Network. cite arxiv:1503.02531Comment: NIPS 2014 Deep Learning Workshop. 2015.URL:http: //arxiv.org/abs/1503.02531. [24] Pavel Izmailov et al. âOn feature learning in the presence of spurious correlationsâ. In:Ad- vances in Neural Information Processing Systems35 (2022), p. 38516â38532. [25]Jared Kaplan et al. âScaling laws for neural language modelsâ. In:arXiv preprint arXiv:2001.08361(2020). [26]Been Kim et al. âInterpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)â. In:International conference on machine learning. PMLR. 2018, p. 2668â2677. [27]Pang Wei Koh et al. âConcept bottleneck modelsâ. In:International Conference on Machine Learning. PMLR. 2020, p. 5338â5348. [28]Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. âImageNet Classification with Deep Convolutional Neural Networksâ. In:Advances in Neural Information Processing Systems25 (2012). [29] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. âImagenet classification with deep convolutional neural networksâ. In:Communications of the ACM60.6 (2017), p. 84â90. [30]Meichen Liu et al. âConformalized Fairness via Quantile Regressionâ. In:Advances in Neural Information Processing Systems. 2022. [31]Scott M Lundberg and Su-In Lee. âA unified approach to interpreting model predictionsâ. In: Advances in neural information processing systems30 (2017). [32]Aravindh Mahendran and Andrea Vedaldi. âUnderstanding deep image representations by inverting themâ. In:Proceedings of the IEEE conference on computer vision and pattern recognition. 2015, p. 5188â5196. [33] Neel Nanda et al. âProgress measures for grokking via mechanistic interpretabilityâ. In:arXiv preprint arXiv:2301.05217(2023). [34]Tuomas Oikarinen and Tsui-Wei Weng. âCLIP-Dissect: Automatic Description of Neuron Representations in Deep Vision Networksâ. In:arXiv preprint arXiv:2204.10965(2022). [35]Chris Olah, Alexander Mordvintsev, and Ludwig Schubert. âFeature Visualizationâ. In:Distill (2017). https://distill.pub/2017/feature-visualization.DOI:10.23915/distill.00007. [36]Chris Olah et al. âZoom In: An Introduction to Circuitsâ. In:Distill(2020). https://distill.pub/2020/circuits/zoom-in.DOI:10.23915/distill.00024.001. [37]Jayneel Parekh, Pavlo Mozharovskyi, and Florence dâAlchĂŠ-Buc. âA framework to learn with interpretationâ. In:Advances in Neural Information Processing Systems34 (2021), p. 24273â 24285. [38] Adam Paszke et al. âPytorch: An imperative style, high-performance deep learning libraryâ. In:Advances in neural information processing systems32 (2019). [39]Alec Radford et al. âLearning transferable visual models from natural language supervisionâ. In:International conference on machine learning. PMLR. 2021, p. 8748â8763. [40]Tilman Räukur et al. âToward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networksâ. In:arXiv e-prints(2022), arXivâ2207. [41]Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. â" Why should i trust you?" Explain- ing the predictions of any classifierâ. In:Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 2016, p. 1135â1144. [42]Ramprasaath R Selvaraju et al. âGrad-cam: Visual explanations from deep networks via gradient-based localizationâ. In:Proceedings of the IEEE international conference on computer vision. 2017, p. 618â626. [43]Wen Shen et al. âInterpretable Compositional Convolutional Neural Networksâ. In:Proceed- ings of the International Joint Conference on Artificial Intelligence. 2021. [44]Dylan Slack et al. âFooling lime and shap: Adversarial attacks on post hoc explanation methodsâ. In:Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society. 2020, p. 180â186. [45]Mingxing Tan and Quoc Le. âEfficientnet: Rethinking model scaling for convolutional neural networksâ. In:International conference on machine learning. PMLR. 2019, p. 6105â6114. [46] Rui Wang, Xiaoqian Wang, and David Inouye. âShapley Explanation Networksâ. In:Interna- tional Conference on Learning Representations. 2021. 12 [47]Kaiyu Yang et al. âTowards fairer datasets: Filtering and balancing the distribution of the people subtree in the imagenet hierarchyâ. In:Proceedings of the 2020 conference on fairness, accountability, and transparency. 2020, p. 547â558. [48] Jason Yosinski et al. âUnderstanding neural networks through deep visualizationâ. In:arXiv preprint arXiv:1506.06579(2015). [49] BIN YU. âStabilityâ. In:Bernoulli(2013), p. 1484â1500. [50]Muhammad Bilal Zafar et al. âFairness constraints: A flexible approach for fair classificationâ. In:The Journal of Machine Learning Research20.1 (2019), p. 2737â2778. [51] Matthew D Zeiler and Rob Fergus. âVisualizing and understanding convolutional networksâ. In:European conference on computer vision. Springer. 2014, p. 818â833. [52]Quanshi Zhang, Ying Nian Wu, and Song-Chun Zhu. âInterpretable convolutional neural networksâ. In:Proceedings of the IEEE conference on computer vision and pattern recognition. 2018, p. 8827â8836. [53]Bolei Zhou et al. âInterpretable basis decomposition for visual explanationâ. In:Proceedings of the European Conference on Computer Vision (ECCV). 2018, p. 119â134. [54]Roland S Zimmermann et al. âHow Well do Feature Visualizations Support Causal Under- standing of CNN Activations?â In:Advances in Neural Information Processing Systems34 (2021), p. 11730â11744. 13 Appendix A Hyperparameters and Training Details This section presents the details of the hyperparamters and training settings used to run our attacks. A.1 Push-Up and Push-Down Attacks We train for 2 epochs over the ImageNet-1k training set with a batch size of 256. We use theAdam optimizer with learning rate 1e-5. RegardingÎą, we employ a dynamic updating rule inspired byAlgorithm 1: Dynamical balancing of Distillation and CKA map lossin appendix A of Davari et alâs [10] in order to have better control over loss in accuracy. We initializeÎąas0.1(except for on the push-down attack forconv-2where useÎą= 0.01had more stable results). If the accuracy loss is greater than 0.5% we halve the current Îą. If it is less than 0.1% we doubleÎą. With this dynamic update, we aim to minimize the loss in accuracy while still ensuring the top images shifts. A.2 Fairwashing Attack Similarly to push-up and push-down attacks, for the maintain loss, we use the ImageNet-1k training set. For the fairwashing attack, we need annotations for the protected attribute. We consider the set of14865images derived from ImageNet-21k for which annotations of labeled demography (gender, race, and age) are available in the ImageNet People Subtree dataset [47]. We use75%of these images (annotated training set) in the maintain loss and use the rest of25%images (annotated testing set) for fairness assessment. We perform the attack with gender as the protected attribute and we binarize this attribute using the majority group defined as âmale in the imageâ. We also use the Adam optimizer with a learning rate of1e-5and we use a batch size of256for losses. No dynamic update forÎąwas needed, and we keep it toÎą= 0.1, corresponding to the initial value ofÎąfor push-up and push-down attacks. Finally, for the attack loss, we attack the last but one layer of AlexNet by considering the neurons (200 in total) whose MILAN [22] descriptions likely relate to humans. We accomplish this by inspecting the neuronsâ MILAN descriptions to get neurons whose descriptions contain one of the following words âfacesâ, âskinâ, âpersonâ, âhumanâ and âpeopleâ. 012 Training Epoch 0 50 100 150 200 250 300 350 Loss Attack Loss conv1 pushdown conv2 pushdown conv3 pushdown conv4 pushdown conv5 pushdown conv5 push-up 012 Training Epoch 2.0 2.2 2.4 2.6 2.8 3.0 Loss Maintain Loss conv1 pushdown conv2 pushdown conv3 pushdown conv4 pushdown conv5 pushdown conv5 push-up Figure 13: Sample training curves for the maintain and attack objectives. Late layers (conv5, conv4) are easier to attack compared to early ones (conv1, conv2 and conv3). The maintain loss is very close to its initial value after two epochs. 14 A.3 Optimization Curves We show in Figure 13 the evolution of attack and maintain losses across two epochs. It can be observed that the attack loss of late layers (conv 4, conv 5) decreases very quickly, and almost monotonically, showing the easiness to attack late layers. In contrast, early layers do not have the same behavior. We can also observe from the training curves that the maintain loss is almost close to its initial value after 2 epochs. This corroborates the observed accuracy preservation as shown in Table 1. B Additional Results This section shows additional illustrations and results for all the attacks. B.1 Additional Results for Push-down Attack on a Single Channel and on all Channels We show additional results for the push-down attacks on a single channel and on all channels simultaneously. B.1.1 Push-up Attack on Single Channel Figure 14 shows the results of initial top-kimages and final ones after running the push-down attack on every single channel. Except for channels 6 and 4 with relatively low CLIP-δscores, it can be observed that all other channels have semantically different final top-kimages compared to the initial ones. This can be also seen by higher values of CLIP-δscores. 15 Channel 0 of Conv 5: Kendall- : -0.030, CLIP- : 0.144 ladybugpomegranate black-footed ferret cucumberpill bottle Initial top-K hencombination lockhenladybug black-footed ferret Final top-K Channel 1 of Conv 5: Kendall- : 0.261, CLIP- : 0.196 shojiwindow screenwindow screenwalletwindow screen Initial top-K crossword puzzlechainlink fencehand-held computercrossword puzzlepillow Final top-K Channel 2 of Conv 5: Kendall- : 0.217, CLIP- : 0.277 mosquecocktail shakerstupateapotmosque Initial top-K mortarboardacademic gownmortarboardbeer bottlemilk can Final top-K Channel 3 of Conv 5: Kendall- : 0.191, CLIP- : 0.283 malamute German short-haired pointer papillonrevolvergreenhouse Initial top-K dogsledsnowmobilepaddlewhite stork electric locomotive Final top-K Channel 4 of Conv 5: Kendall- : 0.221, CLIP- : 0.148 sunscreensoccer ballshieldpinwheeljean Initial top-K indigo buntingpolice vanmaypolemacawshopping basket Final top-K Channel 5 of Conv 5: Kendall- : 0.358, CLIP- : 0.237 monarchmonarchpool tablescabbardmonarch Initial top-K vaultface powderpursewhistletile roof Final top-K Channel 6 of Conv 5: Kendall- : 0.084, CLIP- : 0.198 croquet ballagaricdoughgolf ballhen-of-the-woods Initial top-K golf ballAmerican egretgolf ballgolf ballgolf ball Final top-K Channel 7 of Conv 5: Kendall- : -0.008, CLIP- : 0.482 croquet ballbuckeyebuckeyebuckeyebuckeye Initial top-K lipstickBlenheim spanielwhistlehippopotamusground beetle Final top-K Channel 8 of Conv 5: Kendall- : 0.186, CLIP- : 0.278 strawberrybell peppercucumberbell pepperbell pepper Initial top-K pineappletiger beetlemantistiger beetlecauliflower Final top-K Channel 9 of Conv 5: Kendall- : -0.073, CLIP- : 0.467 Yorkshire terrierfileNorfolk terrierNorwich terrierchiffonier Initial top-K barrelspace heaterrain barrelnailhamster Final top-K Figure 14:Push-down attack on a single-channel ofConv5of AlexNet. All initial images have been replaced by other images. 16 B.1.2 Push-down All-Channel Attack This section presents additional results for the push-down attack on all channels at once. The results are obtained by attacking all the channels of the conv5 layer of AlexNet. We first show visual examples of results obtained from the training set of ImageNet and show its generalization to the validation set. Visual Examples.Figure 15 shows results obtained on 10 randomly chosen channels. It can be observed that all initial top-5 images were completely removed from the set of top-activating images. Additionally, channels with high CLIP-δscores such as channels 102 and 132, present semantically different images (initial vs final) with no overlap classes. In contrast, we observe that channels with low CLIP-δscores such as channels 254 and 227 usually share similar classes in top-activating images. Finally, from Kendall-Ďscores, we observe that channels that have high Kendall-Ď(e.g., channel 108 and 185) do not often have high values of CLIP-δscores, indicating that the weak change in channel behavior assessed by the Kendall-Ďis often related to low semantic change. 17 Channel 102 of Conv 5: Kendall- : 0.291, CLIP- : 0.432 Bernese mountain dog CardiganBrittany spanielAppenzellerPembroke Initial top-K combination lockflatwormAmerican egretAmerican egretdrake Final top-K Channel 108 of Conv 5: Kendall- : 0.610, CLIP- : 0.161 maypolepierumbrellaShetland sheepdogboxer Initial top-K rubber erasertheater curtainGerman shepherdplate rackscrew Final top-K Channel 132 of Conv 5: Kendall- : 0.516, CLIP- : 0.416 window screenwindow screenhoneycombwindow screenwindow screen Initial top-K chain mailbolo tiejackfruitjackfruitstrainer Final top-K Channel 183 of Conv 5: Kendall- : 0.522, CLIP- : 0.263 reelcar wheelwall clockanalog clockmanhole cover Initial top-K limpkintable lamplimpkinthimbleindigo bunting Final top-K Channel 185 of Conv 5: Kendall- : 0.752, CLIP- : 0.266 honeycombchainlink fencecrossword puzzlebarrelhognose snake Initial top-K boxerdingocleaverhare African hunting dog Final top-K Channel 186 of Conv 5: Kendall- : 0.501, CLIP- : 0.247 harmonicaking snakeking snakeacademic gownking snake Initial top-K sturgeonhartebeestgreat white sharkbighornsturgeon Final top-K Channel 216 of Conv 5: Kendall- : 0.596, CLIP- : 0.253 strainerhoneycomblong-horned beetlespider webgrille Initial top-K ski maskechidnaechidnaski maskski mask Final top-K Channel 227 of Conv 5: Kendall- : 0.546, CLIP- : 0.072 lighterremote controlgroenendaelodometerfile Initial top-K groenendaelLabrador retrieverNewfoundlandgroenendael Staffordshire bullterrier Final top-K Channel 232 of Conv 5: Kendall- : 0.568, CLIP- : 0.245 domemosquemosquedomedome Initial top-K domemosquebeer bottlechimechocolate sauce Final top-K Channel 254 of Conv 5: Kendall- : 0.211, CLIP- : 0.128 Siberian huskyPembrokePembrokecolliekit fox Initial top-K Eskimo dogkit foxmalamutelynxlynx Final top-K Figure 15:Push-down all-channel attack ofConv5of AlexNet. All initial top-5 images were completely removed from the new set of top-5 images, demonstrating the success of the attack. Channel indexes were chosen randomly. 18 Generalization on Validation Set.We evaluate the generalization of our attack on the validation set of ImageNet. This gives more insights to the change of feature visualization. Figures 16 and 17 show the initial top-kimages and final ones from training and validation sets for 10 randomly chosen channels. It can be observed that on every channel, from the validation set, at least one image from the initial top-5images is no longer present in final top-5images (for the majority of these channels, the first top-activating is no longer the top one). We also observe a complete replacement of top-5images on the validation set when Kenall-Ďscores and CLIP-δare respectively low and high simultaneously (e.g., channels 37 and 50 of Figure 16). Moreover, the general trends in training and validation are similar suggesting the attack is not just memorizing specific images but leading to a generalized change. 19 Channel 37 of Conv 5 "train": Kendall- : 0.208, CLIP- : 0.340 palacehand-held computerwindow screenpolice vanthimble Initial Training top-K gibbonguenongibbongibbonguenon Final Training top-K Channel 37 of Conv 5 "val": Kendall- : 0.307, CLIP- : 0.084 vending machinepalacetrolleybusbarbershopstreetcar Initial Validation top-K titigibbonguenonorangutanproboscis monkey Final Validation top-K Channel 48 of Conv 5 "train": Kendall- : 0.712, CLIP- : 0.012 mailbagGila monsterlawn mowerpencil boxCD player Initial Training top-K studio couchcribmailbagrulecradle Final Training top-K Channel 48 of Conv 5 "val": Kendall- : 0.691, CLIP- : 0.000 hair slideabacuscellular telephonepencil boxabacus Initial Validation top-K notebooktape playerchestpencil boxprojector Final Validation top-K Channel 50 of Conv 5 "train": Kendall- : 0.435, CLIP- : 0.580 ptarmigancoucalblack grouse red-backed sandpiperptarmigan Initial Training top-K stupamaillotptarmigantotem polebell cote Final Training top-K Channel 50 of Conv 5 "val": Kendall- : 0.560, CLIP- : 0.077 robinblack grousecoucalhouse finchblack grouse Initial Validation top-K stupastupapedestal flat-coated retrievertoy terrier Final Validation top-K Channel 71 of Conv 5 "train": Kendall- : 0.480, CLIP- : 0.223 strainerstrainermanhole coverhandkerchiefmanhole cover Initial Training top-K ocarinaocarinaporcupineocarinaocarina Final Training top-K Channel 71 of Conv 5 "val": Kendall- : 0.388, CLIP- : 0.009 ocarinaocarinawalletmanhole coverflute Initial Validation top-K ocarinaocarinafootball helmetladybugladybug Final Validation top-K Channel 75 of Conv 5 "train": Kendall- : 0.471, CLIP- : 0.267 jaguarthree-toed slothbox turtlemud turtleterrapin Initial Training top-K silky terrierSussex spanielsilky terriersilky terrierwig Final Training top-K Channel 75 of Conv 5 "val": Kendall- : 0.597, CLIP- : 0.084 box turtlebox turtlemud turtleechidnabeaver Initial Validation top-K Sussex spanielpapillonkitepapillonSussex spaniel Final Validation top-K Figure 16:Push-down all-channel attack ofConv5of AlexNet. For each channel, the first two rows are top-k images derived from the training set while the last two are derived from the validation set. 20 Channel 128 of Conv 5 "train": Kendall- : 0.406, CLIP- : 0.083 Scottish deerhoundScottish deerhound Staffordshire bullterrierScottish deerhoundMexican hairless Initial Training top-K Bouvier des FlandresLakeland terrierstandard schnauzer Bouvier des FlandresBorder terrier Final Training top-K Channel 128 of Conv 5 "val": Kendall- : 0.486, CLIP- : 0.032 standard schnauzer miniature schnauzerdingo miniature schnauzerbluetick Initial Validation top-K standard schnauzertoy poodlestandard schnauzerScotch terrierKerry blue terrier Final Validation top-K Channel 144 of Conv 5 "train": Kendall- : 0.605, CLIP- : 0.296 chainpretzelbrain coralchainpretzel Initial Training top-K common newtagamaeftnight snakeeft Final Training top-K Channel 144 of Conv 5 "val": Kendall- : 0.669, CLIP- : -0.005 chaingreen mambapretzelthunder snakegreen snake Initial Validation top-K chainagamathunder snakegreen mambapretzel Final Validation top-K Channel 158 of Conv 5 "train": Kendall- : 0.801, CLIP- : 0.219 barreltile roofchainlink fencehoneycombnight snake Initial Training top-K window screenthunder snakePersian catPersian catLhasa Final Training top-K Channel 158 of Conv 5 "val": Kendall- : 0.793, CLIP- : 0.022 honeycombhoneycombhoneycombhoneycombdishrag Initial Validation top-K digital watchhoneycombnecklaceflytick Final Validation top-K Channel 169 of Conv 5 "train": Kendall- : 0.514, CLIP- : 0.181 miniature poodleporcupineminiature poodlehyenahay Initial Training top-K great grey owl Bouvier des Flandresgreat grey owlgreat grey owlgreat grey owl Final Training top-K Channel 169 of Conv 5 "val": Kendall- : 0.522, CLIP- : 0.012 great grey owlgreat grey owlgreat grey owlgreat grey owl Irish water spaniel Initial Validation top-K great grey owlgreat grey owlgreat grey owllaptopgreat grey owl Final Validation top-K Channel 241 of Conv 5 "train": Kendall- : 0.627, CLIP- : 0.069 gobletred winewindow screengobletbasenji Initial Training top-K radio telescopefolding chairred winehourglassimpala Final Training top-K Channel 241 of Conv 5 "val": Kendall- : 0.588, CLIP- : 0.046 beer glassGreat DaneSalukiflamingodowitcher Initial Validation top-K Great Danebeer glassmalinoischainlink fenceSaluki Final Validation top-K Figure 17:Push-down all-channel attack ofConv5of AlexNet. For each channel, the first two rows are top-k images derived from the training set while the last two are derived from the validation set. 21 B.2 Ablation Study on EfficientNet It is important to show that the proposed attack methodology is not limited to AlexNet. In order to show that the attack can work on newer, more sophisticated neural nets, we have also run an ablation study on EfficientNet [45]. We select the third convolutional block in the Feature 7 layer and perform a push-down attack similar way to AlexNet. The visual results are shown in Appendix A and the metrics for the layer are given in Table 1. We observe similar effects to AlexNet; the top images are changed in terms of the exact images and the semantic concepts. We also observe relatively strong CLIP-δand Kendall-Ďchanges. Having confirmed the generality of our approach in this way, we leave a survey study over all relevant architectures to future work, computation power permitting. B.3 Effect of Depth We vary different layers of AlexNet and evaluate how the attack is affected by depth. Figure 19 shows results obtained on randomly chosen channels from conv1, conv2, conv3, and conv4 of AlexNet. It can be observed that the earliest layers conv1 and conv2 are harder to attack. This is materialized by high values of Kendal-Ďand low values of CLIP-δscores. When increasing the depth (conv3 and conv4) we observe a complete replacement in top-5images in channels 147 (conv3), 121 (conv4) and 124 (conv4), although some of these channels have low values of CLIP-δscores. 22 Channel 8 of features 7 conv block 3: Kendall- : 0.406, CLIP- : 0.103 tripodtripodtricycleiPodminiskirt Initial top-K jacamarcoucalstarfisheartrombone Final top-K Channel 51 of features 7 conv block 3: Kendall- : 0.358, CLIP- : 0.244 snailsnailchitonsnailcustard apple Initial top-K rubber erasercucumberotterhair slideprinter Final top-K Channel 99 of features 7 conv block 3: Kendall- : 0.241, CLIP- : 0.288 Walker houndlaptopwhippetBoston bull African hunting dog Initial top-K bibgasmaskshopping basketChristmas stockingpurse Final top-K Channel 102 of features 7 conv block 3: Kendall- : 0.514, CLIP- : 0.020 prisonbarrowpencil sharpenercontainer shipwreck Initial top-K analog clocktricyclemagnetic compassplastic bagloudspeaker Final top-K Channel 147 of features 7 conv block 3: Kendall- : 0.431, CLIP- : 0.491 Yorkshire terrierYorkshire terriertoy terriersilky terriervizsla Initial top-K lynxpacketjeandumbbellpomegranate Final top-K Channel 167 of features 7 conv block 3: Kendall- : 0.395, CLIP- : 0.492 fountain penwine bottlewine bottleoil filterbeer bottle Initial top-K football helmetamphibiansteel drumstoledock Final top-K Channel 176 of features 7 conv block 3: Kendall- : 0.512, CLIP- : 0.072 red winecarpenter's kitice lollycarpenter's kitletter opener Initial top-K beer glasslotionbeer glasscupvalley Final top-K Channel 193 of features 7 conv block 3: Kendall- : 0.561, CLIP- : 0.264 tennis ballShih-TzuShih-Tzuvacuumtoy terrier Initial top-K cowboy hatchainlink fencehair slidesquirrel monkeypitcher Final top-K Channel 208 of features 7 conv block 3: Kendall- : 0.530, CLIP- : 0.288 walking stickwalking stickwalking stickdeskmonitor Initial top-K wood rabbitpatiomilk canenvelopepuck Final top-K Channel 291 of features 7 conv block 3: Kendall- : 0.353, CLIP- : 0.438 English springerBrittany spanielBrittany spanielEnglish springerEnglish springer Initial top-K shovelmilk canstreet signbarrelbarrel Final top-K Figure 18:Push-down all-channel attack on Feature 7 block 3 of EfficientNet. All initial top-5 images were completely removed from the new set of top-5 images, demonstrating the success of the attack. Channel indexes were randomly chosen. 23 Channel 6 of Conv 1: Kendall- : 0.849, CLIP- : 0.000 brassaccordionfilewindow screenelectric fan Initial top-K accordionbrassfilewindow screenelectric fan Final top-K (a) Layer: Conv1. Channel 11 of Conv 1: Kendall- : 0.735, CLIP- : 0.007 brassspace heateraccordionwindow screenelectric fan Initial top-K brassspace heaterelectric fanwindow screenfile Final top-K (b) Layer: Conv1. Channel 1 of Conv 2: Kendall- : 0.580, CLIP- : 0.127 solar dishwindow screenwindow screengrillesolar dish Initial top-K solar dishsolar dishwindow screencrossword puzzlepick Final top-K (c) Layer: Conv2. Channel 106 of Conv 2: Kendall- : 0.505, CLIP- : 0.059 mailbagshojiblack grouseleafhoppersombrero Initial top-K mailbagleafhoppermonarchanalog clockcroquet ball Final top-K (d) Layer: Conv2. Channel 147 of Conv 3: Kendall- : 0.696, CLIP- : 0.230 window screenwindow screencleaverzebraelectric fan Initial top-K rugby ballbinderanemone fishparachutewall clock Final top-K (e) Layer: Conv3. Channel 214 of Conv 3: Kendall- : 0.622, CLIP- : 0.112 chain mailchaintigerspatulafig Initial top-K tigerfigmegalithchaintripod Final top-K (f) Layer: Conv3. Channel 121 of Conv 4: Kendall- : 0.672, CLIP- : 0.056 flatwormholster typewriter keyboard limpkinelectric ray Initial top-K affenpinscherBedlington terrierWeimaranerScottish deerhoundpolecat Final top-K (g) Layer: Conv4. Channel 124 of Conv 4: Kendall- : 0.509, CLIP- : 0.177 space bar typewriter keyboard slotdial telephoneslot Initial top-K menuearbottlecapdiapermenu Final top-K (h) Layer: Conv4. Figure 19:Push-down all-channel attack of on several layers of AlexNet. Channels indexes were selected randomly. While there are some changes in top-activating images of early layers (conv1 and conv2), they are not significant as materialized by low values of CLIP-δand high values of Kendall-Ď. For conv3 and conv4, we see a complete replacement of top-5 images on channels 147 (conv3), 121 (conv4), and 124 (conv4). 24 B.4 Additional Illustrations for Whack-a-mole This section provides further investigations into the existence of the whack-a-mole problem for the push-down attack on AlexNet. Zoom onto Channel 193 for Whak-a-mole.We begin by showing the full overview of the behavior of channel 193, selected as one "hard" case where similar initial images are found in final (post-attack) top-kimages of another channel. As discussed in Section 4.1, although similar initial images for channel 193 were found in channel 163 after the attack, it appears from the second row of Figure 20 that channel 193 was initially highly correlated with the channel 90 according to CLIP-δscore. Moreover, the fact that the CLIP-δ-W j is0.991<1shows that the nearest post-attack channel (channel 163) is not more correlated than the nearest pre-attack channel (channel 90) according to CLIP scores. This, therefore, limits the existence of the whack-a-mole problem on this channel. Whack-a-mole for channel 193 of Conv 5 Bernese mountain dog Appenzeller Bernese mountain dog AppenzellerEntleBucher Initial top-K for channel 193 RottweilerAppenzellerbluetick Border collie Tibetan mastiff Final top-K for channel 193 honeycombapiaryhoneycombapiaryhoneycomb Nearest pre-Attack channel by KT: 163, Kendall- -W j :0.303 Saint BernardSaint BernardBorder collieAppenzellerSaint Bernard Nearest pre-Attack channel by clip: 90, CLIP-1.00 Chihuahua Bernese mountain dogAppenzellerRottweilerAppenzeller Nearest Post-Attack channel by KT: 163, final top-K, Kendall- -W j :0.185 Chihuahua Bernese mountain dogAppenzellerRottweilerAppenzeller Nearest Post-Attack channel by clip: 163, final top-K, CLIP-0.991 Figure 20:Illustrations for the existence of whack-a-mole on the channel 193, found as one of the "hard" case (as presented in Section 4.1, Figure 5). The first two rows show the initial and final top-kimages for the targeted channel. The third and fourth rows show the initial nearest channels w.r.t. Kendall-Ď-W j and CLIP-δ-W j , respectively. The fifth and sixth rows show the nearest post-attack channel according to Kendall-Ď-W j and CLIP-δ-W j respectively. Additional Investigation of Potential Existence of Whack-a-mole.These randomly selected examples support the general findings reported in figure-6. While certain channels may have similar top images to specific post-attack channels, it is generally the case that even the most similar channels are distinct. In figure-21, the two bottom rows denote the top 5 images of the most similar channels to the pre-attack channel measured by the Kendall-Ďand CLIP-W j respectively. 25 Whack-a-mole for channel 0 of Conv 5 ladybugpomegranate black-footed ferret cucumberpill bottle Initial top-K for channel 0 ice creamladybughenstrawberrybaseball Final top-K for channel 0 honeycombmatchstickapiaryhoneycombhoneycomb Nearest pre-Attack channel by KT: 15, Kendall- -W j :0.110 strawberrybell peppercucumberbell pepperbell pepper Nearest pre-Attack channel by clip: 8, CLIP-1.00 Blenheim spaniel grocery storethimbleorange Blenheim spaniel Nearest Post-Attack channel by KT: 15, final top-K, Kendall- -W j :-0.044 Granny Smithhouse finchbell pepperGranny SmithGranny Smith Nearest Post-Attack channel by clip: 111, final top-K, CLIP-0.952 (a) Targeted channel: 0. Whack-a-mole for channel 121 of Conv 5 grocery storedaisydaisyGranny Smithorange Initial top-K for channel 121 daisydaisy yellow lady's slipper European fire salamander European fire salamander Final top-K for channel 121 honeycombmatchstickapiaryhoneycombhoneycomb Nearest pre-Attack channel by KT: 15, Kendall- -W j :0.106 croquet balllemonorangeGranny Smithlipstick Nearest pre-Attack channel by clip: 13, CLIP-1.00 Blenheim spaniel grocery storethimbleorange Blenheim spaniel Nearest Post-Attack channel by KT: 15, final top-K, Kendall- -W j :0.008 Loaferbell pepperfigpill bottleGranny Smith Nearest Post-Attack channel by clip: 94, final top-K, CLIP-0.941 (b) Targeted channel: 121. Figure 21:Illustrations for the existence of whack-a-mole on two randomly chosen channels. The first two rows show the initial and final top-kimages for the targeted channel. The third and fourth rows show the initial nearest channels w.r.t. Kendall-Ď-W j and CLIP-W j , respectively. The fifth and sixth rows show the nearest post-attack channel according to Kendall-Ď-W j and CLIP-W j , respectively. 26 B.5 Additional Illustrations for the Push-up Attack This section provides additional visual illustrations of the push-up all-channel attack on the layer conv5 of AlexNet. Visual Examples.We first provide additional visual illustrations in Figure 22 of the attack on 10 randomly chosen channels. As a reminder, this push-up attack aims to make images of the Goldfish class appear in the top-kimages of every channel on the targeted layer. From Figure 22, a first observation is the fact that out of these 10 randomly chosen channels, only two channels (channel 15 and channel 23) do not show an image with the Goldfish class. On the rest of the channels, an image with Goldfish was successfully inserted in the final top images. Furthermore, in several cases (channels 110, 125, 145, 180, 183, and 50) is the majority class of final top-5 images, demonstrating the success of this attack. It is also important to note the complete replacement of images with the Goldfish class in some channels (e.g., channel 125). 27 Channel 11 of Conv 5: Kendall- : 0.695, CLIP- : 0.159 brassmenumenubrassmenu Initial top-K goldfishmenugoldfishmenubrass Final top-K Channel 110 of Conv 5: Kendall- : 0.683, CLIP- : 0.296 cranemagpiebald eaglevulturecardoon Initial top-K goldfishcranegoldfishgoldfishcardoon Final top-K Channel 125 of Conv 5: Kendall- : 0.689, CLIP- : 0.307 shower capjigsaw puzzlebath towelshower curtainpencil box Initial top-K goldfishgoldfishgoldfishgoldfishgoldfish Final top-K Channel 145 of Conv 5: Kendall- : 0.600, CLIP- : 0.122 chain mailanemone fishsea anemonecoral reef typewriter keyboard Initial top-K goldfishgoldfishgoldfishgoldfishchain mail Final top-K Channel 15 of Conv 5: Kendall- : 0.553, CLIP- : 0.048 honeycombmatchstickapiaryhoneycombhoneycomb Initial top-K honeycombmatchstickhoneycombrubber eraserhoneycomb Final top-K Channel 180 of Conv 5: Kendall- : 0.607, CLIP- : 0.125 Airedalesea urchinbloodhoundgolden retrieverAfrican grey Initial top-K Airedalegoldfishgoldfishlorikeetgoldfish Final top-K Channel 183 of Conv 5: Kendall- : 0.658, CLIP- : 0.197 reelcar wheelwall clockanalog clockmanhole cover Initial top-K goldfishgoldfishreelcar wheelgoldfish Final top-K Channel 23 of Conv 5: Kendall- : 0.600, CLIP- : 0.126 domedomeballoondomedome Initial top-K domeballoondomedomecar wheel Final top-K Channel 232 of Conv 5: Kendall- : 0.665, CLIP- : 0.096 domemosquemosquedomedome Initial top-K goldfishdomemosquedomebeer bottle Final top-K Channel 50 of Conv 5: Kendall- : 0.709, CLIP- : 0.151 ptarmigancoucalblack grouse red-backed sandpiper ptarmigan Initial top-K ptarmigangoldfishgoldfishgoldfishblack grouse Final top-K Figure 22:Push-up all-channel attack ofConv5of AlexNet. Channel indexes were taken randomly. 28 Generalization for the Push-Up attack.After demonstrating the success of achieving target manipulability of top-kfeature visualization through the push-up attack on training images, it is also important to evaluate whether this success generalizes to unseen data. Figure 23 shows not only top-k images from the training but also from the validation set of ImageNet. We can observe that on all the 10 randomly chosen channels not only at least one image of the Goldfish class is present in the final top-5 images of the training but also at least one image of the Goldfish class is in the final top-5 images from the validation set. Moreover, we also observe a similar number of images of the Goldfish class present in top-5 images from both training and validation sets. This indicates the ability of the push-up attack to generalize on the same distribution from where training examples were drawn. 29 Channel 111 of Conv 5 "train": Kendall- : 0.724, CLIP- : 0.304 entertainment centermousehome theatertelevision entertainment center Initial Training top-K goldfishgoldfishgoldfishgoldfishmouse Final Training top-K Channel 111 of Conv 5 "val": Kendall- : 0.764, CLIP- : 0.087 laptop entertainment centerdesktop computeriPodtelevision Initial Validation top-K goldfishgoldfishlaptop entertainment centergoldfish Final Validation top-K Channel 132 of Conv 5 "train": Kendall- : 0.630, CLIP- : 0.271 window screenwindow screenhoneycombwindow screenwindow screen Initial Training top-K goldfishhoneycombwindow screengrillewindow screen Final Training top-K Channel 132 of Conv 5 "val": Kendall- : 0.660, CLIP- : 0.034 window screenhoneycombshopping basketplanetariumpuffer Initial Validation top-K window screenshopping basketgoldfishhoneycombplanetarium Final Validation top-K Channel 155 of Conv 5 "train": Kendall- : 0.717, CLIP- : 0.140 churchvending machinesteel drumchurchcassette Initial Training top-K goldfishchurchvending machinegoldfishgoldfish Final Training top-K Channel 155 of Conv 5 "val": Kendall- : 0.755, CLIP- : 0.083 monasterybookcasechurchchurchdigital watch Initial Validation top-K goldfishmonasterybookcasechurchgoldfish Final Validation top-K Channel 183 of Conv 5 "train": Kendall- : 0.658, CLIP- : 0.197 reelcar wheelwall clockanalog clockmanhole cover Initial Training top-K goldfishgoldfishreelcar wheelgoldfish Final Training top-K Channel 183 of Conv 5 "val": Kendall- : 0.757, CLIP- : 0.062 manhole coverstopwatchstopwatchtoilet seatsaltshaker Initial Validation top-K goldfishgoldfishgoldfishmanhole coverstopwatch Final Validation top-K Channel 197 of Conv 5 "train": Kendall- : 0.707, CLIP- : 0.260 sarongFrench hornrugby ballseashoresax Initial Training top-K goldfishgoldfishgoldfishsarongFrench horn Final Training top-K Channel 197 of Conv 5 "val": Kendall- : 0.783, CLIP- : 0.131 academic gownskimissilestretcheracademic gown Initial Validation top-K goldfishgoldfishgoldfishacademic gowngoldfish Final Validation top-K Figure 23:Push-up all-channel attack of Conv5 of AlexNet. For each channel, the first two rows are top-k images derived from the training set while the last two are derived from the validation set. 30 Channel 20 of Conv 5 "train": Kendall- : 0.709, CLIP- : 0.087 lipsticklipsticksyringeAngorahair spray Initial Training top-K lipstickgoldfishsyringegoldfishlipstick Final Training top-K Channel 20 of Conv 5 "val": Kendall- : 0.781, CLIP- : -0.008 juncosunglassPomeraniansea slugbrambling Initial Validation top-K goldfishjuncogoldfishsea slugPomeranian Final Validation top-K Channel 207 of Conv 5 "train": Kendall- : 0.758, CLIP- : 0.111 European gallinuleAmerican cootindigo buntingpillowEuropean gallinule Initial Training top-K goldfishgoldfishgoldfishgoldfishgoldfish Final Training top-K Channel 207 of Conv 5 "val": Kendall- : 0.805, CLIP- : 0.030 little blue heronLabrador retrieverDutch ovenAfrican greygreat grey owl Initial Validation top-K goldfishgoldfishgoldfishlittle blue heronLabrador retriever Final Validation top-K Channel 215 of Conv 5 "train": Kendall- : 0.624, CLIP- : 0.410 woolsea urchinstoleponchotennis ball Initial Training top-K goldfishgoldfishgoldfishwoolgoldfish Final Training top-K Channel 215 of Conv 5 "val": Kendall- : 0.722, CLIP- : 0.114 bonnetmittensombreropomegranatetennis ball Initial Validation top-K goldfishgoldfishbonnetmittengoldfish Final Validation top-K Channel 244 of Conv 5 "train": Kendall- : 0.802, CLIP- : 0.383 mailboxmicrowave entertainment centerchest entertainment center Initial Training top-K goldfishgoldfishgoldfishgoldfishgoldfish Final Training top-K Channel 244 of Conv 5 "val": Kendall- : 0.820, CLIP- : 0.110 entertainment centerperfumeperfume entertainment centerchiffonier Initial Validation top-K goldfishgoldfish entertainment centerperfumegoldfish Final Validation top-K Channel 248 of Conv 5 "train": Kendall- : 0.653, CLIP- : 0.095 puckmanhole covermanhole covermanhole covercoil Initial Training top-K puckgoldfishgoldfishmanhole coverbarometer Final Training top-K Channel 248 of Conv 5 "val": Kendall- : 0.662, CLIP- : 0.041 barometershieldbarometermaze typewriter keyboard Initial Validation top-K barometershieldbarometergoldfishbarometer Final Validation top-K Figure 24:Push-up all-channel attack of Conv5 of AlexNet. For each channel, the first two rows are top-k images derived from the training set while the last two are derived from the validation set. 31 B.6 Additional Illustrations for Synthetic Feature Visualization This section provides additional illustrations of the decorrelation between synthetic and natural (through top-kimages) feature visualization. Figure 25 shows the natural and synthetic feature visualization before and after the attack on 4 randomly chosen channels of conv5 of AlexNet. As stated in Section 4.1.1, from this figure, we observe a lack of change in the synthetic optimal image (even when top images have been completely replaced by images of the Goldfish class, e.g., in channel 54). We, therefore, reemphasize that attacking the natural feature visualization does not transpose to attacking the synthetic feature visualization. This indicates a decorrelation between the synthetic feature visualization and the top-k images. 32 top-k vs synthetic for channel 1 of Conv 5 shojiwindow screenwindow screenwalletSynthetic Initial feature visualization crossword puzzlecrossword puzzlechainlink fencetobacco shopSynthetic Final feature visualization for push-down walletshojigoldfishwindow screenSynthetic Final feature visualization for push-up top-k vs synthetic for channel 158 of Conv 5 barreltile roofchainlink fencehoneycombSynthetic Initial feature visualization window screenthunder snakePersian catPersian catSynthetic Final feature visualization for push-down goldfishgoldfishgoldfishgoldfishSynthetic Final feature visualization for push-up top-k vs synthetic for channel 179 of Conv 5 Indian cobrahorizontal bartable lamphorizontal barSynthetic Initial feature visualization quailanalog clockruffed grousebalance beamSynthetic Final feature visualization for push-down goldfishgoldfishgoldfishgoldfishSynthetic Final feature visualization for push-up top-k vs synthetic for channel 188 of Conv 5 banjoacoustic guitarred winepirateSynthetic Initial feature visualization harvestmanpickelhaubewater towerfrying panSynthetic Final feature visualization for push-down goldfishgoldfishbanjored wineSynthetic Final feature visualization for push-up top-k vs synthetic for channel 215 of Conv 5 woolsea urchinstoleponchoSynthetic Initial feature visualization meerkatgreat grey owlgreat grey owlstarfishSynthetic Final feature visualization for push-down goldfishgoldfishgoldfishwoolSynthetic Final feature visualization for push-up top-k vs synthetic for channel 251 of Conv 5 coral funguscoral fungusanemone fishpretzelSynthetic Initial feature visualization Dandie DinmontDandie Dinmonttoy poodleminiature poodleSynthetic Final feature visualization for push-down goldfishcoral fungusgoldfishcoral fungusSynthetic Final feature visualization for push-up top-k vs synthetic for channel 41 of Conv 5 ringneck snakeringneck snakebarometerwall clockSynthetic Initial feature visualization boa constrictorking snakethunder snakeking snakeSynthetic Final feature visualization for push-down goldfishringneck snakebarometergoldfishSynthetic Final feature visualization for push-up top-k vs synthetic for channel 54 of Conv 5 toy terrierbarometergreat grey owldigital watchSynthetic Initial feature visualization toy terrierotterhoundwaffle ironhorned viperSynthetic Final feature visualization for push-down goldfishgoldfishgoldfishgoldfishSynthetic Final feature visualization for push-up Figure 25:Synthetic Feature Visualization attack after push-down and push-up attacks on Conv5 of AlexNet. Channels indexes were taken randomly. We observe a decorrelation between natural top-activating images and synthetic optimal images. 33 B.7 Additional Fairwashing Results. This section presents the results obtained after the fairwashing attack on the last but one layer of AlexNet. Example of the Paper.We begin by showing in Figure 26, the top-30images before and after the attack from both training and testing annotated data. As a reminder, we assume that the interpreter has access to (testing) non-annotated data with a protected attribute (here the gender) and the attacker uses annotated training data to fairwash (making the top-klook fairer) feature visualization. Training annotated data (first row of Figure 26) is shown only for illustration. Initial (training) top-30 images for unit 800:balance: 0.154 Final (training) top-30 images for unit 800:balance: 0.429 Initial (testing) top-30 images for unit 800:balance: 0.250 Final (testing) top-30 images for unit 800:balance: 0.579 Figure 26:Results on channel 800 for training and testing annotated data. Note that the initial top images are shown here only for illustration as these images are used in the fairwashing loss. Only annotated testing data simulate the visualization seen by the interpreter or the regulator. More examples.Figure 27 and 28 simulate what the interpreter or regulator may see on testing annotated data before and after the fairwashing attack on 4 randomly chosen units. We can observe from this figure that when the balance (fairnessmeasure on top-30images) is relatively low the fairwashing attack makes the top images look fairer (e.g., units 943, 1412, 3051, 3135). In particular, we observe (e.g., unit 3051) that the fairwashing attack is usually very effective in cases of severe bias in top-kimages with not many people in each image. 34 Initial (testing) top-30 images for unit 835:balance: 0.667 Final (testing) top-30 images for unit 835:balance: 0.579 Initial (testing) top-30 images for unit 943:balance: 0.364 Final (testing) top-30 images for unit 943:balance: 0.500 Initial (testing) top-30 images for unit 1412:balance: 0.250 Final (testing) top-30 images for unit 1412:balance: 0.429 Initial (testing) top-30 images for unit 2354:balance: 0.304 Final (testing) top-30 images for unit 2354:balance: 0.429 Figure 27:Results on several channels for the fairwashing attack. Units were randomly chosen. Balance (fairness) is usually improved in cases of severe bias, in particular when there are not many people in images. 35 Initial (testing) top-30 images for unit 2495:balance: 1.000 Final (testing) top-30 images for unit 2495:balance: 0.500 Initial (testing) top-30 images for unit 3051:balance: 0.200 Final (testing) top-30 images for unit 3051:balance: 1.000 Initial (testing) top-30 images for unit 3135:balance: 0.154 Final (testing) top-30 images for unit 3135:balance: 0.364 Initial (testing) top-30 images for unit 3872:balance: 0.875 Final (testing) top-30 images for unit 3872:balance: 0.875 Figure 28:Results on several channels for the fairwashing attack. Units were randomly chosen. The balance (fairness) is usually improved in cases of severe bias, in particular when there are not many people in images. 36