Paper deep dive
An Interpretability Illusion for BERT
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda ViĂŠgas, Martin Wattenberg
Models: BERT-base
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/12/2026, 8:12:42 PM
Summary
The paper identifies an 'interpretability illusion' in BERT, where individual neurons or linear combinations of activations appear to encode simple, human-interpretable concepts based on top-activating sentences. However, these interpretations are often inconsistent across different datasets, revealing that the model's embedding space contains dataset-specific patterns rather than universal concept directions. The authors propose a taxonomy of concept directions (local, global, and dataset-level) and recommend validating interpretability findings across multiple datasets.
Entities (5)
Relation Signals (3)
BERT â exhibits â Interpretability Illusion
confidence 95% ¡ We describe an âinterpretability illusionâ that arises when analyzing the BERT model.
Interpretability Illusion â tracedto â Dataset idiosyncrasy
confidence 90% ¡ We propose that the illusion can be traced to three sources: 1. Dataset idiosyncrasy
BERT â trainedon â Quora Question Pairs
confidence 90% ¡ We base our experiments on four different text corpora... Quora Question Pairs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We describe an "interpretability illusion" that arises when analyzing the BERT model. Activations of individual neurons in the network may spuriously appear to encode a single, simple concept, when in fact they are encoding something far more complex. The same effect holds for linear combinations of activations. We trace the source of this illusion to geometric properties of BERT's embedding space as well as the fact that common text corpora represent only narrow slices of possible English sentences. We provide a taxonomy of model-learned concepts and discuss methodological implications for interpretability research, especially the importance of testing hypotheses on multiple data sets.
Tags
Links
- Source: https://arxiv.org/abs/2104.07143
- Canonical: https://arxiv.org/abs/2104.07143
Trouble viewing inline? Open PDF directly â
Full Text
43,775 characters extracted from source content.
Expand or collapse full text
An Interpretability Illusion for BERT Tolga Bolukbasi * 1 Adam Pearce * 1 Ann Yuan * 1 Andy Coenen 1 Emily Reif 1 Fernanda Vi Ě egas 1 Martin Wattenberg 1 Abstract We describe an âinterpretability illusionâ that arises when analyzing the BERT model. Acti- vations of individual neurons in the network may spuriously appear to encode a single, simple con- cept, when in fact they are encoding something far more complex. The same effect holds for linear combinations of activations. We trace the source of this illusion to geometric properties of BERTâs embedding space as well as the fact that common text corpora represent only narrow slices of pos- sible English sentences. We provide a taxonomy of model-learned concepts and discuss method- ological implications for interpretability research, especially the importance of testing hypotheses on multiple data sets. 1. Introduction An outstanding problem in the field of neural networks is understanding how they represent meaning. One simple hypothesis is that the activation level of an individual unit in a network encodes the presence or absence of a meaningful concept. A generalization of this idea suggests that concepts are encoded by linear combinations of neural activations. This point of view has proved fruitful in analyzing image networks (Bau et al., 2017; Olah et al., 2020; Kim et al., 2018). Could the same type of analysis uncover the key concepts that are learned by neural language networks? This paper describes a surprising phenomenon, a kind of âin- terpretability illusion,â that arose from such an experiment on BERT, a model for general language representations (Devlin et al., 2018). The basic set-up for our experiment (described in detail in Section 3.2) was designed to probe whether individual neurons in BERT might have human- interpretable meaning. We gave BERT a large dataset of sentences as inputs, chose a target neuron from the final layer, and examined the sentences that maximally activated * Equal contribution 1 Google Research, Cambridge, MA, USA. Correspondence to: Tolga Bolukbasi<tolgab@google.com>, Adam Pearce<adampearce@google.com>, Ann Yuan<an- nyuan@google.com>. it. Our intuition was that if these maximally-activating sen- tences shared a common pattern, it would be an indication that the target neuron had learned to detect this pattern. Indeed, many of the neurons we probed did show strong, consistent patterns of activation. For example, of the 164,246 sentences in the Quora Question Pairs dataset (Iyer et al., 2017), here are three of the sentences which activate neuron 221 in layer 12 most strongly: â˘âWhat is the meaning behind the song âAngelâ by Eric Clapton?â ⢠âWhatâs the meaning of Johnny Cashâs song âKing of the Hillâ?â â˘âWhat is the meaning behind the Tears for Fears song âMad Worldâ, such as the lyric, âAll around me are familiar facesâ?â These strongly suggest that neuron 221 encodes a concept related to song titles, or perhaps the very specific syntactic structure of these sentences. Based on this evidence alone, it might be tempting to conclude that one can easily interpret the meanings of individual neurons in the final layer of BERT. The plot thickened, however, when we tried the same set of experiments with the same model (no fine-tuning or other modifications) but a different dataset. Using as input a question answering dataset drawn from Wikipedia (58,645 sentences), the top activating sentences for neuron 221 had nothing to do with the meanings of song titles. Instead, they included: ⢠On 16 June 2006, it was announced that Everton had entered into talks with Knowsley Council and Tesco over the possibility of building a new 55,000 seat sta- dium, ex-pandable to over 60,000, in Kirkby. â˘On 15 September 1940, known as the Battle of Britain Day, an RAF pilot, Ray Holmes of No. 504 Squadron RAF rammed a German bomber he believed was going to bomb the Palace. ⢠On 20 August 2010, Queenâs manager Jim Beach put out a Newsletter stating that the band had signed a new contract with Universal Music. Here we are forced to reconsider our initial interpretation. Based on this dataset, it would seem that neuron 221 encodes arXiv:2104.07143v1 [cs.CL] 14 Apr 2021 An Interpretability Illusion for BERT historical events, or perhaps sentences beginning with a date. In search of tie-breaking evidence, we repeated our analysis on a sample (198,085 sentences) from the Toronto Book- Corpus dataset. The top activating sentences there included: â˘Lara pulled out the document Reed had supplied from Greshamâs briefcase. â˘I take Kellanâs business card from my pocket and stretch it over to Realm. â˘Pilcher took a walkie-talkie out of his coat and spoke into the receiver. The picture is now more complicated. The BookCorpus dataset suggests yet a third interpretation for neuron 221. Note that thereâs nothing special about neuron 221: many other neurons show similar behavior. What we have seen is that for each data set, looking at maximally-activating sen- tences produces consistent, interpretable patterns for many neurons. These patterns, however, arenotconsistent across datasets. We consider this to be aninterpretability illusion. The fact that a seemingly consistent pattern can turn out to be a mirage has clear implications for interpretability research. In this paper, we provide evidence that this illusion is a general, reproducible phenomenon of the BERT model. We also consider several possible explanations for the illusion, suggesting that it is useful to separate the notion ofdataset- levelconcept directions from global concept directions. To summarize our main contributions: 1.We identify an interpretability illusion that arises when analyzing a language modelâs activation space. 2.We provide the recommendation that interpretabil- ity researchers conduct their experiments on multiple datasets. 3.We investigate possible causes of the illusion, using a taxonomy of the geometric properties of model-learned concepts: local, global, and dataset-level concept di- rections. 2. Related work There is a large body of prior work exploring the embedding spaces of language models. These embedding spaces ex- hibit both local and global structure. There is local structure in that the nearest neighbors of a datapointâs embedding are similar to it (Bengio et al., 2003). Global directions in embedding spaces have also been found to encode specific concepts (Mikolov et al., 2013; Bolukbasi & Chang, 2016; Raghu et al., 2017; Olah et al., 2020; Vig et al., 2020; Li et al., 2015). In the case of language models, these directions combine to form representations that enable sophisticated language processing (Manning et al. (2020)). Indeed, fol- lowing the probing method of Tenney et al. (2019), Durrani et al. (2020) finds that different elements of linguistic under- standing can be localized to individual or small groups of neurons. There has been substantial work (both for images and text) on determining what a specific neuron encodes by consider- ing which inputs maximally activate it. There are two main ways of doing this: first, by generating such inputs, which can provide clues regarding the encoded pattern (Nguyen et al., 2016; Poerner et al., 2018; B Ě auerle & Wexler), and second, by looking for patterns among real samples that maximally activate a neuron, where the samples are drawn from some preexisting dataset (Na et al. (2019)). Our paper focuses on the latter approach. To review a few more examples, Zhou et al. (2015) and Bau et al. (2017) use this technique to find neurons that respond to particular objects in natural scenes. Dalvi et al. (2019) look at lan- guage models, comparing the concepts for a neuron found by this technique to those found by probing. Szegedy et al. (2014) find that convolutional networks contain neurons that activate in response to semantically related inputs. Zeiler & Fergus (2014) obtain a similar result by visualizing the top activating images patches for a given feature. Olah et al. (2017) compare the maximally activating images from a dataset with images generated to maximize that neuron. There is also a body of work using concept directions for measuring and mitigating unwanted biases in models (Kaneko & Bollegala, 2019; Manzini et al., 2019; Boluk- basi & Chang, 2016). Our experiments suggest that using these techniques for sentence models without validating the directions on multiple datasets could have unintended effects. 3. Establishing the illusion Motivated by this body of work, we set out to find mean- ingful neurons and concept directions in BERTâs activation space. Our pilot experiments with a single test dataset sug- gested clear, consistent meanings for many neurons. How- ever, as described in the introduction, many of those inter- pretations disappeared when we tried to confirm them with a different data set. 3.1. Datasets To test the extent of the illusion, we base our experiments on four different text corpora. 1. Quora Question Pairs (QQP): QQP contains ques- tions from the question-answering website Quora, with 164,246 datapoints. (Iyer et al., 2017). 2.Question-answering Natural Language Inference (QNLI): QNLI contains passages from Wikipedia, with 58,645 datapoints. (Wang et al., 2019). An Interpretability Illusion for BERT 3.Wikipedia (Wiki): Wiki contains a random subset of English Wikipedia as prepared in (Devlin et al., 2018), 203,736 datapoints. 4.Toronto BookCorpus (Books): Books contains sen- tences from online novels. We sampled 198,085 sen- tences from the original data set. (Zhu et al., 2015). 3.2. Experiments We began by creating embeddings for the 624,712 sentences in our four datasets. To do this, we used the BERT-base uncased model from the HuggingFace Transformers library with no fine tuning or dataset specific modifications. We used the final layer hidden state of each sentenceâs[CLS] token as its embedding. Several methods for extracting aggregate sequence represen- tations from BERT can be found in the literature (Reimers & Gurevych, 2019). However our method remains the default for HuggingFace pipelines, and is used in the original BERT paper. We also built an exploratory visualization (Figure 1) demonstrating that these embeddings give rise to highly co- herent clusters across scales, giving us additional confidence in their representational validity. We then randomly picked a set of neurons and looked at their top activating sentences from each dataset. For convenience, we identify a neuron with a basis vector in BERTâs 768- dimensional embedding space; that is, a one-hot vector x(d)âR k where: x(d) l = 1,ifl=d 0,otherwise (1) For each neuron we find its top activating sentences by sorting sentence embeddings according to their activation level, meaning the dot product with this vector. We also find top activating sentences for a set of random directions, rather than basis vectors. In this case we sort sentence embeddings according to their inner product with the random direction. Specifically, for a datasetSand a vectorv: Top activating sentence forv= arg max xâS ăx, vă(2) In this paper, we will refer to the dot product between a sen- tence embedding and a direction as theirprojection score. Next we built an annotation interface that optionally shows: (1) the top ten activating sentences for a neuron, (2) the top ten activating sentences for a random direction, or (3) a random set of ten sentences. We annotated whether each set of ten sentences contained a pattern, and if so, which sentences demonstrated the pattern. During annotation we knew which dataset the sentences were drawn from, butnot Figure 1.Visualization of sentence embeddings from the QQP, QNLI, Wiki, and Books datasets using UMAP showing that the four datasets form distinct clusters. Our methodology for extract- ing these embeddings is described in Section 3.2. whether the sentences were top activating (conditions (1), (2)) or randomly drawn (condition (3)). In total we anno- tated 25 neurons (randomly selected), 33 random directions, and 29 random sets of sentences (Table 1). To define our notion of pattern: a pattern is simply a property shared by a set of sentences. The property may be structural, i.e. the sentences are all the same length. It may also be lexical, i.e. the sentences all contain some variant of the phrase âcoat of armsâ. We use these patterns as proxies for learned concepts by the model. 4. The Illusion Each set of sentences was annotated by two annotators. Ta- ble 1 shows how often both annotators found at least one pattern.Conflictingindicates that one annotator found a pattern and the other did not. To establish a baseline, we also looked for patterns among randomly drawn sets of ten sentences. We found 14% of random sets of sentences to contain patterns, suggesting that our datasets contain intrin- sic topic biases. However, these biases are not sufficient to explain our results. We found that more than 80% of top activating sentences contained patterns (Table 1). The pat- terns found among top activating sentences were also much stronger in that they contained more positive examples than those found in random sets of sentences (Figure 2). Finally, annotators were more in agreement about whether top acti- vating sentences were meaningful: only 8% of annotators were split over whether a neuron was meaningful, and 18% were split over whether a random direction was meaningful, An Interpretability Illusion for BERT QQPQNLIWikiBooksAll Contains patterns?YesNoYesNoYesNoYesNoYesNoConflicting Neurons3 (60%)1 (20%)8 (100%) 06 (100%) 03 (50%) 2 (33%)20 (80%) 3 (12%) 2 (8%) Random direction10 (100%) 010 (83%) 05 (100%) 02 (33%) 027 (82%) 06 (18%) Random sentences05 (100%)2 (22%) 3 (33%)1 (11%) 3 (33%)1 (17%) 3 (50%)4 (14%) 14 (48%) 11 (38%) Table 1.Annotation results for each dataset. For individual datasets, the remaining percentage is for conflicting annotations. QQPQNLIWikiBooks Nested quotes, Colors, Mathe- matics, Military conflict, Pop- ulation statistics, Relationship advice, School exam questions, Questions of comparison, Pro- gramming Biology, Geography, Technol- ogy, Numbers and dates, Mili- tary conflict, Population statis- tics, War history, Windows 8, Etymology Direct statement of fact, Mu- sic, Sporting, Age distribution, Television shows, Olympic facts, Legalese, Measurements, School districts Interpersonal relationships, Na- ture, Quoted speech, Spanish, Sentence fragments, Medieval Europe, Very long sentences, Flirtation Table 2.Sample annotated patterns for each dataset. compared to 38% in the case of random sets of sentences. Examples of our annotations are found in Table 2 (a full list can be found in the Appendix - Table 7). Many of the patterns we found are quite general, for example that a set of sentences contain quoted speech, or that they concern nature. Nevertheless, the annotations of top activating sentences often changed dramatically depending on which dataset the sentences were drawn from. In fact, for each neuron we measured on average 2.5distinctpatterns across QQP, QNLI, Wiki, and Books (Figure 2). Our results suggest that the illusion of meaningfulness we observed in neuron 221 (Section 1) is a general, reproducible phenomenon of BERT. They also show that the illusion is equally pervasive among neurons and random directions. Based on an informal investigation of layers 2 and 7 we believe the illusion occurs for earlier layers as well, although a full analysis is beyond the scope of this paper. 5. Explaining the illusion What gives rise to this illusion? How could the same direc- tion seemingly encode completely different concepts? We propose that the illusion can be traced to three sources: 1. Dataset idiosyncrasy 2.Local semantic coherence in BERTâs embedding space 3. Annotator error Next, we discuss each source in detail. 5.1. Datasets are idiosyncratic First, we consider the hypothesis: Hypothesis 1.QQP, QNLI, Wiki, and Books occupy distinct regions of BERTâs embedding space. This would contribute to the illusion because if the four datasets occupy non-overlapping slices of BERTâs embed- ding space, then in any direction the top activating sentences from each dataset will come from distinct regions of the em- bedding space (Figure 3). Experiments To test our hypothesis, we performed two experiments: ConditionMeanStdev Neurons6.802.37 Random Directions 6.892.13 Random Sentences5.051.96 Figure 2.Annotation statistics. (Top) The number of distinct pat- terns found for each annotated neuron across all datasets. We manually looked over the annotations and counted the number of unique patterns per neuron, grouping semantically equivalent anno- tations together. (Bottom) The number of sentences belonging to a pattern for the different experimental conditions for the sentence groups that are found to be meaningful. We required each pattern to have at least three positive examples. At most a pattern could have ten positive examples, because we only showed the top ten activating sentences for any given direction. An Interpretability Illusion for BERT Figure 3.Schematic illustration of how top activating sentences from geometrically distinct datasets could be semantically unre- lated. The arrow represents a direction in embedding space. Red dots represent sentences from dataset A, and blue dots represent sentences from dataset B. The dots outlined in black represent the the top activating sentences. 1. We built an exploratory visualization of the QQP, QNLI, Wiki, and Books dataset embeddings using the UMAP dimensionality reduction algorithm. 2. We trained a linear SVM classifier to distinguish be- tween the four datasets based on their sentence embed- dings. Results Our exploratory visualization shows that sentences cluster neatly by dataset (Figure 1). Our linear SVM is also able to distinguish between the datasets with high accuracy (Figure 4). These results suggest that QQP, QNLI, Wiki and Books represent relatively idiosyncratic slices of language, thus the top activating sentences for a given neuron from one dataset do not necessarily resemble those from another dataset de- spite having similar activation valuesâlooking at all the pairs of datasets across all neurons, 38% of the ranges of the top 10 activations overlap. Our results align with previ- ous work demonstrating that BERT representations can be used to disambiguate datasets (Aharoni & Goldberg, 2020). However, while this explains why we would find different patterns in different datasets, why should there be patterns at all? We address this question in the next section. 5.2. Local semantic coherence We hypothesize that another source of the illusion is local semantic coherence in BERTâs embedding space geometry. Before testing this, we observe: Observation.Top activating sentences manifest patterns from both local semantic coherence and global directions. This observation may seem obvious. However, suppose that a neuron encodes a certain concept: one might assume that its top activating sentences would clearly point to the en- Figure 4.Confusion matrix for a linear classifier trained to separate QQP, QNLI, Wiki, and Books sentence embeddings. Most datasets are easily separable in the embedding space. Figure 5.Global versus local concepts.Left: Global concept illustration.Blue circles represent sentences containing a global concept. As one moves in the concept direction indicated by the arrow, the density of sentences containing the global concept increases.Right: Local concepts illustration.The red, yellow, and blue shapes represent sentences containing three different concepts. Sentences containing the same concept cluster together, such that there is local concept coherence throughout the space. However, there is no global direction along which any particular concept becomes more dominant. coded concept. Instead, we observe that the sentences may manifest patterns that do not match the encoded concept. Before further analysis, we describe three ways in which BERT may learn to represent concepts (Figure 5): Global concepts:Global concepts become increasingly prevalent as one moves through the embedding space along a linear trajectory. For example, ifmathis a global concept, then there is a direction in the embedding space such that, starting from any point and moving along that direction, one will tend to find more and more sentences relating to math. Or if the concept ispositivity, words likehappy, sunny, bliss, awesomemight increase in frequency. In our analysis we use the presence of certain tokens as a proxy for concepts. By extension, if a direction in embedding space is corre- lated with the occurrence of certain tokens, it suggests the existence of an underlying concept direction. Dataset-level concepts:Like global concepts, adataset- levelconcept is associated with a direction in the embedding An Interpretability Illusion for BERT space. However, unlike global concepts, this direction is only meaningful within the region of the embedding space where samples from the dataset tend to be found. Thus, dataset-level concepts do not generalize to arbitrary inputs. Local concepts:On the other hand,localconcepts emerge only as clusters in the embedding space, and lack a direction. For example, suppose thatmathis a local concept. Then if one looks in the neighborhood of the sentencee=mc 2 , one may find many other sentences containing numbers and symbols because the model groups math-related sentences together. However, there will not be any particular direction along which such sentences become increasingly prevalent. Figure 6.(Top left)The frequency of the word âquickâ monotoni- cally decreases as we look at sentences that increasingly activate neuron 275 in QQP. This suggests that neuron 275 encodes a global concept (Section 5.2) that relates to âquickâ.(Top right)Most neurons, such as neuron 266, do not correlate with the frequency of âquickâ. These are illustrative examples, Table 3 shows over a quarter of neuron/token pair are monotonic. We find evidence for both global and dataset-level concept directions in BERTâs embedding space. Figure 6 illus- trates the change in certain token frequencies as one moves along various neuron directions. Some tokens, such as âweirdâ, monotonically change in frequency as one moves along neuron 266, while other token frequencies are un- changed. Table 3 illustrates that this is not an isolated phe- nomenon, rather there are many tokens that monotonically increase or decrease along different neurons. Table 3 fur- ther shows that these patterns can be unique to a particular dataset, suggesting the existence of dataset-level concept directions. We note that the baseline probability of a token (e.g. âweirdâness) being monotonic with respect to a neuron DatasetsMonotonicIncreasingDecreasing Books27.0%13.4%13.6% QNLI22.7%11.3%11.4% Wiki29.6%14.7%15.0% QQP27.8%13.9%13.9% BooksQNLI7.4%3.0%2.9% BooksWiki9.9%4.3%4.3% QNLIWiki10.9%5.3%5.4% BooksQQP9.2%3.9%3.9% QNLIQQP7.7%3.2%3.2% WikiQQP9.8%4.0%4.1% BooksQNLIWiki4.2%1.8%1.8% BooksQNLIQQP3.0%1.2%1.2% BooksWikiQQP3.9%1.7%1.6% QNLIWikiQQP4.2%1.8%1.9% BooksQNLIWikiQQP1.9%0.8%0.8% Table 3.The table shows how often the token counts across quin- tiles of activations for each neuron/token pair are monotonically increasing or decreasing in each combination of datasets. Combina- tions with more datasets have darker backgrounds. There are some linear directions in the embedding space that are correlated with the same tokens across all datasets (for the 915 tokens appearing at least 100 times in each of the four datasets). Tokens whose frequencies change monotonically: â(125),can(120),is(99),are(98),was(97),that(91),if (88),were(86),to(85),would(84),a(82),it(80),they (78),not(77),god(76),for(75),which(73),more(73), she(70),of(68) Table 4.This table lists the tokens from all four datasets that are most often monotonically changing across the embedding space. For each token, we indicate in parenthesis the number of neurons for which the token changes monotonically in frequency as one moves along the neuron axis. is much lower than the measured rates 1 . Table 4 shows the most monotonic tokens across datasets, hinting that BERT may have learned to encode certain pervasive concepts as global directions, e.g. pronouns and common verbs. Despite the fact that global and dataset-level concept direc- tions exist, we posit that they will be difficult to identify from top activating sentences. First, sentences typically engage with multiple concepts. Thus even if a neuron en- codes a concept, its top activating sentences may have other concepts in common as well. Second, the directions we annotated are likely to themselves align with multiple con- cepts, and the top activating sentences for each concept may look unrelated when listed together. As an example, suppose 1 We use quintiles to measure monotonicity. The probability of a random set of quintiles being monotonic is 2/5! (1.7%). An Interpretability Illusion for BERT we have a dataset in which every sentence represents exactly one concept and we pick a direction that is a linear combi- nation of ten global concept directions. When we look at the ten sentences that most highly activate this direction, we may not see any patterns at all. In the extreme case, we may only see one sentence representing each concept, making it impossible to identify any patterns across sentences. Formally we define a concept distribution as:C= [c 1 , c 2 , ..., c N ]where â N i=1 c i = 1.0 . The concept purity of a sentence or a direction is defined by the skew of its concept distribution vector. A sentence that aligns with exactly one concept would have a one-hot concept distribution vector. 5.2.1. MEASURINGLOCALCONCEPTS Having discussed the difficulty of identifying concept direc- tions from top activating sentences, in this section we aim to show that: Hypothesis 2.When annotating top activating sentences, people identify concepts emerging from local semantic co- herence. This would explain the illusion because local concepts are not necessarily associated with a direction in the activation space, but rather with a semantic cluster. We propose the following analysis to test our hypothesis. The main intuition is that if a concept we annotate emerges from local semantic coherence, then for sentences manifest- ing the concept, their neighborhood in the original embed- ding space should look very similar to their neighborhood when projected onto the concept direction 2 . Rather than measuring exact equality of sentences in the two neighbor- hoods (which is very sensitive), we compare their pairwise distance distributions. More formally, letN k (s)be theknearest neighbors of a sentences: N k (s) = arg max S ⲠâS,|S Ⲡ|=k â ËsâS Ⲡe s ¡e Ës wheree s denotes the embedding of a sentence in the original embedding space. Consistent with our annotation protocol we usek= 10. LetS p,k define the top activatingksentences for a direction p. For each sentencesinS p,k , we measure the dot prod- uct betweensand its nearest neighborss ⲠâN k (s). We 2 Our analysis measures local coherence among top activating sentences. If both neighborhoods match, our analysis cannot deter- mine whether the local coherence comes from a local or a global concept. However we know that for the directions whose meaning changes across datasets, the source of local coherence cannot be a global concept. On the other hand, a neighborhood mismatch would mean something besides local coherence is responsible for any pattern. call this set of distancesD p,nearest . We then measure the dot product between all pairs of top activating sentences (S p,k ), calling this setD p,top . Finally, we measure the dot product between each top activating sentence andkrandom sentences in the dataset to establish a baseline, calling this setD p,random . The histograms in Figure 7 shows howD p,nearest ,D p,top , andD p,random compare. On top is a neuron our annotators deemed meaningful. The pairwise distances between the top activating sentences (D p,top ) overlap with those between the top activating sentences and their nearest neighbors in the original embedding space (D p,nearest ). By contrast, on the bottom is a meaningless neuron, where the pair- wise distances between the top activating sentences (D p,top ) match those between the top activating and random sen- tences (D p,random ). We extend this analysis across all the directions we anno- tated and propose a metric based on the Jaccard similarity: L(h 1 , h 2 ) = â i min(h 1 (i), h 2 (i)) â i max(h 1 (i), h 2 (i)) L measures the intersection over union between any two histogramsh 1 andh 2 . We callLthe locality score for a direction when applied to the histograms ofD p,nearest andD p,top . If Hypothesis 2 is true,Lshould be high for meaningful directions and low for meaningless directions. Table 5 shows that the locality scores for meaningless neu- rons are indeed significantly lower than for meaningful neurons. This means that the patterns found among top activating sentences arise primarily fromlocalgeometry. Locality (Mean)AllQQPQNLIWikiBooks Meaningful0.0260.0120.0140.0370.042 Meaningless0.0100.0030.0080.0160.012 p-value0.00040.0090.0900.0640.014 Table 5.Per sentence locality scores (based on the histogram inter- section over union measure). 5.2.2. ANALYZING OUTLIER SENTENCES We present our preceding analysis of local semantic coher- ence as general, but our experiment targets the edges of the activation space. Perhaps the properties of local semantic co- herence are unique to a small set of outlier sentences which maximally activate across the directions we annotated. We consider this possibility for the QQP dataset. Below are the three QQP sentences that maximally activate the most neurons: â˘âwhat does snoop dogg mean by âlolosâ in the line âhit the corners in them lolos girl âin dr. dreâs 2001 hit âstill d.r.eâ?â(56 top activations) An Interpretability Illusion for BERT Figure 7.The distribution of distances in the original embedding space for QQPâs top activating sentence for neurons 65, 44. Or- ange lines show the distances between the neuronsâ top activating sentence and its nearest neighbors in the original embedding space. Blue lines show the distances between the top activating sentence and other top activating sentences. Green lines show the distances between the top activating sentence and randomly drawn sentences. (Top)A neuron deemed meaningful by annotators. Observe that the blue line significantly overlaps with the orange line, suggesting that the meaning comes from the top activating sentenceâs local neighborhood.(Bottom)A neuron deemed meaningless by anno- tators. Observe the high degree of overlap between the distribution of distances between the top activating sentences and the distances between random sentences. â˘âwhat are some songs like âthe dying of the light - noel gallagharâ?â(40 top activations) â˘âwhat inspired the tv show âskinsâ?â(32 top acti- vations) Taking the pairwise Euclidean distances between every QQP sentence, these three sentences have the highest mean dis- tances. This pattern holds across the entire distribution: the 20 most distant QQP sentences all have at least 11 top ac- tivations and the most distant 1% of sentences account for 48% of all top activations. These numbers suggest a lack of diversity among QQPâs out- lier sentences. Indeed, looking at our experimental results, it appears that we annotated the same sentences across mul- tiple directions: several neurons are maximally activated by nested quotation marks, several others by movie plots and several others by math equations. However most sentences are only top activating once or twice; across the 7,680 top ten activating sentences, there are 4,551 unique sentences. And the illusionpersistswhen the most distant 1% and 10% of sentences are omitted. 5.3. Annotator error A final source of the illusion is our ability to see patterns where they may not exist. As noted in Section 4, annotators often disagreed over whether a set of sentences constituted a pattern. Additionally, annotators differ significantly in their tendency to find patterns (Table 6). This suggests certain patterns may be less an objective property of the sentences than products of an individual annotatorâs (overactive) imag- ination. AnnotatorDirections annotatedPatterns found 03027 (0.90) 13327 (0.82) 22423 (0.92) 34023 (0.56) 45230 (0.58) 52540 (1.6) Table 6.The total number of annotations and the number of pat- terns found per annotator. The number of patterns can exceed the number of annotations because a single direction may exhibit multiple patterns. 6. Conclusion In this work we investigated how BERT represents mean- ing by looking for patterns among top activating sentences along directions in the modelâs activation space. In doing so, we uncovered a counterintuitive âillusionâ: seemingly consistent patterns that turn out to be contingent on the test dataset. We then identified several possible sources of this behavior: the idiosyncratic nature of widely used natural language datasets, the geometry of BERTâs activation space, and annotator bias. Furthermore, despite these âillusoryâ findings, we also provide evidence that there do exist mean- ingful global directions in BERTâs activation space. Several avenues for future work suggest themselves. It would be interesting to replicate our analysis in other set- tings, such as with other sentence models, with token-level embeddings, and with earlier layers of BERT. It would also be natural to look for similar illusions in other types of in- put: images, graphs, and so forth. In any of these contexts, there is clearly much work to be done in characterizing the geometry of concept representation. Our work also draws needed attention to the role of datasets in interpretability research. It is a classic problem in ma- chine learning that a model trained on one dataset may not perform well on another. The results we have described illustrate something slightly different: how aninterpreta- An Interpretability Illusion for BERT tionthat seems valid for one dataset may not generalize to other contexts. Examining how interpretations function across different data distributions may be a helpful step in establishing their validity. AcknowledgementsWe would like to thank Ian Tenney, Jasmijn Bastings, Katherine Lee and Been Kim for their re- views and helpful discussions during the process of writing this paper. References Aharoni, R. and Goldberg, Y. Unsupervised domain clus- ters in pretrained language models. InAssociation for Computational Linguistics, p. 7747â7763, 2020. Bau, D., Zhou, B., Khosla, A., Oliva, A., and Torralba, A. Network dissection: Quantifying interpretability of deep visual representations. InProceedings of the IEEE conference on computer vision and pattern recognition, p. 6541â6549, 2017. Bengio, Y., Ducharme, R., Vincent, P., and Janvin, C. A neu- ral probabilistic language model.The journal of machine learning research, 3:1137â1155, 2003. Bolukbasi, T. and Chang, K.-W. Man is to computer pro- grammer as woman is to homemaker? debiasing word embeddings. InProceedings of the 30th International Conference on Neural Information Processing Systems, p. 4356â4364, 2016. B Ě auerle, A. and Wexler, J.What does bert dream of?URLhttps://pair-code.github. io/interpretability/text-dream/ explainable/. Dalvi, F., Durrani, N., and Sajjad, H. What is one grain of sand in the desert? analyzing individual neurons in deep nlp models. InProceedings of the 33rd AAAI Conference on Artificial Intelligence, p. 6309â6317, 2019. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for lan- guage understanding.arXiv preprint arXiv:1810.04805, 2018. Durrani, N., Dalvi, F., Sajjad, H., and Belinkov, Y. Ana- lyzing individual neurons in pre-trained language mod- els. 2020. URLhttps://arxiv.org/pdf/2010. 02695.pdf. Iyer,S.,Dandekar,N.,andCsernai,K. Firstquoradatasetrelease:Questionpairs, 2017.URLhttps://data.quora.com/ First-Quora-Dataset-Release-Question-Pairs . Kaneko, M. and Bollegala, D.Gender-preserving de- biasing for pre-trained word embeddings.CoRR, abs/1906.00742, 2019. URLhttp://arxiv.org/ abs/1906.00742. Kim, B., Wattenberg, M., Gilmer, J., Cai, C., Wexler, J., Viegas, F., et al. Interpretability beyond feature attribu- tion: Quantitative testing with concept activation vectors (tcav). InInternational conference on machine learning, p. 2668â2677, 2018. Li, J., Chen, X., Hovy, E. H., and Jurafsky, D. Visualiz- ing and understanding neural models in NLP.CoRR, abs/1506.01066, 2015. URLhttp://arxiv.org/ abs/1506.01066. Manning, C., Hewitt, J., Clark, K., Khandelwal, U., and Levy, O. Emergent linguistic structure in artificial neural networks trained by self-supervision.Proceedings of the National Academy of Sciences of the United States of America, 2020. Manzini, T., Lim, Y. C., Tsvetkov, Y., and Black, A. W. Black is to criminal as caucasian is to police: Detecting and removing multiclass bias in word embeddings.CoRR, abs/1904.04047, 2019. URLhttp://arxiv.org/ abs/1904.04047. Mikolov, T., Chen, K., Corrado, G., and Dean, J. Efficient estimation of word representations in vector space, 2013. Na, S., Choe, Y. J., Lee, D.-H., and Kim, G. Discovery of natural language concepts in individual units of cnns. 2019. In the Proceedings of ICLR. Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., and Clune, J. Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. InAd- vances in Neural Information Processing Systems, p. 3387â3395, 2016. Olah, C., Mordvintsev, A., and Schubert, L. Feature vi- sualization.Distill, 2017. doi: 10.23915/distill.00007. https://distill.pub/2017/feature-visualization. Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S. Zoom in: An introduction to cir- cuits.Distill, 2020. doi: 10.23915/distill.00024.001. https://distill.pub/2020/circuits/zoom-in. Poerner, N., Roth, B., and Sch Ě utze, H.Interpretable textual neuron representations for NLP. InProceed- ings of the 2018 EMNLP Workshop BlackboxNLP: An- alyzing and Interpreting Neural Networks for NLP, p. 325â327, Brussels, Belgium, November 2018. Associ- ation for Computational Linguistics. doi: 10.18653/ v1/W18-5437. URLhttps://w.aclweb.org/ anthology/W18-5437. An Interpretability Illusion for BERT Raghu, M., Gilmer, J., Yosinski, J., and Sohl-Dickstein, J. Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability, 2017. Reimers, N. and Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. InProceed- ings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, p. 3982â3992, 2019. Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Er- han, D., Goodfellow, I., and Fergus, R.Intriguing properties of neural networks. InInternational Confer- ence on Learning Representations, 2014. URLhttp: //arxiv.org/abs/1312.6199. Tenney, I., Xia, P., Chen, B., Wang, A., Poliak, A., McCoy, R. T., Kim, N., Durme, B. V., Bowman, S., Das, D., and Pavlick, E. What do you learn from context? probing for sentence structure in contextualized word representations. InInternational Conference on Learning Representations, 2019. Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Sakenis, S., Huang, J., Singer, Y., and Shieber, S. Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020. Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R. GLUE: A multi-task benchmark and anal- ysis platform for natural language understanding. 2019. In the Proceedings of ICLR. Zeiler, M. D. and Fergus, R. Visualizing and understand- ing convolutional networks. InEuropean conference on computer vision, p. 818â833. Springer, 2014. Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., and Tor- ralba, A. Object detectors emerge in deep scene cnn. In International Conference on Learning Representations, 2015. Zhu, Y., Kiros, R., Zemel, R., Salakhutdinov, R., Urta- sun, R., Torralba, A., and Fidler, S. Aligning books and movies: Towards story-like visual explanations by watch- ing movies and reading books, 2015. An Interpretability Illusion for BERT QQPQNLIWikiBooks School exam questions, High school math and physics, Com- parison questions, Indian jobs, Jobs for students, Large body physics, Movies, Music, Math, Physics, Indian cities, Race, Science, Technology and con- sulting, Containing multiple quoted phrases, Hygiene, Tech support, Drugs, India, Pro- gramming, Mathemtical and chemical formulas, Cosmol- ogy, Containing quoted sub- jects, Internet companies, Cor- porate jargon, Quoted songs, Indian paperwork Militaryconflict,Ethnic groups (language,culture, history), Numbers and dates, Standards,Namedplaces (state, university, etc.), DNA, Text book material, Population statistics / census, War history, Military moves, Geographic locations,Africa(mostly Somalia), Engineering, Gov- ernmental power, Geography, Language / linguistics, Group theory,Politics,Military, Governmentsandleaders and revolutions, Displaced people and slavery, Animals, Science,Renovationsand infrastructure changes, US southern cities and especially in NC, Countries, Bible quotes and locations, Resettling and boundaries, War, Histories of ethnic groups, Math, Census data,Militarymaneuvers, Somalia / Eritrea, Law and legislative bodies, Historical events, Etymology, Windows 8,Mathematicalgroups, Years, Political conflict and revolution, Political history, Animals, Family, Geography, Ecology, Census, Chemistry, Weather history, Municipal facts, Indian, Demographics, Weather, Raleigh Settlement history, Direct state- ments of fact, Named locations, Voting stats of cities, Birth family occupations of individ- uals, Locations, British his- tory, Age distribution, Census data, TV shows, Olympians, Daughters, War, Law / con- tracts, International sports, Me- dia, School districts, National borders, Music, Properties of villages, Anatomical descrip- tions, Indiginous communities, Chemistry, Printing presses, Math, Statistics, Television shows, Legalese, French / Ger- man municipalities, Measure- ments, Rules, Minerals, Soccer, Medicine, School rules, Bands Spanish, Description then talk- ing, Snippets describing peo- ple, Fantasy travel, Quotes, Body parts, Plans, Person do- ing a small physical action (laughing, rubbing hands, etc), French or Spanish, Different forms of âwhat Ě s up? how Ě s it going?â, Sentence fragments (long noun phrases), Short statements about a character performing an action, About medieval European islands / cold places, Long blocks of quoted text, Short fragments separated by commas, Long blocks of text, Flirtation, Mil- itary, Non-English, Questions, Violence Table 7.Annotated patterns for each dataset. Edited for brevity, duplicates removed. An Interpretability Illusion for BERT A. Normalization We decided to conduct our analysis on raw embeddings, which are un-normalized. We found that the top activating sentences for directions in the embedding space were in general the same with and without normalization, therefore we decided to keep the vector space as close to how the neural network utilizes it as possible by not normalizing. In addition, we observed that most vector norms are con- centrated aroundâź14which is aboutsqrt(768)/2. This is about the norm of a vector which has all elements equal to 0.5. We also observed that the outputs of the neurons typically fall between -1.0 and 1.0.