Paper deep dive
Emergent Linear Representations in World Models of Self-Supervised Sequence Models
Neel Nanda, Andrew Lee, Martin Wattenberg
Models: Othello-GPT (custom 8-layer transformer)
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/12/2026, 8:24:34 PM
Summary
This paper investigates the internal representations of OthelloGPT, a transformer model trained on Othello move sequences. The authors demonstrate that the model learns a linear representation of the board state relative to the current player (MINE, YOURS, EMPTY) rather than absolute colors (BLACK, WHITE). They validate this finding through causal interventions using vector arithmetic and explore additional linear features like 'FLIPPED' tiles, providing insights into the model's emergent world model and mechanistic interpretability.
Entities (5)
Relation Signals (3)
Neel Nanda ā authored ā Emergent Linear Representations in World Models of Self-Supervised Sequence Models
confidence 100% Ā· Neel Nanda, Andrew Lee, Martin Wattenberg. Proceedings of the 6th BlackboxNLP Workshop
OthelloGPT ā learns ā Linear Representation
confidence 95% Ā· In this work, we provide evidence of a closely related linear representation of the board.
Linear Probes ā validates ā World Model
confidence 90% Ā· To validate that our probes have found a true world model, we confirm that the model uses the encoded board state for its predictions.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Neel Nanda, Andrew Lee, Martin Wattenberg. Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. 2023.
Tags
Links
Full Text
51,555 characters extracted from source content.
Expand or collapse full text
Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, pages 16ā30 December 7, 2023. Ā©2023 Association for Computational Linguistics 16 Emergent Linear Representations in World Models of Self-Supervised Sequence Models Neel Nanda ā Independent Andrew Lee ā University of Michigan Martin Wattenberg Harvard University Abstract How do sequence models represent their decision-making process? Prior work suggests that Othello-playing neural network learned nonlinear models of the board state (Li et al., 2023a). In this work, we provide evidence of a closely relatedlinearrepresentation of the board. In particular, we show that prob- ing for āmy colourā vs. āopponentās colourā may be a simple yet powerful way to inter- pret the modelās internal state. This precise understanding of the internal representations allows us to control the modelās behaviour with simple vector arithmetic.Linear rep- resentations enable significant interpretability progress, which we demonstrate with further exploration of how the world model is com- puted. 1 1 Introduction How do sequence models represent their decision- making process? Large language models are ca- pable of unprecedented feats, yet largely remain inscrutable black boxes. Yet evidence has accu- mulated that models extract features ā articulable properties of the input 2 ā and represent them in its internal activations (Geva et al., 2021; Bau et al., 2020; Gurnee et al., 2023; Belinkov, 2022; Burns et al., 2022; Goh et al., 2021; Elhage et al., 2022a). A key first step for interpreting them is understand- ing how these features are represented. Mikolov et al. (2013c) introduce thelinear representation hypothesis: that features are represented linearly as directions in activation space. This would be highly consequential if true, yet this remains con- troversial and without conclusive empirical justifi- cation. In this work, we present novel evidence of * Equal contribution.neelnanda27@gmail.com, ajyl@umich.edu 1 Code available athttps://github.com/ajyl/mech_ int_othelloGPT 2 Note that our use of the term refers to a higher-level notion than its more common use in deep learning terminology, i.e., an individual neuron. 2 3 4 5 6 7 8 2 3 4 5 6 7 8 Transformer Block Groundtruth Board-States Projected Board-States EmbedTokens Residual Stream Unembed Logits E6 F4 D3 D6 E3 C4 E6 F4 D3 D6 E3 C4 B3 Figure 1: The emergent world models of OthelloGPT are linearly represented. We find that the board states are encoded relative to the current playerās colour (MINEvs. YOURS) as opposed to absolute colours (BLACKvs. WHITE). linear representations, and show that this hypothe- sis has real predictive power. We build on the work of Li et al. (2023a), who demonstrate the emergence of aworld modelin sequence models. Namely, the authors train Oth- elloGPT, an autoregressive transformer model, to predict legal moves in a game of Othello given a sequence of prior moves (Section 2.2). They show that the model spontaneously learns to track the correct board state, recovered usingnon-linear probes, despite never being told that the board ex- ists. They further show a causal relationship be- tween the modelās inner board state and its move predictions using model edits. Namely, they show that the edited network plays moves that are legal in the edited board state even if illegal in the orig- inal board, and even if the edited board state is unreachable by legal play (i.e., out of distribution). Critically, the original authors claim that Othel- loGPT usesnon-linearrepresentations to encode the board state, by achieving high accuracy with non-linear probes, but failing to do so using linear 17 probes. In our work, we demonstrate that a closely related world model is actuallylinearlyencoded. Our key insight is that rather than encoding the coloursof the board (BLACK, WHITE, EMPTY), the sequence model encodes the boardrelativeto the current player of each timestep (MINE, YOURS, EMPTY). In other words, for odd timesteps, the model considersBLACKtiles asMINEandWHITE tiles asYOURS, and vice versa for even timesteps (Section 3). Using this insight, we demonstrate that alinearprojection can be learned with near perfect accuracy to derive the board state. We further demonstrate that we can steer the se- quence modelās predictions by simply conducting vectoral arithmetics using our linear vectors (Sec- tion 4). Put differently, by pushing the modelās activations in the directions ofMINE, YOURS, or EMPTY, we can alter the modelās belief state of the board, and change its predictions accordingly. Our intervention method is much simpler and in- terpretable than that of Li et al. (2023a), which rely on gradients to update the modelās activations (Section 4.1). Our results confirm that our inter- pretation of each probe direction is correct, but also demonstrates that a mechanistic understanding of model representations can lead to better con- trol. Our results do not contradict that of Li et al. (2023a), but add to our understanding of emergent world models. We provide additional interpretations of the se- quence model using linear operations. For example, we provide empirical evidence of how the model derives empty tiles of the board, and find additional linear representations, such as tiles beingFLIPPED at each timestep. Finally, we provide a short discussion of our thoughts. How should we think of linear versus non-linear representations? Perhaps most interest- ingly, why do linear representations emerge? 2 Preliminaries In this section we briefly describe Othello, Othel- loGPT, and our notations. 2.1 Othello Othello is a two player game played on a 8x8 grid. Players take turns playing black or white discs on the board, and the objective is to have the majority of oneās coloured discs by the end of the game. The board always starts with the middle 4 tiles filled with black and white tiles. At each turn, when a tile is played, all of the opponentās discs that are enclosed in a horizontal, vertical, or diagonal row between two discs of the current player are flipped. The game ends when there are no more valid moves for both players. 2.2 OthelloGPT OthelloGPT is a 8-layer GPT model (Radford et al., 2019), each layer consisting of 8 attention heads and a 512-dimensional hidden space. We use the model weights provided by Li et al. (2023a), de- noted there as the synthetic model. The vocabulary space consists of 60 tokens, 3 each one correspond- ing to a playable move on the board (e.g., A4). The model is trained in an autoregressive manner, meaning for a given sequence of movesm <t , the model must predict the next valid movem t . Note that no a priori knowledge of the game nor its rules are provided to the model. Rather, the model is only given move sequences with a training objective to predict next valid moves, by randomly sampling sequences of games from a game tree. This training objective differs from that of models like AlphaZero (Silver et al., 2018), which are trained to play strategic moves to win games. 2.3 Notations Transformers. Our transformer architecture (Vaswani et al., 2017) consists of embedding and unembedding layersEmbandUnembwith a se- ries ofLtransformer layers in-between. Each trans- former layerlconsists ofHmulti-head attentions and a multilayer perception (MLP) layer. A forward pass in the model first embeds the input token at timesteptusing embedding layer Embinto a high dimensional spacex 0 t āR D . We refer tox 0 tāT as the start of theresidual stream. Then each attention headAtt h l ,āhāHand MLP block at layerladd to the residual stream: x l_mid t =x l t + ā hāH Att h l (x l t ) x l+1 t =x l_mid t +MLP l (x l_mid t ) Each attention headAtt h l computes value vec- tors by projecting the residual stream to a lower dimension usingAtt h l .V, linearly combines value 3 The game always starts with 4 tiles in the center of the board already filled. 18 x 0 x 1 x 2 x 3 x 4 x 5 x 6 x 7 Randomized3735.133.935.534.834.734.434.5 Probabilistic61.8 Linear BLACK, WHITE, EMPTY62.274.874.975.075.074.974.874.4 Non-Linear BLACK, WHITE, EMPTY63.488.693.396.397.598.398.798.3 Linear MINE, YOURS, EMPTY90.994.897.298.39999.499.699.5 Table 1: Probing accuracy for board states. OthelloGPT linearly encodes the board state relative to the current player at each timestep (MINEvs. YOURS, as opposed to colours BLACKor WHITE. vectors usingAtt h l .A, and projects back to the residual stream usingAtt h l .O: h(x) = (Attn h l .AāAttn h l .OĀ·Attn h l .V)Ā·x whereānotates a tensor product. A final pre- diction is made by applyingUnembonx Lā1 , fol- lowed by a softmax. Probe Models. We notate linear and non-linear probes asp Ī» andp ν . Our linear probes are sim- ple linear projections from the residual stream: p Ī» (x l t ) =softmax(Wx l t ),WāR DĆ3 . The di- mensionDĆ3comes from doing a 3-way classifi- cation. 4 Non-linear probes are 2-layer MLP mod- els:p ν (x l t ) =softmax(W 1 ReLU(W 2 x l t )),W 1 ā R HĆ3 ,W 2 āR DĆH . Li et al. (2023a) classify the colour at each tile (BLACK, WHITE, EMPTY). Our insight is to classify the coloursrelativeto the current turnās player (MINE, YOURS, EMPTY). 3 Linearly Encoded Board States In this section we describe our experiments to find linear board state representations. 3.1 Experiment Setup Rather than encoding the colour of each tile (BLACK, WHITE, EMPTY), OthelloGPT encodes each tilerelativeto the player of each timestep (MINE, YOURS, EMPTY) ā foroddtimesteps, we considerBLACKto beMINEandWHITEto be YOURS, and vice versa foreventimesteps. In order to learn the weights of our linear probe, we train on random game sequences until a valida- tion loss on a set of 512 games converges according to a patience value of 10. In practice, our linear probes converge after around 100,000 training sam- ples. We test our probes on a held out set of 1,000 games. 4 In practice, because we are predicting the state of all 64 tiles, the shape of our probe isDĆ64Ć3. Residual Stream Residual Stream Residual Stream Unembed P EMPTY (D3) P MINE (D4) P YOURS (D3) P EMPTY (D3) P MINE (D4) P YOURS (D3) Original Board-State x -1 x i+2 x i x i+1 Figure 2: Intervening methodology: we intervene by adding either EMPTY, MINE, or YOURSdirections into each layer of the residual stream. Red squares in each board indicate the tiles that have been intervened, teal tiles indicate new legal moves post-intervention that the model predicts. We train a different probe for each layerl. Hy- perparameters are provided in the Appendix. 3.2 Results Table 1 shows the accuracy for various probes. We include four baselines. The first is a linear probe trained on a randomly initialized GPT model. We also include a probabilistic baseline, in which we always choose the most likely colour per tile at each timestep, according to a set of 60,000 games from training data. The next two baselines are probe models used in Li et al. (2023a): a linear and non-linear probe trained to classify amongst BLACK, WHITE, EMPTY. Our linear probes achieve high accuracy by layer 4. Unbeknownst previously, we show that the emerged board state is linearly encoded. 19 4 Intervening with Linear Directions In this section we demonstrate how we intervene on OthelloGPTās board state using linear probes. 4.1 Method An inherent issue with probing is that it is corre- lational, not causal (Belinkov, 2022). To validate that our probes have found a true world model, we confirm that the model uses the encoded board state for its predictions. To verify this, we conduct the same intervention experiment as Li et al. (2023a). Namely, given an input game sequence (and its corresponding board stateB), we intervene to make the model believe in an altered board stateB ā² . We then observe whether the modelās prediction reflects the made-believe board stateB ā² or the original board stateB. Our intervention approach is simple (Figure 2): we add our linear vectors to the residual stream of each layer: x ā² āx+αp Ī» d (x) wheredindicates a direction amongst MINE, YOURS, EMPTY andαis a scaling factor. In other words, to flip a tile fromYOURStoMINE, we simply push the residual stream at every layer in theMINEdirection, or to āeraseā a previously played tile, we push in the EMPTYdirection. 5 6 Note that this intervention is much simpler than that of Li et al. (2023a). Namely, Li et al. (2023a) edits the activation space (x) of OthelloGPT using their non-linear probes. More specifically, they use non-linear probes to predict board stateB, then compute gradients had the correct board state been the target board stateB ā² , and finally use the gradi- ents to update theactivation spaceof OthelloGPT rather than the weights of the probe model. Instead, we perform a single vector addition. 4.2 Experiment Setup For our intervention experiment, we adopt the same setup and metrics as Li et al. (2023a). We use an evaluation benchmark consisting of 1,000 test cases. Each test case consists of a partial game sequence (B) and a targeted board stateB ā² . 5 We experiment with intervening on different layers. See Appendix for more details. 6 We use the TransformerLens library:https://github. com/neelnanda-io/TransformerLens. Flipping coloursAvg. # Errors Null Intervention Baseline2.723 Non-Linear Intervention0.12 Linear Probe Addition0.10 ErasingAvg. # Errors Null Intervention Baseline2.73 Non-Linear Intervention0.11 Linear Probe Addition0.02 Table 2: Error rates from interventions. We measure the number of false positives and false negatives in the top-N predictions post-intervention, where N is the number of legal moves in the target board stateB ā² . We measure the efficacy of our intervention by treating the task as a multi-label classification prob- lem. Namely, we compare the top-Npredictions post-intervention against the groundtruth set of le- gal moves at stateB ā² , whereNis the number of legal moves atB ā² . We then compute error rate, or the number of false positives and false negatives. Li et al. (2023a) only considers the scenario of flipping the colour of a tile. To also validate our EMPTYdirection, we also experiment with āeras- ingā a previously played tile by making it empty. 4.3 Results Table 2 shows the average error rates after our inter- ventions. A null intervention measures the number of errors by comparing pre-intervention predictions on post-intervention groundtruths. Our interven- tions are equally effective as that of gradient-based editing (Li et al., 2023a), and confirms that our in- terpretation of each linear direction matches how the model uses such directions. 5 Additional Linear Interpretations The linear representation hypothesis is of interest to the mechanistic interpretability community be- cause it provides a foothold into understanding a system. The internal state of the transformer, the residual stream, is the sum of the outputs of all pre- vious components (heads, layers, embeddings and neurons) (Elhage et al., 2021). Albeit the residual stream consisting of linear and non-linear trans- formations, linear functions of the residual stream allow us to identify where a computation of inter- est takes place, or trace how a representation of interest evolves over a forward pass. In this section we leverage our newfound linear representation of board state to provide additional 20 interpretations of OthelloGPT, as proof of concept of how discovering linear representations unlocks downstream interpretability applications. 5.1 Interpreting Empty Tiles Here we interpret how OthelloGPT derives the sta- tus of empty tiles. The EMPTYCircuit.A key insight forEMPTY is that input tokens each correspond to a tile on the board (i.e., A4), and once played, the tile can only change colour but remains non-empty. We view OthelloGPT as using attention heads to ābroadcastā which moves have been played: given a move at timestept, attention heads write this information into other residual streams. This infor- mation (PLAYED) can be represented as following. First, each movem(A4) is embedded:Emb[m]. Then the model writes this information to other residual streams using linear projectionsAtt.Vand Att.O(Section 2.3): PLAYED h (m) =Emb[m]@Att h .V@Att h .O For each attention head in the first layer, 7 we compute the cosine similarity betweenPLAYED and thep Ī» EMPTY direction: max hāH CosSim(PLAYED h (m),p Ī» EMPTY (m)) Since the two terms encodeoppositeinformation, we expect a high negative cosine similarity. We observe an average similarity score of-0.862 across all 60 squares, 8 , confirming thatp EMPTY is encodingNOTPLAYED. This tells us thatp EMPTY is a linear function of the token embeddings. This also implies that OthelloGPT knows which tiles are empty byx 0_mid : after the first attention heads but before the MLP layer. On a binary clas- sification task ofEMPTYvs.NOT-EMPTYfrom 1,000 games in our test split, our probe achieves an accuracy of76.8%and98.9%, when project- ing from the residual stream before and after the attention heads from the first layer. 7 Knowing which moves werePLAYED(i.e. show up in the input sequence), should not depend on any other computa- tion, and thus we expect this information to be written by the attention heads in the first layer. 8 The center 4 squares can never be empty. Figure 3: Difference in probability of A4 being empty, between our clean and corrupt sequences, measured in each attention head. Figure 4: Examples of attention heads from the first layer attending to moves that are YOURS(left) or MINE (right). Logit Attribute for EMPTY. The previous anal- ysis is based on theweightsof the model. Here we provide an alternative analysis by studying the activationsduring inference. First, we select a movem(A4) that we wish to explain. We then construct a ācleanā and ācor- ruptā set of partial game sequences (N=4,569). Our clean set always includesm, while our corrupt set replaces all timesteps withmin the clean set with an alternative move. We ensure that all games in our corrupt set remain legal sequences. Finally, we study thedifference in probabilitythatmis empty, according to our probes, in our two sets. Namely, we project the outputs from each attention head onto the EMPTYdirection and apply a softmax: P EMPTY[m] (Ļ) =Softmax(Ļāp Ī» EMPTY[m] ) whereĻis the output from each attention head. Figure 3 shows the difference in probability that A4 is empty, between our clean and corrupt inputs, measured in each attention head of the first layer. 21 x 0 x 1 x 2 x 3 x 4 x 5 x 6 x 7 Linear FLIPPED, NOT-FLIPPED74.7685.7591.6294.8296.4497.1396.8296.3 Table 3:F1score for probing on FLIPPEDtiles. In addition to the board state, the model also linearly encodes concepts such as flipped tiles per timestep. The figure decomposes two scenarios: when A4 was originally played asMINEorYOURS. This is because some attention heads only attend to moves that areMINE(4, 7), while some only attend to YOURS(1, 3, 8), which we show below. 5.2 Attending to MY& YOURTimesteps We find that some attention heads only attend to eitherMYorYOURmoves. Figure 4 shows two examples: at each timestep, each headalternates between attending to even or odd timesteps. Such behavior further indicates how the model computes its world model based onMINEandYOURSas opposed to BLACKand WHITE. 5.3 Additional Linear Concepts: FLIPPED In addition to linearly representing the board state, we find that OthelloGPT also encodes which tiles are being flipped, or captured, at each timestep. To test this, we modify our probing task to classify be- tweenFLIPPEDvs.NOT-FLIPPED, with the same training setup described above. Given the class im- balance, for this experiment we reportF1scores. Table 3 demonstrates highF1scores by layer 3. We also conduct a modified version of our inter- vention experiment, in which we always randomly select a flipped tile at the current timestep to in- tervene on. Then, instead of adding eitherp Ī» MINE , p Ī» YOURS , orp Ī» EMPTY , wesubtractp Ī» FLIPPED . This tests whether theFLIPPEDfeature is causally relevant for computing the next move, by exploring whether this is sufficient to cause the model to play valid moves in the new board state. We get an average error rate of0.486, compared to a null intervention baseline rate of1.686. One can considerFLIPPEDtiles as the differ- ence between the previous and current board state. One might naturally think that a recurrent com- putation could derive the current board state by iteratively applying such differences. However, transformer models donotmake recursive com- putations! 9 Also, the derivative property of cap- tured tiles being encoded in later layers might be 9 Doing so would require our transformer model to have the same number of layers as the maximum game sequence length of 60. analogous to observations from previous studies of language models that show low-level lexical prop- erties being encoded in lower layers and syntax and semantics being mostly captured in higher layers (Tenney et al., 2019). 5.4 Multiple Circuits Hypothesis Although we find board state representations and their causality on move predictions, we find that they do not explain the entire model. Namely, if our understanding is correct, we expect the model to compute the board state before computing valid moves. However, we find that in end games, this is not the case. To check for the correct board state, we apply our linear probes on each layer, and check the earliest layer in which all 64 tiles are correctly predicted. 10 To check for correct move predictions, we project from each layer using the unembedding layer, and check the earliest layer in which the top-N move predictions are all correct, where N is the number of groundtruth legal moves. Figure 5 plots the proportion of times the board state is computed before (or after) valid moves (first y-axis). We also overlay the average earliest layer in which board or moves are correctly com- puted (second y-axis, aqua and lime curves). To our surprise, we find that in end games, the model often computes legal movesbeforethe board state (black bars). We henceforth refer to this behavior as MOVEFIRST, and share some thoughts. End Game Circuits.First,MOVEFIRSTstarts to occur around move 30, which is the mid-point of the game. Second,MOVEFIRSToccurs more fre- quently as we near the end of the game (increasing black bars). Interestingly, in Othello, starting from the mid-point, there are progressively fewer empty tiles than there are filled tiles as the board fills up. Also note that as the game progresses, it becomes more likely for every empty tile to be a legal move. One possible explanation for this phenomenon is that in the end game, it may be possible to pre- 10 It might be the case that legal moves could be predicted without 100% accuracy of the board state. We try variants (see Appendix), but observe similar trends. 22 01020304050 0 0.2 0.4 0.6 0.8 1 0 1 2 3 4 5 6 7 Earliest Layer (Board) Earliest Layer (Moves) Incorrect (Board) After Same Before Moves (Timestep) % of Games Figure 5: Proportion of times the board state is computed before/after move predictions are made (First y-axis). Light Grey:Boards are computed in an earlier layer than moves.Dark Grey, Black:Boards are computed in the same or later layer than moves.Red:Model never computes the correct board state.Aqua, Lime (Curves): Average earliest layer in which the board or moves are correctly computed (Second y-axis). Starting from the mid- game, we start observing the model compute moves before boards (black bar), and this occurs more frequently as the game progresses. dict legal moves with simpler circuits that do not require the entire board state. For instance, perhaps it combinesEMPTYwith other features such asIS- SURROUNDED-BY-MINEorIS-BORDERand so on. Multiple Circuits.Interestingly, the model still uses the board state at end games. To demon- strate this, we run our intervention experiment on 1,000end games, 11 and still achieve a low error rate of0.112. 12 We thus hypothesize that Othel- loGPT (and more broadly, sequence models) con- sist of multiple circuits. Another hypothesis is that residual networks make āiterative inferencesā (Sec- tion 5.5), and for end games, OthelloGPT uses simpler circuits in the early layers and refines its predictions at late layers using board state. End Game Board Accuracy. We observe that board state accuracy drops near end games. This can be seen by the growing red bars, but also by measuring per-timestep accuracy of our probes (see Appendix). It is unclear whether 1) the model does not bother to compute the perfect board state, as alternative circuits allow the model to still correctly predict legal moves, or 2) the model learns an alter- native circuit because it struggles to compute the correct board state at end games. Memorization. Note that in the first few timesteps, the board and legal moves are some- times both computed in the same layer (dark grey bars). This may be due to memorization: 1) these 11 We intervene on a timestep > 30 12 Non-intervention baseline: 1.988. predictions both occur at the first layer, and 2) there are only so many openings in an Othello game. 5.5 Iterative Feature Refinements Figure 6 visualizes OthelloGPTās āiterative infer- enceā (Jastrzebski et al., 2018; Belrose et al., 2023; Veit et al., 2016; nostalgebraist, 2020), or itera- tive refinement of features. For each layer, we plot the projected board states using our probes, and projected next-move predictions using the un- embedding layer. Multiple evidence of iterative refinements are provided in the Appendix. 6 Discussions 6.1 On Linear vs. Non-Linear Interpretations One challenge with probing is knowing which features to look for. 13 For instance, classifying BLACK, WHITE versus MINE, YOURS leads to different takeaways, which illustrates the danger of projecting our preconceptions. What might seem āsensibleā to a human interpreter (BLACK, WHITE) may not be for a model. In hindsight, given the symmetric game-play of Othello, encoding MINE, YOURSis perfectly sensible for the model (For more examples of non-obvious, sensible features, see (McCoy et al., 2019; Nanda et al., 2023)). More broadly, what is sensible, or alternatively, how we choose to interpret linear or non-linear en- codings, can be relative to how we see the world. Suppose we had a perfect world model of our phys- ical world. Further suppose that if and when it 13 For a longer discussion on probing, see Appendix. 23 12 3 45 6 7 8 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 1 2 34 5 6 78 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A 12 3 45 6 7 8 H G F E D C B A Layer 0Layer 1Layer 2Layer 3Layer 4Layer 5Layer 6Layer 7Board State Figure 6: Iterative refinements: the top row shows each layer projected using our linear probes. The bottom row shows the modelās predictions for legal moves at each layer, by applying the unembedding layer on each layer. computes a gravitational force between two ob- jects (Newtonās law), we discover a neuron whose square root was the distance between two objects. Is this a non-linear representation of distance? Or, given the form of Netwonās law, is the square of the distance a more natural way for the model to represent the feature, and thus considered a linear representation? As this example shows, what con- stitutes a natural feature may be in the eye of the beholder. 6.2 On the Emergence of Linear Representations Linear representations in sequence models have been observed before: iGPT (Chen et al., 2020), which was autoregressively trained to predict next pixels of images, lead to robust linear image rep- resentations. The question remains, why do linear feature representations emerge? What linear repre- sentations are currently encoded in large language models? One reason might be simply that matrix multiplication can easily extract a different subset of linear features for each neuron. However, we leave a complete explanation to future work. 7 Related Work We discuss three broad related areas: understanding internal representations, interventions, and mecha- nistic interpretability. 7.1 Understanding Internal Representations Multiple researchers have studied concept represen- tations in sequence models. Li et al. (2021) train sequence models on a synthetic task, and uncover world models in their activations. Patel and Pavlick (2022) demonstrate that language models can learn to ground concepts (e.g., direction, colour) to real world representations. Burns et al. (2022); Marks and Tegmark (2023) find linear vectors that en- code ātruthfulnessā. Probing techniques have also been used to extract linguistic characteristics in sen- tence embeddings (Conneau et al., 2018; Tenney et al., 2019). Researchers have also usedstruc- turalprobes to uncover syntactic structures in word embeddings (Hewitt and Manning, 2019) and lan- guage models (Eisape et al., 2022). Prior to current day language models, word embeddings (Mikolov et al., 2013b,a) built vectoral word representations. Linear representations are found outside of lan- guage models as well. Merullo et al. (2022) demon- strate that image representations from vision mod- els can be linearly projected into the input space of language models. McGrath et al. (2022) and Lover- ing et al. (2022) find interpretable representations of chess or Hex concepts in AlphaZero. 7.2 Intervening On Language Models A growing body of work has intervened on lan- guage models, by which we mean controlling their behavior by altering their activations. We consider two broad categories. Paramet- ric approaches often use optimizations (i.e. gra- dient descent) to locate and alter activations (Li et al., 2023a; Meng et al., 2022a,b; Hernandez et al., 2023; Hase et al., 2023).Meanwhile, inference-time interventions typically apply linear arithmetics, for instance by using ātruthfulā vec- tors (Li et al., 2023b), ātask vectorsā (Ilharco et al., 2022), or other āsteering vectorsā (Subramani et al., 2022; Turner et al., 2023). 7.3 Mechanistic Interpretability Mechanistic interpretability (MI) studies neural net- works by reverse-engineering their behavior (Olah et al., 2020; Elhage et al., 2021). The goal of MI is to understand the underlying computations and representations of a model, with a broader goal of validating that their behavior aligns with what researchers have intended. Such framework has allowed researchers to better understand grokking 24 (Nanda et al., 2023), superposition (Elhage et al., 2022b; Scherlis et al., 2022; Arora et al., 2018), or even individual neurons (Mu and Andreas, 2020; Antverg and Belinkov, 2021; Gurnee et al., 2023). 8 Conclusion In this work we demonstrated that the emergent world model in Othello-playing sequence models is full of linear representations. Previously unbe- knownst, we demonstrated that the board state in OthelloGPT is linearly represented by encoding the colour of each tilerelativeto the player at each timestep (MINE, YOURS, EMPTY) as opposed to absolute colour (BLACK, WHITE, EMPTY). We showed that we can accurately control the modelās behaviour with simple vector arithmetic on the in- ternal world model. Lastly, we mechanistically interpreted multiple facets of the sequence model, analysing how empty tiles are detected, and linear representations of which pieces are flipped. We find hints that multiple circuits might exist for pre- dicting legal moves in the end game, as well as further evidence that residual networks iteratively refine their features across layers. 9 Acknowledgements We thank the original authors of Li et al. (2023a) for opensourcing their work, making it possible to conduct our research. We thank Chris Olah for invaluable discussion and encouragement, and drawing our attention to the implication of these results for the linear repre- sentation hypothesis. 10 Author Contributions Neel Nanda discovered the linear representation in terms of relative board state, and showed that sim- ple vector arithmetic sufficed for causal interven- tions. He led an initial version of the experiments and write-ups, and advised throughout. Andrew Lee led this write-up and performed all experiments in this paper. He discovered the flipped linear representation, the empty circuit, and the multiple circuit hypothesis results. Martin Wattenberg helped with editing and dis- tilling the paper, and contributed the analogy about a linear vs quadratic representation of distance. References Omer Antverg and Yonatan Belinkov. 2021. On the pitfalls of analyzing individual neurons in language models.arXiv preprint arXiv:2110.07483. Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2018. Linear algebraic struc- ture of word senses, with applications to polysemy. Transactions of the Association for Computational Linguistics, 6:483ā495. David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba. 2020. Understanding the role of individual units in a deep neural network.Proceedings of the National Academy of Sciences. Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances.Computational Lin- guistics, 48(1):207ā219. Nora Belrose, Zach Furman, Logan Smith, Danny Ha- lawi, Igor Ostrovsky, Lev McKinney, Stella Bider- man, and Jacob Steinhardt. 2023. Eliciting latent predictions from transformers with the tuned lens. arXiv preprint arXiv:2303.08112. Collin Burns, Haotian Ye, Dan Klein, and Jacob Stein- hardt. 2022. Discovering latent knowledge in lan- guage models without supervision.ArXiV. Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. 2020. Generative pretraining from pixels. InProceedings of the 37th International Conference on Machine Learning, volume 119 ofProceedings of Machine Learning Research, pages 1691ā1703. PMLR. Alexis Conneau, German Kruszewski, Guillaume Lam- ple, LoĆÆc Barrault, and Marco Baroni. 2018. What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 2126ā2136, Melbourne, Aus- tralia. Association for Computational Linguistics. Tiwalayo Eisape, Vineet Gangireddy, Roger Levy, and Yoon Kim. 2022.Probing for incremental parse states in autoregressive language models. InFind- ings of the Association for Computational Linguis- tics: EMNLP 2022, pages 2801ā2813, Abu Dhabi, United Arab Emirates. Association for Computa- tional Linguistics. Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova Das- Sarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Gan- guli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran- Johnson, Jared Kaplan, Jack Clark, Tom Brown, 25 Sam McCandlish, Dario Amodei, and Christo- pher Olah. 2022a.Softmax linear units.Trans- former Circuits Thread.Https://transformer- circuits.pub/2022/solu/index.html. Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. 2022b. Toy models of superpo- sition.Transformer Circuits Thread. Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Ka- plan, Sam McCandlish, and Chris Olah. 2021. A mathematical framework for transformer circuits. Transformer Circuits Thread. Https://transformer- circuits.pub/2021/framework/index.html. Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021. Transformer feed-forward layers are key-value memories. InProceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 5484ā5495, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics. Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018. Un- der the hood: Using diagnostic classifiers to in- vestigate and improve how language models track agreement information. InProceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and In- terpreting Neural Networks for NLP, pages 240ā248, Brussels, Belgium. Association for Computational Linguistics. Gabriel Goh, Nick Cammarata ā , Chelsea Voss ā , Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. 2021.Multi- modal neurons in artificial neural networks.Distill. Https://distill.pub/2021/multimodal-neurons. Wes Gurnee, Neel Nanda, Matthew Pauly, Kather- ine Harvey, Dmitrii Troitskii, and Dimitris Bert- simas. 2023.Finding neurons in a haystack: Case studies with sparse probing.arXiv preprint arXiv:2305.01610. Peter Hase, Mohit Bansal, Been Kim, and Asma Ghan- deharioun. 2023. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models.arXiv preprint arXiv:2301.04213. Evan Hernandez, Belinda Z Li, and Jacob Andreas. 2023. Measuring and manipulating knowledge rep- resentations in language models.arXiv preprint arXiv:2304.00740. John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word represen- tations. InNorth American Chapter of the Associ- ation for Computational Linguistics: Human Lan- guage Technologies. Association for Computational Linguistics. Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Worts- man, Suchin Gururangan, Ludwig Schmidt, Han- naneh Hajishirzi, and Ali Farhadi. 2022.Edit- ing models with task arithmetic.arXiv preprint arXiv:2212.04089. StanisÅaw Jastrzebski, Devansh Arpit, Nicolas Ballas, Vikas Verma, Tong Che, and Yoshua Bengio. 2018. Residual connections encourage iterative inference. InInternational Conference on Learning Represen- tations. Belinda Z. Li, Maxwell Nye, and Jacob Andreas. 2021. Implicit representations of meaning in neural lan- guage models. InProceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1813ā1827, Online. Association for Computational Linguistics. Kenneth Li, Aspen K Hopkins, David Bau, Fernanda ViĆ©gas, Hanspeter Pfister, and Martin Wattenberg. 2023a. Emergent world representations: Exploring a sequence model trained on a synthetic task. InThe Eleventh International Conference on Learning Rep- resentations. Kenneth Li, Oam Patel, Fernanda ViĆ©gas, Hanspeter Pfister, and Martin Wattenberg. 2023b. Inference- time intervention: Eliciting truthful answers from a language model.arXiv preprint arXiv:2306.03341. Charles Lovering, Jessica Forde, George Konidaris, El- lie Pavlick, and Michael Littman. 2022. Evaluation beyond task performance: Analyzing concepts in al- phazero in hex. InAdvances in Neural Informa- tion Processing Systems, volume 35, pages 25992ā 26006. Curran Associates, Inc. Samuel Marks and Max Tegmark. 2023. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets.arXiv preprint arXiv:2310.06824. Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019. Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference. InProceed- ings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3428ā3448, Florence, Italy. Association for Computational Lin- guistics. Thomas McGrath,Andrei Kapishnikov,Nenad TomaÅ”ev, Adam Pearce, Martin Wattenberg, Demis Hassabis, Been Kim, Ulrich Paquet, and Vladimir Kramnik. 2022. Acquisition of chess knowledge in alphazero.Proceedings of the National Academy of Sciences, 119(47):e2206625119. 26 Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. 2023. The hy- dra effect: Emergent self-repair in language model computations.arXiv preprint arXiv:2307.15771. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022a. Locating and editing factual asso- ciations in GPT.Advances in Neural Information Processing Systems, 36. Kevin Meng, Arnab Sen Sharma, Alex Andonian, Yonatan Belinkov, and David Bau. 2022b. Mass- editing memory in a transformer.arXiv preprint arXiv:2210.07229. Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. 2022. Linearly mapping from image to text space.arXiv preprint arXiv:2209.15162. Tomas Mikolov, Kai Chen, Greg Corrado, and Jef- frey Dean. 2013a.Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781. Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Cor- rado, and Jeff Dean. 2013b. Distributed representa- tions of words and phrases and their compositional- ity. InAdvances in Neural Information Processing Systems, volume 26. Curran Associates, Inc. TomÔŔ Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013c. Linguistic regularities in continuous space word representations. InProceedings of the 2013 conference of the north american chapter of the as- sociation for computational linguistics: Human lan- guage technologies, pages 746ā751. Jesse Mu and Jacob Andreas. 2020. Compositional ex- planations of neurons.Advances in Neural Informa- tion Processing Systems, 33:17153ā17163. Neel Nanda, Lawrence Chan, Tom Liberum, Jess Smith, and Jacob Steinhardt. 2023. Progress mea- sures for grokking via mechanistic interpretability. arXiv preprint arXiv:2301.05217. nostalgebraist. 2020. interpreting gpt: the logit lens. Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020. Zoom in: An introduction to circuits.Distill. Https://distill.pub/2020/circuits/zoom-in. Roma Patel and Ellie Pavlick. 2022. Mapping language models to grounded conceptual spaces. InInterna- tional Conference on Learning Representations. Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. 2020a. Pareto probing: Trading off accuracy for complexity. InProceedings of the 2020 Conference on Empirical Methods in Natural Lan- guage Processing (EMNLP), pages 3138ā3153, On- line. Association for Computational Linguistics. Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020b. Information-theoretic probing for linguistic structure. InProceedings of the 58th Annual Meet- ing of the Association for Computational Linguistics, pages 4609ā4622, Online. Association for Computa- tional Linguistics. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Lan- guage models are unsupervised multitask learners. Naomi Saphra and Adam Lopez. 2019. Understand- ing learning dynamics of language models with SVCCA. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3257ā3267, Minneapolis, Minnesota. Associ- ation for Computational Linguistics. Adam Scherlis, Kshitij Sachan, Adam S Jermyn, Joe Benton, and Buck Shlegeris. 2022. Polysemantic- ity and capacity in neural networks.arXiv preprint arXiv:2210.01892. David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis. 2018. A general reinforcement learning algorithm that masters chess, shogi, and go through self-play.Science, 362(6419):1140ā1144. Nishant Subramani, Nivedita Suresh, and Matthew Pe- ters. 2022. Extracting latent steering vectors from pretrained language models. InFindings of the As- sociation for Computational Linguistics: ACL 2022, pages 566ā581, Dublin, Ireland. Association for Computational Linguistics. Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 4593ā 4601, Florence, Italy. Association for Computational Linguistics. Mycal Tucker, Peng Qian, and Roger Levy. 2021. What if this modified that? syntactic interventions with counterfactual embeddings.InFindings of the Association for Computational Linguistics: ACL- IJCNLP 2021, pages 862ā875, Online. Association for Computational Linguistics. Alex Turner, Monte MacDiarmid, David Udell, lisathiergart, and Ulisse Mini. 2023. Steering gpt- 2-xl by adding an activation vector - ai alignment forum. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Å ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InAdvances in Neural Information Pro- cessing Systems, volume 30. Curran Associates, Inc. 27 Andreas Veit, Michael J Wilber, and Serge Belongie. 2016. Residual networks behave like ensembles of relatively shallow networks.Advances in neural in- formation processing systems, 29. 28 HyperparameterValue OptimizerAdamW Learning Rate1e-2 Weight Decay1e-2 Betas0.9, 0.99 Validation Step 200 Validation Size512 Validation Patience10 Table 4: Hyperparameters used for our linear probes. 1 2 3 4 5 6 7 87654321 0 0.5 1 1.5 2 2.5 Layers Intervened Error Rate First N LayersLast N Layers 0.11 Figure 7: Intervention results depending on layers in- tervened. A Hyperparameters for Linear Probes Table 4 provides hyperparameters used for our lin- ear probes. B Intervening on Different Layers In practice there are a lot of ways to intervene using linear vectors. Figure 7 demonstrates different er- ror rates depending on which layers are intervened. From our experiments, we observe that either a sufficient number of layers need to be intervened for OthelloGPT to alter its predictions. We offer a couple of hypotheses for this. First, we hypothesize that this is because of the residual structure of trans- former models, and while each layer may write additional information into the residual streams, there may still be information from earlier layers that the model uses. A somewhat related hypothe- sis is that OthelloGPT might be demonstrating the Hydra effect (McGrath et al., 2023), in which lan- guage models demonstrate the ability to self-repair its computations after an intervention. C Multiple Circuits In Section 5.4, we find hints that OthelloGPT some- times computes moves before boards at end games. Namely, we check the earliest layers in which the board is correctly predicted with 100% accuracy. Could it be that at end games, legal moves can be predicted without needing the entire board? To this point, we experiment with variations of this exper- iment. In Figure 8, we check the earliest layer in which at least 90% of the board is first correctly computed. In Figure 9, we check the earliest layer in which the āminimum setā of tiles are correctly computed, where the minimum set is set of tiles that make each legal move playable (see Figure 10 for example). Despite a looser criteria for board state, we still see OthelloGPT computing moves before boards at end games. Interestingly, our probes lose accuracy starts to drop in the end game as well (Figure 11). It is unclear whether 1) the model does not bother to compute the perfect board state, as alternative cir- cuits might exist at end games, or 2) the model learns an alternative circuit because it struggles to compute the correct board state at end games. D Evidence of Iterative Feature Refinements As mentioned in Section 5.5, OthelloGPT demon- strates multiple evidence of iterative feature re- finements: 1) Board state accuracy (as well as FLIPPED) improves from layer to layer (Table 1, 3). 2) Next-move predictions also improve from layer to layer. Table 5 reports the top-1 error rate when applying the unembedding layer on every layer using our test set from Section 3. As a base- line, we apply the same unembedding layer from OthelloGPT to the residual streams of a randomly initialized GPT model. 3) Linear probes across layers share similar directions. Figure 12 plots the cosine similarity between all linear probes, av- eraged across all 64 tiles and directions (MINE, YOURS, EMPTY). E On Principled Ways of Probing Probing has produced both excitement and skepti- cism amongst researchers (Belinkov, 2022). Here we provide our learnings regarding probing. One criticism of probes is whether the discov- ered features are actually used by the model, i.e., correlation vs. causation. Intervention is com- monly used to study causality (Giulianelli et al., 2018; Tucker et al., 2021), but have often reached mixed conclusions (Belinkov, 2022). While both linear and non-linear probes have demonstrated 29 01020304050 0 0.2 0.4 0.6 0.8 1 0 1 2 3 4 5 6 7 Earliest Layer (Board) Earliest Layer (Moves) Incorrect (Board) After Same Before Moves (Timestep) % of Games Figure 8: Percentage of times90%of the board state is computed before/after move predictions are made. 01020304050 0 0.2 0.4 0.6 0.8 1 0 1 2 3 4 5 6 7 Earliest Layer (Board) Earliest Layer (Moves) Incorrect (Board) After Same Before Moves (Timestep) % of Games Figure 9: Percentage of times theāminimum setāof necessary board state is computed before/after move predic- tions are made. 12345678 H G F E D C B A Figure 10: Example of āminimum setā of tiles that make move G2 legal. successful interventions (Li et al., 2023b; Turner et al., 2023), linear probes are much easier to inter- pret, as they imply that features simply correspond to vectoral directions. Another challenge is knowing which features to probe for, which can lead to pitfalls. Taking OthelloGPT as an example, classifying BLACK, WHITE versus MINE, YOURS leads to different 01020304050 0.75 0.8 0.85 0.9 0.95 1 Layer 0 Layer 1 Layer 2 Layer 3 Layer 4 Layer 5 Layer 6 Layer 7 Timestep (Moves) A ccuracy Figure 11: Accuracy per timestep for our linear probes. takeaways, which illustrates the danger ofproject- ing our preconceptions. Speaking of incorrect takeaways, our last point concerns the expressivity of probe models. With an expressive-enough probe, there is a danger of the probe computing or memorizing the desired fea- ture that one is looking for, rather than extracting (Pimentel et al., 2020a; Saphra and Lopez, 2019). Still, some researchers view linear classification 30 Baseline: Randomx 0 x 1 x 2 x 3 x 4 x 5 x 6 x 7 0.8560.2150.1520.1120.0790.0490.0150.0040.001 Table 5: Top-1 error rates when applying the unembedding layer to earlier layers. As a baseline we apply Othel- loGPTās unembedding layer on a randomly initialized GPT model. Figure 12:Cosine similarity scores between linear probes across layers. as inadequate (Pimentel et al., 2020b; Saphra and Lopez, 2019). We view our work as evidence that linear probes do have interpretable and controllable power, and anticipate these findings to generalize to larger language models.