Paper deep dive
Monotonic Representation of Numeric Properties in Language Models
Benjamin Heinzerling, Kentaro Inui
Models: Llama-2-13B, Llama-2-7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 7:25:08 PM
Summary
The paper investigates how language models (LMs) encode numeric properties (e.g., birth years, population) in their internal representations. The authors demonstrate that these properties are encoded in low-dimensional, monotonic linear subspaces. By using partial least squares (PLS) regression, they identify specific 'directions' in the activation space that correlate with numeric values. Furthermore, they prove a causal relationship by showing that patching activations along these directions allows for predictable, monotonic manipulation of the model's output.
Entities (5)
Relation Signals (3)
Language Models â encode â Numeric Property
confidence 95% · LMs represent numeric properties in a way that reflects their natural structure
Activation Patching â manipulates â Language Models
confidence 94% · by patching activations along a 'birthyear' direction we can make the LM express an increasingly late birthyear
Partial Least Squares Regression â identifies â Property-Encoding Direction
confidence 93% · PLS identifies small sets of predictive components
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Language models (LMs) can express factual knowledge involving numeric properties such as Karl Popper was born in 1902. However, how this information is encoded in the model's internal representations is not understood well. Here, we introduce a simple method for finding and editing representations of numeric properties such as an entity's birth year. Empirically, we find low-dimensional subspaces that encode numeric properties monotonically, in an interpretable and editable fashion. When editing representations along directions in these subspaces, LM output changes accordingly. For example, by patching activations along a "birthyear" direction we can make the LM express an increasingly late birthyear: Karl Popper was born in 1929, Karl Popper was born in 1957, Karl Popper was born in 1968. Property-encoding directions exist across several numeric properties in all models under consideration, suggesting the possibility that monotonic representation of numeric properties consistently emerges during LM pretraining. Code: this https URL
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
80,080 characters extracted from source content.
Expand or collapse full text
Preprint Monotonic Representation of Numeric Properties in Language Models Benjamin Heinzerling RIKEN / Tohoku University benjamin.heinzerling@riken.jp Kentaro Inui MBZUAI / Tohoku University / RIKEN kentaro.inui@mbzuai.ac.ae Abstract Language models (LMs) can express factual knowledge involving numeric properties such asKarl Popper was born in 1902. However, how this informa- tion is encoded in the modelâs internal representations is not understood well. Here, we introduce a simple method for finding and editing represen- tations of numeric properties such as an entityâs birth year. Empirically, we find low-dimensional subspaces that encode numeric properties monotoni- cally, in an interpretable and editable fashion. When editing representations along directions in these subspaces, LM output changes accordingly. For example, by patching activations along a âbirthyearâ direction we can make the LM express an increasingly late birthyear:Karl Popper was born in 1929,Karl Popper was born in 1957,Karl Popper was born in 1968. Property- encoding directions exist across several numeric properties in all models under consideration, suggesting the possibility that monotonic represen- tation of numeric properties consistently emerges during LM pretraining. Code:https://github.com/bheinzerling/numeric-property-repr 1 Introduction: Do LMs represent numeric properties âappropriatelyâ? Language models (LMs) are capable of expressing factual knowledge (Petroni et al., 2019; Jiang et al., 2020; Roberts et al., 2020; Heinzerling & Inui, 2021; Kassner et al., 2021). For example, when queriedIn which year was Karl Popper born?Llama 2 (Touvron et al., 2023) gives the correct answer1902. While the question if and to what degree LMs can be said to âknowâ anything at all is subject of ongoing debate (Bender & Koller, 2020; Hase et al., 2023b; Mollo & Milli ` ere, 2023; Lederman & Mahowald, 2024), empirical work has progressed from behavioral analysis focused on the accuracy and robustness of knowledge expression (Shin et al., 2020; Jiang et al., 2021; Zhong et al., 2021; Youssef et al., 2023) to representational analysis aimed at understanding how factual knowledge is encoded 1 in model parameters (De Cao et al., 2021; Mitchell et al., 2021; Meng et al., 2022) and activations (Hernandez et al., 2023; Merullo et al., 2023; Geva et al., 2023; Gurnee & Tegmark, 2023). However, representational analysis has so far mainly targeted entity-entity relations such asWarsaw is the capital of Poland(Merullo et al., 2023). If and how LM representations encode factual knowledge involving numeric properties, such as an entityâs birthyear, is less understood. Unlike entity-entity relational knowledge, numeric properties have natural ordering and monotonic structure, such as earlier/later and smaller/larger scales or geographic coordinate systems. While this kind of structure is natural and intuitive for humans, LMs encounter numeric properties only in form of largely unordered and unstructured textual mentions. This raises the question if LMs learn to represent numeric properties appropriately, i.e., according to their natural structure. 1 We say âX is encoded in Yâ as shorthand for âX can be easily extracted from Yâ. See caveats in§5. 1 arXiv:2403.10381v1 [cs.CL] 15 Mar 2024 Preprint Figure 1: Sketch of our main finding. Patching entity representations along specific direc- tions in activation space yields corresponding changes in model output. Here, we devise a simple method for identifying and manipulating representations of numeric properties in LMs. We find low-dimensional subspaces that strongly correlate with numeric properties across models and numeric properties, thereby confirming and extending prior observations of representations of numeric properties in LMs (Li Ì etard et al., 2021; Faisal & Anastasopoulos, 2023; Gurnee & Tegmark, 2023; Godey et al., 2024). Going beyond prior work (see§4), we show that by causally intervening along certain directions in these subspaces, LM output changes correspondingly. That is, we find a monotonic relationship between the intervention and the quantity expressed by the LM. For example, an entitityâs year of birth shifts according to the strength and sign of the intervention along a âbirthyearâ direction (Figure 1). Taken together, our findings suggest that LMs represent numeric properties in a way that reflects their natural structure and that such monotonic representations consistently emerge during LM pretraining. Terminology.Before moving to the main part, we briefly clarify important terms. A quantityconsists of a scalar numericvaluepaired with aunitof measurement. Anumeric propertyis a property that can naturally be described by a quantity, e.g., birthyear, popu- lation size, geographic latitude. Anumeric attributeis an instance of a numeric property, associated with a particular entity. For example, Karl Popper has the numeric attribute birthyear:1902. Bylinear representationwe denote the idea that a numeric attribute is encoded in a linear subspace of a LMâs activation space. Finally, amonotonic representation is a linear representation characterized by a monotonic relationship between directions in activation space and the value of the encoded numeric attribute. That is, as activations shift along a particular direction the value of the corresponding numeric attribute increases or decreases monotonically. 2 Finding Property-Encoding Directions 2.1 Motivation: Representation learning, linear representation hypothesis While numeric properties generally can be mapped naturally onto simple canonical struc- tures, such as number lines or coordinate systems, it is not immediately obvious that pretraining on largely unstructured data enables LMs to appropriately represent such struc- tures. Our main goal is to find out if and how numeric properties, such as an entityâs birthyear, are encoded in the geometry of LM representations. How could such an encoding look like? Based on two arguments, we hypothesize that numeric properties are encoded in low-dimensional linear subspaces of activation space. The first argument rests on a key principle in representation learning: a model generalizes if and only if its representations reflect the structure of the data (Conant & Ashby, 1970; Liu et al., 2022). To the degree that current LMs generalize, in the sense of achieving non-trivial 2 Preprint PropertyProp. IDEntityEntity IDPromptValueUnit birthyearP569Nina FochQ235632In what year was Nina Foch born?1924annum death yearP570Johannes R. BecherQ58057In what year did Johannes R. Becher die?1958annum populationP1082AkhisarQ209905What is the population of Akhisar?1730261 evelationP2044SondrioQ6274How high is Sondrio?360metre longitudeP625.longKorean EmpireQ28233What is the longitude of Korean Empire?126.98degree latitudeP625.latK Ì usnachtQ69216What is the latitude of K Ì usnacht?47.32degree Table 1: Random sample of the entities used in our experiments, along with corresponding numeric attributes and prompts. See details and more samples in Appendix A. performance on benchmarks involving knowledge of numeric properties (Petroni et al., 2019), we can expect that their representations reflect the structure of numeric properties. And since the natural structures of many numeric properties are low-dimensional linear spaces such as number lines, we expect to find low-dimensional linear structure in the representations of a well-performing model. As second argument we adduce the linear representation hypothesis, which posits a cor- respondence between concepts and linear subspaces (Elhage et al., 2022; Park et al., 2023; Nanda et al., 2023). If the linear representation hypothesis is true, 2 this would imply that nu- meric properties are encoded in linear subspaces. For brevity, we will call a low-dimensional linear subspace of a LMâs activation space adirection, regardless of whether it is one- or multi-dimensional. 2.2 Method: Partial least squares regression on entity representations Motivated by the hypothesis that numeric properties are encoded as directions in activation space, we now devise an experimental setup for finding out if such directions exist. An obvious choice for identifying linear structure would be principal component analysis (PCA; Pearson, 1901). However, PCA looks for directions of maximum variance and is unsupervised in the sense that we cannot cannot provide model outputs to steer the algorithm towards particular directions. What we would like to do instead is to provide supervision in order to find directions in activation space that maximally covary with model outputs. A standard method applicable to this kind of problem is partial least squares regression. Partial least squares regression (PLS; Wold et al., 2001) is a supervised method for finding relationships between two data matricesXandY. Briefly summarizing the detailed exposi- tion by Wold et al. (2001), the objective of PLS regression is to find a rank-kdecomposition of the data matrixXinto âX-scoresâT=XWâR nĂk and âX-loadingsâPâR dĂk so that XâTP T =XWP T , with weightsWâR dĂk . The decomposition is subject to the constraint that theX-scores (or conversely: the correspondingkcomponents) can be used to predictY, that isYâXWC T , with âY-loadingsâCâR k . TheY-loadings can be seen as coefficients that quantify how much each of thekcomponents contributes to the prediction of the regression model. There exist efficient methods for solving this optimization problem. 3 In our caseXcorresponds to (a sample of) the activation space of a LM andYto corre- sponding LM outputs. Concretely, for a given numeric property, such as birthyear, we collectnentities that have this property. For each entityewe prepare a suitable prompt and encode this prompt with a LM to obtain an entity representationx e of dimensiond. That is,X= [ x 1 ·e n ] T âR nĂd . Put simply, we encode prompts such asWhen was Karl Popper born?and take the hidden state of a particular token at a particular layer as the LMâs entity representation. We also collect the quantityy e expressed by the LM when queried for the numeric attribute of entitye, that is,Y=y e |eâ [ 1 . .n ] âR n . In the case of our running example query for Karl Popperâs birthyear, Llama 2 expresses the quantity 2 For positive evidence, see Marks & Tegmark (2023); Merullo et al. (2023); Tigges et al. (2023); Jiang et al. (2024), i.a. 3 We use the Scikit-learn (Pedregosa et al., 2011) implementation of NIPALS (Wold, 1966). 3 Preprint 01020304050 #components 0.4 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (a) Birthyear 01020304050 #components 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (b) Death year 01020304050 #components 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (c) Population 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (d) Elevation 01020304050 #components 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (e) Latitude 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (f) Longitude Figure 2: Low-dimensional subspaces of Llama-2-13Bâs 5120-dimensional activation space are predictive of the quantity expressed by the LM when queried for a numeric attribute of an entity, across six different numeric properties. Each subfigure shows the performance of a regression model fitted to predict the expressed quantities from LM-internal entity representations (in layerl=0.3), as a function of the number of PCA/PLS components used for prediction. Unlike regression on PCA components (dashed orange), partial least squares regression (PLS, solid blue) identifies a small set of predictive components. Controls with shuffled labels (dotted green, dash-dotted red) and random entity representations (long-dash-dot purple, dash-dot-dot brown) fail to find predictive subspaces. y e =1902. Having collected entity representationsXand associated LM outputsY, we fit a k-component PLS regression model to predictYfrom ak-dimensional subspace ofX. To measure how well numeric attributes can be predicted from low-dimensional subspaces of activation space, we vary the number of PLS regression componentsk, i.e., the dimen- sionality of the subspace, and record goodness of fit using the coefficient of determination R 2 =1â â n e=1 (y e â Ë y e ) 2 / â n e=1 (y e â Ì y) 2 , where Ë y e is the prediction of the regression model and Ì ythe sample-mean value of the numeric property in question. 4 2.3 Results: Low-dimensional subspaces predictive of numeric attribute expression After selecting six of the most frequent numeric properties 5 in Wikidata (Vrande Ë ci Ì c & Kr Ì otzsch, 2014), for each property we randomly samplen=1000 popular 6 entities and prompt the LM (in English) for the corresponding attribute. Samples of entities and prompts are shown in Table 1 and Appendix A. To obtain entity representations we take thed- dimensional hidden state of the entity mentionâs last subword token at layerlto obtain a sample of the activation spaceX, choosinglon a development set as described in Ap- pendix E. We record thenquantities parsed from the LM output to obtainYand fit PLS models with varying numbers of components to find property-encoding directions. 4 R 2 = 1 for a perfect model,R 2 =0 for the mean (constant) model, andR 2 <0 for models that perform worse than the mean model. 5 These include longitude, which is non-monotonic due to the wrap-around at±180 ⊠. In defense of this choice we note that longitude is locally monotonic almost everywhere. 6 We define popular entities as those in the top decile of the rank mean of Wikidata node degree and Wikipedia article length. 4 Preprint 20020 PLS component 1 20 10 0 10 PLS component 2 -3760 1832 1905 1937 1962 2468 (a) Birthyear 20020 PLS component 1 30 20 10 0 10 20 PLS component 2 -2566 1826 1932 1974 2003 2375 (b) Death year 20020 PLS component 1 40 20 0 20 PLS component 2 0 19135 53931 161694 867305 4915489280 (c) Population 20020 PLS component 1 20 10 0 10 20 PLS component 2 -213 30 114 242 493 8849 (d) Elevation 20020 PLS component 1 30 20 10 0 10 20 PLS component 2 -90.00 34.58 40.77 46.24 51.49 90.00 (e) Latitude 4020020 PLS component 1 20 10 0 10 20 30 40 PLS component 2 -178.00 -77.03 -0.35 9.92 30.39 178.43 (f) Longitude Figure 3: Projection onto the top two components of per-property partial least squares regressions reveals monotonic structure in LM representations. We first fit a PLS model on Llama 2 13B entity representations from our training split for each property, project entity representations from the test split, and then plot the resulting 2-d projections. Each dot represents one entity and color saturation represents the value of the corresponding entity attribute. See units for each property in Table 1. PLS regression results for Llama 2 13B representations are shown in Figure 2 and results for additional models in Appendix B. All numeric properties can be predicted well (R 2 â„0.79), with the exception of the elevation property (R 2 =0.43). Across all six properties, PLS identifies small sets of predictive components. For example, a PLS model withk=7 components achieves a goodness of fit ofR 2 =0.91 when predicting birthyear attributes from entity representations. Generally, all LMs appear to encode almost the entirety (95% of maximumR 2 ) of their stored numeric attribute information in two- to six-dimensional subspaces and a large fraction (80% of maximumR 2 ) appears to be encoded in two to three dimensions (see Appendix C). To further illustrate the low dimensionality of numeric property representation, in Figure 3 we plot the projection of entity representations onto the top two components of PLS regres- sions. The plots for each property exhibit monotonic structure. Concomitant with poor regression fit, monotonic structure is least visible for the elevation property. All other plots show clearly visible directions along which attribute values increase, reflecting the good fit of low-dimensional PLS regression models for these properties. 3 Causal Effect of Property-Encoding Directions 3.1 Motivation: Do property-encoding directions affect model output? So far, we have found correlative evidence for the existence of directions in activation space that monotonically encode numeric properties. Concretely, partial least squares regression on entity representations found directions that are predictive of what quantity the LM will express when queried for a particular numeric attribute of a particular entity. However, representation is not a sufficient criterion for computation (Lasri et al., 2022). In our case this means that numeric properties might be encoded in representations without affecting model 5 Preprint output. In order to make the stronger claim that numeric properties are not only encoded monotonically, but that these representations have a monotonic effect on LM output, we now perform interventions to establish causality. Intuitively, we want to find out if making âsmallâ interventions leads to small changes in model output, if âlargeâ interventions lead to large changes, and if the sign of the intervention matches the sign of the change. We now formalize this intuition by adapting the definition of linear representation proposed by Park et al. (2023) and Jiang et al. (2024). Definition 1(Linear representation of numeric properties, adapted from Jiang et al. (2024)). A numeric property is represented linearly if for all pairs of attribute instancesi,jwith quantitiesq i Ìž=q j and their representations â x i , â x j , there exists asteering vector â uso that â x i â â x j âCone( â u), whereCone( â v) = α â v:α>0 is the cone of vector â v, i.e., its positive span. Linearity of representations only requires that representations lie in a cone, but says nothing about their ordering. To model the natural structure of numeric properties, we introduce the constraint that the ordering of quantities is preserved in representation space. Definition 2(Monotonic representation of numeric properties).A numeric property is represented monotonically if it is represented linearly inCone( â u)and for all triples of attribute instancesh,i,jwith quantitiesq h >q i >q j and representations â x h , â x i , â x j the following holds: â x h â â x j =α hj â uand â x i â â x j =α ij â uif and only ifα hj >α ij . The coefficientsα ij andα hj exist since the definition stipulates linear representation. There are many ways to operationalize this definition. One way is to prepare a âsyntheticâ series of monotonic representations inCone( â u)by varyingαand then testing if these representations result in monotonic output changes, which is what we will do now. 3.2 Method: Activation patching Viewing the LM computation graph as causal graph (Meng et al., 2022; McGrath et al., 2023), we intervene on model activations via activation patching (Vig et al., 2020; Wang et al., 2022; Zhang & Nanda, 2024) and observe the effect on model output. Unlike the common activation patching setup in which one replaces activations resulting from one input with activations from a different input, we create patches by editing activations along particular directions, similar to the activation manipulation method of Matsumoto et al. (2022). Specifically, for each of the topKdirections â u k âR d ,kâ[1. .K]found by PLS as described in§2, we prepare patches â p s,k =α s â u k with edit weightsα s and edit step indexsâ[1. .S]. 7 In the view proposed by Park et al. (2023) and Jiang et al. (2024), we can interpret PLS components as potential steering vectors and edit weights as their coefficients. Lacking a principled method for choosing edit weightsα s , we set their range to the minimum and maximum PLS loadings on each propertyâs training split. This choice yields patches covering the full empirical range of activation projections onto direction â u k . After sampling n train =1000 popular entities for each of the six numeric properties we first fit PLS models for each property, then apply activation patches â p s,k to the representations ofn test =100 held-out 8 entities and for each entitiy record the LMâs expressed quantityy s,k . To evaluate monotonicity, i.e., the notion that small (large) edit weightsα s ,kshould have a small (large) effects and that negative (positive) weights should decrease (increase) the expressed quantity y s,k , we quantify the intervention effect via the ranked Spearman correlationÏ(α s,k ;y s,k ). A question left open so far is where activation patching should be performed. While automatic methods for localizing model components and subnetworks of interest have been proposed (Conmy et al., 2023; Kram Ì ar et al., 2024), for simplicity we perform a coarse, non-exhaustive search across layers and token positions based on one numeric property and use the found setting for all experiments (results shown in Appendix E). In addition to this edit locus, we also search for an edit window, whose purpose is to counteract self-repair 7 We chooseS=80 edits steps to balance step size and computational cost. 8 The held-out entities are entities which were not used to fit PLS models. 6 Preprint α s y s,1 y s,2 y s,3 y s,4 y s,5 y s,6 1.00194119551980198020121929 0.90194119551955198420121929 0.80194119551955198420121929 0.70194119551955198019681929 0.60193219551935195819681929 0.50193219401935195819641929 0.40193219301917195819571902 0.30192919301906195819291902 0.20190219021902193419291902 0.10190219021902190219021902 0.00190219021902190219021902 -0.10188719021902190218821902 -0.20188219021902188718821902 -0.30188319021902188718821902 -0.40161919021906188718821901 -0.50161919021906188718821906 -0.60161919021906188718821906 -0.70161919021906188718801906 -0.80188819021902188718801906 -0.90181519021902185818801906 -1.00181519021902185818801906 Ï(α s ,y s,k )0.910.870.720.970.980.39 (a) Birthyear of Karl Popper α s y s,1 1.007.5 billion 0.907.5 billion 0.807.5 billion 0.707.5 billion 0.607.5 billion 0.507.5 billion 0.401.3 billion 0.301.3 billion 0.201.3 billion 0.1010 million 0.0040,000 -0.1040,000 -0.2025,000 -0.3025,000 -0.4020,000 -0.5020,000 -0.6020,000 -0.7012,000 -0.8012,000 -0.9012,000 -1.0012,000 Ï(α s ,y s,k )0.98 (b) Population of Zittau Table 2: The quantityy s,k expressed by a LM changes as a result of directed activation patching along directionkwith (normalized) edit weightα s , withα s =0.00 corresponding to unedited model activations. Warm colors indicate values larger than and cold colors values smaller than the true value, which, if output by the LM, is printed black. Table (a) shows how one-dimensional directed patches along each of the top six âbirthyearâ PLS components change the answer given by Llama 2 13B to the prompt:In what year was Karl Popper born? One word answer only. It is apparent that the most-correlated component (k=1) does not necessarily correspond to the direction in which model behavior exhibits highest monotonicity, which in this case is componentk=5 with a Spearman correlation of 0.98. Table (b) shows the effect of patching along the top âpopulationâ component on Llama 2 13B when prompted:What is the population of Zittau? One word answer only. (McGrath et al., 2023) and iterative inference effects (Rushing & Nanda, 2024). Layer-wise we find that a window of±2 layers around the edit locus is most effective, which is smaller than the±5 layers used in prior work (Meng et al., 2022; Hase et al., 2023a). We also implement a token-wise window (Monea et al., 2024), finding that in addition to the last entity mention token, patching up to two token representations to the left and one token representation to the right works best for the prompts in our experiments. Typically, this token window size covers the entity mention and the main verb or last token of the prompt, depending on the numeric property (see prompts in Appendix A). In summary, we patch activations in a 5-layer window centered on layerl=0.3 and an up-to 4-token window surrounding the last entity mention subword token. To improve output format adherence, we append the instructionOne word answer onlyto all prompts. 3.3 Results: Property-encoding directions have effects and side-effects We are interested in the effects and side effects on model output when patching activations along property-specific directions. Looking at effects first, Table 2 gives examples of how numeric attribute expression changes as a result of directed activation patching. Patching along âbirthyear â directions results in the expression of different years, although the degree of monotonicity, as quantified by Spearman correlationÏ, varies. Patching along the top âpopulationâ direction causes the model to generate a range of outputs that can be interpreted as population sizes, although the largest values are more suited to a planetary than a 7 Preprint 1.00.50.00.51.0 Edit weight (scaled) 0 25 50 75 100 125 150 Model output change (years) (a) Birthyear 1.00.50.00.51.0 Edit weight (scaled) 50 25 0 25 50 75 100 Model output change (years) (b) Death year 1.00.50.00.51.0 Edit weight (scaled) 0 1 2 3 4 Model output change (people) 1e9 (c) Population 1.00.50.00.51.0 Edit weight (scaled) 0 2000 4000 6000 Model output change (meters) (d) Elevation 1.00.50.00.51.0 Edit weight (scaled) 200 0 200 400 600 Model output change (degrees) (e) Latitude 1.00.50.00.51.0 Edit weight (scaled) 0 2000 4000 6000 8000 10000 12000 14000 Model output change (degrees) (f) Longitude Figure 4: Effect of activation patching along property-specific directions across several numeric properties. Each subplot shows the change in the numeric attribute value expressed by Llama 2 13B, as a function of the edit weightα s . Dark red lines indicate means across 100 entities sampled from held-out test sets and bands show standard deviations. municipal scale. The sequence of outputs has rather sudden jumps, e.g., from40,000 (unedited model,α s =0.00) to10 millionafter taking the first step in the âlarger populationâ direction (α s =0.10). The pattern of jumps and plateaus is plausibly connected to several factors such as tokenization effects and the likely high frequency of certain numerals (1.3 billion: population of China at some point in time;7.5 billion: population of Earth, etc.) in the training data, but we leave a detailed investigation to future work. The pattern also indicates that activation space, while apparently monotonic, is not linear in this direction. The intervention also induces a switch from positional notation (40,000) to named numbers (million,billion), which showcases effects beyond single tokens. Moving to a more systematic analysis, we plot mean-aggregated effects of directed activation patching across six numeric properties in Figure 4. We see that there are properties for which directed activation patching has highly monotonic effects, e.g., birthyear (Ï=0.84), elevation (Ï=0.88), or work period start (Ï=0.90), suggesting that these properties have highly monotonic representations. Other properties exhibit a much smaller degree of monotonic editability, e.g., longitude (Ï=0.55) and population (0.65), suggesting that LM representations do not encode these properties as well. Figures for additional models lead to similar to conclusions and are shown in Appendix F. Having observed the effects of our interventions we now turn to analyzing their side effects on the expression of unrelated, non-targeted numeric properties, i.e., properties that were not the target of the intervention. For example, if we fitted a PLS regression to find and patch along âbirthyearâ directions, birthyear is our targeted property and all other properties, such as death year, population, or longitude are non-targeted properties. Using the same directions found in§2, we prompt LMs for non-targeted attributes, perform directed activation patching with weightα s in a direction found for the targeted property and record expressed quantitiesy âČ s,k . To see if non-targeted properties are affected in a similar monotonic fashion as targeted ones, we quantify the side-effect of directed activation patching as the mean Spearman correlationÏ(α s ,y âČ s,k ), aggregated over 100 entities per property. We perform this procedure for all combinations of targeted and non-targeted properties, including three additional properties, and show results in Figure 5. In this figure, entries on the diagonal show the mean effect size for targeted properties and off-diagonal 8 Preprint area birthyear death year elevation inception latitude longitude population work period start Measured property area birthyear death year elevation inception latitude longitude population work period start Targeted property 0.490.670.640.650.580.520.520.520.56 0.650.820.720.530.680.620.440.700.75 0.500.710.730.590.630.410.340.460.70 0.520.610.490.800.560.430.390.460.52 0.230.380.460.410.530.360.340.560.46 0.510.650.630.630.560.630.540.490.58 0.490.470.610.500.470.400.490.490.52 0.560.710.610.590.610.450.490.650.57 0.460.600.600.640.500.440.420.420.81 0.0 0.2 0.4 0.6 0.8 1.0 Edit effect (Spearman correlation) (a) Llama 2 7B area birthyear death year elevation inception latitude longitude population work period start Measured property area birthyear death year elevation inception latitude longitude population work period start Targeted property 0.540.730.690.700.680.680.600.340.67 0.570.840.660.420.640.560.430.360.84 0.620.680.820.440.660.520.540.630.52 0.450.620.520.880.610.430.400.510.59 0.410.670.590.500.700.640.560.440.60 0.570.550.610.580.510.760.350.500.49 0.710.520.540.570.560.570.550.580.58 0.540.760.710.670.700.690.510.650.68 0.410.770.740.490.650.540.480.520.90 0.0 0.2 0.4 0.6 0.8 1.0 Edit effect (Spearman correlation) (b) Llama 2 13B Figure 5: Mean-aggregated effects and side effects when performing activation patching along property-specific directions in activation space. Diagonal entries (top-left to bottom right) show the effect on the targeted property in terms of mean Spearman correlation between edit weightalpha s ,kand expressed quantityy s ,k. For example, patching an entity representation along a âbirthyear â direction results in a corresponding change in the quantity expressed by Llama 2 13B with a correlation strength of 0.84. Off-diagonal entries show the side-effects of activation patching, e.g., âbirthyear â patches affect LM output when queried for an entityâs death year with a correlation strength of 0.68. entries the size of side-effects. For Llama 2 7B, the mean effect size Ì Ï=0.65±0.12 (i.e., the mean of diagonal entries), is not much larger than the mean side-effect size Ì Ï=0.53±0.11 (mean of off-diagonal entries). In contrast, for Llama 2 13B the effect size of Ì Ï=0.85±0.07 is considerably larger than the size of side effects ( Ì Ï=0.58±0.18). A plausible explanation for these results is that in case of Llama 2 7B different properties share a subspace which encodes generic numeric ranges or generic small-large ranges that are translated into quantities depending on context, while the representational space of Llama 2 13B is more akin to a mixture of generic numeric and property-specific, âmore orthogonalâ subspaces. Clearly, more work is needed to verify this hypothesis. The analysis of side-effects is complicated by actual correlations between properties: Birthyear and death year distances are bounded by the human life span, latitude and population are correlated since the Earthâs northern hemisphere is more populous, locations with higher elevation tend to have smaller populations, etc. Consequently, one might argue that, say, editing an entityâs birthyear should also affect LM output when querying the entityâs death year. 4 Related Work Shaped by the locality of physical reality, the locality of human experience (Prystawski et al., 2023) gives rise to distributional patterns of language use. Such patterns include patterns of geographic and temporal coherence (Heinzerling et al., 2017), which reflect spatiotemporal proximity of real-world entities. These patterns can be picked up by statistical models and allow, e.g., to predict geographic information from co-occurrence statistics of cities mentioned in news articles (Louwerse & Zwaan, 2009). Probing static word vector represen- tations for numeric attributes of geopolitical entities, Gupta et al. (2015) obtain good relative rankings, but do not evaluate absolute values nor analyze the geometry of representations. Continuing this line of research, Li Ì etard et al. (2021) probe LM representations for GPS coordinates. Perhaps due to theâby current standardsâsmall scale of the studied LMs, they find only limited success but report that larger models appeared to encode more geographic information. Faisal & Anastasopoulos (2023) measure how well the geographic proximity of countries can be recovered from LM representations but differ from our work in their focus on the impact of politico-cultural factors. 9 Preprint Closest to our work is the analysis of geo-temporal information encoded in Llama 2 rep- resentations by Gurnee & Tegmark (2023). Our work corroborates their finding of linear subspaces of activation space which are predictive of numeric attributes, but is distinct in three important aspects. First, as we show in§2, the subspaces found PCA, as used by Gurnee & Tegmark, are of considerably higher dimensionality (50â100) than the subspaces found by partial least-square regression (2â17). Our finding thus tightens the upper bound on the complexity of numeric property representation in recent LMs. Second, we make explicit and formalize the notion of monotonic representation. Third, our interventions via directed activation patching (§3) found one-dimensional directions with fine-grained effects on the expression of numeric attributes, across all numeric properties and models we analyzed, thereby establishing a causal relationship between monotonic representations and LM behavior. 5 Limitations 5.1 General limitations of representational analysis None of the language models studied in this work are embodied agents or otherwise capable of embodied cognition. Lacking direct sensorimotor grounding (Harnad, 1990; Mollo & Milli ` ere, 2023; Harnad, 2024), LMs cannot directly perceive, let alone precisely measure, the numerical attributes of which we claim to have found monotonic representations. It follows that any such representations are an artifact of distributional patterns in their training data, and that the best one can hope for is isomorphy between model representations and the properties of the real-world entities to which we tie those representations. Leaving the groundedness of representations aside, the idea that concepts, knowledge, or behavior are âencodedâ in neural representations might seem intuitively appealing, but has been strongly critized, on theoretical grounds in the context of biological and artificial neural networks in general (Brette, 2019), and on empirical grounds in the context of pretrained language models in particular (Hase et al., 2023a; Niu et al., 2024). Analysis of LM representations also has well-known limitations. Under the mild assump- tion that there exists a bijection between inputs and their representations, all information extractable from the input, i.e., the natural language prompt, can also be extracted from the LMâs representation of that sequence (Pimentel et al., 2020b). Hence the question to be answered by representational analysis is not whether a feature of interest can be extracted or not, but how easy it is to extract. How to best quantify âease of extractionâ (Pimentel et al., 2020b) is an open question, although methods have been proposed (Pimentel et al., 2020a; Voita & Titov, 2020). 5.2 Specific limitations of the representational analysis conducted in this work The low-dimensional linear subspaces found in this work allow relatively âeasyâ extraction when compared to the nominally high dimensionalities of activation space, but are still higher-dimensional than necessary, since the represented structures (e.g., years, geographic coordinates) are canonically one- to two-dimensional. Furthermore, activation space is nominally high-dimensional but its intrinsic dimension is believed to be much lower (Li et al., 2018; Aghajanyan et al., 2021; Razzhigaev et al., 2024). For example Razzhigaev et al. (2024) provide estimates for the intrinsic dimension of various LMs, ranging from about 10 to 70 dimensions (the models used in our experiments are not covered). If we view a non-linear, non-monotonic representation of full intrinsic dimensionality as the most complex encoding with worst-case ease of extraction, and one- to two-dimensional linear monotonic encodings as the simplest representation with optimal ease of extraction, then the low-dimensional subspaces we found fall somewhere between these bounds. Whether they are low-dimensional relative to the modelsâ intrinsic dimension is currently unknown. Put differently, if the intrinsic dimension of Llama 2 7B turns out to be, say, 10, then finding, a 10-dimensional subspace that encodes all latitude information (see Appendix C) is not surprising, but necessary. 10 Preprint While we found evidence for monotonic representation of numeric properties, it is likely that our causal interventions via activation patching along one-dimensional directions are too simplistic, considering the fact that according to our PLS regression results, numeric properties are encoded in low- but not one-dimensional subspaces. Hence it is possible that a more refined editing method operating on higher-dimensional directions will allow more precise control over LM output. Furthermore, our analysis is limited to popular entities, frequent numeric properties, and English queries, i.e., the combination most likely to be well-represented in the LM training data. 6 Conclusions and Open Questions We used partial least-squares regression to identify low-dimensional subspaces of activa- tion space that are predictive of the quantity an LM expresses when queried for numeric attributes such as an entityâs birthyear. We then performed activation patching along one- dimensional directions in these subspaces and observed corresponding changes in model output. Our results suggest that LMs learn monotonic representations of numeric properties and that these monotonic representations exist in all of the language models we examined. While this work takes a step towards a better understanding of numeric property represen- tation in LMs, it leaves many questions unanswered. Some of these are: âąWhat exactly do the subspaces we found encode? Large side effects of activation patching suggest that there exist directions that encode generic numeric ranges. At the same time, the pronounced difference between large effect size and smaller side- effects we observed for the largest model in our analysis constitutes evidence for the existence of property-specific subspaces. Can we find more specific directions, possibly by adding constraints to minimize side effects into the regression objective, similar to (Meng et al., 2022)? âąIf there are subspaces encoding generic numeric ranges, are changes along directions in these subspaces absolute or relative to typical value ranges? E.g., does a relative large change in birthyears, say 100 years, translate into a relatively large change in elevation, say 1000 meters, or into a change with the same value, i.e., 100 meters? âąIs the quality of numeric attribute representations in any way connected to LM performance on tasks involving the attributes in question? For example, does the model answer queries likeWho was born earlier, Albert Einstein or Karl Popper? correctly if the corresponding entity representations are appropriately positioned along a âbirthyearâ direction, and incorrectly if they are not? Acknowledgements.This work was supported by JST CREST Grant Number JP- MJCR20D2 and JSPS KAKENHI Grant Number 21K17814. References Armen Aghajanyan, Sonal Gupta, and Luke Zettlemoyer. Intrinsic dimensionality ex- plains the effectiveness of language model fine-tuning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.),Proceedings of the 59th Annual Meeting of the Asso- ciation for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 7319â7328, Online, August 2021. As- sociation for Computational Linguistics. doi: 10.18653/v1/2021.acl-long.568. URL https://aclanthology.org/2021.acl-long.568. Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxan- dra Cojocaru, M Ì erouane Debbah, Ì Etienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guil- herme Penedo. The falcon series of open language models, 2023. Emily M. Bender and Alexander Koller. Climbing towards NLU: On meaning, form, and understanding in the age of data. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel 11 Preprint Tetreault (eds.),Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 5185â5198, Online, July 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.463. URLhttps://aclanthology.org/2020.acl-main. 463. Romain Brette. Is coding a relevant metaphor for the brain?Behavioral and Brain Sciences, 42: e215, 2019. Roger C. Conant and W. Ross Ashby. Every good regulator of a system must be a model of that system.International journal of systems science, 1(2):89â97, 1970. Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adri ` a Garriga-Alonso. Towards automated circuit discovery for mechanistic interpretability, 2023. Nicola De Cao, Wilker Aziz, and Ivan Titov. Editing factual knowledge in language models. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 6491â6506, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.522. URLhttps: //aclanthology.org/2021.emnlp-main.522. Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy models of superposition, 2022. Fahim Faisal and Antonios Anastasopoulos. Geographic and geopolitical biases of lan- guage models.In Duygu Ataman (ed.),Proceedings of the 3rd Workshop on Multi- lingual Representation Learning (MRL), p. 139â163, Singapore, December 2023. Asso- ciation for Computational Linguistics. doi: 10.18653/v1/2023.mrl-1.12. URLhttps: //aclanthology.org/2023.mrl-1.12. Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. Dissecting recall of factual associations in auto-regressive language models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.),Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, p. 12216â12235, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.emnlp-main.751. URLhttps://aclanthology.org/ 2023.emnlp-main.751. Nathan Godey, Ì Eric de la Clergerie, and Beno Ë Ä±t Sagot. On the scaling laws of geographical representation in language models, 2024. Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Pad Ì o. Distributional vectors encode referential attributes. In Llu Ì Ä±s M ` arquez, Chris Callison-Burch, and Jian Su (eds.), Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, p. 12â21, Lisbon, Portugal, September 2015. Association for Computational Linguistics. doi: 10.18653/v1/D15-1002. URLhttps://aclanthology.org/D15-1002. Wes Gurnee and Max Tegmark. Language models represent space and time, 2023. Stevan Harnad. The symbol grounding problem.Physica D: Nonlinear Phenomena, 42(1-3): 335â346, 1990. Stevan Harnad. Language writ large: Llms, chatgpt, grounding, meaning and understand- ing, 2024. Charles R. Harris, K. Jarrod Millman, St Ì efan J. van der Walt, Ralf Gommers, Pauli Vir- tanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fern Ì andez del R Ì Ä±o, Mark Wiebe, Pearu Peterson, Pierre G Ì erard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, 12 Preprint Christoph Gohlke, and Travis E. Oliphant. Array programming with NumPy.Na- ture, 585(7825):357â362, September 2020. doi: 10.1038/s41586-020-2649-2. URLhttps: //doi.org/10.1038/s41586-020-2649-2. Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun. Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models, 2023a. Peter Hase, Mona Diab, Asli Celikyilmaz, Xian Li, Zornitsa Kozareva, Veselin Stoyanov, Mohit Bansal, and Srinivasan Iyer. Methods for measuring, updating, and visualizing factual beliefs in language models. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, p. 2714â2731, Dubrovnik, Croatia, May 2023b. Association for Computational Linguistics. URLhttps://aclanthology.org/ 2023.eacl-main.199. Benjamin Heinzerling and Kentaro Inui. Language models as knowledge bases: On entity representations, storage capacity, and paraphrased queries. In Paola Merlo, Jorg Tiede- mann, and Reut Tsarfaty (eds.),Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, p. 1772â1791, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.153. URLhttps://aclanthology.org/2021.eacl-main.153. Benjamin Heinzerling, Michael Strube, and Chin-Yew Lin. Trust, but verify! better entity linking through automatic verification. In Mirella Lapata, Phil Blunsom, and Alexander Koller (eds.),Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 1, Long Papers, p. 828â838, Valencia, Spain, April 2017. Association for Computational Linguistics. URLhttps://aclanthology.org/E17-1078. Evan Hernandez, Belinda Z. Li, and Jacob Andreas. Inspecting and editing knowledge representations in language models. InArxiv, 2023. URLhttps://arxiv.org/abs/2304. 00740. J. D. Hunter. Matplotlib: A 2d graphics environment.Computing in Science & Engineering, 9 (3):90â95, 2007. doi: 10.1109/MCSE.2007.55. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L Ì elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timoth Ì e Lacroix, and William El Sayed. Mistral 7b, 2023. Yibo Jiang, Goutham Rajendran, Pradeep Ravikumar, Bryon Aragam, and Victor Veitch. On the origins of linear representations in large language models, 2024. Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423â 438, 2020. doi: 10.1162/tacla00324. URLhttps://aclanthology.org/2020.tacl-1.28. Zhengbao Jiang, Jun Araki, Haibo Ding, and Graham Neubig. How can we know when language models know? on the calibration of language models for question answering. Transactions of the Association for Computational Linguistics, 9:962â977, 2021. doi: 10.1162/ tacla00407. URLhttps://aclanthology.org/2021.tacl-1.57. Nora Kassner, Philipp Dufter, and Hinrich Sch Ì utze. Multilingual LAMA: Investigating knowledge in multilingual pretrained language models. In Paola Merlo, Jorg Tiedemann, and Reut Tsarfaty (eds.),Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, p. 3250â3258, Online, April 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.eacl-main.284. URL https://aclanthology.org/2021.eacl-main.284. J Ì anos Kram Ì ar, Tom Lieberum, Rohin Shah, and Neel Nanda. Atp*: An efficient and scalable method for localizing llm behaviour to components, 2024. 13 Preprint Karim Lasri, Tiago Pimentel, Alessandro Lenci, Thierry Poibeau, and Ryan Cotterell. Prob- ing for the usage of grammatical number. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio (eds.),Proceedings of the 60th Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), p. 8818â8831, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.603. URL https://aclanthology.org/2022.acl-long.603. Harvey Lederman and Kyle Mahowald. Are language models more like libraries or like librarians? bibliotechnism, the novel reference problem, and the attitudes of llms, 2024. Chunyuan Li, Heerad Farkhoor, Rosanne Liu, and Jason Yosinski. Measuring the intrinsic dimension of objective landscapes.arXiv preprint arXiv:1804.08838, 2018. Bastien Li Ì etard, Mostafa Abdou, and Anders SĂžgaard. Do language models know the way to Rome? In Jasmijn Bastings, Yonatan Belinkov, Emmanuel Dupoux, Mario Giu- lianelli, Dieuwke Hupkes, Yuval Pinter, and Hassan Sajjad (eds.),Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, p. 510â517, Punta Cana, Dominican Republic, November 2021. Association for Computational Lin- guistics. doi: 10.18653/v1/2021.blackboxnlp-1.40. URLhttps://aclanthology.org/2021. blackboxnlp-1.40. Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams. Towards understanding grokking: An effective theory of representation learn- ing. In S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (eds.), Advances in Neural Information Processing Systems, volume 35, p. 34651â34663. Curran Associates, Inc., 2022. URLhttps://proceedings.neurips.c/paperfiles/paper/2022/ file/dfc310e81992d2e4cedc09ac47eff13e-Paper-Conference.pdf. Max M. Louwerse and Rolf A. Zwaan. Language encodes geographical information.Cogni- tive Science, 33(1):51â73, 2009. doi: https://doi.org/10.1111/j.1551-6709.2008.01003.x. Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets, 2023. Yuta Matsumoto, Benjamin Heinzerling, Masashi Yoshikawa, and Kentaro Inui. Tracing and manipulating intermediate values in neural math problem solvers. In Deborah Ferreira, Marco Valentino, Andre Freitas, Sean Welleck, and Moritz Schubotz (eds.),Proceedings of the 1st Workshop on Mathematical Natural Language Processing (MathNLP), p. 1â6, Abu Dhabi, United Arab Emirates (Hybrid), December 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.mathnlp-1.1. URLhttps://aclanthology.org/2022. mathnlp-1.1. Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg. The hydra effect: Emergent self-repair in language model computations, 2023. Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. Locating and editing factual associations in GPT.Advances in Neural Information Processing Systems, 36, 2022. Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. A mechanism for solving relational tasks in transformer language models, 2023. Eric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn, and Christopher D. Manning. Fast model editing at scale.CoRR, 2021. URLhttps://arxiv.org/pdf/2110.11309.pdf. Dimitri Coelho Mollo and Rapha Ì el Milli ` ere. The vector grounding problem, 2023. Giovanni Monea, Maxime Peyrard, Martin Josifoski, Vishrav Chaudhary, Jason Eisner, Emre Kıcıman, Hamid Palangi, Barun Patra, and Robert West. A glitch in the matrix? locating and detecting language model grounding with fakepedia, 2024. 14 Preprint Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. In Yonatan Belinkov, Sophie Hao, Jaap Jumelet, Najoung Kim, Arya McCarthy, and Hosein Mohebbi (eds.),Proceedings of the 6th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP, p. 16â30, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/ 2023.blackboxnlp-1.2. URLhttps://aclanthology.org/2023.blackboxnlp-1.2. Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn. What does the knowledge neuron thesis have to do with knowledge? InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=2HJRwwbV3G. The Pandas development team. pandas-dev/pandas: Pandas, February 2020. URLhttps: //doi.org/10.5281/zenodo.3509134. Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models, 2023. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. Pytorch: An im- perative style, high-performance deep learning library.Advances in neural information processing systems, 32, 2019. Karl Pearson. On lines and planes of closest fit to systems of points in space.The London, Edinburgh, and Dublin Philosophical Magazine and Journal of Science, 2(11):559â572, 1901. doi: 10.1080/14786440109462720. URLhttps://doi.org/10.1080/14786440109462720. F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in Python.Journal of Machine Learning Research, 12:2825â2830, 2011. Fabio Petroni, Tim Rockt Ì aschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan (eds.),Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), p. 2463â2473, Hong Kong, China, November 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1250. URL https://aclanthology.org/D19-1250. Tiago Pimentel, Naomi Saphra, Adina Williams, and Ryan Cotterell. Pareto probing: Trading off accuracy for complexity.In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.),Proceedings of the 2020 Conference on Empirical Methods in Nat- ural Language Processing (EMNLP), p. 3138â3153, Online, November 2020a. Associ- ation for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.254. URL https://aclanthology.org/2020.emnlp-main.254. Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. Information-theoretic probing for linguistic structure. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.),Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 4609â4622, Online, July 2020b. Association for Computational Linguistics. doi: 10.18653/v1/2020.acl-main.420. URL https://aclanthology.org/2020.acl-main.420. Ben Prystawski, Michael Y. Li, and Noah D. Goodman. Why think step by step? reasoning emerges from the locality of experience, 2023. Anton Razzhigaev, Matvey Mikhalchuk, Elizaveta Goncharova, Ivan Oseledets, Denis Dimitrov, and Andrey Kuznetsov. The shape of learning: Anisotropy and intrinsic dimensions in transformer-based models, 2024. Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? In Bonnie Webber, Trevor Cohn, Yulan He, 15 Preprint and Yang Liu (eds.),Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 5418â5426, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.437. URLhttps:// aclanthology.org/2020.emnlp-main.437. Cody Rushing and Neel Nanda. Explorations of self-repair in language models, 2024. Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. Au- toPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.),Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 4222â4235, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.346. URLhttps://aclanthology.org/2020.emnlp-main. 346. Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. Linear representa- tions of sentiment in large language models, 2023. Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom. Llama 2: Open foundation and fine-tuned chat models, 2023. Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. Investigating gender bias in language models using causal mediation analysis.Advances in neural information processing systems, 33:12388â12401, 2020. Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cour- napeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St Ì efan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, An- drew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, Ì Ilhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Ant Ë onio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors. SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python.Nature Methods, 17:261â272, 2020. doi: 10.1038/s41592-019-0686-2. Elena Voita and Ivan Titov. Information-theoretic probing with minimum description length. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (eds.),Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 183â196, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/ 2020.emnlp-main.14. URLhttps://aclanthology.org/2020.emnlp-main.14. Denny Vrande Ë ci Ì c and Markus Kr Ì otzsch. Wikidata: a free collaborative knowledgebase. Communications of the ACM, 57(10):78â85, 2014. Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small, 2022. Michael L. Waskom. seaborn: statistical data visualization.Journal of Open Source Software, 6 (60):3021, 2021. doi: 10.21105/joss.03021. URLhttps://doi.org/10.21105/joss.03021. 16 Preprint Herman Wold. Estimation of principal components and related models by iterative least squares.Multivariate analysis, p. 391â420, 1966. Svante Wold, Michael Sj Ì ostr Ì om, and Lennart Eriksson. Pls-regression: a basic tool of chemometrics.Chemometrics and intelligent laboratory systems, 58(2):109â130, 2001. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the-art natural language processing. In Qun Liu and David Schlangen (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, p. 38â45, Online, October 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-demos.6. URLhttps://aclanthology.org/ 2020.emnlp-demos.6. Paul Youssef, Osman Koras ̧, Meijie Li, J Ì org Schl Ì otterer, and Christin Seifert. Give me the facts! a survey on factual knowledge probing in pre-trained language models. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.),Findings of the Association for Computational Linguistics: EMNLP 2023, p. 15588â15605, Singapore, December 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-emnlp.1043. URLhttps: //aclanthology.org/2023.findings-emnlp.1043. Fred Zhang and Neel Nanda. Towards best practices of activation patching in language models: Metrics and methods, 2024. Zexuan Zhong, Dan Friedman, and Danqi Chen. Factual probing is [MASK]: Learning vs. learning to recall. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tur, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou (eds.),Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 5017â5033, Online, June 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.naacl-main.398. URLhttps://aclanthology.org/2021.naacl-main.398. 17 Preprint A Data sample PropertyProp. IDEntityEntity IDPromptValueUnit birthyearP569Nina FochQ235632In what year was Nina Foch born?1924annum birthyearP569Geoffrey HolderQ945691In what year was Geoffrey Holder born?1930annum birthyearP569Harriette L. ChandlerQ5664432In what year was Harriette L. Chandler born?1937annum birthyearP569Gabriel Garc Ì Ä±a M Ì arquezQ5878In what year was Gabriel Garc Ì Ä±a M Ì arquez born? 1927annum birthyearP569Norman Schwarzkopf Jr.Q310188In what year was Norman Schwarzkopf Jr. born? 1934annum birthyearP569Paul de VosQ2610964In what year was Paul de Vos born?1590annum birthyearP569Nicolas CarnotQ181685In what year was Nicolas Carnot born?1796annum birthyearP569Steve HarveyQ2347009In what year was Steve Harvey born?1957annum birthyearP569Tommy LawtonQ726272In what year was Tommy Lawton born?1919annum birthyearP569Hans von B Ì ulowQ155540In what year was Hans von B Ì ulow born?1830annum death yearP570Johannes R. BecherQ58057In what year did Johannes R. Becher die?1958annum death yearP570Friedrich Georg Wilhelm von Struve Q57164In what year did Friedrich Georg Wilhelm von Struve die? 1864annum death yearP570Pierre BoulezQ156193In what year did Pierre Boulez die?2016annum death yearP570Giovanni da PalestrinaQ179277In what year did Giovanni da Palestrina die?1594annum death yearP570Abdurrauf FitratQ317907In what year did Abdurrauf Fitrat die?1938annum death yearP570Lucian FreudQ154594In what year did Lucian Freud die?2011annum death yearP570Akseli Gallen-KallelaQ170068In what year did Akseli Gallen-Kallela die?1931annum death yearP570SpockQ16341In what year did Spock die?2263annum death yearP570William OrpenQ922483In what year did William Orpen die?1931annum death yearP570Carlos Santiago M Ì eridaQ1043100 In what year did Carlos Santiago M Ì erida die?1984annum populationP1082AkhisarQ209905What is the population of Akhisar?1730261 populationP1082TripuraQ1363What is the population of Tripura?36659581 populationP1082AlbertQ30940What is the population of Albert?99301 populationP1082High WycombeQ64116What is the population of High Wycombe?1202561 populationP1082Pl Ì onQ497060What is the population of Pl Ì on?89141 populationP1082Republika SrpskaQ11196What is the population of Republika Srpska?12284231 populationP1082LebaneseQ2606511What is the population of Lebanese?80000001 populationP1082GeraardsbergenQ499532What is the population of Geraardsbergen?334031 populationP1082Gorz Ì ow WielkopolskiQ104731What is the population of Gorz Ì ow Wielkopol- ski? 1242951 populationP1082HarranQ199547What is the population of Harran?476061 evelationP2044SondrioQ6274How high is Sondrio?360metre evelationP2044Rio BrancoQ171612How high is Rio Branco?158metre evelationP2044DemminQ50960How high is Demmin?8metre evelationP2044CetinjeQ173338How high is Cetinje?650metre evelationP2044Highland ParkQ576671How high is Highland Park?503metre evelationP2044GozoQ170488How high is Gozo?195metre evelationP2044Saint-Jean-de- Maurienne Q208860How high is Saint-Jean-de-Maurienne?566metre evelationP2044ButteQ467664How high is Butte?1688metre evelationP2044CottbusQ3214How high is Cottbus?76metre evelationP2044Mahilio Ì u RegionQ189822How high is Mahilio Ì u Region?191metre longitudeP625.longKorean EmpireQ28233What is the longitude of Korean Empire?126.98degree longitudeP625.longPine BluffQ80012What is the longitude of Pine Bluff?-92.00degree longitudeP625.longTegernseeQ260130What is the longitude of Tegernsee?11.76degree longitudeP625.longOld C Ì ollnQ269622What is the longitude of Old C Ì olln?13.40degree longitudeP625.longCambridgeQ49111What is the longitude of Cambridge?-71.11degree longitudeP625.longStrynQ5223What is the longitude of Stryn?6.86degree longitudeP625.longCiudad Real ProvinceQ54932What is the longitude of Ciudad Real Province? -4.00degree longitudeP625.longSanta CatarinaQ41115What is the longitude of Santa Catarina?-50.49degree longitudeP625.longWake Forest UniversityQ392667What is the longitude of Wake Forest Univer- sity? -80.28degree longitudeP625.longWest LothianQ204940What is the longitude of West Lothian?-3.50degree latitudeP625.latK Ì usnachtQ69216What is the latitude of K Ì usnacht?47.32degree latitudeP625.latMount Jerome CemeteryQ917854What is the latitude of Mount Jerome Ceme- tery? 53.32degree latitudeP625.latDayton Childrenâs Hospi- tal Q5243510What is the latitude of Dayton Childrenâs Hos- pital? 39.77degree latitudeP625.latLe Flore CountyQ495944What is the latitude of Le Flore County?34.90degree latitudeP625.latCzechoslovakiaQ33946What is the latitude of Czechoslovakia?50.08degree latitudeP625.latPembroke CollegeQ956501What is the latitude of Pembroke College?52.20degree latitudeP625.latHaywardQ491114What is the latitude of Hayward?37.67degree latitudeP625.latBanaskantha districtQ806125What is the latitude of Banaskantha district?24.17degree latitudeP625.latCorbeil-EssonnesQ208812What is the latitude of Corbeil-Essonnes?48.61degree latitudeP625.latElbasanQ114257What is the latitude of Elbasan?41.11degree Table 3: Random sample of the entities used in our experiments, along with corresponding numeric attributes and prompts. Entities, their English labels, and numeric attributes for each property are extracted from an April 2022 dump of Wikidata (wikidata-20220421-all). In many cases Wikidata contains multiple values for a given numeric attribute, e.g., reflecting chronological change such as the population of a city, or owing to conflicting sources. In such cases we take the mode of the distribution as gold value. We also filter out quantities with non-standard units, such as elevations measured in feet. 18 Preprint B Regression on entity representations: Additional figures 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (c) Birthyear 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (d) Death year 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (e) Population 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (f) Elevation 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (g) Latitude 01020304050 #components 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (h) Longitude Figure 6: Regression curves for Llama 2 7B. See explanation in Figure 2. 01020304050 #components 0.4 0.2 0.0 0.2 0.4 0.6 0.8 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (a) Birthyear 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (b) Death year 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (c) Population 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (d) Elevation 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (e) Latitude 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (f) Longitude Figure 7: Regression curves for Falcon 7B. See explanation in Figure 2. 19 Preprint 01020304050 #components 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (a) Birthyear 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (b) Death year 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (c) Population 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (d) Elevation 01020304050 #components 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (e) Latitude 01020304050 #components 1.0 0.5 0.0 0.5 1.0 Goodness of fit ( R 2 ) Method PLS PCA PLS (shuffled labels) PCA (shuffled labels) PLS (random reprs.) PCA (random reprs.) (f) Longitude Figure 8: Regression curves for Mistral 7B. See explanation in Figure 2. C Regression on entity representations: Additional analysis PropertyModelR 2 C maxR 2 C â„0.95R 2 C â„0.90R 2 C â„0.80R 2 C â„0.70R 2 C â„0.60R 2 C â„0.50R 2 birthyear (P569)Falcon 7B0.754222111 birthyear (P569)Llama 2 13B0.917432221 birthyear (P569)Llama 2 7B0.9011643221 birthyear (P569)Mistral 7B0.894322211 death year (P570)Falcon 7B0.612222111 death year (P570)Llama 2 13B0.8412432211 death year (P570)Llama 2 7B0.8211443211 death year (P570)Mistral 7B0.804332211 latitude (P625.lat)Falcon 7B0.676333222 latitude (P625.lat)Llama 2 13B0.8210543322 latitude (P625.lat)Llama 2 7B0.8310532222 latitude (P625.lat)Mistral 7B0.799433222 longitude (P625.long)Falcon 7B0.747533222 longitude (P625.long)Llama 2 13B0.7917653322 longitude (P625.long)Llama 2 7B0.839533222 longitude (P625.long)Mistral 7B0.786533221 population (P1082)Falcon 7B0.674332111 population (P1082)Llama 2 13B0.795442211 population (P1082)Llama 2 7B0.735432211 population (P1082)Mistral 7B0.765422111 evelation (P2044)Falcon 7B0.232222211 evelation (P2044)Llama 2 13B0.433222221 evelation (P2044)Llama 2 7B0.372222211 evelation (P2044)Mistral 7B0.413322211 Table 4: Number of partial least squares regression componentsC [ T ] required for a given goodness of fitT, found using the experimental setup described in§2. For example, the C â„0.95R 2 column shows the number of components required to reach 95 percent of the maximum goodness of fit for the respective property and model. From this column we can read that, e.g., two components of Falcon 7Bâs activation space are sufficient to reach 95 percent of the maximum goodness of fit when predicting the birthyear of entities, indicating that this property is almost entirely encoded in a two-dimensional subspace of this modelâs activation space. 20 Preprint D PLS projections of entity representations: Additional figures 20020 PLS component 1 20 10 0 10 PLS component 2 -3760 1832 1905 1937 1962 2468 (a) Birthyear 20020 PLS component 1 150 125 100 75 50 25 0 PLS component 2 -2566 1826 1932 1974 2003 2375 (b) Death year 20020 PLS component 1 40 30 20 10 0 10 20 30 PLS component 2 0 19135 53931 161694 867305 4915489280 (c) Population 20020 PLS component 1 20 10 0 10 PLS component 2 -213 30 114 242 493 8849 (d) Elevation 050100 PLS component 1 120 100 80 60 40 20 0 20 PLS component 2 -90.00 34.58 40.77 46.24 51.49 90.00 (e) Latitude 050100 PLS component 1 60 40 20 0 20 PLS component 2 -178.00 -77.03 -0.35 9.92 30.39 178.43 (f) Longitude Figure 9: PLS projections of Llama 2 7B entity representations. See explanation in Figure 3. 4020020 PLS component 1 30 20 10 0 10 PLS component 2 -3760 1832 1905 1937 1962 2468 (a) Birthyear 20020 PLS component 1 30 20 10 0 10 20 PLS component 2 -2566 1826 1932 1974 2003 2375 (b) Death year 20020 PLS component 1 30 20 10 0 10 20 30 PLS component 2 0 19135 53931 161694 867305 4915489280 (c) Population 1001020 PLS component 1 15 10 5 0 5 10 15 PLS component 2 -213 30 114 242 493 8849 (d) Elevation 20020 PLS component 1 30 20 10 0 10 20 PLS component 2 -90.00 34.58 40.77 46.24 51.49 90.00 (e) Latitude 200 PLS component 1 20 10 0 10 20 30 PLS component 2 -178.00 -77.03 -0.35 9.92 30.39 178.43 (f) Longitude Figure 10: PLS projections of Falcon 7B entity representations. See explanation in Figure 3. 21 Preprint 20020 PLS component 1 20 10 0 10 PLS component 2 -3760 1832 1905 1937 1962 2468 (a) Birthyear 20020 PLS component 1 20 10 0 10 20 PLS component 2 -2566 1826 1932 1974 2003 2375 (b) Death year 20020 PLS component 1 40 20 0 20 PLS component 2 0 19135 53931 161694 867305 4915489280 (c) Population 100102030 PLS component 1 20 10 0 10 20 PLS component 2 -213 30 114 242 493 8849 (d) Elevation 201001020 PLS component 1 30 20 10 0 10 20 PLS component 2 -90.00 34.58 40.77 46.24 51.49 90.00 (e) Latitude 4020020 PLS component 1 20 10 0 10 20 30 40 PLS component 2 -178.00 -77.03 -0.35 9.92 30.39 178.43 (f) Longitude Figure 11: PLS projections of Mistral 7B entity representations. See explanation in Figure 3. E Choice of probing and edit locus In what year was [mention last] born ? 1.00 0.80 0.60 0.40 0.20 0.00 Layer (relative) .00.00.00 .89.46.00 .94.96.87 .99.99.89 .99.97.90 .93.97.97 0.0 0.2 0.4 0.6 0.8 1.0 Correlation of edit strength and model output change Figure 12: Results of a cursory search for the best probing and edit locus, using Llama 2 7B. Varying token position and layer, we edit the hidden state at this locus as described in§3 and record the Spearman correlation between edit strength and the change in the quantity (here: birthyear) expressed by the model. Correlation is highest (0.99) in the region between layers 0.2 and 0.4 and the last subword token of the entity mention and the following token. Based on this, we choose the last mention token and the middle point at layerl=0.3 as locus for the regression experiments in§2 and activation patching experiments in§3, across all numeric properties and LMs, but acknowledge that a more exhaustive search would likely find better probing and edit loci. 22 Preprint F Edit curves for additional language models 1.00.50.00.51.0 Edit weight (scaled) 0 20 40 60 80 100 120 140 Model output change (years) (a) Birthyear 1.00.50.00.51.0 Edit weight (scaled) 500 400 300 200 100 0 100 Model output change (years) (b) Death year 1.00.50.00.51.0 Edit weight (scaled) 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Model output change (people) 1e9 (c) Population 1.00.50.00.51.0 Edit weight (scaled) 5000 0 5000 10000 15000 Model output change (meters) (d) Elevation 1.00.50.00.51.0 Edit weight (scaled) 0 2000 4000 6000 8000 10000 Model output change (degrees) (e) Latitude 1.00.50.00.51.0 Edit weight (scaled) 0 2000 4000 6000 Model output change (degrees) (f) Longitude Figure 13: Effect of activation patching along property-specific directions across several numeric properties with Llama 2 7B. See explanation in Figure 4. 1.00.50.00.51.0 Edit weight (scaled) 150 100 50 0 50 100 Model output change (years) (a) Birthyear 1.00.50.00.51.0 Edit weight (scaled) 600 500 400 300 200 100 0 Model output change (years) (b) Death year 1.00.50.00.51.0 Edit weight (scaled) 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Model output change (people) 1e7 (c) Population 1.00.50.00.51.0 Edit weight (scaled) 40 20 0 20 40 60 80 Model output change (meters) (d) Elevation 1.00.50.00.51.0 Edit weight (scaled) 40 30 20 10 0 10 Model output change (degrees) (e) Latitude 1.00.50.00.51.0 Edit weight (scaled) 30 20 10 0 10 20 30 40 Model output change (degrees) (f) Longitude Figure 14: Effect of activation patching along property-specific directions across several numeric properties with Falcon 7B (Almazrouei et al., 2023). See explanation in Figure 4. 23 Preprint 1.00.50.00.51.0 Edit weight (scaled) 50 25 0 25 50 75 100 125 Model output change (years) (a) Birthyear 1.00.50.00.51.0 Edit weight (scaled) 1000 800 600 400 200 0 200 Model output change (years) (b) Death year 1.00.50.00.51.0 Edit weight (scaled) 0 2 4 6 Model output change (people) 1e9 (c) Population 1.00.50.00.51.0 Edit weight (scaled) 10000 7500 5000 2500 0 2500 5000 7500 Model output change (meters) (d) Elevation 1.00.50.00.51.0 Edit weight (scaled) 4 2 0 2 4 Model output change (degrees) (e) Latitude 1.00.50.00.51.0 Edit weight (scaled) 0 1000 2000 3000 4000 5000 Model output change (degrees) (f) Longitude Figure 15: Effect of activation patching along property-specific directions across several numeric properties with Mistral 7B (Jiang et al., 2023). See explanation in Figure 4. G Software The following is a list of the main libraries used in this work: âą Numpy (Harris et al., 2020) âą Scikit-learn (Pedregosa et al., 2011) âą Pytorch (Paszke et al., 2019) âą Transformers (Wolf et al., 2020) âą seaborn (Waskom, 2021) âą Matplotlib (Hunter, 2007) âą SciPy (Virtanen et al., 2020) âą Pandas (Pandas development team, 2020) We thank all authors and the open source community in general for creating and maintaining publicly and freely available software. 24