Paper deep dive
Steering Large Language Models using Conceptors: Improving Addition-Based Activation Engineering
Joris Postmus, Steven Abreu
Models: GPT-J-6B, GPT-NeoX-20B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/12/2026, 7:21:33 PM
Summary
This paper introduces 'conceptors' as a novel activation engineering technique for steering Large Language Models (LLMs). Unlike traditional additive steering, which uses a single vector, conceptors act as soft projection matrices that capture complex activation patterns as ellipsoidal regions. The authors demonstrate that conceptors outperform additive methods on various function-steering tasks in GPT-J and GPT-NeoX, and show that Boolean operations on conceptors allow for more effective combination of steering goals.
Entities (5)
Relation Signals (3)
Conceptors ā appliedto ā GPT-J
confidence 100% Ā· we apply this mechanism to function vectors on GPT-NeoX and GPT-J
Conceptors ā appliedto ā GPT-NeoX
confidence 100% Ā· we apply this mechanism to function vectors on GPT-NeoX and GPT-J
Conceptors ā outperforms ā Additive Steering
confidence 95% Ā· Our experiments demonstrate that conceptors outperform traditional methods across multiple steering tasks.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models have transformed AI, yet reliably controlling their outputs remains a challenge. This paper explores activation engineering, where outputs of pre-trained LLMs are controlled by manipulating their activations at inference time. Unlike traditional methods using a single steering vector, we introduce conceptors - mathematical constructs that represent sets of activation vectors as ellipsoidal regions. Conceptors act as soft projection matrices and offer more precise control over complex activation patterns. Our experiments demonstrate that conceptors outperform traditional methods across multiple steering tasks. We further use Boolean operations on conceptors for combined steering goals that empirically outperform additively combining steering vectors on a set of tasks. These results highlight conceptors as a promising tool for more effective steering of LLMs. Our code is available on this http URL.
Tags
Links
Trouble viewing inline? Open PDF directly ā
Full Text
67,787 characters extracted from source content.
Expand or collapse full text
Steering Large Language Models using Conceptors: Improving Addition-Based Activation Engineering Joris Postmus University of Groningen Groningen, Netherlands j.postmus@student.rug.nl Steven Abreu University of Groningen Groningen, Netherlands s.abreu@rug.nl Abstract Large language models have transformed AI, yet reliably controlling their outputs remains a challenge. This paper explores activation engineering, where outputs of pre-trained LLMs are controlled by manipulating their activations at inference time. Unlike traditional methods using a single steering vector, we introduce conceptorsāmathematical constructs that represent sets of activation vectors as ellipsoidal regions. Conceptors act as soft projection matrices and offer more precise control over complex activation patterns. Our experiments demonstrate that conceptors outperform traditional methods across multiple steering tasks. We further use Boolean operations on conceptors for combined steering goals that empirically outperform additively combining steering vectors on a set of tasks. These results highlight conceptors as a promising tool for more effective steering of LLMs. Our code is available on github.com/jorispos/conceptorsteering. 1 Introduction Large language models (LLMs) have rapidly advanced AI capabilities [1], but their potential to spread misinformation [2], reinforce biases [3], and develop harmful behaviors [4] highlights the urgent need for methods to understand and control their outputs. Various methods, including reinforcement learning from human feedback (RLHF) [5], supervised fine-tuning [6], and prompt engineering [7], have been proposed to steer LLM outputs toward desired patterns. However, RLHF and fine-tuning are computationally expensive and struggle with generalization [8,9], while prompt engineering often produces inconsistent results [10]. Activation engineering[11,12] has recently been proposed as a new steering method which works by directly modifying the modelās activations at inference time without changing the modelās parameters and without expensive optimization. A steering vector that represents desired behavior can be computed directly or (more commonly) contrastively from positive and negative examples [13]. However, finding contrastive prompts to identify complex patterns is not always possible and, more importantly, the performance of activation addition for steering is not reliable [12]. This paper introduces an alternative to the predominant approach for steering LLMs using activation engineering. Instead of averaging or subtracting a set of activation vectors to form a steering vector, we use the cached activations to compute aconceptor[14], which we refer to as a āsteering matrixā. Instead of manipulating the LLMās activations using vector addition, the activations are (softly) projected using a matrix-vector multiplication with the steering matrix. We contribute the following: (1) we introduce a novel application of conceptors [14] as steering mechanisms for LLMs, (2) we apply this mechanism to function vectors [15] on GPT-NeoX and GPT-J, and (3) we show how a Boolean algebra on conceptors [16] can be used for combining steering targets on GPT-J. 38th Conference on Neural Information Processing Systems (NeurIPS 2024), MINT Workshop. arXiv:2410.16314v4 [cs.NE] 12 May 2025 Figure 1: Illustration showing the basic geometric difference between additive and conceptor steering using a set of activations for the antonym task. Additive steering acts as a translation of the activation vectors by a fixed steering vector. Conceptor steering acts as a (soft) projection onto a target ellipsoid. 2 Background Adding steering vectors to the residual stream has been used to control the output of LLMs across various domains [12,13,17]. The use case that will mainly be focused on here are the findings from the paper by Toddet al.[15]. Their work showed that a steering vector can be extracted from the residual stream that captures the activation space of an input-output function (e.g. a function that takes a word and returns its antonym). This steering vector can then be added to the residual stream at inference time to steer the model toward performing the captured function. See Figure 6 in Appendix A.1.1 for an illustration of function vector tasks. Their baseline method works as follows. First, a set of in-context learning (ICL) promptsP f that demonstrate a particular taskf(the execution of an input-output function) are compiled. Then, for each promptp f i āP f (e.g.,p antonym 1 =hot:cold,old:), the final tokenās activationsh ā (p f i ) are cached at a specific layerāfrom the residual streamh. The cached activation vectors are then averaged into the steering vector Ģ h f ā for taskfat layerā: Ģ h f ā = 1 |P f | X p f i āP f h ā (p f i )(1) To steer the model towards performing this function, the function (steering) vector Ģ h f ā can be added (without additional re-normalization) to the residual stream at layerāwhen the model would be completing a prompt containing a previously unseen input: h ā² ā =β add Ģ h f ā +h ā (2) whereh ā² ā is the steered activation andβ add >0is a hyperparameter. The performance of additive steering can further be improved by a technique calledmean-centering[18], see Appendix A.2.1. 3 Conceptors as Steering Matrices Conceptors can broadly be defined as a neuro-computational mechanism designed to encapsulate and manipulate the state space of neural activations [14]. A conceptor matrixCis a positive semi-definite matrix that captures the principal directions and variances of a set of neural activation vectors. This structure can be visualized as a high-dimensional ellipsoid that describes the overall shape and spread of the activationsā āunderlying patternā, or state space region. See Figure 2 for a visual illustration. Because conceptors are computed from the cloud of activation vectors and encode the correlations between activations, conceptors can better capture the activation space of complex patterns compared to simple point representations, which discard information about correlations. This difference is illustrated geometrically in Figure 1. Conceptors have been used to control pattern-generating RNNs effectively across various behaviors [16], prevent catastrophic forgetting and enhance continual learning in feedforward networks [19], 2 Conceptor C 1 Conceptor C 2 Conceptor C 3 Figure 2: Illustration of three conceptors as ellipsoids that capture the state space region of different sets of neural activations in 3D space (black points). Reproduced from Jaeger [14]. remove bias subspaces in LLMs like BERT and GPT [20], and distill linguistic abstractions into knowledge graphs from contextual embeddings [21, 22]. One way to formalize the conceptor matrixC, is through an optimization that minimizes the recon- struction error while using a regularization term. The objective function to be minimized is: min C ā„XāXCā„ 2 F +α ā2 ā„Cā„ 2 F whereXis a matrix of neural activation vectors (stacked as rows),ā„Ā·ā„ F is the Frobenius norm, and αis the regularization parameter also referred to as the conceptorāsaperture. This aperture parameter αbalances the trade-off between accurately representing the activation pattern and maintaining a generalized representation. The closed-form solution to this problem is given by: C(R,α) =R R+α ā2 I ā1 withR= X T X n (3) wherenis the number of samples, andIis the identity matrix of the same dimensionality asR. The eigenvaluesμ i of the conceptor matrixCare defined as: μ i =              Ī» i Ī» i +α ā2 for0< Ī» i <1and0< α <ā 0for0< Ī» i <1andα= 0 1for0< Ī» i <1andα=ā 0forĪ» i = 0and0ā¤Ī±ā¤ā 1forĪ» i = 1and0ā¤Ī±ā¤ā whereĪ» i represents the eigenvalues of the correlation matrixR. These eigenvaluesμ i fall within the interval[0,1]and are influenced by the aperture parameterα. Whenαis large, the eigenvalues μ i approach 1 andCapproaches the identity matrix, causing the conceptor to allow for more signal components to pass through the projection of the states with the conceptor matrixCx. Conversely, whenαis small, the eigenvaluesμ i approach 0, causing the conceptor to allow for less variability. In the extreme case ofαā0, the conceptor tends to the zero mapping. We can use the conceptor for steering by collecting activationsh ā (p f i )intoXand then compute the associated conceptorC f ā using Equation 3, and finally steer new hidden activationsh ā with: h ā² ā =β c C f ā h ā (4) whereh ā² ā is the steered activation andβ c >0is a hyperparameter. We can think of this as a āsoft projectionā. A projection matrix has eigenvalues that are either zero or unity, but the conceptor matrix has āsoftā eigenvalues between zero and unity. Thus, the conceptor āsoftly projectsā the activation vectorh ā toward the pattern represented byC f ā by scaling its components according to the patternsā principal directions. 3.1 Boolean Operations on Conceptors We can combine multiple steering matrices using the conceptor Boolean operations as defined by Jaeger [16]. We begin with the OR operation on conceptors, which can be interpreted as merging the 3 data from which each conceptor is computed by adding the covariance matrices on whichC 1 andC 2 were computed. Given thatC 1 was computed with the covariance matrixR 1 andC 2 was computed with the covariance matrixR 2 , the conceptor that is computed on the sum of the two covariance matricesR 1 +R 2 is defined asC 1 āØC 2 : C 1 āØC 2 = (R 1 +R 2 )(R 1 +R 2 +α ā2 I) ā1 (5) C 1 āØC 2 = I+ C 1 (IāC 1 ) ā1 +C 2 (IāC 2 ) ā1 ā1 ā1 (6) The NOT operation on a conceptorCis defined as the conceptor¬Cthat is computed on a covariance matrixR ā1 that is the inverse of the original covariance matrixRfor conceptorC. Intuitively,¬C can be interpreted as the conceptor that arises from data that which co-vary inversely to the data giving rise toC: ¬C=R ā1 (R ā1 +α ā2 I) ā1 (7) ¬C=IāC(8) For our experiments, we use the AND operation which can now be obtained from the NOT and OR operations using de Morganās lawaā§b=¬(aāØb)such that the conceptorC 1 ā§C 2 is computed using the correlation matrix(R ā1 1 +R ā1 2 ) ā1 . This leads to: C 1 ā§C 2 = (R ā1 1 +R ā1 2 ) ā1 (R ā1 1 +R ā1 2 ) ā1 +α ā2 I ā1 (9) C 1 ā§C 2 = (C ā1 1 +C ā1 2 āI) ā1 (10) 3.2 Computational Complexity of Conceptor Steering The cost of computing a conceptor steering matrix is dominated by the matrix inversion and matrix- matrix multiplication of the activation correlation matrixR=X T /n(see Equation 3). This correlation matrix is anĆn-dimensional matrix wherenis the dimension of the activation vectors (typically <4096 for the model sizes we presented, or up to 8192 for larger models such as Llama-2- 70B), so the complexity of the conceptor computation isO(n 3 ). This computation is done entirely offline and the cost is amortized over all future applications of the steering method. The final conceptorCāR nĆn takesO(n 2 )memory ā the same amount as a weight matrix acting on the activation vectors. For 32-bit floating point numbers, this amounts to 17MB forn= 2048, 67MB for n= 4096, or 268MB forn= 8192. During inference, conceptor steering adds an extra matrix-vector multiplicationCxwith the activation vectorx. However, the additional memory and inference cost for applying the conceptor can be eliminated by fusing the conceptor with the succeeding weight matrices for the query, key and value weight matrices. This is equivalent to replacing the existing weight matrixW x with the conceptor- fused weight matrixW C x =W x C . This fusing of operations is standard practice when optimizing networks for inference. We note that there may be an overhead cost for switching the conceptor steering on and off which amounts to the cost of changing the networkās computational graph during inference. We believe this overhead to be negligible during auto-regressive generation on a single data sample, but it must be considered when using batch sizes larger than one. 4 Experiments For our experiments, we will use EleutherAIās GPT-J 6B and GPT-NeoX 20B models, as done in previous works on activation steering [15,18]. For all experiments, we find optimal hyperparameters for each steering method at every layer. The details of our grid search forαandβ c for conceptor-based steering andβ add for additive steering can be found in Appendix A.1.2. 4.1 Function Steering We compare conceptor-based and additive steering mechanisms on their ability to steer a given model towards correctly executing a set of functions. We test both methods on GPT-J with 6B parameters and GPT-NeoX with 20B parameters. For each function, the described experiment will be repeated five times with different random seeds, and all reported results are averaged across across these five 4 runs. The examples of the input-output functions come from the dataset by Toddet al.[15]. We use the following subset of five functions [18]: antonyms (e.g. goodābad), present-past (e.g. goāwent), English-French (e.g. helloābonjour), singular-plural (e.g. mouseāmice), country-capital (e.g. NetherlandsāAmsterdam), and capitalize (e.g. wordāWord). To ensure comparability of our results, we follow [15] as closely as possible. For more details, see Appendix A.1.1. Performance of different steering mechanisms across different functions Accuracy Accuracy Figure 3: Comparison of the accuracy on all six function tasks for conceptor-based steering against additive steering across all layers for GPT-J and GPT-NeoX. For explanation, see main text. The results in Figure 3 show that conceptor-based steering outperforms additive steering (the baseline method reported in Ref. [15]) for every task on both tested models. In line with previous findings [15, 18], steering is most effective across layers 9-16 for GPT-J and layers 10-30 for GPT-NeoX. Table 1 and Figure 4 show that mean-centering (as outlined in Appendix A.2.1) provides a small improvement for both addition-based and conceptor-based steering. Mean-centering improves the performance of additive steering by as much as 2x (on the country-capital task). For conceptor-based steering the improvements of mean-centering are relatively smaller ā at most 5% on the country- capital task. Conceptor-based steering outperforms additive steering on all tasks, even comparing additive steering with mean-centering against conceptor-based steering without mean-centering. Table 1: The effect of mean centering on conceptor-based and addition-based steering on the GPT-J (6B) model, across simple function vector tasks. Results show the best performance across all hyperparameters and across all layers. antonymscapitalizecountry-capitalenglish-frenchpresent-past Addition20.54%93.16%32.04%18.88%69.66% Addition (MC)31.20%95.00%63.90%34.32%83.32% Conceptor52.14%96.68%81.62%59.02%91.56% Conceptor (MC)52.82%96.26%85.32%61.32%91.88% 0510152025 layer 0% 25% 50% 75% 100% GPT-J (6B) antonyms 0510152025 layer capitalize 0510152025 layer country-capital 0510152025 layer english-french 0510152025 layer present-past 0510152025 layer singular-plural conceptor conceptor (MC) addition addition (MC) Performance of mean-centering on different steering mechanisms Figure 4: The effect of mean centering on conceptor-based and addition-based steering on the GPT-J (6B) model across all layers, computed on five different function vector tasks (% accuracy). The line shows the best average performance across five runs for the best hyperparameters for the given layer. 5 4.2 Steering Composite Functions We further conducted experiments where two conceptors, each representing one of three different functions, were combined using the AND operator. The input-output example dataset for this function was generated using GPT-4o. To present the baseline for how well non-combined steering mechanisms perform, we show results for the conceptorC 1,2 and the steering vector Ģ h 1,2 ā that were each computed on the compound function directly. We then combine the conceptors computed on the individual functionsC 1 andC 2 using the AND operation asC 1 ā§C 2 , and we combine the steering vectors Ģ h 1 ā and Ģ h 2 ā using their arithmetic mean 1 2 ( Ģ h 1 ā + Ģ h 2 ā ). Accuracy Figure 5: Performance of additive steering and conceptor steering on composite functions. For explanation of the figure caption, see text. Dashed lines represent the ābaselineā where the steering mechanism is computed on the composite task. Solid lines show task arithmetic. Figure 5 shows the performance of all compared methods across all layers of the GPT-J model. In line with the results from Section 4.1, the conceptor baseline outperformed the additive baseline on all three tasks. The AND-combined conceptor outperformed the mean-combined steering vectors. On one of the three tasks, english-french & antonyms, the AND-combined conceptor even outperforms the additive baseline. 5 Conclusion In our experiments, conceptor-based steering generally outperformed addition-based methods. Further research should be conducted to assess the mechanismsā impact on the modelās overall capabilities, the performance on more complex behaviors/tasks, and the scalability to larger models. A limitation of conceptors is their reliance on more data points to build accurate representations. Additionally, the inherent mathematical structure and additional required computations makes it more computationally expensive compared to simple addition-based methods. However, while more expensive than addition-based approaches, they are still much cheaper than alternatives like RLHF and fine-tuning. Conceptors also introduce a new hyperparameter, the apertureα, that may require tuning for optimal performance. In our experiments, we found a single aperture value,α= 0.1, yields the best performance across all experiments 1 , but this finding must be verified for new models and steering tasks. Despite these challenges, conceptor-based steering methods could offer a more precise and effective way to steer LLMs compared to traditional addition-based methods, proposing a fundamental shift in what is possible with activation engineering. Our experiments on conceptor-based steering further suggest that region-based representations may allow for more flexible and nuanced steering compared to point-based representations. The proposed method could have significant positive implications for debiasing models, aligning models with human values, and overall AI safety. 1 More precisely, the aperture valueα= 0.1is within10%of the best-performing aperture value across all experiments and models. In most experiments, it is the single best value. See Appendix A.3 for more details. 6 References [1]Bo Xu and M. Poo. Large language models and brain-inspired general intelligence.National Science Review, 10, 2023. doi: 10.1093/nsr/nwad267. [2]Yikang Pan, Liangming Pan, Wenhu Chen, Preslav Nakov, Min-Yen Kan, and William Yang Wang. On the risk of misinformation pollution with large language models. InThe 2023 Conference on Empirical Methods in Natural Language Processing, 2023. URLhttps: //openreview.net/forum?id=voBhcwDyPt. [3] Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and fairness in large language models: A survey.arXiv, abs/2309.00770, 2024. URLhttps://arxiv.org/abs/ 2309.00770. [4]Toby Shevlane, Sebastian Farquhar, Ben Garfinkel, Mary Phuong, Jess Whittlestone, Jade Leung, Daniel Kokotajlo, Nahema Marchal, Markus Anderljung, Noam Kolt, Lewis Ho, Divya Siddarth, Shahar Avin, Will Hawkins, Been Kim, Iason Gabriel, Vijay Bolina, Jack Clark, Yoshua Bengio, Paul Christiano, and Allan Dafoe. Model evaluation for extreme risks, 2023. URLhttps://arxiv.org/abs/2305.15324. [5] Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feed- back. InProceedings of the 36th International Conference on Neural Information Processing Systems, NIPS ā22, Red Hook, NY, USA, 2024. Curran Associates Inc. ISBN 9781713871088. [6]Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, editors,Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171ā4186, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics. doi: 10.18653/v1/N19-1423. URLhttps://aclanthology. org/N19-1423. [7]Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Comput. Surv., 55(9), jan 2023. ISSN 0360-0300. doi: 10.1145/3560815. URLhttps://doi.org/10.1145/3560815. [8]LĆ©on Bottou, Frank E. Curtis, and Jorge Nocedal. Optimization methods for large-scale machine learning.arXiv, abs/1606.04838, 2018. URLhttps://arxiv.org/abs/1606.04838. [9]Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan ManĆ©. Concrete problems in AI safety.arXiv, abs/1606.06565, 2016. URLhttps://arxiv.org/ abs/1606.06565. [10]Banghao Chen, Zhaofeng Zhang, Nicolas Langrenāe, and Shengxin Zhu. Unleashing the potential of prompt engineering in large language models: a comprehensive review.ArXiv, abs/2310.14735, 2023. doi: 10.48550/arXiv.2310.14735. [11]Kenneth Li, Oam Patel, Fernanda ViĆ©gas, Hanspeter Pfister, and Martin Wattenberg. Inference- time intervention: Eliciting truthful answers from a language model. InThirty-seventh Con- ference on Neural Information Processing Systems, 2023. URLhttps://openreview.net/ forum?id=aLLuYpn83y. [12] Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Activation addition: Steering language models without optimization, 2024. URLhttps://arxiv.org/abs/2308.10248. [13] Nina Rimsky, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Turner. Steering llama 2 via contrastive activation addition. In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15504ā15522, Bangkok, Thailand, August 2024. Association for Computational Linguistics. URLhttps://aclanthology. org/2024.acl-long.828. 7 [14]Herbert Jaeger. Conceptors: an easy introduction.arXiv, abs/1406.2671, 2014. URLhttps: //arxiv.org/abs/1406.2671. [15]Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau. Function vectors in large language models. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=AwyxtyMwaG. [16]Herbert Jaeger. Controlling recurrent neural networks by conceptors.arXiv, abs/1403.3369, 2017. URLhttps://arxiv.org/abs/1403.3369. [17]Teun van der Weij, Massimo Poesio, and Nandi Schoots. Extending activation steering to broad skills and multiple behaviours, 2024. URLhttps://arxiv.org/abs/2403.05767. [18]Ole Jorgensen, Dylan Cope, Nandi Schoots, and Murray Shanahan. Improving activation steering in language models with mean-centring.arXiv, abs/2312.03813, 2023. URLhttps: //arxiv.org/abs/2312.03813. [19]Owen He.Continual lifelong learning in neural systems: overcoming catastrophic forgetting and transferring knowledge for future learning. PhD thesis, University of Groningen, 2023. [20]Li S. Yifei, Lyle Ungar, and JoĆ£o Sedoc. Conceptor-aided debiasing of large language models. InThe 2023 Conference on Empirical Methods in Natural Language Processing, 2023. URL https://openreview.net/forum?id=M6BJfQ9oup. [21]Jesper Kuiper. Using conceptors to extract abstraction hierarchies from corpora of natural text: Combatting word polysemy using word sense disambiguation techniques. Masterās thesis / essay, University of Groningen, Groningen, Netherlands, January 2024. [22]Paul Bricman. Nested state clouds: Distilling knowledge graphs from contextual embeddings. Bachelorās Project Thesis, University of Groningen, Supervisors: Prof. Dr. Herbert Jaeger, Dr. Jacolien van Rij-Tange, July 2022. URLhttps://fse.studenttheses.ub.rug.nl/ 27840/. [23]Aaron Gokaslan, Vanya Cohen, Ellie Pavlick, and Stefanie Tellex. Openwebtext corpus. http://Skylion007.github.io/OpenWebTextCorpus, 2019. 8 A Appendix A.1 Experimental Details All experiments were run on NVIDIA GPUs. The GPT-NeoX model was run on one NVIDIA RTX A6000 with 48GB of VRAM, and the GPT-J model was run on one NVIDIA GeForce RTX 4090 with 24GB of VRAM. Each hyperparameter sweep took less than 18 hours of compute time per model and per task. A.1.1 Function Steering All the experimental configurations (number of experiments, number of ICL prompts and examples per prompt, accuracy metric, etc.) were, unless mentioned otherwise, adopted from Ref. [15] to ensure comparability of results. For each experiment, to generate the 4 steering mechanisms, we first compileN p = 100(ICL) prompts that demonstrate the respective input-output function. The prompts are formed by randomly samplingN= 10input-output pairs from the function pairs dataset. If for a specific function, the dataset contains less thanN p ĆN= 1000input-output examples, this sampling is done with replacement. For each promptp f i , the last input-output pair has the output stripped, resulting in the format: p f i = āx 1 :y 1 ,x 2 :y 2 ,...,x Nā1 :y Nā1 ,x N : ā wherexrepresents the input tokens of a randomly sampled (input, output) pair,yrepresents the corresponding output tokens,Nrepresents the number of sampled input-output pairs, andiā 1,...,N p . A very simple example whereN p = 3andN= 3can be seen in Figure 6a. old:young, vanish:appear, dark: awake: asleep, future:past, joy: top: bottom, tall:short, accept: h f l (a) Extraction of the antonym function (steering) vector Ģ h f ā at layerlusing 3 ICL prompts. simple: + = complex h f l encode: + = decode h f l (b) Antonym steering vector in 2 zero-shot contexts. Figure 6: Visualization of how an antonym function (steering) vector can be extracted and applied. Example from [15] . Formally, for each functionfāFin our set of in-context learning (ICL) tasks, we have compiled a setP f of ICL promptsp f i āP f . Each promptp f i is a sequence of tokens withNinput-output exemplar pairs(x,y)that demonstrate the functionfmapping betweenxandy. For each experiment, we generateN p such prompts. Now that the ICL prompts have been generated, we need to extract the relevant activations. Toddet al. [15] showed that the neural representations of the functions are encoded in the activation vector of the last token (":") of the prompt, right before the transformer would auto-regressively start generating the output token(s). Moreover, the point in the residual streamhat which the functions were most strongly encoded was shown to be at the beginning of layersL=9,...,16, right before MHA and FFN [15]. Formally, for each functionfāFand each promptp f i āP f , the activation vectorsh f ā (p f i )are extracted from the residual streamhat each relevant layerlāLfrom the last tokenās (":") activation vector. For each functionfāFand each layerlāL, we now haveN p cached activation vectorsh f ā (p f i ) aimed to encode the neural representation offat layerl. Using this, we can generate the layer-specific steering mechanisms for each function as follows: ⢠The standard additive steering mechanism Ģ h f ā is generated by averaging over all the cached activation vectorsh f ā (p f i )respectively as described in Equation 1. 9 ā¢The additive steering mechanism with mean-centering Ģ h f,mc ā is computed by taking the previously generated steering mechanism Ģ h f ā and subtractingμ train as described in Equation 11. ā¢The regular conceptor steering mechanismCis computed as described in Equation 3 using the aperture valueα reg . The correlation matrixRis computed asR= X T X N p , whereXis the matrix of allh f ā (p f i )stacked activation vectors. ā¢The mean-centered conceptor steering mechanismC mc is computed with some minor adjustments. The matrixXis formed by subtractingμ train from the activation vectorsh f ā (p f i ) before stacking them. This results in an adjusted correlation matrixR: R= (Xāμ train ) T (Xāμ train ) N p The mean-centered conceptor matrixC mc can then be calculated as described in Equation 3 using the aperture valueα mc and the adjusted correlation matrixR. To test the performance of the generated steering mechanisms, new sets ofN t = 1000input-output pairs are randomly sampled from the function pairs dataset for each experiment. This is done with replacement for functions where the dataset contains less thanN t pairs. An input promptp t is formatted asp t = āx: ā, wherexis a tokenized input from an input-output pair. The tokenized outputyfrom the pair is left out fromp t as it will be used to test the accuracy of the steering mechanisms. For each experiment, we now haveN t test input promptsp t . To test the accuracy of the steering mechanisms, we apply the layer-specific steering mechanisms on independent forward passes and record their subsequent output. This means that for our experimental configuration, across the functionsfāF, the 5 experiments, the 4 steering mechanisms (excluding the baseline), theN t number of test prompts, and the number of layerslāL, there will be 6Ć5Ć4Ć1000Ć8 = 960,000forward passes, each with a steering intervention. Each steering intervention will consist of a layer-specific steering mechanism modifying the residual streamhat the mechanismsā respective layerl. This modification can be defined as transforming the unmodified residual stream activation vectorh ā into the steered activation vectorh ā² ā . The steering mechanismsā modification can be described as follows: ā¢For the standard additive steering mechanism, the averaged activation vector Ģ h f ā is multiplied by the injection coefficientβ add and added to the residual stream activation vectorh ā : h ā² ā =β add Ģ h f ā +h ā ⢠For the additive steering mechanism with mean-centering, the mean-centered average activation vector Ģ h f,mc ā is multiplied by the injection coefficientβ add and added to the residual stream activation vectorh ā : h ā² ā =β add Ģ h f,mc ā +h ā ā¢For the regular conceptor steering mechanism, the residual stream activation vectorh ā is multiplied using the conceptor matrixCand further multiplied with the rescaling coefficient β c : h ā² ā =β c C h ā ā¢For the mean-centered conceptor steering mechanism, the residual stream activation vector h ā is first adjusted by subtractingμ train . This adjusted vector is then multiplied with the mean-centered conceptor matrixC mc and further multiplied with the rescaling coefficient β c . Finally,μ train is added back to the result: h ā² ā =β c C mc (h ā āμ train ) +μ train ⢠For the baseline condition, no modifications are made to the residual stream. h ā² ā =h ā 10 After the respective modifications have been made to the residual stream, the forward passes will continue as usual. At the end of each forward pass, the final logits are converted into probabilities using a softmax, and the token with the highest probability is selected. This means that at the end of one experiment, we haveN t single-token outputs for each layer-specific steering mechanism. These tokens can now be compared with the first token of outputythat corresponds with the inputxof the initial promptp t . Based on how many of theN t outputs were correctly identified, a top-1 accuracy is calculated for each layer-specific steering mechanism. This experiment is repeated 5 times for each functionfāFto account for variability caused by the random sampling for the generation of the steering mechanisms and test sets. A.1.2 Hyperparameter optimization The performance of the steering mechanisms in the function vector experiments was optimized through a grid search over all hyperparameters. Firstly, we try steering at each layer of the model. For conceptor-based steering, we do a grid search for the aperture valueαwith possi- ble values from0.001,0.0125,0.05,0.1and the scaling coefficientβ c with possible values from 0.5,1.0,2.0,3.0,4.0,5.0. For additive steering, we run a grid search over the scaling coeffi- cientβ add with possible values from0.5,1.0,1.5,2.0,2.5,3.0,4.0,5.0. The results from these hyperparameter sweeps are shown in Appendix A.3 A.2 Additional Experimental Results A.2.1 Mean centering An important improvement for additive steering is a technique calledmean-centering, put forward by Jorgensenet al.[18]. This method enhances the effectiveness of steering vectors by reducing the inherent bias present in the activation space of LLMs. Activation vectors in LLMs tend to be anisotropic, meaning that they are not evenly distributed around the origin, but are instead offset in a consistent direction. This can negatively impact the steering vectorās performance as the bias vectorb representing this offset, does not encode any specific task-related information, diluting the steering vectorās effectiveness. First, the steering vector Ģ h f ā for a specific functionfis computed by averaging the activations at layer āon a set of ICL prompts demonstrating the input-output functionP f (as defined in Equation 1). Ģ h f ā now encodes the task-specific behavior but may still be affected by biases in the modelās overall activation space. Mean-centering attempts to mitigate this by subtracting the mean activation of a broader dataset that represents the general activation space of the model. This is done by computing the mean activation vectorμ train over a large, representative set of promptsD train from the modelās training data. The mean activation vectorμ train was calculated using the same procedure described by Jorgensenet al.[18]: A subset from the dataset used to train GPT-2 was compiled [23]. The subset was constructed by storing all entries from the foldersurlsf_subset01-1/dataandurlsf_subset01-182/data. After this, only entries that contained less than 500 tokens (using the GPT-2 Tokenizer) were retained. This resulted in 210 entries from which the final 10 were removed, leaving a dataset of 200 entries. The mean activation vectorμ train was then computed by averaging the activations over this dataset. Implementing the mean-centering performance enhancement for steering toward the execution of functions can be done as follows: Ģ h f,mc ā = Ģ h f ā āμ train withμ train = 1 |D train | X dāD train h ā (d)(11) where Ģ h f ā is as described in Equation 1, andD train is the dataset for which the mean-centered vector μ train is computed. This refinement leads to a steering vector that can more effectively guide the model toward the specific task and has been shown to have a positive impact on the overall steering effectiveness [18]. 11 A.3 Hyperparameter Sweep Results In the following section, we present results from the hyperparameter optimization described in Appendix A.1.2, in order to assess the sensitivity of both steering mechanisms (additive and conceptor- based) to the hyperparameters. A.3.1 Conceptor Steering Figure 7 shows that the optimal choice of aperture and beta parameters for the conceptor steering mechanism is constant atα= 0.05andβ C = 2.0across all tasks for the GPT-J model (for the layer with the maximum performance). Figure 8 shows similar behavior for the GPT-NeoX model, although the optimal beta parameter isβ= 1and the optimal aperture parameter changes toα= 0.0125 for the country-capital task, andα= 0.1for the english-french task, andα= 0.05for all other tasks. This shows that hyperparameter choices are robust for conceptor steering, but still benefit from task-specific and model-specific optimization. We further show the performance of conceptor-based steering across all layers and different beta values (taking the best-performing aperture value) for the GPT-J model in Figure 9 and for the GPT-NeoX model in Figure 10. For the GPT-J model, the best-performing layers are typically layers 12-14 with some variability (present-past being a few layers later at 14-17, and capitalize working well across layers 9-19). For the GPT-NeoX model, conceptor steering reaches (near-)maximum performance at layer 15 across all tasks, with layer 15 being at around one third of the depth of the model. Figures 11 and 12 show the performance of conceptor-based steering across all layers and different aperture values (taking the best-performing beta value) for the GPT-J model and the GPT-NeoX model, respectively, and show a similar pattern as described above. aperture 0.5 1.0 2.0 3.0 4.0 5.0 beta 0.00.10.20.2 0.00.30.40.4 0.00.40.50.5 0.00.50.50.5 0.00.50.40.4 0.00.50.40.3 antonyms aperture beta 0.00.80.90.9 0.00.90.90.9 0.00.91.01.0 0.01.01.01.0 0.01.01.01.0 0.21.01.01.0 capitalize aperture beta 0.00.50.80.8 0.00.80.80.8 0.00.80.80.8 0.00.80.70.7 0.00.80.60.5 0.00.70.50.5 country-capital 0.0010.01250.050.1 aperture 0.5 1.0 2.0 3.0 4.0 5.0 beta 0.00.10.30.3 0.00.30.50.5 0.00.50.60.6 0.00.60.50.4 0.00.60.40.3 0.00.60.30.3 english-french 0.0010.01250.050.1 aperture beta 0.00.60.70.8 0.00.80.90.9 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 present-past 0.0010.01250.050.1 aperture beta 0.10.60.80.8 0.10.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.90.80.8 singular-plural 0.1 0.2 0.3 0.4 0.5 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 0.1 0.2 0.3 0.4 0.5 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 Conceptor steering of GPT-J (6B) Figure 7: Performance results of the grid search across aperture and beta values (for the optimal layer) for the GPT-J (6B) model, using conceptor-based steering. A.3.2 Additive Steering Additive steering only has two hyperparameters that were being optimized: the layer on which steering was done, and the beta value that determines the āsteering strengthā. Figure 13 shows the performance of additive steering on the GPT-J model across all layers and beta values. Similarly to the results of conceptor-based steering, additive steering works best across layers 9-14 with peak performance always between layers 12-14. The best-performing beta values are 2.0, 3.0, and 4.0, although 2.0 is sufficient to reach peak performance for all tasks. Figure 14 shows the performance of additive steering on the GPT-NeoX model across all layers and beta values. Similar to the best- performing conceptor-based steering hyperparameters, additive steering works best on layers 12-16. The optimal beta values are 1.5 and 2.0. 12 aperture 0.5 1.0 2.0 3.0 4.0 5.0 beta 0.00.10.10.1 0.00.60.60.6 0.50.50.50.5 0.60.40.30.3 0.60.20.20.2 0.60.20.20.2 antonyms aperture beta 0.00.60.70.6 0.51.01.01.0 1.01.01.01.0 1.01.01.01.0 1.00.90.90.9 1.00.60.60.6 capitalize aperture beta 0.00.70.80.8 0.31.00.90.9 0.90.90.90.9 0.90.70.70.6 1.00.40.30.3 0.90.20.20.2 country-capital 0.0010.01250.050.1 aperture 0.5 1.0 2.0 3.0 4.0 5.0 beta 0.00.10.10.1 0.00.60.60.6 0.50.30.30.3 0.60.20.20.2 0.50.20.20.2 0.40.10.10.1 english-french 0.0010.01250.050.1 aperture beta 0.00.70.60.5 0.31.01.01.0 0.90.90.90.9 1.00.90.90.8 1.00.70.70.7 0.90.50.40.4 present-past 0.0010.01250.050.1 aperture beta 0.00.70.60.6 0.51.01.01.0 1.00.90.90.9 1.00.90.90.9 1.00.70.70.7 1.00.50.40.4 singular-plural 0.1 0.2 0.3 0.4 0.5 0.6 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 0.1 0.2 0.3 0.4 0.5 0.6 0.2 0.4 0.6 0.8 0.2 0.4 0.6 0.8 Conceptor steering of GPT-NeoX (20B) Figure 8: Performance results of the grid search across aperture and beta values (for the optimal layer) for the GPT-NeoX (20B) model, using conceptor-based steering. 13 beta 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.20.30.40.30.3 0.00.20.40.40.40.4 0.00.20.30.40.40.4 0.10.40.50.50.50.5 0.00.40.50.50.50.5 0.00.40.50.40.40.4 0.00.40.40.40.30.3 0.20.40.30.30.20.2 0.10.10.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.00.0 0.10.10.10.00.00.0 0.10.10.10.00.00.0 0.10.10.00.00.00.0 0.10.10.00.00.00.0 0.10.10.00.00.00.0 0.10.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 antonyms beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.10.20.20.3 0.00.00.10.10.10.2 0.00.10.40.70.80.9 0.00.00.40.70.80.9 0.00.00.50.80.90.9 0.00.10.70.90.91.0 0.00.50.91.01.01.0 0.00.81.01.01.01.0 0.00.91.01.00.90.9 0.30.91.01.00.90.9 0.10.90.90.90.90.8 0.10.90.90.80.70.6 0.80.90.80.70.60.4 0.90.90.70.60.40.4 0.90.80.60.40.30.3 0.70.70.50.40.30.2 0.70.60.40.30.20.2 0.40.30.30.20.20.2 0.40.20.20.10.10.2 0.10.10.10.10.10.1 0.00.00.00.00.00.0 capitalize beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.10.1 0.00.00.00.00.10.1 0.00.00.10.10.10.1 0.00.10.10.10.10.1 0.00.10.10.10.20.2 0.00.10.50.60.60.5 0.00.10.60.60.50.6 0.20.70.70.70.70.7 0.00.80.80.80.80.7 0.20.80.80.80.70.7 0.50.80.80.70.60.5 0.70.80.70.60.40.2 0.80.70.50.40.20.1 0.80.60.40.20.10.1 0.50.50.20.10.10.1 0.50.40.20.10.10.0 0.30.30.20.10.10.0 0.20.20.10.10.10.1 0.20.20.10.10.10.1 0.20.20.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.10.1 country-capital 0.51.02.03.04.05.0 beta 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.1 0.00.00.10.20.20.2 0.00.10.10.20.20.2 0.10.10.20.20.20.2 0.00.10.20.30.30.2 0.00.10.40.40.40.4 0.10.30.40.40.50.4 0.10.50.60.60.60.6 0.10.40.40.50.40.3 0.00.50.50.40.30.2 0.30.40.40.20.10.0 0.20.20.10.00.00.0 0.10.10.00.00.00.0 0.10.10.00.00.00.0 0.20.10.00.00.00.0 0.20.10.00.00.00.0 0.10.10.00.00.00.0 0.10.10.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 english-french 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.10.20.3 0.00.00.20.30.40.4 0.00.30.40.40.40.4 0.10.20.40.40.40.4 0.10.20.40.40.40.4 0.00.20.40.40.40.4 0.00.20.40.40.40.4 0.40.40.40.40.40.5 0.00.40.50.50.60.6 0.00.80.90.90.90.9 0.00.80.90.90.90.9 0.30.90.90.90.90.8 0.80.90.80.80.70.6 0.70.80.70.50.40.2 0.80.80.50.30.20.1 0.70.60.30.20.10.1 0.60.50.20.10.10.1 0.30.40.20.10.10.1 0.30.30.10.10.10.1 0.20.10.10.10.10.0 0.10.10.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 present-past 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.20.40.5 0.00.10.30.50.60.6 0.10.40.60.60.60.6 0.10.10.60.60.60.6 0.10.40.60.60.60.6 0.00.30.60.60.70.7 0.00.30.60.60.70.7 0.50.60.70.70.70.7 0.00.70.90.80.90.9 0.00.90.90.90.90.9 0.00.90.90.90.90.8 0.80.90.90.90.80.8 0.80.90.80.70.50.4 0.70.80.70.50.40.3 0.80.80.60.40.30.2 0.70.70.40.30.30.2 0.70.60.40.30.20.2 0.60.50.30.20.20.2 0.60.40.30.20.20.2 0.30.20.20.20.20.2 0.20.20.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.10.1 singular-plural 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.2 0.4 0.6 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 Conceptor steering of GPT-J (6B) Figure 9: Performance results of the grid search across layers and beta values (for the optimal aperture value) for the GPT-J (6B) model, using conceptor-based steering. 14 beta 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.10.00.00.0 0.00.40.40.20.20.3 0.10.50.50.30.40.5 0.00.60.50.50.60.6 0.00.60.50.60.60.6 0.10.60.50.60.60.5 0.00.60.50.60.50.4 0.00.50.50.40.30.2 0.10.50.50.40.20.1 0.10.40.30.20.10.1 0.10.20.20.20.10.1 0.00.20.20.10.10.1 0.00.20.20.10.10.0 0.10.20.10.10.10.0 0.00.10.10.10.00.0 0.00.10.10.10.00.0 0.00.10.10.10.00.0 0.00.10.10.00.00.0 0.00.10.10.00.00.0 0.00.10.10.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 antonyms beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.1 0.00.91.01.00.90.9 0.00.81.01.00.80.9 0.11.01.01.01.01.0 0.11.01.01.01.01.0 0.51.01.01.01.01.0 0.61.01.01.01.01.0 0.21.01.01.01.00.9 0.51.01.01.01.00.9 0.71.01.01.00.90.9 0.61.01.01.00.90.9 0.31.00.90.90.90.7 0.31.00.90.90.80.6 0.51.00.90.80.60.3 0.31.00.90.70.40.2 0.50.90.70.40.20.1 0.30.90.70.30.20.1 0.10.80.50.20.10.1 0.30.70.50.20.20.1 0.00.60.30.20.20.1 0.00.50.20.20.20.1 0.00.40.20.20.20.1 0.00.30.20.20.10.1 0.00.30.20.20.20.1 0.00.20.20.10.10.1 0.00.20.10.10.10.1 0.00.20.10.10.10.1 0.00.10.10.10.10.1 0.00.10.10.10.10.1 0.00.10.10.10.10.1 0.00.10.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.10.1 0.10.00.10.10.10.1 0.00.00.10.00.00.0 capitalize beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.10.0 0.00.00.00.10.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.1 0.10.30.30.10.10.2 0.00.80.60.20.20.5 0.00.90.70.30.80.9 0.20.90.80.80.90.9 0.21.00.90.91.00.9 0.81.00.80.90.90.9 0.60.90.90.90.90.8 0.40.90.90.90.70.6 0.40.90.90.80.80.6 0.80.90.90.80.70.5 0.60.90.90.70.50.3 0.60.90.90.70.30.2 0.50.80.70.40.20.1 0.50.80.70.30.20.1 0.70.80.50.20.10.1 0.60.70.30.20.10.1 0.50.60.30.20.10.1 0.60.50.30.20.10.1 0.50.40.20.20.10.1 0.50.40.20.10.10.1 0.40.40.20.10.10.1 0.30.40.20.10.10.1 0.30.30.20.10.10.1 0.30.30.20.10.10.0 0.30.30.20.10.10.1 0.20.20.20.10.10.1 0.20.20.20.20.20.1 0.20.20.20.20.20.2 0.20.20.20.20.20.2 0.20.20.20.20.20.1 0.10.10.10.10.10.1 0.10.20.20.20.20.2 0.20.20.10.10.10.1 0.10.10.10.10.10.1 country-capital 0.51.02.03.04.05.0 beta 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.10.10.00.0 0.00.20.20.20.20.2 0.00.20.20.20.20.2 0.00.30.20.20.30.3 0.00.30.30.20.30.3 0.10.40.20.30.40.3 0.00.60.30.60.50.4 0.00.50.30.50.40.3 0.00.60.40.50.30.2 0.00.60.50.50.30.1 0.10.60.50.40.10.1 0.10.50.50.20.10.0 0.10.40.40.10.00.0 0.10.30.20.10.00.0 0.00.30.20.10.00.0 0.10.20.10.00.00.0 0.10.10.00.00.00.0 0.00.10.00.00.00.0 0.00.10.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 english-french 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.10.20.40.30.10.0 0.00.40.50.50.60.5 0.10.40.60.70.60.4 0.10.70.90.70.80.9 0.21.00.90.91.00.9 0.71.00.91.00.90.9 0.50.90.90.90.90.9 0.11.00.90.90.90.8 0.30.90.90.90.80.8 0.30.90.90.80.60.6 0.20.90.90.70.50.4 0.10.90.70.40.20.1 0.10.80.60.30.20.1 0.30.80.50.20.10.0 0.10.80.40.20.00.0 0.30.70.30.20.00.0 0.10.60.30.20.00.0 0.10.50.20.20.00.0 0.10.40.20.20.10.0 0.00.30.20.20.10.0 0.00.30.20.20.10.1 0.00.30.20.20.10.1 0.00.20.20.20.10.1 0.00.20.20.20.10.1 0.00.20.20.10.10.1 0.00.20.20.10.10.1 0.00.20.10.10.10.1 0.00.20.10.10.10.1 0.00.10.10.10.10.1 0.10.10.10.10.10.1 0.10.10.10.10.10.1 0.00.10.10.00.00.0 0.00.10.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 present-past 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.10.10.00.00.0 0.00.00.00.00.00.0 0.00.00.00.10.00.0 0.00.00.10.00.00.0 0.10.30.60.60.50.2 0.00.60.80.80.70.6 0.20.70.90.80.70.8 0.10.90.90.90.90.9 0.21.00.91.01.01.0 0.71.00.91.01.01.0 0.51.01.01.01.00.9 0.21.01.00.90.90.9 0.51.01.00.90.90.8 0.41.01.00.90.80.7 0.40.90.90.80.60.4 0.20.90.80.60.30.2 0.30.90.80.40.20.2 0.60.90.70.30.20.1 0.30.80.50.30.20.1 0.40.70.40.20.20.0 0.20.70.40.30.20.1 0.10.60.30.30.20.1 0.20.50.30.30.20.1 0.10.40.30.30.20.1 0.10.40.30.20.20.2 0.10.40.30.20.20.2 0.10.40.30.20.20.2 0.10.40.30.20.20.2 0.10.30.30.20.20.2 0.10.30.20.20.20.2 0.10.30.20.20.20.2 0.10.30.20.20.10.1 0.10.30.20.20.20.1 0.10.20.20.10.10.1 0.20.20.10.10.10.1 0.20.20.10.10.10.1 0.10.20.10.10.10.1 0.10.20.10.10.10.1 0.00.10.10.10.10.1 singular-plural 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 Conceptor steering of GPT-NeoX (20B) Figure 10: Performance results of the grid search across layers and beta values (for the optimal aperture value) for the GPT-NeoX (20B) model, using conceptor-based steering. 15 aperture 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.30.40.4 0.00.40.40.4 0.00.40.40.4 0.00.50.50.5 0.00.50.50.5 0.00.40.50.5 0.00.40.40.4 0.00.30.30.4 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 0.00.00.00.0 0.00.00.00.0 antonyms aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.10.30.3 0.00.10.20.2 0.00.30.90.9 0.00.40.90.9 0.00.70.90.9 0.00.91.00.9 0.01.01.01.0 0.01.01.01.0 0.01.01.01.0 0.01.01.01.0 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.80.90.9 0.00.80.90.8 0.00.60.70.7 0.00.50.70.7 0.00.30.30.4 0.20.20.30.4 0.00.00.00.1 0.00.00.00.0 capitalize aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.10.1 0.00.00.10.1 0.00.10.10.1 0.00.10.10.1 0.00.10.20.2 0.00.50.60.6 0.00.60.60.6 0.00.70.70.7 0.00.80.80.8 0.00.80.80.8 0.00.80.80.8 0.00.80.80.8 0.00.70.80.8 0.00.60.70.8 0.00.30.50.5 0.00.10.40.5 0.00.10.20.3 0.00.10.20.2 0.00.10.20.2 0.00.10.10.2 0.00.10.10.1 0.00.10.10.1 0.00.10.10.1 country-capital 0.0010.01250.050.1 aperture 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.10.1 0.00.10.20.2 0.00.10.20.2 0.00.20.20.2 0.00.20.30.3 0.00.40.40.4 0.00.50.40.4 0.00.60.60.6 0.00.50.40.4 0.00.50.50.5 0.00.40.40.4 0.00.10.20.2 0.00.00.10.1 0.00.10.10.1 0.00.10.10.2 0.00.10.10.2 0.00.00.00.1 0.00.00.00.1 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 english-french 0.0010.01250.050.1 aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.30.3 0.00.10.40.4 0.00.40.40.4 0.00.40.40.4 0.00.40.40.4 0.00.40.40.4 0.00.40.40.4 0.00.40.50.5 0.00.50.60.5 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.80.90.9 0.00.70.80.8 0.00.70.70.8 0.00.50.60.7 0.00.40.50.6 0.00.10.20.4 0.00.00.20.3 0.00.00.00.2 0.00.00.00.1 0.00.00.00.0 0.00.00.00.0 present-past 0.0010.01250.050.1 aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.50.5 0.00.20.60.6 0.00.60.60.6 0.10.60.60.6 0.00.60.60.6 0.00.60.70.7 0.00.60.70.7 0.00.70.70.7 0.00.90.90.8 0.00.90.90.9 0.00.90.90.9 0.00.90.90.9 0.00.80.80.9 0.00.60.70.8 0.00.60.70.8 0.00.40.60.7 0.00.30.60.7 0.00.10.30.6 0.00.10.30.6 0.00.00.10.3 0.00.00.10.2 0.00.00.10.1 0.00.00.10.1 singular-plural 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.2 0.4 0.6 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 Conceptor steering of GPT-J (6B) Figure 11: Performance results of the grid search across layers and aperture values (for the optimal beta value) for the GPT-J (6B) model, using conceptor-based steering. 16 aperture 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.10.10.1 0.30.40.40.4 0.50.50.50.4 0.60.60.60.5 0.60.60.60.6 0.60.60.60.6 0.60.60.60.6 0.50.50.50.5 0.50.40.50.5 0.30.30.40.4 0.20.20.20.2 0.20.20.20.2 0.20.20.20.2 0.10.10.20.2 0.10.10.10.1 0.10.10.10.1 0.10.10.10.1 0.10.10.10.1 0.00.10.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.00.1 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 antonyms aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.10.10.1 0.91.01.01.0 0.91.01.00.9 1.01.01.00.9 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 0.91.01.01.0 0.91.01.01.0 0.90.91.01.0 0.90.91.01.0 0.70.80.90.9 0.70.80.90.9 0.50.70.80.8 0.50.60.70.7 0.30.40.60.6 0.20.40.50.5 0.10.30.40.4 0.10.20.30.3 0.10.20.30.3 0.00.10.20.2 0.00.10.20.2 0.00.10.20.2 0.00.10.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.00.1 capitalize aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.10.0 0.10.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.10.10.1 0.20.30.30.3 0.50.60.70.8 0.90.90.90.9 0.90.90.90.9 1.01.00.90.9 0.91.00.90.9 0.90.90.90.9 0.90.90.90.9 0.90.90.90.9 0.90.90.90.9 0.90.80.90.9 0.90.80.90.9 0.70.70.80.8 0.70.70.80.8 0.50.50.70.8 0.30.50.60.7 0.20.30.50.6 0.10.20.50.6 0.10.20.40.5 0.10.20.40.5 0.10.20.40.4 0.10.20.30.4 0.10.10.30.3 0.10.10.30.3 0.10.10.20.3 0.10.10.20.2 0.10.10.20.2 0.10.10.20.2 0.10.10.20.2 0.10.10.20.2 0.10.10.10.1 0.10.10.10.2 0.10.10.10.2 0.10.10.10.1 country-capital 0.0010.01250.050.1 aperture 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.10.10.1 0.20.20.20.2 0.20.20.20.2 0.30.30.30.2 0.30.30.30.3 0.40.40.30.3 0.60.60.60.6 0.50.50.50.5 0.50.50.60.6 0.50.50.60.6 0.50.50.60.6 0.50.40.50.5 0.40.30.40.4 0.20.20.30.3 0.20.20.30.3 0.10.10.10.2 0.00.10.10.1 0.00.00.10.1 0.00.00.00.1 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 english-french 0.0010.01250.050.1 aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.10.40.30.3 0.40.60.60.5 0.40.70.60.6 0.90.90.80.6 1.01.00.90.9 1.01.01.00.9 0.90.90.90.9 0.90.90.91.0 0.90.90.90.9 0.90.90.90.9 0.90.80.90.9 0.70.70.90.9 0.60.60.80.8 0.50.40.80.8 0.30.30.70.8 0.30.30.50.7 0.10.10.50.6 0.10.10.30.5 0.00.00.30.4 0.00.00.30.3 0.00.00.20.3 0.00.00.20.3 0.00.00.20.2 0.00.00.20.2 0.00.00.20.2 0.00.00.20.2 0.00.00.10.2 0.00.00.10.2 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.10.1 0.00.00.00.1 0.00.00.00.0 0.00.00.00.0 present-past 0.0010.01250.050.1 aperture layer 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.00.00.00.0 0.10.00.00.0 0.00.00.00.0 0.00.10.00.0 0.00.10.10.0 0.10.60.60.6 0.60.80.80.7 0.80.90.80.8 0.90.90.90.9 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 1.01.01.01.0 0.90.90.90.9 0.80.80.90.9 0.80.80.90.9 0.70.70.80.9 0.50.50.70.8 0.20.30.60.7 0.10.20.50.7 0.10.10.50.6 0.10.10.40.5 0.00.10.30.4 0.10.10.30.4 0.10.10.30.4 0.10.10.30.4 0.00.10.30.4 0.10.10.30.3 0.10.10.30.3 0.10.10.30.3 0.00.10.20.3 0.10.10.20.3 0.10.10.20.2 0.00.10.20.2 0.00.10.20.2 0.00.00.20.2 0.00.10.10.2 0.00.00.10.1 singular-plural 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.2 0.4 0.6 0.8 0.0 0.2 0.4 0.6 0.8 Conceptor steering of GPT-NeoX (20B) Figure 12: Performance results of the grid search across layers and aperture values (for the optimal beta value) for the GPT-NeoX (20B) model, using conceptor-based steering. 17 beta 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.10.10.10.1 0.00.10.20.20.20.2 0.00.10.20.20.20.2 0.00.10.20.20.20.2 0.00.10.20.20.20.1 0.00.10.20.20.10.1 0.00.00.10.10.10.0 0.00.00.10.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 antonyms beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.10.10.2 0.00.00.00.00.10.2 0.00.00.70.90.90.9 0.00.00.70.90.90.9 0.00.10.80.90.90.9 0.00.10.80.90.90.8 0.00.40.90.90.90.8 0.00.60.90.90.80.8 0.00.70.80.70.60.5 0.00.70.70.60.40.4 0.10.60.60.50.30.3 0.00.40.40.20.20.1 0.00.30.30.20.20.1 0.00.30.30.30.20.2 0.00.10.10.10.10.1 0.00.10.20.20.10.1 0.00.00.10.10.10.1 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 capitalize beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.10.1 0.00.00.00.00.10.1 0.00.00.00.10.10.1 0.00.00.00.10.10.1 0.00.00.10.10.10.1 0.00.00.20.20.20.2 0.00.00.10.20.20.2 0.00.00.20.20.20.2 0.00.00.20.20.20.2 0.00.00.30.20.10.1 0.00.00.20.20.10.1 0.00.00.10.10.10.0 0.00.00.10.10.10.0 0.00.00.10.10.10.0 0.00.00.10.10.10.0 0.00.00.10.10.10.0 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 country-capital 0.51.02.03.04.05.0 beta 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.1 0.00.00.00.00.10.1 0.00.00.00.10.20.2 0.00.00.00.10.20.2 0.00.00.20.20.20.2 0.00.10.20.20.20.2 0.00.10.20.20.20.2 0.00.00.20.20.10.1 0.00.00.20.20.10.1 0.00.00.10.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 english-french 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.20.40.4 0.00.00.10.30.40.4 0.00.00.30.40.40.4 0.00.00.30.40.40.4 0.00.10.40.50.40.4 0.00.30.50.50.50.5 0.00.30.40.40.50.4 0.00.20.40.40.40.4 0.00.30.50.50.50.5 0.00.50.70.70.60.4 0.00.30.60.40.20.1 0.00.30.60.30.20.1 0.00.30.50.20.10.1 0.00.10.30.10.10.0 0.00.00.20.10.00.0 0.00.10.30.10.00.0 0.00.00.10.10.00.0 0.00.00.10.10.00.0 0.00.00.10.10.00.0 0.00.00.10.10.00.0 0.00.00.00.10.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 present-past 0.51.02.03.04.05.0 beta layer 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.00.00.0 0.00.00.00.30.60.6 0.00.00.10.50.60.6 0.00.00.50.60.60.6 0.00.00.50.60.60.6 0.00.20.60.60.60.6 0.00.40.60.70.70.7 0.00.50.60.60.60.6 0.00.40.60.60.60.5 0.00.50.70.70.70.5 0.10.70.70.70.60.5 0.00.50.70.50.30.3 0.00.50.60.40.30.2 0.10.40.50.30.20.1 0.00.30.40.20.10.1 0.00.30.30.20.10.1 0.10.30.40.20.20.1 0.00.10.20.20.10.1 0.00.10.30.20.20.1 0.00.10.20.20.10.1 0.00.10.20.20.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.1 0.00.00.10.10.10.0 singular-plural 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 0.0 0.2 0.4 0.6 0.8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 Additive steering of GPT-J (6B) Figure 13: Performance results of the grid search across layers and beta values for the GPT-J (6B) model, using additive steering. 18 beta 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.10.10.10.10.10.10.1 0.00.10.20.20.20.10.10.1 0.00.20.30.30.20.20.10.1 0.00.30.30.30.20.20.10.1 0.00.20.20.20.20.20.10.1 0.00.10.10.20.20.20.10.1 0.00.00.10.10.10.10.10.0 0.00.00.10.10.10.10.10.0 0.00.00.00.10.10.10.10.0 0.00.00.00.10.10.10.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 antonyms beta layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.1 0.00.10.40.70.80.70.50.4 0.00.10.40.70.80.60.40.3 0.00.20.60.80.90.80.40.3 0.00.30.80.90.90.80.40.3 0.00.50.80.90.90.70.30.3 0.00.60.90.90.80.70.40.3 0.00.70.90.90.80.60.40.0 0.10.70.90.80.70.50.30.0 0.00.70.80.80.70.50.00.0 0.10.70.80.80.60.50.00.0 0.10.70.60.60.50.00.00.0 0.10.50.50.40.40.00.00.0 0.10.30.30.30.10.00.00.0 0.00.20.20.20.00.00.00.0 0.00.20.20.20.00.00.00.0 0.00.10.20.20.00.00.00.0 0.00.10.20.20.00.00.00.0 0.00.10.20.10.00.00.00.0 0.00.10.20.10.10.00.00.0 0.00.10.20.20.10.10.00.0 0.00.10.20.10.10.10.00.0 0.00.10.10.10.10.10.00.0 0.00.10.10.10.10.10.00.0 0.00.10.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.00.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 capitalize beta layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.10.1 0.00.00.10.10.10.10.10.1 0.00.00.10.10.10.10.10.1 0.00.00.10.10.10.10.10.1 0.00.00.20.30.20.20.20.1 0.00.10.30.30.30.20.10.1 0.00.10.30.20.20.20.20.1 0.00.00.20.30.20.20.20.1 0.00.10.30.20.20.20.10.1 0.00.00.20.20.20.20.10.1 0.00.10.20.20.20.20.10.0 0.00.10.20.20.20.20.10.0 0.00.10.20.20.20.10.00.0 0.00.10.20.20.10.10.00.0 0.00.10.20.20.10.00.00.0 0.00.10.20.20.00.00.00.0 0.00.00.20.20.00.00.00.0 0.00.00.10.10.00.00.00.0 0.00.00.10.10.10.00.00.0 0.00.00.10.10.20.10.00.0 0.00.00.10.10.10.10.10.0 0.00.00.10.20.20.20.10.0 0.00.00.10.20.20.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.20.20.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.00.10.10.10.00.0 0.00.00.10.10.10.10.10.1 0.00.00.10.10.10.10.10.1 0.00.00.10.10.10.10.10.1 0.00.00.10.10.10.10.10.1 0.00.00.00.10.10.10.10.1 0.00.00.00.00.00.00.00.0 country-capital 0.51.01.52.02.53.04.05.0 beta 0 2 4 6 8 10 12 14 16 18 20 22 24 26 28 30 32 34 36 38 40 42 layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.10.10.10.10.10.10.1 0.00.10.20.10.10.10.10.0 0.00.10.20.20.20.10.10.1 0.00.10.20.20.10.10.10.0 0.00.10.10.10.10.10.10.0 0.00.10.20.20.10.10.00.0 0.00.10.10.10.10.10.00.0 0.00.10.10.10.10.00.00.0 0.00.10.10.10.10.00.00.0 0.00.00.10.10.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 english-french 0.51.01.52.02.53.04.05.0 beta layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.20.40.40.40.40.30.2 0.00.30.40.40.50.40.20.1 0.00.40.40.40.50.50.30.2 0.10.40.60.70.70.70.30.1 0.10.40.60.70.70.60.20.1 0.20.50.70.70.70.50.20.0 0.30.60.70.70.60.40.10.0 0.20.60.70.60.50.30.10.0 0.20.60.60.50.30.10.00.0 0.20.60.50.40.20.10.00.0 0.20.40.30.20.20.00.00.0 0.10.20.20.10.10.00.00.0 0.00.10.10.10.00.00.00.0 0.00.10.10.10.00.00.00.0 0.00.10.10.10.00.00.00.0 0.00.10.10.10.00.00.00.0 0.00.00.10.10.00.00.00.0 0.00.00.10.10.00.00.00.0 0.00.00.10.10.00.00.00.0 0.00.00.10.10.10.00.00.0 0.00.00.00.10.00.00.00.0 0.00.00.00.10.00.00.00.0 0.00.00.00.10.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.10.00.00.00.0 0.00.00.00.10.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 present-past 0.51.01.52.02.53.04.05.0 beta layer 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.10.1 0.00.20.60.60.70.50.30.2 0.00.50.60.60.60.40.20.1 0.00.60.60.70.60.50.20.2 0.00.70.70.70.70.60.20.2 0.10.70.80.80.70.60.20.1 0.20.80.80.80.80.50.20.1 0.20.70.80.80.60.40.10.1 0.20.70.80.70.50.30.10.0 0.30.70.70.50.30.20.10.0 0.30.70.60.40.20.10.10.0 0.20.50.40.20.20.10.00.0 0.10.40.30.20.10.10.00.0 0.10.20.20.10.10.00.00.0 0.10.10.10.10.10.00.00.0 0.10.10.10.10.10.00.00.0 0.10.10.10.20.00.00.00.0 0.10.10.10.10.00.00.00.0 0.00.10.10.20.10.00.00.0 0.00.10.10.20.10.00.00.0 0.00.10.20.20.20.10.00.0 0.00.10.20.20.20.20.10.0 0.00.10.10.20.10.10.10.0 0.00.10.10.20.10.10.10.0 0.00.10.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.10.10.10.10.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 0.00.00.00.00.00.00.00.0 singular-plural 0.00 0.05 0.10 0.15 0.20 0.25 0.0 0.2 0.4 0.6 0.8 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 0.16 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Additive steering of GPT-NeoX (20B) Figure 14: Performance results of the grid search across layers and beta values for the GPT-NeoX (20B) model, using additive steering. 19