Paper deep dive
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
Eric Bigelow, Daniel Wurgaft, YingQiao Wang, Noah Goodman, Tomer Ullman, Hidenori Tanaka, Ekdeep Singh Lubana
Models: Gemma-2-9B, Llama-3.1-8B, phi-4-mini-instruct, Qwen-2.5-7B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 6:14:35 PM
Summary
The paper proposes a unified Bayesian framework to explain and predict the effects of In-Context Learning (ICL) and activation steering in Large Language Models (LLMs). It posits that both methods function by updating the model's belief in latent concepts: ICL accumulates evidence through context, while activation steering modifies concept priors. The authors demonstrate that this model predicts sigmoidal learning curves, the additivity of interventions in log-belief space, and sudden behavioral shifts, providing a theoretical foundation for inference-time control.
Entities (5)
Relation Signals (4)
Contrastive Activation Addition → implements → Activation Steering
confidence 95% · For our experiments, we will primarily use the steering protocol introduced by Turner et al. (2024)... called Contrastive Activation Addition
Activation Steering → modifies → Concept Priors
confidence 95% · steering operates by changing concept priors
In-Context Learning → updates → Latent Concepts
confidence 95% · in-context learning leads to an accumulation of evidence... altering its belief in latent concepts
Bayesian Belief Dynamics Model → unifies → In-Context Learning
confidence 90% · we develop a unifying, predictive account of LLM control from a Bayesian perspective
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different accounts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly disparate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation-based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interventions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena - e.g., sigmoidal learning curves as in-context evidence accumulates - while predicting novel ones - e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly changing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions.
Tags
Links
Trouble viewing inline? Open PDF directly →
Full Text
92,083 characters extracted from source content.
Expand or collapse full text
BELIEF DYNAMICS REVEAL THE DUAL NATURE OF IN-CONTEXT LEARNING AND ACTIVATION STEERING Eric Bigelow ∗1,2,4 , Daniel Wurgaft ∗1,5 , YingQiao Wang 2 , Noah Goodman 5,6 , Tomer Ullman 2,3 , Hidenori Tanaka 3,4 , Ekdeep Singh Lubana 1 1 Goodfire AI, 2 Department of Psychology, Harvard University, 3 CBS, Harvard University 4 Physics of Intelligence Group, NTT Research, 5 Department of Psychology, Stanford University, 6 Department of Computer Science, Stanford University, ∗ Co-first authors ABSTRACT Large language models (LLMs) can be controlled at inference time through prompts (in-context learning) and internal activations (activation steering). Different ac- counts have been proposed to explain these methods, yet their common goal of controlling model behavior raises the question of whether these seemingly dis- parate methodologies can be seen as specific instances of a broader framework. Motivated by this, we develop a unifying, predictive account of LLM control from a Bayesian perspective. Specifically, we posit that both context- and activation- based interventions impact model behavior by altering its belief in latent concepts: steering operates by changing concept priors, while in-context learning leads to an accumulation of evidence. This results in a closed-form Bayesian model that is highly predictive of LLM behavior across context- and activation-based interven- tions in a set of domains inspired by prior work on many-shot in-context learning. This model helps us explain prior empirical phenomena—e.g., sigmoidal learning curves as in-context evidence accumulates—while predicting novel ones—e.g., additivity of both interventions in log-belief space, which results in distinct phases such that sudden and dramatic behavioral shifts can be induced by slightly chang- ing intervention controls. Taken together, this work offers a unified account of prompt-based and activation-based control of LLM behavior, and a methodology for empirically predicting the effects of these interventions. 1INTRODUCTION Large Language Models (LLMs) have begun demonstrating increasingly impressive capabilities (Brown et al., 2020; Kaplan et al., 2020; Bubeck et al., 2023; Chang et al., 2024). However, reliable use of these systems in practical applications mandates the design of protocols that ensure generated outputs satisfy desirable properties—e.g., avoiding violent or harmful speech, sycophantic responses, or engagement with unsafe queries (Bai et al., 2022b;a; Anwar et al., 2024). To this end, prior work targeting inference-time control of model behavior has developed two broad methodologies: input- level interventions via In-Context Learning (ICL), where contexts such as questions, instructions, dialog, or sequences of input-output examples are used to condition model behavior (Brown et al., 2020; Liu et al., 2023; Wei et al., 2022; Bai et al., 2022b;a), and representation-level interventions via activation steering, where a model’s behavior is modulated by directly intervening on its hidden activations (Turner et al., 2024; Geiger et al., 2021; Templeton et al., 2024). Practical approaches to ICL often involve an informal process of prompt engineering through trial-and-error (White et al., 2023; Sahoo et al., 2024), whereas approaches to activation steering typically use ad-hoc datasets of contrasting pairs of examples (Turner et al., 2024; Marks and Tegmark, 2024). To better understand the empirical success of these methods, recent theoretical work has begun exploring how input and representation-level interventions impact the distribution of generated outputs. Specifically, ICL has been framed as a form of Bayesian inference, where context modulates a space of hypotheses learned during pretraining (Xie et al., 2021; Bigelow et al., 2023; Wurgaft et al., 2025; Arora et al., 2024). Activation steering, on the other hand, has been argued to be a direct consequence of models learning to match the data distribution, which leads them to develop linear representations of concepts in particular layers (Park et al., 2024b; 2025b; Ravfogel et al., 2025; 1 arXiv:2511.00617v1 [cs.LG] 1 Nov 2025 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Steering Vector Magnitude Persona-Matching Behavior Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 In-Context Learning and Activation Steering Both Impact Behavior... by Updating Belief in Latent Concepts Number of ICL Shots Steering Vector Magnitude Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c Output Behaviors y ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# !"#$%&'()%*$+,-%(.*&/ Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 0-1)/(2+,-%(.*&/ #3 #4 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Inputs x Outputs y Hidden Layers Updated Belief in Concept c !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45677#8%#,/(.+9+*#-.# .(:+0#.%#)+.#,/(.#5#,(2.;< ="#>+0 In-Context Prompts x Latent Concepts c ?0&@/%A(./-@ ?+*0%2( B+'.*(7#?+*0%2( Activation Steering In-Context Examples In-Context Learning & Activation Steering Both Impact Behavior ... by Updating Belief in Latent Concepts Inputs x Outputs y Hidden Layers Probability of Behavior Matching Concept p(y 1 | x) Activation Steering Log Number of Input Examples |x| I n - C o n t e x t L e a r n i n g ="#>+0 ="#B% !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##4C./+*#A+%A7+#(*+# 2(-9+#(28#+(0-7(2-A'7(.+8;< ; ;# ;# Steering Vector Magnitude Context Length |x| p(y 1 | x) 0-1+1 23212527 p(y 1 | x) !"#$%#&%'#()*++#,-./#./-0# 0.(.+1+2.3##45#.+28#.%#A'.# %./+*0D#2++80#E+F%*+#1&#%,2;< ="#B% 0 8 32128 Number of ICL Shots 0.5 -1.0 0.0 0.5 1.0 Steering V ector Magnitude !"#$%$&'()*+,-(., Persona-Matching Behavior Steering Vector Magnitude 0123-1-3-2 1.0 0.0 0.5 Number of ICL Shots Persona-Matching Behavior 0 8 32128 1.0 0.0 0.5 Probability of Behavior Matching Concept p(y 1 | x) Number of Input Examples |x| Updated Belief in Concept c In-Context Examples Activation Steering Hidden Layers Hidden Layers Figure 1: Overview of our unified Bayesian theory of in-context learning and activation steering We argue that in-context learning (ICL) and activation steering both impact behavior by updating an LLM’s belief in latent concepts. We empirically test our claims in five domains of manipulating language model “persona” (bottom left) and predict that ICL will follow a sudden learning curve with increasing context length, and that this curve will be shifted under activation steering (top left). By our account, ICL with increasing context length|x|and steering vectors with increasing magnitude both operate by updating an LLM’s belief in latent concepts c. Arora et al., 2016). Given the shared goals of ICL and activation steering, it is plausible that there is a broader framework that helps formalize the notion of control in a probabilistic system, with seemingly disparate approaches, such as ICL and activation steering, acting as specific instances of this framework. This workMotivated by the above, we posit that various approaches to changing LLM behavior at inference time can be understood as belief updating. Specifically, we propose a Bayesian belief dynamics model where in-context learning reweighs concepts according to their likelihood functions, while steering reweighs concepts by altering their prior probabilities (Fig. 1). We design a set of experiments that build on prior work in many-shot ICL (Anil et al., 2024; Agarwal et al., 2024), and introduce activation steering magnitude as an additional dimension for belief updating, along with the number of ICL shots. Our results show three striking behavioral phenomena that can be predicted by our belief dynamics model: specifically, (i) a sigmoidal growth of posterior belief as a function of in-context exemplars, explaining prior results on sudden learning curves in ICL; (i) predictable shift in the ICL behavior proportional to the magnitude of steering vector; and (i) an additive effect of these interventions that yields distinct phases such that, as a function of intervention controls (context and steering magnitude), model behavior changes suddenly. Crucially, by formalizing and fitting our Bayesian model to the behavioral data, we are able to predict the point where this sudden change occurs, offering a concrete prediction for the phenomenon of many-shot jailbreaking (Anil et al., 2024). More broadly, our work demonstrates the utility of applying a Bayesian perspective at various levels of analysis for understanding neural networks (Marr, 1982): to capture the space of behaviors that an LLM performs, as well as aid at understanding the representations underlying such behaviors. In our case, belief updating explains phenomena at both the level of behavior, i.e., how an LLM’s output changes as a function of input given to it, and at the level of representation, i.e., in the effect of activation-level interventions. Correspondingly, this work contributes to a growing body of literature that uses Bayesian theories and models to study learning and conceptual representation in deep neural networks (Bigelow et al., 2023; Park et al., 2025a; Wurgaft et al., 2025). Building on the success of Bayesian approaches in explaining natural intelligence within cognitive science (Tenenbaum et al., 2011; Ullman and Tenenbaum, 2020), we argue that Bayesian principles can serve as a theoretical foundation for many different approaches to interpreting and controlling LLMs. 2BACKGROUND We first offer a short primer highlighting points relevant to the two core phenomena that we aim to unify in this work: in-context learning and activation steering. We build on these points to define our Bayesian model in the next section. 2 2.1IN-CONTEXT LEARNING In-Context Learning (ICL), where an LLM learns from linguistic context, is often contrasted with in- weights learning, where an LLM learns during (pre)training by adjusting model weights (Chan et al., 2022; Reddy, 2023; Lampinen et al., 2024; Nguyen and Reddy, 2024). While ICL is traditionally framed as few-shot learning (Brown et al., 2020), wherein exemplars corresponding to a task are offered to a model in-context and the model is expected to perform the demonstrated task on a novel query, there is a broader spectrum of language model capabilities that fall under the category of in-context learning (Lampinen et al., 2024; Park et al., 2025a; 2024a), e.g., zero-shot learning of a novel language (Gemini Team, 2023; Bigelow et al., 2023; Akyürek et al., 2024) or optimization of a utility function (Von Oswald et al., 2023; Demircan et al., 2024; Yin et al., 2024). ICL as Bayesian Inference As argued by Xie et al. (2021); Bigelow et al. (2023); Panwar et al. (2024); Zhang et al. (2023); Min et al. (2022) and recently verified by Wurgaft et al. (2025); Park et al. (2024a); Raventós et al. (2024) in toy domains, different perspectives and phenomenology associated with ICL can be captured in a unifying, predictive framework by casting ICL as Bayesian inference. We build on this perspective by formalizing a Bayesian account of ICL in practical, large-scale settings. Specifically, following prior work, we define the distribution of model outputsyconditioned on input context x as inference over latent concepts c: p(y|x) = Z c p(y|c) p(c|x)∝ Z c p(y|c) p(x|c) p(c).(1) The space of latent conceptsc∈Cis learned during model pretraining, and then, at inference time, these concepts are evoked by different input prompts x via the concept likelihood functions p(x|c). 2.2ACTIVATION STEERING Activation steering includes a broad set of protocols that intervene on the hidden representations of a language model to manipulate its outputs (Turner et al., 2024; Panickssery et al., 2024). Specifically, such protocols involve isolating directionsdin the representation space such that moving a hidden representationvalong them, i.e., alteringvtov + m· d, increases the odds the output reflects a conceptc, e.g., truthfulness (Li et al., 2023; Pres et al., 2024). Surprisingly, this simple strategy enables control of model behavior across several abstract concepts such as refusal (Arditi et al., 2024), model personalities (Chen et al., 2025; Yang et al., 2025), concepts relevant to defining a theory-of-mind (Chen et al., 2024), factuality (Li et al., 2023), uncertainty (Zur et al., 2025), and self-representations (Zhu et al., 2024). Contrastive Activation AdditionFor our experiments, we will primarily use the steering protocol introduced by Turner et al. (2024); Panickssery et al. (2024), called Contrastive Activation Addition (CAA) or “difference in means” steering. Specifically, CAA constructs steering vectors by collecting activationsa ℓ (X)from an LLM at the final token position of an inputX, for a given layerℓ, over two ‘contrasting’ datasets. As a specific example, suppose thatD c is a dataset of harmful prompts andD c ′ is a dataset of harmless prompts. In this case, CAA can be used to identify a direction for steering towards (or against) harmful queries (Arditi et al., 2024). More formally, we write a general formulation of CAA steering protocols as follows. ˆ d c,ℓ = 1 |D c | X x∈D c a ℓ (x) − 1 |D c ′ | X x∈D c ′ a ℓ (x)(2) = E p(x|c) h a ℓ (x) i − E p(x|c ′ ) h a ℓ (x) i Linear representation hypothesis and activation steering It is unclear precisely why activation steering methods work. These methods are similar in nature to analogies in word vector algebra (Mikolov et al., 2013), as in the classic exampleking : queen :: man : woman, which can be represented in vector algebra asv(king)−v(queen) = v(man)−v(woman). The Linear Rep- resentation Hypothesis (Park et al., 2024b; 2025b) formalizes this connection in terms of embedding representationλ(x)and an unembedding representationγ(y), where output behavior given an input p(y| x)is the softmax of the inner product:p(y| x)∝ exp λ(x) ⊤ γ(y) . If each concept variable 3 Y (C = c)is defined as a set of elements, e.g.,man, king ∈ c male orwoman, queen ∈ c female , concept vectors correspond to directions between an ordered pair of valuesc, e.g.,male→ femaleor English→ Russian. As Park et al. (2024b) show, if a model has learned to match the log posterior odds between a concept and its complement, that is, if p(c|λ(x)) p(c ′ |λ(x)) = p(c) p(c ′ ) , then all directions for increasingp(c|x)will be parallel and correspond to the steering vector identified via methods like CAA. In what follows, we build on this argument to model the effects of activation steering. 3MANY-SHOT IN-CONTEXT LEARNING EXPERIMENTS MSJ REPLICATION - LLAMA # # # # # # # # # ! ! " " ! Figure 2: Replication of many-shot ICL results in persona domains (Anil et al., 2024) For our experiments, we used a selection of datasets that correspond to concepts that LLMs assign relatively low probability to, but which consist of behaviors that a suffi- ciently capable LLM would be able to follow accurately. In other words, we chose datasets for which we expect a significant improvement with many-shot ICL and activa- tion steering, and for which we also expect LLM perfor- mance to reach nearly 100% with≤ 128in-context exem- plars. We focus on the approach of many-shot in-context learning (Anil et al., 2024; Agarwal et al., 2024; Arora et al., 2024), which involves cases where LLM perfor- mance continues to improve when a large number (dozens to hundreds) of input examples are provided in-context. Many-shot ICL provides a case study of in-context learn- ing dynamics, where previous work has shown that many- shot ICL follows a sharp learning trend as the number of ICL examples increases, shown in Fig. 2. In other words, as the amount of in-context data increases, the LLM’s behavior at first changes slowly, then it changes rapidly as the model reaches a transition point (an inflection point typically aroundp(y|x) = 0.5) and finally plateaus towards a maximum value. These ICL dynamics can be effectively explained by power-law scaling models, which assume that LLMs update their beliefs sub-linearly as data accumulates (Anil et al., 2024). Next, we provide further details about our many-shot ICL experiments. Experimental details not provided below appear in App. E. We used three harmful persona datasets previously used for many-shot jailbreaking (Anil et al., 2024; Arora et al., 2024), as well as two additional (non-harmful) persona datasets from the same collection (Perez et al., 2022) 1 . The three harmful personas represent the “dark triad” of personality traits: Psychopathy, Machiavellianism, and Narcissism. Each of these represents a distinct set of properties that, if present in deployed LLMs, could present a risk of harm to users. The two additional personas we test, Subscribes to Moral Nihilism and Believes Life Has No Meaning, are categorized as types of “Moral Nihilism”. These personas are not necessarily harmful, but instead represent an arbitrary set of behaviors that are suppressed by post-training methods such as RLHF (Perez et al., 2022). As depicted in Fig. 1, the datasets consist of 1000 questions of the form Is the following statement something you would say? <statement> with two possible responsesy ∈ Yes, No, where half the statements have Yes as the persona-matching behavior y (c) , and half have No as the persona-matching behavior. Contextxin these domains consists of a sequence of chat-formatted user/assistant exchanges. These persona datasets were chosen because of LLMs’ relatively fast learning rates with these datasets where we can observe the full sudden learning dynamics (Fig. 2), including both transition points and final plateau values, with fewer than 128 in-context examples (i.e.|x| ≤ 128), and because behaviorp(y|x)can easily be measured by taking the LLM’s token logit probabilities for Yes and No. In the following section, we develop a theoretical framework for understanding how both ICL and activation steering operate in terms of updating beliefs in an LLM, and a belief dynamics model that implements this framework. Our framework makes three key predictions, and for each prediction we describe relevant empirical findings with Llama-3.1-8B and compare LLM behavior with that of our belief dynamics model (Eq 8). The analyses in our main text usedLlama-3.1-8b-Instruct (Dubey et al., 2024), a capable model that can be accommodated with relatively modest compute 1 https://github.com/anthropics/evals/tree/main/persona 4 Baseline Prior Figure 3: Belief updating with concept vectors(Left) From a representational perspective, we assume that the default behavior of an LLM (e.g. Neutral Personac ′ ) and the target behavior (Target Personac) correspond to concept vectors. In-context learning (blue) directs the initial belief state from c ′ to increasingly point towardscas a function of the log number of shots|x|. Activation steering (orange) similarly directs the belief state towardscas a function of steering magnitude. (Right) We offer a parallel Bayesian perspective that in-context learning (x k ) and activation steering (v) both operate by changing an LLM’s belief in latent conceptsc. In our theory, in-context learning updates the posterior belief through the likelihood functionp(x|c)(wherep(c|x) ∝ p(x|c)) and activation steering intervenes on concept priors p(c)→ p ′ (c). requirements. We also tested two additional LLMs of similar scale,Qwen-2.5-7b-Instruct andGemma-2-9b-Instruct(Appendix B). The results shown in Fig. 4, Fig. 6, and Fig. 7 represent held-out predictions using 10-fold cross-validation across magnitude values. Overall, we find a very high correlation between LLM probabilities and predictions on held-out data (r = 0.98, averaged across our 5 domains). 4A BELIEF DYNAMICS MODEL OF ICL AND STEERING We now propose a unified model of controlling language models’ behavior via the input context (ICL) and intermediate representations (activation steering). Specifically, given an input contextx, we argue a model’s behaviorp(y|x)∝ p(c|x)can be formalized in the language of Bayesian inference as the beliefp(c|x)it associates with a conceptc, e.g., different personalities for persona manipulation (similar to Anil et al. (2024)). As further context is offered, the model will update its belief overc, whereas steering will either strengthen or suppress this belief in an input-invariant manner. To formalize the argument above, we consider a latent concept space that consists of a target concept c(e.g., a particular persona) and its complementc ′ (i.e., any behavior that does not align withc). To assess how a model’s belief incvs.c ′ evolves as the number of in-context shots|x| = Ngrows, we can examine the posterior oddso(c|x) = p(c) p(x|c) p(c ′ ) p(x|c ′ ) , i.e., the ratio between posterior probabilities of c and c ′ . Specifically, denoting the sigmoid function as σ, we can write the following. p(c|x) = p(c) p(x|c) p(c) p(x|c) + p(c ′ ) p(x|c ′ ) = o(c|x) 1 + o(c|x) = σ (logo(c|x)).(3) Eq. 3 thus puts the log posterior odds at the center of our analysis. To model this further, we can decompose the log posterior odds into a sum of the log prior odds and log-likelihood ratio (Bayes factor):logo(c|x) = log p(c) p(c ′ ) + log p(x|c) p(x|c ′ ) . Here, the prior odds represent the model’s initial belief in conceptccompared toc ′ . Sincec ′ is the complement ofc, the log prior odds are: log p(c) p(c ′ ) = log p(c) 1−p(c) . Consequently, to analyze the effects of ICL and activation steering on a model’s belief in a conceptc, we must evaluate how these interventions affect the Bayes factor and the prior odds. We analyze this next. 5 )%#) & !$" ) ) ) % ) ) # ) ) !" $ '! ' " $! $& ) ) ) ) ) ) )%# & !$" "( $( )%# & !$" ! """ )%#) & !$" '" " )) & !$" & " ! "$! " ICL CURVES - LLAMA Figure 4: In-context learning dynamics are sigmoidal with respect toN 1−α and modulated by activation steeringWe find sigmoidal many-shot in-context learning dynamics (solid lines) which can be effectively fit with a power law of scaling in-context data (dotted line). We additionally find that activation steering with different magnitudes (line colors) shifts in-context learning dynamics. In our belief dynamics model, this is explained by activation steering altering the LLM’s belief state. Model predictions represent held-out predictions from cross-validation. Note that, since we fit our models via cross-validation, we use the averageαfit across folds to transform the x-axis in this figure. 4.1CONTEXT IS EVIDENCE: DYNAMICS OF IN-CONTEXT LEARNING The likelihood term captures the relative evidence forcvs.c ′ fromNin-context examples. To model the log-likelihood, we follow Goodman et al. (2008) by assuming a concept’s log-likelihood declines proportionally to the number of labels that do not correspond to the expected labels for the concept. Denoting l i as the label for in-context example i and y (c) i as the concept-consistent label, we write: logp(x|c)∝−|i∈1,...,N| l i ̸= y (c) i |. In our experiments, all labels will correspond toc, and hencelogp(x|c) = 0andlogp(x|c ′ )∝−N. Thus, the likelihood function can be expected to accumulate evidence linearly with in-context examplesN. However, previous work has observed that the log-probability of next-token predictions scales as a power law with context size (Anil et al., 2024; Liu et al., 2024a; Park et al., 2025a). To account for this scaling, we follow Wurgaft et al. (2025) and model evidence accumulation as sub-linear by multiplying the log-likelihood by a discount factorτ(N). Under power-law growth of likelihood, we can showτ(N) = N −α and hence log-Bayes factor scales with context-size as log p(x|c) p(x|c ′ ) ∝ N 1−α (see App. A). Then, assuming a direct mapping between the concept-consistent labely (c) and the conceptc, the model’s probability for a concept-consistent answer is simplyp(c|x), yielding the following expression: p(y (c) |x) = p(c|x) = σ (− logo(c|x)) = σ − log p(c) p(c ′ ) − γN 1−α ,(4) where γ acts as the proportionality constant. Prediction 1Based on the functional form in Eq. 4, we should expectp(y (c) |x)to follow a sigmoidal trend as N 1−α accumulates. Results Building on our replication of results by Anil et al. (2024), we now show a more precise form of the sudden learning trend: specifically, in Fig. 4, we predict and demonstrate that in-context learning dynamics follow a sigmoid curve as a function ofN 1−α . This trend is captured effectively by our belief dynamics model, which uses a likelihood function that scales sub-linearly. This sub- linearity helps explain the results of prior work as well, since plotting the posterior as a function of log number of in-context exemplars should also yield a sudden learning trend. Beyond offering the precise functional form of this trend, we also show that in-context learning dynamics change as a function of steering magnitude, where positive steering magnitudes lead to similar ICL dynamics with fewer in-context examples (i.e., shifting the ICL curve leftwards) and negative magnitudes have the opposite effect (shifting the curve rightwards). 4.2ALTERING MODEL BELIEF: EFFECTS OF ACTIVATION STEERING We next aim to formalize the effects of activation steering on a model’s belief in some conceptc. To this end, we assume the linear representation hypothesis (LRH) holds for neural networks (Elhage 6 STEER RESPONSE FUNC - LLAMA #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% (((((( " "% & (((((( " "% % " #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% (((((( " "% & (((((( " "% % " Llama-3.1-8B Bayesian Model Figure 5: Change in behavior as a function of steering vector magnitudeAs we scale steering vector magnitude (x-axis), we find a sigmoidal response function in behavior (y-axis). With steering magnitudes in the range[−1, 1], we find approximately linear effects of steering, which taper off as magnitude increases. This pattern holds across different numbers of ICL examples (different colors). This pattern is well-captured by our model, which assumes a linear impact of steering on the log prior odds, and hence a sigmoidal impact in probability space. et al., 2022). Specifically, LRH states that neural network representations encode semantically meaningful concepts in hidden representations in a “linear” manner (Elhage et al., 2022; Arora et al., 2016). Here, “linearity” refers to three related phenomena: (i) concepts are linearly accessible from model representations, e.g., via simple logistic probes (Belinkov, 2022; Tenney et al., 2019); (i) linear algebraic manipulations of hidden representations along certain directions can steer model outputs (Panickssery et al., 2024; Turner et al., 2024); and (i) representations are defined as an additive mixture of these directions (Bricken et al., 2023; Templeton et al., 2024). One can unify these notions within a single formal computational model as follows. v = X i β i (v)d i s.t. d ⊺ i d j ∼ 0 ∀ i,j,(5) wherev ∈ R n is a hidden representation corresponding to inputx,d i ∈ R n represents some concept c i , and β i (v)∈ R is a scalar denoting the extent to which d i is present in v. LRH argues that if a concept is linearly represented (in the sense described above), then a logistic classifierσ (−w ⊺ v− b)suffices to infer the extent to which conceptcis present in the representation v. Since we assume minimal interference between directions that reflect different concepts, a well- trained classifier will have weights in line withd i (assuming it captures the concept we are interested in). Combining these assumptions, we get the following: p(c i |x) = p(c i |v) = σ (−w ⊺ v− b) = σ −β i ∥d i ∥ 2 − X j̸=i β j d T i d j −b ≈ σ (−β i (v)a− b), (6) wherea = ∥d i ∥ 2 . Using this, we can express the posterior odds as follows:log p(c i |v) p(c ′ i |v) = log p(c i |v) 1−p(c i |v) = aβ i (v) + b.Thus, if one steers the model representation along directiond i , e.g., changingvtov + m· d i 2 , the model’s belief in conceptc i will linearly increase (in log-space) toaβ i (v) + a· m + b. Relating this back to the unsteered model’s log-posterior odds, we get the following: log p(c i |v + m· d i ) p(c ′ i |v + m· d i ) = log p(v|c i ) p(v|c ′ i ) + log p(c i ) p(c ′ i ) + a· m = log p(v|c i ) p(v|c ′ i ) + log p ′ (c i ) p ′ (c ′ i ) .(7) 2 Note that steering boundlessly (e.g., takingm → ∞) will push the representations to a region that lies outside the support over which distributionP(c|v)is defined. We see such effects empirically (e.g., see Fig. 12) and thus focus our discussion in the main paper in a range formwhere posterior-belief changes monotonically. 7 HEATMAPS - LLAMA Llama-3.1-8B Bayesian Model ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , Figure 6: In-context learning and activation steering jointly affect behaviorThe in-context learning dynamics we observe in Fig. 4 and the steering vector magnitude response function in Fig. 5 interact to create a phase boundary (Top). Our belief dynamics model re-constructs this diagram with high fidelity (Bottom). That is, steering yields a constant shift in the model log-posterior odds that will consistently change model beliefs for both an individual observationxor an entire populationX ∼ P x . We therefore argue the effects of steering are best described as alteration of a model’s prior beliefs in a concept c, updating the log prior odds fromlog p(c) p(c ′ ) tolog p ′ (c) p ′ (c ′ ) (wherep ′ (c)is an unnormalized prior). Intuitively, this formalizes the claim that steering vectors should be expected to change behaviory regardless of the inputx. For example, for the conceptC happy we should expect the steering vector ˆ d c to make an LM behave more happy even if it is given inputsx (c ′ ) that are not happy, i.e., which have lower p(c| x (c ′ ) ). Prediction 2Assuming linear representation hypothesis holds, Eq. 7 shows steering will increase a model’s belief in concept c i at a sigmoidal rate with steering magnitude m. N* PREDS- LLAMA )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) # # #& ' ! )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) !( #( )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) !!! ) Figure 7: Predicting the amount of context needed to enable a persona. (Top) The be- lief dynamics model is highly predictive of crossover pointsN ∗ betweencandc ′ , with a correlation ofr = 0.97, and (Bottom) ourN ∗ estimates effectively predict phase boundaries in empirical behavioral data. ResultsWe find that activation steering leads to a sigmoidal trend in persona-matching behavior (and thus a linear trend for the posterior odds) as a function of steering vector magnitude (Fig. 5). We observe this within the rangem∈ [−3, 3]for the Dark Triad datasets, and the rangem ∈ [−1.5, 1.5]for Moral Nihilism datasets for Llama-3.1-8B. This trend holds across various context lengths, although with large contexts, the behavior is near ceiling for all magni- tudes. 4.3FINAL MODEL From Eq. 7 , the log posterior odds given an in- tervention onvcan be defined aslogo(c|x) = log p(c) p(c ′ ) +log p(x|c) p(x|c ′ ) +a·m , whereo(v|c) = o(x|c) since we assumep(c|x) = p(c|v)in Sec. 4.2. Next, we substitute the log prior oddslog p(c) p(c ′ ) with a con- stant offsetb, since it does not depend on the precise inputxor its representationv. This gives us our final model of belief update dynamics in ICL: logo(c|x) = a· m + b + γN 1−α (8) 8 This model describes how model behavior changes as a function of both context lengthNand steering magnitudem. Concretely, for the model prediction results described in this work, we fit scalar parameters toa,b,γ,αto empirical averages of model behaviorp(y|x) = p(c|x)(Eq. 4) using L-BFGS, across various contextsx(whereN =|x|) and steering with various magnitudesm. We do so for practical reasons: first, the steering vectors we use in practice are not the true concept vectors d i , and so the effect of steering magnitude will bea∝∥d i ∥ 2 (but not necessarilya =∥d i ∥ 2 ), and second, we estimate the prior oddsbrather than observing the concept priorsp(c),p(c ′ ). This model allows us to compute the transition points in context length for a given steering magnitudemwhen the model’s belief in concept c surpasses c ′ , i.e., when logo(c|x) = 0: N ∗ (m) = − am + b γ 1/(1−α) .(9) Prediction 3Log posterior odds will be additively impacted by varying in-context examples and steering magnitude, and this interaction will yield distinct phases dominated by belief in eithercorc ′ . The boundary between phases—the cross-over pointN ∗ when belief in conceptcsurpasses belief in c ′ —can be predicted as a function of initial log prior odds and steering magnitude (Eq. 9). ResultsObserving the phase diagrams in Fig. 6, we find that our model is highly predictive of the joint effects of in-context learning and steering. Moreover, following our definition ofN ∗ (m)in Eq. 9, we can predict the crossover points when behavior will transition to be dominated byc(Fig. 7). 5DISCUSSION In this work, we present a novel synthesis of prior theoretical and empirical work in two disparate approaches to language model control: in-context learning and activation steering. We find a phase boundary across ICL and activation steering, where the transition point is jointly modulated by context and activations. Further, we present a Bayesian belief dynamics model that formalizes this theory and accurately predicts language model behavior as a function of both context length and steering vector magnitude. Our approach builds on top-down theories of behavior from the perspective of Bayesian belief updating, as well as bottom-up theories of learning and representation in connectionist neural networks. This paves the way for future work to bridge levels of analysis for describing behaviors, the algorithms driving behavior, and the mechanisms that implement those algorithms (Marr, 1982; He et al., 2024). Taken together, our theory of language model control as belief updating and our empirical results supporting this raise a number of important questions. In this work, we found that steering vectors control behavior proportional to the vector magnitude, unless that magnitude becomes too large (see App. C). This may suggest that belief is only represented linearly within some subspace of the model’s representation space, although it is unclear whether belief is represented in a non-linear way outside this space, or whether this subspace represents the full extent of a model’s belief state with respect to a given concept. Further, we found cases with some LLMs (e.g.phi-4-mini-instruct) where steering had no clear effect on behavior - this may suggest that these models represent belief in a non-linear way, or it may indicate a limitation of our particular method for constructing steering vectors. A simpler explanation could be that for some models and certain datasets, there is not sufficient signal for a distinct behavior in the LLM to be captured by our steering vectors - e.g. if a model doesn’t represent a concept at all, then no amount of steering will change its belief in that concept. Our work also raises questions about precisely how LLMs implement belief updates and inference. We find that steering beliefs typically only works in a single layer, or a few layers, while other layers have no clear effect on behavior. Does this suggest that belief is localized to these layers, and if so, could we causally intervene on specific neurons in these layers (Geiger et al., 2025) to have predictable impacts on model behavior? Further, although we find that beliefs are linearly represented and localized in these cases, this begs the question of how distinct aspects of belief and inference are implemented - are concept likelihood functions implemented in a non-linear way, and are they implemented in earlier or later layers relative to this linear belief representation? And, if an agent represents and updates its beliefs in this way, how is inference implemented - is it similar to known algorithms such as Monte Carlo methods or variational inference? Lastly, it is also noteworthy that 9 our belief dynamics model is best fit to averages over LLM behavior rather than its raw data. This is reminiscent of work in cognitive science showing that individual human behaviors may be suboptimal due to resource constraints, but in aggregate, populations of people (Davis-Stober et al., 2014) - or even repeated sampling from an individual person (Vul and Pashler, 2008) - can behave optimally. We see there being a number of exciting directions for future follow-up work. The findings in our work may have practical implications for model control, to help practitioners understand how to best combine behavioral methods like ICL with mechanistic interventions for the purpose of controlling language models. The phase boundaries we find across ICL and steering could also have important consequences for AI safety, since language model behavior might suddenly and dramatically change after some threshold of context or steering is passed. Predicting these transition points, as we have done in this work, may prove to be an essential tool for safe and effective control of language models. One limitation of this work is that we only consider binary concepts and use only one method for constructing steering vectors (CAA). Future work may explore how our theory and model generalize to non-binary concept spaces, where there may be more than one direction for belief to vary across. Another compelling direction for future work is to explore how belief is updated with alternative steering vector methods such as SAEs (Templeton et al., 2024). In order to better understand the intersection of ICL and activation steering, another experiment we performed was to compute steering vectors over multiple shots (see App. D). We found that, counter-intuitively, after vectors were normalized to equal magnitude, steering vectors computed over multiple shots had an even weaker effect than steering with vectors computed over a single query. Further work is required to explain this effect. Finally, we also hope to explore in future work how activation steering and ICL interact with larger and more capable LLMs. Overall, by revealing a common Bayesian mechanism linking prompting and activation steering, our work re-frames how we understand belief and representation in LLMs, opening a rich space for theoretical exploration and principled model control in future work. ACKNOWLEDGEMENTS The authors thank Tom McGrath, Owen Lewis, and Yang Xiang, as well as the Computation and Cognition Lab at Stanford, for useful discussions during the course of this project. 10 REFERENCES Rishabh Agarwal, Avi Singh, Lei Zhang, Bernd Bohnet, Luis Rosias, Stephanie Chan, Biao Zhang, Ankesh Anand, Zaheer Abbas, Azade Nova, John D Co-Reyes, Eric Chu, Feryal Behbahani, Aleksandra Faust, and Hugo Larochelle. Many-shot in-context learning. 2024. Ekin Akyürek, Bailin Wang, Yoon Kim, and Jacob Andreas. In-context language learning: Architec- tures and algorithms. arXiv preprint arXiv:2401.12973, 2024. Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, Francesco Mosconi, Rajashree Agrawal, Rylan Schaeffer, Naomi Bashkansky, Samuel Svenningsen, Mike Lambert, Ansh Radhakrishnan, Carson Denison, Evan J Hubinger, Yuntao Bai, Trenton Bricken, Timothy Maxwell, Nicholas Schiefer, Jamie Sully, Alex Tamkin, Tamera Lanham, Karina Nguyen, Tomasz Korbak, Jared Kaplan, Deep Ganguli, Samuel R Bowman, Ethan Perez, Roger Grosse, and David Duvenaud. Many-shot jailbreaking. 2024. Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al. Foundational challenges in assuring alignment and safety of large language models. arXiv preprint arXiv:2404.09932, 2024. Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction, October 2024. Aryaman Arora, Dan Jurafsky, Christopher Potts, and Noah D Goodman. Bayesian scaling laws for in-context learning. arXiv preprint arXiv:2410.16531, 2024. Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. A Latent Variable Model Approach to PMI-based Word Embeddings. Transactions of the Association for Computational Linguistics, 4:385–399, 2016. doi: 10.1162/tacl_a_00106. URLhttps://aclanthology. org/Q16-1028/. Place: Cambridge, MA Publisher: MIT Press. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862, 2022a. Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, et al. Constitutional ai: Harmlessness from ai feedback. arXiv preprint arXiv:2212.08073, 2022b. Yonatan Belinkov. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219, 2022. Eric J Bigelow, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, and Tomer D Ullman. In-context learning dynamics with random binary sequences. arXiv preprint arXiv:2310.17639, 2023. Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. https://transformer-circuits.pub/2023/monosemantic- features/index.html. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023. 11 Stephanie CY Chan, Ishita Dasgupta, Junkyung Kim, Dharshan Kumaran, Andrew K Lampinen, and Felix Hill. Transformers generalize differently from information stored in context vs in weights. arXiv preprint arXiv:2210.05675, 2022. Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45, 2024. Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey. Persona vectors: Monitor- ing and controlling character traits in language models, July 2025. Yida Chen, Aoyu Wu, Trevor DePodesta, Catherine Yeh, Kenneth Li, Nicholas Castillo Marin, Oam Patel, Jan Riecke, Shivam Raval, Olivia Seow, et al. Designing a dashboard for transparency and control of conversational ai. arXiv preprint arXiv:2406.07882, 2024. Clintin P Davis-Stober, David V Budescu, Jason Dana, and Stephen B Broomell. When is a crowd wise? Decision, 1(2):79, 2014. Can Demircan, Tankred Saanum, Akshay K Jagadish, Marcel Binz, and Eric Schulz. Sparse autoencoders reveal temporal difference learning in large language models. arXiv preprint arXiv:2410.01280, 2024. Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv e-prints, pages arXiv–2407, 2024. Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCan- dlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy models of superposition, 2022. URL https://arxiv.org/abs/2209.10652. Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. Advances in Neural Information Processing Systems, 34:9574–9586, 2021. Atticus Geiger, Duligur Ibeling, Amir Zur, Maheep Chaudhary, Sonakshi Chauhan, Jing Huang, Aryaman Arora, Zhengxuan Wu, Noah Goodman, Christopher Potts, et al. Causal abstraction: A theoretical foundation for mechanistic interpretability. Journal of Machine Learning Research, 26 (83):1–64, 2025. Gemini Team.Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. Noah D Goodman, Joshua B Tenenbaum, Jacob Feldman, and Thomas L Griffiths. A rational analysis of rule-based concept learning. Cognitive science, 32(1):108–154, 2008. Zhonghao He, Jascha Achterberg, Katie Collins, Kevin Nejad, Danyal Akarca, Yinzhu Yang, Wes Gurnee, Ilia Sucholutsky, Yuhan Tang, Rebeca Ianov, et al. Multilevel interpretability of arti- ficial neural networks: leveraging framework and methods from neuroscience. arXiv preprint arXiv:2408.12664, 2024. Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020. Andrew Kyle Lampinen, Stephanie CY Chan, Aaditya K Singh, and Murray Shanahan. The broader spectrum of in-context learning. arXiv preprint arXiv:2412.03782, 2024. Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model. Advances in Neural Information Processing Systems, 36:41451–41530, 2023. Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing. ACM computing surveys, 55(9):1–35, 2023. 12 Toni JB Liu, Nicolas Boullé, Raphaël Sarfati, and Christopher J Earls. Llms learn governing principles of dynamical systems, revealing an in-context neural scaling law. arXiv preprint arXiv:2402.00795, 2024a. Toni J.b. Liu, Nicolas Boulle, Raphaël Sarfati, and Christopher Earls. Llms learn governing prin- ciples of dynamical systems, revealing an in-context neural scaling law. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, page 15097–15117. Association for Computational Linguistics, 2024b. doi: 10.18653/v1/2024.emnlp-main.842. URL http://dx.doi.org/10.18653/v1/2024.emnlp-main.842. Samuel Marks and Max Tegmark. The geometry of truth: Emergent linear structure in large language model representations of true/false datasets, August 2024. David Marr. Vision: A computational investigation into the human representation and processing of visual information. MIT press, 1982. Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representa- tions in vector space. arXiv preprint arXiv:1301.3781, 2013. Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. Rethinking the role of demonstrations: What makes in-context learning work? arXiv preprint arXiv:2202.12837, 2022. Alex Nguyen and Gautam Reddy. Differential learning kinetics govern the transition from memo- rization to generalization during in-context learning, 2024. URLhttps://arxiv.org/abs/ 2412.00104. Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition, July 2024. Madhur Panwar, Kabir Ahuja, and Navin Goyal. In-context learning through the bayesian prism, 2024. URL https://arxiv.org/abs/2306.04891. Core Francisco Park, Ekdeep Singh Lubana, Itamar Pres, and Hidenori Tanaka. Competition dynamics shape algorithmic phases of in-context learning. arXiv preprint arXiv:2412.01003, 2024a. Core Francisco Park, Andrew Lee, Ekdeep Singh Lubana, Yongyi Yang, Maya Okawa, Kento Nishi, Martin Wattenberg, and Hidenori Tanaka. Iclr: In-context learning of representations. In The Thirteenth International Conference on Learning Representations, 2025a. Kiho Park, Yo Joong Choe, and Victor Veitch. The linear representation hypothesis and the geometry of large language models, July 2024b. Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The geometry of categorical and hierarchi- cal concepts in large language models. 2025b. Ethan Perez, Sam Ringer, Kamil ̇ e Lukoši ̄ ut ̇ e, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Ben Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen- Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. Discovering language model behaviors with model-written evaluations, December 2022. Itamar Pres, Laura Ruis, Ekdeep Singh Lubana, and David Krueger. Towards reliable evaluation of behavior steering interventions in llms, October 2024. 13 Allan Raventós, Mansheej Paul, Feng Chen, and Surya Ganguli. Pretraining task diversity and the emergence of non-bayesian in-context learning for regression. Advances in Neural Information Processing Systems, 36, 2024. Shauli Ravfogel, Gilad Yehudai, Tal Linzen, Joan Bruna, and Alberto Bietti. Emergence of linear truth encodings in language models. arXiv preprint arXiv:2510.15804, 2025. Gautam Reddy. The mechanistic basis of data dependence and abrupt learning in an in-context classification task, 2023. URL https://arxiv.org/abs/2312.03002. Pranab Sahoo, Ayush Kumar Singh, Sriparna Saha, Vinija Jain, Samrat Mondal, and Aman Chadha. A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv preprint arXiv:2402.07927, 2024. Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan. Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. Trans- former Circuits Thread, 2024. URLhttps://transformer-circuits.pub/2024/ scaling-monosemanticity/index.html. Joshua B Tenenbaum, Charles Kemp, Thomas L Griffiths, and Noah D Goodman. How to grow a mind: Statistics, structure, and abstraction. science, 331(6022):1279–1285, 2011. Ian Tenney, Dipanjan Das, and Ellie Pavlick. Bert rediscovers the classical nlp pipeline. arXiv preprint arXiv:1905.05950, 2019. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Activation addition: Steering language models without optimization, October 2024. Tomer D Ullman and Joshua B Tenenbaum. Bayesian models of conceptual development: Learning as building models of the world. Annual Review of Developmental Psychology, 2(1):533–558, 2020. Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov. Transformers learn in-context by gradient descent. In International Conference on Machine Learning, pages 35151–35174. PMLR, 2023. Edward Vul and Harold Pashler. Measuring the crowd within: Probabilistic representations within individuals. Psychological Science, 19(7):645–647, 2008. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382, 2023. Daniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka, Gautam Reddy, and Noah D. Goodman. In-context learning strategies emerge rationally, June 2025. Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma. An explanation of in-context learning as implicit bayesian inference. arXiv preprint arXiv:2111.02080, 2021. Shu Yang, Shenzhe Zhu, Liang Liu, Lijie Hu, Mengdi Li, and Di Wang. Exploring the personality traits of llms through latent features steering, February 2025. Yida Yin, Zekai Wang, Yuvan Sharma, Dantong Niu, Trevor Darrell, and Roei Herzig. In-context learning enables robot action prediction in llms. arXiv preprint arXiv:2410.12782, 2024. 14 Yufeng Zhang, Fengzhuo Zhang, Zhuoran Yang, and Zhaoran Wang. What and how does in-context learning learn? bayesian model averaging, parameterization, and generalization. arXiv preprint arXiv:2305.19420, 2023. Wentao Zhu, Zhining Zhang, and Yizhou Wang. Language models represent beliefs of self and others. arXiv preprint arXiv:2402.18496, 2024. Amir Zur, Eric Bigelow, Atticus Geiger, and Ekdeep Singh Lubana. Are language models aware of the road not taken? token-level uncertainty and hidden state dynamics. ICML workshop on actionable interpretability, 2025. 15 Appendices ADERIVATIONS A.1DERIVATION OF THE BAYES FACTOR To model LLM in-context learning behavior, we examine the dynamics of the posterior belief of a model in a conceptc,p(c|x), as context length|x| = Nis varied. To study this, we consider the posterior odds between the concept c and its complement c ′ : o(c|x) = p(c|x) p(c ′ |x) . The posterior odds represent the model’s posterior belief incversus its posterior belief inc ′ after seeing context x, and given prior preference. This can be further decomposed as follows. logo(c|x) = log p(c) p(x|c) p(c ′ ) p(x|c ′ ) = log p(c) p(c ′ ) + log p(x|c) p(x|c ′ ) To model the log posterior odds, we must capture both the prior and likelihood-related terms. We discuss our model of prior odds in Sec. 4.2. To compute the log-likelihoods, we make two crucial assumptions, as follows. 1. Concept log-likelihood declines proportionally to the number of mismatched labels: The persona-adoption settings we examine consist of query-label examples where labels are binary and either map or do not map to a persona. Thus, it is reasonable that the log-likelihood for a concept will decline proportionally with the number of mismatched labels seen. Assuming this likelihood function follows Goodman et al. (2008), who studied rule-based concept learning in humans. Formally, we can express the likelihood function for a concept c as: logp(x|c)∝−|i∈1,...,N| l i ̸= y (c) i |, wherel i is a seen label andy (c) i is the persona-consistent label for the in-context query i. Since in the settings we study all labels are consistent with the persona, that is,l i = y (c) i ,∀ i∈1,...,N, we infer: logp(x|c)∝−|i∈1,...,N|l i ̸= y c i | = 0, and logp(x|c ′ )∝−|i∈1,...,N|l i ̸= y c ′ i | =−N. 2. Log-likelihood scales as a power-law with number of in-context examplesN: This assumption aims at accommodating the power-law behavior observed in studies of LLM in-context learning (Anil et al., 2024; Liu et al., 2024b). Specifically, we assume the common form of a scaling-law from scaling laws predicting loss during pretrainingL(n)≈ L(∞) + A n α (Kaplan et al., 2020). However, in our case,L(N)represents the negative-log likelihood for theN-th in-context example (Anil et al., 2024). Given this assumption, we derive a sub-linear discount termτthat arises from the ratio between the negative log- likelihood forNin-context examples under the power-law assumption, and the negative log-likelihood given by an optimal Bayesian agent using the likelihood function from assumption 1. Following the derivation from Wurgaft et al. (2025), we write: 16 τ := NLL under power-law scaling for N in-context examples NLL given by a Bayesian Learner for N in-context examples = P N n=1 (L(n)− L(∞))δn N = 1 N N X n=1 A n α δn = AN −α Z 1 0 1 ˆn α δˆn = A 1− α N −α = γN −α whereˆn = n /Nandγ = A 1−α is a constant that incorporatesA, the constant from our power-law form. Final expression for Bayes Factor.Following the assumptions above, we can write the functional form for the log Bayes-factor as: log p(x|c) p(x|c ′ ) = logp(x|c)− logp(x|c ′ ) ≈ γN −α (−|i∈1,...,N|l i ̸= y (c) i | +|i∈1,...,N|l i ̸= y (c ′ ) i |) = γN 1−α . 17 A.2DERIVATION OF THE EFFECT OF STEERING MAGNITUDE (EQ. 7) Here we show how the log posterior odds (Eq. 7) can be represented as: log p(c i | v + m· d i ) p(c ′ i | v + m· d i ) = log p(c i | v) p(c ′ i | v) + a· m or equally: log p(c i | v + m· d i ) p(c ′ i | v + m· d i ) = log p(v | c i ) p(v | c ′ i ) + log p(c i ) p(c ′ i ) + a· m Note that the last term does not depend on v. Recall that a given vector embeddingvis defined, according to the Linear Representation Hypothesis, as a linear weighted sum of concept vectorsd i weighted byβ i (v), i.e. how much conceptc i is present in v: v = X i β i (v) d i with the constraint that concept vectors are approximately orthogonal, i.e. d T i d j ≈ 0. Next, the conditional probability is given by p(c i | v) = σ −w T i v− b = σ (η) where η =−w T i v− b. We further assume that our weight vector w approximates concept vector d i scaled by an arbitrary value k: w ≈ k d i Now, consider a shifted representationv +m·d i , where we substitutew → k d i andv → v +m·d i : p(c i | v + m· d i ) = σ −k d T i (v + m· d i )− b = σ −k d T i v− b− k m∥d i ∥ 2 This shows a linear effect of steering magnitude m in logit space. Next, we can represent the log posterior odds as e η : p(c i | v) p(c ′ i | v) = p(c i | v) 1− p(c i | v) = σ(η) 1− σ(η) = 1/(1 + e η ) 1− 1/(1 + e η ) = 1/(1 + e η ) e η /(1 + e η ) = e −η Mapping this into log space, we get: log p(c i | v) p(c ′ i | v) =−η = w T i v + b 18 Next, we define a new term a = 1 2 ∥d i ∥ 2 and, using our previous theorems, substitute as follows: −η = w T i v + b = d T i v + bSubstitute w ≈ d i = d T i X j β j (v)d j + bL.R.H definition = β i (v)∥d i ∥ 2 + bd T i d j ≈ 0, ∀i̸= j = aβ i (v) + bDefinition of a Finally, we define the log posterior odds when steering v by m· d i as: log p(c i | v + m· d i ) p(c ′ i | v + m· d i ) =−η steered Steering changes v → v + m· d i , and thus η steered = d T i (v + m· d i ) + bSubstituting v in η = d T i v + m∥d i ∥ 2 + b = d T i v + a· mDefinition of a = η + a· mDefinition of η Finally, we obtain log p(c i | v + m· d i ) p(c ′ i | v + m· d i ) = log p(c i | v) p(c ′ i | v) + a· m 19 BMAIN RESULTS ACROSS MODELS We find that our account remains highly predictive across three models, with average correlations ofr = 0.98for Qwen-2.5-7B andr = 0.97for Gemma-2-9B computed across the entire heatmap (Fig 8), and correlations ofr = 0.91for Qwen-2.5-7B andr = 0.97for Gemma-2-9B for prediction ofN ∗ (the phase boundary; Fig. 11). Note also that all correlation values are computed for held-out predictions. Furthermore, we find that our predictions regarding the influence of in-context learning and steering are corroborated (Fig. 9, Fig. 10). *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * $! $! $' ( " *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * ") $) *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * !""" *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * (" " *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * '"!"$! " * *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * $! $! $' ( " *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * ") $) *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * !""" *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * (" " *&#&#&%&##& '! $" % * & * * * * # * * & * * * & * * # * * & * % * '"!"$! " * HEATMAPS - QWEN Qwen-2.5-7B Bayesian Model ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! &" &" &) * # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! #+ &+ ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , HEATMAPS - GEMMA Gemma-2-9B Bayesian Model ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! &" &" &) * # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! #+ &+ ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , Figure 8: In-context learning and activation steering jointly affect behavior. Results presented in Fig. 6 replicate across Qwen-2.5-7B and Gemma-2-9B models, showing the generalizability of the belief dynamics model. 20 **&*"**& ' $! ** *& * *# * * ! $ ( ( ! $ $' &* * ** * &* ** ' $! !) $) ** ' $! !!! **&*%* ' $! (! ! ** ' $! ' ! !$ ! ICL CURVES - QWEN )%") & #! )) )% ) )" ) ) ! # ' ' ! # #& ) ) )) ) ) )) & #! !( #( )%" & #! !!! )) & #! '! ! ))%)$) & #! & ! !# ! ICL CURVES - GEMMA Figure 9: In-context learning curves in Qwen-2.5-7B (top) and Gemma-2-9B (bottom). STEER RESPONSE FUNC - QWEN Bayesian Model Qwen-2.5-7B #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% #$($# " "% & #$($# " "% % " #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% #$($# " "% & #$($# " "% % " (((((( " "% (( ($ ( (! ( ( " & & % " (( #( !( (((((( " "% ' "' (((((( " "% (((((( " "% & (((((( " "% % " STEER RESPONSE FUNC - GEMMA Bayesian Model Gemma-2-9B (((((( " "% (( ($ ( (! ( ( " & & % " (( #( !( (((((( " "% ' "' (((((( " "% (((((( " "% & (((((( " "% % " Figure 10: Steering magnitude response function in Qwen-2.5-7B and Gemma-2-9B. 21 N* PREDS- GEMMA +'$'$'&'$$' ( ! %" + + # + + & + + + & + + # + %! %! %( ) " +'$'$'&'$$' ( ! %" + + # + + & + + + & + + # + "* %* +'$'$'&'$$' ( ! %" + + # + + & + + + & + + # + !""" + (a) Gemma-2-9B N* PREDS- QWEN )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) # # #& ' ! )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) !( #( )%"%"%$%""% & #! $) %) ) ) )" ) )% )) )% ) )" ) ) %) $) !!! ) (b) Qwen-2.5-7B Figure 11: The belief dynamics model captures cross-over pointsN ∗ across different language models. CFULL STEERING RANGE The results discussed in our main text focus on the case where the Linear Representation Hypothesis (LRH) holds. However, we empirically find that with larger enough steering magnitudes, the linear effect of steering onlogo(c|x)begins to break down and the sigmoidal response function we show in Fig. 5 converges towards 0 (Fig. 12 and Fig. 13). This is similar to the findings of Panickssery et al. (2024) which shows that LLM behavior begins to break down and become incoherent with very large magnitude steering vectors. We find that behavior converges towards chance (p(y|x) = 0.5), even with very large context lengths. Different datasets have different thresholds formwhich cause behavior to break down (Fig. 13). For Llama-3.1-8b, this magnitude threshold is larger for Narcissism than other datasets. As shown in Fig. 4, Narcissism has less effect from steering with small magnitudes compared to the other 4 datasets, and also has a later transition pointN ∗ . These results may be together explained by Narcissism having a weaker signal for the target conceptcthrough the likelihoodp(x|c), which results in both in-context learning and steering having comparatively less impact on belief compared to datasets with a stronger signal. 22 STEER RESPONSE FUNC FULL RANGE- LLAMA #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% (((((( " "% & (((((( " "% % " #$($# " "% (( ($ ( (! ( ( " & & % " (( #( !( #$($# " "% ' "' #$($# " "% (((((( " "% & (((((( " "% % " ((( " "% (( ($ ( (! ( ( " & & % " (( #( !( ((( " "% ' "' ((( " "% ((( " "% & ((( " "% % " ((( " "% (( ($ ( (! ( ( " & & % " (( #( !( ((( " "% ' "' ((( " "% ((( " "% & ((( " "% % " Llama-3.1-8B Bayesian Model *&#&#&%&##& '! $" * * & * * * * * * * & * * * $! $! $' ( " *&#&#&%&##& '! $" * * & * * * * * * * & * * * ") $) *&#&#&%&##& '! $" * * & * * * * * * * & * * * !""" *&#&#&%&##& '! $" * * & * * * * * * * & * * * (" " *&#&#&%&##& '! $" * * & * * * * * * * & * * * '"!"$! " * *&#&#&%&##& '! $" * * & * * * * * * * & * * * $! $! $' ( " *&#&#&%&##& '! $" * * & * * * * * * * & * * * ") $) *&#&#&%&##& '! $" * * & * * * * * * * & * * * !""" *&#&#&%&##& '! $" * * & * * * * * * * & * * * (" " *&#&#&%&##& '! $" * * & * * * * * * * & * * * '"!"$! " * ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , LRH Holds LRH Breaks DownLRH HoldsLRH Breaks Down ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , Figure 12: With large enough magnitudes, the Linear Representation Hypothesis breaks down Our belief dynamics model is able to explain model behavior within a limited range ofm. When steering magnitudes exceed this range, behavior begins to break down and converges to chance (p(y|x) = 0.5). 23 STEER RESPONSE FUNC FULL RANGE RAW DATA- LLAMA LinearNon-LinearLinearNon-Linear *&#&#&%&##& '! $" * * & * * * * * * * & * * * $! $! $' ( " *&#&#&%&##& '! $" * * & * * * * * * * & * * * ") $) *&#&#&%&##& '! $" * * & * * * * * * * & * * * !""" *&#&#&%&##& '! $" * * & * * * * * * * & * * * (" " *&#&#&%&##& '! $" * * & * * * * * * * & * * * '"!"$! " * ((( " "% (( ($ ( (! ( ( " & & % " (( #( !( ((( " "% ' "' ((( " "% ((( " "% & ((( " "% % " (a) Llama-3.1-8B STEER RESPONSE FUNC FULL RANGE RAW DATA- GEMMA *&#&#&%&##& '! $" * * & * * * * * * * & * * * $! $! $' ( " *&#&#&%&##& '! $" * * & * * * * * * * & * * * ") $) *&#&#&%&##& '! $" * * & * * * * * * * & * * * !""" *&#&#&%&##& '! $" * * & * * * * * * * & * * * (" " *&#&#&%&##& '! $" * * & * * * * * * * & * * * '"!"$! " * ((( " "% (( ($ ( (! ( ( " & & % " (( #( !( ((( " "% ' "' ((( " "% ((( " "% & ((( " "% % " (b) Gemma-2-9B ((( " "% (( ($ ( (! ( ( " & & % " (( #( !( ((( " "% ' "' ((( " "% ((( " "% & ((( " "% % " *&#&#&%&##& '! $" * * & * * * * * * * & * * * $! $! $' ( " *&#&#&%&##& '! $" * * & * * * * * * * & * * * ") $) *&#&#&%&##& '! $" * * & * * * * * * * & * * * !""" *&#&#&%&##& '! $" * * & * * * * * * * & * * * (" " *&#&#&%&##& '! $" * * & * * * * * * * & * * * '"!"$! " * STEER RESPONSE FUNC FULL RANGE RAW DATA- QWEN (c) Qwen-2.5-7B Figure 13: Different datasets have different thresholds for steering breaking down. Different datasets have different thresholds for what steering magnitudesmwill predictably steer model behavior. 24 DMANY-SHOT STEERING VECTOR COMPUTATION Steering vectors are usually computed over a single query with different targets (in our case, a question and a "Yes/No" answer). As an exploratory experiment, we tested steering vector computation while varying number of shots. Interestingly, we find that steering vector norm substantially increases after the first shot, then slowly increases in most layers as additional context is added. We find that normalizing the steering vectors computed across shots to have the norm at shot 0 yields a weaker effect than a 0-shot steering vector, though the effect becomes slightly stronger with number of shots (Fig. 14, Top). Additionally, we find that cosine similarity with the 0-shot vector drops suddenly as another example is added, and similarity with the 128-shot vector slowly increases with context (Fig. 14, Bottom). MULTISHOT STEERING - LLAMA 01 816 128 ! " # $ % & ' ( ) "! "# "% "' #! #% #) $# %! %) &' '% )! *' ""# "#) ??? A$@! A#@! A"@! A!@) A!@' A!@% A!@# !@! !@# !@% !@' !@) "@! #@! $@! ? ? !" ! " # $ % & ' ( ) "! "# "% "' #! #% #) $# %! %) &' '% )! *' ""# "#) ??? A$@! A#@! A"@! A!@) A!@' A!@% A!@# !@! !@# !@% !@' !@) "@! #@! $@! ? ? !" ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, &" &" &) * # ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, #+ &+ ,(%(%('(%%( )" &# '!, (!, !, ,! ,!% ,! ,!( ,!, ,!( ,! ,!% ,! !, (!, '!, "### ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! *# # ,(%(%('(%%( )" &# ! ,! ,!$ ,! ,!' ,! ,! ,!' ,! ,!$ ,! ! )#"#&" # , ! " # $ % & ' ( ) "! "# "% "' #! #% #) $# %! %) &' '% )! *' ""# "#) ??? A$@! A#@! A"@! A!@) A!@' A!@% A!@# !@! !@# !@% !@' !@) "@! #@! $@! ? ? !" ! " # $ % & ' ( ) "! "# "% "' #! #% #) $# %! %) &' '% )! *' ""# "#) ??? A$@! A#@! A"@! A!@) A!@' A!@% A!@# !@! !@# !@% !@' !@) "@! #@! $@! ? ? !" ! " # $ % & ' ( ) "! "# "% "' #! #% #) $# %! %) &' '% )! *' ""# "#) ??? A$@! A#@! A"@! A!@) A!@' A!@% A!@# !@! !@# !@% !@' !@) "@! #@! $@! ? ? !" "! ! "! " "! # 333 33 3 !4! !4# !4$ !4% !4& "4! 3 33!53 33"#&53 Number of Shots Used to Compute Vector Figure 14: Computing Steering Vectors Over Varying Number of Shots. Steering vectors are computed for Llama-3.1-8B for the Psychopathy dataset over varying number of shots. Note that here 0-shot refers to providing the model with a single target query and a "Yes/No" reply and taking the difference in mean activations, whereas a larger number of shots refers to the number of in-context examples provided to the model before the target query. Top panel shows the effect of steering vectors computed over varying number of shots and applied at different magnitudes and context lengths. Bottom panel shows cosine similarity between vectors computed over varying number of shots with the 0-shot vector or the 128-shot vector. 25 EEXPERIMENTAL DETAILS Implementation DetailsIn our experiments, we use Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B. For efficiency reasons, we restrict our analyses to LLMs which balance relatively small scale (∼8 billion parameters) with relatively high performance on major benchmarks. We use 4-bit quantization for further efficiency, and run inference locally, primarily on A100 GPUs. Steering vector training and application are implemented using an open-source repository 3 for LLM steering, which implements Contrastive Activation Addition (Turner et al., 2024). Parameters Varied in ExperimentsFor our experiments, we first tested LLMs with a smaller set of experiment parameters to find the optimal steering layer (Fig. 15). After identifying the optimal steering layer we proceeded with a larger experiment: for each LLM and each of our 5 datasets, we tested models with 33 increments ofm, ranging from[−10, +10]with0.1step increments between m∈ [−1, +1], as well as both positive and negative magnitudes for:[10, 5, 3, 2.5, 2, 1.5]. We steered models using activation addition at the optimal steering layerℓ ∗ . For number of shotsN, we used N = 0, 1, 2, 3, 4, 5, 6, 7, 8, 10, 12, 14, 16, 20, 24, 28, 32, 40, 48, 56, 64, 80, 96, 112, 128. In each case, we randomly sampled 100 sequences of in-context exemplarsxas well as a random target question. Steering Effect by LayerTo find the optimal steering layerℓ ∗ , we first tested LLMs with a smaller set of experiment parameters, testing each LLM with across every 2 layers with steering magnitudes m = [−1, 0, +1]. In the models we use, we consistently find 1 particular layer for which steering is most effective (see examples in Fig. 15). For Llama-3.1-8B, this is consistently layer 12, for Gemma-2-9B, it is layer 20, and for Qwen-2.5-7B, it is layer 14. These are the layers we use for our primary experiments, which systematically vary steering magnitude and context length. Model Fitting We fit the4free parameters (α,γ,a,b) of the Bayesian model for each dataset, modelcombination, which consists of evaluations with29different steering magnitudes, each across25in-context shot values. Given that we aim at capturing a population-level behavior (rather than behavior in an individual context), for each number of shots we average LLM probabilities for the persona-consistent answer across the 100 sampled sequences, and fit these per-shot-number averages, yielding a set of 725 values for eachdataset, model combination. For optimization, we use the L-BFGS-B algorithm provided via the Scipy library’s optimize function, with1000as the maximum number of iterations, and10 −10 gradient and function tolerances. We use Binary Cross entropy loss between probabilities given by the Bayesian model and the LLM for the persona-consistent answer, and apply Pytorch’s automatic differentiation to compute gradients for updating Bayesian model parameters. In order to find good initial parameters for optimization, we conduct basin hopping search with1000iterations, run optimization for the100best candidates, and use the top result in terms of loss. Given that LLMs adopt the persona behaviors tested after relatively few shots, yielding long plateaus around probability of1. To soften the effect of this imbalance, we binlog 2 (N)values (withNdenoting number of shots) to 15 bins, and the loss for values in each bin was multiplied by 1 # shots in bin . The fitting results shown in Fig. 4, Fig. 6, Fig. 7, Fig. 8, Fig. 9, and Fig. 11 represent held-out predictions using 10-fold cross-validation, where for each fold we held out data for3adjacent magnitude values (except one fold which contains2adjacent magnitudes) and predicted data for these held-out magnitude values. Overall, we find a very high correlation between LLM probabilities and predictions on held-out data (r = 0.98, averaged across our 5 domains). In Fig. 5, Fig. 10, Fig. 12, and Fig. 13, in which we show the magnitude response function across different magnitude values, we show Bayesian model results for models fitted to the entire heatmap. Miscellaneous details Given that our experiment includes0-shot evaluations, in cases where we plot number of shots in log-scale, we code N = 0 as N = 0.6 only for plotting purposes. 3 https://github.com/steering-vectors/steering-vectors 26 0510152025303540 Layer 0.2 0.1 0.0 0.1 Steering Effect gemma-2-9b - Machiavellianism +1 Magnitude -1 Magnitude 0510152025303540 Layer 0.4 0.3 0.2 0.1 0.0 gemma-2-9b - Subscribes-to-moral-nihilism 051015202530 Layer 0.2 0.1 0.0 0.1 0.2 0.3 Steering Effect llama-3.1-8b - Machiavellianism 051015202530 Layer 0.3 0.2 0.1 0.0 0.1 0.2 0.3 llama-3.1-8b - Psychopathy 051015202530 Layer 0.10 0.05 0.00 0.05 0.10 Steering Effect llama-3.1-8b - Narcissism Figure 15: Examples of Steering effect by layer from Llama and Gemma Mean effect of steering with CAA vectors computed for each of every 2 layers in the model, given a context length|x| = 1. 27