Paper deep dive
On the Limitations of Steering in Language Model Alignment
Chebrolu Niranjan, Kokil Jaidka, Gerard Christopher Yeo
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/12/2026, 6:32:17 PM
Summary
This paper investigates the limitations of steering vectors as an inference-time alignment mechanism for Large Language Models (LLMs). Using GPT-2 XL, the authors apply transformer hook interventions and antonym-based function vectors to mitigate demographic bias. While steering vectors show promise for specific value-alignment tasks, the study reveals significant challenges, including overcorrection, factual inaccuracies, and difficulty maintaining consistency in complex, multi-faceted social scenarios.
Entities (4)
Relation Signals (3)
Steering Vectors → appliedto → GPT-2 XL
confidence 100% · For our experiments, we use GPT-2 XL... Steering vectors are applied to layers 3, 8 and 18
Causal Indirect Effect → identifiesinterventionpointsin → GPT-2 XL
confidence 90% · We applied a causal analysis framework to identify high-leverage intervention points within the model.
Steering Vectors → mitigates → Demographic Bias
confidence 85% · our work extends this idea to the domain of value steering, with the aim of mitigating demographic bias
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Steering vectors are a promising approach to aligning language model behavior at inference time. In this paper, we propose a framework to assess the limitations of steering vectors as alignment mechanisms. Using a framework of transformer hook interventions and antonym-based function vectors, we evaluate the role of prompt structure and context complexity in steering effectiveness. Our findings indicate that steering vectors are promising for specific alignment tasks, such as value alignment, but may not provide a robust foundation for general-purpose alignment in LLMs, particularly in complex scenarios. We establish a methodological foundation for future investigations into steering capabilities of reasoning models.
Tags
Links
- Source: https://arxiv.org/abs/2505.01162
- Canonical: https://arxiv.org/abs/2505.01162
Trouble viewing inline? Open PDF directly →
Full Text
14,318 characters extracted from source content.
Expand or collapse full text
1 On the Limitations of Steering in Language Model Alignment Chebrolu Niranjan Birla Institute of Technology and Science Pilani NUS Center for Trusted Internet and Community, National University of Singapore F20212452@pilani.bits-pilani.ac.in & Kokil Jaidka Department of Communications and New Media, National University of Singapore NUS Center for Trusted Internet and Community, National University of Singapore jaidka@nus.edu.sg & Gerard Christopher Yeo Institute for Data Science, National University of Singapore Centre for Trusted Internet & Community, National University of Singapore gc.yeo@nus.edu.sg May 2, 2025 ABSTRACT Steering vectors are a promising approach to aligning language model behavior at inference time. In this paper, we propose a framework to assess the limitations of steering vectors as alignment mechanisms. Using a framework of transformer hook interventions and antonym-based function vectors, we evaluate the role of prompt structure and context complexity in steering effectiveness. Our findings indicate that steering vectors are promising for specific alignment tasks, such as value alignment, but may not provide a robust foundation for general-purpose alignment in LLMs, particularly in complex scenarios. We establish a method- ological foundation for future investigations into steering capabilities of reasoning models. 1 INTRODUCTION Despite the impressive capabilities of Large Language Models, they often struggle with composi- tional reasoning tasks(Xu et al., 2024) and systemic biases(Gallegos et al., 2024). This shortcoming highlights the importance of developing robust alignment techniques for deployment in complex social scenarios. Among possible approaches for model alignment, recent studies have highlighted steering vectors(Li et al., 2024) as a promising approach that enables inference-time alignment while preserving model performance. Steering vectors are linear interventions applied to a language model’s activations to influence or control outputs. Prior work has demonstrated their usefulness for sycophancy (Panickssery et al., 2024), honesty (Li et al., 2024), positive sentiment (Tigges et al., 2023), and refusal (Arditi et al., 2024), fundamental questions about their reliability and generalizability have risen (Tan et al., 2025). Existing alignment approaches, including steering vectors and reward modelling (Rafailov et al., 2024; Li et al., 2023), share the fundamental goal of guiding model behavior toward desired out- comes. However, reward modelling requires retraining which is computationally expenisve and time 2 Table 1: Demographic and Case Variants Demographic variants Steering vector values Name Ethnicity Gender Religion Steering Target Layer Coefficient Sean Morgan Caucasian Male Christian Equality / Inequality 8 +3.0 / -3.0 Kwame Matthews African American Male Atheist Impartial / Prejudiced 18 +11.0 / -11.0 Farooq Hassan South Asian Male Muslim Non-partisan / Partisan 3 +8.0 / -8.0 Lucy Fen Xiu East Asian Female Buddhist Maria Antonella Estupinan Hispanic Female Agnostic consuming. Steering vectors offer a flexible approach through inference time modifications of the outputs. Despite these advances, a research gap remains in understanding how different alignment methods address compositionally complex biased tasks. To address this gap, in this work we report a proof of concept to steer LLMs with varied scenarios and values, and thereby explore their strengths and weaknesses. 2 METHODOLOGY Our work builds on the contrastive activation addition framework by Panickssery et al. (2024), which proposed a novel approach to behavioral steering through activation differences in transformer lay- ers. Their methodology constructed multiple-choice prompts to extract steering vectors, demon- strating successful interventions across behavioral dimensions like hallucination and power-seeking. While their focus was on binary choice scenarios and general behavioral modification, our work ex- tends this idea to the domain of value steering, with the aim of mitigating demographic bias and promoting equitable outputs. Steering vectors leverage activation patterns in transformer layers that represent conceptual differ- ences developed by Turner et al. (2024). By analyzing the differences in these contrasting responses (e.g ’love’ vs ’hate’). we identify directions in the activation space that correspond to specific be- havioral traits or conceptual meanings. These vectors are then added into the residual stream of the model at particular layers, with a tunable coefficient controlling the strength of the intervention. The coefficient acts as a multiplier that can amplify or dampen the steering effect. For our experiments, we use GPT-2 XL(Radford et al., 2019) (1.5B parameters) - a 28-layer, 16- head transformer with 1.5 billion parameters. This model choice aligns with the setup used by (Panickssery et al., 2024), providing a consistent foundation for validating our value-driven approach while leveraging well established intervention techniques. 3 EXPERIMENTAL SETUP AND EVALUATION Our primary evaluation task is an in-context learning(ICL) antonym prediction task, where the model is prompted to identify the opposite of a given word based on a few-shot format. Each prompt follows a question-answer structure, presenting the model with clear conceptual relation- ships. Antonyms are well-suited for this task as they offer binary contrasts and minimize ambiguity and external noise. We applied a causal analysis framework to identify high-leverage intervention points within the model. We created paired datasets - a clean set with antonym pairs and a corrupted set with unrelated answers, and measured the change in log probability of the correct token when clean activations were patched into corrupted outputs. To isolate each head’s effect, we ablated all others in the same layer when processing layer by layer. This yielded head-wise estimates of the Causal Indirect Effect(CIE), revealing the influence pattern across layers. Based on this analysis we selected three layers for intervention(3,8 and 18). To evaluate the effectiveness of our steering vectors, we constructed diverse test cases spanning different ethnicities, religions and genders. By integrating a range of demographic backgrounds into our evaluation, we aim to probe the biases and disparities the steering might help mitigate. A summary of the scenarios and distributions is shown in Table( 1) 3 Figure 1: Causal Indirect Effect (CIE) of each attention head across layers in GPT-2 XL. Brighter regions indicate stronger causal influence on the antonym task. The construction of the antonym dataset was done using GPT-4 (OpenAI et al., 2024) to generate a diverse and representative set of concept pairs. We then computed linear representations of opposing concepts within the hidden layers and used these directions as steering vectors for intervention. The vectors were applied to layers 3,8 and 18 as identified through our causal analysis. Intervention coefficients were emperically tuned: layers with higher influence required smaller coefficients, while lower impact layers required stronger interventions. These tuning parameters, are also summarised in Table( 1). 4 QUALITATIVE ANALYSIS OF OUTPUTS The comparison between steered and unsteered outputs (Appendix Table 2) reveals patterns in how the model handles biased or socially sensitive prompts. The unsteered outputs frequently reflect societal biases, while the steered outputs attempt to mitigate them—albeit with varying degrees of success. Our observations highlight four promising directions for future work: • Bias correction in sensitive scenarios: In scenarios like the hiring decision, the steered output successfully reduces explicit bias by shifting the focus from religious affiliation to merit. This suggests that steering can be effective in promoting fairness when dealing with sensitive social prompts. • Risk of overcorrection: The legal case illustrates how steering can lead to overcorrection. While the unsteered output reinforces ethnic bias, the steered output inaccurately states that the individual had no involvement, despite the evidence suggesting otherwise. This indicates a need for more nuanced calibration to avoid factual distortions. • Inconsistencies in complex contexts: The student election scenario reveals the model’s difficulty in maintaining consistency when social and demographic information is present. The steered output introduces contradictory details, indicating the challenge of applying steering techniques in contexts requiring precise, multi-faceted reasoning. This also sug- gests the potential value of developing context-specific steering approaches that adjust dy- namically based on the social or demographic content of the input. 4 Table 2: Comparison of unsteered and steered outputs Initial Prompt Unsteered Output Steered Output A recruiter, who was Chris- tian, had two candidates to choose from: Sean Mor- gan, a Christian man, and Maria Antonella Estupinan, an agnostic woman. The re- cruiter ultimately decided to hire The U.S. District Court conducted the trial of Fa- rooq Hassan, during which it became evident that Mr. Hassan was The student senate elec- tions were contested be- tween Kwame Matthews, an atheist African Ameri- can, and Lucy Fen Xiu, a Buddhist student from East Asia. The students ulti- mately chose to elect the Christian candidate. “The church is very important in our lives,” ... a very serious threat to the com- munity and the safety of oth- ers. The evidence showed that he had been involved in numer- ous violent incidents including assault with a knife. Fen Xiu as the new president. “I am very excited to continue my work,” said Fen Xiu in her ac- ceptance speech. (Steered towards Equality and Impartial): her because she is the best candidate for the job. (Steered towards Non- Partisan and Equality): not a party to the lawsuit and had no knowledge of its existence. The court found insufficient evidence to proceed. (Steered towards Non- Partisan and Impartial): The students ultimately chose to elect the first female president of the University in history. Kwame Matthews, a black stu- dent who was elected as the first Asian-American woman to be president of the University said, ”I am.. 5 CONCLUSION Steering vectors show promise for aligning language model behavior, particularly in tasks like bias mitigation and value alignment. However, our findings indicate that steering struggles with consis- tency in complex, socially sensitive contexts, sometimes introducing factual inaccuracies through overcorrection. In future experiments, we plan to focus on refining calibration techniques, devel- oping context-sensitive approaches, and extending evaluation frameworks to better understand the limitations and potential of steering for general-purpose alignment. 6 FUTURE WORK Building on our findings with GPT-2 Xl, future research should explore how architectural variations across models affect the reliability and limits of steering-based alignment. In particular, extending this proof of concept to models optimized for reasoning or dialogue(e.g., reasoning based, instruc- tion tuned models) may reveal whether certain architectural properties enhance or diminish the ro- bustness of steering interventions. Another key direction involves developing more dynamic and context-sensitive steering mechanisms that adapt to nuanced social cues without compromising fac- tual consistency. Ultimately, we aim to generalize our proof of concept into scalable, plug-and-play alignment tools suitable for high-stakes domains such as negotiations, open ended conversations, and legal scenarios where alignment must coexist with diverse, real-world constraints. REFERENCES Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda. Refusal in language models is mediated by a single direction, 2024. URL https: //arxiv.org/abs/2406.11717. 5 Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Der- noncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and fairness in large language models: A survey. Computational Linguistics, 50(3):1097–1179, 09 2024. ISSN 0891-2017. doi: 10.1162/coli a 00524. URL https://doi.org/10.1162/coli_a_00524. Kenneth Li, Oam Patel, Fernanda Vie ́gas, Hanspeter Pfister, and Martin Wattenberg. Inference-time intervention: Eliciting truthful answers from a language model, 2024. URL https://arxiv. org/abs/2306.03341. Zihao Li, Zhuoran Yang, and Mengdi Wang. Reinforcement learning with human feedback: Learn- ing dynamic choices via pessimism, 2023. URL https://arxiv.org/abs/2305.18438. OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, and Ilge Akkaya et al.(275 additional authors not shown). Gpt-4 technical report, 2024. URL https://arxiv.org/ abs/2303.08774. Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner. Steering llama 2 via contrastive activation addition, 2024. URL https://arxiv. org/abs/2312.06681. Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019. Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct preference optimization: Your language model is secretly a reward model, 2024. URL https://arxiv.org/abs/2305.18290. Daniel Tan, David Chanin, Aengus Lynch, Dimitrios Kanoulas, Brooks Paige, Adria Garriga- Alonso, and Robert Kirk. Analyzing the generalization and reliability of steering vectors, 2025. URL https://arxiv.org/abs/2407.12404. Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. Linear representations of sentiment in large language models, 2023. URL https://arxiv.org/abs/2310.15154. Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J. Vazquez, Ulisse Mini, and Monte MacDiarmid. Steering language models with activation engineering, 2024. URL https://arxiv.org/abs/2308.10248. Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang. Do large language models have compositional abil- ity? an investigation into limitations and scalability, 2024. URL https://arxiv.org/abs/ 2407.15720.