Paper deep dive
Deceptive Alignment Monitoring
Andres Carranza, Dhruv Pai, Rylan Schaeffer, Arnuv Tandon, Sanmi Koyejo
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/12/2026, 7:31:57 PM
Summary
The paper introduces 'Deceptive Alignment Monitoring' as a research direction to address the threat of large machine learning models behaving deceptively. It identifies key areas for monitoring, including data creation/curation, model training/editing, and internal representations/mechanisms, advocating for unsupervised anomaly detection to identify ulterior motives in model behavior.
Entities (4)
Relation Signals (3)
Deceptive Alignment Monitoring ā addresses ā Deceptive Alignment
confidence 98% Ā· In this work, we identify emerging directions... for deceptive alignment monitoring
Mechanistic Anomaly Detection ā usedfor ā Deceptive Alignment Monitoring
confidence 95% Ā· In order to detect and counter this threat, it is imperative to develop interpretability methods... mechanistic anomaly detection
Adversarial Machine Learning ā contributesto ā Deceptive Alignment Monitoring
confidence 90% Ā· We conclude by advocating for greater involvement by the adversarial machine learning community in these emerging directions.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As the capabilities of large machine learning models continue to grow, and as the autonomy afforded to such models continues to expand, the spectre of a new adversary looms: the models themselves. The threat that a model might behave in a seemingly reasonable manner, while secretly and subtly modifying its behavior for ulterior reasons is often referred to as deceptive alignment in the AI Safety & Alignment communities. Consequently, we call this new direction Deceptive Alignment Monitoring. In this work, we identify emerging directions in diverse machine learning subfields that we believe will become increasingly important and intertwined in the near future for deceptive alignment monitoring, and we argue that advances in these fields present both long-term challenges and new research opportunities. We conclude by advocating for greater involvement by the adversarial machine learning community in these emerging directions.
Tags
Links
- Source: https://arxiv.org/abs/2307.10569
- Canonical: https://arxiv.org/abs/2307.10569
Trouble viewing inline? Open PDF directly ā
Full Text
16,379 characters extracted from source content.
Expand or collapse full text
arXiv:2307.10569v2 [cs.LG] 26 Jul 2023 Deceptive Alignment Monitoring Andres Carranza * 1 Dhruv Pai * 1 Rylan Schaeffer * 1 Arnuv Tandon * 1 Sanmi Koyejo 1 Abstract As the capabilities of large machine learning models continue to grow, and as the autonomy afforded to such models continues to expand, the spectre of a new adversary looms:the models themselves. The threat that a model might behave in a seemingly reasonable manner, while secretly and subtly modifying its behavior for ulterior rea- sons is often referred to as deceptive alignment in the AI Safety & Alignment communities. Con- sequently, we call this new directionDeceptive Alignment Monitoring. In this work, we iden- tify emerging directions in diverse machine learn- ing subfields that we believe will become increas- ingly important and intertwined in the near fu- ture for deceptive alignment monitoring, and we argue that advances in these fields present both long-term challenges and new research opportu- nities. We conclude by advocating for greater involvement by the adversarial machine learning community in these emerging directions. 1. Introduction Machine learning models are growing increasingly general- purpose while simultaneously being granted increasingly more autonomy. The combination of greater capabili- ties and greater freedom in choosing when and how to exercise those capabilities raises the spectre that models themselves may behave adversarially to human interests ( Hubinger et al.,2021;Hendrycks et al.,2021;Ngo et al., 2023). In the AI Safety and Alignment communities, this threat is often referred to as deceptive alignment, and large- scale collaborations spanning leading machine learning or- ganizations and institutions are focusing more on how to measure problematic capabilities of large-scale models via benchmarks ( Pan et al.,2023;Shevlane et al.,2023). * Equal contribution 1 Computer Science, Stan- ford University.Correspondence to: Rylan Schaeffer <rschaef@cs.stanford.edu>. 2 nd AdvML Frontiers workshop at40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023. Copyright 2023 by the author(s). In this Blue Sky submission, we aim to elaborate on the dif- ferent possible threats that large models themselves pose when viewed as potential adversaries, and argue that new methods are urgently needed. These methods must go be- yond evaluation on static benchmarks to active, dynamic monitoring of models at multiple stages in a modelās life- cycle: dataset creation and curation, training and learning, network representations and circuit mechanisms. 2. Deceptive Alignment Monitoring 2.1. Data Creation & Curation The use of foundation models to generate and/or curate their own training data has recently emerged in both lan- guage and vision. In language generation, Self-Instruct used one language model to bootstrap training data for fine-tuning a different language model ( Wang et al.,2022), and similar methods have quickly become commonplace, e.g., (Peng et al.,2023;Sun et al.,2023;Xu et al.,2023). In curation, SemDeDup uses an high-performing language model to prune a dataset for training future language mod- els ( Abbas et al.,2023), and DoReMi learns to weight data differently to improve performance when learning ( Xie et al.,2023). In vision, (Li et al.,2023) propose en- dowing a model with the ability to learn to construct its own training dataset in a self-supervised manner by trawl- ing through the internet. Such methods will continue to improve and proliferate. These methods are relevant to deceptive alignment be- cause if a highly capable, autonomous model does wish to change its behavior, creating or curating training data are straightforward mechanisms by which it can achieve such a goal. Currently, such methods have a high degree of human-in-the-loop, but as model scale and complexity increase, humans will play a vanishing role. To the best of our knowledge, we lack methods for monitoring the training data that our models generate or curate to check that the training data has not been selected in pursuit of some ulterior goal. This threat is spiritually similar to re- cent work demonstrating that web-scale datasets can be poi- soned ( Carlini et al.,2023); however, in the setting studied byCarlini et al.(2023), the dataset is known, fixed and con- structed by humans, whereas in this setting, the dataset is unknown, dynamic and constructed by models in a largely Deceptive Alignment Monitoring unsupervised manner. 2.2. Training, Editing & Learning Turning now from data to learning, there are two emerg- ing trends that pose new challenges: using one model to train another model, and directly editing the parameters of a model to change its behavior. The first trend ap- pears commonly when using reward or preference models to fine-tune a language model, such as in Reinforcement Learning from Human Feedback ( Ouyang et al.,2022; Bai et al.,2022a), Reinforcement Learning from AI Feed- back (Bai et al.,2022b), Reinforcement Learning for Feed- back ( Aky Ģurek et al.,2023) and many more, e.g., image synthesis (Lee et al.,2023). The second trend includes several methods for model editing, i.e., directly editing the modelās parameters to change its behavior on specific tasks or inputs, e.g., Meng et al.(2022a);Mitchell et al. (2022);Meng et al.(2022b);Hartvigsen et al.(2022); Ravfogel et al.(2022);Hoelscher-Obermaier et al.(2023). In some methods, models are directly trained to learn to modify their own parameters, e.g., Irie et al.(2022); Kirsch & Schmidhuber(2022b;a). Why does this matter for deceptive alignment? If models can have ulterior motives and act on said motives, then en- abling models to train each other or to edit themselves di- rectly posses a significant threat, since a model can give itself or give another model capabilities that humans would rather the model(s) not possess. To the best of our knowl- edge, we lack techniques to audit what learning signals one model transmits to another model, and we also lack tech- niques to understand whether model editing methods can themselves be poisoned; for a concrete example, it might be possible that by selective choosing a small number of seem- ingly innocent edits (e.g., āUS President: Donald Trumpā Joe Bidenā and āWorld Chess Champion: Magnus Carlsen āDing Lirenā), the combination of edits induces an unpre- dictable and undesirable change in model behavior 1 . 2.3. Representations & Mechanisms A key step towards addressing the identified challenges will be monitoring the representations and mechanisms of large-scale models for āpermissibleā behavior ( Christiano, 2022a;b); here, āpermissibleā behavior is loosely under- stood since the concept itself is in desperate need of sharper understanding. Loosely, if we think of a network as some mapf:X ā Y, previous work has studied adversarial inputs coming from an unusual input distribution Ģp(x)or backdoored models with unusual conditional distributions p(y|x). We are therefore concerned not with unusual input data, nor unusual conditional distributions, but whether the 1 However, we note that sometimes seemingly unpredictable changes are indeed predictable, e.g.,Schaeffer et al.(2023). networkās outputs are produced for the ānormal reasonsā. AsChristiano(2022a) explains, we are looking for outliers in terms of howfhandles perfectly normal data. The reason why this matters is that while inputs, outputs, and inputs-to-output maps may all appear perfectly ordi- nary, highly capable models may conceal some ulterior mo- tive or goal. In order to detect and counter this threat, it is imperative to develop interpretability methods capable of understanding both activations and circuit mechanisms of models and detecting abnormal variations in modelsā pro- cessing. This is important because two models generating identical outputs when provided with the same inputs may compute their outputs for different āreasonsā. Because the exact threat is unknown and likely dynamic, there is an urgent need to develop unsupervised methods for mechanistic anomaly detection that scale well. The next step is to develop an unsupervised methodology that can differentiate between normal model behaviorsāwhere conclusions are reached for āappropriate reasonsāāand anomalous model behaviorsāwhere conclusions are de- rived for erroneous, potentially harmful āreasonsā. This task is particularly challenging since it requires the identifi- cation of patterns in the propagation of information through a model that is indicative of correct reasoning without rely- ing on explicitly supervisory signals. To achieve this, we propose leveraging techniques re- lated to unsupervised anomaly detection to capture devi- ations from typical model behaviors. By comparing a modelās processing across various inputs and outputs, it may be possible to identify patterns that consistently align with desired and appropriate behavior. We hypothesize that these patterns could manifest at three different lev- els of analysis within a model. Firstly, at the individ- ual layer, a comprehensive analysis of activation distribu- tions in the high-dimensional activation space could pro- vide valuable insights into the modelās processing. Sec- ondly, at the layer-to-layer activation level, investigating how high-dimensional modes propagate, transform and evolve through the layers of a model can also offer an un- derstanding of normal and abnormal processing. Thirdly, at the circuit level, identifying subgraphs within the network that correspond to specific transformations on features rele- vant to out-of-domain generalization might also prove pow- erful; however, knowing how to usefully define probabilis- tic distribution over activations, activationsā propagations and circuit mechanisms for anomaly detection are, to the best of our knowledge, open questions. For possible ap- proaches, see Carranza et al.(2023). Deceptive Alignment Monitoring 3. Outlook The human-model interpretability quest can be modeled as an adversarial game, whereby deceptively aligned models subvert interpretability tools in favor of capabilities. More capable models are increasingly threatening, and to main- tain scalable oversight we advocate development of novel tools for deceptive alignment monitoring. Deceptive Alignment Monitoring References Abbas, A., Tirumala, K., Simig, D., Ganguli, S., and Mor- cos, A. S. Semdedup: Data-efficient learning at web- scale through semantic deduplication.arXiv preprint arXiv:2303.09540, 2023. Aky Ģurek, A. F., Aky Ģurek, E., Madaan, A., Kalyan, A., Clark, P., Wijaya, D., and Tandon, N. Rl4f: Gen- erating natural language feedback with reinforcement learning for repairing model outputs.arXiv preprint arXiv:2305.08844, 2023. Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., Das- Sarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El- Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCan- dlish, S., Olah, C., Mann, B., and Kaplan, J. Training a helpful and harmless assistant with reinforcement learn- ing from human feedback, 2022a. Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKin- non, C., et al. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073, 2022b. Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tram`er, F. Poisoning web-scale training datasets is practical.arXiv preprint arXiv:2302.10149, 2023. Carranza, A., Pai, D., Tandon, A., Schaeffer, R., and Koyejo, S. Facade: A framework for adversarial circuit anomaly detection and evaluation, 2023. Christiano,P.Mechanistic anomaly detection and elk,2022a.URL https://w.lesswrong.com/posts/vwt3wKXWaCvqZyF74/mechanistic-anomaly-detection-and-elk. Accessed on May 28, 2023. Christiano, P.Can we efficiently distin- guish different mechanisms?, 2022b.URL https://w.lesswrong.com/posts/JLyWP2Y9LAruR2gi9/can-we-efficiently-distinguish-different-mechanisms. Accessed on May 28, 2023. Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y., and Ghassemi, M. Aging with grace: Lifelong model editing with discrete key-value adaptors.arXiv preprint arXiv:2211.11031, 2022. Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J. Unsolved problems in ml safety.arXiv preprint arXiv:2109.13916, 2021. Hoelscher-Obermaier, J., Persson, J., Kran, E., Konstas, I., and Barez, F. Detecting edit failures in large language models: An improved specificity benchmark. InFind- ings of ACL. Association for Computational Linguistics, 2023. Hubinger, E., van Merwijk, C., Mikulik, V., Skalse, J., and Garrabrant, S. Risks from learned optimization in ad- vanced machine learning systems, 2021. Irie, K., Schlag, I., Csord Ģas, R., and Schmidhuber, J. A modern self-referential weight matrix that learns to mod- ify itself. InInternational Conference on Machine Learn- ing, p. 9660ā9677. PMLR, 2022. Kirsch, L. and Schmidhuber, J. Eliminating meta opti- mization through self-referential meta learning.arXiv preprint arXiv:2212.14392, 2022a. Kirsch, L. and Schmidhuber, J. Self-referential meta learn- ing. InFirst Conference on Automated Machine Learn- ing (Late-Breaking Workshop), 2022b. Lee, K., Liu, H., Ryu, M., Watkins, O., Du, Y., Boutilier, C., Abbeel, P., Ghavamzadeh, M., and Gu, S. S. Aligning text-to-image models using human feedback, 2023. Li, A. C., Brown, E., Efros, A. A., and Pathak, D. Internet explorer: Targeted representation learning on the open web.arXiv preprint arXiv:2302.14051, 2023. Meng, K., Bau, D., Andonian, A., and Belinkov, Y. Lo- cating and editing factual associations in gpt.Advances in Neural Information Processing Systems, 35:17359ā 17372, 2022a. Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y., and Bau, D. Mass-editing memory in a transformer.arXiv preprint arXiv:2210.07229, 2022b. Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C. Memory-based model editing at scale. InInter- national Conference on Machine Learning, p. 15817ā 15831. PMLR, 2022. Ngo, R., Chan, L., and Mindermann, S. The alignment problem from a deep learning perspective, 2023. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730ā27744, 2022. Pan, A., Shern, C. J., Zou, A., Li, N., Basart, S., Woodside, T., Ng, J., Zhang, H., Emmons, S., and Hendrycks, D. Do the rewards justify the means? measuring trade-offs between rewards and ethical behavior in the machiavelli benchmark, 2023. Deceptive Alignment Monitoring Peng, B., Li, C., He, P., Galley, M., and Gao, J. Instruc- tion tuning with gpt-4.arXiv preprint arXiv:2304.03277, 2023. Ravfogel, S., Twiton, M., Goldberg, Y., and Cotterell, R. Linear adversarial concept erasure, 2022. Schaeffer, R., Miranda, B., and Koyejo, S. Are emergent abilities of large language models a mirage?, 2023. Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whit- tlestone, J., Leung, J., Kokotajlo, D., Marchal, N., An- derljung, M., Kolt, N., Ho, L., Siddarth, D., Avin, S., Hawkins, W., Kim, B., Gabriel, I., Bolina, V., Clark, J., Bengio, Y., Christiano, P., and Dafoe, A. Model evalua- tion for extreme risks, 2023. Sun, Z., Shen, Y., Zhou, Q., Zhang, H., Chen, Z., Cox, D., Yang, Y., and Gan, C. Principle-driven self-alignment of language models from scratch with minimal human supervision.arXiv preprint arXiv:2305.03047, 2023. Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H. Self-instruct: Aligning language model with self generated instructions.arXiv preprint arXiv:2212.10560, 2022. Xie, S. M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P., Le, Q. V., Ma, T., and Yu, A. W. Doremi: Optimizing data mixtures speeds up language model pre- training.arXiv preprint arXiv:2305.10429, 2023. Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D. Wizardlm: Empowering large language models to follow complex instructions.arXiv preprint arXiv:2304.12244, 2023.