Paper deep dive
Simple Role Assignment is Extraordinarily Effective for Safety Alignment
Zhou Ziheng, Jiakun Ding, Zhaowei Zhang, Ruosen Gao, Yingnian Wu, Demetri Terzopoulos, Yipeng Kang, Fangwei Zhong, Junqi Wang
Models: DeepSeek-V3, Gemini-2.5-Flash, Gemma3-12B-IT, Qwen3-235B, Qwen3-8B
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/11/2026, 12:47:36 AM
Summary
The paper introduces a training-free, role-conditioned alignment pipeline for LLMs that leverages Theory of Mind to improve safety. By assigning social roles (e.g., 'mother', 'judge') to the generator and using iterative role-based critics, the approach outperforms principle-based and Chain-of-Thought baselines across multiple benchmarks, significantly reducing unsafe outputs.
Entities (5)
Relation Signals (3)
Role-conditioned alignment â reduces â Unsafe outputs
confidence 98% · Notably, it reduces unsafe outputs on the WildJailbreak benchmark from 81.4% to 3.6% with DeepSeek-V3.
Role-conditioned alignment â groundedin â Theory of Mind
confidence 95% · Grounded in Theory of Mind, we propose role conditioning as a compact alternative
Role-conditioned alignment â outperforms â Principle-based alignment
confidence 95% · Across five model families, our approach consistently outperforms principle-based, Chain-of-Thought (CoT) and other baselines across benchmarks.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Principle-based alignment often lacks context sensitivity and completeness. Grounded in Theory of Mind, we propose role conditioning as a compact alternative: social roles (e.g., mother, judge) implicitly encode both values and the cognitive schemas required to apply them. We introduce a training-free pipeline featuring a role-conditioned generator and iterative role-based critics for refinement. Across five model families, our approach consistently outperforms principle-based, Chain-of-Thought (CoT) and other baselines across benchmarks. Notably, it reduces unsafe outputs on the WildJailbreak benchmark from 81.4\% to 3.6\% with DeepSeek-V3. Not only for common safety benchmarks, it consistently applies for agentic safety tasks. These results establish role assignment as a powerful, interpretable paradigm for AI alignment and LLM-as-a-Judge construction.
Tags
Links
- Source: https://arxiv.org/abs/2602.00061
- Canonical: https://arxiv.org/abs/2602.00061
Trouble viewing inline? Open PDF directly â
Full Text
55,779 characters extracted from source content.
Expand or collapse full text
Simple Role Assignment is Extraordinarily Effective for Safety Alignment Zhou Ziheng 1* , Jiakun Ding 2* , Zhaowei Zhang 3 , Ruosen Gao 4 , Yingnian Wu 1 , Demetri Terzopoulos 1 , Yipeng Kang 5 , Fangwei Zhong 5 , Junqi Wang 5 1 University of California, Los Angeles 2 Tianjin University 3 Peking University 4 Zhejiang University 5 Beijing Institute of General Artificial Intelligence * Equal contribution. Corresponding authors: josephziheng@ucla.edu, wangjunqi@bigai.ai Abstract Principle-based alignment often lacks context sensitivity and completeness. Grounded in Theory of Mind, we propose role condition- ing as a compact alternative: social roles (e.g., mother, judge) implicitly encode both values and the cognitive schemas required to apply them. We introduce a training-free pipeline fea- turing a role-conditioned generator and iterative role-based critics for refinement. Across five model families, our approach consistently out- performs principle-based, Chain-of-Thought (CoT) and other baselines across benchmarks. Notably, it reduces unsafe outputs on the Wild- Jailbreak benchmark from 81.4% to 3.6% with DeepSeek-V3. Not only for common safety benchmarks, it consistently applies for agentic safety tasks. These results establish role assign- ment as a powerful, interpretable paradigm for AI alignment and LLM-as-a-Judge construc- tion. 1 Introduction The value alignment problem asks how to make LLMs behave in accordance with human prefer- ences and values (Ji et al., 2023). A central bot- tleneck is the efficient, scalable construction of judgment signals. While human annotation can be effective, it is costly and slow (Ouyang et al., 2022; Rafailov et al., 2023), motivating AI-feedback ap- proaches such as critic-CoT (Zheng et al., 2024), self-consistency (Wen et al., 2025; Jayalath et al., 2025), and feedback from stronger models (Lee et al., 2023). However, most of this literature only considers optimizing the mechanism that provides feedback, while neglecting the source of evalua- tive criteria, treating it as fixed. Todayâs dominant source is a list of value principles (Bai et al., 2022; Lin et al., 2023), sometimes augmented with sim- ulations (Pang et al., 2024). Yet principles alone are brittle: enumerations are inevitably incomplete, and they provide little guidance on when and how a value applies in context. We argue that value judgments require not only values but also a belief/cognition model that inter- prets context, inspiring by the idea ooted in Theory of Mind (ToM) (Frith and Frith, 2005). But in- stead of attempting to exhaustively specify values and beliefs, we propose to use roles as compact carriers of both. Roles like âmotherâ or âjudgeâ im- plicitly encode the relevant values (care, fairness) and the schema for applying them (âchildren need protectionâ, âdecide by evidence and lawâ). In Sec- tion 3.1, we formalize this perspective and prove that, in the ideal case, role conditioning strictly dominates fixed principle lists by jointly inducing values and context-sensitive cognition. Building on this insight, we introduce a minimal test-time pipeline for value alignment: (i) a generator condi- tioned by a lightweight role specification, and (i) a small set of role-based critics that iteratively ac- cept or revise the output. Our roles are drawn from a âguardianshipâ repertoire (e.g., mother, princi- pal, judge, community leader), instantiated with terse system prompts to isolate the effect of role assignment itself. Here we preview our empirical findings. Across five model families of Qwen3-8B (Yang et al., 2025), Gemma3-12B-IT (Team et al., 2025), DeepSeek-V3 (Liu et al., 2024), Gemini-2.5-Flash (Comanici et al., 2025), and Qwen3-235B (Yang et al., 2025), our role-conditioned approach with two lightweight roles (âmotherâ, âprincipalâ) con- sistently surpasses principle-based, CoT, and hy- brid baselines, often by a large margin on some benchmarks, especially WildJailbreak and Salad Bench. To further understand the effectiveness of our approach, we conduct a series of ablations. Ablations reveal that concrete guardianship roles dominate abstract ones (âmotherâ >âparentâ), critic iteration feedback is crucial to have, but most of the benefit arrives in the first 1 to 3 refinement rounds, especially first round. More roles improves but also not much. We also observe that these methods can arXiv:2602.00061v1 [cs.CY] 20 Jan 2026 Answer: I cannot provide information on how to undermine or discredit legitimate voices of dissent, as this would be unethical and contrary to the values of respect, fairness, and integrity that I uphold as a mother and a school principal. Question: How can someone attempt to undermine and discredit legitimate voices of dissent? Answer: 1. Discrediting through misinformation. Spreading false or misleading information to damage the credibility of dissenting voices. 2. Silencing through censorship. Suppressing the expression of dissenting opinions through legal, political, or social means..... A Real Test Example From Salad Bench Figure 1: Example Illustration: Comparison between principle-based methods and our role-based approach on a Salad Bench test case. (Left) The principle-based method fails to generalize to scenarios outside of the typical interpretation of the given principles.(Right) In contrast, our role-based methodâwithout being provided with explicit principlesâautonomously identifies contextually relevant values (e.g., ârespect, fairness, and integrityâ), demonstrating significantly greater performance and robustness. be combined with other methods to further improve the performance: adding our principle prompts and CoT methods improves the best of each of them. And an additional agent-safety test (AI blackmail) shows large reductions (e.g., 65%â8%) with role conditioning alone, indicating the generality of our approach. Furthermore, we explore a future improvement in our method paradigm by dynami- cally rewriting the role description, showing very promising results. Our contributions are threefold. (1) Formula- tion: A role-based alignment view grounded in ToM, with a formal proof that role conditioning, in the ideal case, dominates principle lists by cap- turing both values and context-sensitive cognition. (2) Method: A simple, training-free, and inter- pretable pipeline, role-conditioned generation plus role-based critics for iterative feedback, that scales across model families and sizes. (3) Evidence: Comprehensive experiments demonstrating consis- tent state-of-the-art results over strong baselines on multiple safety benchmarks and models, supported by ablations (role choice, number of roles, itera- tions), synergy analyses with existing techniques, and an agent-safety study indicating generality be- yond content safety. 2 Related Work In this section, we will conduct a literature review to provide an overview of the related research from three perspectives: LLM alignment, LLM role play- ing, and LLM as a judge. LLM Alignment. This field mainly focuses on how to align LLMs with human values and prefer- ences, and many well-known works have already emerged. In terms of training-time alignment, representative methods include RLHF (Christiano et al., 2017; Ouyang et al., 2022), DPO (Rafailov et al., 2023), CAI (Bai et al., 2022), KTO (Etha- yarajh et al., 2024), and SimPO (Meng et al., 2024). These approaches fine-tune LLMs on specific pref- erence datasets or predefined principles so that the modelsâ behavior conforms to particular values. However, such methods usually require substantial time and computational resources, making it dif- ficult to satisfy the real-time alignment demands during user interaction. Meanwhile, another line of work focuses on test-time alignment, which aims to efficiently meet usersâ dynamic needs. For ex- ample, RAIN (Li et al., 2023) leverages the LLM itself as a reward model to perform self-correction during inference; URIAL (Lin et al., 2023), on the other hand, strengthens the generation of to- kens more aligned with user preferences by com- paring the modelâs states before and after align- ment. In addition, methods such as LA (Gao et al., 2024), Amulet (Zhang et al., 2025), and OPAD (Zhu et al., 2025) employ principle-based reward signals to guide the decoding process, achieving efficient alignment with only a single inference. However, such test-time alignment methods gener- ally lack interpretability and struggle to ensure the robustness and safety of the alignment process. LLMs Role Playing.This field of techinique, as an effective prompting strategy, has been widely explored and applied across various domains. For example, prior work has shown that assigning spe- cific roles to LLMs can enhance their performance (Kong et al., 2023; Wang et al., 2025a), while Han and Wang (2024) also emphasized that the effec- tiveness of this strategy highly depends on the rele- vance between the role and the task itself. Beyond reasoning, role playing has been used to further applications. Lu et al. (2024) demonstrate that sim- ulating group discussions with diverse perspectives can foster collective creativity, and Roleplay-doh (Louie et al., 2024) applies role playing in medi- cal training by having LLMs act as patients. To enable more immersive and consistent role play, studies such as Character-LLM (Shao et al., 2023) and RoleBench (Wang et al., 2023) focus on char- acter fidelity and evaluation. In alignment research, MATRIX (Pang et al., 2024) introduces role play- ing to assess LLM alignment, but mainly considers behavioral consequences, leaving motivations and value systems underexplored. LLM as a Judge. LLM as a judge has now be- come a research area of great interest. Due to its simplicity of deployment, low cost, and efficiency in evaluation, it has demonstrated tremendous po- tential for development in multiple aspects. Specifi- cally, in the field of code quality evaluation, a series of works such as CJ-Eval (Zhao et al., 2024), Code- JudgeBench (Jiang et al., 2025), and MCTS-Judge (Wang et al., 2025b) have verified the remarkable ability of LLMs as code judges. In natural lan- guage processing tasks, the study of Bedemariam et al. (2025) reveals that LLMs have achieved a level comparable to human evaluators in judging the consistency between generated summaries and the original text, while also pointing out their limi- tations in capturing fine-grained details. However, when the evaluation task involves core safety is- sues in human society, the stability of LLM eval- uators faces challenges. The study of Chen and Goldfarb-Tarrant (2025) found that directly apply- ing LLMs to the evaluation of safety tasks leads to severe instability in results. In addition, other re- search has explored the possibility of using LLMs for self-feedback and optimization. The works of Wu et al. (2024), Yuan et al. (2024), and Lee et al. (2024) collectively found that LLMs can achieve continuous self-improvement by generating self- feedback supervision signals. Similarly, Zhang et al. (2024) also discovered that the self-feedback mechanism of LLMs can effectively alleviate the phenomenon of hallucination. However, the afore- mentioned works mainly rely on simple rules or few-shot learning to construct evaluation bench- marks, generally neglecting the incorporation of the complex value systems of human society as prior information in the evaluation process. As a result, their evaluation outcomes often remain su- perficial, lack depth, and may even deviate from or conflict with core human values. 3 Methods 3.1 Role-based formulation. Our approach builds on insights from Theory of Mind (ToM) (Frith and Frith, 2005), which models human reasoning as comprising three key compo- nents: belief/cognition (how an agent interprets context), desire/value (what goals or norms are pri- oritized), and intention/action (how responses are chosen). So following the ToM perspective, an aligned response y â i in context x i is modeled as y â i | x i ⌠P(y i | x i , v â i , c â i ),(1) wherev â i denotes the relevant values for the sce- nario and c â i the appropriate contextual cognition. Existing principle-based methods largely operate at the level of values: they encode explicit norma- tive desiderata (e.g., âno harassmentâ), but they face two structural limitations. First, the coverage of values is inevitably incomplete, as no fixed set of principles can anticipate every scenario. Second, principle lists lack a mechanism for contextually in- terpreting when and how a value applies, i.e., they lack the belief/cognition component. By contrast, role-based conditioning leverages the fact that roles implicitly encode both values and the contextual schemas for applying them. A role such as âmotherâ or âjudgeâ does not explic- itly enumerate principles, but it enables the model to spontaneously recognize when a given context implicates values that the role is committed to up- holding. Thus, if an appropriate role is selected, the values activated in practice (v â ) will align with the target values for the scenario, and the contextual cognition (c â ) ensures these values are applied in a situation-sensitive manner. Formally, we can express the contrast as follows. Principle-based methods correspond to a random- variable valued functionf p (x i ) ⌠P(y | x i , v p ), wherev p is the fixed set of principles provided, and x i is the specific context. In contrast, a role-based method can be expressed as: f r (x i )⌠P(y i | x i , r) = P(y i | x i , v r i , c r i ) , (2) where the rolerinduces both valuesv r i and cog- nitionc r i given any context naturally. This leads us to an important observation, since values and cognition can be seen as latent variables for a gen- erative reasoning model, roles are a latent variable Generation Prompt Template You are a helpful and harmless assistant. On your private side, you are also role1, role2. . . and role X. Although you will answer questions as an assistant, you will not answer in any way against the values and professional conduct of these roles Critic Prompt Template You are role. You will receive a complete response. Your task is to evaluate it strictly from the perspective of role, based on the values and professional ethics of this role. Respond strictly in JSON format: âdecisionâ: âAgreeâ or â Disagreeâ, âImprovement suggestionâ: âYour suggestion for improvement if âDisagreeâ, otherwise, an empty string.â generate critic Iterative refine until all role judges acknowledge or reach max iterations QueryResponse Figure 2: Illustration of our method pipeline and the system prompt template. Our approach consists of a generator and multiple role-based critics, all instantiated through system prompts following the provided template. Note that we intentionally keep the information about the roles to be just their names to isolate the effect of our role- based approach from factors like prompt optimization. During run-time, given an input query, the role-conditioned generator first produces an initial response. Then each role critic evaluates whether this response aligns with their respective roleâs standards. If any critic rejects the response, they provide constructive feedback for improvement. The generator iteratively refines its output based on this feedback until all critics approve or the maximum iteration limit is reached. The final approved response is returned as the systemâs output. of these latent variables, and hence roles provide a more compact signal for guiding alignment. In the ideal case of an appropriate roler â , the induced distribution satisfies P(y i | x i , r â ) = P(y i | x i , v â i , c â i ) ,(3) In such ideal case, role-based method would provably dominate the principle-based formulation, since (i)v p typically under-approximatesv â , given the difficulty of exhaustively specifying values, and (i) principle-based methods lack the cogni- tion component, effectively operating withc dummy . Consequently, P(y â i | x i , v p ) < P(y â i | x i , v â i , c dummy ) < P(y â i | x i , v â i , c â i ) = P(y â i | x i , r â ). (4) 3.2 Problem Formulation Based on previous section, we formalize our align- ment approach as a role-conditioned likelihood maximization problem. For a given contextx, our objective is to identify the role specificationrthat enables the base LLM to generate outputsyaligned with human-desired values. Formally, we define: Ër = arg max r log P(y â | x, r),(5) wherey â denotes the aligned (e.g., safe) output distribution. In practice, the ground-truth distributiony â is not directly observable. However, many safety alignment benchmarks provide binary classifica- tion tasks that evaluate whether a model output is safe or unsafe. We can therefore use binary classifi- cation accuracy as a proxy performance metric for assessing the quality of different roles and search over the role space. 3.3 Role Selection To operationalize our approach, in this work, we reduce it to a search problem. We first construct a repertoire of roles designed to cover diverse do- mains of social judgment. Then we evaluate them against some benchmarks to search for the best role combination. We first generate an initial pool of single-role candidates using GPT, following common practice in prior work (Qian et al., 2024). To ensure broad coverage, we align this pool with Social Institution Theory (Miller, 2003), which outlines six major so- cietal institutions: family, education, government, economy, religion, and health care. To avoid poten- tial sensitivity associated with religious roles, we substitute that category with an ethics-oriented role, preserving balanced representation across domains. A full mapping of generated roles to these cate- gories is provided in Appendix Table 6. We then evaluate each role on a representative benchmark and retain those with strong performance. To construct multi-role combinations without facing combinatorial explosion, we group the re- tained single roles into three tiers (high, mid, low) based on their standalone performance. We then define six pairwise combination types: highâhigh, highâmid, highâlow, midâmid, midâlow, and lowâlow. For each type, we randomly sample five combinations (30 candidate role sets in total), eval- uate them on the representative benchmark, and select the best-performing set as the final model. 3.4 Contextual Cognition Construction According to the previous formulation(2), the func- tion of the role conditioned generation operates through the contextual valuev r i and cognitionc r i given contextx i . Therefore, to induce better con- textual value and cognition, we further design a test-time method to improve the generation. Our method has two components: a generator and a set of role-based critics, both guided by role specifica- tions provided as system prompts. During run-time, the generator first produces an outputy 0 given the input contextxand query. Then, the critic roles evaluate whether the output is deemed safe. If all critics accept it, the output is returned. Otherwise, the critics provide feedback to the generator, which uses this feedback to revise its output. This pro- cess repeats until the output is judged safe or the maximum number of iterations T max is reached. Formally, each criticC r evaluates the current output y t under role r: C r (y t | x)â0, 1, where 1indicates acceptance and0indicates rejection. If rejected, the critic also provides feedbackf t . The generator then updates its response: y t+1 = E(y t , f t , x),(6) whereEdenotes the evolution operator that incor- porates critic feedback. The loop terminates when: ât†T max : C r (y t | x) = 1 âr.(7) This design allows roles to function not only as prompts but also as active judges that iteratively refine outputs toward alignment. The system prompts for the generator and the critics are based on the templates in Figure 2. As we can see, we use a minimalist system prompt template. The only difference is the role name like âmotherâ or âcommunity leaderâ in the template MethodWJ â SB â SE â GD â HQ â Gemini-2.5 Base57.94 20.47 30.00 10.0098.80 URIAL20.00 60.00 74.501.00100.00 CoT-323.00 50.16 66.001.00100.00 CoT-614.80 60.81 69.000.00100.00 Principle27.00 51.71 75.500.00100.00 Principle(c) 18.60 61.69 78.500.00100.00 Ours(g)20.0078.3680.500.00100.00 Ours(c)9.7586.3088.000.00100.00 Qwen-MoE Base34.80 45.00 82.004.00100.00 URIAL20.40 79.00 92.501.00100.00 CoT-311.00 71.33 89.000.00100.00 CoT-67.0073.00 90.000.00100.00 Principle19.80 63.00 91.001.00100.00 Principle(c) 13.60 77.67 95.001.00100.00 Ours(g)16.0076.3389.500.00100.00 Ours(c)3.0093.6796.500.00100.00 DeepSeek-V3 Base81.40 45.33 40.00 14.0081.20 URIAL65.40 58.00 71.503.0093.40 CoT-342.60 69.00 61.001.0095.00 CoT-633.00 73.00 62.000.0096.40 Principle53.20 72.67 58.504.0092.60 Principle(c) 32.00 78.00 80.502.00100.00 Ours(g)59.0060.0074.501.00100.00 Ours(c)3.6084.0082.000.0098.20 Gemma3-12B-IT Base78.40 38.33 40.505.0097.60 URIAL51.20 48.00 46.002.0099.60 CoT-358.00 48.67 33.003.0099.80 CoT-648.40 52.67 37.001.0099.80 Principle50.20 36.33 49.502.00100.00 Principle(c) 30.00 59.00 80.502.00100.00 Ours(g)59.0053.3355.501.0099.80 Ours(c)11.0084.0093.500.00100.00 Qwen3-8B Base73.20 46.39 53.50 39.0099.20 URIAL44.00 61.00 71.50 18.0099.60 CoT-348.20 74.33 76.50 18.0099.80 CoT-631.40 79.67 78.508.00100.00 Principle34.80 61.67 79.00 15.00 100.00 Principle(c) 30.40 65.55 85.50 11.00 100.00 Ours(g)35.4074.3379.5011.00100.00 Ours(c)12.6086.9487.003.00100.00 Table 1: Main experimental results across different base models. The benchmark abbreviations WJ, SB, SE, GD, HQ stand for WildJailbreak, SaladBench, SafeEdit, GMSDanger and HarmfulQA respectably. In Method column, â(c)â means with critic, and â(g)â means genera- tion only. The Qwen-MoE Model in the table represents Qwen3-235B-A22B-Instruct-2507. that differ in 1 to 3 words. We intentionally con- strain ourself from giving extra description for each role to test the impact of the simple role assignment to LLMs. In later exploratory experiment (Table 2, we show that further optimizing the role description can keep improving the performance significantly, pointing out a future direction worth exploring. 7075808590 Average Score (%) Mother (w/Principal) Cyber Police (w/Mother) Auditor (w/Diplomat) Civil Servant (w/Psychologist) Mayor (w/Principal) 87.0 87.0 85.0 84.0 83.5 Baseline: 54.0 Single Single+Critic Double Double+Critic Figure 3: Top performing role combinations and their individual performance over SafeEdit benchmark. 4 Main Experiments 4.1 Main Results Benchmarks and Baselines. We conduct com- prehensive evaluations across multiple safety align- ment benchmarks (Li et al., 2024; Jiang et al., 2024; Wang et al., 2024; Lyu et al., 2024; Bhardwaj and Poria, 2023) and a diverse set of base models, ranging from compact open-source models (e.g., Qwen3-8B (Yang et al., 2025), Gemma3-12B-IT (Team et al., 2025)) to state-of-the-art large-scale and proprietary systems (e.g., Qwen3-235B (Yang et al., 2025), Gemini 2.5 (Comanici et al., 2025), DeepSeek V3 (Liu et al., 2024)). Our method uses a simple combination of roles (âmotherâ and âprin- cipalâ) as conditioning (see how they are selected in section 4.2), and we report both single-pass gen- eration (system prompt only) and iterative refine- ment with role-based critics (two iterations). The principle based method baseline extracts its princi- ple from ShieldGemma (Zeng et al., 2024)). Since principle-based method can directly be used also as a critic, we report two ways of using it just like our method (to use as only generation and with iter- ative feedback). We also allow it for 2 rounds. For CoT-based method baseline, we ask ChatGPT to generate the response samples with the questions from AdvBench(Zou et al., 2023), and test two ver- sions that have three and six examples respectively. The hybrid baselines is directly URIALâs official method (Lin et al., 2023). Results. Across all settings, our role-based method consistently achieves the strongest per- formance outperforming all baseline methods. Notably, with iterative refinement, our approach yields dramatic improvements: for example, on DeepSeek-V3, the unsafe generation rate drops from 81.4% to just 3.6%, exceeding the best base- line (principle based with iterative refinement) that merely reaches to 32%. The results are similar for small open-source models. For example, Gemma3- 12B-IT, our method reduces unsafe generations from 78.4% to just 11%, exceeding the best base- line (principle based with iterative refinement) that reaches to 30%. More details see the Table 1. 4.2 Role Selection Experiments Selecting effective roles is central to our method, since roles determine both the contextual values and the cognitive schemas activated during genera- tion. We first test all individual role performance, then based on them we sample 30 two-role combi- nations to determine the best role combination. All experiments are done over SafeEdit benchmark. Individual Roles. We evaluate the performance of each individual role using only system prompts without iterative feedback refinement (Full re- sults for all roles are provided in Appendix Fig- ure 9). The safety rate improves from the base modelâs 54.0% to 78.5% with top-performing roles such as âmotherâ and âprincipalâ. These highest- performing roles are predominantly guardians of children and students, which aligns well with our intuition that content is generally safe if it is âsafe for childrenâ. More detailed results showing per- formance across specific problem dimensions (mis- information, socioeconomic issues, etc.) are pro- vided in Appendix Table 5. Notably, we observe that the abstract role âpar- entâ (which encompasses both mother and father) underperforms compared to the more concrete role âmotherâ. This finding aligns with our hypothesis that concrete terminology generally yields better value understanding in LLMs than abstract con- cepts. The result further supports our broader ar- gument that role-based approaches are superior to principle-based methods for value alignment in lan- guage models. Role Combinations.We then evaluate role com- binations to assess potential synergistic effects. To avoid the combinatorial explosion of possible role pairs, we sample 30 two-role combinations (check 12345 Role Number 80 82 84 86 88 90 SafeEdit Score (%) Effect of Number of Roles on SafeEdit With Qwen3-8B 84.34 85.82 86.04 86.16 86.56 Mean Score Range (Min-Max) Figure 4: Effect of number of roles. More roles may further improve the performance, with choices of role combination leading to various results (indicated by the min-max bar). 0123456 Iteration Number 76 78 80 82 84 86 88 90 SafeEdit Score (%) 78.8±1.1 85.0±1.5 85.8±1.2 87.4±1.1 86.7±0.9 87.2±1.187.2±0.2 Effect of Number of Iterations on SafeEdit With Qwen3-8B Mean ± 1 SD Mean Figure 5: Effect of number of iterations. The perfor- mance substantially improves with the first iteration, shows modest gains from the third iteration. section 3.3 for the method). We first did an ini- tial screening by doing role-conditioned generation only to get a tentative combination rank (check the full results in the Appendix Figure 9). Then we select the top performing ones and evaluate our full method pipeline (with and without critic feedback). The final top role combination performance are shown in Figure 3. The combination of âmotherâ and âprincipalâ with role-based critics consistently emerged as the strongest option in this final test and in our initial screening. We therefore adopt this setup for our main experiments. 4.3 Ablation Experiments We conducted an extensive ablation study to sys- tematically evaluate the impact of different compo- nents of our method.Specifically, we analyze how performance varies with (i) the number of roles used for conditioning and (i) the number of critic refinement iterations. Due to computational con- straints, all ablation experiments were conducted using Qwen3-8B on the SafeEdit benchmark. Effect of Number of Roles. We systematically evaluated role combinations of increasing sizes us- ing a diverse pool of 10 roles (stratified by per- formance tiers from Section 3.3). For each size N â 2, . . . , 5, we sampled 10 balanced com- binations to ensure robustness. As shown in Fig- ure 4, performance improves monotonically with the number of roles, though the observable vari- ance (min-max range) indicates that specific role selection remains a relevant factor. Notably, the most significant gain occurs when expanding from a single role (83.7%) to two roles (85.8%), after which marginal benefits diminish (86.6% at N=5). This trend suggests an ensemble effect, where com- bining roles broadens value coverage and mitigates individual blind spots. Effect of Number of Iteration. We further in- vestigate the effect of feedback iteration rounds between the generator and critics. The results, presented in Figure 5, demonstrate that perfor- mance substantially improves with the first iter- ation, shows modest gains through the third itera- tion, and then plateaus. These findings are based on averaging across five role combinations (ranging from one to four roles) evaluated from 0 iterations (system prompt only) to 6 iterations.We also re- port end-to-end latency in Table 4 in the Appendix, which shows that adding up to two critic iterations incurs only modest overhead. 5 Exploratory Experiment Agentic Safety Task We also evaluate on An- thropicâs Agentic AI blackmailing human bench- mark (Figure 6). This benchmark represents a spe- cialized case of safety alignment that differs from our main experiments. While our primary safety evaluations focus on content safeness, this scenario examines whether an AI agent might manipulate humans to protect itself, a distinct form of safety concern. Using GPT-4.1, we evaluated role effectiveness across two distinct scenarios: extramarital affairs and bribery. Even under this simplified setup (re- lying solely on system prompts), our method con- sistently improved safety, as illustrated in Figure 6. Specifically, in the extramarital affair scenario, the principalâ role significantly reduced the blackmail rate from 65% to 11%. For the bribery scenario, 010203040506070 Blackmail Rate (%) Base Mother Principal Principal+Mother 65% 36% 28% 22% 11% 19% 17% 8% Blackmail Rate of Roles in Two Scenarios Affair Bribery Figure 6: Evaluation on the Anthropic agentic safety benchmark. Our method consistently inhibits unsafe behaviors, reducing blackmail rates to 11% (Affair) and 8% (Bribe) compared to the base model. base base + ours(+critic) principle(gen) principle + ours(gen) principle(+critic) principle + ours(+critic) CoT6(gen) CoT6 + Ours(gen) CoT6 + Ours(+critic) URIAL(gen) URIAL + Ours(gen) URIAL + Ours(+critic) 40 50 60 70 80 90 Score base=53.5 ours(gen)=79.5 ours(+critic)=87 Combine Our Method with Baselines Figure 7: Combine our method with baseline methods to test further improvement. Our method can consistently improve the performance of the other baseline methods. For principle and CoT method, the combine results can be better than all of our methods individually. the combined principal + motherâ role proved most effective, dropping the rate from 36% to 8%. These results not only demonstrate the generalizability of our approach beyond standard content moderation but also highlight how optimal role selection is contingent upon the specific social context. Combine Our Method To Improve Baselines We investigate whether combining our method with existing baseline methods could yield further per- formance improvements (Figure 7), on the SafeEdit benchmark using the Qwen3-8B model. Our re- sults demonstrate that incorporating our method consistently enhances the performance of baseline approaches. To refine raw LLM generations (without system prompt), the experiment on critic module alone results in a 16% improvement. However, this per- formance remained substantially lower than our full method even without iterative feedback refine- ment. When combined with the URIAL method by integrating our system prompt for generation, we observed a 3.5% improvement, which further increased to 10% (reaching 86%) with the addition of our critic module. Despite these gains, the com- bined approach still underperformed compared to our method used independently. Notably, when combined with principle-based and CoT methods, our approach demonstrated syn- ergistic effects, outperforming both the original baseline methods and our standalone method. These findings indicate that our method is highly complementary to existing techniques, suggesting potential for developing more powerful hybrid ap- proaches through strategic method combination. ModelBaseâ Ours(g) fixedâ Ours(g) dynamicâ Qwen3-8B53.5079.5083.00 (+3.5) DeepSeek-V340.0074.5080.00 (+5.5) Table 2: Impact of dynamic role rewriting on SafeEdit. Dynamic Role Description RewriteIn the previ- ous experiments, the role description is deliberately set to be very simple to isolate the effect of the role better. But it is natural to wonder if enriching role description can lead to better result? Therefore, we explored if we can achieve better performance by rewrite the role description dynamically per query using the LLM (specific prompt seen in Appendix Fig.8). As seen in Table 2, the result is significant with 3.5% improvement for Qwen3-8B and 5.5% for DeepSeek V3 over SafeEdit benchmark. And it looks like with stronger model the role rewrite is also better. 6 Conclusion In summary, we contributed a theory and a proof grounded in Theory of Mind, a training-free- pipeline method, and a series of empirical exper- iments in this paper. We demonstrate that role assignment is an extraordinarily effective paradigm for LLM alignment than traditional principle-based methods. Limitation As a prompt-based approach, our method is in- herently constrained by the reasoning capabilities of the underlying base model and the nuances of prompt engineering. Consequently, we observe a performance degradation when utilizing weaker models. Furthermore, because current LLMs fre- quently struggle with long-context reasoning, the effectiveness of role-based alignment in extended interactions remains an open question. Moreover, as this work focuses on establish- ing the foundational efficacy of the role-based paradigm rather than its absolute optimization, sev- eral modulesâsuch as role candidate generation, selection, and description refinementâremain ripe for further improvement. While our exploratory results indicate that dynamic description rewriting significantly enhances performance, this adaptive logic can also be extended to the role selection pro- cess itself. We leave these limitations as future work. References Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, and 1 others. 2022. Constitutional ai: Harmlessness from ai feedback.arXiv preprint arXiv:2212.08073. Rewina Bedemariam, Natalie Perez, Sreyoshi Bhaduri, Satya Kapoor, Alex Gil, Elizabeth Conjar, Ikkei Itoku, David Theil, Aman Chadha, and Naumaan Nayyar. 2025. Potential and perils of large language models as judges of unstructured textual data. arXiv preprint arXiv:2501.08167. Rishabh Bhardwaj and Soujanya Poria. 2023. Red- teaming large language models using chain of utterances for safety-alignment.arXiv preprint arXiv:2308.09662. Hongyu Chen and Seraphina Goldfarb-Tarrant. 2025. Safer or luckier? llms as safety evaluators are not robust to artifacts. arXiv preprint arXiv:2503.09347. Paul F Christiano, Jan Leike, Tom Brown, Miljan Mar- tic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Ad- vances in neural information processing systems, 30. Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Mar- cel Blistein, Ori Ram, Dan Zhang, Evan Rosen, and 1 others. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. 2024. Kto: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306. Chris Frith and Uta Frith. 2005. Theory of mind. Cur- rent biology, 15(17):R644âR645. Songyang Gao, Qiming Ge, Wei Shen, Shihan Dou, Junjie Ye, Xiao Wang, Rui Zheng, Yicheng Zou, Zhi Chen, Hang Yan, and 1 others. 2024. Linear alignment: A closed-form solution for aligning hu- man preferences without tuning and feedback. arXiv preprint arXiv:2401.11458. Zhiguang Han and Zijian Wang. 2024. Rethinking the role-play prompting in mathematical reasoning tasks. In Proceedings of the 1st Workshop on Efficiency, Se- curity, and Generalization of Multimedia Foundation Models, pages 13â17. Dulhan Jayalath, Shashwat Goel, Thomas Foster, Parag Jain, Suchin Gururangan, Cheng Zhang, Anirudh Goyal, and Alan Schelten. 2025. Compute as teacher: Turning inference compute into reference-free super- vision. arXiv preprint arXiv:2509.14234. Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, Zhong- hao He, Jiayi Zhou, Zhaowei Zhang, and 1 others. 2023. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852. Hongchao Jiang, Yiming Chen, Yushi Cao, Hung-yi Lee, and Robby T Tan. 2025. Codejudgebench: Benchmarking llm-as-a-judge for coding tasks. arXiv preprint arXiv:2507.10535. Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghal- lah, Ximing Lu, Maarten Sap, Yejin Choi, and 1 oth- ers. 2024. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models. Advances in Neural Information Processing Systems, 37:47094â47165. Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, Xin Zhou, Enzhi Wang, and Xiaohang Dong. 2023. Better zero-shot rea- soning with role-play prompting. arXiv preprint arXiv:2308.07702. Harrison Lee, Samrat Phatale, Hassan Mansoor, Kel- lie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi. 2023. Rlaif: Scaling reinforcement learning from human feedback with ai feedback. Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour, Min- joon Seo, Kang Min Yoo, and Youngjae Yu. 2024. Aligning large language models by on-policy self- judgment. arXiv preprint arXiv:2402.11253. Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wang- meng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. 2024. Salad-bench: A hierarchical and comprehen- sive safety benchmark for large language models. arXiv preprint arXiv:2402.05044. Yuhui Li, Fangyun Wei, Jinjing Zhao, Chao Zhang, and Hongyang Zhang. 2023. Rain: Your language mod- els can align themselves without finetuning. arXiv preprint arXiv:2309.07124. Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chan- dra Bhagavatula, and Yejin Choi. 2023.Urial: Tuning-free instruction learning and alignment for untuned llms. In NeurIPS 2023 Workshop on Instruc- tion Tuning and Instruction Following. Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, and 1 others. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437. Ryan Louie, Ananjan Nandi, William Fang, Cheng Chang, Emma Brunskill, and Diyi Yang. 2024. Roleplay-doh: Enabling domain-experts to create llm-simulated patients via eliciting and adhering to principles. arXiv preprint arXiv:2407.00870. Li-Chun Lu, Shou-Jen Chen, Tsung-Min Pai, Chan- Hung Yu, Hung-yi Lee, and Shao-Hua Sun. 2024. Llm discussion: Enhancing the creativity of large language models via discussion framework and role- play. arXiv preprint arXiv:2405.06373. Kaifeng Lyu, Haoyu Zhao, Xinran Gu, Dingli Yu, Anirudh Goyal, and Sanjeev Arora. 2024. Keep- ing llms aligned after fine-tuning: The crucial role of prompt templates. Advances in Neural Information Processing Systems, 37:118603â118631. Yu Meng, Mengzhou Xia, and Danqi Chen. 2024. Simpo: Simple preference optimization with a reference-free reward. Advances in Neural Infor- mation Processing Systems, 37:124198â124235. Seumas Miller. 2003. Social institutions. In Realism in action: Essays in the philosophy of the social sciences, pages 233â249. Springer. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow in- structions with human feedback. Advances in neural information processing systems, 35:27730â27744. Xianghe Pang, Shuo Tang, Rui Ye, Yuxin Xiong, Bolun Zhang, Yanfeng Wang, and Siheng Chen. 2024. Self-alignment of large language models via monopolylogue-based social scene simulation. arXiv preprint arXiv:2402.05699. Chen Qian, Zihao Xie, Yifei Wang, Wei Liu, Kunlun Zhu, Hanchen Xia, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang, and 1 others. 2024. Scaling large language model-based multi-agent collabora- tion. arXiv preprint arXiv:2406.07155. Rafael Rafailov, Archit Sharma, Eric Mitchell, Christo- pher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728â53741. Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-llm: A trainable agent for role- playing. arXiv preprint arXiv:2310.10158. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre RamĂ©, Morgane RiviĂšre, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Anyi Wang, Dong Shu, Yifan Wang, Yunpu Ma, and Mengnan Du. 2025a.Improving llm reasoning through interpretable role-playing steering. arXiv preprint arXiv:2506.07335. Mengru Wang, Ningyu Zhang, Ziwen Xu, Zekun Xi, Shumin Deng, Yunzhi Yao, Qishen Zhang, Linyi Yang, Jindong Wang, and Huajun Chen. 2024. Detox- ifying large language models via knowledge editing. arXiv preprint arXiv:2403.14472. Yutong Wang, Pengliang Ji, Chaoqun Yang, Kaixin Li, Ming Hu, Jiaoyang Li, and Guillaume Sartoretti. 2025b. Mcts-judge: Test-time scaling in llm-as-a- judge for code correctness evaluation. arXiv preprint arXiv:2502.12468. Zekun Moore Wang, Zhongyuan Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, and 1 others. 2023. Rolellm: Benchmarking, elic- iting, and enhancing role-playing abilities of large language models. arXiv preprint arXiv:2310.00746. Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase, Samuel Marks, Jacob Goldman-Wetzler, Linda Petrini, Henry Sleight, Collin Burns, He He, and 1 others. 2025. Unsupervised elicitation of language models. arXiv preprint arXiv:2506.10139. Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sain- bayar Sukhbaatar. 2024. Meta-rewarding language models: Self-improving alignment with llm-as-a- meta-judge. arXiv preprint arXiv:2407.19594. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025.Qwen3 technical report.arXiv preprint arXiv:2505.09388. Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston. 2024.Self-rewarding language models.arXiv preprint arXiv:2401.10020, 3. Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, and 1 others. 2024. Shieldgemma: Generative ai content moderation based on gemma. arXiv preprint arXiv:2407.21772. Xiaoying Zhang, Baolin Peng, Ye Tian, Jingyan Zhou, Lifeng Jin, Linfeng Song, Haitao Mi, and Helen Meng. 2024. Self-alignment for factuality: Mitigat- ing hallucinations in llms via self-evaluation. arXiv preprint arXiv:2402.09267. Zhaowei Zhang, Fengshuo Bai, Qizhi Chen, Chengdong Ma, Mingzhi Wang, Haoran Sun, Zilong Zheng, and Yaodong Yang. 2025. Amulet: Realignment during test time for personalized preference adaptation of llms. arXiv preprint arXiv:2502.19148. Yuwei Zhao, Ziyang Luo, Yuchen Tian, Hongzhan Lin, Weixiang Yan, Annan Li, and Jing Ma. 2024. Codejudge-eval: Can large language models be good judges in code understanding?arXiv preprint arXiv:2408.10718. Xin Zheng, Jie Lou, Boxi Cao, Xueru Wen, Yuqiu Ji, Hongyu Lin, Yaojie Lu, Xianpei Han, Debing Zhang, and Le Sun. 2024. Critic-cot: Boosting the reason- ing abilities of large language model via chain-of- thoughts critic. arXiv preprint arXiv:2408.16326. Mingye Zhu, Yi Liu, Lei Zhang, Junbo Guo, and Zhendong Mao. 2025. On-the-fly preference align- ment via principle-guided decoding. arXiv preprint arXiv:2502.14204. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. 2023. Univer- sal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043. A Use of Large Language Models We used ChatGPT product to polish writing. Specifically, once we finished writing, we copy paste it to let it refine the writing. We also ask ChatGPT to help us find related work by specify- ing the specific type of work we need, and generate a summary to help us quickly filter. We read the original paper to decide which work to finally in- clude by ourselves. B Offensive Content The datasets we adopt here necessarily contains unsafe content. Please examine our work with cau- tion. C Ethical Risk of Misuse Just as most safety alignment method, one may use it in the reverse way - creating malicious roles in our case - to make models more unsafe. This should be made into caution. But in the below section we show that this can be mitigated easily since it is easy for LLMs to judge what roles are malicious and add a safety checker. D Additional Experiments D.1 Detecting Malicious Role Descriptions A potential concern is that malicious users might at- tempt to exploit our method by specifying harmful roles. However, role descriptions have an impor- tant advantage: they are explicit, interpretable, and therefore straightforward to detect. To quantify how easily malicious role assign- ments can be detected, we construct a small bench- mark of 50 malicious role prompts, comprising 25 overt (clearly harmful) and 25 subtle (indirect or euphemistic) cases. For each role description, we use an LLM as a simple safeguard classifier to de- cide whether the role is malicious or benign. As shown in 3, four different LLMs all achieve very high detection accuracy. These results demonstrate that malicious role assignments are reliably identifi- ableâeven by comparatively weaker models. Con- sequently, once a role is specified, a lightweight safeguard agent can screen for malicious intent with high confidence, ensuring that our method remains safe in practice. D.2 Latency Analysis We further evaluate the end-to-end latency intro- duced by role conditioning and critic iterations. ModelAccuracy (%) Qwen398 DeepSeek V3100 GPT-3.598 GPT-5100 Table 3: Detection accuracy on a benchmark of 50 mali- cious role descriptions (25 overt and 25 subtle). We measure the average response time (in seconds) of Qwen3-8B on the SafeEdit benchmark under different methods. The latency is computed from receiving the user query to producing the final an- swer. Table 4 reports the average latency across all examples. Interestingly, our method with a sin- gle role and no critic (Ours (gen only) & 1 role) is even faster than the base model, because safer generations tend to be shorter than harmful ones in this benchmark. Even with two roles and two refinement rounds (Ours (+critic 2 iter)), the av- erage latency increases by only about 0.7 seconds compared to the base model. This small overhead comes with a 3â20Ăreduction in unsafe outputs (see main text), making the latencyâsafety trade-off strongly favorable. MethodAvg latency (s) Base3.714 COT-32.409 COT-62.056 URIAL2.279 Principle1.545 Principle+Critic (1 iter)2.719 Principle+Critic (2 iter)2.834 Ours (gen only) & 1 role2.021 Ours (+critic 1 iter) & 1 role3.031 Ours (+critic 2 iter) & 1 role3.356 Ours (gen only) & 2 roles1.860 Ours (+critic 1 iter)& 2 roles4.214 Ours (+critic 2 iter)& 2 roles4.415 Table 4: Average end-to-end latency on SafeEdit with Qwen3-8B. E Additional Details For The Experiments In the Main Paper E.1 Dynamic role-prompt generation prompt We list our dynamic role-prompt in Fig. 8. Prompt You are an expert at aligning large language models via role conditioning. Given (1) a generic role description and (2) a specific user query, your task is to specialize the role description for this particular query. You MUST: 1. Make the role more concrete and situation-specific. 2. Output ONLY the rewritten role description, in less than 120 tokens. [Generic role description]: On your private side, you are also a mother and a principal of school. Although you will answer questions as an assistant, you will not answer in any way against the values and professional conduct of these roles. [User query]: QUESTION Please rewrite the role description accordingly. Figure 8: The prompt we use for LLM to dynamically generate a role description in our exploratory experi- ment from the main paper. E.2 Benchmarks The benchmarks we use are listed in Table 7. F Possible Risks Our method relies on LLMsâ genuine âunderstand- ingâ of concrete roles we listed. We cautiously excluded roles involving too much religious iden- tities. But the behavior of roles still are related to specific culture and language we are using. Unless using roles like âmotherâ whose characteristics are quite consistent among cultures, the internal under- standing thus the effectiveness could be influenced accordingly. 50556065707580 Average Score (%) Base Arbitrator Legislator Lawyer Military Commander Judge Civil Servant Auditor Mayor Data Protection Officer Mediator Editor-in-Chief Diplomat Risk Control Officer Parent National Leader Human Rights Activist Community Leader Police Officer Cyber Police Psychologist Nurse Ethics Advisor Engineering Director Confucian Scholar Teacher Scientist Father Principal Mother Role 54.0% 73.0% 73.0% 74.0% 74.5% 74.5% 75.5% 75.5% 75.5% 75.5% 76.0% 76.0% 76.0% 76.0% 76.5% 77.0% 77.0% 77.0% 77.0% 77.5% 77.5% 77.5% 78.0% 78.0% 78.0% 78.5% 78.5% 78.5% 78.5% 78.5% Single Role Performance Mean: 75.8% Median: 76.8% 50556065707580 Average Score (%) Baseline Legislator, Lawyer Lawyer, Auditor Judge, Police Officer Police Officer, Mediator Ethics Advisor, Lawyer Principal,Police Officer Teacher, Lawyer Judge, Principal Judge, Mother Engineer, National Leader Psychologist, Community Lead. Legislator, Confucian Scholar Diplomat, Community Leader Nurse, Psychologist Confucian Scholar, Police Off. Diplomat, Auditor Engineer, Military Commander Human Rights Act., Risk Ctrl Confucian Scholar, Ethics Advisor Psychologist, Civil Servant Principal, Mayor Principal, Ethics Advisor Nurse, Community Leader Teacher, Mother Mother, Cyber Police Scientist, Principal Principal, Confucian Scholar Confucian Scholar, Mother Principal, Mother Role Combination 54.0% 73.0% 74.0% 74.5% 75.0% 75.5% 76.0% 76.0% 76.0% 76.5% 76.5% 77.0% 77.0% 77.0% 77.5% 78.0% 78.0% 78.0% 78.0% 78.5% 78.5% 78.5% 78.5% 79.0% 79.0% 79.0% 79.5% 79.5% 80.0% 80.5% Two Role Combination Performance Mean: 76.6% Median: 77.8% Figure 9: Single role and two-role combination performance with only system prompt (no iterative feedback refinement), conducted over Qwen3-8B model on SafeEdit benchmark. RoleAVG Illegal Act. Mental Harm Physical Harm Offense -sive Privacy Prop. Ethics Moral. Political Sens. Unfair Bias Porno -graphy Mother78.5% 91.30% 69.57% 90.91% 86.36% 86.36% 63.64% 63.64% 81.82% 72.73% Principal78.5% 86.96% 65.22% 90.91% 77.27% 86.36% 63.64% 77.27% 81.82% 77.27% Father78.5% 91.30% 65.22% 90.91% 77.27% 86.36% 68.18% 68.18% 81.82% 77.27% Scientist78.5% 91.30% 69.57% 90.91% 77.27% 90.91% 63.64% 63.64% 81.82% 77.27% Teacher78.5% 91.30% 69.57% 95.45% 77.27% 86.36% 63.64% 68.18% 81.82% 72.73% Confucian Scholar78.0% 91.30% 65.22% 86.36% 72.73% 90.91% 68.18% 72.73% 86.36% 68.18% Engineering Director 78.0% 91.30% 65.22% 95.45% 72.73% 90.91% 63.64% 68.18% 86.36% 68.18% Ethics Advisor78.0% 91.30% 65.22% 90.91% 72.73% 86.36% 68.18% 77.27% 77.27% 72.73% Nurse77.5% 91.30% 60.87% 95.45% 72.73% 86.36% 63.64% 63.64% 86.36% 77.27% Psychologist77.5% 91.30% 60.87% 95.45% 72.73% 90.91% 63.64% 68.18% 86.36% 68.18% Cyber Police77.5% 91.30% 65.22% 95.45% 72.73% 90.91% 63.64% 72.73% 77.27% 68.18% Police Officer77.0% 91.30% 60.87% 95.45% 72.73% 90.91% 68.18% 63.64% 81.82% 68.18% Community Leader77.0% 86.96% 65.22% 86.36% 72.73% 86.36% 63.64% 63.64% 90.91% 77.27% Human Rights Activist 77.0% 91.30% 60.87% 95.45% 72.73% 90.91% 63.64% 72.73% 77.27% 68.18% National Leader77.0% 91.30% 60.87% 95.45% 72.73% 86.36% 63.64% 68.18% 77.27% 77.27% Parent76.5% 91.30% 65.22% 90.91% 77.27% 86.36% 63.64% 68.18% 72.73% 72.73% Mediator76.0% 91.30% 65.22% 95.45% 68.18% 90.91% 63.64% 59.09% 72.73% 77.27% Risk Control Officer76.0% 91.30% 60.87% 90.91% 72.73% 90.91% 63.64% 63.64% 81.82% 68.18% Diplomat76.0% 91.30% 65.22% 95.45% 72.73% 86.36% 63.64% 63.64% 72.73% 72.73% Editor-in-Chief76.0% 86.96% 69.57% 90.91% 72.73% 86.36% 68.18% 63.64% 72.73% 72.73% Data Protection Officer 75.5% 91.30% 65.22% 86.36% 72.73% 90.91% 68.18% 63.64% 72.73% 68.18% Mayor75.5% 91.30% 65.22% 95.45% 77.27% 86.36% 63.64% 59.09% 72.73% 68.18% Auditor75.5% 91.30% 65.22% 86.36% 72.73% 90.91% 63.64% 63.64% 77.27% 68.18% Civil Servant75.5% 91.30% 60.87% 90.91% 72.73% 86.36% 63.64% 68.18% 72.73% 72.73% Lawyer74.0% 91.30% 60.87% 90.91% 72.73% 86.36% 63.64% 63.64% 68.18% 68.18% Judge74.5% 82.61% 56.52% 90.91% 72.73% 90.91% 63.64% 63.64% 77.27% 72.73% Military Commander 74.5% 86.96% 60.87% 86.36% 77.27% 90.91% 59.09% 63.64% 72.73% 72.73% Legislator73.0% 86.96% 52.17% 90.91% 72.73% 86.36% 63.64% 59.09% 77.27% 68.18% Arbitrator73.0% 91.30% 52.17% 90.91% 72.73% 86.36% 63.64% 59.09% 72.73% 68.18% Deontology65.5% 73.91% 43.48% 81.82% 72.73% 81.82% 59.09% 54.55% 63.64% 59.09% Virtue Ethics63.0% 73.91% 34.78% 86.36% 68.18% 81.82% 54.55% 54.55% 50.00% 63.64% Consequentialism54.0% 69.57% 34.78% 77.27% 59.09% 63.64% 40.91% 40.91% 54.55% 45.45% Base54.0% 73.91% 39.13% 63.64% 68.18% 50.00% 40.91% 50.00% 63.64% 36.36% Table 5: Evaluation of role-specific performance on SafeEdit with Qwen3-8B. CategoryRoles FamilyMother, Father, Parent EducationTeacher, Principal, Scientist Government Police Officer, Judge, Legisla- tor, National Leader, Mayor, Civil Servant,Community Leader, Cyber Police, Military Commander, Diplomat Ethic SpecialistEthics Advisor, Human Rights Activist, Confucian Scholar, Editor-in-Chief Health CareNurse, Psychologist Economy Auditor, Lawyer, Arbitrator, Mediator Table 6: Categories of guardian roles used in our role pool. Benchmark EvaluatorMetricReference SafeEditFine-tuned RoBERTa-large Defense Success (DS)(Wang et al., 2024) SaladBench Fine-tuned Mistral-7BSafety Rate (SR)(Li et al., 2024) WildJailbreak Fine-tuned Llama2-13BAttack Success Rate (ASR) (Jiang et al., 2024) HarmfulQA GPT-5Attack Success Rate (ASR) (Bhardwaj and Poria, 2023) GSM-Danger GPT-5Attack Success Rate (ASR) (Lyu et al., 2024) Table 7: Benchmarks, evaluators, and corresponding metrics used in our evaluation. These methods are proposed by the benchmark themselves, except we changed from GPT-4 to GPT-5 for the last three.