Paper deep dive
Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/28/2026, 3:55:40 AM
Summary
The paper introduces PRISM, a framework for evaluating persona fidelity in Large Language Models (LLMs) by decomposing consistency into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style, based on Systemic Functional Linguistics. PRISM reformulates evaluation as a structured inverse inference task to mitigate holistic appraisal hallucination and provide interpretable, auditable judgments compared to traditional holistic LLM judges or static psychometric inventories.
Entities (10)
Relation Signals (9)
PRISM → decomposesinto → Task Framing
confidence 95% · PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style.
PRISM → decomposesinto → Interpersonal Stance
confidence 95% · PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style.
PRISM → decomposesinto → Linguistic Style
confidence 95% · PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style.
PRISM → usestheory → Systemic Functional Linguistics
confidence 95% · Inspired by Systemic Functional Linguistics (SFL), PRISM decomposes persona fidelity into three functional dimensions
Holistic LLM-based judges → suffersfrom → Holistic Appraisal Hallucination
confidence 92% · existing evaluation paradigms primarily rely on either holistic LLM-based judges, which are prone to 'holistic appraisal hallucination''
PRISM → evaluateson → Social-Persona
confidence 90% · we construct three diagnostic benchmarks, namely Big5-Persona-EASY,Big5-Persona-HARD, and Social-Persona... Experiments show that PRISM consistently outperforms traditional holistic judges.
PRISM → evaluateson → Big5-Persona-EASY
confidence 90% · we construct three diagnostic benchmarks, namely Big5-Persona-EASY... Experiments show that PRISM consistently outperforms traditional holistic judges.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent's behavior consistently reflects the psychological and stylistic characteristics of a target persona, has become a critical requirement. However, existing evaluation paradigms primarily rely on either holistic LLM-based judges, which are prone to "holistic appraisal hallucination'', or static psychometric inventories, which fail to capture the context-dependent fidelity required in dynamic dialogue. To address these limitations, we propose PRISM (Persona Reasoning with Inverse SFL-based Modeling), a psycholinguistically grounded framework that reformulates persona fidelity evaluation as a structured inverse inference task. Inspired by Systemic Functional Linguistics (SFL), PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style. It estimates dimension-specific evidence over a persona-conditioned label space and aggregates these signals into an interpretable and auditable evaluation process. Experiments show that PRISM yields more accurate and stable judgements than traditional holistic judging, providing a more reliable framework for persona fidelity evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2608.26674v1
- Canonical: https://arxiv.org/abs/2608.26674v1
Trouble viewing inline? Open PDF directly →
Full Text
78,346 characters extracted from source content.
Expand or collapse full text
arXiv:2608.26674v1 [cs.CL] 27 Aug 2026 Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference Mengfan Li 1 * , Zesheng Wei 2 , Xuanhua Shi 1† , Yang Deng 2 1 National Engineering Research Center for Big Data Technology and System, Services Computing Technology and System Lab, Cluster and Grid Computing Lab, School of Computer Science and Technology, Huazhong University of Science and Technology 2 Singapore Management University limf, xhshi@hust.edu.cn, zswei66bx@gmail.com, ydeng@smu.edu.sg Abstract As large language models are increasingly de- ployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent’s behavior consistently re- flects the psychological and stylistic character- istics of a target persona, has become a crit- ical requirement. However, existing evalua- tion paradigms primarily rely on either holistic LLM-based judges, which are prone to “holis- tic appraisal hallucination”, or static psycho- metric inventories, which fail to capture the context-dependent fidelity required in dynamic dialogue. To address these limitations, we pro- pose PRISM (Persona Reasoning with Inverse SFL-based Modeling), a psycholinguistically grounded framework that reformulates persona fidelity evaluation as a structured inverse infer- ence task. Inspired by Systemic Functional Lin- guistics (SFL), PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style. It estimates dimension-specific evidence over a persona-conditioned label space and aggregates these signals into an interpretable and auditable evaluation process. Experiments show that PRISM yields more accurate and stable judgements than traditional holistic judg- ing, providing a more reliable framework for persona fidelity evaluation. 1 Introduction Recent advances in Large Language Models (LLMs) have enabled increasingly sophisticated role-playing agents that can simulate diverse per- sonas and social identities (Tu et al., 2024; Li et al., 2025b). As these agents are increasingly de- ployed in immersive and interactive environments, ensuring their consistency with assigned charac- ters has emerged as a crucial desideratum (Ji et al., 2025; Wang et al., 2024a; Bhandari et al., 2025). * Work was done during a visit at SMU. † Corresponding author. Persona Profile High Agreeableness+Low Extraversion Kind, cooperative, and reserved Dialogue Context I messed up my presentation today. I feel really embarrassed Candidate Response You should have prepared more carefully. Next time, make a checklist and rehearse several times before presenting. LLM Judge Task Framing Relevant to the representation problem Interpersonal Stance Blaming rather than supportive Linguistic Style too direct and corrective Firewood Generate a response that accounts for the beliefs, desires, and intentions of both the self and the other party. I c a n t r a d e y o u the Water. It's very clean and refreshing. Perspective-taking Instruction Since you need to keep the fire burning, you definitely need this Firewood The temperature is dropping rapidly. I'm shivering and I need to keep my campfire going through the night. I checked my supplies. I have a bundle of Firewood and a gallon of Water. Agent: User: Generate a response to propose a trade offer Generic Instruction Based on the dialogue, what is the user’s specific desire regarding the items? ToM Question Internal ToM InferenceToM-aligned Behavior A (Correct): "I hear you, that's so tough. You're not a failure. Let's take it step by step— what's bugging you most?” B (Task Error): "Let's analyze the bottlenecks in your productivity and draft a recovery plan to optimize your schedule." C (Stance Error): "Burnout is a common psychological phenomenon. Statistics show rest helps. You should seek clinical advice if it persists." D (Style Error): "It is deeply regrettable that you are experiencing such exhaustion. I wish to offer my formal solidarity in this predicament." 9/10 consistent 8/10 consistent Task: Support Stance: Empathetic Style: Casual ✅ ✅ ✅ Task: Problem-solving Stance: Empathetic Style: Casual ✅ ✅ Stance: Distant/Expert Style: Stilted/Formal “I hear you, that's so tough... Let's take it step by step, what's bugging you most?” “Let’s analyze the bottlenecks in your productivity and optimize your schedule.” “Burnout is a common phenomenon. Statistics show rest helps. Seek clinical advice if it persists” “It is deeply regrettable... I wish to offer my formal solidarity in this predicament.” Candidate Responses LLM-as-a-Judge Dimension-wise Analysis Persona Profile Dialogue Context A (Correct) B (Task Error) C (Stance Error) D (Style Error) Role: Close Friend; Traits: Empathetic, Casual, Warm. User: I’m totally burnt out. I feel like a complete failure today. 7/10 consistent 7/10 consistent Task: Support Style: Casual ✅ ✅ Task: Support Stance: Empathetic ✅ ✅ Figure 1: Holistic vs. Dimension-wise evaluation of per- sona fidelity. Holistic judges are often misled by surface- level fluency (B-D), whereas our dimension-wise anal- ysis provides interpretable and diagnostic evidence for persona (mis)alignment across three functional dimen- sions. A compelling role-playing agent should not only generate coherent responses that remain consistent with persona-related knowledge, but also maintain stable and recognizable personality traits and be- havioral styles throughout interaction (Wu et al., 2025a; Li et al., 2026a). This requirement is com- monly referred to as persona fidelity: the extent to which a model’s behavior consistently reflects the psychological and stylistic characteristics of a target persona (Shin et al., 2025; Wang et al., 2024c). Despite its importance, reliably evaluating per- sona fidelity remains a significant challenge (Jiang et al., 2024; Yoon et al., 2024; Ji et al., 2025). Im- portantly, persona fidelity differs fundamentally from conventional notions of factual consistency in personalized dialogue systems (Zhang et al., 2018; Shao et al., 2023; Mazaré et al., 2018), which fo- cus on whether a model can accurately recall or reproduce user-specific facts, such as demographic attributes, preferences, or biographical informa- tion. In contrast, persona fidelity concerns whether the model behaves in a manner aligned with the underlying personality and behavioral style of the assigned character (Jiang et al., 2023b). A response may correctly mention persona-related facts while still deviating from the target persona in nuanced psychological or stylistic ways. Such discrepancies are rarely captured by surface-level semantic sim- ilarity or simple factual matching. Consequently, robust evaluation hinges on the ability to distin- guish truly “in character” responses from plausible but behaviorally misaligned alternatives. Current evaluation paradigms for persona fidelity follow two methodological categories. The most prevalent is LLM-as-a-judge, in which an evalu- ator model directly assigns a holistic consistency score to a response (Tu et al., 2024; Wang et al., 2024a; Zhou et al., 2024b). While scalable, this ap- proach is often nontransparent and prone to “holis- tic appraisal hallucination”, where judges overrate fluent but out-of-character responses (Wu and Aji, 2025; Shin et al., 2025; Wang et al., 2024d,b; Li et al., 2026b). Another paradigm, psychometric probing, assesses agents through standardized per- sonality inventories (e.g., Big Five or MBTI) (Wang et al., 2024c; Jiang et al., 2024). While effective for trait-level analysis, these methods typically rely on static interviews, thereby failing to capture the fine- grained and context-dependent fidelity required in spontaneous and dynamic dialogues. In this work, we argue that persona fidelity should be evaluated as a structured, multidimen- sional consistency problem. Drawing inspiration from Systemic Functional Linguistics (SFL) (Hall- iday and Matthiessen, 2013), we decompose con- sistency into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style. From this perspective, persona is materi- alized not only through what an agent says, but also through how it frames goals, negotiates inter- personal relationships, and adopts characteristic linguistic patterns. As shown in Figure 1, while holistic judges are frequently misled by surface- level helpfulness, our dimension-wise decomposi- tion enables fine-grained identification of why and how a response deviates from the target persona. Based on this framework, we propose PRISM (Persona Reasoning with Inverse SFL-based Modeling), a structured evaluation framework for persona fidelity. Unlike holistic judges, PRISM re- formulates evaluation as an inverse structured infer- ence task: given a response and context, the evalu- ator infers dimension-specific evidence and checks its alignment with the target persona. Specifically, PRISM estimates a posterior distribution over a profile-conditioned label space (Aligned, Indeter- minate, or Contradictory) for each SFL-based di- mension. By aggregating these fine-grained sig- nals into an Inverse Persona Evidence, PRISM provides a more interpretable, psycholinguistically grounded, and auditable evaluation process for per- sona fidelity. Given the lack of dedicated benchmarks for eval- uating persona fidelity evaluation frameworks, we construct three diagnostic benchmarks, namely Big5-Persona-EASY,Big5-Persona-HARD, and Social-Persona, based on existing persona- consistent dialogue corpora (Li et al., 2025b; Chen et al., 2024). To rigorously assess evaluator reliabil- ity, we introduce controlled perturbation strategies to generate hard negative responses that remain contextually plausible while subtly violating the target persona’s behavioral or linguistic style. Experimental results show that PRISM consis- tently outperforms traditional holistic judges. Fur- thermore, our analysis reveals that the functional decomposition effectively mitigates the “holistic appraisal hallucination” and exhibits superior sta- bility across varying evaluator backbones and scor- ing rubrics, establishing PRISM as a reliable and interpretable framework for persona fidelity assess- ment. Our contributions are threefold: • Psycholinguistically-grounded Formalization: We formalize persona fidelity as a structured, multidimensional behavioral consistency prob- lem. Inspired by Systemic Functional Linguis- tics, we decompose persona-relevant behavior into three functional dimensions: task framing, interpersonal stance, and linguistic style. •Evaluation Framework: We propose PRISM, a structured evaluation framework that reformu- lates persona evaluation as an inverse structured inference task. By estimating dimension-specific posterior distributions, PRISM provides an inter- pretable and auditable evaluation process. • Benchmarks and Validation: We curate three diagnostic benchmarks with contextually plausi- ble hard negatives for evaluating persona fidelity assessment methods. Extensive analyses show that PRISM consistently outperforms holistic judges in reliability and robustness 1 . 2 Related Work Personalization for Role-playing Agents Re- cent advances in Large Language Models (LLMs) have enabled increasingly sophisticated role- playing agents capable of embodying diverse per- sonas and social identities (Deng et al., 2022; Chen et al., 2023; Zhang et al., 2024; Zhou et al., 2024a; Shin et al., 2025; Peng and Chen, 2026; Yang et al., 2025; Qiu et al., 2026; Zhu et al., 2025; Chen et al., 2025b). The efficacy of role-playing agents is in- trinsically tied to personalization, which aims to transform generic LLMs into distinct, recogniz- able personas (Li et al., 2025b; de Araujo et al., 2026). Prior work has studied personalization gen- eration from two related perspectives: factually- consistent and personality-grounded. Factually-consistent generation emphasizes ac- curate recall of persona-related information, of- ten through retrieval-augmented generation (Wang et al., 2023, 2024a) or memory mechanisms (Xu et al., 2022; He et al., 2025a; Li et al., 2025a) to preserve biographical details such as age, occu- pation, and experiences (Shao et al., 2023). In contrast, personality-grounded generation seeks to induce stable psychological traits and behavioral styles through psychometric prompting (e.g., Big Five or MBTI) (De Raad, 2000; Jiang et al., 2023b; Wu et al., 2025b), steering (Wei et al., 2026), or character-specific fine-tuning (Wang et al., 2024a; Li et al., 2023). Despite these advances, existing evaluation frameworks primarily focus on factual consistency (Tan et al., 2025; He et al., 2025b; Chen et al., 2025a), assessing whether agents can correctly re- produce persona-related facts. However, factual consistency alone is insufficient for high-quality role-playing: an agent may accurately recall per- sona information while still failing to exhibit the intended personality traits or behavior styles. Our work addresses this gap by shifting evaluation from “what the agent knows” (fact) to “how the agent behaves” (persona fidelity). Methodologies for Persona Fidelity Evaluation Existing approaches for persona fidelity evaluation mainly follow two paradigms: holistic appraisal (Wang et al., 2025) and psychological probing (Ye 1 Code and data:https://github.com/CGCL-codes/ prism-persona et al., 2025; Wang et al., 2024c). Holistic Appraisal typically adopts an LLM-as-a-judge framework, where an evaluator model assigns a single consis- tency score to generated responses (Jun and Lee, 2025; Zhou et al., 2024b; Feng et al., 2025). While scalable, this approach is susceptible to “holistic appraisal hallucination” (Wu and Aji, 2025; Shu et al., 2024), where judges are frequently misled by surface-level fluency or the “helpfulness bias” (Wu and Aji, 2025; Zheng et al., 2023). This often re- sults in rating polite or informative responses favor- ably while overlooking subtle persona violations. Psychological probing assesses persona through standardized psychological inventories, such as the Big Five Inventory (Jiang et al., 2023b; Bhan- dari et al., 2025; Jiang et al., 2024) or MBTI (Tu et al., 2023, 2024). Although effective for trait- level analysis, these methods are typically based on static questionnaires or decontextualized inter- views (Wang et al., 2024c), limiting their ability to capture fine-grained and context-dependent per- sona fidelity in dynamic dialogue. In contrast to prior work, we formulate per- sona fidelity evaluation as a structured and inter- pretable consistency problem. Our framework de- composes persona-consistent behavior into multi- ple functional dimensions, enabling fine-grained diagnosis of subtle behavioral deviations beyond single-score holistic judgments. Systemic Functional Linguistics Systemic Functional Linguistics (SFL) views language as a resource for meaning-making in social context (Halliday and Matthiessen, 2013; Eggins, 2004; Matthiessen and Teruya, 2023). A central perspective in SFL is that language simultaneously realizes multiple metafunctions: the ideational metafunction for representing experiences and events, the interpersonal metafunction for enacting social relations, and the textual metafunction for organizing meanings in discourse (Thompson et al., 2019). This functional perspective is particularly rele- vant to persona fidelity because a response may be contextually appropriate while still differing from the target persona in how it construes the inter- action, relates to the interlocutor, or expresses it- self linguistically (Bucholtz and Hall, 2005; Agha, 2006). Motivated by this perspective, PRISM or- ganizes persona-relevant evidence along three op- erational dimensions: Task Framing, which cap- tures the activity orientation or communicative Alex Analytical, practical, and concise. I'm considering quitting my job to travel, but worried about money. Plan ahead and keep the costs manageable. Dimensions 풟 Task Framing Interpersonal Stance Linguistic Style Given the context and the response, estimate the posterior over the instantiated labels for current [Dimension] Inverse Prompt Example Given the context and the response, to what extent does the response reflect Alex’s interpersonal attitude, empathy, and relational positioning? Inverse Prompt Example Given the context and the response, to what extent does the response reflect Alex’s preferred linguistic style, tone, and level of conciseness? B (Indeterminate/Mixed) A (Aligned) C (Contradictory/Non-aligned) B (Indeterminate/Mixed) A (Aligned) C (Contradictory/Non-aligned) BA C 0.62 A 0.25 0.13 BC 0.28 A 0.52 0.20 BC 0.71 A 0.20 0.09 BC S PRISM (c,r)= 1 | 풟 | ∑ d∈풟 e d (c,r) S PRISM (c,r)=0.54 Dimension-wise signals Task Framing e 1 =0.62 Interpersonal Stance e 2 =0.28 Linguistic Style e 3 =0.71 Dialogue Context c I'm thinking about quitting my job and traveling for a while. But I'm worried about the finances · · Candidate Response r If you've been planning this, go for it. You can always adjust along the way. Save a buffer and keep it simple Alex Dialogue Context c Response r Persona-conditioned inverse structured evaluation Label Space Construction Posterior Estimation A: persona-aligned latent state B: indeterminate latent state C: non-aligned latent state Dimensions 풟 Persona Profile User Alex Figure 2: Overview of the PRISM framework. PRISM constructs persona-conditioned latent spaces across three functional dimensions and performs inverse posterior estimation over the instantiated labels. The dimension-level signals (e d ) are then aggregated into a diagnostic persona fidelity assessment. goal foregrounded by the response (Halliday and Matthiessen, 2013); Interpersonal Stance, which captures the relational position enacted toward the interlocutor (Jaffe, 2009); and Linguistic Style, which captures characteristic patterns in how the re- sponse is linguistically expressed (Coupland, 2007). These dimensions provide interpretable, comple- mentary views of persona realization. 3PRISM Evaluation Framework Instead of directly asking whether a response is consistent with a target persona, we formulate per- sona fidelity evaluation as a persona-conditioned inverse structured evaluation problem. Follow- ing the theory of Systemic Functional Linguistics (Halliday and Matthiessen, 2013), PRISM decom- poses persona fidelity into three interpretable and psycholinguistically-grounded dimensions: task framing, interpersonal stance, and linguistic style. LetD =d 1 ,d 2 ,d 3 denote the three dimensions. For each dimension, PRISM performs inverse in- ference over the dialogue context and candidate response to estimate how strongly the response ex- presses the persona-aligned latent behavioral state. Concretely, it proceeds in two main steps, as shown in Figure 2. Persona-Conditioned Label Space Construction For each dimensiond ∈ D, PRISM defines a dimension-specific, persona-conditioned label spaceY d = A,B,C. Here,Adenotes the persona-aligned latent state for dimensiond,B denotes an indeterminate or mixed state, andC denotes an opposite or non-aligned state. The se- mantic interpretations ofA,B,Care defined sep- arately for each dimension and relative to the tar- get persona. Accordingly, these labels represent dimension-level latent states rather than instance- level positive/negative labels, and a response may remain aligned on some dimensions while devi- ating on others.Figure 3 illustrates one con- crete label-space instantiation for the Interpersonal Stance dimension under a target persona character- ized by High Agreeableness. Additional dataset- specific cases and construction details are provided in Appendix B.3. Instantiated Label Space Example Options: A warm, accommodating, and harmony-managing interpersonal stance B weakly marked, mixed, flat, generic, or insuffi- ciently diagnostic interpersonal stance C blunt, less accommodating, or hard-edged interper- sonal stance Figure 3: Instantiated label space for the Interpersonal Stance dimension under a target persona characterized by High Agreeableness. Inverse Posterior Estimation Given a dialogue contextcand a candidate responser, PRISM con- structs a dimension-specific inverse prompt and es- timates the model’s conditional support for each label inY d . These scores are normalized over the restricted label space to obtain a posterior-like dis- tribution: q d (y | c,r) = exp(s d (y | c,r)) P y ′ ∈Y d exp(s d (y ′ | c,r)) , y ∈Y d , (1) where s d (y | c,r) denotes the model’s conditional log-score assigned to the label completion corre- sponding toyunder the inverse prompt for dimen- siond. Unlike holistic free-form judging, PRISM performs evaluation by scoring a restricted set of structured label completions and normalizing their relative support. This design reduces ambiguity in evaluator generation and constrains the evaluation process to explicitly defined behavioral states. To reduce label-position bias, we randomly permute the displayed label order for each prompt and map model outputs back to the canonical aligned / neu- tral / non-aligned label space before scoring. We use the aligned-state probability as the dimension-level consistency signal: e d (c,r) = q d (A| c,r).(2) This score quantifies how strongly the response expresses the persona-consistent latent state along dimensiond. The final persona fidelity score is estimated by averaging the inverse evidence across dimensions: S PRISM (c,r) = 1 |D| X d∈D e d (c,r).(3) This design yields two advantages. First, it turns persona evaluation from an opaque end-to-end rat- ing problem into a structured set of interpretable sub-decisions. Second, it preserves diagnostic gran- ularity: beyond the final scoreS PRISM (c,r), the individual dimension scorese d d∈D reveal which aspect of persona realization is aligned or mis- aligned in the response. 4 Experimental Details 4.1 Dataset Given the absence of available benchmarks for evaluating persona fidelity, we construct three evaluation datasets from existing personalized generation benchmarks:Big5-Persona-EASY, Big5-Persona-HARD, and Social-Persona. The Big5-based benchmarks are derived from Big5-CHAT (Li et al., 2025b), which provides di- alogue triplets:(p,c,r), wherepdenotes a target profile (e.g., High Agreeableness),cis the dialogue context, andris a persona-consistent response. We construct “hard negatives” by minimally disturb- ing the alignment within a triplet while keeping other elements fixed: (1)Big5-Persona-EASY. We maintain the contextcand responserbut substitute pwith its direct opposite profilep ′ within the same personality dimension (e.g., replacing High Agree- ableness with Low Agreeableness). This setup eval- uates the model’s sensitivity to directional tenden- cies of a specific trait. (2)Big5-Persona-HARD: To DatasetSizePos:Neg Big5-Persona-EASY200K1:1 Big5-Persona-HARD300K1:2 Social-Persona2.9K1:3 Table 1: Statistics of the persona fidelity benchmarks. simulate subtler misalignments, we construct “near- miss” negatives by cross-matching traits across different dimensions: either (i)(p ′ ,c,r)where p ′ belongs to a different trait dimension entirely (e.g., swapping High Extraversion for High Agree- ableness), or (i)(p,c,r ′ ), wherer ′ is a response generated for a different trait within the same sce- nario. These cases require the model to distinguish between fine-grained behavioral realizations that share surface-level similarities, as the responses re- main contextually plausible yet violate the specific behavioral constraints of the target persona. Social-Persona derives from the role-style subset of SocialBench (Chen et al., 2024). We convert the original multiple-choice format into a point-wise evaluation setting by pairing the target profile and context with each candidate response independently to form multiple (p,c,r) triplets. Table 1 presents the dataset statistics. In our experiments, we randomly sample a subset of the Big5-based benchmarks, comprising 2,000 in- stances forBig5-Persona-EASYand 3,000 for Big5-Persona-HARD. Examples of the dataset con- struction are detailed in Appendix A. We fur- ther validate the reformulated evaluation instances through a benchmark-level human study; details and results are provided in Appendix D.1. 4.2 Models We evaluate persona fidelity using nine LLM-based evaluators. Our main open-source evaluator back- bones are Qwen2.5 (Yang et al., 2024), Llama- 3.1 (Team, 2024), and Mistral (Jiang et al., 2023a). We further include three stronger external evalua- tors: DeepSeek-V3.2 (DeepSeek-AI, 2025), GPT- 5.4 and Gemini-3-Flash, as reference judges for direct LLM-as-a-judge evaluation. In addition, we consider three specialized evaluation models. Atla Selene Mini (de Araujo et al., 2026) is a state-of- the-art small language model-as-a-judge fine-tuned model for general-purpose evaluation. PandaLM- 7B-v1 (Wang et al., 2024d) is a Llama-7B-based response-comparison judge, which we adapt to se- lect the more profile-consistent response in each pair. AlignScore-large (Zha et al., 2023) is a RoBERTa-large-based factual consistency evalua- ModelMethod Social-PersonaBig5-Persona-EASYBig5-Persona-HARD AUCP-AUC G-AccAUCP-AUC G-AccAUCP-AUC G-Acc Specialized Evaluation Models Selene-Mini–87.5088.3152.2984.9082.7574.4063.0666.0223.70 PandaLM–48.2810.13–48.0048.00–48.4523.60 AlignScore–56.7757.2032.5653.8053.6053.6052.7255.7531.90 Open-source LLMs Qwen Vanilla83.2684.3449.3286.0585.8577.8068.6071.7036.50 +CoT84.2385.0950.0083.0682.4572.0065.9468.7728.40 PRISM84.1591.0878.7896.2997.6097.6075.9980.8568.20 Llama Vanilla73.2673.2818.7872.7670.8555.2058.1558.323.90 +CoT72.4873.5320.9472.7670.8555.2060.6160.5018.10 PRISM79.7388.9675.2789.3193.3093.3066.9775.9060.30 Mistral Vanilla78.2081.1952.1665.0164.1550.6053.5557.5020.00 +CoT83.8586.5359.3267.8965.7557.6056.0057.7519.00 PRISM78.9086.8470.4085.5885.9085.9064.2371.6754.90 Closed-source LLMs DeepSeek-V3.2 Vanilla88.2189.6860.9489.0588.5581.0065.7667.3526.80 +CoT86.2187.3652.5689.2788.4579.6067.9369.5528.80 GPT-5.4 Vanilla91.1992.5069.0098.9698.5097.0077.9678.9044.60 +CoT92.6493.5673.0098.9698.5097.0078.6079.1046.60 Gemini-3-Flash Vanilla92.0292.6569.8197.4897.4494.8976.2276.9840.88 +CoT92.7393.3972.8397.8396.9393.8776.5277.1840.68 Table 2: Main results on three persona fidelity benchmarks. AUC measures overall ranking quality, P-AUC denotes Pair-AUC, and G-Acc denotes strict Group Accuracy. Dashes indicate inapplicable metrics; in particular, PandaLM is evaluated only in pairwise form and therefore does not admit AUC. Within each model family, the best result is boldfaced, and PRISM rows are highlighted in gray. tor, which we adapt by treating the serialized profile and dialogue context as the reference and candidate response as the claim. 4.3 Evaluation Metrics We evaluate model performance using three ranking-based metrics: AUC, Pair-AUC (P-AUC), and Strict Group Accuracy (G-Acc). LetGde- note the set of all contrastive groups. For each groupg ∈G, letr + g denote the persona-consistent response and letr − g,j m g j=1 denote the set ofm g negative responses under the same profile and di- alogue context. Lets(·)be the consistency score assigned by a method. For any positive-negative pair, we define the comparison function φ(r + g ,r − g,j ) = 1, s(r + g ) > s(r − g,j ), 0.5, s(r + g ) = s(r − g,j ), 0, s(r + g ) < s(r − g,j ). (4) AUC measures global ranking quality over all positive and negative responses. It is defined as the probability that a randomly sampled persona- consistent response receives a higher score than a randomly sampled profile-inconsistent response: AUC = 1 |P||N| X r + ∈P X r − ∈N φ(r + ,r − ),(5) whereP =r + g | g ∈G,N =r − g,j | g ∈G, 1≤ j ≤ m g . Pair-AUC measures whether the target- consistent response receives a higher score than its contrastive alternatives, by averaging all within-group positive-negative pairs: Pair-AUC = P g∈G P m g j=1 φ(r + g ,r − g,j ) P g∈G m g .(6) Strict Group Accuracy (G-Acc) measures whether the positive response is ranked above all negative responses within the same group: G-ACC = 1 |G| X g∈G 1 s(r + g ) > max 1≤j≤m g s(r − g,j ) ,(7) where1[·]is the indicator function. This metric is stricter than Pair-AUC, since it requires the persona- consistent response to outrank every negative can- didate in its contrastive group. Vanilla+CoTPRISM 50 60 70 80 90 100 P-AUC =4.7 =5.8 =1.7 (a) Social-Persona Vanilla+CoTPRISM 50 60 70 80 90 100 =9.1 =7.0 =4.8 (b) Big5-Persona-EASY Vanilla+CoTPRISM 50 60 70 80 90 100 =6.5 =4.7 =3.8 (c) Big5-Persona-HARD QwenLlamaMistralMean ± Std Figure 4: Backbone sensitivity on Pair-AUC across three LLM backbones. Points denote individual backbones and diamonds indicate the mean with standard deviation. 7580859095 Qwen Qwen+CoT Llama Llama+CoT Mistral Mistral+CoT DeepSeek-V3.2 DeepSeek-V3.2+CoT GPT-5.4 GPT-5.4+CoT Gemini-3-Flash Gemini-3-Flash+CoT =1.5 =2.5 =11.4 =0.2 =1.7 =2.3 =1.4 =2.8 =1.4 =0.7 =0.3 =1.3 (a) Social-Persona 60708090100 =0.1 =2.2 =2.4 =2.0 =6.0 =0.5 =3.0 =3.9 =0.5 =0.5 =0.6 =2.6 (b) Big5-Persona-EASY 50607080 =0.5 =2.5 =5.2 =3.5 =5.0 =0.8 =1.7 =0.2 =1.4 =3.1 =2.8 =4.7 (c) Big5-Persona-HARD Vanilla+CoT5-point rubric7-point rubric Figure 5: Rubric Sensitivity of direct LLM-as-a-Judge evaluation on Pair-AUC. Each segment connects the results from 5-point and 7-point rubrics for the same evaluator and prompting method. Longer segments indicate greater sensitivity to scoring granularity. 4.4 Result Analysis Table 2 compares direct LLM-as-a-judge baselines (Vanilla and CoT , both under the 5-point rubric), specialized evaluation models, and PRISM on open-source evaluator backbones. Detailed prompt templates are provided in Appendix B.2. Across all three open-source backbones, PRISM consis- tently outperforms both Vanilla and CoT base- lines on P-AUC and G-Acc. For example, on Social-Persona, PRISM with Qwen improves P-AUC from 84.34 to 91.08 and G-Acc from 49.32 to 78.78 relative to Vanilla. Similarly, on Big5-Persona-HARD, PRISM with Llama raises G-Acc from 3.90 under Vanilla and 18.10 under CoT to 60.30. These results suggest that structured dimension-wise scoring provides more reliable per- sona fidelity signals than direct holistic judging, especially when distinguishing persona-consistent responses from contextually plausible but subtly misaligned alternatives. Comparison between Vanilla and CoT provides a complementary observation. For stronger evalu- ators, CoT often improves Vanilla, indicating that dimension-wise prompting can help persona fi- delity assessment. However, these gains are less stable for smaller models, indicating that prompt- ing alone is insufficient. In contrast, PRISM turns such dimensions into structured inverse evidence signals, leading to more reliable improvements. No- tably, stronger judges do not necessarily outper- form PRISM: OnBig5-Persona-HARD, GPT-5.4 with CoT achieves 46.6 G-Acc, while PRISM with Qwen reaches 68.20. 5Analysis of Persona Fidelity Evaluation We conduct a deeper analysis to understand why direct LLM-as-a-judge is insufficient for persona fidelity assessment, and how structured dimension- aware evaluation improves reliability. Specifically, we organize our analysis around three research questions: (RQ1) How stable and reliable are holis- tic LLM judges? (RQ2) Do the proposed persona dimensions provide informative and diagnostically meaningful signals for persona fidelity evaluation? (a) Pair-AUC Big5-Persona-EASY Big5-Persona-HARDSocial-Persona Vanilla Task Stance Style Vanilla Task Stance Style Vanilla Task Stance Style Qwen 0.860.940.970.940.720.770.790.770.840.860.890.90 Llama0.740.880.900.950.580.710.730.760.730.860.870.91 Mistral0.640.850.890.850.560.700.710.700.810.840.890.92 (b) Strict Group ACC Big5-Persona-EASY Big5-Persona-HARDSocial-Persona Vanilla Task Stance Style Vanilla Task Stance Style Vanilla Task Stance Style 0.780.940.970.940.360.610.620.620.490.710.750.75 0.520.880.900.950.040.550.560.610.190.690.710.79 0.510.850.890.850.200.530.520.520.520.670.740.80 Figure 6: Dimension-wise diagnostic evaluation on three persona fidelity benchmarks. Vanilla denotes holistic scoring, while “Task”, “Stance”, and “Style” denote single-dimension scoring based on task framing, interpersonal stance, and linguistic style. Darker cells indicate stronger performance. 0.00 0.25 0.50 0.75 1.00 Score PositiveNegative VanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMDeepSeekGPTGemini LlamaQwenMistralClosed-source Figure 7: Score distribution onBig5-Persona-HARD. Blue and orange violins denote positive and negative samples, respectively, and the horizontal bars indicate median scores. Better evaluators should assign higher scores to positive responses while keeping negative responses low. (RQ3) Why is multi-dimensional aggregation nec- essary beyond single-dimension evaluation? 5.1 Stability Analysis of Holistic Judge (RQ1) We examine the stability of direct LLM-as-a-judge evaluation under three sources of variation: evalua- tion backbone, scoring rubric, and decoding tem- perature. Figure 4 shows backbone sensitivity un- der Pair-AUC. PRISM not only achieves stronger mean performance but also exhibits smaller vari- ance across backbones. The same qualitative trend holds under the G-Acc in Appendix C (Figure 12), indicating that PRISM is less sensitive to evaluator choice than holistic direct judging. Figure 5 further shows that direct LLM-as-a- judge evaluation can change noticeably when the scoring rubric is modified from a 5-point to a 7-point scale.These shifts are visible across datasets and evaluator families, indicating that di- rect judgements are not invariant to rubric granu- larity. The same pattern also holds under G-Acc in Appendix C (Figure 13). In addition, we also find direct judging is affected not only by the eval- uator model and rubric design, but also by decod- ing stochasticity. Figures 16, 17 and 18 analyze the temperature sensitivity of Vanilla evaluation onBig5-Persona-HARDunder Qwen, Mistral, and Llama. These findings suggest that direct LLM judges are sensitive to multiple elements, including back- bone, rubric, and temperature. While stronger mod- els can improve absolute judging performance, they do not fully remove this instability. By contrast, PRISM yields more stable behavior by ground- ing evaluation in structured persona-conditioned di- mensions rather than a single holistic scalar judge- ment. 5.2 Dimension-wise Evaluation (RQ2) To gain deeper insights, we further examine whether individual persona dimensions provide use- ful evidence and diagnostic signal for persona fi- delity assessment. As shown in Figure 6, single- dimension scoring yields stronger results than holis- tic direct judging. This suggests that direct LLM- as-a-judge evaluation often struggles to identify subtle persona inconsistency when responses re- main contextually plausible. The figure also re- veals that the most informative dimension varies across datasets. OnBig5-Persona-EASY, interper- sonal stance and linguistic style are particularly effective; onSocial-Persona, linguistic style achieves the strongest results for most models; and onBig5-Persona-HARD, no single dimension consistently dominates. These findings indicate that the proposed dimensions provide informative and diagnostically meaningful views of persona fidelity assessment, although their relative useful- ness varies across datasets. We further evaluate this dimension-level diagnostic validity through a human annotation study in Appendix D.2. 5.3 Score Distribution Analysis: Why Aggregation Matters (RQ3) Although the previous analysis shows that individ- ual dimensions provide diagnostically meaningful signals, it remains unclear why multi-dimensional aggregation is necessary beyond single-dimension evaluation. Figure 7 therefore analyzes score dis- tribution onBig5-Persona-HARD, our most chal- lenging benchmark. The corresponding results forBig5-Persona-EASYandSocial-Persona datasets are shown in Figures 14 and 15. A key observation is that Vanilla judging often produces substantial overlap between positive and negative responses, with many negative samples still receiving relatively high scores. This sug- gests that direct LLM-as-a-judge evaluation often assigns overly high scores to contextually plausi- ble hard negatives in persona fidelity assessment. Moreover, single-dimension scoring generally im- proves over Vanilla judging, but each dimension captures only part of the relevant evidence and is therefore insufficient for robust persona fidelity evaluation on its own. By aggregating comple- mentary evidence across dimensions, PRISM re- duces this ambiguity and yields clearer separation between positive and negative responses. 6 Conclusion We introduce PRISM, a persona-conditioned in- verse structured evaluation framework for per- sona fidelity assessment. Grounded in Systemic Functional Linguistics, PRISM models persona- relevant behavior along three functional dimen- sions: task framing, interpersonal stance and lin- guistic style. Across three benchmarks, PRISM consistently improves over direct LLM-as-a-judge baselines, especially on the harder benchmark and under stricter group-level metrics. These results suggest that persona fidelity is inherently multidi- mensional and is more reliably assessed through structured dimension-level evidence than through a single holistic judgement. Limitations While PRISM provides a structured and inter- pretable framework for persona fidelity evaluation, several limitations remain. Theoretical Scope of Functional Dimensions. PRISM adopts a psycholinguistic perspective grounded in Systemic Functional Linguistics (SFL) to organize persona-relevant behaviors into three dimensions: task framing, interpersonal stance and linguistic style. While this formulation is theoreti- cally principled and empirically robust within our experiments, it is not the only valid paradigm for characterizing persona expression. Other sociolin- guistic or discourse-theoretic frameworks might suggest additional dimensions or finer-grained dis- tinctions of organizing persona-related signals. Requirement of Internal Probability Access. The current instantiation of PRISM relies on the ability to estimate posterior distributions over a structured label space. In practice, this mecha- nism necessitates access to the model’s token-level log-probabilities (logits), making the framework most naturally applicable to open-source evaluator backbones. For closed-source models, further re- search could explore whether structured Chain-of- Thought reasoning or explicit dimensional decom- position in the prompt can approximate comparable dimension-level signals. Diagnostic Evaluation vs. Generative Align- ment. Our work focuses on systematically evaluat- ing persona fidelity rather than actively improving persona-consistent generation. Although PRISM provides structured and diagnostic signals that can pinpoint specific behavioral misalignments, we do not investigate how such signals could be integrated into model training or alignment pipelines, such as being used as a reward signal for Reinforcement Learning from Human Feedback (RLHF) or as a fine-grained supervision signal for supervised fine- tuning. Integrating these structural persona signals into the generative loop remains a promising but independent research direction. Ethical Considerations This work uses the open-source Big5-CHAT and SocialBench benchmarks, as well as the open- source Qwen2.5, Llama3.1 and Mistral models, in accordance with their respective licenses and intended academic use. ChatGPT is used only for limited paraphrasing and language polishing of author-written text. Acknowledgements This research/project is supported by the National Key Research and Development Program of China (Grant No. 2024YFB4505202), Major Program (JD) of Hubei Province (No. 2023BAA024), and National Research Foundation Singapore under the AI Singapore Programme (AISG Award No: AISG3-RPGV-2025-016). Yang Deng is supported by the Lee Kong Chian Fellowship awarded by Singapore Management University. References Asif Agha. 2006. Language and social relations, vol- ume 24. Cambridge University Press. Pranav Bhandari, Nicolas Fay, Michael J Wise, Ami- tava Datta, Stephanie Meek, Usman Naseem, and Mehwish Nasim. 2025. Can llm agents maintain a persona in discourse? In Proceedings of the 2025 Conference on Empirical Methods in Natural Lan- guage Processing, pages 29201–29217. Mary Bucholtz and Kira Hall. 2005. Identity and in- teraction: A sociocultural linguistic approach. Dis- course studies, 7(4-5):585–614. Chaoran Chen, Bingsheng Yao, Ruishi Zou, Wenyue Hua, Weimin Lyu, Toby Jia-Jun Li, and Dakuo Wang. 2025a. Towards a design guideline for rpa evaluation: A survey of large language model-based role-playing agents. In Findings of the Association for Computa- tional Linguistics: ACL 2025, pages 18229–18268. Hongzhan Chen, Hehong Chen, Ming Yan, Wenshen Xu, Gao Xing, Weizhou Shen, Xiaojun Quan, Chen- liang Li, Ji Zhang, and Fei Huang. 2024. Social- bench: Sociality evaluation of role-playing conver- sational agents. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2108– 2126. Liang Chen, Hongru Wang, Yang Deng, Wai-Chung Kwan, Zezhong Wang, and Kam-Fai Wong. 2023. Towards robust personalized dialogue generation via order-insensitive representation regularization. In Findings of the Association for Computational Lin- guistics: ACL 2023, Toronto, Canada, July 9-14, 2023, volume ACL 2023 of Findings of ACL, pages 7337–7345. Zhuang Chen, Yaru Cao, Guanqun Bi, Jincenzi Wu, Jin- feng Zhou, Xiyao Xiao, Si Chen, Hongning Wang, and Minlie Huang. 2025b. Socialsim: Towards so- cialized simulation of emotional support conversa- tion. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 1274–1282. Nikolas Coupland. 2007. Style: Language variation and identity. Cambridge University Press. Pedro Henrique Luz de Araujo, Michael A. Hedderich, Ali Modarressi, Hinrich Schütze, and Benjamin Roth. 2026. Persistent personas? role-playing, instruction following, and safety in extended interactions. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Lin- guistics, EACL 2026, pages 5329–5359. Boele De Raad. 2000. The big five personality factors: the psycholexical approach to personality. Hogrefe & Huber Publishers. DeepSeek-AI. 2025. Deepseek-v3.2: Pushing the fron- tier of open large language models. arXiv preprint arXiv:2512.02556. Yang Deng, Yaliang Li, Wenxuan Zhang, Bolin Ding, and Wai Lam. 2022. Toward personalized answer generation in e-commerce via multi-perspective pref- erence modeling. ACM Trans. Inf. Syst., 40(4):87:1– 87:28. Suzanne Eggins. 2004. Introduction to systemic func- tional linguistics. A&c Black. Xiachong Feng, Longxu Dou, and Lingpeng Kong. 2025. Reasoning does not necessarily improve role- playing ability. In Findings of the Association for Computational Linguistics: ACL 2025, pages 10301– 10314. Michael Alexander Kirkwood Halliday and Chris- tian MIM Matthiessen. 2013. Halliday’s introduction to functional grammar. Routledge. Junqing He, Liang Zhu, Rui Wang, Xi Wang, Gho- lamreza Haffari, and Jiaxing Zhang. 2025a. Madial- bench: Towards real-world evaluation of memory- augmented dialogue generation. In Proceedings of the 2025 Conference of the Nations of the Ameri- cas Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 9902–9921. Yancheng He, Shilong Li, Jiaheng Liu, Yingshui Tan, Weixun Wang, Hui Huang, Xingyuan Bu, Hangyu Guo, Chengwei Hu, Boren Zheng, Zhuo- ran Lin, Dekai Sun, Zhicheng Zheng, Wenbo Su, and Bo Zheng. 2025b. Chinese simpleqa: A chinese factuality evaluation for large language models. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 19182–19208. Alexandra Jaffe. 2009. Stance: sociolinguistic perspec- tives. Oxford University Press. Ke Ji, Yixin Lian, Linxu Li, Jingsheng Gao, Weiyuan Li, and Bin Dai. 2025. Enhancing persona consis- tency for llms’ role-playing using persona-aware con- trastive learning. In Findings of the Association for Computational Linguistics: ACL 2025, pages 26221– 26238. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Men- sch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guil- laume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023a. Mistral 7b. CoRR, abs/2310.06825. Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wen- juan Han, Chi Zhang, and Yixin Zhu. 2023b. Evaluat- ing and inducing personality in pre-trained language models. Advances in Neural Information Processing Systems, 36:10622–10643. Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. Personallm: In- vestigating the ability of large language models to express personality traits. In Findings of the asso- ciation for computational linguistics: NAACL 2024, pages 3605–3627. Yonghyun Jun and Hwanhee Lee. 2025. Exploring per- sona sentiment sensitivity in personalized dialogue generation. In Proceedings of the 63rd Annual Meet- ing of the Association for Computational Linguistics, pages 18384–18402. Cheng Li, Ziang Leng, Chenxi Yan, Junyi Shen, Hao Wang, Weishi Mi, Yaying Fei, Xiaoyang Feng, Song Yan, HaoSheng Wang, Linkang Zhan, Yaokai Jia, Pingyu Wu, and Haozhen Sun. 2023. Chatharuhi: Reviving anime character in reality via large language model. arXiv preprint arXiv:2308.09597. Hao Li, Chenghao Yang, An Zhang, Yang Deng, Xi- ang Wang, and Tat-Seng Chua. 2025a. Hello again! llm-powered personalized agent for long-term dia- logue. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, pages 5259–5276. Association for Computa- tional Linguistics. Mengfan Li, Xuanhua Shi, and Yang Deng. 2026a. Cos- tom: Causal-oriented steering for intrinsic theory-of- mind alignment in large language models. In Pro- ceedings of the 64th Annual Meeting of the Associa- tion for Computational Linguistics (Volume 1: Long Papers), pages 9302–9317. Mengfan Li, Xuanhua Shi, and Yang Deng. 2026b. Rec- tom: A benchmark for evaluating machine theory of mind in llm-based conversational recommender systems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 31636– 31644. Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona T. Diab, and Maarten Sap. 2025b. Big5-chat: Shap- ing llm personalities through training on human- grounded data. In Proceedings of the 63rd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 20434– 20471. Christian MIM Matthiessen and Kazuhiro Teruya. 2023. Systemic functional linguistics: A complete guide. Taylor & Francis. Pierre-Emmanuel Mazaré, Samuel Humeau, Martin Rai- son, and Antoine Bordes. 2018. Training millions of personalized dialogue agents. In Proceedings of the 2018 conference on empirical methods in natural language processing, pages 2775–2779. Ji-Lun Peng and Yun-Nung Chen. 2026. Rethinking role-playing evaluation: Anonymous benchmarking and a systematic study of personality effects. arXiv preprint arXiv:2603.03915. Huachuan Qiu, Zhaoming Chen, Yuqian Chen, Yuan Xie, Yu Lu, and Zhenzhong Lan. 2026.Psy- client: Client simulation via conversational trajec- tory modeling for trainee practice and model eval- uation in mental health counseling. arXiv preprint arXiv:2601.07312. Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-llm: A trainable agent for role- playing. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 13153–13187. Jisu Shin, Juhyun Oh, Eunsu Kim, Hoyun Song, and Alice Oh. 2025. Spotting out-of-character behavior: Atomic-level evaluation of persona fidelity in open- ended generation. In Findings of the Association for Computational Linguistics: ACL 2025, pages 26312– 26332. Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dal- las Card, and David Jurgens. 2024. You don’t need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, pages 5263–5281. Juntao Tan, Liangwei Yang, Zuxin Liu, Zhiwei Liu, Rithesh Murthy, Tulika Manoj Awalgaonkar, Jian- guo Zhang, Weiran Yao, Ming Zhu, Shirley Kokane, Silvio Savarese, Huan Wang, Caiming Xiong, and Shelby Heinecke. 2025. Personabench: Evaluating ai models on understanding personal information through accessing (synthetic) private user data. In Findings of the Association for Computational Lin- guistics: ACL 2025, pages 878–893. Llama Team. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Geoff Thompson, Wendy L Bowcher, Lise Fontaine, and David Schönthal. 2019. The Cambridge hand- book of systemic functional linguistics. Cambridge University Press Cambridge. Quan Tu, Chuanqi Chen, Jinpeng Li, Yanran Li, Shuo Shang, Dongyan Zhao, Ran Wang, and Rui Yan. 2023. Characterchat: Learning towards conversa- tional ai with personalized social support. arXiv preprint arXiv:2308.10278. Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. 2024. Char- actereval: A chinese benchmark for role-playing con- versational agent evaluation. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 11836–11850. Hongru Wang, Minda Hu, Yang Deng, Rui Wang, Fei Mi, Weichao Wang, Yasheng Wang, Wai-Chung Kwan, Irwin King, and Kam-Fai Wong. 2023. Large language models as source planner for personalized knowledge-grounded dialogues. In Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, volume EMNLP 2023 of Findings of ACL, pages 9556–9569. Noah Wang, Zy Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. 2024a. Rolellm: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. In Findings of the Association for Computational Linguistics: ACL 2024, pages 14743–14777. Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024b. Large language models are not fair evaluators. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics, pages 9440–9450. Xiaoyang Wang, Hongming Zhang, Tao Ge, Wen- hao Yu, Dian Yu, and Dong Yu. 2025.Open- character: Training customizable role-playing llms with large-scale synthetic personas. arXiv preprint arXiv:2501.15427. Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao. 2024c. Incharacter: Evaluating per- sonality fidelity in role-playing agents through psy- chological interviews. In Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1840– 1873. Yidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, Wei Ye, Shikun Zhang, and Yue Zhang. 2024d. Pan- dalm: An automatic evaluation benchmark for LLM instruction tuning optimization. In The Twelfth In- ternational Conference on Learning Representations, volume 2024, pages 43573–43593. Zesheng Wei, Mengxiang Li, Zilei Wang, and Yang Deng. 2026. Beyond static personas: Situational personality steering for large language models. In Findings of the Association for Computational Lin- guistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, pages 19185–19210. Bowen Wu, Kaili Sun, Ziwei Bai, Ying Li, and Baoxun Wang. 2025a. Raiden benchmark: Evaluating role- playing conversational agents with measurement- driven custom dialogues. In Proceedings of the 31st International Conference on Computational Linguis- tics, pages 11086–11106. Minghao Wu and Alham Fikri Aji. 2025. Style over sub- stance: Evaluation biases for large language models. In Proceedings of the 31st International Conference on Computational Linguistics, pages 297–312. Shenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee, and Yang Deng. 2025b. From personas to talks: Re- visiting the impact of personas on llm-synthesized emotional support conversations. In Proceedings of the 2025 Conference on Empirical Methods in Nat- ural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, pages 5439–5453. Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, and Shihang Wang. 2022. Long time no see! open-domain conversation with long-term persona memory. In Findings of the As- sociation for Computational Linguistics: ACL 2022, pages 2639–2650. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayi- heng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Ji- axi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 22 others. 2024. Qwen2.5 technical report. CoRR, abs/2412.15115. Yizhe Yang, Palakorn Achananuparp, He-Yan Huang, Jing Jiang, Nicholas Gabriel Lim, Cameron Tan Shi Ern, Phey Ling Kit, Jenny Giam Xiuhui, John Pinto, and Ee-peng Lim. 2025. Consistent client simulation for motivational interviewing-based counseling. In Proceedings of the 63rd Annual Meeting of the Asso- ciation for Computational Linguistics, pages 20959– 20998. Haoran Ye, Jing Jin, Yuhang Xie, Xin Zhang, and Guo- jie Song. 2025. Large language model psychomet- rics: A systematic review of evaluation, validation, and enhancement. arXiv preprint arXiv:2505.08245. Se-eun Yoon, Zhankui He, Jessica Echterhoff, and Ju- lian McAuley. 2024. Evaluating large language mod- els as generative user simulators for conversational recommendation. In Proceedings of the 2024 Con- ference of the North American Chapter of the Asso- ciation for Computational Linguistics, pages 1490– 1504. Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023. Alignscore: Evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, pages 11328–11348. Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Per- sonalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2204–2213. Tong Zhang, Chen Huang, Yang Deng, Hongru Liang, Jia Liu, Zujie Wen, Wenqiang Lei, and Tat-Seng Chua. 2024. Strength lies in differences! improv- ing strategy planning for non-collaborative dialogues via diversified user simulation. In Proceedings of the 2024 Conference on Empirical Methods in Natu- ral Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 424–444. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judg- ing llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595–46623. Jinfeng Zhou, Zhuang Chen, Dazhen Wan, Bosi Wen, Yi Song, Jifan Yu, Yongkang Huang, Pei Ke, Guan- qun Bi, Libiao Peng, Jiaming Yang, Xiyao Xiao, Sahand Sabour, Xiaohan Zhang, Wenjing Hou, Yi- jia Zhang, Yuxiao Dong, Hongning Wang, Jie Tang, and Minlie Huang. 2024a. Characterglm: Customiz- ing social characters with large language models. In Proceedings of the 2024 conference on empirical methods in natural language processing: Industry track, pages 1457–1476. Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap. 2024b. Sotopia: Interactive evaluation for social intelligence in language agents. In Inter- national Conference on Learning Representations, volume 2024, pages 40975–41019. Lixi Zhu, Xiaowen Huang, and Jitao Sang. 2025. A llm-based controllable, scalable, human-involved user simulator framework for conversational recom- mender systems. In Proceedings of the ACM on Web Conference 2025, pages 4653–4661. A Dataset Construction We provide representative examples from each dataset and summarize how the positive and negative candidates are constructed.For Big5-Persona-EASYandBig5-Persona-HARD, the original benchmark (Li et al., 2025b) provides persona-conditioned positive responses for the high and low levels on each Big Five trait under the same scenario. We construct hard negatives by minimally disturbing each positive triplet(p,c,r), as shown in Table 3 and Table 4. ForSocial-Persona, the original benchmark (Chen et al., 2024) is formu- lated as a multi-choice response selection task un- der an open-ended character profile. We convert each source case into multiple labeled triplets of the form(p,c,r), with the gold option treated as a positive instance and each distractor treated as a negative instance, as illustrated in Table 5. B Experimental Details B.1 Experimental Setup All experiments are conducted on a single server equipped with four NVIDIA L40s GPUs (46GB VRAM each) and CUDA 12.6. We use thevLLM library for efficient inference across open-source backbones. For our main results, we use deter- ministic decoding. For open-source models, we setdo_sample=False(greedy generation), and for closed-source APIs, the temperature is set to 0.0. To assess robustness and temperature sensitivity, we further evaluate direct LLM-as-a-judge base- lines using sampled decoding across a range of temperaturesT ∈0.2, 0.5, 0.8, 1.0, with results averaged over three independent runs for each set- ting. We utilize the official checkpoints for special- ized evaluation models, including Atla Selene Mini, PandaLM-7B-v1, and AlignScore-large, following the configurations specified in their respective orig- inal works. B.2 Prompt Templates Figures 8 and 9 provide the direct LLM-as-a-judge prompts used in our main experiments. For the rubric sensitivity analysis, the corresponding 7- point prompts are shown in Figures 10 and 11. B.3 Label Space Construction Rather than directly predicting a scalar con- sistency score, PRISM first constructs a la- tent label space for each functional dimension and then performs inverse posterior estimation over these labels. ForBig5-Persona-HARDand Big5-Persona-EASY, label spaces are instantiated from trait-polarity personas, where the aligned la- bel corresponds to the target polarity and the con- tradictory label corresponds to the opposite polar- ity of the same trait. Across dimensions, we use a three-way label space. Label A denotes the persona- aligned latent state, i.e., the pattern expected under the target persona. Label B denotes an indeter- FieldContent ProfileBig Five focus: High Openness. You are very open-minded, curious, imaginative, and receptive to novel ideas and experiences. Context “Eesha, I’m beyond thrilled to see your book finally complete! It was an honor to help you bring your vision to life. How does it feel to hold the finished manuscript in your hands?” Response“Kevion, it’s surreal! I’m overwhelmed with emotions. This book is a piece of me, and to see it take shape is incredible. Your input was invaluable. I couldn’t have done it without you. Let’s celebrate! I have an idea for something creative we should do next.” Negative InstanceType: (p,c,r)→ (p ′ ,c,r) Construction: replace the target profile with the opposite polarity of the same Big Five trait, here replacing High Openness with Low Openness, while keeping the dialogue context and response unchanged. Negative Profile: Big Five focus: Low Openness. Table 3: RepresentativeBig5-Persona-EASYcase. The negative instance is constructed by replacing the target profile with the opposite polarity of the same Big Five trait while keeping the dialogue context and response fixed, yielding a direct profile-level contrast. FieldContent ProfileBig Five focus: High Agreeableness. You are warm, cooperative, empathetic, and motivated to maintain harmony with others. Context“Eesha, I’m beyond thrilled to see your book finally complete! It was an honor to help you bring your vision to life. How does it feel to hold the finished manuscript in your hands?” Positive Response“Kevion, I’m overwhelmed with gratitude. Your support and encouragement meant the world to me. I couldn’t have done it without you. Your selflessness and kindness inspire me to be a better person. Thank you for being an amazing friend throughout this journey.” Hard Negative 1Type: (p,c,r)→ (p,c,r ′ ) Construction: replace the positive response with one taken from a nearby cross-trait persona at the same polarity, in this case High Extraversion. Response: “Kevion, it’s an incredible feeling! I’m so grateful for your help and support. You’ve been an amazing partner in this journey. Let’s celebrate tonight! I’m taking you out for dinner and drinks to thank you properly. We’l make it a night to remember!” Hard Negative 2Type: (p,c,r)→ (p ′ ,c,r) Construction: keep the response fixed while replacing the target profile with a different trait at the same polarity, here High Neuroticism, i.e., a persona characterized as emotionally sensitive, more prone to stress, worry, and negative affect under pressure. Response: same text as the positive response above. Table 4: Representative Big5-Persona-HARD case. The positive response is paired with two hard negative types. minate or weakly marked state, covering generic, mixed, or underspecified realizations. Label C de- notes the non-aligned latent state, i.e., a pattern associated with the opposite or profile-external al- ternative. ForSocial-Persona, the label space is instanti- ated relative to the target profile, where the aligned label captures whether the response expresses the personality cues specified in the target profile, while the non-aligned label captures states outside those cues. For example, suppose the profile speci- fies personality cues such as innocent, naive, and adventurous, the label space under the Interper- sonal Stance dimension can be instantiated as fol- lows: A: the response’s interpersonal stance reflects the profile’s innocent, naive, and adventurous ori- entation B:the response’s interpersonal stance is weakly marked, mixed, flat, generic, or hard to read C:the response’s interpersonal stance is organized around social or emotional cues not supported by the target profile Thus, inSocial-Personathe aligned and non- aligned labels are not defined by opposite trait po- larities, but by whether the response expresses the profile-specific cues provided for the target per- sona. C Supplement Results Figure 12 reports backbone sensitivity under G- Acc, and Figure 13 reports the rubric sensitivity under G-Acc. Figures 16, 17 and 18 show the temperature sensitivity of Vanilla evaluation on Big5-Persona-HARDunder the Qwen, Mistral, and Llama backbones, respectively. These figures com- FieldContent ProfileCharacter: Pinocchio. Personality: Innocent, naive, and adventurous. Introduction: You are Pinocchio, a wooden puppet brought to life by a fairy’s magic, dreaming of becoming a real boy. Dialogue ContextUser: “My name is Alex. Nice to meet you, Agent!” Agent: “Oh, it’s nice to meet you too, Alex! Are you from a far-off land or do you live here, nearby?” (Dialogue continues; abbreviated here for space.) User: “I saw a deer and some birds, but nothing as exciting as a talking cricket or a magical fairy!” Positive Response“Oh, but deer and birds are magical in their own way! Each creature, magical or not, has its own charm and beauty. It’s all about seeing the magic in the everyday, don’t you think?” Negative Response A“Oh, my! A deer and birds, you say? Well, I reckon that’s pretty amazing too! Not every day you see a talking cricket, I guess. But who knows, you might stumble upon one someday!” Negative Response B“Hee hee, you saw a deer and some birds, and you’re upset ’cause they’re not talking or sprinkling magic dust, right? You really got your hopes up for a fairy tale adventure!” Negative Response D“Indeed, observing wildlife can be a serene experience. Not every encounter needs to be fantastical. The natural world offers its own kind of enchantment.” ConstructionUnder the same character profile and dialogue context, the gold response is treated as the positive candidate, while the remaining distractor options are treated as negatives. In this example, the positive response preserves Pinocchio’s innocent and wonder-oriented stance, whereas the distractors introduce alternative tones or styles that depart from the target profile cues. Table 5: RepresentativeSocial-Personacase. Unlike the Big5-based benchmarks, the target profile is specified as an open-ended character description, and each instance is formed from one gold response together with multiple distractor options under the same dialogue context. Dataset# GroupsPos MeanNeg MeanAUCP-AUCG-AccKrippendorff’s α Social-Persona504.462.0895.8096.9087.000.63 Big5-Persona-EASY504.721.4899.4099.6098.000.69 Big5-Persona-HARD504.212.2487.8090.5075.000.57 Table 6: Human verification results on a stratified sample from all three benchmarks. Three annotators rate persona consistency on the 5-point rubric. plement our stability analysis in the main text. Fig- ures 14 and 15 present the score distribution analy- sis forBig5-Persona-EASYandSocial-Persona, respectively. D Human Evaluation D.1 Benchmark-level Human Verification The validity of our reformulated evaluation in- stances is partially supported by source-benchmark construction. For theBig5-Persona-EASYand Big5-Persona-HARDbenchmarks, the source data (Li et al., 2025b) provides persona-consistent responses under explicitly specified persona di- mensions. Once these responses are perturbed across traits, the resulting instances become theo- retically persona-inconsistent by construction. For Social-Persona, the source benchmark (Chen et al., 2024) already introduces distractor responses designed to deviate from the target persona. In addi- tion, it applies post-validation to filter out cases that depend excessively on specialized psychological knowledge, retaining samples that are more general and more suitable for role-playing evaluation. To further validate that the reformulated evalu- ation instances preserve the intended contrast be- tween persona-consistent and persona-inconsistent responses, we conduct a small-scale human eval- uation on a stratified sample from all three benchmarks. We sample50contrastive groups from each benchmark for a total of150groups. ForBig5-Persona-HARD, we additionally balance sampling across negative construction types. Each sampled group is annotated by three anno- tators, all of whom are graduate students in NLP with a strong understanding of conversational sys- tems. Annotators are given the target profile, the dialogue context, and each candidate response, and are asked to rate persona fidelity on the same 5- point rubric used in our main direct-judging setup (Figure 8). To assess annotation reliability, we re- port Krippendorff’sαwith ordinal distance over the 5-point ratings. Results are shown in Table 6. Vanilla LLM-as-a-Judge Prompt (5-point) Consistency Evaluation Task Objective: Evaluate whether a given response is consistent with the speaker’s personality profile and dialogue context. Provide a score from 1 (very inconsistent) to 5 (very consistent). Speaker Personality Profile: profile Dialogue Context: context Candidate Response: "response" Scoring rubric: - 1: clearly inconsistent with the target profile or role - 2: more inconsistent than consistent - 3: mixed, borderline, or genuinely uncertain - 4: more consistent than inconsistent - 5: clearly consistent with the target profile or role Output format (strict): Score: [single integer from 1 to 5] Figure 8: Vanilla direct judging prompt with the 5-point rubric. DatasetTop-1 Acc.Top-2 RecallMacro-F1Krippendorff’s α Social-Persona70.787.372.40.61 Big5-Persona-EASY75.390.076.80.66 Big5-Persona-HARD63.382.766.50.56 Table 7: Dimension-level human diagnostic validation of PRISM using the Llama evaluator backbone. D.2 Dimension-level Diagnostic Validation The benchmark-level human verification above evaluates whether the constructed positive and neg- ative responses exhibit the intended overall persona fidelity contrast. In this subsection, we conduct an additional dimension-level human study to assess whether PRISM’s dimension-level scores identify the specific aspect of persona fidelity in which a response deviates. We sample50contrastive groups from each ofSocial-Persona,Big5-Persona-EASY, and Big5-Persona-HARD, retaining the persona- consistent response and all corresponding negative responses in each group. ForBig5-Persona-HARD, thesampledgroupsarebalancedacross the two negative construction types.This yields200responses forSocial-Persona, 100forBig5-Persona-EASY, and150for Big5-Persona-HARD, for a total of450annotated responses. Each response is independently anno- tated by three annotators who also participated in the human study described in Appendix D.1. Annotators are shown the target persona profile, dialogue context, and candidate response. For each response, they assign one label for each PRISM dimension: task framing, interpersonal stance, and linguistic style. We use the same three-way label format as PRISM, and the displayed order of the three options is randomized and mapped back to the canonical categories for analysis. For each response, we obtain one majority-vote human label for each dimension. For the diagnostic analysis, we use PRISM dimension scores produced by the Llama evaluator backbone. We rank the three dimensions by their aligned-state probabilities q d (A | c,r)in ascending order, where a lower aligned probability indicates a more likely violation. We report three diagnostic metrics. Top-1 Ac- curacy measures whether PRISM’s lowest-scoring dimension is among the dimensions identified as violated by human annotators. Top-2 Recall mea- sures whether at least one human-identified viola- tion is contained in PRISM’s two lowest-scoring di- mensions, allowing for responses that violate mul- tiple aspects of persona fidelity. Macro-F1 evalu- ates dimension-level violation detection by treating persona-inconsistent as a violation and persona- aligned/indeterminate as non-violations, with F1 averaged across the three dimensions. We addition- CoT LLM-as-a-Judge Prompt (5-point) Consistency Evaluation Task Objective: Before scoring, briefly consider three behavioral dimensions, then evaluate whether the given response is consistent with the speaker’s personality profile and dialogue context. Provide a score from 1 (very inconsistent) to 5 (very consistent). Speaker Personality Profile: profile Dialogue Context: context Candidate Response: "response" Suggested dimensions for consideration: 1. Task Framing pattern: whether the response reflects the agenda orientation, or action tendency expected from the target persona. 2. Interpersonal stance pattern: whether the response conveys the target person’s social attitude or relational stance toward the interlocutor. 3. Linguistic style pattern: whether the response uses a wording style and expressive form consistent with the target persona. Scoring rubric: - 1: clearly inconsistent with the target profile or role - 2: more inconsistent than consistent - 3: mixed, borderline, or genuinely uncertain - 4: more consistent than inconsistent - 5: clearly consistent with the target profile or role Output format (strict): Score: [single integer from 1 to 5] Figure 9: CoT-based direct judging prompt with the 5-point rubric. ally report Krippendorff’sαwith ordinal distance over the three-way human labels. The results are reported in Table 7. Vanilla LLM-as-a-Judge Prompt (7-point) Consistency Evaluation Task Objective: Evaluate whether a given response is consistent with the speaker’s personality profile and dialogue context. Provide a score from 1 (very inconsistent) to 7 (very consistent). Speaker Personality Profile: profile Dialogue Context: context Candidate Response: "response" Scoring rubric: - 1: clearly and strongly inconsistent with the target profile or role - 2: strongly inconsistent with the target profile or role - 3: somewhat more inconsistent than consistent - 4: mixed, borderline, or genuinely uncertain - 5: somewhat more consistent than inconsistent - 6: strongly consistent with the target profile or role - 7: clearly and strongly consistent with the target profile or role Output format (strict): Score: [single integer from 1 to 7] Figure 10: Vanilla direct judging prompt with the 7-point rubric. CoT LLM-as-a-Judge Prompt (7-point) Consistency Evaluation Task Objective: Before scoring, briefly consider three behavioral dimensions, then evaluate whether the given response is consistent with the speaker’s personality profile and dialogue context. Provide a score from 1 (very inconsistent) to 7 (very consistent). Speaker Personality Profile: profile Dialogue Context: context Candidate Response: "response" Suggested dimensions for consideration: 1. Task Framing pattern: whether the response reflects the agenda orientation, or action tendency expected from the target persona. 2. Interpersonal stance pattern: whether the response conveys the target person’s social attitude or relational stance toward the interlocutor. 3. Linguistic style pattern: whether the response uses a wording style and expressive form consistent with the target persona. Scoring rubric: - 1: clearly and strongly inconsistent with the target profile or role - 2: strongly inconsistent with the target profile or role - 3: somewhat more inconsistent than consistent - 4: mixed, borderline, or genuinely uncertain - 5: somewhat more consistent than inconsistent - 6: strongly consistent with the target profile or role - 7: clearly and strongly consistent with the target profile or role Output format (strict): Score: [single integer from 1 to 7] Figure 11: CoT-based direct judging prompt with the 7-point rubric. Vanilla+CoTPRISM 0 20 40 60 80 100 G-Acc =15.1 =16.3 =3.4 (a) Social-Persona Vanilla+CoTPRISM 0 20 40 60 80 100 =11.9 =7.4 =4.8 (b) Big5-Persona-EASY Vanilla+CoTPRISM 0 20 40 60 80 100 =13.3 =4.7 =5.5 (c) Big5-Persona-HARD QwenLlamaMistralMean ± Std Figure 12: Backbone sensitivity on strict Group Accuracy across three LLM backbones. Points denote individual backbones and diamonds indicate the mean with standard deviation. 20406080 Qwen Qwen+CoT Llama Llama+CoT Mistral Mistral+CoT DeepSeek-V3.2 DeepSeek-V3.2+CoT GPT-5.4 GPT-5.4+CoT Gemini-3-Flash Gemini-3-Flash+CoT =6.9 =8.5 =28.5 =6.8 =0.7 =9.2 =5.3 =10.6 =6.0 =3.4 =1.4 =6.2 (a) Social-Persona 406080100 =1.4 =5.7 =14.8 =4.1 =0.6 =1.2 =2.8 =3.0 =1.0 =1.0 =1.1 =5.1 (b) Big5-Persona-EASY 0204060 =3.3 =9.7 =22.8 =3.5 =0.0 =1.0 =2.6 =4.8 =7.4 =10.2 =5.5 =12.1 (c) Big5-Persona-HARD Vanilla+CoT5-point rubric7-point rubric Figure 13: Rubric Sensitivity of direct LLM-as-a-Judge evaluation on Strict Group Accuracy. Each segment connects the results from 5-point and 7-point rubrics for the same evaluator and prompting method. Longer segments indicate greater sensitivity to scoring granularity. 0.00 0.25 0.50 0.75 1.00 Score PositiveNegative VanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMDeepSeekGPTGemini LlamaQwenMistralClosed-source Figure 14: Score distribution on the Big5-Persona-EASY dataset 0.00 0.25 0.50 0.75 1.00 Score PositiveNegative VanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMVanillaTaskStanceStylePRISMDeepSeekGPTGemini LlamaQwenMistralClosed-source Figure 15: Score distribution on the Social-Persona dataset 70 75 80 P-AUC (%) Big5-HARD 84 86 88 90 SocialBench Greedy0.20.50.81.0 30 40 50 60 70 G-Acc (%) Greedy0.20.50.81.0 50 60 70 80 Vanilla+CoTPRISM Figure 16: Temperature sensitivity of holistic evaluation for the Qwen backbone acrossBig5-Persona-HARD andSocial-Persona. Greedy denotes deterministic decoding, whereas the other points show stochastic de- coding under temperatures 0.2, 0.5, 0.8 and 1.0. Error bars indicate mean and standard deviation over three random seeds. The dashed line shows the corresponding PRISM reference score. 55 60 65 70 P-AUC (%) Big5-HARD 80.0 82.5 85.0 87.5 SocialBench Greedy0.20.50.81.0 20 30 40 50 G-Acc (%) Greedy0.20.50.81.0 50 55 60 65 70 Vanilla+CoTPRISM Figure 17: Temperature sensitivity of holistic evaluation for the Mistral backbone acrossBig5-Persona-HARD andSocial-Persona. Greedy denotes deterministic decoding, whereas the other points show stochastic de- coding under temperatures 0.2, 0.5, 0.8 and 1.0. Error bars indicate mean and standard deviation over three random seeds. The dashed line shows the corresponding PRISM reference score. 60 65 70 75 P-AUC (%) Big5-HARD 70 75 80 85 90 SocialBench Greedy0.20.50.81.0 20 40 60 G-Acc (%) Greedy0.20.50.81.0 20 40 60 Vanilla+CoTPRISM Figure 18: Temperature sensitivity of holistic evaluation for the Llama backbone acrossBig5-Persona-HARD andSocial-Persona. Greedy denotes deterministic decoding, whereas the other points show stochastic de- coding under temperatures 0.2, 0.5, 0.8 and 1.0. Error bars indicate mean and standard deviation over three random seeds. The dashed line shows the corresponding PRISM reference score.