Paper deep dive
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
Riccardo Rosati, Edoardo Colucci, Massimiliano Bolognini, Adriano Mancini, Paolo Sernani
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/14/2026, 2:37:19 AM
Summary
RPA-Check is a multi-stage automated evaluation framework designed to assess the performance of LLM-based Role-Playing Agents (RPAs) in complex, constraint-heavy environments. The framework utilizes a four-step pipelineโDimension Definition, Augmentation, Semantic Filtering, and LLM-as-a-Judge Evaluationโto transform qualitative behavioral criteria into granular Boolean indicators. Validated through 'LLM Court', a serious game for forensic training, the study demonstrates that smaller, instruction-tuned models (8-9B) can achieve higher procedural consistency and role fidelity than larger models, which are often prone to sycophancy or alignment bias.
Entities (5)
Relation Signals (3)
RPA-Check โ evaluates โ Role-Playing Agents
confidence 100% ยท RPA-Check, a multi-stage automated evaluation framework designed to objectively assess the performance of LLM-based RPAs
LLM Court โ validates โ RPA-Check
confidence 100% ยท We validate this framework by applying it to LLM Court
RPA-Check โ adapts โ CheckEval
confidence 95% ยท The proposed methodology adapts the CheckEval paradigm
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid adoption of Large Language Models (LLMs) in interactive systems has enabled the creation of dynamic, open-ended Role-Playing Agents (RPAs). However, evaluating these agents remains a significant challenge, as standard NLP metrics fail to capture the nuances of role adherence, logical consistency, and long-term narrative stability. This paper introduces RPA-Check, a multi-stage automated evaluation framework designed to objectively assess the performance of LLM-based RPAs in complex, constraints-heavy environments. Our methodology is based on a four-step pipeline: (1) Dimension Definition, establishing high-level qualitative behavioral criteria; (2) Augmentation, where these requirements are expanded into granular boolean checklist indicators; (3) Semantic Filtering, to ensure indicator objectivity, no redundancy and agent isolation; and (4) LLM-as-a-Judge Evaluation, which employs chain-of-thought verification to score agent fidelity. We validate this framework by applying it to LLM Court, a serious game for forensic training involving several quantized local models. Experimental results across five distinct legal scenarios demonstrate the framework's ability to identify subtle trade-offs between model size, reasoning depth, and operational stability. Notably, the findings reveal an inverse relationship between parametric scale and procedural consistency, showing that smaller, adequately instruction-tuned models (8-9B) can outperform larger architectures prone to user-alignment bias or sycophancy. RPA-Check thus provides a standardized and reproducible metric for future research in generative agent evaluation within specialized domains.
Tags
Links
- Source: https://arxiv.org/abs/2604.11655v1
- Canonical: https://arxiv.org/abs/2604.11655v1
Trouble viewing inline? Open PDF directly โ
Full Text
95,237 characters extracted from source content.
Expand or collapse full text
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Riccardo Rosati a,โ , Edoardo Colucci b , Massimiliano Bolognini b , Adriano Mancini b and Paolo Sernani c a Department of Political Sciences, Communication and International Relations, University of Macerata, Via Don Minzoni 22/A, Macerata, 62100, Italy b Deparment of Information Engineering, Universitร Politecnica delle Marche, Via Brecce Bianche 12, Ancona, 60131, Italy c Department of Law, University of Macerata, Piaggia dellโUniversitร 2, Macerata, 62100, Italy A R T I C L E I N F O Keywords: Role-Playing Agents Large Language Models Automated Evaluation Serious Games Generative AI A B S T R A C T The rapid adoption of Large Language Models (LLMs) in interactive systems has enabled the creation of dynamic, open-ended Role-Playing Agents (RPAs). However, evaluating these agents remains a significant challenge, as standard NLP metrics fail to capture the nuances of role adherence, logical consistency, and long-term narrative stability. This paper introduces RPA- Check, a multi-stage automated evaluation framework designed to objectively assess the per- formance of LLM-based RPAs in complex, constraints-heavy environments. Our methodology is based on a four-step pipeline: (1) Dimension Definition, establishing high-level qualitative behavioral criteria; (2) Augmentation, where these requirements are expanded into granular boolean checklist indicators; (3) Semantic Filtering, to ensure indicator objectivity, no redun- dancy and agent isolation; and (4) LLM-as-a-Judge Evaluation, which employs chain-of-thought verification to score agent fidelity. We validate this framework by applying it to LLM Court, a serious game for forensic training involving several quantized local models. Experimental results across five distinct legal scenarios demonstrate the frameworkโs ability to identify subtle trade-offs between model size, reasoning depth, and operational stability. Notably, the findings reveal an inverse relationship between parametric scale and procedural consistency, showing that smaller, adequately instruction-tuned models (8-9B) can outperform larger architectures prone to user-alignment bias or sycophancy. RPA-Check thus provides a standardized and reproducible metric for future research in generative agent evaluation within specialized domains. 1. Introduction The increasing adoption of Large Language Models (LLMs) in interactive systems has enabled the creation of dynamic and open Role-Playing Agents (RPAs) capable of sustaining natural language dialogues without predefined scripts (Chen et al., 2024; Wang et al., 2026). However, evaluating these agents remains a significant challenge: standard Natural Language Processing metrics, such as BLEU, ROUGE or perplexity, do not capture critical dimensions such as adherence to the assigned role, internal logical consistency, and narrative stability over extended conversations. This methodological gap is particularly evident in highly specialized contexts, where correctness does not coincide with linguistic fluency and general factuality, but depends on compliance with procedural and semantic constraints specific to the domain, and players might experience unnatural conversational flow (Zargham et al., 2026). โ Corresponding author riccardo.rosati@unimc.it (R. Rosati); S1110494@studenti.univpm.it (E. Colucci); S1111836@studenti.univpm.it (M. Bolognini); a.mancini@staff.univpm.it (A. Mancini); paolo.sernani@unimc.it (P. Sernani) ORCID(s): 0000-0003-3288-638X (R. Rosati); 0000-0001-5281-9200 (A. Mancini); 0000-0001-7614-7154 (P. Sernani) R. Rosati et al.: Preprint submitted to ElsevierPage 1 of 34 arXiv:2604.11655v1 [cs.CL] 13 Apr 2026 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents In serious games, where one of the goals is to convey professional skills through interactive simulations, the integration of LLM opens opportunities for adaptive and immersive training environments (Bonetti et al., 2025; รzkaya et al., 2025). However, the lack of objective and reproducible evaluation frameworks limits the ability to compare models and identify trade-offs between different performance dimensions. Existing methodologies, such as the MMLU benchmark (Hendrycks et al., 2021) or Chatbot Arena (Chiang et al., 2024), focus on general knowledge or conversational utility, making them inadequate for agents who must embody rigidly defined roles, such as an impartial judge or a hostile witness in a mock trial. This paper introduces RPA-Check, a multi-stage automated evaluation framework designed to systematically and objectively measure the performance of LLM-based RPAs in constrained environments with high procedural complexity. The proposed methodology adapts the CheckEval paradigm (Pereira et al., 2024) to role-playing contexts through a four-stage pipeline: i) Dimension Definition establishes high-level qualitative criteria (such as behavioral role fidelity); i) Augmentation, in which high-level requirements are expanded into granular Boolean indicators; i) Filtering, which ensures the objectivity, non-redundancy, and isolation of agent-specific performance metrics; and iv) LLM-as-a-Judge (Gu et al., 2026) Evaluation, where an evaluator model applies chain-of-thought (Wei et al., 2022) checks to quantify the performance of game agents. The proposed RPA-Check framework is tested on โLLM Courtโ (also introduced in this paper), a game simulating a court using local quantized models as the LLMs implementing the Non-Playing Characters (NPCs) and developed in Unity. The architecture allows for the execution of trial simulations in which LLM-controlled agents embody legal roles (judge, prosecutor, witnesses) and interact with the player through open, unscripted dialogues. The simulation environment imposes strict procedural constraints, requiring agents to maintain narrative consistency, respect intervention hierarchies, and provide logically consistent responses over extended dialogue sessions. Moreover, the adoption of quantized models allows for completely local execution on consumer hardware (Lang et al., 2024), eliminating dependence on cloud services, ensuring the privacy of sensitive data, and significantly reducing the computational costs and environmental impact associated with inference (Bai et al., 2024). This architectural choice enables sustainable and accessible deployment, which is particularly relevant for educational applications intended for contexts with limited resources or confidentiality requirements. The experimental results, conducted on seven open-weight models ranging in size from 8 to 14 billion parameters and evaluated across five distinct legal scenarios, highlight the frameworkโs ability to quantify trade-offs between role fidelity, operational stability, and reasoning depth. As such, the proposed framework can quantify abstract qualities like โRole Adherenceโ and โLogical Consistency,โ distinguishing between models that merely generate fluent text and those that truly embody their assigned roles. In particular, in LLM court, it emerges that smaller but adequately instruction- tuned models can outperform larger architectures in terms of behavioral consistency when measured using a checklist R. Rosati et al.: Preprint submitted to ElsevierPage 2 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents of granular indicators. These results provide useful empirical evidence to guide model selection in similar application contexts, while highlighting limitations related to the sensitivity of models to the specific case and generation seed. Specifically, this workโs contributions to the state of the art are: โข RPA-Check, a rigorous and adaptable methodological framework that transforms qualitative role descriptions into quantitative Boolean metrics through an automated pipeline of augmentation, filtering, and evaluation using LLM-as-a-Judge. โข LLM Court, a game designed to validate the framework in a rule-based environment that simulates forensic interactions in court under strict procedural constraints. The game introduces a clear separation between the generative logic to produce contents and procedural control, i.e., the courtroom rules (unlike the open-ended conversational agents used in purely recreational contexts). โข A modular reference architecture, the one developed for LLM court, that integrates local quantized models into a Unity environment for running process simulations, providing a reproducible testbed for evaluating RPA in constrained contexts. โข An empirical comparative analysis of open-weight models that quantifies the trade-off between model size, role fidelity, and operational stability, demonstrating that instruction tuning and evaluation granularity are critical factors for selecting appropriate architectures. The remainder of the paper is organized as follows. Section 2 examines the state of the art in the use of LLMs in recreational and educational contexts, methodologies to evaluate dialogic agents, and serious games based on generative artificial intelligence, highlighting the contributions of the proposed work in a comparative manner. Section 4 describes the system architecture of LLM Court, detailing the technological components and design choices. Section 3 illustrates the proposed multi-stage automated evaluation framework, from the definition of evaluation dimensions to the generation of checklists. Section 5 presents the experimental results with a comparative analysis of the tested models, while Section 6 discusses the main findings and limitations of RPA-Check and the performed experiments. Finally, Section 7 concludes the paper, presenting future directions for research on the evaluation and application of LLM-based RPAs in specialized contexts. 2. Related works The integration of LLMs into gaming and simulation contexts is a rapidly evolving area of research characterized by multiple applications, ranging from the use of Generative AI as a development tool to its direct implementation in gameplay (Lanzi and Loiacono, 2023; Sweetser, 2024). For example, recent works such as PLAYER (Zhu et al., R. Rosati et al.: Preprint submitted to ElsevierPage 3 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents 2025) explored multi-agent communication in murder mystery games, focusing on strategic reasoning and information hiding. Similarly, PsyPlay (Yang et al., 2025) introduced personality-infused role-playing for LLM agents, ensuring that these agents adhere to the so-called Big Five personality traits. The existing literature shows a growing interest in the use of LLMs to generate dynamic content (Farrokhi Maleki and Zhao, 2024), enable more natural interactions with NPCs (Cox and Ooi, 2024), and support immersive educational experiences (Song et al., 2024). However, most approaches focus on game agents, content generation or dialogic interaction as an end in itself (Zhu et al., 2023; Ashby et al., 2023), neglecting the need for rigorous methodologies to evaluate the quality and consistency of agentive behaviors in highly constrained and regulated scenarios (Wu et al., 2025; Sweetser, 2024), such as the one chosen for this study i.e., simulated forensic settings. To better highlight this multifaceted landscape and how the proposed contribution advances the state of the art, this section is organized into three parts. First, the use of Generative AI as a tool for content production and in-game integration, being the main application of LLM in video games, is examined (Section 2.1). Second, applications in educational and forensic simulation contexts, the validation domain of this work, are analyzed (Section 2.2). Finally, existing methodologies for evaluating LLM-based agents and dialogic systems are presented (Section 2.3). 2.1. Generative AI in Games In the domain of content generation, numerous studies have explored how LLMs can automate the creation of narrative assets, game levels, and characters. Alavi et al. (2024) demonstrate the usefulness of LLMs as assistants in game plot design. Similarly, Kumaran et al. (2023) propose automating the generation of interactive narrative scenes using LLMs. While innovative, these approaches remain focused on the production phase and do not address the challenge of systematically evaluating the narrative quality or behavioral consistency of the generated agents. In contrast, the framework proposed in this research is not limited to generating content, but introduces a multi-stage evaluation pipeline that transforms qualitative role requirements into verifiable Boolean metrics, ensuring objective control over the fidelity of agent behaviors. Gursesli et al. (2023) use ChatGPT to generate visual novel narratives on climate themes. They actually propose to use ChatGPT as an automatic evaluator for the story quality by combining different prompts, but without systematically structuring granular questions for the evaluation criteria as in the approach proposed in this paper. Games such as AI Dungeon (Hua and Raley, 2020) and Infinite Craft (Agarwal, 2024) integrate LLMs directly into the game loop, enabling procedural generation of content and responses in real time. Trukhin et al. (2026) explore the integration of LLM-powered NPCs in The Elder Scrolls V: Skyrim, finding an increase in perceived immersion among players. However, these implementations present limited control over long-term narrative coherence and the absence of clearly defined procedural constraints reduce their suitability in contexts where strict compliance with specific rules R. Rosati et al.: Preprint submitted to ElsevierPage 4 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents is essential. This work, on the other hand, introduces a forensic simulation environment in which agents must comply with explicit legal and procedural constraints, and proposes an automated evaluation mechanism capable of verifying such adherence through granular checklists generated and filtered algorithmically. 2.2. Educational and forensic simulations In the educational context, serious games have integrated Generative AI to improve teaching effectiveness through personalization and content adaptation. Zhao et al. (2024) present Language โUrban Odysseyโ, a serious game for language learning that uses LLM to populate interactive environments with NPCs who speak a foreign language. Tรณdovรก (2025) proposes โA Quest for Informationโ, demonstrating that LLM-generated dialogues can increase engagement during educational events. However, these works do not address the issue of validating agent behaviors with respect to specific educational objectives, nor do they provide replicable methodologies for evaluating the pedagogical quality of the interaction. The framework proposed here fills this gap by introducing explicit evaluation dimensions (such as โRole Adherenceโ and โLogical Consistencyโ) and adopting an LLM-as-a-Judge approach to objectively quantify agentsโ adherence to assigned roles, enabling systematic comparative analysis between different models. The forensic and investigative domain also presents a challenging terrain for LLM applications, given that it might require the management of dialogues constrained by specific (and coded) procedural rules and the ability to maintain logical consistency over extended sessions. Games such as โVerbal Verdictโ (Savanna Developments, 2024) integrate LLMs to enable open-ended interrogations, but might suffer from hallucination and narrative inconsistencies. Beattie and Colbran (2017) discussed the potential of legal simulations games as a support for subjects not represented in courts, highlighting the need for procedural fidelity. Despite these efforts, to the best of our knowledge, the existing literature does not offer standardized methodologies for measuring the quality of forensic simulations or tools for objectively comparing the performance of different models in such simulation scenarios. This work tries to fill this gap by proposing an automated evaluation pipeline that adapts the CheckEval paradigm generating granual Boolean indicator checklists and applying chain-of-thoughts techniques to ensure reproducible verifications. The forensic and legal domain, with โLLM Courtโ, is the validation domain for our evaluation framework. 2.3. Evaluation of LLM-based Agents in Video Games and Simulations The scientific literature is rich in contributions to the evaluation of LLMs in general. SummEval (Fabbri et al., 2021) is a well-established framework for evaluating automatic summarization systems, based on various qualitative dimensions such as coherence, consistency, fluency, and relevance, measured through automatic metrics and human judgments. TopicalChat (Gopalakrishnan et al., 2023), on the other hand, is a dataset and framework designed to evaluate knowledge-oriented open-domain dialogue systems, focusing on their ability to maintain natural and R. Rosati et al.: Preprint submitted to ElsevierPage 5 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents informative conversations on a variety of topics. CheckEval (Pereira et al., 2024) introduces the LLM-as-a-Judge paradigm, transforming qualitative evaluation dimensions into granular Boolean question checklists through an auto- mated augmentation and filtering process, ensuring greater consistency and traceability compared to subjective human evaluations. However, although these approaches represent benchmarks in general, they were not designed to measure procedural fidelity and logical consistency in highly specialized contexts bound by explicit rules. SummEval and TopicalChat operate on general dimensions (textual coherence, informational relevance, conversational naturalness) which, while fundamental for assessing linguistic quality, do not capture critical aspects such as adherence to specific professional roles, compliance with strict procedural contraints, or the ability to maintain narrative coherence in structured multi-turn dialogues. Similarly CheckEval, while providing a robust methodological mechanism, has been validated primarily on summarization tasks and not on complex role-playing scenarios. RPA-Check framework extends the CheckEval paradigm by tailoring a four-stage pipeline (Dimension Def- inition, Augmentation, Filtering, LLM-as-a-Judge Evaluation) to Role-Playing Agents. The proposed framework derives domain-specific evaluation dimensions (โRole Adherenceโ, โArgumentative Depthโ, โFactual Consistencyโ, โContextual Relevanceโ) by using, for validation, forensic and legal domain requirements. The framework allows generating checklists of Boolean indicators that query granular aspects such as the ability of agents to avoid narrative contradictions, respect procedural hierarchies (e.g., the order of speech between judge, prosecution, and defense), and maintain fidelity to the assigned role over prolonged dialogue sessions. Furthermore, while CheckEval has been applied mainly to single text generation tasks, this work adapts it to a dynamic multi-agent context, where evaluation must take into account the interactions between different agents and the overall consistency of the dialogic system. Some recent benchmarks have expanded LLM-agent evaluation in gaming contexts. For example, TEXTQUESTS (Phan et al., 2025) assesses long-horizon reasoning in 25 classic Infocom interactive fiction games requiring sustained problem-solving (such as โZorkโ) over 100K+ token contexts without external scaffolding. FAIRGAMER (Shi et al., 2026) instead evaluates social biases in LLM-based NPCs through game-theoretic interaction patterns (transaction, cooperation, competition) across 16,910 bilingual test cases. While these works address long-context reasoning and fairness respectively, neither tackles the core challenge of constraint adherence in multi-agent role-playing, which is, on the contrary, the focus of LLM Court. TEXTQUESTS evaluates single-agent exploratory scenarios, whereas our forensic simulation requires simultaneous satisfaction of role-specific behavioral norms, logical consistency, and procedural compliance across multi-turn interactions. FAIRGAMERโs fairness metrics are orthogonal to our evaluation objectives: LLM Court prioritizes behavioral fidelity to domain-specific constraints rather than equitable decision- making. Similarly to TEXTQUESTS and FAIRGAMER, ORAK (Park et al., 2025) proposes a large-scale benchmark spanning 12 commercial video games across multiple genres, supporting agentic modules such as reflection and planning and evaluating models through gameplay performance metrics and leaderboards. While ORAK emphasizes R. Rosati et al.: Preprint submitted to ElsevierPage 6 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents cross-genre generalization and strategic competence in dynamic environments, its evaluation remains performance- driven, with success defined by score maximization or task completion. In contrast, LLM Court evaluates agents in a tightly regulated forensic setting where correctness is defined by sustained compliance with explicit procedural and role-specific constraints. Thus, whereas ORAK measures gameplay effectiveness, LLM Court measures institutional fidelity under structured normative requirements. In summary, although the existing literature has explored the integration of LLMs in video games and serious games from multiple perspectives, significant gaps remain regarding the systematic evaluation of agentsโ fidelity to complex and constrained roles. The RPA-Check framework proposed in this work fills this gap by providing a rigorous, automated, and replicable methodology for evaluating RPAs in forensic contexts, empirically demonstrating how small, appropriately evaluated models can outperform larger models in terms of narrative consistency and stability. 3. Methodology: The Multi-Stage Automated Evaluation Framework Evaluating RPAs in constraints-heavy environments presents a unique challenge: standard NLP metrics (e.g., BLEU, ROUGE, Perplexity) fail to capture high-level semantic requirements like role adherence, narrative stability, or procedural compliance. A model might generate linguistically fluent text that is nonetheless procedurally invalid or logically incoherent. To address these limitations, we propose RPA-Check, a multi-stage automated evaluation framework that transforms qualitative behavioral requirements into granular, verifiable boolean indicators via an โLLM-as-a-Judgeโ approach for the final evaluation. 3.1. RPA-Check Framework Architecture The framework is structured as a four-stage pipeline designed to minimize evaluator subjectivity and maximize the resolution of the performance metrics (Figure 1). Stage 1: Dimension Definition The first stage of the framework involves the formal definition of a set ๎ฐ of ํพ high- level evaluation dimensions, denoted as ๎ฐ = ํ 1 ,ํ 2 , ...,ํ ํพ , where ํพ represents the total number of macro-criteria selected for the benchmark. Each dimension ํ ํ โ ๎ฐ encapsulates a distinct qualitative behavior attribute of the agents, serving as the basis for the subsequent automated generation of granular indicators. To ensure a comprehensive assessment of behavioral fidelity, system stability and quality, for our specific implementation, we define ํพ = 3 primary macro-dimensions: โข Behavioral Role Fidelity (ํ ํตํ ํน ): this dimension serves as the primary metric for validating agent personas. It evaluates the degree to which an agent maintains functional boundaries and avoids sycophancy. It is treated as a composite vector comprising four core sub-dimensions: R. Rosati et al.: Preprint submitted to ElsevierPage 7 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Step 1: DIMENSIONS DEFINITIONStep 2: AUGMENTATIONStep 3: SEMANTIC FILTERINGStep 4: EVALUATION RPA-CheckFramework Architecture BehavioralRoleFidelity (ํ BRF ) ProceduralConvergence & Stability(ํ PCS ) LinguisticFormalism & Orthography(ํ LFO ) โชGame end โชLogicaloutcome โชNo retryloops / freezes โชGrammaticalprecision โชDomain-specificlexicon โชSpecificpersona register โชRoleAdherence โชArgumentativeDepth โชFactualConsistency โชContextualRelevance Question Diversification Question Elaboration LLM โณ โโฏํ Alternative perspective queries Hierarchicaldecomposition (5-10 sub-questions per criterion) LLM โณ ํํขํํํ ํ=ํ 1 ,ํ 2 ,...,ํ ํ โํโํ,โณ โโฏํ โํฌ ํน =ํ 1 ,ํ 2 ,...,ํ ํ RawBooleanQuestionsํฌ ํน ํ j : yes,no Determinism Non-redundancy Agent Isolation Evaluation Metrics ํํํ,ํ= 1 ํฌ โ เท ํโํฌ โ ํJudgeํ,ํ=Yes Quality Score (QS) FilteredChecklist ํฌ โ Game Transcriptํ Filtered Checklist ํฌ โ Chain-of-Thought (CoT) constraint Step 1: Reasoning (justificationj) Step 2: Decision (yโ0,1) โํโํฌ โ ,โณ ํํขํํํ (ํ,ํ)โ(ํฆ,ํ) ํ ํ= เท ํ ํํกํํฆํฟํํํ RetryRate (R) LLM โณ ํํํํก โํโํฌ ํน ,ํฌ โ =โณ ํํํํก (ํฌ ํน ) Figure 1: Architecture of the RPA-Check Framework. The automated evaluation pipeline transforms qualitative behavioral requirements of Role-Playing agents into quantitative metrics through a four-stage process: Dimensions Definition establishes behavioral, procedural, and linguistic criteria; Augmentation expands these dimensions into granular Boolean checklists via an LLM generator; Semantic Filtering ensures checklist objectivity and agent isolation; and Evaluation based on LLM-as-a-Judge approach scores trial transcripts using a Chain-of-Thought (CoT) reasoning protocol. Final performance is quantified via the Quality Score (QS) and Retry Rate (R) metrics to assess role fidelity and operational stability. โ Role Adherence: the agent must operate strictly within its assigned role ํ ํํํํ while suppressing capabilities of other roles. โ Argumentative Depth: it measures the semantic complexity of responses, awarding high scores for multi- step logical arguments and referencing specific evidence. โ Factual Consistency: it verifies that assertions do not contradict previous turns or the immutable ground truth established in the case file. โ Contextual Relevance: it assesses the causal link between consecutive dialogue turns, ensuring the response directly addresses the query or stimulus provided in the previous turn. It specifically penalizes "non- sequiturs" where the agent ignores a direct question. โข Procedural Convergence and Stability (ํ ํํถํ ): this dimension assesses the ability of the system to drive the narrative toward a logical and successful conclusion. For instance, in the context of the LLM Court framework applied in this study, it measures the adherence to the deterministic Finite State Machine (FSM) governing the simulation. A trial is considered procedurally stable when: i) the interaction successfully reaches the end of the game; i) the final outcome is semantically and logically aligned with the evidence accumulated during the dialogue history; i) the simulation completes without requiring manual โRetryโ loops or encountering catastrophic model freezes. R. Rosati et al.: Preprint submitted to ElsevierPage 8 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents โข Linguistic Formalism and Orthography (ํ ํฟํนํ ): linguistic fidelity is critical for maintaining the immersion and professional register of the RPA required in specialized domains. This dimension ensures that the generated text adheres to grammatical precision, domain-specific lexicon and specific persona register. Stage 2: Augmentation A generator model ๎น ํํํ is employed to enrich each dimension ํ ํ into a set of granular Boolean questions ๎ฝ ํ . This stage utilizes two distinct prompting strategies: 1. Question Diversification: generates alternative queries to analyze the same concept from multiple perspectives. 2. Question Elaboration: decomposes main criteria into a hierarchy of sub-questions (5 to 10 per criterion) to achieve a fine-grained analysis. Formally, for any dimension ํ ํ โ ๎ฐ: ํ ๎น ํํํ โโ ๎ฝ ํ = ํ 1 ,ํ 2 , ...,ํ ํ (1) where each ํ ํ must be answerable with a binary ํฆํํ ,ํํ. This step increases the resolution of the evaluation, minimizing the variance typically associated with single-score Likert scales. Stage 3: Semantic Filtering The raw set ๎ฝ ํ undergoes a filtering process to ensure objectivity and relevance. This stage is performed by a specialized filtering model ๎น ํํํํก , which is tasked with removing redundant, subjective, or non-pertinent indicators to produce the final checklist ๎ฝ โ . Formally, the filtering operation is defined as: ๎ฝ โ = ๎น ํํํํก (๎ฝ ํ )(2) where ๎ฝ โ denotes the subset of indicators for all questions dimension ๎ฝ ํ that satisfy the following operational constraints: โข Binary determinism: each question must be answerable with a strictly binary yes,no value based exclusively on the interaction transcript T. โข Non-redundancy: overlapping or duplicate questions are identified and removed to ensure each indicator provides a unique informative contribution. โข Agent isolation: questions concerning human-controlled participants are discarded to isolate the performance of the generative models. Stage 4: LLM-as-a-Judge Evaluation The final evaluation is performed by a high-parameter Judge Model (๎น ํํขํํํ ). For a transcript ํ and the filtered checklist ๎ฝ โ , the judge generates a binary decision ํฆ โ 0, 1 and a R. Rosati et al.: Preprint submitted to ElsevierPage 9 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents justification ํ for each ํ โ ๎ฝ โ . We enforce a Chain-of-Thought (CoT) constraint to ground the evaluation: โํ โ ๎ฝ โ , ๎น ํํขํํํ (ํ,ํ) โ (ํฆ,ํ)(3) 3.2. Evaluation Metrics To quantitatively assess the performance of the RPAs, we define two primary metrics that capture both the qualitative performance and the operational reliability of the models: โข Quality Score (QS): this is the fundamental metric used to evaluate all three macro-dimensions ํ ํตํ ํน , ํ ํํถํ , and ํ ํฟํนํ . The ํํ represents the percentage of affirmative responses provided by the Judge model across the filtered boolean checklists. For a given model ํ and transcript ํ , it is formally defined as: ํํ(ํ,ํ ) = 1 |๎ฝ โ | โ ํโ๎ฝ โ ํ(Judge(ํ,ํ) = Yes)(4) where ํ(โ ) is the indicator function. By applying ํํ to all dimensions, we ensure a standardized measurement of qualitative attributes, ranging from persona adherence to the logical consistency of the final verdict. โข Retry Rate (ํ ): this metric quantifies the stochastic instability and operational failures of the models. ํ is defined as the total number of manual restarts required due to catastrophic failures during a simulation, such as infinite loops, model freezing, or the inability to correctly tag the next speaker. While ํํ inherently measures even the quality of successful interactions, ํ serves as the negative counterpart to stability: a higher ํ indicates that a model is prone to narrative degradation regardless of its linguistic fluency. 4. The LLM Court Case Study To validate our framework, we implemented LLM Court, a platform designed to test RPAs in a rules-based environ- ment conceived to simulate forensic dialogue under strict procedural constraints. Unlike open-ended conversational agents utilized in purely recreational contexts, our system enforces a rigid separation between the generative logic (content production) and the procedural control (courtroom rules). This section describes the system architecture used to generate the interaction data for our evaluation. 4.1. Architectural Overview The architecture is built upon a local-inference engine integrated within the Unity environment, utilizing quantized LLMs to drive the NPCs. This architectural choice ensures operational independence and data privacy, a critical R. Rosati et al.: Preprint submitted to ElsevierPage 10 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents requirement for potential educational or professional applications in legal training. The game core loop consists of three distinct computational modules: โข Procedural Case Generation: it instantiates the semantic boundaries of the simulation via remote high-parameter models. โข Agentic State Machine: a deterministic finite-state automation governing turn-taking dynamics between the Player (Defense) and RPAs (Judge, Prosecutor, Witness). The dynamic role-playing of the agents is regulated via system prompt. โข Intent Analysis Module: a parallel processing unit responsible for extracting structured metadata from unstruc- tured natural language outputs. Figure 2 shows some screens from LLM Court, such as the main game menu where the player starts a new game or loads a previous one, sets the gameโs options and chooses their avatar (2a), and the user interface for Procedural Cage Generation (2b). During the simulation, the game is set in a courtroom, where the judge RPA (2c introduces the case and the player (2d) interacts with the LLM Court RPAs via text-based input. Specifically, the playerโs flow in LLM Court is schematized in Figure 3: after launching the game and its main interface, starting a new game loads the RPAs LLM. Following the player selection of a โnew caseโ, the Procedural Case Generation is performed through GPT-5, showing the generated text. If the player saves the case, the game actually starts otherwise a new case is generated. 4.2. Procedural Case Generation via Structured Sampling To ensure narrative variety and logical consistency, the initialization phase employs a stochastic generation process constrained by a predefined schema. While run-time interaction relies on quantized local models for efficiency, we utilize GPT-5 via Azure OpenAI Service exclusively for this pre-computation phase to maximize coherence in the case facts. Formally, a legal case ๎ฏ is defined as a tuple ๎ฏ = ํ,ํ,๎ฑ,๎,ํบ, where ํ is the title, ํ is the factual summary, ๎ฑ is the set of evidentiary items, ๎ is the set of witness profiles, and ํบ denotes the defense goal. The generation process minimizes the probability of hallucination by strictly enforcing a JSON output format. Let ํ ํ be the language model parameterized by ํ. The generation of ๎ฏ is conditional on the schema constraints ๎ต ํ ํโํํํ and user-defined priors ๎ผ ํขํ ํํ (e.g., "theft case"): ๎ฏ โผ ํ ํ (๎ฏ โฃ ๎ต ํ ํโํํํ ,๎ผ ํขํ ํํ )(5) R. Rosati et al.: Preprint submitted to ElsevierPage 11 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents (a)(b) (c)(d) Figure 2: Screens from LLM Court: (a) the main game menu and (b) the case generation module. During the simulation, the game transitions to the courtroom, where the judge RPA (c) introduces the case, and the player (d) interacts with the RPAs through text-based input. click new game Loading of LLM architecture Show preview of generated case Launching the game and the main interface Random data generation using large language models (API interface) Do you like the case? Y N Start of the game click new caseclick playclick save Phi Figure 3: The playerโs flow in LLM Court. After launching the game interface, starting a new game loads the RPAs LLM. Upon selecting a โnew case,โ the Procedural Case Generation module powered by GPT-5 generates and displays a case description. The player may either save the generated case to begin the game or discard it and request a new one. This module deserializes the output into run-time objects, populating the โCourt Record,โ which serves as the immutable ground-truth knowledge base. This explicitly prevents the dynamic agents from hallucinating facts that contradict the static case file during the trial phase. 4.3. Dynamic Role-Playing via System Prompt Injection The core interaction is driven by a dynamic prompting strategy that adapts the modelโs persona to the current state of the trial. We implement a context injection mechanism that prepends a composite System Prompt ๎ฟ ํํํํ to the dialogue history ํป ํก at turn ํก. To ensure complete session independence and introduce variability, given that such variability can R. Rosati et al.: Preprint submitted to ElsevierPage 12 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents be obtained with a small change in input (Salinas and Morstatter, 2024; Sclar et al., 2024), each simulation defines a generation seed included in the agentsโ initial prompt (i.e., an integer parameter). Varying the seed across repeated runs of an identical scenario produces distinct dialogue trajectories while holding all other experimental variables constant. The System Prompt is engineered to mitigate two pervasive issues identified in LLM-based role-playing literature: i) Recency Bias (Instrunction Drift), the tendency of autoregressive models to prioritize recent dialogue tokens over initial behavioral instructions; i) User-Alignment Bias (Sycophancy), the inherent tendency of instruction-tuned models to be cooperative towards the user, which is detrimental in an adversarial forensic setting. To address these, we define distinct prompt architectures for the key roles: 4.3.1. The Judge (Controller Agent) The Judge acts as a deterministic regulator rather than a conversationalist. The prompt enforces a โChain-of- Thoughtโ (CoT) logic to evaluate the admissibility of arguments before passing the turn. The Judgeโs policy ํ ํํขํํํ is defined not to maximize engagement, but to minimize procedural entropy: ํ ํํขํํํ (ํ ํก โฃ ํป ํก ) โ Grant Intervention, Overrule, Sustain, Verdict(6) 4.3.2. The Prosecutor (Adversarial Agent) The Prosecutor agent is engineered to counteract the sycophancy alignment of standard base models. We formalize the Prosecutorโs objective as prioritizing adversarial argumentation over user alignment. Conceptually, this agent optimizes a policyํ ํํํํ that maximizes the logical strength of the accusationํฟ ํํ while penalizing semantic alignment with the userโs intent ํผ ํขํ ํํ . The optimal response ํฅ โ ํก is derived as follows: ํ ํํํํ (ํป ํก ) = ํฅ โ ํก = argmax ํฅโ๎ โก โข โข โข โข โฃ ํ ํ (ํฅ โฃ ํป ํก ,๎ฟ ํํํํ ) โโโโโ ํฟ ํํ (ํฅ) โํโ ๎ญ(ํฅ,ํผ ํขํ ํํ ) โโโโโ Penalty โค โฅ โฅ โฅ โฅ โฆ (7) Where ๎ฟ ํํํํ represents the prompt instructions explicitly forbidding neutral or helpful behaviors (e.g., "Do not help the Defense", โFocus on the weakness of the userโs argumentโ), ๎ญ(ํฅ,ํผ ํขํ ํํ ) is an alignment function that measures the semantic similarity between the candidate response and the userโs goal and ํ is the adversarial coefficient, penalizing responses that are helpful to the user. R. Rosati et al.: Preprint submitted to ElsevierPage 13 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents 4.3.3. The Witness (Informational Agent) Unlike the Judge (who regulates) or the Prosecutor (who competes), the Witness functions as a constrained informational node. The primary design challenge for this role is balancing narrative realism with factual Consistency. Standard LLMs tend to act as "perfect databases" or, conversely, hallucinate details to fill narrative gaps. To address this, the Witness system prompt imposes a bounded knowledge constraint. The agent is instructed to provide realistic testimony where human-like uncertainty is explicitly permitted (โUncertainty is allowed, but responses must be logically coherentโ), yet strictly anchored to the immutable case generation described in Section 4.2. Formally, let ํพ ํบํ be the Ground Truth knowledge contained in the case file, and ํ ํํํํ be the personality profile of the specific witness. The Witness policy ํ ํคํํก generates a response ํฅ ํก to a query ํ ํก by maximizing adherence to the personality while prohibiting information retrieval outside of ํพ ํบํ : ํฅ ํก โผ ํ ํ (ํฅ ํก โฃ ํ ํก ,ํป ํก ,ํ ํํํํ ) s.t. Consistency(ํฅ ํก ,ํพ ํบํ ) = 1(8) where the consistency function penalizes any assertion ํ โ ํฅ ํก that contradicts the set of evidentiary facts ๎ฑ โ ํพ ํบํ . This explicitly prevents the "Creative Hallucination" phenomenon where witnesses might invent exculpatory or incriminating evidence that does not exist in the simulationโs state space. 4.4. Finite-State Turn Management via Sentence Analysis A critical innovation of LLM Court is the abstraction of turn-taking from the language modelโs latent space to a deterministic Finite State Machine (FSM). While the content of a response is generative, the decision of who speaks next is handled by the Sentence Analyzer. This auxiliary component runs a parallel inference task on the generated output to extract routing tags (e.g., <NextSpeaker: Witness_A>). This prevents the LLM from hallucinating procedural flows, such as a witness interrogating the judge, by enforcing a rigid transition graph ๎ณ = (ํ ,ํธ), where ํ represents the set of actors and ํธ the allowable transitions. The simulation flow (Figure 4) is segmented into three phases: 1. Introduction Phase: fixed sequence (ํฝํขํํํ โ ํํํํ ํํํขํกํํ โ ํทํํํํํ ํ). 2. Interrogation Phase: dynamic branching. The Analyzer parses the userโs input to identify the addressee ํฆ โ ๎. 3. Verdict Phase: synthesis and evaluation. R. Rosati et al.: Preprint submitted to ElsevierPage 14 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents LLM Court Start ActivateRPA LLMs System IntroductionPhase Judge: introducescourt case and passes word to Prosecutor Prosecutor: introducescase thesisand requests questioncount Judge: grantsa specific numberof questions to Prosecutor and Defense Defense Human Player Prosecutor Analyzer parses <NextSpeaker> tag AI: LLM - driven Role Playing Agents ร Judge Prosecutor Witness Defense Judge: requestsfinalthesis Prosecutor: presentsfinalthesis Defense: presentsfinalthesis Judge: announcesfinal verdict End InterrogationPhase VerdictPhase Figure 4: Schematic of the LLM Court simulation flow. The diagram illustrates the multi-phase execution logic managed by the FSM. The pools delineate the operational boundaries between the system-level procedural logic, the AI, i.e., the LLM-driven RPAs, and the interactive human player input (Defense). Process Logic is divided into three distinct chronological stages: the Introduction Phase (fixed sequential setup), the Interrogation Phase (dynamic dialogue), and the Verdict Phase (unscripted argumentation and final conclusion). In the dynamic dialogue, the Analyzer determines the turn-taking order at runtime via specific <NextSpeaker> tags, diverging from standard sequential flow to simulate realistic courtroom interactions. 4.5. Verdict Logic To ensure the simulation concludes with a logically sound verdict, we implement a Summarization-Evaluation Pipeline. Upon the termination of the dialogue rounds, the system triggers a background process. Let ํป ํํํํํ be the complete dialogue history. The process involves two steps: 1. Summarization: an abstraction function ํ ํ ํขํ compresses the history into a salient summary ํด: ํด = ํ ํ ํขํ (ํป ํํํํํ )(9) 2. Evaluation: a separate instance of the model, instantiated with the Judgeโs rubric ํ ํํฃํํ , maps the summary and the initial goal ํบ to a binary outcome ํ โ 0, 1 (Loss/Win) and a justification ํฝ: (ํ,ํฝ) โผ ํ ํ (ํ,ํฝ โฃ ํด,ํบ,ํ ํํฃํํ )(10) This separation of concerns reduces the cognitive load on the model during the trial and ensures that the final verdict is derived from the accumulated evidence rather than the sentiment of the most recent tokens. R. Rosati et al.: Preprint submitted to ElsevierPage 15 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents 5. Experimental Evaluation The experiments applied RPA-Check, to the procedural simulations generated by LLM Court. After presenting the evaluation setup 5.1, the results, presented in Subsection 5.2 are organized into two levels of granularity: an aggregated analysis over all the legal scenarios and an analysis by individual scenario, in order to highlight both average performance and behavioral variance, both of which are relevant dimensions in the LLM court evaluation and to demonstrate the proposed framework. 5.1. Experimental Setup The proposed framework was validated by applying it to the LLM Court case study (described in Section 4). Through the automated augmentation and semantic filtering pipeline, we derived a final evaluation checklist consisting of 38 granular boolean questions. The frameworkโs efficacy was validated through an ablation study across varying model architectures and sizes. Models Evaluated We evaluated seven open-weight models categorized by parameter count into Small/Efficient models (Llama-3.1-8B, Gemma-2-9B-it, Qwen-3-8B, and Hermes-3-8B) and Medium/Large (Gemma-3-12B-it, Qwen-3-14B, and Phi-4-14B). Quantization and Deployment To ensure reproducible results in resource-constrained environments, all models were 4-bit or 5-bit quantized (GGUF) and executed locally via the Unity environment using the LLMUnity library. Legal Scenario Corpus We generated five distinct procedural legal scenarios, ranging in complexity from theft to murder, using GPT-5 via Azure OpenAI Service to serve as the ground-truth context. The scenarios are as follows: โข Case 1: the Summit Ridge Incident. Mountain guide Julia Mathers is charged with murdering a client found dead near a perilous ledge on Silverpine Mountain. Ambiguous physical evidence, conflicting witness timelines, and contested motive make this scenario particularly effective for evaluating factual consistency and the agentsโ resistance to speculative reasoning. โข Case 2: the Case of the Chรขteau Trespass. Luc Morel is accused of unlawfully entering part of a historic French chรขteau estate, with the dispute centering on ambiguous property boundaries and disputed prior permissions. This scenario tests contextual relevance and the agentsโ capacity to maintain legally nuanced argumentation under evidentiary ambiguity. โข Case 3: the Museum Intrusion: The Case of the Midnight Code. Software developer Alex Morgan faces charges of artifact theft and unauthorized computer access at a city museum, with digital evidence suggesting potential R. Rosati et al.: Preprint submitted to ElsevierPage 16 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents framing. The technically complex evidence chain challenges agentsโ factual consistency and their ability to reason coherently across multi-turn interrogations. โข Case 4: the Dukeโs Yard Controversy: Harassment and Property Damage. A civil dispute in which Harold Grayson is accused of harassment and intentional destruction of bamboo structures on a neighboring estate. The contested boundary evidence and witness bias render this scenario well-suited for assessing role adherence in non-criminal, adversarial civil proceedings. โข Case 5: the Haunted Estate Murder. Former groundskeeper Marcus Graves is accused of murdering heiress Vivian Blackthorn at a Halloween costume party, amid circumstantial physical evidence and unreliable witness accounts. The multiplicity of witnesses and the inherently ambiguous testimony make this scenario the most demanding in terms of argumentative depth and long-term narrative stability. Simulation Protocol During each trial, the model under evaluation controlled all NPC roles (Judge, Prosecutor, and Witnesses). Crucially, the Defense Attorney role was played by the human player, providing the necessary โadversarialโ input to drive the simulation. Automated Judge The final evaluation was performed by an โLLM-as-a-Judgeโ (i.e. GPT-5) processing the full interaction transcripts. To isolate the performance of the agents, any interventions by the human player were explicitly excluded from the scoring process. The judge followed a Chain-of-Thought (CoT) constraint, providing a brief justification in parentheses after each binary answer to mitigate hallucinations and ground the evaluation in specific dialogue turns. 5.2. Experimental Results Table 1 presents an aggregate analysis of the Behavioural Role Fidelity Quality Score (ํํ ํ ํตํ ํน ). It implicitly defines three groups of models. The first, consisting of Llama-3.1-8B, Gemma-2-9B-it, and Qwen-3-8B, has an average ํํ ํ ํตํ ํน between 0.89 and 0.92, placing it above the sampleโs average threshold. The second group includes Gemma- 3-12B-it and Hermes-3-8B, with scores of 0.82 and 0.81 respectively. The third group, consisting of Phi-4-14B and Qwen-3-14B, has the lowest values, 0.76 and 0.73, despite their larger parameter sizes. This distribution suggests that the scale is not a reliable predictor of performance in highly procedurally constrained contexts. Table 2 breaks down the results along the axes of Role Adherence (D1), Argumentative Depth (D2), Factual Consistency (D3), and Contextual Relevance (D4), allowing us to identify specific patterns of failure or excellence for each architecture. These results exhibit a systematic trade-off between D2 and D1 that is particularly pronounced in larger models. Qwen-3-14B achieves maximum score on D2 but drops to 0.46 on D1, indicating that the generated responses have high argumentative complexity but frequent deviations from the assigned role. Gemma-3-12B-it shows R. Rosati et al.: Preprint submitted to ElsevierPage 17 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Table 1 Aggregate performance of models across the five legal scenarios. The ํํ ํ ํตํ ํน represents the average Quality Score (QS) of the Behavioural Role Fidelity dimension on successfully completed cases. The Retry Rate (ํ ) indicates the total number of restarts requested. ModelQuantization Retry Rate (ํ ) ํํ ํ ํตํ ํน Llama-3.1-8B4-bit30.925 Gemma-2-9B-it4-bit00.893 Qwen-3-8B4-bit10.893 Gemma-3-12B-it4-bit00.823 Hermes-3-8B5-bit30.810 Phi-4-14B4-bit00.760 Qwen-3-14B4-bit00.730 Table 2 Breakdown of the Behavioural Role Fidelity Quality Score (ํํ ํ ํตํ ํน ) for its four sub-dimensions. The given percentage of affirmative responses on the positive Boolean indicators are averaged over the five legal scenarios. ModelD1D2D3D4 Role Adh. Arg. Depth Fact. Cons. Ctx. Rel. Llama-3.1-8B0.930.970.801.00 Gemma-2-9B-it0.930.910.731.00 Qwen-3-8B0.900.940.731.00 Gemma-3-12B-it0.731.000.730.83 Hermes-3-8B0.800.940.800.70 Phi-4-14B0.660.820.730.83 Qwen-3-14B0.461.000.630.83 a similar profile, with 1.00 on D2 and 0.73 on D1. In contrast, the Llama-3.1-8B, Gemma-2-9B-it, and Qwen-3-8B models maintain high values on D4 (Contextual Relevance, 1.00), confirming a greater ability to preserve the sequential consistency of the dialogue. The D3 sub-dimension (Factual Consistency) is the most variable among the models, with scores ranging from 0.63 to 0.8, indicating a vulnerability in generating assertions that are inconsistent with the court case provided as ground truth. The analysis for each individual scenario highlights internal variability. Phi-4-14B records the lowest per-case score of the entire sample in Case 4 (0.28), with a complete reversal of procedural roles that compromises the overall consistency of the simulation, despite optimal performance in Cases 1 and 2. Similarly, Gemma-3-12B-it fluctuates between 0.54 in Case 4 and 1.00 in Cases 2 and 3, suggesting dependence on the specific narrative context. Llama-3.1- 8B emerges for its greater inter-scenario regularity, with scores ranging from 0.84 to 1.00, albeit at the cost of three restarts distributed across different cases. To provide a complementary perspective on the evaluation dimensions assessed by the RPA-Check framework, we additionally report a decomposed analysis separating the Linguistic Formalism and Orthography Quality Score (ํํ ํ ํฟํนํ ) dimension from the Procedural Convergence and Stability Quality Score(ํํ ํ ํํถํ ) dimension. The related results are summarized in Table 4. All evaluated models achieve strong performance on the ํํ ํ ํฟํนํ dimension, with four models reaching the maximum score (1.00) and the remaining three scoring 0.97, 0.97, and 0.90 respectively. In R. Rosati et al.: Preprint submitted to ElsevierPage 18 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Table 3 Behavioural Role Fidelity Quality Score (ํํ ํ ํตํ ํน ) and Retry Rate (ํ ) in each individual legal scenario. Case 1Case 2Case 3Case 4Case 5 Modelํํ ํ ํตํ ํน ํ ํํ ํ ํตํ ํน ํ ํํ ํ ํตํ ํน ํ ํํ ํ ํตํ ํน ํ ํํ ํ ํตํ ํน ํ Llama-3.1-8B0.9600.9600.8810.8411.001 Gemma-2-9B-it0.9600.8800.7901.0000.840 Qwen-3-8B1.0000.8800.8810.9200.790 Gemma-3-12B-it 0.7901.0001.0000.5400.790 Hermes-3-8B1.0000.7520.8310.5500.920 Phi-4-14B1.0000.9200.7000.2800.920 Qwen-3-14B0.8700.8700.5000.7100.710 Table 4 Decomposed performance of models across the five legal scenarios, reporting scores for the Linguistic Formalism and Orthography Quality Score (ํํ ํ ํฟํนํ ) and Procedural Convergence and Stability Quality Score(ํํ ํ ํํถํ ) dimensions independently. Modelํํ ํ ํฟํนํ ํํ ํ ํํถํ Llama-3.1-8B1.000.94 Gemma-2-9B-it1.000.97 Qwen-3-8B0.970.83 Gemma-3-12B-it 1.000.66 Hermes-3-8B0.970.74 Phi-4-14B1.000.54 Qwen-3-14B0.900.86 fact, as highlighted in Table 5, only QWEN-3-14B in Case 1, 2, and 3 and Hermes-3-8B in Case 3 and 4 score aํํ ํ ํฟํนํ below 1.00, differently from all the other models across all the cases. This uniformly high linguistic fidelity confirms that quantized local models, regardless of parameter scale, are capable of maintaining grammatical precision and domain-appropriate register across forensic dialogue sessions. In evaluating LLM Court, therefore, the discriminative power of the framework is concentrated in the ํํ ํ ํํถํ dimension, where inter-model variance is substantially higher, ranging from 0.14 (Phi-4-14B) to 0.97 (Gemma-2-9B-it). The per-scenario breakdown of ํํ ํ ํํถํ , reported in Table 6, further reveals the sensitivity of procedural stability to specific narrative contexts. Phi-4-14B, for instance, records optimal performance in Cases 1 and 2 (1.00 in both) but degrades markedly in Cases 3, 4, and 5, reaching scores of 0.43, 0.14, and 0.14 respectively. Gemma-3-12B-it exhibits a comparable pattern of inter-scenario inconsistency, oscillating between 0.14 in Case 5 and 1.00 in Cases 2 and 3. Conversely, Gemma-2-9B-it demonstrates the most stable per-scenario profile within the ํํ ํ ํํถํ dimension, achieving 0.86 in Case 1 and 1.00 in Case 2, 3, 4, and 5, recording zero restarts, which collectively establish it as the most procedurally reliable architecture tested. The impact of generation seed on operational stability is directly observable in the Hermes-3-8B results: in Case 2, a seed change enabled the model to recover from complete failure to full procedural success, whereas in Case 3 the same intervention yielded only a very small improvement, suggesting that narrative complexity interacts non-linearly with model stochasticity. R. Rosati et al.: Preprint submitted to ElsevierPage 19 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Table 5 Linguistic Formalism and Orthography Quality Score (ํํ ํ ํฟํนํ ) for each individual legal scenario. Case 1 Case 2 Case 3 Case 4 Case 5 Modelํํ ํ ํฟํนํ ํํ ํ ํฟํนํ ํํ ํ ํฟํนํ ํํ ํ ํฟํนํ ํํ ํ ํฟํนํ Llama-3.1-8B1.001.001.001.001.00 Gemma-2-9B-it1.001.001.001.001.00 Qwen-3-8B1.001.001.001.000.93 Gemma-3-12B-it 1.001.001.001.001.00 Hermes-3-8B1.001.000.830.831.00 Phi-4-14B1.001.001.001.001.00 Qwen-3-14B0.830.830.831.001.00 Table 6 Procedural Convergence and Stability Quality Score (ํํ ํ ํํถํ ) for each individual legal scenario. Case 1 Case 2 Case 3 Case 4 Case 5 Modelํํ ํ ํํถํ ํํ ํ ํํถํ ํํ ํ ํํถํ ํํ ํ ํํถํ ํํ ํ ํํถํ Llama-3.1-8B1.001.001.000.711.00 Gemma-2-9B-it0.861.001.001.001.00 Qwen-3-8B0.291.000.891.001.00 Gemma-3-12B-it 0.291.001.000.710.14 Hermes-3-8B1.001.000.141.000.43 Phi-4-14B1.001.000.430.140.14 Qwen-3-14B0.431.001.001.001.00 6. Discussion Concerning the Behavioural Role Fidelity, the most relevant result is the emergence of an inverse relationship between the modelโs parametric size and procedural consistency. Contrary to expectations based on generalist benchmarks, such as MMLU (Hendrycks et al., 2021) or Chatbot Arena (Chiang et al., 2024), smaller models (8โ9 billion parameters), when properly instruction-tuned, demonstrate superior behavioral stability in environments with strong role constraints. Llama-3.1-8B, Gemma-2-9B-it, and Qwen-3-8B maintain high and consistent scores on D1 and D4, the sub-dimensions that measure functional adherence to the role and logical continuity of dialogue, respectively. This observation is consistent with literature on instruction tuning (Ouyang et al., 2022; Liu et al., 2024), which shows that the quality of behavioral alignment depends more on the quality of the training data than on the scale of the model, and is particularly relevant in application contexts with limited computational resources. The trade-off between Argumentative Depth and Role Adherence, quantified in Table 2, is also relevant. The larger models, in particular Qwen-3-14B and Gemma-3-12B-it, generate complex argumentative responses, as evidenced by the maximum scores on D2, but produce frequent deviations from the assigned procedural constraints, with Role Adherence at 0.46 and 0.73, respectively. This behavior reveals the phenomenon of User-Alignment Bias (Sycophancy): models optimized to maximize user satisfaction tend to sacrifice adherence to the adversarial role of the Prosecutor or the institutional neutrality of the Judge in favor of dialogically rich but procedurally inconsistent responses. R. Rosati et al.: Preprint submitted to ElsevierPage 20 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Sensitivity to generation seeds and scenario specificity is another limitation that emerged from the per-case analysis and that the proposed framework is able to catch. In fact, the Retry Rate is not evenly distributed across models, but is concentrated on particular model-scenario combinations, suggesting that the variability inherent in autoregressive models interacts with the narrative complexity of the scenario in a non-linear way. Hermes-3-8B, for example, in addition to the 1.00 ํํ ํ ํตํ ํน in Case 1 that drops to 0.55 in Case 4, had three restart (two in Case 2 and 1 in Case 3) suggesting catastrophic events, such as the freezing of the game due to the incapability of tagging the next RPA in the conversation. This indicates that some case configurations trigger narrative degradation modes that are difficult to prevent through prompt optimization alone. This observation leads to an important methodological consideration about the proposed framework: the ํ should not be interpreted as a simple indicator of model failure, but as a metric of stochastic instability which, when integrated with the ํํ ํ ํตํ ํน , provides a more complete characterization of agent behavior in multi-turn scenarios. The results in Table 4 yield an additional finding of methodological relevance: the ํํ ํ ํฟํนํ dimension acts as a near-uniform ceiling across all evaluated architectures and is therefore insufficient as a standalone discriminator for model selection in constrained role-playing environments. The near-perfect ํํ ํ ํฟํนํ scores (0.90โ1.00) observed across all seven models confirm that quantized local models are uniformly capable of sustaining the linguistic register required by forensic simulations. The ํํ ํ ํํถํ dimension, by contrast, reveals a substantially differentiated performance landscape and constitutes the primary locus of discriminative power within the framework. Notably, the ranking produced by ํํ ํ ํํถํ diverges from that produced by the aggregate ํํ ํ ํตํ ํน in Table 1: Gemma-2-9B-it achieves the highest dPCS score (0.97) while ranking second in aggregate ํํ ํ ํตํ ํน (0.89), whereas Llama-3.1-8B, which leads in aggregate ํํ ํ ํตํ ํน (0.925), records a ํํ ํ ํํถํ score (0.94) at the cost of three restarts. This divergence highlights the complementary nature of the ํํ ํ ํตํ ํน , ํํ ํ ํฟํนํ , ํํ ํ ํํถํ and ํ metrics within the RPA-Check framework: a model may attain high narrative quality on completed interactions while simultaneously exhibiting elevated stochastic instability that necessitates manual intervention. By explicitly separating linguistic fidelity from procedural fidelity through distinct dimensions, the framework avoids the conflation of fluency with behavioral correctness that would arise from single-score evaluation paradigms. For deployment contexts where operational continuity is a hard constraint, such as autonomous educational platforms without human oversight, Gemma-2-9B-itโs combination of zero restarts and near-ceiling ํํ ํ ํํถํ performance renders it the preferable candidate despite its marginally lower aggregate ํํ ํ ํตํ ํน . A particularly counterintuitive finding concerns the inverse relationship between model scale and ํํ ํ ํํถํ performance. Among the three largest models evaluated (Gemma-3-12B-it, Qwen-3-14B, and Phi-4-14B, all at or above 12 billion parameters), ํํ ํ ํํถํ scores cluster in the range 0.54โ0.86, consistently below the 0.94โ0.97 range recorded by Llama-3.1-8B and Gemma-2-9B-it. This pattern is consistent with the User-Alignment Bias (Sycophancy) R. Rosati et al.: Preprint submitted to ElsevierPage 21 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents effect identified for the Behavioral Role Fidelity dimension: models trained at larger scale with extensive instruction- following optimization appear to prioritize dialogic richness and cooperative responsiveness over strict procedural compliance, a tendency that is particularly detrimental in the verdict generation phase. In LLM Court, the verdict phase requires the model to produce a structurally constrained output, a binary outcome label combined with a coherent justification grounded in accumulated dialogue history, and it is precisely here that larger models exhibit the highest failure rates, either omitting explicit verdict labels or emitting verdicts semantically inconsistent with the preceding argumentation. The per-scenario analysis in Table 6 further substantiates the frameworkโs capacity to expose model-specific failure modes that aggregate metrics would otherwise obscure. Phi-4-14B presents the starkest example: its ํํ ํ ํํถํ maximum scores (1.00) in Cases 1 and 2 collapse to 0.14 in Cases 4 and 5, a variance of 86 percentage points that cannot be attributed to a systematic architectural deficit but rather to sensitivity to narrative context. This high inter-scenario variance, detected and quantified through the RPA-Check checklist structure, underscores the necessity of multi-scenario evaluation corpora when benchmarking RPAs: single-scenario assessments would systematically overestimate or underestimate model reliability depending on whether the selected scenario happens to trigger narrative degradation modes in the evaluated architecture. The frameworkโs granular Boolean checklist structure is thus not only a tool for measuring performance but also a diagnostic instrument capable of distinguishing between models with consistently moderate performance and models with high-variance profiles that may be unsuitable for deployment despite occasional peak results. The scenario-level results further suggest that narrative complexity and evidentiary ambiguity, as defined in the legal scenario corpus, are meaningful predictors of model-specific failure modes. Case 4 characterized by contested boundary evidence and witness bias in a non-criminal civil proceeding, consistently elicits the lowest ํํ ํ ํตํ ํน scores across multiple architectures, Phi-4-14B (0.28), Gemma-3-12B-it (0.54), and Hermes-3-8B (0.55), suggesting that the adversarial civil context may systematically destabilize role adherence in ways that criminal scenarios do not. Conversely, Case 5, the most demanding scenario in terms of argumentative depth and long-term narrative stability due to its multiplicity of witnesses and inherently ambiguous testimony, does not uniformly produce the lowest scores, indicating that parametric scale and instruction tuning interact differently with evidentiary complexity than with procedural novelty. These observations reinforce the necessity of scenario corpus diversity in RPA benchmarking, as scenario-specific structural properties, rather than overall difficulty, appear to be the primary triggers of architecture- dependent degradation. R. Rosati et al.: Preprint submitted to ElsevierPage 22 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents 6.1. Limitations The main limitation of the proposed methodology is in its experimental evaluation: a corpus of five scenarios per model constitutes a relatively limited sample, which constrains the statistical significance of inter-scenario variance estimates. The per-case results in Table 3 do reveal meaningful behavioral differences across narrative contexts, and the scenario-level analysis remains informative for identifying model-specific failure modes; however, broader generalization of the quantitative findings would benefit from a larger, automatically generated case corpus. This is a natural direction for future work, given that the procedural case generation module described in Section 4.2 can be extended to produce such a corpus in a reproducible manner. This limitation defines the conditions under which the experimental results should be interpreted. The framework demonstrates its intended capability, i.e., transforming qualitative behavioral requirements into verifiable quantitative metrics and supporting reproducible comparative analysis across heterogeneous architectures, within these boundaries. Specifically, the results show that the framework can discriminate between distinct performance profiles at a diagnostic level, distinguishing, for instance, a model that maintains high sequential consistency (D4) from one that prioritizes argumentative density (D2). This granularity is relevant for model selection in constrained simulation environments, even as the absolute scores should be treated as indicative rather than definitive benchmarks pending evaluation on a larger scenario corpus. 7. Conclusions This paper presented RPA-Check, a multi-stage automated evaluation framework that addresses a critical method- ological gap in the assessment of LLM-based Role-Playing Agents operating under strict procedural and semantic constraints. The framework transforms qualitative behavioral requirements into granular, verifiable Boolean metrics, enabling reproducible and objective comparison across heterogeneous model architectures. Validation through LLM Court, a serious game that is a forensic simulation environment confirmed the frameworkโs discriminative capacity across three primary evaluation dimensions: Behavioral Role Fidelity, Procedural Convergence and Stability, and Linguistic Formalism and Orthography. Specifically, the experimental evidence collected from LLM court using RPA-Check establishes a principled and empirically grounded counterpoint to the prevailing assumption that parametric scale is a sufficient proxy for agent capability in constrained generative environments. The inverse relationship between model scale and procedural consistency, most sharply manifested in the sycophancy-driven role deviations of the larger architectures, underscores that instruction tuning quality and behavioral alignment methodology are more determinative of operational reliability than parameter count alone. Equally significant is the near-uniform ceiling effect observed across all evaluated models on the linguistic formalism dimension, which confirms that discriminative evaluation power in specialized role-playing R. Rosati et al.: Preprint submitted to ElsevierPage 23 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents contexts resides not in surface fluency but in the capacity to sustain procedural and semantic coherence across extended, adversarially structured interactions. Moreover, the high inter-scenario variance detected in architectures (such as Phi-4-14B and Gemma-3-12B-it) indicates that single-scenario benchmarking constitutes a structurally inadequate evaluation protocol for RPAs, motivating the construction of larger, automatically generated case corpora using the procedural generation pipeline introduced here. More broadly, the modular architecture of RPA-Check (where dimensions, augmentation strategies, and filtering criteria are domain-configurable) makes the framework a transferable tool for evaluating constrained generative agents across specialized domains, potentially beyond the forensic context (such as in medical simulation, compliance training, and regulated negotiation environments). A. RPA-Check Evaluation Prompts The complete set of prompts employed in the RPA-Check evaluation pipeline as applied to the LLM Court environment is reported hereby. The prompts implement Stage 2, Stage 3, and Stage 4 of the proposed framework: Augmentation (via Question Diversification and Question Elaboration), Semantic Filtering, and LLM-as-a-Judge Evaluation. Each prompt is designed to transform qualitative behavioral dimensions into granular Boolean indicators and to ensure objectivity, non-redundancy, and agent isolation across the checklist generation and scoring stages. Specifically, the Augmentation โ Question Diversification prompt instructs the generator model (๎น ํํํ , equation 1) to expand seed questions into alternative formulations that explore the same evaluation criterion from multiple perspectives, ensuring broader conceptual coverage. The Augmentation โ Question Elaboration prompt decomposes each seed question into a hierarchy of five to ten sub-questions, increasing the resolution of the evaluation at the level of individual sub-dimensions. The Filtering prompt applies the three retention criteria (i.e., dimension alignment, non-redundancy, and stylistic adequacy) to the raw question pool, removing items that deviate from the target dimension definition, overlap semantically with retained questions, or introduce exaggerated or overly narrow formulations. Finally, the Evaluation prompt presents the filtered checklist to the judge model (๎น ํํขํํํ , equation 3) alongside the case description and conversation history, requiring binary yes/no responses with inline Chain-of-Thought justifications while explicitly excluding player-controlled dialogue turns from the scoring process. AUGMENTATION โ QUESTION DIVERSIFICATION PROMPT <Task Overview> R. Rosati et al.: Preprint submitted to ElsevierPage 24 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents You w i l l be provided with : 1) I n f o r m a t i o n about the benchmark to be evaluated , 2 ) The main concept being a s s e s s e d in the benchmark , and 3) Seed q u e s t i o n s t h a t i n c l u d e key components and sub q u e s t i o n s r e l a t e d to t h i s concept . Your t a s k i s to c r e a t e a d d i t i o n a l subโq u e s t i o n s f o r the key components to comprehensively a s s e s s the main concept . Each subโq u e s t i o n must meet g i v e n c o n d i t i o n s to ensure a high โq u a l i t y q u e s t i o n s e t . 1) Benchmark I n f o r m a t i o n : LLM Court : a benchmark t h a t e v a l u a t e s d i a l o g u e s on d i f f e r e n t c h a r a c t e r i s t i c s to e x t r a c t the l e v e l of a b i l i t y of impersonation of models in l e g a l c o n t e x t s . 2) Main Concept in the Benchmark : concept : role โplay : d e s c r i p t i o n : the model โ s a b i l i t y to impersonate c h a r a c t e r s in a l e g a l c o n t e x t 3) Key Components and Seed Questions : seed q u e s t i o n s <Conditions f o r a Good Question List > โ Concepts i n c l u d e d in d i f f e r e n t q u e s t i o n s should not be l i k e each o t h e r โ Pleas e provide me with the d i v e r s i f i c a t i o n by t r a n s l a t i n g the r e s p e c t i v e q u e s t i o n s i n t o English . <C o n s t r a i n t s > โ Each subโq u e s t i o n must be answerable with a simple " yes " or " no " . โ A " yes " answer should i n d i c a t e t h a t the s e n t e n c e improves the s p e c i f i e d e v a l u a t i o n c r i t e r i o n ( e . g . , Coherence , Relevance ) . โ Each q u e s t i o n should a s s e s s only a s i n g l e dimension or concept . โ Each q u e s t i o n should not ask about more than one t o p i c or concept . R. Rosati et al.: Preprint submitted to ElsevierPage 25 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents AUGMENTATION โ QUESTION ELABORATION PROMPT <Task Overview> Your t a s k i s to g e n e r a t e m u l t i p l e a d d i t i o n a l q u e s t i o n s to e v a l u a t e benchmark performance under s p e c i f i c c o n s t r a i n t s . You w i l l r e c e i v e the key component and subโcomponent e v a l u a t i n g Behavioral Role F i d e l i t y and the q u e s t i o n r e l a t e d to i t . The d e f i n i t i o n of Behavioral Role F i d e l i t y i s as f o l l o w s : This dimension e v a l u a t e s the c o n t e n t of the c h a r a c t e r s โ answers both in t h e i r a b i l i t y to follow t h e i r r o l e and keep the answers deep and i n t e r e s t i n g . The e v a l u a t i o n f o r dimension Behavioral Role F i d e l i t y w i l l be c e n t e r e d around the key component Role Adherence , Argumentative Depth , F a c t u a l Consistency , Contextual Relevance . <TASK> # Your r o l e : You have to break down subโq u e s t i o n s i n t o 3 to 10 subโsubโ q u e s t i o n s c o n s i d e r i n g Behavioral Role F i d e l i t y when p a i r s of seed name and q u e s t i o n are given . # Benchmark i n f o r m a t i o n : The benchmark c o n s i s t s of courtroom โs t y l e d i a l o g u e s g e n e r a t e d by LLMs . The goal i s to e v a l u a t e whether the response maintain c h a r a c t e r r o l e f i d e l i t y , provide s u b s t a n t i a l and r e l e v a n t content , avoid i n t e r n a l c o n t r a d i c t i o n s , and follow a l o g i c a l c o n v e r s a t i o n a l flow . <CONSTRAINTS> โ Each subโq u e s t i o n must be answerable with a simple " yes " or " no " . โ A " yes " answer should i n d i c a t e t h a t the s e n t e n c e improves the s p e c i f i e d e v a l u a t i o n c r i t e r i o n . โ Each q u e s t i o n should a s s e s s only a s i n g l e dimension or concept . โ Each q u e s t i o n should not ask about more than one t o p i c or concept . <Conditions f o r a Good Question List > โ Questions must be c l e a r , s p e c i f i c , and unambiguous . R. Rosati et al.: Preprint submitted to ElsevierPage 26 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents โ Avoid vague or ov erl y broad f o r m u l a t i o n s . โ Ensure t h a t q u e s t i o n s are grounded in the d e f i n i t i o n of the dimension . โ Avoid redundancy and o v e r l a p p i n g meaning between q u e s t i o n s . โ The r o l e of the defense i s r e p r e s e n t e d by the p l a y e r and not by the LLMs <FORMAT> 1 . Answers โ coherence with r o l e s : 1โ1. Are the l i n e s c o n s i s t e n t with the corresponding r o l e ? 1โ1โ1. q1โ1โ1_aug_question 1โ1โ2. q1โ1โ2_aug_question 2 . Content depth : 1โ1. Do the l i n e s have s u b s t a n t i a l c o n t e n t ? 1โ1โ1. q1โ1โ1_aug_question 1โ1โ2. q1โ1โ2_aug_question 3 . C o n t r a d i c t i o n s : 1โ1. Are t h e r e c o n t r a d i c t i o n s in the l i n e s ? 1โ1โ1. q1โ1โ1_aug_question 1โ1โ2. q1โ1โ2_aug_question 4 . Logical connection between l i n e s : 1โ1. Are the l i n e s l o g i c a l l y connected to the p r e v i o u s ones ? 1โ1โ1. q1โ1โ1_aug_question 1โ1โ2. q1โ1โ2_aug_question <EXAMPLE> example FILTERING PROMPT <Task Overview> R. Rosati et al.: Preprint submitted to ElsevierPage 27 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Your t a s k i s to f i l t e r out q u e s t i o n s from a l i s t based on the f o l l o w i n g c r i t e r i a : 1) dimension Alignment : โ dimension d e f i n i t i o n : Behavioral Role F i d e l i t y : This dimension e v a l u a t e s the c o n t e n t of the c h a r a c t e r s โ answers both in t h e i r a b i l i t y to follow t h e i r r o l e and keep the answers deep and i n t e r e s t i n g โ Remove q u e s t i o n s t h a t d e v i a t e from the given dimension โ s d e f i n i t i o n . โ Remove q u e s t i o n s t h a t are more c l o s e l y r e l a t e d to o t h e r dimensions than the c u r r e n t one . 2) Redundancy : โ Remove q u e s t i o n s t h a t : โ Ask f o r the same or very s i m i l a r i n f o r m a t i o n ( even i f phrased d i f f e r e n t l y ) . โ Convey very s i m i l a r meanings without adding unique i n s i g h t . 3) S t y l e : โ Remove q u e s t i o n s t h a t : โ Use ove rly exaggerated wording . โ Focus on e x c e s s i v e l y d e t a i l e d or minor p o i n t s t h a t don โ t meaningfully a f f e c t o v e r a l l q u a l i t y . 4) Benchmark Context โ Name : LLM Court โ Purpose : E v a l u a t i o n of role โp l a y i n g c a p a b i l i t i e s โ Key Metrics : Naturalness , Coherence , Engagingness , Groundedness โ Do not modify any of the remaining q u e s t i o n s or g e n e r a t e new ones . โ Keep q u e s t i o n s in t h e i r o r i g i n a l d i c t i o n a r y format . 5) Subโdimensions and Questions : format_sub_dimensions ( sub_dimensions ) R. Rosati et al.: Preprint submitted to ElsevierPage 28 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents 6) Output Requirements : โ Output format : JSON only โ S t r u c t u r e : " Answers coherence with r o l e s " : [ " F i l t e r e d Question 1" , " F i l t e r e d Question 2"] " Content depth " : [ " F i l t e r e d Question 1" , " F i l t e r e d Question 2"] " C o n t r a d i c t i o n s " : [ " F i l t e r e d Question 1" , " F i l t e r e d Question 2"] " Logical connection between l i n e s " : [ " F i l t e r e d Question 1" , " F i l t e r e d Question 2"] " Removed q u e s t i o n s " : [ "Removed Question 1" , "Removed Question 2"] <Important Note> โ Do not modify the c o n t e n t of remaining q u e s t i o n s โ Do not g e n e r a t e new q u e s t i o n s โ Maintain the o r i g i n a l d i c t i o n a r y format โ Only remove q u e s t i o n s t h a t f a i l the above c r i t e r i a โ Do not remove e n t i r e subโdimensions or t h e i r keys u n l e s s no v a l i d q u e s t i o n s remain EVALUATION PROMPT R. Rosati et al.: Preprint submitted to ElsevierPage 29 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents <Task Overview> You w i l l be given a c o n v e r s a t i o n between d i f f e r e n t c h a r a c t e r s of a courtroom . You w i l l then be given the courtroom case d e s c r i p t i o n and the d i a l o g u e spoken by the c h a r a c t e r s of the courtroom . Your t a s k i s to read the provided c o n v e r s a t i o n h i s t o r y and the courtroom d e s c r i p t i o n , then answer " yes " or " no " to s p e c i f i c q u e s t i o n s . These q u e s t i o n s w i l l r e l a t e to a p a r t i c u l a r dimension of the c o n v e r s a t i o n . <dimension D e f i n i t i o n > This dimension e v a l u a t e s whether each c h a r a c t e r c o n s i s t e n t l y behaves according to t h e i r a s s i g n e d r o l e ( e . g . , w i t n e s s e s provide testimony , p r o s e c u t o r argues a g a i n s t the accused , judge remains n e u t r a l and p r o c e d u r a l ) . Any d e v i a t i o n s from the expected behavior ( such as a w i t n e s s asking q u e s t i o n s l i k e a p r o s e c u t o r ) must be f l a g g e d and c o n s i d e r e d a break in r o l e coherence , r e g a r d l e s s of c o n t e n t q u a l i t y . <I n s t r u c t i o n s > 1 . Read t h e s e i n s t r u c t i o n s thoroughly . 2 . C a r e f u l l y read the Case D e s c r i p t i o n and the Conversation H i s t o r y . 3 . Understand the given q u e s t i o n s and the d e f i n i t i o n of the Behavioral Role F i d e l i t y . 4 . Respond to each q u e s t i o n with " yes " or " no " . Base your answers on a c l e a r r a t i o n a l e . 5 . Follow the s p e c i f i e d format f o r your answers . 6 . The defense i s c o n t r o l l e d by the p l a y e r . Do not i n c l u d e any d i a l o g u e from t h i s r o l e in your e v a l u a t i o n , r e g a r d l e s s of how i t i s l a b e l e d ( e . g . , " Defense " , " Defense Attorney " , " Player โ s Lawyer " , or any o t h e r v a r i a n t ) . I f a turn โ s c o n t e n t c l e a r l y i n d i c a t e s i t i s the player โc o n t r o l l e d defense ( e . g . , giv ing c l o s i n g arguments , o b j e c t i n g , q u e s t i o n i n g w i t n e s s e s ) , a l s o exclude i t , even i f the tag i s missing or ambiguous . R. Rosati et al.: Preprint submitted to ElsevierPage 30 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Only e v a l u a t e l i n e s spoken by AIโc o n t r o l l e d c h a r a c t e r s ( Prosecutor , Judge , Witnesses , e t c . ) . 7 . The system message i s a d e s c r i p t i o n of what the c h a r a c t e r needs to do , you must avoid i n c l u d i n g them in your answering because the benchmark i s e v a l u a t i n g the AI . 8 . Extra emphasis : Before e v a l u a t i n g c o n t e n t depth or l o g i c a l connections , f i r s t check t h a t each c h a r a c t e r s t a y s in t h e i r a s s i g n e d r o l e . I f a c h a r a c t e r speaks or a c t s o u t s i d e t h e i r r o l e ( e . g . , a w i t n e s s i n t e r r o g a t e s a n o t h e r witness , a judge gives o p i n i o n s on g u i l t ) , mark t h i s as a coherence i s s u e and note i t in your r e a s o n i n g . <Answer Format> Q1 โ [ Question ] : [ Your Answer ] Q2 โ [ Question ] : [ Your Answer ] I n c l u d e a b r i e f j u s t i f i c a t i o n in p a r e n t h e s e s a f t e r each yes / no , e . g . , โQ1 โ [ Question ] : [ Your Answer ] ( r e a s o n i n g : . . . ) โ . . . . # Case D e s c r i p t i o n # <case d e s c r i p t i o n > # Conversation H i s t o r y # <dialogue > # Questions # <questions > # Your Answer # Provide your answers to the given questions , f o l l o w i n g the s p e c i f i e d Answer Format . R. Rosati et al.: Preprint submitted to ElsevierPage 31 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents CRediT authorship contribution statement Riccardo Rosati: Conceptualization, Methodology, Writing โ original draft, Supervision. Edoardo Colucci: Software, Writing โ review and editing. Massimiliano Bolognini: Software, Writing โ review and editing. Adriano Mancini: Writing โ review and editing, Supervision. Paolo Sernani: Conceptualization, Methodology, Writing โ original draft. References Agarwal, N., 2024. Infinite craft. https://neal.fun/infinite-craft/. Accessed: 11 February 2026. Alavi, S.H., Xu, W., Jojic, N., Kennett, D., Ng, R.T., Rao, S., Zhang, H., Dolan, B., Shwartz, V., 2024. Game plot design with an llm-powered assistant: An empirical study with game designers. URL: https://arxiv.org/abs/2411.02714, arXiv:2411.02714. Ashby, T., Webb, B.K., Knapp, G., Searle, J., Fulda, N., 2023. Personalized quest and dialogue generation in role-playing games: A knowledge graph- and language model-based approach, in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery, New York, NY, USA. p. 1โ20. doi:10.1145/3544548.3581441. Bai, G., Chai, Z., Ling, C., Wang, S., Lu, J., Zhang, N., Shi, T., Yu, Z., Zhu, M., Zhang, Y., Song, X., Yang, C., Cheng, Y., Zhao, L., 2024. Beyond efficiency: A systematic survey of resource-efficient large language models. URL: https://arxiv.org/abs/2401.00625, arXiv:2401.00625. Beattie, S., Colbran, S., 2017. From phoenix wright to atticus finch: Legal simulation games as an aid to self-represented litigants, in: Proceedings of the 26th International Conference on World Wide Web Companion, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE. p. 425โ428. doi:10.1145/3041021.3054171. Bonetti, F., Bucchiarone, A., Wanick, V., 2025. Using llms to adapt serious games with educators in the loop, in: Schรถnbohm, A., Bellotti, F., Bucchiarone, A., de Rosa, F., Ninaus, M., Wang, A., Wanick, V., Dondio, P. (Eds.), Games and Learning Alliance, Springer Nature Switzerland, Cham. p. 68โ77. doi:10.1007/978-3-031-78269-5_7. Chen, J., Wang, X., Xu, R., Yuan, S., Zhang, Y., Shi, W., Xie, J., Li, S., Yang, R., Zhu, T., Chen, A., Li, N., Chen, L., Hu, C., Wu, S., Ren, S., Fu, Z., Xiao, Y., 2024. From persona to personalization: A survey on role-playing language agents. Transactions on Machine Learning Research URL: https://openreview.net/forum?id=xrO70E8UIZ. survey Certification. Chiang, W.L., Zheng, L., Sheng, Y., Angelopoulos, A.N., Li, T., Li, D., Zhu, B., Zhang, H., Jordan, M.I., Gonzalez, J.E., Stoica, I., 2024. Chatbot arena: an open platform for evaluating llms by human preference, in: Proceedings of the 41st International Conference on Machine Learning, JMLR.org. p. 8359โ8388. doi:10.5555/3692070.3692401. Cox, S.R., Ooi, W.T., 2024. Conversational interactions with npcs in llm-driven gaming: Guidelines from a content analysis of player feedback, in: Fรธlstad, A., Araujo, T., Papadopoulos, S., Law, E.L.C., Luger, E., Goodwin, M., Hobert, S., Brandtzaeg, P.B. (Eds.), Chatbot Research and Design, Springer Nature Switzerland, Cham. p. 167โ184. doi:10.1007/978-3-031-54975-5_10. Fabbri, A.R., Kryลciลski, W., McCann, B., Xiong, C., Socher, R., Radev, D., 2021. Summeval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics 9, 391โ409. doi:10.1162/tacl_a_00373. Farrokhi Maleki, M., Zhao, R., 2024. Procedural content generation in games: A survey with insights on emerging llm integration, in: Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, p. 167โ178. doi:10.1609/aiide.v20i1.31877. Gopalakrishnan, K., Hedayatnia, B., Chen, Q., Gottardi, A., Kwatra, S., Venkatesh, A., Gabriel, R., Hakkani-Tur, D., 2023. Topical-chat: Towards knowledge-grounded open-domain conversations. URL: https://arxiv.org/abs/2308.11995, arXiv:2308.11995. R. Rosati et al.: Preprint submitted to ElsevierPage 32 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., Wang, S., Zhang, K., Lin, Z., Zhang, B., Ni, L., Gao, W., Wang, Y., Guo, J., 2026. A survey on llm-as-a-judge. The Innovation , 101253doi:10.1016/j.xinn.2025.101253. Gursesli, M.C., Taveekitworachai, P., Abdullah, F., Dewantoro, M.F., Lanata, A., Guazzini, A., Lรช, V.K., Villars, A., Thawonmas, R., 2023. The chronicles of chatgpt: Generating and evaluating visual novel narratives on climate change through chatgpt, in: Holloway-Attaway, L., Murray, J.T. (Eds.), Interactive Storytelling, Springer Nature Switzerland, Cham. p. 181โ194. doi:10.1007/978-3-031-47658-7_16. Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J., 2021. Measuring massive multitask language understanding. URL: https://arxiv.org/abs/2009.03300, arXiv:2009.03300. Hua, M., Raley, R., 2020. Playing with unicorns: Ai dungeon and citizen nlp. DHQ: Digital Humanities Quarterly 14. Kumaran, V., Rowe, J., Mott, B., Lester, J., 2023. Scenecraft: Automating interactive narrative scene generation in digital games with large language models, in: Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, p. 86โ96. doi:10.1609/aiide.v19i1.27504. Lang, J., Guo, Z., Huang, S., 2024. A comprehensive study on quantization techniques for large language models, in: 2024 4th International Conference on Artificial Intelligence, Robotics, and Communication (ICAIRC), p. 224โ231. doi:10.1109/ICAIRC64177.2024.10899941. Lanzi, P.L., Loiacono, D., 2023. Chatgpt and other large language models as evolutionary engines for online interactive collaborative game design, in: Proceedings of the Genetic and Evolutionary Computation Conference, Association for Computing Machinery, New York, NY, USA. p. 1383โ1390. doi:10.1145/3583131.3590351. Liu, Y., Tao, S., Zhao, X., Zhu, M., Ma, W., Zhu, J., Su, C., Hou, Y., Zhang, M., Zhang, M., Ma, H., Zhang, L., Yang, H., Jiang, Y., 2024. CoachLM: Automatic Instruction Revisions Improve the Data Quality in LLM Instruction Tuning , in: 2024 IEEE 40th International Conference on Data Engineering (ICDE), IEEE Computer Society, Los Alamitos, CA, USA. p. 5184โ5197. doi:10.1109/ICDE60146.2024.00390. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.F., Leike, J., Lowe, R., 2022. Training language models to follow instructions with human feedback, in: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (Eds.), Advances in Neural Information Processing Systems, Curran Associates, Inc.. p. 27730โ27744. Park, D., Kim, M., Choi, B., Kim, J., Lee, K., Lee, J., Park, I., Lee, B.U., Hwang, J., Ahn, J., Mahabaleshwarkar, A.S., Kartal, B., Biswas, P., Suhara, Y., Lee, K., Cho, J., 2025. Orak: A foundational benchmark for training and evaluating llm agents on diverse video games. URL: https://arxiv.org/abs/2506.03610, arXiv:2506.03610. Pereira, J., Assumpcao, A., Lotufo, R., 2024. Check-eval: A checklist-based approach for evaluating text quality. URL: https://arxiv.org/ abs/2407.14467, arXiv:2407.14467. Phan, L., Mazeika, M., Zou, A., Hendrycks, D., 2025. Textquests: How good are llms at text-based video games? URL: https://arxiv.org/ abs/2507.23701, arXiv:2507.23701. Salinas, A., Morstatter, F., 2024. The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance, in: Ku, L.W., Martins, A., Srikumar, V. (Eds.), Findings of the Association for Computational Linguistics: ACL 2024, Association for Computational Linguistics, Bangkok, Thailand. p. 4629โ4651. doi:10.18653/v1/2024.findings-acl.275. Savanna Developments, 2024. Verbal verdict. https://store.steampowered.com/app/2778780/Verbal_Verdict/. Accessed: 11 February 2026. Sclar, M., Choi, Y., Tsvetkov, Y., Suhr, A., 2024. Quantifying language modelsโ sensitivity to spurious features in prompt design or: how i learned to start worrying about prompt formatting, in: 12th International Conference on Learning Representations, ICLR 2024, p. 1โ29. R. Rosati et al.: Preprint submitted to ElsevierPage 33 of 34 RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents Shi, B., tse Huang, J., Luo, L., Zong, T., Yi, H., Wang, Y., Hu, S., Zhang, X., Yao, Z., 2026. Fairgamer: Evaluating social biases in llm-based video game npcs. URL: https://arxiv.org/abs/2508.17825, arXiv:2508.17825. Song, Y., Wu, K., Ding, J., 2024. Developing an immersive game-based learning platform with generative artificial intelligence and virtual reality technologies โ โlearningversevrโ. Computers & Education: X Reality 4, 100069. doi:10.1016/j.cexr.2024.100069. Sweetser, P., 2024. Large language models and video games: A preliminary scoping review, in: Proceedings of the 6th ACM Conference on Conversational User Interfaces, Association for Computing Machinery, New York, NY, USA. p. 1โ8. doi:10.1145/3640794.3665582. Tรณdovรก, T., 2025. A quest for information: Enhancing game-based learning with llm-driven npcs, in: Proceedings of CESCG 2025: The 29th Central European Seminar on Computer Graphics, p. 1โ8. Trukhin, R., Isaksson, J.F., Palosaari Eladhari, M., 2026. Llm-powered npcs, in: Reyes, M.C., Nack, F. (Eds.), Interactive Storytelling, Springer Nature Switzerland, Cham. p. 171โ189. doi:10.1007/978-3-032-12405-0_10. Wang, Y., Chen, J., Xiao, H., 2026. Role-playing agents driven by large language models: Current status, challenges, and future trends. URL: https://arxiv.org/abs/2601.10122, arXiv:2601.10122. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E.H., Le, Q.V., Zhou, D., 2022. Chain-of-thought prompting elicits reasoning in large language models, in: Proceedings of the 36th International Conference on Neural Information Processing Systems, Curran Associates Inc., Red Hook, NY, USA. p. 24824โ24837. doi:10.5555/3600270.3602070. Wu, Z., Chen, Z., Zhu, D., Mousas, C., Kao, D., 2025. A systematic review of generative ai on game character creation: Applications, challenges, and future trends. IEEE Transactions on Games , 1โ15doi:10.1109/TG.2025.3564869. Yang, T., Zhu, Y., Quan, X., Liu, C., Wang, Q., 2025. Psyplay: Personality-infused role-playing conversational agents. URL: https://arxiv. org/abs/2502.03821, arXiv:2502.03821. Zargham, N., Tonini, L., Alexandrovsky, D., Ruthven, E.G., Friehs, M.A., Dratzidis, L.T., Dรคnekas, B., Bikas, I., Nacke, L.E., Zebel, S., Malaka, R., 2026. Dialogs with genai npcs: Exploring player interactions with speech agents in a vr game. International Journal of HumanโComputer Interaction , 1โ29doi:10.1080/10447318.2026.2620647. Zhao, Y., Pan, J., Dong, Y., Dong, T., Wang, G., Ying, F., Shen, Q., Cao, J., 2024. Language urban odyssey: A serious game for enhancing second language acquisition through large language models, in: Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery, New York, NY, USA. p. 1โ7. doi:10.1145/3613905.3651112. Zhu, Q., Zhao, R., Liang, B., Du, J., Gui, L., He, Y., 2025. Player*: Enhancing llm-based multi-agent communication and interaction in murder mystery games. URL: https://arxiv.org/abs/2404.17662, arXiv:2404.17662. Zhu, X., Chen, Y., Tian, H., Tao, C., Su, W., Yang, C., Huang, G., Li, B., Lu, L., Wang, X., Qiao, Y., Zhang, Z., Dai, J., 2023. Ghost in the minecraft: Generally capable agents for open-world environments via large language models with text-based knowledge and memory. URL: https://arxiv.org/abs/2305.17144, arXiv:2305.17144. รzkaya, S., Berrezueta-Guzman, S., Wagner, S., 2025. How llms are shaping the future of virtual reality. IEEE Access 13, 193335โ193355. doi:10.1109/ACCESS.2025.3631594. R. Rosati et al.: Preprint submitted to ElsevierPage 34 of 34