Paper deep dive
Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning
Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, Vered Shwartz
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/20/2026, 11:11:42 PM
Summary
This study compares the persuasive and educational effectiveness of three information delivery formatsâstatic essays, conversational chatbots, and narrative text-based gamesâon sustainability topics (recycling and public transit). Using identical factual content across conditions, the research finds that chatbots consistently outperform other modes in subjective measures like perceived importance and engagement. However, a dissociation exists between subjective experience and objective learning: while game participants reported lower perceived learning, they achieved higher scores on a delayed 24-hour knowledge quiz compared to essay readers. The study highlights that engagement proxies like verbosity correlate more with subjective experience than actual retention.
Entities (9)
Relation Signals (9)
Seyed Hossein Alavi â affiliatedwith â University of British Columbia
confidence 95% · SEYED HOSSEIN ALAVI, University of British Columbia, Canada
Seyed Hossein Alavi â affiliatedwith â Vector Institute for AI
confidence 90% · SEYED HOSSEIN ALAVI, University of British Columbia, Canada and Vector Institute for AI, Canada
Chatbot â outperforms â Static Essay
confidence 90% · Across subjective measures, the chatbot condition consistently outperformed the other modes
Chatbot â outperforms â Text-based Game
confidence 90% · Across subjective measures, the chatbot condition consistently outperformed the other modes
Text-based Game â yieldshigherdelayedretentionthan â Static Essay
confidence 90% · participants in the text-based game condition reported learning less than those reading essays, yet achieved higher scores on a delayed (24-hour) knowledge quiz.
Text-based Game â yieldslowerperceivedlearningthan â Static Essay
confidence 90% · participants in the text-based game condition reported learning less than those reading essays
PersuLab â supports â Static Essay
confidence 80% · We introduce PersuLab, a system designed to support both non-interactive and interactive information delivery
PersuLab â â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Interactive systems such as chatbots and games are increasingly used to persuade and educate on sustainability-related topics, yet it remains unclear how different delivery formats shape learning and persuasive outcomes when content is held constant. Grounding on identical arguments and factual content across conditions, we present a controlled user study comparing three modes of information delivery: static essays, conversational chatbots, and narrative text-based games. Across subjective measures, the chatbot condition consistently outperformed the other modes and increased perceived importance of the topic. However, perceived learning did not reliably align with objective outcomes: participants in the text-based game condition reported learning less than those reading essays, yet achieved higher scores on a delayed (24-hour) knowledge quiz. Additional exploratory analyses further suggest that common engagement proxies, such as verbosity and interaction length, are more closely related to subjective experience than to actual learning. These findings highlight a dissociation between how persuasive experiences feel and what participants retain, and point to important design trade-offs between interactivity, realism, and learning in persuasive systems and serious games.
Tags
Links
- Source: https://arxiv.org/abs/2602.17905v2
- Canonical: https://arxiv.org/abs/2602.17905v2
Trouble viewing inline? Open PDF directly â
Full Text
129,214 characters extracted from source content.
Expand or collapse full text
Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning SEYED HOSSEIN ALAVI, University of British Columbia, Canada and Vector Institute for AI, Canada ZINING WANG, University of British Columbia, Canada and Vector Institute for AI, Canada SHRUTHI CHOCKKALINGAM, University of British Columbia, Canada RAYMOND T. NG, University of British Columbia, Canada VERED SHWARTZ, University of British Columbia, Canada and Vector Institute for AI, Canada Interactive systems such as chatbots and games are increasingly used to persuade and educate on sustainability-related topics, yet it remains unclear how different delivery formats shape learning and persuasive outcomes when content is held constant. Grounding on identical arguments and factual content across conditions, we present a controlled user study comparing three modes of information delivery: static essays, conversational chatbots, and narrative text-based games. Across subjective measures, the chatbot condition consistently outperformed the other modes and increased perceived importance of the topic. However, perceived learning did not reliably align with objective outcomes: participants in the text-based game condition reported learning less than those reading essays, yet achieved higher scores on a delayed (24-hour) knowledge quiz. Additional exploratory analyses further suggest that common engagement proxies, such as verbosity and interaction length, are more closely related to subjective experience than to actual learning. These findings highlight a dissociation between how persuasive experiences feel and what participants retain, and point to important design trade-offs between interactivity, realism, and learning in persuasive systems and serious games. Additional Key Words and Phrases: Persuasive Technology; Knowledge Retention; Games for Change; Interactive Systems; Conversa- tional Agents; Game-Based Interaction; Sustainability 1 Introduction Persuasive and educational technologies have been widely explored for environmental sustainability, often aiming to raise awareness, shift attitudes, or encourage pro-environmental behavior [1,14]. Recent advances in large language models (LLMs) have further accelerated this trend by reducing the barriers for developing conversational agents and interactive games. Enabling dynamically generation of persuasive content at scale makes these interactive systems alternatives preferable to traditional static formats such as essays or informational webpages. Despite this growing adoption, there remains limited empirical understanding of how different modes of delivery shape persuasive experience and learning outcomes when the underlying informational content is held constant. Many evaluations of persuasive and educational systems rely heavily on self-reported outcomes such as engagement, enjoyment, perceived learning, or attitude change [19,67]. While these measures capture important aspects of user experience, prior work has shown that subjective perceptions of learning do not always align with objectively measured learning or longer-term knowledge retention [8,51]. This misalignment is particularly relevant for interactive systems, where engagement and interactivity may shape how effective an experience feels without necessarily improving what users ultimately retain. This question is especially relevant for designers of games-for-change and serious games, which are widely used for sustainability education and persuasion [24,44,48]. While interactive narratives can increase engagement, they may Authorsâ Contact Information: Seyed Hossein Alavi, salavis@cs.ubc.ca, University of British Columbia, Canada and Vector Institute for AI, Canada; Zining Wang, University of British Columbia, Canada and Vector Institute for AI, Canada; Shruthi Chockkalingam, University of British Columbia, Canada; Raymond T. Ng, University of British Columbia, Canada; Vered Shwartz, University of British Columbia, Canada and Vector Institute for AI, Canada. 1 arXiv:2602.17905v2 [cs.HC] 24 Feb 2026 2Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz also introduce trade-offs in realism, trust, or cognitive load that affect persuasive impact [9,10,17]. In contrast, static formats such as essays may appear clearer or more credible but encourage passive consumption. Controlled comparisons are therefore needed to disentangle the effects of content from interaction structure. In this work, we address this gap through a controlled comparison of three common information delivery modes: a static essay, a conversational chatbot, and a narrative-driven text-based game. Crucially, all three conditions were grounded in the same set of persuasive arguments and numerical facts, allowing us to examine how delivery format alone shapes subjective experience, perceived attitude change, and objective knowledge retention. We focus on two sustainability-related topics â recycling and public transit â chosen for their real-world relevance and suitability for persuasion-oriented interventions. We conducted a between-subjects user study in which participants experienced exactly one delivery mode and one topic. We collected pre-study measures, post-study subjective ratings and perceived change measures, and administered a delayed (24-hour) objective knowledge quiz. We also analyze the interaction logs for the interactive conditions. This combination enables us to examine not only differences across modes, but also the relationship between interaction behavior, subjective impressions, and learning outcomes. Our findings reveal several important tensions. Conversational interaction consistently improved subjective ex- perience and increased perceived importance of the topic, but perceived learning does not reliably predict objective retention: participants in the text-based game condition reported learning less, yet retained more factual information than those who read an essay. Exploratory analyses of interaction logs further suggest that common engagement proxies â such as verbosity, turn count, or session duration â are more strongly associated with subjective impressions than with actual learning. Together, these results highlight a dissociation between how persuasive experiences feel and what participants ultimately retain. By clarifying the trade-offs between interactivity, realism, subjective experience, and learning, this work aims to inform the design and evaluation of future persuasive chatbots, narrative games, and games-for-change. 2 Background Given the abundance of research on interactive technologies for educational and persuasive purposes, we primarily focus our discussion on technologies developed for the topic of sustainability and climate change. We describe various types of systems and what makes them effective (§2.1), discuss ways to measure learning in such systems (§2.2), and finally more broadly discuss persuasive technologies (§2.3). 2.1 Educational Technology for Sustainability-Related Topics Environmental sustainability is a prominent topic in HCI research and is frequently addressed through persuasive approaches [14]. Persuasive techniques have been applied in educational technologies, where they are used to help learners form, alter, or reinforce attitudes and behaviors that support learning [45]. Under this framing, two interactive modalities have emerged as particularly effective for promoting sustainability-related learning and persuasion. Conversation is especially effective at persuading people [42,43]. Accordingly, conversational agents (or chatbots) have emerged as a simple, widely used interactive modality, which is especially popular in recent years with the advent of LLMs. Prior work has examined the role of chatbots in persuasion [29] as well as in educational contexts [31,50]. LLM-based chatbot engaging in conversation might be substantially more persuasive than single-message interactions. They can promote more active engagement, quickly address user concerns, and tailor arguments to Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning3 individual preferences [40,59]. When designed responsibly, chatbots can serve as enablers of positive individual and social change [18, 27]. An alternative interactive technology is serious games: games that are designed for educational rather than en- tertainment purposes [47]. The interactivity of serious games makes them one of the most effective strategies for providing educational content in an engaging manner [44]. Serious games are thus a popular approach to teach people about sustainability [16,22,48,63]. Numerous empirical studies demonstrate that games and gamification are effective in engaging with climate change in educational contexts [24]. For example, digital games has been demonstrated to improve studentsâ climate literacy as well as their behaviors towards energy related topics [30, 35, 48, 49]. In particular, narrative is a common component of climate change games [22]. Serious games that presented a story in a virtual location based on a real city increased playersâ concern about climate issues[7]. This is also our focus in this paper; for the game interactive modality, we specifically look at a text-based narrative games. Despite this body of work, relatively little research has systematically compared how different interactive systems â such as games and conversational agents â differ from one another in both learning and persuasion outcomes, which we do in this work. 2.2 Measuring Learning in Interactive Systems Prior work on interactive learning system uses both subjective self-reports of perceived learning and objective performance-based assessments, reflecting different assumptions about what constitutes learning and how it should be evaluated. Some studies operationalize learning outcomes as perceived learning effects, capturing the extent to which learners experience an interaction as educational [67]. In contrast, others adopt objective evaluation strategies, assessing learning through changes in knowledge or performance [49]. The choice of evaluation metrics can lead to divergent conclusions about system effectiveness. For example, Nussbaum et al.[49] developed a serious game exploring water-level decline and evaluated learning using pre-study, post-study, and 11-day delayed post-study tests to measure the immediate and long-term knowledge retention. Participants who played the game demonstrated effective long-term knowledge retention, as evidenced by maintained marginal gains on both the post-study and delayed post-study test compared to a control group that engaged in reading a static website. While our work reaffirms this conclusion, it also shows that games underperform in perceived learning metrics. Our work aligns with prior research showing that subjective perceptions of learning may not align with objectively measured learning outcomes in interactive educational systems. For instance, Persky et al.[51] found no correlation between perceived and actual knowledge gains among 277 college and pharmacy students. Similarly, a study comparing perceived and actual academic performance in an online learning environment with 382 participants found no significant relationship between the two measures [8]. This misalignment highlights the need to evaluate interactive learning systems using both subjective and objective measures to better capture true learning outcomes. 2.3 Persuasive Technology for Sustainability Given that our study focuses on sustainability issues, it ties together both learning about the topic and belief change. While we primarily focus on learning outcomes, we also review relevant literature on persuasive technology. De Vries et al.[19] defines persuasive technology as âtechnologies aimed at changing peopleâs attitudes or behaviors through persuasion and social influence, but not through coercion or deceptionâ. 4Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Classic persuasive technology interventions for sustainability typically aim to raise awareness, personalize inter- ventions, and target specific behavior changes [1]. Such technologies have been implemented across a wide range of media, including desktop applications [2], mobile applications [39], and serious games [13,25]. Among these, serious games have gained particular popularity within sustainability-focused persuasive technology research, promoting environmental sustainability goals such as energy conservation and waste reuse [1]. Recent advances in LLMs have shifted some of the focus towards chatbots. LLMs can engage in sophisticated interactive dialogue, making them powerful tools for shaping attitudes, preferences, and behaviors [20,29,40,55]. In particular, their ability to deliver adaptive and personalized persuasive content at unprecedented scale makes LLM- based systems hyper-persuasive technologies [37]. LLMsâ persuasive power has been demonstrated across multiple domains, including consumer marketing [46], healthcare [5,34], politics [11,23,28,59], life-style decisions [18], and pro-environmental appeals [40]. Multiple studies found LLM-generated content was perceived as at least as persuasive, and often more persuasive than human-written content [11,18,34,40]. Other studies have investigated factors that influence LLM persuasiveness, including prompt design and message characteristics [5,29,54,59]. In particular, evidence-based persuasion â where models structure dialogue around fact-checkable, argument-relevant information â has been shown to outperform approaches that rely primarily on rhetorical or emotional language [29,59]. LLM-based serious games are capable of all of the above and are also highly conducive to personalization such as personalized NPCs [3,21] and narratives [4,58], making them effective in persuasion. LLMsâ persuasiveness is typically evaluated against human persuaders [11,56,57]. LLM-based systems with particular mechanisms for increasing persuasiveness are often compared against other LLM-based baselines [27]. Finally, some studies compare interactive LLM-based systems against less interactive settings such as reading static LLM-generated passages [5,11]. However, there remains limited understanding of how LLM persuasion outcomes differ across interaction modalities. Our work addresses this gap by systematically comparing persuasive effects across three modalities that vary in their level and form of interactivity: essays, chatbots, and serious games. 3 Study Design This study examines how information delivery modality influences knowledge retention and persuasive effectiveness in sustainability education. We focus on two topics pertaining to sustainability (§3.1). We conducted a between-subjects user study to compare the knowledge retention and persuasive effectiveness of three information delivery modes: essay, chatbot, and a text-based game (§3.2). While the three conditions differed in their interaction style and delivery format, all were implemented within a shared system (PersuLab) designed to ensure consistent content exposure and comprehensive interaction logging (§3.3). Each participant (§3.4) experienced exactly one mode and one topic. They were asked to complete a pre-study questionnaire before the assigned experience and a post-study questionnaire and delayed knowledge assessment afterwards, as we describe in the next section. 3.1 Topics, Arguments, and Facts Participants were assigned to one of two topics: recycling and public transit. We selected these topics because they represent common, personally actionable sustainability behaviors that are frequently targeted by educational and persuasive technologies [62]. For each topic, we used a fixed set of five arguments in favor of taking actions towards sustainability, each paired with a supporting factual statement (Table 1). Topic assignment was balanced across experimental conditions. To ensure internal validity and enable fair comparison across delivery modes, we strictly Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning5 TopicArgument (without the premise)Supporting fact used in the study Recycling (1) Environmental Preservation: Protecting Natural Resources and EcosystemsRecycling one ton of paper saves approximately 17 trees. (2) Energy Conservation and Lower Carbon EmissionsRecycling aluminum saves up to 95% of the energy compared to pro- ducing new aluminum. (3) Economic Benefits: Cost Savings and Job CreationRecycling creates about 10 jobs for every landfill job. (4) Reducing Pollution: Cleaner Air, Water, and SoilUnrecycled plastic breaks into microplastics that harm over 700 marine species and enter the food chain. (5) Setting a Positive Example and Meeting Social Responsibility GoalsCompanies with recycling programs report up to 20% higher employee engagement and retention. Public transit (1) Environmental Benefits: Reduced Carbon FootprintSwitching from personal cars to public transit can reduce an individualâs transportation emissions by up to 50%. (2) Cost Savings for Individuals and BusinessesOn average, commuters can save $10,000+ per year by using public transit instead of driving. (3) Reduced Traffic CongestionA 10% shift from personal cars to public transit can cut commute times by up to 40%. (4) Health and Safety ImprovementsTraveling by public transit is approximately 10Ăsafer than driving a personal car. (5) Economic and Social Benefits for CommunitiesEvery $1 invested in public transit generates approximately $4 in com- munity benefits. Table 1. Fixed set of persuasive arguments and supporting facts used in the study for each topic. All experimental modes used the same five argumentâfact pairs for each topic, ensuring that participants across experimental modes had access to the same information. Each concrete argument is associated with a premise (e.g., âRecycling uses less energy than making new materialsâ) which we omit here for brevity. See Appendix B for the complete arguments. controlled informational exposure: in all conditions, the system ensured that all five predefined facts were presented before the interaction could be completed (see Appendix B for more details). 3.2 Experiment Conditions The study employed a single-factor between-subjects design with three delivery mode corresponding to varying degrees of interactivity: essay, chat, and a text-based game. 3.2.1 Essay. Given the fixed set of arguments and facts, we asked GPT-4.1 to generate persuasive essays solely based on the facts and arguments (see Appendix C for the prompts). To ensure that all arguments and facts are covered in the essay, we inputted the essay and the arguments into the LLM to verify the coverage. We also randomly sampled essays and manually verified the coverage of arguments. Since the essay mode is not interactive, we opted for speeding up the computation time for generation and argument coverage verification by generating the essays in advance. However, to ensure the generality of the findings and for the sake of fair comparison to the interactive modes, rather than generating one essay per topic, we generated twenty essays and randomly sampled an essay for each participant. 3.2.2 Chatbot. In the chatbot condition, participants engaged in an interactive, free-form conversation with an LLM about the assigned sustainability topic. The chatbot presented the same fixed set of arguments and supporting facts used in the other conditions, but delivered them conversationally in response to participant inputs (see Table 2 for a chat example). Unlike the essay condition, participants could actively steer the interaction by asking questions, requesting clarifica- tion, or reacting to the information presented, allowing them to control the pacing and order in which content was encountered. To ensure comparability across conditions and enable later assessment of knowledge retention, the interaction was structured so that all predefined arguments and factual statements were eventually presented during the session. 6Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Participants could freely interact with the chatbot, and the conversation continued until all target facts had been covered or the participant chose to end the interaction. The interface indicated when all required content had been presented and allowed participants to conclude the session at that point. Appendix D provides details about the content coverage verification mechanism. 3.2.3 Text-based Game. In the text-based game condition, participants interacted with a narrative-driven, text-based game centered on the assigned sustainability topic (see Table 3 for a game example). The game embedded the same fixed set of persuasive arguments and supporting facts used in the essay and chatbot conditions, but presented them within an interactive story world. Participants assumed the role of a protagonist and progressed through the game by reading narrative descriptions and entering free-form text responses at each turn. The system responded by advancing the story, introducing new situations, characters, or dilemmas, while weaving in the predefined arguments and factual information in a contextualized manner. As in the chatbot condition, participants could control the pacing and flow of the experience through their inputs. Similarly to the chatbot, the interface indicated when all required content had been covered and only then allowed participants to conclude the session, though participants could continue interacting until the ending of the game if they wished. Full implementation details of the text-based game, including plot generation and moderation prompts, are provided in Appendix E. Table 2. Example conversation between player and chatbot. Bolded lines highlight factual information embedded into the chatbotâs response. Chatbot: Have you ever thought about how using public transit could make a real differenceâfor you and your community? Making the switch has more benefits than most people realize. Would you be open to hearing some of them? Player: Not really, I would like to hear something new from you. Chatbot: Absolutely, letâs start with something practicalâmoney. Did you know the average commuter saves over $10,000 a year just by choosing public transit instead of driving? Thatâs a lot of extra cash back in your pocket for things youâd actually enjoy. What would you do with those savings? Player: Wow, that is a really shocking number, I think I should buy myself a new laptop and a new phone! Chatbot:Thatâs an awesome way to use the extra money! Beyond the financial perks, thereâs another upside: taking public transit actually cuts your transport emissions by 50%. Itâs a big, simple step toward a cleaner environmentâand itâs as easy as hopping on a bus or train. Would you feel good knowing your daily commute could help clear up the air? [The chat continues until all the facts are presented and the user chooses to end the session.] 3.3 PersuLab We introduce PersuLab, a system designed to support both non-interactive and interactive information delivery while ensuring consistent exposure to a fixed set of persuasive arguments and factual statements. 1 Figure 1 shows the user interfaces across the three conditions. Upon entering the study, participants were assigned a topic and condition and were redirected to the corresponding interface. Participants were unaware of alternative conditions or topics. Each participant was assigned a unique user ID, enabling us to link between interaction logs and questionnaire responses while preserving anonymity. Across all conditions, an End Session button was visible but activated only after condition-specific completion criteria were satisfied. Figure 2 presents an overview of PersuLabâs system architecture and content generation workflows across the three experimental conditions. In the essay condition, content delivery follows a single-pass pipeline. Because factual coverage is guaranteed during pre-generation (see Section 3.2.1), no runtime fact tracking or verification is required. 1 The code is available at: https://github.com/salavi/persulab Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning7 Table 3. Example conversation between player and game system. Bolded lines highlight factual information seamlessly embedded into the storyâs progression. At each turn, the player may choose one of three predefined actions (as in the first turn) or propose a custom response (as in the second turn). Game: Rain streaks across Sandbourne City Hall as you clutch the heavy âSandbourne: 2040â folder outside the mayorâs office, your heartbeat echoing the drip outside. In the hallway, murmurs about record smog, traffic, and wet footprints braid into the cityâs restless story. Your phone buzzes: six days until a decisive transportation vote. Data alone wonât save Sandbourneâyou must turn numbers into a narrative. The door opens. âAlex? Ready to shape the future?â What do you do? (1) Enter confidently, ready to introduce bold ideas right away. (2) Pause for a moment, listening to the worries in the hallway for clues about the public mood. (3) Flip open your folder to review your talking points, wanting to be extra prepared before you go in. (4) Or suggest your own move. Player: 3 Game: You flip open your thick folder, eyes scanning cost projections and air quality charts. A bold, highlighted line catches your gaze: âSwitching to transit can cut your transport emissions by 50%.â Maya Alvarez, bright in her neon vest, rallies volunteers beneath a hazy skyline. âWe need stories, not just science, Alex. Help us make people believe clean air is possible.â Nearby, Officer Don Briggs watches kids cough in the heat. Jared Liu leans on his car, smirking, âSo, you turning us all into bus riders?â Between hope and skepticism, the city waits for your answer. What do you do? (1) Join Maya, helping plan her âClean Air Nowâ event and share the 50% emissions cut with the crowd. (2) Engage Jared, linking pollutionâand his own coughâto the transit facts you just reviewed. (3) Ask Officer Briggs for a ride-along to see firsthand how smog affects the community. (4) Or suggest your own move. Player: I help the volunteers to gather trash for a bit. Game: Gloves on, you join Maya picking up litter by the swings. Kids look grateful; Maya beams. âActions speak, Alex,â she says. Nearby, Jared pauses mid-chuckle, watching you. Officer Briggs nods approval. As Maya helps you up, she whispers, âHelp me get folks to the Clean Air rally tonightâand let them know transit could cut our emissions in half.â What do you do? ... [The game continues with options and interactions until all the facts are presented and the user chooses to end the session.] Essay ChatGame SystemUser End session User System Topic:1 UserId:x Fig. 1. The user interface of PersuLab, illustrating the three system variants corresponding to different assignment conditions. Each interface displays a unique User ID assigned to the participant at the top. An âEnd Sessionâ button becomes active once all facts have been presented in the chat and game conditions, and after three minutes of reading in the essay condition. The minimum exposure period of 60 seconds was introduced to ensure participants had sufficient time to read the essay before proceeding. In contrast, the chatbot and text-based game conditions follow an interactive generation loop. Participant inputs are routed to a moderator module that constructs prompts and invokes the language model to generate the next system response. After each generated turn, an automated fact-checking module evaluates which of the remaining target facts have been covered and updates the systemâs internal record of uncovered content. This updated information is 8Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Fig. 2. System architecture and interaction workflow for PersuLab across experimental modes. The essay condition (orange) follows a single-pass content delivery pipeline using pre-generated essays that already cover all target facts. The chat and game conditions (blue) operate through an interactive loop that includes turn generation, fact coverage tracking, and memory updates. Black paths denote shared system components and data flow common to all modes (e.g., session control and logging). All interactions are logged for subsequent analysis. 0102030 18â24 25â34 35â44 45â54 28 11 3 1 Number of participants (a) Age-group distribution (n=43). ChatEssayGameTotal Recycling87722 Public transit67821 Total14141543 (b) Assignment counts by mode and topic Fig. 3. Participant demographics and assignment distribution. incorporated into subsequent prompts to guide future generations and ensure complete coverage of all predefined facts. In our experiment, after all predefined facts had been presented, participants were free to end the session or continue interacting until either the interaction naturally concluded or the maximum interaction time of 25 minutes was reached. Both interactive conditions maintain a shared session state that includes interaction history, coverage status of target facts, and timing metadata. All interactions across conditions are logged for analysis, including turn-level content, timestamps, and fact coverage information. Detailed prompt structures, generation parameters, and fact-checking mechanisms for each condition are provided in Appendices C, D, and E, respectively. 3.4 Participants We recruited 45 participants through university advertisements and word of mouth. Of these, 43 participants were included in the analysis; one participant did not complete all study steps, and one participated in a dry run. Most participants were between 18â34 years old (Fig. 3a). Participants were approximately evenly distributed across the three experimental conditions (essay, chat, and game) and across the two topics (recycling and public transit) as shown in Fig. 3b. 4 Evaluation Measures We evaluate the knowledge retention and persuasive effectiveness of the three experiment modes through a mixed- methods user study. We compare participantsâ attitudes towards the assigned topic collected before the study (§4.1) Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning9 with their subjective experience and perceived attitude changes after the study (§4.2). To measure objective knowledge retention, participants had to complete a delayed knowledge assessment task (§4.3). Finally, we analyze the interaction logs to probe potential relationships between interaction dynamics and subjective and objective learning outcomes (§4.4). 4.1 Pre-Study Questionnaire Prior to the experimental session, participants completed a pre-study questionnaire capturing baseline attitudes and behaviors relevant to the assigned topic. These measures were used to characterize participantsâ initial positions. We collected baseline measures of importance, behavioral intention, and epistemic confidence, which are detailed below (see Appendix A.1 for the complete pre-study questionnaire). Participants were asked to rate the importance of the topic to them on a 5-point Likert scale. For behavioral intention, we inquired about existing habits pertaining to the topic, such as frequency of recycling or commuting by public transport. Responses were recoded into ordinal scales (1â5) to reflect increasing levels of baseline engagement (e.g., never recycle to always recycle). These measures provide descriptive context about participantsâ pre-study attitudes and behaviors. We also collected a pre-study measure of participantsâ self-perceived confidence in their knowledge of the topic using a 5-point Likert scale. We treat epistemic confidence as a descriptive indicator of participantsâ perceived familiarity with the topic, which provides context for interpreting objective knowledge retention and subjective learning outcomes. Additional demographic and contextual questions (e.g., age, access to recycling services, access to a personal car) were collected for descriptive purposes but were not used in the primary analyses. 4.2 Post-Study Questionnaire Immediately after completing the intervention, participants rated their subjective experience using a set of 5-point Likert-scale items (1 = strongly disagree, 5 = strongly agree). We included questions directly targeting learning and persuasion outcomes as well as questions about subjective experience common in user studies, as we detail below. See Appendix A.2 for the complete set of questions. Perceived Change and Persuasion. To assess persuasive impact, participants reported perceived changes in their attitudes relative to before the session. For each topic, participants indicated whether the interaction led to a change in importance, behavioral intention, and belief in effectiveness, using ordered categorical responses (Less, Same, More, Not sure). Participants additionally rated the convincingness of the arguments presented using 5-point Likert scales. Finally, we prompted participants to describe their qualitative reflections in free-text format. Subjective Experience. We assessed participantsâ perceived ease of following the provided content, engagement, enjoyment, trust in the information, self-reported learning, satisfaction, convincingness of arguments, in- creased motivation to act, overall influence on thinking, as well as their willingness to recommend others and re-encounter in future similar experiences using a set of 5-point Likert-scale items. 4.3 Delayed Objective Knowledge Assessment To measure objective knowledge retention, we reached out to participants 24 hours after the experiment and asked them to complete a multiple-choice knowledge quiz based on the facts covered during the experiment. The 24-hour 10Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz delay is motivated by prior research in psychology and HCI that suggests that an overnight consolidation period facilitates measuring knowledge retention rather than immediate knowledge recall [6, 15, 33, 53]. For each topic, the quiz included five content-covered questions and two control questions referencing information not provided in the experience. Control questions were included to discourage guessing or external lookup; selecting âI have not seen this information beforeâ was treated as the correct response for these items. Participantsâ knowledge scores were computed only as the number of correct responses to the five content-covered questions. Confidence ratings were collected after each question but were not analyzed in this work. The full set of knowledge quiz items and correct answers for each topic is reported in Appendix A.3. 4.4 Interaction Log Metrics (Exploratory) For the interactive conditions (chat and game), we additionally recorded participantsâ interaction logs. These included the number of user/system turns, total and mean user/system words per turn, user-to-system word ratios, session duration, and mean user reaction time (time between received system message and userâs next response). Because these measures were not pre-registered and involved multiple comparisons, analyses of interaction logs is treated as exploratory. We use them to probe potential relationships between interaction dynamics and subjective experience or persuasive outcomes, rather than to test confirmatory hypotheses. Accordingly, findings from these analyses are reported descriptively and interpreted with appropriate caution. 5 Analysis Methods Our analyses are designed to examine the following aspects. First, we look at differences between experimental modes in subjective experience, answering the question âwhich experimental mode do participants prefer?â (§5.1). Second, we examine the perceived changes in persuasive constructs (e.g., importance, behavioral intention, and belief in effectiveness), essentially answering the question âwhich experimental mode do participants perceive as more persuasive?â (§5.2). We then report the results of the objective knowledge retention test, answering the question âwhich experimental mode is the most effective for teaching participants about the topic?â (§5.3). Finally, we perform an exploratory analysis of the relationships between properties of the participant-system interactions (e.g., user word counts) and persuasive outcomes (§5.4). Throughout the analysis, we report descriptive statistics alongside inferential results to support transparent in- terpretation. All significance tests are clearly labeled as confirmatory or exploratory, and findings are interpreted accordingly. 5.1 Subjective Experience To compare post-study subjective experience ratings (e.g. ease of following, engagement, etc.) across the three experi- mental modes (essay, chat, game), we used nonparametric tests appropriate for ordinal Likert-scale data. Specifically, we first applied KruskalâWallis [36] tests to assess overall differences across the three modes. When a significant omnibus effect was observed, we conducted pairwise comparisons using MannâWhitney U [38] tests (essay vs game, essay vs chat, chat vs game). These tests were chosen because Likert-scale responses are ordinal and may violate the assumptions of normality required for parametric tests. All tests were two-sided, and statistical significance was assessed atíŒ= .05. Effect directions and descriptive statistics (means and standard deviations) are reported alongside í-values to aid interpretation. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning11 5.2 Perceived Change and Persuasion To analyze perceived changes in importance, behavioral intention, and belief in effectiveness, we modeled participantsâ self-reported change using ordinal logistic regression [41]. Perceived change responses (Less / Same / More) were encoded as ordinal change scores (â1, 0,+1), with âNot sureâ responses treated as missing. Separate models were fit for each construct (importance, intention, belief ), both within-topic (recycling-only and transit-only) and using a combined-topic dataset in which equivalent constructs were aligned across topics. For each outcome, we first fit an omnibus three-condition model of the form: Change í ⌠Mode í whereChange í is the ordinal change score for participantí,Mode í is the experimental condition (essay as the baseline, with dummy variables for chat and game). To further probe differences between specific modes, we additionally conducted pairwise ordinal logistic regressions between each pair of conditions using the same model specification but restricted to the relevant subset of data. All ordinal logistic regression models (ordered logit / proportional-odds) were estimated via maximum likelihood using a logit link function (BFGS optimizer). Predictor coefficients are reported on the log-odds scale, and statistical significance was assessed using Wald í§-tests [65]. Baseline Robustness Checks. For importance and behavioral intention, conceptually aligned pre-study measures were available. As a robustness check, we re-ran the analyses including the corresponding pre-study scores as covariates to account for potential ceiling effects. The inclusion of baseline measures did not qualitatively change the pattern of results or the relative differences between modes (see Appendix H). 5.3 Objective Knowledge Retention Analyses To compare the participantsâ knowledge scores across experimental modes, we again used KruskalâWallis tests followed by MannâWhitney U tests for pairwise comparisons, due to the ordinal and bounded nature of the score distribution and the modest sample size. Control questions and confidence ratings were excluded from the score computation and analyses. 5.4 Exploratory Interaction Log Analyses For the interactive conditions (chat and game), we explored relationships between interaction log metrics (§4.4) and outcome measures: subjective experiences (§4.2) and objective knowledge scores (§4.3), using Spearman rank-order correlations [68]. Because these analyses involved multiple comparisons and were not pre-registered, they were treated as exploratory. No formal correction for multiple comparisons was applied; instead, results are reported descriptively, with emphasis on effect direction and consistency rather than statistical significance alone. These analyses are intended to generate hypotheses for future work rather than to support confirmatory claims. 6 Findings Unless otherwise noted, we report results pooled across the two topics (recycling and public transit) and compare outcomes by delivery mode. This decision was motivated by the study design: splitting the data by both topic and 12Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz mode would yield small cell sizes (approximately 6â7 participants per condition), limiting interpretability and statistical stability. Additionally, pooling across topics is supported by the observation that topic-specific ratings were highly similar across key measures. In particular, perceived convincingness of arguments was comparable between recycling (mean =3.82) and public transit (mean=3.76), suggesting that topic differences did not meaningfully influence participantsâ evaluations. Accordingly, we focus on mode-level comparisons throughout this section. This section is structured to mirror Sec. 5. We report the participantsâ subjective experience with the various experimental modes (§6.1); the perceived changes in persuasive constructs (§6.2); the results of the objective knowledge test (§6.3); and finally an exploratory analysis of the relationships between properties of the participant-system interactions and persuasive outcomes (§6.4). 6.1 Subjective Experience Fig. 4. Mean subjective outcome ratings (5-point Likert scale) across delivery modes, aggregated across topics (n = 43). Asterisks (*) denote measures with statistically significant differences across conditions. Participantsâ subjective experiences across experimental conditions are summarized in Figure 4. Overall, subjective ratings were relatively high across all three conditions, with mean scores generally above the midpoint of the scale (3) for most measures. Across dimensions, the chatbot condition was consistently rated highest, while the game condition tended to receive lower ratings. The essay condition generally fell between the two interactive conditions, depending on the measure. Ease of following. The essay and chatbot conditions received the highest ease-of-following ratings (both Mean =4.64), whereas the game condition was rated lower (Mean=4.20). A KruskalâWallis test revealed a significant effect of delivery mode on ease of following (í=0.0497). Post-hoc MannâWhitney U tests indicated that participants found the essay condition significantly easier to follow than the game condition (í=0.03). Additionally, the difference between Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning13 the chatbot and game conditions approached statistical significance (í=0.07). It is possible that the additional cognitive demands introduced by narrative structure, role-based interaction, or the verbosity of the system in the game mode contributed to this finding. Self-reported learning. Participants were asked to rate how much they believed they learned during the experiment. Ratings were highest for the chatbot condition (mean=4.29), followed by the essay condition (mean=4.07), and lowest for the game condition (mean=3.30). A KruskalâWallis test indicated a significant effect of delivery mode on self-reported learning (í=0.018). Post-hoc pairwise MannâWhitney tests revealed that participants reported significantly higher learning in the chatbot condition compared to the game condition (í=0.005). The difference between the essay and chatbot conditions approached significance (í= 0.055). Engagement. As expected, given the interactive nature of the conditions, engagement ratings were higher for the chatbot (Mean=4.14) and game (Mean=4.00) compared to the essay condition (Mean=3.64), suggesting that interactivity may enhance participantsâ sense of involvement relative to passive content consumption. Among the three modes, the chatbot received the highest engagement ratings overall, consistent with more extensive user contributions observed in the chatbot condition (§ 6.4) Enjoyment. Interestingly, enjoyment ratings showed a more modest separation across conditions. While the chatbot condition received the highest mean enjoyment rating (Mean=4.00), the game (Mean=3.53) and essay (Mean=3.50) conditions were rated similarly. One possible explanation for the slightly lower enjoyment ratings in the game condition in comparison to the chat is its text-heavy format and lack of visual elements. Trust and Perceived Persuasion. Measures related to perceived persuasion and behavioral impact â including convincingness of arguments, motivation to act, and influence on thinking â showed a consistent numerical advantage for the chatbot condition (means of 4.08, 4.14, and 4.14, respectively). The essay condition followed closely (means of 3.92, 3.71, and 3.93), while the game condition received lower ratings on average (means of 3.38, 3.40, and 3.33). A similar pattern was observed for trust, with participants reporting higher trust in the chatbot condition (mean=3.64) than in the essay (mean= 3.50) and game (mean= 3.27) conditions. Recommendation and Future Re-encounter. Following all the above ratings, participants rated the chatbot condition higher in terms of their willingness to recommend the experience to others and to re-encounter similar experiences in the future (means of 3.79 and 3.93, respectively), compared to the essay condition (means of 3.71 and 3.71). The game condition received the lowest ratings on both measures (means of 3.40 and 3.40). 6.2 Perceived Change and Persuasion We examined perceived changes in participantsâ attitudes toward the target behaviors using post-study self-reports of perceived change across three constructs: importance, intention to act, and belief in effectiveness. These items explicitly asked participants to assess how their attitudes had changed relative to before the session (Less / Same / More), and were analyzed directly as ordinal outcomes. Figure 5 shows the distribution of perceived change responses across experimental conditions; no participant reported a decrease. Descriptively, participants in the chatbot condition more frequently reported increases in perceived importance (79%) and belief in effectiveness (69%) than those in the essay (29% and 57%) and game (14% and 36%) 14Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Fig. 5. Stacked bar chart showing the proportion of participants experiencing perceived change across three constructs: Importance, Intention, and Belief. Each bar represents one of three experiment modes (Essay, Chat, Game), with color-coded segments. Percentages within each segment show the proportion of participants in that category. conditions. In contrast, perceived changes in intention to act were more evenly distributed across conditions (29% chatbot, 29% essay, and 21% game). To formally assess mode effects we fit ordinal logistic regression models predicting perceived change scores from delivery mode (see Sec 5.2). In the combined-topic model, delivery mode had a significant effect on perceived change in importance. Participants in the chatbot condition reported significantly greater increases in importance than those in the essay condition (íœ=2.22,í=0.012), while the game condition did not differ significantly from the essay (íœ=â0.88, í=0.37). Pairwise models confirmed that chatbot interactions led to larger perceived increases in importance than both essay (íœ= 2.22, í= 0.012< 0.05) and game (íœ=â3.09, í= 0.002< 0.05). No significant mode effects were observed for perceived changes in intention or belief in effectiveness. These results indicate that delivery mode primarily influenced participantsâ perceptions of issue importance, rather than directly shaping beliefs about effectiveness or stated behavioral intentions. A baseline robustness analysis incorpo- rating pre-study measures of importance and behavioral intention yielded consistent results (Appendix H), suggesting that these effects are not driven by pre-existing differences between participants. One interpretation of this pattern is that recognizing a topic as effective does not necessarily translate into perceiving the issue as personally important, nor does either guarantee changes in intended behavior. The findings in this section therefore highlight a distinction between belief updating, issue salience, and behavioral intention. While providing information appears sufficient to strengthen beliefs about effectiveness across modes, making an issue feel personally important is more sensitive to interaction format. At the same time, changes in behavioral intention were limited, likely reflecting both the inherent difficulty of motivating action and participantsâ high baseline engagement with the target behaviors. Indeed, pre-study measures show that most participants already reported frequent recycling or regular public transit use (Appendix G), leaving limited room for further increases in stated intention. Overall, these results indicate that conversational interaction can meaningfully increase the perceived importance of sustainability-related behaviors beyond static text, even when informational content is held constant. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning15 Fig. 6. Delayed (24-hour) objective knowledge retention scores by delivery mode (essay, chatbot, text-based game), shown separately for the recycling topic, public transit topic, and the combined dataset. Scores indicate the number of correctly answered content- covered questions (out of five). Diamonds denote mean scores; boxplots show the distribution across participants. 6.3 Objective Knowledge Retention Figure 6 shows participantsâ performance on the delayed (24-hour) objective knowledge quiz, reported separately for each topic and for the combined dataset. Scores reflect the number of correctly answered content-covered questions (out of five). Across both topics and in the combined analysis, participants in the chatbot condition achieved the highest mean knowledge scores, followed by the text-based game, with the essay condition scoring lowest on average. In the combined dataset, mean scores were í= 2.93 for chatbot, í= 2.60 for game, and í= 2.07 for essay. Surprisingly, participants in the game condition consistently outperformed those in the essay condition despite reporting lower self-assessed learning in the post-study questionnaire (§6.1). This pattern was observed for both recycling and public transit topics, as well as in the pooled analysis. However, variability within conditions was substantial, and differences across modes did not reach statistical significance. These results suggest that interactive modes may support stronger factual retention over time, even when participants do not explicitly report greater learning. 6.4 Exploratory Interaction Log Analyses As shown in Table 4, the game condition involved substantially more back-and-forth interaction, with higher numbers of both user and system turns and longer session durations. However, user contributions in the game condition were considerably shorter, while system responses were longer and more verbose, resulting in a much lower userâsystem word ratio compared to the chat condition. Participants in the game condition also responded more quickly on average. This pattern reflects the gameâs interaction structure, in which many turns involved system-posed multiple-choice prompts. Although participants could suggest their own actions, they more often selected one of the predefined options, resulting in shorter user responses. To examine whether interaction patterns were associated with subjective and objective outcomes, we conducted exploratory Spearman rank correlations between interaction metrics and outcome variables. Correlations were computed both across combined chat and game participants (§6.4.1) and separately within each condition (§6.4.2). See Sec 5.4 for the experimental setup. 16Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz MetricChatGame Avg. User Turns7.2513.62 User Total Words150.5040.54 User Mean Words / Turn22.452.92 System Total Words533.921892.31 System Mean Words / Turn63.82131.05 UserâSystem Word Ratio0.3030.020 Session Duration (sec)576.35783.12 Mean Reaction Time (sec)67.9348.62 Table 4. Interaction log statistics by experimental condition. The table summarizes turn-taking behavior, message length, temporal characteristics, and relative userâsystem contribution across chat and game modes. Several associations reached nominal significance (uncorrected), particularly within the combined and chat conditions. However, after controlling for multiple comparisons using the BenjaminiâHochberg false discovery rate (FDR) procedure across all tests, none of the correlations remained statistically significant. We therefore interpret these associations cautiously as exploratory trends rather than confirmatory evidence of systematic relationships. Outcome Variable Interaction Metricíí Interpretation (Descriptive) Self- Reported Learning User turns-0.396 0.0498 More user turns are associated with lower perceived learning. User total words+0.477 0.0159 Higher user word count is associated with higher perceived learning. User avg. words/turn+0.501 0.0107 Longer user messages are associated with higher perceived learning. System total words-0.530 0.0065 Greater system verbosity is associated with lower perceived learning. System avg. words/turn-0.444 0.0262 Longer system messages are associated with lower perceived learning. Userâsystem word ratio+0.519 0.0079 Higher relative user contribution is associated with higher perceived learning. Session duration-0.439 0.0279 Longer sessions are associated with lower perceived learning. TrustMean reaction time+0.399 0.0481 Longer response latencies are associated with higher trust ratings. Table 5. Exploratory Spearman correlations between interaction metrics and outcome variables for combined chat and game participants. After applying the BenjaminiâHochberg false discovery rate (FDR) [12] correction across all tested correlations, none of the associations remained statistically significant. Only correlations with raw í-value< 0.05 are reported in this table. 6.4.1 Combined game and chat logs. As shown in Table 5, several interaction metrics exhibited nominal associations with self-reported knowledge when considering combined participants. Greater user contribution (reflected in higher total user word counts (í= .477,í= .016), longer average user messages (í= .501,í= .011), and a higher userâsystem word ratio (í= .519,í= .008)) was positively associated with perceived learning. In contrast, higher system verbosity (system total words:í=â.530,í= .007; system mean words per turn:í=â.444,í= .026) was negatively associated with self-reported learning. These findings may suggest that users felt they were learning more from more interactive settings in which they were engaged. Other factors that were negatively associated with perceived knowledge were a greater number of user turns (í= â.396,í= .050) and longer session durations (í= â.439,í= .028). These patterns mirror the structural differences observed between conditions in Table 4, where the game condition involved more turns, longer sessions, and more verbose system responses, while the chat condition involved longer user messages and higher relative user contribution. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning17 Outcome Variable Interaction Metricíí Interpretation (Descriptive) Easy to follow User avg. words/turn-0.717 0.0087 Longer user messages are associated with lower ease of following the experience. Userâsystem word ratio-0.615 0.0335 Higher relative user contribution is associated with lower convenience ratings. EngagementUser avg. words/turn-0.656 0.0205 Longer user messages are associated with lower reported engagement. EnjoymentUser avg. words/turn-0.700 0.0113 Longer user messages are associated with lower satisfaction. Encounter Again User total words-0.683 0.0144 Higher user word count is associated with lower desire to repeat the experience. Motivation to Act User turns+0.598 0.0402 More user turns are associated with higher motivation to act. User avg. words/turn-0.702 0.0110 Longer user messages are associated with lower motivation to act. Userâsystem word ratio-0.670 0.0172 Higher relative user contribution is associated with lower motivation to act. Table 6. Exploratory Spearman correlations between interaction metrics and subjective outcomes for the chat condition. Rawí-values are reported. After applying the BenjaminiâHochberg false discovery rate (FDR) correction across all tested correlations, none of the associations remained statistically significant. Finally, mean reaction time showed a modest positive association with trust (í= .399,í= .048). Notably, response latencies were longer on average in the chat condition compared to the game condition (67.93 vs. 48.62 seconds), consistent with the higher trust ratings observed for chat in Section 6.1. We ran Spearman rank correlations between mean reaction time and interaction metrics to disentangle whether longer response latencies reflected increased reading load or greater user contribution. Mean reaction time was positively associated with mean user message length (í= .47,í= .019) and the userâsystem word ratio (í= .45,í= .025), but not with system verbosity (trend only,í â .08). This suggests that longer response time â particularly in the chat condition â primarily reflect increased effort in composing responses rather than processing longer system outputs. 6.4.2 Within chat and game logs. When examining correlations separately within each condition, nominal asso- ciations were observed primarily in the chat condition (Table 6). Longer user messages (user avg. words/turn) were negatively associated with several subjective outcomes, including convenience (í= â.717,í= .009), engagement (í= â.656,í= .021), satisfaction (í= â.700,í= .011), and motivation to act (í= â.702,í= .011). Similarly, a higher userâsystem word ratio was negatively associated with convenience (í=â.615,í= .034) and motivation to act (í=â.670,í= .017). Greater total user word count was also negatively associated with the desire to encounter the experience again (í=â.683, í= .014). In contrast, a greater number of user turns was positively associated with motivation to act (í= .598,í= .040), suggesting that more back-and-forth interaction aligned with higher reported motivation in the chat condition. No nominally significant associations were observed within the game condition. As with the combined analysis, none of the correlations within either condition remained statistically significant after applying the BenjaminiâHochberg FDR correction. These results are therefore interpreted as exploratory trends rather than confirmatory evidence of systematic within-condition relationships. 7 Discussion Our goal in this work was to examine how three different modes of information delivery (essay, chatbot, and text-based game) shape persuasive experience and learning outcomes when informational content is held constant. By controlling for arguments and facts across conditions, our study isolates the role of interaction format and presentation style from content differences. The findings reveal several systematic tensions: interactive formats can enhance engagement and 18Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz perceived importance without uniformly increasing behavioral intention; subjective learning judgments can diverge from objective knowledge retention; and narrative game-based experiences may support memory while simultaneously raising concerns around realism and trust. In this section, we reflect on these results, relate them to prior work on persuasion and interactive systems, and derive implications for the design and evaluation of persuasive technologies and serious games. Delivery format strongly shaped subjective experience even when content was identical. Despite using the same set of arguments and numerical facts across conditions (§3.1), participantsâ subjective evaluations differed systematically by mode (§6.1). The chatbot condition tended to score highest across experience measures (e.g., engagement, enjoyment, trust, and perceived learning), while the essay condition often remained competitive on clarity and ease of following, and the game condition generally trailed. This pattern indicates that differences in subjective experience are not merely a function of informational content, but also of how that content is staged, paced, and interacted with. For persuasive-system designers, this underscores that interface and interaction design can be consequential even under strict content control [60]. Perceived importance does not immediately lead to behavioral change. Among the perceived change and persuasion outcomes (i.e. perceived importance, intention to act, and belief in effectiveness in §6.2), the clearest mode effect emerged for perceived importance. Across the combined-topic models, chat led to significantly stronger increases in importance than essay (and also exceeded game in pairwise models). Although there was a similar weak trend in belief in effectiveness, no comparable effects were observed for intention to act. This pattern suggests that conversational interaction may be especially effective at reframing issue salience â making a topic feel more personally important â without immediately altering downstream behavioral intentions. This finding is aligned with classic persuasion accounts in which engagement and perceived relevance can precede, and sometimes decouple from, stable attitude or intention change [52]. Perceived learning and objective retention diverged, especially for the game condition. A key finding is a dissociation between subjective learning judgments and delayed factual recall. Participants in the game condition reported substantially lower self-reported learning than those in essay (§6.1), yet achieved higher delayed quiz scores than essay in our objective assessment (§6.3). One plausible explanation is that interactive formats may support deeper processing and better encoding through active response generation and sustained attention, even when users feel that the experience was less informative or less serious [69]. In contrast, essays may feel clearer and more âinstructionalâ while encouraging comparatively passive consumption. Regardless, this mismatch cautions against relying on self-reported learning as a stand-alone indicator of educational effectiveness in interactive systems. Realism and trust as a trade-off in narrative game-based persuasion. Participantsâ open-ended explanations for changes in perceived importance, intention, and belief largely converged on a common theme: concerns about realism. This theme was mentioned more frequently in the story condition (7 participants) than in the chat (4 participants) and essay (0 participants) conditions. Participants in the story mode specifically noted that the outcomes of each game choice felt âpositive and smooth which is a little unrealisticâ, and does not require their âproblem-solvingâ mindset. These explanations suggest that some participants interpreted the text-based game as fictional or less grounded in real-world complexity, which in turn may have reduced their trust in the system. This perceived lack of realism not only helps explain the relative lowered persuasiveness scores in the story mode, but also sheds light on the lowered persuasion-related ratings such as convincingness, motivation to act, and influence on thinking, relative to the chat and Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning19 essay (§6.1). Prior work in narrative persuasion highlights perceived realism and narrative credibility as important moderators of persuasion and counter-arguing [9,10,17]. Our findings align with this literature in suggesting that narrative gamification can support memory and retention, while simultaneously risking a loss of perceived credibility when the framing signals fiction rather than real-world relevance. For designers of serious games, this points to a potential trade-off between the cognitive benefits of interactive narrative and the persuasive benefits of credibility. One implication is that grounding strategies (such as real-world settings, references to credible sources, or reflective prompts linking the experience to playersâ own lives) may help preserve trust without sacrificing the benefits of interaction [7]. Interaction traces should not be used as proxies for actual learning. Beyond delivery mode effects, our exploratory correlation analyses suggest that interaction structure is more closely related to how participants experience an interaction than to what they ultimately retain. In the interaction log analyses (§6.4), none of the interaction metrics (e.g., turn counts, word counts, session duration) showed meaningful associations with delayed objective knowledge scores. In contrast, in the combined chat and game analyses (§6.4.1), several interaction properties exhibited consistent directional relationships with subjective outcomes. Greater relative user contributionâreflected in longer user messages and higher userâsystem word ratiosâwas positively associated with self-reported learning, whereas longer session duration was negatively associated with perceived learning. These patterns mirror structural differences between modes (Table 4): the game condition involved more extensive turn-taking and more verbose system output, while the chatbot condition elicited more substantial user contributions. Within-condition analyses (§6.4.2) further clarify this distinction. Nominal associations were observed primarily in the chatbot condition, where interaction was open-ended and user-driven, whereas no comparable within-condition associations emerged in the game condition, where interaction structure was more system-controlled. This suggests that interaction traces may be most informative when users have meaningful agency over conversational flow. Although exploratory and not surviving correction for multiple comparisons, these results offer a clear design signal: behavioral engagement indicators such as verbosity, turn count, or session length should not be assumed to proxy objective learning, even when they shape usersâ subjective impressions. Within-chat interaction dynamics offer actionable design signals. While interaction traces should not be treated as indicators of learning outcomes, the within-chat analyses provide useful guidance for the design of interactive persuasive systems. Compared to the game conditionâwhere user input was constrained by narrative structureâ the chatbot afforded participants greater control over conversational pacing, resulting in meaningful variability in interaction patterns. Within the chatbot condition, longer user messages were nominally associated with lower ratings of ease of following, engagement, enjoyment, and motivation to act, whereas a greater number of conversational turns was positively associated with motivation. Together, these patterns suggest that perceived conversational quality depends less on how much is said and more on how interaction is paced and distributed across turns. Sustained back-and-forth may feel motivating, whereas overly long contributions or system-dominated exchanges may increase cognitive load and reduce perceived fluency. For designers of chat-based persuasive systems, these exploratory findings point to the importance of interaction balance rather than maximal engagementâfavoring shorter, more frequent exchanges that support incremental user par- ticipation. More broadly, they highlight that user agency in conversational pacing may be as important as informational content itself in shaping subjective experience. 20Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Design recommendations for persuasive systems and serious games. Lastly, our findings motivate outcome- driven design choices rather than a blanket preference for âmore interactivityâ. If the goal is to improve subjective experience and increase perceived importance, conversational delivery appears particularly promising (§6.1, §6.2). If the goal is longer-term factual retention, interactive formatsâincluding narrative game experiencesâmay provide advantages even when users report learning less (§6.3). For serious-game developers, two concrete implications follow: (1) balance user agency and system agency so that the system does not dominate the interaction through excessive verbosity, and (2) incorporate credibility and realism cues (e.g., explicit sourcing, grounded settings, reflection steps) to mitigate the trust penalty that can accompany fictional framing. More broadly, evaluations of persuasive interactive systems should jointly consider self-reported, objective outcomes, and behavioral traces, rather than equating engagement proxies (e.g. turns, verbosity, etc.) with effectiveness. Limitations Topic scope. Our study focuses on interactive methods in educational technology for sustainability-related topics. We selected two common topics in this domain â recycling and public transport. However, there are many other relevant topics, such as renewable energy, waste reduction, and biodiversity. Future studies could examine one or more of these additional topics to provide a more holistic understanding of how interactive methods can be effectively utilized in sustainability education and further benefit society. Participant population and sample size for Measuring Persuasion. We had a total of 45 participants. Because we used a between-subject design, the number of participants within each group was relatively smaller. Participants were recruited through university posters. This may have introduced demographic bias, as the sample consisted primarily of younger and more educated individuals, two demographics which have been shown to hold stronger pre-existing beliefs in climate change [26,32]. This pattern was reflected in participantsâ open-ended responses in the post-study questionnaire. Across all conditions, participants reported prior awareness of the importance and effectiveness of recycling and public transit, as well as pre-existing behavioral intentions. Many indicated that the study simply âreinforced the ideaâ. Repeating this study with a different population â for example, individuals who are more skeptical about and have opposing views on sustainability topics â may yield different outcomes. Future work looking to expand the population of participants may have to involve other means in the design, such as debunking misinformation about climate change [64]. Knowledge retention and Belief Integration. We evaluated knowledge retention after 24 hours, which reflects a short-term retention window. To further validate whether the knowledge is retained in the long term, future work should adopt a longitudinal design and administer follow-up knowledge assessments after a longer delay (e.g., several days or weeks). While remembering information is likely necessary for persuasion to last, it is not sufficient on its own. Remembering more facts does not automatically lead to stronger belief change or behavioral action [66]. In fact, individuals are often unable to escape the pull of their existing attitudes, which guide how new information is processed and evaluated [61]. Future research could therefore examine longer-term retention and test whether retained knowledge translates into durable belief change and behavior. Ethical Statement Access. The code base for PersuLab is publicly available. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning21 Participant Selection and Compensation. Participants were adult volunteers recruited through university advertise- ments. All participants were fluent in English, ensuring clear communication throughout the study. Participants were compensated $20 USD for their time, which exceeds local minimum wage standards. Participant Consent and Data Usage. Prior to participation, all participants provided informed consent. The consent form explained the study procedure, data collection, and data usage. Participants were informed that they would be interacting with AI-generated content, advised not to share personal or sensitive information, and assured that all responses would be anonymized and used solely for research purposes. Ethics Review. This study involved a voluntary user study with adult participants. The study protocol was reviewed and approved by the authorsâ institutional ethics review board. All participants provided informed consent prior to participation and were informed that they could withdraw from the study at any time without penalty. Acknowledgments This work was funded, in part, by the Vector Institute for AI, Canada CIFAR AI Chairs program, Accelerate Foundation Models Research Program Award from Microsoft, and an NSERC discovery grant. We thank Dongwook Yoon for helpful feedback and discussion. References [1]Ifeoma Adaji and Mikhail Adisa. 2022. A Review of the Use of Persuasive Technologies to Influence Sustainable Behaviour. In Adjunct Proceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization (Barcelona, Spain) (UMAP â22 Adjunct). Association for Computing Machinery, New York, NY, USA, 317â325. doi:10.1145/3511047.3537653 [2]Ifeoma Adaji, Nafisul Kiron, and Julita Vassileva. 2020. Evaluating the Susceptibility of E-commerce Shoppers to Persuasive Strategies. A Game-Based Approach. In International Conference on Persuasive Technology. https://api.semanticscholar.org/CorpusID:215640239 [3] Seyed Hossein Alavi, Sudha Rao, Ashutosh Adhikari, Gabriel A DesGarennes, Akanksha Malhotra, Chris Brockett, Mahmoud Adada, Raymond T Ng, Vered Shwartz, and Bill Dolan. 2024. Mcpdial: A minecraft persona-driven dialogue dataset. arXiv preprint arXiv:2410.21627 (2024). [4]Seyed Hossein Alavi, Weijia Xu, Nebojsa Jojic, Daniel Kennett, Raymond T Ng, Sudha Rao, Haiyan Zhang, Bill Dolan, and Vered Shwartz. 2024. Game Plot Design with an LLM-powered Assistant: An Empirical Study with Game Designers. arXiv preprint arXiv:2411.02714 (2024). [5] Sacha Altay, Anne-Sophie Hacquin, Coralie Chevallier, and Hugo Mercier. 2023. Information delivered by a chatbot has a positive impact on COVID-19 vaccines attitudes and intentions. Journal of Experimental Psychology: Applied 29, 1 (2023), 52. [6] Fraser Anderson, Tovi Grossman, Daniel Wigdor, and George Fitzmaurice. 2013. Gesture-based interfaces: Learning, performance, and retention. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. 247â256. [7]Jeannette Angel, Alicia LaValle, Deepti Mathew Iype, Stephen Sheppard, and Aleksandra Dulic. 2015. Future delta 2.0 an experiential learning context for a serious game about local climate change. In SIGGRAPH Asia 2015 Symposium on Education (Kobe, Japan) (SA â15). Association for Computing Machinery, New York, NY, USA, Article 12, 10 pages. doi:10.1145/2818498.2818512 [8]Parmjit Singh Aperapar Singh and Lilian Anthonysamy. 2023. The dichotomization of objective and subjective outcome measures of academic performance in an online learning environment. Malaysian Journal of Learning and Instruction (MJLI) 20, 1 (2023), 63â92. [9]Markus Appel and Martina Mara. 2013. The persuasive influence of a fictional characterâs trustworthiness. Journal of Communication 63, 5 (2013), 912â932. [10] Markus Appel and Tobias Richter. 2007. Persuasive Effects of Fictional Narratives Increase Over Time. Media Psychology 10 (12 2007), 113â134. doi:10.1080/15213260701301194 [11]Hui Bai, Jan G Voelkel, Shane Muldowney, Johannes C Eichstaedt, and Robb Willer. 2025. LLM-generated messages can persuade humans on policy issues. Nature Communications 16, 1 (2025), 6037. [12]Yoav Benjamini and Yosef Hochberg. 1995. Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B (Methodological) 57, 1 (1995), 289â300. [13] S , tefan Boncu, Octav-Sorin Candel, and Nicoleta Laura Popa. 2022. Gameful Green: A Systematic Review on the Use of Serious Computer Games and Gamified Mobile Apps to Foster Pro-Environmental Information, Attitudes and Behaviors. Sustainability 14, 16 (2022). doi:10.3390/su141610400 [14]Hronn Brynjarsdottir, Maria HĂ„kansson, James Pierce, Eric Baumer, Carl DiSalvo, and Phoebe Sengers. 2012. Sustainably unpersuaded: how persuasion narrows our vision of sustainability. In Proceedings of the sigchi conference on human factors in computing systems. 947â956. [15] Nicholas J. Cepeda, Harold Pashler, Edward Vul, John T. Wixted, and Doug Rohrer. 2006. Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin 132, 3 (2006), 354â380. 22Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz [16]Raluca Chisalita, Markus Murtinger, and Simone Kriglstein. 2022. Grow Your Plant: A Plant-Based Game For Creating Awareness About Sustainability Behaviour by Using Renewable Energy. In Extended Abstracts of the 2022 Annual Symposium on Computer-Human Interaction in Play (Bremen, Germany) (CHI PLAY â22). Association for Computing Machinery, New York, NY, USA, 177â182. doi:10.1145/3505270.3558344 [17] Hyunyi Cho, Lijiang Shen, and Kari Wilson. 2014. Perceived realism: Dimensions and roles in narrative persuasion. Communication research 41, 6 (2014), 828â851. [18]Shruthi Chockkalingam, Seyed Hossein Alavi, Raymond T. Ng, and Vered Shwartz. 2025. Should I go vegan: Evaluating the Persuasiveness of LLMs in Persona-Grounded Dialogues. In Proceedings of the Third Workshop on Social Influence in Conversations (SICon 2025), James Hale, Brian Deuksin Kwon, and Ritam Dutt (Eds.). Association for Computational Linguistics, Vienna, Austria, 65â72. doi:10.18653/v1/2025.sicon-1.4 [19] Peter W De Vries, Harri Oinas-Kukkonen, Liseth Siemons, Nienke Beerlage-de Jong, and Lisette van Gemert-Pijnen. 2017. Persuasive Technology: Development and Implementation of Personalized Technologies to Change Attitudes and Behaviors: 12th International Conference, PERSUASIVE 2017, Amsterdam, The Netherlands, April 4â6, 2017, Proceedings. Vol. 10171. Springer. [20]Esin Durmus, Liane Lovitt, Alex Tamkin, Stuart Ritchie, Jack Clark, and Deep Ganguli. 2024. Measuring the Persuasiveness of Language Models. https://w.anthropic.com/news/measuring-model-persuasiveness [21] Hoda Elmgadmi and Khalid Nafil. 2025. Large Language Models and Non-Player Characters in Gaming: A Bibliometric Overview. In 2025 International Conference on Intelligent Systems: Theories and Applications (SITA). IEEE, 1â8. [22]Daniel FernĂĄndez Galeote and Juho Hamari. 2021. Game-based Climate Change Engagement: Analyzing the Potential of Entertainment and Serious Games. Proc. ACM Hum.-Comput. Interact. 5, CHI PLAY, Article 226 (Oct. 2021), 21 pages. doi:10.1145/3474653 [23] Philip M. Fernbach, Todd Rogers, Craig R. Fox, and Steven A. Sloman. 2013. Political Extremism Is Supported by an Illusion of Understanding. Psychological Science 24, 6 (2013), 939â946. arXiv:https://doi.org/10.1177/0956797612464058 doi:10.1177/0956797612464058 PMID: 23620547. [24] Daniel FernĂĄndez Galeote, Mikko Rajanen, Dorina Rajanen, Nikoletta-Zampeta Legaki, David J Langley, and Juho Hamari. 2021. Gamification for climate change engagement: review of corpus and future agenda. Environmental Research Letters 16, 6 (jun 2021), 063004. doi:10.1088/1748-9326/abec05 [25]Cassandra Folkins, Emily Read, Jeff Mundee, Max V. Birk, and Scott Bateman. 2020. A Serious Game for Promoting Positive Attitudes Towards Nursing Homes Among Youth. In Proceedings of the Annual Symposium on Computer-Human Interaction in Play (Virtual Event, Canada) (CHI PLAY â20). Association for Computing Machinery, New York, NY, USA, 484â498. doi:10.1145/3410404.3414253 [26]Adrian Furnham and Charlotte Robinson. 2022. Correlates of belief in climate change: Demographics, ideology and belief systems. Acta Psychologica 230 (2022), 103775. doi:10.1016/j.actpsy.2022.103775 [27] Kazuaki Furumai, Roberto Legaspi, Julio Cesar Vizcarra Romero, Yudai Yamazaki, Yasutaka Nishimura, Sina Semnani, Kazushi Ikeda, Weiyan Shi, and Monica Lam. 2024. Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval. In Findings of the Association for Computational Linguistics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 11224â11249. doi:10.18653/v1/2024.findings-emnlp.656 [28] Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2024. How persuasive is AI-generated propaganda? PNAS nexus 3, 2 (2024), pgae034. [29] Kobi Hackenburg, Ben M Tappin, Luke Hewitt, Ed Saunders, Sid Black, Hause Lin, Catherine Fist, Helen Margetts, David G Rand, and Christopher Summerfield. 2025. The levers of political persuasion with conversational artificial intelligence. Science 390, 6777 (2025), eaea3884. [30]Inez EP Harker-Schuch, Franklin P Mills, Steven J Lade, and Rebecca M Colvin. 2020. CO2peration â Structuring a 3D interactive digital game to improve climate literacy in the 12-13-year-old age group. Computers & Education 144 (2020), 103705. doi:10.1016/j.compedu.2019.103705 [31] Sebastian Hobert and Raphael Meyer von Wolff. 2019. Say hello to your new automated tutorâa structured literature review on pedagogical conversational agents. (2019). [32]Anne G. Hoekstra, Kjell Noordzij, Willem de Koster, and Jeroen van der Waal. 2024. The educational divide in climate change attitudes: Understanding the role of scientific knowledge and subjective social status. Global Environmental Change 86 (2024), 102851. doi:10.1016/j.gloenvcha.2024.102851 [33]Shilpa S. Kantak and Carole J. Winstein. 2012. Learning-performance distinction and memory processes for motor skills. Current Opinion in Neurology 25, 6 (2012), 590â597. [34]Elise Karinshak, Sunny Xun Liu, Joon Sung Park, and Jeffrey T Hancock. 2023. Working with AI to persuade: Examining a large language modelâs ability to generate pro-vaccination messages. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1 (2023), 1â29. [35] Erik Knol and Peter W De Vries. 2011. EnerCities-A serious game to stimulate sustainability and energy conservation: Preliminary results. eLearning Papers 25 (2011). [36] William H Kruskal and W Allen Wallis. 1952. Use of ranks in one-criterion variance analysis. J. Amer. Statist. Assoc. 47, 260 (1952), 583â621. [37] Floridi Luciano. 2024. HypersuasionâOn AIâs persuasive power and how to deal with it. Philosophy & Technology 37, 2 (2024), 64. [38] Henry B Mann and Donald R Whitney. 1947. On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics (1947), 50â60. [39]John Matthews, Khin Than Win, Harri Oinas-Kukkonen, and Mark Freeman. 2016. Persuasive technology in mobile applications promoting physical activity: a systematic review. Journal of medical systems 40, 3 (2016), 72. [40]Sandra C Matz, Jacob D Teeny, Sumer S Vaid, Heinrich Peters, Gabriella M Harari, and Moran Cerf. 2024. The potential of generative AI for personalized persuasion at scale. Scientific Reports 14, 1 (2024), 4692. [41]Peter McCullagh. 1980. Regression models for ordinal data. Journal of the Royal Statistical Society Series B: Statistical Methodology 42, 2 (1980), 109â127. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning23 [42] Hugo Mercier. 2020. Not Born Yesterday: The Science of Who We Trust and What We Believe. Princeton University Press, Princeton, NJ. [43] H. Mercier and D. Sperber. 2019. The Enigma of Reason. Harvard University Press. https://books.google.ca/books?id=6JeDvQEACAAJ [44]Martha C Monroe, Richard R Plate, Annie Oxarart, Alison Bowers, and Willandia A Chaves. 2019. Identifying effective climate change education strategies: A systematic review of the research. Environmental Education Research 25, 6 (2019), 791â812. [45]Fernanda Murillo-Muñoz, Christian Navarro-Cota, Reyes JuĂĄrez-RamĂrez, Samantha JimĂ©nez, Juan Ivan Nieto HipĂłlito, Ana I. Molina, and Mabel Vazquez-Briseno. 2021. Characteristics of a Persuasive Educational System: A Systematic Literature Review. Applied Sciences 11, 21 (2021). doi:10.3390/app112110089 [46]Jimin Nam, Reed Orchinik, and David G Rand. 2025. LLMs as Scalable Tools for Interactive Consumer Behavior Experiments: Comparing Persuasion Strategy Effectiveness. (2025). [47]N Nazrina M Nazry and Daniela M Romano. 2017. Mood and learning in navigation-based serious games. Computers in Human Behavior 73 (2017), 596â604. [48]Isabel Newsome. 2020. An Educational Game Bringing Awareness to Declining Insect Populations. In Extended Abstracts of the 2020 Annual Symposium on Computer-Human Interaction in Play (Virtual Event, Canada) (CHI PLAY â20). Association for Computing Machinery, New York, NY, USA, 326â329. doi:10.1145/3383668.3419912 [49]E Michael Nussbaum, Marissa C Owens, Gale M Sinatra, Abeera P Rehmat, Jacqueline R Cordova, Sajjad Ahmad, Fred C Harris Jr, and Sergiu M Dascalu. 2015. Losing the Lake: Simulations to promote gains in student knowledge and interest about climate change. International Journal of Environmental and Science Education 10, 6 (2015), 789â811. [50] Chinedu Wilfred Okonkwo and Abejide Ade-Ibijola. 2020. Python-bot: A chatbot for teaching python programming. Engineering Letters 29, 1 (2020). [51] Adam M. Persky, Edward Lee, and Lauren S. Schlesselman. 2020. Perception of Learning Versus Performance as Outcome Measures of Educational Research. American Journal of Pharmaceutical Education 84 (2020). https://api.semanticscholar.org/CorpusID:216154214 [52]Richard E. Petty and John T. Cacioppo. 1986. The Elaboration Likelihood Model of Persuasion. In Advances in Experimental Social Psychology, Leonard Berkowitz (Ed.). Vol. 19. Academic Press, 123â205. doi:10.1016/S0065-2601(08)60214-2 [53]Henry L. Roediger and Jeffrey D. Karpicke. 2006. Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science 17, 3 (2006), 249â255. [54] Alexander Rogiers, Sander Noels, Maarten Buyl, and Tijl De Bie. 2024. Persuasion with large language models: a survey. arXiv preprint arXiv:2411.06837 (2024). [55]Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2025. On the conversational persuasiveness of GPT-4. Nature Human Behaviour (2025), 1â9. [56]Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2024. On the conversational persuasiveness of large language models: A randomized controlled trial. arXiv preprint arXiv:2403.14380 (2024). [57] Philipp Schoenegger, Francesco Salvi, Jiacheng Liu, Xiaoli Nan, Ramit Debnath, Barbara Fasolo, Evelina Leivada, Gabriel Recchia, Fritz GĂŒnther, Ali Zarifhonarvar, et al.2025. Large Language Models Are More Persuasive Than Incentivized Human Persuaders. arXiv preprint arXiv:2505.09662 (2025). [58]Samuel Shields, Celine Lafosse, Shi Johnson-Bey, Daeun Hwang, Noah Wardrip-Fruin, and Edward F Melcer. 2025. Could vs Should: Exploring Prompting Strategies and Writer Perspectives Towards LLM Assistance in Storylet Authoring. IEEE Transactions on Games (2025). [59] Almog Simchon, Matthew Edwards, and Stephan Lewandowsky. 2024. The persuasive effects of political microtargeting in the age of generative artificial intelligence. PNAS nexus 3, 2 (2024), pgae035. [60]S. Shyam Sundar. 2008. The MAIN Model: A Heuristic Approach to Understanding Technology Effects on Credibility. In Digital Media and Learning, Miriam J. Metzger and Andrew J. Flanagin (Eds.). MacArthur Foundation/MIT Press, 73â100. https://w.issuelab.org/resource/the-main-model- a-heuristic-approach-to-understanding-technology-effects-on-credibility.html [61] Charles S Taber and Milton Lodge. 2006. Motivated skepticism in the evaluation of political beliefs. American journal of political science 50, 3 (2006), 755â769. [62]Anja Thieme, Rob Comber, Julia Miebach, Jack Weeden, Nicole Kraemer, Shaun Lawson, and Patrick Olivier. 2012. "Weâve bin watching you": designing for reflection and social persuasion to promote sustainable lifestyles. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Austin, Texas, USA) (CHI â12). Association for Computing Machinery, New York, NY, USA, 2337â2346. doi:10.1145/2207676.2208394 [63]Giovanni Maria Troiano, Dylan Schouten, Michael Cassidy, Eli Tucker-Raymond, Gillian Puttick, and Casper Harteveld. 2020. Ice Paddles, CO2 Invaders, and Exploding Planets: How Young Students Transform Climate Science Into Serious Games. In Proceedings of the Annual Symposium on Computer-Human Interaction in Play (Virtual Event, Canada) (CHI PLAY â20). Association for Computing Machinery, New York, NY, USA, 534â548. doi:10.1145/3410404.3414256 [64]Sander Van der Linden, Anthony Leiserowitz, Seth Rosenthal, and Edward Maibach. 2017. Inoculating the public against misinformation about climate change. Global challenges 1, 2 (2017), 1600008. [65]Abraham Wald. 1943. Tests of statistical hypotheses concerning several parameters when the number of observations is large. Trans. Amer. Math. Soc. 54, 3 (1943), 426â482. [66]Elena M Galeano Weber, Lisa Lehnen, Doug Lombardi, and Garvin Brod. 2025. Practice testing enhances learning but not attitude change from persuasive texts. Scientific Reports 15, 1 (2025), 32935. 24Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz CategoryQuestionResponse format DemographicsWhat is your age?Categorical (18â24, 25â34, 35â44, 45â54, 55â64, 65+) Topic assignmentWhat is the topic that you were assigned?Multiple choice Context (Recycling)Do you have convenient access to recycling services at home or in your building?Yes / No Importance (Recycling)How important is recycling to you personally?Likert (1â5) Behavioral intention (Recycling)How often do you recycle materials such as paper, plastic, glass, or metal? Categorical; recoded to ordinal 1â5 (Neverâ Always) Epistemic conficence (Recycling)How confident are you in your knowledge about how recycling works and its impact on reducing waste and protecting the environment? Likert (1â5) Context (Transit)Do you have regular access to a personal car?Yes / No Importance (Transit)How important is choosing public transit instead of personal cars to you personally?Likert (1â5) Behaviroal intention (Transit)On a typical week, how many days do you commute outside the home? Ordinal Categorical; recoded to ordinal 1â4 (0, 1â2, 3â4, 5+) Epistemic confidence(Transit) How confident are you in your knowledge about the environmental and social benefits of using public transit instead of personal cars? Likert (1â5) Table 7. Pre-study questionnaire items collected for descriptive and contextual purposes. [67] Hairu Yang, Minghan Cai, Yongfeng Diao, Rui Liu, Ling Liu, and Qianchen Xiang. 2023. How does interactive virtual reality enhance learning outcomes via emotional experiences? A structural equation modeling approach. Frontiers in Psychology 13 (2023), 1081372. [68]Jerrold H. Zar. 2005. Spearman rank correlation. Vol. 7. Encyclopedia of Biostatistics. https://citebay.com/how-to-cite/spearmans-rank-correlation- coefficient/ Wiley. [69]Yu Zhonggen. 2019. A meta-analysis of use of serious games in education over a decade. International Journal of Computer Games Technology 2019, 1 (2019), 4797032. A User Study Questionnaires A.1 Pre-Study Questionnaire Table 7 summarizes the pre-study questionnaire administered prior to the experiment. The questionnaire collected demographic and contextual information, as well as participantsâ self-reported attitudes, behaviors, and confidence related to the assigned topic (recycling or public transit). A.2 Post-Study Questionnaire Table 8 presents the post-study questionnaire administered immediately after the experimental session. The question- naire captured participantsâ subjective experience (e.g., easy to follow, engagement, self-reported learning, enjoyment, trust, influence), and topic-specific perceived changes in importance, intention to act, and belief in effectiveness. Subjec- tive experience items were measured using 5-point Likert scales. Although âconvincingness of argumentsâ was collected as a topic-specific item using a slightly different Likert wording (Not at allâVery), we report it alongside other subjective outcomes because it captures participantsâ experiential judgment of the persuasive content rather than objective belief change. Perceived-change items used ordered categorical responses (Less / Same / More / Not sure) and were recoded to ordinal values (-1, 0, +1) for analysis, with Not sure responses treated as missing. Open-ended prompts were included to collect brief qualitative reflections. A.3 Objective Knowledge Quiz Tables 9 and 10 list the multiple-choice questions used to assess objective knowledge retention. For each topic, partici- pants answered five questions whose answers were explicitly covered during the experiment session, along with two Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning25 CategoryQuestionResponse formatApplies to Subjective experience Ease of followingThe information was presented in a way that was easy to follow.Likert (1=SD, 5=SA)All topics EngagementI felt engaged while taking part in this experience.Likert (1=SD, 5=SA)All topics Self-reported learningThis experience increased my understanding/knowledge of the topic.Likert (1=SD, 5=SA)All topics EnjoymentI enjoyed the way this experience delivered the information.Likert (1=SD, 5=SA)All topics TrustI trusted the information provided in this experience.Likert (1=SD, 5=SA)All topics Would RecommendI would recommend this kind of experience to other people.Likert (1=SD, 5=SA)All topics Re-encounter in futureI would like to encounter this type of experience again in the future.Likert (1=SD, 5=SA)All topics Increased Motivation to actThis experience increased my motivation to act more sustainably.Likert (1=SD, 5=SA)All topics Influenced thinkingOverall, this experience influenced the way I think about this topic.Likert (1=SD, 5=SA)All topics Topic assignmentWhat is the topic that you were assigned?Multiple choiceAll topics Perceived change (Recy- cling) Importance change Compared to before this session, how has the personal importance of recycling changed for you? Less / Same / More / Not sureRecycling only Behavioral intention changeCompared to before this session, how has your intention to recycle whenever possible changed? Less / Same / More / Not sureRecycling only Belief in effectiveness changeCompared to before this session, how has your belief in the effectiveness of recycling changed? Less / Same / More / Not sureRecycling only Convincingness (Recycling)How convincing did you find the arguments about recycling?Likert (1=Not at all, 5=Very)Recycling only Perceived change (Transit) Importance change Compared to before this session, how has the personal importance of choosing public transit changed? Less / Same / More / Not sureTransit only Behavioral intention changeCompared to before this session, how has your intention to use public transit instead of a personal car changed? Less / Same / More / Not sureTransit only Belief in effectiveness change Compared to before this session, how has your belief in the effectiveness of public transit changed? Less / Same / More / Not sureTransit only Convincingness (Transit)How convincing did you find the arguments about using public transit instead of personal cars? Likert (1=Not at all, 5=Very)Transit only Open-ended reflectionsWhy? Please provide brief explanation (2â3 sentences).Free textAfter each perceived change/convincing item Please describe how (if at all) this experience influenced your thoughts, feelings, or actions related to the topic (3â4 sentences). Free textAll topics Table 8. Post-study questionnaire used to acquire subjective experience of participants. control questions referencing information not presented. Control questions were included to discourage guessing or external lookup; selecting âI have not seen this information beforeâ was treated as the correct response. Knowledge scores were computed as the number of correct responses to the five content-covered questions only. Confidence ratings collected after each question (via âHow confident are you in your answer to the previous question?â question) were not analyzed in this work. B Key Arguments For each topic (recycling and public transit), we compiled five persuasive arguments, each paired with a concrete numerical fact. The complete argument sets are listed in Tables 11. All prompts referencing<key_arguments>loaded these files verbatim. C Essay Generation Table 12 shows the template prompt used for generating essays in our experiment. Essay generation was performed using GPT-4.1 with a maximum completion length of 33k tokens and default decoding parameters. We generated 20 essays per topic in advance and sampled from this fixed set during the experiment to reduce interaction delays and avoid unforeseen technical issues during data collection. 26Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz QuestionAnswer optionsCorrect answerType Recycling one ton of paper saves approximately how many trees? 10; 13; 17; 20; I have not seen this information before17Content Recycling aluminum saves up to what percentage of energy compared to producing new aluminum? 80%; 85%; 90%; 95%; I have not seen this information before 95%Content Recycling just one aluminum can can power a TV for ap- proximately how long? 30 minutes; 1 hour; 3 hours; 5 hours; I have not seen this information before I have not seen this informa- tion before Control How many more jobs does the recycling industry generate compared to landfill management? 10Ămore; 5Ămore; 2Ămore; 7Ămore; I have not seen this information before 10Ă moreContent Reducing landfill waste can lower disposal costs by up to what percentage? 10%; 20%; 30%; 40%; I have not seen this information before I have not seen this informa- tion before Control Unrecycled plastic breaks into microplastics that harm how many marine species? Over 70; Over 300; Over 500; Over 700; I have not seen this information before Over 700Content Companies with recycling programs see up to what per- centage increase in employee engagement and retention? 10%; 15%; 20%; 25%; I have not seen this information before 20%Content How confident are you in your answer to the previous question? Likert 1-5 (Not confident at allâExtremely Confi- dent) N/AN/A Table 9. Objective knowledge quiz items for the recycling topic. Each participant answered five content-covered questions and two control questions. Control questions referenced information not presented during the experiment; selecting âI have not seen this information beforeâ was considered the correct response. After each question (content or control), particpants were asked a confidence question âHow confident are you in your answer to the previous questionâ. QuestionAnswer optionsCorrect answerType How much can switching from personal cars to public tran- sit reduce an individualâs transportation emissions? 20%; 35%; 50%; 75%; I have not seen this information before 50%Content A 10% shift from personal cars to public transit can cut commute times by up to: 15%; 25%; 40%; 60%; I have not seen this information before 40%Content How much do businesses typically spend per employee on parking alone each year? $500-$1,000; $1,000-$1,500; $1,500-$2,000; $2,500- $3,000; I have not seen this information before I have not seen this informa- tion before Control How many people can a single lane accommodate when used by buses or trains instead of cars? 2,500; 10,000; 15,000; 20,000; I have not seen this in- formation before I have not seen this informa- tion before Control On average, how much money can a commuter save per year by not driving and using public transit instead? $3,000; $5,000; $7,500; $10,000+; I have not seen this information before $10,000+Content How much safer is traveling by public transit compared to driving a personal car? 5Ăsafer; 10Ăsafer; 20Ăsafer; 40Ăsafer; I have not seen this information before 10Ă saferContent Every $1 invested in public transit generates approximately how much in community benefits? $4; $5; $6; $7; I have not seen this information before$4Content How confident are you in your answer to the previous question? Likert 1â5 (Not confident at allâExtremely confi- dent) N/AN/A Table 10. Objective knowledge quiz items for the public transit topic. Each participant answered five content-covered questions and two control questions. Control questions referenced information not presented during the experiment; selecting âI have not seen this information beforeâ was considered the correct response. After each question (content or control), participants were asked a confidence question. D Chatbot Implementation The chatbot condition was implemented using a large language model with structured system and user prompts. Table 13 illustrates the construction of the system and user messages, along with the initial static prompts used for the recycling and public transit topics. Throughout the experiment we used GPT-4.1 for generating responses with max generation token limit set to 2048 and âUser:â as the primary stop word. In the system message, a placeholder (<key arguments>) was populated with the topic-specific argument text shown in Table 11. Because the study included a delayed (24-hour) objective knowledge assessment, it was necessary to ensure that participants were exposed to all predefined arguments and factual statements during the interaction. The user message prompt included the full interaction history, consisting of all prior system and user messages concatenated into a single context string. Messages were prefixed with explicit role tags (e.g.,âUser: âand âSystem: â) to preserve turn structure. The history string included the initial static system prompt followed by Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning27 Topic A: Recycling Here are five compelling arguments to persuade individuals and businesses to adopt better recycling habits: 1. Environmental Preservation: Protecting Natural Resources and Ecosystems Premise: Recycling reduces raw material extraction, preserving forests, water, and ecosystems. Fact: Recycling one ton of paper saves 17 trees. 2. Energy Conservation and Lower Carbon Emissions Premise: Recycling uses much less energy than making new materials, cutting emissions. Fact: Recycling aluminum saves up to 95% of the energy compared to new production. 3. Economic Benefits: Cost Savings and Job Creation Premise: Recycling creates more jobs and saves businesses money on waste. Fact: Recycling creates 10 jobs for every landfill job. 4. Reducing Pollution: Cleaner Air, Water, and Soil Premise: Recycling reduces pollution from landfill waste and incineration. Fact: Unrecycled plastic breaks into microplastics that harm over 700 marine species and enter the food chain. 5. Setting a Positive Example and Meeting Social Responsibility Goals Premise: Recycling shows commitment and inspires sustainable behavior. Fact: Companies with recycling programs have up to 20% higher employee engagement and retention. By highlighting these arguments, you can effectively appeal to the environmental, economic, and social benefits of better recycling habits, making the case for why change is both necessary and rewarding. Topic B: Public Transit Here are five key arguments for convincing individuals and businesses to use public transit instead of personal cars: 1. Environmental Benefits: Reduced Carbon Footprint Premise: Public transit cuts emissions and helps clean the air. Fact: Switching to transit can cut your transportation emissions by 50%. 2. Cost Savings for Individuals and Businesses Premise: Public transit saves money for commuters and employers. Fact: The average commuter saves $10,000+ per year by not driving. 3. Reduced Traffic Congestion Premise: Transit reduces traffic and speeds up commutes. Fact: A 10% shift to transit can cut commute times by up to 40%. 4. Health and Safety Improvements Premise: Transit is safer and helps reduce pollution-related illness. Fact: Transit is 10Ă safer than driving. 5. Economic and Social Benefits for Communities Premise: Transit boosts local economies and access for all. Fact: Every $1 spent on transit returns $4 in community benefits. By emphasizing these arguments, individuals and businesses can see the tangible benefits of public transit, not just for themselves, but for their communities and the environment as a whole. Table 11. Fixed persuasive arguments and supporting facts for the recycling (top section) and transit (bottom section). Each argument consists of a high-level premise paired with a numerical fact. This content was made by the authors of this paper and used verbatim across all experiment conditions (i.e. essay, chat, and game). alternating user inputs and system responses. All interactions were logged, including turn-level timestamps, message content, the set of facts covered, and the remaining uncovered facts at each turn. To enforce content coverage, each chatbot response was evaluated using an LLM-based fact-checking module. After every system-generated message, the fact checker identified which of the target facts had been covered and returned their corresponding indices. The list of remaining uncovered facts was then updated and injected back into the system prompt via a dedicated placeholder (<remaining uncovered facts>). 28Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz System Message: You are an expert essay writer. For each key argument write a paragraph and focus on the provided fact. The essay must follow the following formatting: <article> <h3><strong>Essay Title</strong></h3> <p>paragraph 1</p> <p>paragraph 2</p> ... <p> Paragraph n</p> </article> User Message: Write a persuasive and informative essay based on the following key arguments: <key arguments> Table 12. Template prompt used for generating essays. We generated 20 prompts for each topic (i.e. recycling and transit), and sampled from them during our user study. See Appendix B for more details on key arguments. The conversation was allowed to continue until all predefined facts had been covered, at which point the interface enabled the participant to end the session. Participants could also continue interacting beyond this point or terminate the interaction earlier if they chose. System Message You are a persuasive chatbot in a turn-based, text-based conversation. The conversation is based on the following key arguments: <key arguments> Your job is to guide the user through a persuasive conversation, making the experience engaging, natural, and informative. Rules for conversation: - Present arguments one at a time. Never present more than one fact in a single turn. - Keep each response short, conversational, and persuasive. - Address the userâs concerns directly and adapt your response to their questions. - If the user is skeptical, acknowledge their perspective and offer a counterargument rooted in the provided key arguments. - Integrate at most one fact per turn. Make the fact stand out clearly but present it organically in the flow of conversation. - Once all facts are covered and the conversation has run its course, end naturally. Constraints: - Never mention that you are an AI or assistant. - Avoid phrases like âSure,â âOf course,â etc. - Do not over-explain rules to the user. Facts to integrate: Here are the remaining facts not yet presented: <remaining uncovered facts> User Message Here is a history of whatâs happened so far: <history> User: <user input> Start your response in a conversational tone (no labels like âSystem:â or âAssistant:â). Remember: - Keep responses short and persuasive. - Integrate at most one remaining fact, and make it stand out naturally in the conversation. - Adapt arguments to the userâs concerns. - You may end the conversation naturally once all facts are covered. System: Initial static prompt for Recycling Imagine making a real difference for the planet every day, just by changing a few habits. Recycling isnât just about sorting wasteâitâs about protecting what matters most. Have you ever thought about the impact your recycling can actually have? Initial static prompt for Transit Have you ever thought about how using public transit could make a real differenceâfor you and your community? Making the switch has more benefits than most people realize. Would you be open to hearing some of them? Table 13. This table shows how we construct the system message and user message prompts for the chatbot in our study along side initial static prompts generated by LLM for Recycling and Transit topics. Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning29 E Text-Based Game Implementation The text-based game consists of two LLM-powered parts: game plot generation and a game moderator. The game plot defines the narrative structure and thematic progression of the experience, while the game moderator generates turn-by-turn game content in response to player input during the interaction. E.1 Game Plot Generation The text-based game condition was driven by a pre-generated narrative plot designed to incorporate the same fixed set of persuasive arguments and factual statements used in the other conditions. Table 14 presents the prompt used to generate the game plot. Given a set of topic-specific key arguments (see Appendix B), the prompt guides the language model to produce a structured narrative outline that is later used by the game moderator to generate interactive gameplay. We used GPT-4.1 with a temperature of 1, top-íof 1, and a maximum generation length of 32k tokens, with all other parameters set to default. Due to space constraints, only a truncated example of a generated plot for the recycling topic is shown in Table 15. In addition to the plot, we pre-generated the first turn of the game for all participants using the prompt shown in Table 16. This initial turn served to introduce the narrative context and player role. For this step, we again used GPT-4.1 with a maximum generation length of 2048 tokens. The primary stop sequence was set toâPlayer:âto ensure that generation terminated before the playerâs response. E.2 Game Moderator During gameplay, LLM acted as a game moderator, generating each subsequent turn based on the pre-generated plot, the full interaction history (including prior game and player turns), and the playerâs most recent input. Table 17 presents the prompt used for the game moderator. For the game moderator, we used GPT-4.1 with a maximum generation length of 2048 tokens, temperature 1, top-í 1, and default values for all other parameters. The stop sequenceâPlayer:âwas again used to ensure clean turn boundaries. Similar to the chatbot condition (see Appendix D), content coverage was enforced using an LLM-based fact-checking mechanism. After each game-generated response, the fact checker (Appendix F) evaluated which of the remaining target facts had been covered and updated the list of uncovered facts accordingly. The dynamically updated list of remaining facts (<fact list str>) was included in subsequent game moderator prompts to guide future generations. All game interactions were logged, including turn-level content, timestamps, and the set of facts covered at each turn. F Fact Checker The prompt used for fact checking and coverage tracking is provided in Table 18. The fact checker receives two inputs: (1) the most recent systemâs (chatbot or game) response and (2) a dynamically constructed list of factual statements that have not yet been covered in the interaction. This list is generated by concatenating all remaining facts into a newline-separated string, excluding argument titles and premises. If no facts remain, the input explicitly indicates that all facts have been covered. The fact checker returns the subset of facts identified as present in the response, which is then used to update the remaining fact list for subsequent turns. 30Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz System Message You are a text-based game generator. Your output will be used by another LLM to guide players through an interactive story. The purpose of the game is to teach and persuade players about the provided arguments in a story-driven, engaging way. The game must follow this structure, in order: 1. Game Title 2. Game Premise 3. Opening Scene 4. Acts â one Act per argument, each including: - Setting (where the act takes place) - NPCs (list of named characters with short descriptions and distinct motivations) - Gameplay (missions, tasks, or dialogue that present the argumentâs fact naturally) - Choices (at least 3 explicit ones; also remind the player they may act outside of the listed choices) 5. Finale â 3 endings (e.g., success, partial success, failure). Endings must reflect how the player engaged with the facts, and include epilogue variations depending on which arguments resonated. Important constraints: - The game must feel like a compelling story first, with facts woven naturally into interactions. - NPCs should embody or challenge the facts (not simply recite them). - All facts must appear in the game world, regardless of choices, but the *consequences and tone* should vary with decisions. Mini Example (for style only): Act Example (abbreviated) - Setting: A busy marketplace at the city gates. - NPCs: - Lira the Merchant (worried about trade tariffs). - Captain Dorn (guardsman enforcing the rules). - Gameplay: The player overhears a debate about tariffs, learns that âlowering tariffs increases regional trade by 20%â [fact], and must decide whether to intervene. - Choices: Side with Lira, support Dorn, or propose a compromise. (Player is reminded they may choose another action beyond these options.) User Message Design a story-driven, text-based game that teaches the following key arguments. The game must follow this structure: 1. Game Title 2. Game Premise 3. Opening Scene 4. Acts (one per argument, with Setting, NPCs, Gameplay, and at least 3 explicit Choices + reminder of open-ended choices) 5. Finale (at least 3 endings: success, partial success, failure, with variations depending on which arguments resonated) Key arguments to include: <key arguments> Table 14. This table shows the game plot generator prompt. The only input the prompt requires is <key arguments> (See B), and the output will be used by the game moderator. G Pre-study Baselines Figure 7 shows the pre-study baseline distributions for importance, behavioral intention, and epistemic confidence across both topics. Participants entered the study with relatively high baseline engagement for behavioral intention, leaving limited room for post-study increases. This baseline pattern may help explain the weak and non-differentiated intention changes observed across modes (Figure 5). In contrast, perceived importance showed greater sensitivity to delivery mode, suggesting that conversational interaction may be more effective at increasing issue salience than motivating additional behavioral change in already engaged populations. H Baseline Robustness Analysis To examine whether pre-existing differences in participantsâ attitudes or behaviors influenced the perceived change results, we conducted additional analyses incorporating pre-study baseline measures where conceptually aligned items were available. Specifically, baseline measures of importance and behavioral intention were included as covariates in ordinal logistic regression models predicting perceived change for these constructs. For each outcome, we fit an omnibus Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning31 Generated Recycling Game (Truncated) 1. Game Title The Cycle of Change â 2. Game Premise You are Jay Lin, a young operations manager who moves back to your hometown of Greenfieldâa city striving to transform its wasteful past into a model of sustainable living. As Jay, you must collaborate with community members, businesses, and town officials to build a recycling program that benefits both the environment and the people. Along your journey, youâl uncover the real-world impacts of recycling through interactions and tough decisions, all while balancing innovation and tradition. â 3. Opening Scene Night falls softly over Greenfield as trucks haul away bags of trash under flickering streetlights. You step off the bus, suitcase in hand, and breathe in the crisp, rural air. A faded posterâa tree circled in arrowsâhangs from the old town hall. Your phone buzzes: âWelcome back, Jay! See you at the council meeting tomorrow. Big plans. âMayor Chen.â As you walk toward your childhood home, you wonder: can one person kick-start real change? Or will Greenfield remain lost in old habits? â 4. Acts â ACT 1: Roots and Remnants Setting: The sun-lit forest edge of Greenfieldâs community park, where a sawmill once stood. NPCs: â Maya Rivera â Veteran science teacher; tree-planting activist who mourns the rapid loss of ancient oaks. â Gary Polk â Sawmill owner; pragmatic about production, skeptical about recycling (âsaves pennies, costs timeâ). Gameplay: â At a town tree-planting event, you overhear Mayaâs plea: recycling paper could save Greenfieldâs treesââDid you know one ton of recycled paper saves 17 trees?â Gary counters that production must go on. Choices: 1. Align with Maya and help launch the recycling pilot at the school. 2. Hear Gary out and attend the logging demo. 3. Suggest a compromiseâstart with recycling for school paper waste, monitor results, and review with local businesses. 4. (Or choose your own approachâperhaps propose an entirely new idea or ask questions...) â [. . . ACT 2: Sparks of Change . . . ] [. . . ACT 3: Prosperity in Refuse . . . ] [. . . ACT 4: Currents Below . . . ] [. . . ACT 5: Leading By Example . . . ] â [. . . Finale and Possible Endings . . . ] Table 15. Truncated version of the recycling game plot. three-condition model of the form: Change í ⌠Mode í + PreScore í whereChange í is the ordinal perceived change score for participantí,Mode í is the experimental condition (essay as the reference category, with dummy variables for chat and game), andPreScore í is the corresponding pre-study baseline measure. We additionally conducted pairwise ordinal logistic regressions between each pair of conditions (essay vs. chat, essay vs. game, and chat vs. game) using the same model specification. Baseline measures captured participantsâ initial self-reported importance of the topic and frequency of engaging in the target behavior (see Fig. 7). Including baseline covariates did not qualitatively change the pattern of results reported in Section 6.2. In baseline- controlled models, the chatbot condition remained associated with significantly greater perceived increases in importance relative to the essay condition (íœ=2.18,í= .014), while the game condition did not differ significantly from the essay (íœ= â0.85,í= .38). Pairwise comparisons similarly showed larger perceived increases in importance for chatbot relative to both essay (íœ=2.34,í= .011) and game (íœ=â3.05,í= .003). No significant mode effects emerged for perceived changes in behavioral intention. Overall, coefficient directions and relative differences between modes were consistent with the baseline-free models. No conceptually equivalent pre-study measure was collected for belief in effectiveness. Accordingly, perceived change in belief was analyzed directly, without baseline adjustment. 32Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz System Message You are a game moderator in a turn-based, text-based story game. The game is based on the following plot: <game plot> Your job is to guide the player through the plot one Act at a time, making the experience immersive and interactive. Rules for moderation: - Progress through Acts chronologically**. Never skip an Act. - Do not reference future events or spoil the story. - Keep each response short, vivid, and immersive. - Present the events of the current Act only. - Include interactive options (at least 3). Always add: âor suggest your own moveâ*. - Player decisions must affect how events unfold. - Introduce at most one fact per turn. Make the fact stand out clearly but present it organically through story (dialogue, events, discoveries). - Do not present multiple uncovered facts at once. - Once all facts are covered and the plot reaches the end, end the game naturally. Constraints: - Never mention that you are an AI or assistant. - Avoid phrases like âSure,â âOf course,â etc. - Do not over-explain rules to the player. Facts to integrate: Here are the remaining facts not yet presented: <fact list str> User Message Here is a history of whatâs happened so far: [none yet] Player: (beginning of the game) Start with an Opening Scene, and present it to me as the player. For this turn, do NOT present any facts. Remember to: - Keep your response short. - Give me at least 3 choices in addition to the last choice: âOr suggest your own move.â Game: Generated First Turn: Recycling The bus doors groan shut behind you as you set foot in Greenfield for the first time in years. The night air smells of damp pine and distant smoke, shadows thrown long by lone streetlights. Across the square, the old town hall looms, its crumbling bricks wrapped in a faded poster: a tree circled by green arrows, edges curled and flapping in the light breeze. Your phone buzzesâa message from Mayor Chen: âWelcome back, Jay! See you at the council meeting tomorrow. Big plans.â As you walk along the cracked sidewalk toward your childhood home, memories flickerâfields where factories now stand, creek water running clear. Can you spark transformation in this place, or will Greenfieldâs habits outlast you? You head home, pondering your first move at sunrise: 1. Visit the community park at dawn, where a volunteer event is rumored. 2. Stop by the closed sawmill, remembering the townâs working-class roots. 3. Drop by the corner cafĂ© to catch up with old friends and local gossip. 4. Or suggest your own move. Generated First Turn: Transit Rain beads and streaks across the grand windows of Sandbourneâs City Hall, the drip-drip mirroring your own heartbeat as you steady yourself outside the mayorâs office. In your hands, the thick âSandbourne: 2040â folder grows heavier with every passing minute. The hallway hums with anxious energy: a city worker mutters about last weekâs record smog, two businesspeople grumble over traffic, and a janitor tries to sweep up wet shoe printsâeach conversation weaving into Sandbourneâs bigger story. Your phone buzzes: the councilâs critical transportation vote is in six days. Your challenge isnât just dataâitâs turning numbers into a narrative that can move this city. The mayorâs assistant leans out. âAlex? Weâl see you now. Ready to shape Sandbourneâs future?â What do you do? 1. Enter confidently, ready to introduce bold ideas right away. 2. Pause for a moment, listening to the worries in the hallway for clues about the public mood. 3. Flip open your folder to review your talking points, wanting to be extra prepared before you go in. 4. Or suggest your own move. Table 16. In our experiment, always system initiates the conversation. The prompt for generating first turn of the game is brought here, along side recycling and transit gamesâ generated first turn. Overall, these robustness analyses suggest that the observed between-mode differences in perceived change are not driven by pre-study baseline variation and that baseline differences did not materially affect the studyâs conclusions. Received 18 February 2026 Games That Teach, Chats That Convince: Comparing Interactive and Static Formats for Persuasive Learning33 System Message You are a game moderator in a turn-based, text-based story game. The game is based on the following plot: <game plot> Your job is to guide the player through the plot one Act at a time, making the experience immersive and interactive. Rules for moderation: - Progress through Acts chronologically**. Never skip an Act. - Do not reference future events or spoil the story. - Keep each response short, vivid, and immersive. - Present the events of the current Act only. - Include interactive options (at least 3). Always add: âor suggest your own moveâ*. - Player decisions must affect how events unfold. - Introduce at most one fact per turn. Make the fact stand out clearly but present it organically through story (dialogue, events, discoveries). - Do not present multiple uncovered facts at once. - Once all facts are covered and the plot reaches the end, end the game naturally. Constraints: - Never mention that you are an AI or assistant. - Avoid phrases like âSure,â âOf course,â etc. - Do not over-explain rules to the player. Facts to integrate: Here are the remaining facts not yet presented: <fact list str> User Message Here is a history of whatâs happened so far: <history> Player: <user input> Keep your response very short Start every response by narrating how the playerâs last action affects the current scene, the NPCs, or the story world. Remember: - Do not skip ahead. - Keep narration short. - Integrate at most one remaining fact, and make it stand out naturally in the story. - Provide at least 3 explicit choices, plus âor suggest your own move.â - You may end the game naturally and show the epilogue if all facts are covered. Game: Table 17. Prompt employed to generate game responses. System Message You are an assistant designed to check which facts are covered by a piece of text. User Message Below is a narrative chunk followed by a list of factual statements. Your task is to return a list of the facts that are covered in this text. Respond with a Python list of the indices of the covered facts â Text: <text> Facts: <fact list str> Which facts are covered in the text? Return their indices in a Python list. Table 18. Prompt used for fact coverage checking in the chatbot and the text-based game conditions. The prompt identifies which predefined factual statements are covered by the systemâs most recent response. We used the OpenAIo4-minimodel with a maximum generation length of 2,048 tokens and default decoding parameters. 34Seyed Hossein Alavi, Zining Wang, Shruthi Chockkalingam, Raymond T. Ng, and Vered Shwartz Fig. 7. Pre-study baseline characteristics of participants across recycling and public transit domains. Top row shows recycling measures: (left) distribution of personal importance ratings (1-5 scale, mean=3.95, median=4.0), (center) frequency of recycling behavior (Never to Always, median=often), and (right) confidence in knowledge about recycling (1-5 scale, mean=3.14, median=3.0). Bottom row shows transit measures: (left) importance of public transit use (1-5 scale, mean=3.43, median=4.0), (center) typical weekly commute days outside home (0 to 5+, median=5+), and (right) confidence in knowledge about transit benefits (1-5 scale, mean=3.33, median=3.0). Red and blue dashed lines indicate mean and median respectively. These baseline measures represent participantsâ pre-intervention attitudes and behaviors across both intervention topics.