Paper deep dive
Mixed-Agent Museum Tour Guide Design Improves Gendered Learning Outcomes and Visitor Preferences
Annette M. Masterson, Wonse Jo, Helena C. Sieh, Lionel P. Robert,, Dawn Tilbury
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/18/2026, 1:05:17 AM
Summary
This paper introduces and evaluates a mixed-agent museum tour guide system that integrates a physical robot (Toyota HSR) with a projected virtual avatar. Through a within-subjects study with 30 participants, the system was tested across three conditions: a single robot with cheerful dialogue, a mixed-agent with storytelling dialogue, and a mixed-agent with bantering dialogue. Results show that while engagement and quality of experience remained consistent across conditions, learning performance significantly improved for female participants in the mixed-agent bantering condition. Participants universally preferred the mixed-agent configuration, highlighting the potential of dyadic conversational styles to enhance gender-specific learning outcomes in human-robot interaction.
Entities (10)
Relation Signals (10)
Mixed-Agent System → comprises → Virtual Avatar
confidence 95% · projected virtual avatar generated by a laser projector
Mixed-Agent System → comprises → Toyota HSR
confidence 95% · comprising a physical service robot platform (Toyota HSR) and a projected virtual avatar
Mixed-Agent System → improves → Learning performance
confidence 95% · mixed-agent conditions improved learning performance for female participants
Learning performance → issignificantlyimprovedfor → Female Participants
confidence 95% · Learning performance revealed a significant gender-moderated difference: the mixed-agent conditions improved learning performance for female participants
Storytelling Dialogue → isastyleof → Mixed-Agent System
confidence 90% · mixed-agent team featuring storytelling conversational style
Bantering Dialogue → isastyleof → Mixed-Agent System
confidence 90% · mixed-agent with bantering dialogue
Mixed-Agent System → ispreferredby → Participants
confidence 90% · participants reported a greater preference for mixed-agent teams regardless of gender
Mixed-Agent System → utilizes → ArUco Markers
confidence 90% · utilized ArUco fiducial markers affixed to the safety helmet
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Robots are increasingly integrated into everyday contexts, including museums, where they can both entertain and educate visitors. To enhance visitor experience and engagement, we present a novel mixed-agent tour guide system that combines a physical robot with a projected virtual agent that actively participates in the tour through conversation and interaction, achieving the interaction richness of two mobile agents from a single platform. We validate the system through a within-subjects study with 30 participants to assess engagement, quality of experience, and learning performance. Participants experienced different conversational styles and agent configurations, and data were collected via surveys, behavioral sensors, and interviews. Results showed that engagement and quality of experience remained consistent across conditions. Learning performance revealed a significant gender-moderated difference: the mixed-agent conditions improved learning performance for female participants. This suggests that the proposed dyadic conversational style in this paper influenced learning performance differently by gender. Nonetheless, in interviews, participants reported a greater preference for mixed-agent teams regardless of gender, citing interaction as a key factor in their experience.
Tags
Links
- Source: https://arxiv.org/abs/2607.14468v1
- Canonical: https://arxiv.org/abs/2607.14468v1
Trouble viewing inline? Open PDF directly →
Full Text
43,122 characters extracted from source content.
Expand or collapse full text
Mixed-Agent Museum Tour Guide Design Improves Gendered Learning Outcomes and Visitor Preferences Annette M. Masterson 4,† , Wonse Jo 5,†,∗ , Helena C. Sieh 1 , Lionel P. Robert, Jr. 1,4 , and Dawn Tilbury 1,2,3 Abstract— Robots are increasingly integrated into everyday contexts, including museums, where they can both entertain and educate visitors. To enhance visitor experience and engagement, we present a novel mixed-agent tour guide system that combines a physical robot with a projected virtual agent that actively participates in the tour through conversation and interaction, achieving the interaction richness of two mobile agents from a single platform. We validate the system through a within- subjects study with 30 participants to assess engagement, quality of experience, and learning performance. Participants experienced different conversational styles and agent configura- tions, and data were collected via surveys, behavioral sensors, and interviews. Results showed that engagement and quality of experience remained consistent across conditions. Learning per- formance revealed a significant gender-moderated difference: the mixed-agent conditions improved learning performance for female participants. This suggests that the proposed dyadic con- versational style in this paper influenced learning performance differently by gender. Nonetheless, in interviews, participants reported a greater preference for mixed-agent teams regardless of gender, citing interaction as a key factor in their experience. I. INTRODUCTION Museums combine entertainment and education, but they often rely on static exhibits or audio tours that limit ac- cessibility and engagement [1]. To enhance experiences, tour guide robots have been proposed as an engaging and financially feasible alternative [1]. Developing robot tour guides capable of delivering immersive, socially responsive interactions that can be evaluated with comprehensive behav- ioral and self-report metrics remains an ongoing challenge. In contrast to the dual-robot frameworks discussed by Velentza et al. [2], our design employs a mixed-agent approach that integrates a gimbal-mounted projector with inverse kinemat- ics and path planning, enabling a projected virtual agent to dynamically reposition and actively participate in the tour through conversation alongside the physical robot. This achieves the interaction richness of two mobile agents from a single robotic platform. We validate that this mixed-agent system, especially when paired with different conversational 1 Department of Robotics, University of Michigan, Ann Arbor, MI 48109, United States 2 Department of Mechanical Engineering, University of Michigan, Ann Arbor, MI 48109, United States 3 Department of Electrical Engineering and Computer Science, University of Michigan, Ann Arbor, MI 48109, United States 4 School of Information, University of Michigan, Ann Arbor, MI 48109, United States 5 A+HRI Laboratory, Department of Information and Telecommuni- cation Engineering, Incheon National University, Incheon, South Korea. jow@inu.ac.kr †: equal contribution ∗: corresponding author styles, yields superior learning performance for women and increased visitor preference over a single-robot configuration. Our design replaces the headset requirement of VR tours with a unified robotic system [3], offering a scalable solution without the massive investment in individual hardware, mim- icking human-guided tours which may not be affordable for all museums [1]. To determine the value of our system on en- gagement and learning performance, we conducted a within- subject user experiment with three conditions: 1) a single physical robot with cheerful dialogue; 2) a mixed-agent (physical plus virtual) with storytelling dialogue; and 3) a mixed-agent with bantering dialogue. The scenario represents a robot-guided tour in an indoor environment (e.g., a museum gallery). We measured engagement and learning performance through behavioral, survey, and interview methodologies. This paper makes several contributions. First, it introduces and evaluates a mixed-agent (physical and virtual) tour- guide system with multiple conversational styles. Second, we provide preliminary evidence of gender-moderated learning benefits for mixed-agent teams, qualitative feedback on the preference for mixed-agent teams, and an observed tension between engagement and learning. These findings advance human-robot interaction within multi-modal systems. This paper examines how robots and mixed-agent teams influence engagement and learning outcomes. We introduce a multi-application tour guide system with built-in behav- ioral metrics with the potential for real-time engagement assessment. To conclude, we validate the system’s impact on learning and engagement through a targeted user experiment. I. BACKGROUND A. Entertaining and Learning with Tour Guide Robots Static displays and audio museum tours have long been the industry standard; however, autonomous systems offer a level of dynamic interaction that correlates with superior visitor retention and satisfaction compared to traditional, non-adaptive methods [4], [5]. Social signals–facial expres- sions, movement, and communicative features–consistently increase perceived sociability and likability of museum robots, particularly when designs emphasize personable traits and positive affect [6]. While not often addressed in tour guide work [7], learning can be a subjective or gendered process potentially influenced by social conditioning [8]. Considering learning with robot guides, more engaging robots can be viewed more positively [9], [10]. Therefore, demographic differences should be acknowledged in system evaluations to better inform future development. arXiv:2607.14468v1 [cs.RO] 16 Jul 2026 Recent work integrates tour guide robots [4] and aug- mented reality (AR) agents [11] to amplify engagement in learning environments. For example, embodied AR agents and emotionally expressive robots outperform audio-only presentations on measures of preference and perceived ex- pressiveness [5], [7]. While autonomous or semi-autonomous movements help robots mimic the experience of a human tour guide [12], their true advantage lies in their lack of constraint. Unlike human guides, robotic systems can seam- lessly integrate multi-modal interactions, creating novel and engaging experiences that transcend traditional tours. Still, limited work has been done on mixed-agent teams and their impact on engagement and learning performance, particularly with ambulatory robot guides. B. Mixed-Agent Engagement Physical robots can enhance engagement in informal learn- ing environments [5], [7]. Similarly, virtual robots, either projected onto the wall [13], as an augmented reality agent [14], or through a mobile application [15], have been found to increase learning performance. In response to enhanced learning and engagement, we designed a tour experience with a combination of agents and integrating multi-modal interactions to expand tour guide offerings. Integrating augmented reality bots that can ma- nipulate video clips with a physical robot allows for more effective interactions with participants [16]. Additional work reports that mixed configurations can sustain engagement [2], [17] and, in some cases, improve learning performance by distributing content across agents [17]. However, Velentza et al. [2], [18] found a complex relationship between robot con- figurations and learning performance. A single-robot setup improved knowledge recall, but a two-robot setup led to lower retention of the tour content, as assessed in the post- experiment quiz. Curiously, despite the diminished learning, participants in the two-robot condition rated their experience as more pleasant. This highlights the importance of testing learning performance across different robot configurations to balance learning performance [19], particularly with a novel mixed-agent system. In this study, we propose the following hypotheses to guide our design and evaluation: H1: Mixed-agent teams will have increased quality of experience, engagement, and learning performance over a single robot. H2: Mixed-agent teams bantering conversational style will have higher quality of experience, engagement, and learning performance over the storytelling style. I. MIXED-AGENT TOUR GUIDE SYSTEM As illustrated in Fig. 1, we propose a novel mixed- agent team comprising a physical service robot platform (Toyota HSR) and a projected virtual avatar generated by a laser projector mounted on the HSR’s back (see Fig. 2). The system’s key technical contributions include: (1) multi- location virtual agent positioning through gimbal-mounted projection, enabling the virtual agent to appear at exhibit locations, (2) synchronized multi-agent coordination across <HSR’s head display> <Interactive button> <Virtual Avatar> Fig. 1: Details of the robot platform, including the robot’s head display and forward-facing depth camera (top left), the interactive shoulder button (bottom right), the rear-facing RealSense Depth Camera (D345i, bottom left), and examples of projections (e.g., image, video, and our virtual avatar). RealSense D435i Fig. 2: Customized projector to augment the virtual avatar. processors (desktop RTX 4090, Jetson Nano, Raspberry Pi 4) to coordinate speech and movements of both agents, and (3) integrated behavioral measurements embedded within the robot’s control loop, enabling continuous data collection (reaction times, physical distance, head orientation). The overall system of the mixed-agent team communicates via the Robot Operating System (ROS) [20], which enables data sharing across a distributed system. The system includes three key ROS packages: Tour Script, which enables a mixed- agent team to explain the exhibits; Navigation, which drives to predefined exhibit locations; and Engagement Tracking, which measures participants’ distance, head pose, and reac- tion time in response to the robot’s prompts. Distinct voices were chosen for the physical and virtual agents [21]. To ensure audio consistency across participants, generated audio files were then downloaded and uploaded directly to the robot’s system. Start/End 2 3 4 5 6 1 : Session 1 (12): Session 2 (34): Session 3 (56) (a) [Main Experiment] StartIntroduction C 1 C 2 C 3 Survey InterviewEnd *Random order (b) Fig. 3: Overview of the experimental setup and protocol: (a) Experiment environment and (b) Procedures for the user experiment, in which the order of the condition (C) was randomized. A. Physical Robot Platform (HSR) The HSR is equipped with the Navigation package that uses 2D Lidar and two head-depth cameras (front and rear views), enabling a safe navigation system that automat- ically avoids obstacles. We implemented the engagement tracking package on the HSR that maintains eye contact with participants throughout the tour [22]. The package includes PID-controlled pan/tilt motors on the HSR’s head continuously track helmet-mounted ArUco markers worn by participants, keeping them centered in the head camera frame for face-to-face orientation. To robustly and directly estimate the 6D pose of the human participant’s head during the interaction, we utilized ArUco fiducial markers affixed to the safety helmet rather than relying on markerless vision-based frameworks like MediaPipe. This tracking capability not only enables us to measure the distance between the participant and the HSR, but also to calculate the head reaction time (RT) based on the robot prompts during the tour. B. Virtual Avatar The virtual avatar (VA) is designed as a cartoon-style vehicle with anthropomorphic features, a design choice in- tended to enhance user engagement [4] and mitigate potential discomfort [23]. We chose not to include a humanoid virtual agent in order to better distinguish the physical and virtual entities as distinct agents. The VA is projected via a gimbal- mounted projector, actuated by two servo motors providing a pan range of±55 ◦ and a tilt range from−90 ◦ to +45 ◦ . To ensure precise visual delivery, projection locations are pre- programmed relative to the robot’s pose at each tour stop. Gimbal orientations are computed using inverse kinematics, while smooth motion trajectories are planned via convex quadratic programming. The VA’s animations are rendered using the PyGame framework, supporting frame-by-frame playback of animated sprites at a refresh rate of approximately 15 Hz. For realistic interaction, the avatar’s mouth movements are synchronized in real-time with OpenAI’s text-to-speech engine [21]. In addition to the avatar, the projector serves as a versatile display for supplementary visual media, such as images and videos, which are projected onto walls during exhibit explanations and onto the floor to assist with navigation. IV. VALIDATION EXPERIMENT To validate the efficacy of the proposed mixed-agent team, we conducted a user study focusing on how this configuration influences participants’ engagement, quality of experience, and learning performance. During the experiment, subjects participated in a robot-led museum tour where the mixed- agent team delivered interactive presentations about various exhibits. This study was conducted with the approval of the Institutional Review Board (#IRB-HUM00278018). A. Participants For this user experiment, we recruited 30 participants (14 female and 16 male; mean age = 32, SD = 13) based on a post-hoc power analysis conducted using GPower [24], with an effect size of f = 0.24, a significance level (α) of .05, and a power of .80. All participants confirmed their eligibility to participate in the study as follows: (1) 18 years old, (2) have normal or corrected vision and hearing, and (3) have the ability to follow the robot for an hour. B. Study Design The experiment employed a within-subjects design com- prising three conditions (60 minutes total), presented in a random order to minimize practice effects [25]. The condi- tions differ based on the type of information projected by the robot’s projector or the style of conversation within the mixed-agent team. Factual verbal (information about the ex- hibits) and visual (videos and images related to the exhibits) content remained identical across all conditions. There are three conditions: C 1 ) a single physical robot with cheerful dialogue; C 2 ) a mixed-agent team featuring storytelling conversational style, with the agents not engaging with each other; and C 3 ) a mixed-agent team interacting humorously in bantering conversation style with each other. The baseline consisted of a single robot speaking in a cheerful manner; Condition 1 presented the same videos as the other conditions but was not accompanied by a virtual agent. Experiments were conducted in a museum-like environ- ment (see Fig. 3a) with participants wearing a helmet- mounted camera for tracking their behavior. Exhibits for this study included six distinct academic posters: departmental history, bipedal robotics, collaborations, robot training facil- ities, autonomous vehicles, and aquatic integration. The scripts in each exhibit were written by researchers and then fed into ChatGPT-4 to generate ideas for cheerful and humorous styles, designated the storytelling–without dialogue between agents–and bantering–humorous dialogue between agents–conditions, respectively. The purpose of the stylistic change is to determine whether communication style helps improve engagement and learning performance. Researchers compiled scripts to align with the conditions, then presented them in randomized order. As an example, scripts for the UGVs exhibit included: Single HSR robot: “The robot was being trained as a UGV — that means Unmanned Ground Vehicle. No driver needed.” Storytelling: HSR robot– “The robot was being trained as a UGV, that means Unmanned Ground Vehicle. No driver needed.” virtual agent– “I’m Scout, a UGV, but I’m way more updated than these bots in 1992.” Bantering: HSR robot–“The robot was being trained as a UGV — that means Unmanned Ground Vehicle. No driver needed. Wait a minute, are you a UGV, Scout?” virtual agent– “That’s right, Remy! But I’m way more updated than these bots in 1992.” C. Survey Procedures Fig. 3b details the overall procedure used in this study. The introduction survey included the pre-quiz for learning performance and demographic questions: age, gender, eth- nicity, and education. The introduction survey also asked participants about their prior experience and familiarity with service robots. Participants completed two questionnaires to assess their engagement and quality of experience at the end of each condition: 1) User Engagement Short Scale (UES) [26], and 2) Quality of Experience (QoE) scale [12] to assess satisfaction. In addition to the self-reported questionnaires, Learning Performance (LP) was measured by comparing pre- experiment quiz scores with post-session quiz scores at the end of each condition, using consistent questions and an- swers across conditions. Quizzes were written by researchers from the information in the scripts. Quiz questions were inputted into ChatGPT-4 to evaluate the complexity. Quizzes were finally assessed during pilot tests to balance difficulty and achievability. Quiz answers were multiple choice and were answered in the exhibits. An example of a quiz question is "What is the maximum speed attainable by the bi-pedal treadmill?" D. Behavioral Measures Our mixed-agent system enables real-time tracking of par- ticipant behavior throughout the tour (Fig. 1). As illustrated in Fig. 4, the behavioral data included: (1) Physical distances – mean distance between participant and robot measured in meters, indicating spatial comfort and physical distance pref- erences; (2) Cognitive reaction time (RT) – mean response latency in seconds from robot’s prompt to participant’s actual actions, measured four times during experiments; and (3) Head angle – angular difference between participant and robot heading directions (degrees), measured in the global frame. Physical distance Elevation Angle HSR Head Angle Measurement (Top View) Angular difference between robot & participant headings (measured in global frame) y x Global Frame <Head> Θ =Head Angle <HSR> GU1 Fig. 4: Behavioral measurement setup. ArUco marker tracking provides: (1) physical distance (green line, meters), (2) elevation angle (red arc, degrees), and (3) head angle – orientation difference between robot heading (yellow) and participant heading (purple) (green arc, θ in degrees). For the physical dimensions, the robot uses both front and rear cameras mounted on its head to measure the distance and angle between the robot’s location and the participant’s head location. This physical distance is not linked to gaze [27], but physical distance. Once the robot detects the ArUco markers attached to the participant’s helmet, it estimates their position in the global map and calculates (1) the distance from the robot to the participant and (2) the bearing angle from the robot’s forward direction to the participant’s location. We also recorded two cognitive RT metrics: (1) Head RT: the time it takes for the participant to turn their head to a specific poster after the robot gives a directional prompt and (2) Button RT: the time between the button prompt and the participant’s push of the button to initiate video playback on the robot’s head-mounted display. E. Interview Procedures To gather in-depth data, the experiment concluded with brief semi-structured interviews. Questions were designed to align with the quantitative scales, such as asking about the most interesting components of the tour to relate to the UES and about the tour experience to relate to QoE. This provided context for the survey and behavioral data. All 30 participants completed the interview, with sessions ranging from 4–14 minutes and an average of 7 minutes. We used randomly generated four-digit participant identification numbers (e.g., P1234) for anonymity. Interviews transcribed using Otter.AI, and transcripts were reviewed for accuracy. Data analysis was conducted inductively to allow themes to emerge directly from the data by assessing each interview comment as a data point [28], [29]. Codes were developed through a reflective thematic analysis, an iterative process that identifies themes across the aggregate data [30]–[32]. Through the data analysis, significant themes emerged fo- cused on knowledge development, robotic capabilities and design, interaction, and mixed-agent preference. V. RESULTS Our user study (N = 30) revealed that our mixed-agent team with a bantering conversational style can enhance learning performance for women, and qualitative results il- lustrate visitor preference across participants. We also reveal a negative correlation between engagement and learning performance. Behavioral results demonstrate women stand closer to the robot, particularly during the bantering condi- tion. Cognitive reaction times illustrate high compliance with prompts. A. Survey Results We evaluated three measures: (1) Sum of Engagement (UES) scores, (2) Mean of Quality of Experience (QoE) scores, and (3) Learning Performance (LP; post–pre quiz difference normalized to a 0–1 range). Repeated-measures ANOVA (rmANOVA) was conducted via Pingouin [33]. From the rmANOVA results, we found a significant effect of condition on LP (F(2, 58) = 3.95,p = 0.025,η 2 p = 0.12), where C 3 (M = 0.77,SD = 0.25) outperformed C 1 (M = 0.61,SD = 0.31). Post-hoc pairwise comparisons using t- tests revealed that participants in C 3 achieved significantly higher learning scores than C 1 , with a mean difference of 0.16 (t(29) = −2.52,p = 0.018). While the difference between C 1 and C 2 (M = 0.75,SD = 0.25) was marginal (t(29) = −2.04,p = 0.051), no significant difference was observed between C 2 and C 3 (t(29) = −0.32,p = 0.752). However, no significant differences were found for Sum UES (p = 0.479) and Mean of QoE scores (p = 0.614), indicating that mixed-agent teams improved overall learning performance, but not perceived engagement or quality of experience. To examine gender differences, participants were divided into 16 males and 14 females, and rmANOVA was conducted on each group. A strong condition effect emerged for LP in the female group (F(2, 26) = 3.86,p = 0.034,η 2 p = 0.23), where C 3 outperformed C 1 by a mean difference of 0.21 (p = 0.028), as illustrated in Fig. 5. Therefore, we can conclude that the mixed-agent team outperformed the single robot, primarily in learning perfor- mance. While Mean QoE showed a positive numerical trend without reaching statistical significance, gender-specific ben- efits emerged: males showed a preference for the storytelling dialogue (C 2 ) for experience, while females achieved signif- icantly higher learning gains in the bantering dialogue (C 3 ). This suggests that while mixed agents generally enhance learning, dialogue effectiveness may be gender-dependent. B. Behavioral Results Behavioral analyses included physical distance, head an- gle, and RT for button presses and head rotations, as prompted by the robot. Participants demonstrated high com- pliance with interactive prompts: button presses achieved 100% response rate, while head prompt achieved 68% response rate, RT showed no significant effects between conditions. C 1 C 3 C 2 0.25 0.00 0.25 0.50 0.75 1.00 1.25 1.50 Learning Performance (LP) 0.7 0.5 0.80.8 0.8 0.7 ** Gender Male Female Condition Fig. 5: Gender comparison of normalized Learning Performance (LP) across conditions using split violin plots, which is post–pre quiz difference normalized to a 0–1 range. Markers indicate group means; brackets with asterisks show significant pairwise differences (∗ : p < 0.05) for female participants between C3 (bantering, mean = 0.77) and C1 (single robot, mean = 0.5). No significant differences for males. Head Reaction Time Physical Distance 0.0 0.2 0.4 0.6 0.8 1.0 Normalized Value Condition 1 - Male Condition 2 - Male Condition 3 - Male Condition 1 - Female Condition 2 - Female Condition 3 - Female Video Reaction Time Metric Fig. 6: Behavioral measures by gender across conditions. Error bars repre- sent standard error of the mean (N=16 males, 14 females). Males maintained significantly greater physical distance across all conditions (p < .001). As shown in Fig. 6, male participants consistently main- tained greater distances from the robot than female partic- ipants across all conditions. Participants stood closest to the robot during C 3 (bantering) and maintained greatest distance during C 2 (storytelling), suggesting the interactive bantering style encouraged closer physical engagement. Post- hoc pairwise comparisons revealed significant distance dif- ferences between all conditions: C 2 vs. C 1 (MD = 0.0285, p < .001), C 2 vs. C 3 (MD = 0.0161, p < .001), and C 1 vs. C 3 (MD = 0.0124, p < .001). As shown in Fig. 7, gender was negatively correlated with head angle (r = −0.31, p = .003), suggesting that male participants tended to exhibit larger directional changes relative to the robot than female participants. Head angle was also moderately positively correlated with Mean QoE (r = 0.26, p = .014), suggesting that participants who engaged in greater head movement reported higher perceived experience quality. In addition, physical distance was pos- itively correlated with head angle (r = 0.29, p = .005), indicating coordinated spatial and directional adjustment during interaction. C. Correlation Results We examined correlations between demographic, survey, and behavioral variables using Pearson correlation analyses [33]. Demographic variables included gender (1 = male, 2 Gender Sum UES Mean QoE LP Head angle Gender Robot experience Sum UES Mean QoE LP Physical distance Reaction time Head angle -0.16 0.10-0.03 0.01-0.150.04 -0.14 0.30 ** -0.37 *** -0.02 -0.48 *** 0.02-0.010.160.14 0.090.02-0.160.05-0.14-0.10 -0.31 ** -0.04 0.12 0.26 * 0.08 0.29 ** -0.12 0.6 0.4 0.2 0.0 0.2 0.4 0.6 0.8 Robot experience Reaction time Reaction time Fig. 7: Heatmap of Pearson correlation coefficients among study vari- ables. The values in each cell represent correlation coefficients (γ) with significance indicated by asterisks ( ∗ : p < 0.05, ∗ : p < 0.01, ∗ : p < 0.001). = female) and robot experience (0 = no prior experience, 1 = prior experience). Survey measures comprised Sum UES, Mean QoE, and LP (post–pre quiz difference). Several statistically significant correlations of moderate or greater magnitude (|r|≥ .25, p < .05) were identified. Gender was negatively correlated with physical distance (r = −0.48, p < .001) and head angle (r = −0.31, p = .003), indicating that male participants tended to maintain greater distance from the robot and exhibit larger directional shifts. LP was negatively associated with UES (r = −0.37, p < .001), whereas robot experience was positively associated with LP (r = 0.30, p = .004). Mean QoE demonstrated a moderate association with head angle (r = 0.26, p = .014). Physical distance was also positively correlated with head angle (r = 0.29, p = .005), suggesting coordinated spatial and directional adjustments. D. Interview Results Interviews provided deeper context for the quantitative findings, confirming a preference for mixed-agents (17 out of 29 participants favored the team, with one neutral). Par- ticipants noted that the collaborative dynamic—characterized by jokes (P7273) and dialogue (P6342)–emulated human-to- human communication, making the experience feel "closer to a human being than a robot" (P8237). This multi-modal delivery provided "fresh perspectives" (P7198) by shifting attention between agents, a feature deemed essential in to- day’s "attention economy" (P3470). The ability to direct their attention to the robot’s screen between exhibit explanations was also an interactive highlight, with participants enjoying the "multiple ways that it’s showing" information (P6496), helped them answer the quiz (P1025). Regarding gender, while no explicit preference was found, female participants particularly highlighted the "creative" and "cute" aspects of the team (P5086), noting that the virtual avatar’s animated eyes made interaction "easier" and more engaging (P2990). Female participants found the "talking" (P5781) and "interaction" (P7198) between the agents increased their engagement with the content. Two primary design challenges emerged: speed and nav- igation. Although the robot’s speed (0.8 km/h) ensured a sense of security (P8871), participants found it significantly slower than a natural human pace, forcing them to moderate their pace (P7601). These findings suggest that future iter- ations must balance cautious, safe navigation with a more natural walking speed to enhance real-world viability. VI. DISCUSSION A. Summary of Results Perceptions of engagement and quality of experience re- mained consistent across conditions in the survey results. Hypothesis 1 was partially supported with learning perfor- mance increasing for women in mixed-agent teams. Hypoth- esis 2 partially supported as the bantering style increased learning performance for women. We also found a negative correlation between engagement and learning performance, consistent with prior work [7], indicating that greater per- ceived engagement does not necessarily translate to improved learning even in real-world settings. The qualitative interview data offered deeper insights, with participants reporting a strong preference for mixed-agent systems. This favor was at- tributed to the robot’s multi-modal capabilities; features like the projector and shoulder button transformed the interaction that felt more dynamic than traditional setups. Behavioral data revealed significance in head angle and physical distance. Head angle showed no significant con- dition effects but was significantly correlated with gen- der, QoE, and reaction time. Notably, females tended to stand closer to the robot than males, contradicting previous research [34] potentially due to contextual differences in robotic appearance [35] or politeness [36]. B. Mixed-Agents and Learning Performance Our results show that mixed-agent effectiveness is highly demographic-dependent; specifically, female participants achieved higher learning performance. This may be due to an overarching preference for talkative robots, particularly when handling more tasks [10], as collaborative environments may be more aligned with women’s interaction style [8]. Women also tend to rate robots that are more personable and have higher attention switches as better for learning [9]. We believe our mixed-agent system may mimic previous research attuned to women’s learning styles and preference for a more collaborative, talkative, and personable robot experience. However, additional analyses should be con- ducted to reaffirm these findings on conversational style and mixed-agent preferences. Our findings suggest a need for personalized conversational styles to match gendered prefer- ences. While the bantering style did not universally improve all metrics, qualitative data confirmed that the mixed-agent configuration offers a uniquely engaging experience beyond other traditional tours [5]. Thus, this study validates the trade-off between visitor enjoyment and learning performance in a real-world context. Although dual-agent setups can hinder information retention while increasing pleasure [7], our gender findings and inter- view data, in particular, demonstrate that a properly balanced mixed-agent system can enhance both simultaneously. While guided tours reduce cognitive load compared to audio tours [37], the mental demand remains a challenge in robot-led interactions and must be addressed. Managing this tension requires a dynamic interaction design that recalibrates its strategy throughout the tour to maintain an optimal balance between entertainment and education. C. Implications for Design and Implementation Our design enhancements offer three key implications for social robotics: (1) Gendered differences necessitate personalized, inclusive content [6]. Our mixed-agent and behavioral metric system can become a platform for real- time adjustments to content. Systems should utilize real-time metrics—like reaction times—to dynamically adjust infor- mation flow, using learning checks during high-engagement segments to mitigate cognitive overload. Removing the de- pendency on ArUco markers would allow behavioral metrics to be integrated more naturally into real-world environments; (2) Multi-modal features should align with user expecta- tions to avoid the "uncanny valley" [38] or an "expectation gap" [39]. Designers should harmonize agent capabilities, expressivity, and timing to reduce dissonance; (3) Robot navigation must mimic human social conventions. While 0.8 km/h provides safety, it forces users to unnaturally moderate their pace. Future designs should integrate transparent safety cues [40] while striving for a more natural walking speed that maintains the fluidity of human guidance. VII. CONCLUSION AND FUTURE WORK We developed a novel mixed-agent museum tour guide system that achieves the interaction richness of mixed-agents through a single robotic platform and a gimbal-mounted projector. Our user study (N = 30) revealed that this multi- modal configuration significantly enhances visitor preference and improves learning performance, particularly for women, while identifying a subtle trade-off between engagement and knowledge retention. Although the study was limited by its controlled lab setting and the subjective nature of humor [41], it underscores the value of dynamic, ambulatory robot teams. Future work should focus on: (1) integrating LLM-based personalization, allowing for choice in mixed-agent and personality styles, (2) balancing engagement with learning performance by testing cognitive load, and (3) validating the mixed-agent system’s navigation in real-world, high-traffic museum environments to accommodate diverse user demo- graphics and abilities. Future research should leverage these real-time behavioral metrics to further inform tour guide systems, creating a more robust framework for quantifying engagement with objective data across increasingly complex social environments. ACKNOWLEDGMENT This work was supported in part by funding from the Uni- versity of Michigan (U-M), U-M’s Undergraduate Research Opportunity Program (UROP), and U-M’s undergraduate capstone students (ME450 and ROB450). We thank them for their invaluable skills and time in developing the mixed- agent tour guide system. REFERENCES [1] L. Garello, F. Cocchella, A. Sciutti, M. Catalano, and F. Rea, “Next- gen museum guides: Autonomous navigation and visitor interaction with an agentic robot,” arXiv preprint arXiv:2507.12273, 2025. [2] A.-M. Velentza, N. Fachantidis, and S. Pliasa, “Which one? choosing favorite robot after different styles of storytelling and robots’ conver- sation,” Frontiers in Robotics and AI, vol. 8, p. 700005, 2021. [3] S. Aoki, N. Itabashi, R. Adha, S. Sameshima, Y. Kinoshita, and T. Kotoku, “Mixed reality guided museum tour: digital enhancement of museum experience,” in 2023 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW).IEEE, 2023, p. 809–810. [4] U. Maniscalco, A. Minutolo, P. Storniolo, and M. Esposito, “Towards a more anthropomorphic interaction with robots in museum settings: An experimental study,” Robotics and Autonomous Systems, vol. 171, p. 104561, 2024. [5] T. Natori, T. Iio, Y. Yoshikawa, and H. Ishiguro, “Impact of table- top robots on questionnaire response rates at a science museum,” International Journal of Social Robotics, vol. 17, no. 4, p. 729–742, 2025. [6] S. Stals, L. Baillie, and F. Jacob, “A robot tour guide for people with sight loss in a robotic assisted living environment,” in 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 2025, p. 1644–1649. [7] A.-M. Velentza, D. Heinke, and J. Wyatt, “Museum robot guides or conventional audio guides? an experimental study,” Advanced Robotics, vol. 34, no. 24, p. 1571–1580, 2020. [8] J.-P. Onnela, B. N. Waber, A. Pentland, S. Schnorf, and D. Lazer, “Using sociometers to quantify social interaction patterns,” Scientific reports, vol. 4, no. 1, p. 5604, 2014. [9] O. Engwall, J. Lopes, and A. Åhlund, “Robot interaction styles for conversation practice in second language learning,” International Journal of Social Robotics, vol. 13, no. 2, p. 251–276, 2021. [10] C. Wienrich, M. E. Latoschik, and D. Obremski, “Gender differences and social design in human-ai collaboration: Insights from virtual cobot interactions under varying task loads,” in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 2024, p. 1–8. [11] Q. Gan, Z. Liu, T. Liu, Y. Zhao, and Y. Chai, “Design and user experience analysis of ar intelligent virtual agents on smartphones,” Cognitive systems research, vol. 78, p. 33–47, 2023. [12] S. Rosa, M. Randazzo, E. Landini, S. Bernagozzi, G. Sacco, M. Pic- cinino, and L. Natale, “Tour guide robot: a 5g-enabled robot museum guide,” Frontiers in Robotics and AI, vol. 10, p. 1323675, 2024. [13] H. Ro, J.-H. Byun, I. Kim, Y. J. Park, K. Kim, and T.-D. Han, “Projection-based augmented reality robot prototype with human- awareness,” in 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 2019, p. 598–599. [14] K.-F. Lee, Y.-L. Chen, H.-C. Hsieh, and K.-Y. Chin, “Application of intuitive mixed reality interactive system to museum guide activity,” in 2017 IEEE International Conference on Consumer Electronics-Taiwan (ICCE-TW). IEEE, 2017, p. 257–258. [15] S. Pliasa, A. M. Velentza, A. G. Dimitriou, and N. Fachantidis, “Interaction of a social robot with visitors inside a museum through rfid technology,” in 2021 6th International Conference on Smart and Sustainable Technologies (SpliTech). IEEE, 2021, p. 01–06. [16] B.-O. Han, Y.-H. Kim, K. Cho, and H. S. Yang, “Museum tour guide robot with augmented reality,” in 2010 16th International Conference on Virtual Systems and Multimedia. IEEE, 2010, p. 223–229. [17] Y.-L. Chen, C.-C. Hsu, C.-Y. Lin, and H.-H. Hsu, “Robot-assisted language learning: Integrating artificial intelligence and virtual reality into english tour guide practice,” Education Sciences, vol. 12, no. 7, p. 437, 2022. [18] A.-M. Velentza, D. Heinke, and J. Wyatt, “Human interaction and improving knowledge through collaborative tour guide robots,” in 2019 28th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, 2019, p. 1–7. [19] A.-M. Velentza, N. Fachantidis, and I. Lefkos, “Human-robot inter- action methodology: Robot teaching activity,” MethodsX, vol. 9, p. 101866, 2022. [20] M. Quigley, B. Gerkey, K. Conley, J. Faust, T. Foote, J. Leibs, E. Berger, R. Wheeler, and A. Ng, “Ros: an open-source robot operating system,” in Proc. of the IEEE Intl. Conf. on Robotics and Automation (ICRA) Workshop on Open Source Robotics, Kobe, Japan, May 2009. [21] OpenAI, “Openai text-to-speech api,” https://platform.openai.com/ docs/guides/text-to-speech, 2024, accessed: 2026-02-20. [22] K. Kompatsiari, F. Ciardo, D. De Tommaso, and A. Wykowska, “Mea- suring engagement elicited by eye contact in human-robot interaction,” in 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2019, p. 6979–6985. [23] P. Dautzenberg, S. Ladwig, and A. M. Rosenthal-von der Pütten, “Follow me: Anthropomorphic appearance and communication im- pact social perception and joint navigation behavior,” in Proceedings of the 2024 ACM/IEEE International Conference on Human-Robot Interaction, 2024, p. 175–183. [24] F. Faul, E. Erdfelder, A.-G. Lang, and A. Buchner, “G* power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences,” Behavior research methods, vol. 39, no. 2, p. 175–191, 2007. [25] W. Jo, R. Wang, B. Yang, D. Foti, M. Rastgaar, and B.-C. Min, “Cognitive load-based affective workload allocation for multihuman multirobot teams,” IEEE Transactions on Human-Machine Systems, 2024. [26] H. L. O’Brien, P. Cairns, and M. Hall, “A practical approach to measuring user engagement with the refined user engagement scale (ues) and new ues short form,” International Journal of Human- Computer Studies, vol. 112, p. 28–39, 2018. [27] M. Koller, A. Weiss, M. Hirschmanner, and M. Vincze, “Robotic gaze and human views: A systematic exploration of robotic gaze aversion and its effects on human behaviors and attitudes,” Frontiers in Robotics and AI, vol. 10, p. 1062714, 2023. [28] J. Corbin and A. Strauss, “Strategies for qualitative data analysis,” in Basics of Qualitative Research: Techniques and Procedures for Developing Grounded Theory, 3rd ed.Thousand Oaks, CA: SAGE Publications, 2008, p. 65–86. [29] M. Q. Patton, Qualitative Research & Evaluation Methods, 3rd ed. Thousand Oaks, CA: SAGE Publications, Inc, 2001. [30] V. Braun and V. Clarke, “Reflecting on reflexive thematic analysis,” Qualitative research in sport, exercise and health, vol. 11, no. 4, p. 589–597, 2019. [31] —, “Using thematic analysis in psychology,” Qualitative research in psychology, vol. 3, no. 2, p. 77–101, 2006. [32] N. K. Denzin and Y. S. Lincoln, The SAGE Handbook of Qualitative Research, 5th ed.Thousand Oaks, CA: SAGE Publications, Inc, 2001. [33] R. Vallat, “Pingouin: statistics in python,” Journal of Open Source Software, vol. 3, no. 31, p. 1026, 2018. [34] L. Takayama and C. Pantofaru, “Influences on proxemic behaviors in human-robot interaction,” in 2009 IEEE/RSJ international conference on intelligent robots and systems. IEEE, 2009, p. 5495–5502. [35] W. Saeki and Y. Ueda, “Sequential model based on human cognitive processing to robot acceptance,” Frontiers in Robotics and AI, vol. 11, p. 1362044, 2024. [36] S. Zojaji, Y. I. Nakano, and C. Peters, “Impact of cultural differences and politeness on joining small groups of humans, robots, and virtual characters,” in 2025 20th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 2025, p. 479–488. [37] C. M. Van Winkle, “The effect of tour type on visitors’ perceived cognitive load and learning,” Journal of Interpretation Research, vol. 17, no. 1, p. 45–57, 2012. [38] M. Mori, K. F. MacDorman, and N. Kageki, “The uncanny valley [from the field],” IEEE Robotics & automation magazine, vol. 19, no. 2, p. 98–100, 2012. [39] M. Kwon, M. F. Jung, and R. A. Knepper, “Human expectations of social robots,” in 2016 11th ACM/IEEE international conference on human-robot interaction (HRI). IEEE, 2016, p. 463–464. [40] A. Tamai, S. Ono, T. Yoshida, T. Ikeda, and S. Iwaki, “Guiding a person through combined robotic and projection movements,” Inter- national Journal of Social Robotics, vol. 14, no. 2, p. 515–528, 2022. [41] R. A. Martin and T. Ford, The psychology of humor: An integrative approach. Academic press, 2018.