Paper deep dive
A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces
Marcel Heisler, Luca Randecker, Christian Becker-Asano
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Previous research has shown that a human-like robot's acceptance heavily depends on the setting in which it operates and its ability to perform relevant tasks. This paper, first, reports on how our robot processes natural language to generate a multimodal, verbal response integrating emotional expressions based on an emotion simulation backend. Then, it describes how visitors were invited to speak with our robot in their own language at three different, public locations, where the robot was running continuously for several days. The TAM2 questionnaire results reveal that on average users were motivated to use the robot and found it rather useful and easy to use regardless of the specific location. However, public spaces like the tourist information and the city library seem to be a better fit for our interactive, robotic head than an office environment such as the building authority, where the willingness to interact was lower. Overall, the robot's multi-lingual responses were very much appreciated, but every fifth user found the response time too slow impeding the dialog flow, which remains to be improved in future work.
Tags
Links
- Source: https://arxiv.org/abs/2607.24113v1
- Canonical: https://arxiv.org/abs/2607.24113v1
Trouble viewing inline? Open PDF directly →
Full Text
36,319 characters extracted from source content.
Expand or collapse full text
A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces Marcel Heisler 1 , Luca Randecker 1 , and Christian Becker-Asano 1 Abstract— Previous research has shown that a human-like robot’s acceptance heavily depends on the setting in which it operates and its ability to perform relevant tasks. This paper, first, reports on how our robot processes natural language to generate a multimodal, verbal response integrating emotional expressions based on an emotion simulation backend. Then, it describes how visitors were invited to speak with our robot in their own language at three different, public locations, where the robot was running continuously for several days. The Technology Acceptance Model Version 2 (TAM2) questionnaire results reveal that on average users were motivated to use the robot and found it rather useful and easy to use regardless of the specific location. However, public spaces like the tourist information and the city library seem to be a better fit for our interactive, robotic head than an office environment such as the building authority, where the willingness to interact was lower. Overall, the robot’s multi-lingual responses were very much appreciated, but every fifth user found the response time too slow impeding the dialog flow, which remains to be improved in future work. I. INTRODUCTION Whether human-like robots will become a part of everyday life in the future depends on several factors, especially their acceptance by humans [1], [2] in different roles [3], [4]. This, in turn, can be influenced by their appearance, their usefulness in the respective setting, their ability to perform the required tasks, as well as the expectations and attitudes of the people who interact with these robots [1], [2], [5]–[12]. Research also highlights that cultural norms significantly shape their acceptance [1], [13]. Among other factors, human-like robots are accepted by society if they are considered useful in the respective setting [1], [10]–[12]. Several use cases with human-like robots have already been investigated, especially with people in need of a specific service [14]. To further investigate how they are perceived by people in different public spaces, our android robot head, named Kim, spent three weeks interacting with visitors at the tourist information (five days), the building authority (four days), and the city library (five days) in Stuttgart, Germany. Some visitors provided data to evaluate the acceptance, in particular the usefulness of the robot at each location based on a German translation of the Technology Acceptance Model Version 2 (TAM2) [15], [16]. Therefore, the remainder of this paper is structured as follows: Section I outlines previous work related to human- like robots employed in public spaces, Section I describes 1 All authors are with Stuttgart Media University, Germany heisler|randecker|becker-asano @hdm-stuttgart.de the methods used to gather the data, Section IV presents the results followed by Section V discussing potential lim- itations. Section VI summarizes the findings and highlights opportunities for future research. I. RELATED WORK Humanoid robots are increasingly being deployed in pub- lic spaces such as airports, hotels, and touristic places to enhance customer service and explore new forms of Human- Robot Interaction (HRI) [17]. While humanoid robots rep- resent a broader category with a wide range of applica- tions, this work focuses specifically on android robots, a subclass of humanoid robots designed to closely resemble humans [18]. Hence, there are fewer studies with android robots than studies involving humanoid robots. However, in public spaces, android robots have been employed in a variety of roles, such as salespeople in clothing stores [19], receptionists in public and private institutions [20]–[27], with [25] specifically examining a library setting, HR experts [28], [29], interviewers in recruitment contexts [30]–[32], restaurant servants [33]; customer service agents in insurance companies [34] information kiosks in museums [35], and in general service roles [36]. Nevertheless, none of these studies have conducted a comparative analysis with an android robot across different public places. Furthermore, no previous study has used a validated questionnaire based on the Technology Acceptance Model Version 2 (TAM2) [16], [37] to measure and compare the acceptance of an android robot in several real-world settings. I. METHODS This section first describes the hardware and software components of the robot head, followed by the experimental setup at each location, the study procedure, and the measures that were applied. A. Hardware and Software As appearance and functionality are important factors for a human-like robot [1], [38], we will provide information about Kim’s hardware and software in detail below. 1) Hardware: The robot was produced by the same Japanese company that also produced the android robots Andrea [35], Nikola [39], Erica [40] and Geminoid HI-1 [41]. It features 14 pneumatic actuators that control the movements of its face, enabling expressive facial animations. To control its movements and expressions, 14 integer values between 0 arXiv:2607.24113v1 [cs.RO] 27 Jul 2026 Fig. 1. While the microphone in front of the robot is turned on, the robot head Kim performs a “listening” animation. Then, OpenAI’s Whisper speech-to- text model is used to transcribe “request audio” into text. The resulting text is sent to an OpenAI ChatGPT 4.1 assistant, which uses Retrieval Augmented Generation (RAG) together with a location-specific system prompt to generate a textual answer. The text-to-speech software XTTS transforms the text into a “response audio”, which is used by FaceXHubert to encode a lip-sync animation of the robot’s face. Finally, the animation is played back together with the audio. The video stream provided by the camera behind the robot’s head is analyzed for the presence of human faces, of which the closest is selected as an eye gaze target, whenever Kim is answering or waiting for user input. and 255 are being sent via RS-485 with 25 hertz frequency by our custom developed software framework, see also [42]. Unlike some of the aforementioned robots, Kim does not have cameras integrated into its eyes. Instead, a webcam is installed behind it to detect and look at the closest person. For audio interaction, an external microphone records sound input, while a loudspeaker provides output, ensuring basic conversational capabilities. Our software framework runs on an Nvidia Jetson Orin, while compressed air and electric power are used to drive the robot’s actuators. The robot head is usually dressed in a black T-shirt and sometimes wears a baseball cap to hide the gray plate at the back of its head. 2) Software components: The software components are an extended version of [35], [42], which combines four separate, open-source Machine Learning (ML) models, see also Fig. 1: 1) ForSpeech-to-Text(STT)(stepone) whisper-large-v3-turbo is used supporting multiple languages as input. 2) Multilingual output (step 3) is generated using XTTS [43] for speech synthesis. This ML model supports 17 languages. A specific design-congruent voice [44] was cloned and used consistently in all languages. 3) To speed up the generation of lip-sync animations (step 4) the former ML model used as described in [45] was replaced by FaceXHubert [46]. 4) The people tracking is based on a realtime analysis of the video stream provided by the webcam using posenet [47] (not shown in Fig. 1). To start a conversation users have to turn on the mi- crophone, which also activates a listening animation. We opted against an active listening solution, because of a high probability of false activations due to the noisy environment in the public spaces. After the user turns off the microphone, the robot switches to a thinking animation, signaling that it is processing the input. As investigated in [48], thinking faces have a positive impact on a more natural HRI. To further enhance its human-likeness the robot blinks at random and yawns if no interaction takes place for some time. To generate context-sensitive textual responses, three location-specific OpenAI assistants are used with their gpt-4.1 model as the basis. The assistants’ prompts contain information about the robot itself and about its current location. In addition, Retrieval Augmented Gener- ation (RAG) is used to provide additional information (see Section I-B). The robot’s emotional states are simulated using WASABI [49]. To trigger the emotion dynamics of WASABI, the OpenAI assistant is prompted to calculate the valence of the last input sentence spoken by the user, which is then sent to WASABI as a valenced impulse ranging between -100 and +100. In effect, WASABI, which is running as a concurrent process in the same computer, returns the emotion likelihood of the emotions “happy”, “sad”, “angry”, “fearful”, “disgusted”, “surprised”, or “neutral”. The internal emotion dynamic of WASABI lets the robot’s emotional state automatically return to “neutral” after some time without any inputs. These emotions are expressed using validated static facial expressions [50]. The animations for thinking and speech have a higher priority than the emotional expressions for the relevant actuators. Thus, while the robot is speaking its mouth area is animated according to the speech signal, but before or after speech it is animated according to the emotion, e.g. smiling. B. Experimental Setup Overall, the responses of the robot were shaped by the location-specific, publicly available information that was provided to the ChatGPT-based backend in advance. For the tourist information, the database contained information about the opening hours, tours, bus and train plans, tickets, services, and tips for visiting the city. At the building authority, the robot was provided with respective laws. The setting at the city library contained information about the opening hours, the building itself, and the locations of different book categories. 1) Tourist Information: The robot was present at the Stuttgart tourist information from May 26 to 30, 2025, between 10 AM and 6 PM each day. The tourist information is at a location with a lot of walk-in customers. A visible roll- up banner with information about Kim was placed directly to the left of the entrance. The robot head was placed on the front desk, as shown in Fig. 2, so that it was visible to visitors when they usually reached human staff to ask questions or make payments. Next to the head was a table display and underneath was a poster with information and instructions on how to interact with Kim. During its deployment at this location, Kim wore a Stuttgart-branded cap and an employee t-shirt in black. 2) Building Authority: From June 3 to 6, 2025, Kim was at the building authority with the following schedule: Tuesday and Wednesday from 2 PM to 4 PM, Thursday from 9 AM to 6 PM, and Friday from 9 AM to 12 PM. Unlike the other two locations, the building authority is not located in a busy area. In addition, the opening hours for visitors were heavily restricted, so that people that interacted with Kim were more likely to be employees in the same building than visitors with building concerns, which is also reflected in the number of completed questionnaires. However, there was also a roll-up, a table display and a poster inviting visitors to take notice of the robot. Kim itself was placed on a table approximately at the same height as the users, as shown in Fig. 3. 3) City Library: From June 10 to 14, 2025, Kim was at the city library available to the public each day from 9 AM to 6 PM, located on the ground floor of the building, in an almost empty huge square space. Here, too, we made sure that the robot head was clearly visible and provided the relevant information, see Fig. 4. A lot of people have interacted with Kim, which is also reflected in the number of completed surveys. C. Study Procedure and Measurement 1) Study Procedure: After the interactions users were in- vited to complete a questionnaire that contained an informed consent and items in both English and German. Additionally, participants were informed that they could cancel the survey at any time. The number of participants at each location varies with 19 (N 1 = 19) completed questionnaires at the tourist information, seven (N 2 = 7) at the building authority, and 23 (N 3 = 23) at the city library. The TAM2 questionnaire from [16] was used as a basis, as it provided the exact German translations of the original items [15]. Our questionnaire can be divided into the fol- lowing categories: general information such as location, date, time, and age; questions evaluating participants’ prior knowl- edge of the robot’s availability at this place and whether it influenced their decision to visit the location; TAM2 items measuring intention to use (ITU), perceived usefulness (PU), and perceived ease of use (PEOU); questions regarding the emotions; as well as additional questions, regarding for how long participants had interacted with the robot, which languages they used, and any suggestions they had. Finally, participants were asked how useful they considered the robot in this specific location. 2) Measurement: Whether participants knew about the robot’s availability prior to their visit, and whether they had specifically come to see the robot was evaluated with yes/no questions. The TAM2 items and emotion-related questions were evaluated using a 7-point Likert scale, where 1 indicated ‘strongly disagree’ and 7 indicated ‘strongly agree’. ITU was measured with the items “Assuming I have access to the robot, I intend to use it.”, and “Given that I have access to the robot, I predict that I would use it.” PU was measured with four items “Using the robot improves my performance.”, “Using the robot increases my productivity.”, “Using the robot enhances my effectiveness.”, and “I find the robot to be useful.” PEOU was measured with four items “My interaction with the robot is clear and understandable.”, “My interaction with the robot does not require a lot of my mental effort.”, “I find the robot to be easy to use.”, and “I find it easy to get the robot to do what I want it to do.” The mean value for each category per participant was calculated. Additionally, the Shapiro-Wilk test was used to determine the W statistic and p-value, in order to assess whether the data is normally distributed. The perceived emotionality of the robot head was mea- sured with the items “I have the impression that the robot displays emotional reactions.”, “I find the robot’s emotional responses appropriate to the situation.” and “The robot ap- pears moody or shows emotional mood swings.” on the same 7-point Likert scale. The final question “How useful do you think would it be to use this robot here today?” was measured by a distinct Likert scale ranging from 0 ‘not at all’ to 10 ‘very much’ to compare our setup with a previous result of a study with the robot Andrea in a public museum [35]. IV. RESULTS All statistics were calculated using Jamovi [51]. A. Previous knowledge of the robot Eighteen of the total of 49 interviewees knew about the availability of the robot prior to their visit (tourist informa- tion: five, building authority: six, city library: seven). Nine Fig. 2. Robot head Kim at the tourist information in Stuttgart, Germany, available from May 26 to 30, 2025. Fig. 3. Robot head Kim at the building authority in Stuttgart, Germany, available from June 3 to 6, 2025. Fig. 4.Robot head Kim at the city library in Stuttgart, Germany, available from June 10 to 14, 2025. 1 2 3 4 5 6 7 Tourist information Building authority City library Location Intention to use Fig. 5.Box plot showing the distribution of intention to use (ITU) by location with indicated median values. 1 2 3 4 5 6 7 Tourist information Building authority City library Location Perceived usefulness Fig. 6.Box plot showing the distribution of perceived usefulness (PU) by location with indicated median values. 1 2 3 4 5 6 7 Tourist information Building authority City library Location Perceived ease of use Fig. 7. Box plot showing the distribution of perceived ease of use (PEOU) by location with indicated median values. 0 2 4 6 8 10 Tourist information Building authority City library Location Usefulness today Fig. 8.Box plot showing the distribution of the usefulness today by location with indicated median values on a scale from zero to ten. TABLE I DESCRIPTIVE STATISTICS OF intention to use (ITU), perceived usefulness (PU), perceived ease of use (PEOU), AND usefulness today BY LOCATION. TABLE I RESULTS OF THE KRUSKAL–WALLIS TESTS FOR intention to use (ITU), perceived usefulness (PU), perceived ease of use (PEOU), AND usefulness today: FOR NON OF THEM p < 0.05 HOLDS TRUE Kruskal-Wallis χ²dfpε² Intention to use0.21220.9000.00441 Perceived usefulness 1.38120.5010.02877 Perceived ease of use 3.32620.1900.06929 Usefulness today 2.02020.3640.04297 out of these 18 specifically came to the location to experience the robot live (tourist information: three, building authority: four, city library: two). B. Differences by location An analysis of the data categorized by location is presented in Figures 5, 6, 7, and 8. Table I presents the descriptive statistics for each variable at each location, including the number of participants (N), the mean, the median, and the standard deviation (SD). Normalityofthedatawasevaluatedusingthe Shapiro–Wilk test for each variable and location. For several groups, the test indicated significant deviations from normal distribution (p < 0.05), including tourist information for “ITU”, “PEOU”, and “How useful do you think would it be to use this robot here today?” as well as city library for “ITU”, “PEOU”, and “How useful do you think would it be to use this robot here today?”. As the assumption of normality was not fully met in all dimensions, a non- parametric Kruskal–Wallis test was applied to compare the locations. Table I shows the obtained p-values. All of them exceeded the significance threshold (α = 0.05), indicating that there are no statistically significant differences between the locations. TABLE I RESULTS OF THE THREE EMOTION-RELATED QUESTIONS. Descriptives Experimental conditionNMeanMedianSD Displaying emotional reactionsTourist information184.174.001.72 Building authority 75.5761.72 City library 234.6552.04 Approriate emotional responsesTourist information185.005.001.64 Building authority 74.8662.27 City library 235.1351.66 Showing mood swingsTourist information182.672.501.71 Building authority 72.8622.12 City library 233.1321.94 C. Perceived emotionality, interaction time, languages used, and suggestions The results for all three emotion-related questions do not differ significantly for the three locations, as shown in Table I with the mean, the median and the standard deviation for all items in all locations. The robot head was perceived as rather emotional (median value greater than four for the first two questions) and as not suffering from mood swings (median values less than four for the last question). The self-reported interaction time varied from “not at all” to thirty minutes with an average interaction time of seven minutes. Most participants reported five minutes of previous interaction with the robot head, which is also the median of the distribution. The multilingualism of the robot was used by many people during their interactions with the robot. The participants interacted with the robot in the following languages, with the frequency of use indicated in parentheses (some visitors tried out more than one language): German (40), English (10), French (8), Spanish (3), Italian (3), Turkish (2), Arabic (1), Chinese (1), Croatian (1), Korean (1), Latvian (1), Persian (1), Polish (1), Portuguese (1), Russian (1), Swabian (a local German dialect) (1), and Vietnamese (1). Finally, the robot’s long response time was frequently noted in the survey, with 10 out of 49 mentioning it. Others commented that the robot head appeared to be uncanny (five times) or funny/nice/innovative (seven times). V. DISCUSSION AND LIMITATIONS Our study has a couple of limitations that might impact its results. One possible explanation for the absence of sta- tistically significant differences between the three locations is the relatively low number of participants, particularly at the building authority. This may have led to an insufficient, statistical power to detect differences across locations. Still, the descriptively lower Intention to use and Usefulness today at the building authority compared to both other locations are in line with a recent survey [52] that found robot deploy- ments to be more acceptable in cultural and public leisure environments like libraries and museums than in government and administrative services. Compared to a previous study with the full-body android robot Andrea employed in a public museum [35], the robot head was perceived as similarly useful in the tourist information and the city library, but less useful in the building authority location. Together, this suggests that those findings from the above mentioned survey might generalize to real-world deployments of android robots although further investigation is needed for verification. The fact that Andrea in the museum and Kim at the library and tourist information center are rated quite similarly raises questions about the effectiveness of the extensions made to the software. While the multilingualism was verbally appreciated by many users, it seems not to be reflected in the perceived usefulness of the robot. However, to draw clear conclusions in this regard other factors like appearance, location, and further extensions that had been implemented in the robot’s software would need to be controlled for. The uncontrolled interaction dynamics in such public environments with multiple people watching and interacting with the robot at the same time can be seen as another limitation. Sometimes visitors started to discuss their answers to the questionnaire items with other people around them. Therefore, the results of the present analysis should be complemented by similar studies in more controlled envi- ronments. Also, a long-term deployment would help assess sustained user acceptance over time. The choice of locations was influenced by practical con- siderations, e.g. their availability and willingness to accom- modate an android robot. Similarly, the robot’s clothing was influenced by suggestions of the hosting locations. In future studies different locations could be compared based on more theoretically grounded assumptions and stricter control of the robot head’s appearance would be necessary, to eliminate possible uncontrolled influences of user perceptions through differing appearances of the robot head. VI. CONCLUSIONS This exploratory study gathered the general public’s opin- ion about a fully autonomous, very humanlike robot head that was set up for verbal interaction for several days each in three public spaces in Germany. Approximately one-third of the visitors (18 out of 49) knew about the availability of the robot in advance and half of them (9 out of 18) specifically came to experience the robot. This indicates a general interest in interacting with very human-like robots in public spaces. Overall, the robot head was evaluated at each location similarly positively on all three TAM2 subscales with median values greater than four on the standard seven point scale. It also became apparent that the system would benefit from a smoother interaction as was mentioned by every fifth visitor. The multilingualism of the robot was welcomed by everyone, who interacted with it, as is reflected in the wide variety of languages used. Previous research suggests that in context-free interaction a more anthropomorphic design of a robot negatively impacts its likeability and perceived safety, but increases its attributed intelligence [8]. Therefore, investigating the technological acceptance of a less anthropomorphic, social robot in the same, public environments seems to be a logical next step. In summary, the study underlines the potential of very human-like, robotic heads in service-oriented settings and contributes to the growing body of research on real-world HRI in public places. REFERENCES [1] M. de Graaf and S. Ben Allouch, “Exploring influencing variables for the acceptance of social robots,” Robotics and autonomous systems, vol. 61, no. 12, p. 1476–1486, 2013. [2] E. Broadbent, R. Stafford, and B. Macdonald, “Acceptance of health- care robots for the older population: Review and future directions,” I. J. Social Robotics, vol. 1, p. 319–330, 11 2009. [3] N. Mavridis et al., “Opinions and attitudes toward humanoid robots in the middle east,” AI and Society, vol. 27, no. 4, p. 517–534, Nov. 2012. [4] T. Hoshikawa, K. Ogawa, and H. Ishiguro, “Future roles for android robots: Survey and trial,” in Procs of the 3rd Intl. Conf. on Human-Agent Interaction, ser. HAI ’15.New York, NY, USA: Association for Computing Machinery, 2015, p. 41–48. [Online]. Available: https://doi.org/10.1145/2814940.2814960 [5] J. Goetz, S. Kiesler, and A. Powers, “Matching robot appearance and behavior to tasks to improve human-robot cooperation,” in The 12th IEEE Intl. Workshop on Robot and Human Interactive Communication, 2003, p. 55–60. [6] E. Broadbent et al., “Robots with display screens: A robot with a more humanlike face display is perceived to have more mind and a better personality,” PLOS ONE, vol. 8, no. 8, p. 1–9, 08 2013. [Online]. Available: https://doi.org/10.1371/journal.pone.0072589 [7] M. Mara and M. Appel, “Science fiction reduces the eeriness of android robots: A field experiment,” Computers in Human Behavior, vol. 48, p. 156–162, 2015. [Online]. Available: https: //w.sciencedirect.com/science/article/pii/S0747563215000199 [8] K. S. Haring et al., “How people perceive different robot types: A direct comparison of an android, humanoid, and non-biomimetic robot,” in 2016 8th Intl. Conf. on Knowledge and Smart Technology (KST), 2016, p. 265–270. [9] J. Wirtz et al., “Brave new world: service robots in the frontline,” Journal of Service Management, vol. 29, no. 5, p. 907–931, 2018. [Online]. Available: https://w.sciencedirect.com/science/article/pii/ S1757581818000063 [10] A. K. Hall et al., “Acceptance and perceived usefulness of robots to assist with activities of daily living and healthcare tasks,” Assistive Technology, vol. 31, no. 3, p. 133–140, 2019. [Online]. Available: https://doi.org/10.1080/10400435.2017.1396565 [11] S. Naneva et al., “A systematic review of attitudes, anxiety, acceptance, and trust towards social robots,” Intl. Journal of Social Robotics, vol. 12, p. 1179–1201, December 2020. [12] C. S. Song and Y.-K. Kim, “The role of the human-robot interaction in consumers’ acceptance of humanoid retail service robots,” Journal of Business Research, vol. 146, p. 489–503, 2022. [Online]. Available: https://w.sciencedirect.com/science/article/pii/ S014829632200323X [13] D. Li, P. L. P. Rau, and Y. Li, “A cross-cultural study: Effect of robot appearance and task,” Intl. Journal of Social Robotics, vol. 2, no. 2, p. 175–186, 2010. [14] M. Mende et al., “Service robots rising: How humanoid robots influence service experiences and elicit compensatory consumer responses,” Journal of Marketing Research, vol. 56, no. 4, p. 535–556, 2019. [Online]. Available: https://doi.org/10.1177/ 0022243718822827 [15] V. Venkatesh and F. D. Davis, “A theoretical extension of the technol- ogy acceptance model: Four longitudinal field studies,” Management Science, vol. 46, no. 2, p. 186–204, 2000. [16] T. Olbrecht, Akzeptanz von E-Learning. Eine Auseinandersetzung mit dem Technologieakzeptanzmodell zur Analyse individueller undsozialerEinflussfaktoren.Jena:Univ.,2010.[On- line]. Available: http://w.db-thueringen.de/servlets/DerivateServlet/ Derivate-21996/Olbrecht/Dissertation.pdf [17] Y. Tong, H. Liu, and Z. Zhang, “Advancements in humanoid robots: A comprehensive review and future prospects,” IEEE/CAA Journal of Automatica Sinica, vol. 11, no. 2, p. 301–328, 2024. [18] H.Ishiguro,“Androidscience:consciousandsubconscious recognition,” Connection Science, vol. 18, no. 4, p. 319–332, 2006. [Online]. Available: https://doi.org/10.1080/09540090600873953 [19] M. Watanabe, K. Ogawa, and H. Ishiguro, “Can androids be salespeople in the real world?” in Procs of the 33rd Annual ACM Conf. Extended Abstracts on Human Factors in Computing Systems, ser. CHI EA ’15.New York, NY, USA: Association for Computing Machinery, 2015, p. 781–788. [Online]. Available: https://doi.org/10.1145/2702613.2702967 [20] —, “Field study: Can androids be a social entity in the real world ?” in ACM/IEEE Intl. Conf. on Human-Robot Interaction, 03 2014, p. 316–317. [21] Y. Kondo et al., “Multi-person human-robot interaction system for android robot,” in 2010 IEEE/SICE Intl. Symposium on System Inte- gration, 2010, p. 176–181. [22] C. Liu et al., “Generation of nodding, head tilting and eye gazing for human-robot dialogue interaction,” in 2012 7th ACM/IEEE Intl. Conf. on Human-Robot Interaction (HRI), 2012, p. 285–292. [23] A. Vishwanath et al., “Humanoid co-workers: How is it like to work with a robot?” in 2019 28th IEEE Intl. Conf. on Robot and Human Interactive Communication (RO-MAN), 2019, p. 1–6. [24] T. Hashimoto et al., “Realization and evaluation of realistic nod with receptionist robot saya,” in 16th IEEE Intl. Conf. on Robot and Human Interactive Communication, ser. IEEE Intl. Workshop on Robot and Human Interactive Communication, 2007, p. 326–331. [25] T. Umetani, T. Kikuchi, and N. Saiwaki, “Remote reference-desk service system using android robot for university librarian,” in 2019 IEEE Intl. Conf. on Advanced Robotics and its Social Impacts (ARSO), 2019, p. 25–27. [26] T. Hashimoto, M. Senda, and H. Kobayashi, “Realization of realistic and rich facial expressions by face robot,” in IEEE Conf. on Robotics and Automation, 2004. TExCRA Technical Exhibition Based., 2004, p. 37–38. [27] J. Reis et al., “Service robots in the hospitality industry: The case of henn-na hotel, japan,” Technology in Society, vol. 63, p. 101423, 11 2020. [28] R. Stock et al., “When robots enter our workplace: Understanding employee trust in assistive robots,” Publications of Darmstadt Technical University, Institute for Business Studies (BWL), 2019. [Online]. Available: https://EconPapers.repec.org/RePEc:dar:wpaper: 118841 [29] R. Stock-Homburg, M. Hannig, and L. Lilienthal, “Conversational flow in human-robot interactions at the workplace: Comparing humanoid and android robots,” in Social Robotics: 12th Intl. Conf.Berlin, Heidelberg: Springer-Verlag, 2020, p. 578–589. [30] K. Inoue et al., “Job interviewer android with elaborate follow-up question generation,” in Procs Intl. Conf. on Multimodal Interaction, ser. ICMI ’20.New York, NY, USA: ACM, 2020, p. 324–332. [Online]. Available: https://doi.org/10.1145/3382507.3418839 [31] E. Baka et al., “Social robots and digital humans as job interviewers: A study of human reactions towards a more naturalistic interaction,” in Human-Computer Interaction. Technological Innovation, M. Kurosu, Ed. Cham: Springer Intl. Publishing, 2022, p. 455–474. [32] K. Inoue et al., A Job Interview Dialogue System with Autonomous Android ERICA. Singapore: Springer Singapore, 2021, p. 291–297. [Online]. Available: https://doi.org/10.1007/978-981-15-9323-925 [33] L. Lu, P. Zhang, and T. C. Zhang, “Leveraging “human-likeness” of robotic service at restaurants,” Intl. Journal of Hospitality Manage- ment, vol. 94, p. 102823, 2021. [34] N. Mishra et al., “Can a humanoid robot be part of the organizational workforce? a user study leveraging sentiment analysis,” in 28th IEEE Intl. Conf. on Robot and Human Interactive Communication, 10 2019, p. 1–7. [35] M. Heisler and C. Becker-Asano, “Conversations with andrea: Visi- tors’ opinions on android robots in a museum,” in 34th IEEE Intl. Conf. on Robot and Human Interactive Communication, 2025, p. 112–119. [36] S. Chuah and J. Yu, “The future of service: The power of emotion in human-robot interaction,” Journal of Retailing and Consumer Services, vol. 61, p. 102551, 04 2021. [37] V. Venkatesh and F. D. Davis, “A Theoretical Extension of the Technology Acceptance Model: Four Longitudinal Field Studies,” Management Science, vol. 46, no. 2, p. 186–204, Feb. 2000. [38] N. Minh Trieu and N. Truong Thinh, “A comprehensive review: Interaction of appearance and behavior, artificial skin, and humanoid robot,” Journal of Robotics, vol. 2023, no. 1, p. 5589845, 2023. [39] W. Sato et al., “An Android for Emotional Interaction: Spatiotemporal Validation of Its Facial Expressions,” Frontiers in Psychology, vol. 12, Feb. 2022, publisher: Frontiers. [40] D. F. Glas et al., “ERICA: The ERATO Intelligent Conversational Android,” in 25th IEEE Intl. Symposium on Robot and Human Interactive Communication. New York, NY, USA: IEEE, Aug. 2016, p. 22–29. [41] A. M. von der P ̈ utten et al., “An android in the field,” in Procs of the 6th Intl. Conf. on Human-robot interaction, ser. HRI ’11. New York, NY, USA: Association for Computing Machinery, Mar. 2011, p. 283– 284. [Online]. Available: https://doi.org/10.1145/1957656.1957772 [42] M. Heisler and C. Becker-Asano, “An Android Robot Head as Embodied Conversational Agent,” in ISR Europe 2023; 56th Intl. Symposium on Robotics, Sep. 2023, p. 93–99. [Online]. Available: https://ieeexplore.ieee.org/document/10363058 [43] E. Casanova et al., “XTTS: a Massively Multilingual Zero-Shot Text- to-Speech Model,” in Interspeech 2024, 2024, p. 4978–4982. [44] J. M. Kuch et al., “Your robot, my voice: Enhancing android robot likability through personalization by cloning the user’s voice,” in 2025 34nd IEEE Intl. Conf. on Robot and Human Interactive Communica- tion (RO-MAN). IEEE, Aug. 2025, p. to appear. [45] M. Heisler, S. Kopp, and C. Becker-Asano, “Making an Android Robot Head Talk,” in 2023 32nd IEEE Intl. Conf. on Robot and Human Interactive Communication (RO-MAN).Busan, Korea, Republic of: IEEE, Aug. 2023, p. 1837–1842. [Online]. Available: https://ieeexplore.ieee.org/document/10309532/ [46] K. I. Haque and Z. Yumak, “FaceXHuBERT: Text-less Speech-driven E(X)pressive 3D Facial Animation Synthesis Using Self-Supervised Speech Representation Learning,” in Procs of the 25th Intl. Conf. on Multimodal Interaction, ser. ICMI ’23.New York, NY, USA: Association for Computing Machinery, Oct. 2023, p. 282–291. [Online]. Available: https://doi.org/10.1145/3577190.3614157 [47] A. Kendall, M. Grimes, and R. Cipolla, “PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization,” in 2015 IEEE Intl. Conf. on Computer Vision (ICCV), Dec. 2015, p. 2938–2946, iSSN: 2380-7504. [Online]. Available: https://ieeexplore. ieee.org/document/7410693 [48] S. Namba et al., “How an android expresses “now loading. . . ”: Examining the properties of thinking faces,” Intl. Journal of Social Robotics, vol. 16, no. 8, p. 1861–1877, 2024, publisher: Springer Science and Business Media LLC. [Online]. Available: https://link.springer.com/10.1007/s12369-024-01163-9 [49] C. Becker-Asano, “WASABI for affect simulation in human-computer interaction,” in Proc. on Emotion Representations and Modelling for HCI Systems. Sydney, Australia: Springer, 2014. [50] A. Kassner and C. Becker-Asano, “Comparing an android head with its digital twin regarding the dynamic expression of emotions,” in 2023 11th Intl. Conf. on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), Sep. 2023, p. 1–7. [Online]. Available: https://ieeexplore.ieee.org/document/10388207 [51] The jamovi project, “jamovi,” https://w.jamovi.org, 2024, version 2.6 [Computer Software]. [52] L. Aymerich-Franch et al., “Public acceptance of cybernetic avatarsintheservicesector:evidencefromalarge- scalesurvey,”FrontiersinRoboticsandAI,vol.12, Jan. 2026. [Online]. Available: https://w.frontiersin.org/journals/ robotics-and-ai/articles/10.3389/frobt.2025.1719342/full