Paper deep dive
Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech
Edwin C. Montiel-Vazquez, Christian Arzate Cruz, Stefanos Gkikas, Thomas Kassiotis, Giorgos Giannakakis, Randy Gomez
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/14/2026, 2:34:02 AM
Summary
The paper introduces a lightweight, emotion-aware transformer model for predicting iconic gesture placement and intensity in robot co-speech. By utilizing text and emotion as inputs, the model avoids the latency associated with audio-based systems and outperforms GPT-4o on the BEAT2 dataset, demonstrating suitability for real-time deployment on social robots like Haru.
Entities (5)
Relation Signals (3)
Proposed Transformer Model → deployedon → Haru
confidence 100% · We deploy it on the social robot Haru [32]
Proposed Transformer Model → trainedon → BEAT2
confidence 100% · We build on the body-expression-audio-text (BEAT2) dataset [2] to train a model
Proposed Transformer Model → outperforms → GPT-4o
confidence 95% · The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that derives iconic gesture placement and intensity from text and emotion alone, requiring no audio input at inference time. The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression on the BEAT2 dataset, while remaining computationally compact and suitable for real-time deployment on embodied agents.
Tags
Links
- Source: https://arxiv.org/abs/2604.11417v1
- Canonical: https://arxiv.org/abs/2604.11417v1
Trouble viewing inline? Open PDF directly →
Full Text
26,123 characters extracted from source content.
Expand or collapse full text
Efficient Emotion-Aware Iconic Gesture Prediction for Robot Co-Speech Edwin C. Montiel-Vazquez School of Engineering and Sciences Tecnologico de Monterrey State of Mexico, Mexico edwincmv@exatec.tec.mx Christian Arzate Cruz Honda Research Institute Japan Wako City, Japan christian.arzate@jp.honda-ri.com Stefanos Gkikas Honda Research Institute Japan Wako City, Japan stefanos.gkikas@jp.honda-ri.com Thomas Kassiotis Department of Electronic Engineering Hellenic Mediterranean University Chania, Greece ddk305@edu.hmu.gr Giorgos Giannakakis Department of Electronic Engineering Hellenic Mediterranean University Chania, Greece ggian@hmu.gr Randy Gomez Honda Research Institute Japan Wako City, Japan r.gomez@jp.honda-ri.com Abstract—Co-speech gestures increase engagement and improve speech understanding. Most data-driven robot systems generate rhythmic beat-like motion, yet few integrate semantic emphasis. To address this, we propose a lightweight transformer that derives iconic gesture placement and intensity from text and emotion alone, requiring no audio input at inference time. The model outperforms GPT-4o in both semantic gesture placement classification and intensity regression on the BEAT2 dataset, while remaining computationally compact and suitable for real-time deployment on embodied agents. Index Terms—affective computing, co-speech generation, ges- tures, emotion, transformers, social robots. I. INTRODUCTION Emotional expressiveness is fundamental to natural and engaging communication. Humans convey their internal state through body gestures and facial expressions, ranging from large deliberate movements that depict the meaning of what is said, known as iconic or semantic gestures, to small rhythmic motions that follow speech rhythm, known as beat gestures [1]. Prior work on robot co-speech gesture generation has largely focused on beat gestures, while semantic gestures remain rarely addressed, despite growing interest in the field [2], [3]. A further limitation of existing methods is that they do not explicitly model how emotion shapes movement [4]–[6]. The most related work is by Ishii et al. [7], who condition pose generation on personality traits. However, emotion — not personality — is what most directly drives physical expression. We therefore condition our system on four basic emotions from Plutchik’s wheel [8]: joy, anger, sadness, and fear. The importance of affective modeling in text-driven communication systems extends beyond robotics; even frameworks designed specifically for analyzing human language in social platforms have identified the absence of emotion recognition as a central limitation, highlighting the broader need for emotion-aware approaches to text processing [9]. Given an utterance and a target emotion, our model produces a gesture sequence that integrates both beat and iconic motion, enabling a robot to express not only what it says but also how it feels, in real time. Fig. 1: Task overview. An utterance is separated into words, and the semantic gesture placement and intensity is calculated per word. Most co-speech methods for robots [5], [10]–[12] and artificial agents [4], [6], [7], [13] assume audio is available at inference time to extract prosodic features or synchronize motion. For robots that rely on text-to-speech (TTS), this assumption introduces latency and reduces responsiveness. While LLMs can integrate semantic context effectively, their computational cost makes them impractical for real-time deployment on most robots and embodied agents. We therefore propose a text-only, emotion-aware pipeline that takes two inputs: the utterance the robot will say and the intended emotion. We build on the body-expression-audio-text (BEAT2) dataset [2] to train a model that identifies semantically relevant words in an utterance and quantifies their gesture intensity [3]. Our approach is presented in Figure 1. The contributions of this work are: (i) a text-based model for semantic gesture placement in sentences, (i) an efficient method for iconic gesture intensity regression, and (i) a framework for emotion- aware semantic gestures in social robots. I. RELATED WORK In this section, we review prior work on co-speech gesture generation for artificial agents and robots. arXiv:2604.11417v1 [cs.RO] 13 Apr 2026 A. Artificial Agents Co-speech gesture generation for artificial agents has re- ceived considerable attention, particularly in recent years. Among the most important contributions to the field are the BEAT [14] and BEAT2 [2] datasets. Based on monologues, both combined provide over70hours of motion-capture recordings and mesh representations paired with audio, text, and frame- level semantic labels. Of particular relevance, BEAT2 includes manually annotated iconic gesture intensity values at the word level. We use BEAT2, the extension of BEAT, to train our iconic gesture placement and intensity models. Recent work has shifted toward full-body motion genera- tion [15], moving away from earlier upper-body approaches [4], [6], [7]. Generating accurate, human-like motion requires ac- counting for variability in emotion and semantic emphasis, two aspects that remain underexplored. Computational efficiency is an additional requirement for real-time robotics applications. Model architectures have evolved from recurrent networks such as long short-term memory (LSTM) [16] to attention-based transformers [17]. Adversarial and diffusion-based methods have also been proposed to improve motion realism and diver- sity [2], [14], [18]. We adopt a transformer-based architecture, as the attention mechanism is well-suited for text-driven tasks while remaining computationally efficient. Applications range from realistic gesture synthesis [3] to stylized 3D character animation [13], showing that models trained on human motion data can generalize across different embodiments. This supports the use of a single model trained on human data for social robot applications. B. Robots Embodied co-speech gesture generation has been explored less thoroughly than its counterpart for artificial agents. Ap- proaches for semantic gestures in robots are sparse, ranging from re-targeting human motion [19] to rule-based policies [20]. Most data-driven methods focus on rhythmic motion, using audio and text to generate head and arm movements that follow speech rhythm. The main differences between methods lie in how linguistic and prosodic information is represented and in the learning architectures used. For linguistic features, sentence-level context has been addressed using BERT embeddings [5], [21], while word- level conditioning enables finer lexical control [10]. Gestures are commonly modeled as a regression problem over joint poses using autoregressive architectures [5], [10]–[12], which generate each new joint position from previous poses to maintain motion continuity. GANs have also been used to better match the distribution of human motion while reducing over-smoothing [22], [23]. A key limitation is that the few robotics methods targeting iconic gestures typically trigger pre-defined animations without modeling placement timing or intensity [12]. In contrast, our model explicitly predicts where iconic gestures should occur and how strongly they should be performed. I. PROPOSED APPROACH As presented in Figure 1, our semantic co-speech pipeline encodes the input through an embedding layer and passes it to a transformer model. The embeddings are drawn from established language models [21], [24] and emotion representations [25]. At inference, the system takes a text prompt and a target emotion as input and outputs a time-aligned sequence specifying the placement and intensity of semantic-emphasis gestures. The text is encoded with SBERT [24], while word-level embeddings are obtained using emo2vec [25]. The word embeddings are combined with the utterance’s emotion by averaging the word and emotion-label representations. The model then predicts iconic-gesture placement at the word level and the corresponding intensity, conditioned on the target emotion, thereby reflecting the affective characteristics of the utterance. Training details are described in the following section. IV. ICONIC GESTURE PREDICTION A. Inputs and Supervision The two inputs to our system are the text prompt the robot will say and the emotion of the utterance. Both are obtained from the BEAT2 dataset [2] for training. BEAT2 provides word- level iconic annotations as continuous intensity values, which we use as regression targets for intensity and classification targets for placement. The intensity labels are binarized by setting the activation to1when the value exceeds0.5and to0 otherwise, following the threshold used in previous work [3]. The data is organized into sentence and word-of-interest pairsp n = (s,w n ), wheres = (w 1 ,w 2 ,...,w 40 )represents the full sentence andnidentifies the word of interest within s. We encodesusing SBERT [24] to obtain sentence-level semantic embeddingsh s ∈R 384 , shared across all words w n in the sentence. For word-level representation, we use emo2vec [25] to obtaine w = emo2vec(w n ), wheree w ∈R 100 . To incorporate the overall sentence emotion,e w is augmented withe emo = emo2vec(label), wherelabelis the emotion label of the sentence. The emotion-enhanced word representation is thene n = (e w + e emo )/2. The final input used to predict the iconic label c or intensity i for each word is p n = (h s ,e n ). B. Transformer Architecture The proposed architecture leverages cross- and self-attention to enable efficient global modeling with reduced computational complexity [26], [27]. Rather than applying attention directly to all input embeddings, a compact latent space is introduced as an intermediate representation. The flattened input embeddings are represented asX∈R M×D , while a learnable latent matrix Z 0 ∈R N×d withN ≪ Maggregates information from the input, forming an efficient bottleneck. Positional information is incorporated using Fourier feature encoding: γ(p) = [sin(πω k p), cos(πω k p),p] K k=1 ,(1) wherepdenotes a normalized coordinate andω k the frequency bands. The encoded features are concatenated with the input Fig. 2: High-level overview of the proposed model. embeddings before attention. Cross-attention maps the input into the latent space: Attn cross (Z, X) = softmax Q Z K ⊤ X √ d h V X ,(2) withQ Z ,K X , andV X denoting query, key, and value projections. The latent representations are then processed by self-attention within the latent space, enabling global interactions among latent tokens. Each attention block is followed by a feedforward transformation: FFN(z) = W 2 GELU(W 1 z),(3) whereW 1 andW 2 are learned parameters. Finally, the latent embeddings are mean-pooled and projected through a fully connected layer for prediction. The implemented configuration uses128latent tokens of dimension256,1cross-attention head, and 8 self-attention heads. V. EXPERIMENTAL DESIGN A. Dataset Preparation For training and evaluation, we use the BEAT2 dataset [2], [3], which represents the state of the art for full-body gesture tracking. Despite its extensive use for gesture generation, the semantic emphasis and iconic gesture aspects of the dataset remain underexplored. BEAT2 contains around 2,000 entries of variable length, which, given their nature as monologue speeches, also vary in the number of iconic gestures present. To improve training and allow the model to better handle sparse iconic activations, we segment the data into sentences, yielding approximately 18,000 data points. Each data point consists of an utterance paired with one of the following emotions: sadness, neutral, anger, contempt, surprise, disgust, fear, or happiness. For consistency with Plutchik’s emotion model [8], validated across psychology and affective computing [28], [29], we rename the label ‘happiness’ to ‘joy’. All results are reported on the test set using an 80/20 train-test split. B. Baseline We compare our model against GPT-4o [30] as a baseline, given the demonstrated capability of LLMs in understanding semantic information [31]. The model was prompted to predict the intensity of iconic gestures per word, in the same format as the dataset, conditioned on the utterance’s emotion. The resulting values were binarized using the same threshold as our model to obtain placement predictions. TABLE I: Performance and computational cost across model configurations. Depth SAAccuracy PrecisionF1GFLOPsLatency (ms)↓ 2868.7853.9250.275.798.39 2468.6453.55 47.843.114.46 2268.7553.82 49.721.773.20 2168.6853.76 49.381.092.16 1868.5353.7649.572.904.02 1468.5653.68 48.981.552.45 1268.5953.47 47.590.781.74 1168.6453.5547.840.551.16 SA: self-attention blocks. The selected configuration is highlighted. TABLE I: Macro-averaged results for iconic placement per word. ModelAccuracyPrecisionRecallF1 LLM53.3652.6353.3652.92 Ours68.6453.5568.6447.84 F1 = F1 score. Underlined values are best per column. C. Model Configurations Since computational efficiency is a primary objective, we evaluate two architectural parameters: the number of cross- attention layers (Depth) and the number of sequential self- attention (SA) blocks. We explore the minimum configuration that delivers strong performance. VI. EXPERIMENTAL RESULTS A. Model Size Table I shows the results across all configurations. Classifica- tion performance remains stable, with accuracy ranging between 68.53%and68.78%, indicating that the task does not benefit from additional capacity. Computational cost, however, varies considerably: GFLOPs drop from5.79to0.55and latency from8.39ms to1.16ms as the model shrinks. We therefore select depth 1 with a single SA block for all subsequent experiments, as it achieves competitive performance at the lowest computational cost. This confirms that a minimal architecture is sufficient for this task, which is a desirable property for real-time robot deployment. B. Iconic Placement Table I reports macro-averaged results for iconic gesture placement. Our model outperforms the LLM baseline across all metrics, with a notable improvement in accuracy (68.64% vs.53.36%). The lower F1 score relative to accuracy reflects the class imbalance inherent in the task, as iconic gestures are sparse within utterances, making precise per-word prediction challenging for both models. C. Intensity Regression Table I reports regression results for iconic gesture intensity. Despite the task being challenging, our model outperforms GPT- 4o across all metrics, improving RMSE from0.22to0.15and Pearson correlation from0.09to0.20. The negative R 2 values for both models indicate that intensity prediction remains an Fig. 3: Semantic co-speech implementation on the social robot Haru. TABLE I: Regression results for word intensity of iconic gestures. ModelMAEMSERMSER 2 PRSpearman LLM0.080.050.22-1.230.090.06 Ours0.080.020.15-0.070.200.16 MAE = Mean Absolute Error, MSE = Mean Squared Error, RMSE = Root Mean Squared Error, R 2 = R-squared (coefficient of determination), PR = Pearson’s correlation coefficient, Spearman = Spearman’s rank correlation coefficient. open problem, likely due to the dataset’s subjective and sparse iconic gesture annotations. VII. DISCUSSION The results show that a lightweight, text-only model can outperform GPT-4o on both iconic gesture placement and intensity regression, using only the utterance and the target emotion as input. This suggests that task-specific training on word-level iconic annotations provides a stronger inductive bias for this problem than the general semantic knowledge encoded in large pretrained models. Placement results are strong, with our model achieving 68.64%accuracy against53.36%for the LLM baseline. Intensity regression remains more challenging for both models, as reflected by the negative R 2 values. This is likely due to the dataset’s subjective and sparse iconic gesture annotations, as well as the limited expressiveness of the current word-level embeddings. Exploring richer semantic representations is a natural direction for future work, alongside larger and more diverse datasets for this task, which remains underexplored in affective computing for embodied agents. VIII. ROBOT IMPLEMENTATION Our model can be deployed alongside any co-speech ap- proach that handles rhythmic motion [5], [10], [11], [22], [23], as it operates independently by predicting the placement and intensity of iconic gestures from text. We deploy it on the social robot Haru [32], as shown in Figure 3. Iconic gesture intensity values are mapped to a set of animations corresponding to the detected emotion and intensity level. When the model identifies a word that requires an iconic gesture, the robot executes the corresponding animation in real time. Although further evaluation across different robot platforms is needed, the implementation demonstrates the feasibility of the proposed approach in a real-world setting. IX. CONCLUSIONS We presented a lightweight, emotion-aware transformer for semantic gesture placement and intensity prediction in robot co-speech. Taking only text and a target emotion as input, the model outperforms GPT-4o on both tasks while remaining computationally compact, with a latency of1.16 ms on GPU. The implementation on Haru robot demonstrates its applicability to real-time embodied agents. The need for low-latency, real-time inference is a recognized challenge across a wide range of application domains [33]; our model, with a latency of 1.16 ms, directly addresses this constraint, remaining lightweight enough for deployment on embodied agents in real-world settings [34]. Future work should focus on improving intensity regression using richer embeddings and on generalizing the approach to other robot platforms and co- speech scenarios beyond iconic gestures, including gaze-aware and perceptually-grounded behaviours [35]. REFERENCES [1]S. Nyatsanga, T. Kucherenko, C. Ahuja, G. E. Henter, and M. Neff, “A comprehensive review of data-driven co-speech gesture generation,” in Computer Graphics Forum, vol. 42, no. 2. Wiley Online Library, 2023, p. 569–596. [2]H. Liu, Z. Zhu, G. Becherini, Y. Peng, M. Su, Y. Zhou, X. Zhe, N. Iwamoto, B. Zheng, and M. J. Black, “Emage: Towards unified holistic co-speech gesture generation via expressive masked audio gesture modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, p. 1144–1154. [3]X. Zhang, J. Li, J. Zhang, Z. Dang, J. Ren, L. Bo, and Z. Tu, “Semtalk: Holistic co-speech motion generation with frame-level semantic emphasis,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, p. 13 761–13 771. [4] M. Neff, M. Kipp, I. Albrecht, and H.-P. Seidel, “Gesture modeling and animation based on a probabilistic re-creation of speaker style,” ACM Transactions On Graphics (TOG), vol. 27, no. 1, p. 1–24, 2008. [5] T. Kucherenko, P. Jonell, S. Van Waveren, G. E. Henter, S. Alexandersson, I. Leite, and H. Kjellstr ̈ om, “Gesticulator: A framework for semantically- aware speech-driven gesture generation,” in Proceedings of the 2020 international conference on multimodal interaction, 2020, p. 242–250. [6]B. Wu, C. Liu, C. T. Ishi, and H. Ishiguro, “Probabilistic human-like gesture synthesis from speech using gru-based wgan,” in Companion pub- lication of the 2021 international conference on multimodal interaction, 2021, p. 194–201. [7] R. Ishii, S. Eitoku, and Y. Sato, “Impact of personality on generation of co-speech nonverbal behaviors represented by 3d skeleton pose,” in Proceedings of the 13th International Conference on Human-Agent Interaction, 2025, p. 247–256. [8]R. Plutchik, “The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice,” American scientist, vol. 89, no. 4, p. 344–350, 2001. [9]P. Chatziadam, A. Dimitriadis, S. Gikas, I. Logothetis, M. Michalodim- itrakis, M. Neratzoulakis, A. Papadakis, V. Kontoulis, N. Siganos, D. Theodoropoulos, G. Vougioukalos, I. Hatzakis, G. Gerakis, N. Pa- padakis, and H. Kondylakis, “Twifly: A data analysis framework for twitter,” Information, vol. 11, no. 5, 2020. [10]Y. Yoon, W.-R. Ko, M. Jang, J. Lee, J. Kim, and G. Lee, “Robots learn social skills: End-to-end learning of co-speech gesture generation for humanoid robots,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, p. 4303–4309. [11]X. Li and C. Dondrup, “A learning-based co-speech gesture generation system for social robots,” in Proceedings of the 12th International Conference on Human-Agent Interaction, 2024, p. 453–455. [12]E. Fern ́ andez-Rodicio, J. J. Gamboa-Montero, M. Maroto-G ́ omez, ́ A. Castro-Gonz ́ alez, and M. A. Salichs, “Evaluating the effect of co- speech gesture prediction on human–robot interaction,” International Journal of Human-Computer Studies, p. 103674, 2025. [13]T. Omine, N. Kawabata, and F. Homma, “Co-speech gesture and facial expression generation for non-photorealistic 3d characters,” in Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Posters, 2025, p. 1–2. [14]H. Liu, Z. Zhu, N. Iwamoto, Y. Peng, Z. Li, Y. Zhou, E. Bozkurt, and B. Zheng, “Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis,” in European conference on computer vision. Springer, 2022, p. 612–630. [15] N. Gao, Y. Bao, D. Weng, J. Zhao, J. Li, Y. Zhou, and P. Wan, “Sarges: Semantically aligned reliable gesture generation via intent chain,” in Proceedings of the International Workshop on Generation and Evaluation of Non-verbal Behaviour for Embodied Agents, 2025, p. 13–21. [16] S. Hochreiter, “Long short-term memory,” Neural Computation MIT- Press, 1997. [17]A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017. [18]S. Yang, Z. Wu, M. Li, Z. Zhang, L. Hao, W. Bao, M. Cheng, and L. Xiao, “Diffusestylegesture: Stylized audio-driven co-speech gesture generation with diffusion models,” arXiv preprint arXiv:2305.04919, 2023. [19] D.-S. Go, H.-J. Hyung, D.-W. Lee, and H. U. Yoon, “Andorid robot motion generation based on video-recorded human demonstrations,” in 2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN). IEEE, 2018, p. 476–478. [20]P. Bremner, A. G. Pipe, M. Fraser, S. Subramanian, and C. Melhuish, “Beat gesture generation rules for human-robot interaction,” in RO-MAN 2009-the 18th IEEE international Symposium on Robot and human interactive communication. IEEE, 2009, p. 1029–1034. [21]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, p. 4171–4186. [22]C. Y. Liu, G. Mohammadi, Y. Song, and W. Johal, “Speech-gesture gan: Gesture generation for robots and embodied agents,” in 2023 32nd IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, 2023, p. 405–412. [23]C. Yu and A. Tapus, “Srg 3: Speech-driven robot gesture generation with gan,” in 2020 16th International Conference on Control, Automation, Robotics and Vision (ICARCV). IEEE, 2020, p. 759–766. [24]N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” arXiv preprint arXiv:1908.10084, 2019. [25]P. Xu, A. Madotto, C.-S. Wu, J. H. Park, and P. Fung, “Emo2vec: Learning generalized emotion representation by multi-task training,” arXiv preprint arXiv:1809.04505, 2018. [26]S. Gkikas, I. Kyprakis, and M. Tsiknakis, “Efficient pain recognition via respiration signals: A single cross-attention transformer multi-window fusion pipeline,” in Companion Proceedings of the 27th International Conference on Multimodal Interaction, ser. ICMI Companion ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 70–79. [27]S. Gkikas and M. Tsiknakis, “Synthetic thermal and rgb videos for automatic pain assessment utilizing a vision-mlp architecture,” in 2024 12th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), 2024, p. 4–12. [28]E. C. Montiel-V ́ azquez, J. A. Ram ́ ırez Uresti, and O. Loyola-Gonz ́ alez, “An Explainable Artificial Intelligence Approach for Detecting Empathy in Textual Communication,” Applied Sciences, vol. 12, no. 19, p. 9407, Sep. 2022. [29]E. C. Montiel-V ́ azquez, C. Arzate Cruz, J. A. R. Uresti, and R. Gomez, “Empatheticexchanges: Toward understanding the cues for empathy in dyadic conversations,” IEEE Access, vol. 12, p. 195 097–195 110, 2024. [30]“GPT-4 technical report.” [Online]. Available: http://arxiv.org/abs/2303. 08774 [31]V. Havl ́ ık, “Meaning and understanding in large language models,” Synthese, vol. 205, no. 1, p. 9, 2024. [32]R. Gomez, D. Szapiro, K. Galindo, and K. Nakamura, “Haru: Hardware design of an experimental tabletop robot assistant,” in Proceedings of the 2018 ACM/IEEE international conference on human-robot interaction, 2018, p. 233–240. [33] D. Antonogiorgakis, A. Britzolakis, P. Chatziadam, A. Dimitriadis, S. Gikas, E. Michalodimitrakis, M. Oikonomakis, N. Siganos, E. Tzagkarakis, Y. Nikoloudakis, S. Panagiotakis, E. Pallis, and E. K. Markakis, “A view on edge caching applications,” 2019. [Online]. Available: https://arxiv.org/abs/1907.12359 [34] C. A. Cruz, Y. Sechayk, T. Igarashi, and R. Gomez, “Data augmentation for 3dmm-based arousal-valence prediction for hri,” in 2024 33rd IEEE International Conference on Robot and Human Interactive Communica- tion (ROMAN), 2024, p. 2015–2022. [35]R. S. Hessels and Y. Fang, “A visual perceptual perspective on gaze in social robotics,” Psychonomic Bulletin & Review, vol. 33, no. 4, p. 131, 2026.