Paper deep dive
When the Inference Meets the Explicitness or Why Multimodality Can Make Us Forget About the Perfect Predictor
J. E. DomĂnguez-Vidal, Alberto Sanfeliu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 9:12:31 PM
Summary
This study evaluates human-robot collaboration in a physical object transportation task, comparing implicit intention inference systems (force-based and velocity-based predictors) against explicit communication systems (button interface and voice commands). Using the IVO mobile social robot, the research demonstrates that once technical performance reaches a sufficient threshold, humans do not perceive further improvements in inference systems. Conversely, humans prefer more natural communication methods (voice) despite higher error rates. The optimal strategy combines both inference and explicit communication, offering the highest acceptance and freedom for collaboration.
Entities (7)
Relation Signals (8)
Voice Command Recognition â elicits â Explicit Intention
confidence 95% · explicit communication methods ... elicit human intention
Button Interface â elicits â Explicit Intention
confidence 95% · explicit communication methods ... elicit human intention
IVO â uses â Force Sensor
confidence 95% · equipped with force sensor to detect the force exchange between both agents
IVO â uses â LiDAR
confidence 95% · equipped with ... LiDAR to detect the environment
Human â prefers â Natural Communication Systems
confidence 92% · the human prefers systems that are more natural to them even though they have higher failure rates
Velocity Predictor â infers â Human Intention
confidence 90% · inference systems that try to predict human intentions
Force Predictor â infers â Human Intention
confidence 90% · inference systems that try to predict human intentions
Combined System â outperforms â Single Strategy Systems
confidence 90% · the preferred option is the right combination of both systems
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Although in the literature it is common to find predictors and inference systems that try to predict human intentions, the uncertainty of these models due to the randomness of human behavior has led some authors to start advocating the use of communication systems that explicitly elicit human intention. In this work, it is analyzed the use of four different communication systems with a human-robot collaborative object transportation task as experimental testbed: two intention predictors (one based on force prediction and another with an enhanced velocity prediction algorithm) and two explicit communication methods (a button interface and a voice-command recognition system). These systems were integrated into IVO, a custom mobile social robot equipped with force sensor to detect the force exchange between both agents and LiDAR to detect the environment. The collaborative task required transporting an object over a 5-7 meter distance with obstacles in the middle, demanding rapid decisions and precise physical coordination. 75 volunteers perform a total of 255 executions divided into three groups, testing inference systems in the first round, communication systems in the second, and the combined strategies in the third. The results show that, 1) once sufficient performance is achieved, the human no longer notices and positively assesses technical improvements; 2) the human prefers systems that are more natural to them even though they have higher failure rates; and 3) the preferred option is the right combination of both systems.
Tags
Links
- Source: https://arxiv.org/abs/2602.18850v1
- Canonical: https://arxiv.org/abs/2602.18850v1
Trouble viewing inline? Open PDF directly â
Full Text
53,431 characters extracted from source content.
Expand or collapse full text
Springer Nature 2021 L A T E X template When the Inference Meets the Explicitness or Why Multimodality Can Make Us Forget About the Perfect Predictor J. E. Dom Ìınguez-Vidal 1,2* and Alberto Sanfeliu 1,2 1* Institut de Rob`otica i Inform`atica Industrial (CSIC-UPC), Llorens i Artigas 4-6, Barcelona, 08028, Spain. 2 Universitat Polit`ecnica de Catalunya - BarcelonaTech (UPC), Jordi Girona, 31, Barcelona, 08034, Spain. *Corresponding author(s). E-mail(s): jdominguez@iri.upc.edu; Contributing authors: alberto.sanfeliu@upc.edu; Abstract Although in the literature it is common to find predictors and inference systems that try to pre- dict human intentions, the uncertainty of these models due to the randomness of human behavior has led some authors to start advocating the use of communication systems that explicitly elicit human intention. In this work, it is analyzed the use of four different communication systems with a human-robot collaborative object transportation task as experimental testbed: two intention pre- dictors (one based on force prediction and another with an enhanced velocity prediction algorithm) and two explicit communication methods (a button interface and a voice-command recognition sys- tem). These systems were integrated into IVO, a custom mobile social robot equipped with force sensor to detect the force exchange between both agents and LiDAR to detect the environment. The collaborative task required transporting an object over a 5â7 meter distance with obstacles in the middle, demanding rapid decisions and precise physical coordination. 75 volunteers perform a total of 255 executions divided into three groups, testing inference systems in the first round, communi- cation systems in the second, and the combined strategies in the third. The results show that, 1) once sufficient performance is achieved, the human no longer notices and positively assesses techni- cal improvements; 2) the human prefers systems that are more natural to them even though they have higher failure rates; and 3) the preferred option is the right combination of both systems. Keywords: Physical Human-Robot Interaction, Intent Detection, Human-in-the-Loop, User Study 1 Introduction Originally, robotics began by performing small, routine and repetitive tasks to relieve humans of this burden. As perception systems improved, robots were able to better perceive and under- stand the environment around them allowing them to perform more complex tasks by being able to make the right decision when navigating an urban environment [1] or choosing the right tool [2]. However, when these machines started to interact with humans, perceiving the environment was no 1 arXiv:2602.18850v1 [cs.RO] 21 Feb 2026 2Article Title longer enough and we started to need to under- stand the humanâs intention as well since their behavior was too uncertain [3, 4]. This is when we began to develop predictors to allow the robot to infer human intent with increasingly higher success rates [5, 6]. While their non-negligible error rates are often blamed on lim- itations in computational capacity or a lack of sufficient data for training, the fact that humans can model the same information we perceive in multiple ways [7] making it so that two humans can represent the same environment differently begs the question of whether we will ever have a perfect predictor. This very question has caused some authors [8â 11] to begin to consider it necessary to combine inference engines with communication systems that make it possible to obtain the humanâs inten- tion explicitly. In this way, they hope to achieve robots that do not just look like humans, but act like humans: they ask questions and request infor- mation when the one they have does not allow them to make decisions with sufficient confidence. Thus, this work arises as a continuation of our previous work [11] where in a collaborative transportation task we first confronted on the one hand an inference system to obtain the humanâs intention from the instantaneous force they are exerting against, on the other hand, a button- based communication system to allow the human to express their intention explicitly. That work proved that allowing humans to directly express themselves can achieve the same improvement in multiple aspects of an effective Human-Robot Interaction (HRI) as using an intention predic- tor. However, it also left unanswered questions, such as what happens if both systems are com- bined, and postulated that we should not continue looking for technical improvements in the predic- tor used, since these may go unnoticed by the human, but that we should pivot towards meth- ods that seek to improve human-robot commu- nication, although without demonstrating these assertions. This work comes to test those consid- erations analysing human preferences at the time of working with a robot in a task with fast phys- ical contact and to answer that question through three rounds of experiments with their respective user studies. Thus, our contributions would be as follows. In the first round of experiments, we compare two predictors with different success rates and find that, once the failure rate is reduced to an accept- able value, the human no longer perceives any improvement. In the second round of experiments, we compare two direct communication systems and find that the human prefers the one that is more natural even though it is technically inferior in terms of response delay and failure rate. Finally, we compare the system preferred by the human in each of the first two rounds of experiments as well as the combination of both to verify that this combination is the most accepted by the human as it offers more freedom to collaborate with the robot, being this our third contribution. In the remainder of the document, we present the related work in Section 2. Section 3 includes an explanation of the task selected for this study, the collaborative transport of objects, as well as all the relevant details of all the systems employed in this work. Section 4 presents the hypotheses to be tested, the setup and methodology employed as well as the distribution of participants who per- formed the experiments. Section 5 presents the results obtained in each round of experiments and Section 6 shows a short discussion of these results. Finally, Section 7 contains the conclusions. 2 Related Work Of the two strategies discussed in the previ- ous Section to know humanâs intention, inference engines and direct communication systems, the former is relatively abundant in the literature [12â 15]. Most of these works use different architectures based on Gaussian Mixture Models (GMM) or some kind of Artificial Neural Network (ANN) to obtain a prediction of the humanâs intention whether this is the trajectory they are going to follow, the movement they are going to make with their hand or the next object they are going to pick up. Applied to collaborative transportation tasks or, in general, tasks with physical contact, it has been common to use control techniques to make the robot adapt to the humanâs wishes using both impedance and admittance controllers [16â 20]. More recent work, on the other hand, has to include some kind of predictor, either of the desired trajectory for the object [21] or of its velocity profile [22] or even of the force that the human is going to exert on the object in the short term [23]. None of these works contemplate Article Title3 the possibility of the human communicating with the robot to reduce uncertainty or resolve any problems generated by an error in the robotâs inference. This second strategy, a direct communication system, is less common. In [24] the authors design a smartphone application with which human and robot can communicate bidirectionally over long distances. This app is applied in [25] to collabo- rative search. Thanks to it, the robot can know both the route that the human is following and the one they want to follow even in urban sce- narios with multiple occlusions that do not allow the robot to track the human. This makes it pos- sible to minimize the overlapping of the areas explored by each agent. Another example where the human is allowed to explicitly communicate with the robot is [26]. In this, the robot infers the goal of the task that the human wants the robot to perform, but before executing it, it asks for con- firmation from the human through an augmented reality (AR) based system. In [27] both strate- gies are combined to improve object manipulation between two robots. To this end, both robots com- municate their plans both implicitly through the force they exert on the object and explicitly by exchanging wireless messages. Finally, [28] allows the human to use a combination of gestures and voice commands to tell the robot which object to pick up and where to take it. Regarding voice commands recognition and applied to robotics, although there are old works that tried to allow the human to transmit simple movement commands to a mobile robot [29, 30], it was not until the proliferation of Artificial Neu- ral Networks (ANN) and the emergence of large datasets containing lists of typical commands [31] that satisfactory results began to be achieved by detecting finite lists of commands [32, 33]. With the exception of [26], none of the above works com- pares their communication system with another one. Moreover, all of them assume that the human is willing to use their system. In this article, we do not take that for granted and perform both the comparison of multiple systems and that ver- ification. Additionally, and to the best of our knowledge, this is the first work that tries to combine both strategies applied to this use case. Fig. 1 Experiments setup. Top Left - Human-robot pair collaboratively transporting an aluminium bar. Goal marked with a chequered flag. Top Right - Scheme of the designed setup. At least eight routes to the goal. Control desk on the right with researcher managing the experiment and camera recording the point of view on the left. Bot- tom Left - Handle for better ergonomics with meaning of each button next to it. Only three buttons are used to tell the robot which route the human wants. Bottom Right - RODE Wireless GO microphone used for voice command recognition. 3 Collaborative Object Transportation Collaborative object transportation is a task in which a human and a robot collaboratively trans- port an object from a starting point to a destina- tion point that may or may not be predefined in advance. It is therefore a task that is performed in close proximity and in which there is an exchange of physical forces between the two agents. These two characteristics mean that the robotâs move- ments cannot be abrupt, so as not to harm the human next to it, and that the robotâs response speed must be high in order to be able to adapt to the rapid changes that the human may make. Additionally, this task usually involves mov- ing an object at a short distance using only a robotic arm. In this work, we will transport an object 5-7 m through a scenario with multiple obstacles so that there are multiple routes to the predefined goal or even change the route on the fly. This implies that the robot must move its platform through the scenario adapting to the route desired by the human and that the human must make multiple decisions along this route (see Fig. 1 - Top). 4Article Title To enable the robot to perform this task, we start from the implementation in [34]. In this, tak- ing advantage of the fact that this is a task in which the exchange of information is mainly done through forces, the environment is modeled by means of virtual forces: a repulsive force for each obstacle detected by the robot and an attractive force for the partial goals that the human-robot pair should follow to reach the place where to leave the object. These partial goals are obtained from the waypoints generated by a global planner. The force resulting from modeling the environ- ment is then combined with the force exerted by the human on one end of the transported object and measured by a force sensor on the robotâs wrist to which the other end of the object, in this case an aluminum bar, is attached. This com- bined force is sent to a controller that generates the robotâs speed commands. Additionally, this system is conditioned using the human intention. For this purpose, we start from the distinction between implicit and explicit intention shown in [35]. Thus, we consider as implicit intention that which can be inferred from the actions of the other agent and as explicit intention that which is expressed using a direct communication channel and a code known and agreed upon by both agents. 3.1 Inference of Implicit Intention Two similar inference systems, representative of the trend in the literature towards using predic- tors, will be used to obtain the implicit intention of the human. First, a force predictor which, from the previous values of the force exerted by the human, the velocity of the human-robot pair, the representative force of the environment and the robotâs LiDAR readings, generates a prediction of the force that the human will exert during the next 1 s. This prediction can be processed to obtain an estimate of the route the human wishes to fol- low in the short term, and with this, condition the robotâs planner to match their intention. This force predictor used that way has a fun- damental shortcoming and that is that when predicting the route it only takes into account the contribution of the human through their force and not the contribution performed by the robot, e.g., avoiding getting too close to an obstacle or damping rapid changes in the force exerted by the human. This is the reason why in our previous work [11] we suggested that the predictor used had room for improvement. To solve this, we develop a second predictor that, in addition to the force to be exerted by the human, predicts the velocity of the human-robot pair during the next 1 s. This second prediction can be directly integrated to obtain an estimate of the short-term desired route with which to also condition the robotâs planner. While the techni- cal details of this second predictor are outside the scope of this paper [36], it allows to reduce the L2 error made in estimating the trajectory at 1 s from 0.199 m with the force predictor to 0.138 m with this velocity predictor. 3.2 Direct Communication of Explicit Intention To obtain the explicit intention of the human we also used two systems. First, the same system with three buttons on the objectâs handle used in [11]. By pressing each of them, the human can tell the robot whether to go straight ahead, turn left or turn right. Information that the robot uses to condition its planer at the next intersection. The second system implemented is a voice command recognition model. With this, the human can ver- bally tell the robot their intention using the âGoâ, âLeftâ and âRightâ commands. These, generate the same conditioning as with the previous buttons (see Fig. 1 - Bottom). The system with buttons is infallible in that the only possible error is that the human presses the wrong button and its processing delay is negli- gible. Meanwhile, voice command recognition can make mistakes (hit rate: 94.75%). Moreover, if a Bluetooth microphone is used as in our case, the delay between the human saying a command and it taking effect can go from 0.2 to 1 s if the RF cell is saturated. 4 Evaluation We conducted 3 rounds of experiments to test whether technical improvement in inferring humanâs desired route has any effect, whether the human prefers a more natural system in expressing intent over technical efficiency, and what effect the combination of both types of systems may have. Article Title5 4.1 Experiments Setup and Methodology Both the starting point and the goal to take the object to are preset and are the same in all exper- iments. All of them are carried out indoors on a stage with OptiTrack on the ceiling to track both agents and thus know the covered distance and the duration of each experiment. The robot used is IVO [37], which has a force sensor on each wrist and can perceive obstacles by means of front LiDAR and rear LaserScan. The experiments per- formed last a minimum of 36 s and a maximum of 110 s. In the first round of experiments, the useful- ness of the two predictors mentioned above in achieving an effective HRI is tested. For that, each volunteer performs three experiments: one without any predictor plus one experiment with each predictor 1 . The second round of experiments allows us to compare both explicit communication systems. For this, the same procedure is followed: one experiment without any communication sys- tem plus one experiment with each 2 . Finally, for the third round, the predictor/- communication system best rated by the volun- teers in each previous round is selected and both strategies are compared. For this, each volunteer performs four experiments: a baseline without pre- dictor or explicit communication system plus one experiment with each strategy plus one experi- ment with both strategies available 3 . In the three rounds of experiments, the baseline experiment is run first followed by the remaining two (three) experiments in random order to avoid statistical distortions. At the end of each experiment, each volunteer fills out a handmade questionnaire to assess both numerically from 1 to 7 and by choosing among the different systems tested different aspects asso- ciated to an effective HRI (evaluation questions shown in the Appendix). The numerical ratings are then analyzed by means of different tests. First, the Saphiro-Wilkâs and Leveneâs tests are applied. If the variable analyzed meets the nor- mality condition, an ANOVA test with Bonferroni correction is applied to check whether there is 1 1st round example: https://youtu.be/c4aPo6WRK4M 2 2nd round example: https://youtu.be/6VL41XovKJg 3 3rd round example: https://youtu.be/mL8DQb1bK 4 a statistically significant variation (p<0.05), in which case, a Tukeyâs HSD (Honest Significant Difference) test is applied. If the normality condi- tion is not met, a non-parametric Kruskal-Wallis test is applied followed by a Nemenyiâs test if sta- tistically significant variation is detected between the systems analyzed in each case. After filling in all the questionnaires, a short interview with three open questions is conducted, giving the volun- teers the opportunity to express themselves freely: What did you think of the whole experiment? What were your feelings during each attempt? What would you improve? 4.2 Hypotheses Relative to intention predictors: H1a - Adding a predictor to the robotâs deci- sion making system reduces the humanâs effort. H1b - Once a sufficient hit rate is reached, the human ceases to positively value technical improvements. Relative to intention communication systems: H2a - Adding a way to explicitly indicate the humanâs intention improves multiples aspects of an effective HRI. H2b - The human prefers a natural communi- cation system with a non-negligible error rate over a more robust one. Relative to the comparison of both systems: H3a - A system that allows the human to directly express their intention improves multi- ple aspects of an effective HRI just as much as a system that attempts to infer it. H3b - A combination of an inference system with an explicit communication one is the option best valuated by the human and the preferred one. 4.3 Participants A total of 75 volunteers were recruited from our research institute as well as from different schools of the partner university. 22 volunteers (age: ÎŒ=26.45, Ï=4.02; 23% female) participated in the first round of experiments performing 66 experiments (3 each). 23 volunteers (age: ÎŒ=27.36, Ï=4.87; 26% female) participated in the second performing 69 runs (also 3 each). Finally, 30 volunteers (age: ÎŒ=28.32, Ï=5.12; 27% female) participated in the third round by performing 120 experiments (4 each). 6Article Title Fig. 2 Assessment of objective measurements and election make by volunteers (predictor experiments). Left: Mean force exerted in orange, maximum force exerted in light blue and duration in yellow for the three experiments. Left axis in Newtons (both forces) and right axis in seconds (duration). Bars represent std. dev. Right: Election made by the 22 volunteers instead of valuate aspects numerically. Force predictor in dark red, velocity predictor in light red and draw in yellow. Fig. 3 Assessment of the main aspects involved in the interaction (predictor experiments). Left: Comparison among the baseline experiment (without any predictor) in gray, experiment with force predictor in dark red and with velocity predictor in light red. Valuation from 1 (very low) to 7 (very high). Statistical significance marked with *: p < 0.05, **: p < 0.01, ***: p < 0.001. Bars represent std. dev. Right: Election made by the 22 volunteers with respect to which system they prefer for the task at hand. Force predictor in dark red, velocity predictor in light red. All the experiments reported in this article have been performed after the approval of the ethics committee of the Universitat Polit`ecnica de Catalunya (UPC) in accordance with all the rele- vant guidelines and regulations (ID: 2023.05) and all the volunteers have signed an informed consent form. No volunteers were paid for participating in this study, ensuring that there is no conflict of interest. 5 Results Before analyzing each round of experiments, we perform a post-hoc statistical power test to know what values we can be statistically sure of. Thus, using the criterion of p < 0.05, for the first round (22 volunteers) we can detect effect sizes as low as η 2 =0.138 with a statistical power of 80%. For the second (23 volunteers), as η 2 =0.133; and for the third round (30 volunteers), as η 2 =0.089. All variables analyzed by variance tests are normally distributed according to the Shapiro-Wilkâs test unless otherwise indicated. 5.1 Force Predictor VS. Velocity Predictor To test hypothesis H1a, we performed three objective measures (experiment duration, mean force, and maximum force exerted by the human during the experiment) in the three experiments comprising this first round. Fig. 2 - Left shows the result. No statistically significant variation is observed in any of the measures (p=0.48, p=0.15 and p=0.18 respectively) so it cannot be claimed that adding a predictor reduces the humanâs effort in any way. H1a is rejected. To test H1b, several tests are performed. First, the volunteers are asked to rate from 1 to 7 multiple aspects corresponding to an effective HRI. Fig. 3 - Left shows the result. As can be seen, statistically significant increases occur for all variables with these increases being similar for both predictors: âRobot contribution to per- formanceâ (F (2, 63)=15.89, p<0.001, η 2 =0.335), âHuman responsibilityâ (F (2, 63)=6.88, p=0.002, η 2 =0.179), âTrust in Robotâ (F (2, 63)=5.46, Article Title7 Fig. 4 Assessment of the subjective easiness perceived by the user to indicate their intention. Left: Com- parison among the three experiments performed in the first round: baseline, force predictor and velocity predictor. Middle: Comparison among the three experiments performed in the second round: baseline, with buttons and with voice commands recognition. Right: Comparison among the four experiments performed in the third round: baseline, velocity predictor, voice commands recognition and both systems. Valuation from 1 (very low) to 7 (very high). Statistical significance marked with *: p < 0.05, **: p < 0.01, ***: p < 0.001. Bars represent std. dev. p=0.006, η 2 =0.147), âComfortâ (F (2, 63)=12.91, p<0.001, η 2 =0.291). The exceptions are âRobot contribution to fluencyâ where the velocity pre- dictor generates a greater increase (F (2, 63)=8.44, p<0.001, η 2 =0.211; with force predictor: p=0.011; with velocity predictor: p<0.001) and âRobot con- tribution as Humanâ where the force predictor is the one with the biggest increase (F (2, 63)=10.74, p<0.001, η 2 =0.254; with force predictor: p<0.001; with velocity predictor: p=0.0012). No notably more positive valuations are therefore detected for the second predictor despite reducing the mean error in trajectory estimation by 30.6% (0.138 m versus 0.199 m) except if the subjective aggres- siveness of the robotâs movements (one of the questions associated to the âRobot contribution to fluencyâ block in our questionnaire, see the Appendix) is independently analyzed in which there is an increase in the case of using the first predictor although without being statistically significant (p=0.051). Additionally, volunteers are asked to explicitly choose between the two predictors, accepting the draw also as a valid option. Fig. 2 - Right shows the result. In general, the velocity predictor is con- sidered to be safer and more fluid, and the force predictor is considered to execute the task faster. The draw is the predominant choice as to which one is easier to use or which one makes the robot behave more naturally. Finally, volunteers are asked to choose which predictor they find most appropriate for perform- ing the task at hand, not giving the draw as an option. Fig. 3 - Right shows a complete draw on this question. Therefore, H1b is confirmed, as the volunteers have not indicated a preference for the velocity predictor over the force predic- tor despite being technically superior. Some of the volunteersâ comments in the post-experiments interview confirms this result. Volunteer 1.6 com- mented âIf you have changed anything between the second (velocity) and the third (force) experiment, I havenât noticed itâ. Volunteer 1.13 commented âThe last one (velocity) seemed smoother to me but both work correctlyâ. At the end of the questionnaire that volunteers fill out after each run, we added a task-specific control question to check that they understood that the various methods they were testing were designed to allow them to indicate their intention to the robot. Fig. 4 - Left shows the volun- teersâ subjective ratings of how easy they found it to indicate their intention to the robot in this first round of experiments. There is a statisti- cally significant increase when using any predictor but this increase is no greater when using the most objectively accurate predictor (performing a Kruskal-Wallis test since it does not pass the Shapiro-Wilkâs test followed by a Nemenyi test: H=17.06, p<0.001; with force predictor: p<0.001; with velocity predictor: p=0.0021) reaffirming the hypothesis H1b. 5.2 Buttons VS. Voice Commands Recognition If we perform the same objective measurements that were performed in the first round of exper- iments, we observe a reduction in the mean and maximum force exerted by the human in the case of using buttons although without being statisti- cally significant (see Fig. 5 - Left). This generates 8Article Title Fig. 5 Assessment of objective measurements and election make by volunteers (direct communication experiments). Left: Mean force exerted in orange, maximum force exerted in light blue and duration in yellow for the three experiments. Left axis in Newtons (both forces) and right axis in seconds (duration). Statistical significance marked with *: p < 0.05. Bars represent std. dev. Right: Election made by the 23 volunteers instead of valuate aspects numerically. Buttons in dark blue, voice commands in light blue and draw in yellow. Fig. 6 Assessment of the main aspects involved in the interaction (direct communication experiments). Left: Comparison among the baseline experiment (without any communication system) in gray, experiment with buttons in dark blue and with voice commands in light blue. Valuation from 1 (very low) to 7 (very high). Statistical significance marked with *: p < 0.05, **: p < 0.01, ***: p < 0.001. Bars represent std. dev. Right: Election made by the 23 volunteers with respect to which system they prefer for the task at hand. Buttons in dark blue, voice commands in light blue. an increase in the duration of the experiment that we cannot consider as statistically significant as it lacks sufficient statistical power(F (2, 66)=3.15, p=0.049 η 2 =0.087). It cannot, therefore, be indi- cated that there is a reduction in human effort. To test H2a, we use the numerical assess- ment made by the volunteers using the same previous questionnaire (see Fig. 6 - Left). A statistically significant improvement is observed in all the aspects analyzed, except in âHuman responsibilityâ in which we lack sufficient statis- tical power (F (2, 66)=3.99, p=0.023 η 2 =0.108), being this always higher in the case of the use of voice commands. We highlight âRobot con- tribution to fluencyâ (F (2, 66)=7.66, p=0.001, η 2 =0.188; with buttons: p=0.032; with voice com- mands: p<0.001) and âComfortâ (F (2, 66)=8.67, p<0.001, η 2 =0.208; with buttons: p=0.014; with voice commands: p<0.001). There is also a sta- tistically significant increase in âRobot contribu- tion to performanceâ (F (2, 66)=19.63, p<0.001, η 2 =0.373) using voice commands relative to using buttons (p=0.042) and a statistically significant reduction in âRobot aggressivenessâ (perform- ing a Kruskal-Wallis test since it does not pass the Shapiro-Wilkâs test followed by a Nemenyi test: H=9.59, p=0.005; with voice commands: p=0.003). H2a is therefore confirmed. For the sake of completeness, the rest of results are as follows: âRobot contribution as Humanâ (F (2, 66)=10.05, p<0.001, η 2 =0.233), âTrust in Robotâ (F (2, 66)=7.66, p=0.001, η 2 =0.188). To test H2b, we asked the volunteers to choose between the two explicit communication systems (see Fig. 5 - Right) with the system with but- tons being considered faster when communicating and the system with command recognition com- ing out victorious in all other aspects analyzed. Finally, the volunteers are asked to choose which system seems more appropriate for the task being the system with voice commands chosen by 86.9% of them (see Fig. 6 - Right). H2b is therefore confirmed. Some of the volunteersâ comments shed light on these results. Volunteer 2.8 commented âEven though I sometimes have to repeat the command, Article Title9 Fig. 7 Assessment of objective measurements and election make by volunteers (both systems experiments). Left: Mean force exerted in orange, maximum force exerted in light blue and duration in yellow for the four experiments. Left axis in Newtons (both forces) and right axis in seconds (duration). Statistical significance marked with *: p < 0.05, **: p < 0.01. Bars represent std. dev. Right: Election made by the 30 volunteers instead of valuate aspects numerically. Velocity predictor in light red, voice commands in light blue and both systems in green. Fig. 8 Assessment of the main aspects involved in the interaction (both systems experiments). Left: Com- parison among the baseline experiment (without predictor nor voice commands) in gray, experiment with velocity predictor in light red, with voice commands in light blue and with both in green. Valuation from 1 (very low) to 7 (very high). Sta- tistical significance marked with *: p < 0.05, **: p < 0.01, ***: p < 0.001. Bars represent std. dev. Right: Election made by the 30 volunteers with respect to which system they prefer for the task at hand. Velocity predictor in light red, voice commands in light blue and with both in green. I prefer to be able to talk to the robotâ. Volun- teer 2.17 said âThere is a delay until my command takes effect that with the first one (buttons) it doesnât happen but the second one (voice com- mands) allows me to focus more on exerting forceâ. Analyzing the result of the control question in which volunteers rate the ease with which they can communicate their intention to the robot, Fig. 4 - Middle shows how volunteers consider that command recognition allows them to indicate their intention more easily (with a Kruskal-Wallis test since it does not pass the Shapiro-Wilkâs test followed by a Nemenyi test: H=19.46, p<0.001; with buttons: p=0.0048; with voice commands: p<0.001) reaffirming the hypothesis H2b. 5.3 Complete System We take the velocity predictor (simply because it seems to generate less aggressive movements since there is not any noticeable preference between them) and the voice commands recognition sys- tem as the two preferred systems by the human for inference and direct communication respec- tively. With these, we perform a last round of experiments to compare both of them and their combination. Looking at the same objective measures used above, only in the mean force exerted by the human is observed a statistically significant vari- ation (F (3, 116)=5.71, p=0.0012 η 2 =0.129) (see Fig. 7 - Left). Specifically, there is a statisti- cally significant increase when using the velocity predictor relative to the baseline (p=0.009) and statistically significant decreases when using the voice commands or the complete system relative to using the predictor (p=0.002 and p=0.014) but not relative to the baseline (p=0.963 and p=0.998). To check H3a the numerical results obtained in the previous rounds cannot be directly com- pared as they were performed by different vol- unteers but require the same people to use 10Article Title both systems. Using the same questionnaire used in previous rounds (see Fig. 8 - Left), it can be observed that both systems produce sta- tistically significant increases in all analyzed parameters except for âHuman responsibilityâ where again we do not have enough statisti- cal power (F (3, 116)=2.95, p=0.036 η 2 =0.071). The system with velocity predictor achieves larger increases in âRobot contribution to fluencyâ (F (3, 116)=21.42, p<0.001 η 2 =0.356; with predictor: p<0.001; with voice commands: p=0.0011) and âRobot contribution to per- formanceâ (F (3, 116)=22.89, p<0.001 η 2 =0.372; with predictor: p<0.001; with voice commands: p=0.0022), while the system with voice com- mands outperforms it in âTrust in Robotâ (F (3, 116)=18.88, p<0.001 η 2 =0.328; with predic- tor: p=0.014; with voice commands: p<0.001) and âComfortâ (F (3, 116)=22.00, p<0.001 η 2 =0.362; with predictor: p=0.0061; with voice commands: p<0.001). H3a is therefore confirmed. For the sake of completeness: âRobot contribution as Humanâ (F (3, 116)=28.13, p<0.001 η 2 =0.421). To verify H3b, Fig. 8 - Left indeed shows that the system that makes use of both options is the one that scores better in all the aspects ana- lyzed, being the only one that manages to reduce the perceived aggressiveness in the robotâs move- ments in a statistically significant way (performing a Kruskal-Wallis test since it does not pass the Shapiro-Wilkâs test followed by a Nemenyi test: H=9.16, p=0.011; complete system: p=0.014). If we ask the volunteers to choose between the three systems (see Fig. 7 - Right), they choose the system with voice commands as the easiest to use and the complete system that makes use of both options in all other aspects analyzed. Finally, when the volunteers were asked to choose which system they considered most appropriate for the task, 86.7% of them opted for the complete system. Therefore, H3b is confirmed. Comments from some volunteers confirm these results. Volunteer 3.19 said, âThe predictor makes it more fluid, but being able to talk to the robot gives me extra security and peace of mindâ. Vol- unteer 3.24 commented, âThey are very different approaches that I think can serve different pur- poses [...] give me both and I choose when and how to use eachâ. Finally, in terms of how easily volunteers feel they can communicate their intention to the robot, Fig. 4 - Right shows once again the vol- unteersâ preference for the system that combines both methods as they can use whichever one they feel more comfortable with at any given time (with a Kruskal-Wallis test since it does not pass the Shapiro-Wilkâs test followed by a Nemenyi test: H=46.39, p<0.001; with predictor: p=0.0013; with voice commands: p<0.001, with both: p<0.001) reaffirming hypothesis H3b. 6 General Discussion The first surprising result of this study is that none of the systems tested seem to produce a reduction in human effort, understood as the mean and max- imum force exerted during the task. In the case of using any of the predictors, it seems that there is even an increase that only became statistically significant in the third round of experiments. It is worth mentioning that the volunteers were aware that a predictor was being used in the experiment that this was receiving as input the previously exerted force, although without knowing which predictor specifically was in use. This is because in the second and third round of experiments, whether using buttons or command recognition, the volunteer needed to know of their existence in order to use them, so we decided to inform the vol- unteers of the existence of a predictor when it was being used to make the experiments comparable to each other. It is possible that this would encourage a greater force on the part of the human to make it easier for the predictor to infer their intention, understood as the desired path. At the same time, we do not believe that the non-significant reduc- tion in the force exerted when using the buttons is so much due to the fact that the buttons encourage less effort. Observing these experiments, humans tend to pay attention to the markings on the han- dle before pressing any button to make sure they do not make a mistake, which means that the force exerted during that time naturally tends to be lower. As for the comparison between the two predic- tors, that a considerable technical improvement goes unnoticed in the humanâs subjective assess- ment of it, confirming the hypothesis H1b, goes against us continuing to improve our predictor. It also seems to indicate that a perfect predictor is not necessary, but simply a good enough one. This result is not entirely unexpected, as it fits with the Article Title11 Pareto rule or the Law of Diminishing Returns in economics. As for the two explicit communication systems, the confirmation of the hypothesis H2b as well as comments such as those of volunteers 2.8 and 2.17 seem to indicate that naturalness is more valued over technical aspects such as lower delay and lower failure rate. These two results combined support the idea that we should pivot the current trend of attempting to infer humanâs intention in the best possible way towards methods that seek to improve human-robot communication making it as humane as possible. It is worth mentioning that this work should not be understood as being against the use of pre- dictors. In one of the experiments performed in the third round of experiments, the Wi-Fi net- work used for the exchange of information between the robot and the computer running the control algorithm was saturated, causing delays in the generation of the robotâs speed commands accord- ing to what its sensors were detecting at any given moment. This made the interaction with the robot complex and counter-intuitive. The inclusion of the 1 s prediction made it possible to compen- sate for these delays and allow the human to perform the task satisfactorily. This is an illustra- tive example of the usefulness of using predictors. Another would be any situation in which no direct communication could take place. This is why we do not advocate discarding the use of predictors, but rather their correct combination with explicit communication systems that are as natural as pos- sible, thus taking advantage of the benefits of both types of systems. This explains the title of this article. The anal- ysis of human preferences carried out in this work leads us to consider multimodality, understood as the use of multiple communication channels of dif- ferent and even redundant nature, as the preferred option for humans when interacting with a robot. This allows us not to have to worry about achiev- ing a perfect predictor. In the end, this fits with our behavior as humans: we use our prior knowl- edge and experience to try to predict the behavior of our peers but, when the uncertainty is too high or we simply do not know the other person well, we choose to communicate directly to avoid misun- derstandings that could harm the outcome of the interaction. It is therefore to be expected that the human expects the same behavior from the robot if we want them to be considered our partners and not mere machines. 7 Conclusions In this work we have conducted three rounds of experiments using a collaborative transporta- tion task to test the following: 1) Predictor-based inference systems can improve multiple aspects of effective HRI. However, there is a sufficient per- formance beyond which technical improvements go unnoticed by the human using them. 2) The human prefers explicit communication methods with the robot that are natural even if this is at the cost of a higher failure rate. 3) Both systems separately can achieve the same subjective ratings by the human in multiple aspects, but it is when properly combined that they achieve the best HRI in terms of fluency, trust in the robot and comfort among others. This study should be replicated in other tasks to confirm those results. In any case, these find- ings can serve as a stepping stone to encourage and justify the use and evolution of human-robot communication systems that seek greater natural- ness, such as natural language or gestures, even if these systems may present errors, simply because humans are willing to accept them. Acknowledgments. The authors want to express their gratitude to Sergi Hern Ìandez for their technical support and to all the volunteers who made this work possible. Declarations Funding. Work supported under the Euro- pean project CANOPIES (H2020- ICT-2020-2- 101016906), the MINECO/AEI ROCOTRANSP project (PID2019-106702RB-C21 MCIN/ AEI /10.13039/501100011033) and the JST Moonshot R & D Grant Number JPMJMS2011-85. The first author acknowledges Spanish FPU grant with ref. FPU19/06582. Competing interests. There are none poten- tial conflicts of interest that could bias the evalu- ation or results of our research. Ethics approval. All the experiments reported in this document have been performed under the 12Article Title approval of the ethics committee of the Universi- tat Polit`ecnica de Catalunya (UPC) in accordance with all the relevant guidelines and regulations (ID: 2023.05). Consent to participate. All the volunteers who participated in this study have signed an informed consent form accepting to participate in the study. Consent for publication. All the volunteers who participated in this study have signed an informed consent form accepting to publish the anonymously obtained data. Appendix A Example of the questions used in the questionnaires The hand-crafted questionnaires used include a first section for demographical data (age, gender, dominant hand,...) followed by the next 7-point Likert (Strongly disagree to Strongly agree) ques- tions grouped here by categories but randomly presented to the volunteer. * means negated ques- tion. We provide the Cronbachâs alpha scores for each category. 1. Robot contribution to fluency (α = 0.786): âą The robot positively contributed to the flu- ency of the interaction. âą The robotâs response speed was appropriate. âą The robotâs movements were aggressive*. 2. Robot contribution to performance (α = 0.843): âą The robot positively contributed to the teamâs performance. âą The robot proposed good solutions to com- plete the task. âą The robot helped in completing the task. 3. Robot contribution as Human (α = 0.887): âą The robot contributed equally to the human in completing the task. âą The human-robot team worked equitably. âą The robot contributed equally to the human in the teamâs performance. 4. Human responsibility (α = 0.807): âą I had to bear the responsibility of the task for the human-robot team to function better. âą I was the most important member of the team. âą The teamâs performance depended largely on me. 5. Trust in Robot (α = 0.780): âą I trusted the robot. âą The robot was trustworthy. âą I trusted that the robot would do the right thing at the right time. 6. Comfort (α = 0.838): âą I felt comfortable accompanying the robot. âą The interaction with the robot was pleasant. âą The robotâs movements made me feel uncom- fortable*. The previous questions were followed by the next 7-point Likert (Impossible to Effortless) ques- tion: âHow easy was it for you to indicate your intention to the robot?â. After all the experiments in the same round, the volunteer is asked to choose which system they consider âsaferâ and âeasier to useâ as well as which system they consider makes the interaction âfaster to executeâ, âmore fluidâ, âmore naturalâ and âmore similar to how two humans would exe- cute itâ. Finally, they are asked which system they consider âmost appropriate for the taskâ. References [1] Goldhoorn, A., Garrell, A., Alqu Ìezar, R., Sanfeliu, A.: Searching and tracking people in urban environments with static and dynamic obstacles. Robotics and Autonomous Sys- tems 98, 147â157 (2017) [2] Saito, N., Ogata, T., Funabashi, S., Mori, H., Sugano, S.: How to select and use tools?: Active perception of target objects using mul- timodal deep learning. IEEE Robotics and Automation Letters 6(2), 2517â2524 (2021) [3] Dragan, A.D.: Robot planning with mathe- matical models of human state and action. arXiv preprint arXiv:1705.04226 (2017) Article Title13 [4] Choudhury, R., Swamy, G., Hadfield-Menell, D., Dragan, A.D.: On the utility of model learning in hri. In: 2019 14th ACM/IEEE International Conference on Human-Robot Interaction (HRI), p. 317â325 (2019). IEEE [5] Ord Ìo Ìnez, F.J., Roggen, D.: Deep Convolu- tional and LSTM Recurrent Neural Networks for Multimodal Wearable Activity Recogni- tion. Sensors 16(1), 115 (2016) [6] Schydlo, P., Rakovic, M., Jamone, L., Santos- Victor, J.: Anticipation in Human-Robot Cooperation: A Recurrent Neural Network Approach for Multiple Action Sequences Pre- diction. In: 2018 IEEE International Con- ference on Robotics and Automation, p. 5909â5914 (2018). IEEE [7] Johnson-Laird, P.N.: Mental Models. The MIT Press, ??? (1989) [8] Gildert, N., Millard, A.G., Pomfret, A., Tim- mis, J.: The need for combining implicit and explicit communication in cooperative robotic systems. Frontiers in Robotics and AI 5, 65 (2018) [9] Lee, B.-J., et al.: Perception-Action-Learning System for Mobile Social-Service Robots using Deep Learning. In: Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32 (2018) [10] Dar, S., Bernardet, U.: When agents become partners: A review of the role the implicit plays in the interaction with artificial social agents. Multimodal Technologies and Inter- action 4(4), 81 (2020) [11] Dom Ìınguez-Vidal, J.E., Sanfeliu, A.: Infer- ence VS. Explicitness. Do We Really Need the Perfect Predictor? The Human-Robot Collaborative Object Transportation Case. In: 32nd IEEE International Conference on Robot and Human Interactive Communi- cation (RO-MAN), p. 1866â1871 (2023). IEEE [12] Luo, R.C., Mai, L.: Human Intention Infer- ence and On-Line Human Hand Motion Pre- diction for Human-Robot Collaboration. In: 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 5958â5964 (2019). IEEE [13] Jain, S., Argall, B.: Recursive bayesian human intent recognition in shared-control robotics. In: 2018 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS), p. 3905â3912 (2018). IEEE [14] Maroger, I., Ramuzat, N., Stasse, O., Watier, B.: Human trajectory prediction model and its coupling with a walking pattern genera- tor of a humanoid robot. IEEE Robotics and Automation Letters 6(4), 6361â6369 (2021) [15] Thobbi, A., Gu, Y., Sheng, W.: Using human motion estimation for human-robot coopera- tive manipulation. In: 2011 IEEE/RSJ Inter- national Conference on Intelligent Robots and Systems, p. 2873â2878 (2011). IEEE [16] Agravante, D.J., Cherubini, A., Bussy, A., Gergondet, P., Kheddar, A.: Collaborative human-humanoid carrying using vision and haptic sensing. In: 2014 IEEE International Conference on Robotics and Automation (ICRA), p. 607â612 (2014). IEEE [17] Bussy, A., Gergondet, P., Kheddar, A., Keith, F., Crosnier, A.: Proactive behavior of a humanoid robot in a haptic transportation task with a human partner. In: 2012 IEEE RO-MAN: The 21st IEEE International Sym- posium on Robot and Human Interactive Communication, p. 962â967 (2012). IEEE [18] Tarbouriech, S., Navarro, B., Fraisse, P., Crosnier, A., Cherubini, A., Sall Ìe, D.: Admit- tance control for collaborative dual-arm manipulation. In: 2019 19th International Conference on Advanced Robotics (ICAR), p. 198â204 (2019). IEEE [19] Yu, X., Li, B., He, W., Feng, Y., Cheng, L.,Silvestre,C.:Adaptive-constrained impedance control for humanârobot co- transportation.IEEEtransactionson cybernetics 52(12), 13237â13249 (2021) [20] Li, Z., Liu, J., Huang, Z., Peng, Y., Pu, H., Ding, L.: Adaptive impedance control 14Article Title of humanârobot cooperation using reinforce- ment learning. IEEE Transactions on Indus- trial Electronics 64(10), 8013â8022 (2017) [21] Alevizos, K.I., Bechlioulis, C.P., Kyriakopou- los, K.J.: Physical humanârobot cooperation based on robust motion intention estimation. Robotica 38(10), 1842â1866 (2020) [22] Al-Yacoub, A., Zhao, Y., Eaton, W., Goh, Y.M., Lohse, N.: Improving human robot collaboration through force/torque based learning for object manipulation. Robotics and Computer-Integrated Manufacturing 69, 102111 (2021) [23] Dom Ìınguez-Vidal, J.E., Sanfeliu, A.: Improv- ing Human-Robot Interaction Effective- ness in Human-Robot Collaborative Object Transportation using Force Prediction. In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 7839â7845 (2023). IEEE [24] Dom Ìınguez-Vidal, J.E., Torres-Rodr Ìıguez, I.J., Garrell, A., Sanfeliu, A.: User-Friendly Smartphone Interface to Share Knowledge in Human-Robot Collaborative Search Tasks. In: 30th IEEE International Conference on Robot and Human Interactive Commu- nication (RO-MAN), p. 913â918 (2021). https://doi.org/10.1109/RO-MAN50785. 2021.9515379 [25] Dalmasso,M.,Dom Ìınguez-Vidal,J.E., Torres-Rodr Ìıguez, I.J., Garrell, A., San- feliu, A.: Shared Task Representation for Human-RobotCollaborativeNavigation: The Collaborative Search Case. International Journal of Social Robotics (2023). https: //doi.org/10.1007/s12369-023-01067-0 [26] Mullen, J.F., Mosier, J., Chakrabarti, S., Chen, A., White, T., Losey, D.P.: Communi- cating inferred goals with passive augmented reality and active haptic feedback. IEEE Robotics and Automation Letters 6(4), 8522â 8529 (2021) [27] Gildert, N.: Combining implicit and explicit communication in object manipulation tasks between two robots. PhD thesis, University of York (2022) [28] Lorentz, V., Weiss, M., Hildebrand, K., Boblan, I.: Pointing Gestures for Human- Robot Interaction with the Humanoid Robot Digit. In: 2023 32nd IEEE International Con- ference on Robot and Human Interactive Communication (RO-MAN), p. 1886â1892 (2023). IEEE [29] Rogalla, O., Ehrenmann, M., Zollner, R., Becher, R., Dillmann, R.: Using gesture and speech control for commanding a robot assis- tant. In: Proceedings. 11th IEEE Interna- tional Workshop on Robot and Human Inter- active Communication, p. 454â459 (2002). IEEE [30] Lv, X., Zhang, M., Li, H.: Robot control based on voice command. In: 2008 IEEE International Conference on Automation and Logistics, p. 2490â2494 (2008). IEEE [31] Warden, P.: Speech commands: A dataset for limited-vocabulary speech recognition. arXiv preprint arXiv:1804.03209 (2018) [32] Majumdar, S., Ginsburg, B.: Matchboxnet: 1d time-channel separable convolutional neural network architecture for speech commandsrecognition.arXivpreprint arXiv:2004.08531 (2020) [33] Kim, B., Chang, S., Lee, J., Sung, D.: Broad- casted residual learning for efficient keyword spotting. arXiv preprint arXiv:2106.04140 (2021) [34] Dom Ìınguez-Vidal, J.E., Rodr Ìıguez, N., Sanfe- liu, A.: Perception-Intention-Action Cycle in Human-Robot Collaborative Tasks: the Col- laborative Lightweight Object Transporta- tion Use-Case. International Journal of Social Robotics, (2024) [35] Dom Ìınguez-Vidal, J.E., Rodr Ìıguez, N., San- feliu, A.: Perception-Intention-Action Cycle as a Human Acceptable Way for Improv- ing Human-Robot Collaborative Tasks. In: Companion of the 2023 ACM/IEEE Interna- tional Conference on Human-Robot Interac- tion, p. 567â571 (2023). https://doi.org/10. Article Title15 1145/3568294.3580149 [36] Dom Ìınguez-Vidal, J.E., Sanfeliu, A.: Explor- ing Transformers and Visual Transformers for Force Prediction in Human-Robot Collabo- rative Transportation Tasks. In: 2024 IEEE International Conference on Robotics and Automation (ICRA), p. (2024). IEEE [37] Laplaza, J., Rodr Ìıguez, N., Dom Ìınguez-Vidal, J.E., Herrero, F., Hern Ìandez, S., L Ìopez, A., Sanfeliu, A., Garrell, A.: IVO Robot: A New Social Robot for Human-Robot Collabora- tion. In: Proceedings of the 2022 ACM/IEEE International Conference on Human-Robot Interaction, p. 860â864 (2022). IEEE