Paper deep dive
Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
Jens Grubert, John Dudley, Eyal Ofek, Per Ola Kristensson
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/1/2026, 11:02:57 AM
Summary
The paper proposes Extended Reality (XR) as a mediation layer for situated human control in human-robot teaming. It argues that XR should go beyond displaying robot intent to enabling humans to understand, shape, authorize, and interrupt robot actions within their physical and social context. The authors identify four mediation functions (connecting human intent to robot autonomy, robot plans to human judgment, levels of shared control, and team roles/handover) and derive six design dimensions (joint action possibilities, socio-physical constraints, uncertainty/plan validity, multimodal control/correction, roles/handover/accountability, and anticipatory recovery). These concepts are grounded in three scenarios: robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention.
Entities (17)
Relation Signals (17)
Jens Grubert → authored → Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
confidence 99% · Jens Grubert * Coburg University of Applied Sciences and Arts
Per Ola Kristensson → authored → Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
confidence 99% · Per Ola Kristensson § University of Cambridge
Eyal Ofek → authored → Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
confidence 99% · Eyal Ofek ‡ University of Birmingham
John Dudley → authored → Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
confidence 99% · John Dudley † University of Cambridge
Extended Reality → supports → Situated Human Control
confidence 95% · We argue that XR should also be understood as a mediation layer for situated human control in human-robot teaming.
Design Dimensions → derivedfrom → Mediation Functions
confidence 93% · Building on these functions, we derive six design dimensions...
Extended Reality → mediates → Team Roles
confidence 92% · We identify four mediation functions connecting ... team roles, handover, and recovery.
Extended Reality → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned motion, reachability, and state. We argue that XR should also be understood as a mediation layer for situated human control in human-robot teaming. Situated human control denotes the human collaborator's ability to understand, shape, authorize, and interrupt robot action within the concrete physical, social, and temporal context in which that action unfolds. We ground this perspective in scenarios from robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention. Across these scenarios, robot autonomy must remain inspectable and adjustable as people move, goals change, sensing is incomplete, control roles shift, and plans become invalid. We identify four mediation functions connecting human intent and robot autonomy, robot plans and human judgment, levels of shared control, and team roles, handover, and recovery. Building on these functions, we derive six design dimensions: joint action possibilities, socio-physical constraints, uncertainty and plan validity, multimodal control and correction, roles, handover, and accountability, and anticipatory recovery. The paper outlines a research agenda for XR systems that make robot autonomy more actionable and accountable in dynamic shared environments.
Tags
Links
- Source: https://arxiv.org/abs/2607.25047v1
- Canonical: https://arxiv.org/abs/2607.25047v1
Trouble viewing inline? Open PDF directly →
Full Text
39,769 characters extracted from source content.
Expand or collapse full text
Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming Jens Grubert * Coburg University of Applied Sciences and Arts John Dudley † University of Cambridge Eyal Ofek ‡ University of Birmingham Per Ola Kristensson § University of Cambridge ABSTRACT Extended Reality (XR) is increasingly used in human-robot interac- tion to communicate robot intent, planned motion, reachability, and state. We argue that XR should also be understood as a mediation layer for situated human control in human-robot teaming. Situated human control denotes the human collaborator’s ability to under- stand, shape, authorize, and interrupt robot action within the con- crete physical, social, and temporal context in which that action un- folds. We ground this perspective in scenarios from robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention. Across these scenarios, robot au- tonomy must remain inspectable and adjustable as people move, goals change, sensing is incomplete, control roles shift, and plans become invalid. We identify four mediation functions connecting human intent and robot autonomy, robot plans and human judg- ment, levels of shared control, and team roles, handover, and recov- ery. Building on these functions, we derive six design dimensions: joint action possibilities, socio-physical constraints, uncertainty and plan validity, multimodal control and correction, roles, handover, and accountability, and anticipatory recovery. The paper outlines a research agenda for XR systems that make robot autonomy more actionable and accountable in dynamic shared environments. Index Terms:extended reality, robotics, human-robot teaming, situated human control, human-robot interaction, embodied artifi- cial intelligence, physical artificial intelligence. 1 INTRODUCTION Human-robot interaction (HRI) is increasingly moving toward teams in which people collaborate with autonomous, semi- autonomous, or teleoperated robots in shared physical environ- ments. Coordination in such settings requires reciprocal anticipa- tion: people need to understand robot behavior, while robots must account for human actions, attention, and task state. This becomes especially important in dynamic environments where people move, goals change, sensing is incomplete, and previously valid robot plans may become inappropriate during execution. Extended Reality (XR) is well suited for this challenge be- cause it can present robot state, intent, constraints, and interaction possibilities directly in the workspace where robot actions unfold [16, 17, 10]. Prior work has shown that XR can support HRI by visualizing planned motion, reachability, robot state, task progress, and authorable robot behaviors [4, 5, 7]. These approaches demon- strate the value of making robot behavior spatially legible. How- ever, emerging HRI scenarios require interfaces that go beyond * e-mail: jens.grubert@hs-coburg.de † e-mail: jjd50@cam.ac.uk ‡ e-mail: e.ofek@bham.ac.uk § e-mail: pok21@cam.ac.uk awareness of robot intent. Human collaborators may need to as- sess whether a proposed action remains appropriate, constrain or correct a plan, approve execution, or interrupt action when the situ- ation changes. XR devices can also contribute information about the human col- laborator. Head pose, gaze, hand motion, and explicit gestures can provide spatial and deictic cues that help robots interpret references, estimate attention, or detect requests for intervention. Such signals should not be treated as direct or reliable measurements of human intent; instead, they provide additional, uncertain evidence that can be combined with task context and explicit input [12, 7]. This bidi- rectional role distinguishes XR mediation from interfaces that only display robot-generated information. We argue that XR should be conceptualized as a mediation layer for situated human control in human-robot teaming. We use sit- uated human control to denote the human collaborator’s ability to understand, shape, authorize, and interrupt robot action within the concrete physical, social, and temporal context in which that action unfolds. This notion of control is broader than direct teleoperation: it includes supervisory control, constraint specification, plan ap- proval, role handover, and recovery from emerging misalignments. XR can support such control by making robot autonomy visible in context and by turning human judgment into actionable corrections, constraints, approvals, or interruptions. This framing is motivated by two recurring HRI challenges. First, robot plans are often technically feasible while remaining questionable in context. A trajectory may avoid collisions but still block a worker’s access, pass too close to a vulnerable patient, vi- olate social expectations, or depend on uncertain sensing. Second, human input is increasingly high-level and multimodal. Speech, gaze, and gesture allow efficient task specification, but they also in- troduce ambiguity about referents, intent, and control scope [12, 7]. In multi-person settings, such as care teams or supervised training, these problems are further complicated by role changes and han- dover of control. This paper makes three contributions. First, we define XR medi- ation as a bidirectional, situated coupling between human or team state, robot autonomy, and the shared environment. Second, we use three contrasting HRI scenarios to identify recurring coordina- tion breakdowns involving spatial consequences, ambiguous input, changing plan validity, divided control, and recovery. Third, we derive six scenario-grounded design dimensions and associated re- search and evaluation considerations for XR interfaces that support situated human control. 2 MOTIVATING SCENARIOS We ground our argument in three scenarios that differ in domain, robot configuration, and interaction style, but share a common de- sign problem: robot autonomy must remain understandable and ac- tionable as the situation changes, see Figure 1. In each scenario, XR can support situated human control by making planned robot action visible in context and by enabling humans to correct, con- strain, approve, interrupt, or recover that action. arXiv:2607.25047v1 [cs.HC] 27 Jul 2026 Figure 1: Illustrations of the three motivating scenarios. Left: robot-assisted bedside nursing. Center: multi-arm supervisory control. Right: collaborative assembly. 2.1 Robot-Assisted Bedside Nursing In robot-assisted bedside nursing, a collaborative robot may sup- port physically demanding care tasks such as lifting, stabilizing, or repositioning a patient’s limb during wound care, hygiene, or preparation for a dressing change [3]. Such tasks can impose high musculoskeletal strain on nurses and may currently require a second caregiver. Reviews of nursing robotics highlight both the potential of assistive robots and the need for stronger nurse-centered design and evaluation in realistic care contexts [14, 9, 1]. This scenario illustrates why robot intent must be interpreted rel- ative to the care situation. A planned robot motion may be collision- free while still being unacceptable because it passes close to the patient’s face, obstructs the nurse’s access to the treatment site, ap- proaches an injured region, or causes discomfort. XR can medi- ate this interaction by previewing the planned motion relative to the patient’s body, showing expected contact or support regions, in- dicating uncertainty in patient tracking, and allowing the nurse to approve, constrain, delay, or interrupt execution. Situated human control is grounded here in caregiver authority: the nurse remains responsible for the care activity, while the robot contributes physi- cal support. The scenario also foregrounds multi-person role and control management.Bedside care may involve nurses, physicians, trainees, patients, and relatives. The robot must distinguish au- thorized commands from incidental conversation, teaching expla- nations, or patient speech. XR can support this by making the cur- rently authorized operator, the scope of control, pending approvals, and available override mechanisms visible to the care team. 2.2 Multi-Arm Supervisory Control A second scenario concerns a teleoperator coordinating multiple robot arms in a remote or hazardous workspace. Directly control- ling each arm can be inefficient because it forces the operator to serialize attention and low-level control. A more scalable strategy is supervisory control, where the operator issues high-level instruc- tions and intervenes when ambiguity, risk, or task complexity re- quires closer involvement [15, 11, 12]. XR is well suited to this setting because it can present the re- mote workspace spatially while supporting speech, gaze, gesture, and direct manipulation. For example, an operator might look at an object, point to a target location, and say “move this to the left bin,” while the system infers the object, target, and responsible robot arm. Such input is efficient, but it creates ambiguity: which ob- ject was selected, which arm should act, which path will be used, and whether another arm or human access space will be affected. XR can externalize these interpretations before execution. This scenario highlights transitions along an intent-to-precision continuum. The operator may begin with high-level intent, refine the plan through deictic constraints, directly adjust a virtual path, or temporarily take fine-grained control. Situated human control depends on the interface making these transitions legible: the user should understand when they are supervising autonomous execu- tion, when they are specifying constraints, and when they are di- rectly controlling motion. 2.3 Collaborative Assembly Under Divided Attention A third scenario concerns collaborative assembly in which a hu- man and robot work on interdependent but partially parallel sub- tasks. For example, a robot may fetch a tool, hold a component, or prepare the next object while the human completes a manual step. Unlike bedside nursing, where the caregiver remains focused on patient care and robot action near the body, this scenario empha- sizes divided attention: the human may be occupied with their own subtask and may only intermittently monitor the robot. This creates a different challenge for situated human control. Small mismatches between robot action and task progress can cas- cade into larger failures. The robot may bring the wrong part, ap- proach before the human is ready, block access to the next work area, or continue executing a plan that no longer matches the hu- man’s current sequence. XR can support anticipatory recovery by visualizing future robot actions, timing conflicts, and likely inter- ference points before they become task failures. In this scenario, XR should support low-effort intervention un- der cognitive load. Planned robot actions can be represented as manipulable spatial objects: a target can be reassigned, a timing conflict can be paused, a path can be shifted away from the hu- man’s workspace, or an action can be redirected to a safer alterna- tive. This connects robot intent visualization with direct manipula- tion and shared control [10, 4, 2]. The key question is whether XR allows users to notice and correct emerging misalignments while maintaining continuity in their own task. 3 XR AS A MEDIATION LAYER The scenarios above illustrate a common problem: robot actions may be technically feasible while remaining inappropriate in the concrete situation in which they will unfold. Human collaborators therefore need more than awareness of robot intent. They must be able to relate a robot’s interpretation and planned behavior to the current task, people, workspace, uncertainty, and distribution of control, and to intervene when these relations become misaligned. Intent visualization remains an important foundation for XR- based HRI. Planned paths, goals, reachability, state, and task progress can help users anticipate robot behavior and coordinate with it [10, 16, 17, 4, 5]. Situated human control extends this role from awareness to actionability. An XR interface should not only communicate what the robot plans to do, but also help users as- sess situated consequences, understand the current control relation, compare alternatives, and correct, approve, interrupt, hand over, or recover robot action. We conceptualize the XR mediation layer as a bidirectional coupling between human or team state, robot autonomy, and the shared environment. Robot-generated information flows toward hu- man collaborators through spatial representations of interpretations, plans, uncertainty, control modes, and anticipated consequences. Human input flows toward the robot through speech, gaze, gesture, direct manipulation, constraints, role assignment, approval, and in- terruption. The shared environment grounds both directions by re- lating input and robot behavior to concrete people, objects, regions, tasks, and events. Mediation therefore comprises contextualization and disambiguation as well as visualization, intervention, and re- covery. XR is not uniquely required for every mediation function. Con- ventional displays can communicate robot state, uncertainty, or control modes. XR becomes particularly advantageous when me- diation depends on spatial registration, embodied and deictic input, mobility, divided attention, or a shared view of physically situated action. It can anchor plans and constraints to robots, people, ob- jects, and regions in the workspace; capture gaze, head, and hand activity as spatial input; and support local or remote users without requiring repeated attention shifts to a separate display. The rel- evant design question is therefore not whether XR can display a particular item of information, but whether spatial grounding and embodied interaction improve the human’s ability to understand or influence robot action [16, 17, 12]. The following subsections distinguish four mediation functions: mediating between human intent and robot autonomy, between robot plans and human judgment, between levels of shared con- trol, and between team roles, handover, and recovery. Table 1 op- erationalizes these functions through their information flow, XR- specific leverage, possible evaluation approach, and relation to the design dimensions developed in Section 4. The mediation functions and design dimensions serve different purposes. The mediation functions describe where XR connects hu- man and robot activity. The design dimensions describe the recur- ring interface and evaluation problems that arise when implement- ing these connections. The four functions are also interdependent: human input must be related to robot interpretation, plans must be judged in context, control relations must remain legible, and team members must be able to revise robot behavior as roles or circum- stances change. 3.1 Mediating Between Human Intent and Robot Auton- omy Robots in collaborative environments are increasingly expected to respond to high-level human instructions such as “move this aside,” “support this part,” “prepare the workspace,” or “bring the next tool.” Such instructions allow users to communicate at the level of task goals rather than low-level robot motion, but they are fre- quently underspecified. The robot must infer referents, select ac- tions, account for constraints, and determine the appropriate scope of autonomy. XR can mediate this relation by grounding human input in the shared environment and externalizing the robot’s interpretation be- fore execution. Speech can be associated with gaze, pointing, or gesture to identify objects, regions, and directions. Conversely, the interface can show which referent, goal, target pose, trajectory, or constraint the robot has inferred. The user can then confirm, refine, or correct that interpretation in situ. This is particularly important for multimodal interfaces, where individual signals may remain am- biguous even when their combination appears plausible [12, 7]. The mediation function is therefore bidirectional. XR supplies spatial context for interpreting human input while also exposing the assumptions made by robot autonomy. This allows ambiguity to be resolved before an incorrect interpretation becomes physical robot motion. 3.2 Mediating Between Robot Plans and Human Judg- ment Robot plans are typically generated with respect to technical objec- tives such as reachability, collision avoidance, efficiency, or control feasibility. Human collaborators evaluate the same plans through additional criteria, including access, comfort, workload, social ap- propriateness, task timing, and the expected actions of other team members. These criteria are difficult to encode completely because their importance depends on the current situation. XR can mediate between robot planning and human judgment by placing candidate actions and their anticipated consequences in the context in which they must be evaluated. Relevant representations may include trajectories and target states, but also contact regions, blocked workspaces, timing conflicts, proximity to sensitive areas, interference with another action, or changes in human posture. In bedside nursing, for example, a collision-free movement of a support arm must still be assessed relative to the patient’s body, the nurse’s access to the treatment site, and the patient’s comfort. In collaborative assembly, a technically feasible handover or arm motion may conflict with the worker’s current subtask, reachable workspace, or next intended action. XR supports situated judgment by making these relations spatially inspectable rather than present- ing the plan independently of its consequences. 3.3 Mediating Between Levels of Shared Control Situated human control involves transitions between different forms of human involvement. A user may initially provide a high-level goal, then add a constraint, select between alternatives, approve autonomous execution, directly manipulate a planned path, or in- terrupt the robot. Shared-control research has long recognized that autonomy and human input can be blended or shifted across a task [2, 11]. The XR-specific challenge is to make the current control relation and available transitions legible in the workspace. An XR mediation layer should therefore communicate not only what the robot will do, but also how the current action is being controlled. The interface may need to distinguish autonomous ex- ecution, waiting for approval, following a user-defined constraint, direct trajectory modification, or safety-limited operation under un- certainty. It should also show how a new human input will affect the plan. Making these relations visible can reduce mode confusion and support deliberate transitions between supervisory and direct con- trol. XR can additionally expose plans and constraints as spatially manipulable objects, allowing the user to move from verbal task specification to deictic correction or direct manipulation without losing the relation to the physical workspace. 3.4 Mediating Team Roles, Handover, and Recovery Many HRI settings involve several humans with different exper- tise, responsibilities, and control rights. A care robot may operate around nurses, physicians, trainees, patients, and relatives; a teleop- erator may coordinate with another supervisor; and a collaborative robot may work among several workers. In these settings, not every utterance, gesture, or gaze cue should be interpreted as a command. The system must account for who is currently allowed to command, constrain, approve, or interrupt robot action, and for which robot or subtask that control applies. XR can mediate team roles by making the active operator, con- trol scope, pending approvals, and available override mechanisms visible. It can also support explicit handover by associating control with particular users, robots, subtasks, or spatial regions. For ex- ample, an experienced worker may temporarily transfer control of one task step to a trainee while retaining approval or interruption rights. A supervisor may delegate one manipulator to autonomous execution while directly controlling another. Mediation must continue during execution because an accepted plan may become inappropriate when people move, uncertainty in- creases, the environment changes, or responsibility shifts. XR can support recovery by communicating plan validity, emerging con- flicts, and available revision options. A user may pause execution, update a referent, add or remove a constraint, redirect a trajectory, transfer control, or request an alternative plan. Interruption and recovery thus become normal elements of teamwork rather than mechanisms reserved only for emergency stops. Together, these four mediation functions describe how XR can maintain alignment between human input, robot interpretation, planned action, control relations, and changing team situations. They provide the conceptual basis for the design dimensions and research questions considered next. 4 DESIGN DIMENSIONS AND RESEARCH AGENDA The mediation perspective leads to a set of design dimensions for XR interfaces that support situated human control. These dimen- sions describe what must become visible and actionable when hu- man collaborators are expected to judge, shape, authorize, and re- vise robot autonomy in dynamic environments. They also point to research questions such as: how to represent the current team state, how to support intervention without overloading users, and how to evaluate whether XR improves human influence over robot action? The dimensions were derived by comparing the recurring coor- dination problems across the three scenarios and relating them to the mediation functions discussed above. Joint action possibilities address the relation between human and robot capabilities; socio- physical constraints address the relation between technically feasi- ble and contextually appropriate action; uncertainty and plan valid- ity address changing or incomplete state information; multimodal control addresses the translation of human input into robot behav- ior; roles and handover address multi-person control; and anticipa- tory recovery addresses emerging misalignment during execution. 4.1 Joint Action Possibilities XR interfaces should represent what the human-robot team can achieve together under current constraints. Existing visualizations often focus on the robot’s individual capabilities, such as its reach- able workspace, planned trajectory, or navigation goal [4, 5]. For situated human control, the more relevant question is relational: which actions become possible through the combination of human abilities, robot capabilities, shared objects, timing, and workspace layout? Joint action possibilities may include reachable handover re- gions, feasible lifting or stabilization configurations, safe cooper- ation zones, human access spaces, or coordinated motions between multiple robot arms. In bedside nursing, this could mean visual- izing whether a robot can support a patient’s leg while preserving the nurse’s access to the treatment area. In multi-arm teleoperation, it could mean showing which objects can be handled in parallel, which arm should approach from which side, and where simultane- ous motion would create interference. A central research question is how XR can represent joint action possibilities without turning the workspace into a dense planning display. Future work should investigate compact representations that communicate what the team can do now, what alternatives ex- ist, and which actions require human decision or approval. 4.2 Socio-Physical Constraints Robot action in shared environments is shaped by constraints that combine physical, social, ergonomic, and organizational factors. Collision avoidance is necessary, but many relevant risks emerge before collision occurs. A robot may block access, approach a vulnerable body region, move through a space that feels uncom- fortable, demand excessive attention, or violate expectations about appropriate conduct. Research on socially aware navigation high- lights the importance of interpersonal distance, comfort, legibility, and contextual appropriateness in robot motion [8, 13]. XR can make such socio-physical constraints visible as first- class elements of interaction. Instead of showing a trajectory alone, the interface can relate the trajectory to personal spaces, protected regions, task access zones, preferred approach directions, or areas where human attention is already occupied. In care contexts, these constraints may involve patient dignity, discomfort, treatment ac- cess, and caregiver workload. In industrial contexts, they may in- volve ergonomic posture, shared tool access, and temporal coordi- nation with human subtasks. The research challenge is to encode such constraints in ways that support rapid judgment without giving soft, context-dependent as- sessments an unwarranted sense of precision. We need methods for visualizing qualitative appropriateness, borderline actions, and trade-offs between efficiency, safety, comfort, and task continuity. 4.3 Uncertainty, Plan Validity, and Situational Aware- ness Situated human control depends on the user’s ability to understand whether a robot plan remains valid. Plans can become outdated when people move, objects are occluded, sensor readings conflict, the task goal changes, or the robot’s interpretation of a command is uncertain. XR interfaces should therefore communicate both planned action and the reliability of the assumptions behind it. Uncertainty cues should be actionable. A user benefits less from abstract confidence values than from knowing where attention, con- firmation, or correction is needed. For example, an XR interface might highlight an uncertain object reference, visually degrade path segments that depend on incomplete sensing, mark regions where tracking quality is low, or indicate that a planned motion requires renewed approval after the environment changes. Prior work on uncertainty visualization emphasizes the tension between informa- tiveness and cognitive load [6]; in XR for HRI, this tension is am- plified because uncertainty cues appear in the same space where users coordinate physical action. This dimension also connects directly to situational awareness. Users must perceive relevant changes, understand their implications for the human-robot task, and anticipate what may happen next. Future work should investigate how XR can support these levels of awareness while keeping robot plans actionable. Important ques- tions include when the system should interrupt the user, when it should quietly update the plan, and when it should request renewed confirmation. 4.4 Multimodal Control and Correction Situated human control requires interaction techniques that match the constraints of collaborative work. Speech is efficient for high- level goals and constraints, such as “move this aside” or “approach from the left.” Gaze and gesture help establish spatial reference, disambiguate objects, and indicate regions. Direct manipulation allows users to reshape virtual trajectories, adjust target poses, or move constraints in the workspace. Haptics or physical interaction may support interruption, guidance, or close-proximity correction. The design goal is to match modality to control act. Relevant acts include specifying intent through speech, resolving references through gaze or pointing, defining exclusion zones through ges- ture, comparing alternative plans through spatial previews, adjust- ing paths through direct manipulation, approving execution explic- itly, interrupting action through voice or gesture, and handing over authority through visible role assignment. Research is needed on interaction vocabularies for XR-mediated control. Such vocabularies should allow users to express task- relevant constraints without becoming responsible for low-level robot programming. At the same time, multimodal systems must Table 1: Operationalization of the four XR mediation functions. The table specifies information flow, XR-specific leverage, an example evaluation approach, and related design dimensions. R: robot, H: human, Team: multiple human collaborators. Mediation functionFlowXR-specific leverageOperational testRelated dimensions Human intent and robot autonomy H→R→HCo-register multimodal input with physi- cal referents and externalize the robot’s in- ferred goal and constraints in situ. Introduce ambiguous deictic commands; measure grounding accuracy, correction success, and correction latency. S4.3; S4.4 Robot plans and human judgment R→H→RAnchor planned actions and consequences to people, objects, access regions, and task context. Present feasible but contextually inappro- priate plans; measure detection rate, deci- sion accuracy, and response time. S4.1; S4.2; S4.3 Levels of shared controlH↔RMake the current control mode and ma- nipulable plan elements visible in the workspace. Multimodal input for coarse- to-fine robot control. Require transitions between instruction, constraint, approval, correction, and inter- ruption; measure mode comprehension and intervention performance. S4.4; S4.6 Team roles, handover, and recovery Team↔RAssociate control scope with users, robots, subtasks, and locations; make handover and recovery options jointly visible. Introduce operator changes or plan invali- dation; measure command attribution, han- dover time, and successful recovery. S4.5; S4.6 provide immediate feedback about interpretation: which referent was selected, which constraint was added, how the plan changed, and which control mode is currently active [12, 7]. 4.5 Roles, Handover, and Accountability Many HRI scenarios involve multiple humans with different roles, expertise, and responsibilities. Yet current interaction designs often assume a single active user. Situated human control requires mech- anisms for determining who can issue commands, who can approve execution, who can interrupt, and how control can be handed over during a task. XR interfaces should represent authority as a dynamic part of the team state. This includes the currently authorized operator, the scope and duration of control, pending approvals, and override rights. A lightweight authority model might combine explicit han- dover, spatial context, speaker or user recognition, gaze direction, and role assignment. The goal is to avoid both overly rigid au- thentication procedures and unsafe ambiguity about whose input the robot should follow. This dimension also supports accountability. Approval should communicate what is being approved: a trajectory, a contact point, a target state, an autonomy level, or a bounded time interval. This matters because authority over robot action can shift dynamically during teamwork. XR can help prevent responsibility from becom- ing diffuse by keeping the current control relation visible. 4.6 Anticipatory Recovery Finally, situated human control should support correction before breakdowns become failures. Collaborative tasks often fail through gradual misalignment: a wrong object reference, delayed handover, unsafe approach direction, outdated plan, or robot action that con- flicts with a human’s next step. XR can support anticipatory recov- ery by making emerging conflicts visible early and by providing low-latency ways to revise robot action. This dimension extends intent visualization from prediction to intervention. The interface can show future conflict states, likely timing mismatches, regions where human and robot action will in- terfere, or points where uncertainty may invalidate execution. Users can then pause, redirect, constrain, or reassign the robot while maintaining task continuity. In shared-control settings, such recov- ery may involve moving along the control continuum: from ap- proving an autonomous plan, to specifying a constraint, to directly adjusting a virtual path, to physically interrupting the robot. Evaluation should therefore go beyond task completion and sub- jective usability. Relevant measures include plan comprehension, detection of problematic but technically feasible actions, correc- tion accuracy, intervention latency, recovery time, trust calibra- tion, workload, authority clarity, and task continuity. Scenario- based studies can deliberately introduce ambiguous commands, so- cially inappropriate paths, stale plans, or hidden uncertainty to test whether users notice and respond appropriately. Longitudinal stud- ies are also needed, since users may initially attend to uncertainty and authority cues but later ignore them or over-trust the system. Taken together, these design dimensions suggest that situated hu- man control is an important design and evaluation goal. The central question is how to make robot action inspectable, adjustable, and accountable under the conditions of real teamwork. 5 CONCLUSION XR offers a distinctive interface layer for HRI because robot action is spatial, embodied, and consequential in the physical world. In this position paper, we argued that XR should be understood as a mediation layer for situated human control in human-robot team- ing. This perspective extends the design focus beyond displaying robot intent toward supporting the situated processes through which human collaborators understand, shape, authorize, interrupt, and re- cover robot actions. We grounded this argument in three scenarios: robot-assisted bedside nursing, multi-arm supervisory control, and collaborative assembly under divided attention. Across these scenarios, robot autonomy must remain understandable and adjustable as the task situation changes. People move, goals shift, sensing remains in- complete, control roles and responsibilities may change, and plans that were appropriate during preview may become unsuitable dur- ing execution. We distinguished four mediation functions through which XR can maintain alignment between human input, robot interpretation, planned action, levels of shared control, and changing team situa- tions. These functions describe how XR can ground multimodal in- put, externalize robot interpretations, contextualize planned actions, make control relations legible, and support handover and recovery. Building on these functions, we identified six design dimensions for situated human control: joint action possibilities, socio-physical constraints, uncertainty and plan validity, multimodal control and correction, roles, handover, and accountability, and anticipatory re- covery. Future systems should be evaluated by whether they can meaningfully influence what a joint human-robot team does next. Treating XR as a mediation layer can help make robot autonomy more inspectable, adjustable, and accountable in dynamic shared environments. 6 ACKNOWLEDGEMENTS Generative AI (ChatGPT-5.5, OpenAI) was used to assist with lan- guage editing, to improve the flow and readability of the manuscript and for image generation. The authors reviewed and assume full re- sponsibility for the content of this article. This research is funded through the project MAPPLE: Multimodal assistive robot platform for nursing tasks to support loads and improve ergonomics. Fed- eral Ministry of Research, Technology and Space, Germany. 2025- 2028. REFERENCES [1] G. T. Babalola, J.-M. Gaston, J. Trombetta, and S. Tulk Jesso. A sys- tematic review of collaborative robots for nurses: Where are we now, and where is the evidence? Frontiers in Robotics and AI, 11:1398140, 2024. doi: 10.3389/frobt.2024.1398140 2 [2] A. D. Dragan and S. S. Srinivasa.A policy-blending formalism for shared control. The International Journal of Robotics Research, 32(7):790–805, 2013. doi: 10.1177/0278364913490324 2, 3 [3] J. Grubert, V. M ̈ uller, N. Elkmann, and P. Jahn. Extended reality and artificial intelligence for multimodal robot assistance in bedside nurs- ing. In Proceedings of the 2026 IEEE International Symposium on Mixed and Augmented Reality Adjunct. IEEE, 2026. 2 [4] U. Gruenefeld, L. Pr ̈ adel, J. Illing, T. Stratmann, S. Drolshagen, and M. Pfingsthorn. Mind the arm: Realtime visualization of robot motion intent in head-mounted augmented reality. In Proceedings of Mensch und Computer 2020, p. 259–266. ACM, 2020. 1, 2, 4 [5] S. Hauck, D. Abdlkarim, J. Dudley, P. O. Kristensson, E. Ofek, and J. Grubert. Reachvox: Clutter-free reachability visualization for robot motion planning in virtual reality. In 2025 IEEE International Sym- posium on Mixed and Augmented Reality (ISMAR), p. 1471–1478. IEEE, 2025. doi: 10.1109/ISMAR67309.2025.00151 1, 2, 4 [6] A. Kamal, P. Dhakal, A. Y. Javaid, V. K. Devabhaktuni, D. Kaur, J. Zaientz, and R. Marinier. Recent advances and challenges in uncer- tainty visualization: A survey. Journal of Visualization, 24(5):861– 890, 2021. doi: 10.1007/s12650-021-00755-1 4 [7] R. Lunding, S. Hubenschmid, T. Feuchtner, and K. Grønbæk. Arthur: Authoring human-robot collaboration processes with augmented real- ity using hybrid user interfaces. Virtual Reality, 29(2):73, 2025. doi: 10.1007/s10055-025-01149-6 1, 3, 5 [8] C. Mavrogiannis, F. Baldini, A. Wang, D. Zhao, P. Trautman, A. Stein- feld, and J. Oh. Core challenges of social robot navigation: A survey. ACM Transactions on Human-Robot Interaction, 12(3):1–39, 2023. doi: 10.1145/3583741 4 [9] C. Ohneberg, N. St ̈ obich, A. Warmbein, I. Rathgeber, A. C. Mehler- Klamt, U. Fischer, and I. Eberl. Assistive robotic systems in nursing care: A scoping review. BMC Nursing, 22(1):72, 2023. doi: 10.1186/ s12912-023-01230-y 2 [10] M. Pascher, U. Gr ̈ unefeld, S. Schneegass, and J. Gerken. How to communicate robot motion intent: A scoping review. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, p. 1–17. ACM, 2023. doi: 10.1145/3544548.3580857 1, 2 [11] T. B. Sheridan. Telerobotics, Automation, and Human Supervisory Control. MIT Press, Cambridge, MA, USA, 1992. 2, 3 [12] W. Si, T. Zhong, N. Wang, and C. Yang. A multimodal teleoper- ation interface for human-robot collaboration. In Proceedings of the 2023 IEEE International Conference on Mechatronics, p. 1–6. IEEE, 2023. doi: 10.1109/ICM54990.2023.10102060 1, 2, 3, 5 [13] P. T. Singamaneni, P. Bachiller-Burgos, L. J. Manso, A. Garrell, A. Sanfeliu, A. Spalanzani, and R. Alami. A survey on socially aware robot navigation: Taxonomy and future challenges. The International Journal of Robotics Research, 43(10):1533–1572, 2024. doi: 10.1177/ 02783649241230562 4 [14] G. P. Soriano, Y. Yasuhara, H. Ito, K. Matsumoto, K. Osaka, Y. Kai, R. Locsin, S. Schoenhofer, and T. Tanioka.Robots and robotics in nursing. Healthcare, 10(8):1571, 2022. doi: 10.3390/ healthcare10081571 2 [15] P. Stotko, S. Krumpen, M. Schwarz, C. Lenz, S. Behnke, R. Klein, and M. Weinmann. A vr system for immersive teleoperation and live exploration with a mobile robot. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems, p. 3630–3637. IEEE, 2019. doi: 10.1109/IROS40897.2019. 8968598 2 [16] R. Suzuki, A. Karim, T. Xia, H. Hedayati, and N. Marquardt. Aug- mented reality and robotics: A survey and taxonomy for ar-enhanced human-robot interaction and robotic interfaces. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, p. 1–33. ACM, 2022. doi: 10.1145/3491102.3517719 1, 2, 3 [17] M. Walker, T. Phung, T. Chakraborti, T. Williams, and D. Szafir. Vir- tual, augmented, and mixed reality for human-robot interaction: A survey and virtual design element taxonomy. ACM Transactions on Human-Robot Interaction, 2023. doi: 10.1145/3597623 1, 2, 3