Paper deep dive
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds
Shiting Gong, Jianpeng Yao, Jinfeng Wang, Marco Pavone, Jiachen Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/13/2026, 4:24:24 AM
Summary
The paper proposes a constrained reinforcement learning (CRL) framework for human-following robots in crowded environments. It decomposes the task into a sparse task reward and independent cost constraints for following distance, human safety, and obstacle safety. The method uses PPO-Lagrangian optimization and integrates prediction uncertainty of human motions to enhance safety in unpredictable scenarios. Experiments in simulation (CrowdNav with static obstacles) and real-robot deployment (ROS 2) demonstrate improved proximity-safety balance compared to baselines.
Entities (8)
Relation Signals (6)
Policy → deployedon → ROS 2
confidence 90% · We further deploy the policy on a real robot under ROS 2
CrowdNav → extendedwith → Static Obstacles
confidence 90% · we extend the CrowdNav simulator with static obstacles for evaluation
CNN → extractsfeaturesfrom → Occupancy Grid Maps
confidence 90% · 3D convolutional neural network (CNN) that extracts features from local occupancy grid maps
Transformer → models → Social Interaction
confidence 90% · Our policy network combines an attention-based Transformer for social interaction modeling
PPO-Lagrangian → optimizes → Human Following
confidence 90% · The entire system is jointly optimized under a constrained reinforcement learning (CRL) framework using PPO-Lagrangian
Adaptive Conformal Inference → computes → Prediction Uncertainty
confidence 85% · The uncertainty is computed by an adaptive conformal inference (ACI) module
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedestrians and obstacles. This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unfamiliar or unpredictable. Existing reinforcement learning (RL) methods typically encode these competing objectives into a single dense reward, but the resulting proximity-safety balance is implicit and difficult to adjust across conditions. To address this, we decompose the human-following task into a sparse task reward and independent cost constraints within a multi-constraint RL formulation, where each constraint is managed through cost thresholds with direct behavioral meaning rather than implicit reward weight ratios, allowing explicit and tunable control over the trade-off. We further quantify the prediction uncertainty of human motions and integrate these estimates into the RL costs to enhance safety under unpredictable conditions. Extensive experiments across both in-distribution and out-of-distribution settings demonstrate that our method achieves an effective proximity-safety balance compared to baselines. Real-robot deployment further validates the feasibility of our method in real-world scenarios. More details are available on our project page: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.10056v1
- Canonical: https://arxiv.org/abs/2608.10056v1
Trouble viewing inline? Open PDF directly →
Full Text
53,708 characters extracted from source content.
Expand or collapse full text
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds Shiting Gong 1∗ , Jianpeng Yao 2∗ , Jinfeng Wang 2 , Marco Pavone 3,4 and Jiachen Li 5† Abstract— Following a target human in crowded environ- ments involves an inherent conflict between staying close to the target and navigating safely among surrounding pedes- trians and obstacles. This conflict becomes more severe in dense scenarios, where aggressive following risks collisions and conservative margins lead to target loss, especially when pedestrian behaviors are unfamiliar or unpredictable. Existing reinforcement learning (RL) methods typically encode these competing objectives into a single dense reward, but the result- ing proximity-safety balance is implicit and difficult to adjust across conditions. To address this, we decompose the human- following task into a sparse task reward and independent cost constraints within a multi-constraint RL formulation, where each constraint is managed through cost thresholds with direct behavioral meaning rather than implicit reward weight ratios, allowing explicit and tunable control over the trade-off. We further quantify the prediction uncertainty of human motions and integrate these estimates into the RL costs to enhance safety under unpredictable conditions. Extensive experiments across both in-distribution and out-of-distribution settings demonstrate that our method achieves an effective proximity-safety balance compared to baselines. Real-robot deployment further validates the feasibility of our method in real-world scenarios. More details are available on our project page: https://nav-ps-balance.github.io/. I. INTRODUCTION Robots are increasingly expected to follow specific indi- viduals in applications such as healthcare, companionship, and assistance [1], requiring them to navigate safely among pedestrians and obstacles while maintaining close proximity to the target. This task involves an inherent conflict between proximity and safety. Unlike point-goal navigation, where the robot can freely reroute around obstacles, human following involves a moving target, and any detour to avoid collisions risks increasing the distance to or losing the target entirely. Dense crowds further intensify this conflict, as safe passages shrink and the robot must frequently decide how to balance proximity against safety. Learning-based methods offer advantages such as millisecond-level action generation via single forward in- ference and strong performance in in-distribution settings, making them attractive for real-time robot navigation. How- ever, existing RL approaches typically combine all objectives into a single dense reward through weighted summation [2], [3]. This formulation entangles the proximity-safety balance ∗ Equal Contribution † Corresponding author: jiachen li@gatech.edu 1 University of Pennsylvania, PA, USA 2 University of California, Riverside, CA, USA 3 Stanford University, CA, USA 4 NVIDIA Research, CA, USA 5 Georgia Institute of Technology, GA, USA within the reward weights, where the trade-off is implicitly determined by their relative magnitudes, which interact in complex ways during training, cannot be independently adjusted for each behavioral aspect, and offer no guarantee that the resulting behavior reflects the designer’s intended balance. This problem is exacerbated in real-world deploy- ment, where pedestrian behaviors are diverse and hard to anticipate, and a fixed implicit trade-off learned in training may fail to generalize to these challenging conditions. To address these limitations, we decompose the human- following task into a sparse task reward for target proximity and three independent cost constraints for following distance, human safety, and obstacle safety. Each constraint is mon- itored by a dedicated critic, and the trade-off is managed through cost thresholds with direct behavioral meaning, such as tolerable intrusions or a desired following distance, rather than reward weight ratios whose effect on policy behavior is implicit and cannot be directly verified. To further enhance safety in unpredictable pedestrian environments, we quantify prediction uncertainty of human motions [4], [5], [6] and incorporate these estimates into both the observation space and cost formulation, allowing the policy to adopt more conservative behavior when confidence in its predictions is low. Our policy network combines an attention-based Trans- former for social interaction modeling with a convolutional neural network for spatial constraint reasoning in a unified architecture [7]. The entire system is jointly optimized under a constrained reinforcement learning (CRL) framework using PPO-Lagrangian [8], [9]. To validate our approach, we extend the CrowdNav [10] simulator with static obstacles for evaluation under more realistic conditions. Across both in-distribution and out-of-distribution (OOD) scenarios with varying crowd densities, pedestrian behaviors, and environ- ment layouts, our method achieves higher success rates and lower collision rates than optimization-based baselines and ablation variants. We further deploy the policy on a real robot under ROS 2 to validate its effectiveness in the real world. The main contributions of this paper are as follows: • We propose a novel framework that decomposes human- following objectives into a sparse task reward and independent cost constraints, with the trade-off managed through behaviorally meaningful cost thresholds and jointly optimized using PPO-Lagrangian. • We integrate prediction uncertainty of human motions into the observation space and cost formulation, and design a unified policy network with spatial constraint reasoning to enable safe human following under uncer- tain and dynamic pedestrian environments. arXiv:2608.10056v1 [cs.RO] 10 Aug 2026 • We extend CrowdNav with static obstacles and con- duct extensive evaluations with varying crowd densities, pedestrian behaviors, and environment layouts, and de- ploy on a real robot under the ROS 2 setup to validate real-world effectiveness. I. RELATED WORK A. Human-Following Robot Human-following robots stay near a target human while moving, forming the foundation for many human-robot in- teraction tasks [11], [12]. Existing research focuses on two aspects: perception and decision making. For perception, approaches range from transmitter-based localization [13] to vision-based tracking under partial occlusion [14] and recent vision-language methods [15], though the latter can suffer from ambiguity when language cannot precisely describe the target or under out-of-distribution conditions. For deci- sion making, optimization-based methods embed following objectives into explicit cost formulations [16], [17], while learning-based approaches [2], [3] enable fast inference via end-to-end policies. However, most are evaluated in in- distribution settings or with constrained motion patterns, lacking analysis of dense crowds where humans exhibit complex OOD behaviors. Although we develop a tracking system for the target human, our core focus is on explicitly managing the proximity-safety trade-off in dense crowds through constraint decomposition. B. Constrained Reinforcement Learning CRL has gained increasing attention as an alternative to reward engineering [18], [19] and in safe robot learning [20], [21], [22], as it allows the integration of constraints into RL agents during the learning process, ensuring that the agents follow these constraints. Several methods are com- monly employed, including Constrained Policy Optimization (CPO) [23], Projection-Based Constrained Policy Optimiza- tion (PCPO) [24], and Lagrangian methods [9]. Among these, Lagrangian methods are particularly advantageous as they can be integrated with any existing RL algorithms and are relatively easy to implement. In this work, we build on PPO-Lagrangian [8], [9] and decompose the human- following task into multiple independent cost constraints with behaviorally meaningful thresholds, rather than encod- ing all objectives into a single reward. I. METHOD A. Problem Formulation We consider an environment with H humans and O static obstacles, indexed by h and o, over an episode with horizon T . One human is designated as the target h target , whom the robot must follow while maintaining a distance d follow ≤ d valid . We formulate this task as a Constrained Markov Decision Process (CMDP) [25]. At each timestep t, the agent observes a state S t comprising the robot’s current state, the positions and predicted trajectories of nearby humans from an upstream predictor, and local occupancy grid maps repre- senting static obstacles, and generates an action A t = (v x , v y ) controlling the robot’s velocity. Executing the action returns a reward R t and three distinct cost signals: 1) Following cost C F t , penalizing deviation from the target; 2) Human collision cost C H t , penalizing proximity to other humans; and 3) Obstacle collision cost C O t , penalizing proximity to static obstacles. This decomposition enables independent tuning of the robot’s behavior for each aspect of the task. Our objective is to learn an optimal policy π(A t | S t ) that maximizes cumulative reward while keeping human and obstacle safety costs below their respective limits δ H and δ O , and maintaining the following cost at a desired level δ F , where each threshold carries direct behavioral meaning, such as tolerable intrusions to pedestrians or obstacles, or a desired following distance to the target. B. Method Overview An overview of our method is illustrated in Fig. 1. At each timestep t, the state S t consists of the robot’s physical state, local occupancy grid maps representing static obsta- cles, and the current positions, predicted trajectories, and quantified prediction uncertainties of nearby humans. The state is processed by a unified policy network, where CNN- encoded occupancy features and human and robot tokens are combined into a sequence and processed by self-attention mechanisms [26], producing a feature representation that serves as input to the actor-critic framework. We employ an actor-critic framework for policy learning, where the actor network is guided by multiple critics. As shown in Fig. 2, the four critics include: a reward critic corresponding to the task reward, a following cost critic for maintaining proximity to the target human, an obstacle collision cost critic, and a human collision cost critic, where uncertainty estimates are further incorporated into the cost formulation. Each critic applies Generalized Advantage Estimation (GAE) [27] to compute advantages and returns. These critics jointly guide the actor’s policy updates via PPO-Lagrangian [8], [9], where Lagrangian multipliers are dynamically adjusted to enforce the cost constraints on following and collision avoidance. C. Uncertainty-Aware Policy Network for Safe Following We design a Transformer-based policy network that rea- sons about complex social interactions among agents and spatial constraints from static obstacles for safe human following in dense crowds. Spatial-temporal scene context is encoded using a 3D convolutional neural network (CNN) that extracts features from local occupancy grid maps (OGMs) over the most recent 5 time steps, leveraging the inductive bias of convolutions for capturing local spatial patterns in grid-structured data. For all detected humans, we construct uncertainty-aware feature tokens comprising the current po- sition, a sequence of K predicted future positions, and their quantified uncertainties: h h (t) = [p h (t), p h,1 (t),..., p h,K (t); ˆ δ h,1 (t),..., ˆ δ h,K (t)],(1) where p h (t) is the current position of the h-th human at time t, p h,k (t) = (x h,k (t), y h,k (t)) denotes the k-th predicted position, and ˆ δ h,k (t) is the estimated prediction uncertainty Obstacle OGMs 3D CNN Human Positions Trajectory Prediction ACI Prediction Uncertainty Robot State Token Sequence Self-Attention MLP Critics Actor <latexit sha1_base64="FRbFVM055r8v9DuVfV/kq+7m1Yc=">AAACynicjVHLSsNAFD2Nr1pfVZdugkVwVVIpfeyKbly4qGAfYKsk6bQOzYvJRCilO3/ArX6Y+Af6F94Z06KI6A1Jzpx7zp25c53I47G0rNeMsbS8srqWXc9tbG5t7+R399pxmAiXtdzQC0XXsWPm8YC1JJce60aC2b7jsY4zPlP5zj0TMQ+DKzmJWN+3RwEfcteWRHW8m2kv4rPbfMEqWjrMn6CUggLSaIb5F/QwQAgXCXwwBJCEPdiI6blGCRYi4vqYEicIcZ1nmCFH3oRUjBQ2sWP6jmh1nbIBrVXNWLtd2sWjV5DTxBF5QtIJwmo3U+cTXVmxv9We6prqbBP6O2ktn1iJO2L/8s2V//WpXiSGqOkeOPUUaUZ156ZVEn0r6uTml64kVYiIU3hAeUHY1c75PZvaE+ve1d3aOv+mlYpVazfVJnhXp9QDrteq9XJpMU4acJ2iWl8w7ZNiqVKsXJYLjdN01Fkc4BDHNM8qGjhHEy3d5SOe8GxcGMKYGNNPqZFJPfv4FsbDB201kps=</latexit> l ⇡ || Humans Following Target + Fig. 1. An overview diagram of our method. Components related to static obstacles are highlighted in orange, those related to humans in yellow, the shared feature extraction modules in green, and those concerning the robot’s physical state and decision making in blue. The detailed mechanism by which the critics generate l π is illustrated in Fig. 2. for the k-th step. The uncertainty is computed by an adaptive conformal inference (ACI) module [6], [20] that maintains online error bounds adapted to the actual prediction quality. At each timestep t, we compute the prediction error δ h,k (t) = ∥p h (t)− p h,k (t−k)∥ 2 for the k-th horizon of the h-th human, where p h (t) is the observed position at time t and p h,k (t−k) is the k-step prediction calculated at time t−k. We run M parallel estimators to estimate the error bound, updated as ˆ δ (m) h,k (t+1) = ˆ δ (m) h,k (t)+ γ (m) 1 h δ h,k (t)> ˆ δ (m) h,k (t) i − α , (2) where γ (m) is the learning rate of the m-th estimator and α∈ (0, 1) is the target miscoverage rate. The final uncertainty estimate ˆ δ h,k (t) is selected from the M estimators via adaptive weighted sampling based on their cumulative performance. These per-human, per-horizon estimates provide the policy with an adaptive measure of prediction reliability, and the same ACI bounds are also incorporated into the human safety cost design in Sec. I-D. To combine information from all agents and scene ele- ments, we implement early fusion [28] by embedding the robot state r(t), target information tg(t), extracted obsta- cle features o(t), and human features h(t) into a shared representation space: x(t) = [ r(t), tg(t), o(t), h(t)], where tg(t) = h target , and x(t) ∈ R N×d , with N representing the total number of tokens and d the embedding dimension. We represent each entity, including the robot, the target, perceived obstacles, and all detected humans, as an individual feature embedding, allowing the self-attention mechanism to model all pairwise interactions across the full scene. In particular, obstacle-human and obstacle-robot relationships can be explicitly captured, whereas these interactions are lost when obstacle features are appended after the attention stage [29]. Human tokens are ordered by each human’s distance to the robot, allowing the positional encoding to reflect proximity and enabling the attention mechanism to prioritize nearby agents. We apply sinusoidal positional encoding [26] and a binary mask for undetected humans: x pe (t) = PE(x(t))⊙ M(t), where M(t) ∈ 0, 1 N×d zeroes out tokens of undetected humans. The encoded sequence is processed by a Trans- former encoder as z(t) =T (x pe (t)), and we extract the robot token output z r (t)∈ R d to produce the final representation GAE Human Collision cost Human Collision critic Human Collision advantage Human Collision return GAE Obstacle Collision cost Obstacle Collision critic Obstacle Collision advantage Obstacle Collision return GAE Following cost Following critic Following advantage Following return GAE Reward Reward critic Reward advantage Reward return Update Update Update <latexit sha1_base64="sdyQPgo1SegXDGPyyFEfe7dzsRE=">AAACxHicjVHbSsNAED2Nt1pvVR99CRZBEEoipW3eioL42IK9QC2SbLc1NDeSjVCK/oCv+m3iH+hfOLumRRHRCUnOnplzdmfHiTw3EYbxmtOWlldW1/LrhY3Nre2d4u5eJwnTmPE2C70w7jl2wj034G3hCo/3opjbvuPxrjM5l/nuHY8TNwyuxDTiA98eB+7IZbYgqnVyUywZZUOF/hOYGSghi2ZYfME1hgjBkMIHRwBB2IONhJ4+TBiIiBtgRlxMyFV5jnsUSJtSFacKm9gJfce06mdsQGvpmSg1o108emNS6jgiTUh1MWG5m67yqXKW7G/eM+Upzzalv5N5+cQK3BL7l25e+V+d7EVghLrqwaWeIsXI7ljmkqpbkSfXv3QlyCEiTuIh5WPCTCnn96wrTaJ6l3drq/ybqpSsXLOsNsW7PKUasFWvWRVzMU4asEVRsxZM57RsVsvVVqXUOMtGnccBDnFM86yhgUs00Vbej3jCs3aheVqipZ+lWi7T7ONbaA8f3PiPkw==</latexit> + <latexit sha1_base64="n+pJkb644R7WYPweu780beWgOXo=">AAACyXicjVHLTsJAFD3UF+ILdemmkZi4IsUQoDuiGxM3mMgjAWPaMuBIX7ZTIxJX/oBb/THjH+hfeGcsRGOM3qbtmXPvOTN3rh26PBaG8ZrR5uYXFpeyy7mV1bX1jfzmVisOkshhTSdwg6hjWzFzuc+agguXdcKIWZ7tsrY9OpL59g2LYh74Z2IcsnPPGvp8wB1LENXqCe6x+CJfMIqGCv0nKKWggDQaQf4FPfQRwEECDww+BGEXFmJ6uijBQEjcOSbERYS4yjPcI0fahKoYVVjEjug7pFU3ZX1aS89YqR3axaU3IqWOPdIEVBcRlrvpKp8oZ8n+5j1RnvJsY/rbqZdHrMAlsX/pppX/1cleBAaoqR449RQqRnbnpC6JuhV5cv1LV4IcQuIk7lM+Iuwo5fSedaWJVe/ybi2Vf1OVkpVrJ61N8C5PqQZs1qpmuTQbJw3YpKiaM6Z1UCxVipXTcqF+mI46ix3sYp/mWUUdx2igSd5XeMQTnrUT7Vq71e4+S7VMqtnGt9AePgAla5IY</latexit> ⇥ <latexit sha1_base64="n+pJkb644R7WYPweu780beWgOXo=">AAACyXicjVHLTsJAFD3UF+ILdemmkZi4IsUQoDuiGxM3mMgjAWPaMuBIX7ZTIxJX/oBb/THjH+hfeGcsRGOM3qbtmXPvOTN3rh26PBaG8ZrR5uYXFpeyy7mV1bX1jfzmVisOkshhTSdwg6hjWzFzuc+agguXdcKIWZ7tsrY9OpL59g2LYh74Z2IcsnPPGvp8wB1LENXqCe6x+CJfMIqGCv0nKKWggDQaQf4FPfQRwEECDww+BGEXFmJ6uijBQEjcOSbERYS4yjPcI0fahKoYVVjEjug7pFU3ZX1aS89YqR3axaU3IqWOPdIEVBcRlrvpKp8oZ8n+5j1RnvJsY/rbqZdHrMAlsX/pppX/1cleBAaoqR449RQqRnbnpC6JuhV5cv1LV4IcQuIk7lM+Iuwo5fSedaWJVe/ybi2Vf1OVkpVrJ61N8C5PqQZs1qpmuTQbJw3YpKiaM6Z1UCxVipXTcqF+mI46ix3sYp/mWUUdx2igSd5XeMQTnrUT7Vq71e4+S7VMqtnGt9AePgAla5IY</latexit> ⇥ <latexit sha1_base64="n+pJkb644R7WYPweu780beWgOXo=">AAACyXicjVHLTsJAFD3UF+ILdemmkZi4IsUQoDuiGxM3mMgjAWPaMuBIX7ZTIxJX/oBb/THjH+hfeGcsRGOM3qbtmXPvOTN3rh26PBaG8ZrR5uYXFpeyy7mV1bX1jfzmVisOkshhTSdwg6hjWzFzuc+agguXdcKIWZ7tsrY9OpL59g2LYh74Z2IcsnPPGvp8wB1LENXqCe6x+CJfMIqGCv0nKKWggDQaQf4FPfQRwEECDww+BGEXFmJ6uijBQEjcOSbERYS4yjPcI0fahKoYVVjEjug7pFU3ZX1aS89YqR3axaU3IqWOPdIEVBcRlrvpKp8oZ8n+5j1RnvJsY/rbqZdHrMAlsX/pppX/1cleBAaoqR449RQqRnbnpC6JuhV5cv1LV4IcQuIk7lM+Iuwo5fSedaWJVe/ybi2Vf1OVkpVrJ61N8C5PqQZs1qpmuTQbJw3YpKiaM6Z1UCxVipXTcqF+mI46ix3sYp/mWUUdx2igSd5XeMQTnrUT7Vq71e4+S7VMqtnGt9AePgAla5IY</latexit> ⇥ <latexit sha1_base64="dQvZCaZsiTuS65hoSuwHm4iDvAU=">AAACznicjVHLSsNAFD2Nr1pfVZdugkVwVVKRttkVBXFZwT6grZKk0xqaF5NJoZTi1h9wq58l/oH+hXfGtCgiekOSO+eec+c+7MhzY2EYrxltaXlldS27ntvY3Nreye/uNeMw4Q5rOKEX8rZtxcxzA9YQrvBYO+LM8m2PtezRuYy3xozHbhhci0nEer41DNyB61iCoE7XI2rfuplezG7zBaNoKNN/OqXUKSC1eph/QRd9hHCQwAdDAEG+BwsxPR2UYCAirIcpYZw8V8UZZsiRNiEWI4ZF6Ii+Qzp1UjSgs8wZK7VDt3j0clLqOCJNSDxOvrxNV/FEZZbob7mnKqesbUJ/O83lEypwR+hfujnzvzrZi8AAVdWDSz1FCpHdOWmWRE1FVq5/6UpQhogw6fcpzsl3lHI+Z11pYtW7nK2l4m+KKVF5dlJugndZpVqwWa2Yp6XFOmnBJlnFXCDNk2KpXCxfnRZqZ+mqszjAIY5pnxXUcIk6Gmrij3jCs1bXxtpMu/+kaplUs49vpj18AFpHlCU=</latexit> F <latexit sha1_base64="Uxghfb3q+rY+mKQ3M8VligdO4oQ=">AAACznicjVHLTsJAFD3UF+ILdemmkZi4Iq0xQHdENywxkUcCaNoyYENpm+mUhBDi1h9wq59l/AP9C++MhWiM0du0vXPuOXfuw4l8LxaG8ZrRVlbX1jeym7mt7Z3dvfz+QTMOE+6yhhv6IW87dsx8L2AN4QmftSPO7LHjs5YzupTx1oTx2AuDazGNWG9sDwNv4Lm2IKjT9Ynat29mtfltvmAUDWX6T8dMnQJSq4f5F3TRRwgXCcZgCCDI92EjpqcDEwYiwnqYEcbJ81ScYY4caRNiMWLYhI7oO6RTJ0UDOsucsVK7dItPLyeljhPShMTj5MvbdBVPVGaJ/pZ7pnLK2qb0d9JcY0IF7gj9S7dg/lcnexEYoKJ68KinSCGyOzfNkqipyMr1L10JyhARJv0+xTn5rlIu5qwrTax6l7O1VfxNMSUqz27KTfAuq1QLtipl69xcrpMWbJGVrSXSPCuapWLp6rxQvUhXncURjnFK+yyjihrqaKiJP+IJz1pdm2hz7f6TqmVSzSG+mfbwAV8JlCc=</latexit> H <latexit sha1_base64="ylXLTceAwx61zjyTfpDKxKBA5mk=">AAACznicjVHLSsNAFD2Nr1pfVZdugkVwVVKRttkV3bizgn1AWyVJpzU0LyaTQinFrT/gVj9L/AP9C++MaVFE9IYkd84958592JHnxsIwXjPa0vLK6lp2PbexubW9k9/da8Zhwh3WcEIv5G3bipnnBqwhXOGxdsSZ5dsea9mjcxlvjRmP3TC4FpOI9XxrGLgD17EEQZ2uR9S+dTO9nN3mC0bRUKb/dEqpU0Bq9TD/gi76COEggQ+GAIJ8DxZiejoowUBEWA9Twjh5roozzJAjbUIsRgyL0BF9h3TqpGhAZ5kzVmqHbvHo5aTUcUSakHicfHmbruKJyizR33JPVU5Z24T+dprLJ1TgjtC/dHPmf3WyF4EBqqoHl3qKFCK7c9IsiZqKrFz/0pWgDBFh0u9TnJPvKOV8zrrSxKp3OVtLxd8U6Ly7KTcBO+ySrVgs1oxT0uLddKCTbKKuUCaJ8VSuVi+Oi3UztJVZ3GAQxzTPiuo4QJ1NNTEH/GEZ62ujbWZdv9J1TKpZh/fTHv4AG+wlC4=</latexit> O <latexit sha1_base64="9TXrMLBtTiErRH5kYCmt2MJ9dIU=">AAACxHicjVHLSsNAFD2Nr1pfVZdugkVwY0mktM2uKIjLFuwDapFkOq2heZFMhFL0B9zqt4l/oH/hnTEtiojekOTMufecmTvXiTw3EYbxmtOWlldW1/LrhY3Nre2d4u5eJwnTmPE2C70w7jl2wj034G3hCo/3opjbvuPxrjM5l/nuHY8TNwyuxDTiA98eB+7IZbYgqnVyUywZZUOF/hOYGSghi2ZYfME1hgjBkMIHRwBB2IONhJ4+TBiIiBtgRlxMyFV5jnsUSJtSFacKm9gJfce06mdsQGvpmSg1o108emNS6jgiTUh1MWG5m67yqXKW7G/eM+Upzzalv5N5+cQK3BL7l25e+V+d7EVghLrqwaWeIsXI7ljmkqpbkSfXv3QlyCEiTuIh5WPCTCnn96wrTaJ6l3drq/ybqpSsXLOsNsW7PKUasFWvWRVzMU4asEVRsxZM57RsVsvVVqXUOMtGnccBDnFM86yhgUs00Vbej3jCs3aheVqipZ+lWi7T7ONbaA8f4biPlQ==</latexit> <latexit sha1_base64="4pYbLPZPwnkP9xg3Qr0vI9J9I+I=">AAACyHicjVHLSsNAFD2Nr1pfVZdugkVwVRIpfeyKgoirCqYt1CpJOq1D0yQkE6WUbvwBt/pl4h/oX3hnTIsiojckOXPuPWfmznVCj8fCMF4z2sLi0vJKdjW3tr6xuZXf3mnGQRK5zHIDL4jajh0zj/vMElx4rB1GzB45Hms5wxOZb92xKOaBfynGIeuO7IHP+9y1BVGWdz05nd7kC0bRUKH/BGYKCkijEeRfcIUeArhIMAKDD0HYg42Yng5MGAiJ62JCXESIqzzDFDnSJlTFqMImdkjfAa06KevTWnrGSu3SLh69ESl1HJAmoLqIsNxNV/lEOUv2N++J8pRnG9PfSb1GxArcEvuXblb5X53sRaCPquqBU0+hYmR3buqSqFuRJ9e/dCXIISRO4h7lI8KuUs7uWVeaWPUu79ZW+TdVKVm5dtPaBO/ylGrAtWqlVjLn46QB1ygqtTnTPCqa5WL5olSoH6ejzmIP+zikeVZQxxkasMib4xFPeNbOtVC718afpVom1eziW2gPH+U1kZg=</latexit> l F <latexit sha1_base64="ABn+CWxWUGVArTnXxgNSvQllyQ0=">AAACyHicjVHLTsJAFD3UF+ILdemmkZi4Iq0hPHZEN8QVJhZIEE1bBpxQ2qadaghh4w+41S8z/oH+hXfGQjTG6G3anjn3njNz5zqhx2NhGK8ZbWl5ZXUtu57b2Nza3snv7rXiIIlcZrmBF0Qdx46Zx31mCS481gkjZo8dj7Wd0ZnMt+9YFPPAvxSTkPXG9tDnA+7agijLu542Zjf5glE0VOg/gZmCAtJoBvkXXKGPAC4SjMHgQxD2YCOmpwsTBkLiepgSFxHiKs8wQ460CVUxqrCJHdF3SKtuyvq0lp6xUru0i0dvREodR6QJqC4iLHfTVT5RzpL9zXuqPOXZJvR3Uq8xsQK3xP6lm1f+Vyd7ERigqnrg1FOoGNmdm7ok6lbkyfUvXQlyCImTuE/5iLCrlPN71pUmVr3Lu7V/k1VSlau3bQ2wbs8pRpwrVqplczFOGnANYpKbcG0TopmuVi+KBXqp+moszjAIY5pnhXU0UATFnlzPOIJz9q5Fmr32uSzVMukmn18C+3hA+n3kZo=</latexit> l H <latexit sha1_base64="heSAISS/ZwLkoqX/9+TabaWS/YQ=">AAACyHicjVHLTsJAFD3UF+ILdemmkZi4Iq0hPHZEN8aNmFggQTRtGXBCaZt2qiGEjT/gVr/M+Af6F94ZC9EYo7dpe+bce87MneuEHo+FYbxmtIXFpeWV7GpubX1jcyu/vdOMgyRymeUGXhC1HTtmHveZJbjwWDuMmD1yPNZyhicy37pjUcwD/1KMQ9Yd2QOf97lrC6Is73pyPr3JF4yioUL/CcwUFJBGI8i/4Ao9BHCRYAQGH4KwBxsxPR2YMBAS18WEuIgQV3mGKXKkTaiKUYVN7JC+A1p1UtantfSMldqlXTx6I1LqOCBNQHURYbmbrvKJcpbsb94T5SnPNqa/k3qNiBW4JfYv3azyvzrZi0AfVdUDp55Cxcju3NQlUbciT65/6UqQQ0icxD3KR4RdpZzds640sepd3q2t8m+qUrJy7a1Cd7lKdWAa9VKrWTOx0kDrlFUanOmeVQ0y8XyRalQP05HncUe9nFI86ygjlM0YJE3xyOe8KydaaF2r40/S7VMqtnFt9AePgD6npGh</latexit> l O <latexit sha1_base64="tGwCWP/DMYycS8Thx7eLjvodaXE=">AAACyHicjVHLTsJAFD3UF+ILdemmkZi4Iq0hPHZEN8YVGgskiKYtA04obdNONYSw8Qfc6pcZ/0D/wjtjIRpj9DZtz5x7z5m5c53Q47EwjNeMtrC4tLySXc2trW9sbuW3d5pxkEQus9zAC6K2Y8fM4z6zBBcea4cRs0eOx1rO8ETmW3csinngX4pxyLoje+DzPndtQZTlXU8upjf5glE0VOg/gZmCAtJoBPkXXKGHAC4SjMDgQxD2YCOmpwMTBkLiupgQFxHiKs8wRY60CVUxqrCJHdJ3QKtOyvq0lp6xUru0i0dvREodB6QJqC4iLHfTVT5RzpL9zXuiPOXZxvR3Uq8RsQK3xP6lm1X+Vyd7Eeijqnrg1FOoGNmdm7ok6lbkyfUvXQlyCImTuEf5iLCrlLN71pUmVr3Lu7V/k1VSlau3bQ2wbs8pRpwrVqplcz5OGnANYpKbc40j4pmuVg+LxXqx+mos9jDPg5pnhXUcYoGLPLmeMQTnrUzLdTutfFnqZZJNbv4FtrDBwHQkaQ=</latexit> l R <latexit sha1_base64="FRbFVM055r8v9DuVfV/kq+7m1Yc=">AAACynicjVHLSsNAFD2Nr1pfVZdugkVwVVIpfeyKbly4qGAfYKsk6bQOzYvJRCilO3/ArX6Y+Af6F94Z06KI6A1Jzpx7zp25c53I47G0rNeMsbS8srqWXc9tbG5t7+R399pxmAiXtdzQC0XXsWPm8YC1JJce60aC2b7jsY4zPlP5zj0TMQ+DKzmJWN+3RwEfcteWRHW8m2kv4rPbfMEqWjrMn6CUggLSaIb5F/QwQAgXCXwwBJCEPdiI6blGCRYi4vqYEicIcZ1nmCFH3oRUjBQ2sWP6jmh1nbIBrVXNWLtd2sWjV5DTxBF5QtIJwmo3U+cTXVmxv9We6prqbBP6O2ktn1iJO2L/8s2V//WpXiSGqOkeOPUUaUZ156ZVEn0r6uTml64kVYiIU3hAeUHY1c75PZvaE+ve1d3aOv+mlYpVazfVJnhXp9QDrteq9XJpMU4acJ2iWl8w7ZNiqVKsXJYLjdN01Fkc4BDHNM8qGjhHEy3d5SOe8GxcGMKYGNNPqZFJPfv4FsbDB201kps=</latexit> l ⇡ Fig. 2.The interaction among the critics and the generation of the action loss. More details can be found in Section I-D. f r (t) = MLP(z r (t)), which serves as input to our CRL policy module described in Sec. I-D. D. Task Decomposition for Explicit Behavior Control Human following in dense crowds requires simultaneously minimizing collision risk and maintaining an appropriate following distance, two inherently conflicting objectives. Encoding all these objectives into a single reward function makes it difficult to independently control each behavior, as reward weights implicitly determine the trade-off and cannot be directly mapped to interpretable behavioral outcomes [18]. We instead formulate the problem as a CMDP [25] that decouples task success from behavioral constraints, enabling explicit and independent control over each aspect: max π E τ∼π " T ∑ t=0 R t (S t , A t ) # s.t. E τ∼π " T ∑ t=0 C F t # = δ F , E τ∼π " T ∑ t=0 C H t # ≤ δ H , E τ∼π " T ∑ t=0 C O t # ≤ δ O , (3) where R t represents the reward function, C F t , C H t , and C O t denote the cost functions for following behavior, human safety, and obstacle safety, respectively, and δ F , δ H , δ O are the corresponding cost thresholds. The safety costs use inequality constraints since safer behavior is always prefer- able, while the following cost uses an equality constraint, maintaining a desired distance rather than minimizing it to avoid both target loss and personal-space intrusion. Our reward function provides sparse feedback only for the task outcomes: R t (S t , A t ) = R success ,if S t ∈ S success , R collision ,if S t ∈ S collision , R target lost , if S t ∈ S targetlost , (4) while all behavioral objectives are managed through the three cost constraints. Each cost function isolates a specific behavior, and its constraint threshold δ i provides direct and interpretable control over the desired trade-off. First, the following cost C F t measures the robot’s distance beyond the personal space threshold d personal , encouraging the robot to maintain a moderate following distance rather than either lagging behind or crowding the target: C F t = ( k 1 (d follow,t − d personal ), if d follow,t > d personal , 0,otherwise, (5) where d follow,t is the Euclidean distance between the robot and the target at time t, d personal is the personal space threshold, and k 1 is the penalty coefficient. Second, the human safety cost C H t quantifies the risk of collision with other pedestrians by measuring intrusions into uncertainty-aware safety regions around each human. For each human h, we define a safety region around the current position p h (t) with radius r ego + r h + r buf , where r ego and r h are the radii of the robot and human, and r buf is a fixed buffer margin. To account for the inherent uncertainty in pedestrian motion prediction, we additionally define safety regions around each of the first K ′ predicted positions p h,k (t) with radius r ego + r h + ˆ δ h,k (t), where the ACI uncertainty bound ˆ δ h,k (t) from Sec. I-C scales each safety region according to the local prediction confidence at that horizon, expanding protection when motion is uncertain and contracting it when predictions are reliable. The cost is proportional to the maximum intrusion depth across all humans and all safety regions at each timestep: C H t = k 2 d intru,t ,(6) where k 2 is the penalty coefficient and d intru,t is the maximum intrusion distance. This formulation aims to ensure safety un- der unpredictable pedestrian behaviors by grounding spatial protection in quantified prediction uncertainty. Third, the obstacle safety cost penalizes the robot for getting too close to static obstacles: C O t = ( 0,if min o (d surf o,t )≥ d safe o , k 3 (d safe o − min o (d surf o,t )), if min o (d surf o,t )< d safe o , (7) where min o (d surf o,t ) is the minimum surface distance from the robot to all static obstacles at time t, d safeo is the safe distance threshold, and k 3 is the penalty coefficient. We employ PPO-Lagrangian [9] to optimize a single unified policy under the three cost constraints, with four separate critics estimating the value functions for the reward and each cost independently: l R t = c 1 (V R θ 1 (S t )− V targ,R t ) 2 ,l F t = c 2 (V F θ 2 (S t )− V targ,F t ) 2 , l H t = c 3 (V H θ 3 (S t )− V targ,H t ) 2 ,l O t = c 4 (V O θ 4 (S t )− V targ,O t ) 2 , (8) where c 1 through c 4 are loss weighting coefficients, V (·) θ i (S t ) are the predicted value functions, and V targ,(·) t are the corre- sponding target values computed from collected rollouts. We compute four distinct advantage estimates using GAE [27], including the reward advantage ˆ A R t and three cost advantages ˆ A F t , ˆ A H t , ˆ A O t . These are combined into a single advantage signal: ˆ A ′ t = ˆ A R t − λ F ˆ A F t − λ H ˆ A H t − λ O ˆ A O t 1+ λ F + λ H + λ O ,(9) where λ F , λ H , λ O are Lagrangian multipliers, and the de- nominator ensures the combined advantage has a consistent scale regardless of the number of active constraints. The actor is updated using the PPO clipping objective [30]: l π t =− min ρ t ˆ A ′ t , clip(ρ t , 1−ε, 1+ε) ˆ A ′ t ,(10) where ρ t = π θ (A t |S t )/π θ old (A t |S t ) is the importance sampling ratio and ε is the clipping parameter. The Lagrangian multi- pliers are updated via dual gradient descent after each epoch: λ i ← max 0, λ i + η λ ˆ J C i (π)− δ i ,i∈H, O,(11) λ F ← λ F + η λ ˆ J C F (π)− δ F ,(12) where η λ is the multiplier learning rate and ˆ J C i (π) is the empirical average cost. When the current cost exceeds its limit, the corresponding multiplier increases, amplifying the cost advantage’s contribution to ˆ A ′ t and steering the policy toward constraint satisfaction, and vice versa. This mecha- nism provides automatic and interpretable control over each behavioral constraint, directly linking the cost thresholds δ i to the robot’s safety-performance trade-off. While multiple critics are employed during training, only the actor network is needed at inference, so the computational cost at deploy- ment is identical to single-critic approaches. IV. EXPERIMENTS A. Simulation Settings We construct cluttered indoor environments fully enclosed by walls, measuring up to 20 m in both dimensions, with randomly generated static obstacles of varying sizes and locations to ensure scene-level diversity. The robot and 40 humans are placed in the scene, with the robot initialized at a random position with a maximum speed of 1.2 m/s and a tar- get human sampled within 1.6 m. The remaining 39 humans are randomly positioned and controlled by ORCA [31], with radii sampled between 0.3 m–0.4 m and maximum speeds between 0.7 m/s–1.4 m/s. Once a human reaches its goal, a new goal is assigned. B. Evaluation Metrics Our evaluation metrics include: 1) Success Rate (SR): SR is defined as the ratio of successful following episodes to the total number of test episodes. 2) Collision Rate (CR): CR is the ratio of episodes in which the robot collides with a human or an obstacle. 3) Target Lost Rate (TLR): TLR is the ratio of episodes in which the distance between the robot and the target human d follow exceeds the valid threshold d valid at any time during the episode. 4) Average Following Distance (AFD): AFD is the average distance between the robot and the target human throughout the entire episode. TABLE I IN-DISTRIBUTION TEST RESULTS MethodsSR↑ CR↓ TLR↓ AFD Overall Human Obstacle SG-HA*1.84%81.52% 79.92%1.60%16.64% 1.91 SG-ORCA17.68% 64.88% 27.84%37.04%17.44% 1.65 SG-MPC30.96% 42.72% 30.08%12.64%26.32% 2.64 OGM-HEIGHT 52.32% 34.72% 21.44%13.28%12.96% 2.19 RL68.00% 24.08% 15.76%8.32%7.92%2.24 RL+ACI71.60% 20.72% 12.80%7.92%7.68%2.28 Ours78.08% 16.16% 10.72%5.44%5.76%2.35 C. Baselines and Ablation Models We compare our CRL-based framework against represen- tative baselines. Following the common two-step paradigm of subgoal generation and downstream planning [32], [16], [13], we construct three subgoal-guided baselines using the subgoal generation strategy from [13]: 1) SG-HA*: Hybrid A* [33], a search-based planner with continuous motion primitives; 2) SG-MPC: An optimization-based controller that incorporates following task objectives [16], [17] and an uncertainty-aware cost function for dynamic human and static obstacle avoidance [34]; 3) SG-ORCA: ORCA [31], a classic velocity-based collision avoidance algorithm. We also include the state-of-the-art RL-based method HEIGHT [29]: 4) OGM-HEIGHT : HEIGHT with its point cloud input replaced by OGMs for simulation compatibility, retrained to convergence with a following distance penalty in place of the original navigation reward, where the OGMs preserve the same local obstacle geometry as the point cloud input for a fair comparison. To validate the contributions of our uncertainty integration and cost decomposition, we include two ablation models: 5) RL: Our policy network trained with standard RL and the same reward function as OGM- HEIGHT, without uncertainty estimates in the observations; and 6) RL+ACI: Our full policy network trained with stan- dard RL, where the three behavioral objectives are incorpo- rated as weighted penalty terms into a single scalar reward, with the human safety penalty remaining uncertainty-aware using the same ACI-derived bounds as our method. For both RL+ACI and our method, we report the best result among all tuned configurations (see Table I). D. Implementation Details We train on an NVIDIA RTX 4090 GPU with a batch size of 480 and a clip parameter of 0.02 for PPO-Lagrangian. The valid following threshold d valid = 5.0 m also serves as the perception range, the personal distance is d personal = 1.0 m beyond which following costs are incurred, and the obstacle safety distance is d safe o = 0.50 m. The policy network uses a Transformer encoder with 4 layers and 8 attention heads, and we evaluate 1250 samples across 5 random seeds. Trajectory predictions use a constant velocity (CV) model [35], which is simple yet sufficient for our setting: the online adaptation of ACI does not assume a highly accurate predictor, but instead expands uncertainty bounds as prediction errors grow, TABLE I EFFECT OF CONSTRAINT TUNING VS. REWARD WEIGHT TUNING IntentMethodSR↑ CR↓ TLR↓ AFD Overall Human Safety Ours (δ F =4.0, δ H =3.2) 71.68% 18.24%8.80%10.08% 2.54 RL+ACI (w H × 2)71.60% 20.72% 12.80%7.68%2.28 Following Ours (δ F =3.2, δ H =4.0) 74.40% 22.40% 14.24%3.20%2.23 RL+ACI (w F × 2)63.92% 27.36% 15.44%8.72%2.26 Balanced Ours (δ F =3.6, δ H =3.6) 78.08% 16.16% 10.72%5.76%2.35 RL+ACI (w F =w H =1)70.80% 21.12% 12.88%8.08%2.26 maintaining conservative safety regions under the non-linear pedestrian behaviors in our OOD scenarios. E. In-Distribution Results The test results under the same setting as the training environment are shown in Table I. Traditional subgoal- guided methods (SG-HA*, SG-ORCA, SG-MPC) exhibit substantially lower SR and higher CR and TLR compared to RL-based methods, indicating limited ability to main- tain following behavior in the presence of dense, dynamic pedestrians. Among RL-based methods, our approach sig- nificantly outperforms OGM-HEIGHT, achieving a 25.76% higher SR while reducing overall CR by 18.56%, human CR by 10.72%, and obstacle CR to the lowest 5.44% among learning-based methods. The ablation results further show that RL+ACI improves over RL in human CR, confirming that uncertainty-aware observations enable more cautious be- havior around unpredictable pedestrians, and our full model achieves the best overall performance, outperforming the two ablation models by 10.08% and 6.48% in SR while reducing human CR by 5.04% and 2.08%, obstacle CR by 2.88% and 2.48%, and TLR by 2.16% and 1.92%. These results validate the effectiveness of our cost decomposition in achieving a good balance between proximity and safety. We visualize the behaviors of Ours and OGM-HEIGHT in the same episode in Fig. 3(a) and Fig. 3(b). OGM- HEIGHT moves directly toward the target but becomes trapped as pedestrians begin moving, ultimately colliding. Our method instead responds to expanding uncertainty areas by steering toward less crowded regions while maintaining the general direction and proximity toward the target at steps 4 and 18. At step 92, when a pedestrian suddenly changes direction, our robot adjusts its trajectory based on the updated uncertainty bounds, successfully avoiding collision while maintaining target proximity. This demonstrates that cost de- composition and uncertainty-aware cost formulation enable the robot to effectively balance safety and following under dynamic and unpredictable pedestrian behaviors. F. Cost Limits vs. Reward Weights for Behavioral Control To validate that explicit cost constraints provide more interpretable and predictable behavioral control than reward- based penalty encoding, we compare cost limit tuning against reward weight tuning under three behavioral intents in Ta- ble I. In both cases, the same policy network architecture and ACI-derived uncertainty bounds are used. For RL+ACI, TABLE I OUT-OF-DISTRIBUTION TEST RESULTS Environments Methods SR↑ CR↓ TLR↓ AFD Overallw/ Humansw/ Obstacles Corridor SG-HA*2.64%88.64%88.56% 0.08% 8.72%1.46 SG-ORCA50.64%47.20%25.28%21.92%2.16%1.48 SG-MPC48.88%39.20%37.04%2.16%11.92%1.84 OGM-HEIGHT56.24%37.60%28.96%8.64%6.16%1.50 RL 77.52%21.84%21.04%0.80%0.64%1.50 RL+ACI82.96%16.72%15.12%1.60% 0.32% 1.53 Ours 89.76%8.64%8.48% 0.16%1.60%1.78 15% Rushing Humans SG-HA*0.16%93.68%92.72% 0.96% 6.16%1.93 SG-ORCA12.48%81.76%37.28%44.48%5.76%1.39 SG-MPC15.76%72.48%57.84%14.64%11.76%2.66 OGM-HEIGHT47.04%38.08%26.56%11.52%14.88%2.22 RL 61.60%35.36%27.36%8.00% 3.04% 2.15 RL+ACI68.48%28.48%23.04%5.44% 3.04% 2.18 Ours 70.56%22.72%19.36% 3.36%6.72%2.23 SF Pedestrian Model SG-HA*0.72%84.88%84.08% 0.80% 14.40%1.74 SG-ORCA22.24%68.80%14.40%54.40%8.96%1.53 SG-MPC29.60%45.92%31.20%14.72%24.48%2.43 OGM-HEIGHT35.84%50.40%22.88%27.52%13.76%2.16 RL 56.96%40.96%21.28%19.68%2.08%2.15 RL+ACI60.64%34.48%21.68%12.80%4.88%2.19 Ours 64.32%33.68%20.08% 13.60% 2.00% 2.21 Groups SG-HA*0.08%92.64%92.64%0.00%7.28%1.83 SG-ORCA27.68%46.48%46.48%0.00%25.84%1.86 SG-MPC26.88%64.08%64.08%0.00%9.04%2.43 OGM-HEIGHT36.32%53.28%53.28%0.00%10.40%2.09 RL 49.36%46.64%46.64%0.00% 4.00% 1.87 RL+ACI56.00%39.20%39.20%0.00%4.80%2.17 Ours 59.68%32.64%32.64% 0.00%7.68%2.14 w F , w H , and w O denote the penalty weights for the following distance, human safety, and obstacle safety in the reward function, respectively. Since the primary conflict in pedes- trian crowd following lies between proximity to the target and safety among dynamic pedestrians, we fix δ O = 1.2 and w O across all configurations and focus the comparison on the following and human safety objectives. We configure three constraint profiles that shift the system’s behavioral priority: safety-conservative (δ F =4.0, δ H =3.2), balanced (δ F =3.6, δ H =3.6), and aggressive- following (δ F =3.2, δ H =4.0). As the profile shifts from safety- conservative to aggressive-following, AFD decreases from 2.54 to 2.23 and TLR drops from 10.08% to 3.20%, while human CR rises from 8.80% to 14.24%. Tightening δ F thus produces closer following, whereas tightening δ H yields safer behavior, each achieved at the expense of the other and reflecting predictable control over both objectives. The balanced profile sits between these extremes and achieves the highest SR of 78.08% and the lowest overall CR of 16.16%, offering the best compromise. These results show that each cost threshold maps directly and predictably onto its intended behavioral aspect, allowing the desired proximity- safety trade-off to be specified explicitly. In contrast, reward weight tuning produces inconsistent and counterintuitive outcomes. Doubling w H yields a neg- ligible 0.08% reduction in human CR, from 12.88% to 12.80%, failing to reflect the intended safety priority. More strikingly, doubling w F leaves AFD completely unchanged at 2.26 while simultaneously raising TLR from 8.08% to 8.72% and human CR from 12.88% to 15.44%, making both following and safety worse. These results confirm that cost limit tuning provides a direct and verifiable mapping from designer specification to policy behavior, which reward weight tuning fundamentally cannot offer. G. Out-of-Distribution Results We evaluate the performance of our policy and baselines under four representative OOD scenarios. 1) OOD Scenarios in the Corridor Environment: We con- struct a 26 m × 4.5 m corridor with 20 randomly initialized pedestrians, and the flow naturally splits into a bidirectional pedestrian flow. As shown in Table I, most baseline and ablation methods exhibit markedly higher human-collision rates in this narrow, crowded environment, whereas our method maintains a low human collision rate of 8.48%. It reduces the human-collision rate by 20.48% compared with OGM-HEIGHT and by 12.56% and 6.64% compared with the two ablation models, and also achieves higher success rates, exceeding OGM-HEIGHT by 33.52% and the two ablations by 12.24% and 6.80%. A representative episode is visualized in Fig. 3(c): at step 9, pedestrians advance from the opposite direction while those behind accelerate, yet our policy identifies a feasible escape path and maintains an appropriate following distance to the target. 2) OOD Scenarios Mixed with 15% Rushing Humans: In this setting, 15% of pedestrians are assigned faster walking speeds, with a minimum of 1.4 m/s and a maximum of 1.7 m/s, which corresponds to the typical speed of brisk walking. From the results in Table I, we observe that the SR of all methods drops. Nevertheless, our method continues to achieve the best overall performance. We also visualize an example in Fig. 3(d), showing that even under the influence of rushing pedestrians, our approach can follow the target effectively and safely. a)b) c)d) e)f) Step 4Step 18Step 26Step 36 Step 92Step 95Step 106Step 117 Expanding ACI Avoidance while maintaining following Escaped and continued following Sudden direction change Avoidance Expanding ACI Successfully maintained following for 30 seconds Collision Step 4 Stuck Escaped Narrow corridor with pedestrians oncoming and behind Bypassing via ACI uncertainty Step 10 Step 16Step 18 Step 9Step 13Step 25Step 77Step 85Step 90 Rushing Escaped Step 77Step 102Step 110 Bypass Continued following Step 10Step 17Step 28 Continued following Bypass Expanding ACI Fig. 3.Visualization of test results. Regular pedestrians are shown in blue, the target pedestrian is represented in red, and the robot is shown in orange. Light blue circles surrounding human agents represent spatial safety buffers derived from our uncertainty quantification method. (a) Ours successfully completes the human-following task in an in-distribution testing environment. (b) OGM-HEIGHT fails to complete the same episode; the light blue circles here indicate trajectory predictions rather than safety buffers. (c) Ours in an OOD narrow corridor environment. (d) Ours in an OOD environment with rushing humans. (e) Ours in an OOD environment using the SF pedestrian model. (f) Ours in an OOD environment with pedestrian groups. 3) OOD Scenarios with SF Pedestrian Model: In this setting, all pedestrian agents follow the social force (SF) model. By modeling social forces between agents, humans’ speeds change more markedly, and in denser scenarios, pedestrians tend to move closer together and create cohesive units. As expected, most methods experience a decline in success rate, which indicates limited adaptability to different behavioral models. We visualize a representative test case in Fig. 3(e), showing that our model is still able to avoid collisions and sustain robust following of the target human. 4) OOD Scenarios with Group Dynamics: In this setting, pedestrians move in groups in an open environment, which occupy larger spatial areas and require wider detours, so it also leads to a drop in performance across most methods. As visualized in Fig. 3(f), our robot can navigate around clustered groups while still successfully keeping up with the target human. Through extensive experiments under different conditions, we find that the policy is most likely to break in scenarios with dense and fast-moving crowds, as well as unpredictable pedestrian behaviors such as sudden turns and closely aligned group movements. H. Real-World Robot Experiments We deploy the policy trained in simulation on a ROS- MASTER X3 robot equipped with Mecanum wheels running ROS 2 after basic clipping and smoothing to validate real- world feasibility. Perception relies on a 2D RPLIDAR- A1 LiDAR, with human detection via the pre-trained DR- SPAAM model [36], tracking via SORT [37], and trajectory prediction via a CV predictor [35], over which our ACI module maintains online uncertainty bounds. Across 10 test trajectories in varied settings, the robot successfully completes 7, following the target human while navigating among moving pedestrians and static obstacles. The three failures stem from upstream perception limitations, such as missed detections of dynamic obstacles and inaccurate static boundary estimations, as well as the target human exceeding the robot’s maximum speed, while the navigation policy itself remains robust throughout all runs. Further demonstrations are available on our project page. V. CONCLUSION This paper addresses the proximity-safety conflict in human-following navigation through a multi-constraint RL framework that decomposes behavioral objectives into in- dependent cost constraints with behaviorally meaningful thresholds, enabling explicit and tunable trade-off control. We integrate prediction uncertainty into the policy net- work and cost formulation to improve safety under unpre- dictable pedestrian conditions. Extensive experiments across in-distribution and out-of-distribution scenarios demonstrate improvements over baselines, and direct comparison con- firms that cost threshold tuning provides more predictable behavioral control than reward weight tuning. Real-robot deployment validates real-world feasibility. Future work will explore adaptive policies that incorporate individual human preferences. REFERENCES [1] S. Li, K. Milligan, P. Blythe, Y. Zhang, S. Edwards, N. Palmarini, L. Corner, Y. Ji, F. Zhang, and A. Namdeo, “Exploring the role of human-following robots in supporting the mobility and wellbeing of older people,” Scientific reports, vol. 13, no. 1, p. 6512, 2023. [2] L. K ̈ astner, B. Fatloun, Z. Shen, D. Gawrisch, and J. Lambrecht, “Human-following and-guiding in crowded environments using se- mantic deep-reinforcement-learning for mobile service robots,” in 2022 International Conference on Robotics and Automation (ICRA). IEEE, 2022, p. 833–839. [3] S. Leisiazar, E. J. Park, A. Lim, and M. Chen, “An mcts-drl based obstacle and occlusion avoidance methodology in robotic follow- ahead applications,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2023, p. 221–228. [4] I. Gibbs and E. Candes, “Adaptive conformal inference under dis- tribution shift,” Advances in Neural Information Processing Systems, vol. 34, p. 1660–1672, 2021. [5] I. Gibbs and E. J. Cand ` es, “Conformal inference for online predic- tion with arbitrary distribution shifts,” Journal of Machine Learning Research, vol. 25, no. 162, p. 1–36, 2024. [6] L. Lindemann, M. Cleaveland, G. Shim, and G. J. Pappas, “Safe planning in dynamic environments using conformal prediction,” IEEE Robotics and Automation Letters, vol. 8, no. 8, p. 5116–5123, 2023. [7] H. Wang, A. H. Tan, and G. Nejat, “Navformer: A transformer ar- chitecture for robot target-driven navigation in unknown and dynamic environments,” IEEE Robotics and Automation Letters, vol. 9, no. 8, p. 6808–6815, 2024. [8] J. Ji, J. Zhou, B. Zhang, J. Dai, X. Pan, R. Sun, W. Huang, Y. Geng, M. Liu, and Y. Yang, “Omnisafe: An infrastructure for accelerating safe reinforcement learning research,” Journal of Machine Learning Research, vol. 25, no. 285, p. 1–6, 2024. [9] A. Ray, J. Achiam, and D. Amodei, “Benchmarking safe exploration in deep reinforcement learning,” arXiv preprint arXiv:1910.01708, 2019. [10] C. Chen, Y. Liu, S. Kreiss, and A. Alahi, “Crowd-robot interaction: Crowd-aware robot navigation with attention-based deep reinforce- ment learning,” in 2019 International Conference on Robotics and Automation (ICRA). IEEE, 2019, p. 6015–6022. [11] M. J. Islam, J. Hong, and J. Sattar, “Person-following by autonomous robots: A categorical overview,” The International Journal of Robotics Research, vol. 38, no. 14, p. 1581–1618, 2019. [12] A. Eirale, M. Martini, and M. Chiaberge, “Human following and guidance by autonomous mobile robots: A comprehensive review,” IEEE Access, 2025. [13] C. Scheidemann, L. Werner, V. Reijgwart, A. Cramariuc, J. Chomarat, J.-R. Chiu, R. Siegwart, and M. Hutter, “Obstacle-avoidant leader following with a quadruped robot,” in 2025 IEEE International Con- ference on Robotics and Automation (ICRA). IEEE, 2025, p. 1407– 1413. [14] H. Ye, J. Zhao, Y. Pan, W. Chen, L. He, and H. Zhang, “Robot person following under partial occlusion,” in 2023 IEEE International Conference on Robotics and Automation (ICRA).IEEE, 2023, p. 7591–7597. [15] S. Wang, J. Zhang, M. Li, J. Liu, A. Li, K. Wu, F. Zhong, J. Yu, Z. Zhang, and H. Wang, “Trackvla: Embodied visual tracking in the wild,” in Conference on Robot Learning.PMLR, 2025, p. 4139– 4164. [16] Y. Song, Q. Zhang, Z. Hu, and J. Liu, “Safe and robust human follow- ing for mobile robots based on self-avoidance mpc in crowded corridor scenarios,” in 2023 IEEE International Conference on Robotics and Biomimetics (ROBIO). IEEE, 2023, p. 1–6. [17] W. Situ, H. Ye, J. Peng, Y. Zhan, and H. Zhang, “Adap-rpf: Adaptive trajectory sampling for robot person following in dynamic crowded environments,” arXiv preprint arXiv:2510.11308, 2025. [18] Y. Kim, H. Oh, J. Lee, J. Choi, G. Ji, M. Jung, D. Youm, and J. Hwangbo, “Not only rewards but also constraints: Applications on legged robot locomotion,” IEEE Transactions on Robotics, vol. 40, p. 2984–3003, 2024. [19] J. Lee, L. Schroth, V. Klemm, M. Bjelonic, A. Reske, and M. Hut- ter, “Evaluation of constrained reinforcement learning algorithms for legged locomotion,” arXiv preprint arXiv:2309.15430, 2023. [20] J. Yao, X. Zhang, Y. Xia, Z. Wang, A. K. Roy-Chowdhury, and J. Li, “Towards generalizable safety in crowd navigation via conformal uncertainty handling,” in Conference on Robot Learning (CoRL), 2025. [21] L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe learning in robotics: From learning-based control to safe reinforcement learning,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 5, no. 1, p. 411–444, 2022. [22] J. Yao, X. Zhang, Y. Xia, Z. Wang, A. K. Roy-Chowdhury, and J. Li, “Sonic: Safe social navigation with adaptive conformal inference and constrained reinforcement learning,” arXiv preprint arXiv:2407.17460, 2024. [23] J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained policy op- timization,” in International conference on machine learning. PMLR, 2017, p. 22–31. [24] T.-Y. Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge, “Projection- based constrained policy optimization,” in International Conference on Learning Representations, 2020. [25] E. Altman, Constrained Markov decision processes. Routledge, 2021. [26] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017. [27] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel, “High- dimensional continuous control using generalized advantage estima- tion,” in International Conference on Learning Representations, 2016. [28] N. Nayakanti, R. Al-Rfou, A. Zhou, K. Goel, K. S. Refaat, and B. Sapp, “Wayformer: Motion forecasting via simple & efficient atten- tion networks,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, p. 2980–2987. [29] S. Liu, H. Xia, F. C. Pouria, K. Hong, N. Chakraborty, Z. Hu, J. Biswas, and K. Driggs-Campbell, “Height: Heterogeneous interac- tion graph transformer for robot navigation in crowded and constrained environments,” IEEE Transactions on Automation Science and Engi- neering, vol. 23, p. 1211–1230, 2026. [30] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximalpolicyoptimizationalgorithms,”arXivpreprint arXiv:1707.06347, 2017. [31] J. Van Den Berg, S. J. Guy, M. Lin, and D. Manocha, “Reciprocal n- body collision avoidance,” in Robotics research: the 14th international symposium ISRR. Springer, 2011, p. 3–19. [32] Z. Zhang, J. Yan, X. Kong, G. Zhai, and Y. Liu, “Efficient motion planning based on kinodynamic model for quadruped robots following persons in confined spaces,” IEEE/ASME Transactions on Mechatron- ics, vol. 26, no. 4, p. 1997–2006, 2021. [33] D. Dolgov, S. Thrun, M. Montemerlo, and J. Diebel, “Path planning for autonomous vehicles in unknown semi-structured environments,” The international journal of robotics research, vol. 29, no. 5, p. 485– 501, 2010. [34] Z. Huang, T. Ji, H. Zhang, F. C. Pouria, K. Driggs-Campbell, and R. Dong, “Interaction-aware conformal prediction for crowd naviga- tion,” arXiv preprint arXiv:2502.06221, 2025. [35] C. Sch ̈ oller, V. Aravantinos, F. Lay, and A. Knoll, “What the constant velocity model can teach us about pedestrian motion prediction,” IEEE Robotics and Automation Letters, vol. 5, no. 2, p. 1696–1703, 2020. [36] D. Jia, A. Hermans, and B. Leibe, “Dr-spaam: A spatial-attention and auto-regressive model for person detection in 2d range data,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, p. 10 270–10 277. [37] A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple online and realtime tracking,” in 2016 IEEE international conference on image processing (ICIP). IEEE, 2016, p. 3464–3468.