Paper deep dive
ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
Shashank Rao Marpally, Allan Wang, Atharva Ghotavadekar, Renato Alexandre Ribeiro, Nhat Le, Pilar Bachiller-Burgos, Pranav Goyal, Subham Agrawal, Yasuhiro Nitta, Howard Ziyu Han, Daeun Song, Masaki Kuribayashi, Kohei Uehara, Xiyue Wang, Yangzhe Kong, Duc M. Nguyen, Amirreza Payandeh, Gerardo Pérez-Gonzålez, Alejandro Torrejón-Harto, Jeeho Ahn, Tisha Jain, Andrew Stratton, Elvin Yang, Jorge de Heuvel, Nico Ostermann-Myrau, Sai Anudeep Sajja, Mithilya Raj, Daisuke Sato, Gaston Rouquette, Nikolas Martelaro, Maki Sugimoto, Hironobu Takagi, Chieko Asakawa, Maren Bennewitz, Aaron Steinfeld, Xuesu Xiao, Christoforos Mavrogiannis, Harold Soh
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 87%
Last extracted: 8/2/2026, 1:32:09 PM
Summary
The paper introduces ACME, a large-scale, multi-modal dataset for social robot navigation and pedestrian trajectory prediction. Collected across 8 sites in 5 countries using 7 different robot embodiments, ACME captures diverse cultural, geographical, and interaction contexts. It provides 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data, including 3D/2D scene features, odometry, and human-annotated trajectory labels. The dataset aims to address gaps in existing research by focusing on goal-driven social navigation, explicit robot-crowd interactions via speech, and challenging real-world scenarios.
Entities (18)
Relation Signals (14)
ACME â contains â Pedestrian Tracking Data
confidence 95% · providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data.
ACME â contains â Onboard Robot Data
confidence 95% · ACME is a large and diverse multi-modal dataset... providing 29.35 hours of onboard robot data
ACME â supports â Social Navigation
confidence 95% · ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research
ACME â supports â Pedestrian Trajectory Prediction
confidence 95% · To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides... human-annotated pedestrian trajectory labels.
ACME â collectedat â 8 Sites
confidence 90% · A large-scale data collection effort across 8 sites in 5 countries
ACME â collectedin â 5 Countries
confidence 90% · A large-scale data collection effort across 8 sites in 5 countries
National University of Singapore â contributedto â ACME
confidence 90% · Shashank Rao Marpally, School of Computing, National University of Singapore... ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset
ACME â uses â 7 Robot Embodiments
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in culture, geography, and human-robot interaction-factors that strongly shape appropriate social behavior. To address this gap, we introduce ACME: A Cross-cultural, Multi-Embodiment dataset for social navigation. A large-scale data collection effort across 8 sites in 5 countries, using 7 robot embodiments, ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research, providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data. Unlike prior datasets, it focuses on capturing goal-driven social navigation behavior in complex social scenarios with explicit robot-crowd interaction through robot speech. To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides 3D and 2D scene features, odometry, interaction information, and human-annotated pedestrian trajectory labels. We make ACME easy to use by providing both human-readable data for each sensor modality as well as raw binary data. Our qualitative and quantitative analyses show that our dataset captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets.
Tags
Links
- Source: https://arxiv.org/abs/2607.21964v1
- Canonical: https://arxiv.org/abs/2607.21964v1
Trouble viewing inline? Open PDF directly â
Full Text
110,038 characters extracted from source content.
Expand or collapse full text
Shashank Rao Marpally, School of Computing, National University of Singapore, 13 Computing Dr, Singapore 117417, Email: smarpall@comp.nus.edu.sg ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset Shashank Rao Marpally 1*1*affiliationmark: Allan Wang 2*2*affiliationmark: Atharva Ghotavadekar 11affiliationmark: Renato Alexandre Ribeiro 22affiliationmark: Nhat Le 33affiliationmark: Pilar Bachiller-Burgos 44affiliationmark: Pranav Goyal 55affiliationmark: Subham Agrawal 66affiliationmark: Yasuhiro Nitta 77affiliationmark: Howard Ziyu Han 88affiliationmark: Daeun Song 3, 103, 10affiliationmark: Masaki Kuribayashi 22affiliationmark: Kohei Uehara 22affiliationmark: Xiyue Wang 22affiliationmark: Yangzhe Kong 33affiliationmark: Duc M. Nguyen 33affiliationmark: Amirreza Payandeh 33affiliationmark: Gerardo PĂ©rez-GonzĂĄlez 44affiliationmark: Alejandro TorrejĂłn-Harto 44affiliationmark: Jeeho Ahn 55affiliationmark: Tisha Jain 55affiliationmark: Andrew Stratton 55affiliationmark: Elvin Yang 55affiliationmark: Jorge de Heuvel 66affiliationmark: Nico Ostermann-Myrau 66affiliationmark: Sai Anudeep Sajja 66affiliationmark: Mithilya Raj 88affiliationmark: Daisuke Sato 88affiliationmark: Gaston Rouquette 77affiliationmark: Nikolas Martelaro 88affiliationmark: Maki Sugimoto 77affiliationmark: Hironobu Takagi 22affiliationmark: Chieko Asakawa 22affiliationmark: Maren Bennewitz 66affiliationmark: Aaron Steinfeld 88affiliationmark: Xuesu Xiao 33affiliationmark: Christoforos Mavrogiannis 55affiliationmark: Harold Soh 1,91,9affiliationmark: 11affiliationmark: School of Computing, National University of Singapore, 22affiliationmark: Miraikan - The National Museum of Emerging Science and Innovation, 33affiliationmark: George Mason University, 44affiliationmark: Universidad de Extremadura, 55affiliationmark: University of Michigan, 66affiliationmark: University of Bonn, 77affiliationmark: Keio University, 88affiliationmark: Carnegie Mellon University, 99affiliationmark: NUS Smart Systems Institute, 1010affiliationmark: Ewha Womans University **affiliationmark: These authors contributed equally to this work. Abstract Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in culture, geography, and human-robot interaction-factors that strongly shape appropriate social behavior. To address this gap, we introduce ACME: A Cross-cultural, Multi-Embodiment dataset for social navigation. A large-scale data collection effort across 8 sites in 5 countries, using 7 robot embodiments, ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research, providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data. Unlike prior datasets, it focuses on capturing goal-driven social navigation behavior in complex social scenarios with explicit robot-crowd interaction through robot speech. To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides 3D and 2D scene features, odometry, interaction information, and human-annotated pedestrian trajectory labels. We make ACME easy to use by providing both human-readable data for each sensor modality as well as raw binary data. Our qualitative and quantitative analyses show that our dataset captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets. keywords: Social Navigation, Robot Safety, Human-Robot Interaction, Pedestrian Trajectory Prediction Preprint notice. This arXiv preprint is an exact copy of the manuscript currently under review at The International Journal of Robotics Research (IJRR). This version will not be updated during the review process. Figure 1: ACME comprises of data collected by 8 teams across 5 countries on 7 embodiments across diverse crowd scenarios. 1 Introduction Social robot navigation addresses the challenge of enabling autonomous agents to move efficiently and naturally within dynamic human environments. Unlike conventional navigation, which primarily focuses on reaching goals without collisions, social navigation emphasizes compliance with cultural and context-dependent norms that humans implicitly follow in shared spaces (Singamaneni et al., 2024). This capability is increasingly critical as robots transition from controlled industrial settings to everyday human-centric applications such as domestic service, food delivery, and security robots. Navigating in the presence of humans is also a challenge faced in the autonomous vehicles (AV) domain, and many recent works have focused on achieving safe autonomous navigation (Muhammad et al., 2020). However, the AV setting benefits from widely adopted and well-defined driving rules, driver behavior expectations, and structure in the environment, such as lane markings, traffic lights, and right-of-way laws. Mobile service robots, in contrast, enjoy no such scaffolding, and must reason about and comply with informal, context-sensitive, and often implicit guidelines (Mavrogiannis et al., 2023). These include cultural and societal conventions such as how close is âtoo closeâ when passing someone, which side of a corridor to stay on, or how to yield in tight spaces. These rules often vary across cultures and scenarios and are rarely explicitly documented. Considerable prior work has sought to formalize the norms governing various aspects of human social navigation, addressing proxemics (Kirby, 2010), social signals (Takayama and Pantofaru, 2009; Mead and Mataric, 2012), intent communication & legibility (Taylor et al., 2022) along with social group dynamics (Wang et al., 2022a). Alongside modeling efforts, complementary work has proposed an evaluation infrastructure for socially compliant navigation, ranging from canonical scenario protocols (Pirk et al., 2022) to comprehensive guidelines spanning metrics, benchmarks, and datasets (Francis et al., 2025). Yet, as critically surveyed by (Mavrogiannis et al., 2023), fundamental challenges persist across motion planning, behavior design, and evaluation. A core reason is that social norms often vary with culture, robot embodiment, and situational context, and multiple norms may apply simultaneouslyâleading to conflicting constraints that are challenging to identify, model, and reconcile owing to their context dependence. This difficulty in identifying and quantifying these principles and norms of âcorrectâ social behavior has inspired a perspective shift towards data-driven methods. Specifically, recent works have focused on imitation learning (Karnan et al., 2022a; Hirose et al., 2023; Han et al., 2025), preference learning (Keselman et al., 2023; Wang et al., 2023; De Heuvel et al., 2023), reinforcement Learning (Liu et al., 2023; Xie and Dames, 2023; Zhu et al., 2026), and lifelong learning (Narasimhan et al., 2024) approaches for learning social behaviors from expert robot and pedestrian data. Such methods rely on large, diverse, multi-modal datasets of real-world humanârobot navigation for model training, and existing datasets fall short along the axes that matter the most for generalization: diversity in robot embodiments, cultural and geographical contexts, humanârobot interaction scenarios, and challenging edge cases (Raj et al., 2024). Closely tied to the challenge of socially compliant navigation is the task of understanding and predicting human motion (Rudenko et al., 2020b). Social navigation and human motion prediction share the same core problem of understanding human motion. Recently, social navigation models have increasingly integrated trajectory prediction as an important component in their framework (Wang et al., 2022a; Liu et al., 2023; Samavi et al., 2025). Many datasets have been collected for training and benchmarking human motion prediction models (Table 1). However, while it is well understood that people from different cultures and environmental contexts exhibit distinct movement patterns and social behaviors, such variance is rarely captured in a single, large-scale dataset. As a result, most predictive models are trained and evaluated on narrow distributions that do not generalize well across culturally or geographically diverse environments, as demonstrated in our benchmark evaluation. To address these gaps, we propose the ACME dataset with the following contributions: âą To the best of our knowledge, we release the largest (w.r.t duration) and most diverse (w.r.t geographical location and robot embodiment) human-demonstrated social navigation and pedestrian trajectory prediction dataset from 8 locations across 5 countries and 7 robot embodiments (Fig 1). âą We release the largest human-labeled Birdâs-Eye View pedestrian trajectory prediction dataset, including camera calibration information to transform these trajectories into the metric space. âą We analyze and compare ACME to prior social navigation and pedestrian trajectory prediction datasets with respect to scenario complexity and pedestrian trajectory characteristics. âą Based on ACME, we benchmark and compare the performance of SOTA vision navigation and pedestrian trajectory prediction models. âą For ease of use, we release human-readable synchronized multi-sensor data, raw ROS2 bags, birdâs eye view video, and human-verified pedestrian trajectories. 2 Related Work 2.0.1 Data-Driven Social Navigation The landscape of social navigation research is characterized by an extensive variety of approaches (Singamaneni et al., 2024). Early approaches centered around treating humans as non-reactive obstacles. Social-Force Model (Helbing and Molnar, 1995) based approaches like Svenstrup et al. (2010) use proxemics to guide potential-field based planners, thus focusing on spatial social norms. Concurrently, other methods treated the social navigation problem as one of dynamic obstacle avoidance (Van Den Berg et al., 2011) with constant-velocity models for pedestrians. Reinforcement learning-based approaches utilized such simplified human models to learn navigation in crowds Chen et al. (2017, 2019); Xie and Dames (2023). More recently, the advent of Vision-Language Models (VLMs) (Song et al., 2024; Xiao et al., 2026; Kong et al., 2025; Payandeh et al., 2025) and foundational models for navigation (Shah et al., 2023; Cheng et al., 2024) has marked a paradigm shift, towards model-free data-driven navigation policy learning. However, the generalization of these modern approaches is fundamentally constrained by the limitations of existing datasets (Karnan et al., 2022a; Nguyen et al., 2023; Rudenko et al., 2020a) which do not yet capture diverse in-the-wild social-navigation scenes across cultures, geographical locations, and different robot embodiments in a unified fashion. 2.0.2 Human Motion Prediction Beginning with SocialLSTM (Alahi et al., 2016a), which standardized trajectory prediction benchmarking, interest in human motion prediction has surged. Beyond early pooling-based (Alahi et al., 2016b; Gupta et al., 2018) and graph-centric approaches (Mohamed et al., 2020), recent methods increasingly leverage transformers and diffusion. For instance, transformer architectures unify interaction reasoning and multimodality within a single encoderâdecoder pipeline (Shi et al., 2023), while diffusion models incorporate social physics or bidirectional consistency to better capture multi-modal futures (Chen et al., 2024; Li et al., 2023; Panigrahi et al., 2023). Most recently, a flow-matching-based distillation model has emerged as the new state-of-the-art method (Fu et al., 2025). Despite rapid state-of-the-art progress, many studies still benchmark mainly on the ETH/UCY (Pellegrini et al., 2009a; Lerner et al., 2007) datasets, which are comparatively small-scale and predominantly outdoor, raising robustness and generalization concerns. Human motion understanding and prediction have also been increasingly relevant to social navigation. It is typically integrated into social navigation models in a two-stage fashion: prediction models first predict pedestriansâ future trajectories, and planning modules drive the robot by modeling the predicted trajectories into costs, rewards, or counterfactual information (Stratton et al., 2026b; Wang et al., 2022a; Liu et al., 2023; Hirose et al., 2023). A few works have recently emerged that directly leverage human motion prediction models to directly sample trajectories for robots (Samavi et al., 2025), effectively training social navigation in a supervised fashion, yet these modelsâ effectiveness is limited in the real world due to the limited sizes of suitable existing datasets. 2.0.3 Social Navigation and Human Motion Datasets Capturing nuanced variations in pedestrian social behavior, such as those associated with cultural context or the embodiment of nearby navigating agents, requires datasets that span a broad range of interaction settings, agent appearances, and environmental configurations. Table 7 provides an overview of the distinguishing features between ACME and prior datasets. Core challenges towards learning social navigation behaviors are addressing the intractability of âcoupledâ approaches, the difficulty of designing context-aware social navigation policies, as well as the lack of robust pedestrian behavior prediction models (Mavrogiannis et al., 2023). Central to addressing these challenges is curating large and diverse datasets to provide a foundation for training models that are predictive, context-aware, and capable of robust generalization to unseen social norms and physical configurations encountered in the real world. Alongside these characteristics, a notable gap is the absence of datasets that incorporate diverse robot embodiments and explicit crowd interaction, a natural and effective tool that humans use to interactively and cooperatively navigate crowded spaces. Trajectory prediction has emerged as an important research direction. However, its focus has traditionally been computer vision related (Huang et al., 2025). To support robust generalization, it is important to curate data collected in diverse contexts and social norms. To support integration into downstream social navigation models, the dataset should be collected in the real world, its trajectory annotations should be verified instead of relying solely on (potentially error-prone) automated tracking, and most importantly, its pedestrian trajectory coordinates should be grounded in metric coordinates instead of pixel coordinates in the image space. Table 1 provides an overview of the characteristics of prior datasets that support trajectory prediction. Table 5 further provides dataset size comparisons with prior datasets that satisfy all three key characteristics. Table 1: Characteristics of all pedestrian datasets that contain pedestrian trajectory annotation ETH UCY Edinburgh VIRAT Town Centre Grand Central Pellegrini et al. (2009a) Lerner et al. (2007) Majecka (2009) Oh et al. (2011) Benfold and Reid (2011) Zhou et al. (2012) Location Outdoor Outdoor Outdoor Outdoor Outdoor Indoor # Countries 1 1 1 1 1 1 Real-World Data â â â â â â Verified Trajectories â â â â Metric Coordinates â â â â CFF SDD L-CAS WildTrack JRDB ATC Alahi et al. (2014) Robicquet et al. (2016) Yan et al. (2017) Chavdarova et al. (2018) Martin-Martin et al. (2021) BrĆĄÄiÄ et al. (2013) Location Indoor Outdoor Indoor Outdoor Both Indoor # Countries 1 1 1 1 1 1 Real-World Data â â â â â â Verified Trajectories â â â â Metric Coordinates â â â â THĂR Crowd-Bot TBD SiT THĂR-MAGNI Bi3 ACME Rudenko et al. (2020a) Paez-Granados et al. (2022a) Wang et al. (2024a) Bae et al. (2023) Schreiter et al. (2025) Stratton et al. (2026a) (Ours) Location Indoor Outdoor Indoor Both Indoor Indoor Both # Countries 1 1 1 1 1 2 5 Real-World Data â â â â Verified Trajectories â â â â â â Metric Coordinates â â â â â â â 3 Motivation for ACME ACME is motivated primarily by the observation that contextual factors heavily influence what constitutes acceptable social navigation behavior for robots. Francis et al. (2025) frames eight principles for evaluating social compliance, including âcontextual appropriatenessâ â socially-appropriate robot behavior must consider the various facets of context like culture, environment, task, and interpersonal context. For example, training robots to stay a fixed minimum distance away from all pedestrians might generate favorable behavior in simpler, sparse scenarios while likely resulting in the âfreezing robotâ problem (Trautman and Krause, 2010; Trautman et al., 2015) in overly crowded or geometrically constrained scenarios. Similarly, robots trained to always move to one side of a hallway may appear socially compliant in some countries and noncompliant in others. Additionally, within the same geographical location, different geometrical layouts of navigation environments, crowd densities, and robot embodiments necessitate adaptive navigation strategiesâwhat is perceived as safe and appropriate in a spacious corridor with sparse pedestrian traffic for a vacuum cleaning robot may be interpreted as inefficient or even obstructive in a narrow, densely populated walkway for a robot dog. These nuances highlight the need for datasets that capture multi-cultural and multi-embodiment data in diverse environments, enabling the development of context- and embodiment-aware navigation policies. This motivation also informs our data-collection strategy: rather than treating diversity as an incidental property of the dataset, we explicitly align ACME with prior recommendations for social-navigation dataset curation (Francis et al., 2025), including broad but resource-aware coverage, well-sampled scenarios, real robot behavior, diverse robot platforms, and systematic annotations. We now highlight the motivations behind the unique features of ACME. 3.1 Capturing Cross-Cultural Social Norms Cultural differences, personal preferences, and environmental factors heavily influence pedestrian motion characteristics. For example, a large-scale study (Sorokowska et al. (2017)) found significant variance in the preferred social, personal, and intimate distance for interaction across country, age, and gender. It is also well known that the direction of traffic flow varies across countries. For example, in the United States, Spain, and Germany, people prefer to walk on the right side of a walkway, while in Singapore, Japan, and India, the left side is preferred. Intending to train robots and human motion prediction models to account for such socio-cultural contexts, we curate data collected across 5 countries in 3 continents (8 different data collection teams; 7 of the 8 locations are university campuses, while Miraikan is a science museum). This distributed data collection strategy helps capture differences in navigation and communication styles of both pedestrians and the robot, thus facilitating learning of culturally appropriate behaviors. 3.2 Diversity in Robot Morphology Social robotics studies have extensively documented the effect of robot morphology on the perception, likability, and nature of interaction of the robot with a user (Bradwell et al., 2021; Haring et al., 2016). For example, the formâfunction attribution bias (FFAB) describes how users may infer a robotâs capabilities, intelligence, or social competence from its physical appearance rather than from its actual functional abilities (Haring et al., 2018). This suggests that robot embodiment can shape human expectations even before any interaction occurs. Different robot morphologies also differ in navigational affordances and offer an opportunity to capture the variance in the behaviors of pedestrians towards different robots operating in the wild. Our dataset utilized 5 visually distinct robots with different navigation styles and sensor setups to gather data (Table 2): namely, 2 quadrupeds (NUS, UBonn), 2 suitcase-style differential drive robots physically guided by human operators (CMU, Miraikan, Keio), 1 rover-style differential drive robot (GMU), and 2 mobile service robots (UMich, UEx). Despite this variance in setups, we maintain consistency across these subsets of our data by maintaining a minimum set of sensor data modalities and common teleoperation and data-recording guidelines. 3.3 Curating Local Social Behaviors We distinguish ACME from other valuable datasets (Karnan et al., 2022a; Nguyen et al., 2023) by focusing on the quality of data, and targeting our data collection effort towards capturing challenging social navigation scenarios in-the-wild. Particularly, we ensure that in each trajectory the robot encounters at least one human, while preferring scenarios with varying levels of crowds and challenging location configurations like blind corners, intersections, and narrow corridors. Our data collection teams sought out locations based on the guidelines of Francis et al. (2025) and timings that coincided to provide maximum crowd encounters with the robot. In order to curate useful interactions and scenarios, we additionally augment the teleoperation procedure with guidelines and tools to facilitate the collection of goal-oriented, high-quality robot trajectory data. 3.4 Social Navigation with Robot Speech Pedestrians often use verbal and non-verbal cues when asking for directions, alerting other social actors of their presence, or apologizing when causing disruptions to other pedestrians. Learning such communication behavior is desirable for robots to adapt to various levels of crowds and navigate in a trustworthy and safe manner (Che et al., 2020; Kannan et al., 2021; Hart et al., 2020). However, there is a lack of real-world datasets focused on learning such communication policies from expert behavior. We thus include a set of phrases that the robot can utter and record when and what was spoken by the robot. This places additional constraints on the teleoperator and robot platform, thus only a subset of ACME includes robot speech data, namely, the NUS, UEx, Miraikan, and Keio subsets. Paired with scene information obtained from LiDAR and FPV RGB, utterance data opens avenues for teaching robots to communicate effectively to proactively influence the crowd behavior. 3.5 Understanding Context Conditioned Human Behavior Besides the data from robots manually operated by humans, social navigation has increasingly integrated trajectory prediction models as part of the broader navigation framework. Some works predict trajectories first, and then plan robot controls by assigning costs to the predicted trajectories (Wang et al., 2022a; Poddar et al., 2023). Other works integrate trajectory prediction outputs as additional rewards or context (Liu et al., 2023; Hirose et al., 2023). Explicitly, more accurate trajectory prediction may lead to better downstream navigation performance. Implicitly, studying how humans move in the same environment also offers valuable insights into how robots should move. Although robot data are of higher quality and realism, containing goal annotations, and are not anthropomorphic (Epley et al., 2007), human behavior data are much richer and greater in quantity. In the ACME dataset, we label human trajectories from the overhead Birdâs-Eye View (BEV) cameras and transform the pixel coordinates into metric coordinates via calibrated sensor information. By doing so, we obtain data on how humans navigate given their current surrounding context. Because multiple humans are often present in our dataset sessions, and their combined trajectories cover wider areas in the dataset environments, human trajectory data captures ground truth human navigation behavior more comprehensively, given their greater quantity and diversity. Recent works have emerged, such as SICNav-Diffusion (Samavi et al., 2025), that directly leverage trajectory prediction as a trajectory sampler. Although their effectiveness in the real world is capped by the limited size and variety of existing datasets, they revealed that a comprehensive human motion dataset may enable additional supervised research approaches to social navigation. 4 The ACME Dataset In this section, we describe the methodology and nature of the data captured in ACME 111All software developed in this section will be open-sourced.. An overview of the data duration, hardware, sensors, and locations for each data collection team is summarized in Table 2. Table 2: Overview of data characteristics and hardware setup for each subset of ACME. Information Subset National University of Singapore (NUS) Museum of Emerging Science and Innovation (Miraikan) Carnegie Mellon University (CMU) Keio University (Keio) University of Michigan (UMich) University of Extremadura (UEx) George Mason University (GMU) University of Bonn (UBonn) Robot Platform Unitree GO2111Unitree GO2 EDU - https://w.unitree.com/go2 Suitcase robot222Cabot robot - Kuribayashi et al. (2025) https://w.miraikan.jst.go.jp/en/lab/AIsuitcase/ Suitcase robot222Cabot robot - Kuribayashi et al. (2025) https://w.miraikan.jst.go.jp/en/lab/AIsuitcase/ Suitcase robot222Cabot robot - Kuribayashi et al. (2025) https://w.miraikan.jst.go.jp/en/lab/AIsuitcase/ Stretch 2333Shadow robot TorrejĂłn et al. (2024) - https://robolab.unex.es/en/robots/shadow/ Shadow Robot444Stretch - https://hello-robot.com/stretch-2-product Scout Mini555AgileX Scout mini - https://global.agilex.ai/products/scout-mini Unitree GO1666Unitree GO1 - https://w.unitree.com/go1 Sensors Realsense D435i, Hesai XT-16 Realsense D435i, Velodyne VLP-16 Zed 2, Velodyne VLP-16 Realsense D435i, Velodyne VLP-16 Realsense D435i, Velodyne VLP-16 Zed 2i, Ricoh Theta z1, Robosense BPearl & Helios Zed 2, Velodyne VLP-16 GoPro, Velodyne VLP-16 Pose Estimate LiDAR-Odometry LiDAR Localization LiDAR Localization LiDAR Localization Wheel Odometry Visual-Inertial Odometry, LiDAR localization Wheel Odometry LiDAR-Odometry BEV Camera GoPro Hero 9 GoPro Hero 12 GoPro Hero 10 GoPro Hero 12 GoPro Hero 13 - Logitech C920 GoPro Hero 12 Map-based Localization No Yes No Yes No No No No Data Duration (hrs) 6.54 7.53 1.56 1.35 3.2 3.29 3.45 2.43 BEV Data Duration (hrs) 10.01 12.39 2.65 3.84 5.10 - 2.30 7.27 BEV # Trajectories 31848 21322 1987 5095 3728 - 3186 4892 4.1 Hardware and Software Setup To facilitate consistency and quality across the dataset while accommodating differences in robot platforms and sensing modalities, we adopt a set of common guidelines for data collection and curation. We ensure that all trajectories comprise of locally accurate robot poses (via Odometry), Egocentric RGB Camera feeds (for semantic scene understanding), 3D LiDAR point clouds (for spatial reasoning), extrinsic and intrinsic sensor calibration (for sensor-fusion). A subset of the dataset also includes Robot Speech (timing and utterance) for crowd interaction. Onboard sensor data was collected in ROS2 (Macenski et al., 2022) bag format, then post-processed for anonymization and downsampled to provide human-readable data. Table 2 describes the hardware for data collection in each location. Additionally, several data-collection teams also collected additional sensor information, for example, accurate robot localization with pre-built maps (Miraikan, CMU, and Keio University teams), and egocentric 360 RGB images (GMU and UEx teams). Each subset of data is also accompanied by sensor calibration for robot-onboard sensors performed with off-the-shelf techniques, allowing for multi-modal scene understanding. In addition to robot data, NUS, Miraikan, CMU, Keio, U-Mich, GMU, and UBonn subsets have overhead BEV cameras set up overlooking the data collection areas. The BEV cameras collected human trajectory data from a less constrained perspective. Since humans in the scene are at least partially observable from the overhead cameras, automated tracking methods were used to assist human trajectory annotations (Wang et al., 2024a). Figure 2: Teleoperation interface used to collect data in NUS where the teleoperator can relay robot speech commands and mark start and end of trajectory segments during data collection. Figure 3: Data filtering with Event Markings 4.1.1 Teleoperation Guidelines (a) Distribution of individual trajectory durations. (b) Distribution of distance traversed in individual trajectories. Figure 4: Compared to prior Social Navigation Datasets that collect unconstrained long trajectories, ACME trajectories are more localized and goal-oriented. Recognizing the limitations of prior data collection methodologies, we established the following teleoperator guidelines for maximizing the utility of the dataset and capturing natural human-robot-interactions in the wild: âą Stay as far away from the robot as safely possible: To emulate âWizard of Ozâ conditions, teleoperators are advised to stay away from the robot and provide pedestrians the illusion that the robot is navigating autonomously. (Note: this does not apply to the human-operated CMU, Miraikan, and Keio suitcase robots) âą Use speech when logical: We use a restricted set of speech commands relayed through a speaker mounted on the robot, which can be triggered through the teleoperation interface (or via the teleoperator verbally in case of the suitcase robots) to standardize behaviors across the dataset. The set of verbal commands is Excuse Meâ, âAttention, Robot Hereâ, âPlease Give Wayâ. An example of the interface (used by the NUS team) is shown in Fig 2. âą Fix a navigation goal apriori to starting data collection and record âBeginâ, âDiscardâ, âEndâ signals for marking trajectory segments: To ensure each trajectory describes a logically appropriate path that we expect a robot to follow to reach a predefined navigation goal, teleoperators are asked to fix (implicitly or explicitly via a marking in the environment) a location target for navigation before collecting a trajectory. Each teleoperator also marks events on the respective teleoperation interface that correspond to the start and end of a trajectory, as well as situations where the trajectory must be discarded (for example, if the teleoperator is approached by a pedestrian for conversation or the robot behaves erratically due to hardware issues, etc.). The final dataset is obtained by filtering the raw data with the scheme shown in Fig 3. The difference in data-collection strategy between ACME and prior social navigation datasets like SCAND (Karnan et al., 2022a) and MuSoHu (Nguyen et al., 2023) is immediately apparent from the distribution of individual trajectory duration and traversal distance (Fig 4): while SCAND and MuSoHu are largely in-the-wild datasets where the data-collection agents wander through locations for long durations and distances, ACME focuses on capturing short goal-directed scenarios. 4.1.2 Transformation from BEV Image Plane to Ground Plane In order to obtain ground plane metric coordinates after the human trajectories were annotated in the BEV camera frames, 2D to 2D homography transformations were estimated. For Miraikan, Keio, GMU, and UBonn, this is done via a poster-sized AprilTag (Olson, 2011) placed on the ground plane briefly during data collection (Fig. 5). Given that the tags were laid flat on the ground plane, homographies can be easily obtained via tag detection and pose estimation. While we acknowledge the inherent inaccuracies of tag-based homography estimation, these methods represent a pragmatic compromise between the complexity of full SLAM-based map building in dynamic environments and having no spatial calibration at all. For CMU and U-Mich, ground truth stationary keypoints (e.g., tile intersections) were selected and measured at each location. The keypointsâ corresponding pixel locations were obtained in the BEV images, and the resulting point correspondences were used to estimate homographies. For NUS, the poster-sized AprilTags were held upright and can be seen by both the robot camera and the BEV cameras. However, simply assuming the tags to be perfectly vertical resulted in homography estimation errors. Instead, we first obtained the robotâs pose with respect to the ground plane-based world frame at the time when it sees the upright tag. We then estimated the BEV camerasâ poses in the world frame via tag detection at the synchronized time from both the robot and the BEV cameras. Homographies were then extracted from the camerasâ world frame poses. Figure 5: Position Synchronization with April Tags. The tag is visible to both the robot (suitcase) and the BEV camera. 4.2 Post Processing 4.2.1 Data Anonymization Due to our data collection teams being spread out around the globe, each teamâs data had to abide by the respective Institutional Review Boardâs privacy guidelines for open-sourcing data collected by onboard and BEV sensors. Notably, there were different levels of anonymization required for the humans captured in RGB data across different institutions. These include (in increasing order of information loss) as shown in Fig 6: âą No anonymization (GMU, CMU) âą Face blurring (NUS, UMich, Keio) âą Full body segmentation (UEx, Miraikan, UBonn) (a) No Anonymization (GMU) (b) Face Blurring (NUS) (c) Full Body Segmentation (UEx) Figure 6: Different types of Data Anonymization across ACME specified by the IRB requirements for each institution. In order to retain the lost information, pedestrian tracking information used to anonymize the data (bounding box, segmentation mask, and keypoints) is provided. Although anonymization is essential to preserve privacy, it leads to a loss of identifying features, making it difficult to use the data with off-the-shelf vision models. To compensate for this information loss, we release bounding box, segmentation mask, and pose estimation (keypoint) detections for the anonymized pedestrians using Yolov8 (Jocher et al. (2023)) and ByteTrack222ByteTrack Github - https://github.com/FoundationVision/ByteTrack (Zhang et al., 2022). Figure 7: Application interface for human verification process. It contains a media player and multiple options to correct tracking errors manually. 4.2.2 Pedestrian Trajectory Tracking and Correction in BEV - Prior work on trajectory datasets widely employs semi-automatic annotations, particularly on pedestrian datasetsâmany of these datasets are relatively small in scale (Pellegrini et al., 2009b; Chavdarova et al., 2018; MartĂn-MartĂn et al., 2023) or lack annotations of the pedestrian surroundings (Karnan et al., 2022b; Paez-Granados et al., 2022b)âhowever, recent efforts have begun addressing these gaps (Wang et al., 2024a). Similar to Wang et al. (2024a), we first performed automated tracking of pedestrians from the BEV cameras using ByteTrack. To ensure the annotation quality of our large-scale dataset, human verification of the tracked trajectories was performed at 10Hz. Based on the features of existing tools Wang et al. (2024a) that streamline the human verification process, we designed our own annotation tool to simplify frame-by-frame human annotation. Our human verification tool (Fig 7) includes a media player that allows users to view videos with automatically generated trajectories by ByteTrack. If an error is detected by the human annotator users, users can correct it using the available editing options: âą Relabel: Used to redraw a trajectory from a specific frame when ByteTrack assigns incorrect or imprecise paths. âą Missing: Used to manually add a trajectory when ByteTrack fails to detect a pedestrian. âą Break: Used when ByteTrack incorrectly assigns the same trajectory to two different pedestrians. âą Join: Used to merge two trajectories into a single continuous path. âą Delete: Used to remove a trajectory from the current frame onward. âą Disentangle: Used to swap back two overlapping trajectories after a specific frame to correct misassigned identities. âą Undo: Reverts the most recent trajectory modification, working backwards through the sessionâs history. Figure 8: Distribution of Scenarios captured by ACME. Scenarios dependent on the type of location where the robot and pedestrians navigate are Location-specific scenario tags while those characterized by the type of relative motion between the pedestrians and the robot are Pedestrian-centric Scenario Tags 4.2.3 Scenario Tagging of Robot Data To enable filtering the dataset for specific scenarios, we additionally developed another annotation tool (a modification of VIA (Dutta and Zisserman, 2019)) that allows annotators to manually tag robot trajectories for scenarios that occur within them. We identify scenarios based on previous works (Francis et al. (2025); Karnan et al. (2022b); Nguyen et al. (2023)) and include both location-based and pedestrian-based scenario tags. Fig. 8 shows the distribution of scenarios in ACME, and the description of each tag is provided in Table 3. Table 3: Tags used to characterize each trajectory in ACME Scenario Tag Tag Description # Tags Narrow Corridors Corridors that are 1-2 pedestrian wide 864 Wide Corridors Corridors that are 4-5 pedestrian wide 1468 Intersections Intersection of multiple corridors/walkways 713 Against Traffic Against oncoming pedestrian traffic 2234 With Traffic Alongside pedestrian traffic 1196 Passing Conversational Groups Past a group of 2 or more people that are talking amongst themselves 1854 Blind corner Past a corner which the robot cannot see past 536 Open Area Areas with minimal spatial constraints 1781 Navigating through large crowds Through large unstructured crowds 352 Entry/Exit Through doorways/Elevators 306 Line Formation (like queues) Across people waiting in a line 328 Overtaking Pedestrian(s) Overtaking a person or groups of people 251 4.2.4 Semantic Tagging of BEV Scenes Figure 9: Example BEV tags overlaid on top of pedestrian trajectories. Each color represents a different semantic tag. The yellow areas are open spaces, the blue areas are wide corridors, the red areas are narrow corridors, the green areas are intersection areas, and the orange areas are blind corner areas. The colored lines show the annotated trajectories in the corresponding recording session. BEV data offers an overhead view of the data collection site, and similar to the SiT dataset (Bae et al., 2023), we perform semantic tagging to characterize the various location types within a single scene. As shown in Fig. 9, data collection areas are segmented and categorized using the same labels as the location tags in the robot data scenario tagging (i.e., ânarrow corridorsâ, âwide corridorsâ, âintersectionsâ, âblind cornersâ, âopen spacesâ, and âentries/exitsâ). These semantic annotations can be used to analyze pedestrian trajectories within and across different scene contexts and may be used as auxiliary information when training models that incorporate human motion behavior. 5 Using the ACME Dataset 5.1 Onboard Robot Data The ACME datasetâs onboard robot data is provided in two complementary formats to support a wide range of research applications: raw ROS2 bag files for full-fidelity playback and processing, and pre-processed, uniformly sampled data in human-readable format at 4 Hz. The preprocessed dataset is organized hierarchically: each of the contributing teams has a dedicated folder containing all trajectories collected by that team. Within each teamâs directory, data is separated into folders, each corresponding to a single recorded trajectory. These are obtained by segmenting a continuous recording session with âBeginâ, âEndâ, and âDiscardâ signals as shown in the Fig 3. The folder names encode metadata such as location and recording date. Each trajectory folder contains a set of subfolders, one for each sensor modality. Data across these subfolders is time-synchronized and obtained by sampling the ROS2 bag files at a 4Hz rate. Files are named according to the index of their timestamp in the sequence of sampled timestamps (at 44Hz) in index.extension format to ease accessing cross-modality data with time synchronization. For each trajectory, the following data streams are available, provided in separate files/folders: âą Odometry: Provided in the form of CSV files containing 10 headers corresponding to the timestamp, X-Y position with respect to the odom frame, yaw, and the velocities along these dimensions. âą Egocentric RGB images: From the robotâs onboard cameras, stored as JPEG files. Additionally, we also provide pedestrian bounding box detections corresponding to the RGB image frames in JSON format. âą LiDAR point clouds: Point Clouds from onboard LiDARs are provided in PCD format for easy processing and visualization. âą Robot Speech: Instances where robot speech was used are enumerated with a 3-column CSV file for each data subset containing the timestamp, trajectory name, and string corresponding to the utterance from the robotâs speaker. Along with this, we provide the full ROS2 bag files containing all recorded topics in their original message formats. Each bag file corresponds to a single trajectory and is provided in a separate folder with a similar folder structure as the processed data. We also provide a CSV file containing the trajectory name and the corresponding scenario tags as well as static transforms, extrinsics, and intrinsics information. Researchers can filter trajectories with these scenario tags to benchmark performance or train policies for specific social navigation scenarios (e.g., high-density crowds, intersections, or passing conversational groups). (a) (b) (c) (d) (e) (f) (g) Figure 10: Pedestrian density and traversability analysis of scenes captured from egocentric view across ACME and 2 recent Social navigation datasets: SCAND (Karnan et al., 2022a) and MuSoHu (Nguyen et al., 2023). 5.2 BEV Camera Data The ACME datasetâs BEV data mainly consists of annotated human trajectories in 10Hz pixel coordinates and the homography matrices that transform the pixel coordinates into ground plane metric coordinates. They are saved as toml and txt files, respectively. Similar to the TBD dataset (Wang et al., 2024a), the human trajectories were verified by the author group, and focused mainly on the moving pedestrians. Each trajectory contains an ID, a starting frame number, and a series of trajectory coordinates. The semantic scene tags are additionally provided in mask image formats as shown in Fig. 9. These tags can be adjusted by using our open-source tagging tool. Lastly, BEV RGB videos are provided, adhering to each institutionâs anonymization requirements. Metadata is also included and contains information such as the BEV cameraâs calibrated intrinsics and distortion coefficients. We also provide the transformed trajectories in metric coordinates on the ground plane as toml files. They are downsampled to 2.5Hz and organized into txt files following ETH (Pellegrini et al., 2009a) and UCY (Lerner et al., 2007) dataset formats, so that trajectory prediction models can directly load the BEV trajectory data. To ensure the metric trajectories are clean, the transformed trajectories were further processed as follows: 1. Discard trajectories shorter than 3.2âs3.2s. 2. In rare occasions, if there is a sudden jump in trajectories due to tracking noise (instantaneous velocity >5âm/s>5m/s), perform a linear interpolation-based fix to connect to coordinates at later timestamps. If no coordinates are found, this segment is discarded. 3. Smooth the trajectories using a smoothing window size of 55. 4. Label the trajectory segments that are stationary (average speed over a window of 88 time steps is <0.5âm/s<0.5m/s). 5. Limit the trajectories of the pedestrians to be within (±25âm,±25âm)(± 25m,± 25m) of the poster-sized tag. This mainly affected UBonn and NUS because some of the data collection areas were large, and the cameras were pointed at a narrower angle, so it was difficult to obtain accurate positions for pedestrians who were far away. The BEV trajectory data analyzed in the evaluation sections were processed using these procedures. 6 Dataset Analysis: Social navigation scenes from onboard-robot data In this section, we analyze on-board robot data collected in ACME and compare it to prior datasets focused on social navigation, specifically SCAND (Karnan et al., 2022b) and MuSoHu (Nguyen et al., 2023). We envision ACME to be useful for learning social navigation robot policies as well as pedestrian trajectory prediction models, and therefore aim to capture scenarios that prove to be challenging for current off-the-shelf navigation models. Our metrics for comparison focus on the complexity and diversity of data captured in each dataset. âScenario complexityâ is hard to quantify and highly subjective. Although there have been recent efforts to characterize scenario complexity (Stratton et al., 2025), analyzing large datasets with the relevant factors that constitute such metrics in an automated manner remains challenging. Thus, we design metrics that can be scalably computed on any dataset. Our characterization of scenario complexity comprises 3 factors that address the inherent multidimensional nature of social navigation scenarios: 1. Pedestrian Density: How crowded is the scene that the robot finds itself in? 2. Traversability: How constrained is the robotâs motion due to pedestrians/other spatial constraints? 3. Degree of Social Compliance (Raj et al. (2024)): How would an off-the-shelf geometric planner behave in the situation the robot finds itself in? How different is this trajectory from the expert âsocially compliantâ trajectory? Figure 11: Navigation scenarios from ACME with the same number of detected pedestrians but presenting vast differences in the navigation challenge faced. 6.1 Pedestrian Density We use the same off-the-shelf Multi-Object Tracking model used for anonymizing our data to generate pedestrian detections on egocentric images for SCAND (Karnan et al., 2022b) and MuSoHu (Nguyen et al., 2023). We characterize pedestrian density with two metrics: 1) Number of pedestrians detected in the image, and 2) Proportion of the image covered by pedestrians. The latter is an important statistic to identify how close people are to the robot (as a proxy when depth information and intrinsic camera parameters are unavailable), which is indicative of the crowd density the robot is navigating through. For example, in Fig 11, although both the images have 8 (detected) pedestrians in total, the scenario on the left offers a much greater challenge to navigate safely as opposed to the scenario on the right owing to a larger number of pedestrians close to the robot. Fig 10(b) shows that ACME captures far fewer scenes with no humans and remains competitive with SCAND (roughly 1/4th the size of our dataset by duration) in terms of the number of pedestrians captured per scene. Notably, while SCAND captures a larger number of scenes with a high pedestrian count (Fig. 10(a)), ACME captures a larger proportion of scenes with pedestrians closer to the robot (Fig.10(c)). 6.2 Traversability The area available for the robot to safely traverse the environment in the presence of humans can indicate the complexity of scenarios captured. Larger traversable regions generally provide more feasible paths around pedestrians, while constrained regions increase the likelihood of close-proximity interactions and navigation decisions that require social awareness. To quantify this aspect, we process the SCAND, MuSoHu, and ACME datasets using a finetuned SAM2 model (Wang et al. (2024b)) to predict a traversability mask on the egocentric RGB images. Fig. 10(c,d) compares the proportion of image area classified as traversable across ACME, SCAND, and MuSoHu. SCAND exhibits the largest traversable regions on average, suggesting that many of its scenes provide relatively open planning spaces. In contrast, MuSoHu is shifted toward lower traversability values, indicating more spatially constrained scenes. ACME lies between these two datasets: its scenes contain less free navigation space than SCAND, but more than MuSoHu. This suggests that ACME captures a broad range of moderately constrained navigation settings, where the robot often has feasible alternatives but still encounters reduced planning space that can require socially aware decision-making. Furthermore, we analyze the relationship between pedestrian coverage (the pixel-wise proportion of egocentric images occupied by pedestrians) and traversability. While raw pedestrian counts can be misleading in open spaces, pedestrian coverage serves as a more reliable proxy for agent proximity and the resulting obstruction of the robotâs path. As illustrated in Fig.10 (f), we observe a clear inverse correlation: as pedestrians occupy more of the visual field (indicating closer proximity or higher local density), the available traversable area decreases. Notably, ACME captures a significantly higher density of âsocially constrainedâ scenarios compared to SCAND and MuSoHu (lower-right regions in Fig 10 (f)). While MuSoHu contains many low-traversability scenes, these are primarily narrow, empty environments with few pedestrians. In contrast, ACME focuses on the long-tail of social navigationâscenes where the robot must negotiate constrained spaces, specifically while navigating through or around human crowds 6.3 Robot Speech (a) Scenes with Robot speech utterances in NUS data (b) Scenes with Robot speech utterances in UEx data (c) Scenes with speech utterances in Miraikan & Keio data Figure 12: Robot speech usage in the NUS, UEx, Miraikan, and Keio subsets. Each marker corresponds to a frame in which a robot speech utterance was issued, plotted by pedestrian density and traversability at that timestamp. Gray contour lines show the overall distribution of pedestrian density and traversability within the corresponding data subset. Trajectories collected in NUS, UEx, Miraikan, and Keio also contain annotations of robot speech usage, indicating when and where robot speech was used during crowd navigation. We discovered that the robot operation method impacts robot speech behavior during data collection and thus makes a distinction between the teleoperated NUS and UEx robots and the human-operated Miraikan and Keio suitcase robots. 6.3.1 Tele-operated robots Recall that the NUS and UEx robots were configured with onboard speakers and could actively play one of 3 utterances: âExcuse Meâ, âAttention! Robot Hereâ and âPlease Give Wayâ. In the NUS and UEx subsets, we found âExcuse Meâ and âPlease Give Wayâ were used under similar interaction conditions, and thus we combined them into a single category. Post-collection feedback from teleoperators suggests that the utterances correspond to two broad interaction functions: âą âExcuse Me/Please Give Wayâ: When the robotâs path was hindered by pedestrians or the robot moved in close proximity to/overtook a person or group of people. This typically occurred in crowded scenes, where the robot needed to request passage to avoid collision. âą âAttention Robot Hereâ,: When pedestrians needed to be alerted of the robotâs presence, especially in situations of low visibility (e.g., doorways/blind corners) or the robot overtakes/passes by a distracted pedestrian (e.g., looking at a phone) and there was a possibility of collision. In these cases, speech served primarily as an awareness cue rather than an explicit request for passage. These usage patterns are also reflected in the FPV scenes corresponding to utterance timestamps. Fig 12 compares the proportion of image area inferred as traversable against the proportion of image area occupied by pedestrians at speech timestamps. In both NUS and UEx, âAttention Robot Hereâ is used at relatively lower pedestrian densities, whereas âExcuse Me/Please give wayâ is used in more crowded scenes for space negotiation. The plots also reveal an embodiment-dependent difference between the NUS and UEx robots. In the NUS subset, âAttention Robot Hereâ is distributed over a wider range of pedestrian densities and traversability values. This is consistent with the NUS platform being a small quadruped at a maximum height of 4040 cm from the ground, making it less visually salient in crowds than the 1.331.33 m tall UEx shadow robot. 6.3.2 Human Operated Robots For the suitcase-robot trajectories in Miraikan and Keio, the human operators pressed a button on the suitcase handle interface whenever they verbally said âExcuse Meâ during a trajectory recording. Because the human operators were physically connected to the robots and were adept at crowd navigation, pedestrians respond primarily to the accompanying human rather than the inconspicuous suitcase as an autonomous social actor. As a result, âExcuse Meâ occurred only in the context of causing disturbance to other pedestrians from the judgment of the human operator. In contrast, when the robots were teleoperated in NUS and UEx, the teleoperators were less in control and executed speech more often. As a result, even though the combined duration of the Miraikan and Keio data is comparable to that of the combined NUS and UEx subsets, speech usage is substantially lower in the suitcase-robot data as shown in Fig 12(c). In Miraikan and Keio, when the accompanying human operators caused disturbance to other pedestrians and uttered âExcuse Meâ, the robots often were not in the immediate vicinity of any other pedestrian. This is because disturbances in Miraikan and Keio often took the form of causing other pedestrians to change course before getting close to them, or cutting through empty space where pedestrians were interacting with each other or with objects in the environment (the reasonable way to drive the robot in high-density situations based on the operatorsâ judgment). During these disturbance events, the robot may not be physically close to other pedestrians. This may be the underlying reason that the data in Fig 12(c) on Miraikan and Keio concentrates towards higher traversable area proportions and lower pedestrian image area proportions. Finally, we emphasize that the quantities in Fig 12 are computed in FPV image space with trained models for detecting pedestrians and traversability. They should not be interpreted as direct metric measurements of crowd density or navigable free space. These quantities are affected by camera intrinsics, extrinsics, and segmentation quality. With respect to speech usage patterns, the most reliable interpretation is not that a universal visual threshold determines when speech is used, but rather that robot-speech is a situated, embodiment-dependent navigation action whose use correlates with visually constrained and socially interactive scenes. 6.4 Comparison to a Geometric Planner Figure 13: Although the distribution of Hausdorff distances between the expert trajectory and TEB remains similar across datasets, the higher planner failure rates showcase that off-the-shelf planners have more difficulty navigating in the scenes captured in ACME compared to other datasets. Raj et al. (2024) defined âsocial-complianceâ of motion planners using the Hausdorff distance between the global plan generated by the planner and the expert trajectory. We hypothesize that our dataset, with its focus on capturing challenging scenarios, would reveal more situations where such context-and-norm-unaware planners would be deemed socially noncompliant. We perform an analysis identical to that of Raj et al. (2024), by sampling future odometry goals (5m ahead of the robotâs current position at a 1Hz rate) and measure the Hausdorff distance between the plan generated by move_base (in contrast to Raj et al. (2024), we use the TEB local planner (Rösmann et al., 2017) to emulate real-world testing conditions) and the expert trajectory. Our experiments revealed an interesting phenomenon: as shown in Fig 13.a, the undirected Hausdorff distance between the TEB planner and the expert trajectory remains similar across all datasets. However, analysis of the failure rate of the planners reveals a clear trend (Fig. 13 b): scenes in the ACME dataset are far more challenging to navigate than SCAND and MuSoHu, resulting in 2x more failures per minute and 4x per 5 meters of data compared to the dataset with the second highest failure rate (SCAND). We verify that the primary mode of planner failure across the dataset is the presence of pedestrians at the sampled goal position for the planner, thus directly correlating with the crowd density that the robot must navigate through. 6.5 Benchmarking Vision Navigation models We showcase the utility of ACME for training and testing general-purpose navigation agents by benchmarking two SoTA navigation models. Specifically, we focus on models designed for real-world vision navigation with zero-shot or few-shot transfer to novel environments and embodiments, namely ViNT (Shah et al., 2023) and NoMAD (Sridhar et al., 2024). We evaluate model performance based on alignment to the expert trajectories with Final Destination Error (FDE), Average MSE on waypoints, and MAOE as proposed in Liu et al. (2025), as well as Total AOE (aggregate orientation error across predicted waypoints). For reference values, we also list performance on the SCAND and MuSoHu datasets. Table 4 shows the results of the benchmark. As expected, both models perform well on the SCAND (since ViNT and NoMAD are trained on SCAND data as well) and relatively poorly on the MuSoHu and ACME dataset. While NoMADâs performance on the MuSoHu and ACME dataset is very similar, ViNT performs better on MuSoHu than on ACME. The ease of using our dataset for such benchmarks, paired with the location and embodiment diversity, could bolster future work to investigate scenarios of failure and improve foundational navigation models to traverse human-inhabited spaces. Figure 14: Qualitative Results: Trajectories generated by SoTA foundational navigation models compared with the ground truth trajectories. Table 4: Performance Benchmark of SoTA foundational navigation models FDE (m) MSE Cosine Similarity MAOE (° ) Mean AOE (° ) Total AOE (° ) ViNT ACME 0.850.85 p m 0.60 0.240.24 p m 0.30 0.950.95 p m 0.23 12.5812.58 p m 26.51 8.248.24 p m 19.93 41.2241.22 p m 99.64 SCAND 0.460.46 p m 0.52 0.110.11 p m 0.23 0.980.98 p m 0.13 5.485.48 p m 15.47 3.683.68 p m 11.31 18.4118.41 p m 56.53 MuSoHu 0.830.83 p m 0.49 0.210.21 p m 0.34 0.960.96 p m 0.18 14.3914.39 p m 21.25 10.1610.16 p m 16.26 50.8050.80 p m 81.30 NoMaD ACME 1.441.44 p m 1.00 0.630.63 p m 0.77 0.930.93 p m 0.22 20.6820.68 p m 31.35 11.7411.74 p m 19.75 93.8893.88 p m 158.00 SCAND 0.790.79 p m 0.92 0.300.30 p m 0.59 0.980.98 p m 0.12 6.866.86 p m 16.39 4.114.11 p m 10.73 32.8932.89 p m 85.83 MuSoHu 1.531.53 p m 0.88 0.630.63 p m 0.77 0.940.94 p m 0.19 18.8318.83 p m 26.10 11.8011.80 p m 17.27 94.4294.42 p m 138.16 7 Dataset Analysis: Pedestrian Trajectory Analysis from BEV data Table 5: Dataset statistics comparison, for datasets with human-verified trajectories grounded in metric space. Datasets Duration # Trajectories Freq (Hz) ETH Pellegrini et al. (2009a) 25 min 650 15 UCY Lerner et al. (2007) 16.5 min 786 2.5 Town Centre Benfold and Reid (2011) 5 min 157 2.5 WildTrack Chavdarova et al. (2018) 200 sec 313 2 JRDB Martin-Martin et al. (2021) 62 min ⌠3.5K 7.5 THĂR Rudenko et al. (2020a) 60+ min 600+ 100 TBD Wang et al. (2024a) 626 min 10.3K 10 SiT Bae et al. (2023) 9 min 1861 10 THĂR-MAGNI Schreiter et al. (2025) 1416 min ⌠10K 100 Bi3 Stratton et al. (2026a) 630 min ⌠11K 120 ACME (Ours) 2613 min 72.1K 10 For comparisons among trajectory prediction datasets, we first analyze basic statistics, i.e., duration, speed, and density. We then benchmark state-of-the-art trajectory prediction models to compare with the combined ETH (Pellegrini et al., 2009a) and UCY (Lerner et al., 2007) dataset. Lastly, to examine whether our dataset captures diverse human behavior, we compare ACMEâs sub-datasets from different countries and semantic environment tags, using context-dependent metrics. We revealed several differences in human navigation behavior from different cultures. Table 5 shows the quantity of our ACME dataset compared to prior datasets that contain human-verified pedestrian trajectory annotations in the metric space. Compared to the second largest dataset, the TBD dataset (Wang et al., 2024a), ACME is almost 4 times larger by duration, and contains 7 times more pedestrian trajectories. Additionally, the TBD dataset only contains data at one location. (a) (b) (c) Figure 15: Trajectory statistics of our data compared to prior datasets with ± standard deviation. The statistics on the right side of the dashed lines are data for each of our sub-datasets. (a) duration in seconds. (b) average motion speed in meters per second. (c) minimum distance between any two pedestrians over time in meters. We additionally compare the statistics of our dataset and representative datasets, extending the evaluation by Rudenko et al. (2020a). In particular, we use the following metrics: (1) Tracking Duration (s): average time duration of the trajectories. (2) Motion Speed (m/sm/s): average speed of the trajectories. (3) Minimum Distance Between People (m): minimum Euclidean distance between any two people, averaged over frames, which partially reflects overall dataset density. The perception noise metric from Rudenko et al. (2020a) is not included in this analysis, as it is revealed in practice that this metric is heavily influenced by the implementation of our post-processing pipeline. For example, a longer trajectory smoothing window will lower perception noise significantly, but at the cost of distorting the trajectories. And for reasons similar to Wang et al. (2024a), trajectory curvature is not measured. 7.0.1 Tracking Duration As shown in Figure 15a, our datasetâs mean tracking duration is not the highest, but with a large standard deviation (±45.1± 45.1). The large variation is typically caused by wandering pedestrians and static pedestrians stopping for various activities, such as ordering food or having a conversation with others. These pedestrian trajectoriesâ durations are often long and varied. Other datasets with a large presence of wandering and static pedestrians also have large variances, such as the ATC (±64.7± 64.7) and TBD (±57.1± 57.1) datasets. Among the sub-datasets, NUS, Miraikan, UMich, and UBonn contain a moderate mix of high-duration trajectories, while CMU and Keio contain a high number of high-duration trajectories. 7.0.2 Average Speed As shown in Figure 15b, overall, the average speed in our dataset is the lowest, caused by large numbers of static pedestrians. However, our dataset has the highest standard deviation (±0.75± 0.75), while the second highest is Edinburgh (±0.64± 0.64) and the third highest is TBD (±0.52± 0.52). This shows that our dataset contains great variation in pedestrian motions. Among the sub-datasets, in CMU and Keio, pedestrians have a low average speed due to a large presence of static pedestrians. In Miraikan, pedestriansâ average speed is also low, because the pedestriansâ overall walking speed tends to be slow to enjoy the museum, and there is also a moderate amount of static pedestrians (e.g., stopping to examine exhibitions). NUS (±0.90± 0.90) and UBonn (±0.87± 0.87) have a high variation in average speed. 7.0.3 Minimum Distance Between People As shown in Figure 15c, our dataset is smaller than Edinburgh but larger than others. However, our dataset has the largest standard deviation (±4.76± 4.76), while the second largest is Edinburgh (±3.5± 3.5) and the third largest is THĂR (±1.6± 1.6). This shows that our dataset has great variation in crowd density. Among the sub-datasets, in high population density locations such as NUS, Miraikan, and Keio, the average crowd density is high, while in the USA and Germany, crowd density is lower. Moreover, NUS (±3.29± 3.29), UMich (±5.69± 5.69), and UBonn (±9.10± 9.10) have particularly great variation in crowd density. 7.1 Benchmarking SoTA Human Trajectory Prediction Models Table 6: Performance Benchmark of SoTA trajectory prediction models. The data in each cell is ADE/FDE in meters (m). Models ETH/UCY ACME ACME * SocialGAN Gupta et al. (2018) 0.34 / 0.71 0.40 / 0.80 0.75 / 1.50 AgentFormer Yuan et al. (2021) 0.23 / 0.47 0.30 / 0.58 0.54 / 1.05 SGNet Wang et al. (2022b) 0.22 / 0.48 0.39 / 0.74 0.71 / 1.39 TUTR Shi et al. (2023) 0.22 / 0.44 0.31 / 0.59 0.56 / 1.09 MoFlow Fu et al. (2025) 0.21 / 0.41 0.31 / 0.59 0.54 / 1.06 â* * indicates the benchmark results on input trajectories with average speed greater than 0.5âm/s0.5m/s. An essential use case of the annotated BEV human trajectories is to train trajectory prediction models. To demonstrate this use case, we benchmarked 6 models that represent state-of-the-art trajectory predictors from different eras: SocialGAN (Gupta et al., 2018), AgentFormer (Yuan et al., 2021), SGNet (Wang et al., 2022b), TUTR (Shi et al., 2023), and MoFlow (Fu et al., 2025). Consistent with prior benchmarks, the trajectory annotations were downsampled to 2.5Hz. Then, the models were provided 8 timestamps of the trajectory histories (3.2s) and generated predictions for the future 12 timestamps (4.8s). The models were evaluated stochastically by sampling 20 trajectory predictions, and we took the minimum average displacement errors (ADE) and final displacement errors (FDE) among the sampled trajectories as the results. We trained the models in a cross-validation fashion on the five sub-datasets of the ETH and UCY datasets. Finally, we evaluated the models both on the combined ETH and UCY datasets and on our entire ACME dataset. The results from the cross-validations are aggregated together by averaging over the minimum 20 ADEs and the minimum 20 FDEs. The benchmarking results are shown in Table 6. The state-of-the-art models perform progressively better on the ETH and UCY combined dataset, consistent with the findings from their respective papers. The newer models also generally perform better on our ACME dataset, except for SGNet, which only performs slightly better than SocialGAN, and AgentFormer, whose performance is on par with the most recent Moflow. All models perform worse on our dataset when compared to the results achieved on ETH and UCY, with ADE and FDE performances dropping by 0.1âm0.1m and 0.15âm0.15m on average. Our dataset contains a mixture of static and dynamic pedestrians, while most pedestrians in ETH and UCY are only dynamic. Similar to the evaluation protocol in the TBD dataset (Wang et al., 2024a), we additionally benchmarked the models on dynamic pedestrians only. Because the ADE and FDE of the static pedestrians are near zero, we evaluated the models on the dynamic pedestrian trajectories of our dataset only (average speed >0.5âm/s>0.5m/s), and found that the models perform worse. As shown in Table 6, the ADE and FDE performances drop by 0.38âm0.38m and 0.72âm0.72m on average, respectively. These performance drops suggest that our ACME dataset presents a challenging evaluation setting and captures aspects of human motion that are not sufficiently represented in the small-scale ETH and UCY datasets. This indicates that our dataset can serve as a benchmark for assessing model generalization and robustness beyond existing datasets. 7.2 Comparison across Datasets A core contribution of our dataset is that it was collected at diverse locations from 7 different institutions around the world, allowing it to reflect differences in human behavior under different environmental layouts and cultures. In this section, we dive deeper into the annotated human trajectory data on the metric ground plane to identify behavior differences and whether they are related to the varying contexts. Based on the semantic tags assigned to the BEV data and following Francis et al. (2025), we focus on three common environmental layouts: open areas, wide corridors, and narrow corridors. We next identify which sub-datasets provide substantial trajectory coverage for each environmental layout: UBonn, Keio, NUS, Miraikan, and UMich for open areas; CMU, GMU, NUS, Miraikan, and UMich for wide corridors; and CMU, NUS, and Miraikan for narrow corridors. 7.2.1 Instantaneous Speed Profiles Figure 16: Mean, median, and speed distribution of all trajectory segments capped at 5âm/s5m/s. Only trajectories with average speed >0.5âm/s>0.5m/s are included. The top yellow cluster data are from sub-datasets with significant open area data. The middle blue cluster data are from sub-datasets that contain significant wide corridor data. The bottom red cluster data are from sub-datasets that contain significant narrow corridor data. Figure 17: Relative direction of pedestrians when interacting with each other. Interaction is defined when two pedestrians get to the closest point to each other, and their distance is <4âm<4m. The top yellow cluster data are from sub-datasets with significant open area data. The middle blue cluster data are from sub-datasets that contain significant wide corridor data. The bottom red cluster data are from sub-datasets that contain significant narrow corridor data. Figure 18: Personal space, or average minimum distance of pedestrians when interacting with each other. Interaction is defined when two pedestrians get to the closest point to each other and their distance is <4âm<4m. The top yellow cluster data are from sub-datasets with significant open area data. The middle blue cluster data are from sub-datasets that contain significant wide corridor data. The bottom red cluster data are from sub-datasets that contain significant narrow corridor data. We analyze walking speed at a finer temporal resolution than before. Instead of summarizing each trajectory by its average speed, we use all instantaneous speeds along all the trajectories and group the results by semantic scene tag. To focus on walking behavior, we include only moving pedestrians, defined as trajectories with an average speed greater than 0.5âm/s0.5\,m/s (some instantaneous speeds may still fall below 0.5âm/s0.5\,m/s). Figure 16 shows the resulting speed distributions. We observe that speed distributions remain similar across different environment layouts. NUS and Miraikan contain data across open areas, wide corridors, and narrow corridors, but the locationsâ speed distributions are similar. UMichâs pedestrian speed distributions are also similar between open areas and wide corridors. CMUâs pedestrian speed distributions are also similar between wide corridors and narrow corridors. This shows that culture and general location attributes dictate walking speed more than environmental layout contexts. Pedestrians in UBonn and NUS exhibit higher walking speeds, which may be partly explained by the prevalence of outdoor settingsâall UBonn locations are outdoors, and NUS also includes many outdoor scenes. In contrast, pedestrians in higher-density locations such as Miraikan and Keio tend to walk more slowly, likely due to the limited space available for navigation. Besides, Miraikan is a museum environment, so visitors often slow down or stop completely to view exhibits. 7.2.2 Passing Side Next, we analyze pedestrian behavior in terms of passing side preference. We first identify all pairs of dynamic trajectories that at any point are within 4âm4\,m of each other and record the point of closest approach for each pair, which we call an interaction point. At each interaction point, we measure the relative direction of one pedestrian with respect to the otherâs heading. These directions characterize how pedestrians position themselves during passing, overtaking, and following interactions. Figure 17 summarizes the results using radial bar charts. Our analysis suggests that pedestriansâ preferred passing side tends to align with the local driving side. UBonn and UMichâs open space data show a strong tendency for other interacting pedestrians to approach closest on the left side, meaning people overwhelmingly pass on the right in Germany and in the US. CMU, GMU, and UMichâs wide corridor data, as well as CMUâs narrow corridor data, all demonstrate the conformity of passing on the right in the US. However, in high-density Asian locations in Japan and Singapore, passing side tendencies become less obvious. While according to Miraikan and NUSâs open space and wide corridor data, people tend to pass on the left, this observation is less pronounced in the narrow corridor data from Miraikan and NUS. In other observations, in high-density locations such as Keio, Miraikan, and NUS, significantly more interaction points occur at the front and back of pedestrians. In crowded areas, pedestrians tend to follow other pedestrians closely, forming lanes to navigate traffic. Close following can result in more close encounters at the front and the rear. 7.2.3 Personal Space Finally, we analyze human behavior in terms of personal space. To perform this analysis, we used the same interaction points as the passing side analysis. Rather than counting the number of interactions in each direction, we compute the average closest-approach distance as a function of relative direction. For each direction, we collect all interaction points where one pedestrian is located in that direction relative to another pedestrianâs heading and average their Euclidean distances. Repeating this over all directions (±180â± 180 ) gives an estimate of direction-dependent personal space. The resulting aggregated personal spaces by sub-datasets and BEV semantic tags are shown in Figure 18. Comparisons between open spaces and corridors in the NUS, Miraikan, and UMich data suggest that lateral personal space becomes narrower in corridor environments. This indicates that environmental layout affects personal space, as pedestrians appear to tolerate smaller side clearances in constrained spaces. However, these lateral differences should not be attributed too strongly to culture, since corridor widths may vary even within the same tagged semantic category (wide or narrow). In open spaces, frontal personal space appears similar across various locations, with Germanyâs slightly lower and the USâs slightly larger. In wide and narrow corridors, CMU, GMU, UMich, and NUSâs data show that personal space tends to be larger at the front in the US and Singapore, closer to 2.5âm2.5m. 8 Conclusion and Discussion In this work, we presented ACME, the first dataset to systematically capture social navigation and pedestrian trajectories across multiple cultural contexts and robot embodiments. ACME combines contributions from 8 data collection teams from culturally and geographically unique locations, thus providing a unique opportunity to study how navigation behaviors and social norms vary across geography, environment, and robot morphology. To the best of our knowledge, ACME is the largest multi-modal social navigation dataset featuring human demonstrations. It provides 29.35 hours of onboard robot data (approximately 1.4 times the size of MuSoHu) and 43.5 hours of pedestrian trajectory data (roughly 7 times the volume of comparable datasets like TBD and THâOR-MAGNI). Additionally, ACME introduces unique features, including 5 distinct robot embodiments and explicit robot-crowd interactions through robot speech. To maximize the utility of the dataset, data with no pedestrians is filtered out, tracking information of anonymized pedestrians is included, the dynamic trajectories from top-down BEV videos are human-verified and tagged semantically, and each robot trajectory is tagged with relevant scenarios to ease data lookup. Analysis of the dataset showed that ACME captures more complex scenarios and pedestrian trajectories than prior datasets while also offering interesting insights into pedestrian behaviors across cultures and contexts and embodiment-dependent interaction behaviors. In the future, we aim to collect data in more open settings with full maps for better localization, possibly leveraging pedestrian trajectory annotations to enhance localization robustness to dynamic pedestrians. Robust localization in dynamic environments is needed to achieve the projection of pedestrian trajectory annotations onto the robotâs egocentric perspective, which unlocks additional utility such as first-person view trajectory prediction (Liu et al., 2026). Currently, only robots used in Miraikan and Keio sub-datasets have accurate localization that is robust to dynamic pedestrians. We additionally aim to expand to multiple robots in the same location to further investigate differences in pedestrian behaviors localized to specific cultures, and train culture and embodiment-aware social navigation policies. 9 Ethical Approval The collection of the ACME dataset involved human observational data across multiple international sites. Ethical approval or exemption was obtained by each participating institutionâs respective review board prior to data collection: âą Approved via Full/Expedited Review: Data collection procedures were reviewed and approved by the Institutional Review Boards at Carnegie Mellon University (STUDY2021_00000199), the University of Extremadura (89//2025 and 214//2025), Keio University, Faculty of Science and Technology (2026-039), and Miraikanâs ethical approval was obtained under Keio University and Miraikanâs own legal team. âą Exempt Status: The review boards at the National University of Singapore (NUS-IRB-2024-476), the University of Michigan (HUM00268385), George Mason University (STUDY00000382), and the University of Bonn determined that the observational data collection in public spaces did not constitute human subjects research requiring full review, granting an exempt status. 10 Funding This research project is supported by A*STAR under its National Robotics Programme (NRP) (award M23NBK0053); the JST ASPIRE Program (JPMJAP2501); NSF 2531320 Public Space Robotics: Community-Driven Models for Social Navigation and Communication; the National Science Foundation (grants 2350352 and 2531320); the Spanish Government under grant PID2022-137344OB-C31 (MCIN/AEI/10.13039/501100011033/FEDER, UE); and the German Federal Ministry of Research, Technology and Space (BMFTR) under the Robotics Institute Germany (RIG), grant No. 16ME0999. 11 Data Accessibility Statement The ACME dataset and associated research materials will be hosted at Hugging Face rather than uploaded directly as research data files. Currently, because of privacy concerns, we are sharing the dataset with the reviewers privately via a password-protected Box in the following link: . The repository will provide access to multi-modal data, including egocentric RGB, 3D LiDAR, odometry, calibration information, interaction annotations, scenario tags, and context semantic segmentation and human-verified pedestrian trajectory annotations from overhead views. For the public release, data collected at each location will comply with the privacy, anonymization, and institutional review requirements of the participating data-collection sites. Accordingly, some visual RGB data have been anonymized through face blurring or full-body masking. However, we plan to release the dataset publicly on Hugging Face via a C BY license. Accompanying the dataset, all annotations tools and processing scripts will also be made publicly available under Apache License 2.0. References Agrawal et al. (2026) Agrawal S, Ostermann-Myrau N, Dengler N and Bennewitz M (2026) Peroi: A pedestrian-robot interaction dataset for learning avoidance, neutrality, and attraction behaviors in social navigation. In: Proc. of IEEE International Conference on Robotics and Automation (ICRA). Alahi et al. (2016a) Alahi A, Goel K, Ramanathan V, Robicquet A, Fei-Fei L and Savarese S (2016a) Social lstm: Human trajectory prediction in crowded spaces. In: Proceedings of the IEEE conference on computer vision and pattern recognition. p. 961â971. Alahi et al. (2016b) Alahi A, Goel K, Ramanathan V, Robicquet A, Fei-Fei L and Savarese S (2016b) Social lstm: Human trajectory prediction in crowded spaces. In: Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit. p. 961â971. Alahi et al. (2014) Alahi A, Ramanathan V and Fei-Fei L (2014) Socially-aware large-scale crowd forecasting. In: Proc. IEEE Conf. on Comput. Vis. and Pattern Recognit. p. 2203â2210. Bae et al. (2023) Bae JW, Kim J, Yun J, Kang C, Choi J, Kim C, Lee J, Choi J and Choi JW (2023) Sit dataset: socially interactive pedestrian trajectory dataset for social navigation robots. Advances in neural information processing systems 36: 24552â24563. Benfold and Reid (2011) Benfold B and Reid I (2011) Stable multi-target tracking in real-time surveillance video. In: CVPR 2011. IEEE, p. 3457â3464. Bradwell et al. (2021) Bradwell HL, Winnington R, Thill S and Jones RB (2021) Morphology of socially assistive robots for health and social care: A reflection on 24 months of research with anthropomorphic, zoomorphic and mechanomorphic devices. In: 2021 30th IEEE International Conference on Robot & Human Interactive Communication (RO-MAN). IEEE, p. 376â383. BrĆĄÄiÄ et al. (2013) BrĆĄÄiÄ D, Kanda T, Ikeda T and Miyashita T (2013) Person tracking in large public spaces using 3-d range sensors. IEEE Trans. on Human-Machine Syst. 43(6): 522â534. Chavdarova et al. (2018) Chavdarova T, BaquĂ© P, Bouquet S, Maksai A, Jose C, Bagautdinov T, Lettry L, Fua P, Van Gool L and Fleuret F (2018) Wildtrack: A multi-camera hd dataset for dense unscripted pedestrian detection. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Che et al. (2020) Che Y, Okamura AM and Sadigh D (2020) Efficient and trustworthy social navigation via explicit and implicit robotâhuman communication. IEEE Transactions on Robotics 36(3): 692â707. Chen et al. (2019) Chen C, Liu Y, Kreiss S and Alahi A (2019) Crowd-robot interaction: Crowd-aware robot navigation with attention-based deep reinforcement learning. In: 2019 international conference on robotics and automation (ICRA). IEEE, p. 6015â6022. Chen et al. (2024) Chen H, Ding J, Li Y, Wang Y and Zhang XP (2024) Social physics informed diffusion model for crowd simulation. In: Proceedings of the AAAI Conference on Artificial Intelligence, volume 38. p. 474â482. Chen et al. (2017) Chen YF, Liu M, Everett M and How JP (2017) Decentralized non-communicating multiagent collision avoidance with deep reinforcement learning. In: 2017 IEEE international conference on robotics and automation (ICRA). IEEE, p. 285â292. Cheng et al. (2024) Cheng AC, Ji Y, Yang Z, Gongye Z, Zou X, Kautz J, Bıyık E, Yin H, Liu S and Wang X (2024) Navila: Legged robot vision-language-action model for navigation. arXiv preprint arXiv:2412.04453 . De Heuvel et al. (2023) De Heuvel J, Corral N, Kreis B, Conradi J, Driemel A and Bennewitz M (2023) Learning depth vision-based personalized robot navigation from dynamic demonstrations in virtual reality. In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 6757â6764. Dutta and Zisserman (2019) Dutta A and Zisserman A (2019) The via annotation software for images, audio and video. In: Proceedings of the 27th ACM international conference on multimedia. p. 2276â2279. Epley et al. (2007) Epley N, Waytz A and Cacioppo JT (2007) On seeing human: a three-factor theory of anthropomorphism. Psychological review 114(4): 864. Francis et al. (2025) Francis A, PĂ©rez-DâArpino C, Li C, Xia F, Alahi A, Alami R, Bera A, Biswas A, Biswas J, Chandra R, Chiang HTL, Everett M, Ha S, Hart J, How JP, Karnan H, Lee TWE, Manso LJ, Mirsky R, Pirk S, Singamaneni PT, Stone P, Taylor AV, Trautman P, Tsoi N, VĂĄzquez M, Xiao X, Xu P, Yokoyama N, Toshev A and MartĂn-MartĂn R (2025) Principles and guidelines for evaluating social robot navigation algorithms. ACM Transactions on Human-Robot Interaction 14(2). Fu et al. (2025) Fu Y, Yan Q, Wang L, Li K and Liao R (2025) Moflow: One-step flow matching for human trajectory forecasting via implicit maximum likelihood estimation based distillation. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 17282â17293. Gupta et al. (2018) Gupta A, Johnson J, Fei-Fei L, Savarese S and Alahi A (2018) Social gan: Socially acceptable trajectories with generative adversarial networks. In: Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit. p. 2255â2264. Han et al. (2025) Han JR, Vanniasinghe M, Sahak H, Rhinehart N and Barfoot TD (2025) Ratatouille: Imitation learning ingredients for real-world social robot navigation. arXiv preprint arXiv:2509.17204 . Haring et al. (2016) Haring KS, Silvera-Tawil D, Takahashi T, Watanabe K and Velonaki M (2016) How people perceive different robot types: A direct comparison of an android, humanoid, and non-biomimetic robot. In: 2016 8th international conference on knowledge and smart technology (kst). IEEE, p. 265â270. Haring et al. (2018) Haring KS, Watanabe K, Velonaki M, Tossell C and Finomore V (2018) Ffabâthe form function attribution bias in humanârobot interaction. IEEE Transactions on Cognitive and Developmental Systems 10(4): 843â851. Hart et al. (2020) Hart J, Mirsky R, Xiao X, Tejeda S, Mahajan B, Goo J, Baldauf K, Owen S and Stone P (2020) Using human-inspired signals to disambiguate navigational intentions. In: International Conference on Social Robotics. Springer, p. 320â331. Helbing and Molnar (1995) Helbing D and Molnar P (1995) Social force model for pedestrian dynamics. Physical review E 51(5): 4282. Hirose et al. (2018) Hirose N, Sadeghian A, VĂĄzquez M, Goebel P and Savarese S (2018) Gonet: A semi-supervised deep learning approach for traversability estimation. In: 2018 IEEE/RSJ international conference on intelligent robots and systems (IROS). IEEE, p. 3044â3051. Hirose et al. (2023) Hirose N, Shah D, Sridhar A and Levine S (2023) Sacson: Scalable autonomous control for social navigation. IEEE Robotics and Automation Letters 9(1): 49â56. Huang et al. (2025) Huang R, Xue H, Pagnucco M, Salim FD and Song Y (2025) Vision-based multi-future trajectory prediction: A survey. IEEE Transactions on Neural Networks and Learning Systems . Jocher et al. (2023) Jocher G, Qiu J and Chaurasia A (2023) Ultralytics YOLO. URL https://github.com/ultralytics/ultralytics. Kannan et al. (2021) Kannan S, Lee A and Min BC (2021) External human-machine interface on delivery robots: Expression of navigation intent of the robot. In: 2021 30th IEEE international conference on robot & human interactive communication (RO-MAN). IEEE, p. 1305â1312. Karnan et al. (2022a) Karnan H, Nair A, Xiao X, Warnell G, Pirk S, Toshev A, Hart J, Biswas J and Stone P (2022a) Socially compliant navigation dataset (scand): A large-scale dataset of demonstrations for social navigation. IEEE Robotics and Automation Letters 7(4): 11807â11814. Karnan et al. (2022b) Karnan H, Nair A, Xiao X, Warnell G, Pirk S, Toshev A, Hart J, Biswas J and Stone P (2022b) Socially compliant navigation dataset (scand): A large-scale dataset of demonstrations for social navigation. IEEE Robotics and Automation Letters 7(4): 11807â11814. Keselman et al. (2023) Keselman L, Shih K, Hebert M and Steinfeld A (2023) Optimizing algorithms from pairwise user preferences. In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 4161â4167. Kirby (2010) Kirby R (2010) Social robot navigation. Carnegie Mellon University. Kong et al. (2025) Kong Y, Song D, Liang J, Manocha D, Yao Z and Xiao X (2025) Autospatial: Visual-language reasoning for social robot navigation through efficient spatial reasoning learning. In: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 11298â11304. Kuribayashi et al. (2025) Kuribayashi M, Uehara K, Wang A, Morishima S and Asakawa C (2025) Wanderguide: Indoor map-less robotic guide for exploration by blind people. In: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI â25. New York, NY, USA: Association for Computing Machinery. ISBN 9798400713941. 10.1145/3706598.3713788. URL https://doi.org/10.1145/3706598.3713788. Lerner et al. (2007) Lerner A, Chrysanthou Y and Lischinski D (2007) Crowds by example. Comput. Graph. Forum 26(3): 655â664. Li et al. (2023) Li R, Li C, Ren D, Chen G, Yuan Y and Wang G (2023) Bcdiff: Bidirectional consistent diffusion for instantaneous trajectory prediction. Advances in Neural Information Processing Systems 36: 14400â14413. Liu et al. (2026) Liu J, Zhou J, Ye K, Lin KY, Wang A and Liang J (2026) Egotraj-bench: Towards robust trajectory prediction under ego-view noisy observations. In: 2026 IEEE International Conference on Robotics and Automation (ICRA). IEEE. Liu et al. (2023) Liu S, Chang P, Huang Z, Chakraborty N, Hong K, Liang W, McPherson DL, Geng J and Driggs-Campbell K (2023) Intention aware robot crowd navigation with attention-based interaction graph. In: 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, p. 12015â12021. Liu et al. (2025) Liu X, Li J, Jiang Y, Sujay N, Yang Z, Zhang J, Abanes J, Zhang J and Feng C (2025) Citywalker: Learning embodied urban navigation from web-scale videos. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 6875â6885. Macenski et al. (2022) Macenski S, Foote T, Gerkey B, Lalancette C and Woodall W (2022) Robot operating system 2: Design, architecture, and uses in the wild. Science robotics 7(66): eabm6074. Majecka (2009) Majecka B (2009) Statistical models of pedestrian behaviour in the forum. Masterâs thesis, School of Informatics, University of Edinburgh . Martin-Martin et al. (2021) Martin-Martin R, Patel M, Rezatofighi H, Shenoi A, Gwak J, Frankel E, Sadeghian A and Savarese S (2021) Jrdb: A dataset and benchmark of egocentric robot visual perception of humans in built environments. IEEE Trans. Pattern Anal. Mach. Intell. . MartĂn-MartĂn et al. (2023) MartĂn-MartĂn R, Patel M, Rezatofighi H, Shenoi A, Gwak J, Frankel E, Sadeghian A and Savarese S (2023) Jrdb: A dataset and benchmark of egocentric robot visual perception of humans in built environments. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(6): 6748â6765. 10.1109/TPAMI.2021.3070543. Mavrogiannis et al. (2023) Mavrogiannis C, Baldini F, Wang A, Zhao D, Trautman P, Steinfeld A and Oh J (2023) Core challenges of social robot navigation: A survey. ACM Transactions on Human-Robot Interaction 12(3): 1â39. Mead and Mataric (2012) Mead R and Mataric MJ (2012) A probabilistic framework for autonomous proxemic control in situated and mobile human-robot interaction. In: Proceedings of the seventh annual ACM/IEEE international conference on Human-Robot Interaction. p. 193â194. Mohamed et al. (2020) Mohamed A, Qian K, Elhoseiny M and Claudel C (2020) Social-stgcnn: A social spatio-temporal graph convolutional neural network for human trajectory prediction. In: Proc. IEEE/CVF Conf. on Comput. Vis. and Pattern Recognit. (CVPR). Muhammad et al. (2020) Muhammad K, Ullah A, Lloret J, Del Ser J and De Albuquerque VHC (2020) Deep learning for safe autonomous driving: Current challenges and future directions. IEEE Transactions on Intelligent Transportation Systems 22(7): 4316â4336. Narasimhan et al. (2024) Narasimhan S, Tan AH, Choi D and Nejat G (2024) Olivia-nav: An online lifelong vision language approach for mobile robot social navigation. arXiv preprint arXiv:2409.13675 . Nguyen et al. (2023) Nguyen DM, Nazeri M, Payandeh A, Datar A and Xiao X (2023) Toward human-like social robot navigation: A large-scale, multi-modal, social human navigation dataset. In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 7442â7447. Oh et al. (2011) Oh S, Hoogs A, Perera A, Cuntoor N, Chen C, Lee JT, Mukherjee S, Aggarwal J, Lee H, Davis L et al. (2011) A large-scale benchmark dataset for event recognit. in surveillance video. In: CVPR 2011. IEEE, p. 3153â3160. Olson (2011) Olson E (2011) Apriltag: A robust and flexible visual fiducial system. In: 2011 IEEE international conference on robotics and automation. IEEE, p. 3400â3407. Paez-Granados et al. (2022a) Paez-Granados D, He Y, Gonon D, Jia D, Leibe B, Suzuki K and Billard A (2022a) Pedestrian-robot interactions on autonomous crowd navigation: Reactive control methods and evaluation metrics. In: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). p. 149â156. 10.1109/IROS47612.2022.9981705. Paez-Granados et al. (2022b) Paez-Granados D, He Y, Gonon D, Jia D, Leibe B, Suzuki K and Billard A (2022b) Pedestrian-robot interactions on autonomous crowd navigation: Reactive control methods and evaluation metrics. In: 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). p. 149â156. 10.1109/IROS47612.2022.9981705. Panigrahi et al. (2023) Panigrahi B, Raj AH, Nazeri M and Xiao X (2023) A study on learning social robot navigation with multimodal perception. arXiv preprint arXiv:2309.12568 . Payandeh et al. (2025) Payandeh A, Song D, Nazeri M, Liang J, Mukherjee P, Raj AH, Kong Y, Manocha D and Xiao X (2025) Social-llava: Enhancing social robot navigation through human-language reasoning. In: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 17192â17198. Pellegrini et al. (2009a) Pellegrini S, Ess A, Schindler K and van Gool L (2009a) Youâl never walk alone: Modeling social behavior for multi-target tracking. In: Proc. IEEE Int. Conf. Comput. Vis. p. 261â268. Pellegrini et al. (2009b) Pellegrini S, Ess A, Schindler K and van Gool L (2009b) Youâl never walk alone: Modeling social behavior for multi-target tracking. In: 2009 IEEE 12th International Conference on Computer Vision. p. 261â268. 10.1109/ICCV.2009.5459260. Pirk et al. (2022) Pirk S, Lee E, Xiao X, Takayama L, Francis A and Toshev A (2022) A protocol for validating social navigation policies. arXiv preprint arXiv:2204.05443 . Poddar et al. (2023) Poddar S, Mavrogiannis C and Srinivasa S (2023) From crowd motion prediction to robot navigation in crowds. In: Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). p. 6765â6772. Raj et al. (2024) Raj AH, Hu Z, Karnan H, Chandra R, Payandeh A, Mao L, Stone P, Biswas J and Xiao X (2024) Rethinking social robot navigation: Leveraging the best of two worlds. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, p. 16330â16337. Robicquet et al. (2016) Robicquet A, Sadeghian A, Alahi A and Savarese S (2016) Learning social etiquette: Human trajectory understanding in crowded scenes. In: Leibe B, Matas J, Sebe N and Welling M (eds.) Comput. Vis. â ECCV 2016. ISBN 978-3-319-46484-8, p. 549â565. Rösmann et al. (2017) Rösmann C, Hoffmann F and Bertram T (2017) Integrated online trajectory planning and optimization in distinctive topologies. Robotics and Autonomous Systems 88: 142â153. Rudenko et al. (2020a) Rudenko A, Kucner TP, Swaminathan CS, Chadalavada RT, Arras KO and Lilienthal AJ (2020a) Thör: Human-robot navigation data collection and accurate motion trajectories dataset. IEEE Trans. Robot. Autom. 5(2): 676â682. Rudenko et al. (2020b) Rudenko A, Palmieri L, Herman M, Kitani KM, Gavrila DM and Arras KO (2020b) Human motion trajectory prediction: A survey. The International Journal of Robotics Research 39(8): 895â935. Samavi et al. (2025) Samavi S, Lem A, Sato F, Chen S, Gu Q, Yano K, Schoellig AP and Shkurti F (2025) Sicnav-diffusion: Safe and interactive crowd navigation with diffusion trajectory predictions. IEEE Robotics and Automation Letters . Schreiter et al. (2025) Schreiter T, Rodrigues de Almeida T, Zhu Y, Gutierrez Maestro E, Morillo-Mendez L, Rudenko A, Palmieri L, Kucner TP, Magnusson M and Lilienthal AJ (2025) Thör-magni: A large-scale indoor motion capture recording of human movement and robot interaction. The International Journal of Robotics Research 44(4): 568â591. Shah et al. (2023) Shah D, Sridhar A, Dashora N, Stachowicz K, Black K, Hirose N and Levine S (2023) Vint: A foundation model for visual navigation. arXiv preprint arXiv:2306.14846 . Shi et al. (2023) Shi L, Wang L, Zhou S and Hua G (2023) Trajectory unified transformer for pedestrian trajectory prediction. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 9675â9684. Singamaneni et al. (2024) Singamaneni PT, Bachiller-Burgos P, Manso LJ, Garrell A, Sanfeliu A, Spalanzani A and Alami R (2024) A survey on socially aware robot navigation: Taxonomy and future challenges. The International Journal of Robotics Research 43(10): 1533â1572. Song et al. (2024) Song D, Liang J, Payandeh A, Raj AH, Xiao X and Manocha D (2024) Vlm-social-nav: Socially aware robot navigation through scoring using vision-language models. IEEE Robotics and Automation Letters . Sorokowska et al. (2017) Sorokowska A, Sorokowski P, Hilpert P, Cantarero K, Frackowiak T, Ahmadi K, Alghraibeh AM, Aryeetey R, Bertoni A, Bettache K et al. (2017) Preferred interpersonal distances: A global comparison. Journal of cross-cultural psychology 48(4): 577â592. Sridhar et al. (2024) Sridhar A, Shah D, Glossop C and Levine S (2024) Nomad: Goal masked diffusion policies for navigation and exploration. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). IEEE, p. 63â70. Stratton et al. (2025) Stratton A, Hauser K and Mavrogiannis C (2025) Characterizing the complexity of social robot navigation scenarios. IEEE Robotics and Automation Letters 10(1): 184â191. Stratton et al. (2026a) Stratton A, Singamaneni PT, Goyal P, Alami R and Mavrogiannis C (2026a) Bi3: A biplatform, bicultural, biperson dataset for social robot navigation. In: Proceedings of the IEEE International Conference on Robotics and Automation (ICRA). Stratton et al. (2026b) Stratton A, Singamaneni PT, Goyal P, Alami R and Mavrogiannis C (2026b) How human motion prediction quality shapes social robot navigation performance in constrained spaces. Proceedings of the ACM/IEEE International Conference on Human Robot Interaction (HRI) . Svenstrup et al. (2010) Svenstrup M, Bak T and Andersen HJ (2010) Trajectory planning for robots in dynamic human environments. In: 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, p. 4293â4298. Takayama and Pantofaru (2009) Takayama L and Pantofaru C (2009) Influences on proxemic behaviors in human-robot interaction. In: 2009 IEEE/RSJ international conference on intelligent robots and systems. IEEE, p. 5495â5502. Taylor et al. (2022) Taylor AV, Mamantov E and Admoni H (2022) Observer-aware legibility for social navigation. In: 2022 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, p. 1115â1122. TorrejĂłn et al. (2024) TorrejĂłn A, Zapata N, Bonilla L, Bustos P and NĂșñez P (2024) Design and development of shadow: A cost-effective mobile social robot for human-following applications. Electronics (Switzerland) 13(17). Trautman and Krause (2010) Trautman P and Krause A (2010) Unfreezing the robot: Navigation in dense, interacting crowds. In: 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems. IEEE, p. 797â803. Trautman et al. (2015) Trautman P, Ma J, Murray RM and Krause A (2015) Robot navigation in dense human crowds: Statistical models and experimental studies of humanârobot cooperation. The International Journal of Robotics Research 34(3): 335â356. Van Den Berg et al. (2011) Van Den Berg J, Snape J, Guy SJ and Manocha D (2011) Reciprocal collision avoidance with acceleration-velocity obstacles. In: 2011 IEEE International Conference on Robotics and Automation. IEEE, p. 3475â3482. Wang et al. (2022a) Wang A, Mavrogiannis C and Steinfeld A (2022a) Group-based motion prediction for navigation in crowded environments. In: Conference on Robot Learning. PMLR, p. 871â882. Wang et al. (2024a) Wang A, Sato D, Corzo Y, Simkin S, Biswas A and Steinfeld A (2024a) Tbd pedestrian data collection: Towards rich, portable, and large-scale natural pedestrian data. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). p. 637â644. 10.1109/ICRA57147.2024.10610335. Wang et al. (2022b) Wang C, Wang Y, Xu M and Crandall DJ (2022b) Stepwise goal-driven networks for trajectory prediction. IEEE Robotics and Automation Letters 7(2): 2716â2723. Wang et al. (2024b) Wang J, Liu D, Chen J, Da J, Qian N, Man TM and Soh H (2024b) Genie: A generalizable navigation system for in-the-wild environments. arXiv preprint arXiv:2506.17960 . Wang et al. (2023) Wang W, Wang R, Mao L and Min BC (2023) Navistar: Socially aware robot navigation with hybrid spatio-temporal graph transformer and preference learning. In: 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, p. 11348â11355. Xiao et al. (2026) Xiao L, Song D, Xiao X and Yamasaki T (2026) E-socialnav: Efficient socially compliant navigation with language models. In: ICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, p. 20077â20081. Xie and Dames (2023) Xie Z and Dames P (2023) Drl-vo: Learning to navigate through crowded dynamic scenes using velocity obstacles. IEEE Transactions on Robotics 39(4): 2700â2719. Xu and Zhang (2021) Xu W and Zhang F (2021) Fast-lio: A fast, robust lidar-inertial odometry package by tightly-coupled iterated kalman filter. IEEE Robotics and Automation Letters 6(2): 3317â3324. Yan et al. (2017) Yan Z, Duckett T and Bellotto N (2017) Online learning for human classification in 3d lidar-based tracking. In: 2017 IEEE/RSJ International Conf. on Intell. Robots and Syst. (IROS). IEEE, p. 864â871. Yan et al. (2020) Yan Z, Schreiberhuber S, Halmetschlager G, Duckett T, Vincze M and Bellotto N (2020) Robot perception of static and dynamic objects with an autonomous floor scrubber. Intelligent Service Robotics 13(3): 403â417. Yuan et al. (2021) Yuan Y, Weng X, Ou Y and Kitani KM (2021) Agentformer: Agent-aware transformers for socio-temporal multi-agent forecasting. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 9813â9823. Zhang et al. (2022) Zhang Y, Sun P, Jiang Y, Yu D, Weng F, Yuan Z, Luo P, Liu W and Wang X (2022) Bytetrack: Multi-object tracking by associating every detection box. In: Avidan S, Brostow G, CissĂ© M, Farinella GM and Hassner T (eds.) Computer Vision â ECCV 2022. Cham: Springer Nature Switzerland. ISBN 978-3-031-20047-2, p. 1â21. Zhou et al. (2012) Zhou B, Wang X and Tang X (2012) Understanding collective crowd behaviors: Learning a mixture model of dynamic pedestrian-agents. In: 2012 IEEE Conf. on Comput. Vis. and Pattern Recognit. IEEE, p. 2871â2878. Zhu et al. (2026) Zhu Y, Yang SM, Magnusson M and Wang A (2026) Hicrowd: Hierarchical crowd flow alignment for dense human environments. In: 2026 IEEE International Conference on Robotics and Automation (ICRA). IEEE. Table 7: Comparison with other Social Navigation Datasets Dataset # Trajectories Trajectory Duration (min) Sensors Nav Method # Robots Location (in/out) TBD (Wang et al., 2024a) 20 264 3D LiDAR, RGB-D Camera, 360 camera, IMU Human operated 1 Indoors SCAND (Karnan et al., 2022a) 138 522 3D LiDAR, RGB-D Camera, Wheel Odometry, Visual Odometry Teleop 2 Both MuSoHu (Nguyen et al., 2023) 285 1138 3D LiDAR, RGB-D Camera, IMU, Visual Odometry, 360 Camera, Microphone Human walking 0 Both HuRON (Hirose et al., 2023) 541 4500 2D LiDAR, 360 camera, Fisheye camera, Wheel Odometry Autonomous 1 Indoors L-CAS (Yan et al., 2017) 3 49 3D LiDAR Teleop 1 Indoors FLO-BOT (Yan et al., 2020) 6 27.5 3D LiDAR, RGB-D camera, Stereo Camera, 2D LiDAR, OEM incremental measuring wheel encoder, IMU Autonomous 1 Indoors THOR (Rudenko et al., 2020a) 600 60 3D LiDAR, Motion capture system, Eye-tracking Glasses Autonomous 1 Indoors JRDB (Martin-Martin et al., 2021) 54 64 3D LiDAR, 2D LiDAR, Omnidirectional Stereo Suite, RGB camera, RGB-D stereo camera, 6D IMU Teleop 1 Both Go Stanford 2 (Hirose et al., 2018) N/A 1002 RGB-D Camera, Fisheye camera, Wheel encoders, Teleop 1 Both Crowdbot (Paez-Granados et al., 2022a) 110 200+ 3D LiDAR, RGB-D Camera, Odometry, contact(Force/Torque sensor) Human operated, Autonomous 1 Outdoor SIT (Bae et al., 2023) N/A 20+ 3D LiDAR, 5x RGB Cameras, IMU, GPS-RTK Teleop 1 Both CityWalker (Liu et al., 2025) N/A 900 3D LiDAR, RGB Cameras, GPS Teleop 1 Outdoor PeRoI (Agrawal et al., 2026) 18699 67.2 2D trajectory from birdâs eye camera Teleop 3 Outdoor Bi3 (Stratton et al., 2026a) 185 630 Motion Capture System, RGB Camera, Depth Camera Autonomous 2 Indoor ACME 3013 1761 3D-LiDAR, RGB(D) Camera, Overhead BEV Cameras, Odometry Human operated, Teleop 7 Both 12 Appendix 12.1 Miscellaneous Details 12.1.1 Data Curation Due to the distributed, multi-site, and multi-embodiment nature of ACME, the raw data streams differ slightly across subsets in terms of sensor availability, calibration quality, and recovery procedures. While all subsets follow the common data-collection protocols described in the main paper, some recordings experienced hardware or software issues that affected specific modalities. We document these subset-specific notes here to support reproducibility and to help users select the appropriate subset and modalities for downstream tasks. âą NUS: To account for sensor misalignment as an effect of robot transport to different locations, we provide session-wise manually corrected calibration values, as corrections on top of the default calibration parameters. âą Miraikan and Keio: 2.2 hours of additional data do not contain image streams (due to an image-sensor connection failure). These recordings (which still contain LiDAR and Odometry data) are not included in the total dataset duration reported in the main paper. âą CMU: Robot odometry was lost due to a hardware failure. Robot pose was recovered post-hoc using LiDAR-inertial odometry (Xu and Zhang, 2021). However, velocity information is not available in the released processed data and would require additional post-processing to recover. âą UEx: Velocity information was lost due to an issue with the ZED 2i ROS 2 driver. We recovered velocity estimates during post-processing using ICP over dense LiDAR scans. âą UBonn: A small number of trajectories contain camera motion caused by a mounting issue, which can invalidate the corresponding camera calibration. Another fraction of trajectories contains âengagementâ type interaction â the robot intentionally interacts with pedestrians to elicit reactions (similar to Hirose et al. (2023)). These trajectories have been tagged appropriately in the released metadata. Sensor calibration. We additionally provide calibration information for the available onboard sensors in each subset. Calibration procedures varied across platforms depending on the robot hardware, sensor mounting, and information available from the robot or sensor manufacturers. We summarize the subset-specific calibration procedures below. âą NUS: We use the default calibration provided by the robot manufacturer for the body-to-LiDAR and body-to-camera transforms. Camera intrinsics were estimated using the standard OpenCV camera calibration pipeline. âą Miraikan, Keio, CMU: Camera intrinsics were obtained from the sensor manufacturer. Static extrinsic transforms between the robot body and onboard sensors were set based on physical measurements of the sensor mounting configuration. âą GMU: Sensor extrinsics were manually calibrated using the known dimensions of the 3D-printed mounting components. âą UMich: The Stretch platform provides a default camera transform through its robot model. For the LiDAR, we used motion capture to register the center of the LiDAR and then computed the transform between the Stretch coordinate frame and the LiDAR coordinate frame in the motion-capture system. âą UEx: Sensor extrinsics were obtained from the SolidWorks design of the robot platform. Camera intrinsics were estimated using OpenCV calibration utilities: the standard pinhole camera calibration pipeline for the ZED camera and the fisheye calibration pipeline for the 360-degree camera. âą UBonn: Calibration information is provided for the available onboard sensors. As noted above, trajectories in which the camera moved due to the mounting issue have been tagged, since the corresponding camera extrinsics may be invalid for those sequences. Time-Synchronization between BEV and onboard data In order to time-synchronize the BEV camera feed with the on-board robot data, we used a QR-code-based method, which allowed post-hoc temporal syncing of the two data streams. This accommodated the use of commercial BEV cameras with limited SDKs (where direct integration with our ROS-based setup was infeasible). To achieve temporal synchronization across these disconnected systems, the NUS, CMU, U-Mich, GMU, and UBonn teams displayed a dynamic QR code that encoded the current local timestamp to each image sensor at the beginning of a recording session (Fig. 19). This permitted alignment of the data streams during post-processing, which facilitated temporal consistency across all image-based sensor streams that were not directly integrated into the ROS network. For Miraikan and Keio, time synchronization was achieved with manual frame selection based on recognizable events (e.g., the frame where someone is about to pick up the tag on the floor). Once a frame with an event/tag was found in both data streams, the rest of the data can be time-synchronized by leveraging ROS bag timestamps for egocentric data and BEV dataâs stable frame rate to obtain timestamps for any given frame. Figure 19: Time Synchronization with timestamp embedded QR Codes