Paper deep dive
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, Toshihiko Yamasaki
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/12/2026, 2:38:50 AM
Summary
The paper introduces 360CityArena, a photorealistic virtual urban navigation benchmark for embodied agents, constructed from 602 360-degree video segments covering 85 streets in the Akihabara district of Tokyo. The benchmark comprises 175 human-crafted tasks across three categories: Environment Understanding, Path Reasoning, and Spatial Reasoning. Evaluations using state-of-the-art Large Multimodal Models (LMMs) like Gemini 2.5 Flash reveal significant performance gaps compared to human capabilities, highlighting challenges in city-scale embodied navigation.
Entities (17)
Relation Signals (9)
360CityArena â reconstructs â Akihabara
confidence 98% · 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan
360CityArena â contains â Environment Understanding
confidence 95% · It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning
360CityArena â contains â Path Reasoning
confidence 95% · It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning
360CityArena â contains â Spatial Reasoning
confidence 95% · It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning
Gemini 2.5 Flash â evaluatedon â 360CityArena
confidence 95% · Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level
360CityArena â evaluates â Embodied Agents
confidence 95% · 360CityArena is a benchmark for evaluating the urban exploration capabilities of embodied agents
360CityArena â createdby â Kenta Watanabe
confidence 90% · Kenta Watanabe... The University of Tokyo... We present 360CityArena
360CityArena â createdby â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a considerable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360-degree video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial reasoning, thereby enabling comprehensive evaluation in realistic urban scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing substantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning.
Tags
Links
- Source: https://arxiv.org/abs/2608.08814v1
- Canonical: https://arxiv.org/abs/2608.08814v1
Trouble viewing inline? Open PDF directly â
Full Text
63,674 characters extracted from source content.
Expand or collapse full text
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Kenta Watanabe, Atsuyuki Miyai, Mizuki Takenawa, Kiyoharu Aizawa, and Toshihiko Yamasaki The University of Tokyo, Tokyo, Japan k_watanabe,miyai,yamasaki@cvm.t.u-tokyo.ac.jp takenawa,aizawa@hal.t.u-tokyo.ac.jp https://360m-team.github.io/360CityArena/ Akihabara, reconstructed from 360°video How can I help you? AI ASSISTANT Go to Animate USER Fig. 1: 360CityArena. We introduce a benchmark for evaluating embodied agents in a photorealistic reconstruction of Akihabara, Tokyo, Japan, built from interconnected 360° video trajectories. The benchmark covers realistic urban streets and evaluates agents on diverse tasks requiring environment understanding, path reasoning, and spa- tial reasoning. Abstract. We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360° videos. Existing outdoor benchmarks either lack sufficient photorealism or complexity, resulting in a consid- erable gap from real-world urban environments. 360CityArena is built on a realistic reconstruction of the Akihabara district in Tokyo, Japan, using 602 360° video segments covering 85 streets, and consists of 175 meticulously human-crafted tasks. It encompasses three task categories: Environment Understanding, Path Reasoning, and Spatial Reasoning, covering fundamental abilities required for urban exploration, such as localization, landmark search, path planning, and relational spatial rea- soning, thereby enabling comprehensive evaluation in realistic urban arXiv:2608.08814v1 [cs.CV] 9 Aug 2026 2K. Watanabe et al. scenes. Our evaluation using state-of-the-art LMM-based agents shows that even the strongest model, Gemini 2.5 Flash, performs far below human level (human: 77.3% vs. Gemini 2.5 Flash: 17.1%), revealing sub- stantial challenges that remain in city-scale embodied navigation and reasoning. 360CityArena provides a necessary and challenging testbed for photorealistic urban-district navigation and spatial reasoning. Keywords: Embodied AI· Urban Navigation Benchmark· Photoreal- istic Virtual Environments 1 Introduction Imagine a future where an AI assistant is a natural part of city life. It can help you when you get lost, guide blind people to where they need to go, and provide support whenever someone needs it. This kind of assistant would make cities more accessible, reduce the difficulties of traveling, and help create a more inclusive society. One promising approach toward achieving this vision is Em- bodied AI, in which agents perform complex tasks based on perception of their surrounding environment [1, 8, 10, 26]. Realizing this vision will require major advances in evaluating and testing such systems across a wide range of tasks within realistic urban environments that capture the visual complexity of real streets. A central problem in city exploration research is the lack of benchmarks that evaluate agents in environments that realistically reflect wide-area navigation in real urban spaces. Existing 3D simulators fall short in reconstructing city scenes with sufficient photorealism and complexity [9,12,17,44]. Although environments built from Google Street View provide realistic imagery [6,11,30,49], they lack dynamic elements and do not offer continuous, fully navigable spaces. To address these issues, recent studies have explored converting real-world urban videos into simulation environments, enabling more realistic and interactive experiences [45]. However, the video clips used in this method are short, making them unsuitable for exploring extended real street networks. In parallel, recent work has studied navigation from different modalities and viewpoints, such as aerial VLN over cities and map-only evaluation of route planning [22,34,46]. These lines of work are complementary, but they still do not capture ground-level, egocentric city exploration in realistic street scenes with rich dynamics. To further advance the field, we need a benchmark that covers tasks within highly realistic and dynamic urban street environments, while supporting smooth egocentric navigation. In this paper, we propose 360CityArena, a benchmark for assessing em- bodied agents in a photorealistic real-world urban district represented as an interconnected pose graph of 360° video trajectories. 360CityArena is built upon a Realistic Virtual World environment [41] that recreates Akihabara, a major urban district in Tokyo, and includes 175 manually created tasks. The environ- ment spans approximately 750 meters northâsouth and 650 meters eastâwest, covering 85 interconnected streets. Each task is grounded in real city structures 360CityArena3 Relational Spatial Reasoning (20) Env. Understanding Localization (25) Landmark Text Search (25) Landmark Image Search (25) Map Navigation (25) Vision Language Navigation (25) Object Count (25) Relational Spatial Reasoning (25) Path Reasoning Spatial Reasoning Explore the city and identify your starting grid cell on the map. Go to the âanimateâ. Go to the landmark shown in the image. Move from the BLUE start to the RED goal on the map. Follow the directions to reach your destination : 1. Go straight. 2. Turn left when you see Sukiya 3. Continue straight 4. Stop in front of Burger King How many vending machines are there along the road up to the next crosswalk? What store is three doors to the right of Kitchen Jiro? Answer : Arrived at (229.3, 595.9) Answer : x:1 y:2 Answer : Arrived at (495.7, 268.3) Answer : Arrived at (244.4, 392.8) Answer : Arrived at (580.0, 349.6) Answer : Akiba no X Answer : There are four. Fig. 2: Examples in each task type in 360CityArena. (i) Environment Under- standing tasks include Localization, Landmark Search with Language, and Landmark Search with Image, where the agent infers its location or navigates to a specified land- mark. (i) Path Reasoning tasks evaluate the agentâs ability to plan and execute routes, such as following map-based paths or vision-language navigation. (i) Spatial Reason- ing tasks assess relational understanding and quantitative perception, including iden- tifying spatial relations between landmarks and counting objects in the environment. and landmarks, encouraging agents to perform practical, exploration-based ac- tivities. As shown in Figure 2, 360CityArena consists of three task types: (i) Envi- ronment Understanding, (i) Path Reasoning, and (i) Spatial Reasoning. (i) En- vironment Understanding targets visual recognition and self-localization, mea- suring an agentâs perceptual understanding of the environment. (i) Path Rea- soning focuses on path planning and decision-making ability within the real spatial layout of the environment. (i) Spatial Reasoning evaluates geospatial reasoning through navigation. At a finer level, the benchmark systematically measures seven distinct capabilities. Furthermore, each task is categorized into Easy, Medium, and Hard, making it possible to evaluate agents according to difficulty level. Through these systematically constructed tasks, 360CityArena enables a comprehensive assessment of an Embodied Agentâs ability to perceive, reason, and act in real-world urban environments. Our experiments show that state-of-the-art LMM-based agents remain far below human performance across all tasks, especially on Map Navigation, Ob- ject Count, and Relational Spatial Reasoning. This demonstrates both the value of 360CityArena as a comprehensive benchmark for realistic urban settings and the need for stronger city-scale agents. We also evaluate removing explicit self- location information (e.g. a map). Performance does not degrade consistently and even improves in some tasks, suggesting that the manner of providing loca- tion priors can affect exploration strategies and failure modes, informing future model design. The contributions of our paper are summarized as follows: 4K. Watanabe et al. Table 1: Comparison of urban navigation environments â realism, structural complex- ity, dynamics, interactivity, and exploration capability. EnvironmentCategory Photo- realism Structural complexity Dynamics Interaction District-scale exploration Motion EmbodiedCity [12]3D simulator LowLowMediumâContinuous MetaUrban [44]3D simulator Low Medium MediumâContinuous CARLA [9]3D simulator LowLowLowâContinuous (clip) Vid2Sim [45]Video-to-sim HighHighMediumââContinuous StreetLearn [30]GSV-basedHighHighLowââDiscrete 360CityArena (Ours) 360° videoHighHighHighââ Continuous (trajectory) â Proposing 360CityArena: We introduce 360CityArena, an urban-district benchmark built on a photorealistic Virtual World reconstruction of Tokyoâs Akihabara district from 360° videos. The benchmark provides seven task types across three categories: Environment Understanding, Path Reason- ing, and Spatial Reasoning, enabling comprehensive evaluation of embod- ied agents in a realistic urban environment with an interconnected street network. Beyond evaluation, 360CityArena would also serves as a practical guideline for constructing training data in photorealistic virtual worlds for embodied AI. â Benchmarking Recent LMMs: We build embodied agents based on sev- eral proprietary LMMs and open-source vision-language models, and bench- mark them on 360CityArena. The results reveal substantial performance gaps compared to humans across all tasks and clear performance degrada- tion as task difficulty increases. â Comprehensive Analysis: Through extensive qualitative and quantita- tive analyses, we investigate how input modality (language versus images), the presence of self-location information, and model- and task-specific fail- ure patterns influence agent behavior and performance. These analyses offer insights into the roles of visual cues and self-localization in urban embodied tasks. 2 Related Work Embodied Agent. An Embodied Agent is an AI system that follows language instructions and completes tasks by perceiving the environment through camera views and sensor information. Recent advances in LMMs have greatly improved multimodal integration and reasoning, making them a strong foundation for embodied agents [15,16,24,56]. Embodied-agent research spans real robots [19, 39,52], indoor simulators [21,37,51], and outdoor simulators [6,45]. While real robots are most realistic, data collection is costly and often environment-specific, limiting generalization. Simulation environments therefore attract attention for their diversity and reproducibility [10]. Indoor simulators ease data collection but have limited scale and simple layouts [37], motivating a shift toward larger, more complex outdoor settings with dynamic elements such as pedestrians and vehicles [30,49]. 360CityArena5 Navigation Agents in Outdoor Environments. Recent advances in simula- tion environments have broadened embodied navigation research beyond indoor settings [21,28,35â37,51] to large-scale outdoor environments. Outdoor naviga- tion tasks are commonly distinguished by how the goal is specified: point-goal navigation, where the agent is given target coordinates [25, 30, 44]; image-goal navigation, where the target is provided as a reference image [17,18]; object-goal navigation, where the goal is described in language (e.g. a landmark) [5,14]; and Vision-Language Navigation (VLN), where the agent follows natural-language instructions [6, 23, 49]. Some work instead evaluates route planning directly on maps without embodied simulation [34, 46, 49]. In parallel, CityNav [22] stud- ies real-world navigation from an aerial viewpoint, which is complementary but differs from ground-level exploration with egocentric street scenes. Separately, map understanding has been explored outside embodied navigation, including text-based map representations for grounding LLM planning (Tag Map [54]) and VLM benchmarks for advanced map queries (MAPWise [32]). While point-goal, image-goal, and object-goal tasks emphasize efficient target reaching via visual or semantic cues, VLN additionally demands higher-level rea- soning to resolve linguistic ambiguity and landmark references. However, aerial- VLN datasets and map-only evaluations do not jointly capture egocentric navi- gation and urban reasoning in realistic street environments. To address this gap, we introduce an urban-district benchmark that unifies these tasks with (i) photo- realistic, dynamic street-level observations, (i) multi-task diagnostics spanning environment understanding, path reasoning, and spatial reasoning, and (i) con- trolled comparisons between language- and image-goal landmark search under matched conditions. Real-world Simulation Environments. Prior work has explored 3D scene reconstruction, Google Street Viewâbased environments, and 3D simulators for building real-world simulation settings. For 3D scene reconstruction, methods built on Neural Radiance Fields [29] and Gaussian Splatting [20] have led to nu- merous extensions [4, 47, 48, 53]. Despite these advancements, most approaches remain limited to producing static scenes and do not provide fully interactive environments [45]. City-scale reconstructions in conventional 3D simulators can procedurally generate diverse variations and support rich agentâenvironment in- teractions, but fall short in photorealism and structural complexity [12,17,27,44]. Environments constructed from Google Street View (GSV) enable diverse and highly realistic city-scale representations that are difficult to achieve with tradi- tional 3D simulation [6,11,30,42,49]. However, these environments from GSV do not contain dynamic elements, and the inherent discontinuity between panora- mas creates a gap from real-world continuous navigation [30]. To address this, a recent effort has attempted to convert real urban videos into interactive simula- tion environments [45]. While this approach is promising, this environment relies on short video clips, making them unsuitable for large-scale city exploration. Tak- enawa et al. [41] introduced a Realistic Virtual World (RVW) generated from large collections of 360° videos, providing a photorealistic and dynamic city-scale environment with real pedestrian and vehicle motions. We leverage this RVW to 6K. Watanabe et al. Fig. 3: Example of the visual observations in 360CityArena. The agent moves through the city while scanning its surroundings. At branching points, it must choose a route. build a benchmark for realistic, city-scale exploration by embodied agents. While navigation within each filmed trajectory is continuous and smooth, discontinu- ities arise at trajectory boundaries, and the environment does not support physi- cal interaction because it is constructed from pre-recorded video. We summarize key properties of representative real-world urban environments in Table 1. 3 The 360CityArena We introduce 360CityArena, a novel benchmark meticulously designed to eval- uate an agentâs ability to explore a realistic urban district across a broad range of tasks (see Figure 3). The benchmark is built on a Realistic Virtual World reconstructed from the actual Akihabara district in Japan and consists of 175 tasks organized into three major categories and seven subcategories. We first describe the Akihabara virtual environment in Section 3.1. We then provide detailed explanations of each task category in Section 3.2. The task con- struction process is outlined in Section 3.3, followed by the evaluation protocol in Section 3.4. 3.1 Akihabara Virtual Environment The environment used in 360CityArena is a Realistic Virtual World of Akihabara constructed from 602 360° video segments over 85 streets [41]. The videos are projected onto a spherical surface and organized into a navigable pose graph in 360CityArena7 Unity, enabling agents to traverse the captured area. Agents therefore navigate along prerecorded 360° video trajectories represented as a pose graph, rather than moving freely to arbitrary 3D positions or physically interacting with the environment. The resulting pose graph forms a single connected component with 193 nodes and 305 edges, and has a mean branching degree of 3.16. In terms of geographic coverage, 360CityArena spans approximately 750 m northâsouth and 650 m eastâwest around Akihabara Station. Akihabara, origi- nally an electronics district, has become a major commercial and cultural hub with dense electronics and anime/game-related retail and tourist attractions. The area features diverse streetscapes with narrow sidewalks, complex building layouts, abundant outdoor ads, and electronic signboards, creating a visually dense, information-rich scene. Each street is captured twice (one per direction), reflecting direction-dependent visual changes. These characteristics make Aki- habara well suited for evaluating embodied agentsâ perception and navigation in a realistic urban simulation. Although 360CityArena currently includes only one city, the representation is general. Following the Movie Map paradigm [40], similar city-scale environ- ments can be built by geo-localizing 360° frames with GPS/SLAM and a map, identifying intersections, and connecting snippets into a traversable pose graph. Our city-agnostic graph augmented with 360° observations is directly compatible with such environments. 3.2 Task Categories Our benchmark consists of three major categories and seven subcategories, with examples shown in Figure 2. Each subcategory contains 25 tasks. We describe each category in detail below. 1. Environment Understanding. This category evaluates an agentâs ability to perceive and understand its surroundings, covering visual perception, memory, and language-grounded scene understanding. It includes three subcategories: Lo- calization, Landmark Search with Language (object-goal), and Landmark Search with Image (image-goal). In Localization, the agent infers its initial position on a 5Ă5 grid from visual observations. Unlike GeoGuessr-style geolocalization, which predicts broad regions from scenes [13], our task operates at a much smaller scale and relies on town layout and local landmarks. The two Landmark Search tasks require navigating to a specified landmark, provided either in natural language or as an image; aside from the input modality, they share identical settings, enabling controlled comparison of modality effects. 2. Path Reasoning. This category evaluates an agentâs ability to plan and explore based on spatial understanding. It measures not only the agentâs plan- ning and decision-making abilities, but also its ability to execute those plans in accordance with the spatial layout of the environment. The subcategories in this task are Map Navigation and Vision-Language Navigation (VLN). Map Naviga- tion requires the agent to choose an optimal route using map information and navigate from a specified start point to a designated goal point. Vision-Language 8K. Watanabe et al. Navigation requires the agent to interpret and follow multi-step natural-language instructions and move to the target accordingly. 3. Spatial Reasoning. This category evaluates an agentâs ability to under- stand the structural layout of objects and the positional relationships between landmarks in the environment. It measures core spatial perception and logical reasoning capabilities. The subcategories in this task are Relational Spatial Rea- soning and Object Count. Relational Spatial Reasoning requires the agent to infer relative spatial relationships between landmarks. Given a reference land- mark and a specified relation, the agent must identify the landmark that satisfies that relation. Object Count is a quantity-estimation task in which the agent must accurately count the number of specified objects within a designated area. 3.3 Benchmark Construction Our construction involves eight annotators. We assigned two annotators to each of the seven sub-tasks. To ensure consistency in task quality across different tasks, one of the authors was assigned to all tasks. To ensure the quality of the benchmark, we selected annotators who had actually visited Akihabara in person. They interacted with the environment directly in Unity and created tasks, and verified that each task was solvable. The overall task construction process required approximately 60 hours. Each task is labeled Easy, Medium, or Hard by annotators based on distance, instruction ambiguity, landmark visibility, and required exploration, rather than fixed step-count thresholds. To support these labels, task-relevant statistics show monotonic increases from Easy to Hard: Map Navigation path length and de- cision points are 124/247/347 m and 3.3/6.4/9.0, Object Count ground-truth counts are 3.5/4.6/9.0, and VLN instruction steps are 4.4/5.9/7.4. Borderline cases may still involve subjectivity. 3.4 Evaluation Protocol We define a set of evaluation criteria to assess agent performance. Our benchmark adopts four evaluation protocols, each designed to capture different aspects of task success. We detail each protocol and its corresponding task categories below. 1. exact_match. exact_match is applied to tasks whose outputs are explicitly defined as numerical values (grid coordinates) or textual strings [31,55]. Under this metric, the agentâs final textual output is evaluated, and the prediction is considered correct only when it exactly matches the answer text. In our bench- mark, the Localization task uses this evaluation metric. 2. fuzzy_match. fuzzy_match leverages a language model (GPT-5 in our implementation) to assess whether the output is semantically equivalent to the ground truth [31, 55]. Importantly, the GPT-5 evaluator is fully isolated from the agent and only compares the final output with the ground-truth answer, preventing any information leakage. In the Relational Spatial Reasoning task, we employed this evaluation metric and additionally performed manual verification. The GPT-5 judge achieved 97.9% agreement with human majority votes (Îș = 360CityArena9 0.937), close to the inter-annotator agreement (Îș = 0.984), supporting its use for evaluation automation. 3. coordinate_match. coordinate_match is used for tasks in which the agent must output its final position after moving through the environment. The coordi- nates are obtained directly from Unity. This metric is applied to Map Navigation, VLN, and Landmark Search with Image / Language. For all of these tasks, the agentâs final position is evaluated by measuring the Euclidean distance to the target location, and the prediction is considered correct if it falls within a thresh- old of Δ. For the distance threshold Δ, we set Δ = 10 m for all tasks except Map Navigation, where a larger threshold of Δ = 20 m is used to account for noise caused by the markerâs size on the map. A threshold-sensitivity check showed stable model rankings for 0.75Ăâ1.5Ă thresholds (Kendallâs Ï â„ 0.89), a hu- manâbest LMM gap above 36 p even at Δ = 30m, and Map Navigation changes within ±2 p for Δ = 15â25 m. 4. mean_relative_accuracy (MRA). MRA is a flexible evaluation metric applied to tasks that involve numerical estimation [50]. In our benchmark, this corresponds to the Object Count task. MRA averages the relative accuracy across a range of evaluation thresholds C =0.5, 0.55,..., 0.95: MRA = 1 10 X ΞâC 1 |Ëyâ y| y < 1â Ξ ,(1) where Ëy denotes the predicted value and y the ground truth. MRA accounts for the magnitude of error, so it provides a more appropriate evaluation than a simple binary correct/incorrect classification. 4 Navigation Agents 4.1 Problem Formulation The environment-agent interaction can be modeled as a partially observable se- quential decision process: E = (S,A,âŠ,T), where S represents the set of states, A represents the set of actions, ⊠represents the set of observations. The transi- tion function is defined as T : SĂAâ S, with deterministic transitions between states conditioned on actions. At each time step t, the environment is in some state s t (e.g. a specific position and viewing direction). The agent receives a partial observation o t â âŠ, which consists of the visual input captured at time t and, optionally, the current positional information shown as an image with a marker indicating the agentâs location and orientation on the map. The agent also maintains a memory buffer M t âM that stores important information from previous steps up to tâ 1. The agent then issues an action a t â A conditioned on both o t and the stored memory M t , which results in a new state s t+1 â S and a new observation o t+1 â ⊠from the updated viewpoint. Simultaneously, relevant information from o t and thoughts is written to the memory, updating it to M t+1 . 10K. Watanabe et al. 4.2 Baseline Agents For our experiments, we evaluate multiple LMMs as baseline agents: GPT- 5 [33], Claude Sonnet 4.5 (20250929) [2], Gemini 2.5 Flash [7], Qwen2.5-VL- 32B-Instruct [3], and InternVL3.5-8B and InternVL3.5-38B [43]. Among these models, GPT-5, Claude Sonnet 4.5 and Gemini 2.5 Flash are closed-source, while Qwen2.5-VL-32B-Instruct, InternVL3.5-8B, and InternVL3.5-38B are open-source. For inference with the open-source models, we used eight NVIDIA A100 80GB GPUs. In addition, all tasks were executed in Unity 6000.0.30f1 (Unity 6 LTS) running on macOS. In our setting, the action space A consists of seven discrete actions: moving forward, tilting the viewpoint upward, tilting the viewpoint downward, rotating the viewpoint to the right, rotating the viewpoint to the left, resetting the view- point to align with the current heading direction, and outputting an answer. At branching points, the action space is augmented with movement actions corre- sponding to the available traversable directions, such as going forward, taking the right branch, or taking the left branch. 5 Experiment 5.1 Experimental Results Human Performance. We conducted human evaluations after obtaining ap- proval from our institutionâs Institutional Review Board (IRB). We recruited five participants, including both undergraduate and graduate students. We selected participants who had been to Akihabara before or who visit the area regularly. Since real-world deployment in a specific urban district can benefit from local familiarity, we report this result as a local-expert human baseline, which serves as an in-domain upper-bound reference rather than a generic human baseline. As shown in Table 2, human participants achieved higher accuracy than current LMMs. For tasks such as Map Navigation, Landmark Search with Im- age, Vision-Language Navigation, and Relational Spatial Reasoning, humans reached around 90% accuracy. In contrast, performance on Localization, Land- mark Search with Language, and Object Count was more modest, at 68%, 64%, and 45%, respectively. In Landmark Search, the task configurations for the Lan- guage and Image variants are identical except for the input modality. The fact that participants achieved much higher accuracy in the Image condition indi- cates that visual inputs contain substantially richer information than textual descriptions, including cues about appearance, location, and overall scene con- text. Next, our main findings are as follows: F1: LMMs perform far below human level. As shown in Table 2, all LMMs fall far short of human performance across the benchmark. Gemini 2.5 Flash achieves the highest overall performance among the evaluated models, while the strongest model varies across individual tasks (e.g., GPT-5 on Landmark (Img) 360CityArena11 Table 2: Overall Results (%). Comparison of model and human performance across seven spatial and reasoning tasks, grouped into three major categories. Gemini 2.5 Flash achieves the highest overall performance, while the strongest model varies across individual tasks. All LMMs still fall far short of human performance. Environment Understanding Path Reasoning Spatial Reasoning Loc Landmark (Lang) Landmark (Img) Map Nav VLN Obj Count Rel Reason GPT-58.016.048.00.08.02.432.0 Claude Sonnet 4.54.04.016.04.04.010.88.0 Gemini 2.5 Flash12.028.036.00.08.024.012.0 Qwen2.5-VL-32B-Instruct 4.016.020.00.00.018.84.0 InternVL3.5-8B4.020.020.00.012.02.80.0 InternVL3.5-38B0.016.00.00.012.07.24.0 Human68.064.092.092.088.045.292.0 EasyMediumHard Success rate (%) Environment Understanding Path Reasoning Spatial Reasoning GPT - 5 Gemini - 2.5 Flash Claude Sonnet 4.5 Qwen2.5 - VL - 32B - Instruct InternVL3.5 - 8B InternVL3.5 - 38B Human GPT - 5 Gemini - 2.5 Flash Claude Sonnet 4.5 Qwen2.5 - VL - 32B - Instruct InternVL3.5 - 8B InternVL3.5 - 38B Human GPT - 5 Gemini - 2.5 Flash Claude Sonnet 4.5 Qwen2.5 - VL - 32B - Instruct InternVL3.5 - 8B InternVL3.5 - 38B Human Fig. 4: Success rate by difficulty level across tasks and models (%). This graph shows model and human performance for each task, divided into three difficulty levels: Easy (E), Medium (M), and Hard (H). and Rel Reason). These results indicate that, despite incremental progress, sub- stantial gaps to human-level environment understanding and spatial reasoning remain. F2: Image-based landmark search can be easier than language-based search. Across models, Landmark Search with Image generally outperforms Landmark Search with Language (e.g., GPT-5: 48.0 vs. 16.0; Claude: 16.0 vs. 4.0; Qwen: 20.0 vs. 16.0), suggesting that visual inputs provide concrete cues (appear- ance, textures, and surrounding context) that more directly support navigation than abstract textual descriptions. However, this improvement is not universal. InternVL does not show a clear gain from image inputs (8B: 20.0 vs. 20.0; 38B: 0.0 vs. 16.0), indicating limitations in effectively leveraging landmark images. By including both image- and language-based variants, our benchmark enables analysis of how models leverage different input modalities. F3: Performance decreases as task difficulty increases. As shown in Fig- ure 4, model performance generally declines with higher task difficulty. For exam- ple, GPT-5âs accuracy in Environment Understanding drops from 28.0 â 25.0 â 12K. Watanabe et al. Table 3: Comparison with and without location information. We observe that the performance did not improve consistently across tasks; in some cases, accuracy decreased instead. Environment Understanding Path Reasoning Spatial Reasoning Loc Landmark (Lang) Landmark (Img) Map Nav VLN Obj Count Rel Reason GPT-58.016.048.00.08.02.432.0 GPT-5 (w/o location) -24.044.04.08.00.048.0 GPT-5 InternVL3.5-38B Gemini-2.5 Flash Environment UnderstandingSpatial ReasoningPath Reasoning ActionGrounding Perception ExplorePlanning Failureshare(%) Model Fig. 5: Failure cause breakdown by task category across models (%). This fig- ure shows the distribution of failure causes for each model across three task categories. Failures are categorized into five types: Action, Grounding, Perception, Explore, and Planning, and each stacked bar reports the percentage breakdown within the corre- sponding (model, task category) setting. 18.5, and in Path Reasoning from 11.1 â 0.0 â 0.0 across the Easy, Medium, and Hard settings. This demonstrates that the benchmark allows meaningful comparison of model performance across different difficulty levels. 5.2 Analysis Effect of Location Information. As shown in Table 3, we examined how explicitly providing the agent with its current location affects its spatial under- standing and reasoning abilities, and found that adding a location prior yields inconsistent performance improvements across tasks. In Environment Under- standing, Landmark (Language) accuracy dropped from 24.0% to 16.0%, while Landmark (Image) showed no meaningful change (within the margin of error). Path Reasoning tasks (Map Navigation and VLN) exhibited nearly no differ- ence, suggesting static location information fails to aid path planning or in- struction interpretation. In Spatial Reasoning, Object Count improved slightly, but Relational Reasoning fell from 48.0% to 32.0%. These performance degra- dations suggest agents struggle to align map-based location and relational in- formation with real-world cues derived from visual inputs, rather than location information being inherently unhelpful. While prior work addresses map-based 360CityArena13 reasoning [34,38,46], the impact of the map-to-visual alignment process remains under-examined. Future research should independently evaluate this alignment capability and, where necessary, provide mechanisms for learning it. Distinct Failure Patterns across Models. To better understand failure factors for three major models (GPT-5, Gemini 2.5 Flash, and InternVL3.5- 38B), we used Gemini 3 Flash to analyze failures from viewpoint images, maps, and thought histories, assigning each case to one of five categoriesâAction, Grounding, Perception, Explore, and Planningâand then had humans verify the label and select the single dominant cause. Action denotes low-level ac- tion errors (e.g., overshoot, wrong action, premature answer); Grounding, self- localization/orientation inconsistencies; Perception, object-recognition failures; Explore, stagnation or loops; and Planning, missed instructions or requirements. The resulting breakdown by model and task category is shown in Figure 5. These results should be interpreted as diagnosing integrated embodied urban perfor- mance, rather than isolating pure spatial reasoning: Action and Explore failures indicate that current LMM agents also suffer from controller and prompting brittleness when converting perception and reasoning into navigation decisions. The results reveal distinct failure patterns reflecting model capabilities. First, GPT-5 contrasts sharply with other models: it shows minimal Action failures (12% in Environment Understanding) but dominant Explore failures (55%), indicating effective basic execution but poor exploration strategy. Conversely, Gemini 2.5 Flash and InternVL3.5-38B struggle primarily with Action failures (around 40%), pointing to deficits in low-level control. Second, Perception fail- ures surge in Spatial Reasoning across all models (reaching 38% for Gemini), confirming that overlooking minute visual details is fatal for spatial tasks. Fur- thermore, Gemini displays consistent Grounding errors (15â21%) across cate- gories, indicating persistent difficulty in aligning map data with the first-person perspective. Case Study of Successes and Failures. For Landmark Search tasks in which the Language variant fails but the Image variant succeeds, we analyze individ- ual cases using GPT-5, examining the agentâs visual observations. The example shown here is a task in which the agent must locate Jonathanâs restaurant. In the Landmark Search with Language failure case in Figure 6, the agent passes by Hotto Motto and hypothesizes that Jonathan is located above it, prompting it to look upward. After failing to find the target, it shifts its gaze right and continues scanning. As a result, the agent stays in the same location, repeatedly changing viewpoints until the episode ends due to the step limit. In contrast, as shown in Figure 7, in Landmark Search with Image, the tar- get image provides cues such as its street-corner location and distinctive blue- and-white vertical facade with a red sign. The agent therefore disregards Hotto Motto, moves toward the major intersection, and successfully finds the target building. 14K. Watanabe et al. t = 7. HottoMotto is on the right.t = 8. Look above. t = 40. Look to the right.t = 41. Look below. Fig. 6: Example of the agentâs views for Landmark Search with Language, resulting in failure. The agent is instructed to search for âJonathanâ. From t = 8 to t = 38, the gaze repeatedly moved up and down. At t = 39, the gaze briefly shifted to the right, but from t = 41 onward, it returned to the same upâdown movement. 6 Limitations and Future Work Restricted Exploration and Limited Interaction. Because 360CityArena is constructed from pre-recorded 360° videos, agents can explore only along captured trajectories rather than move freely to arbitrary locations. As a re- sult, transitions across trajectory boundaries may introduce discontinuities that would not arise in fully interactive 3D simulators or in the real world. In our logs, boundary-transition actions accounted for 11.3% of all actions and were not overrepresented among Action or Explore failures; manual inspection also did not reveal clear discontinuity-induced failures. Nonetheless, we argue that situated visual navigation and spatial reasoning in a highly photorealistic urban setting remain challenging and important for current LMM-based agents. Ex- tending the environment to support freer exploration and richer interaction is an important direction for future work. Limited Urban Diversity. A limitation of the current benchmark is its restric- tion to a single urban district, which may limit urban diversity and conclusions regarding cross-regional generalization. However, we consider that 360CityArena provides a uniquely photorealistic, large-scale setting that facilitates more real- istic urban evaluations than existing benchmarks. We leave cross-regional gen- eralization as an important direction for future work. 360CityArena15 t = 4. HottoMotto is on the right.t = 5. Go past. t = 16. Spot blue and white vertical stripes.t = 29. Arrive at the destination. Fig. 7: Example of the agentâs views for Landmark Search with Image, re- sulting in success. The agent is instructed to search for âJonathanâ given in image form. 7 Conclusions We present 360CityArena, a benchmark designed to evaluate the urban ex- ploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Our experimental results reveal that state- of-the-art LMM-based agents still face significant challenges in understanding and reasoning within city environments. We believe that providing a benchmark closer to real-world conditions will further advance the development of embodied agents that are applicable to real urban settings. Acknowledgements This work was supported in part by JSPS KAKENHI Grant Number 25H01164, SIP Smart Disaster Prevention, JST BOOST Grant Number JPMJBS2418, and JST ASPIRE Program Grant No. JPMJAP2303. This study was conducted with approval from the relevant Ethics Committee (Approval No. UT-IST-RE- 250508_6). 16K. Watanabe et al. References 1. Anderson, P., Chang, A., Chaplot, D.S., Dosovitskiy, A., Gupta, S., Koltun, V., Kosecka, J., Malik, J., Mottaghi, R., Savva, M., et al.: On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757 (2018) 2. Anthropic: Claude Sonnet 4.5 system card. Tech. rep., Anthropic, PBC (2025) 3. Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-VL technical report. arXiv preprint arXiv:2502.13923 (2025) 4. Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Mip-NeRF 360: Unbounded anti-aliased neural radiance fields. In: CVPR. p. 5470â5479 (2022) 5. Brahmbhatt, S., Hays, J.: DeepNav: Learning to navigate large cities. In: CVPR. p. 5193â5202 (2017) 6. Chen, H., Suhr, A., Misra, D., Snavely, N., Artzi, Y.: TOUCHDOWN: Natural language navigation and spatial reasoning in visual street environments. In: CVPR. p. 12538â12547 (2019) 7. Comanici, G., et al.: Gemini 2.5: Pushing the frontier with advanced reason- ing, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261 (2025) 8. Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D.: Embodied question answering. In: CVPR. p. 1â10 (2018) 9. Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., Koltun, V.: CARLA: An open urban driving simulator. In: CoRL. vol. 78, p. 1â16 (2017) 10. Duan, J., Yu, S., Tan, H.L., Zhu, H., Tan, C.: A survey of embodied AI: From simulators to research tasks. TETCI 6(2), 230â244 (2022) 11. Feng, J., Zhang, J., Liu, T., Zhang, X., Ouyang, T., Yan, J., Du, Y., Guo, S., Li, Y.: CityBench: Evaluating the capabilities of large language models for urban tasks. In: KDD. p. 5413â5424 (2025) 12. Gao, C., Zhao, B., Zhang, W., Mao, J., Zhang, J., Zheng, Z., Man, F., Fang, J., Zhou, Z., Cui, J., et al.: EmbodiedCity: A benchmark platform for embodied agent in real-world city environment. arXiv preprint arXiv:2410.09604 (2024) 13. Haas, L., Skreta, M., Alberti, S., Finn, C.: PIGEON: Predicting image geolocations. In: CVPR. p. 12893â12902 (2024) 14. Hong, Y., Sun, R., Li, B., Yao, X., Wu, M., Chien, A., Yin, D., Wu, Y.N., Wang, Z., Chang, K.W.: Embodied web agents: Bridging physical-digital realms for inte- grated agent intelligence. In: NeurIPS. vol. 38 (2025) 15. Huang, W., Abbeel, P., Pathak, D., Mordatch, I.: Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In: ICML. vol. 162, p. 9118â9147 (2022) 16. Ichter, B., Brohan, A., Chebotar, Y., Finn, C., Hausman, K., Herzog, A., Ho, D., Ibarz, J., Irpan, A., Jang, E., Julian, R., Kalashnikov, D., Levine, S., Lu, Y., Parada, C., Rao, K., Sermanet, P., Toshev, A.T., Vanhoucke, V., Xia, F., Xiao, T., Xu, P., Yan, M., Brown, N., Ahn, M., Cortes, O., Sievers, N., Tan, C., Xu, S., Reyes, D., Rettinghouse, J., Quiambao, J., Pastor, P., Luu, L., Lee, K.H., Kuang, Y., Jesmonth, S., Joshi, N.J., Jeffrey, K., Ruano, R.J., Hsu, J., Gopalakrishnan, K., David, B., Zeng, A., Fu, C.K.: Do as I can, not as I say: Grounding language in robotic affordances. In: CoRL. vol. 205, p. 287â318 (2023) 360CityArena17 17. Ji, Y., Zhu, Z., Zhao, Y., Liu, B., Gao, C., Zhao, Y., Qiu, S., Hu, Y., Yin, Q.: To- wards autonomous UAV visual object search in city space: Benchmark and agentic methodology. AAAI 40(22), 18342â18350 (2026) 18. Jiao, J., He, J., Liu, C., Aegidius, S., Hu, X., Braud, T., Kanoulas, D.: LiteVLoc: Map-lite visual localization for image goal navigation. In: ICRA. p. 5244â5251 (2025) 19. Kawaharazuka, K., Oh, J., Yamada, J., Posner, I., Zhu, Y.: Vision-language-action models for robotics: A review towards real-world applications. IEEE Access 13, 162467â162504 (2025) 20. Kerbl, B., Kopanas, G., LeimkĂŒhler, T., Drettakis, G.: 3D gaussian splatting for real-time radiance field rendering. ACM TOG 42(4), 139 (2023) 21. Khanna, M., Ramrakhya, R., Chhablani, G., Yenamandra, S., Gervet, T., Chang, M., Kira, Z., Chaplot, D.S., Batra, D., Mottaghi, R.: GOAT-Bench: A benchmark for multi-modal lifelong navigation. In: CVPR. p. 16373â16383 (2024) 22. Lee, J., Miyanishi, T., Kurita, S., Sakamoto, K., Azuma, D., Matsuo, Y., Inoue, N.: CityNav: A large-scale dataset for real-world aerial navigation. In: ICCV. p. 5912â5922 (2025) 23. Li, J., Padmakumar, A., Sukhatme, G., Bansal, M.: VLN-Video: Utilizing driv- ing videos for outdoor vision-and-language navigation. AAAI 38(17), 18517â18526 (2024) 24. Li, M., Zhao, S., Wang, Q., Wang, K., Zhou, Y., Srivastava, S., Gokmen, C., Lee, T., Li, L.E., Zhang, R., Liu, W., Liang, P., Fei-Fei, L., Mao, J., Wu, J.: Embodied Agent Interface: Benchmarking LLMs for embodied decision making. In: NeurIPS. vol. 37, p. 100428â100534 (2024) 25. Liu, X., Li, J., Jiang, Y., Sujay, N., Yang, Z., Zhang, J., Abanes, J., Zhang, J., Feng, C.: CityWalker: Learning embodied urban navigation from web-scale videos. In: CVPR. p. 6875â6885 (2025) 26. Liu, Y., Chen, W., Bai, Y., Liang, X., Li, G., Gao, W., Lin, L.: Aligning cyber space with physical world: A comprehensive survey on embodied AI. IEEE/ASME TMECH 30(6), 7253â7274 (2025) 27. Liu, Z., He, H., Alumootil, V., Pandya, A., Squicciarini, B., Wu, W., Zhou, B.: Side- walkBench: Benchmarking visual navigation on urban sidewalks. arXiv preprint arXiv:2606.16953 (2026) 28. Majumdar, A., Aggarwal, G., Devnani, B., Hoffman, J., Batra, D.: ZSON: Zero- shot object-goal navigation using multimodal goal embeddings. In: NeurIPS. vol. 35, p. 32340â32352 (2022) 29. Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. In: ECCV. vol. 12346, p. 405â421 (2020) 30. Mirowski, P., Grimes, M., Malinowski, M., Hermann, K.M., Anderson, K., Teplyashin, D., Simonyan, K., Kavukcuoglu, K., Zisserman, A., Hadsell, R.: Learn- ing to navigate in cities without a map. In: NeurIPS. vol. 31 (2018) 31. Miyai, A., Zhao, Z., Egashira, K., Sato, A., Sunada, T., Onohara, S., Yamanishi, H., Toyooka, M., Nishina, K., Maeda, R., et al.: WebChoreArena: Evaluating web browsing agents on realistic tedious web tasks. arXiv preprint arXiv:2506.01952 (2025) 32. Mukhopadhyay, S., Rajgaria, A., Khatiwada, P., Shrivastava, M., Roth, D., Gupta, V.: MAPWise: Evaluating vision-language models for advanced map queries. In: NAACL. p. 9348â9378 (2025) 33. OpenAI: GPT-5 system card. Tech. rep., OpenAI (2025) 18K. Watanabe et al. 34. Paz-Argaman, T., Tsarfaty, R.: RUN through the streets: A new dataset and base- line models for realistic urban navigation. In: EMNLP-IJCNLP. p. 6449â6455 (2019) 35. Puig, X., Undersander, E., Szot, A., Cote, M.D., Yang, T.Y., Partsey, R., Desai, R., Clegg, A., Hlavac, M., Min, S.Y., VondruĆĄ, V., Gervet, T., Berges, V.P., Turner, J.M., Maksymets, O., Kira, Z., Kalakrishnan, M., Malik, J., Chaplot, D.S., Jain, U., Batra, D., Rai, A., Mottaghi, R.: Habitat 3.0: A co-habitat for humans, avatars, and robots. In: ICLR (2024) 36. Savva, M., Chang, A.X., Dosovitskiy, A., Funkhouser, T., Koltun, V.: MINOS: Mul- timodal indoor simulator for navigation in complex environments. arXiv preprint arXiv:1712.03931 (2017) 37. Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., Parikh, D., Batra, D.: Habitat: A platform for embodied AI research. In: ICCV. p. 9339â9347 (2019) 38. Schumann, R., Riezler, S.: Generating landmark navigation instructions from maps as a graph-to-text problem. In: ACL-IJCNLP. p. 489â502 (2021) 39. Shah, D., Sridhar, A., Dashora, N., Stachowicz, K., Black, K., Hirose, N., Levine, S.: ViNT: A foundation model for visual navigation. In: CoRL. vol. 229, p. 711â 733 (2023) 40. Sugimoto, N., Ebine, Y., Aizawa, K.: Building Movie Map â a tool for exploring areas in a city â and its evaluations. In: ACM M. p. 3330â3338 (2020) 41. Takenawa, M., Sugimoto, N., Wöhler, L., Ikehata, S., Aizawa, K.: Building and evaluating a realistic virtual world for large scale urban exploration from 360° videos. MTAP 85, 149 (2026) 42. Wang, S., Liang, C., Gao, Y., Yu, E., Li, S., Li, J., Wang, H.: CitySeeker: How do VLMs explore embodied urban navigation with implicit human needs? In: ICLR (2026) 43. Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., Wang, Z., Chen, Z., Zhang, H., Yang, G., Wang, H., Wei, Q., Yin, J., Li, W., Cui, E., Chen, G., Ding, Z., Tian, C., Wu, Z., Xie, J., Li, Z., Yang, B., Duan, Y., Wang, X., Hou, Z., Hao, H., Zhang, T., Li, S., Zhao, X., Duan, H., Deng, N., Fu, B., He, Y., Wang, Y., He, C., Shi, B., He, J., Xiong, Y., Lv, H., Wu, L., Shao, W., Zhang, K., Deng, H., Qi, B., Ge, J., Guo, Q., Zhang, W., Zhang, S., Cao, M., Lin, J., Tang, K., Gao, J., Huang, H., Gu, Y., Lyu, C., Tang, H., Wang, R., Lv, H., Ouyang, W., Wang, L., Dou, M., Zhu, X., Lu, T., Lin, D., Dai, J., Su, W., Zhou, B., Chen, K., Qiao, Y., Wang, W., Luo, G.: InternVL3.5: Advancing open- source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265 (2025) 44. Wu, W., He, H., He, J., Wang, Y., Duan, C., Liu, Z., Li, Q., Zhou, B.: MetaUrban: An embodied AI simulation platform for urban micromobility. In: ICLR (2025) 45. Xie, Z., Liu, Z., Peng, Z., Wu, W., Zhou, B.: Vid2Sim: Realistic and interactive simulation from video for urban navigation. In: CVPR. p. 1581â1591 (2025) 46. Xing, S., Sun, Z., Xie, S., Chen, K., Huang, Y., Wang, Y., Li, J., Song, D., Tu, Z.: Can large vision language models read maps like a human? arXiv preprint arXiv:2503.14607 (2025) 47. Xu, Q., Yi, X., Xu, J., Tao, W., Ong, Y.S., Zhang, H.: Few-shot NeRF by adaptive rendering loss regularization. In: ECCV. vol. 15124, p. 125â142 (2024) 48. Yang, J., Pavone, M., Wang, Y.: FreeNeRF: Improving few-shot neural rendering with free frequency regularization. In: CVPR. p. 8254â8263 (2023) 49. Yang, J., Ding, R., Brown, E., Qi, X., Xie, S.: V-IRL: Grounding virtual intelligence in real life. In: ECCV. vol. 15103, p. 36â55 (2024) 360CityArena19 50. Yang, J., Yang, S., Gupta, A.W., Han, R., Fei-Fei, L., Xie, S.: Thinking in space: How multimodal large language models see, remember, and recall spaces. In: CVPR. p. 10632â10643 (2025) 51. Yang, R., Chen, H., Zhang, J., Zhao, M., Qian, C., Wang, K., Wang, Q., Ko- ripella, T.V., Movahedi, M., Li, M., Ji, H., Zhang, H., Zhang, T.: EmbodiedBench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. In: ICML. vol. 267, p. 70576â70631 (2025) 52. Zaffar, M., Garg, S., Milford, M., Kooij, J., Flynn, D., McDonald-Maier, K., Ehsan, S.: VPR-Bench: An open-source visual place recognition evaluation framework with quantifiable viewpoint and appearance change. IJCV 129, 2136â2174 (2021) 53. Zhang, D., Wang, C., Wang, W., Li, P., Qin, M., Wang, H.: Gaussian in the wild: 3D Gaussian splatting for unconstrained image collections. In: ECCV. vol. 15134, p. 341â359 (2024) 54. Zhang, M., Qu, K., Patil, V., Cadena, C., Hutter, M.: Tag Map: A text-based map for spatial reasoning and navigation with large language models. In: CoRL. vol. 270, p. 2120â2146 (2024) 55. Zhou, S., Xu, F.F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al.: WebArena: A realistic web environment for building autonomous agents. In: ICLR (2024) 56. Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., Vuong, Q., Vanhoucke, V., Tran, H., Soricut, R., Singh, A., Singh, J., Sermanet, P., Sanketi, P.R., Salazar, G., Ryoo, M.S., Reymann, K., Rao, K., Pertsch, K., Mordatch, I., Michalewski, H., Lu, Y., Levine, S., Lee, L., Lee, T.W.E., Leal, I., Kuang, Y., Kalashnikov, D., Julian, R., Joshi, N.J., Irpan, A., Ichter, B., Hsu, J., Herzog, A., Hausman, K., Gopalakrishnan, K., Fu, C., Florence, P., Finn, C., Dubey, K.A., Driess, D., Ding, T., Choromanski, K.M., Chen, X., Chebotar, Y., Carbajal, J., Brown, N., Brohan, A., Arenas, M.G., Han, K.: RT-2: Vision-language-action models transfer web knowledge to robotic control. In: CoRL. vol. 229, p. 2165â2183 (2023) 20K. Watanabe et al. In this supplement, we describe the experimental details in Section A and prompts used for the 360CityArena experiments in Section B. A Experiment Details A.1 Task Settings For each task, a starting point and direction are defined. In addition, as de- scribed in Section B.3, prompts for each task type, as well as the variables and reference images for each task, are specified. These tasks are designed to evaluate complementary aspects of embodied navigation, including localization, landmark grounding, instruction following, and spatial reasoning. A.2 Map Information The maps used in 360CityArena are created by cropping the exploration area from OpenStreetMap. A.3 Stopping Conditions In our experiments, we introduced the following stopping conditions. The details are as follows: â STEP_LIMIT â One action is counted as one step, and the process termi- nates when the number of steps exceeds 50. â STAGNATION â The process terminates when the agent repeats the same action 20 times consecutively without any change in position. â AWAY_FROM_GOAL â The process terminates when the agent moves away from the goal location for five consecutive movements, excluding view- point changes. A.4 Model Settings Claude Sonnet 4.5 and Qwen2.5-VL-32B-Instruct were evaluated with a tem- perature of 1.0, while GPT-5 does not allow temperature configuration. The maximum output length was set to 8,096 tokens. In addition to the prompts described in Section B, the model input consists of the viewpoint image, a map indicating the agentâs current location, the Reflection Memory, and the task- specific reference image. The Reflection Memory is updated at every step and included in the input to the model. However, because Claude imposes a 5 MB input size limit, the current-location map is downscaled in 15% steps as needed to stay under that limit. 360CityArena21 B Prompts for 360CityArena Experiments This section lists all the prompts used in our experiments. The prompts are divided into System Prompt, Reflection Prompt, and Task Prompt. In addition to these prompts, the inputs to the MLLMs include the viewpoint image, the map indicating the current location, the Reflection Memory, and the task-specific reference map. B.1 System Prompt The System Prompt specifies the reasoning procedure, the available actions, and the format of the output. All prompts were designed specifically for this benchmark and were not part of the training data of the evaluated models. System Prompt You are a computer agent that uses the ReACT (Reasoning, Action, Observation) framework with memory to explore a city. For each step, you should: 1. Think: Analyze the current state and decide what to do next 2. Action: Choose one of the following actions: - W: move forward - LEFT/RIGHT/UP: if you see red arrows, you should select one of them. - S: turn camera to the direction of travel - Q: rotate camera upward (look around) - E: rotate camera downward (look around) - A: rotate camera left (look around) - D: rotate camera right (look around) - ANSWER: answer the question NOTE: - If you see a big red arrow in front of you and want to go straight, you should select UP. - You cannot select LEFT, RIGHT or UP unless you see a big red arrow in front of you. - If pressing "W" does not move you forward, look around; red arrows will appear. You cannot move in any direction where a red arrow is not visible. 3. Observation: You will receive the result of your action You will receive two types of images: 1. Camera view: The first-person view of what you can see in the city 2. Map view (when available): A top-down map showing your current location with a red arrow indicating your position and direction 22K. Watanabe et al. Use both images to make better navigation decisions. The map can help you understand your location and plan your route more effectively. Respond in the following JSON format: "thought": "your reasoning about what to do next", "action": "one of the available actions", "memory": "important information to remember for future steps", "answer": "the answer to the question of the task" To not update memory, respond with an empty string. For example: "thought": "I need to move forward", "action": "W", "memory": "1. My short term plan is to find the signboard of the road. 2. I need to move forward to find the signboard.", "answer": "" Move control (W): - Required: set "answer" to one of SMALL / MEDIUM / LARGE - Mapping: SMALL = short move, MEDIUM = normal move, LARGE = long move - Note: do not include anything else in "answer" when action is "W". Rotation control (A/D/Q/E): - Required: set "answer" to one of SMALL / MEDIUM / LARGE - Mapping: SMALLâ 30°, MEDIUMâ 60°, LARGEâ 90° - Default: if "answer" is omitted,â 24° (â 0.5s at ~48°/s) is used. Another example of changing direction: "thought": "I need to turn to the direction of travel", "action": "S", "memory": "", "answer": "" Another example of answering the question: "thought": "", 360CityArena23 "action": "ANSWER", "memory": "The name of the city is Tokyo.", "answer": "x:100 y:100" Do NOT wrap anything in âjsonâ tags, and only respond with the JSON object. Always analyze the screenshot carefully to determine the correct coordinates for your actions. When a map is provided, use it to understand your current position and make more informed navigation decisions. The memory field should contain any important information you want to remember for future steps. B.2 Reflection Prompt The Reflection Prompt is designed to provide a mechanism that enables the agent to maintain long-term consistency while planning and selecting actions. It specifies how information should be stored in the reflection memory. The reflection memory is saved as text and is updated after each action, serving as input to the agent in the next reasoning step. Reflection Prompt You will only see your last few observations and actions, so you will need to remember important goals, objectives, and information that may be relevant. Make sure to read all the text on the screen and use it to update your reflection memory! You will be given a reflection memory that you can update with your current thoughts -- be careful NOT to overwrite your previous reflection with a new one -- make sure to copy the previous reflection and add to it if you want to retain information. Do not be conservative with your memory, you will need to remember everything! Consider reflecting on: - Important city objectives and goals - Strategies that worked or didnât work - Locations youâve visited and what you found there - Current status of the city 24K. Watanabe et al. Fig. 8: Grid map provided in the Localization Task. The task reference image in the Localization Task Prompt corresponds to this image. Map data© OpenStreetMap contributors, ODbL 1.0, https://w.openstreetmap.org/copyright. Think step by step and update your reflection memory with your current thoughts. B.3 Task Prompts The Task Prompts are prompts defined for each task type. For each task, certain variables and the input images are modified accordingly. Localization In the Localization Task, the grid map shown in Figure 8 is pro- vided, and the agent is required to estimate its current position from this map. The expected answer format is also specified within the prompt. 360CityArena25 Localization Task Prompt Your task is to explore the city and determine your initial starting position. A reference map with a grid overlay is provided in the Task Reference Images. You MUST pick your answer from this grid: select the single grid cell that corresponds to your starting location and output its grid coordinate. Before answering, actively explore your surroundings to gain confidence in your estimate: move around (walk a short distance), rotate your view (left/right and up/down), and re-check landmarks from multiple angles. Do not provide your answer until you are confident in your location. You must complete your exploration and give your final answer within 50 steps. When you have determined your starting position, specify "ANSWER" in the "action" field of your JSON response and provide your answer in the "answer" field in the format: "x:[grid_x] y:[grid_y]", where [grid_x] and [grid_y] are the integer indices of the selected grid cell from the provided grid. Do not output continuous coordinates (e.g., meters); only output the discrete grid indices. Landmark Search with Language âLandmarkNameâ is replaced with the name of a specific landmark for each task, such as âDoutor Coffeeâ or âSEGA.â To reduce task difficulty, we also include the information âThe goal is not far from the starting point.â Landmark Search with Language Task Prompt Your task is to go to LandmarkName. When you get in front of LandmarkName, use the ANSWER action to confirm completion. The goal is not far from the starting point. Landmark Search with Image In the Landmark Search with Image task, a task- specific landmark image, such as the one shown in Figure 9, is provided as the task reference image. The landmarks used here are the same ones selected in the âLandmark Search with Languageâ task. 26K. Watanabe et al. Fig. 9: Example of a landmark image provided in the Landmark Search with Image task. The task reference image in the Landmark Search with Image Task Prompt corresponds to this image. Landmark Search with Image Task Prompt Your task is to go to the landmark shown in the task reference image. When you get in front of the landmark, use the ANSWER action to confirm completion. The goal is not far from the starting point. Map Navigation In the Map Navigation task, a task-specific map image such as the one shown in Figure 10 is provided as the reference image. Map Navigation Task Prompt Your task is to navigate from the starting position to the goal destination using the available actions. A navigation reference map is provided in the task description above. On this reference map: - BLUE marker indicates your starting position (initial location) - RED marker indicates your goal destination When you reach the goal area, use the ANSWER action to confirm completion. 360CityArena27 Fig. 10: Example of a map provided in the Map Navigation Task. The task reference image in the Map Navigation Task Prompt corresponds to this image. Map data© OpenStreetMap contributors, ODbL 1.0, https://w.openstreetmap.org/ copyright. Vision Language Navigation âDirectionsâ is replaced with task-specific naviga- tion instructions such as the following. 1. Go straight. 2. Turn left at the first intersection. 3. Go straight. 4. Stop in front of Surugaya Purchase Center. Vision Language Navigation Task Prompt Your task is to follow the directions to reach your destination. Please follow the instructions below: Directions 28K. Watanabe et al. Once you have reached your destination, output the ANSWER action. Object Count âObjectâ is replaced with a task-specific object name such as âvend- ing machineâ or âplanted tree.â Likewise, âRangeâ is replaced with a description of the search area, such as âthe full perimeter of the block to your leftâ or âthe road up to the next crosswalk.â Object Count Task Prompt Your task is to count the number of Object within Range. Only count items that belong to the specified area/side/segment relative to your current position and along the described road segment. Do not include items outside the specified range. If the specified range describes a block, count along the full perimeter of that block (all four sides) unless a specific side is explicitly specified (e.g., "right side only"). Output the answer as a number. Relational Spatial Reasoning âLandmarkNameâ is replaced with a task-specific landmark name such as âAkiba no X,â while âRelationâ is replaced with expres- sions describing spatial relationships, such as âthe store to the left of itâ or âthe store across the street from it.â Relational Spatial Reasoning Task Prompt Find a nearby LandmarkName and tell me the name of Relation. LandmarkName is right nearby. B.4 Fuzzy Match Evaluation Prompt The fuzzy_match metric (Section 3.4 in the main paper) uses an LLM to judge whether the agentâs answer is semantically equivalent to the ground truth. The following system prompt is given to the evaluator LLM (GPT-5), along with a user message containing both answers. 360CityArena29 Fuzzy Match System Prompt You are a helpful assistant. You are given a user answer and an expected answer. Please determine if the user answer is correct. The answers do not have to match exactly word-for-word - If you determine that different names refer to the same landmark, return True. The evaluator receives the following user message, where user_answer and expected_answer are replaced with the agentâs output and the ground-truth answer, respectively. Fuzzy Match User Message User answer: user_answer Expected answer: expected_answer Output the answer in the following format: is_correct: True/False