Paper deep dive
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models
Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu, Yuxiang Zhang, Hang Su, Yubin Wang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/13/2026, 1:05:16 AM
Summary
World2Mind is a training-free spatial intelligence toolkit that enables Multimodal Foundation Models (MFMs) to perform robust 3D spatial reasoning. By constructing structured spatial cognitive maps—specifically the Allocentric-Spatial Tree (AST)—and employing a three-stage reasoning chain (tool invocation, cue collection, and geometry-semantics interwoven reasoning), the toolkit mitigates the 'semantic-geometry gap' in foundation models, allowing them to achieve significant performance gains in spatial tasks even in text-only settings.
Entities (6)
Relation Signals (3)
World2Mind → utilizes → Allocentric-Spatial Tree (AST)
confidence 100% · World2Mind synthesizes an Allocentric-Spatial Tree (AST) that uses elliptical parameters to model the top-down layout of landmarks accurately.
World2Mind → evaluatedon → VSI-Bench
confidence 95% · We conduct evaluations on two challenging spatial reasoning benchmarks: VSI-Bench
World2Mind → improvesperformanceof → GPT-5.2
confidence 95% · Extensive experiments demonstrate that World2Mind boosts the performance of frontier models, such as GPT-5.2, by 5%~18%.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Achieving robust spatial reasoning remains a fundamental challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical shortcuts via 3D grounding data or remain confined to 2D visual perception, limiting both spatial reasoning accuracy and generalization in unseen scenarios. Inspired by the spatial cognitive mapping mechanisms of biological intelligence, we propose World2Mind, a training-free spatial intelligence toolkit. At its core, World2Mind leverages 3D reconstruction and instance segmentation models to construct structured spatial cognitive maps, empowering MFMs to proactively acquire targeted spatial knowledge regarding interested landmarks and routes of interest. To provide robust geometric-topological priors, World2Mind synthesizes an Allocentric-Spatial Tree (AST) that uses elliptical parameters to model the top-down layout of landmarks accurately. To mitigate the inherent inaccuracies of 3D reconstruction, we introduce a three-stage reasoning chain comprising tool invocation assessment, modality-decoupled cue collection, and geometry-semantics interwoven reasoning. Extensive experiments demonstrate that World2Mind boosts the performance of frontier models, such as GPT-5.2, by 5%~18%. Astonishingly, relying solely on the AST-structured text, purely text-only foundation models can perform complex 3D spatial reasoning, achieving performance approaching that of advanced multimodal models.
Tags
Links
- Source: https://arxiv.org/abs/2603.09774v1
- Canonical: https://arxiv.org/abs/2603.09774v1
Trouble viewing inline? Open PDF directly →
Full Text
28,852 characters extracted from source content.
Expand or collapse full text
World2Mind: Cognition Toolkit for Allocentric Spatial Reasoning in Foundation Models Shouwei Ruan 1 , Bin Wang 2 , Zhenyu Wu 1 , Qihui Zhu 1 , Yuxiang Zhang 2 , Hang Su 3 , Yubin Wang 2 * 1 Institute of Artificial Intelligence, Beihang University 1 Huawei Noah’s Ark Lab 3 Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua-Bosch Joint ML Center, THBI Lab, BNRist Center, Tsinghua University Abstract Achieving robust spatial reasoning remains a fundamen- tal challenge for current Multimodal Foundation Models (MFMs). Existing methods either overfit statistical short- cuts via 3D grounding data or remain confined to 2D vi- sual perception, limiting both spatial reasoning accuracy and generalization in unseen scenarios. Inspired by the spatial cognitive mapping mechanisms of biological intelli- gence, we propose World2Mind, a training-free spatial in- telligence toolkit. At its core, World2Mind leverages 3D re- construction and instance segmentation models to construct structured spatial cognitive maps, empowering MFMs to proactively acquire targeted spatial knowledge regarding interested landmarks and routes of interest. To provide ro- bust geometric-topological priors, World2Mind synthesizes an Allocentric-Spatial Tree (AST) that uses elliptical pa- rameters to model the top-down layout of landmarks ac- curately. To mitigate the inherent inaccuracies of 3D re- construction, we introduce a three-stage reasoning chain comprising tool invocation assessment, modality-decoupled cue collection, and geometry-semantics interwoven reason- ing. Extensive experiments demonstrate that World2Mind boosts the performance of frontier models, such as GPT- 5.2, by 5%∼18%. Astonishingly, relying solely on the AST- structured text, purely text-only foundation models can per- form complex 3D spatial reasoning, achieving performance approaching that of advanced multimodal models. 1. Introduction Although multimodal foundation models (MFMs) [1, 12, 24, 28] excel in general visual understanding and cross- modal reasoning [29], they struggle significantly in em- * Corresponding author. bodied AI and complex spatial reasoning tasks requir- ing physical interaction [15, 17, 18, 26, 34]. This defi- ciency stems from their over-reliance on egocentric obser- vations and lacking the capacity to abstract global spatial topology [18, 34], trapping MFMs in an insurmountable “semantic-geometry gap” in tasks like distance estimation, viewpoint transformation, and path planning. Current efforts to enhance MFMs’ spatial reasoning primarily follow two paradigms. Training-based meth- ods [6, 9, 20] fine-tune models on massive 3D-grounded QA pairs. However, this forces models to overfit statis- tical shortcuts [14, 25] rather than acquiring genuine spa- tial cognition, leading to poor generalization in out-of- distribution scenarios [33]. Alternatively, introducing ex- plicit 3D modalities [7, 10, 21, 31] exacerbates inter-modal alignment challenges [35]. Meanwhile, recent tool-based methods [8, 19, 36] rely on active rendering under large- scale 3D reconstruction. These are severely bottlenecked by reconstruction quality and remain tethered to low-level vi- sual perception, failing to abstract geometric data into struc- tured semantics for high-level logical reasoning. Biological Intelligence (BI) offers an ideal blueprint to break the shackles of egocentric observation. Rather than reacting passively to transient visual inputs, the biological brain intrinsically transforms egocentric views into an al- locentric perspective [4]. Supported by place cells in the hippocampus [3, 11, 22] and grid cells in the entorhinal cor- tex [13], mammals construct a global cognitive map entirely independent of their egocentric viewpoint [23, 27]. This map forms the cornerstone for strategic mental simulation and advanced reasoning [2]. To bridge this gap between foundation models and BI in spatial representation and reasoning, we propose World2Mind, a plug-and-play spatial cognition toolkit that equips models with human-like mental simulation ca- pabilities. World2Mind integrates an efficient geometry- 1 arXiv:2603.09774v1 [cs.AI] 10 Mar 2026 Ego-CentricInput Video Multi-viewImages or 12 3 4 Spatial-relatedQuestion Foundationmodel “Measuringfromtheclosestpointofeachobject,whatisthe distancebetweenthedoorandthetelephone(inmeters)?” Step1:ToolInvocationJudgment ReasoningTrace Step2:Modality-DecoupledCueCollection VisualClues:Fromthevisualinput,Icanobserveanofficedesk with...,indicatesthatthetelephoneiswithin3metersofthedoor. MapClues:Igetthefollowingboundingboxdatainthe2Dplane: telephone_01<Center:(-3.187,-0.371),Width:0.109>,... Step3:Geometry-SemanticsInterwovenReasoning Cross-Validation&Calculation:Weneedtofindtheshortest Euclideandistancebetweentheboundingboxoftelephone_01and door_02.......Distance=(MinimumXofdoor_02)-(MaximumX oftelephone_01)=0.507-(-3.1325)=3.6395meters.WhenI reviewedthevisualcues,Ifeltthisresultwasvisuallyconsistent. Answer:3.64m/Ground-truth:3.8m World2MindPipeline ToolParameters InstanceofInteresting: <door><telephone> <officedesk> ... MapType: Landmark Route Returnto visualization: Depth Anything v3 Segment Anything v3 Projecting&Rendering PointCloudSemanticGrid 1.LandmarkCognitiveMap 2.RouteCognitiveMap ASTYaml RouteMap Ineedusetheworld2mindtooltogetprecisespatial measurementsbetweentheseobjects. APICall Figure 1. Overview of foundation models performing allocentric spatial reasoning via the proposed World2Mind toolkit. Given ego- centric video or multi-view observations, the model first assesses the necessity of tool invocation and subsequently passes key parameters (e.g., instances of interest) to World2Mind to drive the generation of spatial cognitive maps. World2Mind integrates an efficient pipeline for 3D reconstruction and semantic-geometry alignment, returning the required structured spatial knowledge through targeted projection and rendering mechanisms. Furthermore, the model conducts geometry-semantics interwoven reasoning based on both the raw visual observations and the geometric cues provided by World2Mind, ultimately yielding highly reliable answers. semantics alignment pipeline, leveraging pre-trained visual geometry [16, 30, 32] and instance segmentation models [5] to extract semantic voxel grids. From these, it constructs two core representations: 1) a Route Cognitive Map for passability prediction, and 2) a Landmark Cognitive Map for object topology. Encapsulated as an accessible toolset, World2Mind enables models to dynamically specify param- eters (e.g., instances of interest, required spatial knowledge, and map visualizations) to proactively acquire targeted al- locentric spatial knowledge on demand. To provide robust geometric-topological priors, we for- mally define the Allocentric-Spatial Tree (AST) as the core spatial representation in World2Mind. The AST is a directed acyclic graph utilizing geometrically stable land- marks (e.g., beds, tables) as core nodes to hierarchically associate surrounding smaller instances. Crucially, to ap- proximate the fuzzy nature of human cognition, the AST models spatial footprints using rectangle-elliptical param- eters (bounding boxes, major/minor axes, eccentricity, and rotation angles). These designs equip models with robust, dense, and highly actionable geometric-topological priors. However, merely offering spatial representation is insuf- ficient to guarantee robust reasoning. In complex physical scenarios, reconstruction quality often suffers severe cor- ruption due to occlusions or restricted viewpoints, leading to conflicts with objective geometric laws and raw visual observations. To mitigate this risk, we integrate a rigorous spatial reasoning chain into World2Mind: 1) Difficulty As- sessment and Tool Invocation, preventing over-computation on simple superficial queries; 2) Modality-Decoupled Cue Collection, independently extracting information from ego- centric vision, AST structured text, and map visualizations; and 3) Geometry-Semantics Interwoven Reasoning, guid- ing the model to resolve cross-modal conflicts proactively and ultimately yield reliable spatial decisions. Extensive evaluations across various spatial reasoning benchmarks demonstrate that World2Mind yields stable performance improvements of 6%–18% for frontier mod- els like GPT-5.2, while maintaining exceptional efficiency and reasoning interpretability. Astonishingly, leveraging the pure, high-density allocentric priors provided by the AST, text-only foundation models can execute complex 3D reasoning directly within their parameter space simply by reading the AST representation, approaching the perfor- mance of advanced multimodal models. Our findings offer a highly promising pathway to overcome the spatial cogni- tion bottleneck in foundation models. 2. Method Overview This section details the technical overview of the proposed world2mind, as illustrated in Fig. 1 2.1. Geometry-Semantic Alignment Pipeline Given an egocentric video sequence or multi-view image set I t T t=1 , our primary objective is to transcend the lim- itations of 2D vision and construct a robust 3D semantic representation of the physical world. ❶ Depth Estimation & Semantic Extraction. We employ Depth Anything V3 [16] for monocular depth estimation, obtaining the depth map D t ∈ R H×W and camera pose T t ∈ SE(3). Concurrently, we utilize SAM3 [5] to ex- 2 Table 1. Main results on the VSI-Bench [34] benchmark (Tiny subset). ModelsAvg. Numerical Answer (%)Multiple-Choice Answer (%) Obj. Count Abs. Dist. Obj. Size Room Size Rel. Dist. Rel. Dir. Route Plan Appr. Order Frontier models w/o. World2Mind GPT-5.246.752.534.967.550.642.040.734.751.0 Claude-4.6-Opus38.446.918.562.126.840.047.234.730.6 Gemini-3-Pro55.247.832.171.355.054.044.857.179.6 Frontier models w./ World2Mind GPT-5.254.0 (↑7.3)47.4 (↓5.1)33.4 (↓1.5)63.3 (↓4.2)52.4 (↑1.8) 64.0 (↑22.0) 41.1 (↑0.4)51.0 (↑16.3)79.6 (↑28.6) Claude-4.6-Opus 56.0 (↑17.7) 59.0 (↑12.0)34.3 (↑15.8) 67.3 (↑5.2)54.8 (↑28.0) 64.0 (↑24.0) 62.7 (↑15.6) 65.3 (↑30.6)40.8 (↑10.2) Gemini-3-Pro61.0 (↑5.8)51.8 (↑4.1) 36.8 (↑4.7) 57.7 (↓13.5) 62.6 (↑7.6)62.0 (↑8.0) 67.7 (↑22.9) 65.3 (↑8.2) 83.7 (↑4.1) Table 2. Results on the MindCube-Tiny [18] benchmark. ModelsAvg.Around Among Rotation Frontier models w/o. World2Mind GPT-5.249.962.445.248.5 Claude-4.6-Opus48.558.850.729.0 Gemini-3-Pro75.177.268.293.0 Frontier models w./ World2Mind GPT-5.254.6 (↑4.7) 60.4 (↓2.0) 47.7 (↑2.5) 68.0 (↑19.5) Claude-4.6-Opus 62.9 (↑14.4) 82.4 (↑23.6) 60.8 (↑10.1) 45.0 (↑16.0) Gemini-3-Pro 81.6 (↑6.5) 86.0 (↑8.8) 75.8 (↑7.6) 93.5 (↑0.5) tract open-vocabulary semantic masks M t based on a user- specified category list C. To suppress the accumulation of long-tail errors inherent in depth estimation, we introduce a dual-level filtering mechanism based on the predicted con- fidence map C t ∈ [0, 1] H×W . Specifically, we formulate a binary validity maskV t ∈0, 1 H×W as follows: V t (u,v) = I (C t (u,v) > τ pixel )· I (μ t > τ frame ),(1) where μ t = 1 HW P H x=1 P W y=1 C t (x,y) denotes the global spatial confidence of frame t, and I(·) is the indicator func- tion that returns 1 if the condition is met and 0 otherwise. The variables τ pixel and τ frame represent the pixel-level and frame-level thresholds, respectively. A pixel is incorporated into the subsequent reconstruction only whenV t (u,v) = 1. ❷ Point Cloud Mapping & Density Filtering. Qualifying 2D pixels are back-projected into the world coordinate sys- tem via the camera intrinsic matrix K, generating a global point cloudP = (p i ,s i , rgb i ) M i=1 carrying semantic la- bels s i ∈ C. Addressing the boundary outliers inherent in depth estimation, we propose a core region extraction strat- egy: for each point, we calculate its K-nearest neighbor lo- cal density ρ i = 1 K P j∈N K (i) ∥p i − p j ∥ −1 , and eliminate low-density ”tail” points based on density percentiles. This yields an exceptionally pure geometry-semantic substrate. 2.2. Allocentric Cognitive Mapping Inspired by the spatial mapping mechanisms of BI, we dis- till the unstructured point cloud into two highly abstract cognitive maps, enabling the model to proactively acquire spatial knowledge on demand via tool invocation. ❶ Landmark Cognitive Mapping. Traditional methods rely on ambiguous relative relations or simplified grid repre- sentation [18, 34]. To overcome this, we formally define the Allocentric-Spatial Tree (AST), which reorganizes spatial entities as a directed acyclic graph within an absolute coor- dinate system. Specifically, we perform adaptive DBSCAN clustering on each semantic category within the point cloud to separate distinct instances. For each instance node, the AST discards traditional bounding boxes and instead fits a minimum bounding ellipse in the top-down view (X-Z plane), extracting the centroid (x c ,z c ), major and minor axes a and b, and rotation angle θ. This parameterization: 1) significantly enhances robustness against reconstruction boundary noise; 2) perfectly aligns with the fuzzy probabil- ity nature of human spatial footprint perception. Output as dense structured text (e.g., YAML), the AST explicitly en- codes hierarchical containment relationships among entities along with multi-dimensional geometric attributes. ❷ Route Cognitive Mapping.For navigation-oriented tasks, World2Mind also enables extracting the masks of traversable categories (e.g., floors), back-projects and vox- elizes them, and subsequently partitions them into an N×N grid map on the top-down plane. Combined with the map- ping of the camera trajectory sequenceT t , this route map provides the model with explicit priors regarding passability and the human observer’s motion trajectory. 2.3. Geometry-Semantics Interwoven Reasoning In physical scenarios, 2D visual observations are suscepti- ble to occlusions and adverse viewpoints, while 3D recon- struction information may contain local errors. To resolve potential contradictions between these two modalities, we design a rigorous three-stage interwoven reasoning chain. 3 Avg. Obj. Count Abs. Dist. Obj. Size Room Size Rel. Dist. Rel. Dir. Route Plan Appr. Order 0 20 40 60 80 Performance Score 6.0 23.1 3.4 7.1 20.0 6.0 19.9 4.1 18.4 18.2 52.2 10.0 4.4 22.0 18.0 24.6 14.3 0.0 GPT-5.2 (Blind) GPT-5.2 w./ World2Mind Claude-4.6-Opus (Blind) Claude-4.6-Opus w./ World2Mind Figure 2. Performance comparison under the text-only model (“blind”) setting. we report the performance gap on the VSI- Bench (Tiny subset) between foundation models relying solely on commonsense reasoning and those leveraging world2mind to ac- quire structured spatial knowledge for allocentric reasoning. Stage 1: Tool Invocation Judgement. To reduce unnec- essary computational overhead, the model must first eval- uate the spatial relevance of the query. The model should proactively invoke World2Mind only when the task explic- itly involves spatial reasoning, such as occlusion inference, distance estimation, or path planning. Stage 2: Modality-Decoupled Cue Collection. We force the model to extract information independently to prevent early modality bias. The model must simultaneously gather corroborating evidence from three independent sources: egocentric vision, the AST text returned by World2Mind, and optional 2D top-down map visualizations. Stage 3: Conflict Resolution and Cross-Validation. This is the crux of the reasoning chain. The model needs to proac- tively coordinate evidence across different modalities and identify cross-modal conflicts, such as missing objects due to visual truncation or coordinate drift caused by depth er- rors. By cross-validating the objective geometric param- eters of the AST against subjective visual appearances, the model can dynamically weigh the credibility between visual illusions and reconstruction artifacts, ultimately outputting highly reliable and logically interpretable decisions. 3. Experiment 3.1. Experimental Setup We conduct evaluations on two challenging spatial reason- ing benchmarks: VSI-Bench [34], which focuses on video- based reasoning in real-world physical scenes, and Mind- Cube [18], which emphasizes multi-view cognitive map- ping and mental simulation. Our evaluation primarily tar- gets the most frontier multimodal foundation models, in- cluding GPT, Claude, and Gemini, given their exceptional proficiency in tool invocation and instruction-following. 3.2. Main Results and Analysis Evaluation on VSI-Bench. As shown in Tab. 1, the seam- less integration of World2Mind yields a substantial leap in the average performance (Avg.) across all frontier mod- “What is the longest length of the chair(in centimeters)?” I see a black office chair in the closet area (and another chair at the desk with a jacket draped over it). The chair’s longest “side”likely refers to the largest dimension of the chair as seen from the side—typically the height from the floor to the top of the backrest. Visually, it looks like a standard office chair, roughly around 90–110 cm tall. To answer in centimeters more reliably, metric estimation from a spatial tool would help. Tool Call: world2mind Arguments: “categories”: [“chair, desk”],"knowledge_type": "landmark" GPT-5.2 Tool Result GPT-5.2 Map Clues: The tool detected two chairs. •chair_01: rectangle width 0.611 m, height 0.45 m (this “height” is footprint dimension in the 2D map, not vertical). It also reports a vertical Height_Range from 1.093 m to 1.628 m, giving vertical extent ≈ 0.535 m. •chair_02: width 0.12 m, height 0.329 m; Height_Range0.648 m to 1.599 m, vertical extent ≈ 0.951 m. The “longest side”of a chair in cm would most plausibly be the largest dimension among these; the largest vertical extent is for chair_02 at about 0.951 m = 95.1 cm. I’l confirm chair_02 corresponds to the visible office chair (likely the tall one). Tool Call: world2mind Arguments: "visualization_type": "landmark_vis" Visualization Clues:The landmark visualization indicates one chair instance aligns with the closet/desk chair region, and the other is smaller/partial. The larger vertical-extent instance (chair_02) corresponds to the office chair shape. Cross-Validation: Visually, the office chair height around ~1 m seems plausible. The tool’s 95.1 cm estimate for the longest dimension is consistent with a typical office chair. The alternative (chair_01) seems to represent a partial detection with smaller vertical extent, so I trust chair_02. Final Answer: 95 cm GPT-5.2 Tool Result Figure 3. Complete reasoning trace under World2Mind. els. Specifically, GPT-5.2 achieves a 7.3% improvement, while Claude-4.6-Opus achieves a staggering 17.7% im- provement. In-depth analysis of the subtasks reveals that the performance gains are most striking in task categories that rely heavily on allocentric priors. For instance, on Relative Direction, Route Planning, and Relative Distance, the per- formance of Claude-4.6-Opus skyrockets by 15.6%, 30.6%, and 24.0%, respectively. This compellingly demonstrates the critical role of the allocentric spatial knowledge pro- vided by World2Mind in bridging the spatial reasoning gap. Evaluation on MindCube. The results in Tab. 1 further corroborate the universality and robustness of our frame- work in sparse multi-view inputs. Even for Gemini-3-Pro, whose native spatial reasoning capability is already top-tier (with a baseline Avg. of 75.1%), World2Mind successfully shatters its performance ceiling, pushing its average accu- racy to 81.6% (+6.5%). Notably, in tasks like “Rotation” that severely test 3D spatial imagination, the model achieves a remarkable performance breakthrough (GPT-5.2 improves by 19.5%) due to its ability to perform logical deduction grounded in the AST. 3.3. Ablation and Case Study To explore the limits of how structured text of AST em- powers the spatial cognition of large language models, we follow [34] to conduct ablation studies under the text- only (“blind”) setting (see Fig. 2). When visual image in- puts are completely stripped away, foundation models that rely solely on commonsense priors degrade to near-random guessing on spatial tasks. Astonishingly, however, when equipped with World2Mind, both GPT-5.2 and Claude- 4.6-Opus exhibit a remarkable performance rebound in the ”blind” state. On core reasoning tasks such as Object Size and Route Planning, their scores closely approach those achieved with full visual inputs. This profound finding in- dicates that pure, high-quality allocentric geometric pri- ors are entirely sufficient to ignite powerful 3D men- 4 tal reconstruction and simulation capabilities of foun- dation models under the text-based reasoning. Further- more, we visualize the complete reasoning traces powered by World2Mind in Fig. 3, which clearly demonstrate that the interwoven reasoning process exhibits exceptional ro- bustness and logical interpretability when resolving cross- modal conflicts. References [1] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al.Qwen3-vl technical report.arXiv preprint arXiv:2511.21631, 2025. 1 [2] Jacob LS Bellmund, Peter G ̈ ardenfors, Edvard I Moser, and Christian F Doeller. Navigating cognition: Spatial codes for human thinking. Science, 362(6415):eaat6766, 2018. 1 [3] Nicola J Broadbent, Larry R Squire, and Robert E Clark. Spatial memory, recognition memory, and the hippocampus. Proceedings of the National Academy of Sciences, 101(40): 14515–14520, 2004. 1 [4] Neil Burgess. Spatial memory: how egocentric and allocen- tric combine. Trends in cognitive sciences, 10(12):551–557, 2006. 1 [5] Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoub- hik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. Sam 3: Segment anything with concepts. arXiv preprint arXiv:2511.16719, 2025. 2 [6] Boyuan Chen, Zhuo Xu, Sean Kirmani, Brain Ichter, Dorsa Sadigh, Leonidas Guibas, and Fei Xia. Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 14455–14465, 2024. 1 [7] Pingyi Chen, Yujing Lou, Shen Cao, Jinhui Guo, Lubin Fan, Yue Wu, Lin Yang, Lizhuang Ma, and Jieping Ye. Sd-vlm: Spatial measuring and understanding with depth-encoded vision-language models. arXiv preprint arXiv:2509.17664, 2025. 1 [8] Siyi Chen, Mikaela Angelina Uy, Chan Hee Song, Faisal Ladhak, Adithyavairavan Murali, Qing Qu, Stan Birchfield, Valts Blukis, and Jonathan Tremblay. Spacetools: Tool- augmented spatial reasoning via double interactive rl. arXiv preprint arXiv:2512.04069, 2025. 1 [9] An-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo, Rui- han Yang, Jan Kautz, Xiaolong Wang, and Sifei Liu. Spatial- rgpt: Grounded spatial reasoning in vision-language mod- els. Advances in Neural Information Processing Systems, 37:135062–135093, 2024. 1 [10] Erik Daxberger, Nina Wenzel, David Griffiths, Haiming Gang, Justin Lazarow, Gefen Kohavi, Kai Kang, Marcin Eichner, Yinfei Yang, Afshin Dehghan, et al. Mm-spatial: Exploring 3d spatial understanding in multimodal llms. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7395–7408, 2025. 1 [11] Howard Eichenbaum. The role of the hippocampus in navi- gation is memory. Journal of neurophysiology, 117(4):1785– 1796, 2017. 1 [12] Google. Gemini 3.1 pro: Best for complex tasks and bring- ing creative concepts to life. https://deepmind. google/models/gemini/pro/, 2026. 1 [13] Torkel Hafting, Marianne Fyhn, Sturla Molden, May-Britt Moser, and Edvard I Moser. Microstructure of a spatial map in the entorhinal cortex. Nature, 436(7052):801–806, 2005. 1 [14] Jiaxin Huang, Ziwen Li, Hanlve Zhang, Runnan Chen, Xiao He, Yandong Guo, Wenping Wang, Tongliang Liu, and Mingming Gong. Surprise3d: A dataset for spatial under- standing and reasoning in complex 3d scenes. arXiv preprint arXiv:2507.07781, 2025. 1 [15] Shuai Huang, Wenxuan Zhao, and Jun Gao.Si-bench: Benchmarking social intelligence of large language mod- els in human-to-human conversations.arXiv preprint arXiv:2510.23182, 2025. 1 [16] Haotong Lin, Sili Chen, Junhao Liew, Donny Y Chen, Zhenyu Li, Guang Shi, Jiashi Feng, and Bingyi Kang. Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647, 2025. 2 [17] Jingli Lin, Runsen Xu, Shaohao Zhu, Sihan Yang, Peizhou Cao, Yunlong Ran, Miao Hu, Chenming Zhu, Yiman Xie, Yilin Long, et al. Mmsi-video-bench: A holistic bench- mark for video-based spatial intelligence. arXiv preprint arXiv:2512.10863, 2025. 1 [18] Fangzheng Liu, Don Derek Haddad, and Joe Paradiso. Mind- cube: an interactive device for gauging emotions. In Adjunct Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, pages 1–2, 2024. 1, 3, 4 [19] Zhanpeng Luo, Ce Zhang, Silong Yong, Cunxi Dai, Qian- wei Wang, Haoxi Ran, Guanya Shi, Katia Sycara, and Yaqi Xie. pyspatial: Generating 3d visual programs for zero-shot spatial reasoning. arXiv preprint arXiv:2603.00905, 2026. 1 [20] Wufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang, Celso de Melo, Jianwen Xie, and Alan Yuille. Spatialreasoner: To- wards explicit and generalizable 3d spatial reasoning. arXiv preprint arXiv:2504.20024, 2025. 1 [21] Zhenhua Ning, Zhuotao Tian, Shaoshuai Shi, Guangming Lu, Daojing He, Wenjie Pei, and Li Jiang.Enhanc- ing spatial reasoning in multimodal large language models through reasoning-based segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 7851–7860, 2025. 1 [22] John O’Keefe and Jonathan Dostrovsky. The hippocampus as a spatial map: preliminary evidence from unit activity in the freely-moving rat. Brain research, 1971. 1 [23] John O’keefe and Lynn Nadel. The hippocampus as a cog- nitive map. Oxford university press, 1978. 1 [24] OpenAI. Gpt-4v(ision) system card. https://cdn. openai.com/papers/GPTV_System_Card.pdf, 2023. 1 [25] Jianing Qi, Jiawei Liu, Hao Tang, and Zhigang Zhu. Be- yond semantics: Rediscovering spatial awareness in vision- language models. arXiv preprint arXiv:2503.17349, 2025. 1 5 [26] Santhosh Kumar Ramakrishnan, Erik Wijmans, Philipp Kraehenbuehl, and Vladlen Koltun.Does spatial cog- nition emerge in frontier models?arXiv preprint arXiv:2410.06468, 2024. 1 [27] Daniela Schiller, Howard Eichenbaum, Elizabeth A Buffalo, Lila Davachi, David J Foster, Stefan Leutgeb, and Charan Ranganath. Memory and space: towards an understanding of the cognitive map. Journal of Neuroscience, 35(41):13904– 13911, 2015. 1 [28] Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, et al. Openai gpt-5 system card. arXiv preprint arXiv:2601.03267, 2025. 1 [29] Zhaochen Su, Peng Xia, Hangyu Guo, Zhenhua Liu, Yan Ma, Xiaoye Qu, Jiaqi Liu, Yanshu Li, Kaide Zeng, Zhengyuan Yang, et al. Thinking with images for multimodal reasoning: Foundations, methods, and future frontiers. arXiv preprint arXiv:2506.23918, 2025. 1 [30] Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. Vggt: Vi- sual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 5294–5306, 2025. 2 [31] Yuxin Wang, Lei Ke, Boqiang Zhang, Tianyuan Qu, Hanxun Yu, Zhenpeng Huang, Meng Yu, Dan Xu, and Dong Yu. N3d-vlm: Native 3d grounding enables accurate spa- tial reasoning in vision-language models. arXiv preprint arXiv:2512.16561, 2025. 1 [32] Yifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Jiangmiao Pang, Chunhua Shen, and Tong He. Permutation-equivariant visual geome- try learning. arXiv preprint arXiv:2507.13347, 2025. 2 [33] Mingrui Wu, Zhaozhi Wang, Fangjinhua Wang, Jiaolong Yang, Marc Pollefeys, and Tong Zhang. From indoor to open world: Revealing the spatial reasoning gap in mllms. arXiv preprint arXiv:2512.19683, 2025. 1 [34] Jihan Yang, Shusheng Yang, Anjali W Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie. Thinking in space: How mul- timodal large language models see, remember, and recall spaces. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 10632–10643, 2025. 1, 3, 4 [35] Weichen Zhang, Ruiying Peng, Chen Gao, Jianjie Fang, Xin Zeng, Kaiyuan Li, Ziyou Wang, Jinqiang Cui, Xin Wang, Xinlei Chen, et al. The point, the vision and the text: Does point cloud boost spatial reasoning of large language mod- els? arXiv preprint arXiv:2504.04540, 2025. 1 [36] Zaibin Zhang, Yuhan Wu, Lianjie Jia, Yifan Wang, Zhongbo Zhang, Yijiang Li, Binghao Ran, Fuxi Zhang, Zhuohan Sun, Zhenfei Yin, et al. Think3d: Thinking with space for spatial reasoning. arXiv preprint arXiv:2601.13029, 2026. 1 6