Paper deep dive
DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
Wenping Yin, Ziqi Liu, Naixia Mou, Weijia Li, Danfeng Hong, Hao Li
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/1/2026, 10:35:12 AM
Summary
The paper introduces DisasterTD, a framework for disaster toponym disambiguation that integrates multimodal large language models (MLLMs) for semantic candidate generation with cross-view geolocalization using remote sensing and street-view imagery. Evaluated on the Hurricane Harvey dataset, DisasterTD significantly improves geolocalization accuracy, particularly for ambiguous toponyms, by reducing candidate dispersion through visual verification.
Entities (9)
Relation Signals (6)
DisasterTD → evaluatedon → Hurricane Harvey
confidence 95% · We evaluate DisasterTD on the Hurricane Harvey dataset
DisasterTD → uses → Multimodal Large Language Models
confidence 95% · DisasterTD... integrates multimodal large language model (MLLMs)-based semantic reasoning
DisasterTD → uses → Cross-View Geolocalization
confidence 95% · DisasterTD... integrates... with cross-view geolocalization.
Social Media Imagery → inputto → DisasterTD
confidence 90% · SMI is augmented with collected RSI and SVI to construct a cross-view benchmark... DisasterTD... extract toponyms from SMI
Remote Sensing Imagery → inputto → DisasterTD
confidence 90% · cross-view matching between SMI, remote sensing imagery (RSI)... is used to verify and refine these candidate results.
DinoV2 → usedin → DisasterTD
confidence 90% · A Vision Transformer (ViT)-based visual foundation model, DINOv2, is used to bridge the domain gap... in DisasterTD
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.
Tags
Links
- Source: https://arxiv.org/abs/2607.24856v1
- Canonical: https://arxiv.org/abs/2607.24856v1
Trouble viewing inline? Open PDF directly →
Full Text
77,770 characters extracted from source content.
Expand or collapse full text
IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20261 DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization Wenping Yin , Ziqi Liu, Naixia Mou, Weijia Li, Member, IEEE, Danfeng Hong, Senior Member, IEEE, Hao Li, Member, IEEE Abstract—Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocaliza- tion challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model (MLLMs)-based semantic rea- soning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. A Vision Transformer (ViT)-based visual foundation model, DINOv2, is used to bridge the domain gap between overhead and ground-level imagery. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambigua- tion, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization. Index Terms—Toponymdisambiguiation,geolocalization, cross-view, multimodal LLM, disaster response This is the accepted version of a paper to appear in IEEE Transactions on Geoscience and Remote Sensing. This work was partly supported by the Start-Up Grant (SUG) project “Geospatial Artificial Intelligence for Climate Resilient Urban Environment” from the National University of Singapore. (Corresponding author: Hao Li) Wenping Yin is with the College of Geodesy and Geomatics, Shandong Uni- versity of Science and Technology, Qingdao 266590, China, and also with the School of Environmental Science and Spatial Informatics, China University of Mining and Technology, Xuzhou 221116, China (email: yin@cumt.edu.cn). Ziqi Liu is with the State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan 430079, China (email: lzq677@whu.edu.cn). Naixia Mou is with the College of Geodesy and Geomatics, Shandong University of Science and Technology, Qingdao 266590, China (email: mounaixia@163.com). Weijia Li is with the Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen 518055, China. (email: liwei- jia@sz.tsinghua.edu.cn). Danfeng Hong is with the School of Automation, Southeast University, Nanjing 210096, China (email: hongdanfeng1989@gmail.com). Hao Li is with the Department of Geography, National University of Singapore, Singapore 117568, Singapore (email: hao.li@nus.edu.sg). I. INTRODUCTION N ATURAL disasters such as hurricanes, floods, wildfires, and earthquakes have become increasingly frequent and severe in recent decades, posing growing threats to human lives, infrastructure, and ecosystems [1–3]. Rapid and accurate geolocalization of disaster-related information is crucial for effective emergency response, resource allocation, and situa- tional awareness [4, 5]. In particular, social media imagery (SMI) has emerged as a valuable and increasingly important complement to traditional remote sensing data, providing timely on-the-ground perspectives that reveal details unavail- able in satellite or aerial observations [6, 7]. Unlike satellite or aerial imagery, which may suffer from latency or limited cov- erage, SMIs are generated rapidly by affected populations and can capture fine-grained, ground-level impacts. Furthermore, these images often include contextual cues such as damaged infrastructure, water levels, or fire spread patterns, which are highly informative for rapid damage assessment. However, fully leveraging these crowdsourced data requires extracting reliable geographic information, making robust geolocalization essential for modern disaster management. To address this need, researchers have explored various methods for text- and image-based geolocalization. Textual approaches typically extract explicit or implicit geographic references from unstructured messages and map them to can- didate geolocations [8, 9]. These methods often rely on named entity recognition, gazetteer matching, or contextual language models to identify place mentions, but they can struggle when information is vague or incomplete. Visual methods rely on matching image content to reference databases such as remote sensing imagery (RSI) or street-view imagery (SVI) [10–13]. Common image-based methods include ground-view image re- trieval, cross-view image retrieval, and two-dimensional–three- dimensional (2D–3D) matching, each designed to handle different types of image data [9, 14–16]. Recent advances in natural language processing (NLP) and computer vision (CV), especially with the advent of multimodal large language models (MLLMs) and deep neural networks, have significantly improved the capability of extracting and aligning geographic cues across modalities [17, 18]. Despite these developments, important challenges remain: textual cues can be vague or ambiguous, while visual data are often affected by occlusion, scene complexity, or domain gaps between ground-level and overhead imagery. Text-only and image-only strategies face significant limitations in disaster geolocalization, where infor- mation is often noisy, incomplete, and highly time-sensitive. One major challenge in this context lies in toponym dis- arXiv:2607.24856v1 [cs.CV] 26 Jul 2026 IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20262 ambiguation [19, 20]. Toponyms in SMIs often correspond to multiple geolocations (e.g., common landmarks, chain stores, or generic street signs), making it difficult to directly link them to unique geographic coordinates [21]. This ambiguity is further exacerbated in disaster scenarios, where users fre- quently provide incomplete, colloquial, or misspelled place references. Traditional disambiguation approaches, such as gazetteer lookups, probabilistic models, or contextual cues from surrounding text, often perform poorly when confronted with the short, informal, and noisy nature of user-generated content on social platforms [22, 23]. More recent efforts attempt to mitigate these issues by combining semantic reason- ing with spatial constraints, user metadata, visual grounding, and integrating MLLMs with geo-knowledge [24, 25]. While these strategies demonstrate improvements, they remain sen- sitive to data sparsity and domain shifts, and a robust solution capable of handling unambiguous and ambiguous toponyms in complex disaster settings is still lacking. In this study, we propose DisasterTD, a novel remote sens- ing–based framework for disaster toponym disambiguation that integrates MLLMs with cross-view geolocalization. As shown in Fig. 1, DisasterTD first uses MLLMs to extract toponyms and generate candidate geolocations, followed by cross-view visual matching between SMIs, RSIs, and op- tionally SVIs to refine and verify the results. By combining semantic reasoning with multi-source Earth observation data, DisasterTD addresses the limitations of text-only and image- only strategies in complex disaster geolocalization scenarios. Unlike our previous MLLM-based disaster geolocalization work [17], which mainly used extracted geoinformation as di- rect geolocalization cues, DisasterTD treats MLLM-generated geolocations as candidate geolocations that require further disambiguation. By integrating semantic candidate generation with cross-view visual verification, DisasterTD provides a new dedicated solution for disaster toponym disambiguation. Fig. 1. A conceptual diagram of DisasterTD for disaster geolocalization, illus- trating how MLLM-based toponym extraction and cross-view geolocalization enable toponym disambiguation and accurate geolocalization within a critical response time window. The remainder of this paper is organized as follows: Section 2 reviews related work on geolocalization and toponym disam- biguation. Section 3 details the proposed DisasterTD, includ- ing MLLM-based initial disaster geolocalization and cross- view matching disambiguation strategies. Section 4 presents experimental settings and results on the extended Hurricane Harvey dataset. Section 5 discusses the findings and limita- tions, and Section 6 concludes with future research directions and broader GeoAI and remote sensing applications. I. RELATED WORK A. Remote Sensing Assisted Geolocalization RSI serves as a valuable foundation for large-scale geolo- calization by providing consistent and up-to-date overhead perspectives [26, 27]. Early studies relied heavily on template matching or hand-crafted feature descriptors such as scale- invariant feature transform (SIFT) and speeded-up robust features (SURF) to align ground-level photographs or SMI with corresponding regions in remote sensing datasets [28, 29]. While these methods demonstrated the feasibility of leveraging RSI for geolocalization, their effectiveness was often limited by variations in scale, illumination, and seasonal conditions. The development of high-resolution satellite imagery and the availability of large-scale benchmarks have since enabled more sophisticated research on remote sensing assisted geolocaliza- tion [5, 30]. For example, Workman et al. [31] proposed wide- area geolocalization using aerial reference imagery, showing that global coverage can significantly expand applicability, while Vo and Hays [32] demonstrated how overhead imagery can be used to localize and orient street-view photographs. The rise of deep learning has brought various neural net- work architectures to the forefront of remote sensing–based geolocalization [4, 11]. These models can extract high-level semantic representations from imagery, which are significantly more robust to variations in viewpoint, illumination, and environmental conditions compared with traditional feature- based methods. Cross-view geolocalization has emerged as a particularly effective strategy, where ground-level images or SVI are matched against overhead RSI to bridge perspec- tive differences [33, 34]. Such methods have demonstrated promising applications in urban mapping, navigation, and fine- grained geolocation recognition, demonstrating the potential of integrating RSI for fine-grained geolocalization [14–16]. However, significant challenges remain, including occlusions, complex scene clutter, and the inherent domain gap between ground-level and aerial perspectives, which can limit matching accuracy in heterogeneous environments. More recent studies have explored multimodal frameworks that integrate RSI with additional complementary signals, such as textual metadata, sensor readings, or social media streams [17, 35]. By exploiting these heterogeneous infor- mation sources, such methods enhance robustness to noise, missing data, and environmental variations, thereby improving geolocalization accuracy in complex, real-world scenarios. Shi et al. [36], for example, introduced a spatial-aware feature aggregation mechanism that encodes geometric relationships, leading to more discriminative cross-view feature matching. In disaster contexts, remote sensing assisted geolocalization has proven particularly valuable, providing a scalable way to cross-validate ground-level observations against overhead evidence [4, 37]. However, the problem of ambiguity and incomplete cues in disaster imagery persists, motivating the need for integrated frameworks that combine remote sensing with semantic reasoning and cross-view disambiguation. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20263 B. Text- or Image-Based Geolocalization Text-based geolocalization has been widely studied in NLP and geographic information science, especially within the context of social media, news articles, and disaster reports where textual descriptions often contain explicit or implicit geographic references [38, 39]. Early approaches primarily re- lied on gazetteer matching and rule-based heuristics, mapping extracted toponyms to candidate geolocations through string similarity and hierarchical geographic filters [40]. Probabilistic models and topic-based methods were later introduced to integrate contextual signals such as co-occurring toponyms, temporal patterns, or linguistic cues, thereby improving dis- ambiguation under sparse or noisy data [41, 42]. More recently, neural architectures, particularly transformer-based models such as bidirectional encoder representations from transformers (BERT) and generative pre-trained transformer (GPT), have greatly advanced this field by capturing semantic nuances, handling informal and ambiguous expressions, and leveraging pretrained knowledge to support robust inference [24, 43, 44]. Despite these advances, challenges remain in resolving colloquial place mentions, handling spelling vari- ations, and disambiguating vague toponyms without strong geographic anchors. Image-based geolocalization aims to determine the geolo- cation of a photo or RSI by analyzing its visual content. Traditional methods employed handcrafted descriptors such as SIFT or SURF to detect distinctive keypoints and match them against geo-referenced databases [45, 46]. With the rise of deep learning, convolutional neural networks (CNNs) and vi- sion transformers (ViTs) have become the dominant paradigm, enabling end-to-end feature learning and large-scale image retrieval for geolocalization [47, 48]. These approaches have achieved remarkable progress in tasks ranging from street- view to remote-sensing image alignment, exploiting architec- tural structures, vegetation patterns, or landscape textures as discriminative cues [32, 49]. However, image-based methods often face difficulties when visual scenes are visually repeti- tive, temporally dynamic, or degraded by occlusion, weather, or disaster-related damage, which can obscure geolocation- specific features and limit geolocalization accuracy. Recently, increasing attention has been devoted to multi- modal geolocalization, which integrates text and image infor- mation to overcome the limitations of unimodal approaches. In this paradigm, textual cues such as toponyms, directions, or contextual descriptions are combined with visual features extracted from images to provide complementary signals for geolocalization [17, 50, 51]. Multimodal fusion methods range from simple late-stage integration to advanced cross-modal at- tention mechanisms that align semantic representations across modalities [52, 53]. This combined framework is particularly valuable in disaster scenarios, where text from eyewitnesses or emergency reports may be incomplete or ambiguous, while images can provide concrete yet visually challenging evidence. By jointly leveraging textual semantics and visual appearance, multimodal geolocalization improves robustness in resolving toponym ambiguity, narrowing candidate search spaces, and enhancing fine-grained geolocalization accuracy. Image-based approaches also remain effective, especially when leveraging cross-view visual consistency, providing a practical alternative when textual information is limited or unavailable [4, 5, 27]. C. Toponym Disambiguation Methods Toponym disambiguation addresses the problem of mapping ambiguous toponyms to their correct geographic referents. Traditional methods have largely relied on gazetteers, rule- based heuristics, and spatial constraints, often combining string matching with hierarchical filters such as country, region, or city [22, 23]. While effective in well-structured text, these methods often struggle with user-generated content, where misspellings, abbreviations, colloquial expressions, and incomplete descriptions are common. Probabilistic models and statistical learning techniques were later introduced to incorporate contextual cues such as co-occurring toponyms, nearby entities, or surrounding text [54, 55]. However, their performance remained limited by data sparsity, weak semantic representations, and lack of adaptability across domains. The emergence of neural language models has significantly advanced toponym disambiguation research. Transformer- based models such as BERT [56] and GPT [57] capture contextual semantics more effectively than traditional embed- dings, improving the recognition of implicit and ambiguous geographic cues. Recent works have enhanced these models with external knowledge sources, such as geographic knowl- edge graphs, structured gazetteers, or knowledge-enhanced embeddings, to strengthen spatial reasoning and interpretabil- ity [58, 59]. For example, DeLozier et al. [60] introduced TopoCluster, a gazetteer-independent method that ranks can- didates using geographic word profiles learned from distribu- tional evidence, reducing reliance on exact string matching when gazetteer coverage is sparse or noisy. Hu et al. [24] demonstrates that lightweight open-source MLLMs augmented with geo-knowledge and voting ensembles can deliver com- petitive resolution across diverse benchmarks. Their pipeline emphasizes efficiency and interpretability, making it suitable for time-sensitive scenarios such as crisis monitoring. In addition, several studies have explored integrating MLLMs with spatial reasoning, user metadata, or temporal information to refine disambiguation results [24, 25, 61]. Ambiguous or common toponyms, such as street names, chain stores, and public facilities, remain difficult to resolve using text alone. Although existing toponym disambiguation studies have made progress with textual context, gazetteers, and geographic knowledge, they still mainly rely on semantic information and rarely exploit visual evidence from remote sensing or street-view imagery. This limits their ability to distinguish between multiple plausible geolocations associated with the same toponym, especially in complex scenarios such as disasters, where contextual information is often noisy or incomplete. To address this limitation, we propose DisasterTD, a novel framework that integrates MLLM-based semantic candidate generation with cross-view visual verification for disaster toponym disambiguation. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20264 Accurate Disaster Geolocalization Nails of America THE JOINT chiropractic Crawford 3100 Little Buddy MLLM Retrieval service MLLM MLLM SMI examples Geolocation information Candidate geolocations A single geolocation Multiple geolocations MLLM A single geolocation Geocoding service MLLM-Based Initial Geolocalization Cross-View Toponym Disambiguation (28.3175, -80.6104) ViT-based cross-view geolocalization Geocoding service Toponym extraction RSI of candidate geolocations ... Patch embedding Positional encoding ... ... (30.0527, -95.1822)(30.0527, -95.1822) (29.7805, -95.3822)(29.7805, -95.3822) (29.7388, -95.3720)(29.7388, -95.3720) Disaster Events MLLM-Based Rapid Response Time Window Cross-View Geolocalization McDonald's Springfield Main Street ...... McDonald's Springfield Main Street ...... Nails of America THE JOINT chiropractic Crawford 3100 Little Buddy MLLM Retrieval service MLLM MLLM SMI examples Geolocation information Candidate geolocations A single geolocation Multiple geolocations MLLM A single geolocation Geocoding service ViT-based cross-view geolocalization Geocoding service Toponym extraction RSI of candidate geolocations ... Patch embedding Positional encoding ... ... (30.0527, -95.1822)(30.0527, -95.1822) (29.7805, -95.3822)(29.7805, -95.3822) (29.7388, -95.3720)(29.7388, -95.3720) Patch embedding + Positional encoding (+ CLS token) Global feature (CLS token) L2 normalization cls ... cls ... cls ... cls ... cls ... cls ... SMI matching with candidate RSIs Cross-view feature encoding Similarity computation and Top-K retrieval C r o s s-V i e w T o p o n y m D i s a m b i g u a t i o n ML L M-B a s e d I n i t i a l G e o l o c a l i z a t i o n Fig. 2. Overview of DisasterTD, a novel RSI-assisted framework for disaster toponym disambiguation. The pipeline integrates MLLM-based semantic candidate generation with cross-view visual verification, where SMIs are matched with candidate RSIs and SVI serves as an optional bridge. I. METHODS A. Overall Workflow This study proposes DisasterTD, a novel framework for disaster toponym disambiguation that integrates MLLM-based initial geolocalization with cross-view matching, as shown in Fig. 2. In the first stage, an MLLM (GPT-4o-2024-05- 13) is used to extract toponyms from SMI and generate candidate geolocations, which are then mapped to geographic coordinates through map service queries. In the second stage, cross-view matching is applied to refine these candidates. Using panoramic SVI as an intermediate bridge for cross-view matching between ground-level SMI and overhead RSI, mit- igating the viewpoint gap and improving the geolocalization accuracy of SMI. For images with a single clear toponym, the corresponding candidate RSIs are compared with the input SMI based on visual similarity to filter out incorrect matches and retain the most plausible geolocation. For images contain- ing vague or multiple toponyms (e.g., common toponyms such as “KFC” or “McDonald’s”), cross-view matching is further employed to identify the most plausible geolocation. By com- bining these two stages, the proposed framework effectively mitigates errors from toponym ambiguity and improves the reliability and recall of cross-view geolocalization. B. MLLM-Based Initial Geolocalization In disaster scenarios, SMIs often contain useful geographic cues, such as street signs, shop names, building numbers, or short textual messages. However, these cues are often incom- plete, occluded, or embedded in complex scenes. To better utilize this information, we use MLLMs to extract text-based toponyms directly from images while considering the sur- rounding visual context. Given an input image I , the MLLM extracts a set of textual tokens or phrases T =t 1 ,t 2 ,...,t m , where each t i may correspond to a potential geographic cue. Compared with traditional optical character recognition pipelines, MLLMs are more robust to noise and partial text, and can capture semantically meaningful expressions beyond exact string matching, enabling the extraction of both explicit toponyms and implicit geolocation-related information. The extracted textual set T is then mapped to a candidate to- ponym set G =G 1 ,G 2 ,...,G n , where each G i represents a possible geographic entity. These entities differ in specificity: a full address unique point of interest (POI) often corresponds to a single geolocation, while generic names (e.g., chain IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20265 stores) may correspond to multiple possible geolocations. Accordingly, images are categorized based on the structure of G. When |G| = 1 and the toponym is distinctive, the mapping is relatively direct. When |G| > 1 or the toponyms are ambiguous, spatial relationships among candidates are considered. This can be formulated as selecting a geolocation l that maximizes consistency: l ∗ = arg max l n X i=1 I(l∈L(G i )),(1) where L(G i ) denotes the set of possible geolocations associ- ated with G i . Each toponym G i is linked to external map services to retrieve candidate geolocations. For single, well-defined to- ponyms, a geocoding function f geo : G → R 2 maps the entity directly to coordinates (lat,lon). For ambiguous or multiple toponyms, a place retrieval function f place : G → l 1 ,l 2 ,...,l k generates a candidate set of geolocations. To reduce irrelevant results, the search can be restricted to a disaster-affected region Ω, i.e., l j ∈ Ω. This stage outputs a candidate geolocation set L =l 1 ,l 2 ,...,l k for each image, which is further refined in the subsequent cross-view matching stage. This initial geolocalization strategy follows our previous work [17], and more details can be found therein. C. Cross-View Toponym Disambiguation 1) Toponym Disambiguation Strategies: We apply different strategies to different types of toponyms during disambigua- tion. For SMIs containing a single clear toponym, we first use the MLLM recognition results to obtain initial candidate geolocations and retrieve the corresponding RSIs. SVI is introduced as a bridge to reduce the difficulty caused by cross- view differences. Specifically, SMI, SVI, and RSI are encoded into high-dimensional feature vectors using a ViT-based visual foundation model [62], DINOv2, and cosine similarity is then used to measure matching scores. If a candidate RSI shows high similarity with the input SMI, its geolocation is confirmed as the disambiguation result; otherwise, the process continues with the remaining candidates until the optimal geolocation is identified. For SMIs containing ambiguous or multiple toponyms (e.g., “KFC” or “McDonald’s”), the search range is relaxed during the initial retrieval stage to generate multiple candidate geolocations. The corresponding RSI for these candidates are then obtained, and with SVI as a bridge, each candidate RSI is matched with the SMI through DINOv2-based feature encoding and similarity computation. The geolocation associated with the RSI that is most similar to the input SMI is finally selected as the disambiguation result. 2) DINOv2-Based Cross-View Matching: In cross-view ge- olocalization, there exists a significant domain gap between SMI and RSI, and direct matching often fails to achieve satisfactory results. To address this issue and improve the ac- curacy of toponym disambiguation, we extend our previously proposed cross-view geolocalization method [5]. This method was originally applied to SMI-based geolocalization in disaster scenarios, and in this study it is introduced for the first time into the task of toponym disambiguation. We introduce SVI as an optional intermediate bridge and construct a three-view joint optimization framework of SMI↔SVI↔RSI. As shown in Fig. 2, we use the ViT-based DINOv2 as the backbone for cross-view feature extraction. Each input image, including the query SMI, optional SVI, and candidate RSIs, is divided into image patches, combined with positional encoding and a CLS token, and processed by Transformer encoder blocks to obtain CLS-based global features. After L2 normalization, cosine similarity is computed between the query SMI and each candidate RSI, and the candidates are ranked for Top- K retrieval. By jointly constraining local semantics and global spatial consistency through a multi-objective loss, the method enhances SMI↔RSI matching and provides robust support for toponym disambiguation in complex scenarios. We use a DINOv2 model based on the ViT architecture [63] for cross-view feature encoding. Given an input image I , the model divides it into patches, embeds them into a sequence, and encodes them through multi-layer self-attention to obtain a high-dimensional representation: z = DINOv2(I;θ), z ∈ R d (2) where θ denotes the trainable parameters and d is the feature dimension. To achieve cross-view alignment, features from the three views are jointly optimized with the following objective: L total = λ 1 L(z m ,z s ) + λ 2 L(z s ,z r ) + λ 3 L(z m ,z r )(3) where I m ,I s ,I r are the candidate SMI, SVI, and RSI, and z m ,z s ,z r are their corresponding feature vectors. L(·) is the InfoNCE contrastive loss, λ 1 ,λ 2 ,λ 3 are weighting coefficients, and L total is the overall optimization objective. During training, we initialize the model with DINOv2 self- supervised pretraining weights [63, 64] and use the model trained on the MultiIan dataset collected in disaster scenarios to improve generalization on large-scale unlabeled images. DINOv2 enables the model to learn semantically rich and domain-invariant features by enforcing consistency between representations of different augmented views of the same image without requiring manual labels. This is achieved by minimizing a cross-view consistency loss between the teacher and student outputs [64], which can be formulated as: L DINO =− N X i=1 P (t) (x i )· logP (s) (x i )(4) Here, L DINO denotes the cross-entropy loss, where P (t) (x i ) and P (s) (x i ) are the teacher and student output distributions for the i-th feature, and N is the number of feature dimensions. This objective encourages the student network to align its outputs with those of the teacher across varying augmentations, resulting in stable and semantically consistent representations. Moreover, to complement the global consistency enforced by the class token, DINOv2 incorporates a patch-level objec- tive inspired by iBOT [65]. The loss is defined as: L iBOT =− M X i=1 P (t) i · logP (s) i (5) IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20266 where P (t) i and P (s) i represent the teacher and student distributions over the masked patch tokens at position i, and M is the number of masked patches. This loss promotes local feature alignment between corresponding regions, improving the model’s capacity to learn fine-grained representations. In DisasterTD, the same encoder architecture is used for both SVI and RSI to ensure consistent feature extraction and facilitates end-to-end joint training, ultimately improving geolocalization performance and robustness in complex scenarios. To increase sensitivity to local details, the oveall training ob- jective combines both image-level and patch-level constraints, and incorporates KoLeo regularization to encourage uniform feature distribution: L KoLeo =− 1 N N X i=1 log min j̸=i ∥z i − z j ∥ (6) where N is the number of samples, z i denotes the feature vector of the i-th sample, z j denotes other feature vectors, and |z i − z j | is the Euclidean distance between features. D. Baselines and Accuracy Evaluation To comprehensively evaluate the effectiveness of Disas- terTD, we designed two types of baseline experiments. 1) MLLM-based geolocalization only: GPT-4o is directly used to extract toponyms from SMI, and geographic coordinates are obtained through Google Maps Geocoding or Places Search services, without any cross-view validation or filtering. This baseline reflects the role of MLLM in toponym recognition and initial geolocalization. 2) Cross-view matching only: ge- olocation is inferred by directly matching SMI with RSI, and we further compare the effect of introducing SVI as an intermediate bridge. This baseline relies entirely on cross-view visual consistency, without MLLM-based toponym recognition or candidate generation. Within this framework, we compare two setting in the cross-view disambiguation stage: 1) direct matching between SMI and RSI; 2) joint cross-view matching among SMI, SVI, and RSI, where SVI serves as an intermedi- ate bridge. To further evaluate whether the performance gains come from the proposed semantic candidate-generation and visual-verification framework rather than from a specific visual encoder, we also include several representative cross-view geolocalization models, including ConvNeXt [66], SAIG-D [67], TransGeo [68], and Sample4Geo [11]. These models are evaluated in two ways: replacing only the second-stage visual matching module of DisasterTD while keeping the MLLM-based candidate generation stage unchanged, and di- rectly applying them to SMI↔RSI geolocalization without the proposed candidate-generation and disambiguation framework. We use Geoloc ACC@k (k = 1000, 500, 250, 100, 50 m) as the core metric to evaluate geolocalization performance at multiple spatial scales, as defined in Equation 7. This metric measures the proportion of samples whose predicted geolocation falls within a radius of k meters from the ground truth. As shown in Table I, different distance error thresholds represent different spatial levels in disaster response, ranging from coarse district- or neighborhood-level awareness to fine- grained street- or POI-level geolocalization. Since cross-view TABLE I GEOLOCALIZATION DISTANCE THRESHOLDS AND CORRESPONDING SPATIAL LEVELS. DistanceSpatial LevelDescriptionGranularity 50 mPOI-levelBuilding or POIFine-grained 100 mStreet-levelStreet segmentFine-grained 250 mBlock-levelUrban blockMedium 500 mNeighborhood-levelNeighborhoodMedium 1000 mDistrict-levelUrban areaCoarse retrieval methods are often evaluated with Recall@k, we unify them under the Geoloc ACC metric: the predicted geoloca- tion (Top-1 or Top-k) is selected from the retrieval results, its geodesic distance to the ground truth is computed, and compared with the threshold. This yields Geoloc ACC@1000 m, Geoloc ACC@500 m, Geoloc ACC@250 m, Geoloc ACC@100 m, and Geoloc ACC@50 m. In addition, we report Top-1 mean and median spatial errors (in meters) to com- plement the threshold-based metrics, providing a continuous measure of geolocalization accuracy and robustness. Geoloc ACC@d m = |D pre | dist(D pre ,D i ) < d| P n j=1 |D j | (7) In this metric, a prediction D pre is regarded as correct if its geodesic distance from the ground truth D i is smaller than a predefined threshold d m. The numerator counts the number of samples that satisfy this criterion, while the denominator corresponds to the total number of evaluated samples. The resulting ratio reflects the proportion of accurate predictions under the specified distance constraint. IV. RESULTS A. Data Preparation The disaster-related SMI used in this study are derived from the Hurricane Harvey Twitter dataset [69]. This dataset contains a large number of publicly shared SMI during Hur- ricane Harvey in the United States in 2017, covering diverse scene content and geographic information. We use 1,000 SMI with clear geographic cues that were randomly selected in our Fig. 3. Examples of different SMI types, including images with single or multiple toponyms referring to clear or ambiguous geolocations. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20267 Fig. 4. Data examples. Spatial distribution of NOAA Hurricane Harvey RSI data (a), and examples of cross-view toponym disambiguation samples, including query SMI, candidate RSI, and auxiliary SVI (b). previous work [17]. The scale is comparable to datasets used in existing research [8, 17] and is sufficient for this study. The SMIs are categorized based on the amount and type of geographic information, as shown in Fig. 3. Images with a single clear toponym (272 samples as examples) include SOS messages, phone numbers, street signs, or a single POI, among which 134 are SOS or phone number cases and 138 contain only a single street sign or POI. Images with a single ambigu- ous toponym (275 samples) mainly consist of the remaining single street sign or single POI cases, where the toponym often corresponds to multiple candidate geolocations with higher ambiguity. Images with multiple clear toponyms (302 samples) contain the co-occurrence of several street signs or POIs, where the toponyms are relatively unambiguous. Finally, images with multiple ambiguous toponyms (151 samples) in- volve several ambiguous toponyms appearing simultaneously, representing complex cases that are difficult to resolve directly. To conduct cross-view toponym disambiguation experi- ments, we retrieved the candidate RSI and SVI corresponding to each SMI, thereby constructing a multi-view aligned dataset. Data examples are shown in Fig. 4. RSI from the National Oceanic and Atmospheric Administration (NOAA) provide information on overall spatial structure and environmental pat- terns, which help capture macro-level geographic features. SVI serve as a bridge between SMI and RSI, playing a key role in modeling fine-grained local semantics and object appearance, and effectively reducing the difficulties caused by cross-view differences. By integrating SMI, SVI, and RSI, we constructed a cross-view image dataset for Hurricane Harvey, providing a solid foundation for subsequent cross-view matching and toponym disambiguation. B. MLLM-Based Initial Geolocalization Performance The MLLM-based toponym recognition results are shown in Table I. The method achieves strong performance across different image types, with an overall accuracy of 88.35%. Building on this, the MLLM-based initial geolocalization method parses toponyms in SMI into candidate geolocations and evaluates accuracy at different spatial thresholds. Table I also reports the geolocalization results under thresholds of 1000 m, 500 m, 250 m, 100 m, and 50 m, with the mean and median distance errors. As expected, the accuracy gradually decreases as the threshold becomes stricter, since tighter spatial constraints reduce the tolerance of candidate sets and may exclude some true geolocations. The effectiveness of this method has also been demonstrated to outperform baselines that rely solely on MLLMs without map services, those that rely only on map services without MLLMs, as well as Geoapify and Nominatim [17]. Across all thresholds, SMI with a single clear toponym achieve the best performance, reaching 60.45% at 1000 m and remaining relatively robust at finer scales (49.02% at 250 m and 41.50% at 50 m). This indicates that in cases with strong directional cues, such as SOS messages or phone numbers, MLLM-based initial geolocalization can effectively generate candidates close to the ground truth. The accuracy for multiple clear toponyms is 59.39%, comparable to single clear toponyms, suggesting that when multiple unambiguous toponyms co-occur, the candidate range can still be well constrained. In contrast, single ambiguous toponyms perform the worst, with an accuracy of only 29.69%, dropping further to 20.21% at 250 m and 10.01% at 50 m. Such vague toponyms (e.g., chain restaurants or common venue names) often correspond to a large number of widely distributed candidates, making fine-grained geolocalization highly chal- lenging. Multiple ambiguous toponyms achieve slightly higher accuracy of 31.60%, showing that although ambiguity exists, the combination of several vague toponyms can provide weak constraints and lead to a modest improvement. At stricter thresholds, the decline in accuracy becomes more pronounced. For single clear toponyms, accuracy drops from 60.45% to 52.10% (500 m), 49.02% (250 m), 45.37% (100 m), and 41.50% (50 m). For single ambiguous toponyms, accuracy IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20268 TABLE I MLLM-BASED TOPONYM RECOGNITION AND GEOLOCALIZATION ACCURACIES ACROSS IMAGE CATEGORIES AND DISTANCE THRESHOLDS. THE TYPES INCLUDE IMAGES WITH SINGLE-CLEAR, SINGLE-AMBIGUOUS, MULTIPLE-CLEAR, OR MULTIPLE-AMBIGUOUS TOPONYMS, AND OVERALL PERFORMANCE ACROSS CATEGORIES. THE OPTIMAL VALUE FOR EACH METRIC IS SHOWN IN BOLD. Type Metric RecogACC (%) Geoloc ACC @1000 m (%) Geoloc ACC @500 m (%) Geoloc ACC @250 m (%) Geoloc ACC @100 m (%) Geoloc ACC @50 m (%) Mean Error (km) Median Error (km) Single-Clear88.5860.4552.1049.0245.3741.5014.630.88 Single-Ambiguous92.1729.6922.4820.2111.2510.0128.291.85 Overall-Single90.3844.9937.2134.5428.2225.6721.501.37 Multiple-Clear86.2959.3950.0747.8845.2140.8013.250.91 Multiple-Ambiguous85.1031.6024.4520.9313.9010.4422.881.46 Overall-Multiple85.8950.1341.5338.9034.7730.6816.461.09 Overall Performance88.3547.3139.1736.5131.1927.9419.221.24 decreases sharply from 29.69% to 22.48%, 20.21%, 11.25%, and 10.01%. Even in single clear toponym cases, not all results can be resolved precisely, as errors may still occur during text recognition and geocoding. In such cases, the predicted candidates may be close to the true geolocation but still fall outside stricter thresholds. For multi-toponym cases, the initial accuracy is usually higher than that of single ambiguous toponyms, but as thresholds tighten, the complexity introduced by multiple candidates also causes accuracy to drop steadily. This suggests that while co-occurrence of multiple toponyms can provide spatial constraints, it cannot fully eliminate error accumulation and may even complicate the geolocalization problem. Overall accuracy decreases from 47.31% at 1000 m to 39.17% (500 m), 36.51% (250 m), 31.19% (100 m), and 27.94% (50 m), highlighting the limitations of relying solely on MLLM-based initial geolocalization under strict spatial constraints and underscoring the necessity of cross- view matching for further disambiguation. As shown in Fig. 5, SMI containing only ambiguous toponyms often produce retrieval results with spatially dis- persed candidates. This observation confirms that MLLM- based initial geolocalization is more suitable as a candidate generation step. Its core value lies in extracting a possible spatial range from complex semantic information. Although ambiguity cannot be completely resolved, the true geolocation is usually included within the candidate set, providing essential input and data support for subsequent cross-view matching and toponym disambiguation. Even when the retrieved candidates are widely distributed, this method still significantly reduce the search space compared to exhaustive geographic retrieval. C. DisasterTD Performance with Cross-View Disambiguation 1) Single-Toponym Disambiguation Performance: For im- ages containing a single toponym, MLLM-based initial geolo- calization can usually provide a candidate set that contains the true geolocation, but the direct geolocalization accuracy at fine scales is limited. Fig. 5 (e)–(h) illustrates representative examples, where each query SMI is associated with several candidate RSIs. The green box indicates the correctly matched RSI, while the others correspond to incorrect candidates with varying spatial deviations. Fig. 5 (e) shows an example con- taining the ambiguous toponym “Little Buddy”, where the candidate RSIs are distributed across different geolocations, with distance errors of 17.9 km and 9.6 km for incorrect matches. Fig. 5 (f) presents another ambiguous case with the toponym “Crawford 3100”, where incorrect candidates exhibit distance errors of 13 km and 9 km. These examples highlight that ambiguous toponyms often correspond to geographically dispersed candidates, making accurate geolocalization difficult without additional cross-view constraints. After introducing cross-view matching, the geolocalization accuracy of SMI improves significantly at all spatial thresh- olds. As shown in Table I, the accuracy of single clear toponyms increases from 60.45% to 78.52% at the 1000 m threshold, and to 70.35%, 66.26%, 61.70%, and 55.97% at the 500 m, 250 m, 100 m, and 50 m thresholds, respectively. This demonstrates that while MLLM-based initial geolocalization already generates reasonable candidates near the ground truth, cross-view matching effectively removes incorrect options. The improvement is even more pronounced for single ambigu- ous toponyms: accuracy rises sharply from 29.69% to 65.23% at 1000 m, and to 55.12%, 49.36%, 42.38%, and 36.08% at 500 m, 250 m, 100 m, and 50 m. This shows that cross- view matching plays a crucial role when handling images with vague toponyms (e.g., chain stores or common venue names), as it can filter a large number of widely distributed candidates and significantly reduce errors caused by ambiguity. Further analysis shows that the impact of cross-view match- ing varies across categories. For single clear toponyms, the improvement is moderate and mainly serves to validate and refine initial candidates. For single ambiguous toponyms, where the initial accuracy is low, the cross-view mechanism brings much larger gains, effectively reducing errors caused by toponym ambiguity. Overall performance in single-toponym scenarios reaches 71.84% at 1000 m and remains at 62.69%, 57.71%, 51.99%, and 45.97% under stricter thresholds, show- ing consistent improvements of over 20% compared to the initial results. The mean error is reduced to 11.06 km, with a median error of 0.70 km, indicating that cross-view matching improves accuracy while reducing large spatial deviations. 2) Multi-Toponym Disambiguation Performance: In multi- toponym scenarios, cross-view matching also leads to signif- icant improvements in geolocalization accuracy. Fig. 5 (g) shows a case containing the toponyms “NAILS OF AMER- ICA” and “THE JOINT chiropractic”, where the incorrect candidate RSIs are also close to the ground truth (e.g., 140 m and 230 m), indicating that multiple relatively distinctive toponyms can effectively constrain the search space. Fig. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 20269 Examples of Optimized Cross-View Toponym Disambiguation ResultsExamples of MLLM-Based Initial Geolocalization Results (a) (29.7805, -95.3822) (29.7184, -95.3133)(29.8832, -95.5245) (b) (c) (d) (28.3175, -80.6104) (29.7388, -95.3720) (33.9177, -89.0003) (30.0656, -95.1245) (30.0527, -95.1822) Query SMIGround-Truth RSI Examples of Candidate RSIs (e) (f) (b) (h) (g) Distance Error: 17.9 kmDistance Error: 9.6 km Distance Error: 13 km Distance Error: 9 km Distance Error: 140 m Distance Error: 230 m Distance Error: 38 km Distance Error: 7.1 km Query SMIExamples of Initial Geolocalization ResultsGoogle SVI Verification Examples of Optimized Cross-View Toponym Disambiguation Results Examples of MLLM-Based Initial Geolocalization Results (a) (29.7805, -95.3822) (29.7184, -95.3133)(29.8832, -95.5245) (b) (c) (d) (28.3175, -80.6104) (29.7388, -95.3720) (33.9177, -89.0003) (30.0656, -95.1245) (30.0527, -95.1822) Query SMIGround-Truth RSI Examples of Candidate RSIs (e) (f) (b) (h) (g) Distance Error: 17.9 kmDistance Error: 9.6 km Distance Error: 13 km Distance Error: 9 km Distance Error: 140 m Distance Error: 230 m Distance Error: 38 km Distance Error: 7.1 km Query SMIExamples of Initial Geolocalization ResultsGoogle SVI Verification Examples of Optimized Cross-View Toponym Disambiguation Results Examples of MLLM-Based Initial Geolocalization Results (a) (29.7805, -95.3822) (29.7184, -95.3133)(29.8832, -95.5245) (b) (c) (d) (28.3175, -80.6104) (29.7388, -95.3720) (33.9177, -89.0003) (30.0656, -95.1245) (30.0527, -95.1822) Query SMIGround-Truth RSI Examples of Candidate RSIs (e) (f) (b) (h) (g) Distance Error: 17.9 kmDistance Error: 9.6 km Distance Error: 13 km Distance Error: 9 km Distance Error: 140 m Distance Error: 230 m Distance Error: 38 km Distance Error: 7.1 km Query SMIExamples of Initial Geolocalization ResultsGoogle SVI Verification Fig. 5. Examples of MLLM-based initial geolocalization results and optimized cross-view toponym disambiguation results. SMIs in (a), (c), (e), and (f) contain single ambiguous toponyms; image (b) contains a single clear toponym; images in (d) and (g) contain multiple clear toponyms; and image (h) contains multiple ambiguous toponyms. Candidate geolocations with green borders are correct, while those with gray borders are incorrect. TABLE I GEOLOCALIZATION ACCURACIES OF DISASTERTD ACROSS IMAGE CATEGORIES AND DISTANCE THRESHOLDS. THE TYPES INCLUDE IMAGES WITH SINGLE-CLEAR, SINGLE-AMBIGUOUS, MULTIPLE-CLEAR, OR MULTIPLE-AMBIGUOUS TOPONYMS, AND OVERALL PERFORMANCE ACROSS CATEGORIES. THE OPTIMAL VALUE FOR EACH METRIC IS SHOWN IN BOLD. Type MetricGeoloc ACC @1000 m (%) Geoloc ACC @500 m (%) Geoloc ACC @250 m (%) Geoloc ACC @100 m (%) Geoloc ACC @50 m (%) Mean Error (km)Median Error (km) Single-Clear78.5270.3566.2661.7055.9710.340.60 Single-Ambiguous65.2355.1249.3642.3836.0811.770.80 Overall-Single71.8462.6957.7151.9945.9711.060.70 Multiple-Clear75.8067.9164.7058.9955.5110.830.72 Multiple-Ambiguous62.4550.0445.4138.6633.8013.290.55 Overall-Multiple71.3561.9558.2752.2148.2711.650.66 Overall Performance71.6262.3657.9952.0947.0111.330.68 5 (h) includes the toponyms “McDonald’s”, “Chili’s”, and “SUBWAY”, where incorrect candidates exhibit large spatial deviations (e.g., 38 km), while the correct geolocation is successfully identified through cross-view matching. These examples show that although MLLM-based initial geolocal- ization can include the true geolocation, the candidates may be widely dispersed, and cross-view matching is essential for filtering out inconsistent results. As reported in Table I, for multiple clear toponyms, accuracy increases from 59.39% to 75.80% at 1000 m, from 50.07% to 67.91% at 500 m, and from 45.21% to 64.70% at 250 m. The improvements remain consistent at finer scales (58.99% at 100 m and 55.51% at 50 m), showing that cross-view matching effectively leverages spatial consistency among multiple toponyms to maintain robust performance under stricter thresholds. SMI with multiple ambiguous toponyms present a more challenging case. Such toponyms often correspond to a large number of widely distributed candidates, making MLLM- based retrieval alone insufficient. With cross-view matching, spatial and visual consistency across candidates can be eval- uated, leading to substantial performance gains. Accuracy increases from 31.60% to 62.45% at 1000 m, from 24.45% to 50.04% at 500 m, and from 13.90% to 38.66% at 100 m, with consistent improvements at 250 m (45.41%) and 50 m (33.80%). These results show that cross-view matching significantly reduces errors caused by ambiguity and improves the reliability of geolocalization. The average accuracy in multi-toponym scenarios reaches 71.35%, 61.95%, 58.27%, 52.21%, and 48.27% at 1000 m, 500 m, 250 m, 100 m, and 50 m, respectively, with improvements comparable to those in single-toponym scenarios. The mean error is reduced to 11.65 km, with a median error of 0.66 km, further indicating improved spatial precision. In multi- toponym cases, cross-view matching plays a dual role: for mul- tiple clear toponyms, it mainly refines and validates candidates, while for multiple ambiguous toponyms, it becomes the key mechanism for identifying the correct geolocation from a large and dispersed candidate set. Fig. 6 provides a visualization of correctly localized points at the 1000 m threshold, further demonstrating the effectiveness of the proposed DisasterTD. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 202610 Fig. 6. Visualization of disaster geolocalization results within a 1000 m error range. Example SMIs in (a) correspond to the geolocalization points shown in the zoomed-in map view in (b). The gap between the mean and median errors suggests a long-tailed error distribution. Most predictions are close to the ground truth, as reflected by the low median error of 0.68 km, while a few severe failures increase the mean error to 11.33 km. These large-error cases mainly occur when candidate geolocations are highly dispersed, the correct geolocation is missing or poorly represented, or visually similar buildings, road segments, and surrounding layouts cause incorrect cross- view matching. This suggests that DisasterTD is effective for most samples, but its reliability can still be affected by candidate completeness and visual ambiguity in difficult cases. D. Baseline Performance Comparison The MLLM-based initial geolocalization method described earlier serves as the baseline for this study. The MLLM-only approach achieves accuracies of 47.31%, 39.17%, 36.51%, 31.19%, and 27.94% at 1000 m, 500 m, 250 m, 100 m, and 50 m, respectively, which are clearly lower than those of cross-view methods. This result indicates that relying solely on semantic candidate generation is insufficient for reliable fine-grained geolocalization. To further examine the role of visual matching, we introduce two direct cross-view geolo- calization baselines, namely SMI↔RSI and SMI↔SVI↔RSI. The SMI↔RSI method achieves 56.84%, 45.17%, 40.05%, 33.92%, and 28.99%, consistently outperforming the MLLM- only baseline across all thresholds. When SVI is introduced as an intermediate representation, the performance further improves to 63.26%, 51.38%, 46.22%, 40.57%, and 35.24%, suggesting that SVI helps reduce cross-view discrepancies and provides more stable visual correspondence. Despite these improvements, both direct cross-view methods still show no- ticeable performance degradation as the distance threshold de- creases, indicating that image-only matching remains sensitive to viewpoint differences and scene complexity. As shown in Fig. 7 (a), the proposed DisasterTD achieves 71.62%, 62.36%, 57.99%, 52.09%, and 47.01% across the same thresholds, maintaining clear advantages over all base- lines. Compared with direct cross-view methods, the improve- ment is consistent rather than concentrated at a specific scale, suggesting that the integration strategy contributes across the entire geolocalization range. The advantage becomes more apparent at finer spatial resolutions, where competing methods show a sharper decline. In particular, the accuracy remains above 50% at 100 m, while both SMI↔RSI and SMI↔SVI↔RSI fall below this level. This behavior indicates that the proposed method is less affected by local ambiguities and can better preserve spatial precision when narrowing down candidate geolocations. The trend is also reflected in Fig. 7 (b), where the improvement margins remain relatively stable across thresholds, instead of fluctuating significantly with distance. As shown in Fig. 7 (c) and (d), we further evaluate several representative cross-view geolocalization models, including ConvNeXt, SAIG-D, TransGeo, and Sample4Geo. Fig. 7 (c) reports the results when only the second-stage visual matching IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 202611 module of DisasterTD is replaced by these models, while the MLLM-based candidate generation stage remains unchanged. Fig. 7 (d) shows the results when these models are directly applied to SMI↔RSI geolocalization without the proposed candidate-generation and disambiguation framework. Across all four models, incorporating them into DisasterTD leads to much higher accuracy than direct SMI↔RSI geolocaliza- tion. For example, Sample4Geo-based DisasterTD achieves 66.08%, 55.23%, 51.50%, 45.37%, and 38.68% from 1000 m to 50 m, whereas direct SMI↔RSI Sample4Geo reaches only 34.33%, 28.82%, 24.91%, 20.50%, and 19.25%. Similar trends are observed for ConvNeXt, SAIG-D, and TransGeo. These results show that the gains of DisasterTD mainly come from combining MLLM-based semantic candidate generation with cross-view visual verification, rather than from a specific visual encoder. The proposed DisasterTD combines candidate generation and filtering in a more structured manner, allowing the system to first narrow down the search space and then perform targeted verification. This design reduces the impact of both semantic ambiguity and visual mismatch, leading to more stable performance across different spatial scales. Compared with direct cross-view matching, which relies entirely on visual similarity, the additional semantic constraint helps elim- inate implausible candidates at an earlier stage, making the subsequent matching process more focused. As a result, even under strict geolocalization constraints, the method maintains relatively high accuracy and avoids the rapid degradation observed in baseline approaches. These results suggest that image-only cross-view matching is insufficient to fully resolve toponym ambiguity, and that effective fine-grained geolocal- ization in disaster scenarios requires the joint use of semantic candidate generation and cross-view disambiguation. V. DISCUSSION This study presents DisasterTD, a framework that integrates multimodal LLM-based semantic reasoning with cross-view geolocalization for disaster toponym disambiguation. The pro- posed two-stage framework first generates candidate geoloca- tions using MLLMs and then refines these candidates through cross-view visual matching. By explicitly combining semantic cues with spatial and visual evidence, DisasterTD reduces candidate ambiguity and improves geolocalization accuracy across multiple distance thresholds. Compared with text-only or image-only approaches, the framework provides a more balanced and robust solution: semantic reasoning narrows the search space, while cross-view matching further enforces spatial consistency and visual correspondence. This synergy is particularly important in disaster scenarios, where information is often incomplete, heterogeneous, and time-sensitive. The use of a ViT-based visual foundation model further enhances feature alignment across heterogeneous views, enabling more reliable matching between SMI, SVI, and RSI. DisasterTD provides a practical and scalable solution for fine-grained geolocalization under severe toponym ambiguity, addressing an important challenge in disaster-related GeoAI applications. Despite these advantages, several limitations remain. The current evaluation is mainly conducted on the Hurricane Harvey dataset, which may limit the generalizability of Disas- terTD to other disaster types, geographic regions, or linguistic contexts. Disaster-related imagery and textual expressions vary considerably across events, and the model may be sensitive to such domain shifts and regional differences. The framework also relies on external map services for candidate generation and SVI retrieval, which may introduce biases due to uneven global coverage, differences in data quality, or missing infor- mation in less-mapped regions. In practice, these limitations may reduce the completeness of candidate sets and propagate errors to the subsequent cross-view matching stage. Temporal inconsistency is another important challenge. SMI collected during disasters often capture rapidly changing scenes, while the corresponding RSI or SVI may be acquired at different times, leading to discrepancies in visual appear- ance and environmental conditions. Such temporal gaps can weaken cross-view matching performance, especially in heav- ily damaged or evolving environments. From a computational perspective, the SMI↔SVI↔RSI retrieval process can become costly when the candidate pool is large, requiring multiple rounds of feature extraction and similarity computation. Since this study does not include a systematic runtime, latency, or throughput benchmark, the current evaluation focuses mainly on geolocalization accuracy rather than operational efficiency. This may limit scalability and delay response time in real- world emergency scenarios where efficiency is critical. These observations suggest that further improvements are still needed in data coverage, temporal alignment, retrieval efficiency, and model robustness to support large-scale practical deployment. Future work can further extend DisasterTD in several di- rections. Evaluating DisasterTD on a wider range of disaster events and multilingual datasets would help better understand its behavior under different conditions and reduce potential bias toward a single benchmark. Incorporating temporal in- formation, such as event timelines or time-aware retrieval strategies, may improve alignment between dynamic disaster scenes and relatively static reference imagery. Efficiency can be further improved by introducing lightweight feature ex- tractors, approximate nearest neighbor search, or hierarchical filtering strategies that progressively narrow down candidate sets. Additional geographic priors, including road networks, POI distributions, or other forms of volunteered geographic in- formation, may provide complementary constraints, especially in cases with highly ambiguous toponyms. It is also worth exploring adaptive strategies that adjust the balance between semantic and visual cues depending on input quality. Integrat- ing uncertainty estimation and human-in-the-loop mechanisms could make the system more transparent and reliable, which is important for high-stakes disaster response tasks where fully automated decisions may not always be sufficient. VI. CONCLUSION This study proposes DisasterTD, a cross-view toponym disambiguation framework that integrates MLLM-based can- didate generation with ViT-enhanced cross-view matching across SMI, SVI, and RSI. Experimental results show that the proposed framework substantially improves disaster geolocal- ization accuracy under diverse toponym ambiguity scenarios. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 202612 (a) Performance comparison of DisasterTD and baseline methods (b) Improvement of DisasterTD over the baselines (c) Performance of DisasterTD with different baseline models (d) Direct geolocalization performance of different baseline models Fig. 7. Geolocalization performance comparison. (a) Performance of DisasterTD and baseline methods. (b) Improvement of DisasterTD over the baselines. (c) Performance of DisasterTD with different cross-view geolocalization models as the visual matching module. (d) Direct SMI-to-RSI geolocalization performance of these baseline models without candidate generation and disambiguation. Across five spatial thresholds (i.e., 1000 m, 500 m, 250 m, 100 m, and 50 m), DisasterTD consistently outperforms all baselines while also reducing mean and median geolocaliza- tion errors. DisasterTD achieves geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, with a mean error of 11.33 km and a median error of 0.68 km. Compared with the MLLM-only baseline, it achieves an average improvement of 21.79%; relative to the direct cross- view matching between SMI and RSI, the gain is 17.22%; and compared with the joint cross-view matching pipeline among SMI, SVI, and RSI, it still delivers a 10.88% improvement. These results demonstrate that combining semantic reasoning with multi-view visual evidence provides a more reliable solution for fine-grained disaster geolocalization. Future work will extend the framework to broader disaster scenarios and geographic contexts while exploring more efficient cross-view retrieval strategies for real-time disaster response applications. REFERENCES [1] M. Dietze and U. Ozturk, “A flood of disaster response chal- lenges,” Science, vol. 373, no. 6561, p. 1317–1318, 2021. [2] D. Nohrstedt, J. Hileman, M. Mazzoleni, G. Di Baldassarre, and C. F. Parker, “Exploring disaster impacts on adaptation actions in 549 cities worldwide,” Nature communications, vol. 13, no. 1, p. 3360, 2022. [3] M. Shirmohammadi, S. Pirasteh, H. Li, M. Akhavan, V. Isazadeh, J. Ji, C. Chen, and Y. Muhammad, “Challenges and opportunities in flood mapping and modeling of next-generation geospatial intelligence: a review,” Geomatics, Natural Hazards and Risk, vol. 17, no. 1, p. 2660866, 2026. [4] H. Li, F. Deuser, W. Yin, X. Luo, P. Walther, G. Mai, W. Huang, and M. Werner, “Cross-view geolocalization and disaster map- ping with street-view and vhr satellite imagery: A case study of hurricane ian,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 220, p. 841–854, 2025. IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 202613 [5] W. Yin, F. Deuser, Z. Liu, J. Wei, X. Luo, M. Werner, H. Li, and Y. Xue, “Triple-objective cross-view geolocalization of disaster- related vgi: the case of hurricane ian,” International Journal of Geographical Information Science, p. 1–23, 2025. [6] Y. Feng, X. Huang, and M. Sester, “Extraction and analysis of natural disaster-related vgi from social media: review, opportu- nities and challenges,” International Journal of Geographical Information Science, vol. 36, no. 7, p. 1275–1316, 2022. [7] X. Hu, Z. Zhou, H. Li, Y. Hu, F. Gu, J. Kersten, H. Fan, and F. Klan, “Location reference recognition from texts: A survey and comparison,” ACM Computing Surveys, vol. 56, no. 5, p. 1–37, 2023. [8] Y. Hu, G. Mai, C. Cundy, K. Choi, N. Lao, W. Liu, G. Lakhan- pal, R. Z. Zhou, and K. Joseph, “Geo-knowledge-guided gpt models improve the extraction of location descriptions from disaster-related social media messages,” International Journal of Geographical Information Science, vol. 37, no. 11, p. 2289– 2318, 2023. [9] G. Huang, Y. Zhou, X. Hu, L. Zhao, and C. Zhang, “A survey of the research progress in image geo-localization,” J. Geo-Inf. Sci, vol. 25, no. 7, p. 1336–1362, 2023. [10] J. Zhuang, M. Dai, X. Chen, and E. Zheng, “A faster and more effective cross-view matching method of uav and satellite im- ages for uav geolocalization,” Remote Sensing, vol. 13, no. 19, p. 3979, 2021. [11] F. Deuser, K. Habel, and N. Oswald, “Sample4geo: Hard neg- ative sampling for cross-view geo-localisation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, p. 16 847–16 856. [12] N. Gong, L. Li, J. Sha, X. Sun, and Q. Huang, “A satellite-drone image cross-view geolocalization method based on multi-scale information and dual-channel attention mechanism,” Remote Sensing, vol. 16, no. 6, p. 941, 2024. [13] D. Hong, C. Li, N. Yokoya, B. Zhang, X. Jia, A. Plaza, P. Gamba, J. A. Benediktsson, and J. Chanussot, “Hyperspectral imaging,” Nature Reviews Methods Primers, vol. 6, no. 1, p. 19, 2026. [14] C. Fang, J. Gao, P. Han, C. Zhao, and B. Gao, “Scof: Supervised contrastive orthogonal fusion for robust cross-view geolocaliza- tion,” IEEE Transactions on Geoscience and Remote Sensing, vol. 63, p. 1–15, 2025. [15] H. Li, C. Xu, W. Yang, L. Mi, H. Yu, H. Zhang, and G.- S. Xia, “Unsupervised multi-view uav image geo-localization via iterative rendering,” IEEE Transactions on Geoscience and Remote Sensing, 2025. [16] J. Liang, M. Bao, H. Dong, L. Xie, R. W. Liu, and N. Chen, “Dstg: Distillation swin transformer for cross-view geo-localization,” IEEE Transactions on Geoscience and Re- mote Sensing, 2025. [17] W. Yin, Y. Xue, Z. Liu, H. Li, and M. Werner, “Llm-enhanced disaster geolocalization using implicit geoinformation from multimodal data: A case study of hurricane harvey,” Interna- tional Journal of Applied Earth Observation and Geoinforma- tion, vol. 137, p. 104423, 2025. [18] D. Hong, C. Li, X. Li, G. Camps-Valls, and J. Chanus- sot, “Foundation models in remote sensing: Evolving from unimodality to multimodality,” IEEE Geoscience and Remote Sensing Magazine, 2026. [19] X. Hu, Y. Sun, J. Kersten, Z. Zhou, F. Klan, and H. Fan, “How can voting mechanisms improve the robustness and generaliz- ability of toponym disambiguation?” International Journal of Applied Earth Observation and Geoinformation, vol. 117, p. 103191, 2023. [20] H. Li, F. Deuser, W. Yin, S. Knoblauch, W. Zhao, F. Biljecki, Y. Xue, and W. Huang, “Towards generative location awareness for disaster response: A probabilistic cross-view geolocalization approach,” ISPRS Journal of Photogrammetry and Remote Sensing, vol. 237, p. 130–145, 2026. [21] R. Grace, “Toponym usage in social media in emergencies,” International Journal of Disaster Risk Reduction, vol. 52, p. 101923, 2021. [22] S. E. Middleton, G. Kordopatis-Zilos, S. Papadopoulos, and Y. Kompatsiaris, “Location extraction from social media: Geop- arsing, location disambiguation, and geotagging,” ACM Trans- actions on Information Systems (TOIS), vol. 36, no. 4, p. 1–27, 2018. [23] J. Fize, L. Moncla, and B. Martins, “Deep learning for toponym resolution: Geocoding based on pairs of toponyms,” ISPRS International Journal of Geo-Information, vol. 10, no. 12, p. 818, 2021. [24] X. Hu, J. Kersten, F. Klan, and S. M. Farzana, “Toponym resolution leveraging lightweight and open-source large lan- guage models and geo-knowledge,” International Journal of Geographical Information Science, p. 1–28, 2024. [25] J. Wang, Y. Hu, and K. Joseph, “Neurotpr: A neuro-net toponym recognition model for extracting locations from social media messages,” Transactions in GIS, vol. 24, no. 3, p. 719–735, 2020. [26] W. Li, R. Dong, H. Fu, J. Wang, L. Yu, and P. Gong, “Integrating google earth imagery with landsat data to improve 30-m res- olution land cover mapping,” Remote Sensing of Environment, vol. 237, p. 111563, 2020. [27] D. Hong, B. Zhang, H. Li, Y. Li, J. Yao, C. Li, M. Werner, J. Chanussot, A. Zipf, and X. X. Zhu, “Cross-city matters: A multimodal remote sensing benchmark dataset for cross-city semantic segmentation using high-resolution domain adaptation networks,” Remote Sensing of Environment, vol. 299, p. 113856, 2023. [28] D. M. Chen, G. Baatz, K. K ̈ oser, S. S. Tsai, R. Vedantham, T. Pylv ̈ an ̈ ainen, K. Roimela, X. Chen, J. Bach, M. Pollefeys et al., “City-scale landmark identification on mobile devices,” in CVPR 2011. IEEE, 2011, p. 737–744. [29] J. Hays and A. A. Efros, “Im2gps: estimating geographic information from a single image,” in 2008 ieee conference on computer vision and pattern recognition. IEEE, 2008, p. 1–8. [30] M. Zhai, Z. Bessinger, S. Workman, and N. Jacobs, “Predicting ground-level scene layout from aerial imagery,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, 2017, p. 867–875. [31] S. Workman, R. Souvenir, and N. Jacobs, “Wide-area image geolocalization with aerial reference imagery,” in Proceedings of the IEEE International Conference on Computer Vision, 2015, p. 3961–3969. [32] N. N. Vo and J. Hays, “Localizing and orienting street views using overhead imagery,” in European conference on computer vision. Springer, 2016, p. 494–509. [33] Y. Shi, X. Yu, L. Liu, T. Zhang, and H. Li, “Optimal feature transport for cross-view image geo-localization,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, no. 07, 2020, p. 11 990–11 997. [34] Z. Zheng, Y. Wei, and Y. Yang, “University-1652: A multi- view multi-source benchmark for drone-based geo-localization,” in Proceedings of the 28th ACM international conference on Multimedia, 2020, p. 1395–1403. [35] M. Wu and Q. Huang, “Im2city: image geo-localization via multi-modal learning,” in Proceedings of the 5th ACM SIGSPA- TIAL International Workshop on AI for Geographic Knowledge Discovery, 2022, p. 50–61. [36] Y. Shi, L. Liu, X. Yu, and H. Li, “Spatial-aware feature aggre- gation for image based cross-view geo-localization,” Advances in Neural Information Processing Systems, vol. 32, 2019. [37] T. Kustu and A. Taskin, “Deep learning and stereo vision based detection of post-earthquake fire geolocation for smart cities within the scope of disaster management: ̇ Istanbul case,” International journal of disaster risk reduction, vol. 96, p. 103906, 2023. [38] J. L. Leidner, “Toponym resolution in text: annotation, eval- uation and applications of spatial grounding,” in ACM SIGIR IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 202614 Forum, vol. 41, no. 2.ACM New York, NY, USA, 2007, p. 124–126. [39] I. Lourentzou, A. Morales, and C. Zhai, “Text-based geolocation prediction of social media users with neural networks,” in 2017 IEEE International Conference on Big Data (Big Data). IEEE, 2017, p. 696–705. [40] E. Amitay, N. Har’El, R. Sivan, and A. Soffer, “Web-a-where: geotagging web content,” in Proceedings of the 27th annual international ACM SIGIR conference on Research and develop- ment in information retrieval, 2004, p. 273–280. [41] B. Wing and J. Baldridge, “Simple supervised document geolo- cation with geodesic grids,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, 2011, p. 955–964. [42] J. Eisenstein, B. O’Connor, N. A. Smith, and E. Xing, “A latent variable model for geographic lexical variation,” in Proceed- ings of the 2010 conference on empirical methods in natural language processing, 2010, p. 1277–1287. [43] E. Kamalloo and D. Rafiei, “A coherent unsupervised model for toponym resolution,” in Proceedings of the 2018 world wide web conference, 2018, p. 1287–1296. [44] Y. S. Bicakci, J. Shingleton, and A. Basiri, “Street- level geolocalization using multimodal large language mod- els and retrieval-augmented generation,” arXiv preprint arXiv:2509.01341, 2025. [45] D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, no. 2, p. 91–110, 2004. [46] H. Bay, T. Tuytelaars, and L. Van Gool, “Surf: Speeded up robust features,” in European conference on computer vision. Springer, 2006, p. 404–417. [47] T. Weyand, I. Kostrikov, and J. Philbin, “Planet-photo ge- olocation with convolutional neural networks,” in European conference on computer vision. Springer, 2016, p. 37–55. [48] Q. Yi and L. Shan, “Geolocsft: Efficient visual geolocation via supervised fine-tuning of multimodal foundation models,” arXiv preprint arXiv:2506.01277, 2025. [49] A. Toker, Q. Zhou, M. Maximov, and L. Leal-Taix ́ e, “Com- ing down to earth: Satellite-to-street view synthesis for geo- localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, p. 6488– 6497. [50] G. Tahmasebzadeh, E. M ̈ uller-Budack, and R. Ewerth, “Mul- timodal geolocation estimation in news documents,” in Event Analytics across Languages and Communities. Springer Nature Switzerland Cham, 2024, p. 17–45. [51] J. Ye, H. Lin, L. Ou, D. Chen, Z. Wang, Q. Zhu, C. He, and W. Li, “Where am i? cross-view geo-localization with natural language descriptions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, p. 5890– 5900. [52] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning. PmLR, 2021, p. 8748–8763. [53] J. Li, D. Li, S. Savarese, and S. Hoi, “Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,” in International conference on machine learning. PMLR, 2023, p. 19 730–19 742. [54] D. Buscaldi and P. Rosso, “A conceptual density-based approach for the disambiguation of toponyms,” International Journal of Geographical Information Science, vol. 22, no. 3, p. 301–313, 2008. [55] K. Roberts, C. A. Bejan, and S. M. Harabagiu, “Toponym disambiguation using events.” in FLAIRS, 2010. [56] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, p. 4171–4186. [57] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al., “Improving language understanding by generative pre-training,” 2018. [58] D. Buscaldi and P. Rosso, “Map-based vs. knowledge-based toponym disambiguation,” in Proceedings of the 5th workshop on geographic information retrieval, 2008, p. 19–22. [59] M. Gritta, M. T. Pilehvar, and N. Collier, “Which melbourne? augmenting geocoding with maps.” Association for Computa- tional Linguistics, 2018. [60] G. DeLozier, J. Baldridge, and L. London, “Gazetteer- independent toponym resolution using geographic word pro- files,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 29, no. 1, 2015. [61] Y. Sun, Y. Ye, J. Kang, R. Fernandez-Beltran, S. Feng, X. Li, C. Luo, P. Zhang, and A. Plaza, “Cross-view object geo- localization in a local region with satellite imagery,” IEEE Transactions on Geoscience and Remote Sensing, vol. 61, p. 1–16, 2023. [62] D. Hong, B. Zhang, X. Li, Y. Li, C. Li, J. Yao, N. Yokoya, H. Li, P. Ghamisi, X. Jia et al., “Spectralgpt: Spectral remote sensing foundation model,” IEEE transactions on pattern analysis and machine intelligence, vol. 46, no. 8, p. 5227–5244, 2024. [63] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Trans- formers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [64] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al., “Dinov2: Learning robust visual features without super- vision,” arXiv preprint arXiv:2304.07193, 2023. [65] M. Caron, H. Touvron, I. Misra, H. J ́ egou, J. Mairal, P. Bo- janowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF interna- tional conference on computer vision, 2021, p. 9650–9660. [66] Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, 2022, p. 11 976–11 986. [67] Y. Zhu, H. Yang, Y. Lu, and Q. Huang, “Simple, effective and general: A new backbone for cross-view image geo- localization,” arXiv preprint arXiv:2302.01572, 2023. [68] S. Zhu, M. Shah, and C. Chen, “Transgeo: Transformer is all you need for cross-view image geo-localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, p. 1162–1171. [69] M. E. Phillips, “Hurricane harvey twitter dataset,” 2024, https://digital.library.unt.edu/ark:/67531/metadc993940.