Paper deep dive
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation
Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh, Seon Ho Kim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/25/2026, 7:14:35 AM
Summary
This study investigates regional geolocalization within Greater Los Angeles using 9,085 street-view images to determine if pretrained CLIP features are sufficient for fine-grained discrimination and what visual cues support performance after adaptation. The authors compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Results show that encoder adaptation significantly improves accuracy (up to 82.10%) and reduces centroid error compared to frozen methods (~39%). Probing via semantic cue removal, appearance reduction (edge/blur), and scene-configuration disruption (patch scrambling) reveals that adapted models become more sensitive to intact scene configuration and achieve higher accuracy on appearance-reduced inputs, but do not rely solely on coarse structure. Environmental cues like vegetation and sky remain influential, and scrambling sensitivity is not unique to geolocalization tasks.
Entities (19)
Relation Signals (8)
CLIP ā usedfor ā Regional Geolocalization
confidence 95% Ā· We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination
LoRA ā achievesaccuracy ā 75.94-82.10%
confidence 90% Ā· encoder adaptation achieves 75.94-82.10%.
Frozen Readouts ā achievesaccuracy ā 39.03%
confidence 90% Ā· Frozen readouts remain near the 39.03% zero-shot accuracy
Full Fine-Tuning ā reducesmeandistance ā 3.86 km
confidence 90% Ā· Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km.
Patch Scrambling ā causespredictionswitch ā 42.92-45.56%
confidence 85% Ā· Adapted models ... switch 42.92-45.56% of predictions after scrambling
Vegetation ā influentialfor ā Regional Geolocalization
confidence 85% Ā· vegetation and sky remain influential.
Sky ā influentialfor ā Regional Geolocalization
confidence 85% Ā· vegetation and sky remain influential.
Caltech101 ā demonstrates ā Scrambling Sensitivity Not Unique
confidence 80% Ā· A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94-82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92-45.56% of predictions after scrambling, compared with 10.79-14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization.
Tags
Links
- Source: https://arxiv.org/abs/2608.21761v1
- Canonical: https://arxiv.org/abs/2608.21761v1
Trouble viewing inline? Open PDF directly ā
Full Text
44,414 characters extracted from source content.
Expand or collapse full text
What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation*Corresponding author: Seon Ho Kim Changyu Lee Affiliation: University of Southern California Los Angeles, USA clee1806@usc.edu Yeonsoo Park Affiliation: USC IMSC San Francisco, USA tripwithsoo@gmail.com Abdullah Alfarrarjeh Affiliation: German Jordanian University Amman, Jordan abdullah.alfarrarjeh@gju.edu.jo Seon Ho Kim* Affiliation: Integrated Media Systems Center (IMSC) University of Southern California Los Angeles, USA seonkim@usc.edu Abstract Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94ā82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92ā45.56% of predictions after scrambling, compared with 10.79ā14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization. Index Terms: visual geolocalization, CLIP adaptation, regional geolocalization, scene configuration, visual interventions, visionālanguage models Fig. 1: Study design and motivating questions. We first evaluate regional geolocalization across increasing levels of CLIP adaptation. The performance comparison motivates an analysis of the visual information associated with adaptation. Complementary interventions measure selected cue dependence, information retained after appearance reduction, and sensitivity to layout disruption. Their joint interpretation considers both scene-configuration sensitivity and appearance-cue dependence. I Introduction Visual geolocalization aims to infer where an image was captured from its visual content. This capability is useful for applications such as image organization, navigation, geographic information retrieval, and analysis of large-scale visual data. Existing approaches formulate the problem in several ways, including image retrieval, coordinate estimation, and geographic classification [11, 23, 22]. For example, PlaNet divides the world into adaptive geographic cells and predicts the cell associated with an input image [25]. More recently, pretrained visionālanguage models have provided a flexible alternative by representing geographic concepts through text. CLIP-based methods such as StreetCLIP, GeoCLIP, PIGEON, and AddressCLIP have demonstrated promising geolocalization performance across multiple geographic scales [19, 9, 22, 10, 27]. Despite this progress, successful geolocalization over broad geographic areas does not necessarily imply that a model can distinguish between nearby places at high location accuracy. At large scales, locations may differ through relatively prominent visual cues such as language in images, climate-related characteristics such as sky and vegetation, road systems, or street-level architectural styles. These cues may become less distinctive when candidate locations lie within the same metropolitan area. Neighboring areas often share many of the same visual characteristics, leaving models to rely on subtler combinations of local cues. Understanding whether pretrained representations contain enough information for such fine-grained discriminationāand, if not, what additional information models learn through adaptationāis therefore important for characterizing the capabilities of modern geolocalization systems. We study this setting through regional geolocalization, which we define as classification among selected areas within a single metropolitan region. We focus on Greater Los Angeles, where nearby areas share many broad geographic characteristics while still exhibiting variations in streetscape appearance, vegetation, urban form, and other local visual patterns. Our benchmark dataset contains eight classes spanning municipalities, neighborhoods, and districts, which we collectively refer to as regions. These regions are intended to provide a focused testbed for fine-grained geographic discrimination; they do not cover all of Los Angeles and should not be interpreted as representing all cities or metropolitan environments. This setting leads to our first question: Is fine-grained regional information already accessible from a pretrained CLIP representation, or does extracting it require modification of the image encoder? To answer this question, we compare adaptation strategies with progressively greater modification of the pretrained model. These include zero-shot CLIP, frozen-encoder text-score (LP-T) and linear-classifier (LP-C) readouts, Partial Update of the final encoder block, LoRA [14], and full fine-tuning (Full-FT). This spectrum allows us to distinguish methods that only learn a new output mapping from those that modify the visual representation itself. We observe a pronounced difference between these regimes: frozen readouts remain close to the zero-shot baseline, whereas encoder-level adaptation increases accuracy from 39.03% to as high as 82.10%. Thus, in our setting, strong regional discrimination emerges primarily when the pretrained visual representation is allowed to adapt. The large improvement from encoder adaptation raises a second, complementary question: What visual information becomes more important after adaptation? High classification accuracy alone does not reveal how a model distinguishes nearby regions. A street scene contains many potentially informative signals, ranging from individual objects and environmental characteristics to the overall arrangement of scene elements. Two regions, for example, may contain similar vehicles, vegetation, buildings, and roads, yet differ in how these elements typically appear together. We therefore investigate whether adaptation primarily changes sensitivity to particular visual cues, to fine appearance information, or to the broader configuration of the scene. To make this question experimentally tractable, we use controlled interventions that selectively alter different forms of visual information. We remove text, vehicles, vegetation, and sky as complementary and semantically interpretable cues whose prevalence varies across the eight regions and whose removal can largely preserve the surrounding scene. We do not individually remove larger scene-defining components such as roads and buildings because doing so would require extensive inpainting and could substantially change scene geometry. In addition, edge and blur transformations reduce fine appearance information while preserving portions of the sceneās coarse organization, whereas patch scrambling preserves local image content but disrupts its global arrangement. Here, we use spatial structure in a deliberately limited sense: the coarse organization and relative distribution of visual elements within the image, rather than 3-D geometry, symbolic spatial relations, or general spatial reasoning. These interventions provide complementary views of model behavior. Cue removal measures dependence on selected semantic elements; edge and blur inputs test how much predictive information remains when detailed appearance is reduced; and patch scrambling tests whether predictions depend on the intact arrangement of local content. We summarize these effects using accuracy together with two normalized behavioral measures. Retention, defined as transformed accuracy divided by original accuracy, measures how much performance survives an intervention, while prediction switch rate measures the fraction of examples whose predicted label changes. We use retention under appearance reduction to examine whether coarse structural information is sufficient to sustain predictions, and switch rate under scrambling to examine sensitivity to scene configuration. Figure 1 summarizes the overall experimental design. Our results reveal an important distinction between sensitivity to structure and sufficiency of structure. After encoder adaptation, models achieve higher absolute accuracy on edge- and blur-transformed images and change their predictions more frequently when patches are scrambled. However, they do not retain a larger fraction of their original performance once detailed appearance is removed. At the same time, environmental appearance cues such as vegetation and sky remain influential. A matched Caltech101 control further shows that sensitivity to scrambling is not unique to geolocalization, since disrupting spatial arrangement also affects generic object recognition. Taken together, these observations suggest that adaptation makes regional predictions more dependent on intact scene configuration, but does not make structural information alone sufficient for accurate prediction. Instead, adapted models appear to benefit from a combination of scene configuration and appearance-based cues. Our conclusions are intentionally limited to the evaluation setting considered here. In particular, the training and test sets include same-coordinate and nearby views. The experiments therefore measure robustness to viewpoint variation around known locations rather than generalization to geographically disjoint, previously unseen areas. Within this scope, the study uses regional geolocalization as a controlled setting for examining not only whether CLIP can be adapted to distinguish visually similar nearby places, but also how the visual evidence supporting its predictions changes after adaptation. We make three main contributions. ⢠We establish a substantial performance gap between frozen-readout methods and encoder-level adaptation for fine-grained regional geolocalization in the evaluated Greater Los Angeles setting. ⢠We develop a complementary intervention-based analysis using semantic cue removal, appearance reduction, and layout disruption, together with accuracy, retention, and prediction switch rate, to distinguish sensitivity to scene structure from the sufficiency of that structure. ⢠We show that encoder adaptation is associated with greater sensitivity to intact scene configuration without evidence that structural information becomes sufficient on its own; environmental and appearance-based cues remain important to regional predictions. More broadly, this work connects fine-grained geographic prediction with the study of representation adaptation. Rather than treating improved accuracy as the endpoint, we use controlled interventions to examine what information accompanies those gains and how model behavior changes when that information is disrupted. Section I reviews related work. Sections IāV describe the regional geolocalization task, adaptation methods, and probing protocol, respectively. Sections VIāIX present the experimental results, discussion, limitations, and conclusions. I Related Work I-A Visual Geolocalization Visual geolocalization uses classification, retrieval, coordinate prediction, or combinations thereof. PlaNet learns a classifier over geographic cells [25]; Deep Img2GPS combines classification features with retrieval and density estimation [23]; and data-centric work examines image scene localization [1]. OSV-5M emphasizes global coverage and strict geographic trainātest separation [2]. Our independently collected LA dataset does not use OSV-5M images; OSV-5M instead provides a methodological precedent for geographic separation. I-B CLIP Adaptation for Geolocalization CLIP provides transferable imageātext representations [19]. StreetCLIP studies open-domain zero-shot geolocalization [9]; GeoCLIP aligns images with continuous GPS representations [22]; PIGEON combines semantic geocells, contrastive pretraining, and retrieval refinement [10]; and AddressCLIP studies city-wide address prediction [27]. These works establish CLIP-based geographic prediction. Our focus is not a new city-scale application, but the visual information associated with adaptation gains under controlled interventions; predictive performance alone does not identify that information. I-C Probing Object Relations and Spatial Arrangement Downstream accuracy does not by itself imply relational understanding. The Attribution, Relation, and Order (ARO) benchmark probes attributes, relations, and word order [28], while SpatialCLIP targets broader 3-D-inspired spatial understanding [24]. Our interventions test only coarse scene-configuration sensitivity and selected appearance-cue dependence; no individual corruption establishes general spatial reasoning. This motivates a controlled setting in which adaptation and prediction sensitivity can be examined together. I Task and Dataset I-A Los Angeles-Specific Dataset and Operational Definition We construct a Los Angeles street-view dataset through the Google Maps API; it is not derived from OSV-5M [2]. Let =(xi,yi,gi)i=1ND=\(x_i,y_i,g_i)\_i=1^N_D contain image xix_i, one of K=8K=8 region labels yiy_i, and coordinates gig_i, with N=9,085N_D=9,085. The classes are Downtown Los Angeles, Pasadena, Beverly Hills, Santa Monica, Hollywood, Long Beach, Venice Beach, and Koreatown. Four are incorporated cities and the others are recognized neighborhoods or districts; we collectively call them regions and formulate the task as sub-regional classification within greater Los Angeles. Figure 2 summarizes regional sampling, same-coordinate trainātest overlap, and representative scenes selected to show both distinctive and geographically ambiguous views. Fig. 2: Dataset geography and representative scenes for the eight regions. Each group shows sampled locations, a distinctive view, and an ambiguous view. The displayed N is the sample count for that region; annotations report test images whose requested coordinate also occurs in training. Let ā°E denote the evaluation set, Nā°=|ā°|N_E=|E|, and y^i,m y_i,m the prediction of model m for image xix_i. Top-1 accuracy is Acc(m)=1Nā°āiāā°[y^i,m=yi].Acc(m)= 1N_E _i I\! [ y_i,m=y_i ]. (1) Let ckc_k denote the geographic centroid of region k and dā”(gi,ck)d(g_i,c_k) the distance in kilometers between image coordinate gig_i and that centroid. We define centroid error as the mean distance to the centroid of the predicted region: CEā”(m)=1Nā°āāiāā°dā”(gi,cy^i,m).CE(m)= 1N_E _i d(g_i,c_ y_i,m). (2) I-B Dataset Reconstruction and Density-Based Sampling Of the 28,550 raw images, 30 were outside the target regions, leaving 28,520 eligible images. Let nkn_k and AkA_k denote the number of eligible images and area of region k, respectively. To reduce regional density imbalance, we sample at the largest common density, Ļā=minkā”(nk/Ak)=194.49Ļ = _k(n_k/A_k)=194.49 images/km2, determined by Koreatown. Sampling approximately ĻāāAkĻ A_k images per region produces the 9,085-image dataset reported in Table I. This procedure equalizes spatial sampling intensity rather than class frequency; larger regions contribute more images. TABLE I: Area-normalized sampling of the LA dataset. Region Area (km2) Available Sampled Downtown LA 5.13 4,000 998 Hollywood 6.16 4,000 1,197 Beverly Hills 8.98 4,000 1,746 Santa Monica 6.42 4,000 1,248 Venice Beach 3.08 3,921 599 Koreatown 3.08 599 599 Pasadena 7.69 4,000 1,496 Long Beach 6.18 4,000 1,202 Total ā 28,520 9,085 I-C Dataset Splits, Integrity, and Spatial Overlap We use a region-stratified 70/15/15 split: 6,359 training, 1,363 validation, and 1,363 test images, with seed 42. Split metadata and the raw image pool are retained for reproducibility. I-C1 Data Collection and Processing Available metadata primarily indicates collection in early 2026. At each coordinate, the API returned separate perspective images at 0ā0 , 90ā90 , 180ā180 , and 270ā270 rather than stitched panoramas. Polygons define region membership, and coordinate/heading metadata supports the overlap audit. The raw pool and split metadata are preserved for resampling. We did not apply geographic or embedding-based near-duplicate removal; density normalization does not replace coordinate-grouped splitting or duplicate filtering. I-C2 Spatial Overlap and Viewpoint Generalization The image-level split is geographically overlapping: among all 1,363 test records, 755 (55.39%) share a requested coordinate with training and 1,086 (79.68%) lie within 50 m of a training coordinate; the remaining 277 (20.32%) are separated by more than 50 m. Orthogonal headings can show different fields of viewāfor example, a residential street versus a commercial cross-streetāand reduce direct pixel correspondence, yet retain local buildings, intersection geometry, vegetation, road surfaces, and capture conditions. The protocol therefore measures viewpoint variation near known locations, not geographically disjoint generalization. Correct classification from an unseen heading may reflect viewpoint tolerance but cannot exclude micro-location memorization. Coordinate-grouped or spatially buffered splits are required for stricter evaluation; OSV-5M provides a methodological contrast [2]. Several later transformations retain context shared across headings, so their results should not be interpreted as geographically disjoint generalization. IV Adaptation Methods IV-A Common Training Configuration All methods use OpenCLIPās pretrained ViT-L-14 image encoder [5] with its default 224Ć224224Ć 224 resize, center crop, and normalization; no stochastic augmentation is applied. Supervised methods train for 10 epochs with seed 42, batch size 8, AdamW, weight decay 0.01, cosine annealing with Tmax=10T_ =10, cross-entropy loss, and automatic mixed precision with dynamic loss scaling. We do not use early stopping; the checkpoint with the highest validation accuracy is retained for test evaluation. Training examples are shuffled without weighted sampling or class-weighted loss. IV-B Adaptation Strategies Table I orders the strategies by the extent of model modification. Zero-shot classification uses imageāprompt similarity without regional training. Linear Probing with Text-based Scores (Linear ProbeāText; LP-T) optimizes only the scale and bias of the text-based class scores, whereas Linear Probing with a Classifier (Linear ProbeāClassifier; LP-C) trains a linear classifier over frozen image features. Partial Update (PU) updates the final transformer block, ln_post, and the classifier. Low-Rank Adaptation (LoRA) inserts rank-8 adapters (α=16α=16, dropout 0.1) into the attention and MLP projections while keeping the pretrained encoder weights frozen [14]. Full Fine-Tuning (Full-FT) updates the complete pretrained image encoder and the classifier during training. TABLE I: Configuration of the adaptation strategies. Method Trainable components Learning rate Zero-shot None ā LP-T Text-feature scale and bias 10ā410^-4 LP-C Linear classification layer 10ā410^-4 PU Final visual block, ln_post, and head 10ā510^-5 LoRA Adapters in attention and MLP projections 10ā410^-4 Full-FT Full image encoder and classifier 10ā510^-5 The comparison tests whether regional discrimination is accessible through the evaluated frozen readouts or benefits from encoder modification; lower probe performance does not establish that frozen CLIP lacks regional information. These strategies measure how performance varies as progressively more of the encoder is modified. Performance alone, however, does not identify which visual information supports those changes, motivating the complementary input interventions described next. V Intervention-Based Probing Protocol Intervention-based probing measures predictions after controlled input changes. The three visual intervention families test selected semantic cues, information retained after appearance reduction, and sensitivity to intact spatial arrangement. Because each transformation changes multiple properties, no single result is treated as selective; we interpret the families jointly and report both transformed accuracy and change relative to original-image performance. V-A Cue Selection and Removal We select text, vehicles, vegetation, and sky as complementary, semantically interpretable cues whose removal largely preserves the surrounding scene. Text and vehicles represent semantic or activity-related evidence, whereas vegetation and sky represent environmental context. Larger scene-defining components such as roads and buildings are not individually removed because their removal would require extensive inpainting that could alter scene geometry. Instead, edge and blur transformations probe appearance-reduced information, while patch scrambling tests sensitivity to scene arrangement. We use one-way analysis of variance (ANOVA) to test whether the mean image area occupied by each cue differs across the eight regions. Image-level mask coverage differs significantly for all four cues, most prominently vegetation (F=35.04F=35.04, p<0.001p<0.001) and sky (F=10.93F=10.93, p<0.001p<0.001), followed by vehicles (F=5.65F=5.65, p<0.001p<0.001) and text (F=2.68F=2.68, p=0.009p=0.009). These differences provide dataset-level context for the removal tests, although differences in mask area and segmentation quality prevent causal ranking among the selected cues. Text regions are detected and recognized using EasyOCR, which employs Character Region Awareness for Text Detection (CRAFT) [3] and a Convolutional Recurrent Neural Network (CRNN) [20]. Vehicles are detected using the nano variant of YOLOv8 (YOLOv8n), pretrained on MS-COCO [15, 17]. Vegetation and sky regions are identified using the SegFormer-B0 semantic-segmentation model pretrained on ADE20K [26, 30]. After the corresponding regions are masked, Telea inpainting [21] is used to fill the removed areas. For each semantic cue mask, we also generate a random mask covering the same number of pixels, providing a control for the effect of removing an equivalent image area irrespective of semantic content. For transformation T, let y^i,m(T) y_i,m^(T) denote the prediction on Tā”(xi)T(x_i) and Accā”(m,T)Acc(m;T) its accuracy. Following feature-removal studies [29, 13], for cue q and removal TqT_q we report intervention-induced accuracy drop (IAD): IADā”(m,q)=Accā”(m)āAccā”(m;Tq),IAD(m,q)=Acc(m)-Acc(m;T_q), (3) and prediction switch rate: SR(m,q)=1Nā°āiāā°[y^i,mā y^i,m(Tq)].SR(m,q)= 1N_E _i I\! [ y_i,mā y_i,m^(T_q) ]. (4) SR is the fraction of labels changed by removal and measures instability rather than correctness [12]. The cue-removal experiment therefore characterizes sensitivity to the implemented removals rather than causal cue importance. V-B Appearance-Reduced Structure Proxies Canny edges [4] and Gaussian macro blur reduce fine appearance while approximately preserving coarse organization. They do not isolate structure: edge images retain text outlines, vehicle shapes, windows, and vegetation boundaries, while blur retains color and coarse texture. Performance therefore cannot be attributed exclusively to spatial structure. Alongside transformed accuracy, we report retention, which normalizes by each modelās original accuracy: Retentionā”(m,T)=Accā”(m,T)Accā”(m).Retention(m,T)= Acc(m;T)Acc(m). (5) V-C Spatial-Layout Disruption Patch scrambling randomly permutes a regular image grid, preserving block content while disrupting global arrangement; Patch-n uses nĆnĆ n-pixel blocks. Larger blocks retain more coherent local content. Scrambling also introduces unnatural boundaries, object fragmentation, distribution shift, and interactions with positional embeddings. Accordingly, we interpret prediction changes as layout sensitivity rather than semantic spatial reasoning or uniquely geographic structure use. V-D Prompt Controls For zero-shot CLIP, the baseline template is āa Google Street View photo in Region, Los Angeles.ā Length-controlled prompts add semantically neutral filler. Visual-only prompts remove the region name and retain only architecture, vegetation, and streetscape descriptors. Descriptor-swapped prompts cyclically assign each descriptor to the next region as a falsification control. Chance accuracy is 12.5%. Because swapping changes both descriptor content and label association, it does not establish individual descriptor grounding. VI Results Fig. 3: Intervention summary: (a) regional gains, (b) cue removal and matched random masks, (c) accuracy and retention after appearance reduction, and (d) layout disruption in geolocalization and Caltech101. Panels (c) and (d) contrast structural sufficiency and sensitivity. Acc: accuracy; Ctrl: control; Geo: geolocalization; S: structural-sufficiency retention. VI-A Overall Adaptation Performance TABLE I: Adaptation performance on original test images. Lower centroid error is better. Strategy Accuracy (%) Centroid Error (km) Zero-shot 39.03 12.30 LP-T 41.45 12.25 LP-C 38.52 11.22 PU 75.94 5.08 LoRA 78.36 4.21 Full-FT 82.10 3.86 LP-T and LP-C remain near zero-shot accuracy, whereas PU, LoRA, and Full-FT gain 36.91, 39.33, and 43.07 points, respectively (Table I). Full-FT also reduces centroid error from 12.30 to 3.86 km. Under the evaluated readouts and spatially overlapping split, a frozen linear classifier does not recover the discrimination achieved by encoder adaptation. This result is conditional on probe selection and does not imply that frozen CLIP contains no regional information. VI-B Variation Across Regions TABLE IV: Per-region accuracy and percentage-point gain after full fine-tuning. Region Zero-shot (%) Full-FT (%) Gain Downtown LA 35.57 90.60 +55.03 Pasadena 35.56 90.22 +54.66 Beverly Hills 60.69 84.35 +23.66 Santa Monica 36.36 82.35 +45.99 Hollywood 16.11 76.67 +60.56 Long Beach 36.67 75.56 +38.89 Venice Beach 68.89 75.56 +6.67 Koreatown 16.67 71.11 +54.44 Zero-shot accuracy is highest in Venice Beach and Beverly Hills and near 16% in Hollywood and Koreatown (Table IV). Full-FT gains range from 6.67 points in Venice Beach to 60.56 points in Hollywood, showing that the aggregate improvement is not uniform. Class-name alignment, landmark prevalence, and sampling density may contribute, but the present data do not isolate these mechanisms; doing so would require class-normalized confusion analysis and spatially blocked evaluation. VI-C Cue-Removal Sensitivity TABLE V: Prediction switch rate (SR, %) after cue removal. Higher values indicate greater prediction instability. Removed cue Zero-shot LoRA Full-FT Vegetation 7.12 7.92 7.63 Sky 4.40 5.65 4.70 Text 0.88 1.83 0.15 Vehicles 0.66 1.91 1.10 Vegetation and sky removal change more predictions than text or vehicle removal, with the ordering preserved after adaptation (Table V). Environmental cues may correlate with climate, coastal proximity, urban density, or streetscape design and are not inherently spurious. However, removed area and detector or segmentation quality prevent causal ranking across cues, particularly because vegetation and sky masks may cover larger image fractions. VI-D Performance on Appearance-Reduced Structure Proxies Figure 3(c) provides a compact comparison across the main adaptation regimes. Among the frozen readouts, LP-T is shown because it attains the higher original-image accuracy, while LoRA and Full-FT are shown as the two strongest encoder-adapted methods. Together with zero-shot CLIP, these methods summarize the principal performance regimes without overcrowding the combined accuracy and retention visualization. LP-C and PU remain included in the complete original-image and patch-scrambling comparisons (Tables I and VII). TABLE VI: Accuracy and relative retention on appearance-reduced structure proxies. Input Model Accuracy (%) Retention Edges Zero-shot 21.5 0.55 LoRA 32.1 0.41 Full-FT 38.3 0.47 Macro blur Zero-shot 22.1 0.57 LoRA 42.0 0.54 Full-FT 45.2 0.55 LoRA and Full-FT raise edge accuracy from 21.5% to 32.1% and 38.3%, respectively, and macro-blur accuracy from 22.1% to 42.0% and 45.2%, but retention does not improve (Table VI). Thus, adapted models make more correct predictions from the information remaining after transformation without preserving a larger fraction of original accuracy. Absolute accuracy and retention are complementary: the former measures usable residual information, whereas the latter normalizes by each modelās original accuracy. Higher transformed accuracy alone therefore does not establish structural sufficiency. VI-E Sensitivity to Patch Scrambling TABLE VII: Prediction switch rate (SR, %) under patch-based layout disruption. Model Patch-16 Patch-32 Patch-64 Zero-shot 12.40 8.73 5.28 LP-T 14.60 9.24 6.46 LP-C 10.79 7.92 4.92 PU 45.56 41.31 36.10 LoRA 43.21 30.45 22.74 Full-FT 42.92 28.69 17.24 Under Patch-16, frozen models switch 10.79ā14.60% of predictions, compared with 42.92ā45.56% for adapted models (Table VII). The gap persists at larger patches, while sensitivity decreases as more coherent objects and local scene content are preserved. The within-task contrast is consistent with increased layout dependence after encoder adaptation. Figure 4 shows illustrative intervention and viewpoint cases. Fig. 4: Qualitative cases from Downtown Los Angeles, Hollywood, Santa Monica, Koreatown, and Beverly Hills. Columns show training/test viewing directions, four cue masks, Canny edges, and Patch-32 scrambling with zero-shot, LoRA, and Full-FT predictions. Rows 1ā4 share coordinates across viewing directions; the Beverly Hills pair is 0.68 km apart. VI-F Non-Geographic Control on Caltech101 We apply the same scrambling implementation to 2,000 seed-42 Caltech101 samples resized to 224Ć224224Ć 224 [16]. Sample indices are recorded for reproducibility. The zero-shot pipeline uses 102 labels, including the loaderās background label, and no Caltech101 training data. For the fixed model m, relative drop is Dropā”(m,T)=Accā”(m)āAccā”(m,T)Accā”(m)Ć100.Drop(m,T)= Acc(m)-Acc(m;T)Acc(m)Ć 100. (6) TABLE VIII: Caltech101 zero-shot accuracy under the same patch-scrambling protocol. Condition Accuracy (%) Relative drop (%) Clean 89.65 0.00 Patch-16 33.90 62.19 Patch-32 61.40 31.51 Patch-64 78.60 12.33 Caltech101 accuracy drops by 62.19%, 31.51%, and 12.33% under Patch-16, -32, and -64 (Table VIII). Scrambling therefore disrupts generic recognition and cannot by itself establish geolocalization-specific sensitivity. These accuracy drops are not directly comparable to geolocalization SR, which measures any label change. The within-geolocalization frozenāadapted contrast remains the relevant evidence for increased layout dependence. VI-G Prompt Dependence TABLE IX: Zero-shot accuracy under prompt controls. Chance accuracy is 12.5%. Prompt condition Accuracy (%) Length-controlled 38.30 Original baseline 39.03 Visual-only descriptors 36.24 Swapped descriptors 11.15 Original, length-controlled, and visual-only prompts perform similarly, whereas descriptor swapping reduces accuracy to 11.15% (Table IX). Class-specific wording therefore matters beyond prompt length or template form. However, swapping changes both descriptor content and label association and does not establish that each descriptor is grounded in corresponding image evidence. This control concerns zero-shot CLIP, not the adapted image encoder. Because no single intervention isolates structure or appearance, we interpret the results jointly. VII Discussion VII-A Effects of Encoder Adaptation Frozen readouts remain near zero-shot accuracy, whereas PU, LoRA, and Full-FT reach 75.94ā82.10% (Table I). This gap indicates that strong nearby-region discrimination is more accessible when the visual representation adapts; weak frozen readouts do not establish that frozen CLIP lacks regional information. Gains remain uneven across regions (Table IV), ranging from 6.67 points for Venice Beach to 60.56 points for Hollywood. VII-B Structural Sensitivity Without Structural Sufficiency Structural sensitivity and sufficiency yield different conclusions. Full-FT raises edge accuracy from 21.5% to 38.3% and blur accuracy from 22.1% to 45.2%, yet its retention is 0.47 and 0.55 versus 0.55 and 0.57 for zero-shot CLIP (Table VI); LoRA shows similarly lower retention. Adapted models therefore use more residual information without retaining a larger share of original accuracy. Conversely, Patch-16 changes 42.92ā45.56% of adapted predictions versus 10.79ā14.60% for frozen methods (Table VII). Together, these results support greater dependence on intact configuration without structural sufficiency; regional predictions still combine configuration and appearance information. VII-C Interpreting Layout Sensitivity Under General Image Distortion Caltech101 accuracy falls from 89.65% to 33.90% under Patch-16 (Table VIII), so scrambling is not geolocalization-specific. Its accuracy drops are not equivalent to geolocalization SR; the relevant evidence is the higher adapted-model SR under the same geolocalization intervention. This supports increased layout dependence, not uniquely geographic sensitivity, semantic spatial reasoning, or a general ability to represent relations. VII-D Reliability of Geographic Evidence Removing vegetation and sky changes more predictions than removing text or vehicles (Table V), and the ordering persists after adaptation. Greater layout sensitivity therefore coexists with, rather than replaces, dependence on environmental appearance. Prompt swapping also reduces zero-shot accuracy to 11.15% (Table IX), indicating class-specific textual dependence. Mask-area and segmentation differences preclude causal ranking of image cues, so the interventions must be interpreted jointly rather than as selective causal tests. This mixed interpretation is bounded by the dataset and intervention design discussed next. VIII Limitations The study covers one metropolitan area and eight mixed municipality/neighborhood region types, so it does not establish behavior in other cities, scales, or image sources. Its overlapping split evaluates viewpoint variation near known locations, not geographic generalization. Edge and blur are imperfect structure proxies, scrambling introduces artifacts, and cue removal depends on mask area and segmentation quality. Model-performance comparisons use a single seed and are not accompanied by significance tests or confidence intervals. The study also lacks human comparison. Street View images cannot be redistributed [7, 8]. Future work should use coordinate-grouped or buffered splits, matched and audited masks, alternative fills, multiple scrambling permutations, aligned accuracy-drop and SR metrics, multiple seeds, stronger frozen-feature probes, and cross-source, temporal, seasonal, and geographically disjoint evaluation. IX Conclusion Within these constraints, the experiments provide a consistent account of CLIP adaptation for regional geolocalization. Frozen readouts remain near the 39.03% zero-shot baseline, whereas encoder adaptation reaches 75.94ā82.10%; full fine-tuning also reduces centroid error from 12.30 km to 3.86 km. Under the evaluated setting, strong nearby-region discrimination therefore emerges primarily when the pretrained visual representation is allowed to adapt. Adapted models also achieve higher edge and blur accuracy and switch more often under layout disruption, but do not retain a larger fraction of original accuracy after appearance reduction. Vegetation and sky remain influential, and Caltech101 shows that scrambling sensitivity is not uniquely geolocalization-specific. Together, the results support greater dependence on intact scene configuration without structural sufficiency; predictions remain consistent with a combination of configuration and appearance cues. Within this scope, regional geolocalization provides a controlled setting for examining how adaptation changes visual evidence, but the findings concern viewpoint variation near known locations and require confirmation with geographically separated evaluation. Acknowledgment OpenAI GPT-5.6 [18] and Google Gemini 3.1 Pro [6] assisted with English translation, language and style editing, code, and figures. The human authors reviewed and approved all AI-assisted content. References [1] A. Alfarrarjeh, S. H. Kim, S. Rajan, A. Deshmukh, and C. Shahabi (2018) A data-centric approach for image scene localization. In Proc. IEEE Int. Conf. Big Data (Big Data), p. 594ā603. External Links: Document Cited by: §I-A. [2] G. Astruc, N. Dufour, I. Siglidis, C. Aronssohn, N. Bouia, S. Fu, R. Loiseau, V. N. Nguyen, C. Raude, E. Vincent, L. Xu, H. Zhou, and L. Landrieu (2024) OpenStreetView-5M: the many roads to global visual geolocation. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 21967ā21977. Cited by: §I-A, §I-A, §I-C2. [3] Y. Baek, B. Lee, D. Han, S. Yun, and H. Lee (2019) Character region awareness for text detection. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 9365ā9374. Cited by: §V-A. [4] J. Canny (1986) A computational approach to edge detection. IEEE Trans. Pattern Anal. Mach. Intell. 8 (6), p. 679ā698. Cited by: §V-B. [5] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) An image is worth 16x16 words: transformers for image recognition at scale. In Proc. Int. Conf. Learn. Representations (ICLR), Cited by: §IV-A. [6] Google DeepMind (2026) Gemini 3.1 pro model card. Note: https://deepmind.google/models/model-cards/gemini-3-1-pro/Published: Feb. 19, 2026; accessed: Aug. 21, 2026 Cited by: Acknowledgment. [7] Google LLC (2026) Google geo guidelines: products and services. Note: https://about.google/brand-resource-center/products-and-services/geo-guidelinesAccessed: Aug. 14, 2026 Cited by: §VIII. [8] Google LLC (2026) Google maps platform terms of service. Note: https://cloud.google.com/maps-platform/termsAccessed: Aug. 14, 2026 Cited by: §VIII. [9] L. Haas, S. Alberti, and M. Skreta (2023) Learning generalized zero-shot learners for open-domain image geolocalization. arXiv preprint arXiv:2302.00275. Cited by: §I, §I-B. [10] L. Haas, M. Skreta, S. Alberti, and C. Finn (2024) PIGEON: predicting image geolocations. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 12893ā12902. Cited by: §I, §I-B. [11] J. Hays and A. A. Efros (2008) IM2GPS: estimating geographic information from a single image. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 1ā8. Cited by: §I. [12] D. Hendrycks and T. Dietterich (2019) Benchmarking neural network robustness to common corruptions and perturbations. In Proc. Int. Conf. Learn. Representations (ICLR), Cited by: §V-A. [13] S. Hooker, D. Erhan, P. Kindermans, and B. Kim (2019) A benchmark for interpretability methods in deep neural networks. In Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 32, p. 9737ā9748. Cited by: §V-A. [14] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen (2022) LoRA: low-rank adaptation of large language models. In Proc. Int. Conf. Learn. Representations (ICLR), Cited by: §I, §IV-B. [15] G. Jocher, A. Chaurasia, and J. Qiu (2023) Ultralytics YOLOv8. Note: Software, version 8.0.0Available: https://github.com/ultralytics/ultralytics Cited by: §V-A. [16] F. Li, R. Fergus, and P. Perona (2004) Learning generative visual models from few training examples: an incremental bayesian approach tested on 101 object categories. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. Workshop Generative-Model Based Vis., Cited by: §VI-F. [17] T. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. DollĆ”r, and C. L. Zitnick (2014) Microsoft COCO: common objects in context. In Proc. Eur. Conf. Comput. Vis. (ECCV), p. 740ā755. Cited by: §V-A. [18] OpenAI (2026) GPT-5.6 system card. Note: OpenAI Deployment Safety Hub, https://deploymentsafety.openai.com/gpt-5-6Published: July 9, 2026; accessed: Aug. 21, 2026 Cited by: Acknowledgment. [19] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever (2021) Learning transferable visual models from natural language supervision. In Proc. Int. Conf. Mach. Learn. (ICML), p. 8748ā8763. Cited by: §I, §I-B. [20] B. Shi, X. Bai, and C. Yao (2017) An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Trans. Pattern Anal. Mach. Intell. 39 (11), p. 2298ā2304. Cited by: §V-A. [21] A. Telea (2004) An image inpainting technique based on the fast marching method. J. Graphics Tools 9 (1), p. 23ā34. Cited by: §V-A. [22] V. Vivanco Cepeda, G. K. Nayak, and M. Shah (2023) GeoCLIP: CLIP-inspired alignment between locations and images for effective worldwide geo-localization. In Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 36, p. 8690ā8701. Cited by: §I, §I-B. [23] N. Vo, N. Jacobs, and J. Hays (2017) Revisiting IM2GPS in the deep learning era. In Proc. IEEE Int. Conf. Comput. Vis. (ICCV), p. 2621ā2630. Cited by: §I, §I-A. [24] Z. Wang, S. Zhou, S. He, H. Huang, L. Yang, Z. Zhang, X. Cheng, S. Ji, T. Jin, H. Zhao, and Z. Zhao (2025) SpatialCLIP: learning 3d-aware image representations from spatially discriminative language. In Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 29656ā29666. Cited by: §I-C. [25] T. Weyand, I. Kostrikov, and J. Philbin (2016) PlaNet: photo geolocation with convolutional neural networks. In Proc. Eur. Conf. Comput. Vis. (ECCV), p. 37ā55. Cited by: §I, §I-A. [26] E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo (2021) SegFormer: simple and efficient design for semantic segmentation with transformers. In Adv. Neural Inf. Process. Syst. (NeurIPS), Vol. 34, p. 12077ā12090. Cited by: §V-A. [27] S. Xu, C. Zhang, L. Fan, G. Meng, S. Xiang, and J. Ye (2024) AddressCLIP: empowering vision-language models for city-wide image address localization. In Proc. Eur. Conf. Comput. Vis. (ECCV), p. 76ā92. Cited by: §I, §I-B. [28] M. Yuksekgonul, F. Bianchi, P. Kalluri, D. Jurafsky, and J. Zou (2023) When and why vision-language models behave like bags-of-words, and what to do about it?. In Proc. Int. Conf. Learn. Representations (ICLR), Cited by: §I-C. [29] M. D. Zeiler and R. Fergus (2014) Visualizing and understanding convolutional networks. In Proc. Eur. Conf. Comput. Vis. (ECCV), p. 818ā833. Cited by: §V-A. [30] B. Zhou, H. Zhao, X. Puig, S. Fidler, A. Barriuso, and A. Torralba (2017) Scene parsing through ADE20K dataset. In Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), p. 633ā641. Cited by: §V-A.