Paper deep dive
LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning
Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/16/2026, 2:20:17 AM
Summary
The paper introduces LongEarth-Bench, a large-scale benchmark for long-horizon Earth observation reasoning, containing ~120k QA samples from 117k unique images with an average sequence length of 15.14 frames. It defines four cognitive dimensions: evolution summarization, spatial reasoning, anomaly identification, and logical prediction. The authors develop LongEarth (SFT-based) and LongEarth-R1 (GRPO-based) models, with LongEarth-R1 achieving state-of-the-art results on the benchmark by utilizing format, temporal, and spatial rewards.
Entities (12)
Relation Signals (10)
LongEarth-R1 → achievesbestresultson → LongEarth-Bench
confidence 95% · LongEarth-R1 achieves the best results on all 12 long-sequence tasks...
LongEarth-R1 → buildson → LongEarth
confidence 95% · Building on LongEarth, LongEarth-R1 applies...
LongEarth-R1 → usesmethod → GRPO
confidence 95% · LongEarth-R1 applies group relative policy optimization...
LongEarth → usesmethod → SFT
confidence 95% · We develop LongEarth through supervised fine-tuning...
LongEarth-Bench → containssamplesfrom → SpaceNet 7
confidence 90% · LongEarth-Bench is derived from SpaceNet 7 (SN7)...
LongEarth-Bench → containssamplesfrom → DynamicEarthNet
confidence 90% · LongEarth-Bench is derived from ... DynamicEarthNet (DynEarth)...
LongEarth-Bench → coverstask → Evolution Summarization
confidence 90% · covering 12 tasks across evolution summarization...
LongEarth-Bench → coverstask → Spatial Reasoning
confidence 90% · covering 12 tasks across ... spatial reasoning...
LongEarth-Bench → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sensing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth-Bench, a benchmark containing approximately 120k question-answering samples derived from 117k unique images. Its sequences average 15.14 frames and extend to 30 frames, covering 12 tasks across evolution summarization, spatial reasoning, anomaly identification, and logical prediction. A 30k-sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks.
Tags
Links
- Source: https://arxiv.org/abs/2608.13344v1
- Canonical: https://arxiv.org/abs/2608.13344v1
Trouble viewing inline? Open PDF directly →
Full Text
47,624 characters extracted from source content.
Expand or collapse full text
LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning Yupan Ding 1 , Jing Xiao 1 , Zhenyuan Zhang 2 , Chaofeng Chen 1 , Liang Liao 3 , Gui-Song Xia 1 , Mi Wang 4 1 School of Artificial Intelligence, Wuhan University, Wuhan, China 2 School of Computer Science, Wuhan University, Wuhan, China 3 Xi’an University of Electronic Science and Technology, Xi’an, China 4 State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University, Wuhan, China Abstract Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences. However, existing remote sens- ing vision-language models mainly focus on isolated images, image pairs, or short sequences, limiting reliable grounding in the relevant frames and regions. We introduce LongEarth- Bench, a benchmark containing approximately 120k question- answering samples derived from 117k unique images. Its se- quences average 15.14 frames and extend to 30 frames, cov- ering 12 tasks across evolution summarization, spatial rea- soning, anomaly identification, and logical prediction. A 30k- sample subset further provides structured reasoning traces linking key frames and changed regions to final answers. We develop LongEarth through supervised fine-tuning with explicit sequence identifiers and structured chain-of-thought supervision. Building on LongEarth, LongEarth-R1 applies group relative policy optimization with format, temporal, and spatial rewards. LongEarth-R1 achieves the best results on all 12 long-sequence tasks while remaining competitive on standard remote sensing benchmarks. Introduction Remote sensing analysis is increasingly moving beyond iso- lated observations toward understanding how geographic re- gions evolve over time (Weng, Pang, and Xia 2025). With the growing availability of high-revisit Earth observation imagery, it is now possible to observe multi-stage processes such as urban construction, disaster evolution, and ecosystem recovery (Liu et al. 2026a; Irvin et al. 2025; Jain et al. 2026). For example, given observations before, during, and after a flood, a model should identify its onset and affected regions, track its evolution, and infer the recovery stage. This requires more than pairwise change detection: the model must re- cover a temporal trajectory, localize supporting frames and regions, and separate geographic evolution from seasonal and acquisition-induced variations (Li et al. 2024; Soni et al. 2025; Luo et al. 2026). We refer to this capability as long- horizon Earth observation reasoning: reasoning about multi- stage geographic evolution from extended Earth observation sequences. Recent remote sensing vision-language models (RSVLMs) have progressively extended language-guided interpretation from individual observations to image pairs and temporal sequences. Single-temporal models connect an individual observation with language through image description, visual question answering, and visual grounding (Zhang et al. 2024a; Hu et al. 2025). Bi-temporal models further compare paired observations to describe and localize geographic changes (Yang et al. 2025; Liu et al. 2024a). More recent temporal RSVLMs accept multiple observations and support sequence-level dialogue or change understanding (Irvin et al. 2025; Soni et al. 2025; Xuan et al. 2025). Nevertheless, these capabilities are still insufficient for long-horizon reasoning. A single image contains no evolution trajectory, an image pair exposes mainly the endpoints of change, and current sequence models generally emphasize recognition or descriptive question answering rather than reconstructing a complete multi-stage process. They consequently remain limited in tracing critical transi- tions, grounding conclusions across frames and regions, and inferring how an evolving process may continue. To address these limitations, we formulate long-horizon Earth observation reasoning as reasoning over extended se- quences. This capability requires models to identify how geographic entities evolve, localize and characterize their changes, assess whether the observed evolution is temporally consistent, and infer unobserved states. We therefore organize it into four progressive cognitive dimensions: 1) Evolution summarization organizes major stages into coherent temporal trajectories; 2) Spatial reasoning localizes changes and mod- els their evolving spatial relations; 3) Anomaly identification detects chronological violations, repetitions, and contextual interference; and 4) Logical prediction infers subsequent or missing states from accumulated evidence. Figure 1(a) illus- trates this hierarchy through 12 fine-grained tasks. Based on the four cognitive dimensions, we construct LongEarth-Bench, a large-scale benchmark for long- horizon remote sensing spatiotemporal reasoning. It con- tains approximately 120k question-answering (QA) samples derived from 117k images, with an average sequence length of 15.14 frames. The benchmark covers diverse geographic regions, land-cover categories, temporal spans, change rates, and multi-stage processes. Its QA samples are generated from segmentation annotations, geometric relations, and temporal change trajectories. Moreover, 30k samples contain struc- arXiv:2608.13344v1 [cs.AI] 13 Aug 2026 (b) Comparison of Remote Sensing Datasets DynamicVL LongEarth-R1 CDVQA GeoLLaVA C-Expert EarthDial TEOChat VLRS GeoReason DisasterM3 Description √ √ √ √ √ √ √ √ √ √ Spatial Loc. - - - - - - - √ - - Temp. Diag. √ √ √ √ √ √ √ √ √ √ Log. Pred. - - - - - - √ √ - - CoT - - - - - - - √ √ - (c) Capability Coverage of Representative RS Methods(a) Long-Sequence Task Taxonomy and Examples t 2 Q: What land-cover change occurs in this sequence? A: Soil changes into impervious surface. · t 1 t 23 t 24 t 2 t 2 t 1 t 22 t 23 · Q: What is the building growth pattern? A: Steady / Burst / Early-only / Multi-stage / None Q: Identify the continuous range with no new buildings. A: Images 1 to 19. · T1 Temporal Phasing T2 Pattern Classification T3 Stagnation Identification t 1 t 23 t 24 Q:Which image is misplaced in time? A: Image 6. T7 Chronological Violations Q: Which four identical frames were inserted? A: Images 1 to 4. T8 Redundancy Detection Q: Which images show actual flooding? A: Images 4 and 6. t 1 t 6 t 5 t 7 · t 2 t 1 t 6 t 7 · t 2 t 1 t 6 t 7 · T9 Contexual Robustness Q: What is the dominant land-cover pattern over time? A: Water gradually shifts to soil. Q: How extensive is the flood in the final image? A: Slight flooding / Limited / Moderate / Extensive / Near-total. Q: Which image shows the most severe flood? A: Image 5. T4 Change Direction T5 Significant Region T6 Extent Assessment · t 23 t 24 t 2 t 1 · t 2 t 5 t 6 t 1 · t 2 t 1 t 5 t 6 Q: What is the flood trend next? A: Slight spread / Slight recession / Rapid spread / Rapid recession / Stable T10 Trend Prediction Q: Where is new building growth least likely next? A: Upper-left of center, as growth is concentrated lower-right. T11 Spatial Location Prediction Q: What dominant land cover appears after Image 22? A: Forest and vegetation. T12 Missing Frame Prediction · t 1 t 23 t 24 t 2 t 2 t 1 t 15 · ? t 1 t 22 t 24 ? t 23 · Evolution Summarization (EvoSum) Anomaly Identification (AnomID)Spatial Reasoning (Spatial) Logical Prediction (LogPred) 2.3× Long Seq. √ - - - √ √ √ √ - - Figure 1: Task taxonomy, benchmark positioning, and model capability comparison of LongEarth-Bench and LongEarth-R1. (a) Represen- tative examples of the 12 tasks grouped into four cognitive dimensions. (b) Average sequence length versus cognitive-dimension coverage; bubble size denotes maximum sequence length. (c) Capability coverage of representative remote sensing methods, including long-sequence understanding, evolution description, spatial localization, temporal diagnosis, logical prediction, and CoT. tured reasoning annotations connecting frame selection, tem- poral localization, spatial evidence, and final conclusions. Figure 1(b) compares LongEarth-Bench with existing bench- marks according to average sequence length and coverage of the four cognitive dimensions. To enable the RSVLMs with long-horizon spatiotempo- ral reasoning, we first develop LongEarth through super- vised fine-tuning with explicit sequence identifiers and struc- tured chain-of-thought (CoT) traces that connect key frames, temporal changes, and changed regions to final answers. The former establishes stable frame-level temporal anchors, while the latter teaches the model to select relevant observa- tions and integrate temporal and spatial evidence across the sequence. Building on LongEarth, LongEarth-R1 applies group relative policy optimization (GRPO) with comple- mentary format, temporal, and spatial rewards to optimize output completeness, evolution-stage localization, temporal ordering, key-frame selection, and changed-region consis- tency beyond final-answer correctness. Together, LongEarth and LongEarth-R1 establish reproducible supervised and reinforcement-learning baselines on LongEarth-Bench. Fig- ure 1(c) compares LongEarth-R1 with representative re- mote sensing methods across six model capabilities, with LongEarth-R1 covering all six. The main contributions of this paper are as follows. • We formulate long-horizon Earth observation reason- ing as reasoning over multi-stage geographic evolution through four cognitive dimensions covering evolution summarization, spatial reasoning, anomaly identification, and logical prediction. • We introduce LongEarth-Bench, a large-scale benchmark with approximately 120k samples. It covers diverse geo- graphic processes and includes 30k samples with struc- tured annotations for evidence-grounded reasoning. • We develop LongEarth through supervised fine-tuning and LongEarth-R1 through GRPO-based temporal and spatial rewards, and systematically evaluate current RSVLMs across sequence lengths, evolution processes, and reasoning dimensions. Related Work Single- to Multi-Temporal RSVLMs RSVLMs have progressed from static image-language alignment to temporal Earth observation understanding. Early models such as RemoteCLIP (Liu et al. 2024b), GeoRSCLIP (Zhang et al. 2024c), GeoChat (Kuckreja et al. 2024), and EarthGPT (Zhang et al. 2024a) focus on align- ing a single observation with language. VHM (Pang et al. 2025) extends this capability to scene classification, vi- sual question answering, and visual grounding, while Earth- VQA (Wang et al. 2024a) advances relational visual question answering and spatial reasoning. Beyond remote sensing, SpaceVLLM (Wang et al. 2026) studies frame-specific spa- tiotemporal video grounding. RSICCformer (Liu et al. 2022), GeoLLaVA (Elgendy et al. 2024), Change-Agent (Liu et al. 2024a), and CCExpert (Wang et al. 2024b) extend language- based analysis to bi-temporal image pairs through change captioning, detection, or difference-aware integration. Dis- asterM3 (Wang et al. 2025) provides a bi-temporal remote sensing vision-language dataset and benchmark for disaster assessment and response. However, existing bi-temporal set- tings cannot explicitly represent the intermediate stages of an evolving geographic process. Recent multi-temporal RSVLMs process extended Earth observation sequences. EarthDial (Soni et al. 2025) supports multispectral, multi-temporal, and multi-resolution conver- sational inputs, TEOChat (Irvin et al. 2025) enables dia- logue and question answering over temporal observations, (c) Sequence-length distribution by source dataset(d) Overall sequence-length distribution and coverage(e) Subtask-specific sequence-length distribution (a) Sample composition by dimension and subtask(b) Global geographic coverage of source datasets Figure 2: LongEarth-Bench composition, geographic coverage, and sequence-length statistics. (a) Sample distribution across four cognitive dimensions (inner ring) and 12 tasks (outer ring). (b) Geographic coverage; colors denote the five source datasets and marker size the local mean sequence length. (c) Source-specific length distributions with annotated means. (d) Overall length distribution; bars show counts and the curve gives the fraction of sequences at least each length. (e) Task-conditioned length, with rows normalized within each subtask. using sequences that average 2.07 frames and extend to at most 8 frames, and UniRS (Li et al. 2024) unifies single- image, bi-temporal, and video tasks. TimeSenCLIP (Jain et al. 2026) learns spectral-temporal representations from Sentinel-2 time series, while DynamicVL (Xuan et al. 2025) benchmarks dynamic city understanding with multi- temporal scenes averaging 6.73 frames and extending to at most 10 frames, while VLRS-Bench (Luo et al. 2026) evaluates cognition, decision, and prediction with sequences averaging 1.59 frames and covering up to eight temporal phases. These studies establish important foundations for multi-temporal remote sensing understanding. However, ex- tended trajectories spanning many evolution stages remain less studied, particularly when evaluation requires anomaly localization, cross-stage spatial reasoning, and identification of the observations that support a conclusion. Reinforcement Learning for Reasoning in RSVLMs Recent RSVLMs have incorporated structured reasoning and reinforcement learning to support more interpretable and evidence-grounded geospatial analysis. Multimodal- CoT separates rationale generation from answer infer- ence (Zhang et al. 2024b), while Geo-CoT constructs per- ceptually grounded reasoning traces and trains RSThinker through supervised fine-tuning followed by GRPO (Liu et al. 2026b). Beyond remote sensing, IAD-R1 (Li et al. 2026b) similarly combines CoT-based supervised fine-tuning with structured GRPO for consistent vision-language reasoning. SAMChat (Köksal and Alatan 2026) adopts CoT supervision and GRPO for small-scale remote sensing analysis, while RemoteReasoner (Yao et al. 2026) applies reinforcement learning to a unified workflow covering object-, region-, and pixel-level geospatial reasoning. GeoReason (Li et al. 2026a) aligns reasoning and answers through logical-consistency reinforcement learning, while GeoChain (Yerramilli et al. 2025) studies multimodal CoT geographic reasoning. In par- allel, VLRS-Bench evaluates complex reasoning over single- temporal and multi-temporal remote sensing inputs (Luo et al. 2026). These studies show the value of intermediate supervision and task-specific rewards for reasoning beyond final-answer prediction. However, existing methods primarily reason over individual observations or restricted temporal settings, with- out explicitly aligning intermediate reasoning with evidence distributed across long Earth observation sequences. The LongEarth-Bench Dataset LongEarth-Bench defines this hierarchy through 12 fine- grained tasks. We abbreviate the four cognitive dimensions as evolution summarization (EvoSum), spatial reasoning (Spa- tial), anomaly identification (AnomID), and logical predic- tion (LogPred). Figure 2(a) reports the sample distribution. The 120,367 samples are broadly balanced across EvoSum (26.6%), Spatial (22.8%), AnomID (26.7%), and LogPred (23.9%), while preserving diversity across the 12 tasks. Dataset Construction Figure 3 presents the construction pipeline of LongEarth- Bench: multi-source data integration, sequence filtering and cognitive task construction, followed by quality-controlled structured reasoning annotation. Specifically, LongEarth- Bench is derived from SpaceNet 7 (SN7) (Van Etten et al. 2021), SDSU MidWest Flood (SDSU) (Jang et al. 2024), DynamicEarthNet (DynEarth) (Toker et al. 2022), FLAIR#2 (FLAIR) (Garioud et al. 2023), and PASTIS-R (PASTIS) (Sainte Fare Garnot, Landrieu, and Chehata 2022). These sources cover urban construction, flood dynamics, nat- ural land-cover evolution, cloud–snow interference, and crop ●filter corruption ●weak variation ●incomplete coverage AnomIDEvoSum LogPredSpatial ●12 tasks ●derive stage, trend, direction, location, extent ●generate QA Q: Identify the continuous image range with no new buildings. A:Image 1 to Image 19 A. Sample Filtering B. Cognitive Dimensions C. Example QA ●coordinates ●resolution ●temporal metadata ●annotation format SDSU DynEarth FLAIR PASTIS SN7 · · · · · urban construction, flood , land-cover , cloud–snow interference, crop phenology 2 Sequence & Task Construction 1 Multi-source RS Data 3 Quality & Reasoning ●rule-based checks ●human review ●Qwen3-Thinking ●auto verification ●human refinement A. Quality Assurance Pipeline B. Structed Reasoning Subset <think> 1. Visual Scanning 2. Feature Recognition 3. Comprehensive Analysis </think> <answer>Image 1 to 19</answer> · Figure 3: Construction pipeline of LongEarth-Bench. phenology. As shown in Figure 2(b), they span North Amer- ica, South America, Europe, Africa, Asia, and Oceania, re- ducing dependence on a single geographic region or evolu- tion process. Coordinate systems are normalized, while spa- tial resolutions, temporal metadata, and annotation formats are standardized to construct temporally ordered and spatially aligned remote sensing image sequences. Samples with se- vere observation corruption, incomplete temporal coverage, registration failures, or insufficient long-term variation are filtered. Task annotations are generated from temporal trajecto- ries together with segmentation, polygon, land-cover, and image-level flood evidence. Stage boundaries, trends, loca- tions, directions, extents, observation quality, and phenolog- ical states are derived from cross-frame changes. Controlled order perturbations, repeated-frame insertions, and contex- tual disturbances are used to construct anomaly tasks, while prediction tasks are formed from historical trajectories and adjacent temporal states. Rule-based validation verifies se- quence integrity, answer consistency, frame indices, spatial labels, and image paths, followed by human review for visual support and ambiguity. To complement answer-level supervision, we construct a balanced subset of 30k structured reasoning samples across the 12 tasks. Initial traces are generated by Qwen3-VL-8B- Thinking (Bai et al. 2025a) in a unified format: the <think> field contains visual scanning, feature identification, and in- tegrated analysis, and the <answer> field retains the ref- erence answer. We automatically verify tag completeness, stage ordering, non-empty reasoning fields, and answer con- sistency; invalid outputs are regenerated. Human experts then assess visual grounding, chain coherence, and potential an- swer leakage. The resulting subset provides reliable process- level supervision for learning temporal stages, spatial dy- namics, anomalies, and logical prediction. Statistical Analysis of LongEarth-Bench Figure 2(c) represents complementary temporal regimes across the five data sources. SDSU provides the shortest sequences, with a mean length of 6.24 frames, whereas Dyn- Earth provides the longest, averaging 21.31 frames. SN7, FLAIR, and PASTIS cover intermediate and long contexts, with mean lengths of 14.15, 15.84, and 18.56 frames, re- spectively. These source-specific distributions preserve the temporal characteristics of disaster events, urban construc- tion, natural-surface evolution, and crop phenology. As shown in Figure 2(d), LongEarth-Bench contains 120,367 samples in total with an average sequence length of 15.14 frames, a median of 16 frames, and a maximum of 30 frames. Compared with TEOChatlas (Irvin et al. 2025), DVL-Instruct (Xuan et al. 2025), and VLRS-Bench (Luo et al. 2026), LongEarth-Bench is 7.3×, 2.2×, and 9.5× longer on average, and supports maximum sequences that are 3.75×, 3.0×, and 3.75× longer, respectively. More- over, 21.1% of the samples contain at least 24 observations. LongEarth-Bench therefore covers a broad spectrum from compact event sequences to extended multi-stage trajecto- ries, providing a controlled setting for evaluating sequence- grounded long-horizon reasoning. Figure 2(e) summarizes the sequence-length distribution for each task of LongEarth-Bench and shows that all 12 tasks cover multiple sequence-length intervals, while their domi- nant temporal horizons differ according to their evidence re- quirements. Change-direction reasoning and spatial-location prediction rely on extended observations, with 53% and 57% of their samples falling within 22–25 frames, respectively. In contrast, contextual robustness and extent assessment con- tain more compact event-centered sequences. The substan- tial overlap across tasks also prevents sequence length from serving as a simple task shortcut. Method This section presents a two-stage framework for long-horizon spatiotemporal reasoning. As shown in Figure 4, the frame- work is built on Qwen2.5-VL-7B (Bai et al. 2025b). Stage 1 applies supervised fine-tuning (SFT) with two complemen- tary forms of supervision. Sequence-aware answer super- vision uses explicit sequence identifiers to establish stable frame-level temporal anchors, while structured CoT super- vision teaches the model to select relevant observations and integrate temporal and spatial evidence across the sequence, yielding LongEarth. Stage 2 initializes from LongEarth and applies GRPO with complementary format, temporal, and spatial rewards to improve reasoning structure and spatiotem- poral consistency, yielding LongEarth-R1. Supervised Sequence Grounding and Reasoning The first stage contains two supervised components. The first component performs sequence-aware answer supervi- sion, which adapts the base model to long-term remote sens- ing inputs and establishes frame-level temporal anchors. The second component injects structured reasoning traces, which further guides the model to organize cross-frame evidence before producing the final answer. Sequence-Aware Answer Supervision. This component adapts the base model to long-term remote sensing data and establishes the basic mapping from image sequences and questions to answers. Given a remote sensing sequence X = I 1 ,...,I T with T temporal observations and a question q, the model is required to generate an answer a based on the full sequence. Since long-term tasks require both single- frame recognition and cross-frame change understanding, we Structured CoT Supervision B <think> 1. Visual Scanning 2. Feature Recognition 3. Comprehensive Analysis </think> <answer>...</answer> Inject explicit intermediate steps Reasoning supervision LongEarth-Bench 30k structured reasoning subset (30k) Sequence-Aware Answer Supervision A Qwen2.5-VL-7B (Base Model) Ordered image sequence (Length: t) Order-aware tokens (explicit time anchors 퐀) Frozen Visual Encoder (ViT) Trainable Language Module (LoRA) Ordered Multimodal input (LongEarth) v 1 퐀 퐀 1 v 2 퐀 퐀 2 v t 퐀 퐀 퐀 퐀 퐀 · 퐀 1 [IMG 1 ] 퐀 2 [IMG 2 ] [IMG t ] 퐀 t · Image 1Image 2Image t · · Stage 1: Supervised Sequence Grounding and Reasoning Spatiotemporal Reward-Guided Policy Refinement B 퐀=퐀 퐀 퐀 퐀 +퐀 퐀 퐀 퐀 +퐀 퐀 퐀 퐀 Task-irrelevant rewards are masked or re-normalized when needed (1) Temporal reward 퐀 퐀 Pred GT 퐀 퐀 퐀 퐀 퐀 퐀㰀 퐀 퐀 퐀 㰀 퐀 㰀 Pred. order GT order 퐀 퐀 퐀 퐀 퐀 㰀 퐀 퐀 퐀 퐀㰀 퐀 㰀 (2) Spatial reward 퐀 퐀 Region match GT Pred. 퐀퐀= 퐀ఀ퐀簀.∪퐀 Pred. GT Direction consistency 퐀㠀퐀氀퐀 퐀,퐀 =0 Grouped rewards 퐀 1 퐀 2 퐀 퐀 · Relative advantage 퐀 2 퐀 1 퐀 퐀 · + 0 - 퐀 퐀 =퐀 퐀 −퐀㠀 퐀 Policy update update 퐀 to prefer higher-reward 퐀 Group-Based Response Sampling A Long-sequence sample Input 퐀 : Identify the time periods of high water period in the sequence. Image 1Image 2Image t-1Image t · Format reward 퐀 퐀 Group sampling: generate G candidate responses · <think> ... </think> <answer>...</answer> 퐀 퐀 퐀 1 <think> ... <ans>...</think> Stage 2: Spatiotemporal Alignment via GRPO Figure 4: Overview of LongEarth and LongEarth-R1. Stage 1 trains LongEarth through sequence-aware supervised fine-tuning with explicit sequence identifiers (Seq. IDs) and structured reasoning traces from the 30k-sample subset. Stage 2 initializes from LongEarth and applies GRPO to obtain LongEarth-R1 with format, temporal, and spatial rewards. assign each image an explicit sequence identifier (Seq. ID), such as Image 1, Image 2, ..., Image T. Let V t = f v (I t ) denote the visual tokens of the t-th frame encoded by the visual encoder f v . The ordered multimodal input is represented as: H = [q; Image 1,V 1 ;... ; Image T,V T ].(1) This representation provides stable frame-level anchors and reduces ambiguity in key-frame reference, temporal inter- val localization, and cross-frame relation modeling. During training, the visual encoder is frozen, while low-rank adap- tation (LoRA) modules are applied to the language model layers for parameter-efficient alignment (Hu et al. 2022). For answer-only samples, this objective mainly constrains the fi- nal answer and does not explicitly supervise how the model selects cross-frame evidence or organizes intermediate rea- soning. Structured CoT Supervision. Within the supervised stage, we use the structured reasoning subset of LongEarth- Bench to supervise evidence-grounded responses. For each reasoning sample (X,q,r,a), the target response contains a reasoning trace r and final answer a, formatted as <think> r </think> <answer> a </answer>. The reasoning trace follows three steps: visual scanning, feature identification, and integrated analysis. Visual scan- ning describes the main land-cover states and spatial layouts across temporal observations. Feature identification extracts key frames, changed regions, directions, or anomalies re- lated to the question. Integrated analysis combines cross- frame evidence and produces the final conclusion. This de- sign shifts the model from direct answer learning to process- level supervision. Specifically, given the target sequence y = (y 1 ,...,y N ), the structured-reasoning objective is: L CoT =− N X n=1 logp θ (y n | y <n ,H).(2) This objective jointly supervises the reasoning trace and final answer, but reference-trace imitation alone cannot directly penalize temporal ordering errors or spatial inconsistencies. The second stage therefore applies GRPO to further optimize structural validity and spatiotemporal grounding. Spatiotemporal Alignment via GRPO Starting from LongEarth, the second stage applies GRPO to obtain LongEarth-R1 and align reasoning with long-term re- mote sensing objectives (Shao et al. 2024; Shen et al. 2025). Unlike SFT, which fits a single reference output, GRPO com- pares a group of candidate responses for the same input and updates the policy using relative rewards. It therefore im- proves answer quality together with reasoning structure and spatiotemporal consistency. For multimodal input H, the old policy π θ old samples G responses y i = (r i ,a i ), each comprising a reasoning trace and final answer. We assign each response a weighted reward: R i = λ f R fmt i + λ t R time i + λ s R space i ,(3) where the three terms measure format validity, temporal grounding, and spatial consistency, respectively. The format reward checks the required reasoning–answer structure and task-specific answer parsability. The temporal reward com- bines matching of referenced frames with their chronological consistency, while the spatial reward evaluates agreement on changed locations, directions, and extents. For tasks with- out temporal or spatial labels, inapplicable terms are omitted CategoryTaskVideo-LLaVA Qwen2.5 Qwen3-Thinking TEOChat TEOChat* EarthDial EarthDial* LongEarth-R1 Evolution Summarization T1 Temporal Phasing24.5341.2249.4432.0264.1631.8446.1877.31 T2 Pattern Classification39.3246.4459.6035.8149.5845.1969.9083.14 T3 Stagnation Identification10.5834.1129.7335.9850.7425.9246.9563.57 Spatial Reasoning T4 Change Direction32.7243.8640.0831.9240.2833.4445.8053.37 T5 Significant Region34.6540.5336.4731.6050.4532.3962.6968.96 T6 Extent Assessment30.7639.2433.4835.0358.4837.0467.5874.32 Anomaly Identification T7 Chronological Violations2.8612.387.877.9015.475.3715.9250.80 T8 Redundancy Detection5.9039.7836.6810.3127.5714.5120.3290.28 T9 Contextual Robustness18.6035.6537.4515.5148.7412.4249.1289.12 Logical Prediction T10 Trend Prediction29.7846.7832.0336.3055.4420.6160.5476.42 T11 Spatial Location Prediction47.6050.6133.7728.4855.9849.8587.6289.87 T12 Missing Frame Prediction29.6933.8347.8638.4058.9436.1962.0464.25 Table 1: Performance comparison on the 12 long-sequence cognitive tasks in LongEarth-Bench (%). An asterisk (*) denotes supervised fine-tuning on LongEarth-Bench. The best result in each row is shown in bold. and the remaining weights are renormalized; detailed reward definitions are provided in the supplementary material. GRPO normalizes rewards within the response group as A i = (R i − μ R )/(σ R + ε) and uses the policy ratio ρ i = π θ (y i | H)/π θ old (y i | H). Let ̄ρ i = clip(ρ i , 1− ε, 1 + ε). The clipped objective is J GRPO = 1 G G X i=1 min(ρ i A i , ̄ρ i A i )−βD KL (π θ ∥π ref ), (4) where ε is the clipping coefficient, β is the Kullback–Leibler (KL) weight, and π ref is the policy obtained after the super- vised stage. The KL term preserves the visual-language abil- ity acquired during SFT. Together, the rewards shift learning from answer imitation toward structured, temporally ordered, and spatially grounded reasoning. Experiments Experimental Setups Implementation Details. We train LongEarth-R1 on 8× NVIDIA A800 GPUs, freezing the visual encoder and tuning language-side LoRA modules with r = 128 and α = 256 in both stages. SFT runs for two epochs with bfloat16, gradient checkpointing, cosine decay, a peak learning rate of 2×10 −5 , a 0.03 warmup ratio, and no weight decay. Inputs are limited to 8,192 tokens and interleave square-resized images with temporal prefixes and instructions. GRPO starts from the SFT checkpoint and optimizes the same LoRA modules with group-sampled responses and the proposed rewards. Evaluation Metrics. We evaluate final answers using Ac- curacy, Temporal F1, and Spatial F1. Accuracy applies to closed-set answers, including categorical judgments, single- frame localization, fixed spatial labels, and discrete extent levels. Temporal and free-form spatial answers use set- based F1 over predicted and reference elements. Temporal F1 operates on parsed frame indices, whereas Spatial F1 usesS =object, change, region, direction, extent. Tempo- ral F1 evaluates T1, T3, T8, and temporal-set variants of T7, T9, and T12; Spatial F1 evaluates free-form T4–T6 and T11. Remaining tasks use Accuracy. Query: Identify the time periods of low water period in the sequence. Please answer in the following format: Image M to Image N. · t 1 t 2 t 20 t 21 t 22 t 23 t 24 EarthDial Answer:Image 20 to Image 24 TEOChat Answer:Image 20 to Image 24 LongEarth-R1 1. Visual Scan: The sequence shows a coastal area with a distinct landmass and a smaller, circular feature in the water. In Images 1-20, the water around the smaller feature appears relatively clear and dark, suggesting higher water levels. Starting from Image 21, the water color shifts to a lighter, more opaque greenish hue, and the smaller feature becomes more prominent and less submerged...... 2. Feature Identification: The key feature for identifying low water periods is the visibility of the smaller circular landmass. In high water periods (Images 1-20)...... In low water periods (Images 22-24), the landmass becomes clearly visible as a distinct, solid feature...... 3. Synthesis: The sequence shows a clear transition from high water (Images 1-20) to low water (Images 22-24), where the smaller circular feature becomes fully exposed and the water color changes to a lighter...... Answer:Image 22 to Image 24 Figure 5: Qualitative comparison on long-horizon temporal rea- soning. EarthDial and TEOChat predict an incorrect interval, while LongEarth-R1 identifies the correct low-water period and grounds its answer in multi-frame visual evidence. Experimental Results Performance on LongEarth-Bench. We evaluate the 12 LongEarth-Bench tasks spanning EvoSum, Spatial, Ano- mID, and LogPred. Table 1 shows that LongEarth-R1 ranks first on all tasks, with particularly strong gains on Ano- mID and tasks requiring long-range temporal evidence. The consistent improvements show that sequence grounding, structured reasoning, and reward-based alignment jointly strengthen temporal understanding, spatial grounding, and prediction. Figure 5 further shows that LongEarth-R1 avoids temporally misaligned intervals by grounding its answer in multi-frame evidence. Performance on Standard Remote Sensing Tasks. We further test whether specialization for long-term reason- ing preserves general remote sensing understanding. Fol- lowing the standard evaluation protocol (Irvin et al. 2025), we evaluate single-image scene recognition on AID and UCM (Xia et al. 2017; Yang and Newsam 2010), bi-temporal change understanding on ABCD-CD, CDVQA-QA, xBD, and S2Looking (Fujita et al. 2017; Yuan et al. 2022; Gupta et al. 2019; Shen et al. 2021), and short-sequence multi- image reasoning on Qfabric and fMoW (Verma, Panigrahi, and Gupta 2021; Christie et al. 2018). For each benchmark, Category DatasetHumanVideo-LLaVAGeoChatQwen2.5EarthDialTEOChatLongEarth-R1 Single image AID–52.4072.0065.6088.6780.9086.03 UCM–46.5084.4070.7080.4886.3090.19 Bi-temporal ABCD-CD 95.2050.00–69.8063.4685.6091.20 CDVQA-QA 63.4029.80–50.5047.4747.2056.60 xBD (Avg.)56.2021.2025.1036.8024.7859.6071.50 S2Looking (Avg.) 26.5024.8032.708.2029.7057.7060.50 Multi-image (≤ 8) Qfabric (Avg.)71.0017.6018.4023.7033.5970.8069.70 fMoW-LR-TSC–4.9026.300.1024.0845.5045.60 fMoW-HR-TSC 65.9016.6059.2028.5060.8675.1072.90 Table 2: Performance on single-image, bi-temporal, and short-sequence multi-image tasks (%). - denotes unavailable results. Component Ablation SFT Seq.IDs CoT GRPOEvoSum Spatial AnomID LogPred ✗40.5941.2129.2743.74 ✓✗60.7760.2729.2672.49 ✓✗ 70.4659.4372.9373.82 ✓✗69.4759.0864.1270.81 ✓74.6765.5576.7376.85 Reward Ablation Format Temporal SpatialEvoSum Spatial AnomID LogPred ✗✓63.7463.9768.0177.33 ✓✗✓59.8364.2267.0276.00 ✓✗ 61.5062.6273.0570.67 ✓74.6765.5576.7376.85 Table 3: Ablation results of the core components (%). The upper block studies cumulative training components, and the lower block removes individual rewards. LongEarth-R1 is fine-tuned on the corresponding training split and evaluated using the task-native protocol; detailed sources and task descriptions are provided in the supplemen- tary material. Table 2 shows that LongEarth-R1 achieves the best result on six of nine datasets, including all bi-temporal change- understanding benchmarks, indicating effective transfer from long-horizon supervision to conventional change analysis. Its performance remains close to the strongest specialist on the remaining benchmarks, showing that the proposed train- ing preserves general remote sensing understanding across single-image, bi-temporal, and short-sequence settings. Temporal Robustness Analysis. Figure 6 separates two sources of temporal difficulty: the number of input frames processed by the model and the temporal span of evidence re- quired by the question. In the top panel, all methods degrade as input sequences become longer, reflecting the increas- ing need to suppress irrelevant observations and localize the informative frames. LongEarth-R1 nevertheless leads every input-length interval by 10.7–31.9%, retaining 51.4 at 26–30 frames after reaching 87.8 on 2–5-frame inputs. This trend indicates that explicit sequence grounding and reward-based alignment improve robustness to long-context distraction, al- though very long inputs remain challenging. The bottom panel shows a different pattern. LongEarth-R1 achieves the best score in six of seven ground-truth evidence- span intervals, and its performance generally improves as the answer can be supported by evidence distributed across a broader temporal span. Thus, a broad evidence span is not necessarily harmful: when multiple observations provide complementary evolution cues, cross-frame reasoning can Figure 6: Task-macro scores by input length (top) and ground-truth evidence span (bottom); * denotes LongEarth-Bench fine-tuning. benefit from them. The only exception is the 23–30-frame interval, where TEOChat* attains a higher score. Ablation Analysis of Core Components. We conduct cu- mulative component and reward ablations across the four cognitive dimensions of LongEarth-Bench. The component study progressively adds SFT, Seq. IDs, R1, and GRPO, whereas the reward study removes one term at a time from the full objective. Table 3 shows that the complete configuration, LongEarth- R1, achieves the strongest macro-average and the most balanced performance across all four dimensions. Rela- tive to LongEarth, adding GRPO yields consistent gains, with the largest improvement on AnomID (+12.61), indi- cating that reward-based optimization strengthens anomaly- sensitive evidence selection and spatiotemporal reasoning. Removing any reward lowers the macro-average by 5.19– 6.68%; although removing the format reward slightly im- proves LogPred, the full objective remains best overall, con- firming the complementarity of the three rewards. Conclusion Long-term remote sensing understanding requires reasoning over evolving geographic evidence rather than isolated ob- servations. We present LongEarth-Bench, which defines this setting with four cognitive dimensions and structured rea- soning supervision. We further develop LongEarth through supervised sequence grounding and LongEarth-R1 through reward-driven spatiotemporal alignment. Results improve performance across the 12 long-sequence tasks while re- taining transfer to conventional remote sensing tasks. These findings support explicit modeling of temporal order, inter- mediate evidence, and spatial consistency for long-horizon Earth observation reasoning. References Bai, S.; Cai, Y.; Chen, R.; et al. 2025a. Qwen3-VL Technical Report. arXiv preprint arXiv:2511.21631. Bai, S.; Chen, K.; Liu, X.; et al. 2025b. Qwen2.5-VL Tech- nical Report. arXiv preprint arXiv:2502.13923. Christie, G.; Fendley, N.; Wilson, J.; and Mukherjee, R. 2018. Functional Map of the World. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 6172–6180. Elgendy, H.; Sharshar, A.; Aboeitta, A.; Ashraf, Y.; and Guizani, M. 2024. GeoLLaVA: Efficient Fine-Tuned Vision- Language Models for Temporal Change Detection in Remote Sensing. arXiv preprint arXiv:2410.19552. Fujita, A.; Sakurada, K.; Imaizumi, T.; Ito, R.; Hikosaka, S.; and Nakamura, R. 2017. Damage Detection from Aerial Im- ages via Convolutional Neural Networks. In 2017 Fifteenth IAPR International Conference on Machine Vision Applica- tions, 5–8. Garioud, A.; De Wit, A.; Poupée, M.; Valette, M.; Giordano, S.; and Wattrelos, B. 2023. FLAIR #2: Textural and Temporal Information for Semantic Segmentation from Multi-Source Optical Imagery. arXiv preprint arXiv:2305.14467. Gupta, R.; Goodman, B.; Patel, N.; Hosfelt, R.; Sajeev, S.; Heim, E. T.; Doshi, J.; Lucas, K.; Choset, H.; and Gas- ton, M. E. 2019. Creating xBD: A Dataset for Assessing Building Damage from Satellite Imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 10–17. Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In ICLR. Hu, Y.; Yuan, J.; Wen, C.; Lu, X.; Liu, Y.; and Li, X. 2025. RSGPT: A Remote Sensing Vision Language Model and Benchmark. ISPRS Journal of Photogrammetry and Remote Sensing, 224: 272–286. Irvin, J. A.; Liu, E. R.; Chen, J. C.; Dormoy, I.; Kim, J.; Khanna, S.; Zheng, Z.; and Ermon, S. 2025. TEOChat: A Large Vision-Language Assistant for Temporal Earth Obser- vation Data. In ICLR. Jain, P.; Marcos, D.; Ienco, D.; Interdonato, R.; and Berchoux, T. 2026. TimeSenCLIP: A Time Series Vision- Language Model for Remote Sensing. ISPRS Journal of Photogrammetry and Remote Sensing, 236: 99–119. Jang, Y.; Kim, D.; Pack, C.; and Won, K. 2024. A Novel Dataset for Flood Detection Robust to Seasonal Changes in Satellite Imagery. In Proceedings of the ACM Research in Adaptive and Convergent Systems Conference. Köksal, A.; and Alatan, A. A. 2026. SAMChat: Introducing Chain-of-Thought Reasoning and GRPO to a Multimodal Small Language Model for Small-Scale Remote Sensing. IEEE Journal of Selected Topics in Applied Earth Observa- tions and Remote Sensing, 19: 795–804. Kuckreja, K.; Danish, M. S.; Naseer, M.; Das, A.; Khan, S.; and Khan, F. S. 2024. GeoChat: Grounded Large Vision- Language Model for Remote Sensing. In CVPR, 27831– 27840. Li, W.; Xiang, X.; Wen, Z.; et al. 2026a. GeoReason: Align- ing Thinking and Answering in Remote Sensing Vision- Language Models via Logical Consistency Reinforcement Learning. arXiv preprint arXiv:2601.04118. Li, Y.; Cao, Y.; Liu, C.; Xiong, Y.; Dong, X.; and Huang, C. 2026b. IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection. In AAAI, 6583–6591. Li, Y.; Xu, W.; Li, G.; Yu, Z.; Wei, Z.; Wang, J.; and Peng, M. 2024. UniRS: Unifying Multi-Temporal Remote Sens- ing Tasks through Vision Language Models. arXiv preprint arXiv:2412.20742. Liu, C.; Chen, K.; Zhang, H.; Qi, Z.; Zou, Z.; and Shi, Z. 2024a. Change-Agent: Toward Interactive Comprehensive Remote Sensing Change Interpretation and Analysis. IEEE Transactions on Geoscience and Remote Sensing, 62: 1–16. Liu, C.; Zhang, J.; Chen, K.; Wang, M.; Zou, Z.; and Shi, Z. 2026a. Remote Sensing Spatiotemporal Vision–Language Models: A Comprehensive Survey. IEEE Geoscience and Remote Sensing Magazine, 14(1): 383–423. Liu, C.; Zhao, R.; Chen, H.; Zou, Z.; and Shi, Z. 2022. Re- mote Sensing Image Change Captioning with Dual-Branch Transformers: A New Method and a Large-Scale Dataset. IEEE Transactions on Geoscience and Remote Sensing, 60: 1–20. Liu, F.; Chen, D.; Guan, Z.; Zhou, X.; Zhu, J.; Ye, Q.; Fu, L.; and Zhou, J. 2024b. RemoteCLIP: A Vision Language Foun- dation Model for Remote Sensing. IEEE Transactions on Geoscience and Remote Sensing, 62: 1–16. Article 5622216. Liu, J.; Sun, L.; Fu, R.; and Yang, B. 2026b. Towards Faith- ful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models. In ICLR. Luo, Z.; Wang, D.; Guo, H.; Zhang, J.; and Du, B. 2026. VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing. arXiv preprint arXiv:2602.07045. Pang, C.; Weng, X.; Wu, J.; Li, J.; Liu, Y.; Sun, J.; Li, W.; Wang, S.; Feng, L.; Xia, G.-S.; and He, C. 2025. VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis. In AAAI, 6381–6388. Sainte Fare Garnot, V.; Landrieu, L.; and Chehata, N. 2022. Multi-Modal Temporal Attention Models for Crop Mapping from Satellite Time Series. ISPRS Journal of Photogramme- try and Remote Sensing, 187: 294–305. Shao, Z.; Wang, P.; Zhu, Q.; Xu, R.; Song, J.; Bi, X.; Zhang, H.; Zhang, M.; Li, Y. K.; Wu, Y.; and Guo, D. 2024. DeepSeekMath: Pushing the Limits of Mathemati- cal Reasoning in Open Language Models. arXiv preprint arXiv:2402.03300. Shen, H.; Liu, P.; Li, J.; Fang, C.; Ma, Y.; Liao, J.; Shen, Q.; Zhang, Z.; Zhao, K.; Zhang, Q.; Xu, R.; and Zhao, T. 2025. VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model. arXiv preprint arXiv:2504.07615. Shen, L.; Lu, Y.; Chen, H.; Wei, H.; Xie, D.; Yue, J.; Chen, R.; Lv, S.; and Jiang, B. 2021. S2Looking: A Satellite Side- Looking Dataset for Building Change Detection. Remote Sensing, 13(24): 5094. Soni, S.; Dudhane, A.; Debary, H.; Fiaz, M.; Munir, M. A.; Danish, M. S.; Fraccaro, P.; Watson, C. D.; Klein, L. J.; Khan, F. S.; and Khan, S. 2025. EarthDial: Turning Multi- Sensory Earth Observations to Interactive Dialogues. In CVPR, 14303–14313. Toker, A.; Kondmann, L.; Weber, M.; et al. 2022. Dynam- icEarthNet: Daily Multi-Spectral Satellite Dataset for Se- mantic Change Segmentation. In CVPR, 21158–21167. Van Etten, A.; Hogan, D.; Martinez-Manso, J.; Shermeyer, J.; Weir, N.; and Lewis, R. 2021. The Multi-Temporal Urban Development SpaceNet Dataset. In CVPR, 6398–6407. Verma, S.; Panigrahi, A.; and Gupta, S. 2021. QFabric: Multi-Task Change Detection Dataset. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Wang, J.; Xuan, W.; Qi, H.; Liu, Z.; Liu, K.; Wu, Y.; Chen, H.; Song, J.; Xia, J.; Zheng, Z.; and Yokoya, N. 2025. Dis- asterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response. In NeurIPS. Wang, J.; Zhang, Z.; Liu, Z.; Li, Y.; Ge, J.; Xie, H.; and Zhang, Y. 2026. SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability. In AAAI, 9912–9920. Wang, J.; Zheng, Z.; Chen, Z.; Ma, A.; and Zhong, Y. 2024a. EarthVQA: Towards Queryable Earth via Relational Reasoning-Based Remote Sensing Visual Question Answer- ing. In AAAI, 5481–5489. Wang, Z.; Wang, M.; Xu, S.; Li, Y.; and Zhang, B. 2024b. CCExpert: Advancing MLLM Capability in Remote Sensing Change Captioning with Difference-Aware Integration and a Foundational Dataset. arXiv preprint arXiv:2411.11360. Weng, X.; Pang, C.; and Xia, G.-S. 2025. Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Per- spectives. IEEE Geoscience and Remote Sensing Magazine, 13(3): 276–323. Xia, G.-S.; Hu, J.; Hu, F.; Shi, B.; Bai, X.; Zhong, Y.; Zhang, L.; and Lu, X. 2017. AID: A Benchmark Data Set for Per- formance Evaluation of Aerial Scene Classification. IEEE Transactions on Geoscience and Remote Sensing, 55(7): 3965–3981. Xuan, W.; Wang, J.; Qi, H.; Chen, Z.; Zheng, Z.; Zhong, Y.; Xia, J.; and Yokoya, N. 2025. DynamicVL: Benchmark- ing Multimodal Large Language Models for Dynamic City Understanding. In NeurIPS. Yang, C.; Li, Z.; Jiao, H.; Gao, Z.; and Zhang, L. 2025. Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning. IEEE Transactions on Image Processing, 34: 7378–7390. Yang, Y.; and Newsam, S. 2010. Bag-of-Visual-Words and Spatial Extensions for Land-Use Classification. In Proceed- ings of the 18th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems, 270–279. Yao, L.; Liu, F.; Lu, H.; Zhang, C.; Min, R.; Xu, S.; Di, S.; and Peng, P. 2026. RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow. In AAAI, 11883–11891. Yerramilli, S.; Pande, N.; Grover, R.; and Tamarapalli, J. S. 2025. GeoChain: Multimodal Chain-of-Thought for Geo- graphic Reasoning. In Findings of EMNLP, 23624–23639. Yuan, Z.; Mou, L.; Xiong, Z.; and Zhu, X. X. 2022. Change Detection Meets Visual Question Answering. IEEE Trans- actions on Geoscience and Remote Sensing, 60: 5630613. Zhang, W.; Cai, M.; Zhang, T.; Zhuang, Y.; and Mao, X. 2024a. EarthGPT: A Universal Multi-Modal Large Language Model for Multi-Sensor Image Comprehension in Remote Sensing Domain. IEEE Transactions on Geoscience and Remote Sensing, 62: 1–20. Zhang, Z.; Zhang, A.; Li, M.; Zhao, H.; Karypis, G.; and Smola, A. 2024b. Multimodal Chain-of-Thought Reasoning in Language Models. Transactions on Machine Learning Research. Zhang, Z.; Zhao, T.; Guo, Y.; and Yin, J. 2024c. RS5M and GeoRSCLIP: A Large-Scale Vision-Language Dataset and a Large Vision-Language Model for Remote Sensing. IEEE Transactions on Geoscience and Remote Sensing, 62: 1–23. Article 5642123.