Paper deep dive
VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction
Wei Zhang, Yihang Wu, Songhua Li, Qi Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/18/2026, 5:58:53 AM
Summary
The paper introduces VGGT-Align, a framework for long-sequence 3D reconstruction that addresses scale drift in chunk-based pipelines. It proposes Scene Geometric Invariant Anchoring (SGIA) to extract geometric invariants (like camera-to-ground distance and road width) from predicted point clouds, using them to constrain scale during Sim(3) alignment, effectively degenerating it to rigid-body transformation. Additionally, it employs a lightweight test-time adaptation strategy to fine-tune normalization layers for improved intra-chunk predictions.
Entities (9)
Relation Signals (8)
VGGT-Align → evaluatedon → KITTI
confidence 99% · Experiments on multiple long-sequence benchmarks... on KITTI Seq 08... on KITTI Seq 05.
VGGT-Align → containsmodule → Scene Geometric Invariant Anchoring
confidence 97% · Our primary contribution is Scene Geometric Invariant Anchoring (SGIA)... As a secondary contribution, we introduce a lightweight test-time adaptation strategy...
Scene Geometric Invariant Anchoring → solvesproblem → Scale Drift
confidence 96% · SGIA... exploits their cross-chunk consistency to establish scale constraints... severing chain-wise scale error propagation at its source.
VGGT-Align → containsmodule → Test-Time Adaptation
confidence 95% · We further introduce a lightweight test-time adaptation strategy that fine-tunes only normalization-layer parameters...
Scene Geometric Invariant Anchoring → transformsprocess → Sim(3) alignment
confidence 94% · explicitly degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation
Scene Geometric Invariant Anchoring → usesinvariant → Camera-to-ground distance
confidence 93% · Two complementary invariant types are exploited: vertical structures (e.g., camera-to-ground distance gk)...
Scene Geometric Invariant Anchoring → usesinvariant → Road width
confidence 93% · horizontal structures (e.g., road width wk)... Both types provide independent scale observations per chunk.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplicatively and distort global trajectories and point cloud geometry. We present a scale-consistency enhancement framework built on a key insight: in structured environments such as driving scenes, geometric quantities arising from environmental regularity remain inherently invariant across temporal segments, and discrepancies in their per-chunk measurements directly expose inter-chunk scale drift. We propose Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric invariants from each chunk's predicted point cloud via coarse-to-fine robust estimation and exploits their cross-chunk consistency to establish scale constraints independent of point cloud registration, explicitly degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation and severing chain-wise scale error propagation at its source. We further introduce a lightweight test-time adaptation strategy that fine-tunes only normalization-layer parameters via multi-objective self-supervision, progressively improving intra-chunk predictions along the sequence. Both modules are plug-and-play and require no offline retraining. Experiments on multiple long-sequence benchmarks demonstrate state-of-the-art performance, reducing absolute trajectory error by up to 32% with significant gains in trajectory stability and reconstruction quality. Code: this https URL
Tags
Links
- Source: https://arxiv.org/abs/2608.15260v1
- Canonical: https://arxiv.org/abs/2608.15260v1
Trouble viewing inline? Open PDF directly →
Full Text
53,765 characters extracted from source content.
Expand or collapse full text
VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D ReconstructionConference: Proceedings of the 34th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, BrazilProceedings of the 34th ACM International Conference on Multimedia (M ’26), November 10–14, 2026, Rio de Janeiro, BrazilDOI: 10.1145/3767308.3836543ISBN: 979-8-4007-2213-4/2026/11mfp9431CCS: Computing methodologies Reconstruction Wei Zhang email: zhangwei707@mail.nwpu.edu.cn Affiliation: School of Computer Science , Northwestern Polytechnical University , Xi’an , China , Yihang Wu email: wuyihang@mail.nwpu.edu.cn Affiliation: Northwestern Polytechnical University , Xi’an , China , Songhua Li email: lisonghua@mail.nwpu.edu.cn Affiliation: Northwestern Polytechnical University , Xi’an , China and Qi Wang email: crabwq@gmail.com Affiliation: Northwestern Polytechnical University , Xi’an , China 2026; © c Figure 1. Comparison of chunk-based long-sequence reconstruction on KITTI. Existing baselines (a, b) suffer from unconstrained scale drift in Sim(3) alignment, manifesting as scale drift (trajectory shrinks into a distorted cluster) and drift accumulation (trajectory diverges from ground truth). (c) VGGT-Align maintains scale consistency by anchoring each chunk via scene geometric invariants, reducing ATE by 30% on Seq 08 and 45% on Seq 05. Right: inter-chunk point cloud overlay shows our scale correction preserves geometric consistency between adjacent chunk (top, ✓), while the baseline exhibits severe drift at chunk boundaries (bottom, ✗). VGGT-Align achieves state-of-the-art among chunk-based methods.A comparison of long-sequence 3D reconstruction methods on KITTI. Baseline trajectories exhibit accumulated scale drift and geometric misalignment at chunk boundaries, whereas VGGT-Align remains close to the ground-truth trajectory and preserves overlap between adjacent reconstructed point-cloud chunks. Abstract. Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critical failure mode. In chunk-based inference pipelines, the scale degree of freedom in sequential Sim(3) alignment is left unconstrained, causing estimation errors to compound multiplicatively and distort global trajectories and point cloud geometry. We present a scale-consistency enhancement framework built on a key insight: in structured environments such as driving scenes, geometric quantities arising from environmental regularity remain inherently invariant across temporal segments, and discrepancies in their per-chunk measurements directly expose inter-chunk scale drift. We propose Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric invariants from each chunk’s predicted point cloud via coarse-to-fine robust estimation and exploits their cross-chunk consistency to establish scale constraints independent of point cloud registration, explicitly degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation and severing chain-wise scale error propagation at its source. We further introduce a lightweight test-time adaptation strategy that fine-tunes only normalization-layer parameters via multi-objective self-supervision, progressively improving intra-chunk predictions along the sequence. Both modules are plug-and-play and require no offline retraining. Experiments on multiple long-sequence benchmarks demonstrate state-of-the-art performance, reducing absolute trajectory error by up to 32% with significant gains in trajectory stability and reconstruction quality. Code: https://github.com/WZ-CS/VGGT-Align. Keywords: Multi-view stereo, Long-Sequence 3D Reconstruction, Chunk-Based Alignment, Test-Time Adaptation, Scale Drift †c-license: by-nc-nd Figure 2. Qualitative analysis of camera pose estimation on KITTI Seq 08. (a) Bird’s-eye view trajectory comparison. The baseline exhibits significant scale drift in the later portion of the sequence, causing the estimated trajectory to deviate substantially from the ground truth, while our method maintains closer agreement throughout by introducing global scale prior constraints during cross-chunk alignment. (b) Per-chunk scale consistency after Sim(3) alignment. Each point represents the local path length ratio (predicted / ground truth) within a chunk of 75 frames. σ denotes its standard deviation across all chunks. The baseline suffers progressive scale compression, dropping to approximately 0.5 by the end of the sequence (σ=0.157), while our method maintains the ratio near the ideal value of 1.0 (σ=0.065). 1. Introduction Feed-forward 3D vision models (Vaswani et al. 2017; Wang et al. 2024a; Wang et al. 2025a; Wang et al. 2024b; Leroy et al. 2024) have demonstrated capability in recovering dense geometry from unposed images (Zhang et al. 2024; Zhang et al. 2025a; Zhang et al. 2025b; Zhang et al. 2026b; Zhang et al. 2026a). By predicting camera parameters, depth maps, and 3D point clouds in a single forward pass, these methods bypass the iterative optimization of classical SfM and SLAM (Triggs et al. 1999; Schonberger and Frahm 2016), achieving efficiency on short sequences. However, their fixed context window, typically limited to tens of frames, prevents direct application to long sequences comprising thousands of frames, as encountered in autonomous driving and large-scale mapping. To bridge this gap, chunk-based methods (Deng et al. 2025; Lee et al. 2025) partition sequences into overlapping segments, process each independently with a feed-forward model, and sequentially align local reconstructions via Sim(3) transformations estimated from overlap regions. While effective locally, this paradigm introduces a critical vulnerability: unconstrained scale drift. The scale factor at each step is estimated solely from noisy overlap consistency, and since each chunk’s coordinate frame is defined only relative to its predecessor, scale errors propagate multiplicatively along the chain, causing progressive trajectory distortion from local inconsistency to global structural collapse (fig. 1). The root cause is that scale is treated as a free variable without independent constraint. Yet we observe that this is unnecessarily permissive: many structured environments, particularly those with rigidly mounted cameras and regular geometry (e.g., driving scenes), exhibit geometric quantities that remain inherently invariant across temporal segments. Although their per-chunk measurements vary due to unknown local scale, the ratios between adjacent chunks directly reveal inter-chunk scale discrepancy, independently of point cloud registration. Building on this observation, we propose VGGT-Align, a scale-consistency enhancement framework for long-sequence reconstruction. Our primary contribution is Scene Geometric Invariant Anchoring (SGIA), which extracts dominant geometric structures from each chunk’s predicted point cloud via coarse-to-fine robust estimation, and exploits their cross-chunk consistency to establish scale constraints independent of overlap-based alignment. By injecting these constraints, the 7-DoF Sim(3) alignment is explicitly degenerated into a 6-DoF rigid-body transformation, severing chain-wise error propagation at its source. The mechanism operates purely on the model’s own predictions, requiring no external sensors and no offline training. As a secondary contribution, we introduce a lightweight test-time adaptation (TTA) strategy that fine-tunes only normalization-layer parameters during inference, using photometric, geometric, and temporal smoothness self-supervision. This enables the model to progressively adapt to the current scene, improving intra-chunk prediction quality and complementing inter-chunk scale anchoring. Our contributions are summarized as follows: • We identify and formally characterize scale drift in chunk-based reconstruction, showing it arises from unconstrained scale in sequential Sim(3) alignment and compounds multiplicatively. • We propose SGIA, which leverages inherent scene structural regularity to provide per-chunk scale constraints, degenerating Sim(3) into rigid-body transformation and eliminating scale drift at its source. • We introduce a complementary test-time adaptation strategy that improves intra-chunk predictions via online normalization layer fine-tuning. • Extensive experiments on long-sequence benchmarks demonstrate that VGGT-Align achieves state-of-the-art performance, reducing average trajectory error by up to 30% over existing methods, with consistent improvements in trajectory stability, scale consistency, and reconstruction quality. 2. Related Work Feed-forward 3D models and long-sequence reconstruction. Feed-forward transformers such as DUSt3R (Wang et al. 2024b), MASt3R (Leroy et al. 2024), and VGGT (Wang et al. 2025a) predict dense 3D quantities in a single pass, but are limited by their fixed context window and run out of memory on long sequences. Fast3R (Yang et al. 2025) extends to larger frame sets but still cannot handle thousands of frames on consumer hardware. To address this, chunk-based methods partition videos into overlapping segments and stitch local reconstructions via Sim(3) alignment. VGGT-Long (Deng et al. 2025) introduced this paradigm with IRLS alignment and loop closure; SwiftVGGT (Lee et al. 2025) accelerated it via reliability-guided sampling and internal token reuse for loop detection (Nistér 2004; Rublee et al. 2011; Sarlin et al. 2020; Lindenberger et al. 2023). Other scaling strategies include real-time SLAM backends (Murai et al. 2025; Maggio et al. 2025; Xiong et al. 2026), persistent 3D states (Wang et al. 2025b), causal KV caching (Zhuo et al. 2025), and dynamic scene extensions and robust reconstruction benchmarks (Hu et al. 2025; Ma et al. [n. d.]). However, all sequential alignment approaches leave the scale degree of freedom unconstrained, leading to multiplicative drift that our work directly addresses. Scale priors and test-time adaptation. Recovering metric scale from monocular has been a longstanding challenge, approached through ground plane constraints (Zhou et al. 2019), known object dimensions and structural priors (Song and Chandraker 2015; Yang and Scherer 2019), IMU fusion (Campos et al. 2021), and metric depth networks (Bhat et al. 2023; Yin et al. 2023; Hu et al. 2024). Self-supervised monocular depth methods (Godard et al. 2019; Guizilini et al. 2023) also exploit vehicle velocity or stereo baselines as scale signals during training. Our method differs by operating on predicted 3D point clouds at inference time and requiring only cross-chunk consistency of geometric invariants rather than known absolute values. For test-time adaptation (Zhao et al. 2026; Jia et al. 2025a; Jia et al. 2025b; Jia et al. 2025c), prior works address domain shift in depth estimation (Tonioni et al. 2019a; Kuznietsov et al. 2021), stereo (Tonioni et al. 2019b), and image classification (Wang et al. 2020; Xiao et al. 2026b; Xiao et al. 2026a; Xiao et al. 2026c; Xiao et al. 2025). We adopt a similar strategy tailored to chunk-based reconstruction, carrying adapted parameters forward across chunks to progressively improve predictions (Lyu et al. 2025c; Lyu et al. 2025b; Lyu et al. 2025a). Figure 3. Per-chunk scale ratio analysis across three KITTI sequences of varying length. The ideal ratio is 1.0 (dashed line); the green band denotes ± 5% tolerance. The baseline (red) exhibits systematic scale drift with high variance, while VGGT-Align (blue) maintains ratios tightly around 1.0 with consistently lower variance (e.g., σ: 0.184→ 0.079 on Seq 10, 0.064→ 0.030 on Seq 04). Inset trajectory plots show corresponding ATE improvements. On Seq 02, scale correction is visibly effective (σ: 0.178→ 0.097), though the trajectory exhibits a slight positional offset due to residual rotational drift in this 5 km loop; see supplementary material for extended visualizations. 3. Method 3.1. Preliminaries and Problem Formulation Chunk-based long-sequence reconstruction. Given a long image sequence ℐ=I1,…,INI=\I_1,…,I_N\, we partition it into M overlapping chunks Ckk=1M\C_k\_k=1^M with window size L and overlap O. Each chunk is processed by a feed-forward model ℱF (e.g., VGGT (Wang et al. 2025a)), outputting a local point cloud k∈ℝL×H×W×3P_k ^L× H× W× 3 with per-pixel confidence kW_k, camera extrinsics k(i)T_k^(i), and intrinsics k(i)K_k^(i). Adjacent chunks are stitched by estimating a similarity transformation (sk,k,k)∈Sim(3)(s_k,R_k,t_k) (3) from the overlap region via confidence-weighted IRLS (Deng et al. 2025): (1) sk∗,k∗,k∗=argmins,,∑i∈wiρ(‖i(k)−(si(k+1)+)‖),s_k^*,R_k^*,t_k^*= s,R,t _i w_i\,ρ (\|p_i^(k)-(sRp_i^(k+1)+t)\| ), where O indexes overlapping frames, wiw_i is point confidence, and ρ(⋅)ρ(·) is a robust loss. The scale drift problem. Transforming chunk CkC_k into the global frame of C1C_1 requires composing all intermediate transforms, yielding a cumulative scale s¯k=∏j=1k−1sj s_k= _j=1^k-1s_j. Since each sj=sj∗(1+ϵj)s_j=s_j^*(1+ _j) carries estimation error ϵj _j, the cumulative scale becomes: (2) s¯k=s¯k∗∏j=1k−1(1+ϵj). s_k= s_k^* _j=1^k-1(1+ _j). Even a modest per-step bias ([ϵj]=0.02E[ _j]=0.02) compounds exponentially: 1.0250≈2.7×1.02^50≈ 2.7× after 50 chunks. We verify this empirically in fig. 3: the baseline’s per-chunk scale ratio deviates systematically from 1.0 with high variance (σ=0.184σ=0.184), confirming compounding bias rather than zero-mean noise. This motivates our core contribution: an independent scale constraint decoupled from overlap-based registration. 3.2. Framework Overview Figure 4. Overview of the VGGT-Align framework. A long RGB sequence is partitioned into overlapping chunks and processed independently by a feed-forward 3D model. Our two plug-and-play contributions are highlighted: (1) Scene Geometric Invariant Anchoring (SGIA, green, top-right) extracts dominant geometric structures from each chunk’s predicted point cloud via region selection with confidence filtering, followed by coarse-to-fine plane estimation (RANSAC + SVD). Two complementary invariant types are exploited: vertical structures (e.g., camera-to-ground distance gkg_k) and horizontal structures (e.g., road width wkw_k), whose cross-chunk measurement ratios yield per-pair scale constraints ssgias_sgia independent of point cloud registration. This constraint is injected into the overlap-based IRLS alignment, explicitly replacing the estimated scale factor (lock symbol) and degenerating the 7-DoF Sim(3) transformation into a 6-DoF rigid-body transform SE(3), thereby severing chain-wise scale drift at its source. (2) Test-time adaptation (TTA, pink, bottom-left) fine-tunes only normalization-layer parameters using three self-supervised objectives: photometric consistency, geometric alignment, and temporal smoothness, with the adapted parameters improving predictions for subsequent chunks (dashed feedback arrow). The aligned chunks are further refined by loop closure optimization to produce the final globally consistent trajectory and dense point cloud. An overview of VGGT-Align is shown in fig. 4. The pipeline proceeds as follows: (1) The input sequence is partitioned into overlapping chunks via a sliding window. (2) Each chunk is processed by a feed-forward 3D model to obtain per-chunk predictions (depth, poses, and 3D points). Prior to inference, TTA optionally adapts the model’s normalization-layer parameters using self-supervised objectives from the previous chunk, progressively improving prediction quality along the sequence. (3) For each chunk, SGIA extracts geometric invariant observations from the predicted point cloud. (4) During pairwise alignment, the SGIA-derived scale constraint is injected into the IRLS result, explicitly replacing the estimated scale factor and degenerating the Sim(3)Sim(3) alignment into a rigid-body SE(3)SE(3) transformation. (5) An optional loop closure module further refines global consistency. The two proposed modules are complementary: SGIA operates at the inter-chunk alignment stage, constraining the scale degree of freedom; TTA operates at the intra-chunk inference stage, improving prediction quality. Both are plug-and-play and require no offline retraining. 3.3. Scene Geometric Invariant Anchoring 3.3.1. Geometric Invariant Observation The key insight behind SGIA is the following: natural scenes contain geometric quantities that are inherently constant in the physical world but whose measured values in each chunk’s local coordinate system vary due to the unknown, chunk-specific scale factor. The discrepancy between measurements from adjacent chunks reveals their relative scale offset, independently of the overlap-based point cloud registration. Formally, let g∈ℝ+g ^+ be a geometric quantity that is constant across all chunks. In chunk CkC_k’s local coordinate system, this quantity is measured as gkg_k, related to the true value through the chunk’s implicit absolute scale σk _k: (3) gk=g/σk.g_k=g\,/\, _k. For adjacent chunks CkC_k and Ck+1C_k+1, the ratio of their measurements gives the true relative scale: (4) gkgk+1=σk+1σk=sk→k+1∗. g_kg_k+1= _k+1 _k=s_k→ k+1^*. Crucially, eq. 4 holds for any geometric invariant g and does not require knowing the absolute value of g, only that it remains constant across chunks. This makes the constraint broadly applicable without external calibration. We identify two complementary classes of geometric invariants prevalent in structured environments: Vertical invariants capture the relationship between the camera and a dominant horizontal plane. The most natural instance is the camera-to-ground distance gkg_k, which remains constant when the camera is rigidly mounted, as in vehicle-mounted settings. Horizontal invariants capture lateral scene structure that persists across chunks. In driving scenarios, road width wkw_k, determined by lane markings and road boundaries, is locally constant over extended stretches. Both types provide independent scale observations per chunk. When multiple sources are available, they can be fused via weighted combination to improve robustness (section 3.3.3). 3.3.2. Robust Invariant Extraction The extraction pipeline consists of three stages: candidate region selection, coarse-to-fine plane estimation, and invariant measurement. Stage 1: Candidate region selection. For vertical invariants, we exploit the spatial prior that ground surfaces project onto the lower portion of the image. For each frame in chunk CkC_k, we select 3D points corresponding to the bottom ρ fraction of image rows (default ρ=0.3ρ=0.3) as ground candidates, and further filter by retaining only points whose confidence exceeds the τ-th percentile (default τ=40τ=40). For horizontal invariants, a similar strategy selects points from a central horizontal band. Stage 2: Coarse-to-fine plane estimation. We fit a plane ⊤+d=0n x+d=0 to the aggregated candidates via a two-stage procedure. In the coarse stage, RANSAC (Fischler and Bolles 1981) identifies an initial inlier set. In the refinement stage, we compute the covariance matrix of the inliers and extract the plane normal as the eigenvector of the smallest eigenvalue via SVD: (5) =(in−¯)⊤(in−¯),∗=min(),C=(X_in- x) (X_in- x), ^*=v_ (C), where inX_in denotes the inlier point matrix, ¯ x is the inlier centroid, and min(⋅)v_ (·) returns the eigenvector of the smallest eigenvalue. This combines the outlier robustness of RANSAC with the statistical optimality of PCA-based fitting. Stage 3: Invariant measurement. Given the estimated plane (∗,d∗)(n^*,d^*), we compute the signed distance from each frame’s camera center f=k(f)[:3,3]c_f=T_k^(f)[:3,3] to the plane: (6) hf=∗⊤f+d∗.h_f=n^* c_f+d^*. The vertical invariant for chunk CkC_k is the median over valid frames: (7) gk=medianhf∣hf>0,f=1,…,|Ck|.g_k=median\h_f h_f>0,\;f=1,…,|C_k|\. The median provides robustness to outlier frames caused by partial occlusion or non-planar regions. For horizontal invariants, wkw_k is computed analogously as the median lateral extent of ground-plane inliers perpendicular to the estimated forward motion. 3.3.3. Scale Injection and Alignment Degeneracy Given invariant observations from adjacent chunks, we compute the prior-based relative scale. For a single invariant source: (8) ssgia=gk/gk+1.s_sgia=g_k\,/\,g_k+1. When multiple sources are available (e.g., both vertical gkg_k and horizontal wkw_k), we fuse them via weighted combination: (9) ssgia=λg⋅gkgk+1+λw⋅wkwk+1,λg+λw=1,s_sgia= _g· g_kg_k+1+ _w· w_kw_k+1, _g+ _w=1, where the weights reflect relative reliability (default λg=0.7 _g=0.7, λw=0.3 _w=0.3). Scale replacement. The IRLS alignment (eq. 1) yields (sirls,,irls)(s_irls,R,t_irls). We replace the estimated scale with the prior-derived value: (10) s=α⋅ssgia+(1−α)⋅sirls,s=α· s_sgia+(1-α)· s_irls, where α∈[0,1]α∈[0,1] controls the prior strength. We find α=1α=1 optimal on KITTI, while α∈[0.7,0.9]α∈[0.7,0.9] works better on datasets with less regular geometry. Since the rotation R is retained from IRLS, the translation must be adjusted to preserve centroid alignment of the overlap region: (11) =irls+(sirls−s)⋅¯2,t=t_irls+(s_irls-s)·R p_2, where ¯2 p_2 is the overlap centroid in chunk Ck+1C_k+1’s frame. Alignment degeneracy: Sim(3)→SE(3)Sim(3) (3). When α=1α=1, the scale is fully determined by the geometric invariant, independent of overlap-based registration. The 7-DoF Sim(3)Sim(3) alignment degenerates into 6-DoF SE(3)SE(3). This breaks the multiplicative error chain in eq. 2: the cumulative scale is now determined by per-chunk invariant ratios rather than compounding noisy IRLS estimates. Since each chunk’s invariant is measured independently, errors do not accumulate. Robustness and fallback. When invariant extraction fails for a particular chunk (e.g., insufficient ground points), the system falls back to the IRLS-only scale by setting α=0α=0. An adaptive blending strategy can further adjust α based on extraction confidence (RANSAC inlier count, per-frame height variance), providing graceful degradation. 3.4. Test-Time Adaptation While SGIA addresses inter-chunk alignment, per-chunk prediction quality can degrade over long sequences due to distributional shift between training data and the test scene. We introduce a lightweight test-time adaptation (TTA) mechanism that enables the model to progressively adapt to the current scene. Parameter selection. To avoid catastrophic forgetting, we restrict adaptation to normalization-layer parameters only (LayerNorm scale and bias). These constitute a negligible fraction of total parameters but modulate feature distributions at every layer, making them effective targets for distribution alignment. All other parameters remain frozen. Self-supervised objectives. We employ three complementary losses computed without ground-truth supervision: (1) Photometric consistency ℒphotoL_photo: we warp source frames to target views using predicted depth and poses, and measure reconstruction error with a combination of SSIM and L1: (12) ℒphoto=1||∑(i,j)∈(0.85⋅ℒSSIM(Iiw,Ij)+0.15⋅‖Iiw−Ij‖1),L_photo= 1|N| _(i,j) (0.85·L_SSIM(I_i^w,I_j)+0.15·\|I_i^w-I_j\|_1 ), where N is the set of adjacent frame pairs and IiwI_i^w denotes frame i warped to view j. (2) Geometric consistency ℒgeoL_geo: we penalize the confidence-weighted centroid discrepancy between point clouds from neighboring frames. (3) Temporal smoothness ℒsmoothL_smooth: we regularize the second-order derivative of predicted camera translations: (13) ℒsmooth=1N−2∑i=1N−2‖i+2−2i+1+i‖2.L_smooth= 1N\!-\!2 _i=1^N-2\|t_i+2-2t_i+1+t_i\|^2. The total loss is ℒtta=λpℒphoto+λgℒgeo+λsℒsmoothL_tta= _pL_photo+ _gL_geo+ _sL_smooth, with default weights λp=1.0 _p\!=\!1.0, λg=0.5 _g\!=\!0.5, λs=0.1 _s\!=\!0.1. Adaptation protocol. For each chunk CkC_k (after an initial warmup), we perform T gradient steps (default T=3T\!=\!3) on ℒttaL_tta with learning rate 10−410^-4, Adam optimization (Kingma and Ba 2014), and gradient clipping. To reduce memory overhead, TTA operates on a subset of n frames (default n=3n\!=\!3) at reduced resolution (336px vs. 518px for inference), consuming approximately 70% less VRAM. The adaptation does not alter the current chunk’s predictions; the updated parameters take effect during the next chunk’s forward pass. Parameters are accumulated across chunks by default, enabling progressive specialization to the scene. This adapt-then-infer protocol ensures each chunk benefits from self-supervised signals of all preceding chunks without re-inference. Methods LC Calibration Recon. Avg. 00 01 02 03 04 05 06 07 08 09 10 seq. frames - - - 2109 4542 1101 4661 801 271 2761 1101 1101 4071 1591 1201 seq. length (m) - - - 2012 3724 2453 5067 561 394 2206 1233 650 3223 1705 920 contains loop - - - - ✓ ✗ ✓ ✗ ✗ ✓ ✓ ✓ ✗ ✓ ✗ DROID-VO (Teed and Deng 2021) ✗ Required Dense 54.19 98.43 84.20 108.80 2.58 0.93 59.27 64.40 24.20 64.55 71.80 16.91 DPVO (Teed et al. 2024) ✗ Required Sparse 53.61 113.21 12.69 123.40 2.09 0.68 58.96 54.78 19.26 115.90 75.10 13.63 DROID-SLAM (Teed and Deng 2021) - Required Dense 100.28 92.10 344.60 107.61 2.38 1.00 118.50 62.47 21.78 161.60 72.32 118.70 DPV-SLAM (Teed et al. 2023) ✓ Required Sparse 53.03 112.80 11.50 123.53 2.50 0.81 57.80 54.86 18.77 110.49 76.66 13.65 DPV-SLAM++ (Teed et al. 2023) ✓ Required Sparse 25.75 8.30 11.86 39.64 2.50 0.78 5.74 11.60 1.52 110.90 76.70 13.70 MASt3R-SLAM (Murai et al. 2025) ✓ No Need Dense / TL TL TL TL TL TL TL TL TL TL TL CUT3R (Wang et al. 2025b) ✗ No Need Dense / OOM OOM OOM 148.07 22.31 OOM OOM OOM OOM OOM OOM Fast3R (Yang et al. 2025) ✗ No Need Dense / OOM OOM OOM OOM OOM OOM OOM OOM OOM OOM OOM VGGT (Wang et al. 2025a) ✗ No Need Dense / OOM OOM OOM OOM OOM OOM OOM OOM OOM OOM OOM VGGT-Long (Deng et al. 2025) ✓ No Need Dense 29.41 9.87 111.06 37.56 4.89 3.75 9.09 7.47 4.02 62.86 47.48 25.49 SwiftVGGT (Lee et al. 2025) ✓ No Need Dense 29.18 8.17 102.53 36.49 8.12 4.88 11.94 8.88 5.01 64.68 44.13 26.18 VGGT-Align (Ours) ✓ No Need Dense 19.99 7.74 65.11 47.04 5.12 2.55 4.97 4.21 4.48 43.98 23.11 11.66 Table 1. Camera tracking results (ATE RMSE [m] ↓ ) on the KITTI Odometry benchmark. LC: loop closure capability. VGGT-Align achieves the best overall accuracy among all calibration-free or calibration-required methods. OOM: CUDA Out-Of-Memory on a single RTX 4090. TL: Tracking Lost. Color: 1st, 2nd, 3rd. Methods Calib. Avg. 163453191 183829460 315615587 346181117 371159869 405841035 460417311 520018670 610454533 Frames / Length (m) - 198 / 173 198 / 160 199 / 42 199 / 165 199 / 351 196 / 273 199 / 86 198 / 266 199 / 135 198 / 63 DROID-SLAM (Teed and Deng 2021) Required 4.396 3.705 0.301 0.447 8.653 9.320 7.621 4.170 TL 0.264 MASt3R-SLAM (Murai et al. 2025) No Need 5.560 4.500 0.556 1.833 12.544 8.601 1.412 5.428 7.910 1.195 CUT3R (Wang et al. 2025b) No Need 9.872 8.781 3.810 5.790 24.015 13.070 7.261 13.206 8.597 3.229 Fast3R (Yang et al. 2025) No Need / OOM OOM OOM OOM OOM OOM OOM OOM OOM VGGT (Wang et al. 2025a) No Need / OOM OOM OOM OOM OOM OOM OOM OOM OOM VGGT-Long (Deng et al. 2025) No Need 3.085 3.086 2.544 2.336 4.007 4.045 3.141 2.867 3.407 2.333 SwiftVGGT (Lee et al. 2025) No Need 2.854 3.106 2.447 2.346 2.719 4.080 3.106 2.841 2.832 2.210 VGGT-Align (Ours) No Need 1.849 1.160 2.540 0.537 3.140 2.850 1.020 1.520 2.170 1.708 Table 2. Camera tracking results (ATE RMSE [m] ↓ ) on the Waymo Open Dataset. VGGT-Align achieves the best average accuracy among all calibration-free methods and surpasses calibration-required DROID-SLAM. Color: 1st, 2nd, 3rd. 4. Experiments 4.1. Experimental Setup We evaluate on three driving benchmarks: KITTI Odometry (Geiger et al. 2012), Waymo Open Dataset (Sun et al. 2020), and Virtual KITTI (Gaidon et al. 2016), comparing against both calibration-required and calibration-free baselines. For tracking we report ATE RMSE (m); for reconstruction on Waymo we report Accuracy, Completeness, and Chamfer Distance against LiDAR ground truth. All experiments use a single NVIDIA RTX 4090 with VGGT (Wang et al. 2025a) as backbone. Full baseline details, per-dataset settings, and implementation specifics are provided in the appendix. 4.2. Camera Tracking Results KITTI Odometry. table 1 presents results on all 11 KITTI sequences. VGGT-Align achieves the best overall average ATE (19.99) among all methods, ranking first on 7 out of 11 sequences including the longest ones: Seq 00 (4,542 frames, ATE: 7.74), Seq 05 (ATE: 4.97), Seq 08 (ATE: 43.98), Seq 09 (ATE: 23.11), and Seq 10 (ATE: 11.66). Compared to VGGT-Long, our method reduces the overall average by 32% (29.41→ 19.99), with up to 54% reduction on individual sequences (Seq 10: 25.49→ 11.66). Our Avg∗ (7.02) is comparable to VGGT-Long (6.91) while substantially outperforming SwiftVGGT (20.73). On Seq 01, which involves high-speed highway driving (2.23 m /frame), all chunk-based methods struggle due to extreme inter-frame displacement. On Seq 02, our ATE (47.04) is higher than the baseline (37.56); this 5 km sequence is dominated by rotational drift in long loops, which scale anchoring alone cannot fully address. fig. 5 provides trajectory and reconstruction visualizations across four sequences. VGGT-Align consistently produces trajectories that most faithfully follow the ground-truth shape. fig. 6 further compares dense point clouds on Seq 05: VGGT-Long exhibits misaligned road surfaces and duplicated buildings at chunk boundaries, while VGGT-Align produces geometrically coherent reconstruction with continuous surfaces and sharp edges. Figure 5. Qualitative comparison of trajectories and dense reconstructions across four KITTI sequences. Columns 1–4: DROID-SLAM, VGGT-Long, SwiftVGGT, and VGGT-Align (Ours); ground truth in gray dashed, predictions in green, red boxes highlight deviations. Column 5: LiDAR pseudo ground-truth point cloud. VGGT-Align achieves the lowest ATE on all four sequences with notable improvements on Seq 06 (62.47→ 4.21), Seq 09 (72.32→ 23.11), and Seq 10 (118.70→ 11.66). Figure 6. Dense point cloud comparison on KITTI Seq 05. Top: reference frames (#2580, #2588). (a) VGGT-Long: misaligned surfaces, duplicated structures, and scattered outliers at chunk boundaries. (b) VGGT-Align: coherent reconstruction with continuous surfaces and sharp edges from SGIA. Waymo Open Dataset. As shown in table 2, VGGT-Align achieves the best average ATE (1.849) among all calibration-free methods, significantly outperforming VGGT-Long (3.085) and SwiftVGGT (2.854) by 40% and 35%, respectively. Our method ranks first on 5 out of 9 segments and achieves particularly strong results on geometrically complex segments such as 405841035 (1.020) and 460417311 (1.520). Remarkably, VGGT-Align also surpasses calibration-required DROID-SLAM (4.396), demonstrating that well-constrained chunk-based alignment can effectively rival traditional SLAM systems even without access to camera intrinsics. Trajectory and reconstruction visualizations are provided in the supplementary material. Virtual KITTI. table 3 evaluates robustness under six appearance conditions on Scene 20, the longest sequence in the dataset (837 frames, 711 m). VGGT-Align achieves first or second place across all conditions, significantly improving over VGGT-Long. Our calibration-free method even surpasses DROID-SLAM (which requires ground-truth intrinsics) on Morning (3.17 vs. 3.73) and Sunset (3.65 vs. 4.91), confirming that geometric invariants are appearance-agnostic. Results on additional scenes are provided in the supplementary material. Calib. Scene 20 837 frames, 711 m Condition - Clone Fog Morning Overcast Rain Sunset DROID-SLAM (Teed and Deng 2021) Required 3.59 5.08 3.73 3.85 3.78 4.91 MASt3R-SLAM (Murai et al. 2025) No Need TL TL TL TL TL TL CUT3R (Wang et al. 2025b) No Need 129.50 76.96 117.95 114.51 66.70 116.53 Fast3R (Yang et al. 2025) No Need OOM OOM OOM OOM OOM OOM VGGT (Wang et al. 2025a) No Need OOM OOM OOM OOM OOM OOM VGGT-Long (Deng et al. 2025) No Need 9.66 8.19 6.34 4.56 6.50 4.85 VGGT-Align (Ours) No Need 3.70 6.20 3.17 4.27 4.68 3.65 Table 3. Camera tracking results (ATE RMSE [m] ↓ ) on Virtual KITTI Scene 20 under six appearance conditions. VGGT-Align significantly narrows the gap to calibration-required DROID-SLAM and surpasses it under Morning and Sunset. 4.3. 3D Reconstruction Quality table 4 evaluates dense reconstruction on Waymo. VGGT-Align achieves the best average Accuracy (1.056) and Chamfer Distance (1.541) among all calibration-free methods, confirming that scale-consistent alignment directly translates into higher-quality geometry. Note that LiDAR ground truth has a narrower vertical FoV than RGB cameras, so metrics should be interpreted alongside the qualitative results in fig. 6. Methods Metric Calib. Avg. 163453191 183829460 315615587 346181117 371159869 405841035 460417311 520018670 610454533 Frames / Length (m) - - 198 / 173 198 / 160 199 / 42 199 / 165 199 / 351 196 / 273 199 / 86 198 / 266 199 / 135 198 / 63 DROID-SLAM (Teed and Deng 2021) Accuracy ↓ Req. 1.201 0.781 1.136 2.247 2.393 1.090 0.539 0.740 TL 0.677 Completeness ↓ 8.540 4.610 10.245 5.540 8.669 8.592 11.144 5.320 TL 14.201 Chamfer ↓ 4.870 2.696 5.691 3.893 5.531 4.841 5.842 3.030 TL 7.439 MASt3R-SLAM (Murai et al. 2025) Accuracy ↓ No 3.772 3.189 2.988 3.787 4.689 4.436 1.166 4.637 6.417 2.637 Completeness ↓ 3.177 1.715 3.284 2.047 2.981 2.679 2.895 2.002 4.429 6.560 Chamfer ↓ 3.474 2.452 3.136 2.917 3.835 3.558 2.031 3.319 5.423 4.599 CUT3R (Wang et al. 2025b) Accuracy ↓ No 3.884 3.580 1.144 2.418 3.712 3.679 4.346 2.012 12.320 1.744 Completeness ↓ 6.801 8.251 9.352 8.748 8.537 5.467 3.393 6.164 2.302 8.999 Chamfer ↓ 5.343 5.916 5.248 5.583 6.125 4.573 3.869 4.088 7.311 5.371 VGGT-Long (Deng et al. 2025) Accuracy ↓ No 1.182 1.002 0.395 0.925 1.668 2.580 0.679 0.784 1.358 1.246 Completeness ↓ 2.860 2.762 3.417 1.738 3.261 2.791 3.216 1.840 4.694 2.022 Chamfer ↓ 2.021 1.882 1.906 1.331 2.465 2.685 1.948 1.312 3.026 1.634 SwiftVGGT (Lee et al. 2025) Accuracy ↓ No 1.339 1.175 1.364 1.102 1.341 2.622 0.526 1.256 1.500 1.165 Completeness ↓ 1.985 1.049 4.331 1.317 1.353 1.716 3.683 1.236 0.963 2.221 Chamfer ↓ 1.662 1.112 2.848 1.210 1.347 2.169 2.104 1.246 1.231 1.693 VGGT-Align (Ours) Accuracy ↓ No 1.056 0.964 0.489 0.777 0.835 1.713 0.537 0.774 1.191 2.228 Completeness ↓ 2.026 1.628 3.551 1.435 1.386 2.475 2.262 1.569 1.720 2.212 Chamfer ↓ 1.541 1.296 2.020 1.106 1.110 2.094 1.400 1.171 1.455 2.220 Table 4. Dense reconstruction results on the Waymo Open Dataset. Metrics are Accuracy, Completeness, and Chamfer Distance (all in meters ↓ ) against LiDAR ground truth. Note that LiDAR has a narrower vertical FoV than RGB cameras, so metrics should be interpreted alongside qualitative results. Color: 1st, 2nd, 3rd. 4.4. Ablation Study We ablate each component on Waymo (table 5). Starting from the VGGT-Long baseline (Avg: 2.154), the ground plane prior alone (GP) does not uniformly help (2.442) due to suboptimal fixed blending. Adaptive blending (AB) reduces this to 2.173. Adding road width prior (RW) brings significant improvement (1.856, −-14%), confirming that multi-source invariant fusion is more robust than either source alone. TTA provides a further modest gain, yielding the full model at 1.849. Results on additional scenes are provided in the supplementary material. GP RW AB TTA Avg. 163.. 183.. 315.. 346.. 371.. 405.. 460.. 520.. 610.. ✗ ✗ ✗ ✗ 2.154 1.780 2.570 0.760 3.730 3.225 1.390 1.690 3.477 0.355 ✓ ✗ ✗ ✗ 2.442 3.930 2.670 0.588 2.790 3.560 1.120 2.460 2.060 2.800 ✓ ✗ ✓ ✗ 2.173 2.540 2.600 0.616 3.360 3.320 1.190 1.830 2.060 2.040 ✓ ✓ ✓ ✗ 1.856 1.290 2.550 0.495 3.280 2.960 0.780 1.530 2.190 1.629 ✓ ✓ ✓ ✓ 1.849 1.160 2.540 0.537 3.140 2.850 1.020 1.520 2.170 1.708 Table 5. Ablation on Waymo (ATE RMSE [m] ↓ ). Row 1: VGGT-Long baseline. GP: Ground plane Prior. RW: Road Width prior. AB: Adaptive Blending. TTA: Test-Time Adaptation. Row 5: full VGGT-Align. Color: 1st, 2nd, 3rd. 4.5. Runtime Analysis We compare runtime against VGGT-Long on KITTI Seq 00 (4,542 frames, 60 chunks, RTX 4090). As shown in table 6, the VGGT forward pass dominates total runtime and is identical across variants; minor fluctuations (± 0.03 min) fall within run-to-run variance. Notably, IRLS alignment is faster with our modules (1.05 min → 0.88 min), as SGIA provides tighter scale initialization and TTA improves point cloud quality, both yielding cleaner inlier sets and faster convergence. SGIA adds ∼ 0.3 s/chunk (RANSAC + SVD), totaling ∼ 18 s (∼ 5%); TTA adds ∼ 1.2 s/chunk (3 steps, 3 frames at 336px), totaling ∼ 72 s (∼ 3%). These costs are largely offset by the IRLS speedup, resulting in net overhead below 3%, a modest cost well justified by the 32% ATE reduction. Method VGGT Align SGIA TTA Total fwd (IRLS) VGGT-Long 0.68 min 1.05 min — — ∼ 2.56 min Ours (SGIA) 0.63 min 0.89 min 0.14 min — ∼ 2.31 min Ours (full) 0.62 min 0.88 min 0.13 min 0.07 min ∼ 2.32 min Table 6. Runtime breakdown on the Waymo Open Dataset. See Appendix for detailed analysis. 5. Conclusion We present VGGT-Align, a scale-consistency enhancement framework for chunk-based long-sequence 3D reconstruction. Our core contribution, Scene Geometric Invariant Anchoring (SGIA), exploits cross-chunk consistency of inherent scene geometric quantities to establish scale constraints independent of overlap-based registration, degenerating 7-DoF Sim(3) alignment into 6-DoF rigid-body transformation and severing multiplicative scale drift at its source. Complemented by a lightweight test-time adaptation strategy, VGGT-Align achieves state-of-the-art results on KITTI, Waymo, and Virtual KITTI with less than 8% runtime overhead. Both modules are plug-and-play and require no offline retraining. Acknowledgements. This work was supported by the National Natural Science Foundation of China under Grants 62571437 and 62471394. References (1) Bhat et al. (2023) Shariq Farooq Bhat, Reiner Birkl, Diana Wofk, Peter Wonka, and Matthias Müller. 2023. Zoedepth: Zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288 (2023). Campos et al. (2021) Carlos Campos, Richard Elvira, Juan J Gómez Rodríguez, José M Montiel, and Juan D Tardós. 2021. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE transactions on robotics 37, 6 (2021), 1874–1890. Deng et al. (2025) Kai Deng, Zexin Ti, Jiawei Xu, Jian Yang, and Jin Xie. 2025. VGGT-Long: Chunk it, Loop it, Align it–Pushing VGGT’s Limits on Kilometer-scale Long RGB Sequences. arXiv preprint arXiv:2507.16443 (2025). Fischler and Bolles (1981) Martin A Fischler and Robert C Bolles. 1981. Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography. Commun. ACM 24, 6 (1981), 381–395. Gaidon et al. (2016) Adrien Gaidon, Qiao Wang, Yohann Cabon, and Eleonora Vig. 2016. Virtual Worlds as Proxy for Multi-Object Tracking Analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Geiger et al. (2012) Andreas Geiger, Philip Lenz, and Raquel Urtasun. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition. IEEE, 3354–3361. Godard et al. (2019) Clément Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. 2019. Digging into self-supervised monocular depth estimation. In Proceedings of the IEEE/CVF international conference on computer vision. 3828–3838. Guizilini et al. (2023) Vitor Guizilini, Igor Vasiljevic, Dian Chen, Rare s , Ambru s , , and Adrien Gaidon. 2023. Towards zero-shot scale-aware monocular depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 9233–9243. Hu et al. (2024) Mu Hu, Wei Yin, Chi Zhang, Zhipeng Cai, Xiaoxiao Long, Hao Chen, Kaixuan Wang, Gang Yu, Chunhua Shen, and Shaojie Shen. 2024. Metric3d v2: A versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 46, 12 (2024), 10579–10596. Hu et al. (2025) Yu Hu, Chong Cheng, Sicheng Yu, Xiaoyang Guo, and Hao Wang. 2025. VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction. arXiv preprint arXiv:2511.19971 (2025). Jia et al. (2025a) Yuyu Jia, Jiabo Li, and Qi Wang. 2025a. Generalized Few-Shot Semantic Segmentation for Remote Sensing Images. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1–10. doi:10.1109/TGRS.2025.3531874 Jia et al. (2025b) Yuyu Jia, Chenchen Sun, Junyu Gao, and Qi Wang. 2025b. Few-shot Remote Sensing Scene Classification via Parameter-free Attention and Region Matching. ISPRS Journal of Photogrammetry and Remote Sensing 227 (2025), 265–275. doi:10.1016/j.isprsjprs.2025.05.026 Jia et al. (2025c) Yuyu Jia, Qing Zhou, Junyu Gao, and Qi Wang. 2025c. Entity-Guided Attention Twisting Network for Referring Remote Sensing Image Segmentation. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1–10. doi:10.1109/TGRS.2025.3615765 Kingma and Ba (2014) Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Kuznietsov et al. (2021) Yevhen Kuznietsov, Marc Proesmans, and Luc Van Gool. 2021. Comoda: Continuous monocular depth adaptation using past experiences. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2907–2917. Lee et al. (2025) Jungho Lee, Minhyeok Lee, Sunghun Yang, Minseok Kang, and Sangyoun Lee. 2025. SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes. arXiv preprint arXiv:2511.18290 (2025). Leroy et al. (2024) Vincent Leroy, Yohann Cabon, and Jérôme Revaud. 2024. Grounding image matching in 3d with mast3r. In European conference on computer vision. Springer, 71–91. Lindenberger et al. (2023) Philipp Lindenberger, Paul-Edouard Sarlin, and Marc Pollefeys. 2023. Lightglue: Local feature matching at light speed. In Proceedings of the IEEE/CVF international conference on computer vision. 17627–17638. Lyu et al. (2025a) Guangtao Lyu, Xinyi Cheng, Qi Liu, Chenghao Xu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2025a. COME: Advancing Representation Learning and Generative Modeling for High-Quality Text-to-Motion Generation. (2025). Lyu et al. (2025b) Guangtao Lyu, Chenghao Xu, Qi Liu, Jiexi Yan, Muli Yang, Fen Fang, and Cheng Deng. 2025b. Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation. arXiv preprint arXiv:2512.18804 (2025). Lyu et al. (2025c) Guangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang, and Cheng Deng. 2025c. Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization. In ICLR. Ma et al. ([n. d.]) Wenzong Ma, Zhuoxiao Li, Jinjing Zhu, Tongyan Hua, Kanghao Chen, Zidong Cao, Da Yang, Peilun Shi, Yibo Zhou, Wufan Zhao, et al. [n. d.]. SkyEvents: A Large-Scale Event-enhanced UAV Dataset for Robust 3D Scene Reconstruction. In The Fourteenth International Conference on Learning Representations. Maggio et al. (2025) Dominic Maggio, Hyungtae Lim, and Luca Carlone. 2025. Vggt-slam: Dense rgb slam optimized on the sl (4) manifold. arXiv preprint arXiv:2505.12549 (2025). Murai et al. (2025) Riku Murai, Eric Dexheimer, and Andrew J Davison. 2025. Mast3r-slam: Real-time dense slam with 3d reconstruction priors. In Proceedings of the Computer Vision and Pattern Recognition Conference. 16695–16705. Nistér (2004) David Nistér. 2004. An efficient solution to the five-point relative pose problem. IEEE transactions on pattern analysis and machine intelligence 26, 6 (2004), 756–770. Rublee et al. (2011) Ethan Rublee, Vincent Rabaud, Kurt Konolige, and Gary Bradski. 2011. ORB: An efficient alternative to SIFT or SURF. In 2011 International conference on computer vision. Ieee, 2564–2571. Sarlin et al. (2020) Paul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, and Andrew Rabinovich. 2020. Superglue: Learning feature matching with graph neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4938–4947. Schonberger and Frahm (2016) Johannes L Schonberger and Jan-Michael Frahm. 2016. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition. 4104–4113. Song and Chandraker (2015) Shiyu Song and Manmohan Chandraker. 2015. Joint sfm and detection cues for monocular 3d localization in road scenes. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 3734–3742. Sun et al. (2020) Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aurelien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. 2020. Scalability in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2446–2454. Teed and Deng (2021) Zachary Teed and Jia Deng. 2021. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems 34 (2021), 16558–16569. Teed et al. (2023) Zachary Teed, Lahav Lipson, and Jia Deng. 2023. Deep patch visual odometry. Advances in Neural Information Processing Systems 36 (2023), 39033–39051. Teed et al. (2024) Zachary Teed, Lahav Lipson, and Jia Deng. 2024. Deep patch visual odometry. Advances in Neural Information Processing Systems 36 (2024). Tonioni et al. (2019a) Alessio Tonioni, Matteo Poggi, Stefano Mattoccia, and Luigi Di Stefano. 2019a. Unsupervised domain adaptation for depth prediction from images. IEEE transactions on pattern analysis and machine intelligence 42, 10 (2019), 2396–2409. Tonioni et al. (2019b) Alessio Tonioni, Fabio Tosi, Matteo Poggi, Stefano Mattoccia, and Luigi DI Stefano. 2019b. Real-time self-adaptive deep stereo. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 195–204. Triggs et al. (1999) Bill Triggs, Philip F McLauchlan, Richard I Hartley, and Andrew W Fitzgibbon. 1999. Bundle adjustment—a modern synthesis. In International workshop on vision algorithms. Springer, 298–372. Vaswani et al. (2017) Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017). Wang et al. (2020) Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell. 2020. Tent: Fully test-time adaptation by entropy minimization. arXiv preprint arXiv:2006.10726 (2020). Wang et al. (2025a) Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, and David Novotny. 2025a. Vggt: Visual geometry grounded transformer. In Proceedings of the Computer Vision and Pattern Recognition Conference. 5294–5306. Wang et al. (2024a) Jianyuan Wang, Nikita Karaev, Christian Rupprecht, and David Novotny. 2024a. Vggsfm: Visual geometry grounded deep structure from motion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 21686–21697. Wang et al. (2025b) Qianqian Wang, Yifei Zhang, Aleksander Holynski, Alexei A Efros, and Angjoo Kanazawa. 2025b. Continuous 3d perception model with persistent state. In Proceedings of the Computer Vision and Pattern Recognition Conference. 10510–10522. Wang et al. (2024b) Shuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii, and Jerome Revaud. 2024b. Dust3r: Geometric 3d vision made easy. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 20697–20709. Xiao et al. (2026a) Xi Xiao, Xingjian Li, Yunbei Zhang, Cheng Han, Tianming Liu, Tianyang Wang, Runmin Jiang, Jihun Hamm, Xiao Wang, and Min Xu. 2026a. Layer-Specific Prompt Fusion Discovery via Differentiable Search in Vision Foundation Models. arXiv preprint arXiv:2606.26379 (2026). Xiao et al. (2026b) Xi Xiao, Chen Liu, Chih-Ting Liao, Yunbei Zhang, Qizhen Lan, Yuxiang Wei, Lin Zhao, Janet Wang, Jianyang Gu, Muchao Ye, et al. 2026b. Staying VIGILant: Mitigating Visual Laziness via Counterfactual Visual Alignment in MLLMs. arXiv preprint arXiv:2606.26387 (2026). Xiao et al. (2026c) Xi Xiao, Chenrui Ma, Yunbei Zhang, Chen Liu, Zhuxuanzi Wang, Yanshu Li, Lin Zhao, Guosheng Hu, Tianyang Wang, and Hao Xu. 2026c. Not all directions matter: Towards structured and task-aware low-rank model adaptation. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2132–2154. Xiao et al. (2025) Xi Xiao, Yunbei Zhang, Xingjian Li, Tianyang Wang, Xiao Wang, Yuxiang Wei, Jihun Hamm, and Min Xu. 2025. Visual instance-aware prompt tuning. In Proceedings of the 33rd ACM International Conference on Multimedia. 2880–2889. Xiong et al. (2026) Zhuang Xiong, Chen Zhang, Qingshan Xu, and Wenbing Tao. 2026. VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency. arXiv preprint arXiv:2602.05508 (2026). Yang et al. (2025) Jianing Yang, Alexander Sax, Kevin J Liang, Mikael Henaff, Hao Tang, Ang Cao, Joyce Chai, Franziska Meier, and Matt Feiszli. 2025. Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass. In Proceedings of the Computer Vision and Pattern Recognition Conference. 21924–21935. Yang and Scherer (2019) Shichao Yang and Sebastian Scherer. 2019. Monocular object and plane slam in structured environments. IEEE Robotics and Automation Letters 4, 4 (2019), 3145–3152. Yin et al. (2023) Wei Yin, Chi Zhang, Hao Chen, Zhipeng Cai, Gang Yu, Kaixuan Wang, Xiaozhi Chen, and Chunhua Shen. 2023. Metric3d: Towards zero-shot metric 3d prediction from a single image. In Proceedings of the IEEE/CVF international conference on computer vision. 9043–9053. Zhang et al. (2025a) Wei Zhang, Qiang Li, and Qi Wang. 2025a. Refined Cascade Cost Volume for Multiview Remote Sensing Image Reconstruction. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1–11. doi:10.1109/TGRS.2025.3595544 Zhang et al. (2024) Wei Zhang, Qiang Li, Yuan Yuan, and Qi Wang. 2024. Visual Consistency Enhancement for Multiview Stereo Reconstruction in Remote Sensing. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–11. doi:10.1109/TGRS.2024.3482697 Zhang et al. (2026a) Wei Zhang, Songhua Li, Yihang Wu, Qiang Li, and Qi Wang. 2026a. VGGT-CD: Training-Free Robust Registration for 3D Change Detection. arXiv:2605.16859 [cs.CV] https://arxiv.org/abs/2605.16859 Zhang et al. (2026b) Wei Zhang, Yihang Wu, Shengkai Yu, Songhua Li, Qiang Li, and Qi Wang. 2026b. GPR-MVS: Global Propagation Regularization for Large Scale Multi-view Stereo. IEEE Transactions on Geoscience and Remote Sensing (2026), 1–1. doi:10.1109/TGRS.2026.3710063 Zhang et al. (2025b) Wei Zhang, Zhigang Yang, Qiang Li, and Qi Wang. 2025b. Semantic-Guided Multiview Stereo Reconstruction for Aerial Image. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1–11. doi:10.1109/TGRS.2025.3585623 Zhao et al. (2026) Bingxuan Zhao, Qing Zhou, Chuang Yang, and Qi Wang. 2026. SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis. arXiv preprint arXiv:2603.21783 (2026). Zhou et al. (2019) Dingfu Zhou, Yuchao Dai, and Hongdong Li. 2019. Ground-plane-based absolute scale estimation for monocular visual odometry. IEEE Transactions on Intelligent Transportation Systems 21, 2 (2019), 791–802. Zhuo et al. (2025) Dong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu, Jie Zhou, and Jiwen Lu. 2025. Streaming 4d visual geometry transformer. arXiv preprint arXiv:2507.11539 (2025).