Paper deep dive
PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
Zhiyuan Yuan, Guanying Chen, Lingteng Qiu, Ruimao Zhang, Shuguang Cui, Xiaochun Cao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/19/2026, 4:05:11 AM
Summary
The paper introduces PXDepth, a discriminative monocular depth estimation model designed to preserve fine-grained structures and sharp boundaries. It addresses the limitation of existing models that combine large-patch ViT encoders with convolutional decoders, which lose pixel-level cues. PXDepth separates global context modeling (using a large-patch ViT) from pixel-level depth prediction (using Context-Modulated Pixel Transformer blocks with coarse-to-fine pixel compaction), achieving high structural fidelity and efficiency.
Entities (11)
Relation Signals (8)
PXDepth → usescomponent → Pixel-Space Depth Predictor
confidence 95% · PXDepth consists of a Global Context Encoder and a Pixel-Space Depth Predictor.
PXDepth → usescomponent → Global Context Encoder
confidence 95% · PXDepth consists of a Global Context Encoder and a Pixel-Space Depth Predictor.
PXDepth → isevaluatedon → KITTI
confidence 90% · We conduct zero-shot evaluation on the MoGe Benchmark... includes ... KITTI
PXDepth → isevaluatedon → NYUv2
confidence 90% · We conduct zero-shot evaluation on the MoGe Benchmark... includes NYUv2
Global Context Encoder → isimplementedas → ViT
confidence 90% · The Global Context Encoder is implemented with a large-patch ViT
Pixel-Space Depth Predictor → usescomponent → Context-Modulated Pixel Transformer
confidence 90% · The Pixel-Space Depth Predictor refines pixel-space features through its Context-Modulated Pixel Transformer (CM-PiT) blocks.
PXDepth → isinitializedwith → MoGe-2
confidence 85% · model weights initialized from MoGe-2.
PXDepth → istrainedon → Hypersim
confidence 85% · We train on multiple synthetic RGB-D datasets, including Hypersim
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.16984v1
- Canonical: https://arxiv.org/abs/2608.16984v1
Trouble viewing inline? Open PDF directly →
Full Text
60,364 characters extracted from source content.
Expand or collapse full text
PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation Zhiyuan Yuan Guanying Chen Lingteng Qiu Affiliation: CUHKSZ Ruimao Zhang [0.25em] Shuguang Cui Affiliation: Shenzhen-FNii Affiliation: CUHKSZ Xiaochun Cao [0.6em] Sun Yat-sen University Abstract Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at https://yuanzhy29.github.io/PXDepth-Page/. **footnotetext: Corresponding author: chenguanying@mail.sysu.edu.cn Figure 1: We present PXDepth, a discriminative monocular depth estimation model that performs pixel-space modeling for structure-preserving depth prediction. Compared with structure-aware methods (63; 5; 58), PXDepth better preserves local structures without sacrificing global accuracy, while requiring less inference time. 1 Introduction Monocular depth estimation (MDE) is a fundamental geometric task that aims to recover depth from a single image. Driven by ViT backbones and large-scale training data, recent depth estimation models (41; 4; 6; 60; 61; 53; 54) have achieved strong robustness and zero-shot generalization. Although their predictions are often visually plausible in 2D, they can still contain structural distortions that become more apparent when converted into point clouds. These distortions degrade geometric fidelity and limit practical applications in 3D reconstruction, free-viewpoint rendering, robotic manipulation, and immersive content creation. Existing MDE methods can be broadly divided into discriminative and generative paradigms. Discriminative models typically pair a large-patch ViT encoder (11) with a convolution-based upsampling decoder (40; 6; 60; 61; 53; 54). The encoder represents the input as a low-resolution token grid and applies global self-attention to capture rich semantics and scene layout. However, this tokenization can weaken pixel-level cues and high-frequency details during encoding. Once these details are lost at an early stage, convolution-based upsampling alone may not fully recover them. To alleviate this limitation, InfiniDepth (63) improves depth details with implicit representations, but it still reconstructs depth by interpolating low-resolution token features, leaving structural distortions unresolved. In contrast, generative methods such as Pixel-Perfect Depth (PPD) (58) perform diffusion directly in pixel space and use semantic guidance with a cascade Diffusion Transformer (DiT) design to produce depth predictions with finer details and better structural fidelity. However, multi-step denoising requires repeated network evaluations, leading to inefficient inference. The success of pixel-space generative modeling (30; 9; 64; 59) suggests that maintaining pixel-space representations is important for recovering fine-grained geometry. However, the pixel-space diffusion depth formulation remains computationally expensive. Motivated by this observation, we rethink the design of discriminative depth models. We retain a large-patch ViT to provide global context, while introducing a prediction network that models depth in pixel space directly. This separation preserves high-frequency cues throughout depth prediction. We propose PXDepth, a discriminative architecture that directly models depth in pixel space to recover fine-grained details and improve structural fidelity. PXDepth consists of a Global Context Encoder and a Pixel-Space Depth Predictor . The Global Context Encoder provides global context, while the Pixel-Space Depth Predictor refines pixel-space features through its Context-Modulated Pixel Transformer (CM-PiT) blocks. Since performing self-attention over the full-resolution pixel grid is computationally expensive, we adopt a pixel-token compaction mechanism (64) to enable efficient feature interaction. We further adopt a coarse-to-fine pixel compaction design to balance computational efficiency with detailed geometry. Experiments show that PXDepth produces depth predictions with finer details and better structural fidelity than strong discriminative baselines, while remaining more efficient than multi-step generative methods (see Figure 1). In summary, our main contributions are as follows: • We identify a key limitation of mainstream discriminative MDE models. Although large-patch ViT encoders capture strong semantic and scene-level context, their coarse tokenization weakens fine-grained spatial cues that are difficult to recover through convolutional upsampling. • We propose PXDepth, a discriminative depth model that integrates a Global Context Encoder with a Pixel-Space Depth Predictor to preserve global structure and recover fine-grained structural details. • Extensive experiments across multiple benchmarks demonstrate that PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. 2 Related Work Monocular Depth Estimation. Early monocular depth estimation relied on hand-crafted monocular cues and graphical models to infer scene layout from a single image (19; 43). Deep learning later turned MDE into a supervised dense prediction task, with CNN-based models such as multi-scale prediction (12), fully convolutional residual networks (28), ordinal regression (13), and local planar guidance (29). Later relative-depth methods focused more on cross-dataset generalization. MiDaS (41) showed that mixing heterogeneous datasets improves zero-shot transfer, while DPT (40) introduced a strong transformer-based dense prediction framework. Recent foundation models, such as Depth Anything (60; 61), further improve open-domain robustness with large-scale training. In parallel, metric-depth methods (2; 4; 62; 20; 39; 38; 6; 54) aim to recover depth with absolute scale and have steadily improved metric accuracy and generalization. Despite these advances, recent discriminative models often use large-patch ViT encoders with convolutional upsampling decoders. Such early tokenization weakens pixel-level cues and high-frequency details, making detailed geometry difficult to recover during decoding. Depth Estimation with Detailed Geometry. Recent methods have increasingly focused on improving depth details and local geometry. Several discriminative methods enhance details through local refinement, high-resolution prediction, or multi-scale decoding (3; 31; 32). DepthPro (6) improves sharp metric depth with a high-resolution design and boundary-aware training. MoGe (53) improves local geometry by predicting affine-invariant point maps with global and local geometry supervision, and MoGe-2 (54) further improves detail preservation with refined training data. MDA (5) models per-pixel depth ambiguity with a mixture-density representation to produce flying-point-free point maps. InfiniDepth (63) represents depth as a neural implicit field and uses a local implicit decoder, enabling arbitrary-resolution querying and sharper depth details. These methods can produce depth maps with sharp details, but structural distortions may still appear after back-projection into point clouds, limiting applications that rely on accurate 3D geometry. Concurrent with our work, SurGe (25) attributes local surface distortions to the difficulty of fixed convolutional kernels in reconstructing high-frequency, high-amplitude signals and introduces a Neighborhood Attention Decoder for adaptive local feature mixing. MoGe-3 (27) recovers fine geometric structures through Self-Guided Sparse Volumetric Refinement, which repeatedly re-voxelizes and refines a coarse point map in sparse 3D space but requires multiple refinement steps at inference. Generative methods provide another direction for detailed depth estimation. Marigold (23) introduces diffusion priors for monocular depth by adapting image diffusion models. Lotus-2 (18; 17) further adapts generative priors to dense geometry with a deterministic two-stage framework. Pixel-Perfect Depth (58) performs diffusion directly in pixel space and uses semantic guidance with a cascade DiT design, improving both depth details and structural details in point clouds. These methods show the importance of pixel-space modeling for detailed geometry. However, diffusion-based methods require iterative inference, and latent-space generative models may suffer from compression-induced artifacts around edges and details. In contrast, our method retains the discriminative formulation while predicting depth directly in pixel space conditioned on global context from the Global Context Encoder, improving depth details and point-cloud structures with efficient inference (see Figure 2). 3 Motivation (a) Discriminative Architecture (b) PPD (c) Ours Figure 2: Architecture comparison of monocular depth estimators. Conventional discriminative models decode depth from low-resolution ViT features using a convolutional decoder (40), PPD performs iterative pixel-space denoising 58, and PXDepth predicts depth with a Pixel-Space Depth Predictor conditioned on global context in a single feed-forward pass. Overview. Given an input image I∈ℝH×W×3I ^H× W× 3, the MDE task is to predict a pixel-aligned depth map D^∈ℝH×W D ^H× W. D^=f(I). D=f(I). (1) Most discriminative MDE methods pair a large-patch ViT encoder (11; 10; 47) with a convolutional decoder paradigm (60; 61; 53; 54; 40). Specifically, a large-patch ViT encoder ℰvitE_ vit first tokenizes the input image with non-overlapping patches of size P, and then performs self-attention to produce low-resolution features FlrF_ lr. A convolution-based decoder convD_ conv then upsamples these features to recover the pixel-aligned depth map. Flr F_ lr =ℰvit(I), =E_ vit(I), Flr F_ lr ∈ℝHP×WP×dlr, HP× WP× d_ lr, (2) D D =conv(Flr), =D_ conv(F_ lr), D D ∈ℝH×W. ^H× W. This architecture is effective in capturing global structure through self-attention. However, high-frequency image cues are inevitably weakened during large-patch tokenization. Moreover, the decoder (40) mainly relies on convolutional upsampling with spatially shared kernels, making it difficult to handle low-frequency regions and high-frequency details simultaneously (25). As a result, the predicted depth may exhibit sharp details visually, while still containing structural distortions. Generative models offer another direction for detailed depth estimation. Many diffusion-based methods build on pretrained latent diffusion models (23) and benefit from strong generative priors. However, they typically compress depth maps into a variational autoencoder (VAE) latent space, which can weaken sharp depth discontinuities and structural fidelity. Main Idea. PPD (58) avoids VAE-based (24; 23) latent compression by performing diffusion directly in pixel space. It further uses semantic guidance and a cascade DiT (37) design to preserve high-frequency information. Motivated by this separation of global context encoding and pixel-space modeling (9; 34; 55; 64), we introduce a feed-forward discriminative framework for depth prediction. Formally, our model is written as follows. Fctx F_ ctx =ℰctx(I), =E_ ctx(I), Fctx F_ ctx ∈ℝHP×WP×dctx, HP× WP× d_ ctx, (3) D D =pix(I,Fctx), =P_ pix(I,F_ ctx), D D ∈ℝH×W. ^H× W. Here, ℰctxE_ ctx and pixP_ pix denote global-context encoding and pixel-space prediction, respectively. Unlike conventional decoders that reconstruct depth solely from low-resolution encoder features, pixP_ pix operates directly on the input image under the guidance of FctxF_ ctx, preserving high-frequency cues while maintaining global geometric consistency. 4 Method Figure 3: Overview of the proposed PXDepth. PXDepth combines a Global Context Encoder with a Pixel-Space Depth Predictor built from CM-PiT blocks with a coarse-to-fine pixel compaction design. Context features FctxF_ ctx generate pixel-wise modulation parameters for CM-PiT blocks. The Pixel-Space Depth Predictor maintains pixel-space features and changes the compaction size from P to P/2P/2 for coarse-to-fine refinement. As illustrated in Figure 3, PXDepth consists of a Global Context Encoder for extracting global context and a Pixel-Space Depth Predictor for predicting depth in pixel space. The Pixel-Space Depth Predictor maintains full-resolution features, refines them through Context-Modulated Pixel Transformer (CM-PiT) blocks conditioned on the global context, and maps the refined features to a depth map with a lightweight depth head. 4.1 Global Context Encoder The Global Context Encoder is implemented with a large-patch ViT to provide global scene context for the Pixel-Space Depth Predictor. Given the input image I, the encoder tokenizes it into non-overlapping patches with patch size P, producing a low-resolution token grid of size HP×WP HP× WP. Self-attention is then applied over this token grid to aggregate global information. Fctx=ℰctx(I),Fctx∈ℝHP×WP×dctx,F_ ctx=E_ ctx(I), F_ ctx HP× WP× d_ ctx, (4) where ℰctxE_ ctx denotes the Global Context Encoder, FctxF_ ctx is the context features, and dctxd_ ctx is the feature dimension. The Global Context Encoder can also inherit strong priors from pretrained vision foundation models, such as DINOv2 (10), DINOv3 (47), MoGe-2 (54). 4.2 Pixel-Space Depth Predictor Pixel Embedding. Given the input image I, the pixel embedding layer ϕin _ in maps each RGB pixel to a dpixd_ pix-dimensional feature. Fpix0=ϕin(I),Fpix0∈ℝH×W×dpix,F_ pix^0= _ in(I), F_ pix^0 ^H× W× d_ pix, (5) where ϕin _ in is implemented by a 1×11× 1 convolution and Fpix0F_ pix^0 is the initial pixel-space features. Unlike conventional decoders that reconstruct depth from low-resolution ViT features alone, this embedding establishes a pixel-space feature stream whose spatial resolution is preserved throughout the Pixel-Space Depth Predictor. Pixel-Space Feature Modeling. To model pixel-space features for fine-grained depth estimation, we process the initial pixel features Fpix0F_ pix^0 with a sequence of N CM-PiT blocks guided by the context feature FctxF_ ctx. Fpixi+1=CM-PiT i(Fpixi,Fctx),i=0,…,N−1.F_ pix^i+1=CM-PiT ^i(F_ pix^i,F_ ctx), i=0,…,N-1. (6) Here, FpixiF_ pix^i denotes the pixel features before the i-th block. Depth Head. A linear layer maps the refined pixel features to the depth map. D^=ϕout(FpixN),D^∈ℝH×W, D= _ out(F_ pix^N), D ^H× W, (7) where ϕout _ out is implemented by a 1×11× 1 convolution. 4.3 Context-Modulated Pixel Transformer (CM-PiT) Block Each CM-PiT block processes pixel-space features by conditioning its attention and feed-forward network on FctxF_ ctx. We first compare alternative guidance strategies, then present the adopted Context-Guided Adaptive Normalization. Alternative Guidance Strategies. Figure 4 compares three strategies for incorporating FctxF_ ctx into pixel-space depth prediction. Direct addition merges FctxF_ ctx with compacted pixel tokens before self-attention, while cross-attention introduces an additional attention layer to learn the correspondence between the two representations. In contrast, Context-Guided Adaptive Normalization, originally introduced in PixelDiT (64), projects FctxF_ ctx into pixel-wise scale, shift, and residual gates, allowing global context to condition pixel-space features by controlling the normalization and residual updates in each CM-PiT block. We ultimately adopt this design as it provides the best balance between performance and efficiency. Unlike cross-attention, which must learn the correspondence between FctxF_ ctx and FpixF_ pix implicitly, its pixel-wise modulation preserves this correspondence explicitly without introducing an additional attention layer. Figure 4: Alternative guidance strategies. Feature flow within the candidate CM-PiT designs. Context-Guided Adaptive Normalization. As shown in Figure 3, FctxF_ ctx generates adaptive normalization parameters for the CM-PiT block. We introduce a slices of learned projection α(⋅)α(·), β(⋅)β(·), and γ(⋅)γ(·). For each context token, the projection produces p2dpixp^2d_ pix values, which are then reshaped into a p×p× p pixel-aligned map and provides the residual gate, shift, and scale for the attention layer. We formulate it as F~pixi F_ pix^i =RMSNorm(Fpixi)⊙γ(Fctx)+β(Fctx), =RMSNorm(F_ pix^i) γ(F_ ctx)+β(F_ ctx), (8) where F~pixi F_ pix^i denotes the intermediate features. To avoid applying self-attention directly over the full-resolution H×WH× W pixel grid, we adopt the pixel token compaction mechanism (64) with the operators p(⋅)C_p(·) and p(⋅)U_p(·), where pC_p compacts the p2p^2 features in each p×p× p local region into one token, and pU_p expands each attended token back to its original pixel region, as illustrated in Figure 3. F¯pixi F_ pix^i =Fpixi+α(Fctx)⊙p(Attn(p(F~pixi),RoPE)), =F_ pix^i+α(F_ ctx) _p (Attn (C_p( F_ pix^i),RoPE ) ), (9) The attention sequence length is therefore reduced from HWHW to L=(H/p)(W/p)L=(H/p)(W/p), giving a p2p^2-fold reduction. Since compaction is temporary and expansion occurs before the residual update, the network maintains pixel-space features throughout prediction. F¯pixi F_ pix^i is then fed into the feed-forward network which follows the same modulation form, and output Fpixi+1F_ pix^i+1 which will be fed into the next CM-PiT block. Coarse-to-Fine Pixel Compaction. To progressively refine depth from coarse structures to fine details, we adopt the coarse-to-fine pixel compaction design illustrated in Figure 3. The early CM-PiT blocks use a compaction size of p=Pp=P, matching the patch size of the Global Context Encoder. This reduces the number of tokens and enables efficient coarse-structure modeling. The later blocks reduce the compaction size to p=P/2p=P/2, increasing the token density for fine-detail refinement. This coarse-to-fine design balances computational efficiency with geometric fidelity. 5 Experiments Figure 5: Visual comparison on diverse scenes. PXDepth better preserves local structures. 5.1 Implementation Details Training Datasets. We train on multiple synthetic RGB-D datasets, including Hypersim (42), MVS-Synth (21), TartanAir (56), TartanGround (36), UnrealStereo4K (49), GTA-SfM (52), Structured3D (65), UrbanSyn (15), ParallelDomain-4D (50), VKITTI2 (8), SceneNetRGBD (35), RobbySim (48), and Synscapes (57). Model Configuration. We use ViT-L/14 as the Global Context Encoder, with P=14P=14 and dctx=1024d_ ctx=1024 and model weights initialized from MoGe-2. The Pixel-Space Depth Predictor contains N=8N=8 CM-PiT blocks with a pixel feature dimension of dpix=16d_ pix=16. In our coarse-to-fine pixel compaction design, the first 44 blocks use a compaction size of p=14p=14, while the remaining 44 blocks use p=7p=7. We additionally introduce a mask head for valid mask prediction (53; 54). Training Configuration. We train the model in two stages, with 500500K iterations at 518×518518× 518, followed by 200200K iterations at varying resolutions with the image area kept equivalent to 1022×7701022× 770. Optimization is performed with AdamW, and the objective combines normalized depth loss, multi-scale depth-gradient loss, and binary mask loss. Additional optimization and loss details are provided in Appendix A. 5.2 Evaluation Setup and Metrics Benchmarks and Baselines. We conduct zero-shot evaluation on the MoGe Benchmark (53; 54) and the MDA Benchmark (5). The MoGe Benchmark includes NYUv2 (46), KITTI (14), ETH3D (44), iBims-1 (26), Sintel (7), DDAD (16), DIODE (51), and HAMMER (22). The MDA Benchmark consists of NRGBD (1), 7Scenes (45), and HiRoom (33). We compare our method with six recent approaches, including Depth Anything V2 (DA V2) (61), DepthPro (6), InfiniDepth (63), MoGe-2 (54), PPD (58), and MDA (5). Evaluation Metrics. Following the standard affine-invariant depth evaluation protocol, we align each prediction to the ground-truth depth by a scale and shift, and report the widely used relative error (Rel↓ ) and δ1 _1 accuracy (δ1↑ _1 ). Both metrics are reported as percentages. Boundary metrics assess the structural accuracy and sharpness of predicted depth maps. Following PPD (58) and MDA (5), we extract boundaries from normalized ground-truth depth using a Canny detector and remove edges adjacent to invalid regions. Predicted and ground-truth depths at the retained boundary pixels are back-projected into 3D, and the predicted boundary point cloud is aligned to the ground truth using point-to-point iterative closest point (ICP) registration. We report accuracy (Acc) and Chamfer distance (CD) in millimeters, where lower values indicate better boundary geometry. 5.3 Evaluation Results Global Depth Accuracy. As shown in Tables 1 & 2 (a), our method achieves better global depth accuracy than structure-aware methods, while remaining comparable to recent methods that focus on global depth prediction. The remaining performance gap with MoGe-2 (54) may be attributed to its large-scale training on a substantially larger and more diverse collection of data. Table 1: Evaluation on the MoGe Benchmark. Zero-shot depth estimation across eight datasets. Category Method NYUv2 KITTI ETH3D iBims-1 Sintel DDAD DIODE HAMMER Mean Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Global Prediction DA V2 4.4 98.0 6.9 95.7 4.6 97.9 3.5 98.3 21.2 72.8 13.6 86.0 5.0 96.4 4.9 99.1 8.0 93.0 DepthPro 3.8 98.0 5.7 96.1 5.1 96.2 3.0 98.7 16.0 79.9 12.7 83.5 4.2 96.7 3.3 99.4 6.7 93.6 MoGe-2 3.1 98.4 4.2 97.9 2.9 99.0 2.4 98.7 14.0 81.4 8.5 92.1 2.9 98.0 3.0 99.3 5.1 95.6 InfiniDepth 4.6 97.8 5.6 96.6 4.2 98.1 3.5 98.4 19.3 79.5 12.6 86.5 4.6 96.9 3.0 99.2 7.2 94.1 PPD 4.1 97.8 6.9 94.6 4.9 97.1 3.4 98.3 18.5 77.1 13.8 81.3 4.9 95.3 3.1 99.3 7.5 92.6 MDA 3.6 97.6 7.0 94.8 5.4 94.5 3.1 98.0 17.4 77.7 21.8 67.9 4.7 95.2 2.6 98.7 8.2 90.5 Structure Aware Ours 3.5 98.0 5.2 96.8 2.9 98.6 2.7 98.5 14.3 82.3 11.5 88.5 3.6 96.7 2.4 99.4 5.8 94.8 Table 2: Quantitative comparison on the MDA Benchmark. Global accuracy is reported using Rel and δ1 _1, while boundary fidelity is reported using Acc and CD in millimeters. Per-image inference time is measured at 518×518518× 518 on an RTX 5880 GPU. Category Method (a) Global Accuracy (b) Boundary Fidelity Time↓ (ms) NRGBD 7Scenes HiRoom Mean NRGBD 7Scenes HiRoom Mean Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Acc↓ CD↓ Acc↓ CD↓ Acc↓ CD↓ Acc↓ CD↓ Global Prediction DA V2 2.6 98.9 7.7 92.6 3.7 98.1 4.7 96.5 118.8 112.1 97.5 115.1 115.0 111.1 110.5 112.8 55.3 DepthPro 2.2 99.0 7.0 92.8 2.5 99.2 3.9 97.0 140.2 131.5 86.2 97.3 101.3 90.4 109.2 106.4 206.8 MoGe-2 2.2 98.6 6.4 93.0 1.9 99.1 3.5 96.9 136.1 128.4 80.3 90.8 88.0 81.6 101.5 100.3 32.9 InfiniDepth 2.9 98.7 7.6 92.7 3.6 98.5 4.7 96.7 85.9 88.5 100.1 117.7 92.1 90.5 92.7 98.9 65.7 PPD 2.5 98.7 7.1 92.4 3.6 98.3 4.4 96.5 74.9 85.9 92.1 103.6 84.0 91.5 83.6 93.7 209.8 MDA 2.0 98.4 6.8 92.7 2.0 99.0 3.6 96.7 50.9 67.6 79.5 91.3 49.1 59.3 59.8 72.8 267.1 Structure Aware Ours 2.0 98.8 6.9 92.7 1.8 99.4 3.6 97.0 60.2 64.7 83.0 92.0 56.1 54.9 66.5 70.6 56.4 Boundary Fidelity. As shown in Table 2 (b), our method achieves the best mean CD and the second-best mean Acc. These metrics jointly measure the structural accuracy of the reconstructed edges and the number of flying points near object boundaries. DA V2, DepthPro, InfiniDepth, and MoGe-2 recover depth from low-resolution features through upsampling, which often introduces flying points and distorted structures. PPD preserves relatively sharp boundaries, but noise near object boundaries. MDA achieves comparable performance because it produces fewer flying points near boundaries but exhibits more distorted structure. In contrast, our method preserves accurate boundary geometry leading to the best boundary fidelity, as shown in Figure 5. Inference Efficiency. Table 2 compares per-image inference time at 518×518518× 518. Our method is slower than MoGe-2, comparable to DA V2 and InfiniDepth, and approximately 3.7×3.7× faster than PPD. This efficiency advantage over generative methods follows from predicting depth in a single feed-forward pass rather than through iterative denoising. 5.4 Ablation and Analysis All ablation variants are trained on Hypersim (42) and UrbanSyn (15) under the same training protocol and evaluated zero-shot on the HiRoom dataset (33). Table 3: Quantitative ablation on model components. Acc and CD are reported in millimeters. Ablation Variant Rel↓ δ1↑ _1 Acc↓ CD↓ Time (ms)↓ w/o Global Context Encoder 17.87 72.87 261.0 385.6 33.5 DINOv3 Initialization 3.02 98.73 74.8 76.9 61.4 (a) Global Context Encoder MoGe-2 Initialization 2.74 99.01 70.3 72.5 56.4 w/o coarse-to-fine 2.72 98.98 74.9 76.8 51.6 (b) Coarse-to-Fine Pixel Compaction Ours 2.74 99.01 70.3 72.5 56.4 Addition 3.19 98.91 79.3 80.2 52.4 Cross-Attn 2.98 98.88 79.7 80.9 61.0 (c) Guidance Strategy Adaptive Norm 2.74 99.01 70.3 72.5 56.4 Figure 6: Qualitative ablation on model components. The panels show (a) our full model, (b) without the Global Context Encoder, (c) without coarse-to-fine pixel compaction, (d) addition guidance, and (e) cross-attention guidance. Effect of Global Context Encoder. Table 6 (a) shows that removing the Global Context Encoder substantially degrades both global and boundary metrics. MoGe-2 initialization further improves these metrics over DINOv3, suggesting that geometry-aware initialization provides a stronger prior for depth prediction. Figure 6 (a & b) shows that removing the Global Context Encoder causes the model to struggle with capturing global structures. Effect of Coarse-to-Fine Pixel Compaction. As shown in Table 6 (b), coarse-to-fine compaction improves boundary quality while maintaining comparable global depth accuracy. The comparison in Figure 6 (a & c) further shows that the coarse-to-fine design preserves the local structure. These results indicate that reducing the compaction size in later blocks helps refine local geometric details. Effect of Context Guidance Strategy. The guidance-strategy ablation compares the three designs introduced in Figure 4. As shown in Table 6 (c) and Figure 6 (a, d & e), the aopted Adaptive Normalization achieves better global depth accuracy and boundary quality than direct addition and cross-attention, with only a modest increase in inference time over direct addition. 6 Conclusion In this paper, we highlight a potential limitation of the common discriminative monocular depth estimation architecture. Large-patch ViT encoders capture global semantics, yet their coarse tokenization can weaken fine-grained spatial cues that convolutional upsampling may not fully recover. To address this limitation, we present PXDepth, which combines a Global Context Encoder with a Pixel-Space Depth Predictor. The former captures global structure, while the latter preserves fine-grained geometry through CM-PiT blocks and coarse-to-fine pixel compaction. Extensive experiments demonstrate that PXDepth substantially improves local geometric fidelity while preserving competitive global depth accuracy. Its single-pass design is also substantially more efficient than multi-step generative methods. Limitations and future work. One limitation of our approach is that depth annotations for transparent objects can be ambiguous or unreliable, which may lead to inaccurate predictions on transparent and reflective surfaces. Another limitation is that our relative-depth formulation cannot recover metric scale without an external reference. In future work, we aim to improve supervision for challenging materials and incorporate metric cues to enable scale-aware depth estimation. References Azinović et al. (2022) D. Azinović, R. Martin-Brualla, D. B. Goldman, M. Nießner, and J. Thies Neural rgb-d surface reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 6290–6301. Cited by: Table 5, §5.2. Bhat et al. (2021) S. F. Bhat, I. Alhashim, and P. Wonka Adabins: depth estimation using adaptive bins. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 4009–4018. Cited by: §2. Bhat et al. (2022) S. F. Bhat, I. Alhashim, and P. Wonka Localbins: improving depth estimation by learning local distributions. In European Conference on Computer Vision, p. 480–496. Cited by: §2. Bhat et al. (2023) S. F. Bhat, R. Birkl, D. Wofk, P. Wonka, and M. Müller Zoedepth: zero-shot transfer by combining relative and metric depth. arXiv preprint arXiv:2302.12288. Cited by: §1, §2. Bian et al. (2026) S. Bian, C. Xu, and J. Gao Modeling depth ambiguity: a mixture-density representation for flying-point-free depth estimation. External Links: 2606.02552, Link Cited by: §C.1, Figure 1, §2, §5.2, §5.2. Bochkovskii et al. (2024) A. Bochkovskii, A. Delaunoy, H. Germain, M. Santos, Y. Zhou, S. R. Richter, and V. Koltun Depth pro: sharp monocular metric depth in less than a second. arXiv preprint arXiv:2410.02073. Cited by: §C.1, §1, §1, §2, §2, §5.2. Butler et al. (2012) D. J. Butler, J. Wulff, G. B. Stanley, and M. J. Black A naturalistic open source movie for optical flow evaluation. In European Conf. on Computer Vision (ECCV), A. Fitzgibbon et al. (Eds.) (Ed.), Part IV, LNCS 7577, p. 611–625. Cited by: Table 5, §5.2. Cabon et al. (2020) Y. Cabon, N. Murray, and M. Humenberger Virtual kitti 2. External Links: 2001.10773 Cited by: Table 4, §5.1. Chen et al. (2026) Z. Chen, J. Zhu, X. Chen, J. Zhang, X. Hu, H. Zhao, C. Wang, J. Yang, and Y. Tai Dip: taming diffusion models in pixel space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 36136–36146. Cited by: §1, §3. Darcet et al. (2023) T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski Vision transformers need registers. Cited by: §3, §4.1. Dosovitskiy et al. (2020) A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §1, §3. Eigen et al. (2014) D. Eigen, C. Puhrsch, and R. Fergus Depth map prediction from a single image using a multi-scale deep network. Advances in neural information processing systems 27. Cited by: §2. Fu et al. (2018) H. Fu, M. Gong, C. Wang, K. Batmanghelich, and D. Tao Deep ordinal regression network for monocular depth estimation. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2002–2011. Cited by: §2. Geiger et al. (2012) A. Geiger, P. Lenz, and R. Urtasun Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, p. 3354–3361. Cited by: Table 5, §5.2. Gómez et al. (2025) J. L. Gómez, M. Silva, A. Seoane, A. Borràs, M. Noriega, G. Ros, J. A. Iglesias-Guitian, and A. M. López All for one, and one for all: urbansyn dataset, the third musketeer of synthetic driving scenes. Neurocomputing 637, p. 130038. External Links: ISSN 0925-2312, Document, Link Cited by: Table 4, §5.1, §5.4. Guizilini et al. (2020) V. Guizilini, R. Ambrus, S. Pillai, A. Raventos, and A. Gaidon 3D packing for self-supervised monocular depth estimation. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 5, §5.2. He et al. (2025) J. He, H. Li, M. Sheng, and Y. Chen Lotus-2: advancing geometric dense prediction with powerful image generative model. arXiv preprint arXiv:2512.01030. Cited by: §2. He et al. (2024) J. He, H. Li, W. Yin, Y. Liang, L. Li, K. Zhou, H. Zhang, B. Liu, and Y. Chen Lotus: diffusion-based visual foundation model for high-quality dense prediction. arXiv preprint arXiv:2409.18124. Cited by: §2. Hoiem et al. (2007) D. Hoiem, A. A. Efros, and M. Hebert Recovering surface layout from an image. International Journal of Computer Vision 75 (1), p. 151–172. Cited by: §2. Hu et al. (2024) M. Hu, W. Yin, C. Zhang, Z. Cai, X. Long, H. Chen, K. Wang, G. Yu, C. Shen, and S. Shen Metric3d v2: a versatile monocular geometric foundation model for zero-shot metric depth and surface normal estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (12), p. 10579–10596. Cited by: §2. Huang et al. (2018) P. Huang, K. Matzen, J. Kopf, N. Ahuja, and J. Huang DeepMVS: learning multi-view stereopsis. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 4, §5.1. Jung et al. (2023) H. Jung, P. Ruhkamp, G. Zhai, N. Brasch, Y. Li, Y. Verdie, J. Song, Y. Zhou, A. Armagan, S. Ilic, A. Leonardis, N. Navab, and B. Busam On the importance of accurate geometry data for dense 3d vision tasks. External Links: 2303.14840, Link Cited by: Table 5, Appendix D, §5.2. Ke et al. (2024) B. Ke, A. Obukhov, S. Huang, N. Metzger, R. C. Daudt, and K. Schindler Repurposing diffusion-based image generators for monocular depth estimation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9492–9502. Cited by: §2, §3, §3. Kingma and Welling (2013) D. P. Kingma and M. Welling Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114. Cited by: §3. Knaebel et al. (2026) K. Knaebel, G. M. Garcia, C. Schmidt, I. Fradlin, L. Nunes, D. de Geus, and B. Leibe SurGe: improved surface geometry in point maps. arXiv preprint arXiv:2605.31577. Cited by: §2, §3. Koch et al. (2018) T. Koch, L. Liebel, F. Fraundorfer, and M. Korner Evaluation of cnn-based single-image depth estimation methods. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, p. 0–0. Cited by: Table 5, Appendix D, §5.2. Kong et al. (2026) L. Kong, R. Li, R. Wang, S. Xu, C. Yao, J. Xiang, and J. Yang Fine-detail monocular geometry estimation with self-guided sparse volumetric refinement. External Links: 2607.17967, Link Cited by: §2. Laina et al. (2016) I. Laina, C. Rupprecht, V. Belagiannis, F. Tombari, and N. Navab Deeper depth prediction with fully convolutional residual networks. In 2016 Fourth international conference on 3D vision (3DV), p. 239–248. Cited by: §2. Lee et al. (2019) J. H. Lee, M. Han, D. W. Ko, and I. H. Suh From big to small: multi-scale local planar guidance for monocular depth estimation. arXiv preprint arXiv:1907.10326. Cited by: §2. Li and He (2026) T. Li and K. He Back to basics: let denoising generative models denoise. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 36115–36125. Cited by: §1. Li et al. (2024a) Z. Li, S. F. Bhat, and P. Wonka Patchfusion: an end-to-end tile-based framework for high-resolution monocular metric depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10016–10025. Cited by: §2. Li et al. (2024b) Z. Li, S. F. Bhat, and P. Wonka Patchrefiner: leveraging synthetic data for real-domain high-resolution monocular metric depth estimation. In European Conference on Computer Vision, p. 250–267. Cited by: §2. Lin et al. (2025) H. Lin, S. Chen, J. Liew, D. Y. Chen, Z. Li, G. Shi, J. Feng, and B. Kang Depth anything 3: recovering the visual space from any views. arXiv preprint arXiv:2511.10647. Cited by: Table 5, §5.2, §5.4. Ma et al. (2026) Z. Ma, L. Wei, S. Wang, S. Zhang, and Q. Tian Deco: frequency-decoupled pixel diffusion for end-to-end image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 43600–43610. Cited by: §3. McCormac et al. (2017) J. McCormac, A. Handa, S. Leutenegger, and A. J. Davison SceneNet rgb-d: can 5m synthetic images beat generic imagenet pre-training on indoor segmentation?. In Proceedings of the IEEE international conference on computer vision, p. 2678–2687. Cited by: Table 4, §5.1. Patel et al. (2025) M. Patel, F. Yang, Y. Qiu, C. Cadena, S. Scherer, M. Hutter, and W. Wang TartanGround: a large-scale dataset for ground robot perception and navigation. arXiv preprint arXiv:2505.10696. Cited by: Table 4, §5.1. Peebles and Xie (2023) W. Peebles and S. Xie Scalable diffusion models with transformers. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), p. 4172–4182. Cited by: §3. Piccinelli et al. (2025) L. Piccinelli, C. Sakaridis, Y. Yang, M. Segu, S. Li, W. Abbeloos, and L. Van Gool Unidepthv2: universal monocular metric depth estimation made simpler. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §2. Piccinelli et al. (2024) L. Piccinelli, Y. Yang, C. Sakaridis, M. Segu, S. Li, L. Van Gool, and F. Yu Unidepth: universal monocular metric depth estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 10106–10116. Cited by: §2. Ranftl et al. (2021) R. Ranftl, A. Bochkovskiy, and V. Koltun Vision transformers for dense prediction. In Proceedings of the IEEE/CVF international conference on computer vision, p. 12179–12188. Cited by: §1, §2, Figure 2, §3, §3. Ranftl et al. (2020) R. Ranftl, K. Lasinger, D. Hafner, K. Schindler, and V. Koltun Towards robust monocular depth estimation: mixing datasets for zero-shot cross-dataset transfer. IEEE transactions on pattern analysis and machine intelligence 44 (3), p. 1623–1637. Cited by: §1, §2. Roberts et al. (2021) M. Roberts, J. Ramapuram, A. Ranjan, A. Kumar, M. A. Bautista, N. Paczan, R. Webb, and J. M. Susskind Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding. In International Conference on Computer Vision (ICCV) 2021, Cited by: Table 4, §5.1, §5.4. Saxena et al. (2007) A. Saxena, M. Sun, and A. Y. Ng Learning 3-d scene structure from a single still image. In 2007 IEEE 11th international conference on computer vision, p. 1–8. Cited by: §2. Schöps et al. (2017) T. Schöps, J. L. Schönberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger A multi-view stereo benchmark with high-resolution images and multi-camera videos. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 5, §5.2. Shotton et al. (2013) J. Shotton, B. Glocker, C. Zach, S. Izadi, A. Criminisi, and A. Fitzgibbon Scene coordinate regression forests for camera relocalization in rgb-d images. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2930–2937. Cited by: Table 5, §5.2. Silberman et al. (2012) N. Silberman, D. Hoiem, P. Kohli, and R. Fergus Indoor segmentation and support inference from rgbd images. In ECCV, Cited by: Table 5, §5.2. Siméoni et al. (2025) O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski DINOv3. External Links: 2508.10104, Link Cited by: §3, §4.1. Tan et al. (2026) B. Tan, C. Sun, X. Qin, H. Adai, Z. Fu, T. Zhou, H. Zhang, Y. Xu, X. Zhu, Y. Shen, et al. Masked depth modeling for spatial perception. arXiv preprint arXiv:2601.17895. Cited by: Table 4, §5.1. Tosi et al. (2021) F. Tosi, Y. Liao, C. Schmitt, and A. Geiger SMD-nets: stereo mixture density networks. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: Table 4, §5.1. Van Hoorick et al. (2024) B. Van Hoorick, R. Wu, E. Ozguroglu, K. Sargent, R. Liu, P. Tokmakov, A. Dave, C. Zheng, and C. Vondrick Generative camera dolly: extreme monocular dynamic novel view synthesis. European Conference on Computer Vision (ECCV). Cited by: Table 4, §5.1. Vasiljevic et al. (2019) I. Vasiljevic, N. Kolkin, S. Zhang, R. Luo, H. Wang, F. Z. Dai, A. F. Daniele, M. Mostajabi, S. Basart, M. R. Walter, and G. Shakhnarovich DIODE: a dense indoor and outdoor depth dataset. External Links: 1908.00463, Link Cited by: Table 5, §5.2. Wang and Shen (2019) K. Wang and S. Shen Flow-motion and depth network for monocular stereo and beyond. External Links: 1909.05452, Link Cited by: Table 4, §5.1. Wang et al. (2025a) R. Wang, S. Xu, C. Dai, J. Xiang, Y. Deng, X. Tong, and J. Yang Moge: unlocking accurate monocular geometry estimation for open-domain images with optimal training supervision. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 5261–5271. Cited by: §1, §1, §2, §3, §5.1, §5.2. Wang et al. (2025b) R. Wang, S. Xu, Y. Dong, Y. Deng, J. Xiang, Z. Lv, G. Sun, X. Tong, and J. Yang Moge-2: accurate monocular geometry with metric scale and sharp details. arXiv preprint arXiv:2507.02546. Cited by: §C.1, §1, §1, §2, §2, §3, §4.1, §5.1, §5.2, §5.3. Wang et al. (2026) S. Wang, Z. Gao, C. Zhu, W. Huang, and L. Wang Pixnerd: pixel neural field diffusion. In International Conference on Learning Representations, Vol. 2026, p. 43559–43580. Cited by: §3. Wang et al. (2020) W. Wang, D. Zhu, X. Wang, Y. Hu, Y. Qiu, C. Wang, Y. Hu, A. Kapoor, and S. Scherer Tartanair: a dataset to push the limits of visual slam. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 4909–4916. Cited by: Table 4, §5.1. Wrenninge and Unger (2018) M. Wrenninge and J. Unger Synscapes: a photorealistic synthetic dataset for street scene parsing. arXiv preprint arXiv:1810.08705. Cited by: Table 4, §5.1. Xu et al. (2025) G. Xu, H. Lin, H. Luo, X. Wang, J. Yao, L. Zhu, Y. Pu, C. Chi, H. Sun, B. Wang, et al. Pixel-perfect depth with semantics-prompted diffusion transformers. arXiv preprint arXiv:2510.07316. Cited by: §C.1, Figure 1, §1, §2, Figure 2, §3, §5.2, §5.2. Xu et al. (2026) H. Xu, R. Wu, P. Henzler, N. Kalischek, M. Oechsle, F. Manhardt, M. Pollefeys, A. Geiger, F. Tombari, and M. Niemeyer PointDiT: pixel-space diffusion for monocular geometry estimation. arXiv preprint arXiv:2607.02515. Cited by: §1. Yang et al. (2024a) L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao Depth anything: unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 10371–10381. Cited by: §1, §1, §2, §3. Yang et al. (2024b) L. Yang, B. Kang, Z. Huang, Z. Zhao, X. Xu, J. Feng, and H. Zhao Depth anything v2. Advances in Neural Information Processing Systems 37, p. 21875–21911. Cited by: §C.1, §1, §1, §2, §3, §5.2. Yin et al. (2023) W. Yin, C. Zhang, H. Chen, Z. Cai, G. Yu, K. Wang, X. Chen, and C. Shen Metric3d: towards zero-shot metric 3d prediction from a single image. In Proceedings of the IEEE/CVF international conference on computer vision, p. 9043–9053. Cited by: §2. Yu et al. (2026a) H. Yu, H. Lin, J. Wang, J. Li, Y. Wang, X. Zhang, Y. Wang, X. Zhou, R. Hu, and S. Peng InfiniDepth: arbitrary-resolution and fine-grained depth estimation with neural implicit fields. arXiv preprint arXiv:2601.03252. Cited by: §C.1, Table 5, Appendix D, Figure 1, §1, §2, §5.2. Yu et al. (2026b) Y. Yu, W. Xiong, W. Nie, Y. Sheng, S. Liu, and J. Luo Pixeldit: pixel diffusion transformers for image generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 14273–14282. Cited by: §1, §1, §3, §4.3, §4.3. Zheng et al. (2020) J. Zheng, J. Zhang, J. Li, R. Tang, S. Gao, and Z. Zhou Structured3D: a large photo-realistic dataset for structured 3d modeling. In Proceedings of The European Conference on Computer Vision (ECCV), Cited by: Table 4, §5.1. Appendix A Additional Implementation Details This appendix provides additional details about the training configuration, model architecture, and evaluation protocol. A.1 Training Data and Configuration Training Datasets. We train our model exclusively on synthetic RGB-D data. Stage 1 denotes low-resolution pretraining, while Stage 2 denotes high-resolution fine-tuning. Table 4 summarizes the data sources and their stage-specific sampling weights. These weights are relative factors used to construct each data mixture rather than percentages or dataset sizes. Table 4: Training dataset summary. Relative sampling weights for both training stages, native RGB resolutions, and scene distributions of the synthetic datasets used for training. Dataset Scene distribution Resolution Sampling weight Stage 1 Stage 2 Hypersim (42) Indoor 1024×7681024× 768 6.0 10.0 MVS-Synth (21) Urban 1920×10801920× 1080 1.2 1.2 TartanAir (56) Indoor / outdoor 640×480640× 480 3.0 3.0 TartanGround (36) Indoor / outdoor, ground robot 640×640640× 640 4.8 4.8 UnrealStereo4K (49) Indoor / outdoor 3840×21603840× 2160 2.4 2.4 GTA-SfM (52) Urban / rural 640×480640× 480 2.8 2.8 Structured3D (65) Indoor 1280×7201280× 720 4.8 4.8 UrbanSyn (15) Urban driving 2048×10242048× 1024 2.1 2.1 ParallelDomain-4D (50) Driving 640×480640× 480 3.5 3.5 VKITTI2 (8) Driving 1242×3751242× 375 2.5 2.5 SceneNetRGBD (35) Indoor 320×240320× 240 3.6 – RobbySim (48) Indoor 1280×9601280× 960 – 5.4 Synscapes (57) Urban driving 1440×7201440× 720 4.0 4.0 Optimization Details. We train our model on NVIDIA RTX H100 GPUs with per-GPU batch sizes of 88 and 44 for Stages 1 and 2, respectively. We configure AdamW with =(0.9,0.95) β=(0.9,0.95) and a weight decay of 0.010.01. The learning rate of the Pixel-Space Depth Predictor is 1×10−41× 10^-4 in Stage 1 and 5×10−55× 10^-5 in Stage 2, while the Global Context Encoder is optimized with a learning rate of 5×10−65× 10^-6 throughout training. Both stages follow a cosine OneCycle schedule with a 2%2\% warm-up, and the global gradient norm is clipped to 1.01.0. RGB images and depth maps undergo the same geometric transformations to maintain pixel-aligned supervision. We apply horizontal flipping with probability 0.50.5. A.2 Training Objective Depth Normalization. We normalize the ground-truth depth to reduce its dynamic range across scenes. Given a valid ground-truth depth value d, we first convert it to log depth and then normalize each depth map as dlog=log(d+1),dnorm=dlog−d0.02d0.98−d0.02,d_ log= (d+1), d_norm= d_ log-d_0.02d_0.98-d_0.02, (10) where d0.02d_0.02 and d0.98d_0.98 denote the 2%2\% and 98%98\% percentiles of the valid log-depth values, respectively. Normalized Depth Loss. Let d d be the predicted normalized depth and V the set of valid depth pixels. We supervise the prediction with an ℓ1 _1 loss, ℒdepth=1||∑u∈|d^(u)−dnorm(u)|.L_depth= 1|V| _u | d(u)-d_norm(u) |. (11) Multi-Scale Depth-Gradient Loss. To preserve local depth variations and sharp discontinuities, we further supervise the gradients of the depth error e=d^−dnorme= d-d_norm. At strides s∈1,2,4,8s∈\1,2,4,8\, the loss is ℒgrad=∑s∈1,2,4,81|s|∑u∈s(|∇xes(u)|+|∇yes(u)|),L_grad= _s∈\1,2,4,8\ 1|V_s| _u _s (| _xe_s(u)|+| _ye_s(u)| ), (12) where ese_s and sV_s denote the error map and valid pixels sampled at stride s, respectively. Binary Mask Loss. The mask head produces a logit for each pixel, which is converted by a sigmoid into the probability m^(u)∈[0,1] m(u)∈[0,1] that the pixel belongs to the valid mask. Given the binary target m(u)m(u) over the labeled pixel set ℳM, we use binary cross-entropy, ℒmask=−1|ℳ|∑u∈ℳ[m(u)logm^(u)+(1−m(u))log(1−m^(u))].L_mask=- 1|M| _u [m(u) m(u)+ (1-m(u) ) (1- m(u) ) ]. (13) At inference time, pixels with m^(u)>0.5 m(u)>0.5 are retained as valid. Final Loss. The three losses are combined as ℒ=λdℒdepth+λgℒgrad+λmℒmask.L= _dL_depth+ _gL_grad+ _mL_mask. (14) We set λd=1 _d=1, λg=0.5 _g=0.5, and λm=0.5 _m=0.5. Appendix B More Architecture Details B.1 Model Configuration Global Context Encoder. We use ViT-L/14 and initialize it from MoGe-2. We extract the features from layers 5, 11, 17, and 23, project each feature map to dctx=1024d_ ctx=1024 channels with a 1×11× 1 convolution, and sum them to obtain the context feature FctxF_ ctx. Pixel-Space Depth Predictor. The depth branch uses the coarse-to-fine pixel compaction design described in Sec. 4.2. The mask head takes the output of the first four CM-PiT blocks and applies two additional CM-PiT blocks with p=14p=14, followed by a 1×11× 1 convolution. We do not use coarse-to-fine pixel compaction in the mask head because valid mask prediction does not require the same level of fine-grained detail as depth prediction. In our experiments, setting p=7p=7 for the mask head provides little improvement but increases computation. We therefore use p=14p=14 in both mask blocks. All CM-PiT blocks use RMSNorm, a SwiGLU feed-forward layer, and 2D RoPE. Compacted pixel tokens are projected to an attention dimension of 1,5361,536 and processed by 2424 attention heads with query/key normalization. Before the output projection, each attention output is modulated by a learned sigmoid gate predicted from its input token. B.2 Pixel Token Compaction Pixel Token Compaction uses learned projections rather than spatial pooling. We partition the pixel feature map into non-overlapping p×p× p regions, flatten the p2p^2 local pixel features, and project each region from p2dpixp^2d_ pix dimensions to the shared attention dimension. Global self-attention is applied to the resulting (H/p)(W/p)(H/p)(W/p) tokens. A second linear layer projects each attended token back to p2dpixp^2d_ pix dimensions before the features are restored to their original pixel locations. Only the attention branch uses compacted tokens, while the SwiGLU feed-forward layer operates independently on each pixel feature after expansion. B.3 Context-Guided Adaptive Normalization The context feature Fctx∈ℝHP×WP×dctxF_ ctx HP× WP× d_ ctx is passed through a SiLU activation and a linear layer. The linear layer expands FctxF_ ctx into six pixel-wise parameter maps, each with dpixd_ pix channels. These maps provide the shift β, scale residual γ, and residual gate α for the attention and feed-forward branches. They are applied directly to the pixel-space features, preserving spatial alignment without interpolation or learned upsampling. Appendix C Evaluation Protocol Details C.1 Baseline Model Configurations We evaluate all baselines using their official implementations and publicly released checkpoints. For DA V2 (61), we adopt the relative-depth model with a ViT-L backbone. DepthPro (6) and InfiniDepth (63) are evaluated with their publicly released monocular models. For MoGe-2 (54), we select the ViT-L variant that jointly predicts geometry and surface normals. PPD (58) is evaluated with the variant conditioned on semantic features from DA V2 and runs for 4 sampling steps. For MDA (5), we adopt the authors’ primary DA3-Giant configuration with a Gaussian mixture-density head, a dedicated sky component, and an ℓ2 _2 objective in log-depth space. C.2 Evaluation Datasets and Resolution Handling Table 5 summarizes the datasets and dataset-specific evaluation resolutions used in our experiments. For each method, we first resize the input image to its target inference resolution. We then restore the predicted depth map to the corresponding evaluation resolution using nearest-neighbor interpolation before computing metrics or generating visualizations. This ensures that all methods are evaluated at the same resolution on each dataset, enabling a fair comparison. Nearest-neighbor interpolation avoids blending depth values across discontinuities, which can introduce spurious 3D points and distort object boundaries. Table 5: Evaluation benchmark summary. Configured evaluation resolutions and scene distributions of the zero-shot benchmarks. Dataset Type Scene distribution Resolution 7Scenes (45) Real Indoor 504×378504× 378 NRGBD (1) Synthetic Indoor 504×378504× 378 HiRoom (33) Synthetic Indoor 504×504504× 504 NYUv2 (46) Real Indoor 640×480640× 480 KITTI (14) Real Driving 1242×3751242× 375 ETH3D (44) Real Indoor / outdoor 1024×6861024× 686 iBims-1 (26) Real Indoor 640×480640× 480 Sintel (7) Synthetic Animated scenes 872×436872× 436 DDAD (16) Real Urban driving 1400×7001400× 700 DIODE (51) Real Indoor / outdoor 1024×7681024× 768 HAMMER (22) Real Indoor tabletop 1664×8321664× 832 Synth4K (63) Synthetic Indoor / outdoor games 896×504896× 504 C.3 Boundary Evaluation Protocol For the MDA boundary evaluation, each dataset is processed at the evaluation resolution in Table 5. Each view is first cropped to the largest region symmetric around the camera principal point and then resized to cover the target shape followed by a center crop. Camera intrinsics are updated after both operations, and depth is resampled with nearest-neighbor interpolation. Following the MDA protocol, ground-truth depth is clipped to [0.1,65][0.1,65] meters, min–max normalized to an 8-bit image, and processed by a Canny detector with thresholds 100100 and 200200. Pixels adjacent to invalid depth are removed by one 2×22× 2 dilation of the invalid mask. Predicted and ground-truth depths at the remaining edge pixels are back-projected with the ground-truth intrinsics. Samples with fewer than ten valid boundary points are excluded from the 3D boundary metrics. The predicted boundary point cloud is registered to the ground truth by point-to-point ICP with a correspondence threshold of 0.10.1 m. Accuracy (Acc) is the mean nearest-neighbor distance from predicted to ground-truth boundary points, completeness is the reverse distance, and Chamfer distance (CD) is their average. We report Acc and CD in millimeters. Algorithm C.3 summarizes the complete procedure. Algorithm 1 Boundary evaluation protocol. Computation of 3D boundary metrics. Input. Predicted depth, ground-truth depth, and camera intrinsics. function EvaluateBoundary() Detect valid ground-truth depth boundaries. Convert predicted and ground-truth depths at these boundaries into 3D point sets. Discard samples without enough valid boundary points. Align the predicted points to the ground truth using point-to-point ICP. Average predicted-to-ground-truth distances as Acc. Average ground-truth-to-predicted distances as completeness. Average Acc and completeness as CD. return Acc and CD in millimeters. end function Appendix D Additional Quantitative Results Evaluation results. To complement the main evaluation, we provide additional quantitative results that probe zero-shot generalization across both synthetic and real-world scenes. Synth4K (63), introduced by InfiniDepth, comprises five subsets of indoor and outdoor game scenes with dense synthetic depth, and is used to evaluate affine-invariant depth accuracy under controlled geometry. We further report 3D boundary metrics on iBims-1 (26) and HAMMER (22), two real-world datasets included in the MoGe Benchmark that cover indoor and tabletop scenes, to assess whether boundary fidelity transfers beyond synthetic data. The results are reported in Tables 6 and 7, respectively. Table 6: Quantitative comparison on the Synth4K Benchmark. Zero-shot affine-invariant depth estimation across its five subsets. Category Method Synth4K-1 Synth4K-2 Synth4K-3 Synth4K-4 Synth4K-5 Mean Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Rel↓ δ1↑ _1 Global Prediction DA V2 17.9 80.1 11.9 87.1 17.7 83.5 8.4 93.7 8.2 92.7 12.8 87.4 DepthPro 13.0 81.8 9.0 87.3 10.0 86.2 7.3 92.6 6.7 92.7 9.2 88.1 MoGe-2 10.0 86.4 7.6 89.9 7.6 91.3 5.4 94.8 4.9 95.5 7.1 91.6 InfiniDepth 13.6 85.4 10.4 87.4 13.0 85.0 7.9 94.5 6.2 95.1 10.2 89.5 PPD 14.5 78.7 13.0 79.8 15.4 76.0 7.6 92.4 10.2 85.2 12.2 82.4 MDA 22.6 69.2 16.3 80.7 13.9 85.1 14.3 84.4 14.9 81.2 16.4 80.1 Structure Aware Ours 9.9 85.7 7.7 89.7 7.6 91.0 5.2 94.9 4.4 96.0 7.0 91.5 Table 7: Quantitative comparison on iBims-1 and HAMMER. Acc and CD are reported in millimeters. Category Method iBims-1 HAMMER Mean Acc↓ CD↓ Acc↓ CD↓ Acc↓ CD↓ Global Prediction DA V2 145.9 203.2 35.9 43.3 90.9 123.3 DepthPro 138.6 196.6 32.9 40.3 85.8 118.4 MoGe-2 116.3 175.2 29.5 36.7 72.9 105.9 InfiniDepth 157.7 247.4 28.5 38.6 93.1 143.0 PPD 126.4 273.1 23.6 33.8 75.0 153.4 MDA 121.7 172.7 22.3 32.5 72.0 102.6 Structure Aware Ours 112.5 162.3 20.8 29.4 66.7 95.9 Appendix E Additional Qualitative Results Figure 7 & 8 extends the qualitative evaluation to diverse scenes. Across both indoor and outdoor examples, our method more faithfully preserves thin structures and object boundaries, producing cleaner point-cloud reconstructions with fewer flying points than competing methods. Figure 7: Additional qualitative comparisons. Our method preserves fine detail structures. Figure 8: Additional qualitative comparisons. Our method preserves fine detail structures.