Paper deep dive
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
Zhengqin Li, Cheng Zhang, Jakob Engel, Zhao Dong
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/10/2026, 3:04:13 AM
Summary
LSRM (Large Sparse Reconstruction Model) is a feed-forward 3D reconstruction framework that utilizes scaled transformer context windows and Native Sparse Attention (NSA) to achieve high-fidelity object-centric reconstruction and inverse rendering. By employing a two-stage coarse-to-fine pipeline, 3D-aware spatial routing, and a custom block-aware sequence parallelism strategy, LSRM significantly improves texture and geometry details compared to prior state-of-the-art methods.
Entities (5)
Relation Signals (3)
LSRM → utilizes → Native Sparse Attention
confidence 98% · To scale effectively, we adapt native sparse attention in our architecture design
LSRM → evaluatedon → GSO
confidence 95% · Comprehensive experiments on standard object-centric 3D reconstruction benchmarks demonstrate substantial gains... on the GSO dataset
LSRM → performstask → inverse rendering
confidence 95% · Furthermore, when extending LSRM to inverse rendering tasks, qualitative and quantitative evaluations on widely-used benchmarks demonstrate consistent improvements
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows impacts feed-forward 3D reconstruction. Although recent object-centric feed-forward methods deliver robust, high-quality reconstruction, they still lag behind dense-view optimization in recovering fine-grained texture and appearance. We show that expanding the context window -- by substantially increasing the number of active object and image tokens -- remarkably narrows this gap and enables high-fidelity 3D object reconstruction and inverse rendering. To scale effectively, we adapt native sparse attention in our architecture design, unlocking its capacity for 3D reconstruction with three key contributions: (1) an efficient coarse-to-fine pipeline that focuses computation on informative regions by predicting sparse high-resolution residuals; (2) a 3D-aware spatial routing mechanism that establishes accurate 2D-3D correspondences using explicit geometric distances rather than standard attention scores; and (3) a custom block-aware sequence parallelism strategy utilizing an All-gather-KV protocol to balance dynamic, sparse workloads across GPUs. As a result, LSRM handles 20x more object tokens and >2x more image tokens than prior state-of-the-art (SOTA) methods. Extensive evaluations on standard novel-view synthesis benchmarks show substantial gains over the current SOTA, yielding 2.5 dB higher PSNR and 40% lower LPIPS. Furthermore, when extending LSRM to inverse rendering tasks, qualitative and quantitative evaluations on widely-used benchmarks demonstrate consistent improvements in texture and geometry details, achieving an LPIPS that matches or exceeds that of SOTA dense-view optimization methods. Code and model will be released on our project page.
Tags
Links
- Source: https://arxiv.org/abs/2604.05182v1
- Canonical: https://arxiv.org/abs/2604.05182v1
Trouble viewing inline? Open PDF directly →
Full Text
62,984 characters extracted from source content.
Expand or collapse full text
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows Zhengqin Li, Cheng Zhang, Jakob Engel, and Zhao Dong Meta Reality Labs Research Abstract. We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows impacts feed-forward 3D reconstruction. Although recent object-centric feed-forward methods de- liver robust, high-quality reconstruction, they still lag behind dense-view optimization in recovering fine-grained texture and appearance. We show that expanding the context window—by substantially increasing the number of active object and image tokens—remarkably narrows this gap and enables high-fidelity 3D object reconstruction and inverse rendering. To scale effectively, we adapt native sparse attention [73] in our architec- ture design, unlocking its capacity for 3D reconstruction with three key contributions: (1) an efficient coarse-to-fine pipeline that focuses compu- tation on informative regions by predicting sparse high-resolution residu- als; (2) a 3D-aware spatial routing mechanism that establishes accurate 2D-3D correspondences using explicit geometric distances rather than standard attention scores; and (3) a custom block-aware sequence paral- lelism strategy utilizing an All-gather-KV protocol to perfectly balance dynamic, sparse workloads across GPUs. As a result, LSRM handles 20× more object tokens and >2× more image tokens than prior state-of-the- art (SOTA) methods. Extensive evaluations on standard novel-view syn- thesis benchmarks show substantial gains over the current SOTA, yield- ing >2.4 dB higher PSNR and >40% lower LPIPS. Furthermore, when extending LSRM to inverse rendering tasks, qualitative and quantitative evaluations on widely-used benchmarks demonstrate consistent improve- ments in texture and geometry details, achieving an LPIPS that matches or exceeds that of SOTA dense-view optimization methods. Code and model will be released on our project page. Keywords: Object-centric feed-forward reconstruction· Sparse atten- tion· 3D foundation model 1 Introduction Recent years have witnessed rapid progress in the application of 3D foundation models—typically built on large-scale transformer architectures [60]—to tackle 3D tasks previously considered intractable. These tasks include joint estimation of geometry and camera parameters [33, 47, 61, 62], dynamic scene reconstruc- tion [44, 46, 49, 69, 78], and sparse-view reconstruction [22, 26, 63, 79, 84] and in- verse rendering [41, 75]. In object-centric reconstruction and inverse rendering arXiv:2604.05182v1 [cs.CV] 6 Apr 2026 2Z. Li et al. Fig. 1: High-fidelity 3D reconstruction. Given 12–18 images (left), LSRM adapts Native Sparse Attention (NSA) to generate explicit meshes and textures in a single feed-forward pass with minimal compute overhead. As an extension, LSRM can also predict BRDF maps for inverse rendering (right). Zoom in for details. in particular, feed-forward models cut reconstruction time by orders of magni- tude compared with optimization-based pipelines [51,56,80,81], while achieving competitive view synthesis and relighting quality, and often exhibiting stronger geometric robustness on challenging specular objects [41]. Yet a key limitation persists: reconstructed textures frequently miss fine-grained detail, leading to blurry text, smeared logos, and distorted facial features. Recent attempts to address this—by swapping the underlying 3D representation [26, 79] or adding post-hoc texture refinement [54]—improve results but still do not match the fidelity of dense-view optimization. Narrowing this gap is essential for deploy- ing feed-forward 3D reconstruction at the quality and reliability required for practical, industrial-scale 3D digital twin creation. We hypothesize that this gap can be narrowed by expanding the context window—specifically, by increasing the number of active object and image to- kens. However, naïvely upscaling the resolution of 3D object or 2D image repre- sentations rapidly becomes impractical, incurring quadratic (and in volumetric settings, cubic) growth in GPU memory and compute. To overcome this bottle- neck, we adapt native sparse attention (NSA) [73], a hardware-friendly algorithm that restricts each token’s attention to a sparse set of token blocks, enabling ef- ficient training and inference at substantially larger context sizes. Building on NSA, we develop an efficient coarse-to-fine training pipeline. In the first stage, a dense reconstruction transformer identifies the most informative volume tokens; in parallel, we prune 2D image tokens by retaining only those covering fore- ground pixels. This coarse model provides a strong initialization for both the network weights and the geometric structure. In the second stage, the sparse re- construction transformer refines the reconstruction by predicting high-resolution High-Fidelity Object-Centric Reconstruction via Scaled Context Windows3 sparse volume residuals, an approach we find considerably more effective than directly regressing a high-resolution volume from scratch. While enlarging the context window with NSA already improves baseline quality, we introduce two critical architectural adaptations to further maximize reconstruction fidelity and hardware efficiency. First, we observe that the default NSA block selection [73]—which relies purely on attention scores from com- pressed tokens—often fails to capture reliable correspondences between voxel and image tokens, especially in early transformer layers. To remedy this, we propose a 3D-aware spatial routing strategy that explicitly uses the geometric distances established in Stage 1, alongside known camera parameters, to retrieve relevant blocks. This 3D-aware routing markedly enhances texture sharpness, rendering previously blurred fine text legible. Second, the dynamic sparse lay- out naturally induces severe token imbalance across GPUs. To address this, we design a custom block-aware sequence parallelism scheme. We shard tokens by their spatial 2D and 3D blocks, ensuring tokens within the same block remain localized on the same GPU to eliminate cross-device communication during local attention. Furthermore, we leverage the high KV compression of Grouped-Query Attention (GQA) [1] to implement an All-gather-KV protocol. This efficiently broadcasts a lightweight KV cache across the node, providing full global context with minimal communication overhead. Together, these designs enable LSRM to process 20× more object tokens and >2× more image tokens than the prior SOTA [41], without a proportional increase in computational resources. Compre- hensive experiments on standard object-centric 3D reconstruction benchmarks demonstrate substantial gains, yielding >2.4 dB higher PSNR and a >40% re- duction in LPIPS on the GSO dataset [16]. We further extend LSRM to the task of inverse rendering by fine-tuning the model to predict material reflectance from images captured under natural illumination. Both qualitative and quantitative experiments show significant improvements in textural and geometric detail. Specifically on the widely used StanfordORB dataset [29], LSRM sets a new standard for feed-forward inverse rendering methods, achieving an LPIPS that matches the SOTA dense-view optimization method [56]. 2 Related Works Object-Centric Feed-Forward Reconstruction 3D foundation models trained on large-scale datasets [13] have driven major progress in object-centric reconstruc- tion. Pioneering models like LRM [22] achieved impressive single-image geom- etry recovery, but they often struggle to reproduce high-fidelity textures. Sub- sequent works have sought to improve appearance quality through architectural refinements [63], alternative 3D representations [7,26,68,79,89], progressive up- dates [41], and post hoc texture refinement [54]. Despite these advances, a sizable performance gap compared to dense-view, optimization-based pipelines persists. A recent approach [83] narrows this gap via test-time training to effectively en- large the context window; however, it lacks an explicit 3D representation and may suffer from view synthesis degradation at unfamiliar camera angles. In contrast, 4Z. Li et al. LSRM leverages scaled context windows to explicitly reconstruct high-fidelity 3D geometry and textures in a single forward pass. Furthermore, by extending LSRM to predict BRDF material maps, we enable seamless downstream editing and realistic relighting within standard graphics pipelines. Inverse Rendering Inverse rendering aims to decompose images into their in- trinsic factors—geometry, materials, and illumination—so that scenes or ob- jects can be edited and re-rendered realistically. Classical measurement-based approaches [2, 12, 19] rely on specialized capture rigs and controlled environ- ments to densely sample view and lighting directions. Recent optimization-based methods [3–5, 17, 21, 25, 27, 45, 51, 53, 56, 71, 80, 81, 85–87] can often work under natural illumination, but they usually require dense multi-view inputs and re- main fragile in ill-posed conditions such as saturated specularities, cast shadows, and complex interreflections. Many learning-based techniques focus on predicting per-pixel BRDF maps [6,14,36,38–40,42,50,76] without producing a complete 3D representation suitable. Recently, propelled by advances in object-centric 3D reconstruction models, pioneering works [41, 75] have shown that feed-forward models can output fully relightable 3D objects with explicit geometry and ma- terials. LSRM builds on this direction and further improves reconstruction and appearance fidelity by enabling substantially longer context windows. Sparsity in 3D Generation and Reconstruction Many prior approaches handle the inherent sparsity of 3D surfaces indirectly, e.g., by compressing geometry into compact representations [20,23,31,34,35,37,64,77,82,88] or adopting localized point-based formats such as 3D Gaussians [9, 32, 59, 70]. More recently, a line of 3D generative work has modeled sparsity explicitly through structured latent representations—such as sparse voxel grids—to scale volumetric resolution sub- stantially [43,65–67]. In these frameworks, sparse volumes are either encoded by sparse 3D VAEs into dense latent codes [8,43,66,67] or, as in Direct3D-S2 [65], downsampled to a lower-resolution sparse volume that can be generated with an NSA-style transformer operating over 3D token blocks. While these meth- ods demonstrate efficiency in processing high-resolution volumes, their focus remains restricted to synthesizing intricate geometric details, often neglecting texture. Even when texture generation is supported [8,66], the resulting texture fidelity remains limited even under clean synthetic inputs. This highlights that high-fidelity textured object reconstruction is a distinct and challenging problem, motivating the specialized architecture design by LSRM. 3 Method This section details the architecture of LSRM. We first give a simplified review of NSA to provide the necessary background. Next, we formulate the Stage 1 dense reconstruction transformer, specifying the model’s inputs, outputs, and notation. Finally, we describe our Stage 2 sparse reconstruction transformer, which is initialized from the dense model. We detail how we leverage NSA to significantly expand the context window, along with two key improvements to maximize reconstruction fidelity and hardware efficiency. High-Fidelity Object-Centric Reconstruction via Scaled Context Windows5 3.1 Background: Native Sparse Attention NSA is a hardware-friendly mechanism designed for highly efficient end-to-end training and inference on modern NVIDIA GPUs. We adopt NSA to scale the context window because, unlike earlier sparse-attention schemes that impose fixed block-to-block patterns [10,74], NSA lets each query adaptively choose the sparse KV-blocks it attends to, which we find essential for high-fidelity recon- struction. Let q ∈ R N x ×h q ×d h and k,v ∈ R N y ×h kv ×d h denote the query, key, and value tensors computed from the input token sequences x ∈ R N x ×d and y ∈ R N y ×d , respectively. Note that in self-attention, x and y are the same ten- sor. Here, N x and N y represent the sequence lengths, h q and h kv denote the number of heads for queries and keys/values, and d h is the head dimension, so d = h q × d h . NSA dynamically selects and recomputes keys k (·) i and values v (·) i for each query q i ∈ R h q ×d h using specific strategies based on original q, k, and v. The slightly simplified variant employed in LSRM is formulated as follows: Compressed (k cmp ,v cmp ): We divide k and v into B non-overlapping blocks, where B ≪ N. The keys and values within each block are then compressed into a single key-value pair k cmp ,v cmp . Note that because k cmp and v cmp are identical for all queries, we omit the subscript i for them. Selected (k sel i ,v sel i ): For each query q i , we compute its attention scores with the compressed representations k cmp and v cmp , and select the top B sel KV-blocks (with B sel ≪ B). We then define k sel i ,v sel i as the set of all original keys and values contained within these B sel selected blocks. Window (k win i ,v win i ): For each query q i , we define k win i ,v win i as all keys and values located in the same block as the original k i and v i . Our NSACrossAttn layer outputs a gated combination of these three atten- tion branches. Let o i denote the output for query q i : o i = w cmp i CmpAttn(q i ,k cmp ,v cmp ) + w sel i SelAttn(q i ,k sel i ,v sel i ) + w win i WinAttn(q i ,k win i ,v win i )(1) w cmp i ,w sel i ,w win i = Sigmoid(Linear gate (x i )) (2) Following Qiu et al. [52], we predict gating weights w cmp i ,w sel i ,w win i with the same dimension as x i to ensure more stable training. These three attention branches are highly complementary: CmpAttn efficiently captures global context due to its reduced sequence length (B ≪ N), SelAttn retrieves fine-grained details from the most relevant regions with low computational cost (B sel ≪ B), and WinAttn models strictly local context with minimal overhead. Among the three branches, CmpAttn and WinAttn can be directly imple- mented with FlashAttn [11]. SelAttn, however, requires a custom Triton ker- nel [57] because each query q i attends to a dynamically chosen sparse set of KV-blocks. Instead of loading a full block of queries and KV pairs into SRAM, NSA leverages the high compression ratio of GQA [1] to load multiple query heads together with a single shared head of k sel i and v sel i . To fully utilize NVIDIA Tensor Cores, the ratio h q /h kv should be a multiple of 16 [30]. In LSRM, we 6Z. Li et al. Sparse Image Tokens ( ) Patchify & Linear 1D Volume Token Embedding Down- sample Down- sample DINOv3 DenseBlock x 24 Cross- Attn Cross- Attn Gate Add & Norm FFN Add & Norm Cross- Attn Cross- Attn Gate Add & Norm FFN Add & Norm Linear & Reshape Token selection 1D Volume Token Embedding Token selection Patchify & Linear DINOv3 NSABlock x 24 NSA CrsAttn NSA CrsAttn Gate Add & Norm FFN Add & Norm NSA CrsAttn NSA CrsAttn Gate Add & Norm FFN Add & Norm Linear & Reshape Up- sample Up- sample Dense Image Tokens ( ) Dense Volume Tokens ( ) Sparse Volume Tokens ( ) Sparse Feature Volume ( ) Dense Reconstruction Transformer Sparse Reconstruction Transformer Input Images ( ) and Plucker Rays ( ) Mask computation Fig. 2: Network architecture of LSRM. Our method employs a two-stage coarse- to-fine pipeline. In Stage 1, a Dense Reconstruction Transformer generates a coarse, low-resolution volume. Stage 2 utilizes this coarse volume to initialize the active sparse volume tokens and predicts high-resolution sparse residuals, constructing the final 3D representation. adapt the Triton implementation from [30, 65] and set h q = 32 and h kv = 2. This high ratio also enables more efficient sequence parallelism (Sec. 3.5). We next describe how we tailor NSA to high-fidelity 3D object reconstruction. The overall LSRM architecture is shown in Fig. 2. 3.2 Stage 1: Coarse Dense Reconstruction for Initialization We first train a dense reconstruction transformer for lower resolution reconstruc- tion; its weights and intermediate predictions are then used to initialize the Stage 2 sparse reconstruction transformer. Image and Volume Tokenization The inputs consist of a sparse set of M posed images (M ∈ [12, 18] similar to [41]), with camera parameters explicitly encoded as Plücker rays. LetI m M m=1 denote the non-overlapping 8×8 patches extracted from the 256×256 input views, yielding a spatial token map of resolution S img d = 32 per view. LetR m M m=1 denote their corresponding Plücker rays. We compute the combined multi-view image tokens, y, by fusing a linear projection of the patches and rays with features extracted from a frozen DINOv3 [55] encoder: y =Linear(I m ,R m ) + Linear(DINOv3(Upsample(I m ))) M m=1 (3) Because DINOv3 natively operates on 16× 16 patches, the 8× 8 input patches are upsampled prior to encoding. Empirically, incorporating DINOv3 features significantly accelerates training convergence and improves the model’s overall generalizability. The volume tokens x are initialized from learned 3D positional embeddings. To avoid wasting GPU memory when later scaling to high-resolution grids in the Stage 2 sparse transformer, we factorize the embeddings along the three spatial axes. Instead of learning a dense parameter volume, we learn three independent 1D embeddings: P x ∈ R S vol d ×d , P y ∈ R S vol d ×d , and P z ∈ R S vol d ×d , where d = 1024, S vol d = 16 are the embedding dimension and the dense token volume’s resolution. High-Fidelity Object-Centric Reconstruction via Scaled Context Windows7 For a token at spatial coordinate (i,j,k), its embedding p i,j,k is simply the element-wise sum: p i,j,k = P x i + P y j + P z k (4) The resulting initialized volume tokens, x =p i,j,k , are then concatenated with image tokens y and processed by the dense transformer blocks. Dense Reconstruction Transformer Following prior works [41, 63], our dense transformer consists of 24 blocks with hidden dimension 1024. One difference is that we set h q = 32 and h kv = 2 (instead of h q = h kv = 16) to match the subsequent sparse transformer. Additionally, rather than standard monolithic self-attention, we use a gated mixture of cross-attention updates to evolve token representations. This decoupled design provides flexible, fine-grained control over attention branches in Stage 2—for example, allowing independent tuning of the number of selected KV-blocks (L) for image and volume tokens. Let x denote volume tokens and y denote image tokens. The attention module within one DenseBlock is: o x = w x-self CrossAttn(x,x) + w x-cross CrossAttn(x,y) (5) o y = w y-self CrossAttn(y,y) + w y-cross CrossAttn(y,x) (6) w x-self ,w x-cross = Sigmoid(Linear x gate (x))(7) w y-self ,w y-cross = Sigmoid(Linear y gate (y))(8) We denote by x d and y d the volume and image token outputs from the final DenseBlock. We decode the volume tokens x d into a dense feature volume via a linear projection. This layer upsamples spatial resolution by 4×, producing a grid of size S vol df = 64 while reducing feature dimension to d f = 32: X d = Linear(x d ), X d ∈ R S vol df 3 ×d f (9) This dense feature volume, X d , can be utilized for either novel view synthesis or explicit textured mesh extraction via lightweight MLP decoders. With a slight abuse of notation, let p ∈ R 3 be a continuous 3D point sampled within the volume. We compute its corresponding properties as follows: f = Trilinear(X d ;p)(10) z = Sigmoid(MLP z (f)), z∈a,r,m or z = c(11) s = MLP s (f) + s bias (p) (12) where f is the trilinearly interpolated feature vector, s is the predicted SDF , and s bias is an effective offset adapted from [41]. c denotes the emitted color for novel view synthesis. Alternatively a, r, and m represent the albedo, roughness, and metallic material properties for inverse rendering. For rendering, we adopt the VolSDF [72] formulation due to its simplicity and effectiveness. 8Z. Li et al. 3.3 Stage 2: Sparse Reconstruction with Long Context Windows For high-fidelity reconstruction in Stage 2, we substantially increase the resolu- tions of both the input images (e.g., S img = 3S img d = 96) and the volume grid (e.g., S vol = 6S vol d = 96). To efficiently handle the resulting expanded token sequence, we first describe how we extract active, spatially-sparse token subsets, and then present the spatial block partitioning required by our NSA mechanism, followed by the architecture of the sparse reconstruction transformer. Informative Token Selection We build a spatially-sparse token representation by retaining only the most informative image and volume tokens. For image tokens, we discard the background, preserving only patches that contain foreground pixels based on the foreground mask M img . For the volume tokens, we leverage the geometric prior from the Stage 1 dense transformer to compute a binary mask M vol . This mask identifies and retains only the voxels near the object surface, as these primarily determine the final appearance. Concretely, within a voxel at spatial coordinate (i,j,k), we uniformly sample a grid of T = 4 3 = 64 points P i,j,k = p t T t=1 . We evaluate the SDF value s t at each point using the frozen Stage 1 dense feature volume X d and set the voxel mask M vol i,j,k based on a distance threshold τ: s t = MLP s (Trilinear(X d ;p t )) M vol i,j,k = I min p t ∈P i,j,k |s t |≤ τ or min p t ∈P i,j,k s t · max p t ∈P i,j,k s t ≤ 0 (13) where I(·) denotes the indicator function. This rule ensures that only voxels close to the underlying surface, or directly intersected by it, are instantiated as tokens for high-resolution processing. Block Partitioning and Compression A naïve 1D sequence partitioning of tokens destroys spatial locality, which is detrimental to fine-grained 3D reconstruction. Therefore, we adopt a spatial block partitioning strategy similar to Direct3D- S2 [65]. We group active image and volume tokens based on their original 2D and 3D coordinates. Specifically, space is divided into discrete blocks of size 8×8 for images and 8× 8× 8 for the volume, reducing the block-level resolutions to S img b = S img /8 and S vol b = S vol /8. Because uninformative tokens are removed, each spatial block contains a variable number of active tokens. To compute the block-level compressed keys (k cmp ) and values (v cmp ) required by NSA, we first pass the individual token keys k t and values v t within a block through a residual module (ResBlock), and then average the resulting features in-block (AvgPool): k cmp = AvgPool(ResBlock(k t )), v cmp = AvgPool(ResBlock(v t )) Sparse Reconstruction Transformer The sparse reconstruction transformer uses the same number of transformer blocks (24) and hidden dimensions (1024) as the Stage 1 model. The key difference is that we replace the standard CrossAttn in Eq. (5) and (6) with NSACrssAttn, detailed in Eq. (1) and (2). This struc- tural alignment lets us initialize sparse-model weights directly from Stage 1, High-Fidelity Object-Centric Reconstruction via Scaled Context Windows9 Sparse Volume Tokens Query Token Attention Score- based Selection NSABlock 0 Attention Score- based Selection NSABlock 23 3D-aware Selection Fig. 3: Attention- score vs. 3D-aware selection. Standard at- tention captures 2D-3D correspondence only in deep layers, whereas our 3D-aware routing ensures stable selection to enhance texture. which significantly accelerates convergence. Furthermore, instead of predicting the high-resolution sparse volume from scratch, we train the sparse transformer to predict the residual over the dense prediction, an approach we found to be em- pirically much more effective. Let x up d and y up d denote the sparse tokens obtained by selecting and upsampling the informative regions of the dense predictions x d and y d . The forward pass through the sparse transformer blocks is defined as: x (m) ,y (m) = NSABlock (m) x (m−1) + Linear (m) (x up d ), y (m−1) + Linear (m) (y up d ) , for m = 1... 24 (14) x s ,y s = x (24) + x up d , y (24) + y up d (15) We project output tokens into a sparse feature volume X s using the same linear layer as Eq. (9), resulting in resolution S vol f = 4S vol = 384. Note that X s only contains features computed from selected active volume tokens. We therefore combine it with the full dense feature volume X d during volume rendering. During ray marching, if a sampled point falls into a voxel (i,j,k) where the mask M vol i,j,k is true, we query high-frequency features from X s . Conversely, if the point lies in empty space, we query the low-frequency features from X d . For points on the boundary between the two, we compute a linear combination of X s and X d to ensure smooth transitions. Finally, the lightweight MLP decoder processes these queried features exactly as in the dense reconstruction transformer, as shown in Eq. (11) and (12). 3.4 3D-Aware Block Routing Although our experiments show that LSRM benefits from an expanded context window and already improves fidelity over SOTA, standard NSA KV-block se- lection remains a bottleneck. In vanilla NSA, token blocks are chosen solely from attention scores against k cmp . As Fig. 3 illustrates, this data-driven routing of- ten works in deeper layers, yet early transformer blocks frequently miss the most spatially relevant blocks, which cascades through the network and reduces final reconstruction quality. To address this, we propose a 3D-aware routing strategy that selects KV- blocks using the coarse geometry established in Stage 1. We first assign an ex- plicit 3D coordinate to each token. A volume token’s coordinate p vol is simply 10Z. Li et al. Fig. 4: Token count vs. training time. Dis- tribution of active to- kens and their corre- sponding training times for resolutions S img = 96 and S vol = 64. Each point represents a single data instance. GPU 0 4 blocks 14 tokens GPU 1 3 blocks 8 tokens GPU 2 1 blocks 2 tokens GPU 0 2 blocks 7 tokens GPU 1 2 blocks 7 tokens GPU 2 4 blocks 10 tokens All-to-All GPU 0 2 blocks 7 KVs GPU 1 2 blocks 7 KVs All-gather All-to-All GPU 0 8 blocks 24 KVs GPU 1 8 blocks 24 KVs GPU 2 8 blocks 24 KVs GPU 2 4 blocks 10 KVs (a) (b) (c) Fig. 5: Custom sequence parallelism. Toy example assuming≤4 tokens per spatial block across 3 GPUs. (a) Tokens are distributed evenly without breaking spatial blocks. (b) Tokens return to their source GPUs for decoding and volume rendering. (c) All- gather-KV ensures global context before every NSACrossAttn layer. its voxel-center position. For an image token corresponding to an 8× 8 patch, its 3D location p img is defined as the foreground surface point with the highest opacity, rendered via the dense feature volume X d . We then define the 3D point sets P img and P vol for image and volume blocks as the collections of token co- ordinates in the block. Based on these geometric descriptors, our block-selection rules are: Image/Volume to Image Blocks For a given query token with 3D coordinate p∈p img ,p vol , we first project p onto each input image plane using the known camera parameters. To limit computation, we retrieve the B i image blocks per view whose 2D centers are closest to this projection. Next, for each candidate block we compute the minimal 3D Euclidean distance between p and the point set P img . Finally, we select the B i2i (image queries) or B v2i (volume queries) candidates with the smallest such distances. Image/Volume to Volume Blocks Given the query coordinate p ∈ p img ,p vol , we compute its 3D Euclidean distance to the spatial centers of all volume token blocks. We then select the B i2v (for image queries) or B v2v (for volume queries) volume blocks with the smallest center distances. Fig. 3 compares our volume-to-image block selection against the original NSA mechanism on a converged model. The visualization confirms that our geometric approach establishes far more accurate 2D-3D correspondences. In Sec. 4, we show this enhanced spatially-aware routing translates to substantial improvements in reconstructed texture details (Fig. 7 and Tab. 1). 3.5 Custom Sequence Parallelism for Workload Balancing LSRM’s sparse tokenization yields highly variable active-token counts across in- stances, producing straggler GPUs that stall training (Fig. 4) and tightly limit High-Fidelity Object-Centric Reconstruction via Scaled Context Windows11 attainable resolution during multi-GPU training, since each step waits for the slowest worker. Early attempts to cap per-GPU tokens with reservoir sampling routinely dropped as many as one third of active tokens, which destabilized optimization. To balance this dynamic workload systematically, we adapt con- text parallelism [28] into a custom block-aware sequence parallelism. Unlike naive splits that break the block structure required by NSA, our scheme spreads tokens evenly across GPUs while strictly keeping each spatial block on a single device. This preserves block locality, so window attention and KV-block com- pression are computed entirely on-device, with zero cross-GPU communication. To execute compressed (CmpAttn) and selected (SelAttn) attention over global context, tokens must still be routed efficiently. We use an All-gather-KV strategy [18,48] that exploits our high KV compression. Concretely, we all-gather to replicate (k cmp ,v cmp ) and (k,v) on every GPU, while keeping queries (q) lo- cally sharded. Since the KV cache is small, the communication cost of this gather is modest. The result is a favorable trade-off that lets us balance compute across 8 GPUs. As shown in Fig. 5, we perform an all-to-all collective before the first NSABlock to distribute token blocks globally, and then invoke all-gather on keys and values before SelAttn and CmpAttn. 3.6 Implementation and Training Details We train LSRM for novel-view synthesis (NVS) on 600K curated 3D models and then fine-tune it for inverse rendering. Each iteration uses 12–16 input images to render 4 target views. NVS training follows three stages: (1) We train the dense reconstruction transformer (S vol d = 16,S img d = 32), yielding a decoded volume resolution of S vol df = 64 for 256× 256 inputs. (2) We freeze the dense model and use its predictions to train the sparse reconstruction transformer on the fly. We start at lower resolutions (S img = 64,S vol = 64) to support larger batch sizes and quicker convergence. (3) We scale sparse resolutions to S img = 96,S vol = 96. This final stage relies on our custom block-aware sequence parallelism for stable training, with active image and volume token counts peaking at roughly 100K and 300K, respectively. Compared to recent SOTA like LIRM [41]—which processes a maximum of 8 images (64× 64× 8 ≈ 32K image tokens) and uses a hexaplane representation (48× 48× 6≈ 14K object tokens)—LSRM scales to > 2× more image tokens and > 20× more object tokens. These three stages require 5, 7, and 3 days, respectively, on 128 H100 GPUs. Across all stages, we optimize rendering and depth losses, adding a numerical normal loss [41] during the final 500–1000 iterations to further boost geometric fidelity. Because our goal is to evaluate context-window scaling, we reduce compute by initializing inverse rendering from the converged NVS weights. We apply two task-specific architectural changes: (1) we encode background images with a lin- ear layer [41] to better disentangle lighting from materials, and (2) we initialize the BRDF prediction head with pre-trained weights from MLP c . We reuse the same 600K 3D models, rendering them under real HDR environment maps with extensive augmentations. Fine-tuning mirrors the three-stage NVS schedule. For efficiency, we initialize the first two stages directly from their NVS counterparts 12Z. Li et al. LIRMOurs (LSRM)GTLIRMOurs (LSRM)GT Fig. 6: Qualitative NVS Results on the GSO Dataset. LSRM significantly im- proves texture fidelity over prior SOTA [41]. Zoomed-in comparisons show that our method successfully recovers legible text (bottom left), complex geometry (top left) and sharp facial structures (right), whereas the baseline produces blurred artifacts. and halve the training iterations, while the final stage initializes from the con- verged second stage. 4 Experiments Object-centric Novel View Synthesis We evaluate the NVS quality of LSRM on the GSO dataset [16]. Inputs are 16 uniformly sampled 768× 768 views ren- dered with the camera configuration of [41], and we render 12 768× 768 novel views for evaluation. Following standard practice, ground-truth novel views are rendered at 512×512, and we downsample our predictions to this resolution be- fore computing metrics. We compute PSNR, SSIM and LPIPS on the rendered RGB images, averaging over all objects and target views, and we use identi- cal evaluation scripts across methods. Quantitative and qualitative results are summarized in Tab. 1 and Fig. 7, respectively. In Tab. 1, we compare LSRM against state-of-the-art feed-forward baselines. LIRM [41] uses the same number of input images via iterative refinement, but operates at a lower 512× 512 reso- lution. By contrast, LSRM increases the active image and volume token counts during reconstruction, yielding a substantial boost in overall fidelity. Our full model improves PSNR by > 2.4 dB and reduces LPIPS by > 40% relative to the strongest baseline. These gains are also evident qualitatively: LSRM pro- duces sharper textures and recovers recognizable facial structure (e.g., the Little Mermaid and the bee). Even for objects with complex geometry (transformers, top left and bottom right), zoomed-in crops are nearly indistinguishable from ground truth. We additionally perform ablations to validate key components, evaluating both context-window scaling and our 3D-aware block routing. Re- sults in Tab. 1 and Fig. 7 show consistent improvements, restoring previously distorted faces. Overall, context scaling provides the larger gain, while 3D-aware routing primarily sharpens fine-grained structures. High-Fidelity Object-Centric Reconstruction via Scaled Context Windows13 Table 1: Quantitative NVS Results on the GSO Dataset. We compare our approach against SOTA feed-forward baselines [41,63,79]. Our LSRM ablations eval- uate the dense representation against sparse representations at different training steps (S img = S vol = 64 and S img = S vol = 64), both with and without our novel 3D-aware block routing. Best, second best, and third best results are highlighted in red, orange, and yellow, respectively. BaselinesOurs (LSRM Ablations) Metric MeshGS LIRM Dense S img = S vol = 64S img = S vol = 96 -LRM -LRMmodel w/o routing w/ routing w/o routing w/ routing PSNR (↑) 28.13 30.52 30.65 28.4531.3431.9332.7233.08 SSIM (↑) 0.923 0.952 0.949 0.9340.9560.9620.9680.971 LPIPS (↓) 0.093 0.050 0.054 0.0790.0450.0400.0320.028 LSRM Full GT LSRM Full GT w/o 3D-aware routing Fig. 7: Qualitative Ablation Study on the GSO Dataset. Visual comparisons demonstrate that both our 3D-aware block routing and increased spatial resolutions independently enhance texture fidelity and overall rendering quality. Object-centric Inverse Rendering We evaluate LSRM’s inverse-rendering exten- sion on three widely used benchmarks: StanfordORB [29], DigitalTwinCatalogue (DTC) [15], and ObjectsWithLighting (OWL) [58]. These datasets span diverse materials, illumination, and capture setups, offering a stringent test of relighta- bility, texture fidelity, and geometric consistency. As summarized in Tab. 2, LSRM consistently clearly outperforms prior feed-forward baselines. However, gains in pixel-wise metrics such as PSNR are comparatively modest. This reflects the inherently ill-posed nature of inverse rendering: although LSRM disentangles shadows and specular highlights to recover high-fidelity details, global ambigui- ties—e.g., slight color shifts (predicting a deep blue versus a light blue roof for the birdhouse; bottom left of Fig. 8)—can disproportionately penalize PSNR. In contrast, on perceptual metrics like LPIPS, which emphasize high-frequency structure, LSRM shows substantial improvements, reaching performance com- parable to SOTA optimization-based methods. Qualitative results in Fig. 8 corroborate these gains in texture fidelity. For example, LSRM recovers legible text (the “Nutrition Facts” and “Calories 0” label on the can, top right) and fine textures (the golden leaves and flowers on the teapot, top right and bottom left). Beyond appearance, LSRM also improves geometry, accurately recovering the small fence in front of the birdhouse. This geometric accuracy is reflected in StanfordORB Chamfer Distance (CD), which drops markedly relative to LIRM and optimization-based approaches. 14Z. Li et al. Table 2: Quantitative Inverse Rendering Results. We evaluate our approach against optimization-based and feed-forward baselines across three benchmarks. No- tably, LSRM consistently achieves the lowest LPIPS across all datasets, a metric highly sensitive to fine texture quality and perceptual fidelity. Method StanfordORB [29]DigitalTwinCatalogue [15]ObjectsWithLighting [58] PSNR-H ↑ PSNR-L ↑ SSIM ↑ LPIPS ↓ CD ↓ PSNR-H ↑ PSNR-L ↑ SSIM ↑ LPIPS ↓ PSNR ↑ SSIM ↑ LPIPS ↓ NVDiffrecMc24.4331.600.972 0.036 0.5127.7834.550.952 0.04219.820.730.389 InvRender23.7630.830.970 0.046 0.4429.5235.980.9610.03723.770.780.369 NeuralPBIR26.0133.260.9790.0230.43N/AN/AN/AN/AN/AN/AN/A LIRM25.0932.450.9720.0250.3827.6534.840.9600.03123.270.770.322 Ours (LSRM)25.4732.850.9770.0220.2929.6735.660.9640.02524.880.790.269 Ours (LSRM)GTGTOurs (LSRM)LIRMNeuralPBIRInvRenderLIRM Fig. 8: Qualitative Inverse Rendering Results. Visual comparisons against optimization-based and feed-forward baselines demonstrate that LSRM consistently recovers higher-fidelity textures and finer geometric details. Generalization to Smaller Number of Views Since the primary focus of LSRM is to significantly scale the context window for both image and volume tokens, we maintain the number of input images between 12 and 16 during training and do not fine-tune our models on fewer images. Surprisingly, both our feed-forward 3D reconstruction and inverse rendering models generalize well to a smaller number of input views, outperforming the prior state-of-the-art, LIRM [41], which was explicitly trained on 4 to 8 images. We evaluate this generalization capability for novel-view synthesis on the GSO dataset [16]. As summarized in Tab. 3 and Fig. 9, providing fewer views causes only a slight drop in texture sharpness compared to the full 16-view input. However, our results remain substantially superior to LIRM whether using 8 or 16 input views, a conclusion clearly supported by the quantitative metrics. Specifically, our 8-view results achieve a 2 dB increase in PSNR compared to the 16-view LIRM, which corresponds to an approximate 37% drop in MSE loss, while the LPIPS loss drops by more than 40%. We observe a similar phenomenon in inverse rendering, where LSRM consistently outperforms LIRM using only 6 input views. Furthermore, evaluating on fewer views enables a fair comparison with RelitLRM [84], another feed-forward re- lighting model. Unlike our approach, RelitLRM does not explicitly predict ma- terial reflectance; rather, it relies on a generative network to synthesize novel lighting appearances, resulting in slower inference. Nevertheless, LSRM outper- High-Fidelity Object-Centric Reconstruction via Scaled Context Windows15 LIRM 18 views LSRM 6 views LSRM 18 views GT LIRM 16 views LSRM 8 views LSRM 16 views GT 3D Reconstruction with Different Number of Input ViewsInverse Rendering with Different Number of Input Views Fig. 9: Qualitative 3D reconstruction and inverse rendering results with different num- ber of input views. Metric LIRM [41]Ours (LSRM) 8 views 16 views 8 views 16 views PSNR (↑) 30.4830.5632.4633.08 SSIM (↑) 0.947 0.9480.9680.971 LPIPS (↓) 0.056 0.0540.0310.028 Table 3: Quantitative novel view synthesis results on the GSO dataset [16] with differ- ent number of input views. Table 4: Quantitative inverse rendering results with fewer number of views. Views Method StanfordORB [29]DigitalTwinCatalogue [15]ObjectsWithLighting [58] PSNR-H ↑ PSNR-L ↑ SSIM ↑ LPIPS ↓ CD ↓ PSNR-H ↑ PSNR-L ↑ SSIM ↑ LPIPS ↓ PSNR ↑ SSIM ↑ LPIPS ↓ 6 RelitLRM24.6731.520.9690.032 N/AN/AN/AN/AN/A23.080.790.284 LIRM24.7632.110.9710.0270.4828.0334.280.9590.03422.700.760.326 Ours (LSRM)25.5732.740.9750.0230.3229.0935.160.9610.02824.430.800.255 18 Ours (LSRM) 25.4732.850.978 0.022 0.2929.6735.660.965 0.02524.880.790.269 forms both RelitLRM and LIRM across two popular relighting benchmarks (see Tab. 4), with qualitative results in Fig. 9 corroborating these improvements. 5 Conclusions In this work, we show that scaling transformer context windows markedly im- proves feed-forward 3D reconstruction and inverse rendering. By scaling object and image tokens by 20× and 2× with NSA, we close much of the texture- and geometry-fidelity gap to optimization. To make this scale practical, we introduce 3D-aware block routing and block-aware sequence parallelism, boosting render- ing quality while improving GPU utilization. Extensive experiments demonstrate that LSRM sets a new SOTA, recovering crisp text and faithful facial structures. Limitations and Future Works LSRM introduces a general coarse-to-fine pipeline that leverages NSA to effectively scale context windows for 3D-related tasks. While this paper validates its effectiveness on object-centric 3D reconstruction and inverse rendering, we hope to extend this framework to large-scale scene re- construction and general video/image generation in future work. Despite achiev- ing SOTA reconstruction quality, several limitations remain. First, we observe that LSRM does not inherently resolve other ill-posed challenges in inverse ren- 16Z. Li et al. dering, such as accurately estimating roughness and metallic material proper- ties. Furthermore, we still encounter cases where extremely fine details—such as the small ingredient lists can-remain illegible. To address this, future iter- ations could expand the context window even further by integrating sequence parallelism techniques like Ulysses attention [24] and ring attention [48]. Acknowledgements We would like to thank David Clabaugh for creating the teaser figures, Dilin Wang and Yuchen Fan for their support with the training infrastructure, and Hao Tan, Sai Bi, and Kai Zhang for technical discussions regarding sparse attention. References 1. Ainslie, J., Lee-Thorp, J., De Jong, M., Zemlyanskiy, Y., Lebrón, F., Sanghai, S.: Gqa: Training generalized multi-query transformer models from multi-head check- points. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. p. 4895–4901 (2023) 2. Bi, S., Xu, Z., Sunkavalli, K., Kriegman, D., Ramamoorthi, R.: Deep 3d capture: Geometry and reflectance from sparse multi-view images. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 5960–5969 (2020) 3. Boss, M., Braun, R., Jampani, V., Barron, J.T., Liu, C., Lensch, H.: Nerd: Neural reflectance decomposition from image collections. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 12684–12694 (2021) 4. Boss, M., Engelhardt, A., Kar, A., Li, Y., Sun, D., Barron, J., Lensch, H., Jampani, V.: Samurai: Shape and material from unconstrained real-world arbitrary image collections. Advances in Neural Information Processing Systems 35, 26389–26403 (2022) 5. Boss, M., Jampani, V., Braun, R., Liu, C., Barron, J., Lensch, H.: Neural-pil: Neural pre-integrated lighting for reflectance decomposition. Advances in Neural Information Processing Systems 34, 10691–10704 (2021) 6. Careaga, C., Aksoy, Y.: Intrinsic image decomposition via ordinal shading. ACM Transactions on Graphics 43(1), 1–24 (2023) 7. Chen, A., Xu, H., Esposito, S., Tang, S., Geiger, A.: Lara: Efficient large-baseline radiance fields. In: European conference on computer vision. p. 338–355. Springer (2024) 8. Chen, X., Chu, F.J., Gleize, P., Liang, K.J., Sax, A., Tang, H., Wang, W., Guo, M., Hardin, T., Li, X., et al.: Sam 3d: 3dfy anything in images. arXiv preprint arXiv:2511.16624 (2025) 9. Chen, Z., Tang, J., Dong, Y., Cao, Z., Hong, F., Lan, Y., Wang, T., Xie, H., Wu, T., Saito, S., et al.: 3dtopia-xl: Scaling high-quality 3d asset generation via prim- itive diffusion. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 26576–26586 (2025) 10. Child, R., Gray, S., Radford, A., Sutskever, I.: Generating long sequences with sparse transformers. arXiv preprint arXiv:1904.10509 (2019) 11. Dao, T.: Flashattention-2: Faster attention with better parallelism and work par- titioning. arXiv preprint arXiv:2307.08691 (2023) 12. Debevec, P., Hawkins, T., Tchou, C., Duiker, H.P., Sarokin, W., Sagar, M.: Ac- quiring the reflectance field of a human face. In: Proceedings of the 27th annual conference on Computer graphics and interactive techniques. p. 145–156 (2000) High-Fidelity Object-Centric Reconstruction via Scaled Context Windows17 13. Deitke, M., Schwenk, D., Salvador, J., Weihs, L., Michel, O., VanderBilt, E., Schmidt, L., Ehsani, K., Kembhavi, A., Farhadi, A.: Objaverse: A universe of annotated 3d objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 13142–13153 (2023) 14. Deschaintre, V., Aittala, M., Durand, F., Drettakis, G., Bousseau, A.: Single-image svbrdf capture with a rendering-aware deep network. ACM Transactions on Graph- ics (ToG) 37(4), 1–15 (2018) 15. Dong, Z., Chen, K., Lv, Z., Yu, H.X., Zhang, Y., Zhang, C., Zhu, Y., Tian, S., Li, Z., Moffatt, G., et al.: Digital twin catalog: A large-scale photorealistic 3d object digital twin dataset. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 753–763 (2025) 16. Downs, L., Francis, A., Koenig, N., Kinman, B., Hickman, R., Reymann, K., McHugh, T.B., Vanhoucke, V.: Google scanned objects: A high-quality dataset of 3d scanned household items. In: 2022 International Conference on Robotics and Automation (ICRA). p. 2553–2560. IEEE (2022) 17. Engelhardt, A., Raj, A., Boss, M., Zhang, Y., Kar, A., Li, Y., Sun, D., Brualla, R.M., Barron, J.T., Lensch, H., et al.: Shinobi: Shape and illumination using neu- ral object decomposition via brdf optimization in-the-wild. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 19636– 19646 (2024) 18. Fang, J., Zhao, S.: Usp: A unified sequence parallelism approach for long context generative ai. arXiv preprint arXiv:2405.07719 (2024) 19. Goldman, D.B., Curless, B., Hertzmann, A., Seitz, S.M.: Shape and spatially- varying brdfs from photometric stereo. IEEE Transactions on Pattern Analysis and Machine Intelligence 32(6), 1060–1071 (2009) 20. Gupta, A., Xiong, W., Nie, Y., Jones, I., Oğuz, B.: 3dgen: Triplane latent diffusion for textured mesh generation. arXiv preprint arXiv:2303.05371 (2023) 21. Hasselgren, J., Hofmann, N., Munkberg, J.: Shape, light, and material decomposi- tion from images using monte carlo rendering and denoising. Advances in Neural Information Processing Systems 35, 22856–22869 (2022) 22. Hong, Y., Zhang, K., Gu, J., Bi, S., Zhou, Y., Liu, D., Liu, F., Sunkavalli, K., Bui, T., Tan, H.: Lrm: Large reconstruction model for single image to 3d. arXiv preprint arXiv:2311.04400 (2023) 23. Hunyuan3D, T., Yang, S., Yang, M., Feng, Y., Huang, X., Zhang, S., He, Z., Luo, D., Liu, H., Zhao, Y., et al.: Hunyuan3d 2.1: From images to high-fidelity 3d assets with production-ready pbr material. arXiv preprint arXiv:2506.15442 (2025) 24. Jacobs, S.A., Tanaka, M., Zhang, C., Zhang, M., Song, S.L., Rajbhandari, S., He, Y.: Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models. arXiv preprint arXiv:2309.14509 (2023) 25. Jiang, Y., Tu, J., Liu, Y., Gao, X., Long, X., Wang, W., Ma, Y.: Gaussianshader: 3d gaussian splatting with shading functions for reflective surfaces. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 5322–5332 (2024) 26. Jin, H., Jiang, H., Tan, H., Zhang, K., Bi, S., Zhang, T., Luan, F., Snavely, N., Xu, Z.: Lvsm: A large view synthesis model with minimal 3d inductive bias. arXiv preprint arXiv:2410.17242 (2024) 27. Jin, H., Liu, I., Xu, P., Zhang, X., Han, S., Bi, S., Zhou, X., Xu, Z., Su, H.: Tensoir: Tensorial inverse rendering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 165–174 (2023) 18Z. Li et al. 28. Korthikanti, V.A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., Catanzaro, B.: Reducing activation recomputation in large transformer models. Proceedings of Machine Learning and Systems 5, 341–353 (2023) 29. Kuang, Z., Zhang, Y., Yu, H.X., Agarwala, S., Wu, E., Wu, J., et al.: Stanford-orb: a real-world 3d object inverse rendering benchmark. Advances in Neural Information Processing Systems 36 (2024) 30. Lai, X., Lu, J.: native-sparse-attention-triton: Efficient triton implementation of Native Sparse Attention. https://github.com/XunhaoLai/native-sparse- attention-triton (2025) 31. Lan, Y., Hong, F., Zhou, S., Yang, S., Meng, X., Chen, Y., Lyu, Z., Dai, B., Pan, X., Loy, C.C.: Ln3diff++: Scalable latent neural fields diffusion for speedy 3d generation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2025) 32. Lan, Y., Zhou, S., Lyu, Z., Hong, F., Yang, S., Dai, B., Pan, X., Loy, C.C.: Gaus- siananything: Interactive point cloud latent diffusion for 3d generation (2025) 33. Leroy, V., Cabon, Y., Revaud, J.: Grounding image matching in 3d with mast3r. In: European conference on computer vision. p. 71–91. Springer (2024) 34. Li, W., Liu, J., Yan, H., Chen, R., Liang, Y., Chen, X., Tan, P., Long, X.: Crafts- man3d: High-fidelity mesh generation with 3d native generation and interactive geometry refiner. arXiv preprint arXiv:2405.14979 (2024) 35. Li, W., Zhang, X., Sun, Z., Qi, D., Li, H., Cheng, W., Cai, W., Wu, S., Liu, J., Wang, Z., et al.: Step1x-3d: Towards high-fidelity and controllable generation of textured 3d assets. arXiv preprint arXiv:2505.07747 (2025) 36. Li, X., Dong, Y., Peers, P., Tong, X.: Modeling surface appearance from a single photograph using self-augmented convolutional neural networks. ACM Transac- tions on Graphics (ToG) 36(4), 1–11 (2017) 37. Li, Y., Zou, Z.X., Liu, Z., Wang, D., Liang, Y., Yu, Z., Liu, X., Guo, Y.C., Liang, D., Ouyang, W., et al.: Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. IEEE Transactions on Pattern Analysis and Machine Intel- ligence (2025) 38. Li, Z., Snavely, N.: Cgintrinsics: Better intrinsic image decomposition through physically-based rendering. In: Proceedings of the European conference on com- puter vision (ECCV). p. 371–387 (2018) 39. Li, Z., Shi, J., Bi, S., Zhu, R., Sunkavalli, K., Hašan, M., Xu, Z., Ramamoorthi, R., Chandraker, M.: Physically-based editing of indoor scene lighting from a single image. In: European Conference on Computer Vision. p. 555–572. Springer (2022) 40. Li, Z., Sunkavalli, K., Chandraker, M.: Materials for masses: Svbrdf acquisition with a single mobile phone image. In: Proceedings of the European conference on computer vision (ECCV). p. 72–87 (2018) 41. Li, Z., Wang, D., Chen, K., Lv, Z., Nguyen-Phuoc, T., Lee, M., Huang, J.B., Xiao, L., Zhu, Y., Marshall, C.S., et al.: Lirm: Large inverse rendering model for progressive reconstruction of shape, materials and view-dependent radiance fields. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 505–517 (2025) 42. Li, Z., Xu, Z., Ramamoorthi, R., Sunkavalli, K., Chandraker, M.: Learning to re- construct shape and spatially-varying reflectance from a single image. ACM Trans- actions on Graphics (TOG) 37(6), 1–11 (2018) 43. Li, Z., Wang, Y., Zheng, H., Luo, Y., Wen, B.: Sparc3d: Sparse representa- tion and construction for high-resolution 3d shapes modeling. arXiv preprint arXiv:2505.14521 (2025) High-Fidelity Object-Centric Reconstruction via Scaled Context Windows19 44. Liang, H., Ren, J., Mirzaei, A., Torralba, A., Liu, Z., Gilitschenski, I., Fidler, S., Oztireli, C., Ling, H., Gojcic, Z., et al.: Feed-forward bullet-time reconstruction of dynamic scenes from monocular videos. arXiv preprint arXiv:2412.03526 (2024) 45. Liang, Z., Zhang, Q., Feng, Y., Shan, Y., Jia, K.: Gs-ir: 3d gaussian splatting for inverse rendering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 21644–21653 (2024) 46. Lin, C.H., Lv, Z., Wu, S., Xu, Z., Nguyen-Phuoc, T., Tseng, H.Y., Straub, J., Khan, N., Xiao, L., Yang, M.H., et al.: Dgs-lrm: Real-time deformable 3d gaussian reconstruction from monocular videos. arXiv preprint arXiv:2506.09997 (2025) 47. Lin, H., Chen, S., Liew, J., Chen, D.Y., Li, Z., Shi, G., Feng, J., Kang, B.: Depth anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647 (2025) 48. Liu, H., Zaharia, M., Abbeel, P.: Ring attention with blockwise transformers for near-infinite context. arXiv preprint arXiv:2310.01889 (2023) 49. Ma, Z., Chen, X., Yu, S., Bi, S., Zhang, K., Ziwen, C., Xu, S., Yang, J., Xu, Z., Sunkavalli, K., et al.: 4d-lrm: Large space-time reconstruction model from and to any view at any time. arXiv preprint arXiv:2506.18890 (2025) 50. Meka, A., Maximov, M., Zollhoefer, M., Chatterjee, A., Seidel, H.P., Richardt, C., Theobalt, C.: Lime: Live intrinsic material estimation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. p. 6315–6324 (2018) 51. Munkberg, J., Hasselgren, J., Shen, T., Gao, J., Chen, W., Evans, A., Müller, T., Fidler, S.: Extracting triangular 3d models, materials, and lighting from images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 8280–8290 (2022) 52. Qiu, Z., Wang, Z., Zheng, B., Huang, Z., Wen, K., Yang, S., Men, R., Yu, L., Huang, F., Huang, S., et al.: Gated attention for large language models: Non- linearity, sparsity, and attention-sink-free. arXiv preprint arXiv:2505.06708 (2025) 53. Shi, Y., Wu, Y., Wu, C., Liu, X., Zhao, C., Feng, H., Liu, J., Zhang, L., Zhang, J., Zhou, B., et al.: Gir: 3d gaussian inverse rendering for relightable scene factoriza- tion. arXiv preprint arXiv:2312.05133 (2023) 54. Siddiqui, Y., Monnier, T., Kokkinos, F., Kariya, M., Kleiman, Y., Garreau, E., Gafni, O., Neverova, N., Vedaldi, A., Shapovalov, R., et al.: Meta 3d assetgen: Text-to-mesh generation with high-quality geometry, texture, and pbr materials. arXiv preprint arXiv:2407.02445 (2024) 55. Siméoni, O., Vo, H.V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khali- dov, V., Szafraniec, M., Yi, S., Ramamonjisoa, M., et al.: Dinov3. arXiv preprint arXiv:2508.10104 (2025) 56. Sun, C., Cai, G., Li, Z., Yan, K., Zhang, C., Marshall, C., Huang, J.B., Zhao, S., Dong, Z.: Neural-pbir reconstruction of shape, material, and illumination. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 18046–18056 (2023) 57. Tillet, P.: Introducing Triton: Open-source GPU programming for neural networks. https://openai.com/index/triton/ (July 2021), openAI Blog 58. Ummenhofer, B., Agrawal, S., Sepulveda, R., Lao, Y., Zhang, K., Cheng, T., Richter, S., Wang, S., Ros, G.: Objects with lighting: A real-world dataset for evaluating reconstruction and rendering for object relighting. In: 2024 Interna- tional Conference on 3D Vision (3DV). p. 137–147. IEEE (2024) 59. Vahdat, A., Williams, F., Gojcic, Z., Litany, O., Fidler, S., Kreis, K., et al.: Lion: Latent point diffusion models for 3d shape generation. Advances in neural infor- mation processing systems 35, 10021–10039 (2022) 20Z. Li et al. 60. Vaswani, A.: Attention is all you need. Advances in Neural Information Processing Systems (2017) 61. Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Visual geometry grounded transformer. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 5294–5306 (2025) 62. Wang, S., Leroy, V., Cabon, Y., Chidlovskii, B., Revaud, J.: Dust3r: Geometric 3d vision made easy. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 20697–20709 (2024) 63. Wei, X., Zhang, K., Bi, S., Tan, H., Luan, F., Deschaintre, V., Sunkavalli, K., Su, H., Xu, Z.: Meshlrm: Large reconstruction model for high-quality mesh. arXiv preprint arXiv:2404.12385 (2024) 64. Wu, S., Lin, Y., Zhang, F., Zeng, Y., Xu, J., Torr, P., Cao, X., Yao, Y.: Direct3d: Scalable image-to-3d generation via 3d latent diffusion transformer. Advances in Neural Information Processing Systems 37, 121859–121881 (2024) 65. Wu, S., Lin, Y., Zhang, F., Zeng, Y., Yang, Y., Bao, Y., Qian, J., Zhu, S., Cao, X., Torr, P., et al.: Direct3d-s2: Gigascale 3d generation made easy with spatial sparse attention. arXiv preprint arXiv:2505.17412 (2025) 66. Xiang, J., Chen, X., Xu, S., Wang, R., Lv, Z., Deng, Y., Zhu, H., Dong, Y., Zhao, H., Yuan, N.J., et al.: Native and compact structured latents for 3d generation. arXiv preprint arXiv:2512.14692 (2025) 67. Xiang, J., Lv, Z., Xu, S., Deng, Y., Wang, R., Zhang, B., Chen, D., Tong, X., Yang, J.: Structured 3d latents for scalable and versatile 3d generation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 21469–21480 (2025) 68. Xu, Y., Shi, Z., Yifan, W., Chen, H., Yang, C., Peng, S., Shen, Y., Wetzstein, G.: Grm: Large gaussian reconstruction model for efficient 3d reconstruction and gen- eration. In: European Conference on Computer Vision. p. 1–20. Springer (2024) 69. Xu, Z., Li, Z., Dong, Z., Zhou, X., Newcombe, R., Lv, Z.: 4dgt: Learning a 4d gaussian transformer using real-world monocular videos. arXiv preprint arXiv:2506.08015 (2025) 70. Yang, H., Dong, Y., Jiang, H., Xu, D., Pavlakos, G., Huang, Q.: Atlas gaussians diffusion for 3d generation. arXiv preprint arXiv:2408.13055 (2024) 71. Yang, Z., Chen, Y., Gao, X., Yuan, Y., Wu, Y., Zhou, X., Jin, X.: Sire-ir: Inverse rendering for brdf reconstruction with shadow and illumination removal in high- illuminance scenes. arXiv preprint arXiv:2310.13030 (2023) 72. Yariv, L., Gu, J., Kasten, Y., Lipman, Y.: Volume rendering of neural implicit surfaces. Advances in Neural Information Processing Systems 34, 4805–4815 (2021) 73. Yuan, J., Gao, H., Dai, D., Luo, J., Zhao, L., Zhang, Z., Xie, Z., Wei, Y., Wang, L., Xiao, Z., et al.: Native sparse attention: Hardware-aligned and natively trainable sparse attention. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). p. 23078–23097 (2025) 74. Zaheer, M., Guruganesh, G., Dubey, K.A., Ainslie, J., Alberti, C., Ontanon, S., Pham, P., Ravula, A., Wang, Q., Yang, L., et al.: Big bird: Transformers for longer sequences. Advances in neural information processing systems 33, 17283–17297 (2020) 75. Zarzar, J., Monnier, T., Shapovalov, R., Vedaldi, A., Novotny, D.: Twinner: Shining light on digital twins in a few snaps. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 5859–5869 (2025) 76. Zeng, Z., Deschaintre, V., Georgiev, I., Hold-Geoffroy, Y., Hu, Y., Luan, F., Yan, L.Q., Hašan, M.: Rgbx: Image decomposition and synthesis using material- and High-Fidelity Object-Centric Reconstruction via Scaled Context Windows21 lighting-aware diffusion models. In: ACM SIGGRAPH 2024 Conference Papers. SIGGRAPH ’24, Association for Computing Machinery, New York, NY, USA (2024). https://doi.org/10.1145/3641519.3657445, https://doi.org/10. 1145/3641519.3657445 77. Zhang, B., Tang, J., Niessner, M., Wonka, P.: 3dshape2vecset: A 3d shape repre- sentation for neural fields and generative diffusion models. ACM Transactions On Graphics (TOG) 42(4), 1–16 (2023) 78. Zhang, C., Moing, G.L., Koppula, S., Rocco, I., Momeni, L., Xie, J., Sun, S., Sukthankar, R., Barral, J.K., Hadsell, R., et al.: Efficiently reconstructing dynamic scenes one d4rt at a time. arXiv preprint arXiv:2512.08924 (2025) 79. Zhang, K., Bi, S., Tan, H., Xiangli, Y., Zhao, N., Sunkavalli, K., Xu, Z.: Gs-lrm: Large reconstruction model for 3d gaussian splatting. In: European Conference on Computer Vision. p. 1–19. Springer (2025) 80. Zhang, K., Luan, F., Li, Z., Snavely, N.: Iron: Inverse rendering by optimizing neu- ral sdfs and materials from photometric images. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 5565–5574 (2022) 81. Zhang, K., Luan, F., Wang, Q., Bala, K., Snavely, N.: Physg: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In: Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 5453–5462 (2021) 82. Zhang, L., Wang, Z., Zhang, Q., Qiu, Q., Pang, A., Jiang, H., Yang, W., Xu, L., Yu, J.: Clay: A controllable large-scale generative model for creating high-quality 3d assets. ACM Transactions on Graphics (TOG) 43(4), 1–20 (2024) 83. Zhang, T., Bi, S., Hong, Y., Zhang, K., Luan, F., Yang, S., Sunkavalli, K., Freeman, W.T., Tan, H.: Test-time training done right. arXiv preprint arXiv:2505.23884 (2025) 84. Zhang, T., Kuang, Z., Jin, H., Xu, Z., Bi, S., Tan, H., Zhang, H., Hu, Y., Hasan, M., Freeman, W.T., et al.: Relitlrm: Generative relightable radiance for large re- construction models. arXiv preprint arXiv:2410.06231 (2024) 85. Zhang, X., Srinivasan, P.P., Deng, B., Debevec, P., Freeman, W.T., Barron, J.T.: Nerfactor: Neural factorization of shape and reflectance under an unknown illumi- nation. ACM Transactions on Graphics (ToG) 40(6), 1–18 (2021) 86. Zhang, Y., Xu, T., Yu, J., Ye, Y., Jing, Y., Wang, J., Yu, J., Yang, W.: Nemf: Inverse volume rendering with neural microflake field. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 22919–22929 (2023) 87. Zhang, Y., Sun, J., He, X., Fu, H., Jia, R., Zhou, X.: Modeling indirect illumination for inverse rendering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 18643–18652 (2022) 88. Zhao, Z., Liu, W., Chen, X., Zeng, X., Wang, R., Cheng, P., Fu, B., Chen, T., Yu, G., Gao, S.: Michelangelo: Conditional 3d shape generation based on shape-image- text aligned latent representation. Advances in neural information processing sys- tems 36, 73969–73982 (2023) 89. Ziwen, C., Tan, H., Zhang, K., Bi, S., Luan, F., Hong, Y., Fuxin, L., Xu, Z.: Long- lrm: Long-sequence large reconstruction model for wide-coverage gaussian splats. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 4349–4359 (2025)