Paper deep dive
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
Jorge Condor, Nicolas Moenne-Loccoz, Merlin Nimier-David, Piotr Didyk, Zan Gojcic, Qi Wu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/2/2026, 3:34:18 AM
Summary
Neural Harmonic Textures is a novel primitive-based neural representation for high-quality reconstruction and novel-view synthesis. It anchors latent feature vectors to virtual scaffolds surrounding primitives, applying periodic activations (harmonic encoding) to interpolated features. This approach enables efficient deferred shading via a lightweight MLP, bridging the gap between primitive-based methods like 3D Gaussian Splatting and continuous neural radiance fields.
Entities (5)
Relation Signals (3)
Neural Harmonic Textures â integrateswith â 3DGUT
confidence 95% · Our method integrates seamlessly into existing primitive-based pipelines such as 3DGUT
Neural Harmonic Textures â uses â Fourier analysis
confidence 95% · Inspired by Fourier analysis, we apply periodic activations to the interpolated features
Neural Harmonic Textures â improves â 3D Gaussian Splatting
confidence 90% · Neural Harmonic Textures yield state-of-the-art results in real-time novel view synthesis while bridging the gap between primitive- and neural-field-based reconstruction.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Primitive-based methods such as 3D Gaussian Splatting have recently become the state-of-the-art for novel-view synthesis and related reconstruction tasks. Compared to neural fields, these representations are more flexible, adaptive, and scale better to large scenes. However, the limited expressivity of individual primitives makes modeling high-frequency detail challenging. We introduce Neural Harmonic Textures, a neural representation approach that anchors latent feature vectors on a virtual scaffold surrounding each primitive. These features are interpolated within the primitive at ray intersection points. Inspired by Fourier analysis, we apply periodic activations to the interpolated features, turning alpha blending into a weighted sum of harmonic components. The resulting signal is then decoded in a single deferred pass using a small neural network, significantly reducing computational cost. Neural Harmonic Textures yield state-of-the-art results in real-time novel view synthesis while bridging the gap between primitive- and neural-field-based reconstruction. Our method integrates seamlessly into existing primitive-based pipelines such as 3DGUT, Triangle Splatting, and 2DGS. We further demonstrate its generality with applications to 2D image fitting and semantic reconstruction.
Tags
Links
- Source: https://arxiv.org/abs/2604.01204v1
- Canonical: https://arxiv.org/abs/2604.01204v1
Trouble viewing inline? Open PDF directly â
Full Text
83,467 characters extracted from source content.
Expand or collapse full text
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction Jorge Condor 1,2 , Nicolas MoĂ«nne-Loccoz 1 , Merlin Nimier-David 1 , Piotr Didyk 2 , Zan Gojcic 1 , and Qi Wu 1 1 NVIDIA nicolasm,mnimierdavid,zgojcic,qiwu@nvidia.com 2 UniversitĂ della Svizzera italiana, Lugano, Switzerland jorge.condor,piotr.didyk@usi.ch Fig. 1: Neural Harmonic Textures for novel view synthesis. We attach learnable feature vectors (right) to the virtual vertices of bounding tetrahedra encapsulating each primitive (center). After harmonic encoding and accumulation along the ray, a small neural network decodes the resulting signal into RGB color in a deferred manner (left). Abstract. Primitive-based methods such as 3D Gaussian Splatting have recently become the state-of-the-art for novel-view synthesis and related reconstruction tasks. Compared to neural fields, these representations are more flexible, adaptive, and scale better to large scenes. However, the limited expressivity of individual primitives makes modeling high- frequency detail challenging. We introduce Neural Harmonic Textures, a neural representation approach that anchors latent feature vectors on a virtual scaffold surrounding each primitive. These features are interpo- lated within the primitive at ray intersection points. Inspired by Fourier analysis, we apply periodic activations to the interpolated features, turn- ing alpha blending into a weighted sum of harmonic components. The resulting signal is then decoded in a single deferred pass using a small neural network, significantly reducing computational cost. Neural Har- monic Textures yield state-of-the-art results in real-time novel view syn- thesis while bridging the gap between primitive- and neural-field-based arXiv:2604.01204v1 [cs.CV] 1 Apr 2026 2J. Condor et al. reconstruction. Our method integrates seamlessly into existing primitive- based pipelines such as 3DGUT, Triangle Splatting, and 2DGS. We fur- ther demonstrate its generality with applications to 2D image fitting and semantic reconstruction. 1 Introduction Since the introduction of 3D Gaussian Splatting [24], (Lagrangian) primitive- based approaches have largely displaced (Eulerian) continuous neural radiance fields [44,50,56] for novel view synthesis [20,24]. This shift is driven not only by significantly improved rendering speed, but also by the structural advantages of explicit representations. Primitive-based methods naturally adapt to scene detail, scale more gracefully, and readily sup- port motion, deformation, and editing [5,62]. Furthermore, they align well with feed-forward reconstruction pipelines [31,72] and closely resemble widely adopted point-map representations [59,61]. Despite these advantages, the limited expressive power of individual primi- tives remains a fundamental bottleneck. Geometry and appearance are tightly coupled within each primitive, forcing high-frequency spatial detail to be repre- sented by increasing the number of primitives, which directly increases memory consumption and hampers rendering speed. Directional appearance modeling is constrained in a similar manner. View-dependent effects are typically repre- sented using low-order Spherical Harmonics, which are spectrally band-limited and poorly suited for high-frequency phenomena such as sharp specular high- lights. Alternative angular parametric functions, such as spherical Gaussians, spherical Beta kernels [33], and spherical Voronoi [9] increase directional band- width, but leave the fundamental coupling between spatial support and appear- ance modeling unchanged. In contrast, neural radiance fields achieve high local representational capacity by combining positional encodings such as Fourier fea- tures [41,57] or multi-resolution hash grids [44] with a learned multi-layer percep- tron (MLP) decoding. Recent works [37,73] attempted to bridge both paradigms by employing primitives as acceleration structures for querying a global neural field. However, these hybrid approaches inherit key limitations of global neu- ral representations, particularly with respect to motion, deformation, editability and scalability. In this work, we introduce Neural Harmonic Textures, a novel primitive-based neural representation that increases the expressive power of individual primitives while retaining the advantages of Lagrangian representations. To this end, we reinterpret primitives as both geometric carriers and local positional encodings. Specifically, we attach learnable feature vectors to a virtual scaffold encapsulat- ing each primitive (e.g., Gaussians or triangles) and interpolate them locally at rayâprimitive intersection points. Compared to globally anchored feature repre- sentations, our formulation moves with the primitives and therefore naturally supports operations such as motion, deformation, and scene editing. In prior work, such interpolated features are typically processed by a lightweight MLP to decode RGB values, followed by alpha compositing across intersected Neural Harmonic Textures3 primitives to produce the final color, requiring numerous neural evaluations. In- stead, we draw inspiration from the Fourier transform, where a complex signal is expressed as a sum of harmonic components (i.e., periodic functions with dif- ferent amplitudes). We therefore apply periodic functions (sine and cosine) to the interpolated feature vectors before compositing them along the ray. Under this formulation, the activated features act as frequency components, while the primitive opacity modulates their amplitude (Fig. 2). The resulting implicit sig- nal in image space is sufficiently rich to be directly decoded using a lightweight MLP in a single evaluation per pixel, in a deferred-shading manner. We validate Neural Harmonic Textures on standard novel view synthesis benchmarks, achieving state-of-the-art performance among both real-time and offline methods (Sec. 5.1). Our proposed formulation is agnostic to the under- lying primitive type: triangles, 2D or 3D Gaussians, tetrahedra, and other ele- ments can be used interchangeably, enabling seamless integration into existing primitive-based pipelines while significantly increasing their representational ca- pacity (Sec. 5.2). Finally, by abstracting the encoded signal, our approach natu- rally extends to higher-dimensional signals. Joint modeling of color and semantic features becomes possible within a unified local representation, allowing us to exploit both spatial and cross-modal correlations for improved efficiency and compactness (Sec. 5.2). 2 Related Work 2.1 Neural Fields Neural fields have become a foundational model for representing complex, multi- scale signals across a wide range of domains [67], including audio [36], images [39, 49], 3D geometry [30, 46, 55, 60], radiance fields [41], scientific data [12, 63, 64], and volumetric video [11,47]. These methods primarily differ in how they encode their inputs and parameterize their outputs. In most formulations, latent fea- tures are stored in a compressed representation and subsequently decoded by a neural network. This storage can be implicit within the network weights [34,35, 41,43,48,51,57], or explicit through dedicated spatial data structures. Explicit representations typically offer improved scalability and performance, though of- ten at increased memory cost. Common examples include voxel grids [39], tri- planes [4, 11], point clouds [14], hash grids [44], and meshes [58]. In addition, explicit spatial structures are frequently used as acceleration mechanisms to skip empty regions and efficiently identify relevant query locations [32, 44, 73]. More recently, neural fields have been combined with primitive-based methods by turning the primitives into tiny neural fields with closed-form antiderivatives, enhancing their individual modelling power [77]. Beyond the choice of spatial representation, prior work has shown that posi- tional encodings play a crucial role in enabling neural networks to represent high- frequency signals in low-dimensional spatial domains [41,44,51,57]. In contrast to globally anchored spatial structures such as hierarchical voxel grids (hashgrid encodings), which follow an Eulerian formulation, our approach is Lagrangian. 4J. Condor et al. We anchor latent feature vectors relative to explicit particles, enabling adaptive spatial resolution and naturally supporting explicit motion, deformation, and scene editing. 2.2 Primitive-based Representations for Neural Reconstruction Primitive-based methods have recently gained significant traction for 3D recon- struction and novel view synthesis, largely driven by the success of 3D Gaussian Splatting [24]. Numerous extensions have since explored variants of Gaussian primitives [7,8,21,28,33,42,62,65] as well as alternative representations, including Voronoi cells [13], triangle soups [20], and connected meshes [37]. While offering high rendering performance, editability, and interpretability, these approaches typically exhibit poor memory scaling when fitting high-frequency detail due to their joint modeling of geometry and appearance. Neural field methods, in contrast, aggregate information across the entire scene, enabling compact representations. Concurrent work [38, 73] attempts to bridge this gap by using geometric primitives as acceleration structures for neu- ral field evaluation. However, these approaches lose some of the advantages of purely primitive-based representations, as the learned features remain globally anchored (e.g. through hashgrids) and do not follow the primitives under motion or deformation. Furthermore, they inherit the limitations of Eulerian encodings, namely poor scaling in large scenes and high-frequency detail modelling. In con- trast, we propose that geometric primitives themselves, when appropriately pa- rameterized, can serve as an effective positional encoding mechanism, and enable more efficient signal decoding (i.e. deferred shading). 2.3 Deferred Shading and Texturing in Radiance fields Alternatives to direct color evaluation in novel view synthesis have been explored to improve expressivity, rendering quality, and efficiency. Early work introduced image-based neural deferred shading conditioned on 3D meshes [58]. Later ap- proaches extended this idea to neural radiance fields by baking appearance fea- tures into extracted geometry for real-time rendering [6, 19, 69], removing the necessity to run neural networks during runtime. However, these methods rely on multi-stage pipelines that first optimize a global field, then extract geometry and bake appearance features. In contrast, our method jointly optimizes primi- tives and neural appearance from scratch, preserving volumetric flexibility and eliminating the need for a separate baking stage, while retaining high quality and speed. Deferred shading has also been explored in the context of radiance field ren- dering using neural networks as deferred shaders. In Han et al. [16], global fea- ture grids are queried, with features accumulated and decoded once in a deferred manner. Similarly, point-based neural renderers have recently mixed local neu- ral decoding with deferred shading. Hahlbohm et al. [15] uses point-clouds as query points for a global hash-grid encoded feature field, decoded locally using a shallow neural network into implicit features. These are splatted to the camera Neural Harmonic Textures5 plane and finally rendered through a convolutional neural network. However, the necessity of large-bandwidth neural networks to uplift the poor scalability and detail afforded by the spatial encoding made them slow and difficult to train; the former requiring guiding networks with regular NeRF-style rendering to steer the optimization, while the latter required expensive perceptual-loss supervision. To improve expressivity without global feature grids, several primitive-based methods introduce per-primitive textures [3, 22, 54, 66, 68, 75]. However, highly expressive primitives can complicate optimization and degrade reconstruction quality. Instead, we decouple geometry and appearance by pairing local per- particle features with a shared lightweight neural decoder, in turn enabling the reduction of queries to the neural field to a single deferred pass. This allows particle density to more closely follow geometric complexity rather than being artificially increased to reproduce high-frequency appearance. 3 Preliminaries We review primitive-based radiance fields, primitive-based volume rendering, and positional encodings, focusing on the aspects most relevant to our method. Primitive-based Radiance Fields. Primitive-based radiance fields represent com- plex scenes using an unstructured collection of 2D or 3D geometric primitives, such as anisotropic Gaussians [24], oriented disks [21], or triangles [20]. In Sec. 4, we present our formulation on the example of 3D Gaussian primitives [24, 65] whose spatial response function Ï(x) = exp â 1 2 (xâÎŒ) T ÎŁ â1 (xâÎŒ) ,(1) is defined by the particle centerÎŒ âR 3 and covariance matrixÎŁ âR 3Ă3 . To ensure positive semi-definiteness during optimization, the covariance is parame- terized asÎŁ = RSS T R T , where R â SO(3) represents rotation and S âR 3Ă3 encodes scaling. In practice, these parameters are stored as a quaternion qâR 4 and a scale vector sâR 3 . Each primitive additionally carries an opacity param- eter Ï âR and a view-dependent radiance function Ï(d), commonly represented using spherical harmonics. Note that in this case, Ï(d) depends only on the ray direction and not on the spatial position of the intersection point. Primitive-based Volume Rendering. Given a camera ray r(Ï ) = o + Ï d with origin o âR 3 and direction d âR 3 , the rendered ray color c âR 3 is obtained via volume rendering over the primitives P c intersecting the ray: c(o, d) = X iâP c Ï(d)· T i · α i , T i = iâ1 Y j=1 1â α j , (2) where the opacity α i is defined as α i = Ï i Ï i (o +Ï d) and T i is the transmittance. 6J. Condor et al. RGB Fig. 2: Neural Harmonic Textures applied to novel-view synthesis. We virtually attach feature vectors f i to the vertices of tetrahedra inscribing the Gaussian primitives 3 . Fol- lowing 3DGUT [65], we evaluate the point along the ray where the projected Gaussian has maximum response. We barycentrically interpolate vertex features at that point, and encode them with sine and cosine functions into different channels. These are then alpha blended along the rest of the ray, until the resulting sum of harmonics is decoded by a shallow MLP in a single image-space pass. Positional Encoding: Neural networks exhibit a natural spectral bias that limits their ability to learn high-frequency functions in low-dimensional domains [57]. Positional encodings mitigate this limitation by mapping low-dimensional spatial and directional coordinates into higher-dimensional representations be- fore feeding them into the network [57]. While early approaches relied on fixed frequency encodings [17, 45], modern methods often use parametric encodings that store learnable feature vectors in auxiliary spatial data structures, such as grids or trees [44,55]. Query coordinates retrieve and interpolate these features, effectively shifting much of the representational burden from the neural network to the encoding itself. This design trades increased memory usage for reduced computation: only a small subset of encoding parameters is updated per sample, allowing the neural network to remain compact and efficient while significantly accelerating convergence and maintaining high reconstruction quality. 4 Method We present the core components of our approach using 3D Gaussian primitives in the context of novel view synthesis, following the formulation of 3DGUT [65]. We begin by introducing our primitive-bound feature embedding in Sec. 4.1. Next, Sec. 4.2 presents harmonic texturing, which increases per-primitive expressive 3 Gaussian particles are bounded by ellipsoids in world space, which become spheres after whitening. Our virtual bounding tetrahedra are defined in this canonical space. Neural Harmonic Textures7 W h i t e n W h i t e n (a) Latent vectors(b) Interpolation s i n c o s Interpolated (c) Harmonic decomposition Fig. 3: Illustrating our method in 2D. Each primitive is bounded by an ellipsoid in world space, which becomes a sphere in whitened canonical space (a). Considering a virtual bounding tetrahedron in this canonical space, we attach one N-dimensional feature vector f j to each vertex. The primitiveâs contribution is evaluated at the point of maximum response p â of the projected Gaussian along the intersecting ray (b). The feature vectors are barycentrically interpolated at p â and encoded with sine and cosine periodic functions (c). power while preserving the locality and explicit form of the representation. Fi- nally, Sec. 4.3 describes a neural deferred decoding scheme that enables efficient signal reconstruction in a single image-space pass. Generalization to other primi- tive types and additional applications will be discussed in Sec. 5.2. Fig. 2 provides a schematic overview of our method. 4.1 Primitive-bound feature embedding Spatial structures such as triplanes [10], voxel grids [53], and multi-resolution hash grids [44] are commonly used in neural fields to hold latent feature vectors that embed query positions into a higher-dimensional space before decoding the signal. Despite their effectiveness, these representations rely on globally defined regular grids, which limits their scalability to large scenes or high-frequency detail. Moreover, because these structures are fixed in space, they struggle to represent scenes undergoing motion, deformation, or editing. Instead, we exploit the Lagrangian nature and adaptivity of primitive-based representations by anchoring latent features to a virtual scaffold that encap- sulates each primitive. Specifically, consider the isosurface ellipsoid defined by an anisotropic 3D Gaussian (Eq. (1)) with bounded support, which becomes a sphere in canonical space after whitening. We define the virtual tetrahedron bounding this sphere and assign a feature vector f j âR N f to each of its four vertices (j â 0..3) (Fig. 3a). For each primitive intersected by a ray, we de- termine the point p â corresponding to the maximum Gaussian response along the ray [65]. The feature vector at this location, f, is obtained via barycentric interpolation of the vertex features f j . A detailed derivation is provided in the supplementary material. Compared to globally anchored feature representations, our formulation moves with the primitives and therefore naturally supports operations such as motion, deformation, or deletion during scene editing. 8J. Condor et al. 4.2 Harmonic Texturing In previous work, feature vectors are typically passed through a lightweight neural network to decode the primitive appearance prior to volume render- ing, requiring dozens of MLP evaluations per ray to produce the final pixel color [37, 44, 74]. In 3DGS, color decoding overhead (without neural networks) is reduced by approximating view-dependent color emission once per primitive, thereby reducing the complexity from the number of rayâprimitive intersections to the number of primitives. However, this approximation assumes that the sig- nal does not vary spatially within each primitive, which makes it unsuitable for our representation. An alternative strategy is to blend the features along each ray and decode the signal in image space, in the spirit of deferred shading [19]. Inspired by the Fourier transform, where a complex signal is expressed as a sum of harmonic components (i.e., periodic functions of different amplitudes), we instead encode the interpolated features using periodic functions before blending them along the ray. This transforms the features into frequencies and amplitudes, yield- ing a harmonic decomposition of the signal, which we term Harmonic Textures (Fig. 3c). -7-7 -7-7-7-7-1-1 +1 +1+1+1 -10 -10 0000+10+10 -20-20-10-10 +1 -1 Fig. 4: Harmonic textures. In this formulation, the interpolation function effectively acts as a frequency modulator: large differences between vertex features within a prim- itive produce rapidly oscillating, spatially vary- ing textures. Each primitive additionally has a kernel-weighted opacity, which acts as the har- monic amplitude. This behavior is illustrated in the inset Fig. 4 (without the amplitude weight- ing, to better visualize the harmonics). 4.3 Neural Deferred Shading The rich high-dimensional signal yielded by the volume rendering of Harmonic Textures enables a more efficient decoding strategy. We concatenate the accumu- lated harmonics with the ray direction d encoded using second-degree spherical harmonics SH 2 (d) âR 9 as done in [44], and decode the final pixel color c us- ing a shallow MLP, without further positional encoding. The resulting rendering equation is c = MLP Ξ X iâG α i T i sin(f i ) cos(f i ) , k· SH 2 (d) ! ,(3) where f i = interpolate f 0 i , f 1 i , f 2 i , f 3 i , p â i denotes the primitive feature interpo- lated at the point p â i of maximum response,G is the set of primitives intersected by the ray, and α i and T i correspond to the primitive opacity and accumulated transmittance defined in Eq. (2). Neural Harmonic Textures9 5 Results We present our results for radiance-field novel-view synthesis. Additionally, we demonstrate our method on two other reconstruction tasks: semantic scene re- construction and high-resolution image fitting. 5.1 Radiance Field Reconstruction Optimization We directly adopt the densification strategy, loss function and regularization terms of 3DGS-MCMC [25]. Specifically, our loss function is a combination of L 1 , D-SSIM, and regularizers on primitive opacity and scale: R α = 1 P P X i=1 α i , R s = 1 P P X i=1 â„s i â„ 1 ,(4) where P is the number of primitives and s i is the scale vector of primitive i. Our final loss function is: L = (1â λ)L L 1 + λL D-SSIM + λ α R α + λ s R s .(5) Similarly to the exponential scheduling applied to primitive positions [24], we use a cosine annealing scheduler for the feature and MLP learning rates. We also apply an Exponential Moving Average (EMA) filter [44] to the MLP weights Ξ. This enhances robustness to noise and avoids overfitting to individual frames: Ì Îž t â Îł Ì Îž tâ1 + (1â Îł)Ξ t ,(6) where Ì Îž t denotes the filtered weights at step t and Îł is the decay factor. Finally, we dedicate the last 3000 training iterations to optimize only the feature and MLP weights, disabling all regularization terms and freezing all other parameters. We found that this refinement step slightly improves color fidelity, particularly in large-scale scenes. All hyperparameters and other smaller details will be included in the supplementary. Implementation details We implement our method in gsplat [70], using the 3DGUT [65] formulation, to which we add custom CUDA kernels for the for- ward and backward logic described in Sec. 4. To reduce register pressure, we use half-precision memory fetches on feature vectors during rasterization, for both forward and backward kernels. Our neural network uses tiny-cuda-nâs [44] JIT-compiled cooperative vector MLPs, trained and evaluated in half precision (FP16). We apply an automatic scaling factor to the loss values in order to improve half-precision training stability [40]. 10J. Condor et al. MipNeRF360Tanks & TemplesDeep Blending MethodPSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ Instant NGP-Big [44]25.59 0.695 0.375 21.92 0.740 0.342 24.96 0.815 0.459 Mip-NeRF 360 [1]27.60 0.788 0.275 22.22 0.754 0.290 29.40 0.8990.306 ZipNeRF [2] 28.55 0.8290.218 23.64 0.836 0.179â 2DGS [21]27.48 0.816 0.263 22.85 0.827 0.244 29.56 0.904 0.325 3DGS-MCMC [25]28.210.8410.214 24.460.8660.174 29.49 0.9120.306 3DGUT-MCMC [25]28.080.837 0.218 24.20 0.861 0.18029.870.913 0.309 Beta Splatting-MCMC [33]28.12 0.831 0.238 24.54 0.866 0.196 29.56 0.907 0.316 Spherical Voronoi [9]28.56 0.835 0.22824.800.8710.17230.340.9140.299 Triangle Splatting [20]27.00 0.808 0.231 23.05 0.843 0.191 28.92 0.891 0.308 Textured Gaussians [3]27.35 0.827â24.26 0.854â28.33 0.891â NeST [74]26.54 0.776 0.260â Radiance Meshes [37]27.15 0.810 0.274 23.13 0.851 0.200 29.39 0.901 0.362 Neural Harmonic Textures (Ours) 29.010.8400.21225.680.8820.14130.940.9190.302 Table 1: Comparison on MipNeRF360 [1], Tanks & Temples [27], and Deep Blend- ing [18], on Neural Field-style methods, primitive-based methods, and mixed neural field/primitive-based methods. For our method, we use 64 features per primitive (16 per vertex), with a 128-wideĂ 3 hidden layers MLP, which results in a roughly similar number of total paramaters as previous primitive-based approaches on average. For indoor scenes, which typically span a smaller spatial volume, we use 2M primitives, while for larger outdoor scenes we increase the primitive count to 4.5M. To regularize training in high-parallax regimes (e.g., indoor datasets), we encode the central camera ray rather than the individual ray direction. Finally, to ensure a better coverage in terms of training epochs, we train indoor datasets longer (45k iterations for ⌠290 im- ages) than outdoor (20k for ⌠175 images), following Spherical Voronoi [9]. Note that convergence of an MLP-based appearance model can be slower than SH-based ones. Results We present the results of our method against a number of previous works using pure neural fields [1], hashgrid-encoded neural fields [2, 44], pure primitive-based methods [20,21,25,33,65] with regular Spherical Harmonics (SH) appearance or more complex appearance models [3,9], and closer to us, methods mixing primitives (e.g. for acceleration) and neural fields [37, 74] (Tab. 1). We consistently outperform all previous works in a varied set of real datasets [1,18, 27]. We include a selection of views in Fig. 6. Our method particularly excels at high-frequency detail modelling, capturing specular highlights and reflections with superior fidelity. In order to completely isolate the impact of our method, we include a con- trolled experiment in Table 2. We compare between 3DGS-MCMC, 3DGUT- MCMC, and different appearance models for the latter: regular SH, current state-of-the-art Spherical Voronoi functions [9] (SV) and ours (NHT). We im- plement all 4 under the same framework (gsplat), under strictly controlled con- ditions to ensure the only difference between them is the choice of appearance model itself. Table 2 shows that our method consistently outperforms both SH and SV across all benchmarks, without sacrificing real-time performance. Per-scene breakdowns and more details are provided in the appendix (Ap- pendix Tab. 9 and Tab. 10). Neural Harmonic Textures11 MipNeRF360Tanks & TemplesDeep Blending Method (w/ MCMC) PSNRâ SSIMâ LPIPSâ FPSâ PSNRâ SSIMâ LPIPSâ FPSâ PSNRâ SSIMâ LPIPSâ FPSâ 3DGS + SH27.940.8290.24625124.250.8610.188294 29.980.9120.317331 3DGUT + SH27.930.8280.247 201 23.99 0.859 0.19224530.210.9130.318282 3DGUT + SV 28.15 0.823 0.24820224.180.8610.187 24230.29 0.912 0.320 267 3DGUT + NHT (Ours)28.460.8300.232 14024.790.8750.169 22630.880.9180.311 240 Table 2: Quantitative results of our approach and baselines on the MipNeRF360 [1], Tanks & Temples [27], and Deep Blending [18] datasets, comparing our method against Spherical Voronoi [9] (SV) and regular SH models. We isolate the effect of our approach by implementing all methods in the same framework (gsplat [71]), using the same number of primitives (1M), training for the same number of iterations (30k), allocating the same number of parameters per primitive for appearance (48) and using the same hyperparameters for all scenes. Our method uses a 128Ă3 hidden MLP. We improve reconstruction quality while still managing real-time performance. Measured in an RTX A6000 Ada. Table 3: Generality of our approach: Neu- ral Harmonic Textures applied to different primitive-based representations, evaluated on MipNeRF360 [1]. MethodPSNRâ SSIMâ LPIPSâ Triangle Splatting27.00 0.808 0.231 Triangle Splatting + Ours 27.52 0.807 0.190 2DGS27.48 0.816 0.263 2DGS + Ours28.27 0.820 0.238 3DGUT-MCMC27.93 0.828 0.247 3DGUT-MCMC + Ours 28.43 0.829 0.230 1101001,000 Primitive Count (thousands) 20 22 24 26 28 PSNR (dB) NHT (Ours) 3DGS-MCMC 3DGUT-MCMC Fig. 5: Our method outperforms 3DGS and 3DGUT at all primitive counts. The improvement is par- ticularly pronounced in the low- primitive regime (†100k). Ablations We ablate the key design decisions of our method on the MipN- eRF360 dataset. Optimization choices. Appendix Tab. 13 isolates the individual contributions of the training strategies detailed in Sec. 5.1: learning rate scheduling, EMA on the MLP weights, loss scaling, the color-refinement phase, and opacity and scale regularization. Primitive count. Fig. 5 illustrates how reconstruction quality scales with the number of primitives, ranging from 1K to 4M. Our method improves over pre- vious works across this entire range, excelling particularly in highly compact regimes with low primitive counts. For instance, we achieve performance compa- rable to previous methods using 1M primitives with only a third of that amount, offering more flexibility in the memory-speed-quality trade-off. Tab. 15 and Tab. 16 in the appendix explore the impact of varying the per- primitive feature dimension and the MLP architecture, respectively, to illustrate the resulting quality-speed trade-offs. 12J. Condor et al. Spherical VoronoiGround truthOurs3DGUT-MCMC Fig. 6: Comparison between our and previous works on radiance field reconstruction on scenes from MipNeRF360 [1], Tanks and Temples [27]. Our method models high frequency detail and view dependent effects to a higher degree than previous works. EncodingPSNRâ SSIMâ LPIPSâ None (identity)28.21 0.819 0.255 ReLU28.19 0.819 0.255 Cosine28.34 0.826 0.234 Cosine & Sinusoid 28.46 0.830 0.232 Table 4: Ablating the choice of feature encoding function (MipNeRF360 dataset). Feature encoding. Tab. 4 compares different encoding functions applied to the interpolated features. 5.2 Other Applications Alternative primitive-based methods We demonstrate the generality of our approach by incorporating Neural Harmonic Textures into 2DGS [21] (as imple- mented in gsplat) and Triangle Splatting [20]. Our method is readily adapted to these 2D primitives by simply replacing the bounding tetrahedra with vir- tual triangles (2DGS) or directly anchoring feature vectors to explicit triangle geometry (Triangle Splatting), therefore using three feature vectors per primi- tive instead of four. The results are shown in Tab. 3, with further details in the Supplementary. Neural Harmonic Textures13 MethodRGB PSNR â RGB SSIM â RGB LPIPS â LSEG P â LSEG C â FPSâ Feature 3DGS [76]26.270.7850.26243.680.9853.02 Ours28.160.8240.22846.900.99328.11 Table 5: Joint RGB and LSEG semantic feature reconstruction on MipNeRF360 [1], measured on an RTX A6000 Ada. LSEG P and LSEG C denote PSNR and cosine simi- larity of the 512-dim semantic features, respectively. We use 80 features per primitive and a 128Ă3 MLP, while Feature 3DGS uses 176 features per primitive, 512Ă128 CNN and 1.5Ă the number of primitives we use, resulting in an approximately 3Ă higher average memory footprint. Semantic Field Reconstruction Our method can easily model higher dimen- sional data as well. By simply expanding the decoding head of our MLP, we can fit signals of arbitrary dimension. As the rest of the MLP layers remain small, and the size of feature vectors remains the same, the model is forced to exploit potential correlations between dimensions. A particularly challenging case is semantic scene reconstruction, where fields are upwards of 512-dimensional. 3DGS-based methods relying on high-dimensional explicit features generally struggle with such scale, even when paired with feature upsampling networks. We thus test our method on joint novel-view synthesis of RGB radiance and LSEG [29] 512-wide semantic features. Due to the limited resolution of LSEG feature maps, we render and predict at the RGB resolution, but bilinearly down- sample the spatial dimension of the resulting feature maps in order to supervise on ground truth features. We compare against Feature 3DGS [76], a hybrid method relying on 128-wide semantic feature vector per primitive plus a CNN to jointly downsample the spatial dimension and upsample the feature dimen- sion of the resulting rasterized image up to the full 512 dimensions, and another set of 48 SH basis weights for regular RGB color. Results for the aggregated MipNeRF360 dataset can be seen in Tab. 5. We afford Feature-3DGS twice our training time budget, up to 7K iterations, slightly more than in the original pa- per (5K), to accommodate for the larger-scale datasets. We further ablate the per-primitive feature dimensionality for this task in Appendix Tab. 11. Our re- sults show substantial improvement over Feature 3DGS, and the capability to easily scale upwards or downwards in performance or memory, in a completely detached manner from the nature of the data being encoded. It is particularly interesting to note how small the drop in RGB quality is, which points to either high correlation among feature vectors or excessive free bandwidth in our RGB experiments. Our features can then be used in downstream tasks such as scene editing or semantic segmentation. 2D Image Reconstruction Finally, we show results on 2D image fitting. To adapt our method to this application, we make two key changes with respect to the primitive-based approaches. First, we employ a connected 2D triangle mesh topology. This mesh is fully opaque: there is no alpha blending, no Gaussian or 14J. Condor et al. Ratio MethodPSNR ÎŒ â PSNR tm â SSIM tm â LPIPS tm â 10Ă JPEG-XL [23]45.31 44.660.9940.001 Instant NGP [44] 37.6338.910.9490.061 NHT (Ours)37.1639.87 0.9630.023 100Ă JPEG-XL [23]38.5335.540.9710.017 Instant NGP [44] 35.1736.390.9230.078 NHT (Ours)35.0636.430.927 0.048 Table 6: High-resolution (45.7 MP) 14-bit HDR RAW image fitting, averaged over 15 images. We report PSNR in ÎŒ-law transformed space (PSNR ÎŒ ), common in HDR pipelines. PSNR, SSIM, and LPIPS are also computed on tonemapped images (sub- script âtmâ). All methods use the same memory budget per compression ratio. window kernel, simply interpolation of local features for a given pixel. Second, we change our barycentric interpolation to Clough-Tocher cubic interpolation. This makes interpolated features continuous over the triangle edges, which removes high frequency artifacts. The rest of the architecture is the same as described in Sec. 4, and we include further details on the optimization strategy in the Supplementary. We evaluate image reconstruction performance on a new curated dataset of high-resolution (45.7 MP) 14-bit HDR RAW images, and compare it against Instant NGP [44] and the JPEG-XL encoder [23], which supports high-dynamic range. We measure reconstruction quality at two compression ratios (10Ă and 100Ă), computing PSNR in ÎŒ-law transformed space (PSNR ÎŒ ), standard in HDR imaging pipelines. We also report PSNR, SSIM, and LPIPS on tonemapped images (subscript âtmâ). Results are shown in Tab. 6. Our method achieves comparable pixel-error quality to Instant NGP in both HDR and linear spaces, while substantially reducing perceptual error: at 100Ă compression, NHT achieves an LPIPS of 0.048 compared to 0.078 for Instant NGP, a 38% relative improvement. JPEG- XL however still comes out on top, particularly in perceptual metrics, which is to be expected given its human-perception-based inductive biases. We can further compress our representation by a factor of 3 without noticeable loss in quality by aggressively reducing floating point precision of the implicit features, among other strategies. We include more details on this in the supplementary. 6 Discussion Limitations and future work Our method focuses on increasing per-primitive expressivity, which has inherent downsides. In particular, our method tends to overfit more to individual views in scenarios with very sparse supervision, leading to lower novel view synthesis performance. Furthermore, while our method con- sistently achieves 140+ frames per second at inference time, the additional neural decoder results in slightly slower rendering than pure 3DGS methods(Tab. 2). Neural Harmonic Textures15 Instant NGP Original Ours Instant NGP Original OursOursOurs Fig. 7: Comparison of our work vs Instant NGP on a 100Ă compression task. Original images are 45.7MP 14-bit HDR RAW files. We achieve substantially superior perceptual quality at equal compression and similar training times. In future work, we would like to investigate the feasibility of extracting au- tomatic Level of Detail from a finer-grained harmonic decomposition. Given the generality of Neural Harmonic Textures, applications beyond novel view synthe- sis such as radiance caching, neural physically-based rendering and geometric reconstruction are also of great interest. Finally, our work opens the possibility of removing kernel functions entirely, as we can still propagate derivatives to explicit geometry without the spatially decaying opacity function, potentially enabling drastic performance gains. Conclusion We presented Neural Harmonic Textures, which combine particle- anchored feature vectors, harmonic activations, and a single image-space decod- ing pass to achieve state-of-the-art results in novel view synthesis. Our method consistently outperforms baselines across different particle counts, trains rapidly, and has high inference performance. Moreover, it is compatible with existing deferred-shading rendering pipelines and supports a wide range of applications, from lifting semantic fields to high- resolution image reconstruction and beyond. Acknowledgments We would like to thank Ruilong Li and Martin Bisson for helpful discussions. Jorge Condor and Piotr Didyk acknowledge funding from the Swiss National Science Foundation (SNSF, Grant 200502) and an academic gift from Meta. References 1. Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Mip- nerf 360: Unbounded anti-aliased neural radiance fields. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022) 10, 11, 12, 13, 21, 23, 24 2. Barron, J.T., Mildenhall, B., Verbin, D., Srinivasan, P.P., Hedman, P.: Zip-NeRF: Anti-aliased grid-based neural radiance fields. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2023) 10 16J. Condor et al. 3. Chao, B., Tseng, H.Y., Porzi, L., Gao, C., Li, T., Li, Q., Saraf, A., Huang, J.B., Kopf, J., Wetzstein, G., Kim, C.: Textured gaussians for enhanced 3d scene ap- pearance modeling. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025) 5, 10 4. Chen, A., Xu, Z., Geiger, A., Yu, J., Su, H.: Tensorf: Tensorial radiance fields. In: European conference on computer vision. p. 333â350. Springer (2022) 3 5. Chen, Y., Chen, Z., Zhang, C., Wang, F., Yang, X., Wang, Y., Cai, Z., Yang, L., Liu, H., Lin, G.: Gaussianeditor: Swift and controllable 3d editing with gaussian splatting. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 21476â21485 (2024) 2 6. Chen, Z., Funkhouser, T., Hedman, P., Tagliasacchi, A.: MobileNeRF: Exploiting the polygon rasterization pipeline for efficient neural field rendering on mobile architectures. In: The Conference on Computer Vision and Pattern Recognition (CVPR) (2023) 4 7. Condor, J., Hermann, N., Yurtsever, M.A., Didyk, P.: Gabor fields: Orientation- selective level-of-detail for volume rendering (2026), https://arxiv.org/abs/ 2602.05081 4 8. Condor, J., Speierer, S., Bode, L., Bozic, A., Green, S., Didyk, P., Jarabo, A.: Donât splat your gaussians: Volumetric ray-traced primitives for modeling and rendering scattering and emissive media. ACM Transactions on Graphics 44(1) (2025) 4 9. Di Sario, F., Rebain, D., Verbin, D., Grangetto, M., Tagliasacchi, A.: Spherical voronoi: Directional appearance as a differentiable partition of the sphere. arXiv preprint arXiv:2512.14180 (2025) 2, 10, 11 10. Duckworth, D., Hedman, P., Reiser, C., Zhizhin, P., Thibert, J.F., LuÄiÄ, M., Szeliski, R., Barron, J.T.: Smerf: Streamable memory efficient radiance fields for real-time large-scene exploration (2023) 7 11. Fridovich-Keil, S., Meanti, G., Warburg, F.R., Recht, B., Kanazawa, A.: K-planes: Explicit radiance fields in space, time, and appearance. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 12479â 12488 (2023) 3 12. Gadirov, H., Wu, Q., Bauer, D., Ma, K.L., Roerdink, J.B., Frey, S.: Hyperflint: Hypernetwork-based flow estimation and temporal interpolation for scientific en- semble visualization. Computer Graphics Forum 44(3), e70134 (2025). https: //doi.org/https://doi.org/10.1111/cgf.70134, https://onlinelibrary. wiley.com/doi/abs/10.1111/cgf.70134 3 13. Govindarajan, S., Rebain, D., Yi, K.M., Tagliasacchi, A.: Radiant foam: Real- time differentiable ray tracing. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). p. 4135â4145 (2025) 4 14. Govindarajan, S., Sambugaro, Z., Shabanov, A., Takikawa, T., Rebain, D., Sun, W., Conci, N., Yi, K.M., Tagliasacchi, A.: Lagrangian hashing for compressed neural field representations. In: European Conference on Computer Vision. p. 183â199. Springer (2024) 3 15. Hahlbohm, F., Franke, L., Kappel, M., Castillo, S., Eisemann, M., Stamminger, M., Magnor, M.: INPC: Implicit Neural Point Clouds for Radiance Field Render- ing . In: 2025 International Conference on 3D Vision (3DV). p. 168â178. IEEE Computer Society, Los Alamitos, CA, USA (Mar 2025). https://doi.org/10. 1109/3DV66043.2025.00021, https://doi.ieeecomputersociety.org/10.1109/ 3DV66043.2025.00021 4 16. Han, K., Xiang, W., Yu, L.: Volume feature rendering for fast neural radiance field reconstruction. In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., Neural Harmonic Textures17 Levine, S. (eds.) Advances in Neural Information Processing Systems. vol. 36, p. 65416â65427. Curran Associates, Inc. (2023), https://proceedings.neurips.c/ paper_files/paper/2023/file/ce182e31662883d4decc84a0255335b6- Paper- Conference.pdf 4 17. Harris, D., Harris, S.L.: Digital design and computer architecture. Morgan Kauf- mann (2010) 6 18. Hedman, P., Philip, J., Price, T., Frahm, J.M., Drettakis, G., Brostow, G.: Deep blending for free-viewpoint image-based rendering. ACM Transactions on Graphics (Proceedings of SIGGRAPH Asia) 37(6), 257:1â257:15 (2018) 10, 11, 21, 23, 24 19. Hedman, P., Srinivasan, P.P., Mildenhall, B., Barron, J.T., Debevec, P.: Bak- ing neural radiance fields for real-time view synthesis. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2021) 4, 8 20. Held, J., Vandeghen, R., Deliege, A., Hamdi, A., Rebain, D., Giancola, S., Cioppa, A., Vedaldi, A., Ghanem, B., Tagliasacchi, A., et al.: Triangle splatting for real- time radiance field rendering. In: Thirteenth International Conference on 3D Vision (3DV) (2025) 2, 4, 5, 10, 12 21. Huang, B., Yu, Z., Chen, A., Geiger, A., Gao, S.: 2D Gaussian splatting for geo- metrically accurate radiance fields. In: ACM SIGGRAPH 2024 Conference Papers (2024). https://doi.org/10.1145/3641519.3657428 4, 5, 10, 12 22. Huang, Z., Gong, M.: Textured-gs: Gaussian splatting with spatially defined color and opacity. arXiv preprint arXiv:2407.09733 (2024) 5 23. Joint Photographic Experts Group: JPEG XL image coding system. https:// jpeg.org/jpegxl/ (2024), accessed: 2024-05-24 14, 32 24. Kerbl, B., Kopanas, G., LeimkĂŒhler, T., Drettakis, G.: 3D gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics (Proceedings of SIGGRAPH) 42(4) (2023) 2, 4, 5, 9 25. Kheradmand, S., Rebain, D., Sharma, G., Sun, W., Tseng, Y.C., Isack, H., Kar, A., Tagliasacchi, A., Yi, K.M.: 3D gaussian splatting as markov chain monte carlo. In: Advances in Neural Information Processing Systems (NeurIPS) (2024) 9, 10 26. Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization (2017) 29 27. Knapitsch, A., Park, J., Zhou, Q.Y., Koltun, V.: Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (Proceedings of SIGGRAPH) 36(4) (2017) 10, 11, 12, 21, 23, 24 28. Kulhanek, J., Rakotosaona, M.J., Manhardt, F., Tsalicoglou, C., Niemeyer, M., Sattler, T., Peng, S., Tombari, F.: Lodge: Level-of-detail large-scale gaussian splat- ting with efficient rendering. arXiv preprint arXiv:2505.23158 (2025) 4 29. Li, B., Weinberger, K.Q., Belongie, S., Koltun, V., Ranftl, R.: Language-driven semantic segmentation. arXiv preprint arXiv:2201.03546 (2022) 13 30. Li, Z., MĂŒller, T., Evans, A., Taylor, R.H., Unberath, M., Liu, M.Y., Lin, C.H.: Neuralangelo: High-fidelity neural surface reconstruction. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 8456â8465 (2023) 3 31. Liang, H., Ren, J., Mirzaei, A., Torralba, A., Liu, Z., Gilitschenski, I., Fidler, S., Oztireli, C., Ling, H., Gojcic, Z., Huang, J.: Feed-forward bullet-time reconstruc- tion of dynamic scenes from monocular videos. arXiv preprint arXiv:2412.03526 (2024) 2 32. Liu, L., Gu, J., Lin, K.Z., Chua, T.S., Theobalt, C.: Neural sparse voxel fields. In: Advances in Neural Information Processing Systems (NeurIPS) (2020) 3 33. Liu, R., Sun, D., Chen, M., Wang, Y., Feng, A.: Deformable beta splatting. In: Proceedings of SIGGRAPH Conference Papers (2025) 2, 4, 10 18J. Condor et al. 34. Lombardi, S., Simon, T., Saragih, J., Schwartz, G., Lehrmann, A., Sheikh, Y.: Neural volumes: Learning dynamic renderable volumes from images. ACM Trans- actions on Graphics (Proceedings of SIGGRAPH) 38(4) (2019) 3 35. Lombardi, S., Simon, T., Schwartz, G., Zollhoefer, M., Sheikh, Y., Saragih, J.: Mixture of volumetric primitives for efficient neural rendering. ACM Transactions on Graphics (Proceedings of SIGGRAPH) 40(4) (2021) 3 36. Luo, A., Du, Y., Tarr, M., Tenenbaum, J., Torralba, A., Gan, C.: Learning neural acoustic fields. Advances in Neural Information Processing Systems 35, 3165â3177 (2022) 3 37. Mai, A., Hedstrom, T., Kopanas, G., Kontkanen, J., Kuester, F., Barron, J.T.: Radiance meshes for volumetric reconstruction. arXiv preprint arXiv:2512.04076 (2025) 2, 4, 8, 10 38. Mai, A., Hedstrom, T., Kopanas, G., Kontkanen, J., Kuester, F., Barron, J.T.: Radiance meshes for volumetric reconstruction (2025), https://arxiv.org/abs/ 2512.04076 4 39. Martel, J.N.P., Lindell, D.B., Lin, C.Z., Chan, E.R., Monteiro, M., Wetzstein, G.: ACORN adaptive coordinate networks for neural scene representation. ACM Transactions on Graphics (Proceedings of SIGGRAPH) 40(4) (2021). https: //doi.org/10.1145/3450626.3459785, https://doi.org/10.1145/3450626. 3459785 3 40. Micikevicius, P., Narang, S., Alben, J., Diamos, G., Elsen, E., Garcia, D., Ginsburg, B., Houston, M., Kuchaiev, O., Venkatesh, G., et al.: Mixed precision training. arXiv preprint arXiv:1710.03740 (2017) 9 41. Mildenhall, B., Srinivasan, P.P., Tancik, M., Barron, J.T., Ramamoorthi, R., Ng, R.: NeRF: Representing scenes as neural radiance fields for view synthesis. In: Proceedings of ECCV (2020) 2, 3 42. Moenne-Loccoz, N., Mirzaei, A., Perel, O., de Lutio, R., Esturo, J.M., State, G., Fidler, S., Sharp, N., Gojcic, Z.: 3d gaussian ray tracing: Fast tracing of particle scenes. ACM Transactions on Graphics and SIGGRAPH Asia (2024) 4 43. Mujkanovic, F., Nsampi, N.E., Theobalt, C., Seidel, H.P., LeimkĂŒhler, T.: Neural gaussian scale-space fields. ACM Trans. Graph. 43(4) (Jul 2024). https://doi. org/10.1145/3658163, https://doi.org/10.1145/3658163 3 44. MĂŒller, T., Evans, A., Schied, C., Keller, A.: Instant neural graphics primitives with a multiresolution hash encoding. ACM Trans. Graph. 41(4) (2022) 2, 3, 6, 7, 8, 9, 10, 14, 29, 32 45. MĂŒller, T., McWilliams, B., Rousselle, F., Gross, M., NovĂĄk, J.: Neural importance sampling. ACM Transactions on Graphics (TOG) 38(5), 1â19 (2019) 6 46. Park, J.J., Florence, P., Straub, J., Newcombe, R., Lovegrove, S.: Deepsdf: Learning continuous signed distance functions for shape representation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 165â 174 (2019) 3 47. Pumarola, A., Corona, E., Pons-Moll, G., Moreno-Noguer, F.: D-nerf: Neural ra- diance fields for dynamic scenes. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 10318â10327 (2021) 3 48. Reiser, C., Peng, S., Liao, Y., Geiger, A.: KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs. In: Proceedings of ICCV (2021) 3 49. Saragadam, V., LeJeune, D., Tan, J., Balakrishnan, G., Veeraraghavan, A., Bara- niuk, R.G.: Wire: Wavelet implicit neural representations. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 18507â 18516 (2023) 3 Neural Harmonic Textures19 50. Sitzmann, V., Martel, J., Bergman, A., Lindell, D., Wetzstein, G.: Implicit neural representations with periodic activation functions. Advances in neural information processing systems 33, 7462â7473 (2020) 2 51. Sitzmann, V., Martel, J.N., Bergman, A.W., Lindell, D.B., Wetzstein, G.: Implicit neural representations with periodic activation functions. In: Proc. NeurIPS (2020) 3 52. Su, R., Dong, H., Jin, H., Chen, Y., Wang, G., Li, S.: Vertex features for neu- ral global illumination. In: Proceedings of the SIGGRAPH Asia 2025 Conference Papers. p. 1â11 (2025) 25 53. Sun, C., Sun, M., Chen, H.: Direct voxel grid optimization: Super-fast convergence for radiance fields reconstruction. In: CVPR (2022) 7 54. Svitov, D., Morerio, P., Agapito, L., Del Bue, A.: Billboard splatting (b- splat): Learnable textured primitives for novel view synthesis. arXiv preprint arXiv:2411.08508 (2024) 5 55. Takikawa, T., Litalien, J., Yin, K., Kreis, K., Loop, C., Nowrouzezahrai, D., Ja- cobson, A., McGuire, M., Fidler, S.: Neural geometric level of detail: Real-time rendering with implicit 3d shapes. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. p. 11358â11367 (2021) 3, 6 56. Takikawa, T., MĂŒller, T., Nimier-David, M., Evans, A., Fidler, S., Jacobson, A., Keller, A.: Compact neural graphics primitives with learned hash probing. In: SIGGRAPH Asia 2023 Conference Papers. SA â23, Association for Computing Machinery, New York, NY, USA (2023). https://doi.org/10.1145/3610548. 3618167, https://doi.org/10.1145/3610548.3618167 2 57. Tancik, M., Srinivasan, P.P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Sing- hal, U., Ramamoorthi, R., Barron, J.T., Ng, R.: Fourier features let networks learn high frequency functions in low dimensional domains. NeurIPS (2020) 2, 3, 6 58. Thies, J., Zollhöfer, M., NieĂner, M.: Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (Proceedings of SIGGRAPH) 38(4) (2019) 3, 4 59. Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Visual geometry grounded transformer. In: Proceedings of the Computer Vision and Pattern Recognition Conference. p. 5294â5306 (2025) 2 60. Wang, P., Liu, L., Liu, Y., Theobalt, C., Komura, T., Wang, W.: Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689 (2021) 3 61. Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.: Ï 3 : Permutation-equivariant visual geometry learning (2025), https://arxiv.org/abs/2507.13347 2 62. Wu, G., Yi, T., Fang, J., Xie, L., Zhang, X., Wei, W., Liu, W., Tian, Q., Wang, X.: 4D gaussian splatting for real-time dynamic scene rendering. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2024) 2, 4 63. Wu, Q., Bauer, D., Doyle, M.J., Ma, K.L.: Interactive volume visualization via multi-resolution hash encoding based neural representation. IEEE Transactions on Visualization and Computer Graphics p. 1â14 (2023). https://doi.org/10. 1109/TVCG.2023.3293121 3 64. Wu, Q., Insley, J.A., Mateevitsi, V.A., Rizzi, S., Papka, M.E., Ma, K.L.: Dis- tributed neural representation for reactive in situ visualization. IEEE Transac- tions on Visualization and Computer Graphics 31(9), 5199â5214 (2025). https: //doi.org/10.1109/TVCG.2024.3432710 3 20J. Condor et al. 65. Wu, Q., Martinez Esturo, J., Mirzaei, A., Moenne-Loccoz, N., Gojcic, Z.: 3dgut: Enabling distorted cameras and secondary rays in gaussian splatting. Conference on Computer Vision and Pattern Recognition (CVPR) (2025) 4, 5, 6, 7, 9, 10 66. Wurster, S., Zhang, R., Zheng, C.: Gabor splatting for high-quality gigapixel image representations. In: ACM SIGGRAPH 2024 Posters. SIGGRAPH â24, Association for Computing Machinery, New York, NY, USA (2024). https://doi.org/10. 1145/3641234.3671081, https://doi.org/10.1145/3641234.3671081 5 67. Xie, Y., Takikawa, T., Saito, S., Litany, O., Yan, S., Khan, N., Tombari, F., Tomp- kin, J., Sitzmann, V., Sridhar, S.: Neural fields in visual computing and beyond. In: Computer graphics forum. vol. 41, p. 641â676. Wiley Online Library (2022) 3 68. Xu, T.X., Hu, W., Lai, Y.K., Shan, Y., Zhang, S.H.: Texture-gs: Disentangling the geometry and texture for 3d gaussian splatting editing. In: European Conference on Computer Vision. p. 37â53. Springer (2024) 5 69. Yariv, L., Hedman, P., Reiser, C., Verbin, D., Srinivasan, P.P., Szeliski, R., Barron, J.T., Mildenhall, B.: BakedSDF: Meshing neural SDFs for real-time view synthesis. In: ACM SIGGRAPH 2023 Conference Proceedings (2023) 4 70. Ye, V., Li, R., Kerr, J., Turkulainen, M., Yi, B., Pan, Z., Seiskari, O., Ye, J., Hu, J., Tancik, M., Kanazawa, A.: gsplat: An open-source library for gaussian splatting. Journal of Machine Learning Research 26(34), 1â17 (2025) 9, 28 71. Ye, V., Li, R., Kerr, J., Turkulainen, M., Yi, B., Pan, Z., Seiskari, O., Ye, J., Hu, J., Tancik, M., Kanazawa, A.: gsplat: An open-source library for gaussian splatting. Journal of Machine Learning Research 26(34), 1â17 (2025) 11 72. Zhang, K., Bi, S., Tan, H., Xiangli, Y., Zhao, N., Sunkavalli, K., Xu, Z.: Gs-lrm: Large reconstruction model for 3d gaussian splatting. European Conference on Computer Vision (2024) 2 73. Zhang, X., Chen, A., Xiong, J., Dai, P., Shen, Y., Xu, W.: Neural shell texture splatting: More details and fewer primitives. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. p. 25229â25238 (2025) 2, 3, 4 74. Zhang, X., Chen, A., Xiong, J., Dai, P., Shen, Y., Xu, W.: Neural shell texture splatting: More details and fewer primitives. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) (2025) 8, 10 75. Zhou, J., Huang, Y., Dai, W., Zou, J., Zheng, Z., Kan, N., Li, C., Xiong, H.: 3dgabsplat: 3d gabor splatting for frequency-adaptive radiance field rendering. arXiv preprint arXiv:2508.05343 (2025) 5 76. Zhou, S., Chang, H., Jiang, S., Fan, Z., Zhu, Z., Xu, D., Chari, P., You, S., Wang, Z., Kadambi, A.: Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. p. 21676â21685 (2024) 13, 21 77. Zhou, X., Nguyen, B.H., Magne, L., Golyanik, V., LeimkĂŒhler, T., Theobalt, C.: Splat the net: Radiance fields with splattable neural primitives. arXiv preprint arXiv:2510.08491 (2025) 3 Neural Harmonic Textures21 MipNeRF360Tanks & TemplesDeep Blending MethodPSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ 3DGS-MCMC + SH28.21 0.841 0.214 24.46 0.866 0.174 29.49 0.912 0.306 3DGUT-MCMC + SH28.08 0.837 0.218 24.20 0.861 0.180 29.87 0.913 0.309 3DGUT-MCMC + NHT (Ours, 48F) 28.65 0.834 0.212 25.41 0.878 0.156 30.67 0.918 0.302 3DGUT-MCMC + NHT (Ours, 64F) 28.68 0.832 0.211 25.64 0.879 0.154 30.69 0.918 0.301 Table 7: Standardized comparison on MipNeRF360 [1], Tanks & Temples [27], and Deep Blending [18]: 3DGS-MCMC, 3DGUT-MCMC, and ours with 48 and 64 features per primitive (48F, 64F). All methods use the same primitive count (matching the high-quality preset in 3DGS-MCMC, which replicates the number of primitives in each scene from original 3DGS), 30k iterations, and the same hyperparameters across all scenes. Extended version of Tab. 2 of the main paper. While our work particularly excels at extracting higher quality from less primitives, we still outperform previous works at high primitive counts across the board. Supplementary Material A Additional Results In Tab. 9 we provide a detailed breakdown of Tab. 2 from the main paper, with per-scene results on MipNeRF360, Tanks and Temples and Deep Blending. For this comparison, we isolate completely the effect of our method by ensuring an even setup for all methods. All four approaches are implemented under the same framework (gsplat), trained for 30k iterations, using 1M primitives for all scenes, sharing training parameters among all scenes, and limiting to 48 features per primitive reserved for appearance. All results presented in this paper compute the LPIPs metric using the VGG model with normalized inputs to the range of [â1, 1]. Furthermore, in Tab. 7 we provide another standardized comparison on all three benchmarks, only this time we match the same primitive count in each scene as the high-quality reported baselines in their respective papers, training for the same amount of iterations. We include our method with both 48 and 64 features per primitive (48F, 64F). Tab. 10 details the per-scene results of main Tab. 1, which uses our split in- door/outdoor training configuration. This is our higher quality setup, which adapts to the different scene scale and dataset sizes of outdoor and indoor datasets. A.1 Semantic Reconstruction Semantic Reconstruction Details For our semantic reconstruction task, following a similar setup to Feature 3DGS [76], we train a joint RGB and LSEG semantic feature field. Joint RGB and LSEG training uses the same approach, densifica- tion, and regularization as standard radiance-field training, with the differences being in the addition of semantic supervision losses and an expanded decod- ing head in our MLP. The deferred MLP is extended to predict both RGB (3 22J. Condor et al. 3DGS-MCMC + SH3DGUT-MCMC + SHNHT (Ours, 48F)NHT (Ours, 64F) ScenePSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ MipNeRF360 â Outdoor bicycle25.870.8010.18325.690.7900.193 25.530.779 0.20225.62 0.776 0.202 garden28.110.8800.107 27.84 0.874 0.11128.330.8780.10228.230.8760.103 stump27.320.8080.20327.150.8020.21527.060.7880.215 27.02 0.786 0.215 treehill 23.290.6700.31023.130.664 0.313 21.96 0.6390.31122.270.6400.308 flowers22.250.6520.31222.250.648 0.31821.900.6370.315 21.77 0.6310.312 MipNeRF360 â Indoor room32.560.9390.240 32.43 0.939 0.24133.820.9450.22933.750.9460.227 counter29.530.9250.21929.55 0.925 0.22030.800.9310.20530.920.9310.203 kitchen 32.100.9380.136 31.96 0.938 0.13633.430.9440.12733.440.9440.124 bonsai32.840.955 0.216 32.76 0.9550.21534.990.9630.20235.070.9620.201 Tanks & Temples train22.480.8320.221 22.19 0.826 0.22923.880.8560.19824.330.8580.194 truck26.430.8990.126 26.21 0.896 0.13126.940.9010.11526.960.9010.114 Deep Blending drjohnson29.250.9100.310 29.13 0.910 0.31530.060.9160.30830.090.9160.306 playroom29.72 0.9130.30130.600.916 0.30231.280.9200.29731.280.9200.296 M360 Outdoor Avg25.370.7620.22325.210.756 0.230 24.960.7440.22924.98 0.7420.228 M360 Indoor Avg31.760.9390.203 31.67 0.939 0.20333.260.9460.19133.300.9460.189 M360 Total Avg28.210.8410.214 28.080.837 0.21828.650.8340.21228.68 0.8320.211 T&T Avg24.460.8660.174 24.20 0.861 0.18025.410.8790.15725.650.8790.154 DB Avg29.49 0.9120.30629.870.913 0.30930.670.9180.30230.690.9180.301 Table 8: Per-scene results for the standardized comparison (Tab. 7). All methods use the same primitive count (matching the baseline default), 30k iterations, and the same hyperparameters across all scenes. channels) and LSEG features (512 dimensions) via addition of an extra Pytorch linear layer. Rendering speed becomes bottlenecked by this linear layer, and more efficient fused alternatives or upsampling strategies could be explored in the futureâwe consider this a proof-of-concept implementation. The total loss adds an LSEG term to the usual L 1 , D-SSIM, opacity and scale regularizers terms: L LSEG = λ LSEG L L 1 Ë f,f â + λ cos 1â cos( Ë f,f â ) , (7) where Ë f and f â are predicted and ground-truth LSEG feature maps, and cos their cosine similarity. We compute these losses at the native LSEG resolution: the model renders at RGB resolution (high); then, we bilinearly downsample the predicted feature map to match the (lower-resolution) LSEG targets. Tab. 12 lists the LSEG-specific hyperparameters used in our final training. We largely kept the same configuration from our pure RGB radiance field experiments; further tuning for this specific use case may render higher performance. Differences to RGB-only training. Aside from the extra output head and loss terms, training proceeds as in the main paper: same optimizer, learning-rate schedules, EMA on the MLP, and color-refinement phase. The only architectural change is the increased output dimension (3 + 512 when fused, or a separate Neural Harmonic Textures23 3DGS-MCMC3DGUT-MCMC + SH 3DGUT-MCMC + SV 3DGUT-MCMC + NHT (Ours) ScenePSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâ LPIPSâ PSNRâ SSIMâLPIPSâ MipNeRF360 â Outdoor bicycle25.610.7740.24625.510.7710.250 25.21 0.758 0.25525.430.7710.243 garden27.290.8550.155 27.21 0.8540.15627.480.8560.15628.040.8720.122 stump 26.930.7880.25026.930.7890.251 25.95 0.757 0.27026.580.7720.249 treehill 23.450.6650.36623.300.6630.36823.16 0.653 0.371 23.130.6600.340 flowers22.040.6280.36222.070.6270.36321.760.616 0.367 21.60 0.6150.351 MipNeRF360 â Indoor room32.330.9380.24532.36 0.937 0.24632.670.9390.24333.030.9430.234 counter29.48 0.9240.223 29.47 0.923 0.22430.690.9320.21330.830.9310.208 kitchen31.670.9350.14431.720.9350.14432.380.9380.14032.980.9420.132 bonsai32.690.953 0.22032.760.9530.22034.070.9590.21534.550.9610.208 Tanks & Temples train22.500.8320.224 22.10 0.827 0.23122.320.8310.22323.110.8510.202 truck 26.000.8910.151 25.88 0.890 0.15326.040.8910.15126.480.8980.135 Deep Blending drjohnson29.500.9090.32029.600.909 0.32129.690.9090.32130.240.9150.317 playroom30.450.9150.31330.820.9160.31530.890.915 0.31831.520.9200.306 M360 Outdoor Avg25.060.7420.27625.000.7410.278 24.71 0.728 0.28424.960.7380.261 M360 Indoor Avg 31.540.9380.20831.58 0.937 0.20932.450.9420.20332.850.9440.196 M360 Total Avg27.940.8290.246 27.930.8280.24728.15 0.823 0.24828.460.8300.232 T&T Avg24.250.8610.188 23.99 0.859 0.19224.180.8610.18724.790.8750.169 DB Avg29.980.9120.31730.210.9130.31830.290.912 0.32030.880.9180.311 Table 9: Per-scene results on MipNeRF360 [1], Tanks & Temples [27], and Deep Blending [18], from Tab. 2. This is the standardized setup limited to 1M primitives, 30k iterations, same number of appearance features per primitive (48) and same hy- perparameters for all scenes, isolating entirely the effect of our approach from other variables. 512-dim linear layer). We do not use a separate training stage for LSEG; RGB and LSEG are optimized jointly from the start. PCA visualization. LSEG feature maps are 512-dimensional and therefore not directly viewable. For qualitative inspection, we project both ground-truth and predicted features to RGB via PCA. The PCA basis is computed from the first image, then applied the to all frames. Fig. 8 shows example PCA visualizations. They allow a direct visual comparison of semantic structure and show that the predicted features align well with the target, while being substantially faster to generate and higher resolution than inferring the RGB frame only and running the LSEG model to obtain its feature map. Semantic scene reconstruction ex- tracts detail from the multiple training views, improving the granularity of the semantic information available for downstream applications. Finally, in Tab. 11 we ablate the feature dimensionality for the joint RGB and LSEG semantic reconstruction task. B Feature interpolation For a spatial coordinate p lying on the surface of our topology, or inside of its volume, we can compute the feature vector by barycentrically interpolating the 24J. Condor et al. ScenePSNRâ SSIMâ LPIPSâ M360 Outdoor (4.5M, 20k iters, per-ray) garden28.33 0.881 0.106 bicycle25.88 0.788 0.216 stump27.06 0.795 0.223 treehill23.36 0.670 0.306 flowers22.08 0.640 0.330 M360 Indoor (2M, 45k iters, center ray) bonsai35.73 0.964 0.192 counter31.00 0.934 0.193 kitchen33.59 0.945 0.123 room34.03 0.947 0.221 Tanks & Temples (2.5M, 40k iters, center ray) truck26.91 0.900 0.112 train24.45 0.865 0.169 Deep Blending (2M, 30k iters, center ray) drjohnson30.43 0.918 0.309 playroom31.45 0.921 0.296 M360 Outdoor Avg 25.34 0.755 0.236 M360 Indoor Avg33.59 0.947 0.183 M360 Total29.01 0.840 0.212 T&T Avg25.68 0.882 0.141 DB Avg30.94 0.919 0.302 Table 10: Per-scene results of NHT (Ours) with split training configurations on MipN- eRF360 [1], Tanks & Temples [27], and Deep Blending [18]. We adapt primitive count and training length per dataset to accommodate different dataset sizes and spatial scales, but keep the same learning rates and other hyperparameters. Indoor datasets benefit from encoding the central camera ray, instead of per-ray directions, due to their high parallax, acting as a regularization on ray orientation. feature vectors at the vertices of the shape. For example, for a 2D triangle with vertices v 1 , v 2 , v 3 , the feature vector f (p) is computed as f (p) = n X i=1 α i f i (8) where f i are the features at the triangle vertices, and α i are the barycentric co- ordinates at the evaluated point. In the case of 2D triangles, barycentric weights are computed as α i = A i A , A = n X j=1 A j (9) Neural Harmonic Textures25 GTOursGTOurs Fig. 8: LSEG feature PCA visualizations on test views from MipNeRF360. Rows from top to bottom: bicycle, garden, bonsai, kitchen. Each pair shows the ground-truth LSEG features (GT) alongside our rendered features (Ours). Our method faithfully recon- structs semantic feature maps while preserving sharp boundaries, at higher resolutions, and real-time. where A is the total area of the triangle and A i is the area of the sub-triangle formed by the evaluation point p and the edge opposite to vertex i. For a triangle with vertices v 1 , v 2 , v 3 : A 1 = 1 2 â„(v 2 â p)Ă (v 3 â p)â„, A 2 = 1 2 â„(v 3 â p)Ă (v 1 â p)â„, A 3 = 1 2 â„(v 1 â p)Ă (v 2 â p)â„.(10) For 2D connected triangles, this is similar to vertex features recently intro- duced in the context of neural global illumination [52]. Our formulation however is more general and extends to higher order topologies and disconnected meshes. For tetrahedra, barycentric weights are similarly computed as α i = V i V , V = n X j=1 V j (11) 26J. Condor et al. Dim PSNRâ SSIMâ LPIPSâ LSEG P â LSEG C â FPSâ 423.81 0.853 0.25444.330.98876 825.40 0.876 0.20047.100.99475 16 26.06 0.885 0.17848.230.99572 32 26.34 0.887 0.16648.680.99672 64 26.42 0.887 0.16348.790.99670 Table 11: Feature dimensionality ablation for joint RGB + LSEG reconstruction on the truck scene (Tanks & Temples). LSEG P and LSEG C denote PSNR and cosine similarity of the 512-dim semantic features, respectively. Measured on an RTX A6000 Ada. ParameterValue LSEG output feature dimension512 LSEG L 1 loss weight λ LSEG 1.0 LSEG cosine similarity weight λ cos 0.1 Per-primitive feature dim. (shared with RGB) 80 Table 12: LSEG-specific hyperparameters used for joint RGB and semantic feature training. where V is the total volume of the tetrahedron and V i is the volume of the sub- tetrahedron formed by the evaluation point p and the face opposite to vertex i. For a tetrahedron with vertices v 1 , v 2 , v 3 , v 4 : V 1 = 1 6 â„(v 2 â p)Ă (v 3 â p)· (v 4 â p)â„, V 2 = 1 6 â„(v 3 â p)Ă (v 4 â p)· (v 1 â p)â„, V 3 = 1 6 â„(v 4 â p)Ă (v 1 â p)· (v 2 â p)â„, V 4 = 1 6 â„(v 1 â p)Ă (v 2 â p)· (v 3 â p)â„.(12) B.1 Optimization choices and hyperparameters Tab. 13 ablates the individual contributions of the training strategies described in the main paper (learning rate scheduling, EMA on the MLP weights, loss scaling, color-refinement phase, and opacity and scale regularization). Tab. 14 lists the full set of hyperparameters and optimization details used for radiance field reconstruction. C 2D Image Reconstruction Details We provide additional details on the high-resolution HDR image fitting exper- iment of Sec. 5.2. We adapt Neural Harmonic Textures to this domain while Neural Harmonic Textures27 ConfigurationPSNRâ SSIMâ LPIPS (Alex)â Our Model28.49 0.8280.153 A) No LR Scheduling28.37 0.8290.147 B) No EMA on MLP Weights28.46 0.8280.152 D) No Color Refinement Phase 28.35 0.8270.154 E) No Opacity Regularization 27.85 0.8060.181 F) No Scale Regularization28.50 0.8280.153 G) No Direction Encoding28.11 0.8160.155 I) No Direction Scaling28.35 0.8220.158 Table 13: Ablation study on training strategies. The rows are not additive, i.e. we only test one strategy at a time. All experiments on the MipNeRF360 dataset, using 64 features per primitive and a 128Ă3 MLP. Note that the scale regularization (F) does not significantly affect reconstruction quality, but does improve render time. preserving the core idea: learnable per-vertex features, interpolated per pixel and activated with periodic functions, and decoded by a shallow MLP. However, there are a number of differences to the 3D and semantic reconstruction tasks, which we detail below. C.1 Representation Triangle mesh. Unlike the volumetric disconnected primitives used in the scene reconstruction tasks, the 2D image fitting experiment employs a connected De- launay triangulation that tiles the image plane without overlap. Every pixel belongs to exactly one triangle; there is no alpha blending, no Gaussian kernel, and no depth ordering. Vertices can be freely learned and positioned, and carry individual learnable feature vectors. Given a query pixel, its containing triangle is found via a tile-accelerated rasterizer, and the features at the triangleâs three vertices are interpolated to produce a per-pixel feature vector. CloughâTocher C 1 interpolation. Standard barycentric (linear) interpolation produces a C 0 feature field: while continuous, the feature gradient is discon- tinuous across triangle edges. This turns out to be problematic when meshes are fully opaque, as it creates high frequency artifacts. Despite being followed by a frequency encoding and a nonlinear MLP, these gradient discontinuities manifest as visible seam artifacts along every edge. To eliminate this, we replace linear interpolation with CloughâTocher cubic interpolation, which constructs a degree-3 BĂ©zier patch on each triangle from vertex values and vertex gradients, guaranteeing C 1 continuity everywhere (and 28J. Condor et al. Training Color refinement phaseLast 3 000 steps (geometry frozen, no reg.) Loss & regularization SSIM weight λ (D-SSIM vs L 1 ) 0.2 Opacity regularization λ α 0.02 Scale regularization λ s 0.01 Learning rates (initial) Positions (means)1.6Ă 10 â4 Scales4Ă 10 â3 Opacities4Ă 10 â2 Quaternions8Ă 10 â4 Deferred features1.9Ă 10 â2 Deferred MLP7.2Ă 10 â4 LR schedule PositionsExponential decay (final factor 0.01) Features & MLPCosine annealing (final factor 0.1) EMA (deferred MLP) Decay Îł0.95 Start step0 Appearance (NHT) Feature dim. per primitive16-64 MLP hidden dim. Ă layers64â 128Ă 2â 3 View encodingSpherical harmonics (2nd degree) Table 14: Hyperparameters and optimization details for radiance field reconstruction (MipNeRF360, Tanks & Temples, Deep Blending). Values follow our gsplat implemen- tation [70]. C â within each triangle). The 10 BĂ©zier control points are: c 300 = f 0 , c 030 = f 1 , c 003 = f 2 , c 210 = f 0 + 1 3 âf 0 · e 01 , c 120 = f 1 â 1 3 âf 1 · e 01 , c 021 = f 1 + 1 3 âf 1 · e 12 , c 012 = f 2 â 1 3 âf 2 · e 12 , c 102 = f 2 + 1 3 âf 2 · e 20 , c 201 = f 0 â 1 3 âf 0 · e 20 , c 111 = 1 6 c 210 +c 120 +c 021 +c 012 +c 102 +c 201 â 1 6 c 300 +c 030 +c 003 ,(13) where f i andâf i are the feature value and gradient at vertex i, and e ij = v j âv i . Vertex gradient estimation. The CloughâTocher scheme requires per-vertex gra- dients. We estimate these via a per-vertex least-squares fit over the one-ring Neural Harmonic Textures29 Table 15: Ablating feature count N. Measured on an RTXA6000 Ada. Dim PSNRâ SSIMâ LPIPSâ FPSâ 427.18 0.809 0.288 165 827.83 0.818 0.261 156 16 28.22 0.826 0.246 142 24 28.31 0.828 0.240 134 32 28.39 0.829 0.237 133 48 28.43 0.829 0.231 116 64 28.48 0.828 0.230 105 80 28.47 0.828 0.22798 Table 16: Ablating MLP architecture (layer width Ă hidden layer count). Mea- sured on an RTXA6000 Ada. Size PSNRâ SSIMâ LPIPSâ FPSâ 16Ă2 28.16 0.824 0.24097 16Ă4 28.03 0.820 0.24498 32Ă2 28.25 0.825 0.23897 32Ă4 28.23 0.824 0.237 101 64Ă2 28.38 0.827 0.234 102 64Ă4 28.36 0.826 0.233 102 128Ă2 28.47 0.828 0.230 106 128Ă4 28.50 0.828 0.229 103 neighborhood: for each vertex v, we gather all neighbor vertices v j jâN(v) and solve âf v = A â1 v b v , A v = X j âp j âp †j , b v = X j âf j âp j ,(14) where âp j = p j â p v and âf j = f j âf v . The 2Ă2 system is solved analytically per vertex in a CUDA kernel. This operation is differentiable and computed at every forward pass, enabling gradients to flow through the vertex topology. Feature encoding and MLP. Feature encoding and MLP decoding are the same as in the main scene reconstruction task. C.2 Training Pixel sampling. At each iteration, we sample a batch of pixels via stratified sampling. Ground-truth values at sub-pixel coordinates are obtained via bilinear interpolation on the target image, implemented as a CUDA texture lookup for efficiency. Optimizers and schedules. We use Adam [26] with separate parameter groups for positions (lr = 10 â4 ), features (lr = 5Ă 10 â3 ), and MLP weights (lr = 5Ă 10 â5 ). An exponential decay schedule is applied after 20 000 iterations with a factor of 0.33 every 10 000 steps, following Instant NGP [44]. Loss function. We train with a pixel-wise MSE loss in ÎŒ-law encoded space. We found this to produce better results than using the relative L2 loss suggested in the Instant NGP paper. Raw 14-bit sensor values x are transformed as f (x) = log(1 + ÎŒx/w)/ log(1 + ÎŒ) with ÎŒ = 5000 and w = 2 14 â 1 = 16,383 (the white level of 14-bit data). This compresses the dynamic range before supervision, placing relatively more emphasis on shadow detail, which is standard in HDR imaging pipelines. 30J. Condor et al. C.3 Coarse-to-Fine Densification Training begins from a sparse mesh initialized via edge-aware sampling: we com- pute Sobel gradient magnitudes on a downsampled version of the target image and sample initial vertex positions proportionally to edge strength, with a small uniform floor for coverage. The result is a Delaunay triangulation of the sampled points, with boundary vertices ensuring full image coverage. The mesh is progressively refined through densification. Every 500 iterations (starting at iteration 1 500 and ending at 15 000), we score each triangle by score t =DSSIM t ·|pixels t | α ,(15) whereDSSIM t is the mean DSSIM (per-pixel structural dissimilarity) of tri- angle t and α = 0.75 is the area-weighting exponent. The area weighting pre- vents repeated splitting of tiny, high-error triangles and distributes vertices more evenly in proportion to their spatial impact. Triangles with fewer than 3 pixels are excluded from scoring. At each densification step, we insert centroid vertices for the top-scoring tri- angles (up to 0.35ĂV new vertices, where V is the current count), initialize their features by averaging the parent triangleâs vertex features, and re-run Delaunay triangulation on the full vertex set. We also detect when the mesh becomes degenerate (e.g. due to vertices becoming colinear) and re-mesh, although this rarely occurs in practice. C.4 CUDA Rasterizer All rasterization and interpolation operations are implemented as custom for- ward and backward CUDA kernels, supporting both full-image rendering and batched pixel-query modes. Training generally takes between 1-3 minutes in an RTXA6000 Ada. C.5 Post-Training Compression We also briefly explore the use of post-training compression techniques to further reduce storage requirements. We compress vertex positions to UInt16, vertex features to uniform Int8 (with a per-channel scale and offset), and MLP weights to FP16. An entropy coder (Zstandard, level 19) is applied to the serialized payload. Combining these techniques, the resulting images can be compressed by an additional factor of approximately 3Ă beyond the training result, with negligible quality loss. Tab. 17 summarizes the results of the post-training compression techniques. C.6 Per-Image Results Tabs. 18 and 19 provide per-image results for the 2D image fitting experiment at 10Ă and 100Ă compression ratios, respectively. We provide visual comparisons Neural Harmonic Textures31 SchemeBytes BPP Ratio PSNRâ (ÎŒ) PSNRâ (lin) PSNRâ (tm) SSIMâ (ÎŒ) SSIMâ (tm) Baseline4,809,236 0.8418 100Ă 33.8938.4534.780.88300.9163 int8 features2,418,189 0.4233 198.4Ă 33.8538.3334.710.88270.9160 int16 positions4,011,138 0.7021 119.6Ă 33.8938.4534.780.88300.9163 int8+int16+fp16mlp 1,607,793 0.2814 298.4Ă 33.8538.3234.700.88270.9160 + Zstandard1,472,273 0.2577 326.0Ă 33.8538.3234.700.88270.9160 + Brotli1,450,542 0.2539 331.0Ă 33.8538.3234.700.88270.9160 Table 17: Impact of post-training compression on NHT at 100Ă. Averages over 15 im- ages. Baseline: uncompressed fp32 features/parameters + fp16 MLP weights. int8 fea- tures: uniform 8-bit quantization with per-channel scale/offset. int16 positions: fixed- point uint16 vertex positions. int8+int16+fp16: combined quantization (features int8, positions int16, MLP fp16). + Zstandard/Brotli: entropy coding on the serialized pay- load. in Fig. 7 and Fig. 9. Our method achieves competitive PSNR in both ÎŒ-law and tonemapped spaces while substantially outperforming Instant NGP in perceptual quality (LPIPS). The advantage is most pronounced at high compression ratios, where the non-overlapping mesh topology and C 1 feature interpolation avoid the ringing and block artifacts that plague hash-grid methods. Tab. 20 provides the full set of hyperparameters for the 2D image fitting experiment. 32J. Condor et al. PSNRâ (ÎŒ-law) PSNRâ (tonemapped) SSIMâ (tonemapped) LPIPSâ (tonemapped) Image NHT NGP JXL NHT NGP JXL NHT NGP JXL NHT NGP JXL DSC_1012 33.98 34.0437.70 36.9237.59 30.38 0.970.96 0.99 0.00900.0137 0.0004 DSC_1764 40.06 39.04 49.03 43.61 41.7332.02 0.980.97 0.99 0.04010.1149 0.0012 DSC_1967 33.60 35.0141.48 36.50 36.1927.92 0.950.93 0.99 0.01080.0178 0.0004 DSC_2172 35.83 36.59 41.74 40.2338.95 40.31 0.970.95 0.99 0.01820.0317 0.0008 DSC_2655 41.8040.91 50.26 44.3842.73 55.95 0.980.981.00 0.02920.0850 0.0008 DSC_3095 38.07 38.7846.05 39.8338.43 52.47 0.960.94 1.00 0.01840.0450 0.0009 DSC_3351 40.00 42.2048.40 44.5342.97 55.37 0.980.981.00 0.00460.0179 0.0002 DSC_4748 33.39 34.6837.97 37.53 37.0831.97 0.950.94 0.99 0.02180.0544 0.0010 DSC_5007 36.09 36.8047.75 37.4636.54 49.40 0.930.90 0.99 0.02500.0430 0.0009 DSC_6718 41.3040.15 51.27 44.5741.64 55.28 0.980.97 1.00 0.04640.1345 0.0010 DSC_6926 31.24 33.1534.93 33.2734.28 27.42 0.930.91 0.98 0.02460.0449 0.0022 DSC_9008 40.7840.10 49.73 41.9339.98 53.74 0.980.96 1.00 0.01110.0493 0.0006 DSC_1658 39.0538.51 50.39 41.6839.57 53.53 0.970.95 1.00 0.04710.1445 0.0018 DSC_4176 35.17 36.89 48.17 36.50 36.6450.77 0.950.93 1.00 0.02150.0597 0.0006 DSC_7685 37.11 37.5444.83 39.08 39.3353.30 0.970.96 1.00 0.01840.0550 0.0004 Average 37.16 37.6345.31 39.8738.91 44.66 0.960.95 0.99 0.02310.0608 0.0009 Table 18: Per-image comparison at 10Ă compression on our 14-bit HDR RAW dataset (45.7 MP). We compare NHT (Ours), Instant NGP [44], and JPEG-XL [23]. All neural methods use the same parameter budget. Metrics on tonemapped sRGB images unless noted. PSNRâ (ÎŒ-law) PSNRâ (tonemapped) SSIMâ (tonemapped) LPIPSâ (tonemapped) Image NHT NGP JXL NHT NGP JXL NHT NGP JXL NHT NGP JXL DSC_1012 31.49 31.6532.58 33.3433.51 30.18 0.930.930.98 0.02430.0255 0.0065 DSC_1764 37.4137.22 41.89 39.95 39.7031.89 0.960.960.98 0.05520.1341 0.0344 DSC_1967 30.65 31.0534.25 32.4132.61 27.66 0.890.88 0.96 0.03280.0341 0.0057 DSC_2172 32.96 33.3136.09 35.59 35.9038.59 0.920.920.97 0.04190.0436 0.0182 DSC_2655 40.3040.16 43.56 41.74 41.3626.05 0.980.980.99 0.03960.0942 0.0121 DSC_3095 36.5235.93 39.03 36.7435.82 43.61 0.930.92 0.98 0.04590.0663 0.0184 DSC_3351 40.79 41.2343.53 42.1242.51 29.01 0.980.980.99 0.00990.0193 0.0020 DSC_4748 31.05 31.2632.83 34.2334.38 31.47 0.910.910.96 0.05160.0724 0.0142 DSC_5007 33.88 34.00 38.84 33.65 33.6836.35 0.850.84 0.95 0.06920.0729 0.0143 DSC_6718 38.55 38.6843.99 40.7440.19 47.59 0.960.960.99 0.08470.1571 0.0362 DSC_6926 29.48 29.72 30.01 29.9130.54 27.02 0.870.86 0.93 0.05870.0629 0.0331 DSC_9008 38.38 38.3943.15 37.9037.92 34.54 0.950.950.98 0.02750.0514 0.0096 DSC_1658 36.5636.45 41.84 37.8737.29 44.60 0.930.930.98 0.08900.1768 0.0489 DSC_4176 33.49 33.5638.28 33.60 33.6340.03 0.900.89 0.97 0.05210.0880 0.0123 DSC_7685 34.42 34.9138.10 36.63 36.8644.49 0.940.940.98 0.03450.0681 0.0081 Average 35.06 35.1738.53 36.43 36.3935.54 0.930.92 0.97 0.04780.0778 0.0183 Table 19: Per-image comparison at 100Ă compression (same dataset and setup as Tab. 18). Neural Harmonic Textures33 Instant NGPNHT (Ours)Reference 42.5 dB / 0.01942.1 dB / 0.010DSC_3351 40.2 dB / 0.15740.7 dB / 0.085DSC_6718 39.7 dB / 0.13439.9 dB / 0.055DSC_1764 30.5 dB / 0.06329.9 dB / 0.059DSC_6926 36.9 dB / 0.06836.6 dB / 0.035DSC_7685 Fig. 9: Visual comparison at 100Ă compression. Each cell shows tonemapped PSNR (dB) and LPIPS. Our method (NHT) consistently achieves lower LPIPS (better perceptual quality) than Instant NGP at comparable or higher PSNR. The images are taken from the dataset we curated. 34J. Condor et al. Representation Mesh topologyDelaunay triangulation (shared vertices) InterpolationCloughâTocher C 1 cubic Feature activationSinCos: [sin(f), cos(f)] Feature dim. per vertex4-16 MLP hidden dim. 64â 128 Ă 2â 4 layers Training Total iterations25 000 Pixel samplingStratified, 160 000 pixels/batch Loss functionMSE Learning rates (initial) Vertex positions1Ă 10 â4 Vertex features5Ă 10 â3 MLP weights5Ă 10 â5 LR schedule All parametersExponential decay: η(t) = η 0 · 0.33 max(0, (tâ20k)/ 10k) Densification StrategyCoarse-to-fine Delaunay re-meshing ScheduleSteps 1 500â15 000, every 500 steps Growth rate0.35Ă current vertex count per step ScoringDSSIM-based, area-weighted: score t =DSSIM t ·|pixels t | 0.75 Min. triangle size3 pixels (skip smaller triangles) InitializationEdge-aware sampling (Sobel) HDR pipeline Input14-bit RAW (Nikon NEF), 45.7 MP Encoding spaceÎŒ-law: f(x) = log(1 + ÎŒx/w)/ log(1 + ÎŒ) ÎŒ parameter5 000 White level w2 14 â 1 = 16 383 GT pixel lookupBilinear interpolation Table 20: Hyperparameters and optimization details for the 2D image fitting experi- ment. The parameter budget (maximum vertices, MLP size) is derived from the target compression ratio and image dimensions and bitrate.