Paper deep dive
Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS
Kangmin Seo, Jae-Pil Heo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/28/2026, 4:28:27 AM
Summary
This paper introduces a training-free filtering procedure for feed-forward 3D Gaussian Splatting (3DGS) models to remove distractors (transient objects) from reconstructions. The method exploits the per-view association of predicted Gaussians: by excluding Gaussians associated with a specific input view and re-rendering, it identifies regions inconsistent with other views. These regions form candidate distractors, which are verified by checking if their removal reduces reconstruction error in other input views. The approach operates on a single frozen prediction without retraining or scene-specific optimization, consistently improving novel-view quality across multiple models and benchmarks while preserving clean scene reconstructions.
Entities (10)
Relation Signals (9)
Feed-Forward 3D Gaussian Splatting → produces → Gaussian Primitives
confidence 95% · Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images
Distractor Filtering → mitigates → Artifacts
confidence 94% · We introduce a training-free filtering procedure... consistently improves novel-view quality
Gaussian Primitives → associatedwith → Input Views
confidence 93% · The reconstruction models considered in this work retain an association between predicted Gaussians and their input views
Transient Objects → cause → Artifacts
confidence 92% · Such content can be encoded into the per-view Gaussians... As a result, it may produce blurred, duplicated, or floating artifacts in novel views.
Distractor Filtering → uses → Rendering-based Verification
confidence 91% · rendering-based verification retains only candidates whose removal reduces reconstruction error
ReSplat → evaluatedon → NeRF On-the-go
confidence 90% · We evaluate our method on... ReSplat... using... NeRF On-the-go
DepthSplat → evaluatedon → RobustNeRF
confidence 90% · We evaluate our method on DepthSplat... using RobustNeRF
Distractor Filtering → uses →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images in one network execution, making 3D reconstruction increasingly accessible for casual captures. However, such captures frequently contain transient objects that appear in only a subset of the views. Such content can be encoded into the per-view Gaussians associated with the inputs that observe it and remain in the combined representation despite being observed by no other input. As a result, it may produce blurred, duplicated, or floating artifacts in novel views. We introduce a training-free filtering procedure that exploits this per-view prediction structure. For each input, we exclude its associated Gaussians and render the same camera using the remaining representation, revealing content that is inconsistent with the other inputs. Feature similarity forms candidate regions, and rendering-based verification retains only candidates whose removal reduces reconstruction error in the other input views. The procedure operates on a single frozen prediction without retraining or scene-specific optimization. Across three reconstruction models and two distractor benchmarks, it consistently improves novel-view quality with varying numbers of input views. On clean scenes, evaluations across four models show that the original reconstructions are largely preserved.
Tags
Links
- Source: https://arxiv.org/abs/2608.26951v1
- Canonical: https://arxiv.org/abs/2608.26951v1
Trouble viewing inline? Open PDF directly →
Full Text
44,518 characters extracted from source content.
Expand or collapse full text
Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS Kangmin Seo Jae-Pil Heo Abstract Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images in one network execution, making 3D reconstruction increasingly accessible for casual captures. However, such captures frequently contain transient objects that appear in only a subset of the views. Such content can be encoded into the per-view Gaussians associated with the inputs that observe it and remain in the combined representation despite being observed by no other input. As a result, it may produce blurred, duplicated, or floating artifacts in novel views. We introduce a training-free filtering procedure that exploits this per-view prediction structure. For each input, we exclude its associated Gaussians and render the same camera using the remaining representation, revealing content that is inconsistent with the other inputs. Feature similarity forms candidate regions, and rendering-based verification retains only candidates whose removal reduces reconstruction error in the other input views. The procedure operates on a single frozen prediction without retraining or scene-specific optimization. Across three reconstruction models and two distractor benchmarks, it consistently improves novel-view quality with varying numbers of input views. On clean scenes, evaluations across four models show that the original reconstructions are largely preserved. Sungkyunkwan University skmskku@skku.edu, jaepilheo@skku.edu Introduction Recent advances in neural radiance fields (Mildenhall et al. 2021) and 3D Gaussian Splatting (Kerbl et al. 2023) have substantially improved novel view synthesis. More recently, feed-forward 3DGS models directly predict a Gaussian representation from a set of input images in one network execution (Charatan et al. 2024; Chen et al. 2024; Xu et al. 2025b; Ye et al. 2025b; Jiang et al. 2025), enabling fast and generalizable reconstruction. These methods commonly assume that the input images depict a static and view-consistent scene. Casual captures, however, often contain pedestrians, vehicles, and other transient objects that appear in only one or a small subset of the images. When incorporated into the predicted representation, such distractors can produce blurry regions, duplicated appearance, and floating artifacts in novel views. Figure 1: Qualitative overview. Distractors in the input views produce artifacts in the feed-forward reconstruction, which our method suppresses on the same frozen prediction. Distractor handling has primarily been studied in optimization-based NeRF and 3DGS pipelines. Existing methods model transient appearance (Martin-Brualla et al. 2021), use robust objectives or uncertainty estimation (Sabour et al. 2023; Ren et al. 2024), leverage pretrained features or residual cues (Kulhanek et al. 2024; Sabour et al. 2025), separate static and transient content (Wang et al. 2025b; Lin et al. 2025; Park et al. 2026), or use consistency and multistage filtering (Li et al. 2025; Seo et al. 2026; Wang et al. 2026). While fitting a scene to all input images, differences between its renderings and individual observations can provide useful evidence of unsupported content. With the emergence of feed-forward reconstruction, recent work has introduced distractor-aware variants through dedicated training, learned mask prediction, reconstruction input selection, or Gaussian pruning (Bao et al. 2024; Gupta et al. 2026; Pan et al. 2026). These approaches address distractors by adapting particular reconstruction backbones, relying on distractor-specific trained components such as mask prediction heads or semantic segmentation priors. Consequently, transferring distractor robustness to a different reconstruction backbone may require model-specific adaptation. In contrast, we ask whether such robustness can be obtained from the reconstruction itself, without modifying the model. An immediate candidate for such a cue is the rendering-input discrepancy exploited in optimization-based pipelines. In the feed-forward setting, however, this cue is obscured. The reconstruction models considered in this work retain an association between predicted Gaussians and their input views (Xu et al. 2025b; Xu et al. 2025a; Ye et al. 2025a; Jiang et al. 2025). A model can therefore preserve content from an individual input even when that content is inconsistent with the remaining observations. The associated Gaussians may reproduce it when rendered from the corresponding camera, concealing the discrepancy with the input image. The same Gaussians nevertheless remain in the shared 3D representation and may produce artifacts from other viewpoints. Figure 2: View-associated prediction in feed-forward 3DGS. Content from individual inputs, including inconsistent distractors, can be encoded in per-view Gaussians and remain in the combined reconstruction. Our key insight is to repurpose this per-view prediction structure as a detection cue. By excluding the Gaussians associated with one input, we expose content that is inconsistent with the remaining observations. The explicit Gaussian representation then allows us to measure whether removing each proposed subset improves agreement with the other inputs. Thus, the structure that can preserve inconsistent distractors also provides the means to identify and filter them. Based on this insight, we propose a training-free procedure for filtering Gaussians associated with distractors from frozen feed-forward reconstruction models. For each input image, we compare it with a rendering obtained after excluding its associated Gaussians. Regions that the remaining inputs do not explain form separate removal candidates. We temporarily remove the Gaussians associated with each candidate and evaluate their effect from selected other input views. Only verified candidates are used to determine the final Gaussian subsets for removal. All operations are performed after one execution of the frozen reconstruction model; the model is neither modified nor executed again. Our method requires no additional training, learned distractor masks, or predefined distractor categories, and operates on the given inputs alone. Distractor robustness is thus added to already trained feed-forward 3DGS models as a drop-in post-processing step. We evaluate our method on DepthSplat (Xu et al. 2025b), ReSplat (Xu et al. 2025a), and YoNoSplat (Ye et al. 2025a) using RobustNeRF (Sabour et al. 2023) and NeRF On-the-go (Ren et al. 2024) with varying numbers of input views. Each + Ours result starts from the same frozen Gaussian prediction as its corresponding baseline, isolating the effect of filtering. We further evaluate preservation on clean scenes across four models, including AnySplat (Jiang et al. 2025), report GenWildSplat (Gupta et al. 2026) as a trained reference, and conduct ablation and runtime analyses. Our contributions are as follows: • We show that the native association between input views and predicted Gaussians, which allows distractors to persist in the reconstruction, can be repurposed as a cue for training-free distractor filtering. • We propose a filtering procedure that forms removal candidates by excluding view-associated Gaussians and accepts only candidates verified against the other input observations, operating on a single frozen prediction. • We demonstrate consistent improvements across three reconstruction models and two distractor benchmarks with varying numbers of input views, while largely preserving reconstruction quality on clean scenes across four models. Related Work Figure 3: Overview of our training-free filtering procedure. For each input, we exclude its associated Gaussian subset to form separate candidate regions. Each candidate is first evaluated from selected other input views, and only candidates that pass the corresponding verification steps are removed. Feed-Forward 3D Gaussian Splatting Following the development of neural radiance fields (Mildenhall et al. 2021; Barron et al. 2022), 3D Gaussian Splatting (Kerbl et al. 2023) has driven rapid advances in novel view synthesis across diverse scene settings and applications (Lu et al. 2024; Matsuki et al. 2024; Kheradmand et al. 2024; Wu et al. 2024; Qin et al. 2024; Liu et al. 2024; Xie et al. 2024). These approaches typically optimize a separate representation for each scene. A recent direction instead develops generalizable feed-forward models that directly predict Gaussian representations from input images, avoiding per-scene optimization. pixelSplat (Charatan et al. 2024) introduced feed-forward Gaussian reconstruction from image pairs, while MVSplat (Chen et al. 2024) extended this formulation to sparse posed views. DepthSplat (Xu et al. 2025b) further combines Gaussian reconstruction with learned depth estimation. Subsequent methods support unposed or otherwise less constrained inputs (Ye et al. 2025b; Jiang et al. 2025; Ye et al. 2025a), recurrent refinement (Xu et al. 2025a), token-aligned prediction (Li et al. 2026), voxel-aligned prediction (Wang et al. 2025a), and adaptive subpixel primitive prediction (Moreau et al. 2026). The reconstruction models considered in this work preserve an association between predicted Gaussians and their input views. Their explicit Gaussian outputs also allow selected subsets to be removed temporarily and rendered again. We use these properties to inspect the contribution of each input after the scene has been predicted and to identify Gaussian subsets whose removal improves the reconstruction. Ignoring Distractors in 3D Reconstruction Distractor-free novel view synthesis has primarily been studied through optimization-based NeRF and 3DGS pipelines. NeRF-based approaches model transient appearance (Martin-Brualla et al. 2021), suppress inconsistent observations with robust objectives (Sabour et al. 2023), or exploit uncertainty in casually captured scenes (Ren et al. 2024). Related 3D Gaussian Splatting methods leverage pretrained features to handle inconsistent observations (Kulhanek et al. 2024; Sabour et al. 2025), explicitly separate static and transient representations (Wang et al. 2025b; Lin et al. 2025; Park et al. 2026), or suppress unstable artifacts through consistency between independently optimized models (Li et al. 2025). Other methods derive cleaner supervision through progressive or two-stage reconstruction (Seo et al. 2026; Wang et al. 2026). These approaches identify distractors while optimizing a scene-specific representation against its input images. Recent work has also introduced feed-forward reconstruction models trained specifically for distractor handling. VGTW (Pan et al. 2026) builds on VGGT and introduces distractor-aware training together with an auxiliary mask head supervised by pixel-level annotations. DGGS (Bao et al. 2024) builds on MVSplat (Chen et al. 2024) and learns mask prediction and refinement from reference images. Its reported inference procedure uses the predicted masks to score and reselect references before pruning Gaussians associated with distractors. GenWildSplat (Gupta et al. 2026), based on AnySplat (Jiang et al. 2025), uses pretrained semantic segmentation for transient-object masking together with curriculum learning on synthetic and real data. These approaches incorporate distractor handling into the reconstruction pipeline itself, relying on additional trained components or backbone-specific adaptation. Our method instead decouples distractor handling from the reconstruction model: it operates on one Gaussian prediction produced from a fixed set of inputs, with no additional training, predefined distractor categories, or learned mask prediction. Under its reported evaluation protocol, DGGS assumes access to a scene image pool around each query view, from which reconstruction inputs are scored and reselected before reconstruction is executed again. Our method assumes only a fixed input set: the reconstruction model is executed once, and filtering operates on its resulting Gaussian representation. Method 4 views 8 views 16 views Method PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ RobustNeRF dataset DepthSplat 20.10 0.738 0.267 19.91 0.706 0.288 18.44 0.644 0.369 + Ours 21.66 0.762 0.229 21.37 0.735 0.242 19.98 0.690 0.296 YoNoSplat 21.70 0.722 0.260 23.12 0.776 0.230 22.70 0.732 0.249 + Ours 22.47 0.729 0.246 24.04 0.791 0.199 23.26 0.743 0.226 ReSplat 22.24 0.778 0.270 22.73 0.763 0.280 24.71 0.797 0.238 + Ours 22.87 0.782 0.264 23.12 0.768 0.273 24.83 0.799 0.235 NeRF On-the-go dataset DepthSplat 17.86 0.549 0.306 17.76 0.524 0.336 16.32 0.447 0.416 + Ours 18.44 0.563 0.285 18.67 0.547 0.298 17.52 0.489 0.355 YoNoSplat 18.35 0.547 0.329 18.66 0.561 0.326 17.95 0.532 0.350 + Ours 19.61 0.577 0.289 19.64 0.583 0.295 19.13 0.564 0.302 ReSplat 18.88 0.618 0.341 19.12 0.610 0.341 19.47 0.625 0.319 + Ours 19.36 0.629 0.330 19.86 0.625 0.325 19.88 0.635 0.309 Table 1: Quantitative results on the RobustNeRF and NeRF On-the-go datasets. Figure 4: Qualitative comparisons across frozen feed-forward reconstruction backbones. Each ++ Ours result filters the corresponding baseline prediction, suppressing distractor artifacts while largely preserving surrounding scene content. Preliminaries 3D Gaussian Splatting. 3D Gaussian Splatting (Kerbl et al. 2023) represents a scene using NgN_g Gaussian primitives =gnn=1NgG=\g_n\_n=1^N_g, each carrying a position, covariance, opacity, and view-dependent color. We denote by ℛ(,P)R(G,P) the image rendered from G under camera parameters P via depth-ordered alpha compositing. Its explicit representation allows any Gaussian subset to be removed and the result rendered immediately. Feed-Forward 3DGS. Given N posed image-camera pairs =(Ii,Pi)i=1N,X=\(I_i,P_i)\_i=1^N, (1) a feed-forward model FθF_θ predicts a Gaussian representation in one network execution (Charatan et al. 2024; Chen et al. 2024; Xu et al. 2025b). We express its output as Gaussian subsets associated with the input views: ii=1N=Fθ(),=⋃i=1Ni.\G_i\_i=1^N=F_θ(X), = _i=1^NG_i. (2) Here, iG_i contains the Gaussian primitives associated with input view i. For pixel-aligned models, i=gi,u∣u∈Ωi,G_i=\g_i,u u∈ _i\, (3) where Ωi _i is the prediction domain and u indexes a prediction location. The center μi,u _i,u of gi,ug_i,u is typically obtained by unprojecting a predicted depth di,ud_i,u: μi,u=Π−1(u,di,u,Pi), _i,u= ^-1(u,d_i,u;P_i), (4) where Π−1 ^-1 denotes unprojection under camera PiP_i. More generally, iG_i follows the association provided by the native prediction structure of the reconstruction model. Problem Setup We consider reconstruction from a fixed set of N posed input images that may contain transient objects visible in only a subset of the views. At inference, no clean images, distractor masks, or additional images for input selection are available. A frozen feed-forward model produces G once, and neither the network nor the predicted representation is optimized. Our goal is to identify Gaussian subsets associated with distractors and to remove a subset only when the input observations themselves justify its removal. The partition ii=1N\G_i\_i=1^N allows us to exclude temporarily the contribution associated with each input. The explicit Gaussian representation then allows us to render the scene before and after removing a selected subset and measure the resulting change directly. Distractor Proposal For a quantity associated with input view i, the superscript −i-i indicates that it is computed after excluding iG_i. We render camera PiP_i before and after excluding this subset: Ri=ℛ(,Pi),Ri−i=ℛ(∖i,Pi).R_i=R(G,P_i), R_i^-i=R(G _i,P_i). (5) Content in IiI_i represented mainly by iG_i, including distractors observed only in view i, is unlikely to be reproduced in Ri−iR_i^-i. We let Φ denote the DINOv3 feature map (Siméoni et al. 2025), computed from each image. The feature similarity at patch u is si(u)=sim(Φ(Ii)(u),Φ(Ri−i)(u)),s_i(u)=sim ( (I_i)(u), (R_i^-i)(u) ), (6) where simsim denotes cosine similarity. Feature similarity is less sensitive than direct RGB comparison to small appearance differences between the input and rendering. Low feature similarity may also arise from poorly reconstructed static content or from regions that no remaining view observes. We therefore use accumulated opacity and depth to restrict proposal generation. Let ai−i(u)a_i^-i(u) denote the patch-average accumulated opacity in Ri−iR_i^-i, and zi(u)z_i(u) and zi−i(u)z_i^-i(u) the patch-average depths before and after excluding iG_i, computed over pixels with valid depth in both renderings; patches without such pixels do not satisfy the depth condition. Let gi(u)g_i(u) denote the support condition ai−i(u)>ηa_i^-i(u)>η and zi(u)<zi−i(u)z_i(u)<z_i^-i(u), with opacity threshold η. Using similarity boundaries τ1<τ2 _1< _2, we define two proposal ranges: Bi1=u∣si(u)<τ1,gi(u),B_i^1=\u s_i(u)< _1,\;g_i(u)\, (7) Bi2=u∣τ1≤si(u)<τ2,gi(u).B_i^2=\u _1≤ s_i(u)< _2,\;g_i(u)\. (8) The opacity condition removes regions that are left largely empty by the remaining views. The depth condition requires exclusion of iG_i to reveal a farther surface, as expected when the removed Gaussians occlude other scene content. We extract 8-connected components independently from the two ranges on the feature grid and upsample each component to the input image resolution using nearest-neighbor interpolation. Components from different ranges are not merged, even when they touch after upsampling, and no minimum component size is imposed. Across all input views, the resulting component masks are denoted by Mkk=1L\M_k\_k=1^L, and iki_k denotes the input view from which MkM_k was obtained. We use rk∈1,2r_k∈\1,2\ to indicate whether the component was extracted from Bik1B_i_k^1 or Bik2B_i_k^2. Each component mask is mapped to a candidate Gaussian subset k⊂ikC_k _i_k. For pixel-aligned models, we use the native correspondence between prediction locations and Gaussians. For ReSplat, we use its native rasterizer association to select front-surface Gaussians from input iki_k whose projected covariances contribute to pixels inside MkM_k. Let Mi1M_i^1 and Mi2M_i^2 denote the unions of the component masks obtained from the first and second similarity ranges, respectively, for input view i. Candidate Verification The proposal masks identify possible distractors rather than final removals. We first evaluate each component independently. For input view i, we project the centers of iG_i into every other input camera. For each camera, we count the projected centers that have positive depth and fall inside the image boundary. The nvn_v cameras with the largest counts form the verification set iV_i; camera PiP_i itself is not included. For candidate kC_k and each j∈ikj _i_k, we render the scene before and after temporarily removing the candidate: Rj=ℛ(,Pj),Rj−k=ℛ(∖k,Pj).R_j=R(G,P_j), R_j^-C_k=R(G _k,P_j). (9) The corresponding reconstruction errors are Ej(p)=‖Rj(p)−Ij(p)‖1,E_j(p)= \|R_j(p)-I_j(p) \|_1, (10) Ej−k(p)=‖Rj−k(p)−Ij(p)‖1.E_j^-C_k(p)= \|R_j^-C_k(p)-I_j(p) \|_1. (11) The error reduction caused by removing candidate k is δj,k(p)=Ej(p)−Ej−k(p). _j,k(p)=E_j(p)-E_j^-C_k(p). (12) A positive value indicates that removing kC_k reduces the reconstruction error at pixel p. We weight each pixel by the rendering change caused by the candidate: wj,k(p)=‖Rj−k(p)−Rj(p)‖1.w_j,k(p)= \|R_j^-C_k(p)-R_j(p) \|_1. (13) Proposal regions are excluded because they may contain inconsistent content: we exclude Mj1M_j^1 for a candidate from the first similarity range, and Mj1∪Mj2M_j^1∪ M_j^2 for one from the second. Let Qj,kQ_j,k denote the corresponding exclusion mask. We also ignore pixels whose rendering change is below ϵε: Ωj,k=p∣p∉Qj,k,wj,k(p)>ϵ. _j,k=\p p∉ Q_j,k,\;w_j,k(p)>ε\. (14) We pool all valid pixels from the selected views into one verification score: Δk=∑j∈ik∑p∈Ωj,kwj,k(p)δj,k(p)∑j∈ik∑p∈Ωj,kwj,k(p). _k= _j _i_k _p∈ _j,kw_j,k(p)\, _j,k(p) _j _i_k _p∈ _j,kw_j,k(p). (15) A positive score indicates that removing kC_k improves agreement with the selected input observations outside the proposal regions. If the denominator is zero, we set Δk=0 _k=0. Candidates with positive individual scores proceed to final Gaussian selection. We dilate their component masks at the input image resolution and intersect the expanded masks again with the support condition gikg_i_k used during proposal generation. Each resulting mask is mapped to a subset ^k⊂ik C_k _i_k using the same model-specific association as above. Candidates from the first similarity range, whose lower similarity already indicates a clear mismatch, require no further test. Candidates from the second range carry weaker evidence, and we verify them further. For each such candidate, we additionally render every input camera after excluding the Gaussian subset associated with that camera: Rj−j=ℛ(∖j,Pj),R_j^-j=R(G _j,P_j), (16) and after additionally removing the expanded candidate: Rj−j,−^k=ℛ(∖(j∪^k),Pj).R_j^-j,- C_k=R (G (G_j∪ C_k ),P_j ). (17) Using the same rendering-change-weighted error reduction as above, we aggregate all valid pixels over all input views while excluding Mj1∪Mj2M_j^1∪ M_j^2. The candidate proceeds only when this additional score is positive. Finally, the second-range candidates that pass both individual tests are evaluated together: starting from the representation with the accepted first-range candidates removed, we additionally remove all remaining second-range candidates and aggregate the same rendering-change-weighted error reduction over all input views outside Mj1∪Mj2M_j^1∪ M_j^2. They are retained when this score is positive and restored together otherwise; first-range candidates are unaffected. Final Reconstruction Let 1A_1 denote candidates from the first similarity range with positive individual scores, and 2A_2 those from the second range that pass the individual and additional steps. If their combined verification fails, we set 2=∅A_2= . The filtered Gaussian representation is out=∖⋃k∈1∪2^k.G_out=G _k _1 _2 C_k. (18) In implementation, selected Gaussians are removed by setting their opacity to zero; all others remain unchanged. Experiments Experimental Setup. We evaluate on the RobustNeRF (Sabour et al. 2023) and NeRF On-the-go (Ren et al. 2024) benchmarks using 4, 8, and 16 input views containing distractors, covering sparse to moderately dense capture settings. For each scene and input count, we construct four view configurations and average the results over them. Each configuration starts from an input view and a test view sharing high COLMAP (Schonberger and Frahm 2016) sparse-point visibility, and views are added to increase the scene coverage jointly observed by both sets. We report PSNR, SSIM (Wang et al. 2004), and LPIPS (Zhang et al. 2018) on the test images. Each reconstruction model is evaluated at the resolution used by its released inference configuration. We use τ1=0.50 _1=0.50, τ2=0.70 _2=0.70, nv=3n_v=3, and η=0.50η=0.50 for all datasets. Components are extracted with 8-connectivity without a minimum-size threshold. Candidates that pass individual verification are dilated by 9 pixels before their final Gaussian subsets are determined. The additional verification for the second similarity range aggregates all input views. We use ϵ=0.002ε=0.002 and accept a verification result only when its score is positive. These settings are fixed across all datasets, input counts, and reconstruction models. No training, fine-tuning, or scene-specific optimization is performed. Experiments were run on NVIDIA RTX 3090 and RTX 4090 GPUs. All runtime measurements are obtained on a single RTX 4090. Variant PSNR↑ SSIM↑ LPIPS↓ Full Method 20.42 0.657 0.279 RGB Difference 20.31 0.654 0.284 Single Threshold (<τ2< _2) 20.10 0.649 0.291 Single Threshold (<τ1< _1) 20.32 0.655 0.282 Without Verification 20.02 0.645 0.296 Vanilla 19.56 0.641 0.301 Table 2: Ablation study. 4 views 6 views Method PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ RobustNeRF dataset AnySplat 13.97 0.457 0.492 15.78 0.506 0.438 + Ours 14.48 0.469 0.468 16.63 0.526 0.400 GenWildSplat 13.99 0.484 0.516 14.39 0.482 0.532 NeRF On-the-go dataset AnySplat 13.66 0.287 0.469 14.96 0.328 0.437 + Ours 13.96 0.293 0.447 15.38 0.336 0.411 GenWildSplat 13.50 0.315 0.544 13.73 0.315 0.546 Table 3: Comparison with feed-forward models in the four- and six-view settings. Our method filters the frozen AnySplat prediction; GenWildSplat uses its released model. Reconstruction Models and Comparisons. We apply our method to the released DepthSplat (Xu et al. 2025b), ReSplat (Xu et al. 2025a), and YoNoSplat (Ye et al. 2025a) models. All three are evaluated with input camera poses, which removes pose estimation as a confounding factor. ReSplat uses its 8-view checkpoint for up to 8 inputs and its 16-view checkpoint for 16. Each baseline result and its corresponding + Ours result use identical images, cameras, model weights, and initial Gaussian prediction; only the selected Gaussian subsets differ, isolating the effect of filtering. We additionally evaluate AnySplat (Jiang et al. 2025) with and without our method, using its estimated cameras for proposal generation and verification, and include the released GenWildSplat (Gupta et al. 2026) model as an AnySplat-based reference. GenWildSplat targets unconstrained image collections and uses semantic masks for predefined transient object categories. Since it reports reconstruction with two to six input views, we use four- and six-view settings. DGGS (Bao et al. 2024) is excluded, as no implementation or weights were publicly available at submission and its inference requires additional scene images beyond the given inputs. Quantitative Results. Table 1 compares each reconstruction before and after filtering. Our method improves PSNR, SSIM, and LPIPS across all evaluated reconstruction models, datasets, and input counts. Since each pair shares the same images, model weights, and Gaussian prediction, these gains are attributable to filtering alone. 4 views 6 views Method PSNR↑ SSIM↑ LPIPS↓ PSNR↑ SSIM↑ LPIPS↓ With GT Pose DepthSplat 23.14 0.804 0.169 22.04 0.769 0.197 + Ours 23.14 0.804 0.170 22.04 0.769 0.197 YoNoSplat 24.98 0.802 0.167 25.24 0.813 0.163 + Ours 24.98 0.802 0.169 25.17 0.811 0.165 ReSplat 24.88 0.809 0.224 24.49 0.799 0.232 + Ours 24.83 0.809 0.224 24.49 0.799 0.232 With Estimated Pose AnySplat 16.40 0.506 0.412 17.69 0.558 0.348 + Ours 16.39 0.506 0.413 17.69 0.558 0.349 GenWildSplat 14.58 0.507 0.483 14.81 0.506 0.487 Table 4: Quantitative results on clean scenes. Figure 5: Qualitative results on clean scenes. For our method, proposal regions (blue) are rejected during verification, leaving the reconstruction unchanged, whereas GenWildSplat’s semantic masking (red) can remove static objects. Qualitative Results. Figure 4 shows that transient people and objects can produce blurred, duplicated, or floating artifacts in the original reconstructions. Our method suppresses these artifacts across multiple backbones while largely preserving nearby static scene content. All methods are compared using the same enlarged image regions. Ablation Studies. Table 2 reports averages over the ten scenes from both datasets and three backbones in 4-view setting. Removing verification causes the largest degradation, confirming proposal regions should not be removed directly. Replacing DINOv3 similarity with a simple RGB difference retains most of the improvement over the vanilla reconstruction; DINOv3 and dual proposal ranges add further gains. Comparison with In-the-Wild Feed-Forward Model. Table 3 compares AnySplat, AnySplat with our method, and GenWildSplat on the two distractor benchmarks using four and six input views. Here all methods operate from the cameras estimated by AnySplat, which takes unposed images as input; our proposal generation and verification use the same cameras. Our method improves the frozen AnySplat prediction across both datasets and input settings, and achieves higher average performance than the released GenWildSplat model under this protocol. GenWildSplat builds on AnySplat and uses semantic masks for predefined transient categories, making it the closest trained counterpart to ours. Method Views Inference Proposal Verification ReSplat 4 0.22 0.23 0.44 8 0.38 0.35 1.10 16 0.78 0.64 3.22 DepthSplat 4 0.09 0.19 0.58 8 0.14 0.42 2.82 16 0.28 1.02 16.95 YoNoSplat 4 0.22 0.12 0.52 8 0.25 0.23 2.83 16 0.39 0.59 14.77 Table 5: Mean processing time in seconds, averaged over the RobustNeRF and NeRF On-the-go scenes. Preservation on Clean Scenes. Table 4 and Figure 5 evaluate the four clean RobustNeRF scenes under posed and estimated-pose settings. Across four reconstruction models, our method leaves the original reconstructions nearly unchanged: proposal regions on clean inputs are often rejected during verification, as shown in Figure 5. GenWildSplat instead removes objects by semantic category, so static objects in its predefined transient categories can also be masked. Overhead Analysis. Table 5 reports the reconstruction model execution time and the additional time of our filtering procedure for each input count. Filtering runs once per scene, after the Gaussian representation has been predicted, and produces a single filtered representation outG_out; rendering a novel view requires no input reselection, re-reconstruction, or other per-view processing. Verification dominates the filtering time, as it renders each candidate from multiple views, and grows with the number of views and candidates. Conclusion We showed that the per-view prediction structure of feed-forward 3DGS models enables training-free distractor filtering: excluding the Gaussian subset associated with each input exposes content that is inconsistent with the remaining observations, and verification by rendering retains only candidates whose removal improves reconstruction. Across multiple reconstruction models and distractor benchmarks, the resulting procedure consistently reduces distractor artifacts while preserving quality on clean scenes, without retraining or scene-specific optimization. Distractor robustness can thus be obtained from the reconstruction itself, with the model left untouched. Verification cost grows with the number of input views and candidates, since each candidate is checked by rendering. The individual checks are independent, leaving room for batched or parallel evaluation in larger collections. Our method also assumes an association between predicted Gaussians and input views, which the reconstruction models considered here natively provide; extending the procedure to architectures without this structure is an open direction. Future work may further extend the framework with object-level reasoning and geometry-aware generative completion, enabling more structured filtering and reconstruction of newly exposed regions. References Bao et al. (2024) Y. Bao, J. Liao, J. Huo, and Y. Gao Distractor-free generalizable 3d gaussian splatting. arXiv preprint arXiv:2411.17605. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction, Reconstruction Models and Comparisons.. Barron et al. (2022) J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman Mip-nerf 360: unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5470–5479. Cited by: Feed-Forward 3D Gaussian Splatting. Charatan et al. (2024) D. Charatan, S. L. Li, A. Tagliasacchi, and V. Sitzmann Pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d reconstruction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 19457–19467. Cited by: Introduction, Feed-Forward 3D Gaussian Splatting, Feed-Forward 3DGS.. Chen et al. (2024) Y. Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T. Cham, and J. Cai Mvsplat: efficient 3d gaussian splatting from sparse multi-view images. In European conference on computer vision, p. 370–386. Cited by: Introduction, Feed-Forward 3D Gaussian Splatting, Ignoring Distractors in 3D Reconstruction, Feed-Forward 3DGS.. Gupta et al. (2026) V. Gupta, C. Lin, S. Wang, A. Bhattad, and J. Huang Generalizable sparse-view 3d reconstruction from unconstrained images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 33217–33226. Cited by: Introduction, Introduction, Ignoring Distractors in 3D Reconstruction, Reconstruction Models and Comparisons.. Jiang et al. (2025) L. Jiang, Y. Mao, L. Xu, T. Lu, K. Ren, Y. Jin, X. Xu, M. Yu, J. Pang, F. Zhao, et al. Anysplat: feed-forward 3d gaussian splatting from unconstrained views. ACM Transactions on Graphics (TOG) 44 (6), p. 1–16. Cited by: Introduction, Introduction, Introduction, Feed-Forward 3D Gaussian Splatting, Ignoring Distractors in 3D Reconstruction, Reconstruction Models and Comparisons.. Kerbl et al. (2023) B. Kerbl, G. Kopanas, T. Leimkühler, G. Drettakis, et al. 3d gaussian splatting for real-time radiance field rendering.. ACM Trans. Graph. 42 (4), p. 139–1. Cited by: Introduction, Feed-Forward 3D Gaussian Splatting, 3D Gaussian Splatting.. Kheradmand et al. (2024) S. Kheradmand, D. Rebain, G. Sharma, W. Sun, Y. Tseng, H. Isack, A. Kar, A. Tagliasacchi, and K. M. Yi 3d gaussian splatting as markov chain monte carlo. Advances in Neural Information Processing Systems 37, p. 80965–80986. Cited by: Feed-Forward 3D Gaussian Splatting. Kulhanek et al. (2024) J. Kulhanek, S. Peng, Z. Kukelova, M. Pollefeys, and T. Sattler Wildgaussians: 3d gaussian splatting in the wild. arXiv preprint arXiv:2407.08447. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Li et al. (2025) C. Li, Z. Shi, Y. Lu, W. He, and X. Xu Robust neural rendering in the wild with asymmetric dual 3d gaussian splatting. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, External Links: Link Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Li et al. (2026) Y. Li, C. Lv, Z. Tang, H. Yang, and D. Huang Tokensplat: token-aligned 3d gaussian splatting for feed-forward pose-free reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 40886–40895. Cited by: Feed-Forward 3D Gaussian Splatting. Lin et al. (2025) J. Lin, J. Gu, L. Fan, B. Wu, Y. Lou, R. Chen, L. Liu, and J. Ye HybridGS: decoupling transients and statics with 2d and 3d gaussian splatting. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 788–797. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Liu et al. (2024) Y. Liu, C. Luo, L. Fan, N. Wang, J. Peng, and Z. Zhang Citygaussian: real-time high-quality large-scale scene rendering with gaussians. In European Conference on Computer Vision, p. 265–282. Cited by: Feed-Forward 3D Gaussian Splatting. Lu et al. (2024) T. Lu, M. Yu, L. Xu, Y. Xiangli, L. Wang, D. Lin, and B. Dai Scaffold-gs: structured 3d gaussians for view-adaptive rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 20654–20664. Cited by: Feed-Forward 3D Gaussian Splatting. Martin-Brualla et al. (2021) R. Martin-Brualla, N. Radwan, M. S. Sajjadi, J. T. Barron, A. Dosovitskiy, and D. Duckworth Nerf in the wild: neural radiance fields for unconstrained photo collections. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 7210–7219. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Matsuki et al. (2024) H. Matsuki, R. Murai, P. H. Kelly, and A. J. Davison Gaussian splatting slam. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 18039–18048. Cited by: Feed-Forward 3D Gaussian Splatting. Mildenhall et al. (2021) B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng Nerf: representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 (1), p. 99–106. Cited by: Introduction, Feed-Forward 3D Gaussian Splatting. Moreau et al. (2026) A. Moreau, R. Shaw, M. Nazarczuk, J. Shin, T. Tanay, Z. Zhang, S. Xu, and E. Pérez-Pellitero Off the grid: detection of primitives for feed-forward 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11756–11766. Cited by: Feed-Forward 3D Gaussian Splatting. Pan et al. (2026) T. Pan, X. Yang, S. Wang, and X. Wang Visual geometry transformer in the wild: distractor-free 3d reconstruction. arXiv preprint arXiv:2606.22787. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Park et al. (2026) W. Park, M. Nam, S. Kim, S. Jo, and S. Lee ForestSplats: deformable transient field for gaussian splatting in the wild. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, p. 6978–6987. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Qin et al. (2024) M. Qin, W. Li, J. Zhou, H. Wang, and H. Pfister Langsplat: 3d language gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 20051–20060. Cited by: Feed-Forward 3D Gaussian Splatting. Ren et al. (2024) W. Ren, Z. Zhu, B. Sun, J. Chen, M. Pollefeys, and S. Peng Nerf on-the-go: exploiting uncertainty for distractor-free nerfs in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8931–8940. Cited by: Introduction, Introduction, Ignoring Distractors in 3D Reconstruction, Experimental Setup.. Sabour et al. (2025) S. Sabour, L. Goli, G. Kopanas, M. Matthews, D. Lagun, L. Guibas, A. Jacobson, D. Fleet, and A. Tagliasacchi Spotlesssplats: ignoring distractors in 3d gaussian splatting. ACM Transactions on Graphics 44 (2), p. 1–11. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Sabour et al. (2023) S. Sabour, S. Vora, D. Duckworth, I. Krasin, D. J. Fleet, and A. Tagliasacchi Robustnerf: ignoring distractors with robust losses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 20626–20636. Cited by: Introduction, Introduction, Ignoring Distractors in 3D Reconstruction, Experimental Setup.. Schonberger and Frahm (2016) J. L. Schonberger and J. Frahm Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 4104–4113. Cited by: Experimental Setup.. Seo et al. (2026) K. Seo, M. Lee, T. Kim, B. Lee, J. An, and J. Heo PDF-gs: progressive distractor filtering for robust 3d gaussian splatting. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, p. 468–477. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Siméoni et al. (2025) O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H. Jégou, P. Labatut, and P. Bojanowski DINOv3. External Links: 2508.10104, Link Cited by: Distractor Proposal. Wang et al. (2025a) W. Wang, Y. Chen, Z. Zhang, H. Liu, H. Wang, Z. Feng, W. Qin, F. Chen, Z. Zhu, D. Y. Chen, et al. Volsplat: rethinking feed-forward 3d gaussian splatting with voxel-aligned prediction. arXiv preprint arXiv:2509.19297. Cited by: Feed-Forward 3D Gaussian Splatting. Wang et al. (2026) X. Wang, Z. Wang, S. Xie, C. Pan, and Y. Chen DualSplat: robust 3d gaussian splatting via pseudo-mask bootstrapping from reconstruction failures. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 4912–4921. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Wang et al. (2025b) Y. Wang, M. Klasson, M. Turkulainen, S. Wang, J. Kannala, and A. Solin DeSplat: decomposed gaussian splatting for distractor-free rendering. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 722–732. Cited by: Introduction, Ignoring Distractors in 3D Reconstruction. Wang et al. (2004) Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing 13 (4), p. 600–612. Cited by: Experimental Setup.. Wu et al. (2024) G. Wu, T. Yi, J. Fang, L. Xie, X. Zhang, W. Wei, W. Liu, Q. Tian, and X. Wang 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 20310–20320. Cited by: Feed-Forward 3D Gaussian Splatting. Xie et al. (2024) T. Xie, Z. Zong, Y. Qiu, X. Li, Y. Feng, Y. Yang, and C. Jiang Physgaussian: physics-integrated 3d gaussians for generative dynamics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 4389–4398. Cited by: Feed-Forward 3D Gaussian Splatting. Xu et al. (2025a) H. Xu, D. Barath, A. Geiger, and M. Pollefeys Resplat: learning recurrent gaussian splats. arXiv preprint arXiv:2510.08575. Cited by: Introduction, Introduction, Feed-Forward 3D Gaussian Splatting, Reconstruction Models and Comparisons.. Xu et al. (2025b) H. Xu, S. Peng, F. Wang, H. Blum, D. Barath, A. Geiger, and M. Pollefeys Depthsplat: connecting gaussian splatting and depth. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 16453–16463. Cited by: Introduction, Introduction, Introduction, Feed-Forward 3D Gaussian Splatting, Feed-Forward 3DGS., Reconstruction Models and Comparisons.. Ye et al. (2025a) B. Ye, B. Chen, H. Xu, D. Barath, and M. Pollefeys YoNoSplat: you only need one model for feedforward 3d gaussian splatting. arXiv preprint arXiv:2511.07321. Cited by: Introduction, Introduction, Feed-Forward 3D Gaussian Splatting, Reconstruction Models and Comparisons.. Ye et al. (2025b) B. Ye, S. Liu, H. Xu, X. Li, M. Pollefeys, M. Yang, and S. Peng No pose, no problem: surprisingly simple 3d gaussian splats from sparse unposed images. In International Conference on Learning Representations, Vol. 2025, p. 54009–54033. Cited by: Introduction, Feed-Forward 3D Gaussian Splatting. Zhang et al. (2018) R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 586–595. Cited by: Experimental Setup..