Paper deep dive
DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction
Hongbo Duan, Pengting Luo, Chengzhi Zhao, Yuanhao Chiang, Fangming Liu, Xueqian Wang
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic environments. The framework incrementally reconstructs a 3D Gaussian scene representation while suppressing motion-corrupted observations through online uncertainty prediction and uncertainty-weighted Gaussian optimization. A key component of DynActiveGS is the explicit decomposition of uncertainty into structural uncertainty and motion-induced uncertainty, which enables the system to distinguish under-reconstructed static regions from dynamically unreliable areas. Based on these uncertainty fields, DynActiveGS performs dynamic-aware viewpoint selection and dynamic-constrained path planning to favor informative yet stable observations during exploration. The resulting system forms a unified closed-loop pipeline for robust active reconstruction in dynamic scenes. Extensive experiments on challenging dynamic benchmarks demonstrate consistent improvements over existing active reconstruction baselines in reconstruction accuracy, completeness, rendering quality, and exploration efficiency.
Tags
Links
- Source: https://arxiv.org/abs/2608.01178v1
- Canonical: https://arxiv.org/abs/2608.01178v1
Trouble viewing inline? Open PDF directly →
Full Text
52,371 characters extracted from source content.
Expand or collapse full text
by DynActiveGS: Active Gaussian Splatting for Dynamic Scene Reconstruction Hongbo Duan Tsinghua University, Shenzhen International Graduate SchoolShenzhenGuangdongChina dhb24@mails.tsinghua.edu.cn 0009-0002-8463-2170 , Pengting Luo Huawei Technologies LtdShenzhenGuangdongChina pengtingluo@gmail.com , Chengzhi Zhao Harbin Institute of Technology, ShenzhenShenzhenGuangdongChina 2021210731@stu.hit.edu.cn , Yuanhao Chiang Tsinghua University, Shenzhen International Graduate SchoolShenzhenGuangdongChina jiang-yh24@mails.tsinghua.edu.cn , Fangming Liu Peng Cheng LaboratoryShenzhenGuangdongChina fangminghk@gmail.com and Xueqian Wang Tsinghua University, Shenzhen International Graduate SchoolShenzhenGuangdongChina wang.xq@sz.tsinghua.edu.cn (2026) Abstract. We present DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting (3DGS) for autonomous exploration in dynamic environments. The framework incrementally reconstructs a 3D Gaussian scene representation while suppressing motion-corrupted observations through online uncertainty prediction and uncertainty-weighted Gaussian optimization. A key component of DynActiveGS is the explicit decomposition of uncertainty into structural uncertainty and motion-induced uncertainty, which enables the system to distinguish under-reconstructed static regions from dynamically unreliable areas. Based on these uncertainty fields, DynActiveGS performs dynamic-aware viewpoint selection and dynamic-constrained path planning to favor informative yet stable observations during exploration. The resulting system forms a unified closed-loop pipeline for robust active reconstruction in dynamic scenes. Extensive experiments on challenging dynamic benchmarks demonstrate consistent improvements over existing active reconstruction baselines in reconstruction accuracy, completeness, rendering quality, and exploration efficiency. Active Reconstruction, 3D Gaussian Splatting, Dynamic Scenes †journalyear: 2026†copyright: c†conference: Proceedings of the 35th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, Brazil†booktitle: Proceedings of the 35th ACM International Conference on Multimedia (M ’26), November 10–14, 2026, Rio de Janeiro, Brazil†doi: 10.1145/3767308.3835011†isbn: 979-8-4007-2213-4/2026/11†ccs: Computing methodologies Reconstruction†ccs: Computing methodologies Vision for robotics†ccs: Computing methodologies Robotic planning Figure 1. DynActiveGS in action. The robot actively explores a dynamic indoor scene and progressively improves reconstruction completeness. From left to right, we show RGB-D inputs with dynamic disturbances, intermediate reconstruction states, and final renderings after 1000 steps. DynActiveGS preserves stable geometry and rendering quality despite dynamic interference. 1. Introduction The ability to reconstruct 3D scenes from visual observations is fundamental to computer vision and robotics (Newcombe et al., 2011; Whelan et al., 2015; Dai et al., 2017). Beyond passive perception, intelligent agents should actively determine where to move and what to observe (Isler et al., 2016) to efficiently acquire informative data and build high-fidelity scene representations (Kirsch et al., 2019). This paradigm, known as active 3D reconstruction (Aloimonos et al., 1988), integrates perception, uncertainty estimation, and motion planning (Chen et al., 2011). Recent advances in 3D Gaussian Splatting (3DGS) (Kerbl et al., 2023) have enabled explicit, differentiable, and efficient scene representations with high-quality rendering and incremental optimization. These advantages have inspired active reconstruction methods that leverage 3DGS with next-best-view (NBV) planning (Peralta et al., 2020; Pito, 2002) and uncertainty-driven exploration (Lee et al., 2022; Pan et al., 2022; Ran et al., 2023). However, existing active reconstruction approaches (Feng et al., 2024; Jin et al., 2025; Li et al., 2025c; Polyzos et al., 2025; Chen et al., 2025; Li et al., 2025b; Jin et al., 2024; Xu et al., 2025; Xie et al., 2025) are primarily developed under the assumption of static environments. Dynamic objects, including pedestrians, vehicles, and articulated agents, are common in real-world scenarios and introduce fundamental challenges to active reconstruction (Huang et al., 2024; Yu et al., 2024). Motion-corrupted observations can degrade Gaussian optimization (Palazzolo et al., 2019; Xu et al., 2024), distort uncertainty estimation (Ren et al., 2024; Kulhanek et al., 2024), and mislead viewpoint planning (Kuang et al., 2024; Jiang et al., 2024b). Meanwhile, existing dynamic 3DGS and Gaussian-based SLAM methods mainly focus on passive mapping and dynamic object filtering (Xu et al., 2024; Sandström et al., 2025; Zheng et al., 2025), leaving active perception and long-horizon exploration in dynamic environments largely unexplored. Therefore, robust active reconstruction under dynamic scenes remains an open challenge. In this paper, we propose DynActiveGS, a dynamic-aware active reconstruction framework for 3D Gaussian Splatting in dynamic environments. As shown in Fig. 1, DynActiveGS first performs uncertainty-aware Gaussian reconstruction by predicting pixel-wise uncertainty and suppressing motion-corrupted observations during online optimization. Based on the reconstructed Gaussian map, we further disentangle structural uncertainty from motion-induced uncertainty to construct map-level uncertainty fields for exploration. These uncertainty fields guide dynamic-aware viewpoint selection through local-global scoring on a Voronoi graph, while a motion-constrained planner generates efficient and stable trajectories. Together, these components establish a unified perception–planning–reconstruction pipeline for closed-loop active reconstruction in dynamic scenes. We evaluate DynActiveGS on diverse dynamic benchmarks. Results demonstrate consistent improvements over existing active reconstruction baselines in reconstruction accuracy, completeness, rendering quality, and exploration efficiency. Our work extends active 3D reconstruction beyond static environments toward robust embodied perception in dynamic real-world scenarios. Our contributions are summarized as follows: • We propose DynActiveGS, a dynamic-aware active reconstruction framework for 3D Gaussian Splatting in dynamic scenes. • We introduce an uncertainty-aware Gaussian reconstruction strategy that predicts online uncertainty, suppresses motion-corrupted observations, and disentangles structural and motion-induced uncertainty. • We develop a closed-loop framework integrating dynamic-aware viewpoint selection and motion-constrained path planning, achieving robust active exploration in challenging dynamic reconstruction scenarios. 2. Related Work 2.1. 3D Scene Representation Scene representation plays a critical role in balancing reconstruction fidelity, efficiency, and online perception capability. Classical 3D reconstruction relies on structure-from-motion (SfM) pipelines (Liu et al., 2022), which estimate camera poses and scene geometry through feature matching (Schonberger and Frahm, 2016; Sandström et al., 2025), geometric verification, and bundle adjustment (Chng et al., 2022). Although effective, these methods are computationally expensive and mainly assume static environments. Neural scene representations, such as NeRFs (Mildenhall et al., 2021) and their variants (Barron et al., 2022; Li et al., 2022; Chen et al., 2022), achieve high-quality view synthesis but remain inefficient for real-time deployment. Recently, 3D Gaussian Splatting (3DGS) (Kerbl et al., 2023) has emerged as an efficient explicit representation with real-time rendering and incremental optimization, enabling promising applications in online reconstruction and robotic perception. Recent Gaussian-based mapping and SLAM systems further demonstrate its potential, yet most remain limited to static scenes (Xu et al., 2025; Keetha et al., 2024; Li et al., 2024). 2.2. Active 3D Reconstruction Active reconstruction focuses on selecting informative viewpoints to improve scene acquisition efficiency (Chen et al., 2024; Li et al., 2025a). Early approaches exploit geometric uncertainty, occupancy reasoning, or information gain for next-best-view (NBV) planning (Zhang et al., 2025; Pan et al., 2022; Jiang et al., 2024b). Recent methods further integrate neural representations and uncertainty estimation to guide exploration (Lee et al., 2022; Feng et al., 2024; Kuang et al., 2024; Yan et al., 2023). With the development of NeRF and 3DGS representations, several works investigate active reconstruction with neural and Gaussian models (Jin et al., 2025; Li et al., 2025c; Jin et al., 2024; Xu et al., 2025). However, existing frameworks are predominantly designed for static environments, where dynamic objects are ignored or treated as noise, leading to unreliable optimization and viewpoint selection. 2.3. Dynamic Scene Reconstruction Dynamic scenes pose significant challenges for reliable 3D reconstruction. Traditional SLAM methods handle moving objects through motion segmentation and robust optimization (Bescos et al., 2018; Jiang et al., 2024a), while recent neural and Gaussian-based approaches model or filter dynamic content for improved robustness (Xu et al., 2024; Sandström et al., 2025; Zheng et al., 2025). Although uncertainty-aware Gaussian mapping improves reconstruction under dynamic conditions, existing approaches mainly focus on passive mapping and tracking rather than active exploration (Keetha et al., 2024; Zhu et al., 2025). Therefore, they cannot determine how an embodied agent should actively acquire informative observations in dynamic environments. Our work bridges this gap by integrating dynamic scene modeling with active viewpoint planning in a unified Gaussian splatting framework. 3. Method Figure 2. Overview of DynActiveGS. Given RGB-D observations in dynamic environments, DynActiveGS first performs dynamic-aware Gaussian reconstruction by predicting a per-pixel uncertainty map and updating the 3D Gaussian representation through uncertainty-weighted optimization. It then constructs structural uncertainty UsU_s and motion uncertainty UmU_m to guide dynamic-aware viewpoint selection via local-global scoring on a Voronoi graph. Finally, a dynamic-constrained path planner computes stable trajectories using motion-aware edge costs, enabling closed-loop active reconstruction under dynamic disturbances. As shown in Fig. 2, DynActiveGS operates as a closed-loop active reconstruction system in dynamic environments. At each timestep t, the embodied agent acquires an RGB-D observation (It,Dt)(I_t,D_t) and an estimated camera pose t∈SE(3)T_t∈ SE(3). The system updates the 3D Gaussian scene representation through uncertainty-aware reconstruction, constructs structural uncertainty UsU_s and motion uncertainty UmU_m, selects the next informative viewpoint, and executes a dynamically stable path toward it. The loop continues until the exploration budget of S steps is exhausted. 3.1. Preliminary: 3D Gaussian Splatting We represent the scene as a set of anisotropic 3D Gaussian primitives =gii=1NG=\g_i\_i=1^N, where each Gaussian gig_i is parameterized by its mean i∈ℝ3 μ_i ^3, color i∈ℝ3c_i ^3, opacity ηi∈[0,1] _i∈[0,1], scale i∈ℝ3×3S_i ^3× 3, and rotation i∈SO(3)R_i (3). Its covariance is given by (1) i=iii⊤i⊤. _i=R_iS_iS_i R_i . The rendered color is obtained via alpha compositing: (2) ^()=∑i=1Nαii∏j=1i−1(1−αj), C(p)= _i=1^N _ic_i _j=1^i-1(1- _j), where (3) αi=ηiexp(−12(−i)⊤i−1(−i)). _i= _i \! (- 12(p- μ_i) _i^-1(p- μ_i) ). This formulation is fully differentiable and supports efficient online optimization. 3.2. Dynamic-aware Gaussian Reconstruction Dynamic environments violate the static-scene assumption: moving objects and transient occlusions introduce unreliable observations, which can corrupt Gaussian optimization and further mislead active planning. To address this issue, we develop an uncertainty-aware reconstruction strategy for dynamic active reconstruction. Specifically, we predict a per-pixel uncertainty map online, use it to suppress dynamically corrupted observations during Gaussian map optimization, and further lift it into map-level structural and motion uncertainty fields for downstream planning. 3.2.1. Frame-level uncertainty prediction. For each input frame ItI_t, we extract dense visual features using a pre-trained DINOv3 feature extractor (Oquab et al., 2023; Siméoni et al., 2025): (4) Ft=ℱ(It),F_t=F(I_t), where ℱF denotes the DINOv3 encoder and FtF_t is the resulting feature map. These features are fed into a lightweight uncertainty prediction network P to estimate a per-pixel uncertainty map (5) βt=(Ft), _t=P(F_t), where βt∈ℝH×W _t ^H× W is bilinearly upsampled to the input resolution. Here, βt() _t(u) measures the reliability of pixel u, and pixels affected by dynamic motion or transient occlusion tend to have higher uncertainty values. The uncertainty predictor is supervised by appearance and depth consistency; the full objective is provided in the supplementary material. 3.2.2. Uncertainty-weighted Gaussian mapping. Let tG_t denote the current Gaussian map. Given the camera pose tT_t and intrinsics K, we render the predicted RGB image I^t I_t and depth map D^t D_t from tG_t. The Gaussian map is optimized using an uncertainty-weighted rendering loss (6) ℒrender=λ1ℒcolor+λ2ℒdepthβt2+λ3ℒiso,L_render= _1L_color+ _2L_depth _t^2+ _3L_iso, where ℒcolorL_color combines pixel-wise ℓ1 _1 and SSIM losses, and ℒisoL_iso denotes isotropic Gaussian regularization. Here, the uncertainty map βt _t acts as a pixel-wise weighting factor, so that unreliable observations contribute less to map optimization (Ren et al., 2024; Zheng et al., 2025; Li et al., 2026). This encourages the optimizer to focus on stable static structures while suppressing dynamic distractors during Gaussian updates. 3.2.3. Map-level structural and motion uncertainty. To support dynamic-aware planning, we maintain two map-level uncertainty fields over a discrete 3D query set Ω : (7) Us:Ω→ℝ≥0,Um:Ω→ℝ≥0,U_s: _≥ 0, U_m: _≥ 0, where Us()U_s(x) measures structural uncertainty and Um()U_m(x) measures motion-induced unreliability at location x. We estimate structural uncertainty from both insufficient observation and reconstruction mismatch. For each ∈Ωx∈ , let nt()n_t(x) denote the effective observation count and let e¯t() e_t(x) denote an aggregated rendering residual. We define (8) Us()=λc1nt()+ϵ+λrclip(e¯t(),0,emax),U_s(x)= _c 1 n_t(x)+ε+ _r\,clip\! ( e_t(x),0,e_ ), where the first term highlights under-observed regions and the second term emphasizes locations that are poorly explained by the current Gaussian map. Motion uncertainty is obtained by lifting the frame-level uncertainty map βt _t into 3D and accumulating it over time. For each query location ∈Ωx∈ , let t=Π(t,)u_t= (T_t,x) denote its projection into frame t, and let t()1_t(x) indicate whether it is observable in that frame. We update motion uncertainty using (9) Um()←(1−α)Um()+α 1t()βt(t),U_m(x)←(1-α)\,U_m(x)+α\,1_t(x)\, _t(u_t), where α∈(0,1)α∈(0,1) controls the update rate. As a result, UsU_s identifies where more static evidence is needed, while UmU_m identifies where future observations are likely to be unreliable due to dynamic disturbances. 3.3. Dynamic-aware Viewpoint Selection Given the map-level structural uncertainty UsU_s and motion uncertainty UmU_m, we select viewpoints that are both informative and dynamically reliable. To improve long-horizon exploration efficiency in multi-room environments, we adopt a hierarchical planning strategy on a Voronoi graph (Wu et al., 2024) G=(,ℰ)G=(V,E), where each node n∈n denotes a reachable region and each edge encodes traversal connectivity. 3.3.1. Dynamic-aware subregion partition. Rather than planning over all nodes uniformly, we partition the graph into subregions using agglomerative clustering. To account for dynamic disturbances, the pairwise distance between nodes nin_i and njn_j is defined as (10) ddyn(ni,nj)=λedE(ni,nj)+λpdP(ni,nj)+λmdM(ni,nj),d_dyn(n_i,n_j)= _ed_E(n_i,n_j)+ _pd_P(n_i,n_j)+ _md_M(n_i,n_j), where dEd_E is the Euclidean distance, dPd_P is the graph travel distance, and dMd_M is the accumulated motion risk along the connecting path. This produces a set of subregions Rk\R_k\, including the current local subregion RlR_l containing the agent. 3.3.2. Local-global viewpoint selection. We prioritize local exploration before switching to global exploration. For each candidate node n∈Rln∈ R_l, we define a local score (11) Slocal(n)=ℐs(n;Us)−λrℛm(n;Um)−λc(n),S_local(n)=I_s(n;U_s)- _rR_m(n;U_m)- _cC(n), where (n)C(n) is the travel cost from the current pose to node n. The structural utility term is (12) ℐs(n;Us)=∑∈ΩVis(n,)Us(),I_s(n;U_s)= _x∈ Vis(n,x)\,U_s(x), and the motion risk term is (13) ℛm(n;Um)=∑∈ΩVis(n,)Um(),R_m(n;U_m)= _x∈ Vis(n,x)\,U_m(x), where Vis(n,)∈[0,1]Vis(n,x)∈[0,1] is a soft visibility score under the current Gaussian map tG_t. The best local node is selected iteratively as long as its score remains above a threshold. When the local region is either sufficiently explored or becomes unstable due to high motion uncertainty, the planner switches to global selection. For each candidate subregion Rk≠RlR_k≠ R_l, we define (14) Sglobal(Rk)=∑n∈Rk(ℐs(n;Us)−λrℛm(n;Um))−λdd(Rl,Rk)−λvPvisit(Rk). splitS_global(R_k)&= _n∈ R_k (I_s(n;U_s)- _rR_m(n;U_m) )\\ & - _d\,d(R_l,R_k)- _v\,P_visit(R_k). split where d(Rl,Rk)d(R_l,R_k) is the inter-region travel cost and Pvisit(Rk)P_visit(R_k) penalizes repeatedly visited regions. The region with the highest score is selected as the next exploration target, and a local candidate view action at⋆=(t⋆,t⋆)a_t =(p_t , θ_t ) is chosen around the selected node. 3.4. Dynamic-constrained Path Planning After selecting the next target viewpoint, we compute a feasible path that jointly accounts for geometry and dynamic risk. Unlike static planning, the objective here is not simply to minimize travel distance, but also to avoid routes likely to produce unstable or motion-corrupted observations. For each graph edge e∈ℰe , we define a dynamic-constrained traversal cost (15) c(e)=ℓ(e)+λmU¯m(e),c(e)= (e)+ _m U_m(e), where ℓ(e) (e) is the geometric edge length and U¯m(e) U_m(e) is the average motion uncertainty along the edge. The cost of a path π is then (16) path(π)=∑e∈πc(e).C_path(π)= _e∈πc(e). This encourages the robot to favor dynamically stable routes instead of purely shortest paths. To improve temporal robustness, motion uncertainty is accumulated over time using an exponential moving average: (17) Um(t)()=(1−α)Um(t−1)()+αβt(),U_m^(t)(x)=(1-α)\,U_m^(t-1)(x)+α\, _t(x), where βt _t is the frame-level uncertainty map predicted in Sec. 3.2. As a result, transient disturbances do not dominate planning, while persistently dynamic regions are consistently down-weighted. By combining dynamic-aware viewpoint selection with motion-constrained path planning, DynActiveGS forms a closed-loop active reconstruction system. At each exploration step, the robot updates the Gaussian map and uncertainty fields from the latest RGB-D observation, selects the next informative target, and executes a dynamically stable path toward it. The loop continues until the exploration step budget S is exhausted, as summarized in the supplementary material Algorithm 1. 4. Experiments Figure 3. Qualitative comparison of 3D reconstruction results on representative scenes from Social-MP3D and Dynamic Gibson. Blue bounding boxes indicate reference areas for easier comparison, while orange ones highlight low-quality reconstruction. 4.1. Experimental Setup Figure 4. Novel view synthesis results on representative scenes from Social-MP3D and Dynamic Gibson. The evaluated viewpoints are held-out views that do not appear in the training trajectories of any compared method. Blue bounding boxes indicate reference regions for easier comparison, while orange bounding boxes highlight low-quality renderings. Simulator and Dynamic Scenes. All experiments are conducted in the Habitat simulator (Savva et al., 2019). To evaluate dynamic active reconstruction, we use indoor scenes with runtime human dynamics in two settings. For Matterport3D (MP3D) (Chang et al., 2018), we adopt the public Social-MP3D benchmark (Gong et al., 2025), which augments MP3D scenes with dynamic humans. For Gibson (Xia et al., 2018), since no public benchmark provides runtime humans for active reconstruction, we construct a reproducible Dynamic Gibson Protocol in Habitat-Sim, following the human simulation paradigm of HabiCrowd (Vuong et al., 2024). In both settings, virtual pedestrians are instantiated at runtime with randomized start–goal waypoints and simulated using ORCA (Van Den Berg et al., 2011) and UPL++ (Karamouzas et al., 2014; Vuong et al., 2024) crowd dynamics. The dynamic humans are not part of the static scene geometry used for reconstruction evaluation. We use fixed scene splits, shared human configurations, and standardized random seeds across all compared methods. Baselines. We compare DynActiveGS against state-of-the-art active reconstruction methods, including 3DGS-based approaches (ActiveGAMER (Chen et al., 2025), ActiveGS (Jin et al., 2025), ActiveSplat (Li et al., 2025c)) and neural active mapping methods (NARUTO (Feng et al., 2024)). Although some baselines employ free-flying 6DoF cameras, we evaluate them using their official implementations under identical observation budgets. We also include dynamic passive mapping baselines without next-best-view planning. Geometric Metrics. We evaluate geometric reconstruction using Accuracy (Acc), Completion (Com), and Completion Ratio (C.R.) with a 5 cm threshold. Metrics are computed by uniformly sampling points from static ground-truth meshes and comparing them with point clouds extracted from reconstructed Gaussian maps. Since dynamic humans are instantiated at runtime and are excluded from the ground-truth meshes, these metrics measure robustness of static scene reconstruction under dynamic disturbances. Rendering Metrics. For novel view rendering evaluation, we follow the protocol of the baseline ActiveGAMER (Chen et al., 2025) and use same predefined novel-view trajectories for testing. Specifically, we compare the rendered RGB images against static ground-truth renderings at identical poses, and report PSNR, SSIM, and LPIPS. For depth rendering performance, we additionally use Depth L1 distance as the evaluation metric. Implementation Details. We consider a realistic ground mobile robot with a fixed camera height. At each step, the agent executes one discrete navigation action and acquires one posed RGB-D observation along the planned viewpoint. The camera field of view is set to 60∘60 vertically and 90∘90 horizontally, and the system processes image sequences online with on-policy planning and incremental reconstruction. Unlike prior works that assume free-flying 6DoF cameras, our embodiment better reflects practical ground robot constraints. All methods are evaluated under an equal observation budget: each method is allowed 1000 exploration steps, and each step produces exactly one RGB-D frame, ensuring identical input image counts across all approaches. 4.2. Comparison Results Results on Social-MP3D. As shown in Tab. 1, the advantage of DynActiveGS becomes even more evident on the larger and more heavily occluded Social-MP3D scenes. Under the same observation budget, static-scene baselines suffer substantial performance degradation in complex layouts with frequent human occlusions, whereas our method maintains stable reconstruction quality. Table 1. Quantitative comparison on the Social-MP3D dataset for 3D reconstruction and novel view synthesis. Method Metric Gdvg gZ6f HxpK pLe4 YmJk Avg. NARUTO (Feng et al., 2024) Acc (cm)↓ 6.62 5.65 10.87 6.35 13.02 8.50 Com. (cm)↓ 7.91 4.26 9.41 5.98 9.73 7.46 C.R. (%)↑ 82.46 85.11 83.27 79.42 76.76 81.40 PSNR↑ 18.63 18.51 15.46 20.24 16.45 17.86 SSIM↑ 0.624 0.548 0.432 0.697 0.386 0.537 LPIPS↓ 0.431 0.453 0.561 0.482 0.496 0.485 ActiveSplat (Li et al., 2025c) Acc (cm)↓ 4.64 3.05 5.98 6.74 9.27 5.94 Com. (cm)↓ 4.72 3.48 6.75 4.22 5.91 5.02 C.R. (%)↑ 88.27 87.16 85.54 86.64 85.22 86.57 PSNR↑ 19.65 19.12 21.85 21.57 20.47 20.53 SSIM↑ 0.662 0.646 0.788 0.763 0.655 0.703 LPIPS↓ 0.516 0.552 0.387 0.423 0.458 0.467 ActiveGAMER (Chen et al., 2025) Acc (cm)↓ 3.71 3.37 3.78 5.43 5.42 4.34 Com. (cm)↓ 5.76 3.91 4.11 6.08 6.64 5.30 C.R. (%)↑ 92.18 92.74 91.46 88.12 86.35 90.17 PSNR↑ 19.76 20.23 22.83 23.07 21.52 21.48 SSIM↑ 0.670 0.693 0.816 0.826 0.721 0.745 LPIPS↓ 0.445 0.420 0.372 0.425 0.386 0.410 Ours Acc (cm)↓ 2.54 2.11 3.06 2.62 3.97 2.86 Com. (cm)↓ 2.87 2.46 3.82 3.02 4.28 3.29 C.R. (%)↑ 94.41 93.02 95.58 94.77 92.63 94.08 PSNR↑ 21.36 20.74 26.23 24.98 21.87 23.04 SSIM↑ 0.722 0.703 0.894 0.862 0.755 0.787 LPIPS↓ 0.316 0.338 0.185 0.362 0.196 0.279 Table 2. Quantitative comparison on the Dynamic Gibson for 3D reconstruction and novel view synthesis. Method Metric Beac Coop Denm Elmi Eudo Grei Home Ribe Avg. NARUTO (Feng et al., 2024) Acc (cm)↓ 5.42 4.75 4.62 6.31 5.14 5.26 4.58 4.38 5.06 Com. (cm)↓ 5.72 5.36 5.18 7.02 5.72 5.88 5.12 4.95 5.66 C.R. (%)↑ 84.38 86.95 87.42 81.64 85.91 84.72 88.62 88.76 86.05 PSNR↑ 18.94 18.68 17.91 16.85 18.36 16.97 19.31 20.82 18.48 SSIM↑ 0.691 0.737 0.682 0.659 0.676 0.586 0.687 0.763 0.685 LPIPS↓ 0.472 0.449 0.468 0.496 0.531 0.488 0.469 0.456 0.479 ActiveSplat (Li et al., 2025c) Acc (cm)↓ 5.61 4.52 4.05 4.78 4.18 4.65 3.96 3.87 4.45 Com. (cm)↓ 6.25 5.03 4.58 5.41 4.72 5.16 4.34 4.26 4.97 C.R. (%)↑ 84.26 88.37 89.94 86.91 89.18 87.05 90.36 90.82 88.36 PSNR↑ 22.22 21.05 19.34 20.17 20.88 22.43 21.81 20.72 21.08 SSIM↑ 0.797 0.833 0.789 0.714 0.761 0.736 0.731 0.782 0.768 LPIPS↓ 0.323 0.369 0.519 0.445 0.531 0.339 0.456 0.451 0.429 ActiveGAMER (Chen et al., 2025) Acc (cm)↓ 5.08 3.72 4.14 4.31 4.08 3.61 3.51 3.42 3.98 Com. (cm)↓ 5.74 4.28 4.71 4.92 4.56 4.15 3.96 3.88 4.53 C.R. (%)↑ 87.84 92.03 90.47 90.12 91.35 92.76 93.11 93.48 91.40 PSNR↑ 23.47 22.11 20.42 21.34 21.96 22.58 23.95 22.08 22.24 SSIM↑ 0.794 0.832 0.747 0.723 0.756 0.729 0.748 0.785 0.764 LPIPS↓ 0.338 0.391 0.486 0.414 0.398 0.366 0.409 0.372 0.397 Ours Acc (cm)↓ 2.14 2.31 2.24 2.68 3.52 2.76 3.45 2.03 2.52 Com. (cm)↓ 2.56 2.78 2.71 3.26 3.04 3.35 3.98 2.48 3.02 C.R. (%)↑ 96.18 95.37 95.84 93.86 94.92 93.94 91.75 96.73 94.82 PSNR↑ 25.28 24.34 21.86 23.91 24.18 24.86 25.65 24.62 24.34 SSIM↑ 0.866 0.858 0.811 0.771 0.842 0.835 0.841 0.852 0.835 LPIPS↓ 0.254 0.243 0.339 0.262 0.298 0.286 0.311 0.307 0.288 Results on Dynamic Gibson. Tab. 2 reports the quantitative results on Gibson scenes augmented with our pedestrian injection protocol. Although Gibson has traditionally been used as a static active reconstruction benchmark, introducing dynamic humans leads to frequent occlusions and motion-corrupted observations, which pose significant challenges to next-best-view planning. DynActiveGS consistently outperforms all baselines in both geometric and photometric metrics. Compared with the strongest static-scene active reconstruction baseline, our method reduces reconstruction error and improves completion ratio. More importantly, the rendering quality is substantially improved, as reflected by higher PSNR and significantly lower LPIPS. These results demonstrate that the proposed motion-aware uncertainty modeling generalizes effectively to active reconstruction in dynamic scenes. While static-scene baselines continue to expand coverage, they cannot filter motion-contaminated observations, often resulting in duplicated structures and texture inconsistencies. By contrast, DynActiveGS explicitly disentangles structural uncertainty from motion-induced uncertainty, enabling the agent to avoid high-risk viewpoints and maintain stable map updates. More importantly, this experiment shows that the gains of our method are not tied to a specific dataset. Even when dynamic agents are introduced through an external augmentation protocol rather than being natively embedded in the dataset, our framework still exhibits strong robustness and consistent reconstruction quality. Figure 5. Reconstruction progress on Social-MP3D HxpK. DynActiveGS achieves the fastest improvement in both PSNR and completion ratio throughout the exploration process. Figure 6. Rendering quality across reconstruction stages on Dynamic Gibson Denm scene. The proposed uncertainty-aware optimization progressively suppresses dynamic artifacts and improves reconstruction quality from the initial online result to the final refined result. Table 3. Ablation of core components. We progressively enable uncertainty-weighted Gaussian mapping, motion-aware viewpoint selection, and dynamic-constrained planning. Each component brings consistent gains, and the full model achieves the best balance between geometric reconstruction and rendering quality. Variant Weighted Map Motion-aware View Dyn. Planner Acc↓ Com↓ C.R.↑ PSNR↑ SSIM↑ LPIPS↓ L1-D↓ Baseline ✗ ✗ ✗ 3.21 3.84 90.62 18.71 0.503 0.505 6.94 + Weighted Mapping ✓ ✗ ✗ 2.98 3.56 91.85 19.83 0.621 0.471 5.12 + Motion-aware View ✓ ✓ ✗ 2.76 3.28 92.93 20.66 0.733 0.409 4.08 Full (Ours) ✓ ✓ ✓ 2.24 2.71 95.84 21.86 0.811 0.339 2.92 Table 4. Ablation of motion-aware viewpoint selection. Variant UsU_s UmU_m Cost Acc↓ C.R.↑ PSNR↑ Random ✗ ✗ ✗ 3.08 91.24 19.37 Static-Gain ✓ ✗ ✗ 2.81 92.15 20.14 Gain+Cost ✓ ✗ ✓ 2.63 93.08 20.73 Ours ✓ ✓ ✓ 2.03 96.73 24.62 Table 5. Ablation of dynamic-constrained planning. Path Ratio Static Hier. + Dyn. Edge Cost Ours Com.↓ C.R.↑ PSNR↑ Com.↓ C.R.↑ PSNR↑ Com.↓ C.R.↑ PSNR↑ 25% 6.12 88.47 18.95 5.86 89.73 19.42 5.54 90.68 20.91 50% 5.36 90.22 19.71 5.08 91.84 20.74 4.28 92.63 21.87 75% 4.25 91.67 20.82 3.96 92.31 21.26 3.61 93.46 22.72 100% 3.89 92.28 21.97 3.66 93.02 22.63 2.96 94.41 24.17 Qualitative Comparison. Fig. 3 and Fig. 4 present qualitative comparisons. When humans move across the field of view, static baselines tend to produce duplicated structures and motion streaks. DynActiveGS effectively suppresses these artifacts while preserving consistent geometry and texture fidelity. These observations further validate the benefit of risk-aware planning under dynamic disturbances. These results suggest that the primary limitation of conventional active reconstruction in dynamic environments is not insufficient view sampling capacity, but the lack of motion-aware decision-making. Reconstruction Progress over Exploration. As shown in Fig. 5, DynActiveGS exhibits the fastest improvement in both rendering quality and reconstruction completeness on the Social-MP3D HxpK scene. Compared with NARUTO (Feng et al., 2024), ActiveSplat (Li et al., 2025c), and ActiveGAMER (Chen et al., 2025), our method achieves higher PSNR and completion ratio across nearly all exploration stages, further validating the effectiveness of the proposed dynamic-aware active reconstruction pipeline. 4.3. Ablation Study We conduct ablation studies on representative dynamic scenes from Social-MP3D and Dynamic Gibson to validate the key components of DynActiveGS. Our analysis covers three aspects: (1) core module ablation, (2) motion-aware viewpoint selection, and (3) dynamic-constrained planning. Core module ablation. We first evaluate the contribution of the main modules, including uncertainty-weighted Gaussian mapping, motion-aware viewpoint selection, and dynamic-constrained planning, using the Dynamic Gibson Denm scene. As shown in Tab. 3, each component consistently improves both reconstruction and rendering quality. In particular, uncertainty-weighted mapping suppresses motion-corrupted observations, motion-aware viewpoint selection improves observation quality, and dynamic-constrained planning further enhances exploration efficiency. The full model achieves the best overall performance across all metrics. Ablation of motion-aware viewpoint selection. We compare random viewpoint sampling (Random), structural-gain-only selection (Static-Gain), structural gain with travel cost (Gain+Cost), and our full motion-aware formulation (Ours) on the Dynamic Gibson Ribe scene. Tab. 4 shows that using structural uncertainty alone is insufficient in dynamic scenes, since the agent is often attracted to regions with high gain but low observation reliability. Introducing travel cost improves efficiency, while explicitly modeling motion uncertainty further improves both reconstruction accuracy and rendering fidelity. These results confirm the importance of motion-aware viewpoint utility under human disturbances. Ablation of dynamic-constrained planning. We further analyze the planner by comparing a static hierarchical planner (Static Hier.), a planner with dynamic edge cost but without temporal accumulation (+ Dyn. Edge Cost), and our full dynamic-constrained planner (Ours). To evaluate long-horizon efficiency, we extend the exploration horizon to 2000 steps on the Social-MP3D YmJk scene and report results at different path ratios in Tab. 5. Hierarchical planning already improves long-horizon exploration over naive shortest-path execution, while dynamic edge cost reduces unstable routes through highly dynamic regions. Temporal accumulation yields the best final performance, suggesting that persistent motion patterns are more informative for planning than instantaneous disturbances. Stage-wise rendering analysis. Fig. 6 visualizes reconstruction quality at different stages of the pipeline. The initial online result (Online Before Opt) still contains strong dynamic artifacts. After uncertainty-guided online optimization (Online After Opt), these artifacts are substantially reduced. The final offline refinement further improves consistency and produces the best visual quality. Together with the uncertainty visualization, this result shows that the learned uncertainty map effectively highlights motion-corrupted regions and supports robust Gaussian optimization. Overall, these ablations consistently validate the design of DynActiveGS. Uncertainty-weighted mapping improves reconstruction robustness, motion-aware viewpoint selection improves observation quality, and dynamic-constrained planning enhances long-horizon exploration efficiency. Together, these components enable robust active reconstruction in dynamic environments. 5. Conclusion and Future Work We introduced DynActiveGS, a dynamic-aware active reconstruction framework based on 3D Gaussian Splatting. By disentangling structural and motion-induced uncertainty, and integrating dynamic-aware viewpoint selection with constrained path planning, our method enables robust closed-loop exploration and high-fidelity reconstruction in dynamic environments. Experiments on Social-MP3D and dynamically augmented Gibson scenes demonstrate consistent improvements over existing baselines in reconstruction accuracy, completeness, and rendering quality. Limitations. Our evaluation is currently conducted in simulation with controllable human dynamics, while real-world environments may contain more diverse and unpredictable motions. Moreover, the framework is not yet optimized for large-scale multi-floor environments with complex spatial structures. Future Work. Future directions include deployment on real robotic platforms, online motion prediction, and extending dynamic-aware active reconstruction toward 4D scene representations (Lin et al., 2026) for modeling persistent scene dynamics. 6. Acknowledgments This work was supported by the National Natural Science Foundation of China under Grant Nos. 62293545 and U21B6002, in part by the Major Key Project of PCL under Grant Nos. PCL2024A06 and PCL2025A10, and in part by the Shenzhen Science and Technology Program under Grant Nos. RCJC20210706091946001, ZDCY2025090-1104207008, and RCJC20231211085918010. References J. Aloimonos, I. Weiss, and A. Bandyopadhyay (1988) Active vision. International journal of computer vision 1 (4), p. 333–356. Cited by: §1. J. T. Barron, B. Mildenhall, D. Verbin, P. P. Srinivasan, and P. Hedman (2022) Mip-nerf 360: unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5470–5479. Cited by: §2.1. B. Bescos, J. M. Fácil, J. Civera, and J. Neira (2018) DynaSLAM: tracking, mapping, and inpainting in dynamic scenes. IEEE robotics and automation letters 3 (4), p. 4076–4083. Cited by: §2.3. A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Niebner, M. Savva, S. Song, A. Zeng, and Y. Zhang (2018) Matterport3D: learning from rgb-d data in indoor environments. In 7th IEEE International Conference on 3D Vision, 3DV 2017, p. 667–676. Cited by: §4.1. A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su (2022) Tensorf: tensorial radiance fields. In European conference on computer vision, p. 333–350. Cited by: §2.1. L. Chen, H. Zhan, K. Chen, X. Xu, Q. Yan, C. Cai, and Y. Xu (2025) Activegamer: active gaussian mapping through efficient rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 16486–16497. Cited by: §1, §4.1, §4.1, §4.2, Table 1, Table 2. S. Chen, Y. Li, and N. M. Kwok (2011) Active vision in robotic systems: a survey of recent developments. The International Journal of Robotics Research 30 (11), p. 1343–1377. Cited by: §1. X. Chen, Q. Li, T. Wang, T. Xue, and J. Pang (2024) Gennbv: generalizable next-best-view policy for active 3d reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 16436–16445. Cited by: §2.2. S. Chng, S. Ramasinghe, J. Sherrah, and S. Lucey (2022) Gaussian activated neural radiance fields for high fidelity reconstruction and pose estimation. In European Conference on Computer Vision, p. 264–280. Cited by: §2.1. A. Dai, M. Nießner, M. Zollhöfer, S. Izadi, and C. Theobalt (2017) Bundlefusion: real-time globally consistent 3d reconstruction using on-the-fly surface reintegration. ACM Transactions on Graphics (ToG) 36 (4), p. 1. Cited by: §1. Z. Feng, H. Zhan, Z. Chen, Q. Yan, X. Xu, C. Cai, B. Li, Q. Zhu, and Y. Xu (2024) Naruto: neural active reconstruction from uncertain target observations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 21572–21583. Cited by: §1, §2.2, §4.1, §4.2, Table 1, Table 2. Z. Gong, T. Hu, R. Qiu, and J. Liang (2025) From cognition to precognition: a future-aware framework for social navigation. In 2025 IEEE International Conference on Robotics and Automation (ICRA), p. 9122–9129. Cited by: §4.1. B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao (2024) 2d gaussian splatting for geometrically accurate radiance fields. In ACM SIGGRAPH 2024 conference papers, p. 1–11. Cited by: §1. S. Isler, R. Sabzevari, J. Delmerico, and D. Scaramuzza (2016) An information gain formulation for active volumetric 3d reconstruction. In 2016 IEEE international conference on robotics and automation (ICRA), p. 3477–3484. Cited by: §1. H. Jiang, Y. Xu, K. Li, J. Feng, and L. Zhang (2024a) Rodyn-slam: robust dynamic dense rgb-d slam with neural radiance fields. IEEE Robotics and Automation Letters 9 (9), p. 7509–7516. Cited by: §2.3. W. Jiang, B. Lei, and K. Daniilidis (2024b) Fisherrf: active view selection and mapping with radiance fields using fisher information. In European Conference on Computer Vision, p. 422–440. Cited by: §1, §2.2. L. Jin, X. Zhong, Y. Pan, J. Behley, C. Stachniss, and M. Popović (2025) Activegs: active scene reconstruction using gaussian splatting. IEEE Robotics and Automation Letters. Cited by: §1, §2.2, §4.1. R. Jin, Y. Gao, Y. Wang, Y. Wu, H. Lu, C. Xu, and F. Gao (2024) Gs-planner: a gaussian-splatting-based planning framework for active high-fidelity reconstruction. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 11202–11209. Cited by: §1, §2.2. I. Karamouzas, B. Skinner, and S. J. Guy (2014) Universal power law governing pedestrian interactions. Physical review letters 113 (23), p. 238701. Cited by: §4.1. N. Keetha, J. Karhade, K. M. Jatavallabhula, G. Yang, S. Scherer, D. Ramanan, and J. Luiten (2024) Splatam: splat track & map 3d gaussians for dense rgb-d slam. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 21357–21366. Cited by: §2.1, §2.3. B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis (2023) 3D gaussian splatting for real-time radiance field rendering.. ACM Trans. Graph. 42 (4), p. 139–1. Cited by: §1, §2.1. A. Kirsch, J. Van Amersfoort, and Y. Gal (2019) Batchbald: efficient and diverse batch acquisition for deep bayesian active learning. Advances in neural information processing systems 32. Cited by: §1. Z. Kuang, Z. Yan, H. Zhao, G. Zhou, and H. Zha (2024) Active neural mapping at scale. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 7152–7159. Cited by: §1, §2.2. J. Kulhanek, S. Peng, Z. Kukelova, M. Pollefeys, and T. Sattler (2024) WildGaussians: 3d gaussian splatting in the wild. Advances in Neural Information Processing Systems 37, p. 21271–21288. Cited by: §1. S. Lee, L. Chen, J. Wang, A. Liniger, S. Kumar, and F. Yu (2022) Uncertainty guided policy for active robotic 3d reconstruction using neural radiance fields. IEEE Robotics and Automation Letters 7 (4), p. 12070–12077. Cited by: §1, §2.2. K. Li, Y. Tang, V. A. Prisacariu, and P. H. Torr (2022) Bnv-fusion: dense 3d reconstruction using bi-level neural volume fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 6166–6175. Cited by: §2.1. M. Li, Z. Zhu, M. Pollefeys, and D. Barath (2026) DROID-slam in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §3.2.2. S. Li, A. Guédon, C. Boittiaux, S. Chen, and V. Lepetit (2025a) NextBestPath: efficient 3d mapping of unseen environments. In ICLR 2025-Thirteenth International Conference on Learning Representations, Cited by: §2.2. Y. Li, Y. Li, and G. H. Lee (2025b) Active3D: Active High-Fidelity 3D Reconstruction via Hierarchical Uncertainty Quantification. arXiv. Note: arXiv:2511.20050 [cs] External Links: Link, Document Cited by: §1. Y. Li, C. Lyu, Y. Di, G. Zhai, G. H. Lee, and F. Tombari (2024) Geogaussian: geometry-aware gaussian splatting for scene rendering. In European conference on computer vision, p. 441–457. Cited by: §2.1. Y. Li, Z. Kuang, T. Li, Q. Hao, Z. Yan, G. Zhou, and S. Zhang (2025c) ActiveSplat: High-Fidelity Scene Reconstruction through Active Gaussian Splatting. IEEE Robotics and Automation Letters 10 (8), p. 8099–8106. Note: arXiv:2410.21955 [cs] External Links: ISSN 2377-3766, 2377-3774, Link, Document Cited by: §1, §2.2, §4.1, §4.2, Table 1, Table 2. C. Lin, Y. Lin, P. Pan, Y. Yu, T. Hu, H. Yan, K. Fragkiadaki, and Y. Mu (2026) MoVieS: motion-aware 4d dynamic view synthesis in one second. In Proceedings of the Computer Vision and Pattern Recognition Conference (CVPR), Cited by: §5. J. Liu, P. Ji, N. Bansal, C. Cai, Q. Yan, X. Huang, and Y. Xu (2022) Planemvs: 3d plane reconstruction from multi-view stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8665–8675. Cited by: §2.1. B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng (2021) Nerf: representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 (1), p. 99–106. Cited by: §2.1. R. A. Newcombe, S. Izadi, O. Hilliges, D. Molyneaux, D. Kim, A. J. Davison, P. Kohi, J. Shotton, S. Hodges, and A. Fitzgibbon (2011) Kinectfusion: real-time dense surface mapping and tracking. In 2011 10th IEEE international symposium on mixed and augmented reality, p. 127–136. Cited by: §1. M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. (2023) Dinov2: learning robust visual features without supervision. arXiv preprint arXiv:2304.07193. Cited by: §3.2.1. E. Palazzolo, J. Behley, P. Lottes, P. Giguere, and C. Stachniss (2019) ReFusion: 3d reconstruction in dynamic environments for rgb-d cameras exploiting residuals. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 7855–7862. Cited by: §1. X. Pan, Z. Lai, S. Song, and G. Huang (2022) Activenerf: learning where to see with uncertainty estimation. In European Conference on Computer Vision, p. 230–246. Cited by: §1, §2.2. D. Peralta, J. Casimiro, A. M. Nilles, J. A. Aguilar, R. Atienza, and R. Cajote (2020) Next-best view policy for 3d reconstruction. In European Conference on Computer Vision, p. 558–573. Cited by: §1. R. Pito (2002) A solution to the next best view problem for automated surface acquisition. IEEE Transactions on pattern analysis and machine intelligence 21 (10), p. 1016–1030. Cited by: §1. K. D. Polyzos, A. Bacharis, S. Madhuvarasu, N. Papanikolopoulos, and T. Javidi (2025) ActiveInitSplat: How Active Image Selection Helps Gaussian Splatting. arXiv. Note: arXiv:2503.06859 [cs] External Links: Link, Document Cited by: §1. Y. Ran, J. Zeng, S. He, J. Chen, L. Li, Y. Chen, G. Lee, and Q. Ye (2023) Neurar: neural uncertainty for autonomous 3d reconstruction with implicit neural representations. IEEE Robotics and Automation Letters 8 (2), p. 1125–1132. Cited by: §1. W. Ren, Z. Zhu, B. Sun, J. Chen, M. Pollefeys, and S. Peng (2024) Nerf on-the-go: exploiting uncertainty for distractor-free nerfs in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 8931–8940. Cited by: §1, §3.2.2. E. Sandström, G. Zhang, K. Tateno, M. Oechsle, M. Niemeyer, Y. Zhang, M. Patel, L. Van Gool, M. Oswald, and F. Tombari (2025) Splat-slam: globally optimized rgb-only slam with 3d gaussians. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 1680–1691. Cited by: §1, §2.1, §2.3. M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, et al. (2019) Habitat: a platform for embodied ai research. In Proceedings of the IEEE/CVF international conference on computer vision, p. 9339–9347. Cited by: §4.1. J. L. Schonberger and J. Frahm (2016) Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 4104–4113. Cited by: §2.1. O. Siméoni, H. V. Vo, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V. Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, et al. (2025) Dinov3. arXiv preprint arXiv:2508.10104. Cited by: §3.2.1. J. Van Den Berg, S. J. Guy, M. Lin, and D. Manocha (2011) Reciprocal n-body collision avoidance. In Robotics research: the 14th international symposium ISRR, p. 3–19. Cited by: §4.1. A. Vuong, T. Nguyen, M. N. Vu, B. Huang, H. T. T. Binh, T. Vo, and A. Nguyen (2024) Habicrowd: a high performance simulator for crowd-aware visual navigation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 5821–5827. Cited by: §4.1. T. Whelan, S. Leutenegger, R. F. Salas-Moreno, B. Glocker, and A. J. Davison (2015) ElasticFusion: dense slam without a pose graph.. In Robotics: science and systems, Vol. 11. Cited by: §1. P. Wu, Y. Mu, B. Wu, Y. Hou, J. Ma, S. Zhang, and C. Liu (2024) VoroNav: voronoi-based zero-shot object navigation with large language model. In International Conference on Machine Learning, p. 53757–53775. Cited by: §3.3. F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese (2018) Gibson env: real-world perception for embodied agents. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 9068–9079. Cited by: §4.1. Y. Xie, Y. Cai, Y. Zhang, L. Yang, and J. Pan (2025) GauSS-MI: Gaussian Splatting Shannon Mutual Information for Active 3D Reconstruction. In Proceedings of Robotics: Science and Systems, LosAngeles, CA, USA. External Links: Document Cited by: §1. Y. Xu, H. Jiang, Z. Xiao, J. Feng, and L. Zhang (2024) Dg-slam: robust dynamic gaussian splatting slam with hybrid pose optimization. Advances in Neural Information Processing Systems 37, p. 51577–51596. Cited by: §1, §2.3. Z. Xu, R. Jin, K. Wu, Y. Zhao, Z. Zhang, J. Zhao, F. Gao, Z. Gan, and W. Ding (2025) HGS-planner: hierarchical planning framework for active scene reconstruction using 3d gaussian splatting. In 2025 IEEE International Conference on Robotics and Automation (ICRA), Vol. , p. 14161–14167. External Links: Document Cited by: §1, §2.1, §2.2. Z. Yan, H. Yang, and H. Zha (2023) Active Neural Mapping. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, p. 10947–10958 (en). External Links: ISBN 979-8-3503-0718-4, Link, Document Cited by: §2.2. Z. Yu, A. Chen, B. Huang, T. Sattler, and A. Geiger (2024) Mip-splatting: alias-free 3d gaussian splatting. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 19447–19456. Cited by: §1. Z. Zhang, F. Xu, and M. Zhang (2025) Peering into the unknown: active view selection with neural uncertainty maps for 3d reconstruction. arXiv preprint arXiv:2506.14856. Cited by: §2.2. J. Zheng, Z. Zhu, V. Bieri, M. Pollefeys, S. Peng, and I. Armeni (2025) Wildgs-slam: monocular gaussian splatting slam in dynamic environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 11461–11471. Cited by: §1, §2.3, §3.2.2. Z. Zhu, W. Zhang, N. Haala, M. Pollefeys, and D. Barath (2025) VIGS-SLAM: Visual Inertial Gaussian Splatting SLAM. arXiv. Note: arXiv:2512.02293 [cs] External Links: Link, Document Cited by: §2.3.