Paper deep dive
Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy
Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/21/2026, 5:48:32 AM
Summary
Vis2Reg is a visibility-aware, landmark-free geometric 3D-2D registration framework for liver laparoscopy. It addresses challenges of severe occlusion and limited visibility by using mask-consistent visible regions for self-supervised learning. The method combines a robust rigid initialization module with an implicit neural deformation field, achieving high accuracy (92.6% Dice score) and efficiency (111 ms inference) on intraoperative datasets.
Entities (10)
Relation Signals (7)
Vis2Reg → achievesmetric → Dice Score 92.6%
confidence 95% · Vis2Reg achieves a Dice score of 92.6%
Vis2Reg → achievesmetric → Chamfer Distance 1.43 mm
confidence 95% · and a Chamfer Distance of 1.43 mm
Vis2Reg → appliesto → Liver Laparoscopy
confidence 95% · Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy
Vis2Reg → evaluatedon → P2I-LReg Dataset
confidence 95% · We evaluate Vis2Reg on the public P2I-LReg dataset
Vis2Reg → outperforms → Self-P2IR
confidence 92% · Compared with the strongest prior method (Self-P2IR), Vis2Reg improves Dice by 13.71 pp
Vis2Reg → uses → Differentiable Point Rasterization
confidence 90% · enabled by differentiable point rasterization and mask-guided back-projection
Vis2Reg → uses → Implicit Neural Deformation Field
confidence 90% · Vis2Reg combines a robust geometric rigid initialization module with an implicit neural deformation field
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision. Existing landmark-free approaches perform partial-to-complete geometric alignment, yet robust self-supervision under extreme partial visibility remains difficult. We propose Vis2Reg, a visibility-aware registration framework that explicitly constrains deformation using mask-consistent visible regions. We introduce a visibility-aware self-supervision that derives a visible-domain 3D supervision signal from intraoperative masks, enabled by differentiable point rasterization and mask-guided back-projection. This formulation improves robustness under severe occlusion while maintaining fully self-supervised learning. Vis2Reg combines a robust geometric rigid initialization module with an implicit neural deformation field for stable alignment. Vis2Reg achieves a Dice score of 92.6\% and a Chamfer Distance of 1.43 mm on real intraoperative datasets, with 111 ms per-frame inference time, demonstrating both accuracy and practical efficiency.
Tags
Links
- Source: https://arxiv.org/abs/2607.17810v1
- Canonical: https://arxiv.org/abs/2607.17810v1
Trouble viewing inline? Open PDF directly →
Full Text
30,999 characters extracted from source content.
Expand or collapse full text
11institutetext: AI in Medicine and Surgery Group, School of Computer Science, University of Leeds, LS2 9JT, Leeds, United Kingdom 11email: vqkc0507, s.s.ali@leeds.ac.uk 22institutetext: Department of Diagnostic Radiology, Li Ka Shing Faculty of Medicine, The University of Hong Kong, Pok Fu Lam, Hong Kong 22email: xukunzhg@hku.hk 33institutetext: Department of HPB and Transplant Surgery, St. James’s University Hospital, Leeds, United Kingdom 33email: s.farid@nhs.net Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D–2D Registration for Liver Laparoscopy Jiaming Feng Xukun Zhang Shahid Farid Sharib Ali(✉) Abstract Accurate 3D–2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision. Existing landmark-free approaches perform partial-to-complete geometric alignment, yet robust self-supervision under extreme partial visibility remains difficult. We propose Vis2Reg, a visibility-aware registration framework that explicitly constrains deformation using mask-consistent visible regions. We introduce a visibility-aware self-supervision that derives a visible-domain 3D supervision signal from intraoperative masks, enabled by differentiable point rasterization and mask-guided back-projection. This formulation improves robustness under severe occlusion while maintaining fully self-supervised learning. Vis2Reg combines a robust geometric rigid initialization module with an implicit neural deformation field for stable alignment. Vis2Reg achieves a Dice score of 92.6% and a Chamfer Distance of 1.43 m on real intraoperative datasets, with 111 ms per-frame inference time, demonstrating both accuracy and practical efficiency. 1 Introduction Minimally invasive laparoscopic and robotic liver resection is increasingly adopted, but surgeons rely on a narrow field-of-view with no direct access to sub-surface anatomy, limiting assessment of tumour–vessel relationships and increasing surgical risk [5, 1]. Augmented reality (AR) can mitigate this by overlaying patient-specific preoperative 3D liver, vascular, or tumour models onto the intraoperative video, improving spatial awareness beyond the visible surface [1, 17, 21]. Accurate, robust preoperative-to-intraoperative fusion of the deformable liver is therefore critical for reliable AR guidance. At its core, AR-guided laparoscopy requires preoperative-to-intraoperative 3D–2D registration of a preoperative liver model to intraoperative views and camera geometry, which is ill-posed under large non-rigid deformation, adverse imaging (occlusion, specularities, smoke, blood), and partial, view-dependent observations [4, 12, 22, 10, 11]. Early approaches relied on landmarks, contours, or biomechanical constraints [4, 2, 13], but are brittle under occlusion and low overlap [11, 22, 13]. This family spans classical iterative closest point-based (ICP) optimization, biomechanical finite element method (FEM) simulation, and shape-matching such as functional maps [2, 16, 15]. Learning-based approaches predict deformation directly from observations [12] yet remain limited by scarce supervision and incomplete views [1, 13]. A recent paradigm shift reformulates laparoscopic registration as partial-to-complete 3D registration, aligning sparse intraoperative surface observations with the full preoperative liver geometry [26, 7, 14]. This formulation enables self-supervised learning using readily available signals such as segmentation masks and geometric consistency, eliminating the need for explicit 3D correspondence supervision [26, 7]. Nevertheless, intraoperative supervision is fundamentally visibility limited: only the mask-consistent visible surface provides reliable geometric constraints, and explicitly modeling such visibility-consistent supervision to guide deformation remains challenging under severe occlusion and limited overlap [6, 3, 22, 11, 14]. Furthermore, rigid initialization is often unstable when intraoperative observations are sparse and noisy [22, 10, 7, 14]. Generic low-overlap point-cloud registration methods (e.g., PREDATOR [8], GeoTransformer [19]) improve correspondence robustness but are not tailored to visibility-limited surgical self-supervision. To address these challenges, we propose Vis2Reg, a visibility-aware, landmark-free, geometry-only framework for self-supervised laparoscopic liver registration under severe partial visibility. It derives mask-consistent visible-domain 3D supervision from intraoperative masks via differentiable point rasterization and mask-guided back-projection [9], focusing learning on truly observable geometry. This improves deformation robustness while preserving anatomical plausibility, and enables geometry-only inference without masks or explicit correspondences. Our contributions are three-fold: (1) We introduce a visibility-aware self-supervision formulation that constructs mask-consistent visible-domain 3D supervision signals for landmark-free laparoscopic liver registration; (2) we develop a differentiable point rasterization and mask-guided back-projection mechanism to derive visibility-consistent geometric supervision from intraoperative observations; and (3) we integrate robust geometric rigid initialization with an implicit neural deformation field, achieving improved registration robustness and near-real-time performance on in-vivo laparoscopic datasets. 2 Method 2.1 Problem formulation Vis2Reg addresses self-supervised intraoperative liver registration under severe partial visibility and without 3D–2D ground-truth supervision. The clinical goal is a 3D-to-2D AR overlay, achieved as partial-to-complete 3D registration aligning the complete preoperative cloud P with the partial intraoperative observation Q. Each view QvQ_v is reconstructed from laparoscopic images via monocular depth (DepthAnything [24]), the liver mask MvM_v, and back-projection with intrinsics KvK_v. Given a preoperative liver point cloud P=pii=1NP=\p_i\_i=1^N and F view-dependent intraoperative point clouds Qvv=0F−1\Q_v\_v=0^F-1 (all in camera coordinates), together with F-view masks Mvv=0F−1\M_v\_v=0^F-1 and intrinsics Kvv=0F−1\K_v\_v=0^F-1, we denote the merged intraoperative observation as Q≜⋃v=0F−1QvQ _v=0^F-1Q_v. We use F=3F=3 as a short local, non-temporal window corresponding to a near-static liver; MvM_v and KvK_v serve as per-view supervision signals, not model inputs. The objective is to estimate a rigid alignment and a non-rigid deformation modeled as an implicit displacement field gϕg_φ, producing a warped point cloud W W that aligns P to the intraoperative observations under visibility-consistent geometric constraints. As shown in Fig. 1, this is achieved using a geometry-only registration network comprising a robust rigid initialization module and the implicit deformation field gϕg_φ. During training, visibility-consistent supervision is constructed from masks via differentiable rasterization technique and mask-gated back-projection, while at inference the model operates using geometric inputs only. Figure 1: Vis2Reg pipeline. Given preoperative P and intraoperative partial Q, the network outputs warped geometry W W via rigid initialization and non-rigid deformation; differentiable rasterization and mask-gated back-projection construct the visible-domain supervision set VvMV_v^M from (D^v,S^v)( D_v, S_v) and MvM_v. Figure 2: Architecture of the geometry-only registration network in Vis2Reg: EdgeConv feature encoding, rigid initialization (TrigidT_rigid), and implicit non-rigid deformation field gϕg_φ producing the warped point cloud W W. Bottom: EdgeConv update block. 2.2 Geometry-Only Registration Network As illustrated in Fig. 1, Vis2Reg predicts the warped geometry W W via (i) geometric feature encoders, (i) a robust rigid initialization module for stable global alignment under partial visibility, and (i) an implicit displacement field gϕg_φ modeling continuous local deformation. Concretely, the rigid module builds on a GeoTransformer matcher and the deformation field is implemented as a multilayer perceptron (MLP) with sinusoidal representation network (SIREN) activations, both detailed below. Geometric Feature Encoders. In Fig. 2, we extract multi-scale geometric features using a multi-layer EdgeConv encoder [23], which captures local shape context through dynamic neighborhood aggregation. Given a point set X=xii=1nX=\x_i\_i=1^n with initial features hi(0)=xih_i^(0)=x_i, each layer updates point features via hi(ℓ+1)=maxj∈k(i)ψℓ([hi(ℓ)∥hj(ℓ)−hi(ℓ)]),h_i^( +1)= _j _k(i) _ \! ( [h_i^( )\,\|\,h_j^( )-h_i^( ) ] ), (1) where ψℓ(⋅) _ (·) is a shared MLP and k(i)N_k(i) denotes the k-nearest-neighbor (k-N) neighborhood. Concatenating intermediate outputs yields multi-scale features F(X)F(X). We employ two encoders with distinct roles: a Siamese matching encoder ℰmE_m that extracts correspondence features for rigid alignment from (P,Q)(P,Q), while a registration encoder ℰrE_r extracts deformation features conditioned on the pair (P,Q)(P,Q) for the implicit field, i.e., FPm=ℰm(P)F_P^m=E_m(P), FQm=ℰm(Q)F_Q^m=E_m(Q), and FP|Qr=ℰr(P,Q)F_P|Q^r=E_r(P,Q), queried at the source points P. Rigid Seed Initialization. To obtain a stable rigid alignment under sparse, noisy observations, we estimate an initial pose via soft correspondence, hypothesis evaluation, and geometric refinement (Fig. 2). GeoTransformer contextualizes FPm,FQmF_P^m,F_Q^m into F~Pm,F~Qm F_P^m, F_Q^m, giving a similarity matrix Aij=⟨f~i,g~j⟩/τA_ij= f_i, g_j /τ; Sinkhorn normalization yields a soft assignment Π with confidence weights wij=Πijw_ij= _ij, and mutual nearest neighbors form a correspondence set C. Progressive Sample Consensus (PROSAC) then generates pose hypotheses Tm\T_m\ from minimal samples in C, ranked by weighted inlier scores; the top hypotheses are refined by trimmed point-to-point and point-to-plane ICP, and the best provides the rigid initialization for gϕg_φ. Implicit Non-Rigid Deformation Field. We represent non-rigid deformation as a continuous implicit displacement field gϕg_φ conditioned on geometric features. For each source point x∈Px∈ P, positional encoding PE(x)∈ℝdpePE(x) ^d_pe is concatenated with its pair-conditioned registration feature f(x)=FP|Qr(x)f(x)=F_P|Q^r(x) to form the field input z(x)=PE(x)⊕f(x)z(x)=PE(x) f(x). A SIREN-based MLP models the implicit field and predicts a displacement Δ(x)=gϕ(z(x))∈ℝ3 (x)=g_φ(z(x)) ^3, which produces a locally deformed point x′=x+Δ(x)x =x+ (x). The final warped geometry is obtained by applying the rigid alignment to the deformed point: x^=sRP,Qx′+tP,Q,W^=x^∣x∈P. x=sR_P,Q\,x +t_P,Q, W=\ x x∈ P\. (2) The resulting warped point cloud W W is then used for visibility-aware supervision. Since both the pair-conditioned feature FP|QrF_P|Q^r and the rigid transform (sRP,Q,tP,Q)(sR_P,Q,t_P,Q) depend on Q, the deformation adapts to each intraoperative input rather than being a fixed function of P. 2.3 Visibility-Aware Self-Supervised Learning The observed clouds Qvv=0F−1\Q_v\_v=0^F-1 are view-dependent visible subsets of the surface, with no 3D registration ground truth. Vis2Reg therefore constructs a mask-consistent visible-domain 3D supervision signal, restricting supervision to geometrically observable regions defined by the masks and guiding deformation only by physically valid observations. Unlike the symmetric rendered-mask consistency of Self-P2IR [26], our supervision is mask-gated and one-way (observation-to-model), so unobserved model regions are never penalized. Visible-domain construction. Let Ω denote the discrete W×HW× H image lattice. From the warped cloud W W, intrinsics Kv\K_v\, and masks Mv\M_v\, we render depth and silhouette per view by differentiable point rasterization: (S^v,D^v)=ℛKv(W^),v=0,…,F−1,( S_v, D_v)=R_K_v( W), v=0,…,F-1, (3) where ℛKvR_K_v is parameterized by KvK_v and D^v(u)=0 D_v(u)=0 marks empty pixels. We define the mask-consistent visible pixel set Uv=u∈Ω∣D^v(u)>0∧Mv(u)=1U_v=\u∈ D_v(u)>0 M_v(u)=1\ and back-project it to an explicit visible-domain 3D supervision set: VvM=πKv−1(u,D^v(u))|u∈Uv,V^M_v= \π^-1_K_v\! (u, D_v(u) )\; |\;u∈ U_v \, (4) where πKv−1π^-1_K_v is the analytic back-projection from a pixel-depth pair to a camera-frame 3D point. Self-Supervised Objectives. We optimize a weighted combination of visibility-aware alignment losses and deformation regularization: ℒ=λ3Dℒ3D+λvisℒvis+λsilℒsil+λdefℒdef+λsmoothℒsmooth+λtopoℒtopo.L= _3DL_3D+ _visL_vis+ _silL_sil+ _defL_def+ _smoothL_smooth+ _topoL_topo. (5) We denote the one-way Chamfer as CD→(A,B)=1|A|∑a∈Aminb∈B‖a−b‖22CD_→(A,B)= 1|A| _a∈ A _b∈ B\|a-b\|_2^2 and the symmetric Chamfer as CD(A,B)=CD→(A,B)+CD→(B,A)CD(A,B)=CD_→(A,B)+CD_→(B,A). The visibility-aware alignment terms are an observation-to-model one-way Chamfer ℒ3D=1F∑vCD→(Qv,W^)L_3D= 1F _vCD_→(Q_v, W) (handling the partial visibility of QvQ_v), a symmetric visible-domain term ℒvis=1F∑vCD(VvM,Qv)L_vis= 1F _vCD(V_v^M,Q_v) aligning the mask-gated set VvMV_v^M with QvQ_v, and a silhouette term ℒsil=1F∑v[BCE(S^v,Mv)+DiceLoss(S^v,Mv)]L_sil= 1F _v [BCE( S_v,M_v)+DiceLoss( S_v,M_v) ]. The displacement field is regularized by deformation magnitude (ℒdef=1N∑i‖Δi‖22L_def= 1N _i\| _i\|_2^2), local smoothness (ℒsmooth=1N∑i1|(i)|∑j∈(i)‖Δi−Δj‖22L_smooth= 1N _i 1|N(i)| _j (i) \| _i- _j\|_2^2), and local distance or topology preservation (ℒtopo=1N∑i1|(i)|∑j∈(i)(‖x^i−x^j‖2−‖xi−xj‖2)2L_topo= 1N _i 1|N(i)| _j (i) (\| x_i- x_j\|_2 -\|x_i-x_j\|_2 )^2). Here, (i)N(i) is the k-N neighborhood of xix_i and Δi=Δ(xi) _i= (x_i). Training schedule. We adopt a two-stage synthetic→ schedule: the matching and rigid-seed modules are pretrained on synthetic data with ground-truth (GT) poses (non-rigid field disabled), then trained on real data (no GT) with a short rigid warm-up followed by joint visibility-aware optimization. 3 Experiments 3.1 Dataset and Implementation Details We evaluate Vis2Reg on the public P2I-LReg dataset, a real intraoperative liver registration benchmark with 346 keyframes from 21 patients [26]. Each sample provides a preoperative 3D liver model, an intraoperative sparse point cloud, a binary segmentation mask, and camera intrinsics. Following the official patient-level protocol, patients are split into 12/4/5 for training/validation/testing, and results are averaged over 5-fold cross-validation. For robust rigid initialization, we additionally use the patient-specific Landmark-Free Synthetic Dataset (red-green-blue plus depth (RGB-D) observations rendered from preoperative models via physics-based Blender simulation with diverse viewpoints and occlusions; 2500 samples per patient; 60%, 20%, and 20% split; training-fold patients only). Inputs are P, Q, and multi-view masks and intrinsics Mv,Kv\M_v,K_v\, all in camera coordinates at real-world scale (no centering and normalization). During training we apply random dropout and jittering to P and statistical denoising to Q; both are resampled or zero-padded to Nmax=6000N_ =6000. We use AdamW (lr=3×10−4lr=3× 10^-4, wd=10−4wd=10^-4) with cosine scheduling and automatic mixed precision (AMP). In Stage-2 real-data training, we train for 60 epochs (batch size 1) with a 20-epoch rigid warm-up (non-rigid branch and deformation regularizers frozen), followed by joint optimization; a 3-frame multi-view input is used by default. Loss weights are λ3D=0.5 _3D=0.5, λvis=1.2 _vis=1.2, λsil=1.0 _sil=1.0 (BCE and Dice weights both 1.0), λdef=0.1 _def=0.1, λsmooth=0.1 _smooth=0.1, and λtopo=0.3 _topo=0.3. For differentiable rasterization, points-per-pixel=16 and the maximum number of visible points is 5000. On synthetic data, we report relative rotation error (RRE) and relative translation error (RTE), defined as RRE=arccos((tr(pred⊤gt)−1)/2)RRE= ((tr(R_pred R_gt)-1)/2) (degrees) and RTE=‖pred−gt‖2RTE=\|t_pred-t_gt\|_2 (m). On real data, we report Chamfer Distance (ℒ3DL_3D) and silhouette Dice. Dice (our primary AR-overlay metric) measures overlap between the rasterized registered silhouette and the liver mask, whereas CD is a complementary surface-distance measure, not a target registration error. All baselines use the same patient-level folds and protocol implemented in PyTorch on NVIDIA L40S GPU. Table 1: Comparison on synthetic rigid initialization and real intraoperative registration (mean ± std). RRE and RTE are reported only for rigid-comparable methods; non-rigid methods are marked as “–” in the synthetic block. Method Synthetic (Rigid Init) Real Intraoperative (Final) RRE (∘)↓ RTE (m)↓ Dice (%)↑ CD (m)↓ GeoTransformer [19] 31.30 78.10 71.16 ± 6.46 3.43 ± 0.96 DPF [18] – – 72.19 ± 5.97 3.31 ± 1.18 PointSetReg [25] – – 75.24 ± 5.92 3.20 ± 1.03 Self-P2IR [26] 0.21 1.32 78.89 ± 6.76 2.97 ± 1.06 Ours (Vis2Reg) 0.08 0.26 92.60 ± 7.24 1.43 ± 1.26 3.2 Comparison and Ablation Study On the synthetic benchmark (with 3D GT), Vis2Reg achieves the best rigid initialization performance in the synthetic block of Table 1, with RRE 0.08∘0.08 and RTE 0.260.26 m, outperforming GeoTransformer (RRE 31.30∘31.30 , RTE 78.1078.10 m) and Self-P2IR (RRE 0.21∘0.21 , RTE 1.321.32 m), indicating a more reliable rigid seed under low overlap. DPF and PointSetReg are non-rigid models and are therefore not included in the synthetic rigid comparison. The GeoTransformer entry uses the standalone matcher with its built-in closed-form pose estimation, with no robust outlier rejection (e.g., RANSAC) and no ICP refinement, hence the higher RRE/RTE (Table 1). Vis2Reg instead uses the full pipeline (GeoTransformer matching, mutual filtering, PROSAC, trimmed ICP), reconciling the gap with the robust-estimator values reported on the source dataset [26]. On the real intraoperative benchmark (Table 1), Vis2Reg achieves the best mean performance in both final registration metrics, with Dice 92.60±7.2492.60± 7.24% and CD 1.43±1.261.43± 1.26 m. Compared with the strongest prior method (Self-P2IR), Vis2Reg improves Dice by 13.71 p (percentage points) and reduces CD by 1.54 m. These gains are consistent with our visibility-aware design, where mask-gated visible-domain supervision constrains optimization to observed regions and mitigates drift under severe occlusion and partial visibility. Importantly, Self-P2IR already attains a near-perfect rigid init (RRE 0.21∘0.21 , RTE 1.321.32 m), comparable to ours, yet trails Vis2Reg by 13.7113.71 p in Dice; as both start from a strong rigid seed, this gain stems from the visibility-aware non-rigid supervision, not initialization. The ablations (Table 2) corroborate this: removing visibility-aware supervision (ℒvisL_vis or mask-gating) alone reduces Dice by 1313–1818 p. Fig. 3 shows qualitative results on representative keyframes: Vis2Reg achieves tighter preoperative–intraoperative alignment than the baselines and better adapts the visible surface to observations while preserving globally plausible anatomy. Vis2Reg runs at 111.38 ms per forward pass with 1.95 GB GPU memory on an NVIDIA L40S, supporting near-real-time intraoperative use. The reported inference time covers forward registration after QvQ_v and MvM_v are available, excluding depth and mask reconstruction. Figure 3: Qualitative comparison on intraoperative keyframes. Col. 1: input. Cols. 2–5: baselines. Col. 6: Vis2Reg point cloud. Col. 7: Vis2Reg mesh overlay. Col. 8: visible-region fitting. Blue: preoperative; purple: intraoperative. Table 2: Ablations on real data. “With gate” refers to Uv=D^v>0∧Mv=1U_v=\ D_v>0 M_v=1\, while “without gate” uses Uv=D^v>0U_v=\ D_v>0\. Variant ℒvisL_vis gate 1-way ℒ3DL_3D seed Dice↑ CD↓ (m) w/o ℒvisL_vis N Y Y Y 74.29 3.12 w/o mask-gating Y N Y Y 79.57 3.04 sym. ℒ3DL_3D Y Y N Y 84.48 2.98 weak rigid init Y Y Y N 69.32 3.64 Full (Vis2Reg) Y Y Y Y 92.60 1.43 Ablation results in Table 2 validate each design choice: removing ℒvisL_vis or mask-gating, replacing one-way ℒ3DL_3D with symmetric Chamfer, and weakening rigid initialization each degrade performance—the last causing the largest drop—supporting mask-gated visible-domain supervision, the observation-to-model formulation, and a strong rigid seed. 4 Conclusion Vis2Reg is a visibility-aware, landmark-free, geometry-only framework for intraoperative liver 3D–2D registration under severe occlusion, coupling robust rigid initialization with an implicit non-rigid deformation field and differentiable point rasterization for mask- and camera-based self-supervision without 3D ground truth. On real intraoperative data, it significantly improves Dice and Chamfer Distance over prior methods while retaining near-real-time performance. We restrict our evaluation and claims to the P2I-LReg benchmark. Evaluating tumour and vascular structures would require internal anatomical ground truth [20], unavailable here, and is left to future work. credits 4.0.1 Acknowledgements The project was funded by the Engineering and Physical Sciences Research Council [Grant No. UKRI914]. 4.0.2 The authors have no competing interests to declare that are relevant to the content of this article. References [1] S. Ali, Y. Espinel, Y. Jin, P. Liu, B. Güttner, X. Zhang, L. Zhang, T. Dowrick, M. J. Clarkson, S. Xiao, Y. Wu, Y. Yang, L. Zhu, D. Sun, L. Li, M. Pfeiffer, S. Farid, L. Maier-Hein, E. Buc, and A. Bartoli (2025) An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI 2022 challenge. Medical Image Analysis 99, p. 103371. External Links: Document Cited by: §1, §1. [2] P. J. Besl and N. D. McKay (1992) A method for registration of 3D shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence 14 (2), p. 239–256. External Links: Document Cited by: §1. [3] Y. Dai, X. Yang, J. Hao, H. Luo, G. Mei, and F. Jia (2025) Preoperative and intraoperative laparoscopic liver surface registration using deep graph matching of representative overlapping points. International Journal of Computer Assisted Radiology and Surgery 20 (2), p. 269–278. External Links: Document Cited by: §1. [4] Y. Espinel, L. Calvet, K. Botros, E. Buc, C. Tilmant, and A. Bartoli (2022) Using multiple images and contours for deformable 3D-2D registration of a preoperative CT in laparoscopic liver surgery. International Journal of Computer Assisted Radiology and Surgery 17 (12), p. 2211–2219. External Links: Document Cited by: §1. [5] M. Feuerstein, T. Mussack, S. M. Heining, and N. Navab (2008) Intraoperative laparoscope augmentation for port placement and resection planning in minimally invasive liver resection. IEEE Transactions on Medical Imaging 27 (3), p. 355–369. External Links: Document Cited by: §1. [6] P. Guan, H. Luo, J. Guo, Y. Zhang, and F. Jia (2023) Intraoperative laparoscopic liver surface registration with preoperative CT using mixing features and overlapping region masks. International Journal of Computer Assisted Radiology and Surgery 18 (8), p. 1521–1531. External Links: Document Cited by: §1. [7] B. Huang, X. Yang, and F. Jia (2025) A landmark-free 3D–2D rigid liver registration via point cloud matching for laparoscopic surgery. Healthcare Technology Letters 12 (1), p. e70030. External Links: Document Cited by: §1. [8] S. Huang, Z. Gojcic, M. Usvyatsov, A. Wieser, and K. Schindler (2021) PREDATOR: registration of 3D point clouds with low overlap. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 4267–4276. External Links: Document Cited by: §1. [9] J. Johnson, N. Ravi, J. Reizenstein, D. Novotny, S. Tulsiani, C. Lassner, and S. Branson (2020) Accelerating 3D deep learning with PyTorch3D. In SIGGRAPH Asia 2020 Courses, p. 10. External Links: Document Cited by: §1. [10] B. Koo, M. R. Robu, M. Allam, M. Pfeiffer, S. Thompson, K. Gurusamy, B. Davidson, S. Speidel, D. Hawkes, D. Stoyanov, and M. J. Clarkson (2022) Automatic, global registration in laparoscopic liver surgery. International Journal of Computer Assisted Radiology and Surgery 17 (1), p. 167–176. External Links: Document Cited by: §1, §1. [11] L. Maier-Hein, P. Mountney, A. Bartoli, H. Elhawary, D. Elson, A. Groch, A. Kolb, M. Rodrigues, J. Sorger, S. Speidel, and D. Stoyanov (2013) Optical techniques for 3D surface reconstruction in computer-assisted laparoscopic surgery. Medical Image Analysis 17 (8), p. 974–996. External Links: Document Cited by: §1, §1. [12] I. Mhiri, D. Pizarro, and A. Bartoli (2025) Neural patient-specific 3D-2D registration in laparoscopic liver resection. International Journal of Computer Assisted Radiology and Surgery 20 (1), p. 57–64. External Links: Document Cited by: §1. [13] A. Neri, V. Penza, C. Baldini, and L. S. Mattos (2025) Surgical augmented reality registration methods: a review from traditional to deep learning approaches. Computerized Medical Imaging and Graphics 124, p. 102616. External Links: Document Cited by: §1. [14] A. Neri, V. Penza, N. Haouchine, and L. S. Mattos (2025) Benchmarking complete-to-partial point cloud registration techniques for laparoscopic surgery. Frontiers in Robotics and AI 12, p. 1702360. External Links: Document Cited by: §1. [15] M. Ovsjanikov, M. Ben-Chen, J. Solomon, A. Butscher, and L. Guibas (2012) Functional maps: a flexible representation of maps between shapes. ACM Transactions on Graphics 31 (4), p. 30:1–30:11. External Links: Document Cited by: §1. [16] E. Özgür, B. Koo, B. Le Roy, E. Buc, and A. Bartoli (2018) Preoperative liver registration for augmented monocular laparoscopy using backward-forward biomechanical simulation. International Journal of Computer Assisted Radiology and Surgery 13 (10), p. 1629–1640. External Links: Document Cited by: §1. [17] G. A. Prevost, B. Eigl, I. Paolucci, T. Rudolph, M. Peterhans, S. Weber, G. Beldi, D. Candinas, and A. Lachenmayer (2020) Efficiency, accuracy and clinical applicability of a new image-guided surgery system in 3D laparoscopic liver surgery. Journal of Gastrointestinal Surgery 24 (10), p. 2251–2258. External Links: Document Cited by: §1. [18] S. Prokudin, Q. Ma, M. Raafat, J. Valentin, and S. Tang (2023) Dynamic point fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), p. 7964–7976. Cited by: Table 1. [19] Z. Qin, H. Yu, C. Wang, Y. Guo, Y. Peng, S. Ilic, D. Hu, and K. Xu (2023) GeoTransformer: fast and robust point cloud registration with geometric transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (8), p. 9806–9821. External Links: Document Cited by: §1, Table 1. [20] N. Rabbani, L. Calvet, Y. Espinel, B. Le Roy, M. Ribeiro, E. Buc, and A. Bartoli (2022) A methodology and clinical dataset with ground-truth to evaluate registration accuracy quantitatively in computer-assisted laparoscopic liver resection. Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization 10 (4), p. 441–450. External Links: Document Cited by: §4. [21] J. Ramalhinho, S. Yoo, T. Dowrick, B. Koo, M. Somasundaram, K. Gurusamy, D. J. Hawkes, B. Davidson, A. Blandford, and M. J. Clarkson (2023) The value of Augmented Reality in surgery—a usability study on laparoscopic liver surgery. Medical Image Analysis 90, p. 102943. External Links: Document Cited by: §1. [22] M. R. Robu, J. Ramalhinho, S. Thompson, K. Gurusamy, B. Davidson, D. Hawkes, D. Stoyanov, and M. J. Clarkson (2018) Global rigid registration of CT to video in laparoscopic liver surgery. International Journal of Computer Assisted Radiology and Surgery 13 (6), p. 947–956. External Links: Document Cited by: §1, §1. [23] Y. Wang, Y. Sun, Z. Liu, S. E. Sarma, M. M. Bronstein, and J. M. Solomon (2019) Dynamic graph CNN for learning on point clouds. ACM Transactions on Graphics 38 (5), p. 146:1–146:12. External Links: Document Cited by: §2.2. [24] L. Yang, B. Kang, Z. Huang, X. Xu, J. Feng, and H. Zhao (2024) Depth anything: unleashing the power of large-scale unlabeled data. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 10371–10381. External Links: Document Cited by: §2.1. [25] M. Zhao, J. Jiang, L. Ma, S. Xin, G. Meng, and D. Yan (2024) Correspondence-free non-rigid point set registration using unsupervised clustering analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 21199–21208. External Links: Document Cited by: Table 1. [26] J. Zhou, B. Gao, K. Wang, J. Pei, P. Heng, and J. Qin (2025) Landmark-free preoperative-to-intraoperative registration in laparoscopic liver resection. IEEE Transactions on Medical Imaging 44 (11), p. 4350–4362. External Links: Document Cited by: §1, §2.3, §3.1, §3.2, Table 1.