Paper deep dive
Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery
Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao, Keshen Lyu, Song Zhou, Yimeng Chen, Haorui Wang, Qingmin Feng, Shenchao Shi, Huan Zhao, Wenbin Chen, Caihua Xiong, Chidan Wan, Jing Samantha Pan, Xiong Cai, Han Ding
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical intent is highly specialized and tacit. Here, we introduce DiffeoAfford, an action-grounded tissue affordance framework that retrospectively derives visual attention supervision from completed surgical procedures. By combining diffeomorphism-constrained tissue tracking with instrument trajectory analysis, DiffeoAfford generates affordance hotspot labels without manual per-frame annotation. A real-time prediction model trained on these labels anticipates relevant surgical regions and enables AffordView, an assistive auto-framing system for laparoscopic visualization. The proposed framework aligns with expert annotations and intraoperative surgeon gaze, and reduces surgeon cognitive workload during real-world evaluations using subjective, physiological, and behavioral measures.
Tags
Links
- Source: https://arxiv.org/abs/2608.02471v1
- Canonical: https://arxiv.org/abs/2608.02471v1
Trouble viewing inline? Open PDF directly β
Full Text
114,449 characters extracted from source content.
Expand or collapse full text
Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery Jiayu Gu 1,* , Yiwei Wang 2,5,* , Jie Zhang 2,* , Guojun Cao 1,* , Keshen Lyu 1 , Song Zhou 2 , Yimeng Chen 1 , Haorui Wang 3 , Qingmin Feng 4 , Shenchao Shi 6 , Huan Zhao 2 , Wenbin Chen 2,5 , Caihua Xiong 2,5 , Chidan Wan 1 , Jing Samantha Pan 3 , Xiong Cai 1 , Han Ding 2 Affiliations: 1. Department of Hepatobiliary Surgery, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China 2. School of Mechanical Science and Engineering, Huazhong University of Science and Technology, Wuhan, China 3. Department of Psychology, Sun Yat-sen University, Guangzhou, China 4. Institute of Biomedical Engineering, Union Hospital, Tongji Medical College, Huazhong University of Science and Technology, Wuhan, China 5. Institute of Medical Equipment Science and Engineering, Huazhong University of Science and Technology, Wuhan, China 6. Department of Hepatobiliary and Pancreatic Surgery, Renmin Hospital of Wuhan University, Wuhan, China Equal Contribution: * These authors contributed equally: Jiayu Gu, Yiwei Wang, Jie Zhang, Guojun Cao. Corresponding authors: Han Ding, Ph.D. Address: No. 1037 Luoyu Road, Wuhan 430022, Hubei Province, China Telephone/Fax number: 86-13886199756 E- mail: dinghan@hust.edu.cn Xiong Cai, MD, Ph.D. Address: No. 1277 Jiefang Ave, Wuhan 430022, Hubei Province, China Telephone/Fax number: 86-15002748078 E- mail: caixiong@hust.edu.cn Jing Samantha Pan, Ph.D. Address: No. 132 Waihuan Donglu, Guangzhou 510006, Guangdong Province, China Telephone/Fax number: 86-18922370280 E- mail: Panj27@mail.sysu.edu.cn Abstract Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that only experts can produce, and that expertise is tacit: surgeons converge on the same interaction loci without being able to state the rules they use. Here we show that such labels can be recovered from completed procedures rather than annotated. DiffeoAfford grounds tissue affordance retrospectively, propagating recorded instrument contacts across deforming tissue using diffeomorphism- constrained tracking, and produces labels that match the accuracy of annotators given full procedural context. A model trained on these soft labels predicts affordance in real time, reaching 95.16% directional consistency with subsequent camera motion and, though never trained on gaze, aligning with intraoperative surgeon gaze more closely, in both time and space, than the camera assistant. Across 12 paired laparoscopic cholecystectomies, the resulting auto-framing application, AffordView, lowered surgeon cognitive workload on converging subjective, physiological, and behavioral measures. Deriving supervision from action rather than annotation offers a scalable route to anticipatory assistance in the operating room. Keywords Laparoscopic Surgery, Laparoscopic Self-regulation, Surgical Data Science, Visual Attention Modeling, Tissue Affordance Introduction The American Medical Association (AMA) conceptualizes augmented intelligence as a framework that amplifies, rather than replaces, human cognitive capabilities 1,2 . Advancing digital medicine from passive data analysis to proactive augmented intelligence requires artificial systems capable of overcoming sensory information overload 3 . Such systems must emulate biological visual attention to anticipate user intent, selectively allocate computational resources to functionally relevant spatial areas, and suppress background distractors 4,5 . Consequently, modeling visual attention and predicting regions of interest (ROI) constitute foundational objectives for bridging the semantic gap between raw sensory data and high- level cognitive reasoning 6,7 . Establishing this perceptual foundation facilitates the development of diverse augmented intelligence technologies, including diagnostic image segmentation, context-aware augmented reality, and human-robot collaboration 8-11 . Integrating visual attention modeling with surgical data science positions laparoscopic surgery as a primary domain for advanced intelligent assistive applications 3,12,13 . However, transitioning these technologies to the operating room reveals significant technical barriers 14,15 . Although the degradation of laparoscopic video by adverse optical conditions was once considered the primary barrier to computational modeling, modern computer vision models can overcome such visual noise given sufficient annotated data 14-16 . Thus, the definitive obstacle in the surgical setting is not the appearance of the image, but the availability of the label. While natural image datasets benefit from the scalable crowdsourcing of nameable objects with defined boundaries, manual annotation in the surgical domain faces profound, domain-specific barriers. In laparoscopy, the spatial ROI is not a static anatomical structure, but a dynamic target dictated by functional surgical intent. Consequently, identifying this target requires highly specialized procedural knowledge. Crucially, this expert judgment is intrinsically tacit; while experienced surgeons reliably converge on consensus interaction loci, they cannot articulate explicit visual or geometric rules for their selection. Compounded by the continuous deformation of soft tissues and indistinct anatomical boundaries, this reliance on tacit knowledge precludes the employment of lay annotators or explicitly coded heuristic models. Reducing the number of required labels, as in one-shot affordance grounding of deformable objects from a single annotated exemplar per category 17 , does not remove this barrier, because such approaches presuppose a nameable part inventory that can be labeled without specialist knowledge. Therefore, generating dense, per-frame annotations across long surgical videos depends entirely on a scarce pool of expert surgeons, rendering manual supervision prohibitively expensive and fundamentally unscalable 13,14,18 . Hence, we propose that the necessary supervision for computational visual attention models can be derived from a source other than manual annotation. Because completed procedures provide an implicit, retrospective record of spatial priority, the recorded instrument trajectories of the surgeon can retroactively ground the functional ROI. This hindsight-oriented approach motivates a conceptual shift from static anatomical landmark identification toward an action-grounded interpretation of surgical intent 19,20 . To formalize this action-grounded interpretation of intent, we introduce tissue affordance prediction as a computational framework for intention-aware visual attention (Fig. 1). Rooted in the ecological perspective that vision guides action, effective perception is thought to prioritize interaction loci that afford manipulation rather than relying on exhaustive semantic labeling 21 . Empirical evidence demonstrates that expert surgeons employ consistent visuomotor strategies, fixating on shared affordance regions prior to executing movements 22,23 . This anticipatory pattern is consistent with the quiet-eye phenomenon documented in surgical and other motor expertise, in which a stable fixation precedes and predicts skilled action 22 . If expert visual attention is substantially shaped by actionability rather than instrument position alone, anticipating that attention plausibly requires predictive affordance detection rather than reactive tool tracking 22,23 . Translating these theoretical visuospatial principles into computational models requires overcoming the limitations of conventional approaches and the prohibitive manual annotation bottleneck inherent to deformable soft tissues 24,25 . To address these complexities, we formulated a manual annotation-free framework to model tissue affordance. Within this overarching architecture, we developed DiffeoAfford (Diffeomorphic Affordance Grounding), an affordance grounding pipeline that employs a hindsight strategy to retrospectively deduce surgical interaction loci (Fig. 2). This pipeline integrates global transformations and diffeomorphic deformations to map historical instrument tips across every single frame, automatically generating dense, continuous soft labels of tissue affordance, successfully bypassing the reliance on expensive, per-frame expert annotation. These automatically generated soft labels define spatial target regions through accumulated actionability instead of rigid anatomical categories. To convert this retrospectively grounded data into proactive augmented intelligence, we subsequently trained a deep learning model to predict real-time affordance hotspots from the surgical scene. We validated grounding against expert consensus on public data and against intraoperative gaze in LCET, a proprietary eye-tracking dataset of laparoscopic cholecystectomies, then benchmarked prediction against an established reactive instrument-tracking baseline. LCET provides two things public corpora do not: because the grounded labels derive from instrument trajectories, recorded gaze serves as an independent behavioral criterion rather than a restatement of the training signal, and the subset with synchronized surgeon and camera-assistant gaze allows anticipation to be measured against a human camera operator. To translate visual attention modeling into clinical utility, DiffeoAfford underpins a proactive spatial framing application that embodies applied augmented intelligence technology to mitigate surgeon cognitive overload (Fig. 1). Because displaced surgical targets conflict with intrinsic physiological attractors, they may exacerbate extraneous cognitive load and contribute to the prefrontal deactivation associated with executive decline 26,27 . To mitigate this cognitive exertion, we engineered AffordView, an anticipatory auto-framing application that provides ergonomically optimized visualization inspired by participant auto-framing technologies utilized in consumer video conferencing 28 . In current practice this alignment is maintained by the camera assistant, who must infer the surgeonβs target from the same image while operating the scope, and whose misreading surface only when the surgeon corrects them verbally. AffordView addresses that dependency by deriving the framing target from predicted tissue affordance rather than from a second personβs reading of intent. Proactively centering the surgical interaction locus harmonizes the digital interface with human visual-motor constraints, aligning the operative field with the optimal viewing position 20,29 . This spatial alignment reduces the neural cost of compensatory mental rotation and is compatible with the quiet eye strategies associated with expert visuomotor coordination 30 . The capacity of the AffordView application to systematically mitigate cognitive exertion was comprehensively verified during paired surgical sessions using subjective workload questionnaires, continuous electroencephalography, and pupillometry metrics 31 . Fig. 1 | Conceptual overview of the study. a, Concept & Clinical Gap: Suboptimal field-of-view selection elevates the surgeon's cognitive workload, which is mitigated through affordance-guided view adjustment. b, Core Methodology: DiffeoAfford, a pipeline that grounds affordances in hindsight via diffeomorphism- constrained tissue tracking to link future actions to past affordances. c, Real-time Deployment: AffordView, an application that predicts affordance hotspots to auto-frame the intraoperative view, providing an assistive view. Fig. 2 | Overview of the proposed affordance-guided view control framework. The framework comprises DiffeoAfford pipeline for automated annotation (a), an affordance prediction model and the AffordView application for auto-framing (b). a, DiffeoAfford pipeline for affordance hotspot dataset generation. The SAM2 is utilized to segment instruments and derive instrument tips and lengths. To address soft-tissue deformation, DiffeoAfford performs diffeomorphism-constrained tissue tracking by combining global transformations with local deformation fields to establish dense pixel-wise correspondences across frames. Instrument tip points are aggregated onto a query frame, propagated to all other frames via displacement fields, filtered, and fitted to a 2D Gaussian distribution to generate AH heatmap labels. Video clips are then automatically segmented and refined based on instrument length and AH visibility using defined temporal thresholds (ν ν , ν ν ). b, Real-time affordance hotspot prediction model and the AffordView application. A Segformer model, trained on the generated AH dataset, predicts AH heatmaps in real-time. The smoothed centroid of the predicted heatmap serves as the robust target for the application, which dynamically crops and upscales the original video to maintain the view on the relevant surgical interaction zone, overcoming the limitations of baseline instrument tracking. AH, affordance hotspot; SAM2, Segment Anything Model 2; FoV, field of view; Global Trans. & Diffeo. Def., global transformation and diffeomorphic deformation; AH GR. & Temp. Ref., AH grounding and temporal refinement. Results We first confirmed that surgeonsβ gaze exhibits a pronounced center bias on laparoscopic monitors, establishing the physiological rationale for center-aligned framing. Building on this finding, we conducted a series of studies organized in two parts (Fig. 3): First, dataset-based evaluations assessed both affordance grounding and prediction. DiffeoAfford grounding was evaluated against expert consensus on Cholec80 and intraoperative surgeon gaze on LCET, whereas AH prediction was evaluated against synchronized surgeon and camera-assistant gaze on LCET and against subsequent camera motion on AutoLaparo. Second, the real-world utility of AffordView was assessed through a multicenter surgeon-preference survey and an intraoperative user-experience study incorporating subjective and objective measures of cognitive workload. Fig. 3 | Evaluation and assessment setup. Overview of the study design for evaluation of the proposed method and assessment of the application. Center bias in laparoscopic surgery To investigate whether the visual center bias reported in previous study extends to the laparoscopic environment, we analyzed gaze data from our laparoscopic cholecystectomy eye-tracking (LCET) dataset (ν=26 procedures). Surgeons' gaze on laparoscopic monitor demonstrated a distinct center bias, as quantified by the Mahalanobis distance (ν· ν =0.05) between the mean fixation point and the geometric center of the screen (Supplementary Fig. 1). Evaluation of DiffeoAfford on Cholec80 dataset Having established the need to center the FoV, we first required a reliable method to automatically identify the surgical affordance hotspots (AHs). To validate whether the automatically generated AHs align with human expertise, we evaluated DiffeoAfford on the public Cholec80 dataset. We extracted procedural clips from the Calot triangle dissection phase and recruited three groups of annotators: senior surgeons, junior doctors, and medical students (novices). Senior surgeon annotations served as the expert ground truth. Inter-annotator consistency: We measured within-group consistency among human annotators using intraclass correlation coefficients (ICC (3, k)) (Fig. 4a). The consistency increased significantly after annotators watched the video clips (hereafter referred to as the 'In-context' condition), indicating that procedural context provides information beneficial for target recognition. Specifically, for the X and Y coordinates, novice consistency improved from 0.86 (95% CI: 0.81-0.89) and 0.77 (95% CI: 0.70-0.83) in the static condition to 0.97 (95% CI: 0.96-0.98) and 0.93 (95% CI: 0.90-0.95) in the In-context condition. This post-context performance closely paralleled the senior surgeons' consensus of 0.97 (95% CI: 0.95- 0.97) for X and 0.93 (95% CI: 0.91-0.95) for Y. Annotation accuracy: Using senior surgeons' annotations as ground truth, we measured annotation error via Euclidean distance (Fig. 4b). To assess statistical differences, we employed linear mixed models (LMMs) with the Novice (In-context) group as the reference (Fig. 4c). The image center baseline exhibited an estimated error difference of 45.52 pixels (95% CI: 36.85-55.08, ν<0.001). Novices who viewed only static images showed a significant deviation from the reference (difference = 35.27 pixels, 95% CI: 29.74- 41.20, ν<0.001), but their performance improved markedly after viewing clips. Notably, DiffeoAfford reproduced this context-aware performance. Its annotation error did not differ significantly from that of the Novice (In-context) group in the LMM (difference = -1.78 pixels, 95% CI: -5.76-2.68, ν=0.418). A median-based bootstrap equivalence analysis further demonstrated that DiffeoAfford was equivalent to the Novice (In-context) group (ν =0.070, 90% CI, -0.259 to 0.322), with the entire CI contained within the equivalence interval from -1 to 1 (Supplementary Fig. 2c). These results indicate that knowledge about the target is embedded in the procedural context, and DiffeoAfford-grounded AHs from this context successfully capture this expert knowledge. Improved affordance-grounding accuracy over global-transform baselines on Cholec80 Previous affordance-grounding methods for natural-scene videos commonly propagate interaction points across frames using global transformations 32,33 . Following this strategy, we constructed similarity- transform and homography baselines (fitted by RANSAC; hereafter similarity RANSAC and homography RANSAC) while keeping the tracking points, aggregation, propagation, and refinement procedures identical to DiffeoAfford. Across 141 paired images, DiffeoAfford achieved a lower median localization error (32.38 pixels) than both similarity RANSAC (39.40 pixels) and homography RANSAC (35.07 pixels) (Supplementary Fig. 3). Both differences were statistically significant (two-sided paired Wilcoxon signed-rank tests; similarity RANSAC, ν<0.001; homography RANSAC, ν=0.003). Evaluation of DiffeoAfford against intraoperative gaze on LCET dataset While offline expert consensus provides a strong foundation, intraoperative gaze reflects a surgeon's immediate intent and has therefore been widely used to guide FoV control 34-36 . To evaluate how well the generated AHs align with intraoperative intent, we used intraoperative surgeonsβ gaze points as ground truth. To mitigate potential bias arising from the arbitrary selection of gaze time windows, we calculated the Euclidean distance between these actual gaze points and the annotations across three distinct temporal windows (Β±0.5 s, Β±1.0 s, and Β±2.0 s) centered around the target frame (Fig. 4d). Statistical differences were assessed using LMMs following a Box-Cox transformation to satisfy residual normality. Using the Junior (In-context) group as the reference, we found that annotations of DiffeoAfford reached comparable accuracy across all three windows. No statistically significant differences were observed between DiffeoAfford and the In-context juniors: the estimated difference of 5.40 pixels at Β±0.5 s (95% CI: -0.60-11.83, ν=0.079), 3.49 pixels at Β±1.0 s (95% CI: -2.58-10.00, ν=0.267), and -0.70 pixels at Β±2.0 s (95% CI: -6.59-5.60, ν=0.823). The median-based bootstrap equivalence analysis likewise demonstrated equivalence between DiffeoAfford and the Junior (In-context) group across all three windows: ν =0.116 (90% CI, -0.013 to 0.274) at Β±0.5 s, ν =0.059 (90% CI, -0.018 to 0.228) at Β±1.0 s, and ν =β0.009 (90% CI, -0.132 to 0.183) at Β±2.0 s (Supplementary Fig. 4g). This indicates that DiffeoAfford-grounded AHs aligns with where surgeons intend to interact and can serve as a proxy for visual attention. The observation aligns with evidence that while free viewing is driven by general informativeness, attention in specific tasks becomes directed toward object parts that afford interaction 37,38 . Fig. 4 | Evaluation of affordance grounding across surgical datasets. a-c, DiffeoAfford on the Cholec80 dataset (ν=141 images). a, Inter-annotator consistency for X and Y coordinates across participant groups, quantified by the ICC. Error bars denote 95% CI. b, Annotation error (in pixels) relative to the expert ground truth for the human annotator groups, the image center baseline, and DiffeoAfford. c, Forest plot of the estimated inter-group differences in annotation error relative to the Novice (In-context) reference group, presented as estimated differences Β± 95% CI from linear mixed models (LMMs) fitted after a Box-Cox transformation to satisfy the assumption of residual normality. d, Annotation error relative to surgeonsβ intraoperative gaze points on the LCET dataset (ν=164 images). The violin plots display distances between annotations and the gaze points, which serve as the ground truth, across temporal windows (Β±0.5 s, Β±1.0 s, and Β±2.0 s) centered around the target frame; statistical differences were calculated using LMMs following a Box-Cox transformation. Inter-group statistical comparisons and LMM residual diagnostics are provided in Supplementary Fig. 2a, b (Cholec80) and Supplementary Fig. 4a-f (LCET); median-based bootstrap equivalence analyses of DiffeoAfford versus Novice (In-context) and versus Junior (In-context) are shown in Supplementary Fig. 2c and Supplementary Fig. 4g, respectively. For all violin plots (b, d), the internal box limits indicate the interquartile range (IQR), orange lines denote the median, and whiskers extend to a maximum of 1.5 Γ IQR, with circles representing individual outliers. ν values were adjusted for multiple comparisons using the Benjamini-Hochberg procedure; ***, ν<0.001 ; n.s., not significant. ICC, intraclass correlation coefficient; CI, confidence interval; LCET, laparoscopic cholecystectomy eye-tracking; LMM, linear mixed model. Evaluation of AH prediction against intraoperative gaze on the LCET dataset Temporal alignment: Having established that DiffeoAfford-grounded AHs were spatially aligned with intraoperative surgeon gaze, we next examined the temporal relationship between real-time AH prediction and surgeon visual attention. The AH prediction model, trained on Cholec80 using AH annotations generated by DiffeoAfford, was applied to 13 LCET procedures with synchronized surgeon and camera assistant eye tracking, and the sequence of raw predicted AH centroids was compared with surgeon gaze using diagonal cross-recurrence profiles (DCRPs) during both Calot triangle dissection and gallbladder dissection, with camera-assistant gaze serving as a human reference sequence (Fig. 5a). The group-mean DCRP for AH prediction peaked on the candidate-leading side, at -0.20 s during Calot triangle dissection and -0.12 s during gallbladder dissection, whereas camera-assistant gaze peaked on the candidate-lagging side at 0.48 s in both phases. At the procedure level, the median centroid lags for AH prediction were -0.033 s and 0.013 s in the two phases, compared with 0.185 s and 0.204 s for camera-assistant gaze (ν=0.014 ννν 0.014). Thus, the centroid lag of AH prediction relative to surgeon gaze was significantly shorter than that of camera-assistant gaze in both phases. No consistent temporal precedence was observed between AH prediction and surgeon gaze. Spatial alignment: We further compared the spatial distances from surgeon gaze to AH prediction, camera-assistant gaze, and the image center at simultaneous valid timestamps, summarized as a case- level median distance for each phase and comparison. The median case-level advantage of AH prediction over camera-assistant gaze was 55.5 pixels during Calot triangle dissection and 28.4 pixels during gallbladder dissection (ν=0.0093 ννν 0.0479); relative to the image center, the median advantages were 60.8 and 99.1 pixels (ν=0.0015 ννν 0.0061) (Fig. 5b). AH prediction was thus significantly closer to surgeon gaze than either camera-assistant gaze or the image center in both surgical phases. Evaluation of AH prediction for camera motion forecasting on AutoLaparo dataset Building on successful validation of DiffeoAfford in cholecystectomy, we sought to evaluate the technique across a broader range of surgical procedures and evaluate its practical utility for proactive view control. To achieve this, we utilized the AutoLaparo dataset, which comprises laparoscopic hysterectomy videos. We trained the real-time prediction model using DiffeoAfford-grounded AHs from Task 1 videos, and subsequently evaluated its performance on Task 2 (ν=124 camera motion events across four directions: up, down, left, and right), which is specifically designed for evaluating camera motion prediction. During this evaluation, we predicted AH locations prior to camera motion events and compared the predictive performance against an instrument-tracking baseline (detailed in the Supplementary Information). Prediction accuracy: Consistent with previous work 39 , we analyzed the consistency between the AH location and camera motion within the image coordinate system across four directions: up, down, left, and right. Coordinates were normalized using the image center as the origin, with half the image height and width set to 1. The AH locations predicted immediately before camera motion exhibited strong consistency with the subsequent camera movement, with directional consistency of 95.16% (Fig. 5c). This confirms that the predicted AH effectively anticipates the target driving the view adjustment. Superior foresight: We compared the model with the instrument-tracking target prediction serving as the baseline. As shown in Fig. 5d, instrument-tracking method achieved equivalent accuracy near the motion event, but its accuracy degraded rapidly as the temporal gap increased. Specific analysis revealed that instrument tips are often already near the target at the moment preceding camera adjustments, accounting for the short-term advantage; however, this assumption no longer holds at earlier time points. In contrast, the AH prediction maintained stable accuracy over longer horizons, offering superior foresight and robustness. Representative examples illustrating the differences between the two approaches are shown in Fig. 5e. Fig. 5 | Evaluation of AH prediction across surgical datasets. a, b, Evaluation of the Cholec80-trained AH prediction model against intraoperative surgeon gaze in 13 LCET procedures with synchronized surgeon and camera-assistant eye tracking, during Calot triangle dissection and gallbladder dissection. a, Temporal alignment. Left, group-mean diagonal cross-recurrence profiles (DCRPs) obtained by comparing surgeon gaze at time ν‘ with each candidate sequence at time ν‘ + ν; negative lags indicate that the candidate sequence leads surgeon gaze, whereas positive lags indicate that it lags surgeon gaze. Solid lines denote the mean recurrence rate across procedures, shaded bands denote mean Β± 1.96 SEM, and colored dotted vertical lines mark the peaks of the group-mean DCRPs. Right, procedure-level centroid lags, with points from the same procedure connected and horizontal bars indicating the medians; negative values indicate that the candidate tends to lead surgeon gaze, positive values indicate that it tends to lag surgeon gaze, and values near zero indicate temporal coupling without a consistent lead or lag. b, Spatial alignment. Each point represents one procedure, plotting the median distance from AH prediction to surgeon gaze (vertical axis) against the corresponding median distance from camera-assistant gaze or the image center to surgeon gaze (horizontal axis). The diagonal line denotes equal distances; points below the line indicate that AH prediction was closer to surgeon gaze. Median advantage denotes the median across procedures of the comparatorβs case-level median distance minus the corresponding AH-prediction median distance. For a and b, exact two-sided paired Wilcoxon signed-rank tests were applied to procedure-level values. c-e, Evaluation of AH prediction against subsequent camera motion in the AutoLaparo dataset (ν=124 clips with camera motion). c, Distribution of the normalized predicted AH X and Y coordinates at the onset of camera motion, categorized by the subsequent camera movement direction (up, down, left, right); the internal box limits indicate the interquartile range (IQR), orange lines denote the median, and whiskers extend to a maximum of 1.5 Γ IQR, with circles representing individual outliers. d, Directional consistency between the target (using the AH prediction model versus the baseline instrument-tracking method) and the subsequent camera motion over time, relative to the camera motion onset at the 5-second mark. e, Representative cases comparing instrument-tracking targets, AH prediction heatmaps, and subsequent camera-motion directions. Heatmap intensities were normalized independently to the maximum value in each panel for visualization. *, ν<0.05 ; **, ν<0.01 ; LCET, laparoscopic cholecystectomy eye- tracking; AH, affordance hotspot; DCRP, diagonal cross-recurrence profile; IQR, interquartile range; SEM, standard error of the mean. Surgeon preference assessment of AffordView application Following the validation of the model's predictive accuracy, we integrated it into our AffordView application to assess its clinical acceptability. We conducted a multicenter questionnaire survey where participants assessed paired laparoscopic video clips sourced from the Cholec80 dataset across two comparisons: (1) the original video versus the AffordView application using our proposed AH prediction model, and (2) the AffordView application using the proposed AH prediction model versus the instrument-tracking baseline. Out of 32 distributed questionnaires, 28 valid responses were collected (an 87.5% response rate) from 21 tertiary-care hospitals. The respondents comprised a diverse range of surgical experience, including 7 junior doctors, 8 intermediate attending surgeons, 8 associate senior surgeons, and 5 senior surgeons (Fig. 6a). As shown in Fig. 6b, the proposed method was significantly preferred over the original unadjusted view for FoV accuracy, favored in 77.1% (216/280) of the paired comparisons (Odds Ratio [OR] = 3.38, ν<0.001, evaluated via exact binomial test). However, no clear advantages were observed in FoV stability (48.6% vs 51.4%, OR = 0.94, ν=0.676) or zoom rationality (53.6% vs 46.4%, OR = 1.15, ν=0.256). In Fig. 6c, when compared to the instrument-tracking baseline, the proposed method demonstrated clear superiority across all dimensions, outperforming it in accuracy (75.4% preference, OR = 3.06, ν<0.001), stability (75.7% preference, OR = 3.12, ν<0.001), and zoom rationality (61.8% preference, OR = 1.62, ν<0.001). The lack of advantage in zooming prompted the integration of an instrument-driven zoom module for the subsequent intraoperative evaluation. Fig. 6 | Subjective assessment of the FoV adjustment. a, Characteristics of the survey participants obtained from a multicenter questionnaire survey (ν=28 valid responses from 21 tertiary-care hospitals), illustrating the distribution of surgical experience in years (top) and professional titles (bottom). b, c, Surgeons' preference percentages comparing the FoV adjusted by the AH-prediction-driven AffordView against the original unadjusted view (b) and the baseline instrument-tracking-driven AffordView (c). Preferences were evaluated across three dimensions: zoom rationality, FoV stability, and FoV accuracy. The exact fraction of votes favoring the AH prediction method over the total comparisons is displayed within the bars. Statistical significance was evaluated via exact binomial tests (***,ν<0.001; n.s., not significant). Intraoperative user experience assessment of AffordView application To comprehensively assess the application's overall ergonomic impact in clinical practice, the AffordView application was deployed during real-world laparoscopic procedures ( Fig. 7). 24 laparoscopic cholecystectomies were performed by 5 operating surgeons and 12 novice camera assistants, comprising 12 matched pairs (each pair consisting of a control session and an experimental session). The real-time AH prediction model was specifically trained on data from the Calot triangle dissection (Calot) and Gallbladder dissection (Diss) phases, making these two phases the primary targets for the view-control intervention. Given the relatively small sample size inherent to such highly controlled, intraoperative dual-monitor evaluations (ν=12 pairs), exact permutation-based paired tests were utilized as the primary inference method to ensure robust, distribution-free hypothesis testing. Linear mixed models (LMMs) were additionally employed to further control for potential confounding effects. All 12 pairs yielded complete data for analysis, with no pairs excluded. Objective metrics: We assessed the surgeons' CWL and operational fluency through electroencephalography (EEG), eye-tracking, and verbal instructions. The results demonstrated surgical phase specificity, aligning with the targeted intervention phases. For the EEG analysis, the theta/alpha ratio (TAR) (Fig. 7a), a physiological marker of CWL 40-42 , decreased in the experimental group during both phases (Calot: Mean difference = -0.40, Cohenβs νν= β0.64, permutation ν=0.023; Diss: Mean difference = -0.42, Cohenβs νν=β0.86, ν=0.012). Relative to controls, occipital alpha power in the experimental group declined during the later surgical phases, with a significant reduction during Gallbladder dissection phase (Mean difference = -0.90, Cohenβs νν=β0.67, ν=0.043) (Fig. 7c). This occipital alpha suppression is consistent with heightened visual engagement rather than reduced arousal 43,44 . Conversely, no clear phase-specific pattern was observed for the EEG engagement index (EI) (Fig. 7b). As EI reflects general alertness and focus 45,46 , this suggests the intervention did not diminish the surgeons' clinical vigilance. Parallel improvements were observed in oculomotor parameters: the index of pupillary activity (IPA) 47 , a physiological marker of CWL, decreased during the two intervention phases (Fig. 7f) (Calot: Mean difference = -0.037, Cohenβs νν=β0.79, ν=0.021; Diss: Mean difference = -0.040, Cohenβs νν=β1.43, ν=0.002). The proportion of fixation time spent on the AffordView monitor also increased during these target phases, indicating that surgeons utilized the assistive FoV (Fig. 7e). Notably, while stationary gaze entropy (SGE) was observed to decrease in the experimental group during these two phases compared to the control group, neither SGE nor gaze transition entropy (GTE) exhibited significant reductions (Fig. 7g, 7h). We hypothesize that related to the dual-monitor experimental setup, which required surgeons to alternate their gaze between the original and adjusted screens, thereby inflating spatial and transitional gaze entropy. Furthermore, the number of verbal instructions issued to the camera assistant was lower in the experimental group during these two phases (Fig. 7d) (Calot: Mean difference = -2.92, Cohenβs νν= β1.20, ν=0.002; Diss: Mean difference = -1.83, Cohenβs νν=β0.90, ν=0.019), reflecting a reduced need for explicit coordination. Subjective metrics: The objective and behavioral findings were corroborated by the subjective assessments. Surgery Task Load Index (SURG-TLX) 48 scores were significantly lower in the experimental group (exact paired permutation test, ν=0.014), with the Mental Demands and Physical Demands subscales primarily contributing to this reduction (Fig. 8a, 8b). Surgeonsβ ratings of the applicationβs utility also showed higher scores during the phases when the application was active (Fig. 8c). Fig. 7 | Setup for the intraoperative user experience evaluation and objective metrics. The central illustration depicts the experimental configuration deployed in the operating room during paired laparoscopic cholecystectomies. The setup features a dual-monitor system displaying both the original view and auto-framed view (AffordView). For the objective assessment of the application's effectiveness, the surgeon is equipped with an EEG cap incorporating EMI shielding layers to reduce radiofrequency noise from energy devices, and a wearable eye-tracker that simultaneously records verbal instructions. a-c, Intraoperative EEG metrics comparing the Ctrl and Exp conditions across discrete surgical phases, specifically the TAR (a) and EI (b), as well as occipital alpha band power (c). d, The number of verbal instructions issued to the camera assistant in the Ctrl versus Exp groups. e-h, Quantitative evaluation of oculomotor parameters across surgical phases for the Ctrl and Exp groups. Metrics include the percentage of fixation duration on the AffordView (e), the IPA (f), SGE (g), and GTE (h). *, ν<0.05 ; **, ν<0.01 ; ***, ν<0.001 ; EEG, electroencephalography; EMI, electromagnetic interference; TAR, theta/alpha ratio; EI, engagement index; Ctrl, control; Exp, experimental; Prep, Preparation; Calot, Calot triangle dissection; Clip, Clipping and cutting; Diss, Gallbladder dissection; Pack, Gallbladder packaging; Clean, Cleaning and coagulation; Retr, Gallbladder retraction; IPA, index of pupillary activity; SGE, stationary gaze entropy; GTE, gaze transition entropy. Fig. 8 | Subjective metrics. a, b, Cognitive workload assessed via the SURG-TLX, presented as radar charts detailing the unweighted (a) and weighted (b) subscale scores. c, Surgeons' intraoperative utility ratings of the AffordView application across distinct surgical phases. SURG-TLX, Surgery Task Load Index. Discussion Hindsight-driven affordance as a supervision source for intention-aware visual attention modeling The proposed framework grounds tissue affordance retrospectively and, from these labels, learns to predict the spatial ROI from surgical context. Rather than passively estimating where the surgeon looks, this intention-aware modeling anticipates where the surgeon is about to act, bridging the conceptual gap between the reactive mechanics of explicit selection and the cognitive foresight of implicit prediction. Affordance grounding has also been extended from rigid objects 49 to deformable ones, but typically remains anchored to a stable part inventory: garment affordances are attached to named components such as a sleeve or a hem, so that one annotated exemplar transfers across instances of a category 17 . Soft tissue admits no comparable decomposition, since interaction loci follow procedural intent rather than a b c object identity, and the same region may afford different actions at different moments. Supervision can alternatively be generated by acting, as in deformable-object manipulation, where candidate actions are scored against a formalized task objective using interaction the system collects itself 50 . Surgical view control admits neither route, because exploratory interaction cannot be performed on a patient and the operative goal resists formalization as a measurable target state. Within surgery, dense spatial guidance has instead been supplied normatively, by annotating where dissection is permissible rather than where it occurs: Go/No-Go models render expert-defined safe and unsafe zones as heatmaps and have been benchmarked against panels of expert surgeons 51 . Such maps prescribe where action is allowed, whereas the affordance grounded here describes where expert action in fact took place, so the two answer different questions about the same scene. To address these complexities, the concept of tissue affordance has recently emerged. A closely related effort, AffordTissue, a concurrent preprint, likewise predicts dense tissue-affordance heatmaps anchored to tissue rather than to anatomical categories 52 . The two diverge less in machinery than in a prior question: where the spatial supervision for what affords interaction should come from, given that this intent is tacit and resists precise explicit articulation. AffordTissue supplies it externally, from expert per-frame annotations under a language specification of the surgery, tool, and action; our framework derives it internally, treating the surgeon's consummated trajectory as evidence for where interaction was afforded, without exhaustive manual labeling. Which source is appropriate follows from the downstream objective, in keeping with the relational nature of affordance 21 , which is defined over the agentβenvironment system rather than as a fixed property of tissue 53 . An externally specified, verifiable target suits a robotic safety constraint, whereas a behaviorally derived one suits the modeling of a surgeon's anticipatory attention for proactive view control, a setting in which soliciting real-time language would itself add to the workload. The two thus address distinct objectives and are complementary rather than competing. Within this formulation, affordance constitutes an epistemic prior rather than a static anatomical attribute: the evidence that a region affords manipulation accumulates over the unfolding operation. Retroactively analyzing the operative trajectory infers interaction goals and grounds hotspots specific to the scene, defining targets through accumulated actionability rather than static anatomical landmarks. This actionability-based definition is supported by the close alignment of the generated hotspots with true intraoperative surgeon gaze. Diffeomorphic deformation priors and anticipatory alignment for region of interest prediction DiffeoAfford models cross-frame correspondence with a diffeomorphic deformation field fitted jointly with a global transform. Prior natural-scene affordance grounding relies on a global transform alone, which captures dominant camera motion but has too few degrees of freedom for the non-rigid distortion of compliant tissue. Adding a diffeomorphic field supplies the deformation freedom while remaining smooth, invertible, and topology-preserving by construction, precluding the crossings and folds that would violate biological continuity. With the rest of the pipeline fixed, replacing this field with a global transform alone raised grounding error (Supplementary Fig. 3). Combined with the global transform handling macroscopic camera motion, it yields the dense, pixel-wise correspondence used to transport instrument-tip observations across a clip (Fig. 2). Conceptually, this aligns with recent surgical work that regularizes tissue deformation toward physical plausibility; Gong et al, for instance, use diffeomorphic constraints to recover non-crossing, realistic warps of manipulated tissue 54 . Our formulation differs in operating directly on the 2D image plane for affordance grounding, rather than recovering dense 3D geometry, which is not required here and is less suited to this setting, being computationally heavier and less accurate under instrument motion. The 2D choice follows from the viewing geometry: the laparoscope is typically oriented to face the working surface, and the out-of- plane component of surface displacement is comparatively small, so within a single maneuver the projected deformation is well approximated by a 2D diffeomorphism (Supplementary Information). Agreement between the grounded labels and human reference was established in the cholecystectomy evaluations. On Cholec80, the annotations were statistically equivalent to those of novice annotators given full procedural context (Supplementary Fig. 2c), and on LCET they aligned with intraoperative surgeon gaze (Fig. 4d). The pipeline also applied to hysterectomy video, where the grounded labels were sufficient to train an AutoLaparo-specific predictor. That experiment evaluated downstream prediction and camera-motion forecasting rather than grounding accuracy, so it supports the transferability of the grounding-to -prediction framework across procedures rather than grounding fidelity in a second procedure. The result is also consistent with the diffeomorphic prior encoding deformation behavior common to soft tissue rather than procedure-specific anatomy. That distinction parallels a documented signature of expertise: experts represent problems by deep structure where novices organize the same material by surface features 55 , and it is conceptual rather than purely procedural knowledge that supports flexible application to unfamiliar cases 56 . In surgery such knowledge is largely automated, which is why experts omit much of it when describing procedures they perform reliably 57 , and why the interaction loci grounded here were not available through rule elicitation. We do not claim that the diffeomorphic prior implements the representation surgeons use. The parallel is that both abstract over how soft tissue behaves rather than over what a particular procedure contains. Beyond spatial coherence, the framework addresses the temporal dimension of attention prediction by aligning model output with the anticipatory behavior of the surgeon. Anticipation in this sense should be distinguished from foresight as the term is used in deformable-object manipulation, where it denotes a long-horizon value over the agent's own prospective actions, learned so that states temporarily closer to the goal but harder to proceed from are avoided 50 . Here the future to be anticipated belongs to a different agent, and affordance grounding for deformable objects outside surgery is more commonly static, resolving where to act in the current frame without a temporal lead 17 . Within surgery, spatial attention has more often been modeled by regressing observed gaze, for example by jointly predicting scanpath and instrument masks with the explicit aim of informing camera guidance 58 . Deriving the target from action instead leaves gaze available as an independent criterion rather than as the training signal. Contemporary approaches model surgical temporal dynamics either by recognizing or anticipating surgical phases 59,60 and action triplets 12 , or by forecasting the timing of camera movements from instrument kinematics. The triplet formalism does carry an explicit interaction target, but in public benchmarks (e.g., CholecT50) that target is typically a semantic class, with no spatial annotation; any target localization is then recovered as a weakly supervised by-product of attention or class-activation mapping 12 , whose activations depend on the chosen query and need not coincide with the true anatomical target. Kinematics-based models, in turn, anticipate when a camera move is imminent but not where the interaction lies 61 . DiffeoAfford instead grounds a dense, continuous AH from instrumentβ anatomy interactions, yielding a fine-grained spatial target; incorporating explicit action-level semantics into this substrate is a natural extension rather than a competing paradigm, as discussed below. In contrast to purely reactive frameworks, the present approach acknowledges that expert surgeons organize interventions through proactive gaze behaviors, including target locking and quiet eye strategies. These behaviors fixate the interaction target before the hand arrives, and expert surgeons sustain markedly longer pre-movement fixations than novices 23,62 , satisfying the feedforward information requirements of the motor system. Training the prediction model on retrospectively aggregated interaction data enables it to estimate where interaction is expected to occur rather than to react to current tool contact. The LCET prediction evaluation examined this output against intraoperative attention directly. The centroid lag of AH prediction relative to surgeon gaze was significantly shorter than that of camera-assistant gaze in both evaluated phases, while group-mean profiles peaked on the leading side for AH prediction and on the lagging side for camera-assistant gaze (Fig. 5a). The peak lags themselves are descriptive, and no consistent temporal precedence over surgeon gaze was established. Spatially, AH prediction was significantly closer to surgeon gaze than either camera-assistant gaze or the image center (Fig. 5b). On AutoLaparo, the separately trained model reached 95.16% directional consistency immediately before camera motion and retained accuracy over longer horizons as the reactive instrument-tracking baseline degraded (Fig. 5c-e). The anticipation claimed here is therefore with respect to instrument action and subsequent camera motion, not with respect to surgeon gaze, against which the model shows close coupling with less lag than a human assistant. This comparison has a practical reading. Because the assistant frames the view according to where they look, divergence between the assistantβs gaze and the surgeonβs attention is the point at which the surgeon must intervene verbally, which is consistent with the lower number of verbal instructions recorded in the separate deployment cohort (Fig. 7d). This likely reflects the structure of the assistantβs task rather than individual competence: the assistant must infer the surgeonβs intent from the same image while simultaneously operating the scope. Physiological and cognitive rationale for anticipatory auto-framing Minimally invasive surgery fundamentally alters the surgeon's perceptual-motor environment. The dissociation of visual and motor axes, loss of tactile input, and restricted two-dimensional endoscopic view create a tunnel vision effect that severely diminishes peripheral spatial awareness. Consequently, precise field-of-view (FoV) control acts as an essential cognitive scaffold, providing ergonomically optimized visualization to alleviate the substantial burden of navigating this restricted space. Originating from gaze-contingent displays that mimic human foveated vision 63 , visualization technologies in surgical robotics have evolved into active camera control. Early directional gaze approaches effectively treated the eye as a joystick, leaving them vulnerable to the classic Midas Touch problem because they relied on continuous raw gaze without a deliberate confirmation step. In a comparison of gaze-control interfaces outside surgery, continuous directional control produced higher workload and lower task performance than embedded or GUI-based alternatives, indicating that sustained directional control carries an appreciable physical and cognitive cost 64 . While the field of Human-Computer Interaction (HCI) has explored object-based gaze interaction strategies to address this challenge, requiring users to explicitly select targets through gaze can disrupt natural visual exploration and introduce additional interaction demands. To mitigate this limitation in laparoscopic surgery, recent frameworks such as GazeScope avoid explicit selection by leveraging implicit gaze-based attention estimation to infer the user's region of interest and dynamically adjust tracking priorities among surgical tools. Furthermore, because these systems can operate with coarse estimates of visual attention, they reduce reliance on extensive calibration procedures and specialized eye-tracking hardware, enabling deployment using commodity RGB webcams 65 . AffordView requires no additional hardware and aligns the displayed field with the visuomotor constraints under which the surgeon operates. Conventional surgical monitors impose a rigid, two- dimensional spatial reference frame that triggers a central fixation bias, which acts as a robust attractor during scene exploration. This persistent bias acts as a spatial prior, influencing fixation behavior even when task-relevant visual saliency cues are present 66 . We observed this phenomenon in our clinical eye- tracking dataset, where the mean surgeon fixation location exhibited close proximity to the geometric center of the display (Mahalanobis distance of 0.05). By proactively centering the predicted surgical interaction locus, AffordView leverages this natural viewing tendency to provide a stable visual reference. Unlike reactive instrument-tracking approaches or explicit gaze-contingent systems, which may require additional calibration or deliberate gaze control, our intention-aware framework maintains task-relevant information near a predictable high-priority coordinate. Consequently, the operator avoids expending top-down executive effort to override the inherent spatial gravity of the monitor. Intraoperative measurements support this account, showing significant reductions in the surgeons' theta/alpha ratio and index of pupillary activity during targeted surgical phases. Subjective workload assessments corroborated these neurophysiological outcomes, with surgeons reporting significant reductions in both mental and physical demands on the SURG-TLX scale. In multicenter surveys, surgeons preferred this proactive centering strategy to both the unadjusted view and a conventional instrument-tracking baseline. By maintaining spatial congruence between the active surgical site and the display center, auto- framing reduces competition between task-relevant and display-centered reference frames 67 . This alignment may decrease the need for top-down attentional control to compensate for persistent spatial biases, allowing visual attention to remain more closely coupled with surgical objectives 67,68 . Intraoperative eye-tracking data support this effect, showing increased fixation duration on the adjusted view during Calot triangle and gallbladder dissection phases, consistent with greater engagement with the task- relevant surgical field. Limitations and future trajectories for semantic and physical autonomy Although the current framework predicts tissue affordance from visual evidence, it lacks semantic reasoning about why specific regions warrant interaction, and uncertain visual cues can present alternative plausible hotspots. For example, while a tissue plane may physically afford dissection, semantic awareness of gallbladder cancer infiltration is required to prevent interaction with malignant margins and avoid tumor seeding. This diverges fundamentally from the methodology of AffordTissue 52 . Whereas AffordTissue relies on language conditioning to explicitly instruct and generate tool-action specific affordance regions, we envision introducing VLMs (e.g., SurgVLM and EnVR-LPKG) strictly to inject this overarching semantic knowledge into our existing visually grounded framework. 69,70 This task, however, constrains how language may enter: view control must not depend on frequent intraoperative input from the surgeon, which would reinstate the very cognitive burden the framework seeks to reduce. Since most VLMs require a prompt at inference, even if it can be minimal or hiddenοΌhow to supply this semantic prior with little or no real-time input from the surgeon is the central open question. Furthermore, because the existing spatial centering function relies on digital cropping, it does not relieve the restricted peripheral awareness imposed by the narrow endoscopic view 29 . Cropping reduces the displayed field further. Holding the optics and scope position fixed did, however, isolate the effect of target selection from any confound of physical camera motion, since the control and experimental sessions differed only in where the display was centered. The resolution and field-of-view trade-off this reflects has motivated dual-view foveated laparoscopes that retain a wide overview alongside a magnified one, on the grounds that events outside the frame, including inadvertent contact by energized instruments, may go unrecognized 71 . A related caution applies to attentional guidance itself: augmented- reality navigation has been reported to improve targeting accuracy while reducing detection of unexpected findings that remain within the visual field 72 , so the reductions in cognitive load observed here should not be read as evidence of preserved peripheral vigilance, which was not measured. A further constraint concerns the comparison group: the camera assistants in the deployment cohort were novices with fewer than five prior procedures, the population in which misreading of surgical intent is most frequent. Whether the observed reduction in workload persists with experienced assistants remains to be established. Although wide-angle endoscopy can provide partial FOV expansion, fully resolving this physical limitation dictates integrating the intention-aware visual attention model into a kinematically redundant 7-degree-of-freedom robotic manipulator. Such hardware integration will facilitate authentic, physically actuated autonomous camera navigation that respects remote center of motion constraints while optimizing viewing angles, effectively emulating the anticipatory competence of an expert camera operator. More broadly, because the affordance hotspot represents where interaction is expected rather than a view-control signal specifically, the same predictions could support other intraoperative functions, each of which would require its own validation. Materials and methods Datasets Cholec80 dataset 59 : A public dataset containing 80 cholecystectomy videos, with surgical phases labeled, is widely used in laparoscopic surgery research. These procedures were performed by 13 different surgeons. Each video was recorded at a frame rate of 25 fps, with a resolution of 854 Γ 480 pixels. AutoLaparo dataset 73 : A public dataset comprises 21 laparoscopic hysterectomy videos, each recorded at 25 fps with a standard resolution of 1920Γ1080 pixels. Three sub-datasets are designed for the three tasks. For task 1, all 21 videos are annotated with surgical phases. For task 2, 300 video clips are selected from Phases 2 to 4, each lasting 10 seconds, with camera motion start at fifth second. Seven motion modes are defined: Static, Up, Down, Left, Right, Zoom-in, and Zoom-out. Private laparoscopic cholecystectomy eye -tracking (LCET) dataset: The LCET dataset comprises 26 laparoscopic cholecystectomy procedures performed in the Department of Hepatobiliary Surgery at Wuhan Union Hospital. Surgeon gaze was recorded in all 26 procedures using Tobii Pro Glasses 3 at a sampling rate of 100 Hz. In 13 procedures, surgeon gaze was recorded alone, and these procedures were used for the evaluation of DiffeoAfford grounding; in the remaining 13 procedures, camera- assistant gaze was recorded simultaneously using Tobii Pro Glasses 2 at 50 Hz, and these dual-gaze procedures were used for the evaluation of AH prediction. To overcome the distance limitations that render standard screen-based eye trackers ineffective during laparoscopic surgeries, we developed a robust framework for intraoperative eye-tracking acquisition and processing. This framework enables synchronized data collection from multiple subjects and ensures precise temporal and spatial alignment of gaze points with the scene and laparoscopic videos. Through this alignment, the gaze data is mapped onto the video frame coordinate system (Supplementary Fig. 5). The same framework was applied to the surgeon and camera-assistant recordings, placing both gaze sequences and the framewise AH predictions on a shared video timeline. DiffeoAfford pipeline for affordance hotspot dataset generation To retrospectively extract surgical AHs and construct a robust training dataset, we developed DiffeoAfford pipeline (Fig. 2). We observed that while novices struggle to predict the interaction locus prospectively, they can localize it in hindsight after viewing the complete operation. This phenomenon indicates that the definition of the target is inherent to the expert's execution, where the relevant affordance hotspot is dynamically revealed through the action itself 32 . Consequently, the surgeon's operation trajectory can serve as a ground-truth annotation for these hotspots. This insight aligns with frontier research in egocentric vision, which derives actionable regions directly from human demonstration videos. By treating the laparoscopic view as a specialized class of egocentric data, we can leverage these data-driven pipelines to extract affordance labels from expert traces automatically. Given the lack of a universally accepted representation for affordance, prediction problems vary by task, ranging from object localization and functional classification to segmentation, end effector pose estimation, and synthesis 74 . For view control tasks that require spatial location to center the camera, we treat functional segmentation as the target form of affordance prediction. DiffeoAfford depicted in Fig. 2 is designed to retrospectively ground AH in laparoscopic videos. To achieve this, the pipeline executes the following steps: (1) It first performs instance segmentation on the active surgical instruments to locate their tips and orientations (Instrument Segmentation & Localization). (2) Based on instrument motion, the video is initially divided into operative demonstration clips. The pipeline then iteratively applies the following sub-steps to subdivide these into single-action demonstration clips and generate corresponding AH labels (Temporal Video Segmentation): for each clip, (2.1) it establishes dense tracking across frames to account for tissue displacement (Spatial Correspondence) , (2.2) aggregates past and future interactions onto a single reference frame and propagates them back to all other frames (Aggregation and Propagation) , and (2.3) models these points into continuous labels (Refinement). Finally, this yields an AH dataset comprising video clips paired with AH annotations. Instrument Segmentation & Localization: Initially, the active surgical instruments were segmented using the Segment Anything Model 2 (SAM2), chosen for its proven effectiveness in surgical scenarios 75 . and the instrument tips and orientations are derived from the resulting segmentation mask for the subsequent steps. Temporal Video Segmentation: Most existing affordance grounding methods rely on datasets comprising manually segmented single-action demonstration clips 33,76-78 . However, applying this manual protocol to surgical videos would be prohibitively time-consuming and fundamentally contradict our goal of achieving an automated workflow. Because surgical intent and interaction hotspots are intrinsically linked to instrument kinematics, these continuous actions can serve as a reliable proxy for manual segmentation boundaries. To capitalize on this, we designed an iterative refinement process that automatically proposes and refines candidate action segments by leveraging instrument presence and the visibility of the AH. Initially, the video is segmented into operative demonstration clips based on the visible length and duration of the active instrument, and corresponding AH annotations are generated for each clip. To refine these into single-action demonstration clips, the pipeline iteratively applies an automated subdivision process. Specifically, any clip is subdivided if it meets either of the following criteria: (1) over 50% of the AH area (defined as the region with heatmap intensity > 0.1) is occluded for longer than the occluded threshold ν ν , or (2) the total clip duration exceeds the maximum length limit ν ν . Affordance heatmaps are then regenerated for each clip. This subdivision process repeats iteratively until no clips meet the criteria, after which any remaining clips shorter than the minimum duration limit ν ν are discarded. Because continuous surgical maneuvers in the datasets typically last between 10 and 30 seconds, we defined our temporal parameters as ν ν =10ν , ν ν =30ν , and ν ν =5ν . Through these settings, the pipeline automatically generates the candidate affordance datasets. For each generated clip, the following steps are executed to generate the corresponding AH labels: Spatial Correspondence: Following prior work on affordance grounding in egocentric settings 32,33,79 , frame-to-frame spatial correspondences are typically established via global transformations (e.g., homographic transforms). Therefore, advanced point trackers robust to occlusions are employed to obtain sparse spatial correspondences 80,81 . Subsequently, global transforms are fitted to these sparse matches to produce dense spatial correspondences. However, prior work typically targets rigid, natural scenes 32,33,79 , limiting their applicability in highly deformable anatomical environments. To address this, we propose a specialized method by incorporating diffeomorphic deformation fields alongside the existing global transforms (Fig. 2 Spatial Correspondence). This diffeomorphic constrained tracking allows better modeling of tissue deformation and reduces tracking noise (see Displacement Estimation in Supplementary Information for details). Within each video clip, we select a query frame to serve as the tracking start frame. To optimize tracking performance by the trackable area and minimizing the fitted displacement, we select a frame from the middle third of the clip with the least instrument occlusion (operationalized as the shortest visible instrument length) as the query frame. The tracking region is defined as the entire frame region excluding the black border and the instruments. Using this setup, we obtain pixel-wise spatial correspondences between the query frame and the other frames in the clip. Aggregation: To aggregate the instrument tip locations, points from each frame are mapped to their corresponding coordinates in the query frame using the displacement field of that specific frame (Fig. 2 Aggregation). Propagation: Once aggregated, these points are propagated back to all other frames via reverse mapping, relying on the corresponding displacement fields (Fig. 2 Propagation). Refinement: To ensure accuracy on each frame, the mapped points are filtered using the Mahalanobis distance to eliminate outliers. The remaining inlier points are then modeled using a 2D Gaussian distribution. Specifically, a probability density function is fitted to the points and normalized to generate the final affordance label ν΄ ( ν₯,ν¦ ) =exp (β 1 2 ( νβν ) ν νΊ β1 (ν βν)), where ν is the pixel coordinate vector, ν is the mean vector of the filtered points, and νΊ is their covariance matrix. Global-transform baselines: Following prior natural-scene affordance-grounding methods 32,33 , we implemented two global-only correspondence models. Similarity and homography transforms were fitted between the query frame and each target frame using RANSAC with the same sparse tracking points used by DiffeoAfford. Instrument tips were then aggregated in the query frame and propagated to all frames using the fitted transforms. The same Mahalanobis-distance filtering (<2) and centroid calculation were subsequently applied. All other pipeline components were unchanged, thereby isolating the contribution of the diffeomorphic deformation field. Real-time affordance hotspot prediction model To translate offline annotations into proactive intraoperative assistance through real-time prediction, we trained the Segformer model on the datasets generated by DiffeoAfford (Fig. 2). Segformer is a simple and efficient transformer-based semantic segmentation architecture 82 , which has proven effective for laparoscopic tissue recognition 83,84 . To adapt the model to our specific task, the output layer was modified to generate a single-channel affordance heatmap. To proactively predict targets without instrument- tracking bias, we trained exclusively on the first half of clips exceeding a minimum AH-to-instrument distance. During prediction, the centroid of the generated heatmap serves as the predicted AH location. Segformer-B3 were employed and achieved an inference speed of 70 FPS on an NVIDIA RTX 4090 GPU. AffordView application To utilize the real-time AH prediction model for autonomous FoV adjustment, we developed AffordView application (Fig. 2). The adjustments follow these specific protocols: Target Centering: The target position ν ν‘ =(ν₯ ν‘ ,ν¦ ν‘ ), determined either by the instrument-tracking baseline or by our proposed real- time AH prediction model, serves as the recommended view center. Cropping and Upscaling: The software preserves the original aspect ratio while cropping a region of width ν€ and height β around ν ν‘ . This cropped region is then upscaled to the original frame size νΓν» to produce the final adjusted FoV. Boundary Constraints: To prevent excessive cropping and magnification when the ν ν‘ is close to the image border, the crop dimensions are bounded such that ν€ β₯0.6ν and ββ₯0.6ν». This establishes a maximum magnification limit of 1.67Γ. To deploy this application for our clinical evaluations, a specific instance of the prediction model was trained. DiffeoAfford was used to generate 520 valid clips specifically from the Calot triangle dissection and Gallbladder dissection phases in videos 01 to 40. These clips were divided into training, validation, and test sets using an 8:1:1 ratio. On the test set, the Segformer-B3 model achieved an accuracy of 45 Β± 30 pixels (measured via Euclidean distance). Based on feedback from the initial surgeon preference evaluation, we also incorporated a smart zoom control mechanism to enhance the user experience. We observed that zoom adjustments are primarily driven by instrument dynamics rather than spatial target positioning. Therefore, we integrated a timing inference method based on instrument motion characteristics (developed in our previous work 85 ) to determine the optimal moments for zooming in and out. By learning from expert demonstrations, this module automatically detects keyframes in the video stream to trigger FoV adjustments. During a zoom-in event, the image is cropped and magnified around the target point (capped at 1.67x magnification); during a zoom-out event, this magnification effect is reverted. In short, our overall framework bridges retrospective data annotation with proactive surgical assistance. DiffeoAfford first extracts dense anatomical affordance labels from laparoscopic videos, translating implicit surgical intent into quantifiable visual targets. The real-time prediction model then learns from this dataset to proactively predict functional interaction hotspots. Finally, the AffordView application leverages these spatial predictions alongside instrument-driven zoom cues to autonomously execute optimal, ergonomically efficient FoV adjustments during live procedures. Evaluation setup Evaluation of DiffeoAfford on Cholec80 dataset: We evaluated the annotation accuracy of DiffeoAfford using the Calot triangle dissection phase from videos 01-10. We selected the tip of the electrosurgical hook as the primary instrument reference point. The pipeline yielded 141 video clips for evaluation. To reduce instrument-induced annotator bias, we randomly sampled one frame per clip where the distance between the pipeline-annotated AH center and the instrument tip exceeded the dataset's mean distance. Three annotator groups were recruited: senior surgeons (ν=3, >10 yearsβ experience, serving as ground truth), junior doctors (ν=3, 2-3 yearsβ experience), and medical students (ν=3, novices). Senior surgeons annotated images after viewing the corresponding video clips. Junior and novice groups initially annotated static images without context; the novice group then re-annotated them after viewing the contextual clips (In-context). Evaluation of DiffeoAfford against intraoperative gaze on LCET dataset: This evaluation used the 13 LCET procedures with surgeon gaze recordings only. The pipeline generated 60 video clips and randomly sampled one frame every 6 second, yielding 164 images. Unlike the Cholec80 evaluation, the surgeon's actual intraoperative gaze points served as the ground truth. Let ν ν be the timestamp of the sampled image. Three temporal ground truths were established by calculating the mean coordinates of gaze points within the windows ν ν Β±0.5s, ν ν Β±1.0s, and ν ν Β±2.0s. We computed the Euclidean distances from these ground truths to the AH annotations provided by both the pipeline and the human annotators (ν=3 junior doctors, before and after watching the 10-second reference clips) . Evaluation of AH prediction against intraoperative gaze on the LCET dataset: This evaluation used the 13 LCET procedures with synchronized surgeon and camera-assistant eye tracking, restricted to the Calot triangle dissection and gallbladder dissection phases. AH prediction sequences were generated using the same Cholec80-trained SegFormer-B3 model used in the AffordView application, taking the raw framewise centroid of each predicted heatmap without trajectory filtering. All gaze coordinates were mapped to the 1920 Γ 1080 laparoscopic-video coordinate system; surgeon and camera-assistant gaze were placed on a common 50-Hz time grid, and timestamped AH predictions were assigned to the nearest grid samples. Non-contiguous intervals belonging to the same phase were retained as separate segments so that temporal pairs were not formed across segment boundaries. Temporal alignment was evaluated using cross-recurrence quantification analysis (CRQA), which has previously been used in dual eye-tracking studies to characterize temporal coordination and leader- follower relationships between paired gaze sequences 86 . Because the laparoscopic field of view changes continuously, the analysis was restricted to a diagonal band around the main diagonal, which focused the analysis on temporally local recurrence within the same evolving surgical scene 86 , and was summarized by the diagonal cross-recurrence profile (DCRP), which describes the direction and magnitude of temporal coupling between two sequences 87 . The band spanned candidate lags from -3.0 to +3.0 s in 20 ms increments, a window selected from previous gaze cross-recurrence research focusing on recurrence close to the main diagonal and the approximately 3 s intervals between camera movements reported for experienced surgeons during camera targeting tasks 86,88 . Let S(t) denote surgeon gaze and C(t) denote either AH prediction or camera assistant gaze. Recurrence at candidate lag Ο was defined when the angular separation between surgeon gaze at time t and the candidate sequence at time t + Ο did not exceed 1.5Β°, following the spatial proximity criterion used in previous gaze CRQA research 86 , and the recurrence rate at each lag was calculated as the number of recurrent pairs divided by the total number of valid within-segment pairs at that lag: ν ( ν‘,ν ) =νΌ οΏ½ ν ν οΏ½ ν ( ν‘ ) ,νΆ ( ν‘+ν ) οΏ½ β€ 1. 5Β° οΏ½ ν ( ν ) = β ν ( ν‘,ν ) ν‘ ν ( ν£νν£ νν£ )( ν ) A negative candidate lag indicates that the candidate sequence leads surgeon gaze. Procedure-level recurrence-rate profiles were averaged with equal weight across procedures to obtain the displayed group profile, and the shaded bands were calculated as the mean plus or minus 1.96 standard errors. For each procedure, the DCRP was treated as a distribution over candidate lags, and temporal alignment was summarized by its recurrence-rate-weighted mean, referred to here as the centroid lag, following the lag-profile distribution approach described in previous visual-attention research 89 : ν ν = β ν R(ν) ν ν=βν β R(ν) ν ν=βν where ν ν is the centroid lag, ν (ν) is the recurrence rate at candidate lag ν, and ν=3.0 ν . Negative centroid-lag values indicate that the candidate tends to lead surgeon gaze, positive values indicate that it tends to lag surgeon gaze, and values near zero indicate temporal coupling without a consistent lead or lag. Spatial alignment was assessed at simultaneous zero-lag samples. Timestamps were retained when surgeon gaze, AH prediction, and camera-assistant gaze were all valid for the camera-assistant comparison, and when surgeon gaze and AH prediction were valid for the image-center comparison, with the image center defined as ( 960,540 ) pixels. Euclidean distances to surgeon gaze were calculated at each retained timestamp and reduced to a median for each procedure and surgical phase. Median advantage was defined as the median across procedures of ν ννν ννννν‘ν ν βν ν΄ , so that positive values indicate a smaller median distance for AH prediction. Evaluation of real-time AH prediction for camera motion forecasting on AutoLaparo dataset: To investigate the relationship between the model-predicted AH and actual camera motion, we evaluated our method on the AutoLaparo dataset. Because the standardized video clips for AutoLaparo Task 2 (which evaluates camera motion prediction) were exclusively extracted from phase 2 (Dividing Ligament and Peritoneum), phase 3 (Dividing Uterine Vessels and Ligament), and phase 4 (Transecting the Vagina), we specifically targeted these three phases for our evaluation. Using the tips of the bipolar forceps and electric hook as primary instruments, DiffeoAfford generated affordance annotations for 435 training clips from videos 1-10 of Task 1. These annotations were used to train an AutoLaparo-specific SegFormer-B3 model. This model was then used to predict AH locations on the unseen video clips of AutoLaparo Task 2. The final prediction result was defined as the temporal mean of the AH locations predicted within a 0.5-second window prior to the camera motion event. This was compared against the instrument-tracking baseline (detailed in the Supplementary Information), which used the same 0.5- second temporal mean protocol. Surgeon preference assessment of AffordView application: We designed two paired comparisons via an online questionnaire: (1) Original video vs. AffordView application using proposed AH prediction model, and (2) AffordView using proposed AH prediction model vs. the instrument-tracking baseline. To maximize the observable differences for evaluation, we selected 30 thirty-second clips for each comparison from the Calot triangle dissection phase of videos 41-50 in the Cholec80 dataset (extracting 3 clips per video). The selection criterion was based on the maximum average spatial distance between the target points generated by the two respective methods within that clip. Notably, for the original video, the geometric center of the frame was designated as its target point. Three versions of the questionnaire were deployed via the Python Flask framework (Supplementary Fig. 6), each containing 20 randomly ordered pairs (10 from each comparison). Intraoperative user experience assessment of AffordView application: 5 operating surgeons and 12 novice camera assistants (each with experience in fewer than 5 procedures) were recruited for this study. The sample size was determined a priori using a power analysis for paired comparisons (two- sided νΌ=0.05, power = 0.80, expected Cohen's νν=1.0 based on effect sizes reported in prior surgical cognitive workload studies 90 ), yielding a minimum requirement of 10 pairs. 12 paired laparoscopic cholecystectomies were performed to evaluate the application's overall ergonomic impact. Each pair consisted of a control session (dual monitors only, original view only) and an experimental session (one original view monitor, one adjusted view monitor). The same surgeon and camera assistant participated in each pair, with the session order strictly alternated to balance sequence effects. This evaluation incorporated subjective CWL assessment via the SURG-TLX questionnaire, alongside objective EEG and eye-tracking metrics. Details regarding EEG data acquisition and preprocessing are provided in the EEG data acquisition and preprocessing section. To evaluate the surgeons' cognitive workload, we calculated the theta/alpha ratio (TAR), Engagement index (EI) and occipital alpha band power (OA) as: νν΄ν = ν Frontal ν ν Parietal νΌ νΈνΌ= ν Parietal ν½ ν Parietal νΌ +ν Parietal ν OA=ν νννννν‘νν£ νΌ where ν represents the spectral power of a specific frequency band. The cortical regions were defined using the following electrode montages: F3, F4, and Fz for the frontal region; P7 and P8 for the parietal region; and O1 and O2 for the occipital region. Eye-tracking data were collected using Tobii Pro Glasses 3 and processed with Tobii Pro Lab and Python. Metrics included the index of pupillary activity (IPA), fixation duration proportions, and spatial entropy. For gaze entropy, the image area was discretized into 20Γ20 bins. Stationary gaze entropy (SGE) and gaze transition entropy (GTE) were calculated as: ν νΈ=βοΏ½ν ν logν ν ν νννΈ=βοΏ½ν ν οΏ½ν νν logν νν νν where ν ν is the positional distribution probability in the ν-th bin, and ν νν is the transition probability from the ν-th to the ν-th bin. Finally, the surgeon's ratings of utility and the number of verbal instructions to the camera assistant were recorded per surgical phase. EEG data acquisition and preprocessing EEG data were acquired using a 64-channel Brain Products LiveAmp system, which was electromagnetically shielded to minimize radio frequency interference from energy devices, and processed using the MATLAB EEGLAB toolbox. Prior to workload analysis, continuous EEG signals underwent a standardized preprocessing pipeline. Non-EEG auxiliary channels (i.e., IMU sensors) were removed, and the data were bandpass filtered between 4 Hz and 30 Hz. Bad channels and high- amplitude artifacts were detected and rejected using the EEGLAB clean_rawdata function incorporating Artifact Subspace Reconstruction (ASR) with a burst criterion of 20. Missing channels were subsequently spherically interpolated, and the data were re-referenced to the common average. To eliminate ocular and myogenic artifacts, extended Infomax Independent Component Analysis (ICA) was applied. The EEGLAB ICLabel function was utilized to automatically identify and remove non-brain independent components with an artifact classification probability of β₯ 90%. Following artifact rejection, the spectral power for the theta (4-8 Hz), alpha (8-12 Hz), and beta (12-30 Hz) frequency bands was calculated across the predefined surgical phases. Statistical analysis To control for confounding effects during the dataset evaluations, we employed linear mixed models (LMMs). For the Cholec80 dataset, the LMM included group as a fixed effect, with image and annotator as random intercepts. The novice (In-context) group served as the reference level to compute estimated differences against other groups. Similarly, for the LCET dataset, we applied the identical LMM structure, utilizing the junior (In-context) group as the reference level. To account for multiple comparisons against the reference level within these models, all resulting p values were adjusted using the Benjamini- Hochberg procedure. To assess whether DiffeoAfford achieved accuracy equivalent to that of context-informed human annotators, we conducted the same median-based, image-level bootstrap equivalence analysis for Cholec80 and LCET. For each image ν, the human reference error β ν was defined as the median error across the three Novice (In-context) annotators for Cholec80 or the three Junior (In-context) annotators for LCET, and the paired difference was calculated as ν ν =ν ν·νν·νν΄ ν·ννν£,ν ββ ν . Within-human disagreement was defined as ν ν =νννν νν ν<ν |ν νν βν νν |, and the data-driven equivalence margin as β= νννννν ν (ν ν ), so equivalence denotes a discrepancy no larger than the typical disagreement between two human annotators. The standardized effect was calculated as ν = νννννν ν (ν ν )/β, a median-based paired statistic whose sign may differ from the LMM mean estimate. Images were resampled with replacement 20,000 times, with both νννννν ν (ν ν ) and β recalculated in each replicate. Equivalence was concluded when the 90% percentile bootstrap CI for ν , corresponding to two one-sided tests at νΌ = 0.05, lay entirely within the interval from -1 to 1. Global-transform baselines were compared with DiffeoAfford using two-sided paired Wilcoxon signed-rank tests. For the LCET prediction evaluation, each procedure was treated as the statistical unit, and framewise samples were used only to derive procedure-level metrics rather than as independent observations. Within each surgical phase, procedure-level DCRP centroid lags and median distances to surgeon gaze were compared using exact two-sided paired Wilcoxon signed-rank tests: AH prediction versus camera- assistant gaze for both metrics, and AH prediction versus the image center for distance. The p values of these six prespecified comparisons were adjusted jointly using the Holm procedure. Group-mean DCRP peak lags were treated as descriptive and were not subjected to statistical testing. For the surgeon preference assessment, the statistical significance of the preference proportions was evaluated using exact binomial tests. For the intraoperative user experience evaluation, permutation- based paired tests served as the primary inference method. LMMs were subsequently used to further control confounding effects, incorporating condition, phase, and their interaction (condition Γ phase) as fixed effects, order as a fixed covariate, and pair as a random intercept. All statistical tests were two- sided. Ethical approval The study protocol was approved by the Medical Ethics Committee of Union Hospital, Tongji Medical College, Huazhong University of Science and Technology (UHCT250178). All participants involved in the intraoperative user experience assessment signed an informed consent form; no financial compensation was provided. Data availability The datasets generated and/or analyzed during the current study are not publicly available at this stage due to ongoing preparation for public release. They will be made available upon publication of this work. Code availability The code used to generate the results reported in this study will be made publicly available upon publication. Acknowledgements H.D., X.C., Y.W., C.W. and H.Z. disclose support for the research of this work from the Hubei Science and Technology Major Program [grant number 2023BCA002] and the National Key Research and Development Program of China [grant number 2022YFC2407402]. Author contributions J.G., X.C., Y.W., J.P. and H.D. conceptualized the project. H.D., X.C., Y.W., C.W. and H.Z. acquired the funding. J.G. designed and developed DiffeoAfford. J.G., J.Z., S.Z. and X.C. designed and developed the AffordView application. J.G. and Y.W. designed the evaluation experiments for grounding and prediction of AHs, while J.G., J.P., X.C. and Q.F. designed the assessment experiments for the AffordView application. J.P. and Q.F. provided the eye tracker and technical support. W.C. and C.X. provided the electroencephalography equipment and technical support. G.C., K.L., Y.C., H.W. and S.S. performed the assessment experiments. J.G., J.Z. and H.W. analyzed the data. J.G., S.Z. and H.W. constructed the dataset. J.G., X.C., Y.W., J.P. and H.D. wrote and edited the paper. References 1. Crigger E, Khoury C. Making Policy on Augmented Intelligence in Health Care. AMA journal of ethics. 2019;21(2):E188-191. 2. Augmented intelligence in medicine. American Medical Association. 2026. 3. Maier-Hein L, Vedula S, Speidel S, et al. Surgical data science for next-generation interventions. Nature Biomedical Engineering. 2017;1(9):691-696. 4. Ben Hamed S. Decoding Covert Visual Attention in Space and Time from Neural Signals. Annual Review of Vision Science. 2025;11(1):495-520. 5. Belardinelli A. Gaze-Based Intention Estimation: Principles, Methodologies, and Applications in HRI. J Hum- Robot Interact. 2024;13(3):31:31-31:30. 6. Duncan DH, Van Moorselaar D, Theeuwes J. Pinging the brain to reveal the hidden attentional priority map using encephalography. Nature Communications. 2023;14(1):4749. 7. Illamperuma NH, Fooken J. Towards a functional understanding of gaze in goal-directed action. Journal of Neurophysiology. 2024;132(3):767-769. 8. Madani A, Namazi B, Altieri MS, et al. Artificial intelligence for intraoperative guidance: using semantic segmentation to identify surgical anatomy during laparoscopic cholecystectomy. Annals of surgery. 2020. 9. Balabadruni Venkata N, D S, Ganesh Reddy GSL, Potturi P. Eye gaze-controlled camera navigation for enhanced robotic surgery with potential cognitive load reduction. Current Problems in Surgery. 2025;70:101846. 10. Cartella G, Cornia M, Cuculo V, et al. Trends, Applications, and Challenges in Human Attention Modelling. 2024. 11. Chen J, Duan H, Zhang X, Gao B, Grau V, Han J. From gaze to insight: Bridging human visual attention and vision language model explanation for weakly-supervised medical image segmentation. IEEE Transactions on Medical Imaging. 2025. 12. Nwoye CI, Yu T, Gonzalez C, et al. Rendezvous: Attention mechanisms for the recognition of surgical action triplets in endoscopic videos. Medical Image Analysis. 2022;78:102433. 13. Maier-Hein L, Eisenmann M, Sarikaya D, et al. Surgical data science β from concepts toward clinical translation. Medical Image Analysis. 2022;76:102306. 14. Zhang Z, Mascagni P, Reinke A, et al. Artificial IntelligenceβBased Analysis of Laparoscopic Imaging for Intraoperative Surgical Decision Support. Annual Review of Biomedical Engineering. 2026;28(Volume 28, 2026):135-162. 15. Carstens M, Vasisht S, Zhang Z, et al. Artificial intelligence for surgical scene understanding: a systematic review and reporting quality meta-analysis. npj Digital Medicine. 2025;9(1):59. 16. Schmidgall S, Opfermann JD, Kim JW, Krieger A. Will your next surgeon be a robot? Autonomy and AI in robotic surgery. Science Robotics. 2025;10(104):eadt0187. 17. Jia W, Yang F, Duan M, et al. One-shot affordance grounding of deformable objects in egocentric organizing scenes. Paper presented at: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)2025. 18. Bar O, Neimark D, Zohar M, et al. Impact of data on generalization of AI for surgical intelligence applications. Scientific Reports. 2020;10(1):22208. 19. Bihlmaier A, Worn H. Learning surgical know-how: Dexterity for a cognitive endoscope robot. Paper presented at: 2015 IEEE 7th International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE Conference on Robotics, Automation and Mechatronics (RAM)2015. 20. Gao Y, Li Z, Zhao J, Li J, Li J. A humanβAI collaborative framework for surgical field-of-view adjustment: design and experimental validation. Journal of Robotic Surgery. 2025;19(1):259. 21. Gibson J. The theory of affordances. The ecological approach to visual perception. The people, place and, space reader. 1979:56-60. 22. Harvey A, Vickers JN, Snelgrove R, Scott MF, Morrison S. Expert surgeon's quiet eye and slowing down: expertise differences in performance and quiet eye duration during identification and dissection of the recurrent laryngeal nerve. The American Journal of Surgery. 2014;207(2):187-193. 23. Abdelaal AE, Van Rumpt R, Zaman S, et al. The quiet eye phenomenon in minimally invasive surgery. International Journal of Computer Assisted Radiology and Surgery. 2025;20(10):2087-2093. 24. Qiu L, Shen L, Liu L, Liu J, Chen Y, Xing L. ST-NeRP: Spatial-Temporal Neural Representation Learning with Prior Embedding for Patient-specific Imaging Study. 2024. 25. Zhang J, Wang Y, Zhou S, et al. Graph-Based Spatial Reasoning for Tracking Landmarks in Dynamic Laparoscopic Environments. IEEE Robotics and Automation Letters. 2024. 26. Liu J, Qiao X, Xiao Y, et al. Physical and mental health impairments experienced by operating surgeons and camera-holder assistants during laparoscopic surgery: a cross-sectional survey. Frontiers in Public Health. 2023;11. 27. Richter HO, Sundin S, Long J. Visually deficient working conditions and reduced work performance in office workers: Is it mediated by visual discomfort? International Journal of Industrial Ergonomics. 2019;72:128- 136. 28. Wong SW, Crowe P. Visualisation ergonomics and robotic surgery. Journal of robotic surgery. 2023;17(5):1873-1878. 29. DeLucia PR, Griswold JA. Effects of camera arrangement on perceptual-motor performance in minimally invasive surgery. Journal of Experimental Psychology: Applied. 2011;17(3):210. 30. Bapna T, Valles J, Leng S, Pacilli M, Nataraja RM. Eye-tracking in surgery: a systematic review. ANZ journal of surgery. 2023;93(11):2600-2608. 31. Zakeri Z, Mansfield N, Sunderland C, Omurtag A. Physiological correlates of cognitive load in laparoscopic surgery. Scientific Reports. 2020;10(1):12927. 32. Liu S, Tripathi S, Majumdar S, Wang X. Joint hand motion and interaction hotspots prediction from egocentric videos. Paper presented at: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition2022. 33. Li G, Tsagkas N, Song J, et al. Learning precise affordances from egocentric videos for robotic manipulation. Paper presented at: Proceedings of the IEEE/CVF International Conference on Computer Vision2025. 34. Altobelli E, Gidaro S, Bove AM, et al. 1405 TELELAP ALF-X: a novel telesurgical system for the 21st century. The Journal of Urology. 2013;189(4S):e575-e576. 35. Fujii K, Gras G, Salerno A, Yang G-Z. Gaze gesture based human robot interaction for laparoscopic surgery. Medical image analysis. 2018;44:196-214. 36. Hirano Y, Kondo H, Yamaguchi S. Robot-assisted surgery with Senhance robotic system for colon cancer: our original single-incision plus 2-port procedure and a review of the literature. Techniques in Coloproctology. 2021;25(4):467-471. 37. Rehrig G, Barker M, Peacock CE, Hayes TR, Henderson JM, Ferreira F. Look at what I can do: Object affordances guide visual attention while speakers describe potential actions. Attention, Perception, & Psychophysics. 2022;84(5):1583-1610. 38. Garrido-VΓ‘squez P, SchubΓΆ A. Modulation of visual attention by object affordance. Frontiers in Psychology. 2014;5:59. 39. Huber M, Ourselin S, Bergeles C, Vercauteren T. Deep Homography Prediction for Endoscopic Camera Motion Imitation Learning. Paper presented at: International Conference on Medical Image Computing and Computer-Assisted Intervention2023. 40. Marchand C, De Graaf JB, JarrassΓ© N. Measuring mental workload in assistive wearable devices: a review. Journal of NeuroEngineering and Rehabilitation. 2021;18(1):160. 41. Mastropietro A, Pirovano I, Marciano A, Porcelli S, Rizzo G. Reliability of Mental Workload Index Assessed by EEG with Different Electrode Configurations and Signal Pre-Processing Pipelines. Sensors (Basel). 2023;23(3). 42. Lim C, Barragan JA, Farrow JM, Wachs JP, Sundaram CP, Yu D. Physiological Metrics of Surgical Difficulty and Multi-Task Requirement during Robotic Surgery Skills. Sensors (Basel). 2023;23(9). 43. Magosso E, De Crescenzio F, Ricci G, Piastra S, Ursino M. EEG alpha power is modulated by attentional changes during cognitive tasks and virtual reality immersion. Computational intelligence and neuroscience. 2019;2019(1):7051079. 44. Cao T, Wan F, Wong CM, da Cruz JN, Hu Y. Objective evaluation of fatigue by EEG spectral analysis in steady-state visual evoked potential-based brain-computer interfaces. Biomedical engineering online. 2014;13(1):28. 45. Pope AT, Bogart EH, Bartolome DS. Biocybernetic system evaluates indices of operator engagement in automated task. Biological Psychology. 1995;40(1):187-195. 46. D'Ambrosia C, Aronoff-Spencer E, Huang EY, et al. The neurophysiology of intraoperative error: An EEG study of trainee surgeons during robotic-assisted surgery simulations. Frontiers in Neuroergonomics. 2023;3. 47. Duchowski AT, Krejtz K, Krejtz I, et al. The index of pupillary activity: Measuring cognitive load vis-Γ -vis task difficulty with pupil oscillation. Paper presented at: Proceedings of the 2018 CHI conference on human factors in computing systems2018. 48. Wilson MR, Poolton JM, Malhotra N, Ngo K, Bright E, Masters RS. Development and validation of a surgical workload measure: the surgery task load index (SURG-TLX). World journal of surgery. 2011;35(9):1961- 1969. 49. Wang Z, Tian G. Task-oriented robot cognitive manipulation planning using affordance segmentation and logic reasoning. IEEE Transactions on Neural Networks and Learning Systems. 2023;35(9):12172-12185. 50. Wu R, Ning C, Dong H. Learning foresightful dense visual affordance for deformable object manipulation. Paper presented at: Proceedings of the IEEE/CVF International Conference on Computer Vision2023. 51. Laplante S, Namazi B, Kiani P, et al. Validation of an artificial intelligence platform for the guidance of safe laparoscopic cholecystectomy. Surgical endoscopy. 2023;37(3):2260-2268. 52. Maksutova A, Seenivasan L, Ding H, et al. AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction. arXiv preprint arXiv:260401371. 2026. 53. Silva-Gago M, Fedato A, Hodgson T. Visual attention reveals affordances during Lower Palaeolithic stone tool exploration. Archaeol Anthropol Sci 13: 145. In:2021. 54. Gong S, Long Y, Chen K, et al. Self-supervised cyclic diffeomorphic mapping for soft tissue deformation recovery in robotic surgery scenes. IEEE Transactions on Medical Imaging. 2024;43(12):4356-4367. 55. Chi MT, Feltovich PJ, Glaser R. Categorization and representation of physics problems by experts and novices. Cognitive science. 1981;5(2):121-152. 56. Stevenson HW, Azuma HE, Hakuta KE. Child development and education in Japan. Paper presented at: This book is based on a conference sponsored by the Center for Advanced Study in the Behavioral Sciences; the Japan Society for the Promotion of Science; and the Joint Committee on Japanese Studies of the American Council of Learned Societies and the Social Science Research Council.1986. 57. Clark RE, Pugh CM, Yates KA, Inaba K, Green DJ, Sullivan ME. The use of cognitive task analysis to improve instructional descriptions of procedures. Journal of Surgical Research. 2012;173(1):e37-e42. 58. Islam M, Vibashan V, Lim CM, Ren H. St-mtl: Spatio-temporal multitask learning model to predict scanpath while tracking instruments in robotic surgery. Medical Image Analysis. 2021;67:101837. 59. Twinanda AP, Shehata S, Mutter D, Marescaux J, De Mathelin M, Padoy N. Endonet: a deep architecture for recognition tasks on laparoscopic videos. IEEE transactions on medical imaging. 2016;36(1):86-97. 60. Rivoir D, Bodenstedt S, Funke I, et al. Rethinking anticipation tasks: Uncertainty-aware anticipation of sparse surgical instrument usage for context-aware assistance. Paper presented at: International conference on medical image computing and computer-assisted intervention2020. 61. Kossowsky H, Nisky I. Predicting the timing of camera movements from the kinematics of instruments in robotic-assisted surgery using artificial neural networks. IEEE Transactions on Medical Robotics and Bionics. 2022;4(2):391-402. 62. Causer J, Harvey A, Snelgrove R, Arsenault G, Vickers JN. Quiet eye training improves surgical knot tying more than traditional technical training: a randomized controlled study. The American Journal of Surgery. 2014;208(2):171-177. 63. Duchowski AT, Cournia N, Murphy H. Gaze-contingent displays: A review. Cyberpsychology & behavior. 2004;7(6):621-634. 64. Sardinha EN, Zook N, Garate VR, Western D, Munera M. Comparison of Three Interface Approaches for Gaze Control of Assistive Robots for Individuals with Tetraplegia. Paper presented at: 2025 IEEE International Conference on Robotics and Automation (ICRA)2025. 65. Zhang J, Wang B, Pan Z, Li M. GazeScope: A Framework of Gaze Attention-Based Automatic Field-of-View Adjustment for Laparoscopic Robots. IEEE Robotics and Automation Letters. 2025. 66. Tseng P-H, Carmi R, Cameron IGM, Munoz DP, Itti L. Quantifying center bias of observers in free viewing of dynamic natural scenes. Journal of Vision. 2009;9(7):4-4. 67. Bindemann M. Scene and screen center bias early eye movements in scene viewing. Vision research. 2010;50(23):2577-2587. 68. Mairon R, Ben-Shahar O. Stimulus Center Bias Persists Irrespective of Its Position on the Display. Journal of Eye Movement Research. 2025;18(6):77. 69. Zeng Z, Zhuo Z, Jia X, et al. SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence. arXiv preprint arXiv:250602555. 2025. 70. Hao P, Wang H, Yang G, Zhu L. Enhancing visual reasoning with LLM-powered knowledge graphs for visual question localized-answering in robotic surgery. IEEE Journal of Biomedical and Health Informatics. 2025. 71. Katz J, Hua H, Lee S, Nguyen M, Hamilton A. A dual-view multi-resolution laparoscope for safer and more efficient minimally invasive surgery. Scientific Reports. 2022;12(1):18444. 72. Dixon BJ, Daly MJ, Chan H, Vescan AD, Witterick IJ, Irish JC. Surgeons blinded by enhanced navigation: the effect of augmented reality on attention. Surgical endoscopy. 2013;27(2):454-461. 73. Wang Z, Lu B, Long Y, et al. Autolaparo: A new dataset of integrated multi-tasks for image-guided surgical automation in laparoscopic hysterectomy. Paper presented at: International Conference on Medical Image Computing and Computer-Assisted Intervention2022. 74. Apicella T, Xompero A, Cavallaro A. Visual Affordances: Enabling Robots to Understand Object Functionality. arXiv preprint arXiv:250505074. 2025. 75. Liu H, Zhang E, Wu J, Hong M, Jin Y. Surgical SAM 2: Real-time Segment Anything in Surgical Video by Efficient Frame Pruning. 2024. 76. Fang K, Wu T-L, Yang D, Savarese S, Lim J. Demo2vec: Reasoning object affordances from online videos. Paper presented at: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition2018. 77. Damen D, Doughty H, Farinella GM, et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision. 2022;130(1):33-55. 78. Nagarajan T, Feichtenhofer C, Grauman K. Grounded human-object interaction hotspots from video. Paper presented at: Proceedings of the IEEE/CVF International Conference on Computer Vision2019. 79. Ju Y, Hu K, Zhang G, Zhang G, Jiang M, Xu H. Robo-abc: Affordance generalization beyond categories via semantic correspondence for robot manipulation. Paper presented at: European Conference on Computer Vision2024. 80. Xiao Y, Wang Q, Zhang S, et al. SpatialTracker: Tracking Any 2D Pixels in 3D Space. Paper presented at: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition2024. 81. Karaev N, Makarov I, Wang J, Neverova N, Vedaldi A, Rupprecht C. Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024. URL https://arxivorg/abs/241011831. 82. Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P. SegFormer: Simple and efficient design for semantic segmentation with transformers. Advances in neural information processing systems. 2021;34:12077-12090. 83. Kolbinger FR, Rinner FM, Jenke AC, et al. Anatomy segmentation in laparoscopic surgery: comparison of machine learning and human expertise - an experimental study. Int J Surg. 2023;109(10):2962-2974. 84. Skinner G, Chen T, Jentis G, et al. Real-time near infrared artificial intelligence using scalable non-expert crowdsourcing in colorectal surgery. npj Digital Medicine. 2024;7(1):99. 85. Zhang J, Shi S, Wang Y, et al. Automatic Keyframe Detection for Critical Actions from the Experience of Expert Surgeons. 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); 2022. 86. Jermann P, Mullins D, NΓΌssli M-A, Dillenbourg P. Collaborative gaze footprints: Correlates of interaction quality. 2011. 87. Darici D, Sieg L, Eismann H, Karsten J. Leader-follower dynamics in medical training: A dual mobile eye- tracking analysis of teacher-student gaze patterns. Medical Teacher. 2026;48(1):122-130. 88. Jarc AM, Curet MJ. Viewpoint matters: objective performance metrics for surgeon endoscope control during robot-assisted surgery. Surgical endoscopy. 2017;31(3):1192-1202. 89. Dale R, Kirkham NZ, Richardson DC. The dynamics of reference and shared visual attention. Frontiers in psychology. 2011;2:355. 90. Moore LJ, Wilson MR, McGrath JS, Waine E, Masters RS, Vine SJ. Surgeonsβ display reduced mental effort and workload while performing robotically assisted surgical tasks, when compared to conventional laparoscopy. Surgical endoscopy. 2015;29(9):2553-2560. Supplementary Information Supplementary Fig. 1 | Center bias in laparoscopic surgery. Analysis of 50,596 fixation centroids from 26 LCET procedures revealed a distinct center bias. a, Fixation distribution heatmap. Gray ellipses represent Mahalanobis distances of 1 and 2; the red line connects the mean fixation point to the geometric center of the monitor (ν· ν =0.05 ). b, Area-normalized fixation density decreased with distance from the monitor center. Shading and dashed lines denote the 1ν, 2ν, and 3ν radial ranges (ν=265.7 pixels). a b c Supplementary Fig. 2 | Inter-group annotation comparisons, residual diagnostics, and equivalence analysis on the Cholec80 dataset. a, b, Statistical differences and model diagnostics evaluated before (a) and after (b) applying a Box-Cox transformation (ν=0.17) to the annotation error data. In each panel, the top forest plot details the estimated differences in annotation errors relative to the Novice (In-context) reference group. The bottom panels display the corresponding residual diagnostics for the linear mixed models (LMMs), including the histogram of LMM residuals (left) and the normal Q-Q plot (right). The Box-Cox transformation applied in (b) effectively normalizes the right-skewed residual distribution observed in (a). c, Median-based bootstrap equivalence analysis comparing DiffeoAfford with Novice (In-context) across 141 paired images. The point and horizontal bar denote the normalized median paired error difference (ν ) and its 90% image-level bootstrap CI, respectively. The shaded region denotes the equivalence interval from β1 to 1; equivalence was concluded when the entire CI lay within this interval. ν and the data-driven equivalence margin β are defined in Methods. The data-driven equivalence margin was β=17.13 pixels. CI, confidence interval; Adj. p, adjusted ν value; LMM, linear mixed model; Q-Q, quantile-quantile. Supplementary Fig. 3 | Comparison with global-transform affordance-grounding baselines. Violin plots show grounding errors for DiffeoAfford, similarity RANSAC, and homography RANSAC across 141 paired Cholec80 images. Brackets indicate paired Wilcoxon signed-rank tests against DiffeoAfford after Benjaminiβ Hochberg correction. **, ν<0.01; ***, ν<0.001. a b c d e f g Supplementary Fig. 4 | Inter-group annotation comparisons, residual diagnostics, and equivalence analysis on the LCET dataset. a-f, Statistical differences and model diagnostics evaluated across temporal windows of Β±0.5 s (a, d), Β±1.0 s (b, e), and Β±2.0 s (c, f), presented before (a-c) and after (d- f) applying a Box- Cox transformation to the annotation error data. The optimal lambda parameters for the transformations are ν=0.29 (d), ν=0.32(e), and ν=0.33 (f). In each panel, the top forest plot details the estimated differences in annotation errors relative to the junior (in-context) reference group. The bottom panels display the corresponding residual diagnostics for the linear mixed models (LMMs), including the histogram of LMM residuals (left) and the normal Q-Q plot (right). The Box-Cox transformations applied in d- f effectively normalize the right-skewed residual distributions observed in a-c. g, Median-based bootstrap equivalence analyses comparing DiffeoAfford with Junior (In-context) at the Β±0.5-s, Β±1.0-s, and Β±2.0-s gaze windows. Points and horizontal bars denote ν and its 90% image-level bootstrap CI, respectively. The shaded region denotes the equivalence interval from β1 to 1; equivalence was concluded when the entire CI lay within this interval. The data-driven equivalence margins were β=32.19,28.98,and 29.62 pixels for the Β±0.5-s, Β±1.0-s, and Β±2.0-s windows, respectively. CI, confidence interval; Adj. p, adjusted ν value; Q-Q, quantile-quantile. Supplementary Fig. 5 | Framework for laparoscopic eye-tracking data collection. a, Overview of the spatial and temporal alignment process linking the scene video from the wearable eye-tracker with the laparoscopic video record, which enables mapping of the surgeon's gaze points onto the laparoscopic video. b, The software developed for performing the spatial and temporal alignment of the videos. Supplementary Fig. 6 | User interface of the online multicenter questionnaire survey. a, The registration and instruction page outlining the evaluation criteria. b, The paired video evaluation page, where participants simultaneously viewed two synchronized video clips (adjusted via different methods) and indicated their preference across three specified dimensions. Displacement estimation: For each frame, a smooth, dense mapping is estimated to map pixels back to the query frame. This mapping is represented as a diffeomorphic deformation obtained by integrating a 2-D velocity field ν: ν ( ν ) =νΌ ( ν )( ν ) and a configurable global transform ν (implemented as either a similarity or a homographic transform) is optimized jointly. Taking the similarity transform as an example, it is formulated as: ν ( ν ) =ν βνΉ ( ν ) ν+ν where ν£ denotes the velocity field, ν is an isotropic scale factor, ν is the rotation angle, and ν is a translation vector. The deformation follows a Demons-style procedure 1 : at each optimization iteration, the velocity field ν is smoothed by a Gaussian filter and then integrated by repeated composition to yield a diffeomorphic sampling grid ν ( ν ) . Finally, the parameters are optimized by minimizing an energy ab function νΈ that measures the average positional error of the sparse tracked points: νΈ=min ν,ν ,ν,ν 1 ν οΏ½ οΏ½ ν οΏ½ ν οΏ½ ν ν ( ν ) οΏ½ βν ν ( ν ) οΏ½ 2 ν ν=1 where ν ν ( ν ) represents the ν-th tracked point in frame ν, ν ν ( ν ) is its corresponding location in the selected query frame, and ν is the total number of tracked points. Instrument-tracking baseline: To establish a baseline for view control, we implemented the instrument- tracking-based target prediction method. Based on previous researches 2-5 , the center of multiple instrument tips is commonly used as the target location. Therefore, let ν ν denote the coordinate of the ν-th detected instrument tip, the target center νΆ ν ν ν is defined dynamically based on the number of active instruments νΌ (Fig. 2): for a single instrument (νΌ=1), νΆ ν ν ν =ν 1 ; for two instruments (νΌ=2), νΆ ν ν ν is the midpoint ν 1 +ν 2 2 ; and for νΌβ₯3, νΆ ν ν ν is the centroid of the polygon formed by all tips. Validity of the two-dimensional diffeomorphic approximation: Intraoperative soft-tissue motion is a three-dimensional deformation ν·: ν β βΒ³ of the visible surface, whereas the pipeline estimates only its image-plane projection ν = ν β ν·. Modeling ν as a two-dimensional diffeomorphism is a deliberate approximation justified by the operative viewing geometry rather than by full 3D reconstruction. Decompose the surface displacement ν’(ν₯) = ν·(ν₯) β ν₯ into a component parallel to the image plane and a component along the optical axis, βν’βΒ² = βν’β₯βΒ² + βν’β₯βΒ². Only the out-of-plane component ν’ β₯ can create the depth-ordering reversals (self-occlusions) that would break a planar model; when βν’β₯β is small relative to βν’β₯β, the projected motion is well described by a smooth, invertible 2D map. The operative viewing geometry keeps ν’β₯ small. The in-plane component ν’β₯ is the portion of tissue motion actually visible in the image, whereas ν’ β₯ is largely invisible. To work under direct vision and maximize the manipulability and visibility of the target tissue, the surgeon orients the scope so that the relevant deformation unfolds within the image plane, maximizing βν’β₯β and hence suppressing βν’β₯β. The same clinical practice that makes a maneuver observable therefore suppresses the deformation component that would break the two-dimensional approximation. The argument is applied per contiguous operative segment. Within one maneuver the viewpoint and depth ordering are stable, so ν’ β₯ remains small; segment boundaries at which the viewpoint or topology changes (gross camera relocation, instrument-induced occlusion, tissue separation) are handled by the temporal segmentation rather than absorbed into a single field. Residual violations, such as transient folding or the independent motion of a secondary instrument, manifest as displacement outliers and are removed by the Mahalanobis filtering step. Modeling ν as a two-dimensional diffeomorphism is thus consistent with both the imaging geometry and the operative workflow, while avoiding the cost and static-scene assumptions of explicit three-dimensional reconstruction. References 1. Vercauteren T, Pennec X, Perchant A, Ayache N. Non-parametric Diffeomorphic Image Registration with the Demons Algorithm. Paper presented at: Medical Image Computing and Computer-Assisted Intervention β MICCAI 2007; 2007//, 2007; Berlin, Heidelberg. 2. Sun Y, Pan B, Zou S, Fu Y. Adaptive fusion-based autonomous laparoscope control for semi-autonomous surgery. Journal of medical systems. 2020;44(1):4. 3. Sandoval J, Laribi MA, Faure JP, BrΓ¨que C, Richer JP, Zeghloul S. Towards an Autonomous Robot-Assistant for Laparoscopy Using Exteroceptive Sensors: Feasibility Study and Implementation. IEEE Robotics and Automation Letters. 2021;6(4):6473-6480. 4. Li L, Li X, Ouyang B, Ding S, Yang S, Qu Y. Autonomous multiple instruments tracking for robot-assisted laparoscopic surgery with visual tracking space vector method. IEEE/ASME Transactions on Mechatronics. 2021;27(2):733-743. 5. Gruijthuijsen C, Garcia-Peraza-Herrera LC, Borghesan G, et al. Robotic endoscope control via autonomous instrument tracking. Frontiers in Robotics and AI. 2022;9:832208.