Paper deep dive
SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video
Guanyi Qin, Xiaozhen Wang, Zhu Zhuo, Chang Han Low, Yuancan Xiao, Yibing Fu, Haofeng Liu, Kai Wang, Chunjiang Li, Yueming Jin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 12:19:43 PM
Summary
The paper introduces ResGo, a multimodal benchmark for laparoscopic cholecystectomy featuring phase-dependent Go Zone annotations and clinician-authored rationales. It proposes SurGo-R1, a Vision-Language Model optimized via RLHF (GRPO) that employs a 'phase-then-go' architecture to improve contextual reasoning and safe zone identification, significantly outperforming generalist VLMs.
Entities (7)
Relation Signals (5)
SurGo-R1 → trainedon → ResGo
confidence 95% · Built on ResGo, a GRPO-optimized VLM with context-aware multi-turn reasoning, named SurGo-R1
SurGo-R1 → uses → RLHF
confidence 95% · SurGo-R1, a model optimized via RLHF with a multi-turn phase-then-go architecture
ResGo → supports → Cholecystectomy
confidence 93% · ResGo, a multimodal dataset for cholecystectomy curated in-the-wild
SurGo-R1 → outperforms → Vision-Language Models
confidence 92% · SurGo-R1 achieves ... a 6.6× improvement over the mainstream generalist VLMs.
ResGo → contains → Go Zone
confidence 90% · ResGo, a benchmark of laparoscopic frames annotated with Go Zone bounding boxes
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Minimally invasive surgery has dramatically improved patient operative outcomes, yet identifying safe operative zones remains challenging in critical phases, requiring surgeons to integrate visual cues, procedural phase, and anatomical context under high cognitive load. Existing AI systems offer binary safety verification or static detection, ignoring the phase-dependent nature of intraoperative reasoning. We introduce ResGo, a benchmark of laparoscopic frames annotated with Go Zone bounding boxes and clinician-authored rationales covering phase, exposure quality reasoning, next action and risk reminder. We introduce evaluation metrics that treat correct grounding under incorrect phase as failures, revealing that most vision-language models cannot handle such tasks and perform poorly. We then present SurGo-R1, a model optimized via RLHF with a multi-turn phase-then-go architecture where the model first identifies the surgical phase, then generates reasoning and Go Zone coordinates conditioned on that context. On unseen procedures, SurGo-R1 achieves 76.6% phase accuracy, 32.7 mIoU, and 54.8% hardcore accuracy, a 6.6$\times$ improvement over the mainstream generalist VLMs. Code, model and benchmark will be available at this https URL
Tags
Links
- Source: https://arxiv.org/abs/2602.21706v1
- Canonical: https://arxiv.org/abs/2602.21706v1
Trouble viewing inline? Open PDF directly →
Full Text
53,722 characters extracted from source content.
Expand or collapse full text
SurGo-R1: Benchmarking and Modeling Contextual Reasoning for Operative Zone in Surgical Video Guanyi Qin * 1 2 Xiaozhen Wang * 3 Zhu Zhuo * 1 Chang Han Low 1 Yuancan Xiao 3 Yibing Fu 1 Haofeng Liu 1 Kai Wang 3 Chunjiang Li 3 Yueming Jin 1 Abstract Minimally invasive surgery has dramatically im- proved patient operative outcomes, yet identify- ing safe operative zones remains challenging in critical phases, requiring surgeons to integrate visual cues, procedural phase, and anatomical context under high cognitive load. Existing AI systems offer binary safety verification or static detection, ignoring the phase-dependent nature of intraoperative reasoning. We introduce ResGo, a benchmark of laparoscopic frames annotated with Go Zone bounding boxes and clinician-authored rationales covering phase, exposure quality rea- soning, next action and risk reminder. We intro- duce evaluation metrics that treat correct ground- ing under incorrect phase as failures, revealing that most vision-language models cannot handle such tasks and perform poorly. We then present SurGo-R1, a model optimized via RLHF with a multi-turn phase-then-go architecture where the model first identifies the surgical phase, then generates reasoning and Go Zone coordinates con- ditioned on that context. On unseen procedures, SurGo-R1 achieves 76.6% phase accuracy, 32.7 mIoU, and 54.8% hardcore accuracy, a 6.6×im- provement over the mainstream generalist VLMs. Code, model and benchmark will be available athttps://github.com/jinlab-imvr/ SurGo-R1. 1. Introduction Minimally invasive surgery (MIS) has demonstrably im- proved patient outcomes and care quality, offering patients reduced postoperative pain, shorter hospital stays, and faster recovery compared to open surgery approaches (Velanovich, 1 National University of Singapore, Singapore 2 Guangzhou Re- search Translation and Innovation Institute, National University of Singapore, China 3 Southern Medical University, China. Corre- spondence to: Yueming Jin<ymjin@nus.edu.sg>. Preprint. February 26, 2026. 2000). Although enhanced visualization through high- resolution imaging is available, the inherent complexity of human anatomy still increases the demands on surgeons, im- posing a heavier intraoperative cognitive load and requiring substantially more training and practice (Way et al., 2003; Zheng et al., 2010; Zegers et al., 2011; Magrina, 2002). As one of the most crucial MIS procedures, cholecystectomy, for instance, still imposes significant cognitive demands de- spite its relatively standardized workflow. Performed over two million times annually in the U.S. and Europe alone (Pucher et al., 2018; Peery et al., 2022), it carries a persis- tent risk of bile duct injury (BDI) which might result in a 2.5-fold increase in long-term mortality and often requires multiple reconstructive surgeries (Brunt et al., 2020). Criti- cally, most of these injuries stem from visual misperception (Way et al., 2003; Strasberg & Brunt, 2010): surgeons tran- sect what they believe to be the cystic duct, only to discover that they have injured the common bile duct. This nature of struggling to interpret when anatomical landmarks are obscured by inflammation or aberrant anatomy, especially under heavy cognitive load, highlights the need for training and intraoperative guidance. Several core concepts have been proposed to ensure sur- gical safety in the community. The most widely adopted strategy is to require surgeons to fulfill anatomical criteria before operation, such as Critical View of Safety (CVS) for cholecystectomy (Brunt et al., 2020). While this has shown promise, its effectiveness depends on subjective vi- sual assessment under challenging conditions. Recently, the concept of safety zones emerged, prioritizing a more di- rect identification of current safe operating regions (Madani et al., 2022). This highlights the need to develop systems that can provide context-aware support to surgeons. Artifi- cial Intelligence (AI) is increasingly transforming surgical care, and the emergence of AI copilots is opening up new possibilities. Conceived as cognitive collaborators for sur- gical education and intraoperative assistance, such systems can augment surgeons’ expertise for higher quality. Early works primarily focused on binary safety verification, such as automated assessment of the CVS (Mascagni et al., 2022) and subsequent graph-based extensions (Murali et al., 1 arXiv:2602.21706v1 [cs.CV] 25 Feb 2026 ResGo & SurGo-R1 1.Preparation Adhesion, Serosa... Perforation, Thermal injury... Grasp and retract gallbladder, Adjust traction to expose triangle, Lyse adhesions... 2.Calot’s Triangle Dissection Fat, Connective tissue... Bile duct injury, Bleeding... Dissect fat and connective tissue, Skeletonize cystic duct/artery, Achieve Critical View of Safety... 3.Clip and Divide Cystic duct, Cystic artery... Bile leakage, Duct stricture... Confirm safety view, Apply clips on duct and artery, Transect with cold scissors... 4.Gallbladder Dissection Serosa, Liver bed plane... Liver bleeding, Perforation... Identify correct plane, Coagulate attachments, Separate from liver bed progressively... 21 patients ResGoDataset Phase-wiseTissue-specificRisk-informedScene-descriptive 8.53-hour Raw VideoFully Human-annotated by Experts Figure 1. Demo of the proposed ResGo dataset, a novel multimodal benchmark that covers phase recognition, Go Zone grounding, and safety reasoning. Built on ResGo, a GRPO-optimized VLM with context-aware multi-turn reasoning, named SurGo-R1, is further introduced, enabling safety-aware reasoning and Go Zone grounding across scenarios. 2023). These approaches demonstrated the feasibility of AI-assisted safety evaluation but offered limited insight into where safe dissection could occur. Building on this founda- tion, researchers (Madani et al., 2022) pioneered the explicit visual grounding of safe zones, reframing surgical safety as a spatial understanding problem. By localizing safe dissec- tion regions within the operative field, this work marked a key transition from binary verification to visual task formu- lation. Subsequent studies validated this paradigm against expert consensus (Laplante et al., 2023) and demonstrated its relevance in real bile duct injury cases (Khalid et al., 2023). Although spatial grounding marks a major advance over binary safety checks, existing methods exhibit limited generalization beyond the specific phase of dissection, and primarily produce basic visual recognition results, lacking the proactive, explainable capabilities to effectively copilot with surgeons. To bridge this gap, we propose to leverage the ability of Vision-Language Models (VLMs) to connect natural lan- guage with complex visual scenes (Liu et al., 2023) and reason safety zones, which can augment human decision- making by advancing from simple visual recognition to proactive and explainable prediction (Li et al., 2023; Jin & Jeong, 2024; Zeng et al., 2025). We, as the first, introduce ResGo, a multimodal dataset for cholecystectomy curated in-the-wild, that provides explainable Go Zone annotations to support understanding. The dataset is organized around clinically meaningful operative moments and offers hierar- chically structured supervision. Specifically, each sample is spatially grounded with Go Zone bounding boxes. More- over, this spatial evidence is paired with clinician-authored, interpretable annotations, including the ongoing surgical phase, a concise description of the current operative ac- tion, key safety points that emphasize anatomical risks and disallowed maneuvers, and an explicit next-step statement. This design enables models to learn not only where the Go Zone lies, but also why it is considered safe and how the operative plan should proceed, facilitating transparency of safety-aware guidance. The rich nature of information of ResGo enables comprehensive support for modern VLMs development, accommodating diverse training and evalua- tion patterns including supervised fine-tuning and RLHF, furthering agentic workflows, while remaining faithfully aligned with real-world practice and surgeon experience. Building on ResGo, a baseline named SurGo-R1 is also pro- posed. Given an intraoperative image, SurGo-R1 is trained with GRPO to generate structured, clinically grounded out- puts that operationalize Go Zone understanding beyond pixel classification. Concretely, the model reasons about four complementary components: (1) Location, a textual localization that describes where the Go Zone lies relative to salient anatomical landmarks; (2) Exposure, an assessment of whether current retraction and dissection have achieved sufficient visualization for safe progression; (3) Next Action, the immediate recommended maneuver consistent with safe advancement toward CVS; and (4) Critical Risk, the pri- mary source of potential misperception or injury risk in the current context. Across held-out procedures, SurGo- R1 achieves strong performance in both spatial localization and structured reasoning tasks. In summary, our primary contributions are as follows: • We introduce ResGo, the first benchmark for cholecys- tectomy that pairs Go Zone localization with clinician- authored rationales for phase-dependent surgical reason- ing annotations, aligning perception for safety-aware learning. • Upon ResGo, a novel practice-aligned contextual reason- ing pattern as phase-then-go is proposed, accompanied by corresponding evaluation protocols, promoting agen- tic or complex reasoning research. •We propose SurGo-R1, a reasoning-based VLM opti- mized using GRPO that goes beyond static detection by producing structured, interpretable guidance, reflecting how surgeons assess and decide safe actions. • SurGo-R1 substantially improves generalization on held- out procedures, achieving 76.6% phase accuracy and 32.7 mIoU, and outperforming generalist VLM base- 2 ResGo & SurGo-R1 lines with 6.6×. 2. Related Work Surgical Safety Intelligence Bile duct injury during laparo- scopic cholecystectomy predominantly results from visual misperception rather than technical error (Way et al., 2003), motivating computational approaches to augment surgeon awareness. The CVS framework provides anatomical crite- ria for safe clip application, with DeepCVS (Mascagni et al., 2022) pioneering automated verification and LG-CVS (Mu- rali et al., 2023) enhancing robustness through graph-based anatomical reasoning. The Endoscapes dataset (Mascagni et al., 2025) further advanced the field with 201 densely annotated cholecystectomy videos. Complementing CVS assessment, Madani et al. (2022) introduced the zone-based paradigm, reframing surgical safety as spatial localization of safe versus dangerous dissection zones, with subsequent validation against expert panels (Laplante et al., 2023) and on actual injury cases (Khalid et al., 2023). However, these methods operate as fixed-output systems requiring explicit labels for each target category, unable to interpret implicit safety queries or provide knowledge-grounded explanations. Reasoning Grounding Vision-language models have en- abled sophisticated connections between natural language and visual content (Liu et al., 2023; Dai et al., 2023). Med- ical adaptations (Li et al., 2023; Jin & Jeong, 2024; Zeng et al., 2025) show promise on clinical visual QA but gener- ate purely textual outputs without spatial grounding. Visual grounding approaches have evolved from phrase-based lo- calization (Plummer et al., 2015) to VLM-integrated meth- ods producing bounding boxes (Peng et al., 2023; Chen et al., 2023) or region-level understanding (You et al., 2023; Rasheed et al., 2024). The reasoning grounding paradigm (Lai et al., 2024) advances further by handling implicit queries requiring world knowledge, with extensions en- hancing multi-target capabilities (Yang et al., 2023; Ren et al., 2024). Emerging medical reasoning grounding ef- forts (Huang et al., 2025; Tong et al., 2025; Liu et al., 2025) target diagnostic imaging rather than procedural guidance. Our work introduces reasoning grounding to intraoperative surgery, where the model localizes safe dissection regions while generating interpretable responses. 3. ResGo Benchmark To advance surgical reasoning capabilities, we introduce ResGo dataset. Unlike existing in-domain benchmarks lim- ited to basic perception tasks such as phase recognition (Nwoye et al., 2023; Twinanda et al., 2016), ResGo anno- tates the underlying rationale behind safe dissection and Go Zone identification. By capturing this cognitive process alongside visual grounding, the dataset supports both intra- operative guidance and surgical education, demonstrating not only what is seen but how safe surgery is performed. 3.1. Source Acquisition and Demographics To ensure clinical grounding and real-world in-vivo vari- ability, we curated 21 laparoscopic cholecystectomy videos from collaborating institutions. Procedures were recorded with informed consent and ethical approval for education, training, and research purposes using 4K imaging systems (Olympus, Storz, and Huanuokang) to preserve fine-grained anatomical details. In accordance with the Declaration of Helsinki, all videos were de-identified to protect patient privacy 1 . Raw footage was standardized to1920× 1080 resolution and sampled at 0.2 fps, yielding 6,138 frames. Notably, the dataset encompasses diverse patient demo- graphics, as shown in Fig. 2. Patients are predominantly female (70%), with ages ranging from 30 to over 70 years (peak∼55 years) and BMI from 14 to 27, consistent with typical in-the-wild cholecystectomy populations. The co- hort also spans varied clinical presentations: 85% were ASA Class 2 (mild systemic disease), with 6 of 21 patients presenting pre-existing conditions including hypertension, diabetes, and viral hepatitis (4 cases). One patient presented with concurrent acute pancreatitis. Clinical diagnoses such as gallstones with chronic cholecystitis were also recorded. By incorporating such a diverse array of patient demograph- ics, comorbidities, and pathological variations, the dataset is positioned to expose models to the heterogeneity inherent in routine clinical practice, thereby establishing a robust foundation for training VLMs capable of adapting to the unpredictable nature of live surgery. 3.2. Annotation Framework and Pipeline To ensure clinical relevance and precision, a rigorous annota- tion pipeline was implemented by six recruited hepatobiliary surgeons. Prior to the commencement of detailed labeling, a preliminary screening was performed to filter out frames deemed operationally irrelevant, and, to ensure procedu- ral consistency, the temporal boundaries of each video clip were standardized to commence with the preparation of Go Zone exposure and conclude with the gallbladder dissection from the liver bed. Consequently, a curated set of 2,686 high-quality frames was yielded for subsequent annotation. Before the annotation process was initiated, a consensus meeting was convened among the experts to establish uni- fied annotation protocols. Particular emphasis was placed on standardizing the criteria for the Go Zone, which, after consensus and agreement, defined as the optimal region for surgical operation within the current field of view in strict adherence to safety guidelines, while, for the preparation 1 Ethical approval details will be provided upon publication. 3 ResGo & SurGo-R1 (a) Gender Distribution(b) Age Distribution(c) ASA Distribution (d) BMI Distribution(e) Distribution of Comorbidities & Medical History(f) Diagnosis Occurrence Figure 2. The left panels detail patient demographics and clinical profiles, including distributions for gender, age, ASA classification, BMI, and relevant medical history. The right panels present the processed surgical phase statistics, illustrating duration distributions, occurrence frequencies, and the phase transition matrix for the four defined phases. Table 1. Annotation Statistics and Properties. 2,686 annotated samples in total, where each sample consists of a structured in- struction spanning four dimensions. desc. for description. AttributeSpecification Data Acquisition EquipmentLaparoscope (Olympus/Storz/H.) Video Resolution1920× 1080 (Downsampled) FPS0.2 fps Total Videos21 (Cholecystectomy) Total Raw Frames6,138 Avg. Duration ∼ 25 mins / video Annotation Scope Annotators3 Senior Surgeons (>3 years exp.) Reviewers3 Chief Surgeons (>15 years exp.) Total Annotated2,686 Samples Annotation Details (4 Dimensions per Sample) a. Surgical PhaseCurrent procedural phase b. Anatomic(1) Text: Grounding desc. of Go Zones (2) Visual: LabelMe Bounding box c. ReasoningAnalysis upon quality of exposure d. PlanningNext-step & critical risk reminders or exposure stage, the Go Zone is defined as the targeted anatomical region requiring clearance to facilitate the safe identification. Subsequently, the workload was distributed by video to three senior surgeons, each possessing over three years of experience. Following this allocation, the annotation was conducted across four distinct dimensions in accordance with the established pipeline, as illustrated in Fig. 3 as: a) Surgical Phase Identification: Selected frames were classified into specific procedural phases, i.e., Phase A - Preparation of the Go Zone, Phase B - Dissection of Phase: Dissection of Calot's Triangle. Reasoning: The Go Zone is located within the Calot’s triangle, specifically targeting the fatty and connective tissue overlying the cystic duct and cystic artery .... Multi-modal Annotation Junior Surgeon 1 Phase: Dissection of Calot's Triangle. Reasoning: The Go Zone is located within the Calot’s triangle, specifically targeting the fatty and connective tissue overlying the cystic duct and cystic artery .... Multi-modal Annotation Junior Surgeon 1 Stage I: Frame Filtering .... .... Phase: Dissection of Calot's Triangle. Reasoning: The Go Zone is located within the Calot’s triangle, specifically targeting the fatty and connective tissue overlying... Planing: Continue blunt dissection... Multi-modal Annotation Senior Surgeon Stage I: Annotation by 3 senior surgeons (per video) Stage I: Consensus-based review by 3 chief surgeons Chief Surgeon Consensus Agreement? Individual Annotations ✅ or ❌ Grounding Gallbladder Dissection Preparation of Go Zone Figure 3. Illustration of the annotation and review pipeline Calot’s Triangle, Phase C - Clip and Divide, and Phase D - Gallbladder Dissection from Liver Bed. b) Grounding of the Go Zone: Bounding boxes were drawn around safe operation areas based on clinical ex- perience adhere to the agreed definition. These annota- tions were executed using theLabelMetool. Further, textual anatomic descriptions of the Go Zone are also included. c)Exposure Quality Reasoning: Reasoning regarding the quality of Go Zone exposure was conducted based on visual cues from specific anatomical tissues. d)Planning - Action & Risk Reminders: Subsequent procedural steps were formulated at a fine-grained level based on temporal context and surgeon judgment. For instance, specific actions, such as applying traction to facilitate exposure, were annotated to reflect intraoper- ative decisions, accompanied by reminders of potential risks arising from incorrect maneuvers. For the textual components (Phase, Anatomic, Reasoning, 4 ResGo & SurGo-R1 and Planning), data were recorded via Excel and indexed by video frame number to ensure alignment with the visuals. A detailed summary is provided in Tab. 1. The annotation process was conducted on the 2,686 previously selected frames, yielding corresponding total samples. To guaranty data quality for benchmarking and maintain high consensus, a rigorous quality control architecture was adopted. Following the initial annotation, a mutual cross- check was first performed by the annotating surgeons for a format check. Subsequently, a comprehensive sample- by-sample review of the entire dataset was conducted by three chief surgeons, each possessing over 15 years of expe- rience. Frames were included only upon unanimous agree- ment among reviewing surgeons. After the review, 90 frame samples of 4 videos were identi- fied for a refinement process. These adjustments were neces- sitated by discrepancies including, bounding boxes slightly exceeding anatomical boundaries and reasoning annotations that were deemed fragmented, requiring a modification into a structured format. Following refinement, full consensus was achieved among all senior and chief surgeons without the need for further review, thereby ensuring the accuracy and adherence to surgical protocols. 3.3. Analysis and Splits of Annotated Samples In this section, the dataset is analyzed based on the four annotated phases (A through D), with detailed distributions presented in Fig. 2. It is observed that the preparation phase exhibits the shortest duration. Given that the exposure of the Go Zone in cholecystectomy is considered procedurally straightforward compared to other complex surgeries, this relative brevity is expected and is found to be aligned with routine clinical practice. Valuable insights are further derived from the transition ma- trix. While the preparation phase is consistently followed by the dissection of Calot’s triangle, an alternating pattern is fre- quently observed between the dissection of Calot’s triangle and the clip and divide phase. This suggests that the pro- cedural workflow is modulated by surgical complexity and the surgeon’s execution strategy. While a linear progression may be achieved in standard cases or by highly experienced practitioners, this iterative alternation reflects the adaptive measures required when increased difficulty is encountered. This variability highlights the necessity of the proposed Rea- soning Go, which is intended to assist surgeons by reducing intraoperative cognitive burden and facilitating a more fluid surgical workflow or provide education. Finally, it is observed that, in terms of duration, the gallblad- der dissection phase records the longest average duration. Conversely, significant variance is noted in the dissection of Calot’s triangle, where the presence of lower extreme values suggests that rapid phase transitions exist in scenarios. Train and Test Splits The dataset was partitioned into training and testing splits at the video level to ensure robust evaluation on unseen patients. Specifically, 4 of the 21 videos were allocated to the testing set, comprising a total of 495 samples. The remaining 17 videos comprising 2,191 frames were designated for the training set. 3.4. Benchmark Problem Formulation Instead of modeling Go Zone reasoning as a static ground- ing task, it is formulated as a conditioned derivative of the evolving surgical context. Precise phase identification serves as a contextual anchor rather than a final output, fa- cilitating phase-conditioned reasoning to deliver effective intra-operative guidance, since, crucially, a failure in phase recognition renders any downstream grounding or reasoning futile, since a Go Zone predicted under an incorrect phase assumption is clinically meaningless. To enforce this dependency and align with the cognitive workflow validated by surgeons, the benchmark is struc- tured not as a one-shot reasoning task, but as a sequential decision process unfolding as phase-then-go. This design necessitates that the phase-level understanding derived from the initial inference is explicitly incorporated into the subse- quent Go Zone reasoning as: P (b,p|I) = P (p|I) |z Phase Recognition ·P (b|I,p,D(p)) |z Contextual Reasoning ,(1) wherebdenotes the Go Zone reasoning andIrepresents the input frame.D(p)signifies the Go Zone definition and operative criteria corresponding to phasep. This formula- tion transforms the problem from implicit pattern matching into a structured inference task, ensuring that the Go Zone is strictly derived from the distinct anatomical rules of the identified phase. 3.5. Evaluation Protocols A hierarchical evaluation strategy is adopted to reflect the sequential nature of the benchmark. While standard ground- ing metrics, such as accuracy at IoU thresholds (Acc@ 0.25 , mA@ 0.25:0.5 ), center distance normalized by image diago- nal (∆ cen ), and mean IoU, are reported, they are deemed in- sufficient where valid grounding is contingent upon correct phase identification. To address this, conditioned metrics are introduced, where spatial reasoning is evaluated solely on samples with correctly predicted phase, to simulate the scenarios when phase metadata is provided for educational purposes. Furthermore, hardcore metrics (HA 0.25 , HmIoU) are proposed to measure end-to-end reliability. In this frame- work, grounding outputs associated with erroneous phase identification are penalized as invalid regardless of spatial 5 ResGo & SurGo-R1 SurGo-R1 Task 1: Phase Recognition A. Preparation of the Go Zone B. Dissection of Calot's Triangle ✅ C. Clip and Divide Cystic Duct and Artery D. Gallbladder Dissection from the Liver Bed Location: The Go Zone is the fibro-fatty tissue within Calot‘s Triangle, specifically around the cystic duct and artery, which are being dissected. The instrument is engaged in this area, and the surrounding tissue appears to be fibro-fatty in nature..... Safety-aware Reasoning Task 2: Grounding Phase-Definition Mapping Tool Go Zone Definition & Reasoning Instruction Generation Target the Safety Window.The Go Zone is strictly limited to the fibro- fatty tissue within the triangle. Please provide safety-aware reasoning about Go Zone. Turn 푰 Turn 푰 Figure 4. Overview of SurGo-R1 and phase-then-go pipeline. accuracy, thereby capturing the joint success of phase recog- nition and Go Zone. For the explicit assessment of rea- soning capabilities, a small-scale, double-blind review by surgeons is recommended, wherein generated outputs are scored based on clinical correctness. 4. SurGo-R1 Baseline Methodology 4.1. Overview of SurGo-R1 To practically instantiate the proposed phase-then-go logic, SurGo-R1, a multi-modal reasoning framework optimized via GRPO, is introduced as a baseline to serve as a reference standard for the community and to facilitate future method- ological advancements. In Turn 1, the model answers a phase Multiple Choice Question (MCQ), compelling the model to prioritize procedural context as a prerequisite be- fore reasoning. In Turn 2 then performs reasoning and grounding conditioned on the predicted phase with tool- provided definitions, as in Fig. 4 4.1.1. PREREQUISITE TURN: CONTEXTUAL PRIMING In the first turn, we establish the procedural context by pre- senting the model with a surgical scene frame and candidate phase labels. Framed as an MCQ, the model identifies the current phase based on visual cues. To incentivise precise phase recognition, we train with GRPO using a strict binary accuracy reward:R acc = ( 1 if ˆp = p 0 otherwise , where the re- ward is 1 if the predicted phaseˆpmatches the ground truth p, and 0 otherwise. Having model explicitly training with phase MCQ ensures the model prioritizing surgical phase recognition as a foundation. 4.1.2. REASONING TURN: SURGICAL ANALYSIS Given the predicted phaseˆp, the model proceeds to Go Zone reasoning and grounding. Following an agentic workflow, it invokes a phase-definition mapping tool that retrieves procedure-specific constraintsD ˆp , providing structured con- text for the visual input. Guided by this definition, the model generates a structured outputb, comprising a reasoning ra- tionale across four dimensions related to surgical safety, as in (1) Location relative to anatomical landmarks, (2) Ex- posure adequacy, (3) Next action, and (4) Critical risk, alongside the final Go Zone coordinate prediction, ground- ing the spatial output in a structured surgical assessment. Phase-Definition Mapping Tool The agentic workflow is implemented through a dynamic contextual prompting mechanism that functions as an intermediate supervisor, with distinct behaviors adopted for training and inference. During the training stage, the predicted phaseˆpis compared against the ground truthy. To ensure optimal conditioning for the subsequent reasoning task, any classification errors are rectified, and the definitionD y corresponding to the correct phase is injected into the prompt. Conversely, during the inference stage, the process operates autonomously: the definitionD ˆp is dynamically retrieved and inserted based solely on the model’s prediction, regardless of its accuracy. Content of this tool is provided in the appendix. This differential strategy provides two complementary bene- fits: (a) Intensified Spatial Alignment: By rectifying phase errors during training, it is ensured that the reasoning mod- ule is consistently exposed to valid procedural contexts. This isolates spatial learning from classification noise, thereby enabling the model to master the challenging text-to-image grounding task without interference from upstream errors. (b) Adaptive Context Integration: This mechanism trains the model to strictly condition its spatial reasoning on the provided textual definition. Consequently, during inference, the model is capable of generating grounding outputs that are logically consistent with its own phase predictions. Reward Modeling In the reasoning turn, we continue to employ GRPO to optimize the generation of the clinical rationale, covering the above-mentioned four dimensions, and Go Zone grounding. We design a composite reward function to provide supervision signals while balancing com- putational efficiency. Reasoning Reward: To align rationales with clinical stan- dards, we introduce a lightweight reward based on seman- tic entity matching. Instead of rigid traditional template comparisons, we employ scispaCy to dynamically extract core entities from the ground truth, including surgical tar- gets, operative actions, and safety constraints. The reward computes recall of these entities in the generated output: R reason = |w|w∈W∩K| |K| , whereKis the set of extracted entities andWthe unique tokens in the prediction. This reward serves as a semantic anchor during GRPO training, incentivizing the model to integrate key clinical concepts into rationales, without the overhead of LLM-based reward models. 6 ResGo & SurGo-R1 Table 2. Quantitative results on the ResGo dataset. The best results are highlighted in bold. Method PhaseGroundingConditionedHardcore AccAcc@ 0.25 mA@ 0.25:0.5 ∆ cen ↓mIoUCA 0.25 CA 0.25:0.5 C∆ cen ↓CmIoUHA 0.25 HmIoU Generalist VLMs Qwen3-VL-4B-Ins36.514.77.6111.712.321.511.58.8415.87.885.79 Qwen3-VL-8B-Ins 39.316.97.8310.913.214.86.249.8312.55.864.94 Qwen3-VL-32B-Ins53.116.38.019.2914.215.57.679.2613.58.287.20 Qwen3-VL-30B-A3B-Ins51.911.95.3512.610.712.85.5111.711.46.675.95 Qwen3-VL-4B-Thi31.53.842.0520.83.716.412.8820.25.462.021.72 Qwen3-VL-8B-Thi34.112.35.7916.78.0412.46.2118.27.184.242.45 InternVL3.5-8B33.57.273.3717.05.816.633.1115.96.212.222.08 InternVL3.5-14B35.34.852.1216.25.465.141.6214.76.331.822.24 Ovis2.5-2B22.61.010.3424.11.290.970.3322.61.040.270.39 Ovis2.5-9B42.81.820.5715.12.081.420.4714.42.050.610.88 Specialist Models EndoChat30.50.810.1721.02.231.320.3320.72.460.400.75 Hulu-Med-7B37.54.041.5510.58.035.382.0610.38.022.023.01 Hulu-Med-14B31.30.810.2410.95.050.650.1111.24.680.201.46 QoQ-Med3-VL-8B37.715.95.9310.212.113.75.449.2512.45.054.69 SurGo-R176.668.339.74.1132.771.540.93.6333.854.825.9 Grounding Rewards: To supervise Go Zone localization, we combine two reward functions. The primary objec- tive is standard Intersection-over-Union (IoU) for bound- ary alignment:R IoU = |b pred ∩b gt | |b pred ∪b gt | . However, relying solely on IoU presents a significant drawback in the context of surgery where early predictions may not overlap with the ground truth at all, yielding zero reward and vanishing gra- dients. To address this, we introduce an auxiliary Center Distance Reward that provides dense supervision by measur- ing proximity between predicted and ground-truth centers: R dist = exp − |c pred −c gt | 2 τ , whereτis a scaling factor adjusted by the grounding nature of VLMs. This term en- courages the model to shift toward the correct anatomical region even without precise overlap, maintaining gradient flow and stabilizing GRPO training. 4.2. Implementation Details Qwen3-VL-8B-Instruct is utilized as the foundation VLM, and a two-stage training pipeline tailored to the sequential task structure is followed. Stage 1: Phase Recognition. The model is trained exclusively on the phase identifica- tion turn using GRPO with 4 generations, a temperature of 0.9, a learning rate of2× 10 −6 , a batch size of 16, and 10 epochs. The reference penaltyβis set to 0.01. Rewards are comprised ofR acc and a format reward, with loss applied only to the first turn to enforce phase recognition. Stage 2: Multi-Turn Reasoning. The full pipeline is then trained with rewardsR acc ,R reason ,R IoU ,R dist , and a format reward. τ = 100is set to match Qwen3’s scale of grounding con- vention. GRPO training is conducted using 4 generations andβ = 0.01over 80 steps. Loss is applied across the full trajectory to reinforce the dependency between phase recog- nition and reasoning. Sequence-level importance sampling is applied throughout. All experiments are executed on 4 NVIDIA H200 GPUs, with 1 GPU allocated for vLLM- based rollout sampling and 3 GPUs utilized for updates. 5. Experiments 5.1. Overall Performance Comparison To evaluate the capabilities of contemporary VLMs (Bai et al., 2025; Wang et al., 2025b; Lu et al., 2025; Wang et al., 2025a; Jiang et al., 2025; Dai et al., 2025) on the proposed benchmark, comprehensive inference experiments were con- ducted. It is noted that the evaluation protocol is strictly aligned with the hierarchical phase-then-go formulation. This includes the utilization of the definition-tool, ensuring that the reasoning process during inference is structurally consistent with the proposed framework. As evidenced in Tab. 2, significant deficiencies are observed in existing generalist and specialist models regarding their capacity to align procedural phase knowledge with Go Zone grounding and safety reasoning. These limitations, particu- larly in ensuring consistency, underscore the critical value and need of the proposed ResGo dataset for these gaps. Conversely, the substantial performance gains achieved by SurGo-R1 validate the feasibility of this challenging task, thereby establishing a foundation to incentivize research. 5.2. Qualitative Analysis and Reasoning Study Visualizations and qualitative results SurGo-R1 generated are in Fig. 5. Noted that SurGo-R1 generates better rea- soning answers with more comprehensive descriptions of locations and risk explanations, while w/o Rea tends to provide less informative and shorter answers. 7 ResGo & SurGo-R1 Ours Qwen3 w/o Rea GT: Location:The current focus is within Calot’s Triangle, specifically at the junction of the cystic duct, cystic artery, surrounded by fatty and fibrous connective tissue. Exposure: Retraction is adequate, with the gallbladder being pulled toward the patient’s right upper quadrant, exposing partial visualization of the cystic duct and cystic artery; the duct is largely bare but requires further clearance ... w/o Reason: Location: The Go Zone is located around the cystic duct and gallbladder neck, within the fibro- fatty tissue of Calot's triangle. Exposure: The retraction appears sufficient to expose the area ... Qwen3: Location: The Go Zone is located at the area where the gallbladder is being dissected from the liver, which is the white/translucent line of adhesion. Exposure: The exposure is sufficient as the gallbladder is clearly visible and the liver bed is exposed ... Ours: Location: The Go Zone is the fibro-fatty tissue within Calot's Triangle, specifically around the cystic duct and artery, which are being actively dissected. The area is bounded by the gallbladder neck, the liver bed, and the cystic duct. Exposure: The retraction is sufficient as the gallbladder and surrounding structures are clearly visible .... GT: Next Action: Apply a laparoscopic clip securely and perpendicularly to the cystic duct, ensuring complete closure and a safe distance from the common bile duct. Critical Risk: Avoid misidentifying the common hepatic duct as the cystic duct or placing the clip too close to the bile duct, which may cause bile duct injury or stricture. w/o Reason: Next Action: The next step would be to divide the gallbladder neck after confirming the clip's secure placement. Critical Risk: Misidentifying the structure or clipping the wrong vessel/duct, which could lead to bleeding or bile leakage. Qwen3:Next Action: Divide the cystic duct with the instrument. Critical Risk: Accidentally clipping the common bile duct (CBD) instead of the cystic duct. Ours: Next Action: The immediate next action would be to secure the clip firmly and then divide the cystic duct. This step is critical to ensure that the gallbladder can be safely removed without causing bile duct injury. Critical Risk: The primary risk of error is misidentifying the cystic duct/artery and clipping the common bile duct, which could lead to bile duct injury. This is a critical risk that must be avoided to prevent complications. Figure 5. Qualitative results of reasoning and grounding. Models trained withoutR reason are denoted as w/o Rea or w/o Reason. Yellow and green bounding boxes represent model predictions and ground truths, respectively. Table 3. Ablations on the ResGo dataset.✓indicates inclusion in the training process, while✗ denotes exclusion. R dist Def-Tool GroundingHardcore Acc@ 0.25 ∆ cen ↓mIoUHA 0.25 HmIoU ✗56.35.7723.243.516.9 ✓✗ 58.45.2323.644.217.5 ✗✓67.84.3532.153.125.0 ✓w/o-rectify64.74.5830.348.721.4 ✓68.34.1132.754.825.9 Table 4. Comparison between Single- and Multi-turn performance. Method PhaseGroundingHardcore AccAcc@ 0.25 ∆ cen ↓mIoUHA 0.25 HmIoU Single-turn69.566.94.2432.149.622.7 Multi-turn (Ours)76.668.34.1132.754.825.9 Table 5. Ablation study on the reasoning reward (R reason ). SettingSelected Ratio (%)↑Acc. Rating↑ Qwen3-VL-8B-Ins0.2710.7 w/oR reason 17.347.2 w/R reason (Ours)79.952.5 To further evaluate the impact ofR reason , reasoning out- puts on the whole test set were collected and de-identified (A/B/C) for a blind review by three clinicians to review se- lect rate, while accuracy rating was noted based on factual correctness (0 for errors/hallucinations, 1 otherwise). As observed in Tab. 5, SurGo-R1 is preferred more often and achieves higher ratings, confirming the reasoning reward. 5.3. Ablation Study To investigate the effectiveness of the proposed components, ablation studies were performed on the ResGo dataset, with results summarized in Tab. 3. It is observed that the in- troduction of the distance reward (R dist ) yielded improve- ments in spatial precision. A more substantial gain was ob- served with the integration of the Phase-Definition Mapping Tool. By explicitly retrieving anatomical definitions based on phases, the semantic gap between visual features and medical knowledge was effectively bridged. Furthermore, the analysis highlights the critical role of the rectification mechanism since its exclusion during training was found to degrade performance by subjecting the grounding module to erroneous procedural contexts. Consequently, the opti- mal performance was achieved when both components were employed, validating the proposed design. Furthermore, the efficacy of the proposed multi-turn frame- work was evaluated against a single-turn baseline, as de- tailed in Tab. 4. To ensure a fair comparison, both models were initialized from the checkpoint following the Stage 1 MCQ training. It is evident that the multi-turn approach outperforms the single-turn counterpart across all metrics, given that the single-turn model attempts to simultaneously resolve phase identification and spatial grounding, which hinders model from learning. 6. Conclusion To address the demand for explainable guidance in MIS, ResGo is introduced as the first multimodal cholecystec- tomy dataset pairing spatial Go Zone localization with clin- ical rationales in this paper. Within this framework, surgi- cal safety and grounding are reformulated as a contextual reasoning task, where perception is integrated with phase- 8 ResGo & SurGo-R1 dependent annotations of risks and procedures. Building on this, SurGo-R1 is proposed, a reasoning-based VLM optimized via GRPO to generate interpretable guidance and ground Go Zones. By transitioning to a multi-turn reason- ing pattern, the gap between visual perception and complex surgical decision-making is effectively bridged, establishing a new foundation for advancing surgical intelligence. Impact Statement This work advances the field of surgical AI, with primary benefits targeting clinical decision support. We acknowl- edge that rigorous validation is a prerequisite for further development and deployment. Regarding societal impact, this specific study does not introduce unique ones necessi- tating distinct discussion. References Bai, S., Cai, Y., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., Ge, C., Ge, W., Guo, Z., Huang, Q., Huang, J., Huang, F., Hui, B., Jiang, S., Li, Z., Li, M., Li, M., Li, K., Lin, Z., Lin, J., Liu, X., Liu, J., Liu, C., Liu, Y., Liu, D., Liu, S., Lu, D., Luo, R., Lv, C., Men, R., Meng, L., Ren, X., Ren, X., Song, S., Sun, Y., Tang, J., Tu, J., Wan, J., Wang, P., Wang, P., Wang, Q., Wang, Y., Xie, T., Xu, Y., Xu, H., Xu, J., Yang, Z., Yang, M., Yang, J., Yang, A., Yu, B., Zhang, F., Zhang, H., Zhang, X., Zheng, B., Zhong, H., Zhou, J., Zhou, F., Zhou, J., Zhu, Y., and Zhu, K. Qwen3-vl technical report, 2025. Brunt, L. M., Deziel, D. J., Telem, D. A., Strasberg, S. M., Aggarwal, R., Asbun, H., Bonjer, J., McDonald, M., Al- seidi, A., Ujiki, M., et al. Safe cholecystectomy multi- society practice guideline and state-of-the-art consensus conference on prevention of bile duct injury during chole- cystectomy. Surgical endoscopy, 34(7):2827–2855, 2020. Chen, K., Zhang, Z., Zeng, W., Zhang, R., Zhu, F., and Zhao, R. Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195, 2023. Dai, W., Li, J., Li, D., Tiong, A., Zhao, J., Wang, W., Li, B., Fung, P. N., and Hoi, S. Instructblip: Towards general- purpose vision-language models with instruction tuning. Advances in neural information processing systems, 36: 49250–49267, 2023. Dai, W., Chen, P., Ekbote, C., and Liang, P. P. Qoq- med: Building multimodal clinical foundation mod- els with domain-aware grpo training. arXiv preprint arXiv:2506.00711, 2025. Huang, Y., Peng, Z., Zhao, Y., Yang, P., Yang, X., and Shen, W. Medseg-r: Reasoning segmentation in medical images with multimodal large language models. arXiv preprint arXiv:2506.10465, 2025. Jiang, S., Wang, Y., Song, S., Hu, T., Zhou, C., Pu, B., Zhang, Y., Yang, Z., Feng, Y., Zhou, J. T., et al. Hulu- med: A transparent generalist model towards holistic medical vision-language understanding. arXiv preprint arXiv:2510.08668, 2025. Jin, J. and Jeong, C. W. Surgical-llava: Toward surgical scenario understanding via large language and vision models. arXiv preprint arXiv:2410.09750, 2024. Khalid, M. U., Laplante, S., Masino, C., Alseidi, A., Jayara- man, S., Zhang, H., Mashouri, P., Protserov, S., Hunter, J., Brudno, M., et al. Use of artificial intelligence for decision-support to avoid high-risk behaviors during la- paroscopic cholecystectomy. Surgical Endoscopy, 37(12): 9467–9475, 2023. Lai, X., Tian, Z., Chen, Y., Li, Y., Yuan, Y., Liu, S., and Jia, J. Lisa: Reasoning segmentation via large language model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 9579– 9589, 2024. Laplante, S., Namazi, B., Kiani, P., Hashimoto, D. A., Al- seidi, A., Pasten, M., Brunt, L. M., Gill, S., Davis, B., Bloom, M., et al. Validation of an artificial intelligence platform for the guidance of safe laparoscopic cholecys- tectomy. Surgical endoscopy, 37(3):2260–2268, 2023. Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon, H., and Gao, J. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36:28541–28564, 2023. Liu, B., Zou, K., Zhan, L.-M., Lu, Z., Dong, X., Chen, Y., Xie, C., Cao, J., Wu, X.-M., and Fu, H. Gemex: A large-scale, groundable, and explainable medical vqa benchmark for chest x-ray diagnosis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 21310–21320, 2025. Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tun- ing. Advances in neural information processing systems, 36:34892–34916, 2023. Lu, S., Li, Y., Xia, Y., Hu, Y., Zhao, S., Ma, Y., Wei, Z., Li, Y., Duan, L., Zhao, J., et al. Ovis2. 5 technical report. arXiv preprint arXiv:2508.11737, 2025. Madani, A., Namazi, B., Altieri, M. S., Hashimoto, D. A., Rivera, A. M., Pucher, P. H., Navarrete-Welton, A., Sankaranarayanan, G., Brunt, L. M., Okrainec, A., et al. Artificial intelligence for intraoperative guidance: using 9 ResGo & SurGo-R1 semantic segmentation to identify surgical anatomy dur- ing laparoscopic cholecystectomy. Annals of surgery, 276 (2):363–369, 2022. Magrina, J. F. Complications of laparoscopic surgery. Clin- ical obstetrics and gynecology, 45(2):469–480, 2002. Mascagni, P., Vardazaryan, A., Alapatt, D., Urade, T., Emre, T., Fiorillo, C., Pessaux, P., Mutter, D., Marescaux, J., Costamagna, G., et al. Artificial intelligence for surgical safety: automatic assessment of the critical view of safety in laparoscopic cholecystectomy using deep learning. An- nals of surgery, 275(5):955–961, 2022. Mascagni, P., Alapatt, D., Murali, A., Vardazaryan, A., Garcia, A., Okamoto, N., Costamagna, G., Mutter, D., Marescaux, J., Dallemagne, B., et al. Endoscapes, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecystectomy. Scientific Data, 12(1):331, 2025. Murali, A., Alapatt, D., Mascagni, P., Vardazaryan, A., Gar- cia, A., Okamoto, N., Mutter, D., and Padoy, N. Latent graph representations for critical view of safety assess- ment. IEEE Transactions on Medical Imaging, 43(3): 1247–1258, 2023. Nwoye, C. I., Alapatt, D., Yu, T., Vardazaryan, A., Xia, F., Zhao, Z., Xia, T., Jia, F., Yang, Y., Wang, H., et al. Cholectriplet2021: A benchmark challenge for surgical action triplet recognition. Medical Image Analysis, 86: 102803, 2023. Peery, A. F., Crockett, S. D., Murphy, C. C., Jensen, E. T., Kim, H. P., Egberg, M. D., Lund, J. L., Moon, A. M., Pate, V., Barnes, E. L., et al. Burden and cost of gastrointestinal, liver, and pancreatic diseases in the united states: update 2021. Gastroenterology, 162(2):621–644, 2022. Peng, Z., Wang, W., Dong, L., Hao, Y., Huang, S., Ma, S., and Wei, F.Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. Plummer, B. A., Wang, L., Cervantes, C. M., Caicedo, J. C., Hockenmaier, J., and Lazebnik, S. Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision, p. 2641– 2649, 2015. Pucher, P. H., Brunt, L. M., Davies, N., Linsk, A., Munshi, A., Rodriguez, H. A., Fingerhut, A., Fanelli, R. D., As- bun, H., Aggarwal, R., et al. Outcome trends and safety measures after 30 years of laparoscopic cholecystectomy: a systematic review and pooled data analysis. Surgical endoscopy, 32(5):2175–2183, 2018. Rasheed, H., Maaz, M., Shaji, S., Shaker, A., Khan, S., Cholakkal, H., Anwer, R. M., Xing, E., Yang, M.-H., and Khan, F. S. Glamm: Pixel grounding large multimodal model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 13009– 13018, 2024. Ren, Z., Huang, Z., Wei, Y., Zhao, Y., Fu, D., Feng, J., and Jin, X. Pixellm: Pixel reasoning with large multimodal model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 26374– 26383, 2024. Strasberg, S. M. and Brunt, M. L. Rationale and use of the critical view of safety in laparoscopic cholecystectomy. Journal of the American College of Surgeons, 211(1): 132–138, 2010. Tong, Q., Lu, Z., Liu, J., Zheng, Y., and Lu, Z.-M. Medisee: Reasoning-based pixel-level perception in medical im- ages. In Proceedings of the 33rd ACM International Conference on Multimedia, p. 2742–2751, 2025. Twinanda, A. P., Shehata, S., Mutter, D., Marescaux, J., De Mathelin, M., and Padoy, N. Endonet: a deep archi- tecture for recognition tasks on laparoscopic videos. IEEE transactions on medical imaging, 36(1):86–97, 2016. Velanovich, V. Laparoscopic vs open surgery: a prelimi- nary comparison of quality-of-life outcomes. Surgical endoscopy, 14(1):16–21, 2000. Wang, G., Bai, L., Wang, J., Yuan, K., Li, Z., Jiang, T., He, X., Wu, J., Chen, Z., Lei, Z., et al. Endochat: Grounded multimodal large language model for endoscopic surgery. arXiv preprint arXiv:2501.11347, 2025a. Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., et al. Internvl3. 5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265, 2025b. Way, L. W., Stewart, L., Gantert, W., Liu, K., Lee, C. M., Whang, K., and Hunter, J. G. Causes and prevention of laparoscopic bile duct injuries: analysis of 252 cases from a human factors and cognitive psychology perspective. Annals of surgery, 237(4):460–469, 2003. Yang, S., Qu, T., Lai, X., Tian, Z., Peng, B., Liu, S., and Jia, J. Lisa++: An improved baseline for reasoning seg- mentation with large language model. arXiv preprint arXiv:2312.17240, 2023. You, H., Zhang, H., Gan, Z., Du, X., Zhang, B., Wang, Z., Cao, L., Chang, S.-F., and Yang, Y. Ferret: Refer and ground anything anywhere at any granularity. arXiv preprint arXiv:2310.07704, 2023. 10 ResGo & SurGo-R1 Zegers, M., de Bruijne, M. C., de Keizer, B., Merten, H., Groenewegen, P. P., van der Wal, G., and Wagner, C. The incidence, root-causes, and outcomes of adverse events in surgical units: implication for potential prevention strategies. Patient safety in surgery, 5(1):13, 2011. Zeng, Z., Zhuo, Z., Jia, X., Zhang, E., Wu, J., Zhang, J., Wang, Y., Low, C. H., Jiang, J., Zheng, Z., et al. Surgvlm: A large vision-language model and systematic evalua- tion benchmark for surgical intelligence. arXiv preprint arXiv:2506.02555, 2025. Zheng, B., Cassera, M. A., Martinec, D. V., Spaun, G. O., and Swanstr ̈ om, L. L. Measuring mental workload during the performance of advanced laparoscopic tasks. Surgical endoscopy, 24(1):45–50, 2010. 11 ResGo & SurGo-R1 Appendix Phase-Definition Mapping Tool The content utilized by the Phase-Definition Mapping Tool is derived from a curated lexicon of phase-specific definitions, as detailed in Table 6. Comprising distinct anatomical targets and safety-driven exclusion criteria, these descriptions are dynamically inserted into the<Surgical Context>in the prompt to explicitly condition the model’s spatial grounding. Table 6. Anatomical Go Zone Definitions for Phase-Guided Reasoning. The content utilized by the Phase-Definition Mapping Tool is presented in this table, supplying specific operative cues due to the deficiency in surgical domain knowledge common to current VLMs. Procedural PhaseGo Zone Definition & Reasoning Instruction Preparation of the Go ZoneTarget the Calot’s triangle area. The Go Zone is defined as the peritoneum overlying the cystic duct and artery junction, where the initial dissection must commence to expose the underlying structures. Dissection of Calot’s TriangleTarget the “Safety Window.” The Go Zone is strictly limited to the fibro-fatty tissue within the triangle. It explicitly excludes the liver bed (cystic plate) and the common bile duct (CBD) to prevent bile duct injury. Clip and DivideTarget the skeletonized cystic duct and artery. The Go Zone is the specific segment of these structures that is free of surrounding tissue and distinct from the CBD, rendering it suitable for safe clip application. Gallbladder DissectionTarget the connective tissue plane (areolar tissue). The Go Zone is identified as the white or translucent line of adhesion separating the gallbladder from the liver bed. Reasoning Turn Prompt We use the following template for the phase-guided reasoning turn that conditions the model on the predicted procedural phase and its corresponding definition retrieved from Phase-Definition Mapping Tool. Curly-braced fields (...) are populated at inference time:phasenameis the predicted phase name from Turn 1 output,choiceletteris the multiple-choice identifier model selected, anddefinitionis the phase-specific Go Zone definition and constraints provided by the Tool. Reasoning Turn Prompt ### Surgical Context ** Identified Phase ** : phase_name (choice_letter) ** Definition ** : definition ### Mission Locate the "Go Zone" strictly based on the ** Definition ** above. ### Required Output Format <thinking> Briefly map the visual landmarks (e.g., liver edge, duct, instrument) to the text definition above. </thinking> <reasoning> 1. Location: (Anatomical location description of the Go Zone) 2. Exposure: (Is retraction sufficient?) 3. Next Action: (Immediate maneuver) 4. Critical Risk: (Primary risk of error) </reasoning> <answer> [xmin, ymin, xmax, ymax] 12 ResGo & SurGo-R1 </answer> 13