Paper deep dive
Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
Mohammad Javad Ahmadi, Hamid D. Taghirad
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/19/2026, 5:03:08 AM
Summary
This paper introduces an explainable AI-powered framework for automated, objective skill assessment in cataract surgery, specifically focusing on the capsulorhexis phase. The authors present the world's largest dataset of cataract surgery videos (2,000 recordings) and a subset of 83 annotated videos (CVSAD-83) rated using the new Capsulorhexis Skill Assessment System (CSAS). The framework uses computer vision to extract motion-based metrics (e.g., path length, jerk, energy) and demonstrates strong correlation with expert subjective ratings, achieving up to 87% accuracy. The approach aims to replace subjective, biased evaluations with objective, explainable quantitative indicators.
Entities (16)
Relation Signals (15)
Hamid D. Taghirad → isauthorof → Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
confidence 99% · and Hamid D. Taghirad
Mohammad Javad Ahmadi → isauthorof → Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
confidence 99% · Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery Mohammad Javad Ahmadi
Mohammad Javad Ahmadi → isaffiliatedwith → K.N. Toosi University of Technology
confidence 98% · Mohammad Javad Ahmadi 1 and Hamid D. Taghirad 1* 1* Applied Robotics and AI Solutions (ARAS), Faculties of Electrical and Computer Engineering, K.N. Toosi University of Technology
Hamid D. Taghirad → isaffiliatedwith → K.N. Toosi University of Technology
confidence 98% · Mohammad Javad Ahmadi 1 and Hamid D. Taghirad 1* 1* Applied Robotics and AI Solutions (ARAS), Faculties of Electrical and Computer Engineering, K.N. Toosi University of Technology
CVSAD-83 → isannotatedwith → Capsulorhexis Skill Assessment System
confidence 98% · 83 preprocessed surgical videos have been independently reviewed and rated by three trained medical professionals using CSAS subjective indicators.
Capsulorhexis Skill Assessment System → isusedfor → Cataract Surgery
confidence 97% · specifically focusing on cataract surgery... introduced the Capsulorhexis Skill Assessment System (CSAS)
Explainable AI-Powered Framework → achievesaccuracy →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.
Tags
Links
- Source: https://arxiv.org/abs/2608.17522v1
- Canonical: https://arxiv.org/abs/2608.17522v1
Trouble viewing inline? Open PDF directly →
Full Text
65,419 characters extracted from source content.
Expand or collapse full text
Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery Mohammad Javad Ahmadi 1 and Hamid D. Taghirad 1* 1* Applied Robotics and AI Solutions (ARAS), Faculties of Electrical and Computer Engineering, K.N. Toosi University of Technology, Tehran, Iran. *Corresponding author(s). E-mail(s): taghirad@kntu.ac.ir; Contributing authors: mjahmadi@email.kntu.ac.ir; Abstract Persistent shortages in the surgical workforce and inherent limitations of tra- ditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by intro- ducing a novel, explainable AI-powered framework for automated skill assess- ment, specifically focusing on cataract surgery. We present the world’s largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced com- puter vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of out- puts, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjec- tive ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion- based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework’s capacity to accurately model surgical expertise. Keywords: Surgical Education, Skill Assessment, Cataract Surgery, Computer Vision, Dataset, Video Analysis, Motion Analysis 1 arXiv:2608.17522v1 [cs.CV] 18 Aug 2026 1 Introduction Global demand for surgical services is significantly increasing; however, access to these procedures remains uneven. Notably, low- and middle-income countries face a shortfall of 143 million essential surgeries each year [1]. Expectations indicate that every country will require at least 5,000 surgical procedures per 100,000 people annually by 2030 [2]. Addressing this gap necessitates a significant expansion in healthcare infrastructure and personnel, including doubling the surgical workforce over the next fifteen years [2]. Nevertheless, current educational and assessment methods in surgical training, such as the Master-Apprentice Model (MAM), are inadequate to achieve these objectives. These conventional methods face several challenges, notably the reliance on subjective performance evaluations rather than objective, measurable standards, which often leads to biased and inconsistent assessments [3]. Furthermore, the high workload of trainers limits the time available for providing detailed and comprehensive feedback [4]. The growing gap between surgical needs and available resources, combined with the limited integration of advanced technology in training, has resulted in a lack of skilled surgeons, longer patient wait times, increased surgical errors, and patient deaths [5]. Research indicates that early-career surgeons are nearly nine times more likely to commit procedural errors compared to experienced surgeons [6], underscoring the need for improved training and assessment methods [7, 8]. By implementing enhanced training programs, healthcare systems could prevent approximately 42,000 hospital readmissions annually, save up to $3.62 million in costs [9], and mitigate 67% of cases leading to adverse events due to inadequate instrument handling skills [10]. In this context, advancements in artificial intelligence (AI) present promising opportunities to enhance surgical procedures and transform training paradigms. One approach is the recording and analysis of surgical videos. Unlike sensor and kinematic data, which require expensive equipment, video recordings are readily available in most operating rooms through devices like surgical microscopes. Beyond their accessibility, surgical videos provide a comprehensive view of the entire process, capturing the sur- geon’s movements, the patient’s condition, and the progress of the surgical workflow [11]. Furthermore, integrating surgical recordings into training programs has become an invaluable resource for enhancing practical skills and academic research in the surgical field [11–13]. Novice surgeons report that reviewing videos and receiving feedback helps them prepare for surgeries, improve their anatomical knowledge, and refine their skills by analyzing past performances [13]. Despite the recognized benefits of surgical video materials, their widespread use remains limited due to challenges such as the time-consuming and expensive review of extensive video archives [14–16]. These findings highlight the need for advanced systems that streamline surgical video management, review, and analysis. Implement- ing intelligent platforms to create annotated video libraries could enhance training programs and improve clinical outcomes. While current research efforts in automated surgical video analysis frequently focus on identifying operative phases and workflows, data-driven methodologies for objective video-based surgical skill assessment remain limited [17]. Current research 2 predominantly utilizes deep neural network architectures to analyze surgical videos by processing them either as individual frames or as short video segments. These neural networks are designed to extract spatiotemporal features from video that correspond to distinct skill levels. Funke et al. [18] proposed an end-to-end deep-learning pipeline that classifies sur- gical skills from short video clips using an inflated 3-D CNN and a Temporal Segment Network. Their evaluation was conducted on JIGSAWS [19], a bench-top simula- tion dataset, not footage from the real operations. Although the reported accuracies exceeded 95%, this performance was achieved only for a three-class classification. Moreover, the reliance on simulated data and the absence of interpretable feedback constrain the approach’s immediate clinical relevance. Kitaguchi et al. [20] trained a 3-D CNN on 40-second clips from 74 videos of laparoscopic colorectal procedures labeled with Endoscopic Surgical Skill Qualifica- tion System (ESSQS) scores. Their model achieved a mean accuracy of 75% when classifying surgical performance into low, medium, or high skill levels. However, the moderate accuracy, use of a small and selectively sampled dataset, and reliance on a black-box three-tier output limit the method’s robustness and its utility for providing meaningful, detailed educational feedback. Zia et al. [21] proposed RP-Net-V2, a CNN–LSTM architecture that segments robotic-assisted radical prostatectomies into 12 steps and evaluates performance through task-based efficiency metrics, achieving a Jaccard index of 0.85 and strong correlations with expert scores. However, the model’s reliance on raw RGB frames leads to frequent confusion between visually similar tasks, the class imbalance is left unaddressed, a simple running-window median filter shifts some predicted boundaries by over a minute, and the ground-truth labels originate from a single annotator. Lajk ́o et al. [22] proposed a 2D video-based skill assessment method using sparse optical flow and standard classifiers on the JIGSAWS dataset, achieving around 80% accuracy. However, performance remained below kinematics-based methods, and the approach relied on manually selected tool regions, limiting automation. Although deep learning models have achieved promising results in medical applica- tions, their widespread use in sensitive clinical settings remains constrained by several issues, including the black box nature of most neural networks, the demand for exten- sive expert-annotated datasets, and limited capacity for delivering detailed, actionable feedback to surgeons. To address these challenges, this paper introduces an AI-powered framework that integrates explainable methods for objective and automated skill evaluation in surgical videos. Although the framework is applicable to a wide range of surgical procedures, its effectiveness should be independently evaluated in each context. As a proof of concept, we focus on cataract surgery, specifically the capsulorhexis phase, as our primary case study. This choice is motivated by the fact that cataract surgery is among the most frequently performed surgeries worldwide [23]. Cataract remains the primary cause of preventable vision loss and significant visual impairment, particularly in developing nations [24, 25]. Phacoemulsification is the globally preferred surgery for cataract removal, making it a vital component of the residency educational curriculum for developing surgical competency [26]. Surgical 3 competency requires deep knowledge, dexterity, micro-surgical skills, and dedication [27]. Creating a circular opening in the anterior surface of the cataracted lens, called capsulorhexis, is a fateful step in phacoemulsification cataract surgery. It’s a significant challenge for first-year residents [28, 29]. Trainee surgeons require years of training and evaluation, with evolving methods and tools for learning and assessment. Rubrics, with their structured approach, are the conventional tools in skill assessments. Several rubrics are available for evaluating cataract surgery, such as the Interna- tional Council of Ophthalmology’s Ophthalmology Surgical Competency Assessment Rubric (ICO-OSCAR) [30], Objective Assessment of Skills in Intraocular Surgery (OASIS) [31], Global Rating Assessment of Skills in Intraocular Surgery (GRASIS) [32], and Objective Structured Assessment of Cataract Surgical Skill (OSACSS) [33]. However, these rubrics typically require direct trainer supervision, leading to high costs, time-consuming processes, and potential biases. Our proposed framework overcomes these limitations by automating the assess- ment process, eliminating the need for direct trainer supervision. This reduces both costs and time, while minimizing the potential for biases, providing a more efficient and objective method for evaluating surgical skills. In the following sections, we present our newly developed cataract surgery video dataset, the largest in the world in terms of the number of data, and comprehensive annotations. We then describe our method for intelligent, automated, and objective skill evaluation, detailing the framework’s architecture and operational steps. In the Results section, we report the performance of the proposed framework on 83 cataract surgery videos, accompanied by extensive analyses that underscore both the signifi- cance of this dataset and the method’s effectiveness. We also discuss each objective metric in depth to illuminate various dimensions of skill assessment. Finally, in the Conclusion, we summarize the key findings and outline potential directions for future research. 2 Methods In this section, we present a large-scale surgical video dataset designed to support interdisciplinary AI research and serve as ground truth for validating our automated skill assessment systems. We then introduce a novel AI-powered framework that seam- lessly processes these videos, generating interpretable metrics of surgical performance and advancing the reliability of automated skill evaluation in clinical practice. 2.1 Dataset Crafting Description Cataract surgery is performed with precise instruments under a surgical microscope, while an attached camera captures the operation from the surgeon’s perspective. The ARAS Group and Farabi Eye Hospital jointly developed an infrastructure for the systematic acquisition of these surgical videos. From 2021 to 2024, more than 2,500 videos were recorded at up to 60 frames per second with a resolution of 1920×1080 pixels. The video selection process involved 4 rigorous exclusion criteria that filtered out videos with technical flaws such as substandard image quality, and low resolution. Having established a high-quality dataset, the next step is to select an appropriate method for skill-based annotation. In the domain of surgical video analysis, two prin- cipal methods have been utilized to annotate surgeon skill. The first method involves identifying individual surgeons and estimating their proficiency based on metrics such as academic rank and accumulated surgical experience, as utilized in the Cataract-101 dataset [34]. In contrast, the second approach relies on expert review of surgical videos, where evaluators assign skill scores according to predefined standards. This expert-based approach generally provides a more accurate and unbiased evaluation of surgical performance than identity-based methods. This study introduces a video-based skill annotation method that segments sur- gical procedures into distinct phases, each evaluated with specialized criteria. In our case study, video segments corresponding to the capsulorhexis phase were extracted from the recorded procedures, and detailed annotations were assigned to these seg- ments using a new system developed to standardize the review process of capsulorhexis videos. The proposed assessment framework was validated through a comprehensive review of existing surgical skill assessment systems [35–39], carried out in collaboration with two experienced surgeons. Based on this review, we introduced the Capsulorhexis Skill Assessment System (CSAS). Table 1 details the 11 indicators of CSAS, incorporating seven subjective and four objective metrics: rhexis position, shape, size, and time. Notably, the indicators “Commencement of Flap” and “Circular Completion” follow the ICO-OSCAR [39] guidelines, whereas “Tissue Handling”, “Instrument Handling”, and “Microscope Use” have been adapted from the GRASIS [40] scoring system. In this study, 83 preprocessed surgical videos have been independently reviewed and rated by three trained medical professionals using CSAS subjective indicators. This review process resulted in crafting the cataract surgical video skill assessment dataset (CVSAD-83) [41]. Each indicator was rated on a five-point scale, where higher scores indicate greater proficiency. The reliability of these annotations was confirmed by an intraclass correlation coefficient (ICC) of 0.79 and a Cronbach’s alpha coefficient of 0.93, indicating high inter-rater reliability and internal consistency. To facilitate the categorization of surgeon skills according to the Dreyfus Model [42], an overall score (OS) was computed for each video by averaging the scores of all indicators. This OS, ranging from 1 to 5, is a reliable metric for skill classification. 2.2 AI-Powered Objective Skill Assessment Framework This paper introduces an innovative framework designed to enhance the explain- ability of AI-powered systems for automatic surgical skill assessment through video analysis. The framework begins with a computer vision algorithm that tracks and segments surgical tools and relevant tissues, extracting pixel-level motion informa- tion related to these elements. Subsequently, this raw pixel data is processed through a post-processing stage, converting it into actionable information for surgical skill assessment. Finally, an objective evaluation system is developed, which evaluates the 5 Indicator Novice Advanced Beginner Competent Description Score Description Score Description Score Instrument Handling Repeated, Abrupt, Awk- ward/Bizarre, and Harsh Movements; Endless Insertion/Entryand Exit 1–2 Selected/Occasional Inappropriate Movements 3–4 Fine and Smooth Movements; No Inappropriate Movements 5 Motion Unsure Surgical Plan; Needless in- doubt Movements 1–2 Certain Surgical Plan; Occasional/s- elected Unnecessary Movements 3–4 Maximum Effective Movements; No Unnecessary Movements 5 Commencement of flap Tentative Chases rather than Con-trolled; Numerous disruptions in thecortex 1–2 Flap Pull-up After 2–3 Tries; Subtle Cortex Disruptions 3–4 Delicate and Controlled Approach; No Cortex disruption 5 Tissue Handling Using Unnecessary Force; Damage to Cornea and Conjunctiva 1–2 Suitable Tissue Interactions; Unin- tentional/Accidental Tissue Damage 3–4 Excellent Tissue Interactions; No Tissue Damage 5 Microscope Use Multiple Recentering; Multiple Refo- cus 1–2 Few attempts to recenter/refocus 3–4 Keeps Eye Centered; Good Focused View 4–5 Circular Completion Not Able to Achieve Circular Rhexis; Entrance to Dangerous Regions(Extend to the Periphery) 1–2 Difficulty Achieving Rhexis 3–4 Rapid, Unaided Completion of Rhexis 5 Adverse Events Rhexis Tear; Entering Hazardous Areas; Unable to Recognize Adverseevents. Inappropriate Reaction ✓ – – – – Position of Rhexis > 1 m 1–2 0.4–1 m 3–4 < 0.4 m 5 Shape of Rhexis Irregular Shape and Margins 1–2 Semi-Circular Shape; Minor Irregu- larities along Edges 3–4 Perfectly Circular 4–5 Size of Rhexis < 4.25 m or > 6.25 m 1–2 4.25–4.75 m or 5.75–6.25 m 3–4 4.75–5.75 m (5–5.5 ± 0.25 m) 5 Time > 75 seconds 1–3 – – ≤ 75 seconds 4–5 Table 1 : Capsulorhexis Skill Assessment System (CSAS) 6 surgeon’s skills based on motion data and formulated criteria extracted from the logic of subjective indices. 2.2.1 Motion Data Extraction and Processing In this section, we present a detailed overview of the methods used to extract and process video-based motion data. To accomplish this, our proposed framework inte- grates state-of-the-art computer vision algorithms for segmentation and tracking, such as SAM-2 [43] or YOLO-11 [44]. In this study, pre-trained models developed in our previous studies [45–47] were utilized to segment the capsulorhexis instrument and the corneal tissue in the videos of CVSAD-83 dataset. The motion data obtained through tracking these components provides the essential data to evaluate surgical skills. Following the segmentation-tracking stage, the frame-by-frame coordinates of the capsulorhexis tool-tip and the corneal center are identified, forming the fundamen- tal motion data for analysis. However, to ensure the resulting measurements are meaningful for skill evaluation, a proper referencing system must be established. With- out appropriate referencing, raw positional measurements may lead to inaccurate conclusions about skill performance. Absolute referencing, in which each component’s position is determined relative to a fixed point in the frame (e.g., the top-left corner), can be misleading. Rapid shifts in a tool’s location with respect to this fixed point may arise from patient eye movements or changes in the microscope’s field of view, rather than a lack of surgical precision. To address this limitation, a relative referencing scheme provides a more reliable measure for skill-related analysis by representing the tool-tip coordinates in relation to the target tissue (in this case, the corneal center). Transitioning to a relative reference system effectively translates sudden absolute movements into smoother, more inter- pretable variations in the tool-tip’s position, allowing for a more accurate assessment of surgical performance. After extracting these relative positions, the data must be scaled to account for varying magnifications in the surgical microscope across different videos. One practical method is to measure the corneal diameter in pixels using computer vision techniques and then convert these measurements to physical units based on the known corneal diameter (11.71 ± 0.42 m [48]). Although this scaled data may not precisely replicate real-world dimensions, it serves its primary purpose of capturing the essential motion trends indicative of surgical proficiency. Finally, higher-order motion features such as velocity, acceleration, and jerk are derived from the scaled relative positions for objective skill evaluation in subsequent analytical steps. By focusing on the dynamics of tool movement rather than on abso- lute positional data, this approach provides a more reliable framework for subsequent analytic processes to distinguish between different levels of surgical expertise. 2.2.2 Objective Metrics for Automated Surgical Skill Assessment Following the discussion on motion data extraction and processing methods, this section presents the objective metrics for automated surgical skill assessment. These 7 metrics were derived through expert consultation and by adapting the principles from subjective assessment systems. 1. Procedure Time (PT): The total time for completing the surgery, T , is computed by subtracting the start time from the end time, as in Equation 1: PT = t end − t start .(1) 2. Path Lenrth (PL): This metric is computed by summing the magnitudes of the tool’s relative positional coordinates. Specifically, if (x i ,y i ) represents the position at the i-th time instant, then the path length is given by Equation 2: PL = n X i=1 q x 2 i + y 2 i ,(2) where n is the total number of sampled points in the recorded trajectory. 3. Number of Alternating (Back-and-Forth) Movements (NAM): This metric reflects the count of tool direction changes. A change in the sign of the velocity along any axis indicates a reversal in tool motion. The total number of such movements is calculated as in Equation 3: NAM = n−1 X i=1 δ(sign(v i+1 )− sign(v i )),(3) where sign(·) extracts the velocity direction along either x or y axis, and δ(·) is 1 if a sign change is detected, otherwise 0. 4. Motion Difficulty (MD): Abrupt changes in acceleration indicate higher jerk, which are associated with reduced smoothness [49]. The overall difficulty is quantified by summing the magnitude of the jerk vector, as given by Equation 4: MD = n X i=1 q j 2 x i + j 2 y i ,(4) where j x i and j y i are the jerk components along the x and y axes, respectively. 5. Energy Consumption (EC): By assuming the needle mass is negligible and con- stant, the kinetic energy concept 1 2 mv 2 can be used to estimate the total energy. Let (v x i ,v y i ) be the velocity components at time i. The EC is estimated using Equation 5: EC = n X i=1 1 2 (v 2 x i + v 2 y i ).(5) 6. Number of Local Peaks (NLP): The local peaks in speed, acceleration, or jerk profiles are counted using computational methods such as the find peaks Python function in scipy.signal. Let NLP denote the total number of detected peaks across these profiles, formalized as in Equation 6: NLP = Peaks(speed) + Peaks(acceleration) + Peaks(jerk),(6) where Peaks(·) returns the count of local peaks in the corresponding signal. 8 7. Total Force (TF): Inspired by the physical relation Force = m× a and assuming negligible needle mass, the total force is estimated by summing the magnitude of the acceleration vector. Equation 7 defines Total Force: TF = n X i=1 q a 2 x i + a 2 y i .(7) 8. Tissue Interaction (TI): Sharp deviations from the mean acceleration can indicate sudden forces that might harm tissue. By comparing local acceleration peaks with the mean, an interaction index (TI) is computed as indicated in Equation 8: TI = n X i=1 q a 2 x i + a 2 y i − μ a ,(8) where μ a is the average magnitude of acceleration. 9. Speed Frequency Smoothness (SFS): The mean frequency magnitude in the fast Fourier transform (FFT) of the speed signal is used to indicate frequency smoothness. Equation 9 details the computation: SFS = 1 n n X i=1 |FFT(v i )|,(9) where FFT(·) denotes the FFT of the velocity at each sampled interval. 10. Spectral Entropy (SE): The distribution of the signal’s energy across frequencies is captured by the spectral entropy, described in Equation 10: SE =− X f P (f ) logP (f ),(10) where P (f ) = |FFT(x f )| 2 P j |FFT(x j )| 2 is the normalized power spectral density of the signal. In the subsequent section, it will be demonstrated how these metrics, which were developed from subjective indices and motion data extracted from surgical videos, correlate with expert surgeons’ subjective assessments in the CVSAD-83 dataset. An overview of our proposed framework is presented in Figure 1. 3 Results This section presents and discusses the findings of three primary analyses conducted on the CVSAD-83 dataset. First, we examine the distribution of subjective skill scores, offering insights into typical performance levels and areas for targeted improvement. Next, we assess the correlation between ten automated computed objective metrics and the overall subjective score (OS) assigned by expert surgeons, thereby evaluating the reliability of automated indicators as proxies for clinical expertise. Finally, we explore each metric’s discriminative power by clustering surgeons into different skill groups, comparing the results with ground-truth expert labels. 9 Video Preprocessing Video Dataset Phase Recognition Video Snippets Objective Skill Assessment Framework Results Trainer Surgeon Validation Study Advanced Vision Model Processed Data Fig. 1: Overview of the proposed framework. The box plots in Figure 2 illustrate the distribution of scores for the subjective indicators of the CSAS evaluation system, which were used to evaluate the 83 videos from the CVSAD dataset, in addition to the overall score (OS). In each plot, the box spans the interquartile range (IQR), with the median and mean values denoted by the horizontal lines and whiskers marking the full data range. Notably, most indicators exhibit median scores near or above 3.75, indicating that the majority of assessed performances cluster within the upper half of the 1–5 scale. This insight into the distribution of performances helps trainers develop targeted interventions that can systematically improve surgical outcomes by addressing specific weak points discovered through the CSAS framework. For instance, the scores for Microscope Use and Tissue Handling appear slightly higher than those for Motion and Circular Completion, indicating that surgeons generally excel in maintaining focus and handling ocular tissues while having more varied performance on tasks requiring precise movement control. Following the subjective score distribution analysis, we explored the relationship between ten objective metrics generated by our automated assessment framework and the expert-assigned Overall Scores (OS). Establishing strong correlations between objective indicators and subjective assessments is crucial for validating the reliability of our framework. Figure 3 presents scatter plots for each metric against OS, with the horizontal axis representing the expert-assigned score (OS) value and the vertical axis showing the corresponding objective metric output. Data points are labeled by their video IDs, and a regression line highlights the general trend in each plot. A consistent inverse rela- tionship emerges across all ten objective indicators, reflecting that surgeons assigned higher OS values typically exhibit lower measured values for objective metrics. A detailed and thorough analysis of the charts corresponding to each of the 10 objective indicators is presented as follows: 10 Instrument Handling 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 1.88 Max: 5.00 Median: 3.88 Mean: 3.87 Motion Plot 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 1.88 Max: 5.00 Median: 3.76 Mean: 3.51 Commencement of Flap 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 2.00 Max: 5.00 Median: 3.88 Mean: 3.68 Tissue Handling 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 2.12 Max: 5.00 Median: 3.88 Mean: 3.94 Microscope Use 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 2.12 Max: 5.00 Median: 3.88 Mean: 3.97 Circular Completion 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 1.36 Max: 5.00 Median: 3.76 Mean: 3.60 Overall Score (OS) 2.5 3.0 3.5 4.0 4.5 5.0 Score Min: 2.39 Max: 4.88 Median: 3.79 Mean: 3.72 IQR (25th-75th) Whiskers (Min-Max) MedianMean MedianMean IQR (25th-75th percentile) Fig. 2: Overview of CSAS scores and overall performance from 83 CVSAD cases. 1. Procedure Time (PT): As depicted in Figure 3(a), a significant inverse correla- tion was observed between Procedure Time (PT) and subjective skill scores, with a regression slope of –60.057. As skill scores increased, PT decreased, indicating that more proficient surgeons complete procedures more efficiently, likely due to enhanced technical dexterity and optimized decision-making. These findings high- light the value of PT as a metric for surgical competence, emphasizing its potential to improve patient outcomes and increase operating room efficiency. 2. Path Length (PL): Analysis of the PL metric underscores a negative regression slope of –272.946 in its correlation with subjective surgical skill scores (Figure 3(b)). The shorter path lengths observed at higher skill levels suggest that expe- rienced surgeons execute movements more efficiently, limiting unnecessary travel. This streamlined motion is particularly valuable during delicate procedures like capsulorhexis, where minimizing extra instrument movement enhances precision and patient outcomes. 3. Number of Alternating Movements (NAM): Figure 3(c) shows that the NAM exhibits a strong negative slope (–1840.795) when regressed against subjective surgical skill scores. Surgeons with higher proficiency perform fewer directional reversals, implying more controlled, deliberate motions and purposeful instrument handling. Reducing rapid back-and-forth movements enhances procedural efficiency 11 2.53.03.54.04.55.0 Subjective Score 0 50 100 150 200 250 300 Procedure Time (PT) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Procedure Time (PT) Surgeon ID Regression Line (Slope: -60.057) (a) Procedure Time (PT) 2.53.03.54.04.55.0 Subjective Score 0 200 400 600 800 1000 1200 Path Lenrth (PL) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Path Lenrth (PL) Surgeon ID Regression Line (Slope: -272.946) (b) Path Length (PL) 2.53.03.54.04.55.0 Subjective Score 0 2000 4000 6000 8000 Number of Alternating Movements (NAM) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Number of Alternating Movements (NAM) Surgeon ID Regression Line (Slope: -1840.795) (c) Number of Alternating Movements (NAM) 2.53.03.54.04.55.0 Subjective Score 0 50 100 150 200 Motion Difficulty (MD) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Motion Difficulty (MD) Surgeon ID Regression Line (Slope: -44.944) (d) Motion Difficulty (MD) 2.53.03.54.04.55.0 Subjective Score 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 Energy Consumption (EC) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Energy Consumption (EC) Surgeon ID Regression Line (Slope: -0.227) (e) Energy Consumption (EC) 2.53.03.54.04.55.0 Subjective Score 0 1000 2000 3000 4000 5000 6000 Number of Local Peaks (NLP) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Number of Local Peaks (NLP) Surgeon ID Regression Line (Slope: -1298.269) (f) Number of Local Peaks (NLP) 2.53.03.54.04.55.0 Subjective Score 0 20 40 60 80 100 120 Total Force (TF) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Total Force (TF) Surgeon ID Regression Line (Slope: -25.394) (g) Total Force (TF) 2.53.03.54.04.55.0 Subjective Score 0.1 0.2 0.3 0.4 0.5 0.6 Tissue Interaction (TI) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Tissue Interaction (TI) Surgeon ID Regression Line (Slope: -0.059) (h) Tissue Interaction (TI) Figure 3 continues on the next page. 12 2.53.03.54.04.55.0 Subjective Score 0.2 0.4 0.6 0.8 1.0 Speed Frequency Smoothness (SFS) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Speed Frequency Smoothness (SFS) Surgeon ID Regression Line (Slope: -0.171) (a) Speed Frequency Smoothness (SFS) 2.53.03.54.04.55.0 Subjective Score 6.0 6.5 7.0 7.5 8.0 8.5 Spectral Entropy (SE) Value ID018 ID008 ID014 ID010 ID002 ID016 ID020 ID003 ID009 ID019 ID001 ID021 ID022 ID004 ID005 ID015 ID007 ID006 ID011 ID023 ID017 ID012 ID013 ID035 ID037 ID042 ID034 ID044 ID024 ID025 ID041 ID043 ID046 ID039 ID031 ID026 ID028 ID027 ID033 ID045 ID030 ID040 ID038 ID029 ID032 ID036 ID049 ID048 ID083 ID071 ID081 ID079 ID073 ID070 ID061 ID072 ID050 ID063 ID047 ID056 ID060 ID068 ID064 ID052 ID062 ID053 ID077 ID054 ID065 ID076 ID051 ID075 ID059 ID078 ID074 ID055 ID067 ID069 ID080 ID082 ID058 ID066 ID057 Correlation between Subjective Score and Spectral Entropy (SE) Surgeon ID Regression Line (Slope: -0.729) (b) Spectral Entropy (SE) Fig. 3: Correlation plots for different metrics. and lessens the potential for tissue damage. These improvements in motor plan- ning and execution are especially valuable in complex operations like capsulorhexis, where even minor inefficiencies can compromise patient outcomes. 4. Motion Difficulty (MD): In Figure 3(d), the analysis of the MD metric demonstrates a significant inverse relationship with subjective surgical skill scores, as indicated by a regression slope of –44.944. This result reveals that higher skill levels are associated with smoother movements, thereby minimizing abrupt changes in accel- eration. This reduction in motion difficulty is critical for decreasing tissue trauma and achieving optimal surgical outcomes [49]. 5. Energy Consumption (EC): The analysis in Figure 3(e) reveals a negative regression slope of –0.227, indicating that higher skill corresponds with lower energy con- sumption. This relationship suggests that proficient surgeons execute more efficient movements that minimize unnecessary kinetic energy expenditure. Consequently, such energy efficiency is a marker of refined motor control. 6. Number of Local Peaks (NLP): The regression analysis depicted in Figure 3(f) shows that the NLP is strongly negatively correlated with subjective surgical skill scores (slope: –1298.269). This finding suggests that more proficient surgeons achieve more fluid and controlled instrument movements, marked by fewer abrupt fluctuations. Enhanced motor control and precision, as evidenced by reduced NLP values, are vital in minimizing tissue trauma and improving surgical outcomes. 7. Total Force (TF): The analysis of the TF in Figure 3(g) reveals a strong inverse correlation (slope: –25.394) with subjective skill scores. Surgeons with higher profi- ciency achieve more fluid and controlled instrument handling, reducing unnecessary force application. 8. Tissue Interaction (TI): The regression analysis in Figure 3(h) shows a negative slope of –0.059, implying that increased surgical proficiency is associated with reduced TI values. Such reduction and consistency in motion reduce the likelihood of abrupt force fluctuations that could damage tissue, underscoring the importance of uniform and controlled acceleration patterns in ensuring successful outcomes in sensitive surgical interventions. 13 9. Speed Frequency Smoothness (SFS): In the frequency domain, abrupt and inade- quately planned movements are typically represented by high-frequency coefficients, while smoother movements correspond to lower-frequency coefficients [50]. The analysis depicted in Figure 3(a) shows that the SFS metric is inversely correlated with subjective surgical skill scores (regression slope: –0.171). Lower SFS values, associated with higher proficiency, indicate smoother speed profiles with fewer high-frequency fluctuations. 10. Spectral Entropy (SE): The regression analysis in Figure 3(b) shows that the spec- tral entropy (SE) metric exhibits a negative slope of –0.729. This result implies that higher-skill surgeons demonstrate lower SE values, meaning that the motion signal’s energy is more tightly focused in particular frequency bands. This energy distribution pattern indicates consistent instrument movements and underscores improved motor control and precision. These correlations confirm the reliability of the proposed framework in replicating expert-based evaluations. By employing video-derived metrics as automated proxies for real-world competence, this approach enables consistent and scalable skill assessments across various surgical contexts. In the final stage of our analysis, multiple clustering algorithms, including AggAv- erage, AggComplete, AggSingle, Agglomerative, BayesianGMM, Birch, BisectingK- Means, GaussianMixture, KMeans, MiniBatchKMeans, and Spectral, were employed to partition surgeons into two distinct groups based on their Overall Score (OS) as determined by expert subjective evaluation. The group exhibiting higher OS values was designated as “Expert (E),” whereas the lower-scoring group was labeled “Inter- mediate (I).” This OS-based clustering served as the ground truth for subsequent comparisons. Subsequently, the same clustering techniques were applied independently to ten objective performance metrics derived from our intelligent assessment framework. Consistent with the inverse relationships observed in Figure 3, lower values of these metrics corresponded to the “Expert” cluster, while higher values mapped to the “Intermediate” cluster. Subsequently, each objective metric’s clustering was compared against the OS- based clustering to assess how accurately it could distinguish between “Expert” and “Intermediate” surgeons. The formula below measures this accuracy by quantifying the overlap of surgeon IDs in the E and I clusters derived from each objective metric with the corresponding E and I clusters obtained via OS-based clustering. Thus, it provides a straightforward indicator of how well each objective metric aligns with the expert-derived groupings: Accuracy = |E OS ∩E metric | +|I OS ∩I metric | |E OS | +|I OS | (11) where E OS and I OS are the sets of surgeons labeled “Expert” and “Intermediate” by the OS-based clustering, andE metric andI metric are the corresponding sets obtained from the metric-based clustering. Table 2 summarizes the highest accuracy achieved for each indicator. For each objective metric, the table lists the clustering method that produced the optimal 14 accuracy when compared with OS-based labels, along with the corresponding threshold ranges (minimum and maximum values) for both the “Expert” and “Intermediate” clusters. Among all the objective indicators, NAM clustered via the Agglomerative method exhibited the greatest alignment with the surgeon groupings derived from subjective scoring. In the case of the Overall Score, the table reports the aggregated thresholds across all methods, with an observed overlap in the range of 3.23 to 3.78. This overlap reflects the inherent variability in subjective assessments when consolidated over multiple clustering approaches. Table 2: Top clustering accuracies and threshold ranges for classifying sur- geons into Expert and Intermediate groups. IndicatorMethodAccuracy (%)Thresholds for [E], [I] PTKMeans82[17-93], [97-288] PLKMeans84[69-505], [544-1327] NAMAgglomerative87[404-3795], [4063-9135] MDSpectral85[14-60], [63-217] ECGaussianMixture82[0.06-0.39], [0.40-1.42] NLPSpectral85[350-1582], [1616-6115] TFSpectral85[8-34], [36-123] TIGaussianMixture75[0.05-0.21], [0.22-0.58] SFSAgglomerative81[0.21-0.52], [0.54-1.10] SESpectral84[5.84-7.37], [7.40-8.66] Overall Score (OS)All Methods-[2.39-3.78], [3.23-4.88] Overall, our findings emphasize the critical role of precise hand-eye coordination and muscle control in surgical performance. These results align closely with pre- vious research highlighting the importance of motion smoothness, controlled force application, and careful instrument handling for successful surgical outcomes [51–54]. Additionally, our study represents a significant advancement beyond prior inves- tigations, such as Kim et al. [55], by examining a more comprehensive set of skill indicators derived from detailed motion data, coupled with a rigorously annotated dataset. Our analysis clearly demonstrates that expert surgical performance depends heavily on the effective coordination of visual perception, hand movements, and muscular precision. 4 Conclusion In this paper, we introduced an explainable AI framework designed to evaluate sur- gical skills in cataract procedures automatically. The objective framework presented here offers an accurate and scalable approach for assessing surgical proficiency. This approach has strong potential to complement or even replace existing subjective assessment systems like ICO-OSCAR. Furthermore, it can play a valuable role in refining surgical training programs by providing targeted, objective feedback, thereby systematically guiding trainees toward improved surgical outcomes. 15 The study stands out for several key achievements. First, we crafted the world’s largest repository of cataract surgery videos, encompassing 2,000 recordings, of which 83 were meticulously annotated using the newly established Capsulorhexis Skill Assess- ment System (CSAS). This dataset enabled a reliable benchmark for subjective (expert-driven) and objective (motion-based) measurements, thereby bridging a long- standing gap in AI-driven surgical education research. Notably, it surpasses JIGSAWS [19] by capturing real surgical scenes rather than simulated tasks, and it improves upon Cataract-101 [34] and Cataract-1K [56] through a structured, multi-indicator annotation process conducted by three independent physicians. Second, our approach proposed a data-driven pipeline that extracts clinically mean- ingful motion signals from the raw video stream. Through advanced computer vision algorithms for segmentation and tracking, we converted pixel-level motion data into quantifiable metrics that thoroughly reflect technical precision, smoothness, and tis- sue handling. This capability addresses the need for objective, reproducible indicators of surgical expertise. Third, and most critically, we demonstrated that our ten computed motion-based metrics, ranging from path length to tissue interaction measures, exhibit robust corre- lations with expert-derived capsulorhexis skill ratings, achieving up to 87% accuracy in skill classification. This level of performance indicates that automated, objective measures can match or even exceed traditional subjective evaluations, challenging the notion that high-quality surgical skill assessments must rely solely on human observers. Fourth, our focus on explainability stands as a cornerstone achievement. By trans- lating raw computational outputs into clinically interpretable indices, we offer not just an accuracy-driven tool but also a transparent framework readily understood by surgical educators and trainees. Such interpretability fosters trust in AI and provides targeted feedback, allowing surgeons to pinpoint, for example, whether excessive force, abrupt acceleration, or instrument misalignment underlies lower performance scores. Despite these promising results, certain limitations persist. The current validation focuses heavily on the capsulorhexis phase, and additional studies are essential to ensure generalizability across other surgical tasks and clinical environments. The vary- ing complexities of multi-institutional settings also underscore the necessity for broader validation. Furthermore, while the proposed metrics operate efficiently in retrospective analyses, the challenges of real-time feedback integration, such as processing speed, user interface design, and integration within the clinical workflow, remain outstanding. Future directions for this research thus include (1) broadening the scope to cap- ture other phases of cataract surgery (e.g., phacoemulsification and intraocular lens placement) and extending to different surgical specialties, (2) conducting large-scale, multi-center validation studies to account for diverse patient populations and health- care infrastructures, (3) developing real-time coaching systems that combine motion analytics with intraoperative feedback to foster faster learning curves, and (4) incor- porating advanced simulator and virtual reality platforms, allowing trainees to benefit from immediate, data-driven metrics in a safe, controlled setting, (5) undertaking a clinical study to assess and model the learning curve of surgical trainees employing this method, and (6) leveraging outputs from objective metrics and threshold-based clus- tering to automatically generate personalized skill reports, complete with strengths, 16 weaknesses, and targeted suggestions, by integrating large language models (LLMs). Ultimately, such efforts may help alleviate global workforce shortages and democratize access to advanced surgical training by providing scalable, objective, and explainable AI-based tools. References [1] Center, V.U.M.: Global Surgery Facts and Figures. Vanderbilt University Medical Center. Accessed: April 28, 2024 (2024) [2] UCSF Department of Surgery: Global Surgery 2030: evidence and solutions for achieving health, welfare, and economic development. Last accessed: March 1, 2025 (Accessed 2025). https://surgery.ucsf.edu/sites/default/files/umbraco/ media/8065752/OverviewGS2030.pdf [3] Augestad, K.M., Butt, K., Ignjatovic, D., Keller, D.S., Kiran, R.: Video-based coaching in surgical education: a systematic review and meta-analysis. Surgical Endoscopy 34(2), 521–535 (2019) https://doi.org/10.1007/s00464-019-07265-0 [4] MBZUAI: AI-Driven Surgical Skill Optimization. https://mbzuai.ac.ae/ the-node/ai-talks/ai-talk/ai-driven-surgical-skill-optimization/.Accessed: 2024-12-18 (2024) [5] Nathan, M., al.: Intraoperative adverse events can be compensated by technical performance in neonates and infants after cardiac surgery: A prospective study. The Journal of Thoracic and Cardiovascular Surgery 142(5), 1098–11075 (2011) https://doi.org/10.1016/j.jtcvs.2011.07.003 [6] Papaspyros, S.C., Javangula, K.C., Prasad Adluri, R.K., O’Regan, D.J.: Briefing and debriefing in the cardiac operating room. analysis of impact on theatre team attitude and patient safety. Interactive CardioVascular and Thoracic Surgery 10(1), 43–47 (2010) https://doi.org/10.1510/icvts.2009.217356 [7] Tevis, S.E., Kennedy, G.D.: Postoperative complications and implications on patient-centered outcomes. Elsevier BV (2013). https://doi.org/10.1016/j.jss. 2013.01.032 . http://dx.doi.org/10.1016/j.jss.2013.01.032 [8] Semel, M.E., Lipsitz, S.R., Funk, L.M., Bader, A.M., Weiser, T.G., Gawande, A.A.: Rates and patterns of death after surgery in the United States, 1996 and 2006. Elsevier BV (2012). https://doi.org/10.1016/j.surg.2011.07.021 . http://dx. doi.org/10.1016/j.surg.2011.07.021 [9] Weiser, T.G., al.: An estimation of the global volume of surgery: a modelling strategy based on available data. The Lancet 372(9633), 139–144 (2008) https: //doi.org/10.1016/s0140-6736(08)60878-8 [10] Regenbogen, S.E., Greenberg, C.C., Studdert, D.M., Lipsitz, S.R., Zinner, 17 M.J., Gawande, A.A.: Patterns of technical error among surgical malpractice claims. Annals of Surgery 246(5), 705–711 (2007) https://doi.org/10.1097/sla. 0b013e31815865f8 [11] Gallant, J., Brelsford, K., Sharma, S., Grantcharov, T., Langerman, A.: Patient perceptions of audio and video recording in the operating room. Annals of Surgery 276(6), 1057–1063 (2021) [12] Levin, M., McKechnie, T., Kruse, C.C., Aldrich, K., Grantcharov, T.P., Langer- man, A.: Surgical data recording in the operating room: a systematic review of modalities and metrics. British Journal of Surgery 108(6), 613–621 (2021) [13] Theator: Three Benefits of Recording Surgical Videos (2024). https://theator.io/ blog/benefits-of-recording-surgical-videos/ [14] Cheikh Youssef, S., Haram, K., No ̈el, J., Patel, V., Porter, J., Dasgupta, P., Hachach-Haram, N.: Evolution of the digital operating room: the place of video technology in surgery. Springer (2023). https://doi.org/10.1007/ s00423-023-02830-7 . http://dx.doi.org/10.1007/s00423-023-02830-7 [15] Awshah, S., Bowers, K., Eckel, D.T., Diab, A.F., Ganam, S., Sujka, J., Docimo, S., DuCoin, C.: Current trends and barriers to video management and ana- lytics as a tool for surgeon skilling. Springer (2024). https://doi.org/10.1007/ s00464-024-10754-6 . http://dx.doi.org/10.1007/s00464-024-10754-6 [16] Eckhoff, J.A., Rosman, G., Altieri, M.S., Speidel, S., Stoyanov, D., Anvari, M., Meier-Hein, L., M ̈arz, K., Jannin, P., Pugh, C., Wagner, M., Witkowski, E., Shaw, P., Madani, A., Ban, Y., Ward, T., Filicori, F., Padoy, N., Talamini, M., Meireles, O.R.: SAGES consensus recommendations on surgical video data use, structure, and exploration (for research in artificial intelligence, clinical quality improvement, and surgical education). Springer (2023). https://doi.org/10.1007/ s00464-023-10288-3 . http://dx.doi.org/10.1007/s00464-023-10288-3 [17] Ahmadi, M.J., Allahkaram, M.S., Abdi, P., Mohammadi, S.-F., D. Taghi- rad, H.a.: Image processing and machine vision in surgery and its train- ing. Journal of Control 17(2) (2023) https://doi.org/10.61186/joc.17.2.25 http://joc.kntu.ac.ir/article-1-999-en.pdf [18] Funke, I., Mees, S.T., Weitz, J., Speidel, S.: Video-based surgical skill assess- ment using 3D convolutional neural networks. Springer (2019). https://doi.org/ 10.1007/s11548-019-01995-1 . http://dx.doi.org/10.1007/s11548-019-01995-1 [19] Gao, Y., Vedula, S.S., Reiley, C.E., Ahmidi, N., Varadarajan, B., Lin, H.C., Tao, L., Zappella, L., B ́ejar, B., Yuh, D.D., Chen, C.C.G., Vidal, R., Khudanpur, S., Hager, G.D.: The jhu-isi gesture and skill assessment working set (jigsaws): A sur- gical activity dataset for human motion modeling. In: Modeling and Monitoring of Computer Assisted Interventions (M2CAI) – MICCAI Workshop (2014) 18 [20] Kitaguchi, D., Takeshita, N., Matsuzaki, H., Igaki, T., Hasegawa, H., Ito, M.: Development and Validation of a 3-Dimensional Convolutional Neu- ral Network for Automatic Surgical Skill Assessment Based on Spatiotem- poral Video Analysis. American Medical Association (AMA) (2021). https: //doi.org/10.1001/jamanetworkopen.2021.20786 . http://dx.doi.org/10.1001/ jamanetworkopen.2021.20786 [21] Zia, A., Guo, L., Zhou, L., Essa, I., Jarc, A.: Novel evaluation of surgical activity recognition models using task-based efficiency metrics. Springer (2019). https://doi.org/10.1007/s11548-019-02025-w . http://dx.doi. org/10.1007/s11548-019-02025-w [22] Lajk ́o, G., Nagyn ́e Elek, R., Haidegger, T.: Endoscopic Image-Based Skill Assess- ment in Robot-Assisted Minimally Invasive Surgery. MDPI AG (2021). https: //doi.org/10.3390/s21165412 . http://dx.doi.org/10.3390/s21165412 [23] Phacoemulsification Devices Market Analysis - US, Canada, Germany, UK, China - Size and Forecast 2024-2028. Technavio. Accessed: April 28, 2024 (2023) [24] Flaxman, S.R., Bourne, R.R.A., Resnikoff, S., Ackland, P., Braithwaite, T., Cicinelli, M.V., Das, A., Jonas, J.B., Keeffe, J., Kempen, J.H., Leasher, J., Limburg, H., Naidoo, K., Pesudovs, K., Silvester, A., Stevens, G.A., Tahhan, N., Wong, T.Y., Taylor, H.R., Bourne, R., Ackland, P., Arditi, A., Barkana, Y., Bozkurt, B., Braithwaite, T., Bron, A., Budenz, D., Cai, F., Casson, R., Chakravarthy, U., Choi, J., Cicinelli, M.V., Congdon, N., Dana, R., Dandona, R., Dandona, L., Das, A., Dekaris, I., Del Monte, M., Deva, J., Dreer, L., Ellwein, L., Frazier, M., Frick, K., Friedman, D., Furtado, J., Gao, H., Gazzard, G., George, R., Gichuhi, S., Gonzalez, V., Hammond, B., Hartnett, M.E., He, M., Hejtman- cik, J., Hirai, F., Huang, J., Ingram, A., Javitt, J., Jonas, J., Joslin, C., Keeffe, J., Kempen, J., Khairallah, M., Khanna, R., Kim, J., Lambrou, G., Lansingh, V.C., Lanzetta, P., Leasher, J., Lim, J., Limburg, H., Mansouri, K., Mathew, A., Morse, A., Munoz, B., Musch, D., Naidoo, K., Nangia, V., Palaiou, M., Parodi, M.B., Pena, F.Y., Pesudovs, K., Peto, T., Quigley, H., Raju, M., Ramulu, P., Rankin, Z., Resnikoff, S., Reza, D., Robin, A., Rossetti, L., Saaddine, J., Sandar, M., Serle, J., Shen, T., Shetty, R., Sieving, P., Silva, J.C., Silvester, A., Sitorus, R.S., Stambolian, D., Stevens, G., Taylor, H., Tejedor, J., Tielsch, J., Tsilimbaris, M., Meurs, J., Varma, R., Virgili, G., Wang, Y.X., Wang, N.-L., West, S., Wiede- mann, P., Wong, T., Wormald, R., Zheng, Y.: Global causes of blindness and distance vision impairment 1990–2020: a systematic review and meta-analysis. Lancet Glob. Health 5(12), 1221–1234 (2017) [25] Khairallah, M., Kahloun, R., Bourne, R., Limburg, H., Flaxman, S.R., Jonas, J.B., Keeffe, J., Leasher, J., Naidoo, K., Pesudovs, K., Price, H., White, R.A., Wong, T.Y., Resnikoff, S., Taylor, H.R., Vision Loss Expert Group of the Global Burden of Disease Study: Number of people blind or visually impaired by cataract worldwide and in world regions, 1990 to 2010. Invest. Ophthalmol. Vis. Sci. 19 56(11), 6762–6769 (2015) [26] Lacmanovi ́c Lonˇcar, V.: The resident surgeon phacoemulsification learning curve at clinical department of ophthalmology, sestre milosrdnice university hospital center. Acta Clin. Croat., 549–554 (2016) [27] Neufeld, A., Hanson, L.L., Pettey, J.: Teaching in the operating room: trends in surgical skills transfer in ophthalmology. Ann. Eye Sci. 2, 41–41 (2018) [28] Bharucha, K.M., Adwe, V.G., Hegade, A.M., Deshpande, R.D., Deshpande, M.D., Kalyani, V.K.S.: Evaluation of skills transfer in short-term phacoemul- sification surgery training program by international council of ophthalmology -ophthalmology surgical competency assessment rubrics (ICO-OSCAR) and assessment of efficacy of ICO-OSCAR for objective evaluation of skills transfer. Indian J. Ophthalmol. 68(8), 1573–1577 (2020) [29] Prakash, G., Jhanji, V., Sharma, N., Gupta, K., Titiyal, J.S., Vajpayee, R.B.: Assessment of perceived difficulties by residents in performing routine steps in phacoemulsification surgery and in managing complications. Can. J. Ophthalmol. 44(3), 284–287 (2009) [30] Golnik, K.C., Beaver, H., Gauba, V., Lee, A.G., Mayorga, E., Palis, G., Saleh, G.M.: Cataract surgical skill assessment. Ophthalmology 118(2), 427–15 (2011) [31] Cremers, S.L., Ciolino, J.B., Ferrufino-Ponce, Z.K., Henderson, B.A.: Objective assessment of skills in intraocular surgery (OASIS). Ophthalmology 112(7), 1236– 1241 (2005) [32] Cremers, S.L., Lora, A.N., Ferrufino-Ponce, Z.K.: Global rating assessment of skills in intraocular surgery (GRASIS). Ophthalmology 112(10), 1655–1660 (2005) [33] Saleh, G.M., Gauba, V., Mitra, A., Litwin, A.S., Chung, A.K.K., Benjamin, L.: Objective structured assessment of cataract surgical skill. Arch. Ophthalmol. 125(3), 363–366 (2007) [34] Schoeffmann, K., Taschwer, M., Sarny, S., M ̈unzer, B., Primus, M.J., Putzgru- ber, D.: Cataract-101: video dataset of 101 cataract surgeries. In: C ́esar, P., Zink, M., Murray, N. (eds.) Proceedings of the 9th ACM Multimedia Systems Conference, MMSys 2018, Amsterdam, The Netherlands, June 12-15, 2018, p. 421–425. ACM, ??? (2018). https://doi.org/10.1145/3204949.3208137 . https: //doi.org/10.1145/3204949.3208137 [35] Adwe, V., Bharucha, K., Hegade, A., Deshpande, R., Deshpande, M., Kalyani, V.S.: Evaluation of skills transfer in short-term phacoemulsification surgery train- ing program by international council of ophthalmology -ophthalmology surgical competency assessment rubrics (ico-oscar) and assessment of efficacy of ico-oscar 20 for objective evaluation of skills transfer. Indian Journal of Ophthalmology 68(8), 1573 (2020) https://doi.org/10.4103/ijo.ijo205819 [36] Dean, W.H., Murray, N.L., Buchan, J.C., Golnik, K., Kim, M.J., Burton, M.J.: Ophthalmic simulated surgical competency assessment rubric for manual small- incision cataract surgery. Journal of Cataract and Refractive Surgery 45(9), 1252– 1257 (2019) https://doi.org/10.1016/j.jcrs.2019.04.010 [37] Puri, S., Sikder, S.: Cataract surgical skill assessment tools. Journal of Cataract and Refractive Surgery 40(4), 657–665 (2014) https://doi.org/10.1016/j.jcrs. 2014.01.027 [38] Saleh, G.M.: Objective structured assessment of cataract surgical skill. Archives of Ophthalmology 125(3), 363 (2007) https://doi.org/10.1001/archopht.125.3.363 [39] Golnik, K.C., Beaver, H., Gauba, V., Lee, A.G., Mayorga, E., Palis, G., Saleh, G.M.: Cataract surgical skill assessment. Ophthalmology 118(2), 427–4275 (2011) https://doi.org/10.1016/j.ophtha.2010.09.023 [40] Cremers, S.L., Lora, A.N., Ferrufino-Ponce, Z.K.: Global rating assessment of skills in intraocular surgery (grasis). Ophthalmology 112(10), 1655–1660 (2005) https://doi.org/10.1016/j.ophtha.2005.05.010 [41] Ahmadi, M.J., Gandomi, I., Abdi, P., et al.: Cataract-lmm large-scale multi- source multi-task benchmark for deep learning in surgical video analysis. Scientific Data 13, 1189 (2026) https://doi.org/10.1038/s41597-026-07464-0 [42] Dreyfus, S.E.: The Five-Stage Model of Adult Skill Acquisition. Accessed: 2025-03-04 (2004). https://w.bumc.bu.edu/facdev-medicine/files/2012/03/ Dreyfus-skill-level.pdf [43] Ravi, N., Gabeur, V., Hu, Y.-T., Hu, R., Ryali, C., Ma, T., Khedr, H., R ̈adle, R., Rolland, C., Gustafson, L., Mintun, E., Pan, J., Alwala, K.V., Carion, N., Wu, C.-Y., Girshick, R., Doll ́ar, P., Feichtenhofer, C.: SAM 2: Segment Anything in Images and Videos (2024). https://arxiv.org/abs/2408.00714 [44] Khanam, R., Hussain, M.: YOLOv11: An Overview of the Key Architectural Enhancements (2024). https://arxiv.org/abs/2410.17725 [45] Ahmadi, M.J., Allahkaram, M.S., Rashvand, A., Lotfi, F., Abdi, P., Motahari- far, M., Mohammadi, S.F., Taghirad, H.D.: Aras-farabi experimental framework for skill assessment in capsulorhexis surgery. In: 2021 9th RSI International Conference on Robotics and Mechatronics (ICRoM), p. 385–390. IEEE, ??? (2021). https://doi.org/10.1109/icrom54204.2021.9663494 . http://dx.doi.org/10. 1109/ICRoM54204.2021.9663494 21 [46] Lafouti, M., Ahmadi, M.J., Allahkaram, M.S., Gandomi, I., Lotfi, F., Moham- madzadeh, M., Abdi, P., Taghirad, H.D.: Surgical instrument tracking for capsulorhexis eye surgery based on siamese networks. In: 2022 10th RSI Interna- tional Conference on Robotics and Mechatronics (ICRoM), p. 196–201. IEEE, ??? (2022). https://doi.org/10.1109/icrom57054.2022.10025355 . http://dx.doi. org/10.1109/ICRoM57054.2022.10025355 [47] Gandomi, I., Vaziri, M., Ahmadi, M.J., Reyhaneh Hadipour, M., Abdi, P., Taghi- rad, H.D.: A deep dive into capsulorhexis segmentation: From dataset creation to sam fine-tuning. In: 2023 11th RSI International Conference on Robotics and Mechatronics (ICRoM), p. 675–681. IEEE, ??? (2023). https://doi.org/10. 1109/icrom60803.2023.10412370 . http://dx.doi.org/10.1109/ICRoM60803.2023. 10412370 [48] R??fer, F., Schr??der, A., Erb, C.: White-to-white corneal diameter: Normal val- ues in healthy humans obtained with the orbscan i topography system. Cornea 24(3), 259–261 (2005) https://doi.org/10.1097/01.ico.0000148312.01805.53 [49] Hwang, H., Lim, J., Kinnaird, C., Nagy, A.G., Panton, O.N.M., Hodgson, A.J., Qayumi, K.A.: Correlating motor performance with surgical error in laparoscopic cholecystectomy. Surgical Endoscopy 20(4), 651–655 (2005) https://doi.org/10. 1007/s00464-005-0370-8 [50] Soleymani, A., Sadat Asl, A.A., Yeganejou, M., Dick, S., Tavakoli, M., Li, X.: Surgical skill evaluation from robot-assisted surgery recordings. In: 2021 International Symposium on Medical Robotics (ISMR), p. 1–6. IEEE, ??? (2021). https://doi.org/10.1109/ismr48346.2021.9661527 . http://dx.doi.org/10. 1109/ISMR48346.2021.9661527 [51] Kletz, S., Schoeffmann, K., Leibetseder, A., Benois-Pineau, J., Husslein, H.: Instrument Recognition in Laparoscopy for Technical Skill Assessment, p. 589– 600. Springer, ??? (2019). https://doi.org/10.1007/978-3-030-37734-2 48 . http: //dx.doi.org/10.1007/978-3-030-37734-248 [52] Nau, P., Worden, E., Lehmann, R., Kleppe, K., Mancini, G.J., Mancini, M.L., Ramshaw, B.: Global assessment of surgical skills (gass): validation of a new instrument to measure global technical safety in surgical proce- dures. Surgical Endoscopy 37(10), 7964–7969 (2023) https://doi.org/10.1007/ s00464-023-10116-8 [53] Aghazadeh, F., Zheng, B., Tavakoli, M., Rouhani, H.: Motion smoothness-based assessment of surgical expertise: The importance of selecting proper metrics. Sensors 23(6), 3146 (2023) https://doi.org/10.3390/s23063146 [54] Hutchinson, K., Chen, K., Alemzadeh, H.: Towards Interpretable Motion-level Skill Assessment in Robotic Surgery. arXiv (2023). https://doi.org/10.48550/ ARXIV.2311.05838 . https://arxiv.org/abs/2311.05838 22 [55] Kim, T.S., O’Brien, M., Zafar, S., Hager, G.D., Sikder, S., Vedula, S.S.: Objec- tive assessment of intraoperative technical skill in capsulorhexis using videos of cataract surgery. International Journal of Computer Assisted Radiology and Surgery 14(6), 1097–1105 (2019) https://doi.org/10.1007/s11548-019-01956-8 [56] Ghamsarian, N., El-Shabrawi, Y., Nasirihaghighi, S., Putzgruber-Adamitsch, D., Zinkernagel, M., Wolf, S., Schoeffmann, K., Sznitman, R.: Cataract-1k dataset for deep-learning-assisted analysis of cataract surgery videos. Scientific Data 11(1) (2024) https://doi.org/10.1038/s41597-024-03193-4 23