Source-linked AI summary

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

Mohammad Javad Ahmadi, Hamid D. Taghirad

arXiv:2608.17522v1cs.CVcs.AI

TL;DR

Traditional cataract-surgery skill assessment relies on lengthy training and subjective evaluation, motivating automated, data-driven alternatives. The paper introduces an explainable AI framework and large video dataset, achieving up to 87% accuracy while correlating robustly with expert ratings.

  • Problem

    Capsulorhexis skill assessment remains challenging for trainees and relies on conventional rubric-based evaluation despite the availability of informative surgical videos.

  • Method

    The study constructs a 2,000-recording cataract-surgery video dataset and develops an explainable AI framework that generates interpretable motion-based performance metrics.

  • Results

    87% accuracy was achieved in skill classification, with ten computed motion-based metrics showing robust correlations with expert-derived capsulorhexis skill ratings.

  • Takeaways & Limitations

    The framework provides transparent, objective indicators that can support or potentially replace subjective assessment and deliver targeted feedback for surgical training.

  • Takeaways & Limitations

    Validation focuses heavily on capsulorhexis, so broader studies are needed across surgical tasks, clinical environments, and multi-institutional settings.

Abstract

from arXiv · show

Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automated, data-driven approaches in surgical education. This study addresses these challenges by introducing a novel, explainable AI-powered framework for automated skill assessment, specifically focusing on cataract surgery. We present the world's largest dataset of cataract surgery videos, comprising 2,000 recordings. Additionally, we propose an AI-powered analytical framework that employs advanced computer vision and signal-processing techniques to automatically evaluate surgical videos to derive objective, quantitative performance indicators that complement or potentially replace subjective scoring methods. A significant advantage of our framework over previous methods lies precisely in its explainability of outputs, elevating it beyond merely an opaque skill classification tool. Through experimental analysis of 83 cataract surgery videos, we demonstrate that the automatically computed metrics exhibit strong correlations with expert-based subjective evaluations, achieving up to 87% accuracy in surgical skill assessment. Each metric was individually examined, and expert surgeons provided subjective ratings using the newly introduced Capsulorhexis Skill Assessment System (CSAS). These subjective assessments were compared with ten objective motion-based metrics extracted through our framework. The results indicated a robust correlation between subjective ratings and automated indicators, underscoring the framework's capacity to accurately model surgical expertise.

1 Introduction

Rising surgical needs and limitations of subjective, trainer-intensive assessment motivate automated video-based skill evaluation. This paper introduces an explainable AI framework, demonstrated on cataract-surgery capsulorhexis, to provide objective assessment while reducing supervision demands, costs, time, and bias.

  • Motivation: Global demand for surgical services is increasing, while low- and middle-income countries face a shortfall of 143 million essential surgeries annually.Every country is expected to require at least 5,000 surgical procedures per 100,000 people annually by 2030.
  • Motivation: Traditional surgical training relies on subjective evaluations and limited trainer feedback, producing biased or inconsistent assessments.The Master-Apprentice Model is constrained by trainers’ high workloads and limited time for detailed feedback.
  • Motivation: Early-career surgeons are nearly nine times more likely to commit procedural errors than experienced surgeons, underscoring the need for improved training and assessment.The shortage of skilled surgeons is associated with longer patient wait times, increased surgical errors, and patient deaths.
  • Video-based assessment: Surgical videos are accessible through operating-room devices and capture the complete surgical process, enabling AI-based analysis without expensive sensor or kinematic equipment.However, reviewing extensive video archives is time-consuming and expensive, motivating systems that streamline video management, review, and analysis.
  • Related work: Existing automated surgical-video research focuses mainly on phases and workflows, while objective skill assessment remains limited and deep-learning outputs are often difficult to interpret.Prior approaches include models achieving over 95% accuracy for three-class simulated-data classification and around 80% accuracy using manually selected tool regions, but they lack interpretable or fully automated feedback.
  • Proposed framework: The proposed framework integrates explainable methods for objective, automated skill evaluation, using cataract-surgery capsulorhexis as a proof-of-concept case study.Capsulorhexis is a challenging step for first-year residents, and the framework aims to eliminate direct trainer supervision while reducing assessment costs, time, and potential bias.

2 Methods

The methods establish a high-quality cataract-surgery video resource and an explainable framework that converts surgical motion into interpretable skill indicators. The framework combines phase-specific expert annotation with computer-vision tracking, relative motion processing, scaling, and higher-order dynamics.

  • Dataset and framework: The study developed a surgical-video dataset and an AI framework that generates interpretable performance metrics for automated skill evaluation.The framework is intended to support interdisciplinary AI research and provide ground truth for validating automated assessment systems.
  • Motion extraction: The framework segments and tracks the capsulorhexis instrument and corneal tissue, then uses tool-tip and corneal-center coordinates as foundational motion data.Pre-trained segmentation models were applied to the CVSAD-83 videos, while the framework is described as combining computer vision with post-processing for actionable skill information.
  • Motion processing: Relative referencing expresses tool-tip motion against the corneal center, reducing distortions caused by eye movements and microscope-field changes.This scheme produces smoother, more interpretable variations than absolute frame-based coordinates for skill-related analysis.
  • Objective metrics: The processed trajectories are scaled for microscope magnification and transformed into velocity, acceleration, jerk, path length, alternating movements, and speed-frequency smoothness metrics.Scaling uses the measured corneal diameter and known corneal diameter of 11.71 ± 0.42 mm; the metrics derive from expert-informed subjective assessment principles.

3 Results

Across 83 CVSAD videos, CSAS subjective scores showed generally strong performance with variation in motion-related tasks, while all ten objective indicators were inversely related to expert Overall Scores. Clustering based on objective metrics reproduced expert-versus-intermediate groupings, with NAM under Agglomerative clustering showing the greatest alignment.

  • Subjective score distribution: Most CSAS indicators had median scores near or above 3.75 across the 83 evaluated CVSAD videos.Microscope Use and Tissue Handling scored slightly higher than Motion and Circular Completion, indicating greater variation in motion-related tasks.
  • Objective–subjective correlations: All ten objective indicators exhibited inverse relationships with expert Overall Scores, with higher-rated surgeons typically showing lower metric values.Figure 3 compared each metric against OS using labeled video points and regression lines.
  • Objective–subjective correlations: –60.057 was the regression slope for Procedure Time (PT), which decreased as subjective skill scores increased.The result associates higher skill with greater procedural efficiency.
  • Objective–subjective correlations: –44.944 was the regression slope for Motion Difficulty (MD), demonstrating a significant inverse relationship with subjective surgical skill scores.The passage states that higher skill levels were associated with this inverse relationship.
  • Objective–subjective correlations: –0.171 was the regression slope for Speed Frequency Smoothness (SFS), with lower SFS values indicating smoother speed profiles and higher proficiency.The frequency-domain interpretation links abrupt movements to high-frequency coefficients and smoother movements to lower-frequency coefficients.
  • Clustering-based skill grouping: NAM clustered via Agglomerative achieved the greatest alignment with surgeon groupings derived from subjective scoring, while lower objective values corresponded to the Expert cluster.Objective-metric clustering consistently mapped lower values to Expert and higher values to Intermediate; Overall Score thresholds overlapped from 3.23 to 3.78.

4 Conclusion

The conclusion presents an explainable AI framework for scalable, objective cataract-surgery skill assessment, supported by a 2,000-video repository and validation achieving up to 87% accuracy. It emphasizes clinically interpretable feedback while identifying generalizability and real-time integration as remaining challenges.

  • Contributions: The study established the world’s largest cataract-surgery video repository, containing 2,000 recordings, including 83 annotated with CSAS.The dataset supports benchmarking expert subjective ratings against objective motion-based measurements.
  • Contributions: The framework extracts clinically meaningful motion signals from video using computer-vision segmentation and tracking, producing quantitative indicators of precision, smoothness, and tissue handling.These indicators are intended to provide objective and reproducible measures of surgical expertise.
  • Results: 87% accuracy was achieved in skill classification using ten motion-based metrics that robustly correlated with expert-derived capsulorhexis skill ratings.The metrics ranged from path length to tissue-interaction measures.
  • Explainability: Explainability translates computational outputs into clinically interpretable indices that can reveal contributors to lower performance, including excessive force, abrupt acceleration, or instrument misalignment.This transparency is intended to foster trust and support targeted feedback for educators and trainees.
  • Limitations: Validation remains concentrated on the capsulorhexis phase, requiring further study across surgical tasks, clinical environments, and multi-institutional settings, alongside real-time feedback integration.The passage notes that the metrics operate efficiently in retrospective analyses but that real-time integration remains challenging.
  • Future directions: Future work includes assessing additional cataract-surgery phases and specialties, conducting large-scale multi-center validation, tracking trainee learning curves, and generating personalized reports with LLMs.These reports are envisioned to include strengths, weaknesses, and targeted suggestions.
Loading 2608.17522v1…