SearcharxivSearch

arXiv subjects

Rebecca Hisey

Publications and source records attributed to Rebecca Hisey.

4 recordsLinked to original sources

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

Vision-based surgical skill assessment has shown strong in-domain results, yet a fundamental question remains unasked: do these models learn transferable representations of surgical proficiency, or do they merely encode dataset-specific visual patterns? This paper systematically analyzes what limits cross-domain skill transfer between the GOALS and OSATS assessment scales using the LASANA and JIGSAWS datasets. Each evaluated method serves a targeted diagnostic purpose: end-to-end training to test whether supervised skill learning transfers directly, Adaptive Sharpness-Aware Minimization (ASAM) to probe whether flatter loss landscapes improve generalization, and augmentation-based self-supervised and contrastive learning to assess whether domain-invariant pretraining decouples skill from visual context. Transfer is evaluated in both directions using a disjoint-participant held-out test set for JIGSAWS. Results reveal an asymmetry: backbones pretrained on JIGSAWS achieve CCC values of 0.77 to 0.80 on LASANA, closely matching the end-to-end baseline, showing cross-rubric transfer is feasible when the target domain provides consistent supervision. Transfer to JIGSAWS fails across all methods, likely due to annotation inconsistencies. Control experiments with a Kinetics-pretrained backbone suggest task-specific heads carry the majority of the skill prediction burden, while the backbone need only provide adequate spatiotemporal features. These findings offer a new perspective on vision-based skill assessment: the central question of whether skill representations transfer across scoring systems has not been previously investigated. Results indicate the visual component is dominant but not solely responsible for skill prediction; further work is needed to conclusively disentangle transferable skill features from those bound to a specific visual domain.

cs.CV

Current validation practice undermines surgical AI development

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation as an important contributing factor. Existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, yielding misleading or clinically irrelevant results. We introduce a comprehensive catalogue of validation pitfalls in AI-based surgical video analysis, derived from a multi-stage Delphi process with 92 international experts. Pitfalls span three categories: (1) data, (2) metric selection/configuration, and (3) aggregation and reporting. A systematic review of surgical AI papers reveals that these pitfalls are widespread. Experiments on surgical video datasets show that ignoring temporal and hierarchical data structures can understate uncertainty, obscure critical failure modes, and alter algorithm rankings. To address these shortcomings, we provide consensus-based best practices compiled. Together, this work provides an evidence-based framework for rigorous validation of surgical video analysis algorithms, guiding benchmarking, reporting, regulatory review, and clinical translation.

q-bio.OT

OSS: Open Suturing Skills Vision-Based Assessment Challenge 2024-2025

Achieving high levels of surgical skill through effective training is essential for optimal patient outcomes. Automated, data-driven skill assessment holds significant potential to improve surgical training. While machine learning-based methods are increasingly popular for assessing skills in minimally invasive surgery, their application to open surgery remains limited. We present the results of a dedicated MICCAI challenge designed to benchmark and advance vision-based skill assessment in open surgery. The challenge dataset comprises videos of an open suturing training task recorded with a static GoPro camera in a dry-lab setting, with instrument trajectories available in addition to the primary video modality. The OSS Challenge was hosted over two consecutive years, comprising two and three independent tasks, respectively: (1) classifying skill level into four classes, (2) predicting the full Objective Structured Assessment of Technical Skills across eight categories, and (3) tracking hands and surgical tools. Participants submitted diverse solutions including deep learning-based video models, tracking-driven methods, and hybrid approaches. General-purpose spatiotemporal video models consistently achieved the strongest performance, though conceptually diverse approaches reached competitive levels when well-executed. Predicting fine-grained OSATS scores remains challenging but benefits substantially from increased training data. Keypoint tracking proves difficult given frequent occlusions and out-of-frame instances, limiting current applicability for motion-based skill analysis. This work benchmarks innovative and diverse solutions for surgical skill assessment, highlighting both the promise and current limitations of video-based evaluation in open surgery and identifying critical directions for advancing automated skill assessment toward clinical impact.

cs.CV

Optimizing Surgical Plans for Parenchyma-Sparing Liver Resections through Contour-Guided Resection and Surface Approximation

Objective: This study introduces a novel method for defining virtual resections in liver cancer surgery, aimed at enhancing the adaptability of parenchyma-sparing resection (PSR) plans. By comparing these with traditional anatomical resection (AR) plans, we explore the potential for optimization in surgical planning. Methods: Leveraging contours and spline surface approximations directly from the liver's surface, our method aligns closely with actual surgical procedures, offering a more realistic representation of curved resection paths. This technique, tested against 14 cases from the OSLO-COMET study, incorporates surface deformation for versatile plan modeling, comparing volumetric outcomes of PSR and AR. Results: The study highlights significant benefits of PSR over AR, including reduced resected volume ($32.71 \pm 13.80$ ml for PSR vs. $249.53 \pm 135.23$ ml for AR, $p <0.0001$) and higher remnant liver volume ($1922.77 \pm 442.86$ ml for PSR vs. $1716.87 \pm 403.00$ ml for AR, $p <0.0001$). PSR also showed a considerably higher remnant percentage ($98.16 \pm 0.81%$) compared to AR ($87.40 \pm 6.49%$, $p <0.0001$). Conclusion: The proposed approach is able to define virtual resections accommodating a wide variety of resections (i.e., PSR and AR). Careful surgical planning using virtual resections can optimize the resection strategy. Significance: This study presents a novel computer-aided planning system for liver surgery, demonstrating its efficacy and flexibility for definition of virtual resections. Virtual surgery planning can be used for optimization of resection strategies leading to increased preservation of healthy tissue.

physics.med-ph