paper-with-me

Papers

Scaling up self-supervised learning for improved surgical foundation models

2025-01-16 · Tim J. M. Jaspers, Ronald L. P. D. de Jong, Yiping Li, Carolus H. J. Kusters, Franciscus H. A. Bakker, Romy C. van Jaarsveld, Gino M. Kuiper, Richard van Hillegersberg, Jelle P. Ruurda, Willem M. Brinkman, Josien P. W. Pluim, Peter H. N. de With, Marcel Breeuwer, Yasmina Al Khalil, Fons van der Sommen

Foundation models have revolutionized computer vision by achieving vastly superior performance across diverse tasks through large-scale pretraining on extensive datasets. However, their application in surgical computer vision has been limited. This study addresses this gap by introducing SurgeNetXL, a novel surgical foundation model that sets a new benchmark in surgical computer vision. Trained on the largest reported surgical dataset to date, comprising over 4.7 million video frames, SurgeNetXL achieves consistent top-tier performance across six datasets spanning four surgical procedures and three tasks, including semantic segmentation, phase recognition, and critical view of safety (CVS) classification. Compared with the best-performing surgical foundation models, SurgeNetXL shows mean improvements of 2.4, 9.0, and 12.6 percent for semantic segmentation, phase recognition, and CVS classification, respectively. Additionally, SurgeNetXL outperforms the best-performing ImageNet-based variants by 14.4, 4.0, and 1.6 percent in the respective tasks. In addition to advancing model performance, this study provides key insights into scaling pretraining datasets, extending training durations, and optimizing model architectures specifically for surgical computer vision. These findings pave the way for improved generalizability and robustness in data-scarce scenarios, offering a comprehensive framework for future research in this domain. All models and a subset of the SurgeNetXL dataset, including over 2 million video frames, are publicly available at: https://github.com/TimJaspers0801/SurgeNet.

📄 PDF Abstract BibTeX arXiv:2501.09436

Code (1)

timjaspers0801/surgenet 공식 구현 pytorch

Tasks

Self-Supervised LearningSemantic Segmentation

Similar Papers 제목 키워드 기반

Surgical Depth Anything: Depth Estimation for Surgical Scenes using Foundation Models

2024-10-09 · Ange Lou, Yamin Li, Yike Zhang, Jack Noble

Monocular depth estimation is crucial for tracking and reconstruction algorithms, particularly in the context of surgical videos. However, the inherent challenges in directly obtaining ground truth depth maps during surg…

Depth EstimationMonocular Depth Estimation

The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes

2024-02-27 · Charles Alba, Bing Xue, Joanna Abraham, Thomas Kannampallil 외

Clinical notes recorded during a patient's perioperative journey holds immense informational value. Advances in large language models (LLMs) offer opportunities for bridging this gap. Using 84,875 pre-operative notes and…

Domain AdaptationMulti-Task LearningWord Embeddings

Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks

2018-05-22 · Gaurav Yengera, Didier Mutter, Jacques Marescaux, Nicolas Padoy

Real-time algorithms for automatically recognizing surgical phases are needed to develop systems that can provide assistance to surgeons, enable better management of operating room (OR) resources and consequently improve…

ManagementSurgical phase recognition

Data-Efficient Surgical Phase Segmentation in Small-Incision Cataract Surgery: A Controlled Study of Vision Foundation Models

2026-04-12 · Lincoln Spencer, Song Wang, Chen Chen arxiv

Surgical phase segmentation is central to computer-assisted surgery, yet robust models remain difficult to develop when labeled surgical videos are scarce. We study data-efficient phase segmentation for manual small-inci…

A generalizable foundation model for intraoperative understanding across surgical procedures

2026-02-14 · Kanggil Park, Yongjun Jeon, Soyoung Lim, Seonmin Park 외 arxiv

In minimally invasive surgery, clinical decisions depend on real-time visual interpretation, yet intraoperative perception varies substantially across surgeons and procedures. This variability limits consistent assessmen…

Representation LearningScene Understanding