paper-with-me

홈 › Papers

SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase Recognition

2024-09-30 · Shu Yang, Zhiyuan Cai, Luyang Luo, Ning Ma, Shuchang Xu, Hao Chen

Capitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably poses significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatial-temporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed as Parameter-Efficient Image-to-Surgical-Video Transfer Learning. Firstly, we develop a parameter-efficient transfer learning benchmark SurgPETL for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Spatial-Temporal Adaptation module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatial-temporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with STA.

📄 PDF Abstract BibTeX arXiv:2409.20083

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionSurgical phase recognitionTemporal Action LocalizationTemporal SequencesTransfer Learning

Methods 이 논문이 사용한 방법론

Adapter 설명 없음

Similar Papers 제목 키워드 기반

Deep learning-enabled prediction of surgical errors during cataract surgery: from simulation to real-world application

2025-03-28 · Maxime Faure, Pierre-Henri Conze, Béatrice Cochener, Anas-Alexis Benyoussef 외

Real-time prediction of technical errors from cataract surgical videos can be highly beneficial, particularly for telementoring, which involves remote guidance and mentoring through digital platforms. However, the rarity…

Domain AdaptationPredictionUnsupervised Domain Adaptation

Long-Term Temporally Consistent Unpaired Video Translation from Simulated Surgical 3D Data

2021-03-31 · ICCV 2021 10 · Dominik Rivoir, Micha Pfeiffer, Reuben Docea, Fiona Kolbinger 외

Research in unpaired video translation has mainly focused on short-term temporal consistency by conditioning on neighboring frames. However for transfer from simulated to photorealistic sequences, available information o…

Neural RenderingTranslation

HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition

2024-05-16 · Kun Yuan, Vinkle Srivastav, Nassir Navab, Nicolas Padoy

Natural language could play an important role in developing generalist surgical models by providing a broad source of supervision from raw texts. This flexible form of supervision can enable the model's transferability a…

Contrastive LearningSurgical phase recognition

Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model

2025-05-12 · Wei Li, Ming Hu, Guoan Wang, Lihao Liu 외

In ophthalmic surgery, developing an AI system capable of interpreting surgical videos and predicting subsequent operations requires numerous ophthalmic surgical videos with high-quality annotations, which are difficult …

Video Generation

SurgMAE: Masked Autoencoders for Long Surgical Video Analysis

2023-05-19 · Muhammad Abdullah Jamal, Omid Mohareri

There has been a growing interest in using deep learning models for processing long surgical videos, in order to automatically detect clinical/operational activities and extract metrics that can enable workflow efficienc…

Self-Supervised Learning