paper-with-me

홈 › Papers

Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures

2023-07-27 · Kun Yuan, Vinkle Srivastav, Tong Yu, Joel L. Lavanchy, Jacques Marescaux, Pietro Mascagni, Nassir Navab, Nicolas Padoy

Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The training code and weights are public.

📄 PDF Abstract BibTeX arXiv:2307.15220

Code (2)

camma-public/peskavlp 공식 구현 pytorch
camma-public/surgvlp 공식 구현 pytorch

Tasks

Automatic Speech RecognitionContrastive LearningFew-Shot LearningRepresentation LearningRetrievalspeech-recognitionSpeech RecognitionTripletVideo CaptioningVideo Retrieval

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding

2025-03-14 · David Gastager, Ghazal Ghazaei, Constantin Patsch

Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions.…

DenoisingDense Video Captioningparameter-efficient fine-tuningTemporal Localization+3

Multimodal and self-supervised representation learning for automatic gesture recognition in surgical robotics

2020-10-31 · Aniruddha Tamhane, Jie Ying Wu, Mathias Unberath

Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities which have multiple, versatile uses. Its a…

DecoderGesture RecognitionRepresentation LearningTransfer Learning

Sound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes

2025-10-28 · Jonas Hein, Lazaros Vlachopoulos, Maurits Geert Laurent Olthof, Bastian Sigrist 외 arxiv

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…

Sound Source LocalizationScene UnderstandingPoint Clouds

Towards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks

2020-01-11 · Duygu Sarikaya, Pierre Jannin

Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across differ…

Activity RecognitionGesture RecognitionSurgical Gesture Recognition

From Phase Grounding to Intelligent Surgical Narratives

2026-03-05 · Ethan Peterson, Huixin Zhan arxiv

Video surgery timelines are an important part of tool-assisted surgeries, as they allow surgeons to quickly focus on key parts of the procedure. Current methods involve the surgeon filling out a post-operation (OP) repor…