Learning Multi-modal Representations by Watching Hundreds of Surgical Video Lectures
Recent advancements in surgical computer vision applications have been driven by vision-only models, which do not explicitly integrate the rich semantics of language into their design. These methods rely on manually annotated surgical videos to predict a fixed set of object categories, limiting their generalizability to unseen surgical procedures and downstream tasks. In this work, we put forward the idea that the surgical video lectures available through open surgical e-learning platforms can provide effective vision and language supervisory signals for multi-modal representation learning without relying on manual annotations. We address the surgery-specific linguistic challenges present in surgical video lectures by employing multiple complementary automatic speech recognition systems to generate text transcriptions. We then present a novel method, SurgVLP - Surgical Vision Language Pre-training, for multi-modal representation learning. Extensive experiments across diverse surgical procedures and tasks demonstrate that the multi-modal representations learned by SurgVLP exhibit strong transferability and adaptability in surgical video analysis. Furthermore, our zero-shot evaluations highlight SurgVLP's potential as a general-purpose foundation model for surgical workflow analysis, reducing the reliance on extensive manual annotations for downstream tasks, and facilitating adaptation methods such as few-shot learning to build a scalable and data-efficient solution for various downstream surgical applications. The training code and weights are public.
Code (2)
Tasks
Automatic Speech RecognitionContrastive LearningFew-Shot LearningRepresentation LearningRetrievalspeech-recognitionSpeech RecognitionTripletVideo CaptioningVideo RetrievalMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Watch and Learn: Leveraging Expert Knowledge and Language for Surgical Video Understanding
Automated surgical workflow analysis is crucial for education, research, and clinical decision-making, but the lack of annotated datasets hinders the development of accurate and comprehensive workflow analysis solutions.…
DenoisingDense Video Captioningparameter-efficient fine-tuningTemporal Localization+3Multimodal and self-supervised representation learning for automatic gesture recognition in surgical robotics
Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities which have multiple, versatile uses. Its a…
DecoderGesture RecognitionRepresentation LearningTransfer LearningSound Source Localization for Spatial Mapping of Surgical Actions in Dynamic Scenes
Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual…
Sound Source LocalizationScene UnderstandingPoint CloudsTowards Generalizable Surgical Activity Recognition Using Spatial Temporal Graph Convolutional Networks
Modeling and recognition of surgical activities poses an interesting research problem. Although a number of recent works studied automatic recognition of surgical activities, generalizability of these works across differ…
Activity RecognitionGesture RecognitionSurgical Gesture RecognitionFrom Phase Grounding to Intelligent Surgical Narratives
Video surgery timelines are an important part of tool-assisted surgeries, as they allow surgeons to quickly focus on key parts of the procedure. Current methods involve the surgeon filling out a post-operation (OP) repor…