Multimodal and self-supervised representation learning for automatic gesture recognition in surgical robotics
Self-supervised, multi-modal learning has been successful in holistic representation of complex scenarios. This can be useful to consolidate information from multiple modalities which have multiple, versatile uses. Its application in surgical robotics can lead to simultaneously developing a generalised machine understanding of the surgical process and reduce the dependency on quality, expert annotations which are generally difficult to obtain. We develop a self-supervised, multi-modal representation learning paradigm that learns representations for surgical gestures from video and kinematics. We use an encoder-decoder network configuration that encodes representations from surgical videos and decodes them to yield kinematics. We quantitatively demonstrate the efficacy of our learnt representations for gesture recognition (with accuracy between 69.6 % and 77.8 %), transfer learning across multiple tasks (with accuracy between 44.6 % and 64.8 %) and surgeon skill classification (with accuracy between 76.8 % and 81.2 %). Further, we qualitatively demonstrate that our self-supervised representations cluster in semantically meaningful properties (surgeon skill and gestures).
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderGesture RecognitionRepresentation LearningTransfer LearningSimilar Papers 제목 키워드 기반
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects rema…
Representation LearningSelf-supervised Learning Matters: A Simple Ensemble Solution for Micro-Gesture Recognition
In this paper, we present XInsight Lab's solution to the micro-gesture classification track of the 4th MiGA Challenge at IJCAI 2026, in which our solution ranked first and achieved a new state-of-the-art result. We propo…
Micro-gesture RecognitionSelf-Supervised LearningRepresentation LearningLearning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
In face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make …
Contrastive LearningDiagnosticRepresentation LearningDeep self-supervised learning with visualisation for automatic gesture recognition
Gesture is an important mean of non-verbal communication, with visual modality allows human to convey information during interaction, facilitating peoples and human-machine interactions. However, it is considered difficu…
Gesture RecognitionSelf-Supervised LearningSelf-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
Gestures are pivotal in enhancing co-speech communication. While recent works have mostly focused on point-level motion transformation or fully supervised motion representations through data-driven approaches, we explore…
Self-Supervised LearningSSIMVideo Generation