CMD: Self-supervised 3D Action Representation Learning with Cross-modal Mutual Distillation
In 3D action recognition, there exists rich complementary information between skeleton modalities. Nevertheless, how to model and utilize this information remains a challenging problem for self-supervised 3D action representation learning. In this work, we formulate the cross-modal interaction as a bidirectional knowledge distillation problem. Different from classic distillation solutions that transfer the knowledge of a fixed and pre-trained teacher to the student, in this work, the knowledge is continuously updated and bidirectionally distilled between modalities. To this end, we propose a new Cross-modal Mutual Distillation (CMD) framework with the following designs. On the one hand, the neighboring similarity distribution is introduced to model the knowledge learned in each modality, where the relational information is naturally suitable for the contrastive frameworks. On the other hand, asymmetrical configurations are used for teacher and student to stabilize the distillation process and to transfer high-confidence information between modalities. By derivation, we find that the cross-modal positive mining in previous works can be regarded as a degenerated version of our CMD. We perform extensive experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD II datasets. Our approach outperforms existing self-supervised methods and sets a series of new records. The code is available at: https://github.com/maoyunyao/CMD
Code (1)
Tasks
3D Action RecognitionAction RecognitionFew-Shot Skeleton-Based Action RecognitionKnowledge DistillationRepresentation LearningSelf-supervised Skeleton-based Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity
We present CrissCross, a self-supervised framework for learning audio-visual representations. A novel notion is introduced in our framework whereby in addition to learning the intra-modal and standard 'synchronous' cross…
Audio ClassificationRetrievalSelf-Supervised Action RecognitionSelf-Supervised Audio Classification+4Self-Supervised Learning by Cross-Modal Audio-Video Clustering
Visual and audio modalities are highly correlated, yet they contain different information. Their strong correlation makes it possible to predict the semantics of one from the other with good accuracy. Their intrinsic dif…
Action RecognitionAudio ClassificationClusteringDeep Clustering+4MaCLR: Motion-aware Contrastive Learning of Representations for Videos
We present MaCLR, a novel method to explicitly perform cross-modal self-supervised video representations learning from visual and motion modalities. Compared to previous video representation learning methods that mostly …
Action DetectionAction RecognitionContrastive LearningRepresentation LearningCross and Learn: Cross-Modal Self-Supervision
In this paper we present a self-supervised method for representation learning utilizing two different modalities. Based on the observation that cross-modal information has a high semantic meaning we propose a method to e…
Action RecognitionOptical Flow EstimationRepresentation LearningSelf-Supervised Learning+1Learning Speech Representations from Raw Audio by Joint Audiovisual Self-Supervision
The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and a…
Acoustic Scene ClassificationAction RecognitionScene ClassificationSelf-Supervised Learning+1