Robust Cross-Modal Knowledge Distillation for Unconstrained Videos
Cross-modal distillation has been widely used to transfer knowledge across different modalities, enriching the representation of the target unimodal one. Recent studies highly relate the temporal synchronization between vision and sound to the semantic consistency for cross-modal distillation. However, such semantic consistency from the synchronization is hard to guarantee in unconstrained videos, due to the irrelevant modality noise and differentiated semantic correlation. To this end, we first propose a \textit{Modality Noise Filter} (MNF) module to erase the irrelevant noise in teacher modality with cross-modal context. After this purification, we then design a \textit{Contrastive Semantic Calibration} (CSC) module to adaptively distill useful knowledge for target modality, by referring to the differentiated sample-wise semantic correlation in a contrastive fashion. Extensive experiments show that our method could bring a performance boost compared with other distillation methods in both visual action recognition and video retrieval task. We also extend to the audio tagging task to prove the generalization of our method. The source code is available at \href{https://github.com/GeWu-Lab/cross-modal-distillation}{https://github.com/GeWu-Lab/cross-modal-distillation}.
Code (1)
Tasks
Action RecognitionAudio TaggingKnowledge DistillationRetrievalVideo RetrievalSimilar Papers 제목 키워드 기반
Cross-modal Contrastive Distillation for Instructional Activity Anticipation
In this study, we aim to predict the plausible future action steps given an observation of the past and study the task of instructional activity anticipation. Unlike previous anticipation tasks that aim at action label p…
Knowledge DistillationLearning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action Detection
In video understanding, most cross-modal knowledge distillation (KD) methods are tailored for classification tasks, focusing on the discriminative representation of the trimmed videos. However, action detection requires …
Action DetectionKnowledge DistillationVideo UnderstandingCross-modal knowledge distillation for action recognition
In this work, we address the problem how a network for action recognition that has been trained on a modality like RGB videos can be adapted to recognize actions for another modality like sequences of 3D human poses. To …
Action RecognitionKnowledge DistillationXKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning
We present XKD, a novel self-supervised framework to learn meaningful representations from unlabelled videos. XKD is trained with two pseudo objectives. First, masked data reconstruction is performed to learn modality-sp…
Action ClassificationClassificationKnowledge DistillationRepresentation Learning+4Seeing your sleep stage: cross-modal distillation from EEG to infrared video
It is inevitably crucial to classify sleep stage for the diagnosis of various diseases. However, existing automated diagnosis methods mostly adopt the "gold-standard" lectroencephalogram (EEG) or other uni-modal sensing …
EEGElectroencephalogram (EEG)