paper-with-me

Papers

Audio-Visual Instance Discrimination with Cross-Modal Agreement

2020-04-27 · CVPR 2021 1 · Pedro Morgado, Nuno Vasconcelos, Ishan Misra

We present a self-supervised learning approach to learn audio-visual representations from video and audio. Our method uses contrastive learning for cross-modal discrimination of video from audio and vice-versa. We show that optimizing for cross-modal discrimination, rather than within-modal discrimination, is important to learn good representations from video and audio. With this simple but powerful insight, our method achieves highly competitive performance when finetuned on action recognition tasks. Furthermore, while recent work in contrastive learning defines positive and negative samples as individual instances, we generalize this definition by exploring cross-modal agreement. We group together multiple instances as positives by measuring their similarity in both the video and audio feature spaces. Cross-modal agreement creates better positive and negative sets, which allows us to calibrate visual similarities by seeking within-modal discrimination of positive instances, and achieve significant gains on downstream tasks.

📄 PDF Abstract BibTeX arXiv:2004.12943

Code (1)

facebookresearch/AVID-CMA 공식 구현 pytorch

Tasks

Action RecognitionAudio ClassificationContrastive LearningSelf-Supervised Action RecognitionSelf-Supervised Audio ClassificationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Robust Audio-Visual Instance Discrimination

2021-03-29 · CVPR 2021 1 · Pedro Morgado, Ishan Misra, Nuno Vasconcelos

We present a self-supervised learning method to learn audio and video representations. Prior work uses the natural correspondence between audio and video to define a standard cross-modal instance discrimination task, whe…

Action RecognitionContrastive LearningSelf-Supervised LearningTransfer Learning

Unsupervised Voice-Face Representation Learning by Cross-Modal Prototype Contrast

2022-04-28 · Boqing Zhu, Kele Xu, Changjian Wang, Zheng Qin 외

We present an approach to learn voice-face representations from the talking face videos, without any identity labels. Previous works employ cross-modal instance discrimination tasks to establish the correlation of voice …

Contrastive LearningRepresentation Learning

Multimodal Input Aids a Bayesian Model of Phonetic Learning

2024-07-22 · Sophia Zhi, Roger P. Levy, Stephan C. Meylan

One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal …

Dual Domain-Adversarial Learning for Audio-Visual Saliency Prediction

2022-08-10 · Yingzi Fan, Longfei Han, Yue Zhang, Lechao Cheng 외

Both visual and auditory information are valuable to determine the salient regions in videos. Deep convolution neural networks (CNN) showcase strong capacity in coping with the audio-visual saliency prediction task. Due …

Domain AdaptationPredictionSaliency PredictionUnsupervised Domain Adaptation

Modality-Aware Contrastive Instance Learning with Self-Distillation for Weakly-Supervised Audio-Visual Violence Detection

2022-07-12 · Jiashuo Yu, Jinyu Liu, Ying Cheng, Rui Feng 외

Weakly-supervised audio-visual violence detection aims to distinguish snippets containing multimodal violence events with video-level labels. Many prior works perform audio-visual integration and interaction in an early …

Anomaly Detection In Surveillance Videosaudio-visual learningMultiple Instance Learning