paper-with-me

Papers

On Negative Sampling for Audio-Visual Contrastive Learning from Movies

2022-04-29 · Mahdi M. Kalayeh, Shervin Ardeshir, Lingyi Liu, Nagendra Kamath, Ashok Chandrashekar

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal a plethora of information about what happens in a scene, make the audio-visual space an intuitive choice for representation learning. In this paper, we explore the efficacy of audio-visual self-supervised learning from uncurated long-form content i.e movies. Studying its differences with conventional short-form content, we identify a non-i.i.d distribution of data, driven by the nature of movies. Specifically, we find long-form content to naturally contain a diverse set of semantic concepts (semantic diversity), where a large portion of them, such as main characters and environments often reappear frequently throughout the movie (reoccurring semantic concepts). In addition, movies often contain content-exclusive artistic artifacts, such as color palettes or thematic music, which are strong signals for uniquely distinguishing a movie (non-semantic consistency). Capitalizing on these observations, we comprehensively study the effect of emphasizing within-movie negative sampling in a contrastive learning setup. Our view is different from those of prior works who consider within-video positive sampling, inspired by the notion of semantic persistency over time, and operate in a short-video regime. Our empirical findings suggest that, with certain modifications, training on uncurated long-form videos yields representations which transfer competitively with the state-of-the-art to a variety of action recognition and audio classification tasks.

📄 PDF Abstract BibTeX arXiv:2205.00073

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAudio ClassificationContrastive LearningFormRepresentation LearningSelf-Supervised Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Watching Too Much Television is Good: Self-Supervised Audio-Visual Representation Learning from Movies and TV Shows

2021-06-16 · NeurIPS 2021 12 · Mahdi M. Kalayeh, Nagendra Kamath, Lingyi Liu, Ashok Chandrashekar

The abundance and ease of utilizing sound, along with the fact that auditory clues reveal so much about what happens in the scene, make the audio-visual space a perfectly intuitive choice for self-supervised representati…

Contrastive LearningRepresentation LearningSelf-Supervised Learning

Enhancing Sound Source Localization via False Negative Elimination

2024-08-29 · Zengjie Song, Jiangshe Zhang, Yuxi Wang, Junsong Fan 외

Sound source localization aims to localize objects emitting the sound in visual scenes. Recent works obtaining impressive results typically rely on contrastive learning. However, the common practice of randomly sampling …

audio-visual learningContrastive Learningobject-detectionObject Detection+1

On Negative Sampling for Contrastive Audio-Text Retrieval

2022-11-08 · Huang Xie, Okko Räsänen, Tuomas Virtanen

This paper investigates negative sampling for contrastive learning in the context of audio-text retrieval. The strategy for negative sampling refers to selecting negatives (either audio clips or textual descriptions) fro…

Audio to Text RetrievalContrastive LearningRetrievalText Retrieval

Active Contrastive Learning of Audio-Visual Video Representations

2020-08-31 · ICLR 2021 1 · Shuang Ma, Zhaoyang Zeng, Daniel McDuff, Yale Song

Contrastive learning has been shown to produce generalizable representations of audio and visual data by maximizing the lower bound on the mutual information (MI) between different views of an instance. However, obtainin…

Contrastive LearningRepresentation LearningVideo Classification

Robust Audio-Visual Instance Discrimination

2021-03-29 · CVPR 2021 1 · Pedro Morgado, Ishan Misra, Nuno Vasconcelos

We present a self-supervised learning method to learn audio and video representations. Prior work uses the natural correspondence between audio and video to define a standard cross-modal instance discrimination task, whe…

Action RecognitionContrastive LearningSelf-Supervised LearningTransfer Learning