paper-with-me

Papers

Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization

2018-06-30 · NeurIPS 2018 12 · Bruno Korbar, Du Tran, Lorenzo Torresani

There is a natural correlation between the visual and auditive elements of a video. In this work we leverage this connection to learn general and effective models for both audio and video analysis from self-supervised temporal synchronization. We demonstrate that a calibrated curriculum learning scheme, a careful choice of negative examples, and the use of a contrastive loss are critical ingredients to obtain powerful multi-sensory representations from models optimized to discern temporal synchronization of audio-video pairs. Without further finetuning, the resulting audio features achieve performance superior or comparable to the state-of-the-art on established audio classification benchmarks (DCASE2014 and ESC-50). At the same time, our visual subnet provides a very effective initialization to improve the accuracy of video-based action recognition models: compared to learning from scratch, our self-supervised pretraining yields a remarkable gain of +19.9% in action recognition accuracy on UCF101 and a boost of +17.7% on HMDB51.

📄 PDF Abstract BibTeX arXiv:1807.00230

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionAudio ClassificationSelf-Supervised Action RecognitionSelf-Supervised Audio ClassificationTemporal Action Localization

Similar Papers 제목 키워드 기반

SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face Representation

2025-10-11 · Zeyu Ling, Xiaodong Gu, Jiangnan Tang, Changqing Zou arxiv

We introduce SyncLipMAE, a self-supervised pretraining framework for talking-face video that learns synchronization-aware and transferable facial dynamics from unlabeled audio-visual streams. Our approach couples masked …

Visual Speech RecognitionAction Recognition

Self-supervised learning for audio-visual speaker diarization

2020-02-13 · Yifan Ding, Yong Xu, Shi-Xiong Zhang, Yahuan Cong 외

Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propo…

Self-Supervised Learningspeaker-diarizationSpeaker DiarizationTriplet+1

Self-Supervised Video Forensics by Audio-Visual Anomaly Detection

2023-01-04 · CVPR 2023 1 · Chao Feng, Ziyang Chen, Andrew Owens

Manipulated videos often contain subtle inconsistencies between their visual and audio signals. We propose a video forensics method, based on anomaly detection, that can identify these inconsistencies, and that can be tr…

Anomaly DetectionDeepFake DetectionVideo Forensics

Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

2018-09-21 · Soo-Whan Chung, Joon Son Chung, Hong-Goo Kang

This paper proposes a new strategy for learning powerful cross-modal embeddings for audio-to-video synchronization. Here, we set up the problem as one of cross-modal retrieval, where the objective is to find the most rel…

Binary ClassificationCross-Modal RetrievalRetrievalspeech-recognition+3

Unified Video-Language Pre-training with Synchronized Audio

2024-05-12 · Shentong Mo, Haofan Wang, Huaxia Li, Xu Tang

Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing pre-training approaches either captured t…