Themes Informed Audio-visual Correspondence Learning
The applications of short-term user-generated video (UGV), such as Snapchat, and Youtube short-term videos, booms recently, raising lots of multimodal machine learning tasks. Among them, learning the correspondence between audio and visual information from videos is a challenging one. Most previous work of the audio-visual correspondence(AVC) learning only investigated constrained videos or simple settings, which may not fit the application of UGV. In this paper, we proposed new principles for AVC and introduced a new framework to set sight of videos' themes to facilitate AVC learning. We also released the KWAI-AD-AudVis corpus which contained 85432 short advertisement videos (around 913 hours) made by users. We evaluated our proposed approach on this corpus, and it was able to outperform the baseline by 23.15% absolute difference.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Deep Video Inpainting Guided by Audio-Visual Self-Supervision
Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video…
audio-visual learningVideo InpaintingClass-aware Sounding Objects Localization via Audiovisual Correspondence
Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localiza…
Objectobject-detectionObject DetectionObject Localization+1AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the disso…
Contrastive LearningDeepFake DetectionFace SwappingRepresentation LearningLearning Representations from Audio-Visual Spatial Alignment
We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based …
Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic SegmentationTelling Left from Right: Learning Spatial Correspondence of Sight and Sound
Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…
audio-visual learning