paper-with-me

홈 › Papers

Themes Informed Audio-visual Correspondence Learning

2020-09-14 · Runze Su, Fei Tao, Xudong Liu, Hao-Ran Wei, Xiaorong Mei, Zhiyao Duan, Lei Yuan, Ji Liu, Yuying Xie

The applications of short-term user-generated video (UGV), such as Snapchat, and Youtube short-term videos, booms recently, raising lots of multimodal machine learning tasks. Among them, learning the correspondence between audio and visual information from videos is a challenging one. Most previous work of the audio-visual correspondence(AVC) learning only investigated constrained videos or simple settings, which may not fit the application of UGV. In this paper, we proposed new principles for AVC and introduced a new framework to set sight of videos' themes to facilitate AVC learning. We also released the KWAI-AD-AudVis corpus which contained 85432 short advertisement videos (around 913 hours) made by users. We evaluated our proposed approach on this corpus, and it was able to outperform the baseline by 23.15% absolute difference.

📄 PDF Abstract BibTeX arXiv:2009.06573

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Deep Video Inpainting Guided by Audio-Visual Self-Supervision

2023-10-11 · Kyuyeon Kim, Junsik Jung, Woo Jae Kim, Sung-Eui Yoon

Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video…

audio-visual learningVideo Inpainting

Class-aware Sounding Objects Localization via Audiovisual Correspondence

2021-12-22 · Di Hu, Yake Wei, Rui Qian, Weiyao Lin 외

Audiovisual scenes are pervasive in our daily life. It is commonplace for humans to discriminatively localize different sounding objects but quite challenging for machines to achieve class-aware sounding objects localiza…

Objectobject-detectionObject DetectionObject Localization+1

AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection

2024-06-05 · CVPR 2024 1 · Trevine Oorloff, Surya Koppisetti, Nicolò Bonettini, Divyaraj Solanki 외

With the rapid growth in deepfake video content, we require improved and generalizable methods to detect them. Most existing detection methods either use uni-modal cues or rely on supervised training to capture the disso…

Contrastive LearningDeepFake DetectionFace SwappingRepresentation Learning

Learning Representations from Audio-Visual Spatial Alignment

2020-11-03 · NeurIPS 2020 12 · Pedro Morgado, Yi Li, Nuno Vasconcelos

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based …

Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic Segmentation

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…

audio-visual learning