paper-with-me

Papers

Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

2021-05-03 · Yan-Bo Lin, Yu-Chiang Frank Wang

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which would degrade the user experience due to the lack of ambient information. To address this issue, we propose an audio spatialization framework to convert a monaural video into a binaural one exploiting the relationship across audio and visual components. By preserving the left-right consistency in both audio and visual modalities, our learning strategy can be viewed as a self-supervised learning technique, and alleviates the dependency on a large amount of video data with ground truth binaural audio data during training. Experiments on benchmark datasets confirm the effectiveness of our proposed framework in both semi-supervised and fully supervised scenarios, with ablation studies and visualization further support the use of our model for audio spatialization.

📄 PDF Abstract BibTeX arXiv:2105.00708

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Deep Video Inpainting Guided by Audio-Visual Self-Supervision

2023-10-11 · Kyuyeon Kim, Junsik Jung, Woo Jae Kim, Sung-Eui Yoon

Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video…

audio-visual learningVideo Inpainting

CASP-Net: Rethinking Video Saliency Prediction from an Audio-VisualConsistency Perceptual Perspective

2023-03-11 · Junwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang 외

Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods a…

DecoderSaliency PredictionVideo Saliency Prediction

CASP-Net: Rethinking Video Saliency Prediction From an Audio-Visual Consistency Perceptual Perspective

2023-01-01 · CVPR 2023 1 · Junwen Xiong, Ganglai Wang, Peng Zhang, Wei Huang 외

Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP metho…

DecoderSaliency PredictionVideo Saliency Prediction

Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing

2025-09-17 · Yaru Chen, Ruohao Guo, Liting Gao, Yang Xiang 외 arxiv

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or …

CLASP: Cross-modal Salient Anchor-based Semantic Propagation for Weakly-supervised Dense Audio-Visual Event Localization

2025-08-06 · Jinxing Zhou, Ziheng Zhou, Yanghao Zhou, Yuxin Mao 외 arxiv

The Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize events in untrimmed videos that occur simultaneously in both the audio and visual modalities. This paper explores DAVEL under a new and m…

audio-visual event localization