paper-with-me

Papers

Self-supervised Audio Spatialization with Correspondence Classifier

2019-05-14 · Yu-Ding Lu, Hsin-Ying Lee, Hung-Yu Tseng, Ming-Hsuan Yang

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-supervised audio spatialization network that can generate spatial audio given the corresponding video and monaural audio. To enhance spatialization performance, we use an auxiliary classifier to classify ground-truth videos and those with audio where the left and right channels are swapped. We collect a large-scale video dataset with spatial audio to validate the proposed method. Experimental results demonstrate the effectiveness of the proposed model on the audio spatialization task.

📄 PDF Abstract BibTeX arXiv:1905.05375

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Auxiliary Classifier Auxiliary Classifiers are type of architectural component that seek to improve the convergence of very deep networks. They are classifier heads we attach to layers before the…

Similar Papers 제목 키워드 기반

Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

2021-05-03 · Yan-Bo Lin, Yu-Chiang Frank Wang

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural au…

Audio GenerationSelf-Supervised Learning

Telling Left from Right: Learning Spatial Correspondence of Sight and Sound

2020-06-11 · CVPR 2020 6 · Karren Yang, Bryan Russell, Justin Salamon

Self-supervised audio-visual learning aims to capture useful representations of video by leveraging correspondences between visual and audio inputs. Existing approaches have focused primarily on matching semantic informa…

audio-visual learning

In-the-wild Audio Spatialization with Flexible Text-guided Localization

2025-06-01 · Tianrui Pan, Jie Liu, Zewen Huang, Jie Tang 외

To enhance immersive experiences, binaural audio offers spatial awareness of sounding objects in AR, VR, and embodied AI applications. While existing audio spatialization methods can generally map any available monaural …

Spatial Reasoning

Deep Multimodal Clustering for Unsupervised Audiovisual Learning

2018-07-09 · CVPR 2019 6 · Di Hu, Feiping Nie, Xuelong. Li

The seen birds twitter, the running cars accompany with noise, etc. These naturally audiovisual correspondences provide the possibilities to explore and understand the outside world. However, the mixed multiple objects a…

Clustering

Self-supervised Learning of Audio Representations from Audio-Visual Data using Spatial Alignment

2022-06-02 · Shanshan Wang, Archontis Politis, Annamaria Mesaros, Tuomas Virtanen

Learning from audio-visual data offers many possibilities to express correspondence between the audio and visual content, similar to the human perception that relates aural and visual information. In this work, we presen…

Acoustic Scene ClassificationAction Recognitionobject-detectionObject Detection+4