paper-with-me

Papers

Visually Informed Binaural Audio Generation without Binaural Audios

2021-04-13 · CVPR 2021 1 · Xudong Xu, Hang Zhou, Ziwei Liu, Bo Dai, Xiaogang Wang, Dahua Lin

Stereophonic audio, especially binaural audio, plays an essential role in immersive viewing environments. Recent research has explored generating visually guided stereophonic audios supervised by multi-channel audio collections. However, due to the requirement of professional recording devices, existing datasets are limited in scale and variety, which impedes the generalization of supervised methods in real-world scenarios. In this work, we propose PseudoBinaural, an effective pipeline that is free of binaural recordings. The key insight is to carefully build pseudo visual-stereo pairs with mono data for training. Specifically, we leverage spherical harmonic decomposition and head-related impulse response (HRIR) to identify the relationship between spatial locations and received binaural audios. Then in the visual modality, corresponding visual cues of the mono data are manually placed at sound source positions to form the pairs. Compared to fully-supervised paradigms, our binaural-recording-free pipeline shows great stability in cross-dataset evaluation and achieves comparable performance under subjective preference. Moreover, combined with binaural recordings, our method is able to further boost the performance of binaural audio generation under supervised settings.

📄 PDF Abstract BibTeX arXiv:2104.06162

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

Cross-modal Generative Model for Visual-Guided Binaural Stereo Generation

2023-11-13 · Zhaojian Li, Bin Zhao, Yuan Yuan

Binaural stereo audio is recorded by imitating the way the human ear receives sound, which provides people with an immersive listening experience. Existing approaches leverage autoencoders and directly exploit visual spa…

AttributeAudio Generation

TTMBA: Towards Text To Multiple Sources Binaural Audio Generation

2025-07-22 · Yuxuan He, Xiaoran Yang, Ningning Pan, Gongping Huang arxiv

Most existing text-to-audio (TTA) generation methods produce mono outputs, neglecting essential spatial information for immersive auditory experiences. To address this issue, we propose a cascaded method for text-to-mult…

Audio Generation

Cyclic Learning for Binaural Audio Generation and Localization

2024-01-01 · CVPR 2024 1 · Zhaojian Li, Bin Zhao, Yuan Yuan

Binaural audio is obtained by simulating the biological structure of human ears which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to …

Audio GenerationObjectObject Localization

Localize to Binauralize: Audio Spatialization From Visual Sound Source Localization

2021-01-01 · ICCV 2021 10 · Kranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. Rajagopalan

Videos with binaural audios provide an immersive viewing experience by enabling 3D sound sensation. Recent works attempt to generate binaural audio in a multimodal learning framework using large quantities of videos …

Audio GenerationSound Source Localization

Sep-Stereo: Visually Guided Stereophonic Audio Generation by Associating Source Separation

2020-07-20 · ECCV 2020 8 · Hang Zhou, Xudong Xu, Dahua Lin, Xiaogang Wang 외

Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with…

Audio Generation