paper-with-me

Papers

Localize to Binauralize: Audio Spatialization From Visual Sound Source Localization

2021-01-01 · ICCV 2021 10 · Kranthi Kumar Rachavarapu, Aakanksha, Vignesh Sundaresha, A. N. Rajagopalan

Videos with binaural audios provide an immersive viewing experience by enabling 3D sound sensation. Recent works attempt to generate binaural audio in a multimodal learning framework using large quantities of videos with accompanying binaural audio. In contrast, we attempt a more challenging problem -- synthesizing binaural audios for a video with monaural audio in a weakly supervised setting and weakly semi-supervised setting. Our key idea is that any down-stream task that can be solved only using binaural audios can be used to provide proxy supervision for binaural audio generation, thereby reducing the reliance on explicit supervision. In this work, as a proxy-task for weak supervision, we use Sound Source Localization with only audio. We design a two-stage architecture called Localize-to-Binauralize Network (L2BNet). The first stage of L2BNet is a Stereo Generation (SG) network employed to generate two-stream audio from monaural audio using visual frame information as guidance. In the second stage, an Audio Localization (AL) network is designed to use the synthesized two-stream audio to localize sound sources in visual frames. The entire network is trained end-to-end so that the AL network provides necessary supervision for the SG network. We experimentally show that our weakly-supervised framework generates two-stream audio containing binaural cues. Through user study, we further validate that our proposed approach generates binaural-quality audio using as little as 10% of explicit binaural supervision data for the SG network.

📄 PDF Abstract BibTeX

Code (1)

kranthikumarr/localize-to-binauralize 공식 구현

Tasks

Audio GenerationSound Source Localization

Similar Papers 제목 키워드 기반

Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation

2021-05-03 · Yan-Bo Lin, Yu-Chiang Frank Wang

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural au…

Audio GenerationSelf-Supervised Learning

Audio-Visual Camera Pose Estimation with Passive Scene Sounds and In-the-Wild Video

2025-12-13 · Daniel Adebi, Sagnik Majumder, Kristen Grauman arxiv

Understanding camera motion is a fundamental problem in embodied perception and 3D scene understanding. While visual methods have advanced rapidly, they often struggle under visually degraded conditions such as motion bl…

Camera Pose EstimationScene Understanding

Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video

2021-11-21 · Rishabh Garg, Ruohan Gao, Kristen Grauman

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in v…

Audio GenerationMulti-Task LearningRoom Impulse Response (RIR)

Self-supervised Audio Spatialization with Correspondence Classifier

2019-05-14 · Yu-Ding Lu, Hsin-Ying Lee, Hung-Yu Tseng, Ming-Hsuan Yang

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-…

Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization

2025-11-24 · Ellie L. Zhang, Duoduo Liao, Callie C. Liao arxiv

Generation of dynamic, scalable multi-species bird soundscapes remains a significant challenge in computer music and algorithmic sound design. Birdsongs involve rapid frequency-modulated chirps, complex amplitude envelop…