paper-with-me

Papers

Geometry-Aware Multi-Task Learning for Binaural Audio Generation from Video

2021-11-21 · Rishabh Garg, Ruohan Gao, Kristen Grauman

Binaural audio provides human listeners with an immersive spatial sound experience, but most existing videos lack binaural audio recordings. We propose an audio spatialization method that draws on visual information in videos to convert their monaural (single-channel) audio to binaural audio. Whereas existing approaches leverage visual features extracted directly from video frames, our approach explicitly disentangles the geometric cues present in the visual stream to guide the learning process. In particular, we develop a multi-task framework that learns geometry-aware features for binaural audio generation by accounting for the underlying room impulse response, the visual stream's coherence with the sound source(s) positions, and the consistency in geometry of the sounding objects over time. Furthermore, we introduce a new large video dataset with realistic binaural audio simulated for real-world scanned environments. On two datasets, we demonstrate the efficacy of our method, which achieves state-of-the-art results.

📄 PDF Abstract BibTeX arXiv:2111.10882

Code (0)

등록된 구현이 없습니다.

Tasks

Audio GenerationMulti-Task LearningRoom Impulse Response (RIR)

Similar Papers 제목 키워드 기반

OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models

2025-09-30 · Subrata Biswas, Mohammad Nur Hossain Khan, Bashima Islam arxiv

Spatial reasoning is fundamental to auditory perception, yet current audio large language models (ALLMs) largely rely on unstructured binaural cues and single step inference. This limits both perceptual accuracy in direc…

Spatial Reasoning

AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic Synthesis

2024-06-13 · Swapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng 외

Novel view acoustic synthesis (NVAS) aims to render binaural audio at any target viewpoint, given a mono audio emitted by a sound source at a 3D scene. Existing methods have proposed NeRF-based implicit models to exploit…

Audio SynthesisNeRF

Mixture-of-Experts Framework for Field-of-View Enhanced Signal-Dependent Binauralization of Moving Talkers

2025-09-16 · Manan Mittal, Thomas Deppisch, Joseph Forrer, Chris Le Sueur 외 arxiv

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to e…

Cyclic Learning for Binaural Audio Generation and Localization

2024-01-01 · CVPR 2024 1 · Zhaojian Li, Bin Zhao, Yuan Yuan

Binaural audio is obtained by simulating the biological structure of human ears which plays an important role in artificial immersive spaces. A promising approach is to utilize mono audio and corresponding vision to …

Audio GenerationObjectObject Localization

Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation

2026-04-02 · Teng Liu, Yinfeng Yu arxiv

Audio-Visual Navigation (AVN) requires an embodied agent to navigate toward a sound source by utilizing both vision and binaural audio. A core challenge arises in complex acoustic environments, where binaural cues become…

Visual Navigation