paper-with-me

Papers

AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis

2023-02-04 · NeurIPS 2023 11

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a first-of-its-kind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset.

📄 PDF Abstract BibTeX arXiv:2302.02088

Code (1)

aluo-x/learning_neural_acoustic_fields pytorch

Tasks

3D geometryAudio GenerationNeRF

Similar Papers 제목 키워드 기반

Aria-NeRF: Multimodal Egocentric View Synthesis

2023-11-11 · Jiankai Sun, Jianing Qiu, Chuanyang Zheng, John Tucker 외

We seek to accelerate research in developing rich, multimodal scene models trained from egocentric data, based on differentiable volumetric ray-tracing inspired by Neural Radiance Fields (NeRFs). The construction of a Ne…

NeRF

DecentNeRFs: Decentralized Neural Radiance Fields from Crowdsourced Images

2024-03-19 · Zaid Tasneem, Akshat Dave, Abhishek Singh, Kushagra Tiwary 외

Neural radiance fields (NeRFs) show potential for transforming images captured worldwide into immersive 3D visual experiences. However, most of this captured visual data remains siloed in our camera rolls as these images…

${M^2D}$NeRF: Multi-Modal Decomposition NeRF with 3D Feature Fields

2024-05-08 · Ning Wang, Lefei Zhang, Angel X Chang

Neural fields (NeRF) have emerged as a promising approach for representing continuous 3D scenes. Nevertheless, the lack of semantic encoding in NeRFs poses a significant challenge for scene decomposition. To address this…

NeRF

NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis

2024-01-23 · Chongke Bi, Xiaoxing Liu, Zhilei Liu

Talking face synthesis driven by audio is one of the current research hotspots in the fields of multidimensional signal processing and multimedia. Neural Radiance Field (NeRF) has recently been brought to this research f…

DisentanglementFace GenerationNeRF

DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

2023-09-14 · Yaoyu Su, Shaohui Wang, Haoqian Wang

In this paper, we present the decomposed triplane-hash neural radiance fields (DT-NeRF), a framework that significantly improves the photorealistic rendering of talking faces and achieves state-of-the-art results on key …

NeRF