paper-with-me

Papers

Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information

2025-10-07 · Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci, Danilo Comminiello arxiv

The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to generate viewpoint-specific content from larger, immersive 360-degree environments. This limitation restricts the creation of audio-visual experiences that are aware of off-camera events. To the best of our knowledge, this is the first work to introduce a framework for controllable audio-visual generation, addressing this unexplored gap. Specifically, we propose a diffusion model by introducing a set of powerful conditioning signals derived from the full 360-degree space: a panoramic saliency map to identify regions of interest, a bounding-box-aware signed distance map to define the target viewpoint, and a descriptive caption of the entire scene. By integrating these controls, our model generates spatially-aware viewpoint videos and audios that are coherently influenced by the broader, unseen environmental context, introducing a strong controllability that is essential for realistic and immersive audio-visual generation. We show audiovisual examples proving the effectiveness of our framework.

📄 PDF Abstract BibTeX arXiv:2510.06060

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

2026-08-24 · Mining Tan, Yinuo Wang, Ziqi Zhou, Weize Quan 외 arxiv

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…

Novel View SynthesisSpatial Reasoning

ViSAGe: Video-to-Spatial Audio Generation

2025-06-13 · Jaeyeon Kim, Heeseung Yun, Gunhee Kim

Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem o…

Audio Generation

SonoWorld: From One Image to a 3D Audio-Visual Scene

2026-03-30 · Derong Jin, Xiyi Chen, Ming C. Lin, Ruohan Gao arxiv

Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual sc…

Scene Generation

Learning Representations from Audio-Visual Spatial Alignment

2020-11-03 · NeurIPS 2020 12 · Pedro Morgado, Yi Li, Nuno Vasconcelos

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based …

Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic Segmentation

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

2026-08-01 · Masaki Yoshida, Ren Togo, Takahiro Ogawa, Miki Haseyama arxiv

3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition…

Audio Generation