Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to generate viewpoint-specific content from larger, immersive 360-degree environments. This limitation restricts the creation of audio-visual experiences that are aware of off-camera events. To the best of our knowledge, this is the first work to introduce a framework for controllable audio-visual generation, addressing this unexplored gap. Specifically, we propose a diffusion model by introducing a set of powerful conditioning signals derived from the full 360-degree space: a panoramic saliency map to identify regions of interest, a bounding-box-aware signed distance map to define the target viewpoint, and a descriptive caption of the entire scene. By integrating these controls, our model generates spatially-aware viewpoint videos and audios that are coherently influenced by the broader, unseen environmental context, introducing a strong controllability that is essential for realistic and immersive audio-visual generation. We show audiovisual examples proving the effectiveness of our framework.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation
Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects i…
Novel View SynthesisSpatial ReasoningViSAGe: Video-to-Spatial Audio Generation
Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem o…
Audio GenerationSonoWorld: From One Image to a 3D Audio-Visual Scene
Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual sc…
Scene GenerationLearning Representations from Audio-Visual Spatial Alignment
We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based …
Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic SegmentationScene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition…
Audio Generation