paper-with-me

Papers

SonoWorld: From One Image to a 3D Audio-Visual Scene

2026-03-30 · Derong Jin, Xiyi Chen, Ming C. Lin, Ruohan Gao arxiv

Tremendous progress in visual scene generation now turns a single image into an explorable 3D world, yet immersion remains incomplete without sound. We introduce Image2AVScene, the task of generating a 3D audio-visual scene from a single image, and present SonoWorld, the first framework to tackle this challenge. From one image, our pipeline outpaints a 360° panorama, lifts it into a navigable 3D scene, places language-guided sound anchors, and renders ambisonics for point, areal, and ambient sources, yielding spatial audio aligned with scene geometry and semantics. Quantitative evaluations on a newly curated real-world dataset and a controlled user study confirm the effectiveness of our approach. Beyond free-viewpoint audio-visual rendering, we also demonstrate applications to one-shot acoustic learning and audio-visual spatial source separation. Project website: https://humathe.github.io/sonoworld/

📄 PDF Abstract BibTeX arXiv:2603.28757

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Similar Papers 제목 키워드 기반

Attentional Graph Convolutional Network for Structure-aware Audio-Visual Scene Classification

2022-12-31 · Liguang Zhou, Yuhongze Zhou, Xiaonan Qi, Junjie Hu 외

Audio-Visual scene understanding is a challenging problem due to the unstructured spatial-temporal relations that exist in the audio signals and spatial layouts of different objects and various texture patterns in the vi…

Scene ClassificationScene RecognitionScene Understanding

AudioScenic: Audio-Driven Video Scene Editing

2024-04-25 · Kaixin Shen, Ruijie Quan, Linchao Zhu, Jun Xiao 외

Audio-driven visual scene editing endeavors to manipulate the visual background while leaving the foreground content unchanged, according to the given audio signals. Unlike current efforts focusing primarily on image edi…

Multi-Modal Gaze Following in Conversational Scenarios

2023-11-09 · Yuqi Hou, Zhongqun Zhang, Nora Horanyi, Jaewon Moon 외

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audi…

Audio-Infused Automatic Image Colorization by Exploiting Audio Scene Semantics

2024-01-24 · Pengcheng Zhao, Yanxiang Chen, Yang Zhao, Zhao Zhang

Automatic image colorization is inherently an ill-posed problem with uncertainty, which requires an accurate semantic understanding of scenes to estimate reasonable colors for grayscale images. Although recent interactio…

ColorizationImage Colorization

Learning Visual Styles from Audio-Visual Associations

2022-05-10 · Tingle Li, Yichen Liu, Andrew Owens, Hang Zhao

From the patter of rain to the crunch of snow, the sounds we hear often convey the visual textures that appear within a scene. In this paper, we present a method for learning visual styles from unlabeled audio-visual dat…

Image Stylization