paper-with-me

홈 › Papers

Generating Moving 3D Soundscapes with Latent Diffusion Models

2025-07-09 · Christian Templin, Yanda Zhu, Hao Wang arxiv

Spatial audio has become central to immersive applications such as VR/AR, cinema, and music. Existing generative audio models are largely limited to mono or stereo formats and cannot capture the full 3D localization cues available in first-order Ambisonics (FOA). Recent FOA models extend text-to-audio generation but remain restricted to static sources. In this work, we introduce SonicMotion, the first end-to-end latent diffusion framework capable of generating FOA audio with explicit control over moving sound sources. SonicMotion is implemented in two variations: 1) a descriptive model conditioned on natural language prompts, and 2) a parametric model conditioned on both text and spatial trajectory parameters for higher precision. To support training and evaluation, we construct a new dataset of over one million simulated FOA caption pairs that include both static and dynamic sources with annotated azimuth, elevation, and motion attributes. Experiments show that SonicMotion achieves state-of-the-art semantic alignment and perceptual quality comparable to leading text-to-audio systems, while uniquely attaining low spatial localization error.

📄 PDF Abstract BibTeX arXiv:2507.07318

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation

2024-10-14 · Peiwen Sun, Sitong Cheng, Xiangtai Li, Zhen Ye 외

Recently, diffusion models have achieved great success in mono-channel audio generation. However, when it comes to stereo audio generation, the soundscapes often have a complex scene of multiple objects and directions. C…

Audio Generationmultimodal generation

Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator

2025-04-25 · Minjae Kang, Martim Brandão

Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed…

Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization

2025-11-24 · Ellie L. Zhang, Duoduo Liao, Callie C. Liao arxiv

Generation of dynamic, scalable multi-species bird soundscapes remains a significant challenge in computer music and algorithmic sound design. Birdsongs involve rapid frequency-modulated chirps, complex amplitude envelop…

ImmerseDiffusion: A Generative Spatial Audio Latent Diffusion Model

2024-10-19 · Mojtaba Heydari, Mehrez Souden, Bruno Conejo, Joshua Atkins

We introduce ImmerseDiffusion, an end-to-end generative audio model that produces 3D immersive soundscapes conditioned on the spatial, temporal, and environmental conditions of sound objects. ImmerseDiffusion is trained …

Descriptivemodel

Few-Shot Concept Unlearning with Low Rank Adaptation

2025-05-18 · Udaya Shreyas, L. N. Aadarsh

Image Generation models are a trending topic nowadays, with many people utilizing Artificial Intelligence models in order to generate images. There are many such models which, given a prompt of a text, will generate an i…

DenoisingImage GenerationMachine Unlearning