paper-with-me

홈 › Papers

MOSPA: Human Motion Generation Driven by Spatial Audio

2025-07-16 · Shuyang Xu, Zhiyang Dou, Mingyi Shi, Liang Pan, Leo Ho, Jingbo Wang, Yuan Liu, Cheng Lin, Yuexin Ma, Wenping Wang, Taku Komura arxiv

Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive Spatial Audio-Driven Human Motion (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human MOtion generation driven by SPatial Audio, termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse, realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation

📄 PDF Abstract BibTeX arXiv:2507.11949

Code (0)

등록된 구현이 없습니다.

Tasks

Motion Synthesis

Similar Papers 제목 키워드 기반

EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space

2026-01-17 · Jing Zhang, Bingjie Fan, Jixiang Zhu, Zhe Wang arxiv

We propose EmoLat, a novel emotion latent space that enables fine-grained, text-driven image sentiment transfer by modeling cross-modal correlations between textual semantics and visual emotion features. Within EmoLat, a…

EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation

2026-02-12 · Bingyuan Wang, Xingbei Chen, Zongyang Qiu, Linping Yuan 외 arxiv

Emotion is important for creating compelling virtual reality (VR) content. Although some generative methods have been applied to lower the barrier to creating emotionally rich content, they fail to capture the nuanced em…

Image Outpainting

MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation

2026-02-11 · Yejin Kim, Wilbert Pumacay, Omar Rayyan, Max Argus 외 arxiv

Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that characterize real environments are vast and un…

Robot Navigation

Text2Performer: Text-Driven Human Video Generation

2023-04-17 · ICCV 2023 1 · Yuming Jiang, Shuai Yang, Tong Liang Koh, Wayne Wu 외

Text-driven content creation has evolved to be a transformative technique that revolutionizes creativity. Here we study the task of text-driven human video generation, where a video sequence is synthesized from texts des…

Video Generation

Fg-T2M: Fine-Grained Text-Driven Human Motion Generation via Diffusion Model

2023-09-12 · ICCV 2023 1 · Yin Wang, Zhiying Leng, Frederick W. B. Li, Shun-Cheng Wu 외

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively con…

Motion GenerationMotion Synthesis