paper-with-me

홈 › Papers

Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue

2026-08-28 · Chengqian Ma, Wei Tao, Haoyu Zhang, Yiwen Guo hf

An avatar that holds a conversation should decide what to say and to move while saying it, yet these abilities live in separate model families: spoken dialogue models produce speech without motion, and co-speech motion models produce motion only from audio handed to them. The standard remedy is a cascade that first generates the spoken response and then runs a motion model over the finished audio, which requires a second full inference pass and precludes any joint optimisation between the two. We present Motion-Omni, an end-to-end framework in which a spoken dialogue model natively outputs explicit facial expression together with hand, upper-body and lower-body motion, generated directly from the hidden states that produce the speech. Joint training is not optional here: with the speech pathway frozen, motion remains misaligned with the audio, and co-adapting the LLM, Speech Generator and Motion Generator under both objectives is what recovers alignment while retaining spoken-dialogue ability. Supervision comes from a scalable, model-agnostic pipeline that pseudo-labels consistent-voice speech responses with a replaceable motion teacher, yielding 422,856 quality-ranked pairs (1,402 hours). We further release SwDA-500 and, to our knowledge, the first public evaluation protocol for stochastic open-ended full-body spoken dialogue, matching audio across motion systems while unifying rendering, automatic metrics, human evaluation, and latency measurement. Instantiated with a Qwen2.5-7B-Instruct backbone, Motion-Omni-Q7 matches the same-audio teacher cascade to within 2% on reference-free motion metrics while responding 5.4 x faster (RTF=0.78, faster than real time), surpasses all non-teacher cascades on beat correlation and diversity, and reaches a 2.62% word error rate, the lowest among the omni-modal systems compared.

📄 PDF Abstract BibTeX arXiv:2609.04250

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

2025-10-22 · Guowei Xu, Yuxuan Bian, Ailing Zeng, Mingyi Shi 외 arxiv

This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequence-to-sequence manner. OmniMotion-X effic…

Empathy Omni: Enabling Empathetic Speech Response Generation through Large Language Models

2025-08-26 · Haoyu Wang, Guangyan Zhang, Jiale Chen, Jingyu Li 외 arxiv

With the development of speech large language models (speech LLMs), users can now interact directly with assistants via speech. However, most existing models only convert response content into speech without fully captur…

Response Generation

OmniEgoCap: Camera-Agnostic Sequence-Level Egocentric Motion Reconstruction

2025-12-22 · Kyungwon Cho, Hanbyul Joo arxiv

The proliferation of commercial egocentric devices offers a unique lens into human behavior, yet reconstructing full-body 3D motion remains difficult due to frequent self-occlusion and the 'out-of-sight' nature of the we…

OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning

2024-06-13 · Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao 외

We present OmniH2O (Omni Human-to-Humanoid), a learning-based system for whole-body humanoid teleoperation and autonomy. Using kinematic pose as a universal control interface, OmniH2O enables various ways for a human to …

Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation

2024-10-01 · ACMMM24 2024 10 · Bohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 외

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, su…

Gesture GenerationMotion Generation