paper-with-me

홈 › Papers

It Takes Two: Real-time Co-Speech Two-person's Interaction Generation via Reactive Auto-regressive Diffusion Model

2024-12-03 · Mingyi Shi, Dafei Qin, Leo Ho, Zhouyingcheng Liao, Yinghao Huang, Junichi Yamagishi, Taku Komura

Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses. Additionally, most existing methods rely on offline sequence-to-sequence frameworks, which are unsuitable for online applications. In this work, we introduce an audio-driven, auto-regressive system designed to synthesize dynamic movements for two characters during a conversation. At the core of our approach is a diffusion-based full-body motion synthesis model, which is conditioned on the past states of both characters, speech audio, and a task-oriented motion trajectory input, allowing for flexible spatial control. To enhance the model's ability to learn diverse interactions, we have enriched existing two-person conversational motion datasets with more dynamic and interactive motions. We evaluate our system through multiple experiments to show it outperforms across a variety of tasks, including single and two-person co-speech motion generation, as well as interactive motion generation. To the best of our knowledge, this is the first system capable of generating interactive full-body motions for two characters from speech in an online manner.

📄 PDF Abstract BibTeX arXiv:2412.02419

Code (0)

등록된 구현이 없습니다.

Tasks

Motion GenerationMotion Synthesis

Similar Papers 제목 키워드 기반

PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models

2026-01-14 · Rajarshi Roy, Jonathan Raiman, Sang-gil Lee, Teodor-Dumitru Ene 외 arxiv

Recent advances in duplex speech models have enabled natural, low-latency speech-to-speech interactions. However, existing models are restricted to a fixed role and voice, limiting their ability to support structured, ro…

FlashLabs Chroma 1.0: A Real-Time End-to-End Spoken Dialogue Model with Personalized Voice Cloning

2026-01-16 · Tanyu Chen, Tairan Chen, Kai Shen, Zhenghua Bao 외 arxiv

Recent end-to-end spoken dialogue systems leverage speech tokenizers and neural audio codecs to enable LLMs to operate directly on discrete speech representations. However, these models often exhibit limited speaker iden…

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

2024-09-10 · Qingkai Fang, Shoutao Guo, Yan Zhou, Zhengrui Ma 외

Models like GPT-4o enable real-time interaction with large language models (LLMs) through speech, significantly enhancing user experience compared to traditional text-based interaction. However, there is still a lack of …

You said that?

2017-05-08 · Joon Son Chung, Amir Jamaludin, Andrew Zisserman

We present a method for generating a video of a talking face. The method takes as inputs: (i) still images of the target face, and (ii) an audio speech segment; and outputs a video of the target face lip synched with the…

DecoderUnconstrained Lip-synchronization

VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

2026-08-26 · Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao 외 arxiv

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel info…