paper-with-me

Papers

Generating coherent spontaneous speech and gesture from text

2021-01-14 · Simon Alexanderson, Éva Székely, Gustav Eje Henter, Taras Kucherenko, Jonas Beskow

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions of both of these types of data: On the speech side, text-to-speech systems are now able to generate highly convincing, spontaneous-sounding speech using unscripted speech audio as the source material. On the motion side, probabilistic motion-generation methods can now synthesise vivid and lifelike speech-driven 3D gesticulation. In this paper, we put these two state-of-the-art technologies together in a coherent fashion for the first time. Concretely, we demonstrate a proof-of-concept system trained on a single-speaker audio and motion-capture dataset, that is able to generate both speech and full-body gestures together from text input. In contrast to previous approaches for joint speech-and-gesture generation, we generate full-body gestures from speech synthesis trained on recordings of spontaneous speech from the same person as the motion-capture data. We illustrate our results by visualising gesture spaces and text-speech-gesture alignments, and through a demonstration video at https://simonalexanderson.github.io/IVA2020 .

📄 PDF Abstract BibTeX arXiv:2101.05684

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture GenerationMotion GenerationSpeech Synthesistext-to-speechText to Speech

Similar Papers 제목 키워드 기반

Unified speech and gesture synthesis using flow matching

2023-10-08 · Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow 외

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associ…

Audio SynthesisMotion Synthesistext-to-speechText to Speech+1

CoordSpeaker: Exploiting Gesture Captioning for Coordinated Caption-Empowered Co-Speech Gesture Generation

2025-11-28 · Fengyi Fang, Sicheng Yang, Wenming Yang arxiv

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Exist…

Gesture Generation

SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain

2025-03-26 · Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao 외

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGe…

Gesture Generation

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion

2025-05-03 · Xingqun Qi, Yatian Wang, Hengyuan Zhang, Jiahao Pan 외

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by individual self-talking, they overlook the practica…

Gesture Generation

SemGes: Semantics-aware Co-Speech Gesture Generation using Semantic Coherence and Relevance Learning

2025-07-25 · Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak arxiv

Creating a virtual avatar with semantically coherent gestures that are aligned with speech is a challenging task. Existing gesture generation research mainly focused on generating rhythmic beat gestures, neglecting the s…

Gesture Generation