paper-with-me

Papers

Unified speech and gesture synthesis using flow matching

2023-10-08 · Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow, Éva Székely, Gustav Eje Henter

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures. This paper presents a novel, unified architecture for jointly synthesising speech acoustics and skeleton-based 3D gesture motion from text, trained using optimal-transport conditional flow matching (OT-CFM). The proposed architecture is simpler than the previous state of the art, has a smaller memory footprint, and can capture the joint distribution of speech and gestures, generating both modalities together in one single process. The new training regime, meanwhile, enables better synthesis quality in much fewer steps (network evaluations) than before. Uni- and multimodal subjective tests demonstrate improved speech naturalness, gesture human-likeness, and cross-modal appropriateness compared to existing benchmarks. Please see https://shivammehta25.github.io/Match-TTSG/ for video examples and code.

📄 PDF Abstract BibTeX arXiv:2310.05181

Code (0)

등록된 구현이 없습니다.

Tasks

Audio SynthesisMotion Synthesistext-to-speechText to SpeechText-To-Speech Synthesis

Similar Papers 제목 키워드 기반

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

2025-01-31 · Pinxin Liu, Luchuan Song, Junhua Huang, Haiyang Liu 외

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spa…

DenoisingGesture Generation

Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction

2025-10-13 · Téo Guichoux, Théodor Lemerle, Shivam Mehta, Jonas Beskow 외 arxiv

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We i…

Gesture Generation

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons

2023-09-13 · Sicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 외

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability acros…

DiversityGesture Generation

Integrated Speech and Gesture Synthesis

2021-08-25 · Siyang Wang, Simon Alexanderson, Joakim Gustafson, Jonas Beskow 외

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline.…

Speech Synthesistext-to-speechText to Speech

SemConFlow: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-Matching

2026-03-27 · Lanmiao Liu, Esam Ghaleb, Aslı Özyürek, Zerrin Yumak arxiv

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, whic…

Semantic RetrievalGesture Generation