paper-with-me

Papers

Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction

2025-10-13 · Téo Guichoux, Théodor Lemerle, Shivam Mehta, Jonas Beskow, Gustav Eje Henter, Laure Soulier, Catherine Pelachaud, Nicolas Obin arxiv

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.

📄 PDF Abstract BibTeX arXiv:2510.12834

Code (0)

등록된 구현이 없습니다.

Tasks

Gesture Generation

Similar Papers 제목 키워드 기반

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons

2023-09-13 · Sicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 외

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability acros…

DiversityGesture Generation

Unified speech and gesture synthesis using flow matching

2023-10-08 · Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow 외

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associ…

Audio SynthesisMotion Synthesistext-to-speechText to Speech+1

Integrated Speech and Gesture Synthesis

2021-08-25 · Siyang Wang, Simon Alexanderson, Joakim Gustafson, Jonas Beskow 외

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline.…

Speech Synthesistext-to-speechText to Speech

DiffMotion: Speech-Driven Gesture Synthesis Using Denoising Diffusion Model

2023-01-24 · Fan Zhang, Naye Ji, Fuxing Gao, Yongping Li

Speech-driven gesture synthesis is a field of growing interest in virtual human creation. However, a critical challenge is the inherent intricate one-to-many mapping between speech and gestures. Previous studies have exp…

Denoising

SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain

2025-03-26 · Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao 외

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGe…

Gesture Generation