paper-with-me

Papers

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons

2023-09-13 · Sicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li, Zhensong Zhang, Qiaochu Huang, Lei Hao, Songcen Xu, Xiaofei Wu, Changpeng Yang, Zonghong Dai

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability across different motion capture standards. In addition, it is a challenging task due to the weak correlation between speech and gestures. To address these problems, we present UnifiedGesture, a novel diffusion model-based speech-driven gesture synthesis approach, trained on multiple gesture datasets with different skeletons. Specifically, we first present a retargeting network to learn latent homeomorphic graphs for different motion capture standards, unifying the representations of various gestures while extending the dataset. We then capture the correlation between speech and gestures based on a diffusion model architecture using cross-local attention and self-attention to generate better speech-matched and realistic gestures. To further align speech and gesture and increase diversity, we incorporate reinforcement learning on the discrete gesture units with a learned reward function. Extensive experiments show that UnifiedGesture outperforms recent approaches on speech-driven gesture generation in terms of CCA, FGD, and human-likeness. All code, pre-trained models, databases, and demos are available to the public at https://github.com/YoungSeng/UnifiedGesture.

📄 PDF Abstract BibTeX arXiv:2309.07051

Code (1)

youngseng/unifiedgesture 공식 구현 pytorch

Tasks

DiversityGesture Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Audio-Driven Co-Speech Gesture Video Generation

2022-12-05 · Xian Liu, Qianyi Wu, Hang Zhou, Yuanqi Du 외

Co-speech gesture is crucial for human-machine interaction and digital entertainment. While previous works mostly map speech audio to human skeletons (e.g., 2D keypoints), directly generating speakers' gestures in the im…

Video Generation

Pretrained Diffusion Models for Unified Human Motion Synthesis

2022-12-06 · Jianxin Ma, Shuai Bai, Chang Zhou

Generative modeling of human motion has broad applications in computer animation, virtual reality, and robotics. Conventional approaches develop separate models for different motion synthesis tasks, and typically use a m…

Motion GenerationMotion SynthesisOpen-Ended Question Answering

Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction

2025-10-13 · Téo Guichoux, Théodor Lemerle, Shivam Mehta, Jonas Beskow 외 arxiv

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We i…

Gesture Generation

Unified speech and gesture synthesis using flow matching

2023-10-08 · Shivam Mehta, Ruibo Tu, Simon Alexanderson, Jonas Beskow 외

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associ…

Audio SynthesisMotion Synthesistext-to-speechText to Speech+1

SHREC 2021: Track on Skeleton-based Hand Gesture Recognition in the Wild

2021-06-21 · Ariel Caputo, Andrea Giachetti, Simone Soso, Deborah Pintani 외

Gesture recognition is a fundamental tool to enable novel interaction paradigms in a variety of application scenarios like Mixed Reality environments, touchless public kiosks, entertainment systems, and more. Recognition…

Action RecognitionGesture RecognitionHand Gesture RecognitionHand-Gesture Recognition+1