paper-with-me

홈 › Papers

Co-Speech Gesture Synthesis by Reinforcement Learning With Contrastive Pre-Trained Rewards

2023-01-01 · CVPR 2023 1 · Mingyang Sun, Mengchen Zhao, Yaqing Hou, Minglei Li, Huang Xu, Songcen Xu, Jianye Hao

There is a growing demand of automatically synthesizing co-speech gestures for virtual characters. However, it remains a challenge due to the complex relationship between input speeches and target gestures. Most existing works focus on predicting the next gesture that fits the data best, however, such methods are myopic and lack the ability to plan for future gestures. In this paper, we propose a novel reinforcement learning (RL) framework called RACER to generate sequences of gestures that maximize the overall satisfactory. RACER employs a vector quantized variational autoencoder to learn compact representations of gestures and a GPT-based policy architecture to generate coherent sequence of gestures autoregressively. In particular, we propose a contrastive pre-training approach to calculate the rewards, which integrates contextual information into action evaluation and successfully captures the complex relationships between multi-modal speech-gesture data. Experimental results show that our method significantly outperforms existing baselines in terms of both objective metrics and subjective human judgements. Demos can be found at https://github.com/RLracer/RACER.git.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons

2023-09-13 · Sicheng Yang, Zilin Wang, Zhiyong Wu, Minglei Li 외

The automatic co-speech gesture generation draws much attention in computer animation. Previous works designed network structures on individual datasets, which resulted in a lack of data volume and generalizability acros…

DiversityGesture Generation

Diffusion-Based Co-Speech Gesture Generation Using Joint Text and Audio Representation

2023-09-11 · Anna Deichler, Shivam Mehta, Simon Alexanderson, Jonas Beskow

This paper describes a system developed for the GENEA (Generation and Evaluation of Non-verbal Behaviour for Embodied Agents) Challenge 2023. Our solution builds on an existing diffusion-based motion synthesis model. We …

Gesture GenerationMotion Synthesis

SARGes: Semantically Aligned Reliable Gesture Generation via Intent Chain

2025-03-26 · Nan Gao, Yihua Bao, Dongdong Weng, Jiayi Zhao 외

Co-speech gesture generation enhances human-computer interaction realism through speech-synchronized gesture synthesis. However, generating semantically meaningful gestures remains a challenging problem. We propose SARGe…

Gesture Generation

Diff-TTSG: Denoising probabilistic integrated speech and gesture synthesis

2023-06-15 · Shivam Mehta, Siyang Wang, Simon Alexanderson, Jonas Beskow 외

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-ve…

DenoisingSpeech Synthesis

BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer

2023-09-07 · Kunkun Pang, Dafei Qin, Yingruo Fan, Julian Habekost 외

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is diff…