MPE4G: Multimodal Pretrained Encoder for Co-Speech Gesture Generation
When virtual agents interact with humans, gestures are crucial to delivering their intentions with speech. Previous multimodal co-speech gesture generation models required encoded features of all modalities to generate gestures. If some input modalities are removed or contain noise, the model may not generate the gestures properly. To acquire robust and generalized encodings, we propose a novel framework with a multimodal pre-trained encoder for co-speech gesture generation. In the proposed method, the multi-head-attention-based encoder is trained with self-supervised learning to contain the information on each modality. Moreover, we collect full-body gestures that consist of 3D joint rotations to improve visualization and apply gestures to the extensible body model. Through the series of experiments and human evaluation, the proposed method renders realistic co-speech gestures not only when all input modalities are given but also when the input modalities are missing or noisy.
Code (0)
등록된 구현이 없습니다.
Tasks
Gesture GenerationSelf-Supervised LearningSimilar Papers 제목 키워드 기반
Chain of Generation: Multi-Modal Gesture Synthesis via Cascaded Conditional Control
This study aims to improve the generation of 3D gestures by utilizing multimodal information from human speech. Previous studies have focused on incorporating additional modalities to enhance the quality of generated ges…
Gesture GenerationRhythmHOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made…
Gesture GenerationRhythmLeveraging Speech for Gesture Detection in Multimodal Communication
Gestures are inherent to human interaction and often complement speech in face-to-face communication, forming a multimodal communication system. An important task in gesture analysis is detecting a gesture's beginning an…
MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization
This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and …
Gesture GenerationQuantizationSpeech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents loo…
Gesture GenerationRhythm