Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters
Recent advances in co-speech gesture and talking head generation have been impressive, yet most methods focus on only one of the two tasks. Those that attempt to generate both often rely on separate models or network modules, increasing training complexity and ignoring the inherent relationship between face and body movements. To address the challenges, in this paper, we propose a novel model architecture that jointly generates face and body motions within a single network. This approach leverages shared weights between modalities, facilitated by adapters that enable adaptation to a common latent space. Our experiments demonstrate that the proposed framework not only maintains state-of-the-art co-speech gesture and talking head generation performance but also significantly reduces the number of parameters required.
Code (1)
Tasks
Face GenerationTalking Face GenerationTalking Head GenerationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Speech-driven 3D Conversational Gestures from Video
We propose the first approach to automatically and jointly synthesize both the synchronous 3D conversational body and hand gestures, as well as 3D face and head animations, of a virtual character from speech input. Our a…
3D Face AnimationGenerative Adversarial NetworkGesture GenerationHand Pose Estimation+1AVI-Talking: Learning Audio-Visual Instructions for Expressive 3D Talking Face Generation
While considerable progress has been made in achieving accurate lip synchronization for 3D speech-driven talking face generation, the task of incorporating expressive facial detail synthesis aligned with the speaker's sp…
Face GenerationHallucinationTalking Face GenerationLearning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face Animation
Expressions are fundamental to conveying human emotions. With the rapid advancement of AI-generated content (AIGC), realistic and expressive 3D facial animation has become increasingly crucial. Despite recent progress in…
Versatile Multimodal Controls for Expressive Talking Human Animation
In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not onl…
Human AnimationEMO2: End-Effector Guided Audio-Driven Avatar Video Generation
In this paper, we propose a novel audio-driven talking head method capable of simultaneously generating highly expressive facial expressions and hand gestures. Unlike existing methods that focus on generating full-body o…
Gesture GenerationVideo Generation