paper-with-me

Papers

Expressive Talking Head Generation With Granular Audio-Visual Control

2022-01-01 · CVPR 2022 1 · Borong Liang, Yan Pan, Zhizhi Guo, Hang Zhou, Zhibin Hong, Xiaoguang Han, Junyu Han, Jingtuo Liu, Errui Ding, Jingdong Wang

Generating expressive talking heads is essential for creating virtual humans. However, existing one- or few-shot methods focus on lip-sync and head motion, ignoring the emotional expressions that make talking faces realistic. In this paper, we propose the Granularly Controlled Audio-Visual Talking Heads (GC-AVT), which controls lip movements, head poses, and facial expressions of a talking head in a granular manner. Our insight is to decouple the audio-visual driving sources through prior-based pre-processing designs. Detailedly, we disassemble the driving image into three complementary parts including: 1) a cropped mouth that facilitates lip-sync; 2) a masked head that implicitly learns pose; and 3) the upper face which works corporately and complementarily with a time-shifted mouth to contribute the expression. Interestingly, the encoded features from the three sources are integrally balanced through reconstruction training. Extensive experiments show that our method generates expressive faces with not only synced mouth shapes, controllable poses, but precisely animated emotional expressions as well.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Head Generation

Similar Papers 제목 키워드 기반

EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

2024-11-23 · CVPR 2025 1 · Haotian Wang, Yuzhe Weng, Yueyan Li, Zilu Guo 외

Diffusion models have revolutionized the field of talking head generation, yet still face challenges in expressiveness, controllability, and stability in long-time generation. In this research, we propose an EmotiveTalk …

Talking Head Generation

AI killed the video star. Audio-driven diffusion model for expressive talking head generation

2025-11-27 · Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 외 arxiv

We propose Dimitra++, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we propose a conditional Motion Diffusion T…

Talking Head Generation

EMOdiffhead: Continuously Emotional Control in Talking Head Generation via Diffusion

2024-09-11 · Jian Zhang, Weijian Mai, Zhijun Zhang

The task of audio-driven portrait animation involves generating a talking head video using an identity image and an audio track of speech. While many existing approaches focus on lip synchronization and video quality, fe…

Portrait AnimationTalking Head GenerationVideo Generation

Dimitra: Audio-driven Diffusion model for Expressive Talking Head Generation

2025-02-24 · Baptiste Chopin, Tashvik Dhamija, Pranav Balaji, Yaohui Wang 외

We propose Dimitra, a novel framework for audio-driven talking head generation, streamlined to learn lip motion, facial expression, as well as head pose motion. Specifically, we train a conditional Motion Diffusion Trans…

Talking Head Generation

EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters

2025-03-25 · Xuli Shen, Hua Cai, Dingding Yu, Weilin Shen 외

Generating emotion-specific talking head videos from audio input is an important and complex challenge for human-machine interaction. However, emotion is highly abstract concept with ambiguous boundaries, and it necessit…

TAG