paper-with-me

Papers

MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion Model

2023-08-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai, Cai Yu, Jiao Dai, Jizhong Han

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlooked. Responsive listening head generation is an important task that aims to model face-to-face communication scenarios by generating a listener head video given a speaker video and a listener head image. An ideal generated responsive listening video should respond to the speaker with attitude or viewpoint expressing while maintaining diversity in interaction patterns and accuracy in listener identity information. To achieve this goal, we propose the \textbf{M}ulti-\textbf{F}aceted \textbf{R}esponsive Listening Head Generation Network (MFR-Net). Specifically, MFR-Net employs the probabilistic denoising diffusion model to predict diverse head pose and expression features. In order to perform multi-faceted response to the speaker video, while maintaining accurate listener identity preservation, we design the Feature Aggregation Module to boost listener identity features and fuse them with other speaker-related features. Finally, a renderer finetuned with identity consistency loss produces the final listening head videos. Our extensive experiments demonstrate that MFR-Net not only achieves multi-faceted responses in diversity and speaker identity information but also in attitude and viewpoint expression.

📄 PDF Abstract BibTeX arXiv:2308.16635

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingDiversity

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Responsive Listening Head Generation: A Benchmark Dataset and Baseline

2021-12-27 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

We present a new listening head generation benchmark, for synthesizing responsive feedbacks of a listener (e.g., nod, smile) during a face-to-face conversation. As the indispensable complement to talking heads generation…

Talking Head GenerationTranslation

Interactive Conversational Head Generation

2023-07-05 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate…

SentenceTalking Head Generation

Diffusion-based Realistic Listening Head Generation via Hybrid Motion Modeling

2025-01-01 · CVPR 2025 1 · Yinuo Wang, Yanbo Fan, Xuan Wang, Guo Yu 외

Listening head generation aims to synthesize non-verbal responsive listening head videos that naturally react to a certain speaker, for which, both realistic head movements, expressive facial expressions, and high vi…

Motion GenerationVideo Generation

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

2024-03-01 · CVPR 2024 1 · Xi Liu, Ying Guo, Cheng Zhen, Tong Li 외

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in v…

Motion GenerationRhythm

Hierarchical Semantic Perceptual Listener Head Video Generation: A High-performance Pipeline

2023-07-19 · Zhigang Chang, Weitai Hu, Qing Yang, Shibao Zheng

In dyadic speaker-listener interactions, the listener's head reactions along with the speaker's head movements, constitute an important non-verbal semantic expression together. The listener Head generation task aims to s…

DecoderTalking Head GenerationVideo Generation