paper-with-me

Papers

DiffListener: Discrete Diffusion Model for Listener Generation

2025-02-05 · Siyeol Jung, Taehwan Kim

The listener head generation (LHG) task aims to generate natural nonverbal listener responses based on the speaker's multimodal cues. While prior work either rely on limited modalities (e.g. audio and facial information) or employ autoregressive approaches which have limitations such as accumulating prediction errors. To address these limitations, we propose DiffListener, a discrete diffusion based approach for non-autoregressive listener head generation. Our model takes the speaker's facial information, audio, and text as inputs, additionally incorporating facial differential information to represent the temporal dynamics of expressions and movements. With this explicit modeling of facial dynamics, DiffListener can generate coherent reaction sequences in a non-autoregressive manner. Through comprehensive experiments, DiffListener demonstrates state-of-the-art performance in both quantitative and qualitative evaluations. The user study shows that DiffListener generates natural context-aware listener reactions that are well synchronized with the speaker. The code and demo videos are available in https://siyeoljung.github.io/DiffListener

📄 PDF Abstract BibTeX arXiv:2502.06822

Code (0)

등록된 구현이 없습니다.

Tasks

model

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Emotional Listener Portrait: Neural Listener Head Generation with Emotion

2023-09-29 · ICCV 2023 1 · Luchuan Song, Guojun Yin, Zhenchao Jin, Xiaoyi Dong 외

Listener head generation centers on generating non-verbal behaviors (e.g., smile) of a listener in reference to the information delivered by a speaker. A significant challenge when generating such responses is the non-de…

Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication

2025-05-08 · Jinhe Huang, Yongkang Cheng, Yuming Hang, Gaoge Han 외

Full-body gestures play a pivotal role in natural interactions and are crucial for achieving effective communication. Nevertheless, most existing studies primarily focus on the gesture generation of speakers, overlooking…

DenoisingGesture GenerationGesture Synchronization

Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion

2025-04-29 · Zesheng Wang, Alexandre Bruckert, Patrick Le Callet, Guangtao Zhai

Generating realistic listener facial motions in dyadic conversations remains challenging due to the high-dimensional action space and temporal dependency requirements. Existing approaches usually consider extracting 3D M…

Action GenerationFADImage GenerationMotion Synthesis

MFR-Net: Multi-faceted Responsive Listening Head Generation via Denoising Diffusion Model

2023-08-31 · Jin Liu, Xi Wang, Xiaomeng Fu, Yesheng Chai 외

Face-to-face communication is a common scenario including roles of speakers and listeners. Most existing research methods focus on producing speaker videos, while the generation of listener heads remains largely overlook…

DenoisingDiversity

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

2024-03-01 · CVPR 2024 1 · Xi Liu, Ying Guo, Cheng Zhen, Tong Li 외

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion.The applications of listener agent generation in v…

Motion GenerationRhythm