paper-with-me

Papers

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

2024-12-05 · CVPR 2025 1 · Yongming Zhu, Longhao Zhang, Zhengkun Rong, Tianshu Hu, Shuang Liang, Zhipeng Ge

Imagine having a conversation with a socially intelligent agent. It can attentively listen to your words and offer visual and linguistic feedback promptly. This seamless interaction allows for multiple rounds of conversation to flow smoothly and naturally. In pursuit of actualizing it, we propose INFP, a novel audio-driven head generation framework for dyadic interaction. Unlike previous head generation works that only focus on single-sided communication, or require manual role assignment and explicit role switching, our model drives the agent portrait dynamically alternates between speaking and listening state, guided by the input dyadic audio. Specifically, INFP comprises a Motion-Based Head Imitation stage and an Audio-Guided Motion Generation stage. The first stage learns to project facial communicative behaviors from real-life conversation videos into a low-dimensional motion latent space, and use the motion latent codes to animate a static image. The second stage learns the mapping from the input dyadic audio to motion latent codes through denoising, leading to the audio-driven head generation in interactive scenarios. To facilitate this line of research, we introduce DyConv, a large scale dataset of rich dyadic conversations collected from the Internet. Extensive experiments and visualizations demonstrate superior performance and effectiveness of our method. Project Page: https://grisoon.github.io/INFP/.

📄 PDF Abstract BibTeX arXiv:2412.04037

Code (0)

등록된 구현이 없습니다.

Tasks

DenoisingMotion Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation

2025-12-23 · Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Joon Son Chung 외 arxiv

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have …

Dialogue Generation

Let's Chorus: Partner-aware Hybrid Song-Driven 3D Head Animation

2025-01-01 · CVPR 2025 1 · Xiumei Xie, Zikai Huang, Wenhao Xu, Peng Xiao 외

Singing is a vital form of human emotional expression and social interaction, distinguished from speech by its richer emotional nuances and freer expressive style. Thus, investigating 3D facial animation driven by si…

EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control

2026-03-19 · Yuzhe Weng, Haotian Wang, Yuanhong Yu, Jun Du 외 arxiv

Audio-driven talking head generation aims to create vivid and realistic videos from a static portrait and speech. Existing AR-based methods rely on intermediate facial representations, which limit their expressiveness an…

Talking Head GenerationVideo Generation

GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian Splatting

2025-12-11 · Madhav Agarwal, Mingtian Zhang, Laura Sevilla-Lara, Steven McDonagh arxiv

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Di…

Image Generation

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels

2026-04-11 · Yuzhe Weng, Haotian Wang, Xinyi Yu, Xiaoyan Wu 외 arxiv

Audio-driven human video generation has achieved remarkable success in monologue scenarios, largely driven by advancements in powerful video generation foundation models. Moving beyond monologues, authentic human communi…

Physical IntuitionVideo Generation