paper-with-me

Papers

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation

2026-06-30 · Baiqin Wang, Sen Chen, Jiankuo Zhao, Xiangyu Liu, Zhen Lei, Xiangyu Zhu arxiv

Conversational talking face generation has recently attracted increasing attention, aiming to synthesize interactive talking videos where characters speak, listen, and respond dynamically to each other. This task presents three core challenges: 1) Flexibility: enabling multi-round dialogues with an arbitrary number of participants; 2) Naturalness: maintaining coherent motion and appropriate non-verbal feedback throughout the interaction; and 3) Efficiency: achieving real-time generation and low computation overhead for long-term continuous online conversation. Despite recent advances, existing methods still fall short in balancing all three requirements. To bridge this gap, we introduce InterTalk, a novel and efficient framework designed for highly interactive conversational talking face generation. Built upon a motion-based architecture, InterTalk supports real-time conversation synthesis. Our method achieves strong flexibility by explicitly modeling multi-round conversational dynamics among each participant, eliminating constraints on their numbers. To enhance interactivity, we incorporate motion feedback from multiple participants and introduce an iterative generation strategy for more natural behaviors. Besides, we disentangle motion into several facial components, enabling targeted refinements for natural response such as precise lip sync and realistic eye blinking. Finally, we construct a new multi-person conversational dataset and enrich it with 3D face-based data augmentation. Extensive experiments demonstrate that InterTalk achieves superior interaction quality while maintaining real-time performance at 30 FPS.

📄 PDF Abstract BibTeX arXiv:2606.31088

Code (0)

등록된 구현이 없습니다.

Tasks

Talking Face GenerationData Augmentation

Similar Papers 제목 키워드 기반

Interactive Conversational Head Generation

2023-07-05 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

We introduce a new conversation head generation benchmark for synthesizing behaviors of a single interlocutor in a face-to-face conversation. The capability to automatically synthesize interlocutors which can participate…

SentenceTalking Head Generation

TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation

2025-12-23 · Ji-Hoon Kim, Junseok Ahn, Doyeop Kwak, Joon Son Chung 외 arxiv

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have …

Dialogue Generation

Multi-human Interactive Talking Dataset

2025-08-05 · Zeyu Zhu, Weijia Wu, Mike Zheng Shou arxiv

Existing studies on talking video generation have predominantly focused on single-person monologues or isolated facial animations, limiting their applicability to realistic multi-human interactions. To bridge this gap, w…

Video Generation

MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted Enhancement

2026-01-05 · Lei Zhu, Lijian Lin, Ye Zhu, Jiahao Wu 외 arxiv

Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listen…

Talking Head Generation

Automatic Generation of Chatbots for Conversational Web Browsing

2020-08-19 · Pietro Chittò, Marcos Baez, Florian Daniel, Boualem Benatallah

In this paper, we describe the foundations for generating a chatbot out of a website equipped with simple, bot-specific HTML annotations. The approach is part of what we call conversational web browsing, i.e., a dialog-b…

Chatbot