paper-with-me

Papers

Talking Together: Synthesizing Co-Located 3D Conversations from Audio

2026-03-09 · Mengyi Shan, Shouchieh Chang, Ziqian Bai, Shichen Liu, Yinda Zhang, Luchuan Song, Rohit Pandey, Sean Fanello, Zeng Huang arxiv

We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often produce disembodied "talking heads" akin to a video conference call, our work is the first to explicitly model the dynamic 3D spatial relationship -- including relative position, orientation, and mutual gaze -- that is crucial for realistic in-person dialogues. Our system synthesizes the full performance of both individuals, including precise lip-sync, and uniquely allows their relative head poses to be controlled via textual descriptions. To achieve this, we propose a dual-stream architecture where each stream is responsible for one participant's output. We employ speaker's role embeddings and inter-speaker cross-attention mechanisms designed to disentangle the mixed audio and model the interaction. Furthermore, we introduce a novel eye gaze loss to promote natural, mutual eye contact. To power our data-hungry approach, we introduce a novel pipeline to curate a large-scale conversational dataset consisting of over 2 million dyadic pairs from in-the-wild videos. Our method generates fluid, controllable, and spatially aware dyadic animations suitable for immersive applications in VR and telepresence, significantly outperforming existing baselines in perceived realism and interaction coherence.

📄 PDF Abstract BibTeX arXiv:2603.08674

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose Generation

2024-06-12 · Jiadong Liang, Feng Lu

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they t…

Face GenerationSelf-Supervised LearningTalking Face Generation

AV-Flow: Transforming Text to Audio-Visual Human-like Interactions

2025-02-18 · Aggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhoefer 외

We introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vi…

Speech Synthesis

Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head Synthesis

2025-03-28 · Shuai Shen, Wanhua Li, Yunpeng Zhang, Weipeng Hu 외

Talking head synthesis has become a key research area in computer graphics and multimedia, yet most existing methods often struggle to balance generation quality with computational efficiency. In this paper, we present a…

Computational EfficiencyTalking Head Generation

Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

2022-11-22 · Jiaxiang Tang, Kaisiyuan Wang, Hang Zhou, Xiaokang Chen 외

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose a…

NeRFTalking Face Generation

Taming Transformer for Emotion-Controllable Talking Face Generation

2025-08-20 · Ziqi Zhang, Cheng Deng arxiv

Talking face generation is a novel and challenging generation task, aiming at synthesizing a vivid speaking-face video given a specific audio. To fulfill emotion-controllable talking face generation, current methods need…

Talking Face Generation