paper-with-me

홈 › Papers

OT-Talk: Animating 3D Talking Head with Optimal Transportation

2025-05-03 · Xinmu Wang, Xiang Gao, Xiyun Song, Heather Yu, Zongfang Lin, Liang Peng, Xianfeng GU

Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech signals and facial dynamics remains a challenge, often resulting in incorrect lip syncing and unnatural facial movements. To address this, we propose OT-Talk, the first approach to leverage optimal transportation to optimize the learning model in talking head animation. Building on existing learning frameworks, we utilize a pre-trained Hubert model to extract audio features and a transformer model to process temporal sequences. Unlike previous methods that focus solely on vertex coordinates or displacements, we introduce Chebyshev Graph Convolution to extract geometric features from triangulated meshes. To measure mesh dissimilarities, we go beyond traditional mesh reconstruction errors and velocity differences between adjacent frames. Instead, we represent meshes as probability measures and approximate their surfaces. This allows us to leverage the sliced Wasserstein distance for modeling mesh variations. This approach facilitates the learning of smooth and accurate facial motions, resulting in coherent and natural facial animations. Our experiments on two public audio-mesh datasets demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively in terms of mesh reconstruction accuracy and temporal alignment. In addition, we conducted a user perception study with 20 volunteers to further assess the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2505.01932

Code (0)

등록된 구현이 없습니다.

Tasks

Temporal Sequences

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

ScanTalk: 3D Talking Heads from Unregistered Scans

2024-03-16 · Federico Nocentini, Thomas Besnier, Claudio Ferrari, Sylvain Arguillere 외

Speech-driven 3D talking heads generation has emerged as a significant area of interest among researchers, presenting numerous challenges. Existing methods are constrained by animating faces with fixed topologies, wherei…

Compressing Video Calls using Synthetic Talking Heads

2022-10-07 · Madhav Agarwal, Anchit Gupta, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 외

We leverage the modern advancements in talking head generation to propose an end-to-end system for talking head video compression. Our algorithm transmits pivot frames intermittently while the rest of the talking head vi…

Face ReenactmentTalking Head GenerationVideo Compression

Embedded Representation Learning Network for Animating Styled Video Portrait

2024-04-29 · Tianyong Wang, Xiangyu Liang, Wangguandong Zheng, Dan Niu 외

The talking head generation recently attracted considerable attention due to its widespread application prospects, especially for digital avatars and 3D animation design. Inspired by this practical demand, several works …

NeRFRepresentation LearningTalking Head Generation

VectorTalker: SVG Talking Face Generation with Progressive Vectorisation

2023-12-18 · Hao Hu, Xuan Wang, Jingxiang Sun, Yanbo Fan 외

High-fidelity and efficient audio-driven talking head generation has been a key research topic in computer graphics and computer vision. In this work, we study vector image based audio-driven talking head generation. Com…

Face GenerationImage ReconstructionTalking Face GenerationTalking Head Generation

Animating Face using Disentangled Audio Representations

2019-10-02 · Gaurav Mittal, Baoyuan Wang

All previous methods for audio-driven talking head generation assume the input audio to be clean with a neutral tone. As we show empirically, one can easily break these systems by simply adding certain background noise t…

Representation LearningTalking Head Generation