paper-with-me

Papers

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

2024-12-11 · Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo, Zixuan Guo, Weijiang Yu, Heng Chang, Fei Ma, Fei Richard Yu

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of training video. However, due to the limited scale of the training data, these methods often exhibit poor performance in audio-lip synchronization and visual quality. In this paper, we propose a novel 3D Gaussian-based method called PointTalk, which constructs a static 3D Gaussian field of the head and deforms it in sync with the audio. It also incorporates an audio-driven dynamic lip point cloud as a critical component of the conditional information, thereby facilitating the effective synthesis of talking heads. Specifically, the initial step involves generating the corresponding lip point cloud from the audio signal and capturing its topological structure. The design of the dynamic difference encoder aims to capture the subtle nuances inherent in dynamic lip movements more effectively. Furthermore, we integrate the audio-point enhancement module, which not only ensures the synchronization of the audio signal with the corresponding lip point cloud within the feature space, but also facilitates a deeper understanding of the interrelations among cross-modal conditional features. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking head synthesis compared to previous methods.

📄 PDF Abstract BibTeX arXiv:2412.08504

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Points2Sound: From mono to binaural audio using 3D point cloud scenes

2021-04-26 · Francesc Lluís, Vasileios Chatziioannou, Alex Hofmann

For immersive applications, the generation of binaural sound that matches its visual counterpart is crucial to bring meaningful experiences to people in a virtual environment. Recent studies have shown the possibility of…

Audio Synthesis

DoppDrive: Doppler-Driven Temporal Aggregation for Improved Radar Object Detection

2025-08-17 · Yuval Haitman, Oded Bialer arxiv

Radar-based object detection is essential for autonomous driving due to radar's long detection range. However, the sparsity of radar point clouds, especially at long range, poses challenges for accurate detection. Existi…

Autonomous DrivingObject DetectionPoint Clouds

EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

2024-07-11 · ZhiYuan Chen, Jiajiong Cao, Zhiquan Chen, Yuming Li 외

The area of portrait image animation, propelled by audio input, has witnessed notable progress in the generation of lifelike and dynamic portraits. Conventional methods are limited to utilizing either audios or facial ke…

Image Animation

Gentlest ascent dynamics on manifolds defined by adaptively sampled point-clouds

2023-02-09 · Juan M. Bello-Rivas, Anastasia Georgiou, Hannes Vandecasteele, Ioannis G. Kevrekidis

Finding saddle points of dynamical systems is an important problem in practical applications such as the study of rare events of molecular systems. Gentlest ascent dynamics (GAD) is one of a number of algorithms in exist…

Point Cloud Audio Processing

2021-05-06 · Krishna Subramani, Paris Smaragdis

Most audio processing pipelines involve transformations that act on fixed-dimensional input representations of audio. For example, when using the Short Time Fourier Transform (STFT) the DFT size specifies a fixed dimensi…

BIG-bench Machine Learning