paper-with-me

홈 › Papers

SE4Lip: Speech-Lip Encoder for Talking Head Synthesis to Solve Phoneme-Viseme Alignment Ambiguity

2025-04-08 · Yihuan Huang, Jiajun Liu, Yanzhen Ren, Wuyang Liu, Juhua Tang

Speech-driven talking head synthesis tasks commonly use general acoustic features (such as HuBERT and DeepSpeech) as guided speech features. However, we discovered that these features suffer from phoneme-viseme alignment ambiguity, which refers to the uncertainty and imprecision in matching phonemes (speech) with visemes (lip). To address this issue, we propose the Speech Encoder for Lip (SE4Lip) to encode lip features from speech directly, aligning speech and lip features in the joint embedding space by a cross-modal alignment framework. The STFT spectrogram with the GRU-based model is designed in SE4Lip to preserve the fine-grained speech features. Experimental results show that SE4Lip achieves state-of-the-art performance in both NeRF and 3DGS rendering models. Its lip sync accuracy improves by 13.7% and 14.2% compared to the best baseline and produces results close to the ground truth videos.

📄 PDF Abstract BibTeX arXiv:2504.05803

Code (0)

등록된 구현이 없습니다.

Tasks

3DGScross-modal alignmentNeRF

Similar Papers 제목 키워드 기반

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

2024-12-11 · Yifan Xie, Tao Feng, Xin Zhang, Xiangyang Luo 외

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize h…

AudioVisual Speech Synthesis: A brief literature review

2021-02-18 · Efthymios Georgiou, Athanasios Katsamanis

This brief literature review studies the problem of audiovisual speech synthesis, which is the problem of generating an animated talking head given a text as input. Due to the high complexity of this problem, we approach…

Speech Synthesistext-to-speechText to Speech

S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head Synthesis

2024-08-18 · Dongze Li, Kang Zhao, Wei Wang, Yifeng Ma 외

Talking head synthesis is a practical technique with wide applications. Current Neural Radiance Field (NeRF) based approaches have shown their superiority on driving one-shot talking heads with videos or signals regresse…

NeRF

READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head Generation

2025-08-05 · Haotian Wang, Yuzhe Weng, Jun Du, Haoran Xu 외 arxiv

The introduction of diffusion models has brought significant advances to the field of audio-driven talking head generation. However, the extremely slow inference speed severely limits the practical implementation of diff…

Talking Head Generation

Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis

2023-07-18 · ICCV 2023 1 · Jiahe Li, Jiawei Zhang, Xiao Bai, Jun Zhou 외

This paper presents ER-NeRF, a novel conditional Neural Radiance Fields (NeRF) based architecture for talking portrait synthesis that can concurrently achieve fast convergence, real-time rendering, and state-of-the-art p…

NeRF