paper-with-me

Papers

Speech Driven Talking Face Generation from a Single Image and an Emotion Condition

2020-08-08 · Sefik Emre Eskimez, You Zhang, Zhiyao Duan

Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an end-to-end talking face generation system that takes a speech utterance, a single face image, and a categorical emotion label as input to render a talking face video synchronized with the speech and expressing the conditioned emotion. Objective evaluation on image quality, audiovisual synchronization, and visual emotion expression shows that the proposed system outperforms a state-of-the-art baseline system. Subjective evaluation of visual emotion expression and video realness also demonstrates the superiority of the proposed system. Furthermore, we conduct a human emotion recognition pilot study using generated videos with mismatched emotions among the audio and visual modalities. Results show that humans respond to the visual modality more significantly than the audio modality on this task.

📄 PDF Abstract BibTeX arXiv:2008.03592

Code (1)

eeskimez/emotalkingface 공식 구현 pytorch

Tasks

Emotion RecognitionFace GenerationTalking Face Generation

Similar Papers 제목 키워드 기반

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

2021-08-09 · Xinsheng Wang, Qicong Xie, Jihua Zhu, Lei Xie 외

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. In this paper, we present an automatic me…

Talking Head Generationtext-to-speechText to Speech

Face2VoiceSync: Lightweight Face-Voice Consistency for Text-Driven Talking Face Generation

2025-07-25 · Fang Kang, Yin Cao, Haoyu Chen arxiv

Recent studies in speech-driven talking face generation achieve promising results, but their reliance on fixed-driven speech limits further applications (e.g., face-voice mismatch). Thus, we extend the task to a more cha…

Talking Face GenerationFace Alignment

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

2024-05-12 · Changpeng Cai, Guinan Guo, Jiao Li, Junhao Su 외

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. Whil…

DisentanglementFace GenerationTalking Face GenerationTalking Head Generation

UniFLG: Unified Facial Landmark Generator from Text or Speech

2023-02-28 · Kentaro Mitsui, Yukiya Hono, Kei Sawada

Talking face generation has been extensively investigated owing to its wide applicability. The two primary frameworks used for talking face generation comprise a text-driven framework, which generates synchronized speech…

DecoderFace GenerationSpeech SynthesisTalking Face Generation+2

See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region Refinement

2025-10-28 · Jinting Wang, Jun Wang, Hei Victor Cheng, Li Liu arxiv

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key…

Talking Face Generation