paper-with-me

Papers

Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videos

2021-12-01 · CVPR 2022 1 · Foivos Paraperas Papantoniou, Panagiotis P. Filntisis, Petros Maragos, Anastasios Roussos

In this paper, we introduce a novel deep learning method for photo-realistic manipulation of the emotional state of actors in "in-the-wild" videos. The proposed method is based on a parametric 3D face representation of the actor in the input scene that offers a reliable disentanglement of the facial identity from the head pose and facial expressions. It then uses a novel deep domain translation framework that alters the facial expressions in a consistent and plausible manner, taking into account their dynamics. Finally, the altered facial expressions are used to photo-realistically manipulate the facial region in the input scene based on an especially-designed neural face renderer. To the best of our knowledge, our method is the first to be capable of controlling the actor's facial expressions by even using as a sole input the semantic labels of the manipulated emotions, while at the same time preserving the speech-related lip movements. We conduct extensive qualitative and quantitative evaluations and comparisons, which demonstrate the effectiveness of our approach and the especially promising results that we obtain. Our method opens a plethora of new possibilities for useful applications of neural rendering technologies, ranging from movie post-production and video games to photo-realistic affective avatars.

📄 PDF Abstract BibTeX arXiv:2112.00585

Code (1)

foivospar/NED 공식 구현 pytorch

Tasks

DisentanglementNeural Rendering

Similar Papers 제목 키워드 기반

EmoInstruct-TTS: Dual-Path Instruction-Guided Emotional Speech Synthesis

2026-06-08 · Minghui Wu, Ganjun Liu, Zikun Fang, Ting Meng 외 arxiv

Instruction-based controllable speech synthesis enables users to specify emotions through natural language. However, existing approaches often rely on coarse emotion labels and lack explicit modeling of fine-grained inte…

Speech Synthesis

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment

2026-04-21 · Chaonan Ji, Jinwei Qi, Sheng Xu, Peng Zhang 외 arxiv

Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular control for expressiveness, while methods…

Towards Authentic Movie Dubbing with Retrieve-Augmented Director-Actor Interaction Learning

2025-11-18 · Rui Liu, Yuan Zhao, Zhenqi Jia arxiv

The automatic movie dubbing model generates vivid speech from given scripts, replicating a speaker's timbre from a brief timbre prompt while ensuring lip-sync with the silent video. Existing approaches simulate a simplif…

Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation

2025-12-22 · Guoli Jia, Junyao Hu, Xinwei Long, Kai Tian 외 arxiv

Image generation based on diffusion models has demonstrated impressive capability, motivating exploration into diverse and specialized applications. Owing to the importance of emotion in advertising, emotion-oriented ima…

Image Generation

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models

2025-10-15 · Yizhou Peng, Yukun Ma, Chong Zhang, Yi-Wen Chao 외 arxiv

While Text-to-Speech (TTS) systems enable emotional control via natural-language instructions, expressiveness, naturalness, and speech quality degrade when the target emotion conflicts with the textual semantics. We prop…