paper-with-me

홈 › Papers

EmojiDiff: Advanced Facial Expression Control with High Identity Preservation in Portrait Generation

2024-12-02 · Liangwei Jiang, Ruida Li, Zhifeng Zhang, Shuo Fang, Chenguang Ma

This paper aims to bring fine-grained expression control to identity-preserving portrait generation. Existing methods tend to synthesize portraits with either neutral or stereotypical expressions. Even when supplemented with control signals like facial landmarks, these models struggle to generate accurate and vivid expressions following user instructions. To solve this, we introduce EmojiDiff, an end-to-end solution to facilitate simultaneous dual control of fine expression and identity. Unlike the conventional methods using coarse control signals, our method directly accepts RGB expression images as input templates to provide extremely accurate and fine-grained expression control in the diffusion process. As its core, an innovative decoupled scheme is proposed to disentangle expression features in the expression template from other extraneous information, such as identity, skin, and style. On one hand, we introduce \textbf{I}D-irrelevant \textbf{D}ata \textbf{I}teration (IDI) to synthesize extremely high-quality cross-identity expression pairs for decoupled training, which is the crucial foundation to filter out identity information hidden in the expressions. On the other hand, we meticulously investigate network layer function and select expression-sensitive layers to inject reference expression features, effectively preventing style leakage from expression signals. To further improve identity fidelity, we propose a novel fine-tuning strategy named \textbf{I}D-enhanced \textbf{C}ontrast \textbf{A}lignment (ICA), which eliminates the negative impact of expression control on original identity preservation. Experimental results demonstrate that our method remarkably outperforms counterparts, achieves precise expression control with highly maintained identity, and generalizes well to various diffusion models.

📄 PDF Abstract BibTeX arXiv:2412.01254

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization

2024-10-18 · Bin Lin, Yanzhen Yu, Jianhao Ye, Ruitao Lv 외

Existing audio-driven facial animation methods face critical challenges, including expression leakage, ineffective subtle expression transfer, and imprecise audio-driven synchronization. We discovered that these issues s…

GPUPortrait Animation

High-Accuracy Facial Depth Models derived from 3D Synthetic Data

2020-03-26

In this paper, we explore how synthetically generated 3D face models can be used to construct a high accuracy ground truth for depth. This allows us to train the Convolutional Neural Networks (CNN) to solve facial depth …

3D ReconstructionDepth EstimationScene UnderstandingVocal Bursts Intensity Prediction

X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation

2025-05-16 · PeiZhen Li, Longbing Cao, Xiao-Ming Wu, Runze Yang 외

The ability to imitate realistic facial expressions is essential for humanoid robots engaged in affective human-robot communication. However, the lack of datasets containing diverse humanoid facial expressions with prope…

Diversity

Uncover Common Facial Expressions in Terracotta Warriors: A Deep Learning Approach

2021-05-11 · Wenhong Tian, Yuanlun Xie, Tingsong Ma, Hengxin Zhang

Can advanced deep learning technologies be applied to analyze some ancient humanistic arts? Can deep learning technologies be directly applied to special scenes such as facial expression analysis of Terracotta Warriors? …

Deep Learning

Expressive Speech-driven Facial Animation with controllable emotions

2023-01-05 · Yutong Chen, Junhong Zhao, Wei-Qiang Zhang

It is in high demand to generate facial animation with high realism, but it remains a challenging task. Existing approaches of speech-driven facial animation can produce satisfactory mouth movement and lip synchronizatio…