paper-with-me

홈 › Papers

Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild

2024-12-17 · Xingjian Wang, Li Chai

In-the-wild dynamic facial expression recognition (DFER) encounters a significant challenge in recognizing emotion-related expressions, which are often temporally and spatially diluted by emotion-irrelevant expressions and global context. Most prior DFER methods directly utilize coupled spatiotemporal representations that may incorporate weakly relevant features with emotion-irrelevant context bias. Several DFER methods highlight dynamic information for DFER, but following explicit guidance that may be vulnerable to irrelevant motion. In this paper, we propose a novel Implicit Facial Dynamics Disentanglement framework (IFDD). Through expanding wavelet lifting scheme to fully learnable framework, IFDD disentangles emotion-related dynamic information from emotion-irrelevant global context in an implicit manner, i.e., without exploit operations and external guidance. The disentanglement process contains two stages. The first is Inter-frame Static-dynamic Splitting Module (ISSM) for rough disentanglement estimation, which explores inter-frame correlation to generate content-aware splitting indexes on-the-fly. We utilize these indexes to split frame features into two groups, one with greater global similarity, and the other with more unique dynamic features. The second stage is Lifting-based Aggregation-Disentanglement Module (LADM) for further refinement. LADM first aggregates two groups of features from ISSM to obtain fine-grained global context features by an updater, and then disentangles emotion-related facial dynamic features from the global context by a predictor. Extensive experiments on in-the-wild datasets have demonstrated that IFDD outperforms prior supervised DFER methods with higher recognition accuracy and comparable efficiency. Code is available at https://github.com/CyberPegasus/IFDD.

📄 PDF Abstract BibTeX arXiv:2412.13168

Code (1)

CyberPegasus/IFDD 공식 구현 pytorch

Tasks

DisentanglementDynamic Facial Expression RecognitionFacial Expression Recognition

Similar Papers 제목 키워드 기반

Dynamic Causal Disentanglement Model for Dialogue Emotion Detection

2023-09-13 · Yuting Su, Yichen Wei, Weizhi Nie, Sicheng Zhao 외

Emotion detection is a critical technology extensively employed in diverse fields. While the incorporation of commonsense knowledge has proven beneficial for existing emotion detection methods, dialogue-based emotion det…

DisentanglementEmotion Recognitionmodel

An Event-comment Social Media Corpus for Implicit Emotion Analysis

2020-05-01 · LREC 2020 5 · Sophia Yat Mei Lee, Helena Yan Ping Lau

The classification of implicit emotions in text has always been a great challenge to emotion processing. Even though the majority of emotion expressed implicitly, most previous attempts at emotions have focused on the ex…

ClassificationEmotion ClassificationEmotion RecognitionGeneral Classification

CASEIN: Cascading Explicit and Implicit Control for Fine-grained Emotion Intensity Regulation

2023-06-27 · Yuhao Cui, Xiongwei Wang, Zhongzhou Zhao, Wei Zhou 외

Existing fine-grained intensity regulation methods rely on explicit control through predicted emotion probabilities. However, these high-level semantic probabilities are often inaccurate and unsmooth at the phoneme level…

Disentanglement

iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for Speech Synthesis based on Disentanglement between Prosody and Timbre

2022-06-29 · Guangyan Zhang, Ying Qin, Wenjie Zhang, Jialun Wu 외

The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when sp…

DisentanglementSpeaker IdentificationSpeech Synthesis

Why disentanglement-based speaker anonymization systems fail at preserving emotions?

2025-01-22 · Ünal Ege Gaznepoglu, Nils Peters

Disentanglement-based speaker anonymization involves decomposing speech into a semantically meaningful representation, altering the speaker embedding, and resynthesizing a waveform using a neural vocoder. State-of-the-ar…

DisentanglementEmotion RecognitionSpeaker anonymization