paper-with-me

Papers

Clip-aware expressive feature learning for video-based facial expression recognition

2022-03-25 · Information Sciences 2022 3 · Yuanyuan Liu, Chuanxu Feng, Xiaohui Yuan, Lin Zhou, Wenbin Wang, Jie Qin, and Zhongwen Luo

Video-based facial expression recognition (FER) has received increased attention as a result of its widespread applications. However, a video often contains many redundant and irrelevant frames. How to reduce redundancy and complexity of the available information and extract the most relevant information to facial expression in video sequences is a challenging task. In this paper, we divide a video into several short clips for processing and propose a clip-aware emotion-rich feature learning network (CEFLNet) for robust video-based FER. Our proposed CEFLNet identifies the emotional intensity expressed in each short clip in a video and obtains clip-aware emotion-rich representations. Specifically, CEFLNet constructs a clip-based feature encoder (CFE) with two-cascaded self-attention and local–global relation learning, aiming to encode clip-based spatio-temporal features from the clips of a video. An emotional intensity activation network (EIAN) is devised to generate emotional activation maps for locating the salient emotion clips and obtaining clip-aware emotion-rich representations, which are used for expression classification. The effectiveness and robustness of the proposed CEFLNet are evaluated using four public facial expression video datasets, including BU-3DFE, MMI, AFEW, and DFEW. Extensive experiments demonstrate the improved performance of our proposed CEFLNet in comparison with the state-of-the-art methods.

📄 PDF Abstract BibTeX

Code (1)

CVLab-Liuyuanyuan/CEFLNet pytorch

Tasks

Dynamic Facial Expression RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)

Similar Papers 제목 키워드 기반

TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles

2023-04-01 · Yifeng Ma, Suzhen Wang, Yu Ding, Bowen Ma 외

Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which m…

2D Semantic Segmentation task 3 (25 classes)Talking Head Generation

CLIP-AUTT: Test-Time Personalization with Action Unit Prompting for Fine-Grained Video Emotion Recognition

2026-03-30 · Muhammad Osama Zeeshan, Masoumeh Sharafi, Benoit Savary, Alessandro Lameiras Koerich 외 arxiv

Personalization in emotion recognition (ER) is essential for accurate interpretation of subtle and subject-specific expressive patterns. Recent advances in vision-language models (VLMs), such as CLIP, demonstrate strong …

Facial Expression RecognitionVideo Emotion Recognition

Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Features

2026-03-20 · Zheng Gao, Debin Meng, Yunqi Miao, Zhensong Zhang 외 arxiv

Current diffusion-based makeup transfer methods commonly use the makeup information encoded by off-the-shelf foundation models (e.g., CLIP) as condition to preserve the makeup style of reference image in the generation. …

Contrastive LearningImage Editing

Context-Aware Academic Emotion Dataset and Benchmark

2025-07-01 · Luming Zhao, Jingwen Xuan, Jiamin Lou, Yonghui Yu 외 arxiv

Academic emotion analysis plays a crucial role in evaluating students' engagement and cognitive states during the learning process. This paper addresses the challenge of automatically recognizing academic emotions throug…

Facial Expression RecognitionEmotion Recognition

VOODOO XP: Expressive One-Shot Head Reenactment for VR Telepresence

2024-05-25 · Phong Tran, Egor Zakharov, Long-Nhat Ho, Liwen Hu 외

We introduce VOODOO XP: a 3D-aware one-shot head reenactment method that can generate highly expressive facial expressions from any input driver video and a single 2D portrait. Our solution is real-time, view-consistent,…

Disentanglement