paper-with-me

Papers

Two-level Attention with Two-stage Multi-task Learning for Facial Emotion Recognition

2018-11-29 · Xiaohua Wang, Muzi Peng, Lijuan Pan, Min Hu, Chunhua Jin, Fuji Ren

Compared with facial emotion recognition on categorical model, the dimensional emotion recognition can describe numerous emotions of the real world more accurately. Most prior works of dimensional emotion estimation only considered laboratory data and used video, speech or other multi-modal features. The effect of these methods applied on static images in the real world is unknown. In this paper, a two-level attention with two-stage multi-task learning (2Att-2Mt) framework is proposed for facial emotion estimation on only static images. Firstly, the features of corresponding region(position-level features) are extracted and enhanced automatically by first-level attention mechanism. In the following, we utilize Bi-directional Recurrent Neural Network(Bi-RNN) with self-attention(second-level attention) to make full use of the relationship features of different layers(layer-level features) adaptively. Owing to the inherent complexity of dimensional emotion recognition, we propose a two-stage multi-task learning structure to exploited categorical representations to ameliorate the dimensional representations and estimate valence and arousal simultaneously in view of the correlation of the two targets. The quantitative results conducted on AffectNet dataset show significant advancement on Concordance Correlation Coefficient(CCC) and Root Mean Square Error(RMSE), illustrating the superiority of the proposed framework. Besides, extensive comparative experiments have also fully demonstrated the effectiveness of different components.

📄 PDF Abstract BibTeX arXiv:1811.12139

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionFacial Emotion RecognitionMulti-Task LearningVocal Bursts Valence Prediction

Similar Papers 제목 키워드 기반

A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party Conversations

2023-07-01 · Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 2023 7 · Wenjie Zheng, Jianfei Yu, Rui Xia, Shijin Wang

Multimodal Emotion Recognition in Multiparty Conversations (MERMC) has recently attracted considerable attention. Due to the complexity of visual scenes in multi-party conversations, most previous MERMC studies mainly fo…

Emotion RecognitionEmotion Recognition in ConversationFacial Expression Recognition (FER)Multimodal Emotion Recognition+1

SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition

2026-07-10 · Hyein Park, Namho Kim, Junhwa Kim arxiv

Ambivalence and hesitancy are subtle behavioral states that are expressed through a combination of verbal content, facial behavior, visual context, and acoustic cues. Effective recognition therefore requires not only ext…

MATT: Multimodal Attention Level Estimation for e-learning Platforms

2023-01-22 · Roberto Daza, Luis F. Gomez, Aythami Morales, Julian Fierrez 외

This work presents a new multimodal system for remote attention level estimation based on multimodal face analysis. Our multimodal approach uses different parameters and signals obtained from the behavior and physiologic…

Facial Landmark DetectionHead Pose EstimationPose Estimation

Two-stage Temporal Modelling Framework for Video-based Depression Recognition using Graph Representation

2021-11-30 · Jiaqi Xu, Siyang Song, Keerthy Kusumam, Hatice Gunes 외

Video-based automatic depression analysis provides a fast, objective and repeatable self-assessment solution, which has been widely developed in recent years. While depression clues may be reflected by human facial behav…

DynamicID: Zero-Shot Multi-ID Image Personalization with Flexible Facial Editability

2025-03-09 · Xirui Hu, Jiahao Wang, Hao Chen, Weizhan Zhang 외

Recent advancements in text-to-image generation have spurred interest in personalized human image generation, which aims to create novel images featuring specific human identities as reference images indicate. Although e…

Contrastive LearningFacial EditingImage GenerationText to Image Generation+1