Facial expression recognition with grid-wise attention and visual transformer
Facial Expression Recognition (FER) has achieved remarkable progress as a result of using Convolutional Neural Networks (CNN). Relying on the spatial locality, convolutional filters in CNN, however, fail to learn long-range inductive biases between different facial regions in most neural layers. As such, the performance of a CNN-based model for FER is still limited. To address this problem, this paper introduces a novel FER framework with two attention mechanisms for CNN-based models, and these two attention mechanisms are used for the low-level feature learning the high-level semantic representation, respectively. In particular, in the low-level feature learning, a grid-wise attention mechanism is proposed to capture the dependencies of different regions from a facial expression image such that the parameter update of convolutional filters in low-level feature learning is regularized. In the high-level semantic representation, a visual transformer attention mechanism uses a sequence of visual semantic tokens (generated from pyramid features of high convolutional layer blocks) to learn the global representation. Extensive experiments have been conducted on three public facial expression datasets, CK+, FER+, and RAF-DB. The results show that our FER-VT has achieved state-of-the-art performance on these datasets, especially with a 100% accuracy on CK + datasets without any extra training data.
Code (1)
Tasks
Facial Expression RecognitionFacial Expression Recognition (FER)Similar Papers 제목 키워드 기반
FERAtt: Facial Expression Recognition with Attention Net
We present a new end-to-end network architecture for facial expression recognition with an attention model. It focuses attention in the human face and uses a Gaussian space representation for expression recognition. We d…
DecoderFacial Expression RecognitionFacial Expression Recognition (FER)General ClassificationReweighting Framewise Attention in Video Transformers for Facial Expression Understanding
Understanding facial expressions in videos requires modeling subtle and localized facial dynamics under unconstrained conditions. Although recent Vision Transformer (ViT)-based video models have shown strong performance …
Facial Expression RecognitionFacial expression recognition based on local region specific features and support vector machines
Facial expressions are one of the most powerful, natural and immediate means for human being to communicate their emotions and intensions. Recognition of facial expression has many applications including human-computer i…
Emotion RecognitionFacial Expression RecognitionFacial Expression Recognition (FER)Triplet Loss-less Center Loss Sampling Strategies in Facial Expression Recognition Scenarios
Facial expressions convey massive information and play a crucial role in emotional expression. Deep neural network (DNN) accompanied by deep metric learning (DML) techniques boost the discriminative ability of the model …
Facial Expression RecognitionFacial Expression Recognition (FER)Metric LearningTripletDesign of an Expression Recognition Solution Employing the Global Channel-Spatial Attention Mechanism
Facial expression recognition is a challenging classification task with broad application prospects in the field of human - computer interaction. This paper aims to introduce the methods of our upcoming 8th Affective Beh…
Facial Expression Recognition