Vision Transformer Equipped with Neural Resizer on Facial Expression Recognition Task
When it comes to wild conditions, Facial Expression Recognition is often challenged with low-quality data and imbalanced, ambiguous labels. This field has much benefited from CNN based approaches; however, CNN models have structural limitation to see the facial regions in distant. As a remedy, Transformer has been introduced to vision fields with global receptive field, but requires adjusting input spatial size to the pretrained models to enjoy their strong inductive bias at hands. We herein raise a question whether using the deterministic interpolation method is enough to feed low-resolution data to Transformer. In this work, we propose a novel training framework, Neural Resizer, to support Transformer by compensating information and downscaling in a data-driven manner trained with loss function balancing the noisiness and imbalance. Experiments show our Neural Resizer with F-PDLS loss function improves the performance with Transformer variants in general and nearly achieves the state-of-the-art performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Facial Expression RecognitionFacial Expression Recognition (FER)Inductive BiasMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
MULLER: Multilayer Laplacian Resizer for Vision
Image resizing operation is a fundamental preprocessing module in modern computer vision. Throughout the deep learning revolution, researchers have overlooked the potential of alternative resizing methods beyond the comm…
image-classificationImage ClassificationImage Quality Assessmentobject-detection+1More comprehensive facial inversion for more effective expression recognition
Facial expression recognition (FER) plays a significant role in the ubiquitous application of computer vision. We revisit this problem with a new perspective on whether it can acquire useful representations that improve …
Facial Expression RecognitionFacial Expression Recognition (FER)Image GenerationLearning to Resize Images for Computer Vision Tasks
For all the ways convolutional neural nets have revolutionized computer vision in recent years, one important aspect has received surprisingly little attention: the effect of image size on the accuracy of tasks being tra…
Image Quality AssessmentLearning Vision Transformer with Squeeze and Excitation for Facial Expression Recognition
As various databases of facial expressions have been made accessible over the last few decades, the Facial Expression Recognition (FER) task has gotten a lot of interest. The multiple sources of the available databases r…
Facial Expression Recognition (FER)Emotion Separation and Recognition from a Facial Expression by Generating the Poker Face with Vision Transformers
Representation learning and feature disentanglement have garnered significant research interest in the field of facial expression recognition (FER). The inherent ambiguity of emotion labels poses challenges for conventio…
DisentanglementFace GenerationFacial Expression RecognitionFacial Expression Recognition (FER)+1