Multi-modal Multi-label Facial Action Unit Detection with Transformer
Facial Action Coding System is an important approach of facial expression analysis.This paper describes our submission to the third Affective Behavior Analysis (ABAW) 2022 competition. We proposed a transfomer based model to detect facial action unit (FAU) in video. To be specific, we firstly trained a multi-modal model to extract both audio and visual feature. After that, we proposed a action units correlation module to learn relationships between each action unit labels and refine action unit detection result. Experimental results on validation dataset shows that our method achieves better performance than baseline model, which verifies that the effectiveness of proposed network.
Code (0)
등록된 구현이 없습니다.
Tasks
Action Unit DetectionFacial Action Unit DetectionSimilar Papers 제목 키워드 기반
MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label Generation
The lack of large-scale, demographically diverse face images with precise Action Unit (AU) occurrence and intensity annotations has long been recognized as a fundamental bottleneck in developing generalizable AU recognit…
Representation LearningMulti-Task Multi-Modal Self-Supervised Learning for Facial Expression Recognition
Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when desi…
Emotion ClassificationEmotion Recognition in ConversationFacial Expression RecognitionSelf-Supervised LearningREACT2023: the first Multi-modal Multiple Appropriate Facial Reaction Generation Challenge
The Multi-modal Multiple Appropriate Facial Reaction Generation Challenge (REACT2023) is the first competition event focused on evaluating multimedia processing and machine learning techniques for generating human-approp…
Media2Face: Co-speech Facial Animation Generation With Multi-Modality Guidance
The synthesis of 3D facial animations from speech has garnered considerable attention. Due to the scarcity of high-quality 4D facial data and well-annotated abundant multi-modality labels, previous methods often suffer f…
Fusing Body Posture with Facial Expressions for Joint Recognition of Affect in Child-Robot Interaction
In this paper we address the problem of multi-cue affect recognition in challenging scenarios such as child-robot interaction. Towards this goal we propose a method for automatic recognition of affect that leverages body…