Multi-modal Facial Action Unit Detection with Large Pre-trained Models for the 5th Competition on Affective Behavior Analysis in-the-wild
Facial action unit detection has emerged as an important task within facial expression analysis, aimed at detecting specific pre-defined, objective facial expressions, such as lip tightening and cheek raising. This paper presents our submission to the Affective Behavior Analysis in-the-wild (ABAW) 2023 Competition for AU detection. We propose a multi-modal method for facial action unit detection with visual, acoustic, and lexical features extracted from the large pre-trained models. To provide high-quality details for visual feature extraction, we apply super-resolution and face alignment to the training data and show potential performance gain. Our approach achieves the F1 score of 52.3% on the official validation set of the 5th ABAW Challenge.
Code (0)
등록된 구현이 없습니다.
Tasks
Action Unit DetectionFace AlignmentFacial Action Unit DetectionSuper-ResolutionSimilar Papers 제목 키워드 기반
Multi-modal Multi-label Facial Action Unit Detection with Transformer
Facial Action Coding System is an important approach of facial expression analysis.This paper describes our submission to the third Affective Behavior Analysis (ABAW) 2022 competition. We proposed a transfomer based mode…
Action Unit DetectionFacial Action Unit DetectionMultimodal Spontaneous Emotion Corpus for Human Behavior Analysis
Emotion is expressed in multiple modalities, yet most research has considered at most one or two. This stems in part from the lack of large, diverse, well-annotated, multimodal databases with which to develop and test al…
Action Unit DetectionMultimodal Channel-Mixing: Channel and Spatial Masked AutoEncoder on Facial Action Unit Detection
Recent studies have focused on utilizing multi-modal data to develop robust models for facial Action Unit (AU) detection. However, the heterogeneity of multi-modal data poses challenges in learning effective representati…
Action Unit DetectionFacial Action Unit DetectionRepresentation LearningMultimodal Representation Learning Techniques for Comprehensive Facial State Analysis
Multimodal foundation models have significantly improved feature representation by integrating information from multiple modalities, making them highly suitable for a broader set of applications. However, the exploration…
Emotion RecognitionRepresentation LearningHierarchical Vision-Language Interaction for Facial Action Unit Detection
Facial Action Unit (AU) detection seeks to recognize subtle facial muscle activations as defined by the Facial Action Coding System (FACS). A primary challenge w.r.t AU detection is the effective learning of discriminati…
Facial Action Unit DetectionRepresentation Learning