AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts
Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for integrating such multimodal information often stumble, leading to less-than-ideal outcomes in the task of facial action unit detection. To overcome these shortcomings, we propose a novel approach utilizing audio-visual multimodal data. This method enhances audio feature extraction by leveraging Mel Frequency Cepstral Coefficients (MFCC) and Log-Mel spectrogram features alongside a pre-trained VGGish network. Moreover, this paper adaptively captures fusion features across modalities by modeling the temporal relationships, and ultilizes a pre-trained GPT-2 model for sophisticated context-aware fusion of multimodal information. Our method notably improves the accuracy of AU detection by understanding the temporal and contextual nuances of the data, showcasing significant advancements in the comprehension of intricate scenarios. These findings underscore the potential of integrating temporal dynamics and contextual interpretation, paving the way for future research endeavors.
Code (0)
등록된 구현이 없습니다.
Tasks
Action Unit DetectionFacial Action Unit DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Change Detection in Multi-temporal VHR Images Based on Deep Siamese Multi-scale Convolutional Networks
Very-high-resolution (VHR) images can provide abundant ground details and spatial geometric information. Change detection in multi-temporal VHR images plays a significant role in urban expansion and area internal change …
Change DetectionAction Unit Detection with Region Adaptation, Multi-labeling Learning and Optimal Temporal Fusing
Action Unit (AU) detection becomes essential for facial analysis. Many proposed approaches face challenging problems in dealing with the alignments of different face regions, in the effective fusion of temporal informati…
Action Unit DetectionMulti-Label LearningProgress Regression RNN for Online Spatial-Temporal Action Localization in Unconstrained Videos
Previous spatial-temporal action localization methods commonly follow the pipeline of object detection to estimate bounding boxes and labels of actions. However, the temporal relation of an action has not been fully expl…
Action Localizationobject-detectionObject DetectionPosition+2Advancing Community Detection with Graph Convolutional Neural Networks: Bridging Topological and Attributive Cohesion
Community detection, a vital technology for real-world applications, uncovers cohesive node groups (communities) by leveraging both topological and attribute similarities in social networks. However, existing Graph Convo…
AttributeCommunity Detection2D bidirectional gated recurrent unit convolutional Neural networks for end-to-end violence detection In videos
Abnormal behavior detection, action recognition, fight and violence detection in videos is an area that has attracted a lot of interest in recent years. In this work, we propose an architecture that combines a Bidirectio…
Action Recognition