paper-with-me

홈 › Papers

Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling

2026-03-12 · Junhyeong Byeon, Jeongyeol Kim, Sejoon Lim arxiv

Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality, such as facial expressions or speech, is often insufficient for capturing these complex emotional cues. To address this limitation, we propose a multimodal emotion recognition framework for the Expression (EXPR) task in the 10th Affective Behavior Analysis in-the-wild (ABAW) Challenge. Our framework builds on large-scale pre-trained models for visual and audio representation learning and integrates them in a unified multimodal architecture. To better capture temporal patterns in facial expression sequences, we incorporate temporal visual modeling over video windows. We further introduce a bi-directional cross-attention fusion module that enables visual and audio features to interact in a symmetric manner, facilitating cross-modal contextualization and complementary emotion understanding. In addition, we employ a text-guided contrastive objective to encourage semantically meaningful visual representations through alignment with emotion-related text prompts. Experimental results on the ABAW 10th EXPR benchmark demonstrate the effectiveness of the proposed framework, achieving a Macro F1 score of 0.32 compared to the baseline score of 0.25, and highlight the benefit of combining temporal visual modeling, audio representation learning, and cross-modal fusion for robust emotion recognition in unconstrained real-world environments.

📄 PDF Abstract BibTeX arXiv:2603.11971

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Emotion RecognitionRepresentation Learning

Similar Papers 제목 키워드 기반

Group Gated Fusion on Attention-based Bidirectional Alignment for Multimodal Emotion Recognition

2022-01-17 · PengFei Liu, Kun Li, Helen Meng

Emotion recognition is a challenging and actively-studied research area that plays a critical role in emotion-aware human-computer interaction systems. In a multimodal setting, temporal alignment between different modali…

Emotion RecognitionMultimodal Emotion Recognition

Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition

2025-01-25 · Junwei Feng, Xueyan Fan

Emotion recognition has a wide range of applications in human-computer interaction, marketing, healthcare, and other fields. In recent years, the development of deep learning technology has provided new methods for emoti…

cross-modal alignmentEmotion ClassificationEmotion RecognitionMarketing+1

LMR-CBT: Learning Modality-fused Representations with CB-Transformer for Multimodal Emotion Recognition from Unaligned Multimodal Sequences

2021-12-03 · Ziwang Fu, Feng Liu, HanYang Wang, Siyuan Shen 외

Learning modality-fused representations and processing unaligned multimodal sequences are meaningful and challenging in multimodal emotion recognition. Existing approaches use directional pairwise attention or a message …

Efficient Neural NetworkEmotion RecognitionMultimodal Emotion Recognition

HCAM -- Hierarchical Cross Attention Model for Multi-modal Emotion Recognition

2023-04-14 · Soumya Dutta, Sriram Ganapathy

Emotion recognition in conversations is challenging due to the multi-modal nature of the emotion expression. We propose a hierarchical cross-attention model (HCAM) approach to multi-modal emotion recognition using a comb…

Emotion ClassificationEmotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognition

MCN-CL: Multimodal Cross-Attention Network and Contrastive Learning for Multimodal Emotion Recognition

2025-11-14 · Feng Li, Ke Wu, Yongwei Li arxiv

Multimodal emotion recognition plays a key role in many domains, including mental health monitoring, educational interaction, and human-computer interaction. However, existing methods often face three major challenges: u…

Multimodal Emotion RecognitionContrastive Learning