paper-with-me

홈 › Papers

Mamba-Enhanced Text-Audio-Video Alignment Network for Emotion Recognition in Conversations

2024-09-08 · Xinran Li, Xiaomao Fan, Qingyang Wu, Xiaojiang Peng, Ye Li

Emotion Recognition in Conversations (ERCs) is a vital area within multimodal interaction research, dedicated to accurately identifying and classifying the emotions expressed by speakers throughout a conversation. Traditional ERC approaches predominantly rely on unimodal cues\-such as text, audio, or visual data\-leading to limitations in their effectiveness. These methods encounter two significant challenges: 1) Consistency in multimodal information. Before integrating various modalities, it is crucial to ensure that the data from different sources is aligned and coherent. 2) Contextual information capture. Successfully fusing multimodal features requires a keen understanding of the evolving emotional tone, especially in lengthy dialogues where emotions may shift and develop over time. To address these limitations, we propose a novel Mamba-enhanced Text-Audio-Video alignment network (MaTAV) for the ERC task. MaTAV is with the advantages of aligning unimodal features to ensure consistency across different modalities and handling long input sequences to better capture contextual multimodal information. The extensive experiments on the MELD and IEMOCAP datasets demonstrate that MaTAV significantly outperforms existing state-of-the-art methods on the ERC task with a big margin.

📄 PDF Abstract BibTeX arXiv:2409.05243

Code (1)

Alena-Xinran/MaTAV 공식 구현 pytorch

Tasks

Emotion RecognitionMambamultimodal interactionVideo Alignment

Similar Papers 제목 키워드 기반

Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models

2026-02-24 · Christian Simon, Masato Ishii, Wei-Yao Wang, Koichi Saito 외 arxiv

Scaling multimodal alignment between video and audio is challenging, particularly due to limited data and the mismatch between text descriptions and frame-level video information. In this work, we tackle the scaling chal…

Audio Generation

AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection

2026-04-02 · Damith Chamalke Senadeera, Dimitrios Kollias, Gregory Slabaugh arxiv

Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Audio can provide complementary evidence for violent events that are difficult to r…

Audio-Enhanced Text-to-Video Retrieval using Text-Conditioned Feature Alignment

2023-07-24 · ICCV 2023 1 · Sarah Ibrahimi, Xiaohang Sun, Pichao Wang, Amanmeet Garg 외

Text-to-video retrieval systems have recently made significant progress by utilizing pre-trained models trained on large-scale image-text pairs. However, most of the latest methods primarily focus on the video modality w…

RetrievalText to Video RetrievalVideo AlignmentVideo Retrieval

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

2026-06-02 · Xuan Wei, Jiahui Chen, Kaiheng Li, Mingyu Shao 외 arxiv

Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applications in talking-head synthesis, co-speech gesture generation, and …

Gesture GenerationVideo Generation

Learning Representations from Audio-Visual Spatial Alignment

2020-11-03 · NeurIPS 2020 12 · Pedro Morgado, Yi Li, Nuno Vasconcelos

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based …

Action RecognitionRepresentation LearningSemantic SegmentationVideo Semantic Segmentation