paper-with-me

Papers

Heterogeneous bimodal attention fusion for speech emotion recognition

2025-03-09 · Jiachen Luo, Huy Phan, Lin Wang, Joshua Reiss

Multi-modal emotion recognition in conversations is a challenging problem due to the complex and complementary interactions between different modalities. Audio and textual cues are particularly important for understanding emotions from a human perspective. Most existing studies focus on exploring interactions between audio and text modalities at the same representation level. However, a critical issue is often overlooked: the heterogeneous modality gap between low-level audio representations and high-level text representations. To address this problem, we propose a novel framework called Heterogeneous Bimodal Attention Fusion (HBAF) for multi-level multi-modal interaction in conversational emotion recognition. The proposed method comprises three key modules: the uni-modal representation module, the multi-modal fusion module, and the inter-modal contrastive learning module. The uni-modal representation module incorporates contextual content into low-level audio representations to bridge the heterogeneous multi-modal gap, enabling more effective fusion. The multi-modal fusion module uses dynamic bimodal attention and a dynamic gating mechanism to filter incorrect cross-modal relationships and fully exploit both intra-modal and inter-modal interactions. Finally, the inter-modal contrastive learning module captures complex absolute and relative interactions between audio and text modalities. Experiments on the MELD and IEMOCAP datasets demonstrate that the proposed HBAF method outperforms existing state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2503.06405

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningEmotion RecognitionSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Contrastive Learning 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

Bimodal Connection Attention Fusion for Speech Emotion Recognition

2025-03-08 · Jiachen Luo, Huy Phan, Lin Wang, Joshua D. Reiss

Multi-modal emotion recognition is challenging due to the difficulty of extracting features that capture subtle emotional differences. Understanding multi-modal interactions and connections is key to building effective b…

DecoderEmotion RecognitionSpeech Emotion Recognition

A Simple Attention-Based Mechanism for Bimodal Emotion Classification

2024-06-28 · Mazen Elabd, Sardar Jaf

Big data contain rich information for machine learning algorithms to utilize when learning important features during classification tasks. Human beings express their emotion using certain words, speech (tone, pitch, spee…

ClassificationDeep LearningEmotion Classification

Bimodal Speech Emotion Recognition Using Pre-Trained Language Models

2019-11-29 · Verena Heusser, Niklas Freymuth, Stefan Constantin, Alex Waibel

Speech emotion recognition is a challenging task and an important step towards more natural human-machine interaction. We show that pre-trained language models can be fine-tuned for text emotion recognition, achieving an…

Emotion RecognitionReinforcement LearningSpeech Emotion Recognition

PARROT: Synergizing Mamba and Attention-based SSL Pre-Trained Models via Parallel Branch Hadamard Optimal Transport for Speech Emotion Recognition

2025-06-01 · Orchid Chetia Phukan, Mohd Mujtaba Akhtar, Girish, Swarup Ranjan Behera 외

The emergence of Mamba as an alternative to attention-based architectures has led to the development of Mamba-based self-supervised learning (SSL) pre-trained models (PTMs) for speech and audio processing. Recent studies…

Emotion RecognitionMambaSelf-Supervised LearningSpeech Emotion Recognition

Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?

2025-06-02 · Mohd Mujtaba Akhtar, Orchid Chetia Phukan, Girish, Swarup Ranjan Behera 외

In this work, we focus on non-verbal vocal sounds emotion recognition (NVER). We investigate mamba-based audio foundation models (MAFMs) for the first time for NVER and hypothesize that MAFMs will outperform attention-ba…

Emotion RecognitionMambaSpeech Emotion RecognitionSynthetic Speech Detection