paper-with-me

홈 › Papers

Attention Driven Fusion for Multi-Modal Emotion Recognition

2020-09-23 · Darshana Priyasad, Tharindu Fernando, Simon Denman, Clinton Fookes, Sridha Sridharan

Deep learning has emerged as a powerful alternative to hand-crafted methods for emotion recognition on combined acoustic and text modalities. Baseline systems model emotion information in text and acoustic modes independently using Deep Convolutional Neural Networks (DCNN) and Recurrent Neural Networks (RNN), followed by applying attention, fusion, and classification. In this paper, we present a deep learning-based approach to exploit and fuse text and acoustic data for emotion classification. We utilize a SincNet layer, based on parameterized sinc functions with band-pass filters, to extract acoustic features from raw audio followed by a DCNN. This approach learns filter banks tuned for emotion recognition and provides more effective features compared to directly applying convolutions over the raw speech signal. For text processing, we use two branches (a DCNN and a Bi-direction RNN followed by a DCNN) in parallel where cross attention is introduced to infer the N-gram level correlations on hidden representations received from the Bi-RNN. Following existing state-of-the-art, we evaluate the performance of the proposed system on the IEMOCAP dataset. Experimental results indicate that the proposed system outperforms existing methods, achieving 3.5% improvement in weighted accuracy.

📄 PDF Abstract BibTeX arXiv:2009.10991

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion ClassificationEmotion Recognition

Methods 이 논문이 사용한 방법론

DCNN Diffusion-convolutional neural networks (DCNN) is a model for graph-structured data. Through the introduction of a diffusion-convolution operation, diffusion-based representations…

Similar Papers 제목 키워드 기반

Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space

2025-04-14 · Kangwei Liu, Junwu Liu, Xiaowei Yi, Jinlin Guo 외

Audio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for compre…

Contrastive LearningDiversity

Relational graph-driven differential denoising and diffusion attention fusion for multimodal conversation emotion recognition

2026-03-22 · Ying Liu, Yuntao Shou, Wei Ai, Tao Meng 외 arxiv

In real-world scenarios, audio and video signals are often subject to environmental noise and limited acquisition conditions, resulting in extracted features containing excessive noise. Furthermore, there is an imbalance…

Emotion Recognition

Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis

2025-11-25 · Zheqi Hu, Xuanjing Chen, Jinlin Hu arxiv

With the rapid growth of AI-generated content (AIGC) across domains such as music, video, and literature, the demand for emotionally aware recommendation systems has become increasingly important. Traditional recommender…

Emotional IntelligenceRecommendation SystemsIntent Recognition

GCM-Net: Graph-enhanced Cross-Modal Infusion with a Metaheuristic-Driven Network for Video Sentiment and Emotion Analysis

2024-10-02 · Prasad Chaudhari, Aman Kumar, Chandravardhan Singh Raghaw, Mohammad Zia Ur Rehman 외

Sentiment analysis and emotion recognition in videos are challenging tasks, given the diversity and complexity of the information conveyed in different modalities. Developing a highly competent framework that effectively…

Emotion RecognitionGraph SamplingSentiment Analysis

HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution Learning

2025-07-09 · Chuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang 외 arxiv

Multi-modal emotion recognition has garnered increasing attention as it plays a significant role in human-computer interaction (HCI) in recent years. Since different discrete emotions may exist at the same time, compared…

Emotion Recognition