paper-with-me

Papers

Learning Alignment for Multimodal Emotion Recognition from Speech

2019-09-06 · Haiyang Xu, HUI ZHANG, Kun Han, Yun Wang, Yiping Peng, Xiangang Li

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech recognition techniques to generate text from speech and then apply natural language processing to analyze the sentiment. Further, emotion recognition will be beneficial from using audio-textual multimodal information, it is not trivial to build a system to learn from multimodality. One can build models for two input sources separately and combine them in a decision level, but this method ignores the interaction between speech and text in the temporal domain. In this paper, we propose to use an attention mechanism to learn the alignment between speech frames and text words, aiming to produce more accurate multimodal feature representations. The aligned multimodal features are fed into a sequential model for emotion recognition. We evaluate the approach on the IEMOCAP dataset and the experimental results show the proposed approach achieves the state-of-the-art performance on the dataset.

📄 PDF Abstract BibTeX arXiv:1909.05645

Code (1)

ZhiqiWang12-hash/text_audio_classification tf

Tasks

Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization

2025-07-25 · Hsuan-Yu Wang, Pei-Ying Lee, Berlin Chen arxiv

In this paper, we investigate the impact of incorporating timestamp-based alignment between Automatic Speech Recognition (ASR) transcripts and Speaker Diarization (SD) outputs on Speech Emotion Recognition (SER) accuracy…

Multimodal Emotion RecognitionSpeech Emotion RecognitionSpeaker DiarizationSpeech Recognition

Qieemo: Speech Is All You Need in the Emotion Recognition in Conversations

2025-03-05 · Jinming Chen, Jingyi Fang, Yuanzhong Zheng, Yaoxuan Wang 외

Emotion recognition plays a pivotal role in intelligent human-machine interaction systems. Multimodal approaches benefit from the fusion of diverse modalities, thereby improving the recognition accuracy. However, the lac…

AllAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Classification+3

Group Gated Fusion on Attention-based Bidirectional Alignment for Multimodal Emotion Recognition

2022-01-17 · PengFei Liu, Kun Li, Helen Meng

Emotion recognition is a challenging and actively-studied research area that plays a critical role in emotion-aware human-computer interaction systems. In a multimodal setting, temporal alignment between different modali…

Emotion RecognitionMultimodal Emotion Recognition

Improving Speech Emotion Recognition Through Cross Modal Attention Alignment and Balanced Stacking Model

2025-05-26 · Lucas Ueda, João Lima, Leonardo Marques, Paula Costa

Emotion plays a fundamental role in human interaction, and therefore systems capable of identifying emotions in speech are crucial in the context of human-computer interaction. Speech emotion recognition (SER) is a chall…

Emotion RecognitionSpeech Emotion Recognition

Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment

2024-12-30 · Xuechen Wang, Shiwan Zhao, Haoqin Sun, Hui Wang 외

Multimodal emotion recognition (MER), leveraging speech and text, has emerged as a pivotal domain within human-computer interaction, demanding sophisticated methods for effective multimodal integration. The challenge of …

cross-modal alignmentEmotion RecognitionMultimodal Emotion Recognition