Deep Multimodal Learning for Emotion Recognition in Spoken Language
In this paper, we present a novel deep multimodal framework to predict human emotions based on sentence-level spoken language. Our architecture has two distinctive characteristics. First, it extracts the high-level features from both text and audio via a hybrid deep multimodal structure, which considers the spatial information from text, temporal information from audio, and high-level associations from low-level handcrafted features. Second, we fuse all features by using a three-layer deep neural network to learn the correlations across modalities and train the feature extraction and fusion modules together, allowing optimal global fine-tuning of the entire structure. We evaluated the proposed framework on the IEMOCAP dataset. Our result shows promising performance, achieving 60.4% in weighted accuracy for five emotion categories.
Code (0)
등록된 구현이 없습니다.
Tasks
Emotion RecognitionSentenceSimilar Papers 제목 키워드 기반
A Multimodal Emotion Recognition System: Integrating Facial Expressions, Body Movement, Speech, and Spoken Language
Traditional psychological evaluations rely heavily on human observation and interpretation, which are prone to subjectivity, bias, fatigue, and inconsistency. To address these limitations, this work presents a multimodal…
DiagnosticEmotion RecognitionMultimodal Emotion RecognitionVISTANet: VIsual Spoken Textual Additive Net for Interpretable Multimodal Emotion Recognition
This paper proposes a multimodal emotion recognition system, VIsual Spoken Textual Additive Net (VISTANet), to classify emotions reflected by input containing image, speech, and text into discrete classes. A new interpre…
Emotion RecognitionMultimodal Emotion RecognitionLow Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars
Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages. This pilot study presents NEST-V1 (Nepali Emotion and Speech Transformer - Version 1…
Sign Language TranslationEmotion ClassificationEmotion RecognitionSpeech RecognitionMMER: Multimodal Multi-task Learning for Speech Emotion Recognition
In this paper, we propose MMER, a novel Multimodal Multi-task learning approach for Speech Emotion Recognition. MMER leverages a novel multimodal network based on early-fusion and cross-modal self-attention between text …
Emotion RecognitionMultimodal Emotion RecognitionMulti-Task LearningSpeech Emotion RecognitionEmotion Recognition in Signers
Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatical and affective facial expressions and the scarcity of data for model training. T…
Emotion Recognition