paper-with-me

홈 › Papers

Speech Emotion Recognition with Dual-Sequence LSTM Architecture

2019-10-20 · Jianyou Wang, Michael Xue, Ryan Culhane, Enmao Diao, Jie Ding, Vahid Tarokh

Speech Emotion Recognition (SER) has emerged as a critical component of the next generation human-machine interfacing technologies. In this work, we propose a new dual-level model that predicts emotions based on both MFCC features and mel-spectrograms produced from raw audio signals. Each utterance is preprocessed into MFCC features and two mel-spectrograms at different time-frequency resolutions. A standard LSTM processes the MFCC features, while a novel LSTM architecture, denoted as Dual-Sequence LSTM (DS-LSTM), processes the two mel-spectrograms simultaneously. The outputs are later averaged to produce a final classification of the utterance. Our proposed model achieves, on average, a weighted accuracy of 72.7% and an unweighted accuracy of 73.3%---a 6% improvement over current state-of-the-art unimodal models---and is comparable with multimodal models that leverage textual information as well as audio signals.

📄 PDF Abstract BibTeX arXiv:1910.08874

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

Sigmoid Activation 설명 없음
Tanh Activation 설명 없음
LSTM An LSTM is a type of recurrent neural network that addresses the vanishing gradient problem in vanilla…

Similar Papers 제목 키워드 기반

Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM

2024-11-14 · Xiaoran Yang, Shuhan Yu, Wenxi Xu

This paper builds upon an existing speech emotion recognition model by adding an additional LSTM layer to improve the accuracy and processing efficiency of emotion recognition from audio data. By capturing the long-term …

Emotion RecognitionSentiment AnalysisSpeech Emotion Recognition

Multi-stream Attention-based BLSTM with Feature Segmentation for Speech Emotion Recognition

2020-10-25 · Interspeech 2020 10 · Yuya Chiba1, Takashi Nose1, Akinori Ito

This paper proposes a speech emotion recognition technique that considers the suprasegmental characteristics and temporal change of individual speech parameters. In recent years, speech emotion recognition using Bidir…

Data AugmentationEmotional Speech SynthesisEmotion RecognitionSpeech Emotion Recognition+1

Speech Emotion Recognition Based on CNN+LSTM Model

2021-10-01 · ROCLING 2021 10 · Wei Mou, Pei-Hsuan Shen, Chu-Yun Chu, Yu-Cheng Chiu 외

Due to the popularity of intelligent dialogue assistant services, speech emotion recognition has become more and more important. In the communication between humans and machines, emotion recognition and emotion analysis …

Emotion RecognitionmodelSpeech Emotion Recognition

Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

2026-04-16 · Adelekun Oluwademilade, Ademola Adedamola, Abiola Abdulhakeem, Akinpelu Azeezat 외 arxiv

Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of …

Speech Emotion Recognition

Speech Emotion Recognition Based on Self-Attention Weight Correction for Acoustic and Text Features

2022-11-08 · IEEE Access 2022 11 · JENNIFER SANTOSO, Takeshi Yamada, Kenkichi Ishizuka, Taiichi Hashimoto 외

Speech emotion recognition (SER) is essential for understanding a speaker’s intention. Recently, some groups have attempted to improve SER performance using a bidirectional long short-term memory (BLSTM) to extract featu…

Emotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognitionspeech-recognition+1