paper-with-me

Papers

Multi-stream Attention-based BLSTM with Feature Segmentation for Speech Emotion Recognition

2020-10-25 · Interspeech 2020 10 · Yuya Chiba1, Takashi Nose1, Akinori Ito

This paper proposes a speech emotion recognition technique that considers the suprasegmental characteristics and temporal change of individual speech parameters. In recent years, speech emotion recognition using Bidirectional LSTM (BLSTM) has been studied actively because the model can focus on a particular temporal region that contains strong emotional characteristics. One of the model’s weaknesses is that it cannot consider the statistics of speech features, which are known to be effective for speech emotion recognition. Besides, this method cannot train individual attention parameters for different descriptors because it handles the input sequence by a single BLSTM. In this paper, we introduce feature segmentation and multi-stream processing into attention-based BLSTM to solve these problems. In addition, we employed data augmentation based on emotional speech synthesis in a training step. The classification experiments between four emotions (i.e., anger, joy, neutral, and sadness) using the Japanese Twitter-based Emotional Speech corpus (JTES) showed that the proposed method obtained a recognition accuracy of 73.4%, which is comparable to human evaluation (75.5%).

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationEmotional Speech SynthesisEmotion RecognitionSpeech Emotion RecognitionSpeech Synthesis

Similar Papers 제목 키워드 기반

High Performance Sequence-to-Sequence Model for Streaming Speech Recognition

2020-03-22 · Thai-Son Nguyen, Ngoc-Quan Pham, Sebastian Stueker, Alex Waibel

Recently sequence-to-sequence models have started to achieve state-of-the-art performance on standard speech recognition tasks when processing audio data in batch mode, i.e., the complete audio data is available when sta…

speech-recognitionSpeech RecognitionVocal Bursts Intensity Prediction

End-to-End Multi-View Lipreading

2017-09-01 · Stavros Petridis, Yujiang Wang, Zuwei Li, Maja Pantic

Non-frontal lip views contain useful information which can be used to enhance the performance of frontal view lipreading. However, the vast majority of recent lipreading works, including the deep learning approaches whic…

General ClassificationLipreading

End-to-End Audiovisual Fusion with LSTMs

2017-09-12 · Stavros Petridis, Yujiang Wang, Zuwei Li, Maja Pantic

Several end-to-end deep learning approaches have been recently presented which simultaneously extract visual features from the input images and perform visual speech classification. However, research on jointly extractin…

ClassificationGeneral Classificationspeech-recognitionSpeech Recognition

Pre-trained Deep Convolution Neural Network Model With Attention for Speech Emotion Recognition

2021-03-02 · Front. Physiol 2021 3 · Hua Zhang, Ruoyun Gou, Jili Shang, Fangyao Shen 외

Speech emotion recognition (SER) is a difficult and challenging task because of the affective variances between different speakers. The performances of SER are extremely reliant on the extracted features from speech sign…

Emotion RecognitionSentenceSpeech Emotion Recognition

Utterance-level end-to-end language identification using attention-based CNN-BLSTM

2019-02-20 · Weicheng Cai, Danwei Cai, Shen Huang, Ming Li

In this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level,…

Language Identification