paper-with-me

Papers

Learning spectro-temporal features with 3D CNNs for speech emotion recognition

2017-08-14 · Jaebok Kim, Khiet P. Truong, Gwenn Englebienne, Vanessa Evers

In this paper, we propose to use deep 3-dimensional convolutional networks (3D CNNs) in order to address the challenge of modelling spectro-temporal dynamics for speech emotion recognition (SER). Compared to a hybrid of Convolutional Neural Network and Long-Short-Term-Memory (CNN-LSTM), our proposed 3D CNNs simultaneously extract short-term and long-term spectral features with a moderate number of parameters. We evaluated our proposed and other state-of-the-art methods in a speaker-independent manner using aggregated corpora that give a large and diverse set of speakers. We found that 1) shallow temporal and moderately deep spectral kernels of a homogeneous architecture are optimal for the task; and 2) our 3D CNNs are more effective for spectro-temporal feature learning compared to other methods. Finally, we visualised the feature space obtained with our proposed method using t-distributed stochastic neighbour embedding (T-SNE) and could observe distinct clusters of emotions.

📄 PDF Abstract BibTeX arXiv:1708.05071

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

CNN-n-GRU: end-to-end speech emotion recognition from raw waveform signal using CNNs and gated recurrent unit networks

2023-03-23 · 21st IEEE International Conference on Machine Learning and Applications (ICMLA) 2023 3 · Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian Mishara

We present CNN-n-GRU, a new end-to-end (E2E) architecture built of an n-layer convolutional neural network (CNN) followed sequentially by an n-layer Gated Recurrent Unit (GRU) for speech emotion recognition. CNNs and RNN…

Emotion RecognitionSpeech Emotion Recognition

Emotion Recognition from Speech

2019-12-22 · Kannan Venkataramanan, Haresh Rengaraj Rajamohan

In this work, we conduct an extensive comparison of various approaches to speech based emotion recognition systems. The analyses were carried out on audio recordings from Ryerson Audio-Visual Database of Emotional Speech…

Emotion ClassificationEmotion RecognitionGeneral Classification

A breakthrough in Speech emotion recognition using Deep Retinal Convolution Neural Networks

2017-07-12 · Yafeng Niu, Dongsheng Zou, Yadong Niu, Zhongshi He 외

Speech emotion recognition (SER) is to study the formation and change of speaker's emotional state from the speech signal perspective, so as to make the interaction between human and computer more intelligent. SER is a c…

Data AugmentationEmotion RecognitionSpeech Emotion Recognition

Modulation spectral features for speech emotion recognition using deep neural networks

2023-01-14 · Premjeet Singh, Md Sahidullah, Goutam Saha

This work explores the use of constant-Q transform based modulation spectral features (CQT-MSF) for speech emotion recognition (SER). The human perception and analysis of sound comprise of two important cognitive parts: …

Emotion RecognitionSpeech Emotion Recognition

Unlocking the Emotional States of High-Risk Suicide Callers through Speech Analysis

2024-03-22 · IEEE 18th International Conference on Semantic Computing (ICSC) 2024 3 · Alaa Nfissi, Wassim Bouachir, Nizar Bouguila, Brian Mishara

Suicide remains a major public health concern worldwide, and early detection of suicidal ideation is crucial for prevention. One promising approach for monitoring symptoms is through the prediction of suicidal speech, as…

Emotion RecognitionSpeech Emotion Recognition