Deep Multilayer Perceptrons for Dimensional Speech Emotion Recognition
Modern deep learning architectures are ordinarily performed on high-performance computing facilities due to the large size of the input features and complexity of its model. This paper proposes traditional multilayer perceptrons (MLP) with deep layers and small input size to tackle that computation requirement limitation. The result shows that our proposed deep MLP outperformed modern deep learning architectures, i.e., LSTM and CNN, on the same number of layers and value of parameters. The deep MLP exhibited the highest performance on both speaker-dependent and speaker-independent scenarios on IEMOCAP and MSP-IMPROV corpus.
Code (1)
Tasks
Deep LearningEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Speech & Song Emotion Recognition Using Multilayer Perceptron and Standard Vector Machine
Herein, we have compared the performance of SVM and MLP in emotion recognition using speech and song channels of the RAVDESS dataset. We have undertaken a journey to extract various audio features, identify optimal scali…
Data AugmentationEmotion RecognitionOn The Differences Between Song and Speech Emotion Recognition: Effect of Feature Sets, Feature Types, and Classifiers
In this paper, we evaluate the different features sets, feature types, and classifiers on both song and speech emotion recognition. Three feature sets: GeMAPS, pyAudioAnalysis, and LibROSA; two feature types: low-level d…
Emotion RecognitionregressionSpeech Emotion RecognitionUnsupervised low-rank representations for speech emotion recognition
We examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptor…
Dimensionality ReductionEmotion RecognitionGeneral ClassificationSpeech Emotion RecognitionF0 Modeling In Hmm-Based Speech Synthesis System Using Deep Belief Network
In recent years multilayer perceptrons (MLPs) with many hid- den layers Deep Neural Network (DNN) has performed sur- prisingly well in many speech tasks, i.e. speech recognition, speaker verification, speech synthesis et…
ClusteringSpeaker Verificationspeech-recognitionSpeech Recognition+1Investigating salient representations and label Variance in Dimensional Speech Emotion Analysis
Representations derived from models such as BERT (Bidirectional Encoder Representations from Transformers) and HuBERT (Hidden units BERT), have helped to achieve state-of-the-art performance in dimensional speech emotion…
Emotion RecognitionSpeech Emotion Recognition