On The Differences Between Song and Speech Emotion Recognition: Effect of Feature Sets, Feature Types, and Classifiers
In this paper, we evaluate the different features sets, feature types, and classifiers on both song and speech emotion recognition. Three feature sets: GeMAPS, pyAudioAnalysis, and LibROSA; two feature types: low-level descriptors and high-level statistical functions; and four classifiers: multilayer perceptron, LSTM, GRU, and convolution neural networks are examined on both song and speech data with the same parameter values. The results show no remarkable difference between song and speech data using the same method. In addition, high-level statistical functions of acoustic features gained higher performance scores than low-level descriptors in this classification task. This result strengthens the previous finding on the regression task which reported the advantage use of high-level features.
Code (1)
Tasks
Emotion RecognitionregressionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Speech & Song Emotion Recognition Using Multilayer Perceptron and Standard Vector Machine
Herein, we have compared the performance of SVM and MLP in emotion recognition using speech and song channels of the RAVDESS dataset. We have undertaken a journey to extract various audio features, identify optimal scali…
Data AugmentationEmotion RecognitionFeature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition
Robust speech emotion recognition relies on the quality of the speech features. We present speech features enhancement strategy that improves speech emotion recognition. We used the INTERSPEECH 2010 challenge feature-set…
Emotion Recognitionfeature selectionSpeech Emotion Recognitionemotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and fr…
Emotion RecognitionSelf-Supervised LearningSentiment AnalysisSpeech Emotion RecognitionEmotion Recognition from Speech
In this work, we conduct an extensive comparison of various approaches to speech based emotion recognition systems. The analyses were carried out on audio recordings from Ryerson Audio-Visual Database of Emotional Speech…
Emotion ClassificationEmotion RecognitionGeneral ClassificationSpeech Emotion Recognition Using Quaternion Convolutional Neural Networks
Although speech recognition has become a widespread technology, inferring emotion from speech signals still remains a challenge. To address this problem, this paper proposes a quaternion convolutional neural network (QCN…
Emotion RecognitionSpeech Emotion Recognitionspeech-recognitionSpeech Recognition