Optimizing Speech Emotion Recognition using Manta-Ray Based Feature Selection
Emotion recognition from audio signals has been regarded as a challenging task in signal processing as it can be considered as a collection of static and dynamic classification tasks. Recognition of emotions from speech data has been heavily relied upon end-to-end feature extraction and classification using machine learning models, though the absence of feature selection and optimization have restrained the performance of these methods. Recent studies have shown that Mel Frequency Cepstral Coefficients (MFCC) have been emerged as one of the most relied feature extraction methods, though it circumscribes the accuracy of classification with a very small feature dimension. In this paper, we propose that the concatenation of features, extracted by using different existing feature extraction methods can not only boost the classification accuracy but also expands the possibility of efficient feature selection. We have used Linear Predictive Coding (LPC) apart from the MFCC feature extraction method, before feature merging. Besides, we have performed a novel application of Manta Ray optimization in speech emotion recognition tasks that resulted in a state-of-the-art result in this field. We have evaluated the performance of our model using SAVEE and Emo-DB, two publicly available datasets. Our proposed method outperformed all the existing methods in speech emotion analysis and resulted in a decent result in these two datasets with a classification accuracy of 97.06% and 97.68% respectively.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationEmotion Recognitionfeature selectionGeneral ClassificationSpeech Emotion RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Speech Emotion Recognition Using Speech Feature and Word Embedding
—Emotion recognition can be performed automatically from many modalities. This paper presents a categorical speech emotion recognition using speech features and word embedding. Text features can be combined with speech f…
Emotion RecognitionSpeech Emotion RecognitionFeature Selection Enhancement and Feature Space Visualization for Speech-Based Emotion Recognition
Robust speech emotion recognition relies on the quality of the speech features. We present speech features enhancement strategy that improves speech emotion recognition. We used the INTERSPEECH 2010 challenge feature-set…
Emotion Recognitionfeature selectionSpeech Emotion RecognitionSpeech Emotion Recognition Based on Multi-feature and Multi-lingual Fusion
A speech emotion recognition algorithm based on multi-feature and Multi-lingual fusion is proposed in order to resolve low recognition accuracy caused by lack of large speech dataset and low robustness of acoustic featur…
Emotion RecognitionSpeech Emotion RecognitionManta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-Sequence
In few-shot action recognition (FSAR), long sub-sequences of video naturally express entire actions more effectively. However, the high computational complexity of mainstream Transformer-based methods limits their applic…
Action RecognitionContrastive LearningFew-Shot action recognitionFew Shot Action Recognition+1Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages
Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…
Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition