paper-with-me

홈 › Papers

Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention

2025-07-04 · HyeYoung Lee, Muhammad Nadeem arxiv

Speech Emotion Recognition (SER) traditionally relies on auditory data analysis for emotion classification. Several studies have adopted different methods for SER. However, existing SER methods often struggle to capture subtle emotional variations and generalize across diverse datasets. In this article, we use Mel-Frequency Cepstral Coefficients (MFCCs) as spectral features to bridge the gap between computational emotion processing and human auditory perception. To further improve robustness and feature diversity, we propose a novel 1D-CNN-based SER framework that integrates data augmentation techniques. MFCC features extracted from the augmented data are processed using a 1D Convolutional Neural Network (CNN) architecture enhanced with channel and spatial attention mechanisms. These attention modules allow the model to highlight key emotional patterns, enhancing its ability to capture subtle variations in speech signals. The proposed method delivers cutting-edge performance, achieving the accuracy of 97.49% for SAVEE, 99.23% for RAVDESS, 89.31% for CREMA-D, 99.82% for TESS, 99.53% for EMO-DB, and 96.39% for EMOVO. Experimental results show new benchmarks in SER, demonstrating the effectiveness of our approach in recognizing emotional expressions with high precision. Our evaluation demonstrates that the integration of advanced Deep Learning (DL) methods substantially enhances generalization across diverse datasets, underscoring their potential to advance SER for real-world deployment in assistive technologies and human-computer interaction.

📄 PDF Abstract BibTeX arXiv:2507.03251

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionEmotion ClassificationData Augmentation

Similar Papers 제목 키워드 기반

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

2025-06-02 · Alef Iury Siqueira Ferreira, Lucas Rafael Gris, Alexandre Ferro Filho, Lucas Ólives 외

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the IN…

Audio TaggingEmotion RecognitionGraph AttentionQuantization+1

Speech Emotion Recognition with Global-Aware Fusion on Multi-scale Feature Representation

2022-04-12 · Wenjing Zhu, Xiang Li

Speech Emotion Recognition (SER) is a fundamental task to predict the emotion label from speech data. Recent works mostly focus on using convolutional neural networks~(CNNs) to learn local attention map on fixed-scale fe…

Emotion RecognitionSpeech Emotion Recognition

Hybrid CNN-Transformer Architecture for Arabic Speech Emotion Recognition

2026-03-28 · Youcef Soufiane Gheffari, Oussama Mustapha Benouddane, Samiya Silarbi arxiv

Recognizing emotions from speech using machine learning has become an active research area due to its importance in building human-centered applications. However, while many studies have been conducted in English, German…

Speech Emotion Recognition

Fine-grained Early Frequency Attention for Deep Speaker Representation Learning

2020-09-03 · Amirhossein Hajavi, Ali Etemad

Deep learning techniques have considerably improved speech processing in recent years. Speaker representations extracted by deep learning models are being used in a wide range of tasks such as speaker recognition and spe…

Deep LearningEmotion RecognitionRepresentation LearningSpeaker Recognition+3

EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification

2025-05-26 · Deok-Hyeon Cho, Hyung-Seok Oh, Seung-bin Kim, Seong-Whan Lee

Speech emotion recognition predicts a speaker's emotional state from speech signals using discrete labels or continuous dimensions such as arousal, valence, and dominance (VAD). We propose EmoSphere-SER, a joint model th…

Emotion RecognitionregressionSpeech Emotion Recognition