paper-with-me

Papers

Modulation spectral features for speech emotion recognition using deep neural networks

2023-01-14 · Premjeet Singh, Md Sahidullah, Goutam Saha

This work explores the use of constant-Q transform based modulation spectral features (CQT-MSF) for speech emotion recognition (SER). The human perception and analysis of sound comprise of two important cognitive parts: early auditory analysis and cortex-based processing. The early auditory analysis considers spectrogram-based representation whereas cortex-based analysis includes extraction of temporal modulations from the spectrogram. This temporal modulation representation of spectrogram is called modulation spectral feature (MSF). As the constant-Q transform (CQT) provides higher resolution at emotion salient low-frequency regions of speech, we find that CQT-based spectrogram, together with its temporal modulations, provides a representation enriched with emotion-specific information. We argue that CQT-MSF when used with a 2-dimensional convolutional network can provide a time-shift invariant and deformation insensitive representation for SER. Our results show that CQT-MSF outperforms standard mel-scale based spectrogram and its modulation features on two popular SER databases, Berlin EmoDB and RAVDESS. We also show that our proposed feature outperforms the shift and deformation invariant scattering transform coefficients, hence, showing the importance of joint hand-crafted and self-learned feature extraction instead of reliance on complete hand-crafted features. Finally, we perform Grad-CAM analysis to visually inspect the contribution of constant-Q modulation features over SER.

📄 PDF Abstract BibTeX arXiv:2301.05868

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

VAW-GAN for Disentanglement and Recomposition of Emotional Elements in Speech

2020-11-03 · Kun Zhou, Berrak Sisman, Haizhou Li

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition…

DecoderDisentanglementGenerative Adversarial NetworkVoice Conversion

Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing

2026-01-26 · Marouane El Hizabri, Abdelfattah Bezzaz, Ismail Hayoukane, Youssef Taki arxiv

Speech Emotion Recognition systems often use static features like Mel-Frequency Cepstral Coefficients (MFCCs), Zero Crossing Rate (ZCR), and Root Mean Square Energy (RMSE). Because of this, they can misclassify emotions …

Speech Emotion RecognitionEmotion Classification

A Novel Trajectory-based Spatial-Temporal Spectral Features for Speech Emotion Recognition

2017-12-01 · ROCLINGIJCLCLP 2017 11 · Chun-Min Chang, Wei-Cheng Lin, Chi-Chun Lee
Emotion RecognitionSpeech Emotion Recognition

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

2025-06-02 · Alef Iury Siqueira Ferreira, Lucas Rafael Gris, Alexandre Ferro Filho, Lucas Ólives 외

Training SER models in natural, spontaneous speech is especially challenging due to the subtle expression of emotions and the unpredictable nature of real-world audio. In this paper, we present a robust system for the IN…

Audio TaggingEmotion RecognitionGraph AttentionQuantization+1

Biologically inspired speech emotion recognition

2021-11-15 · Reza Lotfidereshgi, Philippe Gournay

Conventional feature-based classification methods do not apply well to automatic recognition of speech emotions, mostly because the precise set of spectral and prosodic features that is required to identify the emotional…

Emotion RecognitionSpeech Emotion Recognition