paper-with-me

홈 › Papers

Non-linear frequency warping using constant-Q transformation for speech emotion recognition

2021-02-08 · Premjeet Singh, Goutam Saha, Md Sahidullah

In this work, we explore the constant-Q transform (CQT) for speech emotion recognition (SER). The CQT-based time-frequency analysis provides variable spectro-temporal resolution with higher frequency resolution at lower frequencies. Since lower-frequency regions of speech signal contain more emotion-related information than higher-frequency regions, the increased low-frequency resolution of CQT makes it more promising for SER than standard short-time Fourier transform (STFT). We present a comparative analysis of short-term acoustic features based on STFT and CQT for SER with deep neural network (DNN) as a back-end classifier. We optimize different parameters for both features. The CQT-based features outperform the STFT-based spectral features for SER experiments. Further experiments with cross-corpora evaluation demonstrate that the CQT-based systems provide better generalization with out-of-domain training data.

📄 PDF Abstract BibTeX arXiv:2102.04029

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Towards Progressive Multi-Frequency Representation for Image Warping

2024-01-01 · CVPR 2024 1 · Jun Xiao, Zihang Lyu, Cong Zhang, Yakun Ju 외

Image warping a classic task in computer vision aims to use geometric transformations to change the appearance of images. Recent methods learn the resampling kernels for warping through neural networks to estimate mi…

Image Super-ResolutionMissing ValuesSuper-Resolution

Dysarthria Normalization via Local Lie Group Transformations for Robust ASR

2025-04-16 · Mikhail Osipov

We present a geometry-driven method for normalizing dysarthric speech by modeling time, frequency, and amplitude distortions as smooth, local Lie group transformations of spectrograms. Scalar fields generate these deform…

Robust Speech Recognitionspeech-recognitionSpeech RecognitionZero-shot Generalization

Analysis of constant-Q filterbank based representations for speech emotion recognition

2022-11-29 · Premjeet Singh, Shefali Waldekar, Md Sahidullah, Goutam Saha

This work analyzes the constant-Q filterbank-based time-frequency representations for speech emotion recognition (SER). Constant-Q filterbank provides non-linear spectro-temporal representation with higher frequency reso…

Emotion RecognitionSpeech Emotion Recognition

Vowel Enhancement in Early Stage Spanish Esophageal Speech Using Natural Glottal Flow Pulse and Vocal Tract Frequency Warping

2015-09-01 · WS 2015 9 · Rizwan Ishaq, Dhanananjaya Gowda, Paavo Alku, Begonya Garcia Zapirain
Speech Enhancement

Comparison of Time-Frequency Representations for Environmental Sound Classification using Convolutional Neural Networks

2017-06-22 · M. Huzaifah

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual display…

Audio ClassificationClassificationEnvironmental Sound ClassificationGeneral Classification+3