paper-with-me

홈 › Papers

Analysis of constant-Q filterbank based representations for speech emotion recognition

2022-11-29 · Premjeet Singh, Shefali Waldekar, Md Sahidullah, Goutam Saha

This work analyzes the constant-Q filterbank-based time-frequency representations for speech emotion recognition (SER). Constant-Q filterbank provides non-linear spectro-temporal representation with higher frequency resolution at low frequencies. Our investigation reveals how the increased low-frequency resolution benefits SER. The time-domain comparative analysis between short-term mel-frequency spectral coefficients (MFSCs) and constant-Q filterbank-based features, namely constant-Q transform (CQT) and continuous wavelet transform (CWT), reveals that constant-Q representations provide higher time-invariance at low-frequencies. This provides increased robustness against emotion irrelevant temporal variations in pitch, especially for low-arousal emotions. The corresponding frequency-domain analysis over different emotion classes shows better resolution of pitch harmonics in constant-Q-based time-frequency representations than MFSC. These advantages of constant-Q representations are further consolidated by SER performance in the extensive evaluation of features over four publicly available databases with six advanced deep neural network architectures as the back-end classifiers. Our inferences in this study hint toward the suitability and potentiality of constant-Q features for SER.

📄 PDF Abstract BibTeX arXiv:2211.16363

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Filterbank design for end-to-end speech separation

2019-10-23 · Manuel Pariente, Samuele Cornell, Antoine Deleforge, Emmanuel Vincent

Single-channel speech separation has recently made great progress thanks to learned filterbanks as used in ConvTasNet. In parallel, parameterized filterbanks have been proposed for speaker recognition where only center f…

Speaker RecognitionSpeech Separation

Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation

2022-07-02 · Vikramjit Mitra, Hsiang-Yun Sherry Chien, Vasudha Kowtha, Joseph Yitan Cheng 외

Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech…

Knowledge DistillationMulti-Task LearningValence Estimation

Learnable Frontends that do not Learn: Quantifying Sensitivity to Filterbank Initialisation

2023-02-20 · Mark Anderson, Tomi Kinnunen, Naomi Harte

While much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In th…

Action DetectionActivity DetectionSensitivity

Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

2026-04-25 · Khalid Zaman, Masashi Unoki arxiv

Human-imitated speech poses a greater challenge than AI-generated speech for both human listeners and automatic detection systems. Unlike AI-generated speech, which often contains artifacts, over-smoothed spectra, or rob…

DeepVOX: Discovering Features from Raw Audio for Speaker Recognition in Non-ideal Audio Signals

2020-08-26 · Anurag Chowdhury, Arun Ross

Automatic speaker recognition algorithms typically use pre-defined filterbanks, such as Mel-Frequency and Gammatone filterbanks, for characterizing speech audio. However, it has been observed that the features extracted …

Speaker RecognitionTriplet