Unsupervised low-rank representations for speech emotion recognition
We examine the use of linear and non-linear dimensionality reduction algorithms for extracting low-rank feature representations for speech emotion recognition. Two feature sets are used, one based on low-level descriptors and their aggregations (IS10) and one modeling recurrence dynamics of speech (RQA), as well as their fusion. We report speech emotion recognition (SER) results for learned representations on two databases using different classification methods. Classification with low-dimensional representations yields performance improvement in a variety of settings. This indicates that dimensionality reduction is an effective way to combat the curse of dimensionality for SER. Visualization of features in two dimensions provides insight into discriminatory abilities of reduced feature sets.
Code (0)
등록된 구현이 없습니다.
Tasks
Dimensionality ReductionEmotion RecognitionGeneral ClassificationSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition
There are individual differences in expressive behaviors driven by cultural norms and personality. This between-person variation can result in reduced emotion recognition performance. Therefore, personalization is an imp…
Emotion RecognitionSpeech Emotion RecognitionValence EstimationContrastive Unsupervised Learning for Speech Emotion Recognition
Speech emotion recognition (SER) is a key technology to enable more natural human-machine communication. However, SER has long suffered from a lack of public large-scale labeled datasets. To circumvent this problem, we i…
Emotion RecognitionRepresentation LearningSpeech Emotion RecognitionLeveraging Content and Acoustic Representations for Speech Emotion Recognition
Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. T…
Emotion RecognitionLanguage ModellingLarge Language ModelSpeech Emotion RecognitionDisentangling Prosody Representations with Unsupervised Speech Reconstruction
Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content a…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)DisentanglementEmotion Recognition+6Variational Autoencoders for Learning Latent Representations of Speech Emotion: A Preliminary Study
Learning the latent representation of data in unsupervised fashion is a very interesting process that provides relevant features for enhancing the performance of a classifier. For speech emotion recognition tasks, genera…
Emotion ClassificationEmotion RecognitionGeneral ClassificationSpeech Emotion Recognition