Towards Transferable Speech Emotion Representation: On loss functions for cross-lingual latent representations
In recent years, speech emotion recognition (SER) has been used in wide ranging applications, from healthcare to the commercial sector. In addition to signal processing approaches, methods for SER now also use deep learning techniques which provide transfer learning possibilities. However, generalizing over languages, corpora and recording conditions is still an open challenge. In this work we address this gap by exploring loss functions that aid in transferability, specifically to non-tonal languages. We propose a variational autoencoder (VAE) with KL annealing and a semi-supervised VAE to obtain more consistent latent embedding distributions across data sets. To ensure transferability, the distribution of the latent embedding should be similar across non-tonal languages (data sets). We start by presenting a low-complexity SER based on a denoising-autoencoder, which achieves an unweighted classification accuracy of over 52.09% for four-class emotion classification. This performance is comparable to that of similar baseline methods. Following this, we employ a VAE, the semi-supervised VAE and the VAE with KL annealing to obtain a more regularized latent space. We show that while the DAE has the highest classification accuracy among the methods, the semi-supervised VAE has a comparable classification accuracy and a more consistent latent embedding distribution over data sets.
Code (0)
등록된 구현이 없습니다.
Tasks
ClassificationDenoisingEmotion ClassificationEmotion RecognitionSpeech Emotion RecognitionTransfer LearningSimilar Papers 제목 키워드 기반
Continuous Metric Learning For Transferable Speech Emotion Recognition and Embedding Across Low-resource Languages
Speech emotion recognition~(SER) refers to the technique of inferring the emotional state of an individual from speech signals. SERs continue to garner interest due to their wide applicability. Although the domain is mai…
DenoisingEmotion ClassificationEmotion RecognitionMetric Learning+1Learning Transferable Features for Speech Emotion Recognition
Emotion recognition from speech is one of the key steps towards emotional intelligence in advanced human-machine interaction. Identifying emotions in human speech requires learning features that are robust and discrimina…
Domain AdaptationEmotional IntelligenceEmotion RecognitionSpeech Emotion RecognitionEvaluation of Error and Correlation-Based Loss Functions For Multitask Learning Dimensional Speech Emotion Recognition
The choice of a loss function is a critical part in machine learning. This paper evaluated two different loss functions commonly used in regression-task dimensional speech emotion recognition, an error-based and a correl…
Emotion RecognitionregressionSpeech Emotion RecognitionModality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition
Despite the recent achievements made in the multi-modal emotion recognition task, two problems still exist and have not been well investigated: 1) the relationship between different emotion categories are not utilized, w…
Emotion RecognitionMultimodal Emotion RecognitionWord EmbeddingsABHINAYA -- A System for Speech Emotion Recognition In Naturalistic Conditions Challenge
Speech emotion recognition (SER) in naturalistic settings remains a challenge due to the intrinsic variability, diverse recording conditions, and class imbalance. As participants in the Interspeech Naturalistic SER Chall…
Emotion RecognitionSpeech Emotion Recognition