SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge. We propose a large-scale benchmark to evaluate the robustness and adaptability of state-of-the-art SER models in both in-domain and out-of-domain settings. Our benchmark includes a diverse set of multilingual datasets, focusing on less commonly used corpora to assess generalization to new data. We employ logit adjustment to account for varying class distributions and establish a single dataset cluster for systematic evaluation. Surprisingly, we find that the Whisper model, primarily designed for automatic speech recognition, outperforms dedicated SSL models in cross-lingual SER. Our results highlight the need for more robust and generalizable SER models, and our benchmark serves as a valuable resource to drive future research in this direction.
Code (1)
Tasks
Automatic Speech RecognitionBenchmarkingEmotion RecognitionSelf-Supervised LearningSpeech Emotion Recognitionspeech-recognitionSpeech RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning Transferable Features for Speech Emotion Recognition
Emotion recognition from speech is one of the key steps towards emotional intelligence in advanced human-machine interaction. Identifying emotions in human speech requires learning features that are robust and discrimina…
Domain AdaptationEmotional IntelligenceEmotion RecognitionSpeech Emotion RecognitionEmotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition
Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach …
Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesistext-to-speech+1Improving Cross-Domain Hate Speech Generalizability with Emotion Knowledge
Reliable automatic hate speech (HS) detection systems must adapt to the in-flow of diverse new data to curtail hate speech. However, hate speech detection systems commonly lack generalizability in identifying hate speech…
Hate Speech DetectionExploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…
Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning
Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector M…
AttributeEmotion ClassificationMulti-Task Learningtext-to-speech+2