paper-with-me

홈 › Papers

SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition

2024-08-14 · Mohamed Osman, Daniel Z. Kaplan, Tamer Nadeem

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions remains a challenge. We propose a large-scale benchmark to evaluate the robustness and adaptability of state-of-the-art SER models in both in-domain and out-of-domain settings. Our benchmark includes a diverse set of multilingual datasets, focusing on less commonly used corpora to assess generalization to new data. We employ logit adjustment to account for varying class distributions and establish a single dataset cluster for systematic evaluation. Surprisingly, we find that the Whisper model, primarily designed for automatic speech recognition, outperforms dedicated SSL models in cross-lingual SER. Our results highlight the need for more robust and generalizable SER models, and our benchmark serves as a valuable resource to drive future research in this direction.

📄 PDF Abstract BibTeX arXiv:2408.07851

Code (1)

spaghettiSystems/serval 공식 구현 pytorch

Tasks

Automatic Speech RecognitionBenchmarkingEmotion RecognitionSelf-Supervised LearningSpeech Emotion Recognitionspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Learning Transferable Features for Speech Emotion Recognition

2019-12-23 · Alison Marczewski, Adriano Veloso, Nívio Ziviani

Emotion recognition from speech is one of the key steps towards emotional intelligence in advanced human-machine interaction. Identifying emotions in human speech requires learning features that are robust and discrimina…

Domain AdaptationEmotional IntelligenceEmotion RecognitionSpeech Emotion Recognition

Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

2020-10-26

Neural text-to-speech (TTS) approaches generally require a huge number of high quality speech data, which makes it difficult to obtain such a dataset with extra emotion labels. In this paper, we propose a novel approach …

Emotion RecognitionSpeech Emotion RecognitionSpeech Synthesistext-to-speech+1

Improving Cross-Domain Hate Speech Generalizability with Emotion Knowledge

2023-11-24 · Shi Yin Hong, Susan Gauch

Reliable automatic hate speech (HS) detection systems must adapt to the in-flow of diverse new data to curtail hate speech. However, hate speech detection systems commonly lack generalizability in identifying hate speech…

Hate Speech Detection

Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations

2024-09-26 · Yujia Sun, Zeyu Zhao, Korin Richmond, Yuanchao Li

Emotion recognition from speech and music shares similarities due to their acoustic overlap, which has led to interest in transferring knowledge between these domains. However, the shared acoustic cues between speech and…

Domain AdaptationDomain GeneralizationEmotion RecognitionMusic Emotion Recognition+3

Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep Learning

2022-06-15 · Rui Liu, Berrak Sisman, Björn Schuller, Guanglai Gao 외

Emotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector M…

AttributeEmotion ClassificationMulti-Task Learningtext-to-speech+2