paper-with-me

홈 › Papers

SERAB: A multi-lingual benchmark for speech emotion recognition

2021-10-07 · Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann, Milos Cernak

Recent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation protocols. To facilitate the process, here, we present the Speech Emotion Recognition Adaptation Benchmark (SERAB), a framework for evaluating the performance and generalization capacity of different approaches for utterance-level SER. The benchmark is composed of nine datasets for SER in six languages. Since the datasets have different sizes and numbers of emotional classes, the proposed setup is particularly suitable for estimating the generalization capacity of pre-trained DNN-based feature extractors. We used the proposed framework to evaluate a selection of standard hand-crafted feature sets and state-of-the-art DNN representations. The results highlight that using only a subset of the data included in SERAB can result in biased evaluation, while compliance with the proposed protocol can circumvent this issue.

📄 PDF Abstract BibTeX arXiv:2110.03414

Code (2)

neclow/serab 공식 구현 tf
gasserelbanna/serab-byols pytorch

Tasks

BenchmarkingEmotion RecognitionSpeech Emotion Recognition

Methods 이 논문이 사용한 방법론

qb Experts Team 설명 없음

Similar Papers 제목 키워드 기반

CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition

2023-10-18 · Kari A Noriy, Xiaosong Yang, Marcin Budka, Jian Jun Zhang

Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization across languages. It excels at fosterin…

Audio ClassificationContrastive LearningCross-Lingual TransferData Augmentation+6

Do Speech Emphasis Models Generalize across Languages and Emotions?

2026-06-26 · Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang 외 arxiv

Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech. We introduce MMEE (Multilingual Multi…

Large Language Models Meet Contrastive Learning: Zero-Shot Emotion Recognition Across Languages

2025-03-25 · Heqing Zou, Fengmao Lv, Desheng Zheng, Eng Siong Chng 외

Multilingual speech emotion recognition aims to estimate a speaker's emotional state using a contactless method across different languages. However, variability in voice characteristics and linguistic diversity poses sig…

Contrastive LearningDiversityEmotion RecognitionSpeech Emotion Recognition

METTS: Multilingual Emotional Text-to-Speech by Cross-speaker and Cross-lingual Emotion Transfer

2023-07-29 · Xinfa Zhu, Yi Lei, Tao Li, Yongmao Zhang 외

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotio…

DisentanglementDiversityQuantizationSpeech Synthesis+2

CAMEO: Collection of Multilingual Emotional Speech Corpora

2025-05-16 · Iwona Christop, Maciej Czajka

This paper presents CAMEO -- a curated collection of multilingual emotional speech datasets designed to facilitate research in emotion recognition and other speech-related tasks. The main objectives were to ensure easy a…

Emotion RecognitionSpeech Emotion Recognition