VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.
Code (0)
등록된 구현이 없습니다.
Tasks
Speech Emotion RecognitionText GenerationSimilar Papers 제목 키워드 기반
Seeing What You Say: Expressive Image Generation from Speech
This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At i…
Image GenerationINTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition
We revisit the INTERSPEECH 2009 Emotion Challenge -- the first ever speech emotion recognition (SER) challenge -- and evaluate a series of deep learning models that are representative of the major advances in SER researc…
BenchmarkingEmotion RecognitionSpeech Emotion RecognitionSERAB: A multi-lingual benchmark for speech emotion recognition
Recent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation…
BenchmarkingEmotion RecognitionSpeech Emotion RecognitionSER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions …
Automatic Speech RecognitionBenchmarkingEmotion RecognitionSelf-Supervised Learning+3Speech Emotion Recognition Using Speech Feature and Word Embedding
—Emotion recognition can be performed automatically from many modalities. This paper presents a categorical speech emotion recognition using speech features and word embedding. Text features can be combined with speech f…
Emotion RecognitionSpeech Emotion Recognition