paper-with-me

홈 › Papers

VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs

2026-03-09 · Hezhao Zhang, Huang-Cheng Chou, Shrikanth Narayanan, Thomas Hain arxiv

Speech Large Language Models (LLMs) show great promise for speech emotion recognition (SER) via generative interfaces. However, shifting from closed-set classification to open text generation introduces zero-shot stochasticity, making evaluation highly sensitive to prompts. Additionally, conventional speech LLMs benchmarks overlook the inherent ambiguity of human emotion. Hence, we present VoxEmo, a comprehensive SER benchmark encompassing 35 emotion corpora across 15 languages for Speech LLMs. VoxEmo provides a standardized toolkit featuring varying prompt complexities, from direct classification to paralinguistic reasoning. To reflect real-world perception/application, we introduce a distribution-aware soft-label protocol and a prompt-ensemble strategy that emulates annotator disagreement. Experiments reveal that while zero-shot speech LLMs trail supervised baselines in hard-label accuracy, they uniquely align with human subjective distributions.

📄 PDF Abstract BibTeX arXiv:2603.08936

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Emotion RecognitionText Generation

Similar Papers 제목 키워드 기반

Seeing What You Say: Expressive Image Generation from Speech

2025-11-05 · Jiyoung Lee, Song Park, Sanghyuk Chun, Soo-Whan Chung arxiv

This paper proposes VoxStudio, the first unified and end-to-end speech-to-image model that generates expressive images directly from spoken descriptions by jointly aligning linguistic and paralinguistic information. At i…

Image Generation

INTERSPEECH 2009 Emotion Challenge Revisited: Benchmarking 15 Years of Progress in Speech Emotion Recognition

2024-06-10 · Andreas Triantafyllopoulos, Anton Batliner, Simon Rampp, Manuel Milling 외

We revisit the INTERSPEECH 2009 Emotion Challenge -- the first ever speech emotion recognition (SER) challenge -- and evaluate a series of deep learning models that are representative of the major advances in SER researc…

BenchmarkingEmotion RecognitionSpeech Emotion Recognition

SERAB: A multi-lingual benchmark for speech emotion recognition

2021-10-07 · Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann, Milos Cernak

Recent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation…

BenchmarkingEmotion RecognitionSpeech Emotion Recognition

SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition

2024-08-14 · Mohamed Osman, Daniel Z. Kaplan, Tamer Nadeem

Speech emotion recognition (SER) has made significant strides with the advent of powerful self-supervised learning (SSL) models. However, the generalization of these models to diverse languages and emotional expressions …

Automatic Speech RecognitionBenchmarkingEmotion RecognitionSelf-Supervised Learning+3

Speech Emotion Recognition Using Speech Feature and Word Embedding

2019-11-18 · APSIPA ASC 2019 11 · Bagus Tris Atmaja, Kiyoaki Shirai, and Masato Akagi

—Emotion recognition can be performed automatically from many modalities. This paper presents a categorical speech emotion recognition using speech features and word embedding. Text features can be combined with speech f…

Emotion RecognitionSpeech Emotion Recognition