paper-with-me

홈 › Papers

EMO-SUPERB: An In-depth Look at Speech Emotion Recognition

2024-02-20 · Haibin Wu, Huang-Cheng Chou, Kai-Wei Chang, Lucas Goncalves, Jiawei Du, Jyh-Shing Roger Jang, Chi-Chun Lee, Hung-Yi Lee

Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems. However, 80.77% of SER papers yield results that cannot be reproduced. We develop EMO-SUPERB, short for EMOtion Speech Universal PERformance Benchmark, which aims to enhance open-source initiatives for SER. EMO-SUPERB includes a user-friendly codebase to leverage 15 state-of-the-art speech self-supervised learning models (SSLMs) for exhaustive evaluation across six open-source SER datasets. EMO-SUPERB streamlines result sharing via an online leaderboard, fostering collaboration within a community-driven benchmark and thereby enhancing the development of SER. On average, 2.58% of annotations are annotated using natural language. SER relies on classification models and is unable to process natural languages, leading to the discarding of these valuable annotations. We prompt ChatGPT to mimic annotators, comprehend natural language annotations, and subsequently re-label the data. By utilizing labels generated by ChatGPT, we consistently achieve an average relative gain of 3.08% across all settings.

📄 PDF Abstract BibTeX arXiv:2402.13018

Code (1)

EMOsuperb/EMO-SUPERB-submission 공식 구현 pytorch

Tasks

Emotion RecognitionSelf-Supervised LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?

2024-02-02 · Orchid Chetia Phukan, Gautam Siddharth Kashyap, Arun Balaji Buduru, Rajesh Sharma

Availability of representations from pre-trained models (PTMs) have facilitated substantial progress in speech emotion recognition (SER). Particularly, representations from PTM trained for paralinguistic speech processin…

AllEmotion RecognitionSpeech Emotion Recognition

Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

2024-11-08 · Chien-yu Huang, Wei-Chih Chen, Shu-wen Yang, Andy T. Liu 외

Multimodal foundation models, such as Gemini and ChatGPT, have revolutionized human-machine interactions by seamlessly integrating various forms of data. Developing a universal spoken language model that comprehends a wi…

Emotion Recognition

EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations

2024-07-22 · Wenze Ren, Yi-Cheng Lin, Huang-Cheng Chou, Haibin Wu 외

The neural codec model reduces speech data transmission delay and serves as the foundational tokenizer for speech language models (speech LMs). Preserving emotional information in codecs is crucial for effective communic…

Emotion RecognitionSpeech Emotion Recognition

ML-SUPERB: Multilingual Speech Universal PERformance Benchmark

2023-05-18 · Jiatong Shi, Dan Berrebbi, William Chen, Ho-Lam Chung 외

Speech processing Universal PERformance Benchmark (SUPERB) is a leaderboard to benchmark the performance of Self-Supervised Learning (SSL) models on various speech processing tasks. However, SUPERB largely considers Engl…

Automatic Speech RecognitionLanguage IdentificationSelf-Supervised Learningspeech-recognition+1

Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond

2023-10-09 · Jiatong Shi, William Chen, Dan Berrebbi, Hsiu-Hsuan Wang 외

The 2023 Multilingual Speech Universal Performance Benchmark (ML-SUPERB) Challenge expands upon the acclaimed SUPERB framework, emphasizing self-supervised models in multilingual speech recognition and language identific…

Language Identificationspeech-recognitionSpeech Recognition