paper-with-me

Papers

Decoding Emotions: A comprehensive Multilingual Study of Speech Models for Speech Emotion Recognition

2023-08-17 · Anant Singh, Akshat Gupta

Recent advancements in transformer-based speech representation models have greatly transformed speech processing. However, there has been limited research conducted on evaluating these models for speech emotion recognition (SER) across multiple languages and examining their internal representations. This article addresses these gaps by presenting a comprehensive benchmark for SER with eight speech representation models and six different languages. We conducted probing experiments to gain insights into inner workings of these models for SER. We find that using features from a single optimal layer of a speech model reduces the error rate by 32\% on average across seven datasets when compared to systems where features from all layers of speech models are used. We also achieve state-of-the-art results for German and Persian languages. Our probing results indicate that the middle layers of speech models capture the most important emotional information for speech emotion recognition.

📄 PDF Abstract BibTeX arXiv:2308.08713

Code (1)

95anantsingh/decoding-emotions 공식 구현 pytorch

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

CLARA: Multilingual Contrastive Learning for Audio Representation Acquisition

2023-10-18 · Kari A Noriy, Xiaosong Yang, Marcin Budka, Jian Jun Zhang

Multilingual speech processing requires understanding emotions, a task made difficult by limited labelled data. CLARA, minimizes reliance on labelled data, enhancing generalization across languages. It excels at fosterin…

Audio ClassificationContrastive LearningCross-Lingual TransferData Augmentation+6

Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study

2024-01-23 · W. Ronny Huang, Cyril Allauzen, Tongzhou Chen, Kilol Gupta 외

In the era of large models, the autoregressive nature of decoding often results in latency serving as a significant bottleneck. We propose a non-autoregressive LM-fused ASR system that effectively leverages the paralleli…

Language ModelingLanguage ModellingLarge Language Modelspeech-recognition+1

Do Speech Emphasis Models Generalize across Languages and Emotions?

2026-06-26 · Megan Wei, Deepali Aneja, Jiaqi Su, Yunyun Wang 외 arxiv

Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and evaluated on monolingual neutral read speech. We introduce MMEE (Multilingual Multi…

Towards Generalizable SER: Soft Labeling and Data Augmentation for Modeling Temporal Emotion Shifts in Large-Scale Multilingual Speech

2023-11-15 · Mohamed Osman, Tamer Nadeem, Ghada Khoriba

Recognizing emotions in spoken communication is crucial for advanced human-machine interaction. Current emotion detection methodologies often display biases when applied cross-corpus. To address this, our study amalgamat…

Contrastive LearningCross-corpusData AugmentationZero-shot Generalization

Bi-directional Context-Enhanced Speech Large Language Models for Multilingual Conversational ASR

2025-06-16 · Yizhou Peng, Hexin Liu, Eng Siong Chng

This paper introduces the integration of language-specific bi-directional context into a speech large language model (SLLM) to improve multilingual continuous conversational automatic speech recognition (ASR). We propose…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+3