paper-with-me

홈 › Papers

Multiscale Contextual Learning for Speech Emotion Recognition in Emergency Call Center Conversations

2023-08-28 · Théo Deschamps-Berger, Lori Lamel, Laurence Devillers

Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcity of emotion datasets collected in the wild and the inability to take into account the dialogue context. The CEMO dataset, composed of conversations between agents and patients during emergency calls to a French call center, fills this gap. The nature of these interactions highlights the role of the emotional flow of the conversation in predicting patient emotions, as context can often make a difference in understanding actual feelings. This paper presents a multi-scale conversational context learning approach for speech emotion recognition, which takes advantage of this hypothesis. We investigated this approach on both speech transcriptions and acoustic segments. Experimentally, our method uses the previous or next information of the targeted segment. In the text domain, we tested the context window using a wide range of tokens (from 10 to 100) and at the speech turns level, considering inputs from both the same and opposing speakers. According to our tests, the context derived from previous tokens has a more significant influence on accurate prediction than the following tokens. Furthermore, taking the last speech turn of the same speaker in the conversation seems useful. In the acoustic domain, we conducted an in-depth analysis of the impact of the surrounding emotions on the prediction. While multi-scale conversational context learning using Transformers can enhance performance in the textual modality for emergency call recordings, incorporating acoustic context is more challenging.

📄 PDF Abstract BibTeX arXiv:2308.14894

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

End-to-End Speech Emotion Recognition: Challenges of Real-Life Emergency Call Centers Data Recordings

2021-10-28 · Théo Deschamps-Berger, Lori Lamel, Laurence Devillers

Recognizing a speaker's emotion from their speech can be a key element in emergency call centers. End-to-end deep learning systems for speech emotion recognition now achieve equivalent or even better results than convent…

Deep LearningEmotion RecognitionSpeech Emotion Recognition

Speech Emotion Recognition with Multiscale Area Attention and Data Augmentation

2021-02-03 · Mingke Xu, Fan Zhang, Xiaodong Cui, Wei zhang

In Speech Emotion Recognition (SER), emotional characteristics often appear in diverse forms of energy patterns in spectrograms. Typical attention neural network classifiers of SER are usually optimized on a fixed attent…

Data AugmentationEmotion RecognitionSpeech Emotion Recognition

Exploring Attention Mechanisms for Multimodal Emotion Recognition in an Emergency Call Center Corpus

2023-06-12 · Théo Deschamps-Berger, Lori Lamel, Laurence Devillers

The emotion detection technology to enhance human decision-making is an important research issue for real-world applications, but real-life emotion datasets are relatively rare and small. The experiments conducted in thi…

Decision MakingEmotion RecognitionMultimodal Emotion RecognitionSpeech Emotion Recognition

Dynamic Parameter Memory: Temporary LoRA-Enhanced LLM for Long-Sequence Emotion Recognition in Conversation

2025-07-11 · Jialong Mai, Xiaofen Xing, Yawei Li, Zhipeng Li 외

Recent research has focused on applying speech large language model (SLLM) to improve speech emotion recognition (SER). However, the inherently high frame rate in speech modality severely limits the signal processing and…

4kEmotion RecognitionEmotion Recognition in ConversationLarge Language Model+2

Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought

2025-02-25 · Zhixian Zhao, Xinfa Zhu, Xinsheng Wang, Shuiyuan Wang 외

Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), …

Emotion RecognitionLanguage ModelingLanguage ModellingSpeech Emotion Recognition