paper-with-me

홈 › Papers

MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition

2025-06-24 · Hyo Jin Jon, Longbin Jin, Hyuntaek Jung, Hyunseo Kim, Donghun Min, Eun Yi Kim

This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural speech, including intra- and inter-subject variability, we propose Multi-level Acoustic-Textual Emotion Representation (MATER), a novel hierarchical framework that integrates acoustic and textual features at the word, utterance, and embedding levels. By fusing low-level lexical and acoustic cues with high-level contextualized representations, MATER effectively captures both fine-grained prosodic variations and semantic nuances. Additionally, we introduce an uncertainty-aware ensemble strategy to mitigate annotator inconsistencies, improving robustness in ambiguous emotional expressions. MATER ranks fourth in both tasks with a Macro-F1 of 41.01% and an average CCC of 0.5928, securing second place in valence prediction with an impressive CCC of 0.6941.

📄 PDF Abstract BibTeX arXiv:2506.19887

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeEmotion RecognitionSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling

2026-08-14 · Xueqi Wang, Zhigang Wang, Runqing Zhang, Zhenqi Jia 외 arxiv

Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level ret…

Emotion Recognition in ConversationContrastive LearningSpeech Synthesis

Multiscale Contextual Learning for Speech Emotion Recognition in Emergency Call Center Conversations

2023-08-28 · Théo Deschamps-Berger, Lori Lamel, Laurence Devillers

Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcit…

Emotion RecognitionSpeech Emotion Recognition

Detecting Emotion Carriers by Combining Acoustic and Lexical Representations

2021-12-13 · Sebastian P. Bayerl, Aniruddha Tammewar, Korbinian Riedhammer, Giuseppe Riccardi

Personal narratives (PN) - spoken or written - are recollections of facts, people, events, and thoughts from one's own experience. Emotion recognition and sentiment analysis tasks are usually defined at the utterance or …

Emotion RecognitionNatural Language UnderstandingSentiment Analysis

CFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition

2023-07-28 · Jiang Li, XiaoPing Wang, Yingjian Liu, Zhigang Zeng

Multimodal emotion recognition in conversation (ERC) has garnered growing attention from research communities in various fields. In this paper, we propose a Cross-modal Fusion Network with Emotion-Shift Awareness (CFN-ES…

Emotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognitionmultimodal interaction

DEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models

2026-03-17 · Jiaqi Xiong, Yunjia Qi, Qi Cao, Yu Zheng 외 arxiv

Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based se…