MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
This paper presents our contributions to the Speech Emotion Recognition in Naturalistic Conditions (SERNC) Challenge, where we address categorical emotion recognition and emotional attribute prediction. To handle the complexities of natural speech, including intra- and inter-subject variability, we propose Multi-level Acoustic-Textual Emotion Representation (MATER), a novel hierarchical framework that integrates acoustic and textual features at the word, utterance, and embedding levels. By fusing low-level lexical and acoustic cues with high-level contextualized representations, MATER effectively captures both fine-grained prosodic variations and semantic nuances. Additionally, we introduce an uncertainty-aware ensemble strategy to mitigate annotator inconsistencies, improving robustness in ambiguous emotional expressions. MATER ranks fourth in both tasks with a Macro-F1 of 41.01% and an average CCC of 0.5928, securing second place in valence prediction with an impressive CCC of 0.6941.
Code (0)
등록된 구현이 없습니다.
Tasks
AttributeEmotion RecognitionSpeech Emotion RecognitionSimilar Papers 제목 키워드 기반
S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level ret…
Emotion Recognition in ConversationContrastive LearningSpeech SynthesisMultiscale Contextual Learning for Speech Emotion Recognition in Emergency Call Center Conversations
Emotion recognition in conversations is essential for ensuring advanced human-machine interactions. However, creating robust and accurate emotion recognition systems in real life is challenging, mainly due to the scarcit…
Emotion RecognitionSpeech Emotion RecognitionDetecting Emotion Carriers by Combining Acoustic and Lexical Representations
Personal narratives (PN) - spoken or written - are recollections of facts, people, events, and thoughts from one's own experience. Emotion recognition and sentiment analysis tasks are usually defined at the utterance or …
Emotion RecognitionNatural Language UnderstandingSentiment AnalysisCFN-ESA: A Cross-Modal Fusion Network with Emotion-Shift Awareness for Dialogue Emotion Recognition
Multimodal emotion recognition in conversation (ERC) has garnered growing attention from research communities in various fields. In this paper, we propose a Cross-modal Fusion Network with Emotion-Shift Awareness (CFN-ES…
Emotion RecognitionEmotion Recognition in ConversationMultimodal Emotion Recognitionmultimodal interactionDEAF: A Benchmark for Diagnostic Evaluation of Acoustic Faithfulness in Audio Language Models
Recent Audio Multimodal Large Language Models (Audio MLLMs) demonstrate impressive performance on speech benchmarks, yet it remains unclear whether these models genuinely process acoustic signals or rely on text-based se…