paper-with-me

홈 › Papers

ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations

2026-04-28 · Kexue Wang, Yinfeng Yu, Liejun Wang arxiv

To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means that different people may express emotions differently. In our daily lives, we can see this. When communicating with different people, some express "happiness" through their facial expressions and words, while others may hide their happiness or express it through their actions. Both are expressions of 'happiness,' but such differences in emotional expression are still too difficult for machines to distinguish. Current emotion recognition remains at a 'static' level, using a single recognition model to identify all emotional styles. This "simplification" often affects the recognition results, especially in multi-turn dialogues. To address this problem, this paper introduces a novel Multi-Level Speaker Adaptive Network (ML-SAN), which, specifically, effectively addresses the challenge of speaker identity information confusion. ML-SAN does not simply assign a speaker's ID after recognition; instead, it employs a three-stage adaptive process: First, Input-level Calibration uses Feature-Level Linear Modulation (FiLM) to adjust the raw audio and visual features into a neutral space unrelated to the speaker. Then, Interaction-level Gating re-adjusts the trust level for each modality (e.g., voice or facial features) based on the speaker's identity information. Finally, Output-level Regularization maintains the consistency of speaker features in the latent space. Tests on the MELD and IEMOCAP datasets show that our model (ML-SAN) achieves better results, performs exceptionally well in handling challenging tail sentiment categories, and better addresses the diversity of speakers in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2604.25383

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal Emotion Recognition

Similar Papers 제목 키워드 기반

Comparison of Gender- and Speaker-adaptive Emotion Recognition

2014-05-01 · LREC 2014 5 · Maxim Sidorov, Stefan Ultes, Alex Schmitt, er

Deriving the emotion of a human speaker is a hard task, especially if only the audio stream is taken into account. While state-of-the-art approaches already provide good results, adaptive methods have been proposed in or…

AttributeEmotion ClassificationEmotion RecognitionSpeaker Identification+1

Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition

2025-01-25 · Junwei Feng, Xueyan Fan

Emotion recognition has a wide range of applications in human-computer interaction, marketing, healthcare, and other fields. In recent years, the development of deep learning technology has provided new methods for emoti…

cross-modal alignmentEmotion ClassificationEmotion RecognitionMarketing+1

AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition

2026-03-07 · Yunsheng Wang, Yuntao Shou, Yilong Tan, Wei Ai 외 arxiv

Multimodal dialogue emotion recognition captures emotional cues by fusing text, visual, and audio modalities. However, existing approaches still suffer from notable limitations in modeling emotional dependencies and lear…

Multimodal Emotion Recognition

DNN-HMM based Speaker Adaptive Emotion Recognition using Proposed Epoch and MFCC Features

2018-06-04 · Md. Shah Fahad, Jainath Yadav, Gyadhar Pradhan, Akshay Deepak

Speech is produced when time varying vocal tract system is excited with time varying excitation source. Therefore, the information present in a speech such as message, emotion, language, speaker is due to the combined ef…

Emotion Recognition

AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations

2024-01-26 · Naresh Kumar Devulapally, Sidharth Anand, Sreyasee Das Bhattacharjee, Junsong Yuan 외

Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modaliti…

Emotion Recognition