paper-with-me

홈 › Papers

MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge

2025-05-30 · Xin Jing, Jiadong Wang, Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller

Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have different preferences and may lack the necessary contextual knowledge, which can lead to varied and inaccurate labels. Meanwhile, Large Language Models (LLMs) have emerged as a scalable alternative for annotating text data. However, the potential of LLMs to perform emotional speech data annotation without human supervision has yet to be thoroughly investigated. To address these problems, we apply GPT-4o to annotate a multimodal dataset collected from the sitcom Friends, using only textual cues as inputs. By crafting structured text prompts, our methodology capitalizes on the knowledge GPT-4o has accumulated during its training, showcasing that it can generate accurate and contextually relevant annotations without direct access to multimodal inputs. Therefore, we propose MELT, a multimodal emotion dataset fully annotated by GPT-4o. We demonstrate the effectiveness of MELT by fine-tuning four self-supervised learning (SSL) backbones and assessing speech emotion recognition performance across emotion datasets. Additionally, our subjective experiments\' results demonstrate a consistence performance improvement on SER.

📄 PDF Abstract BibTeX arXiv:2505.24493

Code (1)

keikinn/meltdataset 공식 구현

Tasks

Emotion RecognitionSelf-Supervised LearningSpeech Emotion Recognition

Similar Papers 제목 키워드 기반

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

2025-05-21 · Cheng Yifan, Zhang Ruoyi, Shi Jiatong

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional …

Emotion RecognitionFace DetectionLanguage ModelingLanguage Modelling+6

MuSE: a Multimodal Dataset of Stressed Emotion

2020-05-01 · LREC 2020 5 · Mimansa Jaiswal, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo 외

Endowing automated agents with the ability to provide support, entertainment and interaction with human beings requires sensing of the users{'} affective state. These affective states are impacted by a combination of emo…

Emotion ClassificationGeneral Classification

MuSE-ing on the Impact of Utterance Ordering On Crowdsourced Emotion Annotations

2019-03-27 · Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo 외

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously d…

Emotion Recognition

Multimodal learning of melt pool dynamics in laser powder bed fusion

2025-09-03 · Satyajit Mojumder, Pallock Halder, Tiana Tonge arxiv

While multiple sensors are used for real-time monitoring in additive manufacturing, not all provide practical or reliable process insights. For example, high-speed X-ray imaging offers valuable spatial information about …

Transfer Learning

XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning

2026-02-05 · Hanwen Zhang, Yao Liu, Peiyuan Jiang, Lang Junjie 외 arxiv

Explainable Multimodal Emotion Recognition plays a crucial role in applications such as human-computer interaction and social media analytics. However, current approaches struggle with cue-level perception and reasoning …

Multimodal Emotion RecognitionSemantic Similarity