Contextual Emotion Recognition using Large Vision Language Models
"How does the person in the bounding box feel?" Achieving human-level recognition of the apparent emotion of a person in real world situations remains an unsolved task in computer vision. Facial expressions are not enough: body pose, contextual knowledge, and commonsense reasoning all contribute to how humans perform this emotional theory of mind task. In this paper, we examine two major approaches enabled by recent large vision language models: 1) image captioning followed by a language-only LLM, and 2) vision language models, under zero-shot and fine-tuned setups. We evaluate the methods on the Emotions in Context (EMOTIC) dataset and demonstrate that a vision language model, fine-tuned even on a small dataset, can significantly outperform traditional baselines. The results of this work aim to help robots and agents perform emotionally sensitive decision-making and interaction in the future.
Code (0)
등록된 구현이 없습니다.
Tasks
Decision MakingEmotion RecognitionImage CaptioningLanguage ModelingLanguage ModellingSimilar Papers 제목 키워드 기반
Visual and textual prompts for enhancing emotion recognition in video
Vision Large Language Models (VLLMs) exhibit promising potential for multi-modal understanding, yet their application to video-based emotion recognition remains limited by insufficient spatial and contextual awareness. T…
Emotion RecognitionVideo Emotion RecognitionVisual PromptingIn-Depth Analysis of Emotion Recognition through Knowledge-Based Large Language Models
Emotion recognition in social situations is a complex task that requires integrating information from both facial expressions and the situational context. While traditional approaches to automatic emotion recognition hav…
Emotion RecognitionSteering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
Large-scale audio language models (ALMs), such as Qwen2-Audio, are capable of comprehending diverse audio signal, performing audio analysis and generating textual responses. However, in speech emotion recognition (SER), …
Emotion RecognitionLanguage ModelingLanguage ModellingSpeech Emotion RecognitionMELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
Although speech emotion recognition (SER) has advanced significantly with deep learning, annotation remains a major hurdle. Human annotation is not only costly but also subject to inconsistencies annotators often have di…
Emotion RecognitionSelf-Supervised LearningSpeech Emotion RecognitionKorean Drama Scene Transcript Dataset for Emotion Recognition in Conversations
Understanding emotions in conversation is a challenging task, as the sentences often have an implied meaning that is not generally understood in isolation. Efficient use of contextual information is essential for emotion…
Emotion RecognitionEmotion Recognition in Conversation