paper-with-me

Papers

Contextually-rich human affect perception using multimodal scene information

2023-03-13 · Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shrikanth Narayanan

The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focused on expressions extracted from salient face crops. However, emotions perceived by humans rely on multiple contextual cues including social settings, foreground interactions, and ambient visual scenes. In this work, we leverage pretrained vision-language (VLN) models to extract descriptions of foreground context from images. Further, we propose a multimodal context fusion (MCF) module to combine foreground cues with the visual scene and person-based contextual information for emotion prediction. We show the effectiveness of our proposed modular design on two datasets associated with natural scenes and TV shows.

📄 PDF Abstract BibTeX arXiv:2303.06904

Code (1)

usc-sail/mica-context-emotion-recognition 공식 구현 pytorch

Similar Papers 제목 키워드 기반

The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents

2023-09-26 · Che-Jui Chang, Samuel S. Sohn, Sen Zhang, Rajath Jayashankar 외

Previous studies regarding the perception of emotions for embodied virtual agents have shown the effectiveness of using virtual characters in conveying emotions through interactions with humans. However, creating an auto…

MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception

2024-06-22 · Guanqun Wang, Xinyu Wei, Jiaming Liu, Ray Zhang 외

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant stride…

Common Sense ReasoningLanguage ModellingLarge Language ModelMultimodal Large Language Model+4

A Novel Context-Aware Multimodal Framework for Persian Sentiment Analysis

2021-03-03 · Kia Dashtipour, Mandar Gogate, Erik Cambria, Amir Hussain

Most recent works on sentiment analysis have exploited the text modality. However, millions of hours of video recordings posted on social media platforms everyday hold vital unstructured information that can be exploited…

Multimodal Sentiment AnalysisPersian Sentiment AnalysisSentiment Analysis

MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans

2024-09-30 · Anna Deichler, Jim O'Regan, Jonas Beskow

In this paper, we present a novel dataset captured using a VR headset to record conversations between participants within a physics simulator (AI2-THOR). Our primary objective is to extend the field of co-speech gesture …

Gesture Generation

LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?

2026-01-27 · Zhuang Yu, Lei Shen, Jing Zhao, Shiliang Sun arxiv

Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, and temporally structured educational cont…