Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Consequently, they are limited in modeling the intricate dynamics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dynamics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to curate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Experiments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.io/projects/MMSI.
Code (0)
등록된 구현이 없습니다.
Tasks
coreference-resolutionCoreference ResolutionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
GAViD: A Large-Scale Multimodal Dataset for Context-Aware Group Affect Recognition from Videos
Understanding affective dynamics in real-world social systems is fundamental to modeling and analyzing human-human interactions in complex environments. Group affect emerges from intertwined human-human interactions, con…
Social Genome: Grounded Social Reasoning Abilities of Multimodal Models
Social reasoning abilities are crucial for AI systems to effectively interpret and respond to multimodal human communication and interaction within social contexts. We introduce SOCIAL GENOME, the first benchmark for fin…
Social-BiGAT: Multimodal Trajectory Forecasting using Bicycle-GAN and Graph Attention Networks
Predicting the future trajectories of multiple interacting agents in a scene has become an increasingly important problem for many different applications ranging from control of autonomous vehicles and social robots to s…
Autonomous VehiclesDecoderGenerative Adversarial NetworkGraph Attention+2Towards Multimodal Social Conversations with Robots: Using Vision-Language Models
Large language models have given social robots the ability to autonomously engage in open-domain conversations. However, they are still missing a fundamental social skill: making use of the multiple modalities that carry…
Face-to-Face Contrastive Learning for Social Intelligence Question-Answering
Creating artificial social intelligence - algorithms that can understand the nuances of multi-person interactions - is an exciting and emerging challenge in processing facial expressions and gestures from multimodal vide…
Contrastive LearningGraph Neural NetworkQuestion Answering