Enhancing Social Relation Inference with Concise Interaction Graph and Discriminative Scene Representation
There has been a recent surge of research interest in attacking the problem of social relation inference based on images. Existing works classify social relations mainly by creating complicated graphs of human interactions, or learning the foreground and/or background information of persons and objects, but ignore holistic scene context. The holistic scene refers to the functionality of a place in images, such as dinning room, playground and office. In this paper, by mimicking human understanding on images, we propose an approach of \textbf{PR}actical \textbf{I}nference in \textbf{S}ocial r\textbf{E}lation (PRISE), which concisely learns interactive features of persons and discriminative features of holistic scenes. Technically, we develop a simple and fast relational graph convolutional network to capture interactive features of all persons in one image. To learn the holistic scene feature, we elaborately design a contrastive learning task based on image scene classification. To further boost the performance in social relation inference, we collect and distribute a new large-scale dataset, which consists of about 240 thousand unlabeled images. The extensive experimental results show that our novel learning framework significantly beats the state-of-the-art methods, e.g., PRISE achieves 6.8$\%$ improvement for domain classification in PIPA dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Contrastive Learningdomain classificationRelationScene ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Technical Note: Defining and Quantifying AND-OR Interactions for Faithful and Concise Explanation of DNNs
In this technical note, we aim to explain a deep neural network (DNN) by quantifying the encoded interactions between input variables, which reflects the DNN's inference logic. Specifically, we first rethink the definiti…
Relational visual representations underlie human social interaction recognition
Humans effortlessly recognize social interactions from visual input. Attempts to model this ability have typically relied on generative inverse planning models, which make predictions by inverting a generative model of a…
Graph Neural Network"where is this relationship going?": Understanding Relationship Trajectories in Narrative Text
We examine a new commonsense reasoning task: given a narrative describing a social interaction that centers on two protagonists, systems make inferences about the underlying relationship trajectory. Specifically, we prop…
NavigatePrediction``where is this relationship going?'': Understanding Relationship Trajectories in Narrative Text
We examine a new commonsense reasoning task: given a narrative describing a social interaction that centers on two protagonists, systems make inferences about the underlying relationship trajectory. Specifically, we prop…
NavigatePredictionPLAYER*: Enhancing LLM-based Multi-Agent Communication and Interaction in Murder Mystery Games
We introduce WellPlay, a reasoning dataset for multi-agent conversational inference in Murder Mystery Games (MMGs). WellPlay comprises 1,482 inferential questions across 12 games, spanning objectives, reasoning, and rela…
Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+2