Situation Recognition: Visual Semantic Role Labeling for Image Understanding
This paper introduces situation recognition, the problem of producing a concise summary of the situation an image depicts including: (1) the main activity (e.g., clipping), (2) the participating actors, objects, substances, and locations (e.g., man, shears, sheep, wool, and field) and most importantly (3) the roles these participants play in the activity (e.g., the man is clipping, the shears are his tool, the wool is being clipped from the sheep, and the clipping is in a field). We use FrameNet, a verb and role lexicon developed by linguists, to define a large space of possible situations and collect a large-scale dataset containing over 500 activities, 1,700 roles, 11,000 objects, 125,000 images, and 200,000 unique situations. We also introduce structured prediction baselines and show that, in activity-centric images, situation-driven prediction of objects and activities outperforms independent object and activity recognition.
Code (1)
Tasks
Activity RecognitionGrounded Situation RecognitionSemantic Role LabelingSituation RecognitionStructured PredictionSimilar Papers 제목 키워드 기반
Grounding Semantic Roles in Images
We address the task of visual semantic role labeling (vSRL), the identification of the participants of a situation or event in a visual scene, and their labeling with their semantic relations to the event or situation. W…
Image CaptioningQuestion AnsweringSemantic Role LabelingEffectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Situation recognition refers to the ability of an agent to identify and understand various situations or contexts based on available information and sensory inputs. It involves the cognitive process of interpreting data …
Semantic Role LabelingVideo CaptioningClipSitu: Effectively Leveraging CLIP for Conditional Predictions in Situation Recognition
Situation Recognition is the task of generating a structured summary of what is happening in an image using an activity verb and the semantic roles played by actors and objects. In this task, the same activity verb can d…
Grounded Situation RecognitionSituation RecognitionAttention-Based Context Aware Reasoning for Situation Recognition
Situation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action…
Action RecognitionFine-grained Action RecognitionGrounded Situation RecognitionQuestion Answering+4Mixture-Kernel Graph Attention Network for Situation Recognition
Understanding images beyond salient actions involves reasoning about scene context, objects, and the roles they play in the captured event. Situation recognition has recently been introduced as the task of jointly reason…
Graph AttentionGraph Neural NetworkGrounded Situation RecognitionSituation Recognition