Papers Situation Recognition
“Situation Recognition” 태그가 달린 논문 15편 · 필터 해제
One Identity, Many Roles: Multimodal Entity Coreference for Enhanced Video Situation Recognition
Video Situation Recognition (VidSitu) addresses the challenging problem of "who did what to whom, with what, how, and where" in a video. It tests thorough video understanding by requiring identification of salient action…
Situation RecognitionVisual GroundingKIRETT -- A wearable device to support rescue operations using artificial intelligence to improve first aid
This short paper presents first steps in the scientific part of the KIRETT project, which aims to improve first aid during rescue operations using a wearable device. The wearable is used for computer-aided situation reco…
Situation RecognitionThe Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
Context recognition (SR) is a fundamental task in computer vision that aims to extract structured semantic summaries from images by identifying key events and their associated entities. Specifically, given an input image…
Semantic Role LabelingSituation RecognitionMulti-Label LearningDynamic Scene Understanding from Vision-Language Representations
Images depicting complex, dynamic scenes are challenging to parse automatically, requiring both high-level comprehension of the overall situation and fine-grained identification of participating entities and their intera…
Grounded Situation RecognitionHuman-Human Interaction RecognitionHuman Interaction RecognitionHuman-Object Interaction Detection+2ClipSitu: Effectively Leveraging CLIP for Conditional Predictions in Situation Recognition
Situation Recognition is the task of generating a structured summary of what is happening in an image using an activity verb and the semantic roles played by actors and objects. In this task, the same activity verb can d…
Grounded Situation RecognitionSituation RecognitionCollaborative Transformers for Grounded Situation Recognition
Grounded situation recognition is the task of predicting the main activity, entities playing certain roles within the activity, and bounding-box groundings of the entities in the given image. To effectively deal with thi…
Grounded Situation RecognitionImage ClassificationObject DetectionScene Understanding+3Rethinking the Two-Stage Framework for Grounded Situation Recognition
Grounded Situation Recognition (GSR), i.e., recognizing the salient activity (or verb) category in an image (e.g., buying) and detecting all corresponding semantic roles (e.g., agent and goods), is an essential step towa…
Grounded Situation RecognitionObject RecognitionSituation RecognitionTriplet+1Grounded Situation Recognition with Transformers
Grounded Situation Recognition (GSR) is the task that not only classifies a salient action (verb), but also predicts entities (nouns) associated with semantic roles and their locations in the given image. Inspired by the…
DecoderGrounded Situation RecognitionImage ClassificationObject Detection+4Attention-Based Context Aware Reasoning for Situation Recognition
Situation Recognition (SR) is a fine-grained action recognition task where the model is expected to not only predict the salient action of the image, but also predict values of all associated semantic roles of the action…
Action RecognitionFine-grained Action RecognitionGrounded Situation RecognitionQuestion Answering+4Grounded Situation Recognition
We introduce Grounded Situation Recognition (GSR), a task that requires producing structured semantic summaries of images describing: the primary activity, entities engaged in the activity with their roles (e.g. agent, t…
Grounded Situation RecognitionImage RetrievalRetrievalSituation RecognitionMixture-Kernel Graph Attention Network for Situation Recognition
Understanding images beyond salient actions involves reasoning about scene context, objects, and the roles they play in the captured event. Situation recognition has recently been introduced as the task of jointly reason…
Graph AttentionGraph Neural NetworkGrounded Situation RecognitionSituation RecognitionSituation Recognition with Graph Neural Networks
We address the problem of recognizing situations in images. Given an image, the task is to predict the most salient verb (action), and fill its semantic roles such as who is performing the action, what is the source and …
Grounded Situation RecognitionSituation RecognitionRecurrent Models for Situation Recognition
This work proposes Recurrent Neural Network (RNN) models to predict structured 'image situations' -- actions and noun entities fulfilling semantic roles related to the action. In contrast to prior work relying on Conditi…
Grounded Situation RecognitionHuman-Object Interaction DetectionImage CaptioningPrediction+1Commonly Uncommon: Semantic Sparsity in Situation Recognition
Semantic sparsity is a common challenge in structured visual classification problems; when the output space is complex, the vast majority of the possible predictions are rarely, if ever, seen in the training set. This pa…
Grounded Situation RecognitionSituation RecognitionStructured PredictionSituation Recognition: Visual Semantic Role Labeling for Image Understanding
This paper introduces situation recognition, the problem of producing a concise summary of the situation an image depicts including: (1) the main activity (e.g., clipping), (2) the participating actors, objects, substanc…
Activity RecognitionGrounded Situation RecognitionSemantic Role LabelingSituation Recognition+1