Situation Recognition with Graph Neural Networks
We address the problem of recognizing situations in images. Given an image, the task is to predict the most salient verb (action), and fill its semantic roles such as who is performing the action, what is the source and target of the action, etc. Different verbs have different roles (e.g. attacking has weapon), and each role can take on many possible values (nouns). We propose a model based on Graph Neural Networks that allows us to efficiently capture joint dependencies between roles using neural networks defined on a graph. Experiments with different graph connectivities show that our approach that propagates information between roles significantly outperforms existing work, as well as multiple baselines. We obtain roughly 3-5% improvement over previous work in predicting the full situation. We also provide a thorough qualitative analysis of our model and influence of different roles in the verbs.
Code (1)
Tasks
Grounded Situation RecognitionSituation RecognitionSimilar Papers 제목 키워드 기반
From Photo Streams to Evolving Situations
Photos are becoming spontaneous, objective, and universal sources of information. This paper develops evolving situation recognition using photo streams coming from disparate sources combined with the advances of deep le…
Mixture-Kernel Graph Attention Network for Situation Recognition
Understanding images beyond salient actions involves reasoning about scene context, objects, and the roles they play in the captured event. Situation recognition has recently been introduced as the task of jointly reason…
Graph AttentionGraph Neural NetworkGrounded Situation RecognitionSituation RecognitionKIRETT: Knowledge-Graph-Based Smart Treatment Assistant for Intelligent Rescue Operations
Over the years, the need for rescue operations throughout the world has increased rapidly. Demographic changes and the resulting risk of injury or health disorders form the basis for emergency calls. In such scenarios, f…
Toward more frugal models for functional cerebral networks automatic recognition with resting-state fMRI
We refer to a machine learning situation where models based on classical convolutional neural networks have shown good performance. We are investigating different encoding techniques in the form of supervoxels, then grap…
VideoScoop: A Non-Traditional Domain-Independent Framework For Video Analysis
Automatically understanding video contents is important for several applications in Civic Monitoring (CM), general Surveillance (SL), Assisted Living (AL), etc. Decades of Image and Video Analysis (IVA) research have adv…
Object Recognition