SpatialSim: Recognizing Spatial Configurations of Objects with Graph Neural Networks
Recognizing precise geometrical configurations of groups of objects is a key capability of human spatial cognition, yet little studied in the deep learning literature so far. In particular, a fundamental problem is how a machine can learn and compare classes of geometric spatial configurations that are invariant to the point of view of an external observer. In this paper we make two key contributions. First, we propose SpatialSim (Spatial Similarity), a novel geometrical reasoning benchmark, and argue that progress on this benchmark would pave the way towards a general solution to address this challenge in the real world. This benchmark is composed of two tasks: Identification and Comparison, each one instantiated in increasing levels of difficulty. Secondly, we study how relational inductive biases exhibited by fully-connected message-passing Graph Neural Networks (MPGNNs) are useful to solve those tasks, and show their advantages over less relational baselines such as Deep Sets and unstructured models such as Multi-Layer Perceptrons. Finally, we highlight the current limits of GNNs in these tasks.
Code (0)
등록된 구현이 없습니다.
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
spatialSim: multi-species spatiotemporal size-structured operating model for management strategy evaluation
Spatiotemporal processes have the potential to be one of the most influential factors governing how fisheries targeting sedentary species respond to harvesting. Despite this, management strategy evaluation often fails to…
ManagementVSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph Convolutions
Comprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually. This is the main objective in Human-Object Interaction (HO…
Human-Object Interaction DetectionObjectSpatial ReasoningClosing the Loop: Graph Networks to Unify Semantic Objects and Visual Features for Multi-object Scenes
In Simultaneous Localization and Mapping (SLAM), Loop Closure Detection (LCD) is essential to minimize drift when recognizing previously visited places. Visual Bag-of-Words (vBoW) has been an LCD algorithm of choice for …
Graph MatchingLoop Closure DetectionSimultaneous Localization and MappingEfficient and Interpretable Robot Manipulation with Graph Neural Networks
Manipulation tasks, like loading a dishwasher, can be seen as a sequence of spatial constraints and relationships between different objects. We aim to discover these rules from demonstrations by posing manipulation as a …
Decision MakingGraph Neural NetworkImitation LearningRobot ManipulationLearning a Hierarchical Compositional Shape Vocabulary for Multi-class Object Representation
Hierarchies allow feature sharing between objects at multiple levels of representation, can code exponential variability in a very compact way and enable fast inference. This makes them potentially suitable for learning …
Object