Learning to perceive objects by prediction
The representation of objects is the building block of higher-level concepts. Infants develop the notion of objects without supervision. The prediction error of future sensory input is likely the major teaching signal for infants. Inspired by this, we propose a new framework to extract object-centric representation from single 2D images by learning to predict future scenes in the presence of moving objects. We treat objects as latent causes of which the function for an agent is to facilitate efficient prediction of the coherent motion of their parts in visual input. Distinct from previous object-centric models, our model learns to explicitly infer objects' locations in a 3D environment in addition to segmenting objects. Further, the network learns a latent code space where objects with the same geometric shape and texture/color frequently group together. The model requires no supervision or pre-training of any part of the network. We created a new synthetic dataset with more complex textures on objects and background and found several previous models not based on predictive learning overly rely on clustering colors and lose specificity in object segmentation. Our work demonstrates a new approach for learning symbolic representation grounded in sensation and action.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectPredictionSemantic SegmentationSpecificitySimilar Papers 제목 키워드 기반
SymbioLCD: Ensemble-Based Loop Closure Detection using CNN-Extracted Objects and Visual Bag-of-Words
Loop closure detection is an essential tool of Simultaneous Localization and Mapping (SLAM) to minimize drift in its localization. Many state-of-the-art loop closure detection (LCD) algorithms use visual Bag-of-Words (vB…
Loop Closure DetectionSimultaneous Localization and MappingLearning 3D object-centric representation through prediction
As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans develop the ability of perceiving objec…
ObjectPredictionClouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts
We develop new representations and algorithms for three-dimensional (3D) object detection and spatial layout prediction in cluttered indoor scenes. We first propose a clouds of oriented gradient (COG) descriptor that lin…
3D Object DetectionGeneral Classificationobject-detectionObject Detection+1Active Object Perceiver: Recognition-guided Policy Learning for Object Searching on Mobile Robots
We study the problem of learning a navigation policy for a robot to actively search for an object of interest in an indoor environment solely from its visual inputs. While scene-driven visual navigation has been widely s…
Deep Reinforcement LearningObjectObject RecognitionReinforcement Learning+1Egocentric affordance detection with the one-shot geometry-driven Interaction Tensor
In this abstract we describe recent [4,7] and latest work on the determination of affordances in visually perceived 3D scenes. Our method builds on the hypothesis that geometry on its own provides enough information to e…
Affordance Detection