Learning 3D object-centric representation through prediction
As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans develop the ability of perceiving objects situated in 3D environments without supervision, models that learn the same set of abilities with similar constraints faced by human infants are lacking. Towards this end, we developed a novel network architecture that simultaneously learns to 1) segment objects from discrete images, 2) infer their 3D locations, and 3) perceive depth, all while using only information directly available to the brain as training data, namely: sequences of images and self-motion. The core idea is treating objects as latent causes of visual input which the brain uses to make efficient predictions of future scenes. This results in object representations being learned as an essential byproduct of learning to predict.
Code (0)
등록된 구현이 없습니다.
Tasks
ObjectPredictionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Unsupervised Dynamics Prediction with Object-Centric Kinematics
Human perception involves discerning complex multi-object scenes into time-static object appearance (ie, size, shape, color) and time-varying object motion (ie, location, velocity, acceleration). This innate ability to u…
ObjectPredictionObject-Centric Video Prediction via Decoupling of Object Dynamics and Interactions
We propose a novel framework for the task of object-centric video prediction, i.e., extracting the compositional structure of a video sequence, as well as modeling objects dynamics and interactions from visual observatio…
ObjectPredictionVideo PredictionObject-Centric Learning with Slot Attention
Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed represe…
ObjectObject DiscoveryProperty PredictionObject-centric Video Prediction without Annotation
In order to interact with the world, agents must be able to predict the results of the world's dynamics. A natural approach to learn about these dynamics is through video prediction, as cameras are ubiquitous and powerfu…
ObjectPredictionVideo PredictionCausal-JEPA: Learning World Models through Object-Level Latent Masking
World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-depend…
Visual Question Answering