paper-with-me

Papers

Learning 3D object-centric representation through prediction

2024-03-06 · John Day, Tushar Arora, Jirui Liu, Li Erran Li, Ming Bo Cai

As part of human core knowledge, the representation of objects is the building block of mental representation that supports high-level concepts and symbolic reasoning. While humans develop the ability of perceiving objects situated in 3D environments without supervision, models that learn the same set of abilities with similar constraints faced by human infants are lacking. Towards this end, we developed a novel network architecture that simultaneously learns to 1) segment objects from discrete images, 2) infer their 3D locations, and 3) perceive depth, all while using only information directly available to the brain as training data, namely: sequences of images and self-motion. The core idea is treating objects as latent causes of visual input which the brain uses to make efficient predictions of future scenes. This results in object representations being learned as an essential byproduct of learning to predict.

📄 PDF Abstract BibTeX arXiv:2403.03730

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectPrediction

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Unsupervised Dynamics Prediction with Object-Centric Kinematics

2024-04-29 · Yeon-Ji Song, Suhyung Choi, Jaein Kim, Jin-Hwa Kim 외

Human perception involves discerning complex multi-object scenes into time-static object appearance (ie, size, shape, color) and time-varying object motion (ie, location, velocity, acceleration). This innate ability to u…

ObjectPrediction

Object-Centric Video Prediction via Decoupling of Object Dynamics and Interactions

2023-02-23 · Angel Villar-Corrales, Ismail Wahdan, Sven Behnke

We propose a novel framework for the task of object-centric video prediction, i.e., extracting the compositional structure of a video sequence, as well as modeling objects dynamics and interactions from visual observatio…

ObjectPredictionVideo Prediction

Object-Centric Learning with Slot Attention

2020-06-26 · NeurIPS 2020 12 · Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran 외

Learning object-centric representations of complex scenes is a promising step towards enabling efficient abstract reasoning from low-level perceptual features. Yet, most deep learning approaches learn distributed represe…

ObjectObject DiscoveryProperty Prediction

Object-centric Video Prediction without Annotation

2021-05-06 · Karl Schmeckpeper, Georgios Georgakis, Kostas Daniilidis

In order to interact with the world, agents must be able to predict the results of the world's dynamics. A natural approach to learn about these dynamics is through video prediction, as cameras are ubiquitous and powerfu…

ObjectPredictionVideo Prediction

Causal-JEPA: Learning World Models through Object-Level Latent Masking

2026-02-11 · Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun 외 arxiv

World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-depend…

Visual Question Answering