paper-with-me

홈 › Papers

SIMONe: View-Invariant, Temporally-Abstracted Object Representations via Unsupervised Video Decomposition

2021-06-07 · NeurIPS 2021 12 · Rishabh Kabra, Daniel Zoran, Goker Erdogan, Loic Matthey, Antonia Creswell, Matthew Botvinick, Alexander Lerchner, Christopher P. Burgess

To help agents reason about scenes in terms of their building blocks, we wish to extract the compositional structure of any given scene (in particular, the configuration and characteristics of objects comprising the scene). This problem is especially difficult when scene structure needs to be inferred while also estimating the agent's location/viewpoint, as the two variables jointly give rise to the agent's observations. We present an unsupervised variational approach to this problem. Leveraging the shared structure that exists across different scenes, our model learns to infer two sets of latent representations from RGB video input alone: a set of "object" latents, corresponding to the time-invariant, object-level contents of the scene, as well as a set of "frame" latents, corresponding to global time-varying elements such as viewpoint. This factorization of latents allows our model, SIMONe, to represent object attributes in an allocentric manner which does not depend on viewpoint. Moreover, it allows us to disentangle object dynamics and summarize their trajectories as time-abstracted, view-invariant, per-object properties. We demonstrate these capabilities, as well as the model's performance in terms of view synthesis and instance segmentation, across three procedurally generated video datasets.

📄 PDF Abstract BibTeX arXiv:2106.03849

Code (1)

lkhphuc/simone jax

Tasks

Instance SegmentationObjectSemantic Segmentation

Similar Papers 제목 키워드 기반

Book Review: The Structure of Scientific Articles: Applications to Citation Indexing and Summarization by Simone Teufel

2012-01-01 · CL 2012 1 · Robert E. Mercer
ArticlesInformation Retrieval

Deep Networks Can Resemble Human Feed-forward Vision in Invariant Object Recognition

2015-08-17 · Saeed Reza Kheradpisheh, Masoud Ghodrati, Mohammad Ganjtabesh, Timothée Masquelier

Deep convolutional neural networks (DCNNs) have attracted much attention recently, and have shown to be able to recognize thousands of object categories in natural image databases. Their architecture is somewhat similar …

Object Recognition

Discrete Predictive Representation for Long-horizon Planning

2021-01-01 · Thanard Kurutach, Julia Peng, Yang Gao, Stuart Russell 외

Discrete representations have been key in enabling robots to plan at more abstract levels and solve temporally-extended tasks more efficiently for decades. However, they typically require expert specifications. On the ot…

Deep Reinforcement LearningObjectReinforcement Learning (RL)

Learning Fine-grained View-Invariant Representations from Unpaired Ego-Exo Videos via Temporal Alignment

2023-06-08 · NeurIPS 2023 11

The egocentric and exocentric viewpoints of a human activity look dramatically different, yet invariant representations to link them are essential for many potential applications in robotics and augmented reality. Prior …

Video Understanding

MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System

2025-06-18 · Miaoxin Pan, Jinnan Li, Yaowen Zhang, Yi Yang 외

Object-level SLAM offers structured and semantically meaningful environment representations, making it more interpretable and suitable for high-level robotic tasks. However, most existing approaches rely on RGB-D sensors…

ObjectObject SLAM