paper-with-me

Papers

Object-Centric Representation Learning with Generative Spatial-Temporal Factorization

2021-11-09 · NeurIPS 2021 12 · Li Nanbo, Muhammad Ahmed Raza, Hu Wenbin, Zhaole Sun, Robert B. Fisher

Learning object-centric scene representations is essential for attaining structural understanding and abstraction of complex scenes. Yet, as current approaches for unsupervised object-centric representation learning are built upon either a stationary observer assumption or a static scene assumption, they often: i) suffer single-view spatial ambiguities, or ii) infer incorrectly or inaccurately object representations from dynamic scenes. To address this, we propose Dynamics-aware Multi-Object Network (DyMON), a method that broadens the scope of multi-view object-centric representation learning to dynamic scenes. We train DyMON on multi-view-dynamic-scene data and show that DyMON learns -- without supervision -- to factorize the entangled effects of observer motions and scene object dynamics from a sequence of observations, and constructs scene object spatial representations suitable for rendering at arbitrary times (querying across time) and from arbitrary viewpoints (querying across space). We also show that the factorized scene representations (w.r.t. objects) support querying about a single object by space and time independently.

📄 PDF Abstract BibTeX arXiv:2111.05393

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectRepresentation Learning

Similar Papers 제목 키워드 기반

Feature-Attending Recurrent Modules for Generalization in Reinforcement Learning

2021-12-15 · Wilka Carvalho, Andrew Lampinen, Kyriacos Nikiforou, Felix Hill 외

Many important tasks are defined in terms of object. To generalize across these tasks, a reinforcement learning (RL) agent needs to exploit the structure that the objects induce. Prior work has either hard-coded object-c…

Objectreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling

2024-01-29 · Yuze Hao, Jianrong Zhang, Tao Zhuo, Fuan Wen 외

Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object man…

Object

Compositional Video Synthesis by Temporal Object-Centric Learning

2025-07-28 · Adil Kaan Akan, Yucel Yemez arxiv

We present a novel framework for compositional video synthesis that leverages temporally consistent object-centric representations, extending our previous work, SlotAdapt, from images to video. While existing object-cent…

Scene UnderstandingVideo Generation

Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames

2025-05-30 · Sahithya Ravi, Gabriel Sarch, Vibhav Vineet, Andrew D. Wilson 외

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We …

ObjectSpatial Reasoning

Cross-View Exocentric to Egocentric Video Synthesis

2021-07-07 · Gaowen Liu, Hao Tang, Hugo Latapie, Jason Corso 외

Cross-view video synthesis task seeks to generate video sequences of one view from another dramatically different view. In this paper, we investigate the exocentric (third-person) view to egocentric (first-person) view v…

Generative Adversarial NetworkVideo Generation