paper-with-me

Papers

SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models

2022-10-12 · Ziyi Wu, Nikita Dvornik, Klaus Greff, Thomas Kipf, Animesh Garg

Understanding dynamics from visual observations is a challenging problem that requires disentangling individual objects from the scene and learning their interactions. While recent object-centric models can successfully decompose a scene into objects, modeling their dynamics effectively still remains a challenge. We address this problem by introducing SlotFormer -- a Transformer-based autoregressive model operating on learned object-centric representations. Given a video clip, our approach reasons over object features to model spatio-temporal relationships and predicts accurate future object states. In this paper, we successfully apply SlotFormer to perform video prediction on datasets with complex object interactions. Moreover, the unsupervised SlotFormer's dynamics model can be used to improve the performance on supervised downstream tasks, such as Visual Question Answering (VQA), and goal-conditioned planning. Compared to past works on dynamics modeling, our method achieves significantly better long-term synthesis of object dynamics, while retaining high quality visual generation. Besides, SlotFormer enables VQA models to reason about the future without object-level labels, even outperforming counterparts that use ground-truth annotations. Finally, we show its ability to serve as a world model for model-based planning, which is competitive with methods designed specifically for such tasks.

📄 PDF Abstract BibTeX arXiv:2210.05861

Code (1)

pairlab/slotformer 공식 구현 pytorch

Tasks

ObjectQuestion AnsweringVideo PredictionVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

How do I escalate a complaint with Expedia?*EscalateFastService How do I escalate a complaint with Expedia? Call + 1 ≈ 888 ≈ 829 ≈ 0881 or + 1 || 888 || 829 || 0881 for Fast Resolution & Exclusive Travel Deals! Need to escalate a complaint…
Transformer A Transformer is a model architecture that eschews recurrence and instead relies entirely on an [attention…

Similar Papers 제목 키워드 기반

FSHNet: Fully Sparse Hybrid Network for 3D Object Detection

2025-01-01 · CVPR 2025 1 · Shuai Liu, Mingyue Cui, Boyang Li, Quanmin Liang 외

Fully sparse 3D detectors have recently gained significant attention due to their efficiency in long-range detection. However, sparse 3D detectors extract features only from non-empty voxels, which impairs long-range…

3D Object Detectionobject-detectionObject Detection

SlotGNN: Unsupervised Discovery of Multi-Object Representations and Visual Dynamics

2023-10-06 · Alireza Rezazadeh, Athreyi Badithela, Karthik Desingh, Changhyun Choi

Learning multi-object dynamics from visual data using unsupervised techniques is challenging due to the need for robust, object representations that can be learned through robot interactions. This paper presents a novel …

ObjectObject DiscoveryObject RearrangementSpatial Reasoning

Relevance-Guided Modeling of Object Dynamics for Reinforcement Learning

2020-03-03 · William Agnew, Pedro Domingos

Current deep reinforcement learning (RL) approaches incorporate minimal prior knowledge about the environment, limiting computational and sample efficiency. \textit{Objects} provide a succinct and causal description of t…

Atari GamesDeep Reinforcement LearningMuJoCoObject+5

Towards Scale-Aware, Robust, and Generalizable Unsupervised Monocular Depth Estimation by Integrating IMU Motion Dynamics

2022-07-11 · Sen Zhang, Jing Zhang, DaCheng Tao

Unsupervised monocular depth and ego-motion estimation has drawn extensive research attention in recent years. Although current methods have reached a high up-to-scale accuracy, they usually fail to learn the true scale …

Depth EstimationMonocular Depth EstimationMotion EstimationUnsupervised Monocular Depth Estimation

VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization

2025-08-19 · Jiajing Lin, Shu Jiang, Qingyuan Zeng, Zhenzhong Wang 외 arxiv

The intrinsic dynamics of an object governs its physical behavior in the real world, playing a critical role in enabling physically plausible interactive simulation with 3D assets. Existing methods have attempted to infe…

Bilevel Optimization