paper-with-me

Papers

SCALOR: Generative World Models with Scalable Object Representations

2019-10-06 · ICLR 2020 1 · Jindong Jiang, Sepehr Janghorbani, Gerard de Melo, Sungjin Ahn

Scalability in terms of object density in a scene is a primary challenge in unsupervised sequential object-oriented representation learning. Most of the previous models have been shown to work only on scenes with a few objects. In this paper, we propose SCALOR, a probabilistic generative world model for learning SCALable Object-oriented Representation of a video. With the proposed spatially-parallel attention and proposal-rejection mechanisms, SCALOR can deal with orders of magnitude larger numbers of objects compared to the previous state-of-the-art models. Additionally, we introduce a background module that allows SCALOR to model complex dynamic backgrounds as well as many foreground objects in the scene. We demonstrate that SCALOR can deal with crowded scenes containing up to a hundred objects while jointly modeling complex dynamic backgrounds. Importantly, SCALOR is the first unsupervised object representation model shown to work for natural scenes containing several tens of moving objects.

📄 PDF Abstract BibTeX arXiv:1910.02384

Code (2)

JindongJiang/JindongJiang.github.io
JindongJiang/SCALOR pytorch

Tasks

ObjectRepresentation Learning

Similar Papers 제목 키워드 기반

Benchmarking Unsupervised Object Representations for Video Sequences

2020-06-12 · Marissa A. Weis, Kashyap Chitta, Yash Sharma, Wieland Brendel 외

Perceiving the world in terms of objects and tracking them through time is a crucial prerequisite for reasoning and scene understanding. Recently, several methods have been proposed for unsupervised learning of object-ce…

BenchmarkingClusteringMulti-Object TrackingObject+4

Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

2025-12-01 · Haodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan 외 arxiv

Generating realistic hand-object interactions (HOI) videos is a significant challenge due to the difficulty of modeling physical constraints (e.g., contact and occlusion between hands and manipulated objects). Current me…

Video Generation

Choreographing a World of Dynamic Objects

2026-01-07 · Yanzhe Lyu, Chen Geng, Karthik Dharmarajan, Yunzhi Zhang 외 arxiv

Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we present a universal generative pipeline…

Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

2026-06-20 · Yangtao Chen, Zixuan Chen, Peiyang Wang, Yong-Lu Li 외 arxiv

Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with r…

Learning Object-Centric Representations Based on Slots in Real World Scenarios

2025-09-29 · Adil Kaan Akan arxiv

A central goal in AI is to represent scenes as compositions of discrete objects, enabling fine-grained, controllable image and video generation. Yet leading diffusion models treat images holistically and rely on text con…

Unsupervised Video Object SegmentationVideo GenerationImage Generation