paper-with-me

Papers

Grouped Discrete Representation Guides Object-Centric Learning

2024-07-01 · Rongzhen Zhao, Vivienne Wang, Juho Kannala, Joni Pajarinen

Similar to humans perceiving visual scenes as objects, Object-Centric Learning (OCL) can abstract dense images or videos into sparse object-level features. Transformer-based OCL handles complex textures well due to the decoding guidance of discrete representation, obtained by discretizing noisy features in image or video feature maps using template features from a codebook. However, treating features as minimal units overlooks their composing attributes, thus impeding model generalization; indexing features with natural numbers loses attribute-level commonalities and characteristics, thus diminishing heuristics for model convergence. We propose \textit{Grouped Discrete Representation} (GDR) to address these issues by grouping features into attributes and indexing them with tuple numbers. In extensive experiments across different query initializations, dataset modalities, and model architectures, GDR consistently improves convergence and generalizability. Visualizations show that our method effectively captures attribute-level information in features. The source code will be available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2407.01726

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeObject

Similar Papers 제목 키워드 기반

Grouped Discrete Representation for Object-Centric Learning

2024-11-04 · Rongzhen Zhao, Vivienne Wang, Juho Kannala, Joni Pajarinen

Object-Centric Learning (OCL) can discover objects in images or videos by simply reconstructing the input. For better object discovery, representative OCL methods reconstruct the input as its Variational Autoencoder (VAE…

AttributeObjectObject Discovery

Organized Grouped Discrete Representation for Object-Centric Learning

2024-09-05 · Rongzhen Zhao, Vivienne Wang, Juho Kannala, Joni Pajarinen

Object-Centric Learning (OCL) represents dense image or video pixels as sparse object features. Representative methods utilize discrete representation composed of Variational Autoencoder (VAE) template features to suppre…

ObjectRepresentation Learning

Unsupervised Discovery and Composition of Object Light Fields

2022-05-08 · Cameron Smith, Hong-Xing Yu, Sergey Zakharov, Fredo Durand 외

Neural scene representations, both continuous and discrete, have recently emerged as a powerful new paradigm for 3D scene understanding. Recent efforts have tackled unsupervised discovery of object-centric neural scene r…

Novel View SynthesisObjectScene Understanding

SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation

2025-11-10 · Taisei Hanyu, Nhat Chung, Huy Le, Toan Nguyen 외 arxiv

Inspired by how humans reason over discrete objects and their relationships, we explore whether compact object-centric and object-relation representations can form a foundation for multitask robotic manipulation. Most ex…

Patch-based Object-centric Transformers for Efficient Video Generation

2022-06-08 · Wilson Yan, Ryo Okumura, Stephen James, Pieter Abbeel

In this work, we present Patch-based Object-centric Video Transformer (POVT), a novel region-based video generation architecture that leverages object-centric information to efficiently model temporal dynamics in videos.…

ObjectVideo EditingVideo GenerationVideo Prediction