paper-with-me

Papers

Dual-State Slot Attention: Decoupling Appearance and Identity for Video Object-Centric Learning

2026-06-10 · Sieu Tran, Duc Nguyen, Hao Vo, Khoa Vo, Ngan Le arxiv

Unsupervised video object-centric learning aims to decompose dynamic scenes into persistent, object-level representations without supervision. However, existing slot-based methods struggle to maintain stable object identity in challenging settings such as rapid motion and partial occlusion. First, they typically encode both the per-frame appearance of an object and its identity across frames in a single slot vector, creating an objective conflict that leads to slot swapping: reconstruction requires sensitivity to transient visual changes, whereas temporal consistency requires invariance to them. Second, the token renormalization used in Slot Attention can amplify weakly attending slots, allowing them to absorb tokens from other objects and destabilize slot-to-object correspondence. We propose Dual-State Slot Attention (DSSA), a fully self-supervised framework that addresses these limitations by separating appearance from identity and by reducing spurious updates from weakly matching slots. DSSA decomposes each slot into a local state for per-frame appearance and an identity state for temporally stable object information, thereby aligning reconstruction and temporal consistency with separate representations. The identity state is updated through a learned recurrent transition that acts as a temporal filter on the local state, while competition-modulated aggregation (CMA) down-weights updates from weakly matching slots and prevents them from absorbing tokens from other objects. Experiments on MOVi-C, MOVi-D, and YouTube-VIS demonstrate that DSSA consistently improves segmentation quality and temporal consistency over prior methods, while also yielding stronger downstream object recognition and video dynamics prediction. Code and models will be made publicly available upon acceptance.

📄 PDF Abstract BibTeX arXiv:2606.12601

Code (0)

등록된 구현이 없습니다.

Tasks

Object Recognition

Similar Papers 제목 키워드 기반

TSA: Temporal Slot Activation for Persistent Object-Centric Video Representation

2026-06-10 · Duc Nguyen, Sieu Tran, Hao Vo, Khoa Vo 외 arxiv

Unsupervised video object-centric learning aims to decompose dynamic scenes into temporally persistent entity representations. Existing recurrent video slot-attention methods propagate a fixed set of slots across frames,…

HSA: Hierarchical Slot Attention for Multi-granularity Scene-Decomposition

2026-07-09 · Neelu Madan, Rongzhen Zhao, Andreas Mogelmose, Juho Kannala 외 arxiv

Slot attention is a powerful framework for object-centric learning, decomposing visual scenes into latent slots through iterative competitive attention. However, existing methods share two critical limitations: they deco…

Rethinking Object-Centric Representations for Video Dynamics Modeling

2026-06-22 · Amaury Wei, Ismail Nejjar, Olga Fink arxiv

Unsupervised video object tracking aims to decompose dynamic scenes into persistent, object-centric entities without manual annotations. Many recent approaches rely on slot-based representations, where a fixed set of lat…

Video Object Tracking

MUFASA: A Multi-Layer Framework for Slot Attention

2026-02-07 · Sebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer 외 arxiv

Unsupervised object-centric learning (OCL) decomposes visual scenes into distinct entities. Slot attention is a popular approach that represents individual objects as latent vectors, called slots. Current methods obtain …

Unsupervised Object Segmentation

Unsupervised Structural Scene Decomposition via Foreground-Aware Slot Attention with Pseudo-Mask Guidance

2025-12-02 · Huankun Sheng, Ming Li, Yixiang Wei, Yeying Fan 외 arxiv

Recent advances in object-centric representation learning have shown that slot attention-based methods can effectively decompose visual scenes into object slot representations without supervision. However, existing appro…

Representation Learning