paper-with-me

홈 › Papers

Smoothing Slot Attention Iterations and Recurrences

2025-08-07 · Rongzhen Zhao, Wenyan Yang, Juho Kannala, Joni Pajarinen arxiv

Slot Attention (SA) lies at the heart of mainstream Object-Centric Learning (OCL). Image features can be aggregated into object-level representations by SA \textit{iteratively} refining cold-start query slots. For video, such aggregation proceeds by SA \textit{recurrently} shared across frames, with queries cold-started on the first frame while transitioned from the previous frame's slots thereafter. However, cold-start queries lack sample-specific cues thus hindering precise aggregation on image or video's first frame; Non-first frames' queries are already sample-specific thus requiring aggregation transforms different from the first frame. We address these issues with our \textit{SmoothSA}: (1) To smooth SA iterations on image or video's first frame, we \textit{preheat} cold-start queries with rich input-feature information, by a tiny module self-distilled inside OCL; (2) To smooth SA recurrences across video's first and non-first frames, we \textit{differentiate} the homogeneous aggregation transforms by using full and single iterations respectively. Comprehensive experiments on object discovery, recognition and visual reasoning validate our method's effectiveness. Further visual analyses illuminate the underline mechanisms. Our \textit{source code}, \textit{model checkpoints} and \textit{training logs} are provided on https://github.com/Genera1Z/SmoothSA.

📄 PDF Abstract BibTeX arXiv:2508.05417

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Predicting Video Slot Attention Queries from Random Slot-Feature Pairs

2025-08-02 · Rongzhen Zhao, Jian Li, Juho Kannala, Joni Pajarinen arxiv

Unsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator …

Scene Understanding

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences

2026-04-22 · Neehal Tumma, Noel Loo, Daniela Rus arxiv

To address the increasing long-context compute limitations of softmax attention, several subquadratic recurrent operators have been developed. This work includes models such as Mamba-2, DeltaNet, Gated DeltaNet (GDN), an…

Sliding Window Recurrences for Sequence Models

2025-12-15 · Dragos Secrieru, Garyk Brixi, Yoshua Bengio, Taiji Suzuki 외 arxiv

Multi-hybrid architectures are poised to take over language modeling due to better quality and performance. We introduce a hierarchical decomposition framework for linear recurrences that allows us to develop algorithms …

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

2024-02-29 · Soham De, Samuel L. Smith, Anushan Fernando, Aleksandar Botev 외

Recurrent neural networks (RNNs) have fast inference and scale efficiently on long sequences, but they are difficult to train and hard to scale. We propose Hawk, an RNN with gated linear recurrences, and Griffin, a hybri…

Language ModellingMamba

Adaptive Slot Attention: Object Discovery with Dynamic Slot Number

2024-06-13 · CVPR 2024 1 · Ke Fan, Zechen Bai, Tianjun Xiao, Tong He 외

Object-centric learning (OCL) extracts the representation of objects with slots, offering an exceptional blend of flexibility and interpretability for abstracting low-level perceptual features. A widely adopted method wi…

DecoderObjectObject Discovery