paper-with-me

홈 › Papers

Learning Object-Centric Video Models by Contrasting Sets

2020-11-20 · Sindy Löwe, Klaus Greff, Rico Jonschkowski, Alexey Dosovitskiy, Thomas Kipf

Contrastive, self-supervised learning of object representations recently emerged as an attractive alternative to reconstruction-based training. Prior approaches focus on contrasting individual object representations (slots) against one another. However, a fundamental problem with this approach is that the overall contrastive loss is the same for (i) representing a different object in each slot, as it is for (ii) (re-)representing the same object in all slots. Thus, this objective does not inherently push towards the emergence of object-centric representations in the slots. We address this problem by introducing a global, set-based contrastive loss: instead of contrasting individual slot representations against one another, we aggregate the representations and contrast the joined sets against one another. Additionally, we introduce attention-based encoders to this contrastive setup which simplifies training and provides interpretable object masks. Our results on two synthetic video datasets suggest that this approach compares favorably against previous contrastive methods in terms of reconstruction, future prediction and object separation performance.

📄 PDF Abstract BibTeX arXiv:2011.10287

Code (0)

등록된 구현이 없습니다.

Tasks

Future predictionObjectSelf-Supervised Learning

Similar Papers 제목 키워드 기반

Temporally Consistent Object-Centric Learning by Contrasting Slots

2024-12-18 · CVPR 2025 1 · Anna Manasyan, Maximilian Seitzer, Filip Radovic, Georg Martius 외

Unsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these repre…

Inductive BiasObjectObject Discovery

HumanForge: A Human-Centric Deepfake Video Benchmark with Multi-Agent Forgery Rationales

2026-07-09 · Wenbo Xu, Zhimin Chen, Xiaojie Liang, Hengrui Liu 외 arxiv

Rapid advancements in video diffusion models and temporal editing tools have enabled the generation of highly realistic human-centric videos, presenting unprecedented challenges to digital content forensics. Existing ben…

Zero-shot Generalization

Patch-based Object-centric Transformers for Efficient Video Generation

2022-06-08 · Wilson Yan, Ryo Okumura, Stephen James, Pieter Abbeel

In this work, we present Patch-based Object-centric Video Transformer (POVT), a novel region-based video generation architecture that leverages object-centric information to efficiently model temporal dynamics in videos.…

ObjectVideo EditingVideo GenerationVideo Prediction

Object-Shot Enhanced Grounding Network for Egocentric Video

2025-05-07 · CVPR 2025 1 · Yisen Feng, Haoyu Zhang, Meng Liu, Weili Guan 외

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentr…

Video Grounding

Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning

2024-08-07 · Zi-Yi Dou, Xitong Yang, Tushar Nagarajan, Huiyu Wang 외

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse acti…

Multi-Instance RetrievalRepresentation LearningStyle Transfer