paper-with-me

Papers

Language-Mediated, Object-Centric Representation Learning

2020-12-31 · Findings (ACL) 2021 8 · Ruocheng Wang, Jiayuan Mao, Samuel J. Gershman, Jiajun Wu

We present Language-mediated, Object-centric Representation Learning (LORL), a paradigm for learning disentangled, object-centric scene representations from vision and language. LORL builds upon recent advances in unsupervised object discovery and segmentation, notably MONet and Slot Attention. While these algorithms learn an object-centric representation just by reconstructing the input image, LORL enables them to further learn to associate the learned representations to concepts, i.e., words for object categories, properties, and spatial relationships, from language input. These object-centric concepts derived from language facilitate the learning of object-centric representations. LORL can be integrated with various unsupervised object discovery algorithms that are language-agnostic. Experiments show that the integration of LORL consistently improves the performance of unsupervised object discovery methods on two datasets via the help of language. We also show that concepts learned by LORL, in conjunction with object discovery methods, aid downstream tasks such as referring expression comprehension.

📄 PDF Abstract BibTeX arXiv:2012.15814

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectObject DiscoveryReferring ExpressionReferring Expression ComprehensionRepresentation LearningSemantic SegmentationUnsupervised Object Segmentation

Methods 이 논문이 사용한 방법론

MoNet Mixture model network (MoNet) is a general framework allowing to design convolutional deep architectures on non-Euclidean domains such as graphs and manifolds. Image and…

Similar Papers 제목 키워드 기반

Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

2026-07-28 · Adam Scicluna, Gavin Paul, Alen Alempijevic arxiv

Zero-shot ObjectNav methods increasingly use vision-language priors, but direct object-object similarity in the latent space is often a weak proxy for spatial co-occurrence. We present an analytical, training-free semant…

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

2026-01-18 · Tsan Tsai Chan, Varsha Suresh, Anisha Saha, Michael Hahn 외 arxiv

Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation strategies tend towards an image-centric in…

CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

2025-03-27 · CVPR 2025 1 · Aniket Didolkar, Andrii Zadaianchuk, Rabiul Awal, Maximilian Seitzer 외

Object-centric representation learning aims to decompose visual scenes into fixed-size vectors called "slots" or "object files", where each slot captures a distinct object. Current state-of-the-art object-centric models …

Image GenerationObjectObject DiscoveryQuestion Answering+4

UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learning

2026-06-18 · Wenhao Chi, Arkaprava Sinha, Dominick Reilly, Hieu Le 외 arxiv

Egocentric video understanding is inherently limited by the narrow perspective of wearable cameras: a single viewpoint, a single modality, a single model cannot capture the full richness of human action. We argue that a …

Representation LearningAction SegmentationAction RecognitionVideo Retrieval

Object-Centric World Model for Language-Guided Manipulation

2025-03-08 · Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim

A world model is essential for an agent to predict the future and plan in domains such as autonomous driving and robotics. To achieve this, recent advancements have focused on video generation, which has gained significa…

Autonomous DrivingmodelObjectObject Recognition+1