paper-with-me

홈 › Papers

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

2026-01-29 · Alexandre Chapin, Bruno Machado, Emmanuel Dellandréa, Liming Chen arxiv

The generalization capabilities of robotic manipulation policies are heavily influenced by the choice of visual representations. Existing approaches typically rely on representations extracted from pre-trained encoders, using two dominant types of features: global features, which summarize an entire image via a single pooled vector, and dense features, which preserve a patch-wise embedding from the final encoder layer. While widely used, both feature types mix task-relevant and irrelevant information, leading to poor generalization under distribution shifts, such as changes in lighting, textures, or the presence of distractors. In this work, we explore an intermediate structured alternative: Slot-Based Object-Centric Representations (SBOCR), which group dense features into a finite set of object-like entities. This representation permits to naturally reduce the noise provided to the robotic manipulation policy while keeping enough information to efficiently perform the task. We benchmark a range of global and dense representations against intermediate slot-based representations, across a suite of simulated and real-world manipulation tasks ranging from simple to complex. We evaluate their generalization under diverse visual conditions, including changes in lighting, texture, and the presence of distractors. Our findings reveal that SBOCR-based policies outperform dense and global representation-based policies in generalization settings, even without task-specific pretraining. These insights suggest that SBOCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.

📄 PDF Abstract BibTeX arXiv:2601.21416

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Defending Against Indirect Prompt Injection Attacks With Spotlighting

2024-03-20 · Keegan Hines, Gary Lopez, Matthew Hall, Federico Zarfati 외

Large Language Models (LLMs), while powerful, are built and trained to process a single text input. In common applications, multiple inputs can be processed by concatenating them together into a single stream of text. Ho…

Prompt Engineering

Object-Centric Representations Improve Policy Generalization in Robot Manipulation

2025-05-16 · Alexandre Chapin, Bruno Machado, Emmanuel Dellandrea, Liming Chen

Visual representations are central to the learning and generalization capabilities of robotic manipulation policies. While existing methods rely on global or dense features, such representations often entangle task-relev…

Optical Character Recognition (OCR)Robot Manipulation

Prompt-Driven Dynamic Object-Centric Learning for Single Domain Generalization

2024-02-28 · CVPR 2024 1 · Deng Li, Aming Wu, YaoWei Wang, Yahong Han

Single-domain generalization aims to learn a model from single source domain data to achieve generalized performance on other unseen target domains. Existing works primarily focus on improving the generalization ability …

Domain Generalizationimage-classificationImage ClassificationObject+3

You-Do, I-Learn: Unsupervised Multi-User egocentric Approach Towards Video-Based Guidance

2015-10-16 · Dima Damen, Teesid Leelasawassuk, Walterio Mayol-Cuevas

This paper presents an unsupervised approach towards automatically extracting video-based guidance on object usage, from egocentric video and wearable gaze tracking, collected from multiple users while performing tasks. …

Object

Deep Reinforcement Learning via Object-Centric Attention

2025-04-03 · Jannis Blüml, Cedric Derstroff, Bjarne Gregori, Elisabeth Dillies 외

Deep reinforcement learning agents, trained on raw pixel inputs, often fail to generalize beyond their training environments, relying on spurious correlations and irrelevant background details. To address this issue, obj…

Deep Reinforcement LearningInductive BiasObjectreinforcement-learning+1