paper-with-me

홈 › Papers

Object-Centric Latent Action Learning

2025-02-13 · Albina Klepach, Alexander Nikulin, Ilya Zisman, Denis Tarasov, Alexander Derevyagin, Andrei Polubarov, Nikita Lyubaykin, Vladislav Kurenkov

Leveraging vast amounts of unlabeled internet video data for embodied AI is currently bottlenecked by the lack of action labels and the presence of action-correlated visual distractors. Although recent latent action policy optimization (LAPO) has shown promise in inferring proxy-action labels from visual observations, its performance degrades significantly when distractors are present. To address this limitation, we propose a novel object-centric latent action learning framework that centers on objects rather than pixels. We leverage self-supervised object-centric pretraining to disentangle action-related and distracting dynamics. This allows LAPO to focus on task-relevant interactions, resulting in more robust proxy-action labels, enabling better imitation learning and efficient adaptation of the agent with just a few action-labeled trajectories. We evaluated our method in eight visually complex tasks across the Distracting Control Suite (DCS) and Distracting MetaWorld (DMW). Our results show that object-centric pretraining mitigates the negative effects of distractors by 50%, as measured by downstream task performance: average return (DCS) and success rate (DMW).

📄 PDF Abstract BibTeX arXiv:2502.09680

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningObject

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Causal Object-Centric Models for Planning with Monte Carlo Tree Search

2026-06-12 · Rodion Vakhitov, Leonid Ugadiarov, Alexey Skrynnik, Aleksandr Panov arxiv

We introduce COMET (Causal Object-centric Model for Efficient Tree search), a model-based reinforcement learning algorithm that performs Monte Carlo Tree Search in a slot-structured latent space. COMET pairs a frozen uns…

Reinforcement Learning

When Object-Centric World Models Meet Policy Learning: From Pixels to Policies, and Where It Breaks

2025-11-08 · Stefano Ferraro, Akihiro Nakano, Masahiro Suzuki, Yutaka Matsuo arxiv

Object-centric world models (OCWM) aim to decompose visual scenes into object-level representations, providing structured abstractions that could improve compositional generalization and data efficiency in reinforcement …

Reinforcement Learning

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

2025-02-11 · Angel Villar-Corrales, Sven Behnke

Predicting future scene representations is a crucial task for enabling robots to understand and interact with the environment. However, most existing methods rely on videos and simulations with precise action annotations…

ObjectVideo Prediction

$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation

2026-08-06 · Zhe Li, Zhenzhe Zhang, Yangyang Wei, Wenjie Zhang 외 arxiv

Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies ty…

Causal-JEPA: Learning World Models through Object-Level Latent Masking

2026-02-11 · Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun 외 arxiv

World models require robust relational understanding to support prediction, reasoning, and control. While object-centric representations provide a useful abstraction, they are not sufficient to capture interaction-depend…

Visual Question Answering