paper-with-me

홈 › Papers

PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning

2025-02-11 · Angel Villar-Corrales, Sven Behnke

Predicting future scene representations is a crucial task for enabling robots to understand and interact with the environment. However, most existing methods rely on videos and simulations with precise action annotations, limiting their ability to leverage the large amount of available unlabeled video data. To address this challenge, we propose PlaySlot, an object-centric video prediction model that infers object representations and latent actions from unlabeled video sequences. It then uses these representations to forecast future object states and video frames. PlaySlot allows the generation of multiple possible futures conditioned on latent actions, which can be inferred from video dynamics, provided by a user, or generated by a learned action policy, thus enabling versatile and interpretable world modeling. Our results show that PlaySlot outperforms both stochastic and object-centric baselines for video prediction across different environments. Furthermore, we show that our inferred latent actions can be used to learn robot behaviors sample-efficiently from unlabeled video demonstrations. Videos and code are available on https://play-slot.github.io/PlaySlot/.

📄 PDF Abstract BibTeX arXiv:2502.07600

Code (1)

angelvillar96/PlaySlot pytorch

Tasks

ObjectVideo Prediction

Similar Papers 제목 키워드 기반

Sensorimotor World Models: Perception for Action via Inverse Dynamics

2026-06-18 · Petr Ivashkov, Randall Balestriero, Bernhard Schölkopf arxiv

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compa…

Multistep Inverse Is Not All You Need

2024-03-18 · Alexander Levine, Peter Stone, Amy Zhang

In real-world control settings, the observation space is often unnecessarily high-dimensional and subject to time-correlated noise. However, the controllable dynamics of the system are often far simpler than the dynamics…

All

SCAR: Self-Supervised Continuous Action Representation Learning

2026-05-13 · Hongjia Liu, Fan Feng, Minghao Fu, Xinyue Wang 외 arxiv

Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across emb…

Representation Learning

Iso-Dream: Isolating and Leveraging Noncontrollable Visual Dynamics in World Models

2022-05-27 · Minting Pan, Xiangming Zhu, Yunbo Wang, Xiaokang Yang

World models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios such as autonomous driving, there commonly exists noncontrollable dynamics independent of the action sig…

Autonomous DrivingDecision Making

Model-Based Reinforcement Learning with Isolated Imaginations

2023-03-27 · Minting Pan, Xiangming Zhu, Yitao Zheng, Yunbo Wang 외

World models learn the consequences of actions in vision-based interactive systems. However, in practical scenarios like autonomous driving, noncontrollable dynamics that are independent or sparsely dependent on action s…

Autonomous DrivingmodelModel-based Reinforcement Learningreinforcement-learning+3