GATSBI: Generative Agent-centric Spatio-temporal Object Interaction
We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making scenarios, an agent faces complex high-dimensional observations where multiple entities interact with each other. The agent requires a good scene representation of the visual observation that discerns essential components and consistently propagates along the time horizon. Our method, GATSBI, utilizes unsupervised object-centric scene representation learning to separate an active agent, static background, and passive objects. GATSBI then models the interactions reflecting the causal relationships among decomposed entities and predicts physically plausible future states. Our model generalizes to a variety of environments where different types of robots and objects dynamically interact with each other. We show GATSBI achieves superior performance on scene decomposition and video prediction compared to its state-of-the-art counterparts.
Code (1)
Tasks
Decision MakingObjectRepresentation LearningVideo PredictionSimilar Papers 제목 키워드 기반
GATSBI: Generative Adversarial Training for Simulation-Based Inference
Simulation-based inference (SBI) refers to statistical inference on stochastic models for which we can generate samples, but not compute likelihoods. Like SBI algorithms, generative adversarial networks (GANs) do not req…
Bayesian InferenceFeature-Attending Recurrent Modules for Generalization in Reinforcement Learning
Many important tasks are defined in terms of object. To generalize across these tasks, a reinforcement learning (RL) agent needs to exploit the structure that the objects induce. Prior work has either hard-coded object-c…
Objectreinforcement-learningReinforcement LearningReinforcement Learning (RL)EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…
LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map
A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critical as foundational approaches rooted in a…
Zero-shot GeneralizationRepresentation LearningSpatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control
Human-centric traffic signal control in corridor networks must increasingly account for multimodal travelers, particularly high-occupancy public transportation, rather than focusing solely on vehicle-centric performance.…
Multi-agent Reinforcement Learning