paper-with-me

홈 › Papers

GATSBI: Generative Agent-centric Spatio-temporal Object Interaction

2021-04-09 · CVPR 2021 1 · Cheol-Hui Min, Jinseok Bae, Junho Lee, Young Min Kim

We present GATSBI, a generative model that can transform a sequence of raw observations into a structured latent representation that fully captures the spatio-temporal context of the agent's actions. In vision-based decision-making scenarios, an agent faces complex high-dimensional observations where multiple entities interact with each other. The agent requires a good scene representation of the visual observation that discerns essential components and consistently propagates along the time horizon. Our method, GATSBI, utilizes unsupervised object-centric scene representation learning to separate an active agent, static background, and passive objects. GATSBI then models the interactions reflecting the causal relationships among decomposed entities and predicts physically plausible future states. Our model generalizes to a variety of environments where different types of robots and objects dynamically interact with each other. We show GATSBI achieves superior performance on scene decomposition and video prediction compared to its state-of-the-art counterparts.

📄 PDF Abstract BibTeX arXiv:2104.04275

Code (1)

mch5048/gatsbi 공식 구현

Tasks

Decision MakingObjectRepresentation LearningVideo Prediction

Similar Papers 제목 키워드 기반

GATSBI: Generative Adversarial Training for Simulation-Based Inference

2022-03-12 · ICLR 2022 4 · Poornima Ramesh, Jan-Matthis Lueckmann, Jan Boelts, Álvaro Tejero-Cantero 외

Simulation-based inference (SBI) refers to statistical inference on stochastic models for which we can generate samples, but not compute likelihoods. Like SBI algorithms, generative adversarial networks (GANs) do not req…

Bayesian Inference

Feature-Attending Recurrent Modules for Generalization in Reinforcement Learning

2021-12-15 · Wilka Carvalho, Andrew Lampinen, Kyriacos Nikiforou, Felix Hill 외

Many important tasks are defined in terms of object. To generalize across these tasks, a reinforcement learning (RL) agent needs to exploit the structure that the objects induce. Prior work has either hard-coded object-c…

Objectreinforcement-learningReinforcement LearningReinforcement Learning (RL)

EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT

2025-10-27 · Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He 외 arxiv

Egocentric video reasoning centers on an unobservable agent behind the camera who dynamically shapes the environment, requiring inference of hidden intentions and recognition of fine-grained interactions. This core chall…

LASAR: Towards Spatio-temporal Reasoning with Latent Cognitive Map

2026-05-16 · Jinzhou Tang, Sidi Liu, Waikit Xiu, Weixing Chen 외 arxiv

A fundamental challenge in embodied AI is verifying if agents build internal models of spatial structure or merely learn to mimic task-specific expert trajectories. This is critical as foundational approaches rooted in a…

Zero-shot GeneralizationRepresentation Learning

Spatio-temporal dual-stage hypergraph MARL for human-centric multimodal corridor traffic signal control

2026-02-19 · Xiaocai Zhang, Neema Nassir, Milad Haghani arxiv

Human-centric traffic signal control in corridor networks must increasingly account for multimodal travelers, particularly high-occupancy public transportation, rather than focusing solely on vehicle-centric performance.…

Multi-agent Reinforcement Learning