paper-with-me

홈 › Papers

Improving Generative Imagination in Object-Centric World Models

2020-10-05 · Zhixuan Lin, Yi-Fu Wu, Skand Peri, Bofeng Fu, Jindong Jiang, Sungjin Ahn

The remarkable recent advances in object-centric generative world models raise a few questions. First, while many of the recent achievements are indispensable for making a general and versatile world model, it is quite unclear how these ingredients can be integrated into a unified framework. Second, despite using generative objectives, abilities for object detection and tracking are mainly investigated, leaving the crucial ability of temporal imagination largely under question. Third, a few key abilities for more faithful temporal imagination such as multimodal uncertainty and situation-awareness are missing. In this paper, we introduce Generative Structured World Models (G-SWM). The G-SWM achieves the versatile world modeling not only by unifying the key properties of previous models in a principled framework but also by achieving two crucial new abilities, multimodal uncertainty and situation-awareness. Our thorough investigation on the temporal generation ability in comparison to the previous models demonstrates that G-SWM achieves the versatility with the best or comparable performance for all experiment settings including a few complex settings that have not been tested before.

📄 PDF Abstract BibTeX arXiv:2010.02054

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

VISTAv2: World Imagination for Indoor Vision-and-Language Navigation

2025-11-14 · Yanjia Huang, Xianshun Jiang, Xiangbo Gao, Mingyang Wu 외 arxiv

Vision-and-Language Navigation (VLN) requires agents to follow language instructions while acting in continuous real-world spaces. Prior image imagination based VLN work shows benefits for discrete panoramas but lacks on…

OC-NMN: Object-centric Compositional Neural Module Network for Generative Visual Analogical Reasoning

2023-10-28 · Rim Assouel, Pau Rodriguez, Perouz Taslakian, David Vazquez 외

A key aspect of human intelligence is the ability to imagine -- composing learned concepts in novel ways -- to make sense of new scenarios. Such capacity is not yet attained for machine learning systems. In this work, in…

Data AugmentationOut-of-Distribution GeneralizationVisual Reasoning

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

2026-05-12 · Yifan Yin, Zehao Wen, Suyu Ye, Jieneng Chen 외 arxiv

Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame world modeling as novel-view synthesis or future-frame prediction, emp…

Object-Centric World Models for Causality-Aware Reinforcement Learning

2025-11-18 · Yosuke Nishimoto, Takashi Matsubara arxiv

World models have been developed to support sample-efficient deep reinforcement learning agents. However, it remains challenging for world models to accurately replicate environments that are high-dimensional, non-statio…

Reinforcement Learning

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

2026-04-29 · Wanyue Zhang, Wenxiang Wu, Wang Xu, Jiaxin Luo 외 arxiv

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent…

Spatial Reasoning