paper-with-me

Papers

Imaginative World Modeling with Scene Graphs for Embodied Agent Navigation

2025-08-09 · Yue Hu, Junzhe Wu, Ruihan Xu, Hang Liu, Avery Xi, Henry X. Liu, Ram Vasudevan, Maani Ghaffari arxiv

Semantic navigation requires an agent to navigate toward a specified target in an unseen environment. Employing an imaginative navigation strategy that predicts future scenes before taking action, can empower the agent to find target faster. Inspired by this idea, we propose SGImagineNav, a novel imaginative navigation framework that leverages symbolic world modeling to proactively build a global environmental representation. SGImagineNav maintains an evolving hierarchical scene graphs and uses large language models to predict and explore unseen parts of the environment. While existing methods solely relying on past observations, this imaginative scene graph provides richer semantic context, enabling the agent to proactively estimate target locations. Building upon this, SGImagineNav adopts an adaptive navigation strategy that exploits semantic shortcuts when promising and explores unknown areas otherwise to gather additional context. This strategy continuously expands the known environment and accumulates valuable semantic contexts, ultimately guiding the agent toward the target. SGImagineNav is evaluated in both real-world scenarios and simulation benchmarks. SGImagineNav consistently outperforms previous methods, improving success rate to 65.4 and 66.8 on HM3D and HSSD, and demonstrating cross-floor and cross-room navigation in real-world environments, underscoring its effectiveness and generalizability.

📄 PDF Abstract BibTeX arXiv:2508.06990

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Planning from Imagination: Episodic Simulation and Episodic Memory for Vision-and-Language Navigation

2024-11-30 · Yiyuan Pan, Yunzhe Xu, Zhe Liu, Hesheng Wang

Humans navigate unfamiliar environments using episodic simulation and episodic memory, which facilitate a deeper understanding of the complex relationships between environments and objects. Developing an imaginative memo…

NavigateVision and Language Navigation

GenEx: Generating an Explorable World

2024-12-12 · Taiming Lu, Tianmin Shu, Junfei Xiao, Luoxin Ye 외

Understanding, navigating, and exploring the 3D physical real world has long been a central challenge in the development of artificial intelligence. In this work, we take a step toward this goal by introducing GenEx, a s…

FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction

2025-09-25 · Yixiang Dai, Fan Jiang, Chiyu Wang, Mu Xu 외 arxiv

High-quality 3D world models are pivotal for embodied intelligence and Artificial General Intelligence (AGI), underpinning applications such as AR/VR content creation and robotic navigation. Despite the established stron…

Novel View SynthesisVideo Generation

3D-Belief: Embodied Belief Inference via Generative 3D World Modeling

2026-05-12 · Yifan Yin, Zehao Wen, Suyu Ye, Jieneng Chen 외 arxiv

Recent advances in visual generative models have highlighted the promise of learning generative world models. However, most existing approaches frame world modeling as novel-view synthesis or future-frame prediction, emp…

RoBoSR: Structured Scene Representations for Embodied Robotic Reasoning

2026-06-23 · Kewei Hu, Wanchan Yu, Fangwen Chen, Jing Jiajian 외 arxiv

Despite rapid progress, embodied reasoning under real-world variability remains challenging. Existing approaches rely on demonstration-driven sequential biases, limiting flexibility in open-ended and long-horizon tasks t…

Zero-shot GeneralizationScene Understanding