paper-with-me

Papers

MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation

2025-11-13 · Xun Huang, Shijia Zhao, Yunxiang Wang, Xin Lu, Wanfa Zhang, Rongsheng Qu, Weixin Li, Yunhong Wang, Chenglu Wen arxiv

Embodied navigation is a fundamental capability for robotic agents operating. Real-world deployment requires open vocabulary generalization and low training overhead, motivating zero-shot methods rather than task-specific RL training. However, existing zero-shot methods that build explicit 3D scene graphs often compress rich visual observations into text-only relations, leading to high construction cost, irreversible loss of visual evidence, and constrained vocabularies. To address these limitations, we introduce the Multi-modal 3D Scene Graph (M3DSG), which preserves visual cues by replacing textual relational edges with dynamically assigned images. Built on M3DSG, we propose MSGNav, a zero-shot navigation system that includes a Key Subgraph Selection module for efficient reasoning, an Adaptive Vocabulary Update module for open vocabulary support, and a Closed-Loop Reasoning module for accurate exploration reasoning. Additionally, we further identify the last mile problem in zero-shot navigation determining the feasible target location with a suitable final viewpoint, and propose a Visibility-based Viewpoint Decision module to explicitly resolve it. Comprehensive experimental results demonstrate that MSGNav achieves state-of-the-art performance on the challenging GOAT-Bench and HM3D-ObjNav benchmark. The code will be publicly available at https://github.com/ylwhxht/MSGNav.

📄 PDF Abstract BibTeX arXiv:2511.10376

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

2025-01-01 · CVPR 2025 1 · Zhuoman Liu, Weicai Ye, Yan Luximon, Pengfei Wan 외

Realistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to …

Optical Flow Estimation

PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation

2024-11-21 · Zhuoman Liu, Weicai Ye, Yan Luximon, Pengfei Wan 외

Realistic simulation of dynamic scenes requires accurately capturing diverse material properties and modeling complex object interactions grounded in physical principles. However, existing methods are constrained to basi…

Optical Flow Estimation

Unleashing the Power of Imbalanced Modality Information for Multi-modal Knowledge Graph Completion

2024-02-22 · Yichi Zhang, Zhuo Chen, Lei Liang, Huajun Chen 외

Multi-modal knowledge graph completion (MMKGC) aims to predict the missing triples in the multi-modal knowledge graphs by incorporating structural, visual, and textual information of entities into the discriminant models…

Knowledge Graph CompletionKnowledge GraphsMulti-modal Knowledge Graph

Unleashing Network Potentials for Semantic Scene Completion

2024-03-12 · CVPR 2024 1 · Fengyun Wang, Qianru Sun, Dong Zhang, Jinhui Tang

Semantic scene completion (SSC) aims to predict complete 3D voxel occupancy and semantics from a single-view RGB-D image, and recent SSC methods commonly adopt multi-modal inputs. However, our investigation reveals two l…

Advances in 3D Neural Stylization: A Survey

2023-11-30 · Yingshu Chen, Guocheng Shao, Ka Chun Shum, Binh-Son Hua 외

Modern artificial intelligence offers a novel and transformative approach to creating digital art across diverse styles and modalities like images, videos and 3D data, unleashing the power of creativity and revolutionizi…

NavigateNeural StylizationStyle TransferSurvey