paper-with-me

홈 › Papers

MAG-Nav: Language-Driven Object Navigation Leveraging Memory-Reserved Active Grounding

2025-08-07 · Weifan Zhang, Tingguang Li, Yuzhen Liu arxiv

Visual navigation in unknown environments based solely on natural language descriptions is a key capability for intelligent robots. In this work, we propose a navigation framework built upon off-the-shelf Visual Language Models (VLMs), enhanced with two human-inspired mechanisms: perspective-based active grounding, which dynamically adjusts the robot's viewpoint for improved visual inspection, and historical memory backtracking, which enables the system to retain and re-evaluate uncertain observations over time. Unlike existing approaches that passively rely on incidental visual inputs, our method actively optimizes perception and leverages memory to resolve ambiguity, significantly improving vision-language grounding in complex, unseen environments. Our framework operates in a zero-shot manner, achieving strong generalization to diverse and open-ended language descriptions without requiring labeled data or model fine-tuning. Experimental results on Habitat-Matterport 3D (HM3D) show that our method outperforms state-of-the-art approaches in language-driven object navigation. We further demonstrate its practicality through real-world deployment on a quadruped robot, achieving robust and effective navigation performance.

📄 PDF Abstract BibTeX arXiv:2508.05021

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Navigation

Similar Papers 제목 키워드 기반

MC-GPT: Empowering Vision-and-Language Navigation with Memory Map and Reasoning Chains

2024-05-17 · Zhaohuan Zhan, Lisha Yu, Sijie Yu, Guang Tan

In the Vision-and-Language Navigation (VLN) task, the agent is required to navigate to a destination following a natural language instruction. While learning-based approaches have been a major solution to the task, they …

DiversityNavigateVision and Language Navigation

Can an Embodied Agent Find Your "Cat-shaped Mug"? LLM-Guided Exploration for Zero-Shot Object Navigation

2023-03-06 · Vishnu Sashank Dorbala, James F. Mullen Jr., Dinesh Manocha

We present LGX (Language-guided Exploration), a novel algorithm for Language-Driven Zero-Shot Object Goal Navigation (L-ZSON), where an embodied agent navigates to a uniquely described target object in a previously unsee…

Motion PlanningObjectobject-detectionObject Detection+1

Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation

2024-03-21 · Leyuan Sun, Asako Kanezaki, Guillaume Caron, Yusuke Yoshiyasu

Object-goal navigation is a crucial engineering task for the community of embodied navigation; it involves navigating to an instance of a specified object category within unseen environments. Although extensive investiga…

Common Sense ReasoningLanguage ModelingLanguage ModellingLarge Language Model+2

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

2026-09-02 · Kun-Yang Yu, Yingzhe Li, Hongyu Xu, Shi-Yu Tian 외 arxiv

Vision-Language Navigation (VLN) requires an embodied agent to follow natural-language instructions in unseen environments. Recent progress has been largely driven by Multimodal Large Language Models (MLLMs). Existing me…

Vision-Language Navigation

MCNav: Memory-Aware Dynamic Cognitive Map for Zero-shot Goal-oriented Navigation

2026-05-19 · Jingyu Li, Zhe Liu, Wenxiao Wu, Li Zhang arxiv

Navigating to instance-level targets in complex environments is a challenging problem. Many existing zero-shot methods achieve strong performance by modeling the entire environment and leveraging large language models fo…

Scene Understanding