paper-with-me

홈 › Papers

HyPerNav: Hybrid Perception for Object-Oriented Navigation in Unknown Environment

2025-10-27 · Zecheng Yin, Hao Zhao, Zhen Li arxiv

Objective-oriented navigation(ObjNav) enables robot to navigate to target object directly and autonomously in an unknown environment. Effective perception in navigation in unknown environment is critical for autonomous robots. While egocentric observations from RGB-D sensors provide abundant local information, real-time top-down maps offer valuable global context for ObjNav. Nevertheless, the majority of existing studies focus on a single source, seldom integrating these two complementary perceptual modalities, despite the fact that humans naturally attend to both. With the rapid advancement of Vision-Language Models(VLMs), we propose Hybrid Perception Navigation (HyPerNav), leveraging VLMs' strong reasoning and vision-language understanding capabilities to jointly perceive both local and global information to enhance the effectiveness and intelligence of navigation in unknown environments. In both massive simulation evaluation and real-world validation, our methods achieved state-of-the-art performance against popular baselines. Benefiting from hybrid perception approach, our method captures richer cues and finds the objects more effectively, by simultaneously leveraging information understanding from egocentric observations and the top-down map. Our ablation study further proved that either of the hybrid perception contributes to the navigation performance.

📄 PDF Abstract BibTeX arXiv:2510.22917

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion

2025-05-29 · Chunlong Xie, Jialing He, Shangwei Guo, Jiacheng Wang 외

We present Adversarial Object Fusion (AdvOF), a novel attack framework targeting vision-and-language navigation (VLN) agents in service-oriented environments by generating adversarial 3D objects. While foundational model…

Language ModelingLanguage ModellingObjectService Composition+1

FreqNav: Stage-Wise Frequency Routing for Object-Oriented Aerial Vision-Language Navigation

2026-08-02 · Yin Tang, Jiawei Ma, Jiahao Li, Hao Zhang 외 arxiv

Object-oriented aerial vision-and-language navigation (VLN) requires searching for a described target and landing on it precisely, under long-horizon and closed-loop control. Guided by a target-descriptive instruction du…

Vision-Language NavigationContinuous Control

Designing Privacy-Preserving Visual Perception for Robot Navigation Based on User Privacy Preferences

2026-04-07 · Xuying Huang, Sicong Pan, Delphine Reinhardt, Maren Bennewitz arxiv

Visual navigation is a fundamental capability of mobile service robots, yet the onboard cameras required for such navigation can capture privacy-sensitive information and raise user privacy concerns. Existing approaches …

Visual NavigationRobot Navigation

D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

2025-12-14 · Zihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee Lee arxiv

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose…

Question Answering

IntentReact: Guiding Reactive Object-Centric Navigation via Topological Intent

2026-03-26 · Yanmei Jiao, Anpeng Lu, Wenhan Hu, Rong Xiong 외 arxiv

Object-goal visual navigation requires robots to reason over semantic structure and act effectively under partial observability. Recent approaches based on object-level topological maps enable long-horizon navigation wit…

Visual Navigation