paper-with-me

홈 › Papers

Visual Navigation with Spatial Attention

2021-04-20 · CVPR 2021 1 · Bar Mayo, Tamir Hazan, Ayellet Tal

This work focuses on object goal visual navigation, aiming at finding the location of an object from a given class, where in each step the agent is provided with an egocentric RGB image of the scene. We propose to learn the agent's policy using a reinforcement learning algorithm. Our key contribution is a novel attention probability model for visual navigation tasks. This attention encodes semantic information about observed objects, as well as spatial information about their place. This combination of the "what" and the "where" allows the agent to navigate toward the sought-after object effectively. The attention model is shown to improve the agent's policy and to achieve state-of-the-art results on commonly-used datasets.

📄 PDF Abstract BibTeX arXiv:2104.09807

Code (1)

barmayo/spatial_attention pytorch

Tasks

NavigateObjectreinforcement-learningReinforcement Learning (RL)Visual Navigation

Similar Papers 제목 키워드 기반

Audio Spatially-Guided Fusion for Audio-Visual Navigation

2026-04-02 · Xinyu Zhou, Yinfeng Yu arxiv

Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core chall…

Visual Navigation

Building Category Graphs Representation with Spatial and Temporal Attention for Visual Navigation

2023-12-06 · Xiaobo Hu, Youfang Lin, Hehe Fan, Shuo Wang 외

Given an object of interest, visual navigation aims to reach the object's location based on a sequence of partial observations. To this end, an agent needs to 1) learn a piece of certain knowledge about the relations of …

ObjectVisual Navigation

VTNet: Visual Transformer Network for Object Goal Navigation

2021-05-20 · ICLR 2021 1 · Heming Du, Xin Yu, Liang Zheng

Object goal navigation aims to steer an agent towards a target object based on observations of the agent. It is of pivotal importance to design effective visual representations of the observed scene in determining naviga…

Object

Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships

2020-04-29 · Yunlian Lv, Ning Xie, Yimin Shi, Zijiao Wang 외

Embodied artificial intelligence (AI) tasks shift from tasks focusing on internet images to active settings involving embodied agents that perceive and act within 3D environments. In this paper, we investigate the target…

Deep Reinforcement LearningVisual Navigation

TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation

2026-03-03 · Jiaxing Liu, Zexi Zhang, Xiaoyan Li, Boyue Wang 외 arxiv

Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on static, disembodied vision-language task…

Vision-Language NavigationSpatial Reasoning