paper-with-me

홈 › Papers

VLD: Visual Language Goal Distance for Reinforcement Learning Navigation

2025-12-08 · Lazar Milikic, Manthan Patel, Jonas Frey arxiv

Training end-to-end policies from image data to directly predict navigation actions for robotic systems has proven inherently difficult. Existing approaches often suffer from either the sim-to-real gap during policy transfer or a limited amount of training data with action labels. To address this problem, we introduce Vision-Language Distance (VLD) learning, a scalable framework for goal-conditioned navigation that decouples perception learning from policy learning. Instead of relying on raw sensory inputs during policy training, we first train a self-supervised distance-to-goal predictor on internet-scale video data. This predictor generalizes across both image- and text-based goals, providing a distance signal that can be minimized by a reinforcement learning (RL) policy. The RL policy can be trained entirely in simulation using privileged geometric distance signals, with injected noise to mimic the uncertainty of the trained distance predictor. At deployment, the policy consumes VLD predictions, inheriting semantic goal information-"where to go"-from large-scale visual training while retaining the robust low-level navigation behaviors learned in simulation. We propose using ordinal consistency to assess distance functions directly and demonstrate that VLD outperforms prior temporal distance approaches, such as ViNT and VIP. Experiments show that our decoupled design achieves competitive navigation performance in simulation with strong sim-to-real transfer, providing an alternative and, most importantly, scalable path toward reliable, multimodal navigation policies.

📄 PDF Abstract BibTeX arXiv:2512.07976

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Safe Multi-Agent Navigation guided by Goal-Conditioned Safe Reinforcement Learning

2025-02-25 · Meng Feng, Viraj Parimi, Brian Williams

Safe navigation is essential for autonomous systems operating in hazardous environments. Traditional planning methods excel at long-horizon tasks but rely on a predefined graph with fixed distance metrics. In contrast, s…

BenchmarkingReinforcement Learning (RL)Safe Reinforcement Learning

Green for Go, Red for No: Visual Grounding via Semantic Segmentation for VLA Navigation Policies

2026-07-06 · Adrian Szvoren, Dimitrios Kanoulas, Nilufer Tuptuk arxiv

Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first emp…

Semantic SegmentationRobot NavigationVisual Grounding

Multi-goal Audio-visual Navigation using Sound Direction Map

2023-08-01 · Haru Kondoh, Asako Kanezaki

Over the past few years, there has been a great deal of research on navigation tasks in indoor environments using deep reinforcement learning agents. Most of these tasks use only visual information in the form of first-p…

Deep Reinforcement LearningNavigateVisual Navigation

PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Maps

2026-06-01 · Junlin Long, Zeyu Zhang, Xu Deng, Yiran Wang 외 arxiv

Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household service robotics, assistive robotics, a…

Semantic correspondenceVisual Navigation

SARL*: Deep Reinforcement Learning based Human-Aware Navigation for Mobile Robot in Indoor Environments

2020-01-20 · Keyu Li, Yangxin Xu, Jiankun Wang, Max Q.-H. Meng

In a human-robot coexisting environment, reaching the goal position safely and efficiently is essential for a mobile service robot. In this paper, we present an advanced version of the Socially Attentive Reinforcement Le…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)