paper-with-me

Papers

FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning

2025-09-19 · Naoki Yokoyama, Sehoon Ha arxiv

Enabling robotic assistants to navigate complex environments and locate objects described in free-form language is a critical capability for real-world deployment. While foundation models, particularly Vision-Language Models (VLMs), offer powerful semantic understanding, effectively adapting their web-scale knowledge for embodied decision-making remains a key challenge. We present FiLM-Nav (Fine-tuned Language Model for Navigation), an approach that directly fine-tunes pre-trained VLM as the navigation policy. In contrast to methods that use foundation models primarily in a zero-shot manner or for map annotation, FiLM-Nav learns to select the next best exploration frontier by conditioning directly on raw visual trajectory history and the navigation goal. Leveraging targeted simulated embodied experience allows the VLM to ground its powerful pre-trained representations in the specific dynamics and visual patterns relevant to goal-driven navigation. Critically, fine-tuning on a diverse data mixture combining ObjectNav, OVON, ImageNav, and an auxiliary spatial reasoning task proves essential for achieving robustness and broad generalization. FiLM-Nav sets a new state-of-the-art in both SPL and success rate on HM3D ObjectNav among open-vocabulary methods, and sets a state-of-the-art SPL on the challenging HM3D-OVON benchmark, demonstrating strong generalization to unseen object categories. Our work validates that directly fine-tuning VLMs on diverse simulated embodied data is a highly effective pathway towards generalizable and efficient semantic navigation capabilities.

📄 PDF Abstract BibTeX arXiv:2509.16445

Code (0)

등록된 구현이 없습니다.

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

Tactile Modality Fusion for Vision-Language-Action Models

2026-03-15 · Charlotte Morissette, Amin Abyaneh, Wei-Di Chang, Anas Houssaini 외 arxiv

We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While recent advances in VLA models have introduced robot policies that are both…

PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models

2025-07-23 · Jiansong Wan, Chengming Zhou, Jinkua Liu, Xiangge Huang 외 arxiv

Recent studies have explored pretrained (foundation) models for vision-based robotic navigation, aiming to achieve generalizable navigation and positive transfer across diverse environments while enhancing zero-shot perf…

Representation LearningVisual Navigation

End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering

2024-11-08 · Dylan Goetting, Himanshu Gaurav Singh, Antonio Loquercio

We present VLMnav, an embodied framework to transform a Vision-Language Model (VLM) into an end-to-end navigation policy. In contrast to prior work, we do not rely on a separation between perception, planning, and contro…

Language ModelingLanguage ModellingQuestion AnsweringSpatial Reasoning

Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents

2025-04-04 · Jaymari Chua, Chen Wang, Lina Yao

Generalizable alignment is a core challenge for deploying Large Language Models (LLMs) safely in real-world NLP applications. Current alignment methods, including Reinforcement Learning from Human Feedback (RLHF), often …

Safe Reinforcement Learning

Laser-Based Fabrication of Microstructures on Nickel Thin Films and Its Applications in On-Chip Thin Film Inductors

2023-11-17 · Srikanth Itapu, Vamsi Borra, Daniel G. Georgiev

This work reports on the fabrication of microbump structures on Ni films by single-pulse, localized laser irradiation. Conditions for the reproducible formation of such microstructures have been identified in terms of la…