paper-with-me

홈 › Papers

Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule

2020-09-16 · ICLR 2021 1 · Shuhei Kurita, Kyunghyun Cho

Vision-and-language navigation (VLN) is a task in which an agent is embodied in a realistic 3D environment and follows an instruction to reach the goal node. While most of the previous studies have built and investigated a discriminative approach, we notice that there are in fact two possible approaches to building such a VLN agent: discriminative \textit{and} generative. In this paper, we design and investigate a generative language-grounded policy which uses a language model to compute the distribution over all possible instructions i.e. all possible sequences of vocabulary tokens given action and the transition history. In experiments, we show that the proposed generative approach outperforms the discriminative approach in the Room-2-Room (R2R) and Room-4-Room (R4R) datasets, especially in the unseen environments. We further show that the combination of the generative and discriminative policies achieves close to the state-of-the art results in the R2R dataset, demonstrating that the generative and discriminative policies capture the different aspects of VLN.

📄 PDF Abstract BibTeX arXiv:2009.07783

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingVision and Language Navigation

Similar Papers 제목 키워드 기반

When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation

2026-08-04 · Yinuo Jiang, Yongjie Ye, Zhou Tao, Xiang Zhuang 외 arxiv

On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher signals. Recent selective OPD methods improve this process by prioritizing signals th…

Point What You Mean: Visually Grounded Instruction Policy

2025-12-22 · Hang Yu, Juntu Zhao, Yufeng Liu, Kaiyu Li 외 arxiv

Vision-Language-Action (VLA) models align vision and language with embodied control, but their object referring ability remains limited when relying solely on text prompt, especially in cluttered or out-of-distribution (…

Visual Grounding

In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use

2026-08-06 · Jiarui Yang, Wen Huang, Jiale Zhang, Maowei Hu 외 arxiv

Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expert action chunks conditioned on a static…

Robot Manipulation

Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

2026-09-04 · Yumiao Li, Peixin Liu, Donglin Di, Chen Li 외 arxiv

Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highl…

Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning

2026-06-29 · Peng, Lee, Yin Zhang, Yanglin Zhang 외 arxiv

Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, directly applying RL to rollout multimodal reasoning can lead to instability, due t…

Reinforcement LearningMultimodal Reasoning