paper-with-me

Papers

ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

2023-02-09 · Ying Shen, Daniel Bis, Cynthia Lu, Ismini Lourentzou

The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities.

📄 PDF Abstract BibTeX arXiv:2302.04865

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Navigation

Similar Papers 제목 키워드 기반

Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation

2026-03-31 · Edoardo Zorzi, Francesco Taioli, Yiming Wang, Marco Cristani 외 arxiv

We propose Question-Asking Navigation (QAsk-Nav), the first reproducible benchmark for Collaborative Instance Object Navigation (CoIN) that enables an explicit, separate assessment of embodied navigation and collaborativ…

Good Time to Ask: A Learning Framework for Asking for Help in Embodied Visual Navigation

2022-06-20 · Jenny Zhang, Samson Yu, Jiafei Duan, Cheston Tan

In reality, it is often more efficient to ask for help than to search the entire space to find an object with an unknown location. We present a learning framework that enables an agent to actively ask for help in such em…

Visual Navigation

Deep Learning for Embodied Vision Navigation: A Survey

2021-07-07 · Fengda Zhu, Yi Zhu, Vincent CS Lee, Xiaodan Liang 외

"Embodied visual navigation" problem requires an agent to navigate in a 3D environment mainly rely on its first-person observation. This problem has attracted rising attention in recent years due to its wide application …

Autonomous DrivingDeep LearningNavigateSurvey+1

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

2026-08-31 · Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu 외 hf

Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial pri…

Zero-shot GeneralizationReinforcement LearningInstruction FollowingSpatial Reasoning

Which way is `right'?: Uncovering limitations of Vision-and-Language Navigation model

2023-11-30 · Meera Hahn, Amit Raj, James M. Rehg

The challenging task of Vision-and-Language Navigation (VLN) requires embodied agents to follow natural language instructions to reach a goal location or object (e.g. `walk down the hallway and turn left at the piano'). …

Vision and Language Navigation