paper-with-me

Papers

A modular vision language navigation and manipulation framework for long horizon compositional tasks in indoor environment

2021-01-19 · Homagni Saha, Fateme Fotouhif, Qisai Liu, Soumik Sarkar

In this paper we propose a new framework - MoViLan (Modular Vision and Language) for execution of visually grounded natural language instructions for day to day indoor household tasks. While several data-driven, end-to-end learning frameworks have been proposed for targeted navigation tasks based on the vision and language modalities, performance on recent benchmark data sets revealed the gap in developing comprehensive techniques for long horizon, compositional tasks (involving manipulation and navigation) with diverse object categories, realistic instructions and visual scenarios with non-reversible state changes. We propose a modular approach to deal with the combined navigation and object interaction problem without the need for strictly aligned vision and language training data (e.g., in the form of expert demonstrated trajectories). Such an approach is a significant departure from the traditional end-to-end techniques in this space and allows for a more tractable training process with separate vision and language data sets. Specifically, we propose a novel geometry-aware mapping technique for cluttered indoor environments, and a language understanding model generalized for household instruction following. We demonstrate a significant increase in success rates for long-horizon, compositional tasks over the baseline on the recently released benchmark data set-ALFRED.

📄 PDF Abstract BibTeX arXiv:2101.07891

Code (1)

Homagn/MOVILAN 공식 구현 pytorch

Tasks

Instruction FollowingVision-Language Navigation

Similar Papers 제목 키워드 기반

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation

2022-06-17 · Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, Xin Eric Wang

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill t…

Object

D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

2025-12-14 · Zihan Wang, Seungjun Lee, Guangzhao Dai, Gim Hee Lee arxiv

Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose…

Question Answering

Being-0: A Humanoid Robotic Agent with Vision-Language Models and Modular Skills

2025-03-16 · Haoqi Yuan, Yu Bai, Yuhui Fu, Bohan Zhou 외

Building autonomous robotic agents capable of achieving human-level performance in real-world embodied tasks is an ultimate goal in humanoid robot research. Recent advances have made significant progress in high-level co…

Task Planning

A Navigation Framework Utilizing Vision-Language Models

2025-06-11 · Yicheng Duan, Kaiyu Tang

Vision-and-Language Navigation (VLN) presents a complex challenge in embodied AI, requiring agents to interpret natural language instructions and navigate through visually rich, unfamiliar environments. Recent advances i…

NavigatePrompt EngineeringVision and Language Navigation

Fully Autonomous Real-World Reinforcement Learning with Applications to Mobile Manipulation

2021-07-28 · Charles Sun, Jędrzej Orbik, Coline Devin, Brian Yang 외

We study how robots can autonomously learn skills that require a combination of navigation and grasping. While reinforcement learning in principle provides for automated robotic skill learning, in practice reinforcement …

Continual LearningNavigatereinforcement-learningReinforcement Learning+1