paper-with-me

홈 › Papers

Modular Framework for Visuomotor Language Grounding

2021-09-05 · Kolby Nottingham, Litian Liang, Daeyun Shin, Charless C. Fowlkes, Roy Fox, Sameer Singh

Natural language instruction following tasks serve as a valuable test-bed for grounded language and robotics research. However, data collection for these tasks is expensive and end-to-end approaches suffer from data inefficiency. We propose the structuring of language, acting, and visual tasks into separate modules that can be trained independently. Using a Language, Action, and Vision (LAV) framework removes the dependence of action and vision modules on instruction following datasets, making them more efficient to train. We also present a preliminary evaluation of LAV on the ALFRED task for visual and interactive instruction following.

📄 PDF Abstract BibTeX arXiv:2109.02161

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

SAGA: Open-World Mobile Manipulation via Structured Affordance Grounding

2025-12-14 · Kuan Fang, Yuxin Chen, Xinghao Zhu, Farzad Niroui 외 arxiv

We present SAGA, a versatile and adaptive framework for visuomotor control that can generalize across various environments, task objectives, and user specifications. To efficiently learn such capability, our key idea is …

FALCON: Actively Decoupled Visuomotor Policies for Loco-Manipulation with Foundation-Model-Based Coordination

2025-12-04 · Chengyang He, Ge Sun, Yue Bai, Junkai Lu 외 arxiv

We present FoundAtion-model-guided decoupled LoCO-maNipulation visuomotor policies (FALCON), a framework for loco-manipulation that combines modular diffusion policies with a vision-language foundation model as the coord…

Example-Driven Model-Based Reinforcement Learning for Solving Long-Horizon Visuomotor Tasks

2021-09-21 · Bohan Wu, Suraj Nair, Li Fei-Fei, Chelsea Finn

In this paper, we study the problem of learning a repertoire of low-level skills from raw images that can be sequenced to complete long-horizon visuomotor tasks. Reinforcement learning (RL) is a promising approach for ac…

Model-based Reinforcement Learningreinforcement-learningReinforcement Learning (RL)

MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements

2025-11-17 · SeokJoo Kwak, Jihoon Kim, Boyoun Kim, Jung Jae Yoon 외 arxiv

Graphical User Interface (GUI) grounding - the task of mapping natural language instructions to screen coordinates - is essential for autonomous agents and accessibility technologies. Existing systems rely on monolithic …

myGym: Modular Toolkit for Visuomotor Robotic Tasks

2020-12-21 · Michal Vavrecka, Nikita Sokovnin, Megi Mejdrechova, Gabriela Sejnova 외

We introduce a novel virtual robotic toolkit myGym, developed for reinforcement learning (RL), intrinsic motivation and imitation learning tasks trained in a 3D simulator. The trained tasks can then be easily transferred…

Imitation LearningOpenAI GymReinforcement Learning (RL)