paper-with-me

홈 › Papers

Semantic Skill Grounding for Embodied Instruction-Following in Cross-Domain Environments

2024-08-02 · Sangwoo Shin, SeungHyun Kim, Youngsoo Jang, Moontae Lee, Honguk Woo

In embodied instruction-following (EIF), the integration of pretrained language models (LMs) as task planners emerges as a significant branch, where tasks are planned at the skill level by prompting LMs with pretrained skills and user instructions. However, grounding these pretrained skills in different domains remains challenging due to their intricate entanglement with the domain-specific knowledge. To address this challenge, we present a semantic skill grounding (SemGro) framework that leverages the hierarchical nature of semantic skills. SemGro recognizes the broad spectrum of these skills, ranging from short-horizon low-semantic skills that are universally applicable across domains to long-horizon rich-semantic skills that are highly specialized and tailored for particular domains. The framework employs an iterative skill decomposition approach, starting from the higher levels of semantic skill hierarchy and then moving downwards, so as to ground each planned skill to an executable level within the target domain. To do so, we use the reasoning capabilities of LMs for composing and decomposing semantic skills, as well as their multi-modal extension for assessing the skill feasibility in the target domain. Our experiments in the VirtualHome benchmark show the efficacy of SemGro in 300 cross-domain EIF scenarios.

📄 PDF Abstract BibTeX arXiv:2408.01024

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

Efficient Skill Grounding via Code Refactoring with Small Language Models

2026-06-06 · Sera Choi, Wonje Choi, Saehun Chun, Daehee Lee 외 arxiv

Effective skill grounding is essential for deploying reusable skills in embodied agents, as even minor embodiment or environmental differences can render an entire skill incompatible. This challenge is particularly prono…

Grounding Beyond Detection: Enhancing Contextual Understanding in Embodied 3D Grounding

2025-06-05 · Yani Zhang, Dongming Wu, Hao Shi, Yingfei Liu 외

Embodied 3D grounding aims to localize target objects described in human instructions from ego-centric viewpoint. Most methods typically follow a two-stage paradigm where a trained 3D detector's optimized backbone parame…

LocoVLM: Grounding Vision and Language for Adapting Versatile Legged Locomotion Policies

2026-02-11 · I Made Aswin Nahrendra, Seunghyun Lee, Dongkyu Lee, Hyun Myung arxiv

Recent advances in legged locomotion learning are still dominated by the utilization of geometric representations of the environment, limiting the robot's capability to respond to higher-level semantics such as human ins…

Are We There Yet? Learning to Localize in Embodied Instruction Following

2021-01-09 · Shane Storks, Qiaozi Gao, Govind Thattai, Gokhan Tur

Embodied instruction following is a challenging problem requiring an agent to infer a sequence of primitive actions to achieve a goal environment state from complex language and visual inputs. Action Learning From Realis…

Instruction Followingobject-detectionObject Detection

OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping

2025-08-03 · Danyang Li, Zenghui Yang, Guangpeng Qi, Songtao Pang 외 arxiv

Grounding natural language instructions to visual observations is fundamental for embodied agents operating in open-world environments. Recent advances in visual-language mapping have enabled generalizable semantic repre…