paper-with-me

홈 › Papers

Self-driven Grounding: Large Language Model Agents with Automatical Language-aligned Skill Learning

2023-09-04 · Shaohui Peng, Xing Hu, Qi Yi, Rui Zhang, Jiaming Guo, Di Huang, Zikang Tian, Ruizhi Chen, Zidong Du, Qi Guo, Yunji Chen, Ling Li

Large language models (LLMs) show their powerful automatic reasoning and planning capability with a wealth of semantic knowledge about the human world. However, the grounding problem still hinders the applications of LLMs in the real-world environment. Existing studies try to fine-tune the LLM or utilize pre-defined behavior APIs to bridge the LLMs and the environment, which not only costs huge human efforts to customize for every single task but also weakens the generality strengths of LLMs. To autonomously ground the LLM onto the environment, we proposed the Self-Driven Grounding (SDG) framework to automatically and progressively ground the LLM with self-driven skill learning. SDG first employs the LLM to propose the hypothesis of sub-goals to achieve tasks and then verify the feasibility of the hypothesis via interacting with the underlying environment. Once verified, SDG can then learn generalized skills with the guidance of these successfully grounded subgoals. These skills can be further utilized to accomplish more complex tasks which fail to pass the verification phase. Verified in the famous instruction following task set-BabyAI, SDG achieves comparable performance in the most challenging tasks compared with imitation learning methods that cost millions of demonstrations, proving the effectiveness of learned skills and showing the feasibility and efficiency of our framework.

📄 PDF Abstract BibTeX arXiv:2309.01352

Code (0)

등록된 구현이 없습니다.

Tasks

Imitation LearningInstruction FollowingLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Co-EPG: A Framework for Co-Evolution of Planning and Grounding in Autonomous GUI Agents

2025-11-13 · Yuan Zhao, Hualei Zhu, Tingyu Jiang, Shen Li 외 arxiv

Graphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current method…

Synthetic Data Generation

EvoGround: Self-Evolving Video Agents for Video Temporal Grounding

2026-05-13 · Minjoon Jung, Byoung-Tak Zhang, Lorenzo Torresani arxiv

Video temporal grounding (VTG) takes an untrimmed video and a natural-language query as input and localizes the temporal moment that best matches the query. Existing methods rely on large, task-specific datasets requirin…

R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding

2025-07-08 · Joonhyung Park, Peng Tang, Sagnik Das, Srikar Appalaraju 외 arxiv

Visual agent models for automating human activities on Graphical User Interfaces (GUIs) have emerged as a promising research direction, driven by advances in large Vision Language Models (VLMs). A critical challenge in G…

Object Detection

Beyond Literal Descriptions: Understanding and Locating Open-World Objects Aligned with Human Intentions

2024-02-17 · Wenxuan Wang, Yisi Zhang, Xingjian He, Yichen Yan 외

Visual grounding (VG) aims at locating the foreground entities that match the given natural language expressions. Previous datasets and methods for classic VG task mainly rely on the prior assumption that the given expre…

Visual Grounding

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents

2026-05-19 · Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo 외 arxiv

We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professio…

Visual Grounding