paper-with-me

홈 › Papers

Bridging the Agent-World Gap: Text World Models for LLM-based Agents

2026-06-08 · Yixia Li, Hongru Wang, Peng Lai, Zhiwen Ruan, He Zhu, Youxin Zhu, Ganlong Zhao, Minda Hu, Yun Chen, Sibei Yang, Peng Li, Jeff Z. Pan, Jia Pan, Guanhua Chen, Yang Liu, Guanbin Li arxiv

Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet many remain largely reactive, mapping observations to actions without an explicit model of how these environments are structured and evolve. This motivates text world models (TWMs): transition models over textual states that, given a state and a candidate action, predict the resulting webpage, terminal output, API response, or user reply, thereby supporting planning, efficient learning, and principled evaluation. We systematically review text world models for LLM-based agents, organized around a formal framework and the agent lifecycle: (1) Foundations, defining text world models and characterizing them by state representation and grounding domain; (2) Construction, taxonomizing LLM-as-WM and code-as-WM paradigms and reviewing methods for building them; (3) Application, examining how world models support agents at training time through experience synthesis and at inference time through planning, verification, and adaptation; and (4) Evaluation, covering both evaluation of the world model itself and its use as an evaluation environment for agents. We aim to consolidate this rapidly developing area, clarify its design space, and highlight open challenges for future research.

📄 PDF Abstract BibTeX arXiv:2606.09032

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Implicit Intelligence -- Evaluating Agents on What Users Don't Say

2026-02-23 · Ved Sirdeshmukh, Marc Wetter arxiv

Real-world requests to AI agents are fundamentally underspecified. Natural human communication relies on shared context and unstated constraints that speakers expect listeners to infer. Current agentic benchmarks test ex…

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

2025-01-27 · Wei Chow, Jiageng Mao, Boyi Li, Daniel Seita 외

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have sh…

BenchmarkingCommon Sense ReasoningScene UnderstandingTask Planning

Instruction Agent: Enhancing Agent with Expert Demonstration

2025-09-08 · Yinheng Li, Hailey Hultquist, Justin Wagle, Kazuhito Koishida arxiv

Graphical user interface (GUI) agents have advanced rapidly but still struggle with complex tasks involving novel UI elements, long-horizon actions, and personalized trajectories. In this work, we introduce Instruction A…

A Practical Memory Injection Attack against LLM Agents

2025-03-05 · Shen Dong, Shaocheng Xu, Pengfei He, Yige Li 외

Agents based on large language models (LLMs) have demonstrated strong capabilities in a wide range of complex, real-world applications. However, LLM agents with a compromised memory bank may easily produce harmful output…

Leveraging Evolutionary Surrogate-Assisted Prescription in Multi-Objective Chlorination Control Systems

2025-08-26 · Rivaaj Monsia, Olivier Francon, Daniel Young, Risto Miikkulainen arxiv

This short, written report introduces the idea of Evolutionary Surrogate-Assisted Prescription (ESP) and presents preliminary results on its potential use in training real-world agents as a part of the 1st AI for Drinkin…