paper-with-me

홈 › Papers

Mind the Sim2Real Gap in User Simulation for Agentic Tasks

2026-03-11 · Xuhui Zhou, Weiwei Sun, Qianou Ma, Yiqing Xie, Jiarui Liu, Weihua Du, Sean Welleck, Yiming Yang, Graham Neubig, Sherry Tongshuang Wu, Maarten Sap arxiv

As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals. Yet, these simulations are frequently assumed to be faithful to real human behaviors, often without rigorous verification. We formalize the Sim2Real gap in user simulation and present the first study running the full $τ$-bench protocol with real humans (451 participants, 165 tasks), benchmarking 31 LLM simulators across proprietary, open-source, and specialized families using the User-Sim Index (USI), a metric we introduce to quantify how well LLM simulators resemble real user interactive behaviors and feedback. Behaviorally, LLM simulators are excessively cooperative, stylistically uniform, and lack realistic frustration or ambiguity, creating an "easy mode" that inflates agent success rates above the human baseline. In evaluations, real humans provide nuanced judgments across eight quality dimensions while simulated users produce uniformly more positive feedback; rule-based rewards are failing to capture rich feedback signals generated by human users. Overall, higher general model capability does not necessarily yield more faithful user simulation. These findings highlight the importance of human validation when using LLM-based user simulators in the agent development cycle and motivate improved models for user simulation.

📄 PDF Abstract BibTeX arXiv:2603.11245

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge

2025-06-26 · Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu 외

Agentic search such as Deep Research systems, where large language models autonomously browse the web, synthesize information, and return comprehensive citation-backed answers, represents a major shift in how users inter…

Benchmarking

Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation

2026-02-02 · Jun He, Junyan Ye, Zilong Huang, Dongzhi Jiang 외 arxiv

While text-to-image generation has achieved unprecedented fidelity, the vast majority of existing models function fundamentally as static text-to-pixel decoders. Consequently, they often fail to grasp implicit user inten…

Text-to-Image Generation

When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation

2026-04-01 · Henry Peng Zou, Chunyu Miao, Wei-Chieh Huang, Yankai Chen 외 arxiv

As LLM agents transition from short, static problem solving to executing complex, long-horizon tasks in dynamic environments, the ability to handle user interruptions, such as adding requirement or revising goals, during…

NeuroSkill(tm): Proactive Real-Time Agentic System Capable of Modeling Human State of Mind

2026-03-03 · Nataliya Kosmyna, Eugene Hauptmann arxiv

Real-time proactive agentic system, capable of modeling Human State of Mind, using foundation EXG model and text embeddings model, running fully offline on the edge. Unlike all previously known systems, the NeuroSkill(tm…

VeriMind: Agentic LLM for Automated Verilog Generation with a Novel Evaluation Metric

2025-03-15 · Bardia Nadimi, Ghali Omar Boutaib, Hao Zheng

Designing Verilog modules requires meticulous attention to correctness, efficiency, and adherence to design specifications. However, manually writing Verilog code remains a complex and time-consuming task that demands bo…

ARCCode GenerationText Generation