paper-with-me

홈 › Papers

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations

2026-01-23 · Preethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh, Seraphina Goldfarb-Tarrant arxiv

Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on τ-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore, evaluations using simulated users exhibit systematic miscalibration, underestimating agent performance on challenging tasks and overestimating it on moderately difficult ones. African American Vernacular English (AAVE) speakers experience consistently worse success rates and calibration errors than Standard American English (SAE) speakers, with disparities compounding significantly with age. We also find simulated users to be a differentially effective proxy for different populations, performing worst for AAVE and Indian English speakers. Additionally, simulated users introduce conversational artifacts and surface different failure patterns than human users. These findings demonstrate that current evaluation practices risk misrepresenting agent capabilities across diverse user populations and may obscure real-world deployment challenges.

📄 PDF Abstract BibTeX arXiv:2601.17087

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation

2026-04-07 · Ming Zhu, Juntao Tan, Rithesh Murthy, Jielin Qiu 외 arxiv

LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstrained LLM defaults produce a Formalism Ceiling (style match rates of 6-8% …

Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions

2026-06-04 · Xinnong Zhang, Wanting Shan, Hanjia Lyu, Zhongyu Wei 외 arxiv

Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it remains unclear whether these simulations reflect precise user-specific …

Mind the Sim2Real Gap in User Simulation for Agentic Tasks

2026-03-11 · Xuhui Zhou, Weiwei Sun, Qianou Ma, Yiqing Xie 외 arxiv

As NLP evaluation shifts from static benchmarks to multi-turn interactive settings, LLM-based simulators have become widely used as user proxies, serving two roles: generating user turns and providing evaluation signals.…

LLMs Get Lost In Multi-Turn Conversation

2025-05-09 · Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville

Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define, explore, and refine …

A Practical Recipe Towards Improving Sim-and-Real Correlation for VLA Evaluation

2026-06-09 · Shuo Wang, Hanyuan Xu, Yingdong Hu, Fanqi Lin 외 arxiv

Simulation has become an essential tool for evaluating and improving vision-language-action (VLA) policies, offering scalable, reproducible, and controllable alternatives to costly real-world robot evaluation. Recent sim…