paper-with-me

Papers

Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations

2026-05-04 · Yu Lu Liu, Hyokun Yun, Tanya Roosta, Ziang Xiao arxiv

There is growing interest in exploring user simulation as an alternative to gathering and scoring real user-chatbot interactions for AI chatbot evaluation. For this purpose, it is important to ensure the realism of the simulation, i.e., the extent to which simulated dialogues reflect real dialogues users have with chatbots. Most existing methods evaluating simulation realism produce coarse quality signal and remain solely at the level of individual dialogues. To support more rigorous evaluation in this area, we propose realsim, an evaluation framework that enables practitioners to take a distributional view of real vs. simulated dialogues along 8 dimensions, covering attributes related to the communicative functions of the interaction, user states, and the surface form of user messages. We then instantiate the framework with a curated dataset of 1K multi-turn task-focused real user-chatbot dialogues that cover 16 domains of chatbot applications. Overall, we find that simulated users tend to struggle at capturing communication frictions that real users introduce to interactions, which could make evaluations based on such simulations overly optimistic. We also observe variability in performance across different domains, which may indicate a need for domain-specific user simulators.

📄 PDF Abstract BibTeX arXiv:2605.02624

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SynthTextEval: Synthetic Text Data Generation and Evaluation for High-Stakes Domains

2025-07-09 · Krithika Ramesh, Daniel Smolyak, Zihao Zhao, Nupoor Gandhi 외 arxiv

We present SynthTextEval, a toolkit for conducting comprehensive evaluations of synthetic text. The fluency of large language model (LLM) outputs has made synthetic text potentially viable for numerous applications, such…

Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realistic Coaching Agent Interactions

2025-02-18 · Taedong Yun, Eric Yang, Mustafa Safdari, Jong Ha Lee 외

We present an end-to-end framework for generating synthetic users for evaluating interactive agents designed to encourage positive behavior changes, such as in health and lifestyle coaching. The synthetic users are groun…

Real-time Program Evaluation using Anytime-valid Rank Tests

2025-04-30 · Sam van Meer, Nick W. Koning

Counterfactual mean estimators such as difference-in-differences and synthetic control have grown into workhorse tools for program evaluation. Inference for these estimators is well-developed in settings where all post-t…

counterfactualvalid

Realistic Synthetic Household Data Generation at Scale

2026-02-06 · Siddharth Singh, Ifrah Idrees, Abraham Dauhajre arxiv

Advancements in foundation models have catalyzed research in Embodied AI to develop interactive agents capable of environmental reasoning and interaction. Developing such agents requires diverse, large-scale datasets. Pr…

Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets

2026-05-18 · Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti, Yajat Nagaraj Kiran 외 arxiv

Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information. Generating and evaluating synthetic data across privacy, ut…

Decision Making